Repository navigation
test: test-buffer-creation-regression intermittent failure on smartos #10166
Description
Activity
@nodejs/platform-smartos
Related: #10161
- addedbufferIssues and PRs related to the buffer subsystem.Issues and PRs related to the buffer subsystem.smartosIssues and PRs related to the SmartOS platform.Issues and PRs related to the SmartOS platform.testIssues and PRs related to Node.js core tests and test infrastructure.Issues and PRs related to Node.js core tests and test infrastructure.
on Dec 7, 2016 Fixed in #10161.
Still failed in sequential run here: https://ci.nodejs.org/job/node-test-commit-smartos/5739/nodes=smartos15-64/console
not ok 1271 sequential/test-buffer-creation-regression --- duration_ms: 369.121 severity: fail stack: |- timeout ...Suggestion: mark the test as flaky on smartos and move on? I suspect this is caused by the infamously terrible virtual memory management that smartos/illumos inherited from solaris. It's not the first time we've hit this.
Counter-suggestion: don't debug by superstition or prejudice. Could we please get a little more detail on what it is you're talking about? If there's a SmartOS or illumos issue here, we'd like to understand and fix it. And if you are unable or unwilling to debug it please let us know, and we will find someone to do it.
In the description of the second failure,
duration_ms: 369.121means that it took 369 seconds (not milliseconds) fortools/test.pyto notice that this specific test did not complete.However, the default timeout is
60seconds, andtools/test.pyactively polls its child processes' status every at least as frequently as 1/10th of a second.It seems to suggest that the system was heavily loaded and that the test harness process couldn't be scheduled as frequently as usual, thus making this failure possibly caused by the environment more than by the test itself or a bug in node on SmartOS.
Given that this test doesn't seem to fail that frequently, and that I haven't been able to reproduce that failure manually, I would suggest the following way forward:
-
Pass
--abort-on-timeouttotools/test.pywhen running tests on SmartOS as described in Enable --abort-on-timeout for node-test-commit-smartos Jenkins job build#613. -
Continue adding instrumentation features to the test harness so that we can get processes' and system stats when a test times out.
-
Close this issue without marking this test as flaky, and reopen it when the next failure happens. The reason for not marking this test as flaky is that we currently don't know if the test itself is actually flaky, and that it is possible that failures were due to the environment. Marking it as flaky would hide future actual failures. Hopefully, when the next failure occurs, we'll be able to use the instrumentation described in 1) and 2) to have a better idea of what the problem is.
How does that sound?
-
SGTM
thus making this failure possibly caused by the environment more than by the test itself or a bug in node on SmartOS.
After further testing, it seems that the performance of this test on SmartOS could be related to the version of the platform on which the VM running the test is provisioned.
To make a long story short, after looking at various kstats (notably the
n_pf_throttle_usecin thememory_capmodule that represents the time spent by the process being throttled when page faulting) and microstate counters withprstat -m, I couldn't come to a definite conclusion about what's going on.However, I did provision two VMs using the same image as the current test VMs used to run test Jenkins jobs, using a 8GBs RAM package. One was provisioned on a server using a 6 months old platform, the other on a server using a more recent platform. The VM provisioned on the more recent platform was able to run this test in ~1 second consistently, whereas the other was running the same test in ~10 seconds consistently, with some peaks ~50 seconds.
So instead of moving forward with what I suggested previously, I'll investigate further and update this issue as soon as I have more details.
If needed, we can reduce the
sizeof that testconst size = 8589934592;
to
const size = 4294968296;
I'll submit a PR soon.
If needed, we can reduce the size of that test
It certainly won't hurt to perform computation and I/O that doesn't need to happen. However, just to make sure everyone is on the same page, reducing the size of the buffer created to
4294968296won't systematically lower the rate of failures for this test, and fixing this test on SmartOS will require understanding why it performs poorly on that platform compared to others.@misterdjules I agree. I submitted the PR because we are over-allocating anyway.
4294968296(or anything more than4294967296) should be enough to prove that the patch is working.- added a commit that references this issue
on Mar 3, 2017 - added a commit that references this issue
on Apr 4, 2017 - added a commit that references this issue
on Apr 10, 2017 - added 2 commits that reference this issue
on Apr 18, 2017 - added a commit that references this issue
on Jul 19, 2017 It seems that we may have not seen this failure for a few months now. I'm going to close this but anyone should feel free to re-open (or comment requesting it be re-opened if GitHub does not allow you to re-open) if they feel that's not the right way to go.
Failure on smartos15-64 during regression run for change not related to this test:
https://ci.nodejs.org/job/node-test-commit-smartos/5631/nodes=smartos15-64/console