镜像站点 · 本页由第三方 GitHub 只读镜像提供,非 GitHub 官方站点,不接受任何登录或凭据输入。前往 github.com
Skip to content

test: test-buffer-creation-regression intermittent failure on smartos #10166

Description

@mhdawson
  • Version: master
  • Platform: smartos
  • Subsystem: buffer

Failure on smartos15-64 during regression run for change not related to this test:

https://ci.nodejs.org/job/node-test-commit-smartos/5631/nodes=smartos15-64/console

not ok 96 parallel/test-buffer-creation-regression
  ---
  duration_ms: 61.210
  severity: fail
  stack: |-
    timeout
  ...

Activity

  1. mhdawson commented on Dec 7, 2016

    @mhdawson
    MemberAuthor

    @nodejs/platform-smartos

  2. cjihrig commented on Dec 7, 2016

    @cjihrig
    Contributor

    Related: #10161

  3. added
    bufferIssues and PRs related to the buffer subsystem.
    smartosIssues and PRs related to the SmartOS platform.
    testIssues and PRs related to Node.js core tests and test infrastructure.
    on Dec 7, 2016
  4. Trott commented on Dec 7, 2016

    @Trott
    Member

    Fixed in #10161.

  5. mhdawson commented on Dec 13, 2016

    @mhdawson
    MemberAuthor

    Still failed in sequential run here: https://ci.nodejs.org/job/node-test-commit-smartos/5739/nodes=smartos15-64/console

    not ok 1271 sequential/test-buffer-creation-regression
      ---
      duration_ms: 369.121
      severity: fail
      stack: |-
        timeout
      ...
    
  6. bnoordhuis commented on Dec 13, 2016

    @bnoordhuis
    Member

    Suggestion: mark the test as flaky on smartos and move on? I suspect this is caused by the infamously terrible virtual memory management that smartos/illumos inherited from solaris. It's not the first time we've hit this.

  7. bcantrill commented on Dec 14, 2016

    @bcantrill
    Contributor

    Counter-suggestion: don't debug by superstition or prejudice. Could we please get a little more detail on what it is you're talking about? If there's a SmartOS or illumos issue here, we'd like to understand and fix it. And if you are unable or unwilling to debug it please let us know, and we will find someone to do it.

  8. misterdjules commented on Feb 3, 2017

    @misterdjules

    In the description of the second failure, duration_ms: 369.121 means that it took 369 seconds (not milliseconds) for tools/test.py to notice that this specific test did not complete.

    However, the default timeout is 60 seconds, and tools/test.py actively polls its child processes' status every at least as frequently as 1/10th of a second.

    It seems to suggest that the system was heavily loaded and that the test harness process couldn't be scheduled as frequently as usual, thus making this failure possibly caused by the environment more than by the test itself or a bug in node on SmartOS.

    Given that this test doesn't seem to fail that frequently, and that I haven't been able to reproduce that failure manually, I would suggest the following way forward:

    1. Pass --abort-on-timeout to tools/test.py when running tests on SmartOS as described in Enable --abort-on-timeout for node-test-commit-smartos Jenkins job build#613.

    2. Continue adding instrumentation features to the test harness so that we can get processes' and system stats when a test times out.

    3. Close this issue without marking this test as flaky, and reopen it when the next failure happens. The reason for not marking this test as flaky is that we currently don't know if the test itself is actually flaky, and that it is possible that failures were due to the environment. Marking it as flaky would hide future actual failures. Hopefully, when the next failure occurs, we'll be able to use the instrumentation described in 1) and 2) to have a better idea of what the problem is.

    How does that sound?

  9. Trott commented on Feb 3, 2017

    @Trott
    Member

    SGTM

  10. misterdjules commented on Feb 4, 2017

    @misterdjules

    thus making this failure possibly caused by the environment more than by the test itself or a bug in node on SmartOS.

    After further testing, it seems that the performance of this test on SmartOS could be related to the version of the platform on which the VM running the test is provisioned.

    To make a long story short, after looking at various kstats (notably the n_pf_throttle_usec in the memory_cap module that represents the time spent by the process being throttled when page faulting) and microstate counters with prstat -m, I couldn't come to a definite conclusion about what's going on.

    However, I did provision two VMs using the same image as the current test VMs used to run test Jenkins jobs, using a 8GBs RAM package. One was provisioned on a server using a 6 months old platform, the other on a server using a more recent platform. The VM provisioned on the more recent platform was able to run this test in ~1 second consistently, whereas the other was running the same test in ~10 seconds consistently, with some peaks ~50 seconds.

    So instead of moving forward with what I suggested previously, I'll investigate further and update this issue as soon as I have more details.

  11. thefourtheye commented on Feb 5, 2017

    @thefourtheye
    Contributor

    If needed, we can reduce the size of that test

    const size = 8589934592;

    to

    const size = 4294968296;

    I'll submit a PR soon.

  12. misterdjules commented on Feb 7, 2017

    @misterdjules

    If needed, we can reduce the size of that test

    It certainly won't hurt to perform computation and I/O that doesn't need to happen. However, just to make sure everyone is on the same page, reducing the size of the buffer created to 4294968296 won't systematically lower the rate of failures for this test, and fixing this test on SmartOS will require understanding why it performs poorly on that platform compared to others.

  13. thefourtheye commented on Feb 8, 2017

    @thefourtheye
    Contributor

    @misterdjules I agree. I submitted the PR because we are over-allocating anyway. 4294968296 (or anything more than 4294967296) should be enough to prove that the patch is working.

  14. Trott commented on Jul 26, 2017

    @Trott
    Member

    It seems that we may have not seen this failure for a few months now. I'm going to close this but anyone should feel free to re-open (or comment requesting it be re-opened if GitHub does not allow you to re-open) if they feel that's not the right way to go.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bufferIssues and PRs related to the buffer subsystem.smartosIssues and PRs related to the SmartOS platform.testIssues and PRs related to Node.js core tests and test infrastructure.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions