镜像站点 · 本页由第三方 GitHub 只读镜像提供,非 GitHub 官方站点,不接受任何登录或凭据输入。前往 github.com
Skip to content

Support for file-system based persistent code cache in user-land module loaders #47472

Description

@joyeecheung

This stemmed from a Twitter thread. Specifically I am wondering if there are any concerns over having something similar to what https://github.com/zertosh/v8-compile-cache does in core, the general idea is:

  1. If the user enables this feature (probably should be off by default) e.g. via an environment variable, whenever we compile a module, we produce the code cache for the module, and on process exit, we store any new cache produced in a cache directory on the file system.
  2. The next time the process is launched (with this feature enabled again), whenever we are loading a module, we attempt to load the cache from that directory and use it when compiling the module, in order to speed up the start up (where most of the time is usually spent on compilation).

This is also similar to what Chrome does with the V8 code cache.

The motivation for implementing this in core is that, for a user-land module to do this for CJS, it has to monkey patch the CJS loader, and this increases the compatibility burden (v8-compile-cache has 17M weekly downloads, and it needs to monkey-patch Module.prototype._compile to work. From a glance of its issue tracker it seems some of the issues are not really fixable in the user land either, like piping into the internal source maps cache). For ESM currently the user land can only use --loader to customize the compilation, which has a cost of its own (especially when we move it to a separate thread), creating a disparity from CJS, and also even with --loader I doubt if user-land code can integrate into e.g. the source map cache without asking us to expose too much internals.

The most risky part of this feature might be the growth of the cache, but it seems manageable if:

  1. The feature is opt-in (via an Environment variable, for example, or a method that can be called to enable/disable from user land).
  2. We do some checks for the size of the cache directory when this feature is used, and set a default cache size limit to prevent unbound growth.

This doesn't seem too radical, for example we already persist something like the repl history by default, and we also have features like NODE_V8_COVERAGE that does a similar "writing a lot of data to a directory when enabled" thing. I don't think this would increase the code complexity much either (we might also want a read-only version of this for SEA in the future too). So opening this issue to see if there are any concerns about having this in core before implementing it.

Activity

  1. joyeecheung commented on Apr 7, 2023

    @joyeecheung
    MemberAuthor

    cc @nodejs/startup @nodejs/loaders

  2. added
    moduleIssues and PRs related to the module subsystem.
    loadersIssues and PRs related to ES module loaders.
    on Apr 8, 2023
  3. jakebailey commented on Apr 8, 2023

    @jakebailey
    Member

    For completeness, there's also https://www.npmjs.com/package/v8-compile-cache-lib with another 9 million weekly downloads; this one is used by ts-node and others. @cspotcode

    Given how many short lived node processes there are out there, having this sort of thing enabled globally feels like it could be a good idea (depending on the downsides).

  4. bnoordhuis commented on Apr 8, 2023

    @bnoordhuis
    Member

    Code cache corruption is an issue though. V8 only performs the lightest of sanity checks. Bad inputs will crash the process, or worse. It opens up new attack vectors.

  5. bmeck commented on Apr 8, 2023

    @bmeck
    Member
  6. joyeecheung commented on Apr 8, 2023

    @joyeecheung
    MemberAuthor

    It opens up new attack vectors.

    I can't think of anything new that's out of the scope in our threat model - to feed bad input to the module loader, the attacker needs to have access to the cache directory and corrupt the cache on-disk. But if the attacker has that level of access to the file system, the integrity about the actual source files already can't be trusted - unless policy is enabled, but in that case we could take policy into account in the implementation, whereas it'd be harder for any user-land solutions to do this, no matter how popular they already are. And this seems to be an even better motivation to provide this in core because the existing popular user-land solutions with ~25M weekly downloads already monkey patch Module.prototype._compile in a way that completely drops the policy assertions. We have already explicitly stated that we trust the file system when loading a module in our threat model, anyway. In addition we have features like NODE_EXTRA_CA_CERTS / SSL_CERT_DIR / NODE_REPL_EXTERNAL_MODULE / NODE_ICU_DATA etc., and they are probably much easier/straight-forward to exploit compared to code caches (even in those cases, exploits that depend on altering the inputs to those environment variables are already out of our security scope, because again we simply trust the file system).

    If the feature is opt-in, the mitigation against any new-found venerability that actually is in our security scope (even though I think that'd be unlikely given the reasons stated above) would also be simple - the user can just stop using it (e.g. unset the environment variable), and we can quickly make a security release by making it a noop until the vulnerability is addressed, and it shouldn't result in behavioral regression - the regression would only be on the module loading performance.

  7. bnoordhuis commented on Apr 9, 2023

    @bnoordhuis
    Member

    One obvious angle is running node as setuid root, or otherwise running with elevated privileges (ex. capabilities on Linux.)

  8. joyeecheung commented on Apr 12, 2023

    @joyeecheung
    MemberAuthor

    @bnoordhuis Where is the threat coming from in those cases? How would this be different compared to e.g. NODE_EXTRA_CA_CERTS / SSL_CERT_DIR / NODE_REPL_EXTERNAL_MODULE / NODE_ICU_DATA?

  9. bnoordhuis commented on Apr 12, 2023

    @bnoordhuis
    Member

    We're being careful to ignore those environment variables when running as setuid root or a host of other things. That same caution should be applied when reading files from disk.

  10. joyeecheung commented on Apr 15, 2023

    @joyeecheung
    MemberAuthor

    We're being careful to ignore those environment variables when running as setuid root or a host of other things. That same caution should be applied when reading files from disk.

    Yes I agree though I don't see this as a new threat - I think whatever we need is probably already covered by SafeGetEnv and if there's something missing, we should fix SafeGetEnv for all these variables, and I doubt environment variables for on-disk code cache should be treated any differently compared to these sensitive variables in this regard.

    P.S.: we don't use SafeGetEnv for all the variables I mentioned above...maybe we should...

    P.P.S.: On the other hand we have this popular package with ~25M weekly downloads in the ecosystem that does the something similar without taking elevated privileges into account...I would say implementing it in core could probably help improving the situation with the potential threat in the ecosystem

    P.P.P.S.: we should probably consider exposing safeGetEnv to user-land for this purpose.

  11. 46 remaining items

  12. added a commit that references this issue on Apr 22, 2024
  13. moved this from Awaiting Triage to Done in Node.js feature requestson Jun 29, 2024
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

discussIssues opened for discussion and feedback.feature requestIssues requesting new Node.js features.loadersIssues and PRs related to ES module loaders.moduleIssues and PRs related to the module subsystem.

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions