Repository navigation
Support for file-system based persistent code cache in user-land module loaders #47472
Description
Activity
- addedfeature requestIssues requesting new Node.js features.Issues requesting new Node.js features.
on Apr 7, 2023 - addeddiscussIssues opened for discussion and feedback.Issues opened for discussion and feedback.
on Apr 7, 2023 cc @nodejs/startup @nodejs/loaders
- addedmoduleIssues and PRs related to the module subsystem.Issues and PRs related to the module subsystem.loadersIssues and PRs related to ES module loaders.Issues and PRs related to ES module loaders.
on Apr 8, 2023 For completeness, there's also https://www.npmjs.com/package/v8-compile-cache-lib with another 9 million weekly downloads; this one is used by ts-node and others. @cspotcode
Given how many short lived node processes there are out there, having this sort of thing enabled globally feels like it could be a good idea (depending on the downsides).
Code cache corruption is an issue though. V8 only performs the lightest of sanity checks. Bad inputs will crash the process, or worse. It opens up new attack vectors.
- A singular DB instead of a file per cache entry would likely be better I would imagine. Avoids things like .pyc mismatches and can be used to add extra checks as a container format. Attack vector introduction is real but should likely be discussed since it might be simple to find an adequate mitigation even if it requires opt-in.…On Sat, Apr 8, 2023, 4:50 AM Ben Noordhuis ***@***.***> wrote: Code cache corruption is an issue though. V8 only performs the lightest of sanity checks. Bad inputs will crash the process, or worse. It opens up new attack vectors. — Reply to this email directly, view it on GitHub <#47472 (comment)>, or unsubscribe <https://github.com/notifications/unsubscribe-auth/AABZJI5CVZNMONQDGCR32O3XAEYHFANCNFSM6AAAAAAWW757XE> . You are receiving this because you are on a team that was mentioned.Message ID: ***@***.***>
It opens up new attack vectors.
I can't think of anything new that's out of the scope in our threat model - to feed bad input to the module loader, the attacker needs to have access to the cache directory and corrupt the cache on-disk. But if the attacker has that level of access to the file system, the integrity about the actual source files already can't be trusted - unless policy is enabled, but in that case we could take policy into account in the implementation, whereas it'd be harder for any user-land solutions to do this, no matter how popular they already are. And this seems to be an even better motivation to provide this in core because the existing popular user-land solutions with ~25M weekly downloads already monkey patch
Module.prototype._compilein a way that completely drops the policy assertions. We have already explicitly stated that we trust the file system when loading a module in our threat model, anyway. In addition we have features likeNODE_EXTRA_CA_CERTS/SSL_CERT_DIR/NODE_REPL_EXTERNAL_MODULE/NODE_ICU_DATAetc., and they are probably much easier/straight-forward to exploit compared to code caches (even in those cases, exploits that depend on altering the inputs to those environment variables are already out of our security scope, because again we simply trust the file system).If the feature is opt-in, the mitigation against any new-found venerability that actually is in our security scope (even though I think that'd be unlikely given the reasons stated above) would also be simple - the user can just stop using it (e.g. unset the environment variable), and we can quickly make a security release by making it a noop until the vulnerability is addressed, and it shouldn't result in behavioral regression - the regression would only be on the module loading performance.
Reacted by Chengzhong Wu, Sean Larkin and JòanOne obvious angle is running node as setuid root, or otherwise running with elevated privileges (ex. capabilities on Linux.)
@bnoordhuis Where is the threat coming from in those cases? How would this be different compared to e.g.
NODE_EXTRA_CA_CERTS/SSL_CERT_DIR/NODE_REPL_EXTERNAL_MODULE/NODE_ICU_DATA?We're being careful to ignore those environment variables when running as setuid root or a host of other things. That same caution should be applied when reading files from disk.
We're being careful to ignore those environment variables when running as setuid root or a host of other things. That same caution should be applied when reading files from disk.
Yes I agree though I don't see this as a new threat - I think whatever we need is probably already covered by
SafeGetEnvand if there's something missing, we should fixSafeGetEnvfor all these variables, and I doubt environment variables for on-disk code cache should be treated any differently compared to these sensitive variables in this regard.P.S.: we don't use
SafeGetEnvfor all the variables I mentioned above...maybe we should...P.P.S.: On the other hand we have this popular package with ~25M weekly downloads in the ecosystem that does the something similar without taking elevated privileges into account...I would say implementing it in core could probably help improving the situation with the potential threat in the ecosystem
P.P.P.S.: we should probably consider exposing
safeGetEnvto user-land for this purpose.46 remaining items
- added a commit that references this issue
on Apr 22, 2024 - added a commit that references this issue
on Apr 29, 2024 - added a commit that references this issue
on Jul 30, 2024 - added a commit that references this issue
on Aug 8, 2024
This stemmed from a Twitter thread. Specifically I am wondering if there are any concerns over having something similar to what https://github.com/zertosh/v8-compile-cache does in core, the general idea is:
This is also similar to what Chrome does with the V8 code cache.
The motivation for implementing this in core is that, for a user-land module to do this for CJS, it has to monkey patch the CJS loader, and this increases the compatibility burden (v8-compile-cache has 17M weekly downloads, and it needs to monkey-patch
Module.prototype._compileto work. From a glance of its issue tracker it seems some of the issues are not really fixable in the user land either, like piping into the internal source maps cache). For ESM currently the user land can only use--loaderto customize the compilation, which has a cost of its own (especially when we move it to a separate thread), creating a disparity from CJS, and also even with--loaderI doubt if user-land code can integrate into e.g. the source map cache without asking us to expose too much internals.The most risky part of this feature might be the growth of the cache, but it seems manageable if:
This doesn't seem too radical, for example we already persist something like the repl history by default, and we also have features like
NODE_V8_COVERAGEthat does a similar "writing a lot of data to a directory when enabled" thing. I don't think this would increase the code complexity much either (we might also want a read-only version of this for SEA in the future too). So opening this issue to see if there are any concerns about having this in core before implementing it.