镜像站点 · 本页由第三方 GitHub 只读镜像提供,非 GitHub 官方站点,不接受任何登录或凭据输入。前往 github.com
Skip to content
This repository was archived by the owner on Sep 2, 2023. It is now read-only.
This repository was archived by the owner on Sep 2, 2023. It is now read-only.

Loader Hooks #351

Description

@reasonablytall

Hooking into the dependency loading steps in Nodejs should be easy, efficient, and reliable across CJS+ESM. Loader hooks would allow for developers to make systematic changes to dependency loading without breaking other systems.

It looks like discussion on this topic has died down, but I'm really interested in loader hooks and would be excited to work on an implementation! There's of prior discussion to parse through, and with this issue I'm hoping to reignite discussion and to create a place for feedback.

Some of that prior discussion:


edit (mylesborins)

here is a link to the design doc

https://docs.google.com/document/d/1J0zDFkwxojLXc36t2gcv1gZ-QnoTXSzK1O6mNAMlync/edit#heading=h.xzp5p5pt8hlq

Activity

  1. reasonablytall commented on Jul 17, 2019

    @reasonablytall
    ContributorAuthor

    Some use cases I've encountered:

    I'm working on a custom dependency bundler and loader designed to improve cold-start startup times by transparently loading from a bundle to avoid file-system overhead. Currently, I have to monkey-patch module and reimplement CJS resolution with @soldair's node-module-resolution. I have to deeply understand and often reimplement CJS+ESM internals to work on this.

    I also want to load modules from V8 code-cache similar to v8-compile-cache. Again I have to re-implement Module._compile and manually handle fallback for other extensions.

    Some other use cases that would benefit:

    • Transpilers/compileres (babel, ts-node)
    • Bundlers (webpack, browserify, ncc)
    • Other custom loaders (pnp, tink, loading from zip, etc)
    • Code instrumentation (coverage, logging, testing)
  2. devsnek commented on Jul 17, 2019

    @devsnek
    Member

    I think the current exposed hooks are the right hooks to expose, but we definitely need to work on polishing the API:

    • dynamic modules are a bit rough atm
    • how are hooks registered
    • how do multiple registered hooks interact
  3. MylesBorins commented on Jul 17, 2019

    @MylesBorins
    Contributor

    Very excited to see interest in this! I believe that @bmeck has a POC that has memory leaks that need to be fixed. @guybedford may know about this too

  4. hybrist commented on Jul 17, 2019

    @hybrist
    Contributor

    @devsnek I think there's more things that list is missing. E.g. providing resource content, not just format. Or the question of if we can extend aspects of this feature to CommonJS (e.g. for the tink/entropic/yarn case that currently requires monkey-patching the CommonJS loader or even the fs module itself). The current hooks were a good starting point but I would disagree that they are the right hooks.

  5. devsnek commented on Jul 17, 2019

    @devsnek
    Member

    @jkrems i think cjs loader hooks are outside the realm of our design. cjs can only deal with local files and it uses filenames, not urls.

    Providing resource content is an interesting idea though. I wonder if we could just figure out a way to pass vm modules to the loader.

  6. bmeck commented on Jul 17, 2019

    @bmeck
    Member

    @devsnek we discussed and even implemented a PoC of intercepting CJS in the middle of last year and had a talk on the how/why in both

    These would only allow for files for require since that is what CJS works with but it should be tenable. Interaction with require.cache is a bit precarious but solvable if enough agreement can be reached.

  7. devsnek commented on Jul 17, 2019

    @devsnek
    Member

    @bmeck i don't doubt it can be done, i'm just less convinced it makes sense to include with the esm loader hooks given the large differences in the systems.

  8. guybedford commented on Jul 17, 2019

    @guybedford
    Contributor

    @a-lxe thanks for opening this discussion. It was interesting to hear you say that multiple loaders were one of the features you find important here. The PR at nodejs/node#18914 could certainly be revived. Is this something you are thinking of working on? I'd be glad to collaborate on this work if you would like to discuss it further at all.

  9. reasonablytall commented on Jul 17, 2019

    @reasonablytall
    ContributorAuthor

    @guybedford Yeah! At least to me it seems right now the singular --loader api is insufficient for current loader use cases. For example in my projects I test with ts-node istanbul, mocha, and source-map-support -- each of which hooks into loading in one way or another IIRC. Optimally these could each independently interface with a loader hook api and smoothly compound on each other.

    I think a node loader hook api needs to provide mechanisms for compounding on and falling back to already registered hooks (or the default cjs/esm behavior). I'm not really sure yet where to focus work, but I definitely want to collaborate :)

  10. guybedford commented on Jul 17, 2019

    @guybedford
    Contributor

    @a-lxe agreed we need a way to chain loaders. Would the approach in nodejs/node#18914 work for you, or if not, how would you want to go about it differently? One way to start might be to get that rebased and working again and then to iterate on it from there.

  11. reasonablytall commented on Jul 17, 2019

    @reasonablytall
    ContributorAuthor

    @guybedford I like the way nodejs/node#18914 chains the loaders and provides parent to allow fallback/augmentation of both the resolve + dynamic instantiation steps. I have some ideals for what a loader hook api should look like (particularly wrt supporting cjs) but I don't think those should get in the way of providing multiple --loader for esm. To be honest working on reviving that PR would be really useful for me in getting up to speed with things, so I would be happy to get started on that.

    Some gripes which are more relevant to the initial --loader implementation rather than the multiple --loader feature:

    • Why is there no runtime api for registering loaders? The current mechanisms using --require to register loaders via preloaded modules feel nice to me, and allow for opting to manually register at a particular point and with dynamic parameters.
    • A loader has to implement both hooks (and add fallback overhead) even if it only affects one.
    • Similarly, I feel like the hooks could be more granular. @bmeck 's Resource APIs for Node splits things into locate, retrieve, and translate hooks. With an additional initialize hook for actually creating the module, these match the Allow for multiple --loader flags node#18914 functionality, with an added bonus that only initialize needs to have coupling with cjs/esm. I'm curious what you think on this.
    • Doesn't hook into cjs require :'(

    Also, your last comment on nodejs/node#18914 hints at another loaders implementation by @bmeck. Does this exist in an actionable state?

    @BridgeAR this work also exists as part of the new loaders work which @bmeck started, so that effectively takes over from this PR already. Closing sounds sensible to me.

  12. guybedford commented on Jul 17, 2019

    @guybedford
    Contributor

    Why is there no runtime api for registering loaders?

    Loaders are a higher-level feature of the environment, kind of like a boot system feature. They sit at the root of the security model for the application, so there are some security concerns here. In addition to that, hooking loaders during runtime can lead to unpredictable results, since any already-loaded modules will not get loaders applied. I'm sure @bmeck can clarify on these points, but those are the two I remember on this discussion offhand.

    A loader has to implement both hooks (and add fallback overhead) even if it only affects one.

    There is nothing to say we won't have CJS loader hooks or a generalized hook system, but it's just that our priority to date has been getting the ESM loader worked out. In addition the ESM hooks allow async functions, while CJS hooks would need some manipulation to support async calls. There's also the problem of the loaders running in different resolution spaces (URLs v paths) as discussed. Once we have our base ESM loader API finalized I'm sure we could extend it to CJS with some extra resolution metadata and handling of the resolution spaces, but I very much feel that loader unification is a "nice to have" that is additive over the base-level ESM API which should be the priority for us to consolidate and work towards first. That loader stability and architecture should take preference in the development process. That said, if you want to work on CJS unification first, feel free, but there are no guarantees the loader API will be stable or even unflagged unless we work hard towards that singular goal right now. So what I'm saying is chained loaders, whether the loader is off-thread, whether the API will be abstracted to deal with multi-realm and non-registry based API, and the translate hook all take preference in the path to a stable API to me, overy unifying ESM and CJS hooks. And that path is already very tenuous and unlikely, so that we should focus our combined efforts on the API stability first and foremost.

    Similarly, I feel like the hooks could be more granular.

    Implementing a translate or fetch hook for --loader could certainly be done and was a deliberate omission in the loader API. It is purely a problem of writing the code, making a PR, and the real hard part - getting consensus!

    Doesn't hook into cjs require :'(

    As mentioned above, this work can be done, but I would prefer to get the ground work done first.

  13. reasonablytall commented on Jul 17, 2019

    @reasonablytall
    ContributorAuthor

    That all makes a lot of sense and I appreciate you describing it for me 🙂

    I can start with pulling nodejs/node#18914 and getting that in a working state.

  14. GeoffreyBooth commented on Jul 23, 2019

    @GeoffreyBooth
    Member

    Just to spark some discussion, here’s a wholly theoretical potential API that I could imagine being useful to me as a developer:

    import { registerHook } from 'module';
    import { promises as fs, constants as fsConstants } from 'fs';
    
    registerHook('beforeRead', async function automaticExtensionResolution (module) {
      const extensions = ['', '.mjs', '.js', '.cjs'];
      for (let i = 0; i < extensions.length; i++) {
        const resolvedPathWithExtension = `${module.resolvedPath}${extensions[i]}`;
        try {
          await fs.access(resolvedPathWithExtension, fsConstants.R_OK);
          module.originalResolvedPath = module.resolvedPath;
          module.resolvedPath = resolvedPathWithExtension;
          break;
        } catch {}
      }
      return module;
    }, 10);

    The new registerHook method takes three arguments:

    • The hook name, which is a point in Node’s code where these callbacks will be run. beforeRead, afterRead, etc.
    • The function to call at that point, which takes as input an object with all the properties related to import or require resolution and module loading that developers might want to override. Properties set on this object persist to other callbacks registered to later hooks in the process (e.g. module.foo set during beforeRead would be accessible in a different callback registered to afterRead).
    • (Optional): The priority to call registered callbacks. If multiple callbacks have the same priority level, they are evaluated in the order that they were registered.

    In the first example, my automaticExtensionResolution callback is registered to beforeRead because it’s important to rewrite the path that Node tries to load before Node tries to load any files from disk (because './file' wouldn’t exist but './file.js' might, and we don’t want an exception thrown before our callback can tell Node to load './file.js' instead). I’m imagining the module object here has an unused specifier property with whatever the original string was, e.g. pkg/file, and what Node would resolve that to in resolvedPath, e.g. ./node_modules/pkg/file.

    Another example:

    import { registerHook } from 'module';
    import CoffeeScript from 'coffeescript';
    
    registerHook('afterRead', async function transpileCoffeeScript (module) {
      if (/\.coffee$|\.litcoffee$|\.coffee\.md$/.test(module.resolvedPath)) {
        module.source = CoffeeScript.compile(module.source);
      }
      return module;
    }, 10);

    This hook is registered after Node has loaded the file contents from disk (module.source) but before the contents are added to Node’s cache or evaluated. This gives my callback a chance to modify those contents before Node does anything with them.

    And so on. I have no idea how close or far any of the above is from the actual implementation of the module machinery; hopefully it’s not so distant as to be useless. Most of the loader use cases in our README could be satisfied by an API like this:

    • Code coverage/instrumentation: In an afterRead hook, a callback could count lines of code or the like.
    • Runtime loaders, transpilation at import time: In an afterRead hook, example above.
    • Arbitrary sources for module source text: In a beforeRead hook, our callback could assign content into module.source (and then Node would know to not read from disk for this module).
    • Mock modules (injection): Basically the same as previous, a beforeRead hook could return a different path to load instead, or prefill the source code to use: if (module.specifier === 'request') module.source = mockRequest etc. Ideally Node would handle if source were an actual module rather than just a string to be evaluated.
    • Specifier resolution customization: This is like the extension resolution example above, though if we also want to support specifiers that Node can’t resolve, like import 'https://something', we would need another hook like beforeResolve.
    • Package encapsulation: We’re already implementing this as "exports", but it could just as easily be implemented as a loader, at least for ESM. It would be in beforeRead like the extension resolution example.
    • Conditional imports: In beforeRead or afterRead, based on some condition the module.source could be set to an empty string.

    Anyway this is just to start a discussion of what kind of public-facing API we would want, and the kind of use cases it would support. I’m not at all married to any of the above, I’m just hoping that we come up with something that has roughly the same versatility as this.

  15. 89 remaining items

  16. cspotcode commented on May 2, 2020

    @cspotcode

    Will there eventually be the ability to install loader hooks at runtime? Currently, CLI tools that need to bootstrap an execution environment which includes hooks have a tough time doing so in a cross-platform manner that doesn't impose extra performance overhead.

    It's possible for one node process to spawn another, but there are lots of caveats with that, and it's slower.

    My use-case is ts-node. Our normal interface is ts-node scripts.ts. This is compatible with Linux shebangs.

    Today we need to tell users to node --loader ts-node/esm script.ts. Ideally, users can ts-node script.ts, which launches one and only one node process, and we install hooks at runtime.

  17. guybedford commented on May 2, 2020

    @guybedford
    Contributor

    @cspotcode I'd suggest using ts-node to spawn a new Node.js process with the loader hooks set. There likely will be APIs to spawn subloaders in future in the same Node.js process, but that work has not yet begun. Contributions to loader work is also welcome.

  18. cspotcode commented on May 8, 2020

    @cspotcode

    I have another question / bit of feedback. I hope this is the right place to post. Technically it pertains to something require.extensions hooks must do to properly integrate with ESM.

    node's built-in require.extensions['.js'] implementation checks if the file should be treated as ESM. If so, it throws an error.

    ts-node attaches a custom require.extensions['.ts'] (and .tsx , .jsx, and overwrites the built-in .js hook as needed). We read the file from disk, compile .ts->.js, then pass this string to module._compile.

    Our hook needs to mimic the error-throwing behavior of node's built-in hook. For example, mocha relies on this behavior to correctly support both CommonJS hooks and ESM hooks simultaneously: https://github.com/mochajs/mocha/blob/master/lib/esm-utils.js#L10-L23

    require.extensions['.js'].toString()
    
    > require.extensions['.js'].toString()
    'function(module, filename) {\n' +
      "  if (filename.endsWith('.js')) {\n" +
      '    const pkg = readPackageScope(filename);\n' +
      "    // Function require shouldn't be used in ES modules.\n" +
      "    if (pkg && pkg.data && pkg.data.type === 'module') {\n" +
      '      const parentPath = module.parent && module.parent.filename;\n' +
      "      const packageJsonPath = path.resolve(pkg.path, 'package.json');\n" +
      '      throw new ERR_REQUIRE_ESM(filename, parentPath, packageJsonPath);\n' +
      '    }\n' +
      '  }\n' +
      "  const content = fs.readFileSync(filename, 'utf8');\n" +
      '  module._compile(content, filename);\n' +
      '}'
    

    My plan right now is a hack where I invoke the built-in .js hook, passing it a filename I know does not exist, but that is in the right directory. This costs a failed fs call:

    try {
      require.extensions['.js'](
        // make `module` object
        {_compile(){}},
        filename + 'DOESNOTEXIST.js' // I can make this more robust, appending a random UUID
      );
    } catch(e) {
      // Inspect the thrown error to see if it's an `ERR_REQUIRE_ESM` error
    }
    

    EDIT: actually, won't be using this hack. In the case another require hook has been installed before us, we can't be sure require.extensions['.js'] is going to be node's built-in hook. Instead I'm going to extract the relevant code from node's source into our codebase.

    Ideally node exposes its synchronous CJS/ESM classifier to support our use-case. I realize that require.extensions has been deprecated for years, but that doesn't strictly prohibit node from exposing this API.

  19. GeoffreyBooth commented on May 8, 2020

    @GeoffreyBooth
    Member

    This has been something I've been thinking about as well. I think for a loader to handle both ESM and CommonJS files, it also needs to hook into require.extensions (at least for now). I'd like to proxy/override or wrap require.extensions['.js'] for extensions that should be handled similarly, but without necessarily having the “throw if this is ESM” check. I wonder if there's a not-terrible way to prevent that, aside from copying the source of require.extensions['.js'] into my code.

    Another thing we might want to consider is making the ESM loader hooks simply the loader hooks, to apply to both CommonJS and ESM, finally replacing the deprecated require.extensions. I haven't given thought into how that would work in practice, but there are lots of use cases such as instrumentation where the loader should affect all files, not just ESM ones, and it's more work for loader authors to hook into both of Node's loaders in very different ways.

  20. cspotcode commented on May 8, 2020

    @cspotcode

    @GeoffreyBooth the problem with unifying the hooks is that require() must be synchronous.

    Right now, when a user installs ts-node's ESM hooks, we know they can't have installed any other hooks. Depending how you think about, we are responsible for implementing some basic features that would ideally be implemented by third-party ESM hooks.

    If a third-party library installs require.extensions['.coffee'], for example, do we need special logic in our hooks to resolve() .coffee files and getFormat them the same as .js files? We cannot pass them to defaultGetFormat because it doesn't understand .coffee files.

  21. GeoffreyBooth commented on May 24, 2020

    @GeoffreyBooth
    Member

    the problem with unifying the hooks is that require() must be synchronous.

    Hmmmm. That is a problem. @jkrems or @weswigham is there any hope here, or would hooks that work with CommonJS run into the same issues that Wes’ “require of ESM” PR did?

    I suppose one could write synchronous hooks? CoffeeScript transpilation is synchronous, for example; I dunno if TypeScript’s is? Obviously there couldn’t be a sync HTTP loader but there are plenty of useful loader cases that don’t need async.

    Anyway applying these hooks to CommonJS as well is a long-term maybe goal. AFAIK CommonJS wasn’t designed to be hookable/customizable, and the current ways people do it are monkey-patched hacks more or less. Perhaps it’s best left as is.

  22. bmeck commented on May 26, 2020

    @bmeck
    Member

    @GeoffreyBooth one of the motivations of moving loading off thread/isolated was we have proof of concept that we can use a SharedArrayBuffer to sleep the main thread while the loader does async tasks but looks as if it were blocking via Atomics.wait. It does not have the same effect as require of ESM.

  23. GeoffreyBooth commented on May 26, 2020

    @GeoffreyBooth
    Member

    one of the motivations of moving loading off thread/isolated

    What's the benefit of moving to a threaded loader in terms of user story? Do we get improved performance? Improved security? New functionality that wouldn't have been possible in a single-threaded loader?

  24. hybrist commented on May 27, 2020

    @hybrist
    Contributor

    Do we get improved performance?

    Decreased resource/memory usage is the most expected outcome, especially with multiple contexts or threads that all need hooks. If the APIs for hooks aren't designed to run in isolation, this would be hard or impossible to achieve in userland (since none of the hook implementations would be compatible with those assumptions by default). It's much more realistic to run one TSC or babel instance than 100 per process.

    Another aspect is increased stability: The loader can't accidentally be broken by or break the application, at least not as easily as when they run in the same global scope.

  25. bmeck commented on May 27, 2020

    @bmeck
    Member

    @GeoffreyBooth

    So we have 2 users effectively:

    1. application code
    2. loader code

    I think for application code not much would be visibly affected and/or varies to much per application to state much about it.

    I think for loader code there are a lot of pro/con to consider but I believe the pros outweigh the cons significantly. In theory, a portion of the pros could be left to users to do things like spin up workers inside of their own loader, but a variety of things are not feasible or have higher advantages if done by the runtime itself.

    The overall story is a bit complicated in terms of performance but overall I'd say for simple workflows that putting them on a thread is worse, but for complex workflows and applications it is better.

    • 😄 - loaders only need to be instantiated once instead of per thread/context
    • 😢 - spinning up a loader is much more costly (expect an extra 10ms)
    • 😄 - you can spin up multiple loaders at once
    • 😢 - loaders have to communicate using serialized messages, no object sharing
    • 😄 - doing CPU heavy work in a loader doesn't block the application/other loaders (if multiple loaders are ever allowed)
    • 😢 - you have to do duplicate work sometimes in threads

    Security has a bunch of discussions about what your security model is but for some simple statements that aren't really controversial:

    • 😄 - prototype pollution won't be a way to attack loader hook callbacks
    • 😄 - when auditing, don't have to worry about application code mutating references inside of loader code
    • 😄 - trust boundary is well defined by using a different context, so things like the permissions PR can apply different levels of trust to the loader vs application code
    • 😐 - generated code from a loader is still susceptible to attacks, same as status quo
    • 😐 - loader DOS E.G. ReDOS won't be a DOS for the application itself. Still a concern though.

    Per features that if loaders are not on same thread:

    • 😄 - blocking the main thread to do asynchronous operations
    • 😄 - you can share data to be reused across all threads for loading purposes (code/resolution cache)
    • 😢 - you cannot directly share object references with a loader and application code. This affects patterns like testdouble uses

    I do strongly think we need to solve the object reference problem, but we haven't really spent time looking into it to my knowledge. Even if we don't move off thread we likely need to solve the object reference problem in order to allow a solution for saving variables in getGlobalPreloadCode such that they can be referenced in code generated by a loader without being mutable/globals.

  26. just-boris commented on Jul 4, 2020

    @just-boris

    Another use-case for custom loaders. Stubbing CSS-imports in UI libraries.

    Currently there are similar solutions for CommonJS

    I tried to implement the same functionality using the new Node.js loader API: https://github.com/proxy/gist.github.com/just-boris/b07d66e306c94cf42db41b010231fbbf

    Works well for such cases.

  27. milahu commented on Sep 29, 2021

    @milahu

    solved the "machine-level store and symlinked node_modules" problem
    with LD_PRELOAD and nodejs-hide-symlinks ... like an early marriage of pnpm and /nix/store : )

  28. frank-dspeed commented on Dec 4, 2022

    @frank-dspeed

    Hi i am defining a new packaging module standard it is in Fact a ECMAScript Module based standard for packaging modules to get reused so i want to define some fundamentals and conventions on property to express importent meta like integrity related content hashes and the used hash algorithm and verification algorithm.

    how can i pull that off chained with the loader hooks api to maybe directly support the needed functions

    i call it web-module standard it is designed to form a world wide web scale shared p2p distributed content addressed module build cache and module based package system

    so Web 4.0 compose apps out of modules without even the need for nodejs any web platform works for that.

    and we can directly link remote contexts then no roundtrips via packaging are needed if we agree on a module standard that can get directly used even when the module is relativ Zero Trust. like it is with any npm package.

    what do you think?

  29. GeoffreyBooth commented on Sep 1, 2023

    @GeoffreyBooth
    Member

    Closing we we have https://github.com/nodejs/loaders now devoted to this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions