Repository navigation
module: ESM loaders next steps #36396
Description
Activity
- addedesmIssues and PRs related to the ECMAScript Modules implementation.Issues and PRs related to the ECMAScript Modules implementation.
on Dec 5, 2020 Combining load and transform into a single hook makes me very uncomfortable. It seems like this will result in transform hooks being silently skipped by load hooks if chained in the wrong order.
It seems like this will result in transform hooks being silently skipped by load hooks if chained in the wrong order.
Can you elaborate on this concern? It would be possible to have transform as a 2nd pass but it would create two competing ways to write a compiling loader which seems confusing.
It would also mean that it's really hard to write a loader that can be 100% certain that the code they generate for a (potentially virtual/in-memory) module definitely runs without modifications. Because even if its load hook runs first, some other loader's transform hook may mess with the generated code. And it would require a careful dance to undo that with a transform hook.
- addeddiscussIssues opened for discussion and feedback.Issues opened for discussion and feedback.
on Dec 5, 2020 It seems like this will result in transform hooks being silently skipped by load hooks if chained in the wrong order.
Can you elaborate on this concern? It would be possible to have transform as a 2nd pass but it would create two competing ways to write a compiling loader which seems confusing.
The example https loader has a code branch which does not call
next. In this case if next were babel or ts-node the code would not get transpiled as expected by the user.As far as having two ways to transform a module technically yes but IMO it would be incorrect to implement a transform with the load hook (a bug in that loader).
It would also mean that it's really hard to write a loader that can be 100% certain that the code they generate for a (potentially virtual/in-memory) module definitely runs without modifications. Because even if its load hook runs first, some other loader's transform hook may mess with the generated code. And it would require a careful dance to undo that with a transform hook.
I might be misunderstanding on this point, I would expect all
--loaders to be loaded before any are activated. Will that not be the case?The example https loader has a code branch which does not call
next. In this case if next were babel or ts-node the code would not get transpiled as expected by the user.I think this might be a source of confusion. In that example, the HTTPS loader’s
nextrefers to Node’s internalload—the one that loads source code from disk. Since the URL in question is anhttpsURL, there’s no point in asking Node to try to load it: Node would just error. Note that after theif (url.startsWith('https://')) {block, the loader does callnext, to let Node handle all non-httpsURLs.So if this were the only loader in use, e.g.
node --loader https-loader.mjs file.mjs, any HTTPS URLs infile.mjswould be loaded by the HTTPS loader and non-HTTPS URLs would be loaded by Node, via the call tonextat the end of the HTTPS loader’sloadhook. So basically the HTTPS loader’sloadhook runs first, and sometimes calls Node’sloadhook. The HTTPS loader is therefore always run first, before the equivalent Node hook; this is how the non-chainable loaders currently work (wherenextis currently the equivalent ofdefaultLoad, a reference to Node’s internal version of the hook).When the CoffeeSript loader is added to the chain, its
nextis the HTTPS loader. And if you look at CoffeeScript’sloadhook, it callsnextright at the start: it wants the source as returned by Node’sloador by any other loaders between CoffeeScript and Node (such as the HTTPS loader). So for example:node --loader coffeescript-loader.mjs --loader https-loader.mjs file.coffee- CoffeeScript’s
loadruns first. Its call tonextruns the HTTPS loader’sload. - Within the HTTPS loader’s
load, HTTPS URLs are processed and the source/format returned; and for other URLsnextis called, which runs Node’sloadto ask Node to provide source/format. Either way, the source/format are returned to the CoffeeScript loader as the return value for the CoffeeScript loader’snextcall. - Back in the CoffeeScript
load, we then transpile the source and return the source/format. Since this is the first/outermost loader, this is the “final” return value and what Node then takes to evaluate.
Every loader hook has to return the expected value, whether a string URL for
resolveor{ source, format }forload. If it doesn’t, Node would error. So there’s no chance of a loader author shipping a loader that breaks earlier loaders (HTTPS loader breaking the CoffeeScript loader, in this case). A loader author could short-circuit and prevent later loaders from ever getting run, like how in this example the HTTPS loader sometimes prevents Node’sloadfrom ever evaluating. It depends on the loader whether this is appropriate; basically, if the input tonextis something that could be successfully processed by Node’s version of the hook, the loader probably should callnext(which might be Node’s version or it might be some other compatible loader’s) and modify that, rather than duplicating what Node can do and/or short-circuiting unnecessarily. But if the arguments tonextwould throw when passed to Node, as they would for the HTTPS loader here, the loader shouldn’t callnextwith them because then that would just cause errors.So
nextin this case refers to the next registered loader, as in CoffeeScript referring to HTTPS referring to Node, where Node is always the last loader. Does this make sense?And because the calls are explicit—the chaining happens via calls to
next, not implicitly by running functions in sequence—that’s why we can’t really have a separatetransformSource, as far as I can tell. In atransformSourcehook,nextwould need to be either the next loader’sgetSourceortransformSource; it would get confusing pretty fast.- CoffeeScript’s
When the CoffeeSript loader is added to the chain, its
nextis the HTTPS loader. And if you look at CoffeeScript’sloadhook, it callsnextright at the start: it wants the source as returned by Node’sloador by any other loaders between CoffeeScript and Node (such as the HTTPS loader). So for example:node --loader coffeescript-loader.mjs --loader https-loader.mjs file.coffee-
CoffeeScript’s
loadruns first. Its call tonextruns the HTTPS loader’sload. -
Within the HTTPS loader’s
load, HTTPS URLs are processed and the source/format returned; and for other URLsnextis called, which runs Node’sloadto ask Node to provide source/format. Either way, the source/format are returned to the CoffeeScript loader as the return value for the CoffeeScript loader’snextcall. -
Back in the CoffeeScript
load, we then transpile the source and return the source/format. Since this is the first/outermost loader, this is the “final” return value and what Node then takes to evaluate.
My specific concern is what happens when someone (or something) runs:
node --loader https-loader.mjs --loader coffeescript-loader.mjs file.coffeeThis would cause import of
https://...file.coffeeto fail due to the coffeescript-loader.mjs being skipped. Having https-loader.mjs providegetSourceand coffeescript-loader.mjs providetransformSourcewould eliminate this type of error. In this example the damage might be somewhat limited since it would cause a clear failure. In the nyc use case we would simply fail to capture coverage but the code would still function. At most they would get a coverage threshold error.This becomes more difficult to defend against when you consider
NODE_OPTIONS="--loader https-loader.mjs".So
nextin this case refers to the next registered loader, as in CoffeeScript referring to HTTPS referring to Node, where Node is always the last loader. Does this make sense?And because the calls are explicit—the chaining happens via calls to
next, not implicitly by running functions in sequence—that’s why we can’t really have a separatetransformSource, as far as I can tell. In atransformSourcehook,nextwould need to be either the next loader’sgetSourceortransformSource; it would get confusing pretty fast.nextgiven toresolvealways refers to the nextresolve. The same would be true ofnextgiven togetSourceortransformSource, each stage of the hooks complete before any part of the following stage executes. So you are proposing 2 stages - resolve and load. I'm suggesting we need three stages - resolve, retrieve and transform.-
Should we convert this issue to a discussion?
Should we convert this issue to a discussion?
We could, but this is also an issue in that it has TODO items that pull requests should address (eventually to close the issue). I guess we could create a separate discussion related to this issue? My understanding (correct me if I’m wrong) is that once this becomes a discussion it can’t be converted back into an issue.
Having https-loader.mjs provide
getSourceand coffeescript-loader.mjs providetransformSourcewould eliminate this type of error.If you were to refactor the above
coffeescript-loader.mjsinto agetSource/transformSourcemodel, note that CoffeeScript loader would still need to provide agetSourcehook, in order to returnformat. And since HTTPS loader doesn’t callnextfor HTTPS links, this format returned by the CoffeeScript loader undernode --loader https-loader.mjs --loader coffeescript-loader.mjs file.coffeewould just get lost. So the different API doesn’t solve this problem: either version fails if the loaders are registered in the wrong order.Keep in mind that the reason we’re moving the
getFormatwork intoloadis because sometimes we need the source in order to determine the format. Most (any?) transpilation loaders therefore would need to implementload, to define a format for their custom file types (.ts,.jsx, etc.). Once they do that, it’s natural for them to also want to transform the source in the same function, as there’s a lot less uncertainty there than if they interact with a separate chain oftransformhooks.Having https-loader.mjs provide
getSourceand coffeescript-loader.mjs providetransformSourcewould eliminate this type of error.If you were to refactor the above
coffeescript-loader.mjsinto agetSource/transformSourcemodel, note that CoffeeScript loader would still need to provide agetSourcehook, in order to returnformat. And since HTTPS loader doesn’t callnextfor HTTPS links, this format returned by the CoffeeScript loader undernode --loader https-loader.mjs --loader coffeescript-loader.mjs file.coffeewould just get lost. So the different API doesn’t solve this problem: either version fails if the loaders are registered in the wrong order.Keep in mind that the reason we’re moving the
getFormatwork intoloadis because sometimes we need the source in order to determine the format. Most (any?) transpilation loaders therefore would need to implementload, to define a format for their custom file types (.ts,.jsx, etc.). Once they do that, it’s natural for them to also want to transform the source in the same function, as there’s a lot less uncertainty there than if they interact with a separate chain oftransformhooks.Could we allow
transformSourceto also return/alter a format? I'm not against breaking changes to the transformSource hook if that's what it would take to keep it separate. I do not think thetransformSourceshould even have an argument for thenextfunction, node.js should unconditionally call all registered transformation hooks in sequence. My ultimate desire is that the only way for another hook earlier in the chain to prevent mytransformSourcefrom being called is for the earlier hook to throw (abort the import).Also sorry for my slow response, offline stuff has taken over most of my time lately.
I do not think the
transformSourceshould even have an argument for thenextfunction, node.js should unconditionally call all registered transformation hooks in sequence.My concern with this is that now we have two patterns of chaining, with the
nextstyle in addition to this alternate one. What would this even look like in practice? Thetransformfunction receives source as input, and needs to return source as output—and if it fails to return, that’s akin to failing to callnext? Is the nexttransformfunction called with undefined input, or does it get the input that the previoustransformdidn’t transform? Is the benefit that alltransformfunctions are always called worth the cost of understanding this very different pattern of how chaining works fortransformseparate fromload?Also are we running all the
loadfunctions first, then all thetransformfunctions? So the final source returned at the end of the chainedloadfunctions would be passed into the firsttransformfunction? Or would they be called in pairs, like the output from the first loader’sloadis passed into the first loader’stransform, and then that output is passed into the second loader’sload, and so on?Basically, there are a lot of questions to be worked out when designing the API for how a separate
transformwould work, now that we have chaining. That’s the short version of what I ran into when Jan and I designed the new API in the top post. I’m not opposed to adding a separatetransformhook, but it seems like something that can come later as a follow-up PR to the PRs suggested in the top post; it doesn’t require any changes toload, sotransformcan be added cleanly afterward. I would suggest that we build what’s outlined above first, which will probably undergo changes once actual code is written, and once we see how chained loaders actually turn out then we can revisittransform.I do not think the
transformSourceshould even have an argument for thenextfunction, node.js should unconditionally call all registered transformation hooks in sequence.My concern with this is that now we have two patterns of chaining, with the
nextstyle in addition to this alternate one. What would this even look like in practice? Thetransformfunction receives source as input, and needs to return source as output—and if it fails to return, that’s akin to failing to callnext? Is the nexttransformfunction called with undefined input, or does it get the input that the previoustransformdidn’t transform?Would it be better for node.js to throw a TypeError if a transform gave an undefined return, reference the URL of the offending hook in the message? Honestly how node.js handles undefined return isn't hugely important to me as long as it's documented. I'll never return undefined from a transform and someone else returning undefined from a transform will not cause my hook to be skipped silently.
import CoffeeScript from 'coffeescript'; // CoffeeScript files end in .coffee, .litcoffee or .coffee.md const extensionsRegex = /\.coffee$|\.litcoffee$|\.coffee\.md$/; export function transform(previousResult, context) { // The first check is technically not needed but ensures that // we don’t try to compile things that already _are_ compiled. if (previousResult.format === undefined && extensionsRegex.test(context.url)) { // For simplicity, all CoffeeScript URLs are ES modules. const format = 'module'; const source = CoffeeScript.compile(previousResult.source, { bare: true }); return {format, source}; } // no action so return `previousResult` which came from // the `load` chain / previous `transform` functions. return previousResult; }
Is the benefit that all
transformfunctions are always called worth the cost of understanding this very different pattern of how chaining works fortransformseparate fromload?I think it is. For someone writing a transform hook having that hook skipped is a bug. If I write a transform hook that gets silently skipped because of another hook that the end-user might not even be aware of this would make support more difficult. Keep in mind
--loaderis not just for end users, it will be injected by tooling. This can be done through process arguments orNODE_OPTIONSenvironment, nyc for example will eventually add a loader toNODE_OPTIONSfor child processes (as it currently does to inject a--requireoption).Also are we running all the
loadfunctions first, then all thetransformfunctions? So the final source returned at the end of the chainedloadfunctions would be passed into the firsttransformfunction? Or would they be called in pairs, like the output from the first loader’sloadis passed into the first loader’stransform, and then that output is passed into the second loader’sload, and so on?Yes, the
loadchain would run first to completion, then thetransformchain would run starting with the final result of theloadchain.Basically, there are a lot of questions to be worked out when designing the API for how a separate
transformwould work, now that we have chaining. That’s the short version of what I ran into when Jan and I designed the new API in the top post. I’m not opposed to adding a separatetransformhook, but it seems like something that can come later as a follow-up PR to the PRs suggested in the top post; it doesn’t require any changes toload, sotransformcan be added cleanly afterward. I would suggest that we build what’s outlined above first, which will probably undergo changes once actual code is written, and once we see how chained loaders actually turn out then we can revisittransform.I'm not strongly against this idea but I worry about loader hooks becoming stable without a separate transform hook.
Okay, my current state of thinking it through: I think it's possible to have a separate
transformhook like this, at the expense of making one edge case more complicated. It would effectively split the order constraints for different kinds of loaders into three. My current understanding of "in what order need the getSource hooks run" is:- "Isolated code generation": Any hook that needs complete control over specific specifiers. Think: low-level instrumentation hooks that generate virtual modules like
my-system:super-fragile-code. If anything transforms that generated code, it breaks (e.g. it needs access to certain constructors and transpilation would lead to invalid runtime behavior). - Code instrumentation, e.g. coverage.
- JS-to-JS transformations, e.g. babel. Or WASM-to-WASM transformation - intra-format transformation.
- Non-JS-to-JS transformation, e.g. tsc. Or non-WASM-to-WASM transformation - inter-format transformation.
- Additional resource loaders, e.g. load from zip file or HTTPS.
Adding
transformwould split this into:- Type 1: Hooks that offer both
transformandgetSource. - Type 2: Hooks that offer only
transform. Within this group, in this exact order:
- Code instrumentation.
- JS-to-JS.
- Non-JS-to-JS.
- Type 3: Hooks that offer only
getSource. Within this group, in this exact order:
- Additional resource loaders.
Type 1 hooks would have to added first. Type 2 and 3 hooks can be added interleaved but would have to be ordered within their categories. A JS-to-JS hook could be added before or after a "additional resource" hook but would have to be listed after any code instrumentation hooks and before all non-js-to-js hooks.
In other words: Code instrumentation could still not be blindly added as the first hook, it would have to be added after any and all type 1 hooks.
One consequence of splitting the precedence into these 3 groups is that type 1 hooks would be harder to write: In a system with just one pass, they could just return the exact source they require from
getSourceand otherwise delegate to the chain. In a two-pass system, type 1 loaders need to:- Return a placeholder for
getSource. This value will be ignored later and could be an empty string. - Return the actual source in their
transformstep, ignoring whatever the argument was.
In other words: Writing a loader that ignores coverage instrumentation is still possible, just more awkward. And the sorting requirements are a little relaxed but really just for a single case: A hook for additional protocols or resource loading behaviors has a little more flexibility in where in the chain it's specified.
- "Isolated code generation": Any hook that needs complete control over specific specifiers. Think: low-level instrumentation hooks that generate virtual modules like
These last few posts have made me more certain that we should complete the already-proposed work first before potentially adding a
transformhook. I addedtransformas a potential sixth PR after the five already in the list at top, so it’s part of the roadmap to at least be considered.It seems to me that
transformwill require a fair bit of work to scope out a design that will work for all the use cases we’re trying to enable; but not havingtransformmost likely won’t prevent any use cases, so its addition would be optional.The appeal of
transformseems to be that under @coreyfarrell’s proposal, it would always run, whereas under the plan forresolveandload, sometimes loaders get bypassed ifnextisn’t called. This is already achievable withouttransform, by assuring that no transforming loaders are placed in order after any loaders that fail to callnext. So while I see the appeal of this feature, it’s more of a convenience than something that enables previously unachievable functionality. In thenode --loader coffeescript-loader.mjs --loader https-loader.mjs file.coffeeexample above, the CoffeeScript loader is always run because it’s first. Any loader that always callsnextcould go ahead of it, and the CoffeeScript loader would still always run. It’s only the loaders that come after the HTTPS loader that are at risk—and even then, maybe that’s what we might want. I can imagine that there might be some combination of loaders where we have no choice but to put the transforming loader after a not-next-calling loader, and thereforetransformenables a previously-impossible scenario; but I feel like I need to see an example of what that use case would be and whytransformis the best solution for it. I think we need to get several more examples of use cases and example loaders to know what problems we’re trying to solve, and that should probably come after we’ve done the earlier work to build chainedresolveandloadand see what other issues we have with them.Reacted by Jan Potoms and Charles Samborski68 remaining items
Please 👍 if you are interested in attending and I will send you a link that day
We should be getting started in about 10 minutes. The link is below. Hope to see you all there!
https://meet.google.com/jod-busy-frb
@bengl @JakobJingleheimer @zackschuster @songkeys @GeoffreyBooth @d3x0r @giltayar @Qard @Flarna
The meeting ran for about an hour and a half with additional dialogue scheduled for next week.
Reacted by Jacob SmithFor those interested in the continued dialog, today's meeting has been rescheduled for next week at the same time.
We would like to have our conclusions from the last meeting reified in #37468, which is currently making good progress.
Oh, missed it again...some proper calendar invites would help. 🤔
Reacted by Jacob SmithNo problemo @Qard, hopefully a Google Calendar invitation will help everyone with adding it to their calendars.
📅 Event: Node.js Module Loaders Working Group Meeting 2021-04-23
If there is some other calendar invitation format you had in mind, please let me know.
Reacted by Jacob SmithCould not find the requested event.
Hm, here is the whole calendar then. I've made it public, so it should work.
Reacted by Jacob Smith- addedloaders-agendaIssues and PRs to discuss during Loaders team meetings.Issues and PRs to discuss during Loaders team meetings.loadersIssues and PRs related to ES module loaders.Issues and PRs related to ES module loaders.
on Apr 20, 2021 Hey folks. I’ve been working on setting us up as a proper team akin to the Modules team. We now have a repo: https://github.com/nodejs/loaders.
There’s also a group: @nodejs/loaders. Please respond with 🚀 if you’d like me to add you to this group, and by extension to the new Loaders team. You’ll start receiving GitHub notifications whenever that group tag is used.
As for meetings, the way those work is that we need to use the Node Zoom account which records and streams to YouTube. I’ve asked for access to this, which won’t come until next week at the earliest. So I think we should probably postpone this week’s meeting (sorry). Also, no two Node teams’ meetings can overlap, because the Zoom/YouTube accounts can only stream one meeting at a time. The Friday 15:00 UTC meeting time overlaps with another team’s meeting, so we need to find a new regular time. Here’s the calendar of Node team meetings. There’s a gap two hours later, at 17:00 UTC / 1 pm ET / 10 am PT; I was thinking we could do this time every two weeks, starting on Friday April 30. If this works for you, please reply with 👍 , otherwise 👎 . If you can’t attend and you’d like to, please login to node-js.slack.com and discuss on the
#esmchannel. Depending on how that discussion goes I might create a Doodle to find a new time.Thanks!
Reacted by Jacob Smith, Bryan English, Bradley Farias, Derek Lewis and Michaël ZassoReacted by Jo and Jacob SmithReacted by Jacob Smith, Bryan English, Bradley Farias, Derek Lewis, Gil Tayar, Michaël Zasso, Stephen Belanger, Zack Schuster, Dominik Gruber, Gerhard Stöbich and 3 moreHey @nodejs/loaders, everyone who 🚀 ‘ed the previous message should be in @nodejs/loaders now. Please also take a look at https://github.com/nodejs/loaders.
I’m still waiting on the Zoom credentials, so I guess we can’t meet tomorrow. I’ll propose a new time once I have the credentials, probably at the same time either a week or two weeks from tomorrow. Following the pattern of other Node teams, the meetings will be announced via issues in the team repo; you should be notified via the @nodejs/loaders tag and if you watch the repo. Issues and PRs on any
nodejsrepo taggedloaders-agendawill be flagged for discussion at the meetings.Reacted by Jacob Smith and Zack Schuster@GeoffreyBooth if the issue is recording the meeting, could we use a google Meet or does it have to be zoom? I have an early adopter account/organisation for gSuite, so I get most of the paid features for free; I remember Meet switched meeting recording to a premium feature a while back—I think I still get it (I can check).
Update: I can't find "recording" in settings (it must be enabled now apparently), so I'm not sure. "G Suite Legacy" is never listed in any of the support documentation, so I dunno if I'm supposed to have it (best I can find is a list of general features for G Suite Legacy).
Hey @nodejs/loaders, I should have everything we need for our next meeting, which will be at 17:00 UTC / 1 pm ET / 10 am PT on May 14 (not tomorrow). nodejs/loaders#1 was a test of the meeting infrastructure, that the scripts that look at the Google calendar would generate the meeting agenda and so on properly, and everything appears to be working. Sorry for all the delays and confusion, and I look forward to seeing all of you next week!
Reacted by Michaël Zasso, Derek Lewis and Jacob SmithReacted by Michaël Zasso, Derek Lewis and Jacob Smith- removedloaders-agendaIssues and PRs to discuss during Loaders team meetings.Issues and PRs to discuss during Loaders team meetings.
on May 28, 2021 Closing this issue as discussion has moved to https://github.com/nodejs/loaders
This issue is meant to be a tracking issue for where we as a team think we want ES module loaders to go. I’ll start it off by writing what I think the next steps are, and based on feedback in comments I’ll revise this top post accordingly.
I think the first priority is to finish the WIP PR that @jkrems started to slim down the main four loader hooks (
resolve,getFormat,getSource,transformSource) into two (resolveToURLandloadFromURL, or should they be calledresolveandload?). This would solve the issue discussed in #34144 / #34753.Next I’d like to add support for chained loaders. There was already a PR opened to achieve this, but as far as I can tell that PR doesn’t actually implement chaining as I understand it; it allows the
transformSourcehook to be chained but not the other hooks, if I understand it correctly, and therefore doesn’t really solve the user request.A while back I had a conversation with @jkrems to hash out a design for what we thought a chained loaders API should look like. Starting from a base where we assume #35524 has been merged in and therefore the only hooks are
resolveandloadandgetGlobalPreloadCode(which probably should be renamed to justglobalPreloadCode, as there are no longer any other hooks namedget*), we were thinking of changing the last argument of each hook fromdefault<hookName>tonext, wherenextis the next registered function for that hook. Then we hashed out some examples for how each of the two primary hooks,resolveandload, would chain.Chaining
resolvehooksSo for example say you had a chain of three loaders,
unpkg,http-to-https,cache-buster:unpkgloader resolves a specifierfooto an urlhttp://unpkg.com/foo.http-to-httpsloader rewrites that url tohttps://unpkg.com/foo.cache-busterthat takes the url and adds a timestamp to the end, so likehttps://unpkg.com/foo?ts=1234567890.These could be implemented as follows:
unpkgloaderhttp-to-httpsloadercache-busterloaderThese chain “backwards” in the same way that function calls do, along the lines of
cacheBusterResolve(httpToHttpsResolve(unpkgResolve(nodeResolve(...))))(though in this particular example, the position ofcache-busterandhttp-to-httpscan be swapped without affecting the result). The point though is that the hook functions nest: each one always just returns a string, like Node’sresolve, and the chaining happens as a result of callingnext; and if a hook doesn’t callnext, the chain short-circuits. I’m not sure if it’s preferable for the API to benode --loader unpkg --loader http-to-https --loader cache-busteror the reverse, but it would be easy to flip that if we get feedback that one way is more intuitive than the other.Chaining
loadhooksChaining
loadhooks would be similar toresolvehooks, though slightly more complicated in that instead of returning a single string, eachloadhook returns an object{ format, source }wheresourceis the loaded module’s source code/contents andformatis the name of one of Node’s ESM loader’s “translators”:commonjs,module,builtin(a Node internal module likefs),json(with--experimental-json-modules) orwasm(with--experimental-wasm-modules).Currently, Node’s internal ESM loader throws an error on unknown file types:
import('file.javascript')throws, even if the contents of that file are perfectly acceptable JavaScript. This error happens during Node’s internalresolvewhen it encounters a file extension it doesn’t recognize; hence the current CoffeeScript loader example has lots of code to tell Node to allow CoffeeScript file extensions. We should move this validation check to be after the format is determined, which is one of the return values ofload; so basically, it’s onloadto return aformatthat Node recognizes. Node’s internalloaddoesn’t know to resolve a URL ending in.coffeetomodule, so Node would continue to error like it does now; but the CoffeeScript loader under this new design no longer needs to hook intoresolveat all, since it can determine the format of CoffeeScript files withinload. In code:coffeescriptloaderAnd the other example loader in the docs, to allow
importofhttps://URLs, would similarly only need aloadhook:httpsloaderIf these two loaders are used together, where the
coffeescriptloader’snextis thehttpsloader’s hook andhttpsloader’snextis Node’s native hook, so likecoffeeScriptLoad(httpsLoad(nodeLoad(...))), then for a URL likehttps://example.com/module.coffee:httpsloader would load the source over the network, but returnformat: undefined, assuming the server supplied a correctContent-Typeheader likeapplication/vnd.coffeescriptwhich ourhttpsloader doesn’t recognize.coffeescriptloader would get that{ source, format: undefined }early on from its call tonext, and setformat: 'module'based on the.coffeeat the end of the URL. It would also transpile the source into JavaScript. It then returns{ format: 'module', source }wheresourceis runnable JavaScript rather than the original CoffeeScript.Chaining
globalPreloadCodehooksFor now, I think that this wouldn’t be chained the way
resolveandloadwould be. This hook would just be called sequentially for each registered loader, in the same order as the loaders themselves are registered. If this is insufficient, for example for instrumentation use cases, we can discuss and potentially change this to follow the chaining style ofload.Next Steps
Based on the above, here are the next few PRs as I see them:
resolve,loadandglobalPreloadCode.resolveandload. Node’s internal loader already has no-ops fortransformSourceandgetGlobalPreloadCode, so all this really entails is merging the internalgetFormatandgetSourceinto one functionload.resolve(on detection of unknown extensions) to withinload(if the resolved extension has no defined translator).default<hookName>becomesnextand references the next registered hook in the chain.loadreturn value offormat: 'commonjs'to work, or at least error informatively. See esm: Modify ESM Experimental Loader Hooks #34753 (comment).transformhook (see below).This work should complete many of the major outstanding ES module feature requests, such as supporting transpilers, mocks and instrumentation. If there are other significant user stories that still wouldn’t be possible with the loaders design as described here, please let me know. cc @nodejs/modules