Repository navigation
memory leak in shadow realms #47353
Description
Activity
@nodejs/realm
Could this memory leak cause performance regression when we run URL tests? I'm seeing a different result in node benchmark compare results, compared to microbenchmarks with mitata.
- addedmemoryIssues and PRs related to Node.js memory management or memory footprint.Issues and PRs related to Node.js memory management or memory footprint.realmIssues and PRs related to the ShadowRealm API and node::Realm.Issues and PRs related to the ShadowRealm API and node::Realm.
on Apr 1, 2023 I think we would eventually need some way to "know" when a ShadowRealm can be collected / closed / destroyed to dispose of any connected libuv handles or other native resources.
Could this memory leak cause performance regression when we run URL tests? I'm seeing a different result in node benchmark compare results, compared to microbenchmarks with mitata.
I don't think so, this is only relevant in shadow realms, and I don't think we have URLs benchmarks in shadow realms (it's only supported starting from two weeks ago and not even released, and still it's behind
--experimental-shadow-realms). Principal realms are unaffected because...you only get one principal realm per Node.js instance, and we actively destroy them when there are no more tasks to run.ShadowRealm is currently behind the runtime flag
--experimental-shadow-realmor--harmony-shadow-realmand it should not affect performance in the principal realm.I think we would eventually need some way to "know" when a ShadowRealm can be collected / closed / destroyed to dispose of any connected libuv handles or other native resources.
If BaseObjects are weak and no active HandleWraps or ReqWraps, the ShadowRealms' context should be unreachable and can be GC-ed. However, as pointed out by @joyeecheung as option 2, this needs to avoid making BaseObjects strong. This can be particularly a problem when we need to nest BaseObjects with c++ member fields by v8 global handles as v8's garbage collector is unable to infer the references between c++ objects. As an alternative, the references can be saved as object internal fields, or as v8's cppgc
TracedReference.I also noticed that the current machinery for managing BaseObjects does not work well with the idea that "a context will become unreachable before the BaseObjects are destructed". For example I am trying to make BindingData weak, and I notice the they get deleted first via the cleanup hooks (as part of the destruction of the Realm) and then via the weak callbacks, whereas in the principal realm we simply assume that weak callbacks, if ever called, must always be called before the cleanup hooks (and we already have guard in the code for double-deletion in this case, but not the other one). There are probably more assumption about this throughout the code base. This also means, IIUC, none of the BaseObjects destructors should access the context or even modify JS states (not even setting the internal fields, that also crashes when the context is unreachable). This can be particularly problematic for complex objects like AsyncWraps (who...calls destroy hooks during destruction. Ouch).
I now wonder if it's too late to make the ShadowRealms API actively dispose themselves (e.g.
realm.dispose()) to avoid the leak. That would be the easiest solution, implementation-wise, to solve this problem. Probably harder TC39-process wise though (or it might not even be on the table - I am not sure if this constraint is just from our implementation).I now wonder if it's too late to make the ShadowRealms API actively dispose themselves (e.g. realm.dispose()) to avoid the leak. That would be the easiest solution, implementation-wise, to solve this problem. Probably harder TC39-process wise though (or it might not even be on the table - I am not sure if this constraint is just from our implementation).
From my point of view, it would be tough for us to deal with any native/asynchronous resources without a way to dispose of the realm deliberately. With the current semantics, if we open a socket inside a ShadowRealm, and the wrapper object is collected, what should we do? There are three options:
- keep the realm going until there are native resources allocated to it
- shut all the associated native resources down
- crash badly
With a
dispose()function, we could decide what behavior we want to put in place.if we open a socket inside a ShadowRealm
I think the ability of opening a socket is probably out of the scope in ShadowRealms ;). But even in those cases, those wrappers are usually weak so would not hold the context alive. The more problematic ones are e.g. caches (e.g. for module imports and serialization, which should be available in shadow realms?), we may need strong references to them to guarantee idempotence.
The
AsyncWrapproblem is a bit separate - they can be made weak, but also we need to make sure that they don't access the context when the weak callbacks are called, which is unsafe if the weak callbacks are called as a result of the context itself becoming reachable. This isn't possible for principal realms, because their the context is always destructed after all BaseObjects are cleaned up and async tasks are finished, but currently shadow realms have a different order. AndAsyncWrapare still used in some APIs that might be available to shadow realms (e.g. anything that uses the thread pool for async work, like compression streams and crypto jobs).Yeah, modules are supposed to be supported in the shadow realm and can be imported with
shadowRealm.importValue()or dynamic imports inevaluate().I think modules are a bit different from
AsyncWrapproblem. IMO, we should keep the realm alive if there are any ongoing async tasks. This avoids the problem that a shadow realm is becoming unreachable before an async task is finished.Reacted by Matteo Collina- added a commit that references this issue
on Apr 4, 2023 5 remaining items
@joyeecheung thanks for the clarification. I'm afraid my comment above may be not clear. If we defer the draining of the realms' cleanup hooks to a safe point to access V8 API (i.e. outside of weak callbacks), we can avoid depending on the order of the weak callbacks of BaseObjects and v8::Context.
If the realm's cleanup hooks are drained in the weak callback, BaseObject's (whose weak callbacks are pending to be invoked) destructor tries to create an instance handle of the wrapper object and set the internal fields. This violates V8's weak callback API contract and leads to a crash.
The second-pass callback is a way that shows how we defer draining the realm cleanup hooks. It can also be replaced with
Environment::SetImmediate: legendecas@4582785#diff-a0d00119410bc8c902ca44fce396732d95fd4bb4f13ad177852b2663ce746ca6R52. This doesn't depend on the order of the weak callback invocations. Whether or not a BaseObject's weak callback is called before the context's weak callback, their persistent handles are reset in BaseObject's weak callback and their destructor is not going to modify the internal fields.It can also be replaced with Environment::SetImmediate: legendecas@4582785#diff-a0d00119410bc8c902ca44fce396732d95fd4bb4f13ad177852b2663ce746ca6R52. This doesn't depend on the order of the weak callback invocations.
I think as long as we use
realm()in the BaseObject destructors and we destruct the BaseObjects in the weak callback, it still assumes that the realm weak callback is called after the BaseObject wrapper weak callbacks (if the realm weak callback is called first, in the BaseObject weak callbacksrealm()returns a dangling pointer). So it's still necessary to make the wrappers weak without giving it a weak callback. Which is okay for BindingData, but not necessarily for others...Actually if the realm weak callback is called first,
data.GetParameter()in the BaseObject callbacks already returns a dangling pointer..I think the second pass trick only enables us to make the BaseObject destructors less strict (now you can call the V8 APIs that do not need a context), but it doesn't solve the issue that if realm weak callbacks are called before BaseObject callbacks, the BaseObject weak callbacks should not even be called (theredata.GetParameter()can already be bogus).(I also noticed that technically speaking BaseObject's weak callbacks should be made two-passed and shouldn't actually touch V8 in the first pass weak callback too...and we should make the cleanup queue for BaseObjects a two-pass process as well to avoid interleaving weak callbacks with cleanup hooks - which might happen if the BaseObject destructors allocates any memory somehow - I am not sure how probably that is, but left a TODO here)
- added a commit that references this issue
on Apr 12, 2023 - added a commit that references this issue
on Apr 13, 2023 - added a commit that references this issue
on Apr 20, 2023 - added a commit that references this issue
on May 2, 2023 - added a commit that references this issue
on Jul 6, 2023 - added a commit that references this issue
on Jul 6, 2023 This issue has been marked as stale due to 210 days of inactivity.
It will be automatically closed in 30 days if no further activity occurs. If this is still relevant, please leave a comment or update it to keep it open.- addedstaleIssues and PRs marked stale due to inactivity and scheduled for automatic closure.Issues and PRs marked stale due to inactivity and scheduled for automatic closure.
on Jun 7, 2026 This issue has been automatically closed after 30 days of inactivity following its stale status (no activity for a total of 240 days).
If this is still relevant, feel free to reopen it or leave a comment with additional details so we can continue the discussion.
Version
all
Platform
all
Subsystem
shadow-realm
What steps will reproduce the bug?
How often does it reproduce? Is there a required condition?
Always
What is the expected behavior? Why is that the expected behavior?
Shouldn't crash from OOM
What do you see instead?
Crash from OOM
Additional information
See analysis #47339 (comment). This was introduced in #46809. The current memory management of shadow realms requires that there should be no strong global references in the graph, but we currently have many of them in the code base - among them are the binding data and the aliased arrays associated with the encoding binding, which is lazily created when
TextEncoderis accessed, for example.I think we have at least two options:
Or, we could use some help from V8 to get notified about the realms being unreachable, and release the realms in some callback instead of in a weak callback of the context. It's not clear to me what's enough for us at this point, though.