Repository navigation
Discussion/Tracking SnapshotCreator support #13877
Description
Activity
- addedlib / srcIssues and PRs involving general changes in the lib/ or src/ directories.Issues and PRs involving general changes in the lib/ or src/ directories.processIssues and PRs related to the process subsystem.Issues and PRs related to the process subsystem.
on Jun 22, 2017 I'm working on first one.
Does this still really matter now with Ignition (and TurboFan)?
@mscdex less so, but allows for rudimentary equivalence of features of
pkgand does have impact since you don't need to run anyrequireon startup.- addedperformanceIssues and PRs related to the performance of Node.js.Issues and PRs related to the performance of Node.js.discussIssues opened for discussion and feedback.Issues opened for discussion and feedback.
on Jun 22, 2017 Some numbers from local investigation: https://twitter.com/bradleymeck/status/875758792805404673
On a Mac 64bit:
- 2MB~ is minimal snapshot size (also why setting max old space < 3MB doesn't work)
- 10e6 globals serialized into a snapshot w/ different hidden classes (
global[i++]={[i]:i}) is 88MB, no significant change in speed when calling Isolate::New w/ this vs 2MB. - Isolate::New time seems to float around 10ms on local machine (consistent w/ v8 builtin default snapshot)
If the
v8::SnapshotCreatorwas able to use the existingIsolate, we could use this for pure JS post-mortem dumps, right?@refack yes, but the snapshot blob is platform and build specific unlike core dumps
But we'll end up with a live system that could be dynamically interrogated. Sounds like an amazing feature, and if V8 stabilize the blob format it seems possible they could implement portability.
/cc @nodejs/post-mortem @nodejs/v8Not sure what you mean by "live", all
v8::Handles need to be closed in order to serialize a blob, can't have JS on the stack.Not sure what you mean by "live", all v8::Handles need to be closed in order to serialize a blob, can't have JS on the stack.
Ok not live, suspended animation, but AFAICT you could reignite the event loop and use the preloaded code for example to rerun a JS function to see what it's doing.
For example if a production server enters an err state, before you restart it you "suspend" it, take the blob to your dev machine and fiddle with the state in JS land. Mind blown!Related: Proposal for snapshots as a Realm API for wider support in the ecosystem and getting libraries to consider the requirements.
I built a mode into Prepack to compile snapshots of basic Node CLIs. One limitation that I hit was that things like the type of a socket becomes hard coded. You might imagine that taking a snapshot of a CLI that prints to the console and then running it with
node mysnapshot > file.txtor something would work but now that stdout is a different type and a lot of the JS streams around stdio has be rewired. Now Prepack can automatically encode multiple branches over abstract values but you'd probably want to manually handle such cases.@sebmarkbage I think the custom serializer deserializer functions in v8 could be enough to handle this. Even with that though, libs could check tty info that gets invalidated when
mainruns after being deserialized, like you said.26 remaining items
- is the only thing I think that might be a problem. However, I don't think Node itself ever changes the underlying buffer once created. Addons might, but if throwing an error on encountering them is sane that might be a good solution.
Line 391 in aa496f4
Local<ArrayBuffer> ab = ArrayBuffer::New(env->isolate(), data, length); It was able to spend some time on this last week updating the snapshot/inflate code in ChakraCore to deal with some of the additional external data references that appear when running in a Node host. At this point the largest problem I am running into is getting Node to reset the list of native pointers before inflate. I worked around this with a pretty hard-coded hack but it would be great if there was a cleaner implementation of the Node host side code.
I'll be at Node Summit this week if anyone else is going to be there and wants to discuss in more detail.
@bmeck, you mentioned that there are instances of ArrayBuffer with external off-heap backing store. Do you have any pointers for me on that? This may be solved simply by adding more logic to the array buffer allocator on Node side.
@hashseed I think most of the cases are fine, but Buffer.allocUnsafe and Buffer.allocUnsafeSlow are probably not able to be made internalized to my knowledge.
Since we never use kExternalized; you can get a list like below easily from grepping the source.
- Allocated but look to act like Internalized
Line 527 in 4796fbc
ArrayBuffer::New(isolate, fields_ptr, fields_count * sizeof(*fields_ptr));
Lines 543 to 546 in 4796fbc
Local<ArrayBuffer> uid_fields_ab = ArrayBuffer::New( isolate, uid_fields_ptr, uid_fields_count * sizeof(*uid_fields_ptr)); - Are these just for Domains ?
Line 1209 in 4796fbc
ArrayBuffer::New(env->isolate(), fields, sizeof(*fields) * fields_count);
Line 1254 in 4796fbc
ArrayBuffer::New(env->isolate(), fields, sizeof(*fields) * fields_count); - No clue how to fix this NAPI call which always makes externals
Lines 2885 to 2886 in 4796fbc
v8::Local<v8::ArrayBuffer> buffer = v8::ArrayBuffer::New(isolate, external_data, byte_length); - No clue how to fix
Buffer.unsafe*
Line 391 in 4796fbc
Local<ArrayBuffer> ab = ArrayBuffer::New(env->isolate(), data, length); fs.statavoids copy into heap, but fundamentally acts like internalized
Lines 1424 to 1426 in 4796fbc
Local<ArrayBuffer> ab = ArrayBuffer::New(env->isolate(), fields, sizeof(double) * 2 * 14); v8module hooks up to v8 C++ Heap statistics. Could repopulate after revival?
Lines 147 to 149 in 4796fbc
ArrayBuffer::New(env->isolate(), env->heap_statistics_buffer(), heap_statistics_buffer_byte_length));
Lines 195 to 197 in 4796fbc
ArrayBuffer::New(env->isolate(), env->heap_space_statistics_buffer(), heap_space_statistics_buffer_byte_length));
- Allocated but look to act like Internalized
It seems
is used to comunicate the state of the Tick queue between JS and C++ lands.Lines 1249 to 1257 in 4796fbc
// Values use to cross communicate with processNextTick. uint32_t* const fields = env->tick_info()->fields(); uint32_t const fields_count = env->tick_info()->fields_count(); Local<ArrayBuffer> array_buffer = ArrayBuffer::New(env->isolate(), fields, sizeof(*fields) * fields_count); args.GetReturnValue().Set(Uint32Array::New(array_buffer, 0, fields_count)); }
But AFAICT it could be de-abstracted sincefields_count === 2always.I’m kind of jumping in here without having read the full context yet, but if it’s about the AB backing store not being on the heap: It should be pretty easy to remedy that by creating the ABs in JS and then passing them to C++ instead of the other way around, right?
@addaleax the issue is the kExternalized ArrayBuffers are not managed by the v8 heap; that is correct. We can still use C++ to make kInternalized ArrayBuffers and do in
node_buffer.cc:
Lines 304 to 307 in 4796fbc
ArrayBuffer::New(env->isolate(), data, length, ArrayBufferCreationMode::kInternalized);
The snapshot serialization does not currently have a way to revive kExternalized ArrayBuffers.Reacted by Anna HenningsenTo reset the discussion on ArrayBuffers...
We have three kinds of ArrayBuffer backing stores:
- on V8's heap and managed by V8's GC
- off-heap, but registered to V8's GC, and therefore managed by V8's GC
- off-heap and external, i.e. not managed by V8's GC.
We can serialize and deserialize the first two kinds just fine. For off-heap backing stores, the deserializer simply calls the array buffer allocator.
The third kind is the one that the serializer currently cannot deal with, since we assume that the embedder is managing its lifetime. Maybe we should just assume that the embedder provided an array buffer allocator that correctly registers the backing store wrt lifetime management?
Maybe we should just assume that the embedder provided an array buffer allocator that correctly registers the backing store wrt lifetime management?
If that is added to the documentation, that seems fine.
closing in favor of #17058
@bmeck & @hashseed - #17058 seems to only be about accelerating Node.js' own loading and not a means for allowing tools like Webpack and similar to get the kinds of gains that the Atom folks got by using V8 snapshotting for Atom's startup, which this issue was supposed to be about. I've looked through the comments on #17058 (and command-line switches for Node.js v12.6, etc.) and it doesn't seem like the scope widened. I see a command-line switch that will generate a snapshot in response to a signal, but nothing about using a snapshot on startup. Am I just missing it? (Seems likely.) Thanks.
@bmeck @hashseed any reply to @tjcrowder ? I'm interested in this topic for environments like AWS Lambda to reduce "cold start" startup times.
@Bessonov I'm not working on anything, after the startup perf gains to my knowledge no one has looked at adding user-land support. I know it is possible to do the entire work in user-land with C++ if you want any guidance. Ideally you would not have any C++ user-land objects though on the heap. If C++ objects are in the heap then it is much more work to support.
I'm not working on anything, after the startup perf gains to my knowledge no one has looked at adding user-land support.
Fwiw, that’s still on my long-term TODO list, partially for Workers but also Node.js in general. If anybody else puts work into this, I’ll be happy, though :)
Reacted by T.J. Crowder and jasperReacted by jasperReacted by jasper
Relevant earlier discussion on #9473 .
v8::SnapshotCreatoris a means to capture a heap snapshot of JS, CodeGen, and C++ bindings and revive them w/o performing loading/evaluation steps that got to there. This issue is discussing what would be needed for tools likewebpackwhich run many times and have significant startup cost need in order to utilize snapshots.CC: @hashseed @refack
intptr_tintptr_ts--make-snapshotand--from-snapshotmain()functions for snapshots (save tov8::Private::ForApi($main_symbol_name)).vm.Contextfor snapshotrequire.cachepaths?The v8 API might be able to have some changes made(LANDED)Right now the v8 API would need a
--make-snapshotCLI flag sincev8::SnapshotCreatorcontrolsIsolatecreation and node would need to use the created isolate.Since all JS handles need to be closed when creating the snapshot, a
main()function would need to be declared during snapshot creation after all possible preloading has occurred. The snapshot could then be taken when node exits if exiting normally (note,unref'd handles may still exist).Some utility like
WarmUpSnapshotDataBlobfrom v8 so that the JIT code is warm when loaded off disk also relevant.