Repository navigation
zlib.inflateRawSync how do I get the number of bytes read? #8874
Description
Activity
- addedzlibIssues and PRs related to the zlib module and its compression dependencies.Issues and PRs related to the zlib module and its compression dependencies.
on Oct 1, 2016 - addedfeature requestIssues requesting new Node.js features.Issues requesting new Node.js features.
on Oct 2, 2016 The bytes consumed count is tracked internally but isn't exposed. It could be retrofitted onto the asynchronous API but it would be very awkward with the synchronous API. Either it would have to be a property on the returned buffer or a side channel such as a function or variable that holds the count of the last operation.
I would love to see this feature added. Having to use a non-async pure-JS implementation like pako, just to get this little bit of information is pretty lame.
It can be hijacked from the private API, by monkey-patching
._handle.writeSyncor._handle.callbackbut that's also not a real solutions, and requires re-implementing part of the module.If we can decide on an API, I might be willing to implement it.
Some ideas:
- Extra callback arg:
zlib.inflate(buff, opts, function(err, buffer, read) { });
- Does not translate to synchronous.
- Do any other callbacks take more than 2 arguments?
- An option for returning an object with buffer and this extra data:
zlib.inflate(buff, {bytesRead: true}, function(err, data) { var buffer = data.buffer; var bytesRead = data.bytesRead; });
- Can similarly work with synchronous.
- Property of buffer:
zlib.inflate(buff, opts, function(err, buffer) { var bytesRead = buffer.bytesRead; });
- Can similarly work with synchronous.
(see idea 4 before)
Reacted by smyt1Reacted by smyt1Option 2 - opting in to a different return value / callback signature - is probably the most acceptable. There are other APIs in core that work the same.
(No panacea though, it destroys composability.)
Options 2 and 3 both sound good to me.
I personally like option 3.
Adding a property changes the object's shape (hidden class), that has performance implications.
One other possible issue, is supporting Node streams. I'm not sure any in my suggested ideas cover that use case.
Another option is to append the length to the end of the buffer (written as an unsigned 32 bit integer) if requested. Default behavior would not write the length (so this would not disrupt existing code), but if you request the length the buffer will be 4 bytes longer. So long as
buffer.slice(0,-4)is a light operation there would be no significant performance penalty besides allocating an extra 4 bytes per request.For Node stream usage, somehow this information should be available here:
const MemoryStream = require('memorystream'); const stream = new MemoryStream(); const engine = zlib.createInflate(); engine.on('data', function(buffer) { // How to know how much data was read? }); stream.pipe(engine); stream.write(buffer);
Some ideas:
- Exposed via
engine.bytesReador similar.
- This value would probably represent the amount read so far, increasing and updating each on data.
- Expose by property on
buffer.
- What value would be best exposed per buffer? The amount for that chunk? Of the whole thing?
- Hidden class issue again.
This also gave me another idea. What if these
zlib.create*objects had async and sync "process" methods like the method ofzlib, where this information could be accessed. Sort-of a more-advanced OOP API:var engine = zlib.createInflate(); // Async: engine.process(buff, opts, function(err, buffer) { var bytesRead = engine.bytesRead; }); // Sync: var buffer = engine.processSync(buff, opts); var bytesRead = engine.bytesRead;
I think
processprobably wouldn't be the best name, a better name could be chosen, but I think you get the idea.- Exposed via
That's overengineering it. We're unlikely to introduce completely new APIs just to expose a simple data property.
As well, it's a variation of the side channel I mentioned earlier. You could accomplish the same thing by adding a
zlib.inflateSync.bytesReadproperty that gets updated after every call.- added a commit that references this issue
on May 10, 2017 5 remaining items
- added a commit that references this issue
on May 31, 2017 - added 2 commits that reference this issue
on Jun 5, 2017 - added a commit that references this issue
on Mar 29, 2018 - added a commit that references this issue
on Apr 9, 2018 - added a commit that references this issue
on Apr 12, 2018 - added a commit that references this issue
on Jul 27, 2026
zlib.inflateRawSyncreturns the decompressed data. AFAICT it does not indicate how many bytes were read from the input buffer. Is there a way to find that?Use case: some ZIP writers use the "data descriptor" feature of the pkzip file format. The CRC-32 checksum of the uncompressed data and the lengths actually appear directly after the deflated data, so you need to know the number of bytes that the deflator read in order to jump ahead to those fields. FWIW zip libraries leveraging zlib, like yauzl, do not bother with the checksum.