Repository navigation
Incorrect result for Buffer#toString #6075
Description
Activity
The issue is that both input buffers contain invalid character sequences that get substituted with the replacement character,
u+FFFD. That's why the UTF-8 strings are the same - the replacements are in the same locations - but the hexadecimal representation is not.I'll close,
Buffer#toString()is working as expected and documented in this case.- addedinvalidIssues and PRs that are invalid.Issues and PRs that are invalid.bufferIssues and PRs related to the buffer subsystem.Issues and PRs related to the buffer subsystem.
on Apr 6, 2016 @bnoordhuis I don't understand where input buffers contain invalid characters? Can you please point out where the input is wrong.
For example, in the sequence
[217, 132, 45, 138, 77...], 138 is not a valid starting point for a UTF-8 character sequence because those always have the two most significant bits set (EDIT: except for single-byte characters, of course.)138 & 192is 128 when it should be 192 (because 192 == 128 + 64.)I apologize that I still don't understand the concept. Can you please provide some reference where I can read about Buffer in detail so that I could understand what you meant.
I might be looking dumb here.
The no-argument version of
buf.toString()interprets the bytes in the buffer as UTF-8 and turns that into a string. One or more bytes make up a character; https://en.wikipedia.org/wiki/UTF-8#Description explains what those byte sequences look like. Not all sequences are valid; those are replaced with aU+FFFDcharacter.Hope that clears it up.
Thanks a lot @bnoordhuis 💙 💛 💚 💜
Hello Team,
I am using older version(
0.10.40) on one my project and facing some issue while grouping of data typeBuffer. Issue also present on latest version.Issue
If you see carefully, then there is difference in the
hexvalues of both the variables but theirutf8string values are exactly which is actually creating problem for us.Actual Use Case
Expected Result
We should have 2 keys as user_uuid for posts are different.
Actual Result
All the posts are grouped under same key, because
#toString()returns same value for both the buffer object.Fix
Buffer#toStringshould have defaulthexencoding output.Please let me know if I am on wrong path or understood incorrectly or it should be fix on other libraries itself. And if you think that I am on right path and it should be fixed here then I can raise a PR for the same.