Conversation
`Chars` already has a specialized `advance_by` that counts char start bytes in 32 byte chunks instead of decoding every char. The reverse direction fell back to the default implementation, so `chars().nth_back(n)`, `chars().rev().skip(n)` and friends decoded each char one by one. Mirror the forward implementation from the back. When the skipped bytes start in the middle of a char, its leading byte was not counted, so its continuation bytes are kept in the iterator.
Collaborator
|
Thanks for the pull request, and welcome! The Rust Project has assigned @Darksonn (or someone else) to review your changes, you should hear from them (or someone else) within the next two weeks. Please see the contribution instructions and our LLM policy for more information. Why was this reviewer chosen?The reviewer was selected based on:
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Charsalready has a specializedadvance_by(added in 40cf1f9, counts char start bytes in 32 byte chunks), butDoubleEndedIterator for Charsused the defaultadvance_back_by, sochars().nth_back(n),chars().rev().skip(n)/.rev().nth(n)decoded every char one by one.This mirrors the forward implementation from the back: count non-continuation bytes in 32 byte chunks via
as_rchunks, and if the skipped bytes begin in the middle of a char (whose leading byte therefore wasn't counted), keep its continuation bytes in the iterator. The remainder is skipped char by char.Benchmarks (
./x bench library/coretests --test-args chars_advance,corpora::ru::LARGE, x86_64):Tests:
test_iterator_advance_back(mirror of the forward test) andtest_iterator_advance_back_matches_next_back, which comparesadvance_back_by(n)/nth_back(n)against repeatednext_back()for alln(including past the end) on mixed 1–4 byte strings spanning several chunks, at several start offsets.r? libs
AI disclosure: This change was assisted by GitHub Copilot and Claude. I reviewed and tested it myself.