Description
ParseChangeID accepts Git change IDs that are not in their canonical form.
The parser uses net/url, which can normalize parts of the URI or drop them entirely. As a result, different input strings can be parsed into the same ChangeID.
For example, inputs with an uppercase scheme, userinfo, query parameters, fragments, or non-canonical percent-encoding are currently accepted even though they don't round-trip to the original string.
The parser also accepts empty segments in the repository path.
Steps to Reproduce
- Call
ParseChangeID with a non-canonical Git change ID, for example:
GIT://git.example.com/uber/monorepo/refs%2Fheads%2Fmain/<sha>
- Observe that parsing succeeds.
- Call
String() on the returned ChangeID.
- The resulting string is different from the original input.
The same issue can be reproduced with userinfo, query parameters, fragments, lowercase percent-encoding, unnecessary percent-encoding, and empty repository path segments.
Expected Behavior
ParseChangeID should reject any input that is not in canonical form.
A valid change ID should round-trip exactly:
ParseChangeID(raw).String() == raw
Actual Behavior
Non-canonical inputs are accepted and net/url normalizes or drops parts of them during parsing.
For example, an uppercase GIT:// scheme is parsed as git://, and query parameters or fragments are not included in the resulting ChangeID.
Environment
- Go version: go1.27.1
- OS: Windows
- Bazel version: 8.4.1
Logs / Screenshots
Not applicable.
Additional Context
This can result in multiple different strings representing the same change ID.
The parser should enforce a single canonical representation so that change IDs are unambiguous.
Description
ParseChangeIDaccepts Git change IDs that are not in their canonical form.The parser uses
net/url, which can normalize parts of the URI or drop them entirely. As a result, different input strings can be parsed into the sameChangeID.For example, inputs with an uppercase scheme, userinfo, query parameters, fragments, or non-canonical percent-encoding are currently accepted even though they don't round-trip to the original string.
The parser also accepts empty segments in the repository path.
Steps to Reproduce
ParseChangeIDwith a non-canonical Git change ID, for example:GIT://git.example.com/uber/monorepo/refs%2Fheads%2Fmain/<sha>String()on the returnedChangeID.The same issue can be reproduced with userinfo, query parameters, fragments, lowercase percent-encoding, unnecessary percent-encoding, and empty repository path segments.
Expected Behavior
ParseChangeIDshould reject any input that is not in canonical form.A valid change ID should round-trip exactly:
ParseChangeID(raw).String() == rawActual Behavior
Non-canonical inputs are accepted and
net/urlnormalizes or drops parts of them during parsing.For example, an uppercase
GIT://scheme is parsed asgit://, and query parameters or fragments are not included in the resultingChangeID.Environment
Logs / Screenshots
Not applicable.
Additional Context
This can result in multiple different strings representing the same change ID.
The parser should enforce a single canonical representation so that change IDs are unambiguous.