Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Noticed while checking how xsd:base64Binary values get validated. A document whose base64Binary value carried trailing junk still validated clean:
Traced it to
JavaBase64Holder.lex. It decodes withBase64.getMimeDecoder().decode(v), and the JDK MIME decoder silently drops every character that is not in the base64 alphabet, not only line separators. SoSGVsbG8=!!!!decodes toHello,SGV!!!sbG8=decodes toHello, and!!!!decodes to an empty array. Each is accepted instead of being reported invalid.lexis the validator the streamingValidatorcalls for base64Binary throughvalidateLexical, andset_textcalls it too, so an out-of-alphabet value passes document validation.The sibling
JavaHexBinaryHolder.lexdoes not have this problem:HexBin.decodereturns null on any non-hex character and the value is reported invalid. base64 had drifted from that behaviour.Fix scans the value first and reports it invalid when a character is neither in the base64 alphabet nor XML whitespace, then decodes as before. Whitespace and line-wrapped values are untouched, so valid input decodes exactly as it did.
Regression test added in
Base64BinaryValidateTest.