Skip to content

reject non-alphabet characters in base64Binary lex - #123

Open
aizu-m wants to merge 1 commit into
apache:trunkfrom
aizu-m:base64-reject-nonalphabet
Open

aizu-m wants to merge 1 commit into
apache:trunkfrom
aizu-m:base64-reject-nonalphabet

Conversation

@aizu-m

@aizu-m aizu-m commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Noticed while checking how xsd:base64Binary values get validated. A document whose base64Binary value carried trailing junk still validated clean:

<t:root>SGVsbG8=!!!!</t:root>   ->   doc.validate() == true

Traced it to JavaBase64Holder.lex. It decodes with Base64.getMimeDecoder().decode(v), and the JDK MIME decoder silently drops every character that is not in the base64 alphabet, not only line separators. So SGVsbG8=!!!! decodes to Hello, SGV!!!sbG8= decodes to Hello, and !!!! decodes to an empty array. Each is accepted instead of being reported invalid. lex is the validator the streaming Validator calls for base64Binary through validateLexical, and set_text calls it too, so an out-of-alphabet value passes document validation.

The sibling JavaHexBinaryHolder.lex does not have this problem: HexBin.decode returns null on any non-hex character and the value is reported invalid. base64 had drifted from that behaviour.

Fix scans the value first and reports it invalid when a character is neither in the base64 alphabet nor XML whitespace, then decodes as before. Whitespace and line-wrapped values are untouched, so valid input decodes exactly as it did.

before: <t:root>SGVsbG8=!!!!</t:root> validates
after:  reported invalid; "SGVsbG8=", "SGVs bG8=" and newline-wrapped values still validate

Regression test added in Base64BinaryValidateTest.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant