Conversation
TOML 1.0 allows grouping underscores in integers (1_000,
0xdead_beef), and tomlkit parses them fine, but api.integer("1_000")
silently normalized the raw text to "1000" because it built the
item via item(int(raw)), discarding the original string.
Validate string input against the TOML v1.0.0 integer grammar and
preserve it verbatim as the item's raw text; invalid strings raise
ValueError. The parse path and integer(int) behavior are unchanged.
Fixes python-poetry#336.
Signed-off-by: Junyi Yao <j.yao@wustl.edu>
for more information, see https://pre-commit.ci
Author
|
Hi @frostming — just a gentle ping on this one. It's a small fix for #336 (integer underscores are dropped when creating integers from strings), with 29 new parametrized cases and the full suite passing (407 tests). Happy to adjust anything you'd like changed — thanks for maintaining tomlkit! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #336.
Invariants
integer()receives a string, the original text is stored verbatim as the item's raw text (Integer.as_string()already returns the stored raw), sodumps(integer("1_000"))now emits1_000instead of silently normalizing to1000._INTEGER_REencodes the TOML v1.0.0 § Integer grammar (https://toml.io/en/v1.0.0#integer): decimal with optional sign and no leading zeros,0x/0o/0bwith lowercase prefix and no sign, underscores only between digits. Its accept set mirrorsParser._parse_number, so anything this function preserves is also parseable — a preserved literal always round-trips.ValueError(the same exceptionint(raw)raised before for garbage input), instead of being emitted verbatim or silently normalized.api.integer(str)constructor path changes.integer(int)keeps the exact old behavior (item(int(raw)), includingbool→Integer(1)), and the parser is not modified.Notes
int(raw, 0)on the already-validated string; base-0 handles the0x/0o/0bprefixes, and Python's underscore placement rules coincide with TOML's for validated input.Integer.__init__'s_signdetection (^[+\-]\d+$) is intentionally left alone: it only affects how arithmetic results re-render an explicit+, which is orthogonal to this fix.Tests
1_000,5_349_221,0xdead_beef,0xDEAD_BEEF,0o7_7,0b1_0,+1_000,-1_000).ValueErrorfor_100,100_,1__000,01_2,0x_1,0X1,+0x1,"","abc","1.0"," 1".parse("n = 1_000\n")).test_toml_tests.py, which needs the uninitialized corpus submodule).