Skip to content

fix: preserve underscores when creating integers from strings - #610

Open
Kylinny wants to merge 2 commits into
python-poetry:masterfrom
Kylinny:fix-integer-underscore-emission
Open

Kylinny wants to merge 2 commits into
python-poetry:masterfrom
Kylinny:fix-integer-underscore-emission

Conversation

@Kylinny

@Kylinny Kylinny commented Sep 24, 2026

Copy link
Copy Markdown

Fixes #336.

Invariants

  1. Emission preserves the author's grouping. When integer() receives a string, the original text is stored verbatim as the item's raw text (Integer.as_string() already returns the stored raw), so dumps(integer("1_000")) now emits 1_000 instead of silently normalizing to 1000.
  2. The grammar validator is the single source of truth for what may be preserved. The new _INTEGER_RE encodes the TOML v1.0.0 § Integer grammar (https://toml.io/en/v1.0.0#integer): decimal with optional sign and no leading zeros, 0x/0o/0b with lowercase prefix and no sign, underscores only between digits. Its accept set mirrors Parser._parse_number, so anything this function preserves is also parseable — a preserved literal always round-trips.
  3. Invalid strings are rejected at construction, so no invalid TOML can be emitted. Anything outside the grammar raises ValueError (the same exception int(raw) raised before for garbage input), instead of being emitted verbatim or silently normalized.
  4. The parse path is untouched. Only the api.integer(str) constructor path changes. integer(int) keeps the exact old behavior (item(int(raw)), including bool → Integer(1)), and the parser is not modified.

Notes

  • The value is computed with int(raw, 0) on the already-validated string; base-0 handles the 0x/0o/0b prefixes, and Python's underscore placement rules coincide with TOML's for validated input.
  • Integer.__init__'s _sign detection (^[+\-]\d+$) is intentionally left alone: it only affects how arithmetic results re-render an explicit +, which is orthogonal to this fix.

Tests

  • Parametrized emission tests over valid groupings (1_000, 5_349_221, 0xdead_beef, 0xDEAD_BEEF, 0o7_7, 0b1_0, +1_000, -1_000).
  • Rejection tests asserting ValueError for _100, 100_, 1__000, 01_2, 0x_1, 0X1, +0x1, "", "abc", "1.0", " 1".
  • Round-trip tests asserting the emission path agrees with the parse path (parse("n = 1_000\n")).
  • Full suite: 407 passed (excluding test_toml_tests.py, which needs the uninitialized corpus submodule).

TOML 1.0 allows grouping underscores in integers (1_000,
0xdead_beef), and tomlkit parses them fine, but api.integer("1_000")
silently normalized the raw text to "1000" because it built the
item via item(int(raw)), discarding the original string.

Validate string input against the TOML v1.0.0 integer grammar and
preserve it verbatim as the item's raw text; invalid strings raise
ValueError. The parse path and integer(int) behavior are unchanged.

Fixes python-poetry#336.

Signed-off-by: Junyi Yao <j.yao@wustl.edu>
@Kylinny

Kylinny commented Sep 26, 2026

Copy link
Copy Markdown
Author

Hi @frostming — just a gentle ping on this one. It's a small fix for #336 (integer underscores are dropped when creating integers from strings), with 29 new parametrized cases and the full suite passing (407 tests). Happy to adjust anything you'd like changed — thanks for maintaining tomlkit!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Emitting underscores in integers

1 participant