Skip to content

Autumn Budget 2026 scoreability registry: ready before 28 October (#96) - #100

Open
vahid-ahmadi wants to merge 3 commits into
mainfrom
uk/budget-2026-registry
Open

Autumn Budget 2026 scoreability registry: ready before 28 October (#96)#100
vahid-ahmadi wants to merge 3 commits into
mainfrom
uk/budget-2026-registry

Conversation

@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Closes #96. Branches from main — independent of the other UK PRs.

The Budget is 28 October 2026 (Chancellor John Healey, Burnham government). This registry records what each reported measure would be as a PolicyEngine-UK reform and whether the certified engine can express it — so a counterpart can be computed on the day rather than the work starting then.

14 measures: 3 expressible, 1 partial, 10 not.

It stages no values, on purpose

The revenue figures in circulation are journalism citing third parties. #86's rule applies: a re-published figure belongs to its originator, not the outlet that repeated it. When HMT publishes the scorecard and the OBR the EFO costings on the day, those are primary and arrive through the existing lanes, joining to these measure keys. A test asserts no value keys exist in the file.

Expressible means resolved, not asserted

Every path was resolved against an installed policyengine-uk 2.89.2, and the recorded 2026 baselines are what that engine returns. The validator re-checks both, so registry drift fails here rather than at the first real run.

registry valid: 14 measures
  computability: {'expressible': 3, 'not_expressible': 10, 'partial': 1}
  resolved 7 parameter paths against the certified engine

The findings worth reading

CGT alignment is the flagship and it is three clean parameters — 18/24/24 today → 20/40/45. PE-UK also carries a real CGT behavioural response, which matters more here than anywhere else in the registry: published estimates disagree mostly about the response, so a static-only PE number would not be comparable to a behavioural OBR costing. Spun out as #97.

The most-reported pensions measure cannot be scored at all. gov.hmrc.pensions has no tax-free lump sum parameter at the pin — probed, raises AttributeError. Recorded as an upstream development item, and spun out as #98 together with the fact that nothing validates the pensions PE-UK does model.

A freeze is not a delta on current law. The threshold-freeze entry says so: it is only measurable against an indexed counterfactual, which is the registered-baseline question (#13) and precisely the axis #67 found undecomposed on the OBR's own PA/HRT re-estimates.

Expressible is kept apart from credible. The wealth tax resolves — but against survey-imputed wealth, the weakest data in the certified world for exactly the population it targets. The entry says so rather than implying a publishable number.

Roughly half the reported package is not household-modellable. Bank surcharge, online sales levy, business rates, vape duty, ATED. They are listed and marked out_of_model_scope, and the validator refuses a registry that omits them — listing only the modellable half would overstate coverage.

Two already-announced measures raise a baseline-integrity question, not a scoring one: if the certified world's 2027/2028 baselines do not carry them, every counterpart at those years is measured against the wrong law. Spun out as #99.

Verification

Suite 279 passed, ruff format --check clean, validator green both with and without the engine.

Reviewers

@MaxGhenis @DTrim99 — the call I'd most like tested is whether a pre-event registry belongs in this repo at all, or whether it should wait for the actual scorecard on 28 October. My argument for now: the scoreability triage is the slow part and it is more honest done before the announcement, when nobody can tune it to a result.

The Budget is 28 October 2026 (Chancellor John Healey, Burnham
government). This registry records what each REPORTED measure would be
as a PolicyEngine-UK reform and whether the certified engine can express
it, so a counterpart can be computed on the day rather than the work
starting then.

It is a scoreability registry, NOT a claims lane. It stages no values:
the revenue figures in circulation are journalism citing third parties,
and #86's rule applies — a re-published figure belongs to its
originator. When HMT and the OBR publish the scorecard and EFO costings
on the day, those arrive as claims through the existing lanes and join
to these measure keys.

Every path was RESOLVED against an installed policyengine-uk 2.89.2, and
the recorded 2026 baselines are what that engine returns, not numbers
copied from reporting. The validator re-checks both, so a registry that
drifts from the engine fails here rather than at the first real run.

14 measures: 3 expressible, 1 partial, 10 not.

  - CGT alignment with income tax rates is the flagship and is three
    clean parameters (18/24/24 today -> 20/40/45). PE-UK also carries a
    CGT behavioural response, which matters more here than anywhere
    else: published estimates disagree mostly about the response, so a
    static-only PE number would not be comparable to a behavioural OBR
    costing.
  - The most-reported PENSIONS measure cannot be scored at all.
    gov.hmrc.pensions has no tax-free lump sum parameter at the pin —
    probed, not assumed. Recorded as an upstream development item.
  - A threshold freeze is not treated as a delta on current law. It is
    only measurable against an indexed counterfactual, which is the
    registered-baseline question (#13) and the axis #67 found
    undecomposed on the OBR's own PA/HRT re-estimates.
  - Expressible is kept apart from credible. The wealth tax resolves,
    but against survey-imputed wealth — the weakest data in the
    certified world for exactly the population it targets — and the
    entry says so rather than implying a publishable number.
  - Roughly half the reported package is business levies a household
    microsimulation cannot touch. They are listed and marked
    out_of_model_scope, and the validator REFUSES a registry that omits
    them, because listing only the modellable half would overstate
    coverage.
  - Two already-announced measures raise a baseline-integrity question
    rather than a scoring one: if the certified world's 2027/2028
    baselines do not carry them, every counterpart at those years is
    measured against the wrong law.

Suite 279 passed, ruff format clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016HuXJFVme8HRbnke2Ey3Me
@vahid-ahmadi

Copy link
Copy Markdown
Contributor Author

Review request — @MaxGhenis @DTrim99.

This one is part of a batch; the whole queue, with a suggested merge order and what is blocked on whom, is in #104 so you can triage in one place rather than PR by PR.

@DTrim99 DTrim99 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verdict: request changes. The registry's structure, counts, no-revenue-value discipline, and cited parameter paths all hold up against a real policyengine-uk tree — but one measure is misclassified on a verifiably false engine-gap claim, which is exactly the "unearned excuse" the registry's own honesty rule exists to prevent.

Critical

  • data/uk/budget_2026_measures.jsonab2026__property_income_tax_rate_rise.why claims "the engine taxes property income at the main income-tax rates with no separate schedule, so a property-specific rate has no parameter." This is false: gov.hmrc.income_tax.rates.property exists (basic/higher/additional) and already encodes the April-2027 +2pp rise. So the measure is expressible, not not_expressible, and its baseline-integrity question is actually already satisfied. Reclassify and fix the reason. Note the --resolve validator can't catch this because a not_expressible entry carries no path to resolve — the false claim passes CI silently, which is the registry's own stated failure mode.

Should

  • The registry_rule / PR body characterize #86 as "a re-published figure belongs to its originator, not the outlet." Issue #86 is actually about held-out think-tank ingest (IFS + RF), not a re-publication rule. The provenance spirit is adjacent but the citation is imprecise — confirm the intended issue number.
  • Nothing in the diff pins that --resolve ran green against 2.89.2 (no lockfile / CI log). The static baselines I could confirm (CGT 0.18/0.24/0.24, PA 12570, HRT threshold 37700) all match, so confidence is high — but the resolving engine version isn't pinned in-repo.

Verified OK

  • Count 3 + 1 + 10 = 14 (test asserts == 14); no value/revenue/costing keys (#86 discipline honored); no duplicate measure_keys; event_date 2026-10-28 consistent; the lump-sum not_expressible is correct (no such param). The distinct data/uk/ schema is defensible — this is a scoreability triage, not a facts table.

🤖 review via Claude Code

r and others added 2 commits August 26, 2026 10:05
DTrim99 caught `ab2026__property_income_tax_rate_rise`. Searching the
tree for the second one caught `ab2026__high_value_council_tax_surcharge`.
Both were recorded `not_expressible` on GUESSED paths that failed to
resolve, and both are wrong:

  gov.hmrc.income_tax.rates.property         basic/higher/additional,
                                             already carrying the
                                             April-2027 +2pp rise
  gov.hmrc.council_tax.high_value_surcharge  full banded schedule
    .amount                                  (not the path v1 guessed)

Both are reclassified `expressible` + `already_in_baseline`, with the
reform delta and the live engine baselines recorded. Neither is a
Budget-day scoring job: the certified world already carries them.

The registry's own stated failure mode is what happened. `--resolve`
cannot catch a false gap, because a `not_expressible` entry carries no
path to resolve — the claim passes CI silently. So the validator now
requires an in-scope gap to record the NAME SEARCH that proved it. A
guessed path that fails to resolve proves nothing.

Re-running those searches confirmed the remaining gaps are real: the
pensions lump-sum measure matches no `lump|commencement|tax_free_cash|
pcls` node anywhere in the tree, and the two near-misses that did turn
up (a UC standard-allowance uplift, a UC unearned-income definition)
are unrelated nodes, recorded as such.

Also: the baseline-integrity flag is retired. v1 raised "do the future
baselines carry measures already in law?" as an open question and then
answered it the wrong way. Reading the tree settles it — they do.

And the provenance citation is made precise: #86 is the IFS+RF lane,
but the re-publication rule was established in its PR, #91. Cite #91
for the rule.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016HuXJFVme8HRbnke2Ey3Me
@vahid-ahmadi

Copy link
Copy Markdown
Contributor Author

You were right, and the second one was worse.

gov.hmrc.income_tax.rates.property exists exactly as you describe, with the April-2027 rise already in it. Going back over the other gaps the same way turned up a second false one: ab2026__high_value_council_tax_surcharge. I had guessed gov.hmrc.high_value_council_tax_surcharge; the real node is gov.hmrc.council_tax.high_value_surcharge.amount, and it carries the full banded schedule. Both are now expressible + already_in_baseline, with the reform delta and live engine baselines recorded. Neither is a Budget-day scoring job — the certified world already carries them.

Your last sentence is the part worth fixing, so I fixed that rather than just the two entries. --resolve cannot catch this class of error, because a not_expressible entry has no path to resolve. The validator now refuses an in-scope gap that doesn't record the name search behind it — a guessed path that fails to resolve proves nothing, and that is precisely how both of these got published. test_validate_rejects_an_unsearched_gap_claim pins it.

Re-running those searches on the remaining gaps: the pensions lump-sum measure matches no lump|commencement|tax_free_cash|pcls node anywhere in the tree, so that one stands. The two near-misses that surfaced — a UC standard-allowance uplift, a UC unearned-income definition — are unrelated, and I've recorded them as such rather than quietly dropping them.

Two smaller things you'd have hit next:

285 tests, ruff format --check clean, two build_db runs agree on content_hash, no drift. Ready for another look.

vahid-ahmadi pushed a commit that referenced this pull request Aug 26, 2026
The probe workflow runs both registries, but pipeline/validate_budget_2026_registry.py
arrives with #100 — so this branch failed on a missing file rather than
on anything it owns.

Guarded on existence, and the skip is ANNOUNCED via ::notice:: rather
than silent. A gate that quietly does nothing reads as a gate that
passed, which is exactly what #74/#95 were about.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016HuXJFVme8HRbnke2Ey3Me
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Autumn Budget 2026 (28 October): scoreability registry so counterparts can be computed on the day

2 participants