Skip to content

UK national calibration run-readiness: doctrine, live parity trio, US-format diagnostics, scorer, staging posture (#623) - #743

Draft
juaristi22 wants to merge 18 commits into
mainfrom
uk-national-first-calibrated-candidate-623
Draft

UK national calibration run-readiness: doctrine, live parity trio, US-format diagnostics, scorer, staging posture (#623)#743
juaristi22 wants to merge 18 commits into
mainfrom
uk-national-first-calibrated-candidate-623

Conversation

@juaristi22

Copy link
Copy Markdown
Collaborator

Part of #623 under #665 (WS-D). Registry #736 item 10 asked to close-or-narrow #623 after #729 delivered the seam core; this PR is the narrowing, per the adjudicated plan (María, 2026-08-21): all calibration tooling lands now — solve doctrine, honest gate evidence, US-format diagnostics, candidate-vs-incumbent scoring, run posture — and the licensed armed run itself stays held until the WS-E spine completes E8 (#684) + E10 (#686), or María names a different base. Builds on merged #735 (the 2025 Ledger surface) and #733 (the FRS 2024-25 line).

What lands

  1. National solve doctrine — new uk_runtime/national_doctrine.py mirroring local_doctrine.py: every solver option the seam previously inherited silently (epochs 256, lr 0.02, max_weight_ratio 10.0, seed 0, loss cap 10.0, l0 0.0, free mass, uniform target weights) is a declared, tamper-tested constant. Doctrine v1 deliberately lifts Add the UK national calibration step over ledger-backed target references #729's values; first-run evidence is the revision path (adjudication 3).
  2. Single compile path + writer-safe materializationUKNationalCalibrationStage consumes the driver's compiled registry at load_uk_frs_release().calibration_year (2025) and materializes bindings through the shared interpreter via a lifecycle subclass of the shared UKFrameTargetAdapter: prepared scratch columns are restored away post-solve, so the staged frame survives the real HDFStore writer (closes Add the UK national calibration step over ledger-backed target references #729 dispositions finding 4, per adjudication 2 — materializer-owned lifecycle). The materialization period is passed explicitly — never the input frame's base-year time_period, which lags the calibration year.
  3. Canonical weight/mass conventionsnational_calibration_mass_reason() names the bound families; a post-solve fence requires the CALIBRATED kind and exactly one appended mass record; the stage manifest carries the rowwise-convention weights block (household_weight_kind_chain, calibration_mass_change, solve block, doctrine echo).
  4. The parity trio evaluates for real on armed builds_stage_parity_evidence builds the uk_export_surface/uk_target_surface/uk_target_fit evidence from independently sourced sides: candidate from the staged frame and solve diagnostics, reference from the frozen eFRS parity instrument (parity_reference parameter, driver passes the committed instrument) and the declared registry at name@period grain. A copied reference can never fabricate a pass; unarmed builds keep the honest evidence_absent, pinned in both postures.
  5. US-format-identical diagnostics + chronicle provenance — the driver's bespoke diagnostics blob is replaced by uk_calibration_diagnostics_payload/write_uk_calibration_diagnostics — the same shared producer the US release path imports (schema v6, per-epoch loss_trajectory, per-target rows, loss attribution), so the calibration dashboard consumes UK and US files identically. A hermetic format-parity pin asserts the UK payload's shared layer equals the shared producer's output with uk_diagnostics strictly additive. The build block carries the chronicle artifact provenance (facts sha256, manifest sha256, profile ids), build id, code pins, and input posture — every target value traceable through ledger_facts_sha256 (Publish Chronicle package IDs in calibration diagnostics targets #661).
  6. Candidate-vs-incumbent scorer (US base v2: one CPS+ACS+PUF-detail pool; datasets labeled by exact record count (dense = full pool; exact-k L0 selection) #578 rule 1) — new tools/score_uk_national_candidate.py: both artifacts rescored on the same frozen register with relative_error_loss (cap 10.0), emitting the June-schema score_vs_enhanced_frs block ({candidate,incumbent}×{train,holdout,full}_loss + target_wins) with per-family wins. Signed differences: holdout_basis: "none_declared" (June's holdout split lives only in the archived pipeline); the diagnostics build block reserves score_vs_enhanced_frs: null until the licensed run merges the real receipt.
  7. Run readiness, runs held — a declared non-certified staging-candidate input posture (--staging-candidate-input-sha256: sha-gated with a mid-read race guard, labeled tier in the build record, refused for release candidates) so the armed run is push-button on any pre-clone spine; docs/uk-national-calibration-runbook-623.md records the exact command, the evidence-dir layout, and the two unblock conditions. Grain basis (adjudication 1): the incumbent's published enhanced_frs_2024_25.h5 is itself pre-clone (52,846 households; clone_and_assign feeds only its local-weights product), so pre-clone scoring is apples-to-apples — with one signed method difference to carry into the score receipt: incumbent national weights are a collapsed local solve, ours is a direct national solve under doctrine.

Explicitly out of scope

Verification

Hermetic: doctrine pinned tests; single-compile-path + writer regression through the real write_uk_national_frame; armed/unarmed parity-trio postures; mass-record fences; US↔UK diagnostics format-parity pin; scorer unit tests on synthetic twins; staging-posture flag matrix (adversarial refusals). Synthetic end-to-end in CI: armed synthetic build → real writer → full battery, US-format diagnostics, Logbook row with ledger_facts bound in input_pins_digest. Full three-shard suite + ruff green locally on the merged tip.

🤖 Generated with Claude Code

juaristi22 and others added 9 commits August 21, 2026 11:39
… re-map, ingest scale ladder

The frozen release object uk/frs_release.json (survey/base 2024, calibration
2025, SN 9563, DOI, UKDS zip sha, HF acquisition pins) drives lockstep asserts
over the re-pinned raw-tab manifest: all 21 frs_table artifacts re-pinned to
the 2024-25 tabs with the SPI-convention keys, six stages' SN 9252 prose
defect fixed, and the typed spec moved in lockstep. TIME_PERIOD and the HMRC
SPI/CGT build periods move to "2024"; the 2023-24 HMRC published surface is
re-mapped as a signed nearest-available-vintage declaration
(period_mapping: latest_published_tax_year) with the frozen original
byte-untouched and the source-contract validator reconstructing the live
canonical payload. take_up_contract build_year 2024 flips the two 2024
date-keyed rates. The ingest driver joins the #627 scale ladder
(--sample-fraction post-frs_spine via the generic frame-sampling helpers)
with receipt-postures on three full-scale fences at sampled rungs, and the
WAS bridge-donor locator defect that refused every full-roster licensed run
is fixed with a hermetic pin-coherence regression test.

Part of #723.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The regenerator pins move to enhanced_frs_2024_25.h5 at the v1.56.14 tag
(= incumbent-at-pin ebf733c) and the committed reference re-freezes at the
new vintage: 145 columns (surface unchanged), period "2024", entity counts
now test-pinned with the re-derived record-count identity
(16,288 raw + 10,000 SPI) x 2 + 270 CGT band donors = 52,846 - no raw
household is dropped at 2024-25 and the donor count follows the HMRC band
file (30 x 9). The release-input coverage manifest regenerates against the
new reference with the certified 2023 candidate unchanged (candidate side
moves at #686); known gaps stay empty and the restored-column receipts hold.

Part of #723.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The efrs-post-calibration input-mass descriptor is replaced in place (same
name, #687's replacement model): enhanced_frs_2024_25.h5 identity and the
totals_sha256 of the regenerated 131-column weighted-totals evidence, with
the registry, gates.json, and the data-shard publication mirrors re-pinned
in the same reviewed change per the gate-battery contract. No thresholds
move - #723 records the re-measured baselines, #686 arms them.
UK_REFERENCE_DATASET_NAME follows the incumbent's 2024-25 dataset name. Both
per-reference reviewed exclusions are re-signed against the new reference
(charitable_investment_gifts; owned_land on a fresh 2024-25 stability
receipt), pending the approver's confirmation on the PR.

Part of #723.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ise channel-blind scope note

UK_REFERENCE_DATASET_NAME drops the _recalibrated suffix (adjudicated
2026-08-20): the pinned reference is the published enhanced_frs_2024_25.h5
itself and no recalibrated variant exists at this vintage, so the gate-report
label now names the artifact exactly (June report strings keep their own
label). The registry scope notes stop claiming the incumbent "structurally
lacks" the SPI clone channel - the 2024-25 artifact carries the synthetic
rows structurally but no admin-restored mass in the channel-exclusive
columns, which is the fact the reviewed exclusions rest on; the approved
exclusion reasons are untouched. Gate digests re-cut over the post-#729
union in the same change.

Part of #723.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…block; self-describing June freeze

Finding 1 (accepted): the family-coverage block rendered the #723 re-mapped
period fields from the canonical manifest while hashing only the frozen
mirror - evidence fields and their hash must name the same bytes. The block
now carries a dual pin (source_manifest for the frozen June identity,
canonical_source_manifest for the bytes the re-mapped fields come from) and
a test binds each field set to the sha256 of the file it actually derives
from. Finding 2 (rejected with armor): the committed replay report is the
June evidence freeze and deliberately keeps mapped_build_period 2023 - it is
evidence for the grandfathered release, not the 2024 line, and retires with
the frozen manifest after #686 per #687; instead of regenerating it, a new
assertion binds it to the FROZEN manifest's declared mapping so the
partition is self-describing.

Part of #723.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…tree

Both parents edit uk/gates.json and uk/country_package.json, so the merged
manifest digests differ from either side's pins. Re-pinned by recomputation
(never by picking a side): spec bundle e12a2cb8…, policy 404968fb…,
gates manifest 59c7808d…, spec fingerprint bfb98736….

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…-format diagnostics, scorer, staging posture

Work items 1-7 of the narrowed #623 plan (runs held per the 2026-08-21
adjudications; everything here is synthetic/hermetic):

- uk_runtime/national_doctrine.py: every solver constant the national
  calibration inherits becomes a declared, tamper-tested doctrine value
  (epochs 256, lr 0.02, ratio 10.0, seed 0, loss cap 10.0, l0 0.0,
  free mass, uniform weights).
- Single compile path: the stage consumes the driver's compiled registry
  at the release calibration year and materializes bindings through the
  shared target_materialization interpreter; prepared scratch columns are
  restored post-solve, so the staged frame survives the real HDFStore
  writer (closes #729 dispositions finding 4).
- Canonical mass/weight conventions mirroring the rowwise path: declared
  mass reason, post-solve fence (CALIBRATED kind, exactly one appended
  record), household_weight_kind_chain + calibration_mass_change manifest.
- The parity trio evaluates for real on armed builds: candidate side from
  the staged frame and solve diagnostics, reference side from the frozen
  eFRS parity instrument and the declared registry at name@period grain —
  never a copied reference; unarmed builds keep evidence_absent.
- calibration_diagnostics.json is now produced by the same shared
  diagnostics producer the US release path uses (per-epoch loss
  trajectory, per-target rows), with a pinned US<->UK format-parity test;
  the build block carries the chronicle artifact provenance
  (facts/manifest shas, profile ids) and reserves score_vs_enhanced_frs.
- tools/score_uk_national_candidate.py: #578 rule-1 scoring, both
  artifacts rescored on one frozen register, June-schema score block with
  a declared none_declared holdout basis.
- Declared non-certified staging-candidate input posture (sha-gated with
  a mid-read race guard, refused for release candidates) and the held-run
  runbook, unblocking on WS-E E8+E10 or an explicitly named base.

Implemented by Codex from the reviewed plan; review pass fixed the
materialization period (declared calibration year, never the frame's
base-year time_period) and the parity-evidence sourcing above.

Part of #623 under #665.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d adapter

Both PRs merged upstream, including María's adaptations of parts of this
branch's work (c534517 routes calibration through the shared materializer;
1d066e5 is the cherry-picked #733 review fix). Reconciliation:

- target_materialization.py and verify_uk_identity_stability.py: main's
  reviewed versions taken wholesale.
- national_calibration.py: this branch's registry+period+doctrine stage kept
  (all tests and the driver target it); its private frame adapter replaced by
  a lifecycle subclass of main's shared UKFrameTargetAdapter — one adapter,
  now writer-safe (prepared scratch columns restored away post-solve, per the
  adjudicated materializer-owned lifecycle).
- Main's packaged-binding stage tests ported to the registry+period API; the
  persistence assertion inverted to the adjudicated writer-clean invariant
  (result columns exactly equal the input columns); the materialization stub
  adapter adopts the shared count-variable convention.
- Digest pins auto-merged to main's post-union re-cut and verified green by
  the pin suites — no re-measurement needed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ain merge

The post-#735/#733 merge at 9a8bd46 resurrected the standalone
`except ValueError` block that #733's review commit 1d066e5 had
folded into the single generic handler. Because the narrower clause
catches first, the named-edge branch inside `except Exception` became
unreachable, silently reverting a reviewed structural fix.

This branch has no business touching the spine driver at all — its
scope is the national calibration seam — so the file returns to main's
version exactly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
juaristi22 added a commit that referenced this pull request Aug 23, 2026
An armed UK national build crashed before its first stage: _source_pins
stored the full Ledger provenance block under the 'ledger_facts' role, but
role_pins_digest requires each role to carry exactly sha256 and size_bytes,
so every armed run raised ValueError at input_pins_digest.

The two halves landed independently — the strict pin contract with Logbook
adoption (#666) and the Ledger role pin with the compile-parity wiring — and
neither PR branch fires it alone, because only an armed run (--ledger-facts)
reaches this path and the licensed runs were held. It reproduces on both
#743 and #747 branches and on main.

The pin now carries the feed's verified digest and byte size; the richer
Ledger identity block already travels in safe_artifacts, source_vintages,
and the diagnostics build block, so nothing is lost.

This fix belongs upstream in #743 (the run-readiness PR whose runbook
documents the armed command); it rides the build branch until then.
juaristi22 added a commit that referenced this pull request Aug 23, 2026
… production path

First armed run receipt: 187 of 388 references skipped at materialization,
because they bind simulated tax-benefit outputs (income tax, NICs, UC,
caseloads, payment-band crosstabs) and the calibration stage's frame adapter
reads stored columns only. UKPolicyEngineAdapter — the live-sim adapter —
has no production caller, and the stage's tests feed synthetic frames with
measure columns pre-attached, so every real input (spine or certified June
candidate) fails the 388-target surface the same way. The June release never
hit this because its 149 targets were demographics-only.

The runner now drives the materializer's own skip reports: each round
computes the missing (entity, variable) pairs from a live policyengine-uk
simulation over the same records — native-entity calculate, numeric map_to,
categorical group-to-member broadcast, boolean any-collapse — attaches them
to the frame, and retries until the register binds. The incumbent gets the
same treatment before scoring, so rule 1 compares both datasets under one
yardstick, and the score block declares that symmetry. References this
posture cannot bind (the salary-sacrifice counterfactual deltas need
adapter.counterfactual_delta) are excluded with a per-reference receipt.

The production fix belongs in the #743 lane: either wire
UKPolicyEngineAdapter into the stage or land this materialization as a
declared pre-calibration step.
juaristi22 and others added 8 commits August 24, 2026 12:58
An armed UK national build crashed before its first stage: _source_pins
stored the full Ledger provenance block under the 'ledger_facts' role, but
role_pins_digest requires each role to carry exactly sha256 and size_bytes,
so every armed run raised ValueError at input_pins_digest.

The two halves landed independently — the strict pin contract with Logbook
adoption (#666) and the Ledger role pin with the compile-parity wiring — and
neither PR branch fires it alone, because only an armed run (--ledger-facts)
reaches this path and the licensed runs were held. It reproduces on both
#743 and #747 branches and on main.

The pin now carries the feed's verified digest and byte size; the richer
Ledger identity block already travels in safe_artifacts, source_vintages,
and the diagnostics build block, so nothing is lost.

This fix belongs upstream in #743 (the run-readiness PR whose runbook
documents the armed command); it rides the build branch until then.
243 of the 388 UK references are banded, and none of them were ever
sliced. Every employment-income band materialized 35,351,186 (roughly
everyone with employment income) and every state-pension band 13,518,447
(roughly every state pensioner), so a band's apparent overshoot -- up to
6,759x on the state-pension 1m+ cell -- was an unsliced total compared
against a band value, not a data defect.

The slice was declared but unconsumed: the contract binding names
groupby_variable, each compiled spec carries its own band's lower edge in
Ledger filter metadata, and _prepared_column_values read neither. The
binding's own filters list did work (verified: a family_type == SINGLE
binding correctly masks a COUPLE row), so this is specifically the band
dimension.

Both published encodings reduce to one lower edge -- a numeric
*_lower_bound (HMRC SPI) or a range label in monthly units scaled by
band_period_factor (DWP awards) -- because no reference anywhere declares
an upper bound. A band's upper edge is its sibling's lower edge within the
same contract target, grouped per contract target rather than per
dimension so two measures sharing a dimension cannot slice each other on
the wrong boundaries; the top band runs to infinity. Validated against the
real register: 243 banded references, zero unreadable, and the derived
bounds match the measure names (150000-200000, 1000000-inf, and UC monthly
500.01-600.00 to annual 6000.12-7200.12).

A band whose edge cannot be read now raises, so the measure is skipped and
reported rather than silently reporting the whole population as one band --
the failure mode that hid this. Tests cover the partition property, an
adjacent-bands-differ regression, the monthly label conversion, and the
refusal.

Lives in the shared materialization module, so it cherry-picks into any
branch.
All 15 two-child-limit references failed to materialize. Two naming faults,
both in the contract rather than the provider:

- eight children references declared value_variable "children_count", which
  is not a policyengine-uk variable. They now bind what they actually count:
  uc_is_child_limit_affected for the affected-children rows (mapped to
  household it sums to the count of flagged children) and is_child for the
  children-in-affected-households rows (the total children there).

- fifteen bindings put a prose label in count_of ("affected_households",
  "affected_children", "children_in_affected_households"). count_of is a
  column fallback consulted when value_variable is an entity-count
  indicator, so the provider looked those labels up as columns and raised.
  The labels move to notes, leaving household references on the
  household_count unit indicator as intended.

The provider keeps its documented behaviour; the pre-existing test covering
count_of as a real column still passes.
The stub adapter and national-stage frame both modelled the old
children_count column. Mapped to household, uc_is_child_limit_affected sums
to the number of flagged children, so it serves as both the affected flag
and the affected-children count; the fixtures now carry counts rather than
indicators.

Caught by running the suites properly: the earlier chain piped pytest into
tail, so its exit status was tail's and two real failures rode through into
760fe8b.
UKFrameTargetAdapter.household_condition built a group-membership column
for every non-household entity, so a condition declaring entity "person"
asked for "person_person_id" and raised. People sit directly in a
household; only group entities need that lookup.

Every ONS household-composition reference uses person-level conditions
(is_child, age), so all ten were excluded from calibration. That left
household structure unconstrained while population stayed targeted, and
the solve satisfied population by inflating households rather than
multiplying them: weighted benunits per household reached 1.476 against
the incumbent's 1.132, from an identical unweighted 1.158.

The downstream cost is Universal Credit. 75.5% of the candidate's single
UC claimants end up in multi-benunit households (incumbent: 8.6%), where
only one benunit claims housing costs, so the housing element reaches
47.2% of them against the incumbent's 73.9% -- despite the candidate
having more renters (88.3% vs 79.2%) and higher rent, and despite beating
the incumbent within both strata (92.6% vs 78.9% solo, 32.5% vs 20.4%
multi). Simpson's paradox, driven entirely by composition. Median single
UC award falls to GBP 8,059 against GBP 9,310, emptying the GBP 8-11k and
GBP 14-19k award bands that carry 98% of the caseload shortfall.

Shared UK runtime, so this cherry-picks into the spine lane.
The 1a3274b crash fix landed without a test; this locks the contract it
restored: the ledger_facts role pin carries exactly {sha256, size_bytes}
(both feed layouts) and survives role_pins_digest, while the full
provenance block is rejected — the exact shape that crashed the first
armed run at input_pins_digest.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ine rule into the solve

Re-cut of the build branch's 035ed63 under María's 2026-08-24 ruling:
family_equal enters _ALLOWED_TARGET_WEIGHT_RULES with its armed-run
receipts, and uk_national_target_loss_weights() derives the family-share
vector — but the doctrine DEFAULT stays uniform. She passes family_equal
as an explicit per-run setup while the weighting doctrine is measured;
neither rule is adopted by default (run-9 receipts cut both ways).

Also fixes the latent wiring defect the armed run exposed: the stage
echoed target_weight_rule in its manifest but never passed a weight
vector to calibrate(), so any declared rule silently solved uniform.
The vector now travels explicitly; uniform maps to None, keeping the
shipped identity byte-stable under the default.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The two-child-limit rebinding (760fe8b) replaced the prose count_of
labels with real value_variable columns; the chronicle loader-guarantee
test still demanded count_of on every baseline_flag_crosstab binding.
The guarantee now matches the provider: affected_flag_variable plus
either counted-column spelling. Latent on the build branch too — this
suite was never re-run there after the rebinding.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Automated review pass (Claude Code, high effort, diff only — no verification runs). Posting on the draft since two of these bear on whether the calibration gate means anything; the other two are dead code.

Substantive

1. packages/microcosm-build/src/microcosm/build/uk_runtime/hmrc_restoration.py:851 — the certified-pin check is skipped for the staging tier. _validate_certified_candidate_identity returns early for any identity whose tier == "staging_candidate", verifying only that revision equals a hard-coded constant; the sha256/size comparison against the certified pin never runs. Any UKCertifiedCandidateIdentity constructed with that tier — not only one produced by verify_staging_candidate_uk_input — passes the HMRC replay-base gate with an unverified file. The SHA binding currently lives only in the CLI path, which makes it a convention rather than an invariant; a second caller constructing the identity directly gets no binding at all.

2. tools/score_uk_national_candidate.py:60 — the rule-1 holdout comparison is vacuous by construction. The score block sets candidate_train_loss, candidate_holdout_loss and candidate_full_loss all to the same candidate.final_loss (and likewise for the incumbent), with holdout_basis: "none_declared". Since the candidate was calibrated on this very registry, there is no held-out quantity in the comparison and rule 1 can only ever be won by the candidate — any overfit divergence is silently absorbed. This is the same failure shape as the E7 receipt on #747: a green result that reflects the absence of a check rather than the presence of agreement. Given this is the component deciding whether a calibration run succeeded, it seems worth either declaring a real holdout basis or having the rule refuse when holdout_basis == "none_declared" rather than scoring.

Dead code

3. tools/build_uk_frs_spine.py:871 — the new except ValueError as error: clause is inserted ahead of the pre-existing except Exception as error: whose body branches on isinstance(error, ValueError). That branch is now unreachable, and any ValueError special-casing it performed beyond the rung-abort signature is silently lost. Worth a look given this is the same handler restructured for the #733 review — it may be that the merge of the two paths dropped something.

4. tools/build_uk_national_dataset.py:1202source_vintages["frs"] = load_uk_frs_release().vintage is added two lines above an identical pre-existing assignment; the new line is dead and load_uk_frs_release() is now called twice.

Findings 3 and 4 are trivial. 1 and 2 are the ones I would want resolved before this leaves draft — in both cases a gate reports success without having tested the thing it names.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants