UK national calibration run-readiness: doctrine, live parity trio, US-format diagnostics, scorer, staging posture (#623) - #743
Conversation
… re-map, ingest scale ladder The frozen release object uk/frs_release.json (survey/base 2024, calibration 2025, SN 9563, DOI, UKDS zip sha, HF acquisition pins) drives lockstep asserts over the re-pinned raw-tab manifest: all 21 frs_table artifacts re-pinned to the 2024-25 tabs with the SPI-convention keys, six stages' SN 9252 prose defect fixed, and the typed spec moved in lockstep. TIME_PERIOD and the HMRC SPI/CGT build periods move to "2024"; the 2023-24 HMRC published surface is re-mapped as a signed nearest-available-vintage declaration (period_mapping: latest_published_tax_year) with the frozen original byte-untouched and the source-contract validator reconstructing the live canonical payload. take_up_contract build_year 2024 flips the two 2024 date-keyed rates. The ingest driver joins the #627 scale ladder (--sample-fraction post-frs_spine via the generic frame-sampling helpers) with receipt-postures on three full-scale fences at sampled rungs, and the WAS bridge-donor locator defect that refused every full-roster licensed run is fixed with a hermetic pin-coherence regression test. Part of #723. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The regenerator pins move to enhanced_frs_2024_25.h5 at the v1.56.14 tag (= incumbent-at-pin ebf733c) and the committed reference re-freezes at the new vintage: 145 columns (surface unchanged), period "2024", entity counts now test-pinned with the re-derived record-count identity (16,288 raw + 10,000 SPI) x 2 + 270 CGT band donors = 52,846 - no raw household is dropped at 2024-25 and the donor count follows the HMRC band file (30 x 9). The release-input coverage manifest regenerates against the new reference with the certified 2023 candidate unchanged (candidate side moves at #686); known gaps stay empty and the restored-column receipts hold. Part of #723. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The efrs-post-calibration input-mass descriptor is replaced in place (same name, #687's replacement model): enhanced_frs_2024_25.h5 identity and the totals_sha256 of the regenerated 131-column weighted-totals evidence, with the registry, gates.json, and the data-shard publication mirrors re-pinned in the same reviewed change per the gate-battery contract. No thresholds move - #723 records the re-measured baselines, #686 arms them. UK_REFERENCE_DATASET_NAME follows the incumbent's 2024-25 dataset name. Both per-reference reviewed exclusions are re-signed against the new reference (charitable_investment_gifts; owned_land on a fresh 2024-25 stability receipt), pending the approver's confirmation on the PR. Part of #723. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ise channel-blind scope note UK_REFERENCE_DATASET_NAME drops the _recalibrated suffix (adjudicated 2026-08-20): the pinned reference is the published enhanced_frs_2024_25.h5 itself and no recalibrated variant exists at this vintage, so the gate-report label now names the artifact exactly (June report strings keep their own label). The registry scope notes stop claiming the incumbent "structurally lacks" the SPI clone channel - the 2024-25 artifact carries the synthetic rows structurally but no admin-restored mass in the channel-exclusive columns, which is the fact the reviewed exclusions rest on; the approved exclusion reasons are untouched. Gate digests re-cut over the post-#729 union in the same change. Part of #723. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…block; self-describing June freeze Finding 1 (accepted): the family-coverage block rendered the #723 re-mapped period fields from the canonical manifest while hashing only the frozen mirror - evidence fields and their hash must name the same bytes. The block now carries a dual pin (source_manifest for the frozen June identity, canonical_source_manifest for the bytes the re-mapped fields come from) and a test binds each field set to the sha256 of the file it actually derives from. Finding 2 (rejected with armor): the committed replay report is the June evidence freeze and deliberately keeps mapped_build_period 2023 - it is evidence for the grandfathered release, not the 2024 line, and retires with the frozen manifest after #686 per #687; instead of regenerating it, a new assertion binds it to the FROZEN manifest's declared mapping so the partition is self-describing. Part of #723. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ibrated-candidate-623
…tree Both parents edit uk/gates.json and uk/country_package.json, so the merged manifest digests differ from either side's pins. Re-pinned by recomputation (never by picking a side): spec bundle e12a2cb8…, policy 404968fb…, gates manifest 59c7808d…, spec fingerprint bfb98736…. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…-format diagnostics, scorer, staging posture Work items 1-7 of the narrowed #623 plan (runs held per the 2026-08-21 adjudications; everything here is synthetic/hermetic): - uk_runtime/national_doctrine.py: every solver constant the national calibration inherits becomes a declared, tamper-tested doctrine value (epochs 256, lr 0.02, ratio 10.0, seed 0, loss cap 10.0, l0 0.0, free mass, uniform weights). - Single compile path: the stage consumes the driver's compiled registry at the release calibration year and materializes bindings through the shared target_materialization interpreter; prepared scratch columns are restored post-solve, so the staged frame survives the real HDFStore writer (closes #729 dispositions finding 4). - Canonical mass/weight conventions mirroring the rowwise path: declared mass reason, post-solve fence (CALIBRATED kind, exactly one appended record), household_weight_kind_chain + calibration_mass_change manifest. - The parity trio evaluates for real on armed builds: candidate side from the staged frame and solve diagnostics, reference side from the frozen eFRS parity instrument and the declared registry at name@period grain — never a copied reference; unarmed builds keep evidence_absent. - calibration_diagnostics.json is now produced by the same shared diagnostics producer the US release path uses (per-epoch loss trajectory, per-target rows), with a pinned US<->UK format-parity test; the build block carries the chronicle artifact provenance (facts/manifest shas, profile ids) and reserves score_vs_enhanced_frs. - tools/score_uk_national_candidate.py: #578 rule-1 scoring, both artifacts rescored on one frozen register, June-schema score block with a declared none_declared holdout basis. - Declared non-certified staging-candidate input posture (sha-gated with a mid-read race guard, refused for release candidates) and the held-run runbook, unblocking on WS-E E8+E10 or an explicitly named base. Implemented by Codex from the reviewed plan; review pass fixed the materialization period (declared calibration year, never the frame's base-year time_period) and the parity-evidence sourcing above. Part of #623 under #665. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d adapter Both PRs merged upstream, including María's adaptations of parts of this branch's work (c534517 routes calibration through the shared materializer; 1d066e5 is the cherry-picked #733 review fix). Reconciliation: - target_materialization.py and verify_uk_identity_stability.py: main's reviewed versions taken wholesale. - national_calibration.py: this branch's registry+period+doctrine stage kept (all tests and the driver target it); its private frame adapter replaced by a lifecycle subclass of main's shared UKFrameTargetAdapter — one adapter, now writer-safe (prepared scratch columns restored away post-solve, per the adjudicated materializer-owned lifecycle). - Main's packaged-binding stage tests ported to the registry+period API; the persistence assertion inverted to the adjudicated writer-clean invariant (result columns exactly equal the input columns); the materialization stub adapter adopts the shared count-variable convention. - Digest pins auto-merged to main's post-union re-cut and verified green by the pin suites — no re-measurement needed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ain merge The post-#735/#733 merge at 9a8bd46 resurrected the standalone `except ValueError` block that #733's review commit 1d066e5 had folded into the single generic handler. Because the narrower clause catches first, the named-edge branch inside `except Exception` became unreachable, silently reverting a reviewed structural fix. This branch has no business touching the spine driver at all — its scope is the national calibration seam — so the file returns to main's version exactly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
An armed UK national build crashed before its first stage: _source_pins stored the full Ledger provenance block under the 'ledger_facts' role, but role_pins_digest requires each role to carry exactly sha256 and size_bytes, so every armed run raised ValueError at input_pins_digest. The two halves landed independently — the strict pin contract with Logbook adoption (#666) and the Ledger role pin with the compile-parity wiring — and neither PR branch fires it alone, because only an armed run (--ledger-facts) reaches this path and the licensed runs were held. It reproduces on both #743 and #747 branches and on main. The pin now carries the feed's verified digest and byte size; the richer Ledger identity block already travels in safe_artifacts, source_vintages, and the diagnostics build block, so nothing is lost. This fix belongs upstream in #743 (the run-readiness PR whose runbook documents the armed command); it rides the build branch until then.
… production path First armed run receipt: 187 of 388 references skipped at materialization, because they bind simulated tax-benefit outputs (income tax, NICs, UC, caseloads, payment-band crosstabs) and the calibration stage's frame adapter reads stored columns only. UKPolicyEngineAdapter — the live-sim adapter — has no production caller, and the stage's tests feed synthetic frames with measure columns pre-attached, so every real input (spine or certified June candidate) fails the 388-target surface the same way. The June release never hit this because its 149 targets were demographics-only. The runner now drives the materializer's own skip reports: each round computes the missing (entity, variable) pairs from a live policyengine-uk simulation over the same records — native-entity calculate, numeric map_to, categorical group-to-member broadcast, boolean any-collapse — attaches them to the frame, and retries until the register binds. The incumbent gets the same treatment before scoring, so rule 1 compares both datasets under one yardstick, and the score block declares that symmetry. References this posture cannot bind (the salary-sacrifice counterfactual deltas need adapter.counterfactual_delta) are excluded with a per-reference receipt. The production fix belongs in the #743 lane: either wire UKPolicyEngineAdapter into the stage or land this materialization as a declared pre-calibration step.
An armed UK national build crashed before its first stage: _source_pins stored the full Ledger provenance block under the 'ledger_facts' role, but role_pins_digest requires each role to carry exactly sha256 and size_bytes, so every armed run raised ValueError at input_pins_digest. The two halves landed independently — the strict pin contract with Logbook adoption (#666) and the Ledger role pin with the compile-parity wiring — and neither PR branch fires it alone, because only an armed run (--ledger-facts) reaches this path and the licensed runs were held. It reproduces on both #743 and #747 branches and on main. The pin now carries the feed's verified digest and byte size; the richer Ledger identity block already travels in safe_artifacts, source_vintages, and the diagnostics build block, so nothing is lost. This fix belongs upstream in #743 (the run-readiness PR whose runbook documents the armed command); it rides the build branch until then.
243 of the 388 UK references are banded, and none of them were ever sliced. Every employment-income band materialized 35,351,186 (roughly everyone with employment income) and every state-pension band 13,518,447 (roughly every state pensioner), so a band's apparent overshoot -- up to 6,759x on the state-pension 1m+ cell -- was an unsliced total compared against a band value, not a data defect. The slice was declared but unconsumed: the contract binding names groupby_variable, each compiled spec carries its own band's lower edge in Ledger filter metadata, and _prepared_column_values read neither. The binding's own filters list did work (verified: a family_type == SINGLE binding correctly masks a COUPLE row), so this is specifically the band dimension. Both published encodings reduce to one lower edge -- a numeric *_lower_bound (HMRC SPI) or a range label in monthly units scaled by band_period_factor (DWP awards) -- because no reference anywhere declares an upper bound. A band's upper edge is its sibling's lower edge within the same contract target, grouped per contract target rather than per dimension so two measures sharing a dimension cannot slice each other on the wrong boundaries; the top band runs to infinity. Validated against the real register: 243 banded references, zero unreadable, and the derived bounds match the measure names (150000-200000, 1000000-inf, and UC monthly 500.01-600.00 to annual 6000.12-7200.12). A band whose edge cannot be read now raises, so the measure is skipped and reported rather than silently reporting the whole population as one band -- the failure mode that hid this. Tests cover the partition property, an adjacent-bands-differ regression, the monthly label conversion, and the refusal. Lives in the shared materialization module, so it cherry-picks into any branch.
All 15 two-child-limit references failed to materialize. Two naming faults,
both in the contract rather than the provider:
- eight children references declared value_variable "children_count", which
is not a policyengine-uk variable. They now bind what they actually count:
uc_is_child_limit_affected for the affected-children rows (mapped to
household it sums to the count of flagged children) and is_child for the
children-in-affected-households rows (the total children there).
- fifteen bindings put a prose label in count_of ("affected_households",
"affected_children", "children_in_affected_households"). count_of is a
column fallback consulted when value_variable is an entity-count
indicator, so the provider looked those labels up as columns and raised.
The labels move to notes, leaving household references on the
household_count unit indicator as intended.
The provider keeps its documented behaviour; the pre-existing test covering
count_of as a real column still passes.
The stub adapter and national-stage frame both modelled the old children_count column. Mapped to household, uc_is_child_limit_affected sums to the number of flagged children, so it serves as both the affected flag and the affected-children count; the fixtures now carry counts rather than indicators. Caught by running the suites properly: the earlier chain piped pytest into tail, so its exit status was tail's and two real failures rode through into 760fe8b.
UKFrameTargetAdapter.household_condition built a group-membership column for every non-household entity, so a condition declaring entity "person" asked for "person_person_id" and raised. People sit directly in a household; only group entities need that lookup. Every ONS household-composition reference uses person-level conditions (is_child, age), so all ten were excluded from calibration. That left household structure unconstrained while population stayed targeted, and the solve satisfied population by inflating households rather than multiplying them: weighted benunits per household reached 1.476 against the incumbent's 1.132, from an identical unweighted 1.158. The downstream cost is Universal Credit. 75.5% of the candidate's single UC claimants end up in multi-benunit households (incumbent: 8.6%), where only one benunit claims housing costs, so the housing element reaches 47.2% of them against the incumbent's 73.9% -- despite the candidate having more renters (88.3% vs 79.2%) and higher rent, and despite beating the incumbent within both strata (92.6% vs 78.9% solo, 32.5% vs 20.4% multi). Simpson's paradox, driven entirely by composition. Median single UC award falls to GBP 8,059 against GBP 9,310, emptying the GBP 8-11k and GBP 14-19k award bands that carry 98% of the caseload shortfall. Shared UK runtime, so this cherry-picks into the spine lane.
The 1a3274b crash fix landed without a test; this locks the contract it restored: the ledger_facts role pin carries exactly {sha256, size_bytes} (both feed layouts) and survives role_pins_digest, while the full provenance block is rejected — the exact shape that crashed the first armed run at input_pins_digest. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ine rule into the solve Re-cut of the build branch's 035ed63 under María's 2026-08-24 ruling: family_equal enters _ALLOWED_TARGET_WEIGHT_RULES with its armed-run receipts, and uk_national_target_loss_weights() derives the family-share vector — but the doctrine DEFAULT stays uniform. She passes family_equal as an explicit per-run setup while the weighting doctrine is measured; neither rule is adopted by default (run-9 receipts cut both ways). Also fixes the latent wiring defect the armed run exposed: the stage echoed target_weight_rule in its manifest but never passed a weight vector to calibrate(), so any declared rule silently solved uniform. The vector now travels explicitly; uniform maps to None, keeping the shipped identity byte-stable under the default. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The two-child-limit rebinding (760fe8b) replaced the prose count_of labels with real value_variable columns; the chronicle loader-guarantee test still demanded count_of on every baseline_flag_crosstab binding. The guarantee now matches the provider: affected_flag_variable plus either counted-column spelling. Latent on the build branch too — this suite was never re-run there after the rebinding. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Automated review pass (Claude Code, high effort, diff only — no verification runs). Posting on the draft since two of these bear on whether the calibration gate means anything; the other two are dead code. Substantive1. 2. Dead code3. 4. Findings 3 and 4 are trivial. 1 and 2 are the ones I would want resolved before this leaves draft — in both cases a gate reports success without having tested the thing it names. |
Part of #623 under #665 (WS-D). Registry #736 item 10 asked to close-or-narrow #623 after #729 delivered the seam core; this PR is the narrowing, per the adjudicated plan (María, 2026-08-21): all calibration tooling lands now — solve doctrine, honest gate evidence, US-format diagnostics, candidate-vs-incumbent scoring, run posture — and the licensed armed run itself stays held until the WS-E spine completes E8 (#684) + E10 (#686), or María names a different base. Builds on merged #735 (the 2025 Ledger surface) and #733 (the FRS 2024-25 line).
What lands
uk_runtime/national_doctrine.pymirroringlocal_doctrine.py: every solver option the seam previously inherited silently (epochs 256, lr 0.02, max_weight_ratio 10.0, seed 0, loss cap 10.0, l0 0.0, free mass, uniform target weights) is a declared, tamper-tested constant. Doctrine v1 deliberately lifts Add the UK national calibration step over ledger-backed target references #729's values; first-run evidence is the revision path (adjudication 3).UKNationalCalibrationStageconsumes the driver's compiled registry atload_uk_frs_release().calibration_year(2025) and materializes bindings through the shared interpreter via a lifecycle subclass of the sharedUKFrameTargetAdapter: prepared scratch columns are restored away post-solve, so the staged frame survives the real HDFStore writer (closes Add the UK national calibration step over ledger-backed target references #729 dispositions finding 4, per adjudication 2 — materializer-owned lifecycle). The materialization period is passed explicitly — never the input frame's base-yeartime_period, which lags the calibration year.national_calibration_mass_reason()names the bound families; a post-solve fence requires the CALIBRATED kind and exactly one appended mass record; the stage manifest carries the rowwise-conventionweightsblock (household_weight_kind_chain,calibration_mass_change, solve block, doctrine echo)._stage_parity_evidencebuilds theuk_export_surface/uk_target_surface/uk_target_fitevidence from independently sourced sides: candidate from the staged frame and solve diagnostics, reference from the frozen eFRS parity instrument (parity_referenceparameter, driver passes the committed instrument) and the declared registry atname@periodgrain. A copied reference can never fabricate a pass; unarmed builds keep the honestevidence_absent, pinned in both postures.uk_calibration_diagnostics_payload/write_uk_calibration_diagnostics— the same shared producer the US release path imports (schema v6, per-epochloss_trajectory, per-target rows, loss attribution), so the calibration dashboard consumes UK and US files identically. A hermetic format-parity pin asserts the UK payload's shared layer equals the shared producer's output withuk_diagnosticsstrictly additive. Thebuildblock carries the chronicle artifact provenance (facts sha256, manifest sha256, profile ids), build id, code pins, and input posture — every target value traceable throughledger_facts_sha256(Publish Chronicle package IDs in calibration diagnostics targets #661).tools/score_uk_national_candidate.py: both artifacts rescored on the same frozen register withrelative_error_loss(cap 10.0), emitting the June-schemascore_vs_enhanced_frsblock ({candidate,incumbent}×{train,holdout,full}_loss + target_wins) with per-family wins. Signed differences:holdout_basis: "none_declared"(June's holdout split lives only in the archived pipeline); the diagnosticsbuildblock reservesscore_vs_enhanced_frs: nulluntil the licensed run merges the real receipt.--staging-candidate-input-sha256: sha-gated with a mid-read race guard, labeled tier in the build record, refused for release candidates) so the armed run is push-button on any pre-clone spine;docs/uk-national-calibration-runbook-623.mdrecords the exact command, the evidence-dir layout, and the two unblock conditions. Grain basis (adjudication 1): the incumbent's publishedenhanced_frs_2024_25.h5is itself pre-clone (52,846 households;clone_and_assignfeeds only its local-weights product), so pre-clone scoring is apples-to-apples — with one signed method difference to carry into the score receipt: incumbent national weights are a collapsed local solve, ours is a direct national solve under doctrine.Explicitly out of scope
Verification
Hermetic: doctrine pinned tests; single-compile-path + writer regression through the real
write_uk_national_frame; armed/unarmed parity-trio postures; mass-record fences; US↔UK diagnostics format-parity pin; scorer unit tests on synthetic twins; staging-posture flag matrix (adversarial refusals). Synthetic end-to-end in CI: armed synthetic build → real writer → full battery, US-format diagnostics, Logbook row withledger_factsbound ininput_pins_digest. Full three-shard suite + ruff green locally on the merged tip.🤖 Generated with Claude Code