Skip to content

Baselines registry + descriptive-register gate (issues #13, #9) - #71

Merged
MaxGhenis merged 5 commits into
mainfrom
schema-baselines-gate
Aug 19, 2026
Merged

Baselines registry + descriptive-register gate (issues #13, #9)#71
MaxGhenis merged 5 commits into
mainfrom
schema-baselines-gate

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Step 1 of the UK-foundation plan (dual-review adjudication): #32's two architecture-neutral schema layers, extracted and hardened, with main's vocabulary untouched.

  • Baselines registry (Baseline as a first-class attribute of every score #13): baselines table + baseline_key on scores/results/exhibits via the shared preparers; 15 curated worlds; deliberate-registration gate (fired on Land the TPC T26-0009 per-sheet re-parse on main (135 by-race claims) #31's day-old pre_obbba_law while landing — registered with its distinctness rationale). Committed DB migrated: 42,610 claims → 12 in-use worlds, zero unregistered.
  • View guard, null-hole closed (the adjudication's condition): claims backfilled by exact projection from their own reform_json; legacy results stay NULL and render baseline_unvalidated against non-current-law worlds — no fabricated provenance, no plain agreement across worlds, zero retroactive shifts.
  • Descriptive gate (App design: public scoreboard + mission control, four views #9): methodological_difference class; pe_gap/external_issue raise without a citable action_link (caught and fixed the CPSP erratum's missing citation).

Suite: 179 tests. Next: repair #48 against the adjudication's eight blockers on this base.

🤖 Generated with Claude Code

MaxGhenis and others added 5 commits August 19, 2026 11:38
The two architecture-neutral schema layers from #32, extracted and
hardened as the prerequisite the UK-foundation adjudication called for
— main's vocabulary (Metric/UnitConcept/STANDARD_CONDITIONS) untouched.

Registry (#13): baselines table + baseline_key on external_scores /
pe_results / pe_exhibits, stamped by the shared preparers; curated in
scorecard_db/baselines.py (15 worlds); register_baselines fails loudly
on any key in data the registry doesn't describe — extending it is a
deliberate act. It fired twice while landing this: on RV's
pre_obbba_current_law (as #32 documented) and on pre_obbba_law from
today's #31 T26-0009 re-parse — a world introduced hours ago, now
registered with its distinct-from-pre_obbba_current_law rationale
stated (equivalence is a future deliberate act, never assumed).

Comparisons-view guard: a comparable result whose executed baseline
differs from the claim's renders pe_status_effective='constructed'.
Hardened beyond #32 (the adjudication's null-hole finding): claims'
baseline_key is backfilled by exact projection from each row's own
reform_json (never inference); legacy results keep NULL — the executed
baseline is provenance we don't have — and the view downgrades those to
'baseline_unvalidated' wherever the claim's world is not current law
(the null world's key is injected into the view SQL so the guard works
on unseeded files). On the committed DB: 42,610 claims -> 12 in-use
worlds, zero unregistered, zero retroactive status shifts.

Gate (#9): DiagnosisClass.methodological_difference joins the taxonomy;
db.diagnose raises for pe_gap/external_issue without a citable
action_link. It caught main's CPSP erratum diagnosis missing its
citation (the erratum NOTES anchor — ported from #32's fix).

Suite: 179 tests (14 new: registry, guard both branches, gate,
migration).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
All five review blockers:

1. Producers stamp executed baselines — campaign/platform/urban/solo
   results and exhibits carry current_law (verifiably what those runs
   executed; the jct expiry-reversal's negation stays documented in its
   construction); reform-validation stamps per scoring mode via two new
   deliberately-registered worlds: pre_obbba_expiry_2026 (isolated runs'
   shared pre-OBBBA expiry baseline, per the l0 note) and
   jcx_stack_position (the position-varying stacked family, position in
   the construction string — registered honestly as a family so stacked
   runs are non-NULL and can never render plain agreement against any
   single-world claim). Committed DB re-ingested: 783 stamped results
   (603/144/36), 8,590 legacy NULLs preserved as-is.
2. Guard end-to-end: export_populations mirrors the #13 view guard per
   result (status + status_effective + claim/result baseline labels;
   summary counts by effective status); the app renders effective
   status in the pill/filter and badges per-release downgrades.
3. Migration backfill is strict: _legacy_claim_baseline_key projects
   exactly (missing key / JSON null -> current_law) and fails loudly on
   {} / [] / false / non-string policy / non-object reform_json;
   baseline_key() validates a nonempty string slug.
4. Committed artifacts honor the #9 gate: the CPSP erratum diagnosis
   rewritten through db.diagnose with its citation, and the NOTES.md
   erratum got an anchorable heading so the fragment resolves.
5. Registry provenance corrected against the vendored sources: JCX-29 =
   SFC substitute (JCX-30/31 the manager's pair); the pre-2025-tariffs
   world is keyed by TPC (161) + PWBM (22), not Tax Foundation; Senate
   Title VII cites the vendored T25-0209/T25-0215.

Tests: 189 (backfill projection + 6 malformed-descriptor cases + reopen
idempotency + producer non-NULL for campaign and RV + exporter
downgrade both branches).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The reviewer proved the jcx_stack_position family key laundered 36
distinct executed worlds into one hash — and that the claimed
mitigation was false (the construction string carried no position; the
buildi chain visibly moves through different baselines per position).

Now: stacked runs stamp one registered world per (chain, provision) —
{policy: obbba_stack_below, chain: f0af251|jcx_producer, provision: id}
— enumerated deliberately from the canonical provision map (36 entries;
the family entry is disavowed and its unreferenced registry row
removed), and the construction gains the chain position
(:stack_posNN:). The RV test now pins per-world counts (f0 chain = 2
results per provision, producer chain = 6) instead of entrenching the
collapse. Registry: 51 worlds, zero unregistered, zero effective-status
shifts.

Also: the pre-2025-tariffs provenance now states exactly where the
baseline is explicit (T25-0366's 11 manifest rows) vs assigned from the
documented workbook verification (T26-0010/0112/0113).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Round-3 findings: (1) the position token was stamped on isolated
results and counted scored-row ordinals — f0's CDCC is raw chain
position 17, after the unscored senior-deduction link that demonstrably
feeds its baseline. Stacked constructions now carry chain_posNN as the
RAW ordinal (unscored links counted); isolated rows carry no token —
their baseline is the shared expiry world. (2) register_baselines was
not convergent: an in-place upgrade would keep the disavowed
jcx_stack_position row. The seeder now prunes registry rows no longer
defined — refusing loudly if data still references them (disavowal must
re-key first) — with upgrade regression tests both ways. Stale comments
describing the removed family corrected. 191 tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per the round-4 non-blocking notes: a test now asserts f0's CDCC
carries chain_pos17 (the raw ordinal past the unscored senior link) and
that no isolated result carries a position token; the one remaining
comment describing the removed jcx_stack_position family is corrected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@MaxGhenis
MaxGhenis merged commit 58f5f34 into main Aug 19, 2026
2 checks passed
MaxGhenis added a commit that referenced this pull request Aug 19, 2026
Evidence-based rebuild on the #71 schema foundation. Every relationship
now traces to a consumption surface read at the certified pins
(2026-08-19): pe-uk-data@dd68c73 targets/sources/{obr,dwp,hmrc_spi}.py
and policyengine-uk's takeup.yaml parameters.

- Ledger routing (boundary rule 2026-08-02): 749 admin outturn cells —
  DWP recipient counts/amounts claimed (546), HMRC 2023-24 SPI-outturn
  liabilities (166), OBR FY2024-25 column (37) — route to
  data/ledger/uk_admin_outturns.jsonl with deterministic fact ids and
  the consuming pin named where pe-uk-data literally reads the cell.
  The relationship registry RAISES if an outturn ever reaches it.
- Relationships keyed exactly: OBR consumed = the 12 named Table 4.9
  forecast lines (adapter vocabulary, pe-uk-data labels quoted); PC
  take-up family = seed_source (the engine parameter cites the FYE-2020
  edition); HB = held_out (engine's stated by-definition-1.0 design;
  uk#1813 comparator); HMRC levels = seed_source (shared SPI 3.6/3.7
  base), reckoner = held_out; UKMOD = peer-held. Distribution:
  78 consumed / 1,836 seed / 14,261 held.
- Closed identity registry (scorecard_db/uk_aliases.py): (source, axis,
  value) -> canonical, raising on unknowns; HBAI/UKMOD poverty lines
  canonicalized to (basis, percent, median-vintage) so HBAI-relative
  and UKMOD-60 share a world deliberately; winter_fuel alias; explicit
  DISTINCT pairs (HB-pensioners vs HB; benefit_units vs families).
- Units repaired: gbp value_kind (never 'usd'), GBP_PER_WEEK /
  GBP_PER_MONTH / INDEX_0_1 unit concepts (Gini is an index).
- Atomic: stage + validate everything, then ONE transaction replaces
  the five sources wholesale; rollback verified by injection.
- Exact accounting: 16,175 claims + 749 ledger + 1,829 drops = every
  adapter row (the adjudication's corrected 16,924); drift raises.

Registry: 52 worlds (+hmrc_indexed_baseline_spring_2025, gate-caught).
DB 54.3 MB. Suite: 213 tests (21 UK: triangles, ledger routing,
rollback injection, accounting drift, pinned integration).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MaxGhenis pushed a commit that referenced this pull request Aug 19, 2026
…gates (#48)

UK foundation step 2 (adjudication sequence: #70 archive -> #71 schema layers -> this).

15,851 claims + 1,073 Ledger admin-outturn facts + 1,829 documented drops = all
18,753 adapter rows, drift-gated. Closed identity registry (uk_aliases) with
deliberate DISTINCT pairs; relationships evidence-keyed at pe-uk-data@dd68c73;
HBAI absolute-line anchors per the Notes sheet; reckoner rows as per-tax-head
reform worlds; single-transaction ingest with the baseline-registration gate
INSIDE the transaction.

Dual gate: fable + sol (gpt-5.6-sol, ultra), four rounds on this PR after the
two-model adjudication chose it as the UK foundation. Sol's rounds fixed, among
others: full HMRC 1990-91..2023-24 outturn span to Ledger, gbp-as-usd value
kinds, anchor-break windows as explicit mixed constructions, unit validation
before the reckoner early-return, registration running post-commit (now gates
persistence, with a row-level rollback test), and a NULL-blind semantic pin.
Final verdict MERGE-SAFE at 82faff9; CI green; suite 215.

Follow-ups tracked, not blocking: HBAI adapter rounded-text precision (upstream
#44 fix before HBAI comparisons); DB storage decision (54.95 MB, past GitHub's
50 MB warning) held for Max before long-tail sources.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
MaxGhenis added a commit that referenced this pull request Aug 20, 2026
Rebase completion (the branch forked before the whole UK arc — #43-#48,
#52, #71, #72 all landed under it):

- uk-deductions-frr (merged today, unknown to this branch) was the one
  committed lane without a country tag -> UK.
- The UK ingests' appended lane metas now carry country explicitly
  (ingest_uk_externals' five lanes + ingest_uk_deductions' one), so a
  fresh-feed append passes the new sync_lane_feed guard and never files
  a UK lane under the app's missing-key US default.
- app/public/data copies refreshed from data/ (they had drifted to a
  270-row populations.json vs 284) and pinned: new test asserts the
  committed copies byte-match data/, and a second asserts every
  committed lane carries US|UK — the two drift classes this rebase
  surfaced.

Suite: 241 python + 3 bun; oxlint + vite build clean; committed DB
byte-stable.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
vahid-ahmadi pushed a commit that referenced this pull request Aug 20, 2026
…e the country config with the executed-baseline provenance

The promised one-line GBP resolution (this branch was second to land, so
its duplicate UnitConcept.GBP goes; #48's commented one stays). The
rebase also reconciles #71's executed-baseline machinery with the
per-country config: ENGINE_VERSIONS lives inside COUNTRIES["US"] with
the back-compat alias, and _obbba_results carries both #71's position
argument (chain_pos construction) and the country/run_prefix parameters.

Suite: 248 passed, 6 skipped; ruff format clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
DTrim99 added a commit that referenced this pull request Aug 21, 2026
Rebased onto current main (picks up #74's DB-leaves-git model + #71's
baseline_key columns; clean three-way). Addresses Max's 8/20 re-gate — the
injection/secrets/architecture classes passed; three findings remained.

Finding 1 (critical — producer couldn't start): the Modal image put the
  cloned microcosm shards on PYTHONPATH without installing them, so
  microcosm.fit's eager quantile-forest import failed. Now pip-installs the
  shards (constraints file keeps the release-exact engines pinned), reads the
  installed versions back and asserts they equal the manifest, and runs a
  `backfill.py --plan` smoke before the multi-hour run so a missing dep fails
  fast.

Finding 2 (critical — attestation declarative, not verified):
  - Ingest now VERIFIES the block: h5_sha256 is 64-hex, commit ids are hex,
    and both attested engine versions must equal the artifact's own `engine`
    block — the forged `engine=9.9.9` / `h5_sha256="x"` now fail.
  - Producer rehashes the ACTUAL H5 (cached or downloaded) and stamps the
    observed hash, verifying it against the manifest; merge() requires the
    partials' single surviving revision to EQUAL the attested producer/driver
    (all-old partials can't be relabeled under a new driver); the workflow
    cross-checks the harvested artifact's Modal call id against the spawner's
    recorded id.

Finding 3 (high — post-#74 the workflow aborted): it `git add`ed the now-
  gitignored `data/scorecard.db`. Now commits only the raw artifact +
  source.json, and verifies by building the whole DB to a throwaway path
  (`scorecard_db.build_db`) — the DB is never staged.

Tests: attestation verification (bad sha / non-hex commit / engine mismatch,
  plus the clean fixture still ingests). Full RV suite 45 pass; real ingest on
  a copy of the committed DB holds the non-RV invariant (675 results x 5
  releases). Modal-runtime paths (shard install, H5 rehash, call-id
  cross-check) reasoned + unit-tested where stdlib-reachable; a live dry-run
  before enabling the schedule remains the last mile.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant