Official-calculator oracles (mode 3): closed-set extension, dated-reading provenance, calculator work list - #64
Conversation
Review — Official-calculator oracles (mode 3)Enforcement is real code, not just SCHEMA.md prose:
Should address
Suggestion
Reviewed with Claude Code assistance. |
ec749c8 to
f3e9bb8
Compare
|
Both should-address items plus the suggestion addressed in f3e9bb8 (rebased onto the updated #49 branch, a1d57d7):
Suite: 190 passed / 4 skipped; ruff clean. Force-pushed with lease for the rebase. |
Re-review — addressed ✅
Only open item is the merge order — |
a1d57d7 to
8adbc93
Compare
f3e9bb8 to
942666e
Compare
…d-reading provenance, 10-case calculator work list - Oracle enum grows five calculator ids (govuk_income_tax_estimator, govuk_hicbc_calculator, policy_in_practice_boc, entitledto, turn2us); CALCULATOR_ORACLES marks them as live services whose result rows must carry the reading date (YYYY-MM-DD) in oracle_version, enforced at validation, with the archive-the-reading rule documented in SCHEMA.md. oracle_version is now required non-empty for every oracle. - Three new boundary cases (2025-26 statutory values, verified against gov.uk/LITRG): HICBC full clawback at exactly GBP 80,000, the zero-allowance/additional-rate double boundary at GBP 125,140, and PSA ordering with GBP 1,500 savings interest (expected extra liability exactly GBP 100). Inputs + rationale only, per the battery doctrine. - battery/calculator_set.json + load_calculator_set(): the 10-case starter work list, each case into >= 2 calculators so PE-vs-UKMOD splits always have an adjudicating third reading; entries validate against the battery (unknown case, single oracle, model oracle in the set, duplicate, empty notes all raise). Suite: 174 passed / 4 skipped (was 164); ruff format clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…r rows Review items from the 2026-08-18 pass: - _contains_iso_date now parses every shape-matching chunk with date.fromisoformat, so "2026-13-45" / "0000-00-00" / "2025-02-30" fail; "2024-02-29" passes. - The SCHEMA.md archive rule is validation, not prose: a calculator-oracle row must carry an "archive: <path-or-url>" annotation citing the archived reading, or it raises. Model oracles (UKMOD/TAXSIM) are unaffected. Also rebased onto uk/ukmod-cases-schema @ a1d57d7 (schema_version, variable_class, classify->CaseResult seam), resolving the test-file overlap additively. Suite: 190 passed, 4 skipped. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
942666e to
e29e523
Compare
Gate round 1 — the closed-enum discipline is right; the seam makes every calculator call raiseClean: the Oracle enum extension is properly closed (unknown ids raise), the post-#74 virtual merge is conflict-free, and the mirrored lane feeds agree.
Suggest resolving #49's seam and persistence questions first — this PR inherits both, and its calculator-specific value (dated readings, archive citations) will land cleanly once the substrate exists. 🤖 Generated with Claude Code |
Carries the repaired #49 schema, which resolves the first blocker on its own: from_classification now accepts annotations, so calculator rows can use the required seam instead of raising, and #49's bypasses (optional variable_class, caller-supplied classifications and tolerances) are gone. 3. Dated-reading provenance is structured, not syntactic. A calculator row now carries a CalculatorReading: the reading date as a real calendar date, the exact https page, the tax year the SERVICE computed, the archive location, and sha256 digests of both the archived bytes and the canonical input vector. The oracle_version date must be the reading's own; a bare "archive:" annotation is rejected and the citation must name the reading's own archive_ref; a reading may not postdate its row; and a model oracle may not carry a reading at all. CALCULATOR_POLICY_YEARS declares what each service actually computes — gov.uk's income-tax estimator is 2026-27 only — and both the work list and every row are checked against it, which is the concrete consequence the review named. The three cases pinned to 2025 for verified 2025-26 thresholds move to 2026 (those thresholds are frozen at the same nominal values), so the assignment is honest rather than merely permitted. validate_results() checks a run against its battery: an unknown case_id and an invented variable were both accepted before. 4. Benchmark class and calibration relationship are assigned per oracle with publisher evidence, and land on the row. All seven oracles are different_model + held_out: GOV.UK describes its own tools' outputs as estimates and explicitly calls Entitledto, Turn2us and Policy in Practice "independent" calculators, so none is an authority PE is fitted to. A caller cannot relabel a calculator as authoritative. Also: the two-child multiple-birth case leaves the calculator work list. After #49 it is evaluated in the registered pre_ab2025 world, and a production calculator computes current law only — now enforced, so the abolished case is the only member of that family a calculator can answer. 2. Mode-3 persistence is still absent and is deliberately not guessed at here: it is the same table/writer/export/builder design #49 defers to the pairing Max offered, and the epistemic columns added here are written to travel on the row so they are ready for it. Suite 340 passed / 4 skipped, ruff format clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016HuXJFVme8HRbnke2Ey3Me
1. The classify -> CaseResult seam is enforced, not conventional. __post_init__ now RECOMPUTES classify() and compares. A stored classification is valid only if it equals the classifier's, or if it adjudicates a row the classifier left `unclassified` — drawn from ADJUDICATABLE and carrying a writeup. Every round-2 probe now raises: 100 vs 100 as unclassified/pe_gap, a numeric mismatch as policy_scope_mismatch with nothing said, and a boolean 1 vs 0 as match_within_tolerance with a caller-supplied tolerance of 100. pe_gap and policy_scope_mismatch are two-sided: mechanical on a null side, adjudicated (and explained) on a numeric-vs-numeric row — previously only oracle_difference and rounding demanded a writeup. Tolerances now come from DEFAULT_TOLERANCES alone; from_classification no longer takes a table, and a row carrying a different tolerance is rejected. It DOES now take annotations, which is also what made every #64 calculator call raise. 2. variable_class is required (the validator cannot check a row without it, and the shipped valid-row test omitted it) and schema_version is required rather than defaulting to the current contract. Blank engine or oracle versions are rejected. 4. The two-child-limit family isolates one mechanism per pair. Three 2026 cases: `binding` (three SEPARATE births, so no exception can stand in for the limit) and `multiple-birth` (identical household, one date of birth apart) both in the registered `pre_ab2025` world, and `abolished` — byte-identical household, same year, same rates — under current law. binding/abolished attributes the abolition; binding/multiple-birth attributes the exception. The old pair used twins on both sides, so an engine that wrongly kept the limit but rightly applied the exception still paid three elements and passed, and its year-on-year delta moved ages, rates and policy year at once. Cases can now reference a registered baseline world, which is the instrument that made the same-year pair possible. 5. Identity closure reaches the edges. Regions are a closed per-country registry (region="MARS" raises), expected_focus is closed per country and may not repeat, NaN and infinity are rejected everywhere an amount is read, a boolean-class comparison may only hold 0 or 1, the battery's `schema` path must name a contract this repo defines, and a case baseline must already be registered in baselines.py. UK private-rent cases must now pin a BRMA — LHA is set per Broad Rental Market Area, so "the rent clears the cap in most Yorkshire BRMAs" was not a pinned world — and a BRMA outside its own region raises. 6. sync_lane_feed's `updated` is documented as the caller-supplied LITERAL it is, with the constraint that every caller in a build must pass the same constant, so the next caller does not reintroduce the drift. Deliberately NOT in this commit: finding 3, the case/result table, writer, exporter and build_db step, plus the epistemic columns and the status/diagnosis split that ride on them. That is the piece with repo-wide surface that Max offered to pair on, and guessing at it unilaterally is how it ends up re-litigated. SCHEMA.md and the module docstring now state the open shape explicitly rather than leaving it implied. Suite 306 passed / 4 skipped, ruff format clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016HuXJFVme8HRbnke2Ey3Me
|
Addressed in
Also, falling out of #49: the multiple-birth case leaves the calculator work list. It is now evaluated in the registered
Suite 340 passed / 4 skipped, ruff format clean. |
Same integration finding as uk/ukmod-cases-schema, which this branch stacks on: uk/calculator-oracles forked before #74, so it had no scorecard_db/build_db.py and still tracked the committed database. Its CI was the PRE-#74 workflow, so the determinism check and the no-drift gate had never run here either. Now under the current gates: builds from scratch, two builds agree on content_hash, clean tree afterwards, suite 359 passed (up from 321). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016HuXJFVme8HRbnke2Ey3Me
|
Integration review finding — this branch had never met the post-#74 gates. It forked before #74 (the database leaving git), so it carried no Current Nothing in the branch's own code changed. |
GitHub runs a pull_request workflow from the PULL REQUEST'S OWN branch, not from the base. So a branch that forked before a gate was added keeps running the workflow WITHOUT it — and its green tick looks identical to a full one while meaning strictly less. That is not hypothetical. #74 made the database a derived artifact and added a determinism check and a no-drift gate. Branches forked before it kept the pre-#74 workflow, so neither gate had ever run against them, and they were reviewed and approved on the understanding that both had. An audit of every open PR found four in that state; two were mine (#49, #64, since fixed) and two are still open. gate-freshness.yml runs from the BASE via pull_request_target, so a stale head cannot skip it: the check is defined by main and applies to every PR regardless of what its own .github looks like. The gate set is DERIVED FROM THE BASE rather than hardcoded — it reads main's ci.yml and requires every determinism/no-drift line it finds to be present in the head's — so a gate added later is enforced on every open PR without anyone remembering to update the guard. It also refuses a branch that still tracks data/scorecard.db, whose build cannot have been from-scratch. Security: pull_request_target runs in the base repo's context, so this job NEVER checks out or executes pull-request code. It reads git metadata only and holds contents:read. A test asserts that, and asserts the guard cannot quietly become a pull_request trigger. Verified by running the exact logic against all 16 open PRs: 14 pass and exactly the two known-stale branches fail, with the reasons named. Suite 270 passed, ruff format clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016HuXJFVme8HRbnke2Ey3Me
Stacked on #49 (
uk/ukmod-cases-schema) — retarget tomainwhen #49 merges; the diff shown here is this lane's own work only.What's here
govuk_income_tax_estimator,govuk_hicbc_calculator,policy_in_practice_boc,entitledto,turn2us).CALCULATOR_ORACLESmarks them as live services: theirCaseResult.oracle_versionmust carry the reading date (YYYY-MM-DD) — enforced at validation — and SCHEMA.md documents the provenance rule (archive every reading; manual or terms-compliant access, one case at a time, never a scrape).oracle_versionis now required non-empty for model oracles too.battery/calculator_set.json+load_calculator_set(): the 10-case starter work list from Official-calculator oracles (mode 3): gov.uk and benefits-calculator household cases as a second UK oracle family #63 — each case assigned to ≥ 2 calculators so a PE-vs-UKMOD disagreement always has an adjudicating third reading; the two-child abolition pair gets three production-calculator readings each. Fail-loud validation: unknown case id, fewer than two oracles, a model oracle in the set, duplicates, and empty notes all raise.No CaseResult rows and no expected values are committed — readings come from the real calculators later, dated and archived.
Suite: 174 passed / 4 skipped (was 164);
ruff format --checkclean.Builds #63.
🤖 Generated with Claude Code