Skip to content

Ingest the two harvested-but-unused UK think-tank families: IFS + Resolution Foundation (#86) - #91

Open
vahid-ahmadi wants to merge 2 commits into
mainfrom
uk/thinktank-ingest
Open

Ingest the two harvested-but-unused UK think-tank families: IFS + Resolution Foundation (#86)#91
vahid-ahmadi wants to merge 2 commits into
mainfrom
uk/thinktank-ingest

Conversation

@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Closes #86.

The 2026-08-02 UK sweep staged seven families and five were ingested. IFS (268 rows) and Resolution Foundation (71) sat unused with their NOTES.md and manifests beside them — while the repo referenced both anyway: baselines.py registers ifs_2cl_fp_removal_rolled_out, and produce_campaign_uk declines four archived campaign rows because RF is "a long-tail source (held)".

314 claims land, and the accounting is exact and asserted:

339 staged rows = 314 ingested + 25 tallied drops

Both are independent models — the benchmark class the scorecard exists for — so every row is held_out with the consumption surfaces read at the certified pin. Where the repo-wide permanent-holdout doctrine already answers (poverty rates), it still wins.

A proposal is not a decision

145 rows carried only a harvest-side proposed_metric. DISPOSITIONS turns each of the 21 distinct names into a registered Metric or a tallied drop with a reason, and an unlisted name raises. Those 145 could have shrunk silently to zero; the accounting is what stops that.

Adopted proposals land on eight new metrics, all change or share siblings of levels already in the repo — benefit_cost_change, taxpayer_count_change, average_household_income_change, share_gaining, share_losing, spending_share, benefit_uprating_rate, real_income_growth — following the rule that keeps revenue_change apart from revenue_level.

Four names are dropped as derived: cost_per_child_lifted_out_of_poverty is a ratio of two quantities already staged separately, so PE could only "answer" it by dividing two of its own numbers — agreement would be mechanical rather than evidential.

Eight RF rows are not RF claims

They carry an attribution naming HM Treasury, a UK Parliament impact assessment, or a "Government estimate cited by RF" — and two say outright "not an RF model output". Staging them under resolution_foundation would attribute a government figure to a think tank, and a PE divergence against one would read as disagreement with a model that never produced it.

They are dropped, and the guard runs before the metric disposition — load-bearing, because all eight carry a metric this module otherwise adopts. A test pins the ordering.

Identity, which was the whole job

Both publishers name things in prose. IFS's coverage-restricted analyses get their own geographies rather than aliasing to UK — filing an England-and-Wales figure as UK is the exact fail-open the registry exists to stop — with the publication's verbatim wording kept as provenance. Deciles, quintiles and benefit names are canonical per source and recorded DISTINCT against UKMOD's quintiles and each other's deciles.

Deliberate unblock

test_blocked_targets_have_no_claims_yet fired, exactly as its docstring promised it would. RF is no longer claim-absent, so produce_campaign_uk's two blocked families get a precise reason rather than a categorical one: the archived descriptors speak the harvest's proposal vocabulary while the ingested claims speak the decided one, and the two_child fiscal-cost row targets a claim this ingest deliberately drops. Both are machine-checked now instead of asserted in prose.

Verification

  • suite 286 passed
  • two builds agree on content_hash
  • no-drift gate clean (lanes.json regenerated and mirrored)
  • ruff format --check clean

Reviewers

@MaxGhenis @DTrim99 — the two calls I'd most like challenged are (1) whether eight new sibling metrics is the right vocabulary cost for 314 rows, or whether some should have been drops, and (2) the third-party-attribution drop: I think re-publishing those eight under their originator is a separate harvest decision, but it is a judgement.

The 2026-08-02 UK sweep staged seven families and five were ingested.
IFS (268 rows) and Resolution Foundation (71) sat unused with their
NOTES.md and manifests beside them, while the repo referenced both
anyway — baselines.py registers ifs_2cl_fp_removal_rolled_out, and
produce_campaign_uk declines four archived rows because RF is "a
long-tail source (held)".

314 claims land. The accounting is exact and asserted:

    339 staged rows = 314 ingested + 25 tallied drops

Both are INDEPENDENT models, which is the benchmark class the scorecard
is for, so every row is held_out with the consumption surfaces read at
the certified pin (relationships.py) — except where the repo-wide
permanent-holdout doctrine already answers, which it still wins.

A proposal is not a decision. 145 rows carried only a harvest-side
`proposed_metric`; DISPOSITIONS turns each of the 21 distinct names into
a registered Metric or a tallied drop, and an unlisted name RAISES. The
adopted ones land on eight new metrics that are all CHANGE or SHARE
siblings of levels this repo already carries — benefit_cost_change,
taxpayer_count_change, average_household_income_change, share_gaining,
share_losing, spending_share, benefit_uprating_rate, real_income_growth
— following the rule that already keeps revenue_change apart from
revenue_level. Four names are dropped as DERIVED: a cost-per-child-
lifted-out-of-poverty is a ratio of two quantities already staged
separately, and PE could only "answer" it by dividing two of its own
numbers, which makes agreement mechanical rather than evidential.

Eight Resolution Foundation rows are not Resolution Foundation claims.
They carry an `attribution` naming HM Treasury, a UK Parliament impact
assessment or a "Government estimate cited by RF", and two say outright
"not an RF model output". Staging them under resolution_foundation would
attribute a government figure to a think tank, and a PE divergence
against one would read as disagreement with a model that never produced
it. They are dropped, and the guard runs BEFORE the metric disposition —
load-bearing, because all eight carry a metric this module otherwise
adopts.

Identity is closed, and closing it was the whole job: both publishers
name things in prose. IFS's coverage-restricted analyses get their own
geographies rather than aliasing to UK — filing an England-and-Wales
figure as UK is the exact fail-open this registry exists to stop — with
the publication's verbatim wording kept as provenance. Deciles,
quintiles and benefit names are canonical per source, and recorded
DISTINCT against UKMOD's quintiles and each other's deciles.

Deliberate unblock, as test_blocked_targets_have_no_claims_yet demands:
RF is no longer claim-absent, so produce_campaign_uk's two blocked
families get a PRECISE reason instead of a categorical one — the
archived descriptors speak the harvest's proposal vocabulary while the
ingested claims speak the decided one, and the two_child fiscal-cost row
targets a claim this ingest deliberately drops. Both are machine-checked
now rather than asserted in prose.

Suite 286 passed, two builds agree on content_hash, no-drift clean,
ruff format clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016HuXJFVme8HRbnke2Ey3Me

@DTrim99 DTrim99 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed this closely — checked out the branch, ran stage() + the full ingest() round-trip + test_uk_thinktanks_ingest.py (24 pass) against the real vendored harvest, and traced the producer/identity paths. It's a genuinely disciplined ingest; the accounting-as-enforcement approach does exactly what the PR body claims. Approving, with answers to your two calls and two non-blocking nits.

Verified (not taken on trust)

  • Accounting identity is computed and asserted, not prose: stage() raises "accounting does not close" if read != ingested + dropped, and check_accounting pins _EXPECTED and runs inside ingest(). On real data 339 = 314 + 25 (drops 12+8+2+1+1+1) reconciles.
  • Fail-loud dispositions — deleting a disposition entry makes an undecided proposal raise ("a proposal is not a decision"); confirmed live. All 145 proposal-only rows / 21 names are decided.
  • Attribution guard ordering is load-bearing and correct — the guard precedes both the DROPS and DISPOSITIONS checks, and I confirmed all 8 attributed RF rows carry a metric the module otherwise adopts (reform_fiscal_cost, revenue_change, …), so without the ordering they'd all ingest. The inspect.getsource ordering test is the right way to pin it. IFS has 0 attributed rows, so the guard can't over-drop.
  • Held-out + permanent-holdout precedence — all 314 land HELD_OUT; uk_relationship checks never_calibrate(metric) first, so poverty_rate returns PERMANENT_HOLDOUT_BASIS before the per-source branch.
  • Values never re-derived (float(row["value"]) only; exchequer cost → REVENUE_CHANGE with sign in conditions, not re-signed), geography identity (restricted-coverage IFS rows get UK_excl_* geographies with verbatim provenance — nothing silently becomes UK), and determinism (content_hash order-insensitive, excludes autoincrement id, no timestamps) all check out.

Your call #1 — is 8 new metrics the right vocabulary cost?

Mostly yes. Seven of the eight are clean change/share siblings of levels already in models.py, and each respects the revenue_changerevenue_level rule — no true duplicates. The restraint on the other side (dropping the ratios/gaps like cost_per_child_lifted_out_of_poverty, and refusing to force-map benefit_rate_weekly) is the right discipline; those would have been mechanical agreement.

The one I'd single out is average_household_income_change — it's the thinnest of the eight, sitting very close to the existing AVG_CHANGE_AFTER_TAX_INCOME_USD. The net/household-income-vs-after-tax-income distinction is real and I'd keep it, but I'd add a one-line comment on that metric spelling out what separates it from avg_change_after_tax_income_usd, so a future connector doesn't map an IFS/RF "average household income change" row onto the wrong PE quantity. Small ask, not a blocker.

Your call #2 — the third-party-attribution drop

I fully agree with dropping them, and I don't think it's a close call. Filing an HMT / UK-Parliament / "Government estimate" figure under resolution_foundation would make any PE divergence read as "PE disagrees with RF's model" against a number RF never produced — that's precisely the false signal the scorecard exists to prevent, and two of the rows say outright they aren't RF outputs. Re-publishing those eight under their true originator is a separate harvest decision, and keeping it out of an RF ingest is the correct boundary. The drop reasons preserve the attribution/verbatim, so the follow-up harvest can still find them.

Non-blocking nits

  1. The runtime BLOCKED dict in produce_campaign_uk.py still carries the old prose ("RF long-tail, held") while the precise proposal-vs-decided-vocabulary reasoning only made it into the module docstring — cosmetic/stale, and it's the runtime value a caller sees.
  2. The DISTINCT edges are correct but asymmetric/incomplete (e.g. rf:quintile_5 ↔ ukmod:q5, IFS deciles 2–9 unrecorded). Inert — the _R key is (source, axis, value) so cross-source unification is impossible regardless — but if the DISTINCT set is meant as an audit ledger it reads as half-populated.

Note

The full-DB build check failing is a pre-existing hmrc_reckoner_t2 error on main, not from this PR — worth confirming it's unrelated before reading the red mark as this branch's.

Nice work — the "accounting is what stops the silent shrink-to-zero" framing genuinely holds up in the code.

Reviewed with Claude Code assistance.

@DTrim99 DTrim99 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed this closely — checked out the branch, ran stage() + the full ingest() round-trip + test_uk_thinktanks_ingest.py (24 pass) against the real vendored harvest, and traced the producer/identity paths. It's a genuinely disciplined ingest; the accounting-as-enforcement approach does exactly what the PR body claims. Approving, with answers to your two calls and two non-blocking nits.

Verified (not taken on trust)

  • Accounting identity is computed and asserted, not prose: stage() raises "accounting does not close" if read != ingested + dropped, and check_accounting pins _EXPECTED and runs inside ingest(). On real data 339 = 314 + 25 (drops 12+8+2+1+1+1) reconciles.
  • Fail-loud dispositions — deleting a disposition entry makes an undecided proposal raise ("a proposal is not a decision"); confirmed live. All 145 proposal-only rows / 21 names are decided.
  • Attribution guard ordering is load-bearing and correct — the guard precedes both the DROPS and DISPOSITIONS checks, and I confirmed all 8 attributed RF rows carry a metric the module otherwise adopts (reform_fiscal_cost, revenue_change, …), so without the ordering they'd all ingest. The inspect.getsource ordering test is the right way to pin it. IFS has 0 attributed rows, so the guard can't over-drop.
  • Held-out + permanent-holdout precedence — all 314 land HELD_OUT; uk_relationship checks never_calibrate(metric) first, so poverty_rate returns PERMANENT_HOLDOUT_BASIS before the per-source branch.
  • Values never re-derived (float(row["value"]) only; exchequer cost → REVENUE_CHANGE with sign in conditions, not re-signed), geography identity (restricted-coverage IFS rows get UK_excl_* geographies with verbatim provenance — nothing silently becomes UK), and determinism (content_hash order-insensitive, excludes autoincrement id, no timestamps) all check out.

Your call #1 — is 8 new metrics the right vocabulary cost?

Mostly yes. Seven of the eight are clean change/share siblings of levels already in models.py, and each respects the revenue_changerevenue_level rule — no true duplicates. The restraint on the other side (dropping the ratios/gaps like cost_per_child_lifted_out_of_poverty, and refusing to force-map benefit_rate_weekly) is the right discipline; those would have been mechanical agreement.

The one I'd single out is average_household_income_change — it's the thinnest of the eight, sitting very close to the existing AVG_CHANGE_AFTER_TAX_INCOME_USD. The net/household-income-vs-after-tax-income distinction is real and I'd keep it, but I'd add a one-line comment on that metric spelling out what separates it from avg_change_after_tax_income_usd, so a future connector doesn't map an IFS/RF "average household income change" row onto the wrong PE quantity. Small ask, not a blocker.

Your call #2 — the third-party-attribution drop

I fully agree with dropping them, and I don't think it's a close call. Filing an HMT / UK-Parliament / "Government estimate" figure under resolution_foundation would make any PE divergence read as "PE disagrees with RF's model" against a number RF never produced — that's precisely the false signal the scorecard exists to prevent, and two of the rows say outright they aren't RF outputs. Re-publishing those eight under their true originator is a separate harvest decision, and keeping it out of an RF ingest is the correct boundary. The drop reasons preserve the attribution/verbatim, so the follow-up harvest can still find them.

Non-blocking nits

  1. The runtime BLOCKED dict in produce_campaign_uk.py still carries the old prose ("RF long-tail, held") while the precise proposal-vs-decided-vocabulary reasoning only made it into the module docstring — cosmetic/stale, and it's the runtime value a caller sees.
  2. The DISTINCT edges are correct but asymmetric/incomplete (e.g. rf:quintile_5 ↔ ukmod:q5, IFS deciles 2–9 unrecorded). Inert — the _R key is (source, axis, value) so cross-source unification is impossible regardless — but if the DISTINCT set is meant as an audit ledger it reads as half-populated.

Note

The full-DB build check failing is a pre-existing hmrc_reckoner_t2 error on main, not from this PR — worth confirming it's unrelated before reading the red mark as this branch's.

Nice work — the "accounting is what stops the silent shrink-to-zero" framing genuinely holds up in the code.

Reviewed with Claude Code assistance.

@DTrim99 DTrim99 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed this closely — checked out the branch, ran stage() + the full ingest() round-trip + test_uk_thinktanks_ingest.py (24 pass) against the real vendored harvest, and traced the producer/identity paths. It's a genuinely disciplined ingest; the accounting-as-enforcement approach does exactly what the PR body claims. Approving, with answers to your two calls and two non-blocking nits.

Verified (not taken on trust)

  • Accounting identity is computed and asserted, not prose: stage() raises "accounting does not close" if read != ingested + dropped, and check_accounting pins _EXPECTED and runs inside ingest(). On real data 339 = 314 + 25 (drops 12+8+2+1+1+1) reconciles.
  • Fail-loud dispositions — deleting a disposition entry makes an undecided proposal raise ("a proposal is not a decision"); confirmed live. All 145 proposal-only rows / 21 names are decided.
  • Attribution guard ordering is load-bearing and correct — the guard precedes both the DROPS and DISPOSITIONS checks, and I confirmed all 8 attributed RF rows carry a metric the module otherwise adopts (reform_fiscal_cost, revenue_change, …), so without the ordering they'd all ingest. The inspect.getsource ordering test is the right way to pin it. IFS has 0 attributed rows, so the guard can't over-drop.
  • Held-out + permanent-holdout precedence — all 314 land HELD_OUT; uk_relationship checks never_calibrate(metric) first, so poverty_rate returns PERMANENT_HOLDOUT_BASIS before the per-source branch.
  • Values never re-derived (float(row["value"]) only; exchequer cost → REVENUE_CHANGE with sign in conditions, not re-signed), geography identity (restricted-coverage IFS rows get UK_excl_* geographies with verbatim provenance — nothing silently becomes UK), and determinism (content_hash order-insensitive, excludes autoincrement id, no timestamps) all check out.

Your call #1 — is 8 new metrics the right vocabulary cost?

Mostly yes. Seven of the eight are clean change/share siblings of levels already in models.py, and each respects the revenue_changerevenue_level rule — no true duplicates. The restraint on the other side (dropping the ratios/gaps like cost_per_child_lifted_out_of_poverty, and refusing to force-map benefit_rate_weekly) is the right discipline; those would have been mechanical agreement.

The one I'd single out is average_household_income_change — it's the thinnest of the eight, sitting very close to the existing AVG_CHANGE_AFTER_TAX_INCOME_USD. The net/household-income-vs-after-tax-income distinction is real and I'd keep it, but I'd add a one-line comment on that metric spelling out what separates it from avg_change_after_tax_income_usd, so a future connector doesn't map an IFS/RF "average household income change" row onto the wrong PE quantity. Small ask, not a blocker.

Your call #2 — the third-party-attribution drop

I fully agree with dropping them, and I don't think it's a close call. Filing an HMT / UK-Parliament / "Government estimate" figure under resolution_foundation would make any PE divergence read as "PE disagrees with RF's model" against a number RF never produced — that's precisely the false signal the scorecard exists to prevent, and two of the rows say outright they aren't RF outputs. Re-publishing those eight under their true originator is a separate harvest decision, and keeping it out of an RF ingest is the correct boundary. The drop reasons preserve the attribution/verbatim, so the follow-up harvest can still find them.

Non-blocking nits

  1. The runtime BLOCKED dict in produce_campaign_uk.py still carries the old prose ("RF long-tail, held") while the precise proposal-vs-decided-vocabulary reasoning only made it into the module docstring — cosmetic/stale, and it's the runtime value a caller sees.
  2. The DISTINCT edges are correct but asymmetric/incomplete (e.g. rf:quintile_5 ↔ ukmod:q5, IFS deciles 2–9 unrecorded). Inert — the _R key is (source, axis, value) so cross-source unification is impossible regardless — but if the DISTINCT set is meant as an audit ledger it reads as half-populated.

Note

The full-DB build check failing is a pre-existing hmrc_reckoner_t2 error on main, not from this PR — worth confirming it's unrelated before reading the red mark as this branch's.

Nice work — the "accounting is what stops the silent shrink-to-zero" framing genuinely holds up in the code.

Reviewed with Claude Code assistance.

1. The runtime BLOCKED prose was stale. The precise proposal-vs-decided
   vocabulary reasoning only reached the module docstring, while the
   dict a caller actually reads still said "RF long-tail, held". Both
   blocked families now carry the real reason at runtime.

2. The DISTINCT ledger is exhaustive over what the prose claims, not a
   sample: all ten ifs:decile_N <-> ukmod:q pairs, plus both RF quintile
   and both RF decile edges.

3. ...but NOT against HBAI, and the prose is corrected rather than the
   ledger padded. HBAI registers no decile or quantile vocabulary at all
   (its subgroups are children/pensioners/working_age/total), so there
   is nothing on that side to be distinct FROM — asserting a pair
   against a value nobody registered would be the same overclaiming the
   ledger exists to prevent. A test pins that absence and why.

4. AVERAGE_HOUSEHOLD_INCOME_CHANGE now spells out what separates it from
   avg_change_after_tax_income_usd: this is a change in household NET
   income (post tax AND transfers, the concept IFS and RF publish), not
   the US tables' after-TAX income, so a connector maps an IFS/RF row
   onto the net-income quantity.

Suite 287 passed, ruff format clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016HuXJFVme8HRbnke2Ey3Me
vahid-ahmadi pushed a commit that referenced this pull request Aug 24, 2026
1. The DISTINCT prose overclaimed. It named HBAI (and the adapter named
   HMT) among the groupings ONS quintiles are distinct from, while no
   such edges existed. Corrected in the direction the evidence points:
   the ons_etb edges are now exhaustive against UKMOD (all five),
   Resolution Foundation and IFS — and HBAI is named as a DELIBERATE
   absence, because it registers no decile or quantile vocabulary at all
   (its subgroups are children/pensioners/working_age/total), so there
   is nothing on that side to assert against. Padding the ledger with
   hbai edges would have been the same overclaiming in reverse.

2. The adapter pins encoding="utf-8" on its read. The SHA-256 gate
   fixes the bytes; a platform default could still change how they are
   decoded.

3. .gitattributes forces -text on the vendored .csv, .csv-metadata.json
   and .jsonl artifacts, so a Windows checkout with autocrlf cannot
   rewrite line endings and break every SHA pin confusingly. Thank you
   for hitting that locally rather than leaving it for a contributor.

Also carries the #91 review fixes via merge.

Suite 310 passed, two builds agree on content_hash, ruff format clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016HuXJFVme8HRbnke2Ey3Me
@vahid-ahmadi

Copy link
Copy Markdown
Contributor Author

Thank you — all four addressed in the latest push, and one of them was a real defect in my own fix rather than a nit.

Nit 1 — the stale runtime prose. You were right, and it mattered more than "cosmetic". I updated the module docstring and left the BLOCKED dict alone, which is the value a caller actually reads. Both families now carry the precise reason at runtime: the proposal-vs-decided vocabulary mismatch for uprating_april2026, and for two_child the under-specified descriptor plus the fact that its reform_fiscal_cost row targets a claim this ingest deliberately drops.

Nit 2 — the DISTINCT ledger, and a correction in the other direction. Agreed it should read as an audit ledger rather than a sample, so the ifs↔ukmod edges are now all ten and both RF quintile and both RF decile edges are recorded.

But I did not add ifs/rf ↔ hbai edges, and I think that would have been the same overclaiming with the sign flipped: HBAI registers no decile or quantile vocabulary at all — its subgroups are children/pensioners/working_age/total — so there is nothing on that side to be distinct from. Asserting a pair against a value nobody registered is exactly what the ledger exists to prevent. The prose is corrected to stop naming HBAI, and a test pins both the absence and the reason.

Your call #1 — the metric comment. Added, and made sharper than a spelling note: it now says these are not the same quantity. average_household_income_change is a change in household NET income (post tax and transfers, which is what IFS and RF publish); avg_change_after_tax_income_usd is the US tables' after-tax income. So a connector maps an IFS/RF row onto the net-income quantity, not the after-tax one.

Your note on the failing full-DB build — I could not reproduce it as a main defect, and I think I know what you hit. A fresh build_db on main completes and two builds agree on content_hash; the hmrc_reckoner_t2: 14 line is the resolved-family count in the summary, not an error.

What I can reproduce is this: data/scorecard.db is gitignored (post-#74) and therefore survives branch switches. Build it on one branch, check out another, run the suite, and the campaign-producer tests fail against a database whose claims don't match the working tree. rm data/scorecard.db before running the suite after a checkout and it goes away. Worth knowing generally — it bit me twice in this session before I spotted it.

CI is green on this branch, so nothing red needs discounting.

Suite 287 passed, ruff format clean.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Ingest the two harvested-but-unused UK think-tank families: IFS (268) + Resolution Foundation (71)

2 participants