From 84ffb3a683303779e5f8f88992bd9363294a9237 Mon Sep 17 00:00:00 2001 From: Matt McKay Date: Tue, 1 Sep 2026 08:59:13 +1000 Subject: [PATCH] PLAN: record that the static migration completed 2026-08-18 The status line still said "24 of 41, 17 to go" from the 2026-08-11 audit; the 2026-08-31 audit.json reports 40 of 40 static datasets migrated and repointed, 23 live-API lectures, 24 orphans, 0 legacy refs and 2 URL forms. Tracks B, C and D are marked complete with their landing, flip and validation PRs; Track X carries the current per-repo orphan count and its workspace tracker; Phase 9's repoint box is ticked and the builder-recovery note now distinguishes the three recovered pipelines from the two still unrecovered. Co-Authored-By: Claude Fable 5 --- PLAN.md | 26 +++++++++++++------------- 1 file changed, 13 insertions(+), 13 deletions(-) diff --git a/PLAN.md b/PLAN.md index f53bf48..441fe33 100644 --- a/PLAN.md +++ b/PLAN.md @@ -1,10 +1,10 @@ # PLAN — `data-lectures` (formerly `QuantEcon/data`) -**Status:** active roadmap (last updated 2026-08-12) — **the repo is LIVE**: the first repoint merged 2026-07-17 (P1, `lingcod_msy_recovery.csv` → `msy_fishery`), so published filenames are an API from here on +**Status:** active roadmap (last updated 2026-09-01) — **the repo is LIVE**: the first repoint merged 2026-07-17 (P1, `lingcod_msy_recovery.csv` → `msy_fishery`), so published filenames are an API from here on -**Where the numbers stand (`audit.json`, 2026-08-11):** 24 of 41 static datasets migrated and repointed, 17 to go; 22 lectures still fetch live API data; 26 committed orphans; 0 legacy-repo references; 4 URL forms in use. +**Where the numbers stand (`audit.json`, 2026-08-31):** **40 of 40 static datasets migrated and repointed — the static migration completed 2026-08-18** ([#98](https://github.com/QuantEcon/data-lectures/pull/98), [#99](https://github.com/QuantEcon/data-lectures/pull/99)); 23 lectures still fetch live API data (Track E); 24 committed orphans (Track X); 0 legacy-repo references; 2 URL forms in use. -**`CATALOG.md` and `migrated` both say 24 today, and that agreement is temporary.** They count different things: `migrated` counts datasets whose *consumers* read this repo; the catalog counts datasets that *live* here. They agree only when no wave is in flight. The P3 fold is the worked example — the six `high_dim_data` files landed 2026-08-10 at `status: landed` with `consumers: []`, opening a six-file gap that closed on 2026-08-11 when PR set C repointed them and [#69](https://github.com/QuantEcon/data-lectures/pull/69) flipped the records. Expect that gap again for the duration of any wave that lands ahead of its repoints, which is the convention here. +**`CATALOG.md` and `migrated` both say 40 today, and that agreement holds only while no wave is in flight.** They count different things: `migrated` counts datasets whose *consumers* read this repo; the catalog counts datasets that *live* here. They agree only when no wave is in flight. The P3 fold is the worked example — the six `high_dim_data` files landed 2026-08-10 at `status: landed` with `consumers: []`, opening a six-file gap that closed on 2026-08-11 when PR set C repointed them and [#69](https://github.com/QuantEcon/data-lectures/pull/69) flipped the records. Expect that gap again for the duration of any wave that lands ahead of its repoints, which is the convention here. Every figure on that line comes from `stats` in the generated `audit.json`, and every figure below that restates one is a copy that can drift — as all seven of them had by 2026-08-07, each understating progress by three repoint sets. Re-read them from `audit.json` before quoting them, and prefer citing the generated file over this document. @@ -20,7 +20,7 @@ This repository is being shaped into the **single canonical repository for data | [QuantEcon/meta#338](https://github.com/QuantEcon/meta/issues/338) | Pilot: migrate one dataset per hosting pattern, landing here | | [QuantEcon/QuantEcon.manual#108](https://github.com/QuantEcon/QuantEcon.manual/pull/108) | Draft `styleguide/datasets.md` — the convention's design surface | | [QuantEcon/data#1](https://github.com/QuantEcon/data/issues/1), [#2](https://github.com/QuantEcon/data/issues/2), [#4](https://github.com/QuantEcon/data/issues/4) | Pre-existing execution items (LFS, fold in `high_dim_data`, repoint lectures) | -| [QuantEcon/workspace-lectures#14](https://github.com/QuantEcon/workspace-lectures/issues/14) | Session work plan: pilot kickoff sequencing (meta#337 risks → P1 → P2) | +| [QuantEcon/workspace-lectures#14](https://github.com/QuantEcon/workspace-lectures/issues/14) | Standing cross-repo tracker for the migration — tracks A–E and X–Y, repoint rules, outstanding write-backs (it began as the pilot's session work plan) | ## Where we are @@ -214,18 +214,18 @@ The remaining work decomposes by **consuming series** rather than by hosting pat | Track | Datasets | Coupling | Blocked on | | --- | --- | --- | --- | | **A — `intro` + `wasm`** | 17, **all done**. The last two CSVs landed as wave A4 ([#74](https://github.com/QuantEcon/data-lectures/pull/74), flipped in [#75](https://github.com/QuantEcon/data-lectures/pull/75)); `graph.txt` was never a migration — see below | — | — | -| **B — `python.myst`** | 7, cut into two waves. **B1′ done**: the `ols` trio, `fp.dta` and `NEWQDATA.csv` landed in [#79](https://github.com/QuantEcon/data-lectures/pull/79), flipped in [#80](https://github.com/QuantEcon/data-lectures/pull/80). **B2′**: `hansen_singleton_1982/1983_data.csv` | **three consumers, not one** — `lecture-python.zh-cn` reads by URL *and* holds byte-identical copies of all 7 plus both builders (and is outside `SCAN_REPOS`, so the audit cannot see it); `lecture-python.notebooks` lags a publish tag; `lecture-stats` carried a published-site prose link to `fp.dta` behind a daily linkcheck. B2′ adds a fourth kind: the two builders migrate too, and each lecture names them twice outside its data cell | nothing | -| **C — `advanced.myst`** | 6: `dataBHS.mat`, `acs_data_summary.csv`, `bbh` ×2, `fred_data.csv`, `hansen_jagannathan_1991_data.json` | none | nothing (builder recovery is in-wave work, not a gate) | -| **D — `programming`** | 1: `test_pwt.csv` | none | nothing — a single-PR track | +| **B — `python.myst`** | 7, **all done**, cut into two waves. **B1′**: the `ols` trio, `fp.dta` and `NEWQDATA.csv` landed in [#79](https://github.com/QuantEcon/data-lectures/pull/79), flipped in [#80](https://github.com/QuantEcon/data-lectures/pull/80). **B2′**: `hansen_singleton_1982/1983_data.csv` landed in [#82](https://github.com/QuantEcon/data-lectures/pull/82), flipped in [#83](https://github.com/QuantEcon/data-lectures/pull/83); both waves independently validated in [#84](https://github.com/QuantEcon/data-lectures/issues/84) | **three consumers, not one** — `lecture-python.zh-cn` reads by URL *and* holds byte-identical copies of all 7 plus both builders (and is outside `SCAN_REPOS`, so the audit cannot see it); `lecture-python.notebooks` lags a publish tag; `lecture-stats` carried a published-site prose link to `fp.dta` behind a daily linkcheck. B2′ adds a fourth kind: the two builders migrate too, and each lecture names them twice outside its data cell | nothing | +| **C — `advanced.myst`** | 6, **all done**. **C1**: the `bbh` pair and `hansen_jagannathan_1991_data.json` landed in [#92](https://github.com/QuantEcon/data-lectures/pull/92), flipped in [#95](https://github.com/QuantEcon/data-lectures/pull/95), validated in [#96](https://github.com/QuantEcon/data-lectures/issues/96). **C2**: `fred_data.csv`, `acs_data_summary.csv` and `dataBHS.mat` (converted to `dataBHS.csv`) landed in [#98](https://github.com/QuantEcon/data-lectures/pull/98), flipped in [#99](https://github.com/QuantEcon/data-lectures/pull/99), validated in [#100](https://github.com/QuantEcon/data-lectures/issues/100) | none by URL — but six other org repos hold byte-identical copies of `acs_data_summary.csv` and `dataBHS.mat`, and `lecture-tools-techniques` publishes its own read of the latter, so acceptance was scoped to advanced's own URLs | — | +| **D — `programming`** | 1, **done**: `test_pwt.csv` rode with wave C2 ([#98](https://github.com/QuantEcon/data-lectures/pull/98), [#99](https://github.com/QuantEcon/data-lectures/pull/99)); its four consuming repos are recorded in [#101](https://github.com/QuantEcon/data-lectures/pull/101) | none | — | | **E — dynamic / live-API** | the UNRATE twin, then the 15 incidental API lectures | wasm is the forcing customer | [#14](https://github.com/QuantEcon/data-lectures/issues/14) schema decisions, [#26](https://github.com/QuantEcon/data-lectures/issues/26) fetch layer | -| **X — orphan sweep** | 26 committed orphans across 6 repos — dp 10, programming 5, wasm 5, intro 3, python.myst 2, `continuous_time_mcs` 1 | per repo | that repo's repoints landing first | +| **X — orphan sweep** | 24 committed orphans across 6 repos (`audit.json`, 2026-08-31) — dp 10, programming 5, intro 4, python.myst 2, wasm 2, `continuous_time_mcs` 1; tracked at [QuantEcon/workspace-lectures#57](https://github.com/QuantEcon/workspace-lectures/issues/57) | per repo | nothing — every repo's repoints have landed; `graph.txt` needs its own per-repo reader sweep (below) | | **Y — consumer interface (`qeld`)** | the `qeld` package, Q1–Q7 of `PLAN-QELD-PACKAGE.md` — audit support, the package, pilots, then adoption by win; QEP graduation stays | — | nothing — re-scoped 2026-08-12 (D11): the DNS → custom domain → URL-sweep sequence this row used to carry is retired | **`graph.txt` was closed out as a non-migration (2026-08-12).** It is synthetic teaching data — `provenance: toy`, null in every real provenance field — and the shortest-path exercise teaches its format by quoting the first line, so the data has to stay visible on the page. Hosting it here would have put a toy in a registry that exists to carry provenance. Instead `lecture-wasm` stopped fetching intro's committed copy over the network and embeds it with `%%file` like every sibling ([QuantEcon/lecture-wasm#63](https://github.com/QuantEcon/lecture-wasm/pull/63)), which retired the last cross-repo read of that blob anywhere in the organisation. `graph.txt` consequently no longer appears as a scanned dataset at all. Four repos embed it via `%%file` — intro, dp, jax and wasm — and two of those (intro, dp) also commit a copy the cell overwrites before reading, so those two are shadowed orphans; jax and wasm commit none, which is the cleaner shape. The remaining committed copies (`lecture-intro.zh-cn`, the canary, `lecture-python.zh-cn`, `lecture-dp.monorepo`, `ipynb_pdf_constructor`) are read by nothing. Intro's committed copy is now deletable as Track X — but the same blob sits in **7** repos byte-identically (a further two, `QuantEcon.jl` and `QuantEcon.lectures.code`, hold a 4,692-byte variant differing by one trailing space) and is regenerated at 17 `%%file` sites, including archived `.rst` ancestors that `gh search code` cannot see, so that deletion needs its own per-repo reader sweep rather than an org-wide sweep. `lecture-dp`, `lecture-jax` and `continuous_time_mcs` are **not data consumers** — dp's 10 committed files are inherited orphans, jax embeds `graph.txt` via `%%file`, and continuous_time_mcs has one orphan scratch file. They appear only in Track X. -**Tracks A–D are independent of each other and can run in any order or in parallel.** The only hard dependencies in the whole programme are: `usa-gini-nwealth-tincome-lincome.csv` is built from `SCF_plus_mini.csv` (so it follows the SCF migration inside Track A); Track E's rollout needs its own template proven first; Track X follows its repo's repoints; and Track Y's adoption sweep (qeld Q7) is last. +**Tracks A–D are complete (2026-08-18); E, X and Y remain.** They were independent of each other and ran in the order A → B → C+D. The only hard dependencies left in the programme are: Track E's rollout needs its own template proven first; Track X's `graph.txt` deletion needs a per-repo reader sweep; and Track Y's adoption sweep (qeld Q7) is last. **Track Y was re-scoped 2026-08-12: invest in `qeld`, defer the custom domain indefinitely** (`PLAN-QELD-PACKAGE.md` D11, where the reasoning is recorded in full). The short form: the package delivers everything the domain would have — a stable interface point that survives backend rework — plus tidier lectures and call-site metadata, without the forever-promise of a branded public host, and so without [#35](https://github.com/QuantEcon/data-lectures/issues/35)'s promotion gate ever coming due. It is a commitment decision, not an effort one: the DNS record remains two actions QuantEcon controls (the stale A record was deleted; the name is NXDOMAIN as of 2026-08-10) and [#37](https://github.com/QuantEcon/data-lectures/issues/37) stays open as deferred-not-dead, reopenable at any time because qeld's base URL stays on the raw forms (never `quantecon.github.io` — D11's redirect trap). The classifier constraint this paragraph used to carry moves with the re-scope: `classify_url` must learn the `qeld.url('X')` pattern **before** any consumer adopts it (qeld Q1), for the same structural reason it would have had to learn the canonical host before a URL sweep — otherwise every migrated read classifies as broken and the dashboard inverts. @@ -305,7 +305,7 @@ Full automation: ### Phase 6 — Metadata backfill for existing holdings -- [ ] Manifest per dataset for the files now in `lectures/`: source, license, retrieval date, schema, consumers, provenance class. Schema sketched in `manifest-schema.yml` (Phase 2); backfill is per-file work gated on the license check below. **Largely done** — 27 non-`.yml` files, 24 manifests; the only dataset still lacking one is `business_cycle_data.csv` (`business_cycle_info.md` and `business_cycle_metadata.md` are prose, not datasets, so the gap is 1 file and not 3) +- [ ] Manifest per dataset for the files now in `lectures/`: source, license, retrieval date, schema, consumers, provenance class. Schema sketched in `manifest-schema.yml` (Phase 2); backfill is per-file work gated on the license check below. **Largely done** — 43 non-`.yml` files, 40 manifests (2026-09-01); the only dataset still lacking one is `business_cycle_data.csv` (`business_cycle_info.md` and `business_cycle_metadata.md` are prose, not datasets, so the gap is 1 file and not 3) - [ ] Classify: the 8 static intro files are author-assembled or verbatim; `business_cycle_data.csv` is the one dynamic snapshot and needs its cadence declared - [ ] Licence check **per source**, not per file: the question is *"may this source be cached and served publicly, with attribution?"* — a cheap binary gate (`redistribution: permitted | restricted`, see AGENTS.md "Licensing and attribution"), a fast yes for public data sources. Two sources already answered: World Bank is **CC BY-4.0** (`business_cycle_metadata.md`, the model for what a manifest should capture) and RAM Legacy is **CC BY 4.0** (established against its Zenodo DOI record, P1). The remaining sources need the equivalent established by hand @@ -346,9 +346,9 @@ The first end-to-end deployment: one dataset per hosting pattern, each the harde ### Phase 9 — Adoption (broad sweep — the step that stalled in Feb 2025) -- [ ] Repoint the remaining consuming lectures as datasets land here (data#4) — **17 datasets**, organised as tracks A–D above. Mechanical, but see "Repoint rules": repoint all consumers of a dataset together, and never delete a copy a sibling repo reads -- [ ] Remove lecture repos' duplicate copies as each repoint merges (tracked with the orphan sweep in meta#337) — 26 orphans today, Track X. Note the wasm mirror copies are only safe to delete **after** wasm reads data-lectures directly, not before -- [ ] Intake rule for migrations: constructed datasets arrive **with their builders**; the 5 known constructed-but-unscripted files (`hansen_jagannathan_1991_data.json`, `fred_data.csv`, the two `bbh` extracts, `acs_data_summary.csv`) need their pipelines recovered or rewritten — recorded as QEP follow-ups per meta#338 +- [x] Repoint the remaining consuming lectures as datasets land here (data#4) — **done 2026-08-18**: all 40 static datasets in the corpus are migrated and `repointed` (tracks A–D above; the last four landed in [#98](https://github.com/QuantEcon/data-lectures/pull/98) and flipped in [#99](https://github.com/QuantEcon/data-lectures/pull/99)). The "Repoint rules" stay binding on any future wave: repoint all consumers of a dataset together, and never delete a copy a sibling repo reads +- [ ] Remove lecture repos' duplicate copies as each repoint merges (tracked at [QuantEcon/workspace-lectures#57](https://github.com/QuantEcon/workspace-lectures/issues/57) now that meta#337 is closed) — 24 orphans today (`audit.json`, 2026-08-31), Track X. Note the wasm mirror copies are only safe to delete **after** wasm reads data-lectures directly, not before +- [ ] Intake rule for migrations: constructed datasets arrive **with their builders**; of the 5 known constructed-but-unscripted files, three arrived with recovered builders in Track C (`fred_data.csv` and the two `bbh` extracts — see `builders/README.md`); `hansen_jagannathan_1991_data.json` and `acs_data_summary.csv` still have none — recorded as QEP follow-ups per meta#338 - [ ] Graduate the convention to a QEP and merge manual#108, with the remaining sweep as its rollout checklist ## Open decisions (owned by meta#336 / manual#108, not this repo)