diff --git a/CHANGELOG.md b/CHANGELOG.md index c94a38a..257d127 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -8,51 +8,57 @@ As of April 18, 2026 the user asked for roughly five files per response. Keep th ## Current physical deepening notes -* Completed E002-PW2 with the supported cumulative-energy counter. All 32 - frozen factorial runs completed with exact warm binding, measurement validity - passed, and no invalidators fired. Effective counter period was 91.667 ms; - held-out arms contained 83 to 109 updates. Snapshot support was 59.30 sparse - and 110.06 dense; checkpoint-group support was 124.56 and 176.94. Total - interaction was `2.2416e-5 [2.1746e-6, 3.5305e-5] J/token`, checkpoint group - `5.8845e-6 [3.0774e-6, 8.9671e-6]`, and snapshot - `4.9917e-6 [2.8497e-6, 7.4481e-6]`. All 3 mechanism gates passed. The - sensitivity-only idle-subtracted interaction was +* Completed E002-PW2. This run used the supported cumulative-energy counter, + the energy source PW1 selected after its own counter proved too coarse. All + 32 frozen factorial runs finished with exact warm binding, measurement + validity passed, and no invalidators fired. The counter updated every + 91.667 ms on average; each held-out arm contained 83 to 109 updates. + Snapshot support was 59.30 sparse and 110.06 dense; checkpoint-group support + was 124.56 and 176.94. Total interaction was + `2.2416e-5 [2.1746e-6, 3.5305e-5] J/token`, the checkpoint group + `5.8845e-6 [3.0774e-6, 8.9671e-6]`, and the snapshot + `4.9917e-6 [2.8497e-6, 7.4481e-6]`. All 3 mechanism gates passed. One + caveat: the sensitivity-only idle-subtracted interaction was `3.9825e-6 [-8.0109e-6, 1.2479e-5] J/token` and crossed zero, so the frozen raw-cumulative primary passes without establishing baseline insensitivity. - Sparse continuation passed all 8 gates with NLL upper `0.0085037`, 3.03% attempted- - work saving, 40 opportunity ticks saved, and energy-ratio upper `1.00319`. - Conclusion: `checkpoint_cadence_attributed_sparse_continuation_survives`; + Sparse continuation passed all 8 gates with NLL upper `0.0085037`, a 3.03% + attempted-work saving, 40 opportunity ticks saved, and energy-ratio upper + `1.00319`. Conclusion: + `checkpoint_cadence_attributed_sparse_continuation_survives`; artifact `cfbca215878629bc416f169e5ded80684151d9b2a621548c7fef08207c41f8ee`. - Next is E002-PW3 multi-GPU/rack dependency-safe dephasing with simultaneous + Next is E002-PW3: multi-GPU/rack dependency-safe dephasing with simultaneous per-GPU cumulative, rack-PDU, storage, and cooling telemetry. Rare restore/rejoin estimates remain exploratory; no facility transfer is claimed. -* Executed E002-PW1's frozen 2x2 checkpoint-cadence by survivor-continuation - factorial: all 32 runs completed and the LC3 warm-state binding matched - exactly. The result is preserved as `measurement_invalid`, artifact +* Ran E002-PW1, the frozen 2x2 checkpoint-cadence by survivor-continuation + factorial. All 32 runs completed and the LC3 warm-state binding matched + exactly, but the energy measurement itself failed. The result is preserved + as `measurement_invalid`, artifact `aff76946b26876820cdaa4ca43d0b6160cdc18b2f4c5bacd053cfe92f529d4f5`. - Requested 20 ms instantaneous-power polling produced a 494.693 ms effective - device-update period with selected +250 ms lag at the frozen boundary. The - only active invalidators are `insufficient_evaluation_power_updates` and + We requested 20 ms instantaneous-power polling, but the device only updated + every 494.693 ms in practice, with a selected +250 ms lag at the frozen + boundary. The only active invalidators are + `insufficient_evaluation_power_updates` and `insufficient_pooled_cadence_phase_updates`. Raw LC3-corner energy was `0.789 [0.703, 0.923]`; raw sparse salvage was `0.823 [0.665, 1.019]`, with all non-energy gates passing. Both are - inadmissible, and all three mechanism gates failed. That failure selected the - locally supported cumulative-energy counter for E002-PW2. Its capability observation - recorded 880 polls, 40 counter changes, an 88.44 ms median gap, and 26,920 mJ - cumulative delta. PW2 kept the factorial frozen and produced the valid result - recorded above. - -* Executed the LC2-to-LC3 research redirect without overwriting failed - protocols. LC2 v1 persisted `protocol_failed_warm_start_not_late_stage`; LC2 - v2 established a valid 8,192-tick late-stage checkpoint and exact no-failure - equivalence, then persisted `protocol_failed_calibration_validity` before - held-out evaluation. LC3 replaced raw NLL first crossing with an exact - 524,288-canonical-token frontier and completed 12 held-out observations. The - candidate passed learning noninferiority, attempted-work saving, and - opportunity-tick saving, but failed solely on sampled device energy: median - adaptive/fixed 1.068 and paired 90 percent upper bound 1.134 versus the + inadmissible, and all three mechanism gates failed. That failure is what + selected the locally supported cumulative-energy counter for E002-PW2. Its + capability observation recorded 880 polls, 40 counter changes, an 88.44 ms + median gap, and 26,920 mJ cumulative delta. PW2 kept the factorial frozen + and produced the valid result recorded above. + +* Ran the LC2-to-LC3 research redirect. Failed protocols were kept on record, + not overwritten. LC2 v1 persisted + `protocol_failed_warm_start_not_late_stage`; LC2 v2 established a valid + 8,192-tick late-stage checkpoint and exact no-failure equivalence, then + persisted `protocol_failed_calibration_validity` before held-out evaluation. + LC3 replaced raw NLL first crossing with an exact 524,288-canonical-token + frontier and completed 12 held-out observations. The candidate passed + learning noninferiority, attempted-work saving, and opportunity-tick saving, + but failed on one thing: sampled device energy. Its median adaptive/fixed + ratio was 1.068 and its paired 90 percent upper bound was 1.134, above the frozen 1.05 ceiling. The persisted conclusion is `candidate_falsified_equal_canonical_work`; the result and compact observatory sidecar preserve all gates and the measured/modeled boundary. @@ -63,36 +69,36 @@ As of April 18, 2026 the user asked for roughly five files per response. Keep th * Completed E001-LC1 with 40 real RTX 3060 Laptop GPU runs: 10 calibration and 30 held-out evaluation observations. The frozen conclusion is `candidate_falsified_small_model_calibration` with - `candidate_survives_lc1=False`. Adaptive interrupted reached better median - final held-out NLL, 2.31465 versus fixed restart's 2.34115, but fixed did - exactly 12.5 percent less attempted work. The finite-horizon progress-per- - FLOP objective therefore favored fixed at worse quality, and every policy - crossed the calibration target at the first 32-tick observation. The result, - compact learning sidecar, and observatory preserve the hard split, paired - interval, falsifier outcomes, measured device boundary, and provenance. - The LC1 wave set package version 0.26.0 and selected late-stage fixed-target - E001-LC2, followed by an explicit bridge from observed learning curves to - modeled datacenter mechanics. That historical decision is preserved by the - later LC2 and LC3 records above. - -* Prior wave: executed the first recovery-backed E001 research loop. A - transition-driven runner now compares synchronous wait/restore, fixed-local - checkpoint restart, adaptive recovery, and a future-trace oracle on one - matched two-failure scenario. All policies reach durable frontier 8 with - exact work conservation. - Adaptive beats synchronous on completion time and modeled inter-site bytes, - while fixed-local beats adaptive on time and bytes and adaptive beats - fixed-local on lost work and modeled energy. The content-addressed result and - observatory artifacts preserve the `inconclusive_frontier_hypothesis` status - because learning remains a shared declared prior. The observatory renders the - four-policy failure clock, recovery episodes, work accounting, byte classes, - and missing learning evidence at Freshman, Researcher, and Full trace depth. - End-of-feature read-only verification passed all five pytest, syntax, audit, - demo, and docs-stats gates in `290.16s`. - -* Integrated the ten-step expansion wave (nine of ten branches; the SEMF - plus quark-decomposition branch remains a draft pending test - reconciliation). New surfaces: a sourced DGX H100 power BOM with an + `candidate_survives_lc1=False`. The adaptive interrupted policy reached a + better median final held-out NLL, 2.31465 versus fixed restart's 2.34115, + but fixed did exactly 12.5 percent less attempted work. So the + finite-horizon progress-per-FLOP objective favored fixed even at worse + quality, and every policy crossed the calibration target at the first + 32-tick observation. The result, compact learning sidecar, and observatory + preserve the hard split, paired interval, falsifier outcomes, measured + device boundary, and provenance. The LC1 wave set package version 0.26.0 + and selected late-stage fixed-target E001-LC2, followed by an explicit + bridge from observed learning curves to modeled datacenter mechanics. That + historical decision is preserved by the later LC2 and LC3 records above. + +* Prior wave: ran the first recovery-backed E001 research loop. A + transition-driven runner now compares four recovery policies — synchronous + wait/restore, fixed-local checkpoint restart, adaptive recovery, and a + future-trace oracle — on one matched two-failure scenario. All policies + reach durable frontier 8 with exact work conservation. No policy wins + everything: adaptive beats synchronous on completion time and modeled + inter-site bytes, fixed-local beats adaptive on time and bytes, and adaptive + beats fixed-local on lost work and modeled energy. The content-addressed + result and observatory artifacts keep the `inconclusive_frontier_hypothesis` + status because learning remains a shared declared prior. The observatory + renders the four-policy failure clock, recovery episodes, work accounting, + byte classes, and missing learning evidence at Freshman, Researcher, and + Full trace depth. End-of-feature read-only verification passed all five + pytest, syntax, audit, demo, and docs-stats gates in `290.16s`. + +* Integrated the ten-step expansion wave. Nine of ten branches landed; the + SEMF plus quark-decomposition branch remains a draft pending test + reconciliation. New surfaces: a sourced DGX H100 power BOM with an assumption-labeled full-TCO pack that resolves econ.cost.per_token end to end (3.738e-9 at the EIA 2024 industrial tariff, missing=0, 75 trace steps); sourced Pythia-160M and commercial-tariff packs @@ -106,27 +112,27 @@ As of April 18, 2026 the user asked for roughly five files per response. Keep th and a closed metadata tail (with_sp_units 1428 to 1493, with_references 1324 to 1493, equations_with_references 878 to 959, equations_with_unit_check 799 to 893). The docs-stats gate caught all - ten stale README/site coverage numbers at integration and they were - refreshed from live output. Scenario-audit now reports 99 issues - across 8 packs by design: three open sourced cost frontiers each keep - their ~33 missing economics roots visible while the closure pack - resolves 4 of 4 targets. Full pytest passed `841 passed in 256.30s` + ten stale README/site coverage numbers at integration, and we refreshed + them from live output. Scenario-audit now reports 99 issues + across 8 packs, and that is deliberate: three open sourced cost + frontiers each keep their ~33 missing economics roots visible while the + closure pack resolves 4 of 4 targets. Full pytest passed `841 passed in 256.30s` (one expected RuntimeWarning from the uncertainty failure-count test); full verifier passed `5/5 gates passed in 260.02s`; read-only full verifier passed `5/5 gates passed in 262.07s`; audit gate PASS; impeccable detect on docs/ reports only the known CLI-flag em-dash false positive. -* Finalized the portfolio form-and-deliverable polish wave. The docs site +* Finished the portfolio form-and-deliverable polish wave. The docs site moved to the three-font system from `DESIGN.md` (IBM Plex Sans reading copy, Pixelify Sans chrome and headings, IBM Plex Mono commands), gained absolute Open Graph metadata plus `og:url`, `og:type`, and `twitter:card`, converted leaked markdown backticks into real `code` elements, removed the dead empty `docs/styles.css`, null-guarded `docs/app.js` panel renders, and darkened eyebrow labels to clear 4.5:1 contrast. The impeccable static - detector now reports only a known false positive on `docs/` (it counts the - seven CLI `--flag` tokens in the console sample as em-dashes; prose has - none). README example fixes: the dependency-cone snippet sorts roots by + detector now reports only one known false positive on `docs/`: it counts + the seven CLI `--flag` tokens in the console sample as em-dashes; the prose + has none. README example fixes: the dependency-cone snippet sorts roots by name instead of comparing `Variable` objects, and `evaluate_targets` targets `training.tokens_per_sec`; both re-ran successfully. Ledger reconciliation recorded the Pythia energy-floor wave end state: scope @@ -134,7 +140,7 @@ As of April 18, 2026 the user asked for roughly five files per response. Keep th project files moved 7 to 0, and Pythia `cost_per_token` still reports 33 missing inputs, so cost closure stays on the visible backlog. Full verifier passed `4/4 gates passed in 141.07s` on this base. -* Finalized the live next-work compass and scenario-audit missing-family +* Finished the live next-work compass and scenario-audit missing-family ergonomics wave. Added `gpu_stack.next_work` with `NextWorkPlan`, `NextWorkItem`, and `build_next_work_plan(...)`; added `next-work` and `next-work --json`; added aggregate @@ -148,21 +154,21 @@ As of April 18, 2026 the user asked for roughly five files per response. Keep th `4/4 gates passed in 107.69s`; read-only full verifier passed `4/4 gates passed in 95.58s`; final source-clean check reported `cache_dirs=0 pyc_files=0 pytest_cache_dirs=0 ruff_cache_dirs=0`. -* Finalized the physical root-debt boundary hardening wave. Runtime capped - live workers at six, so bounded write lanes were tracked through a +* Finished the physical root-debt boundary hardening wave. The runtime capped + live workers at six, so we tracked the bounded write lanes through a pseudo-git coordination ledger (now archived at `archive/AGENT_GITLOG.md`). MOSFET, interconnect, lithography source/species, and - medium-response source surfaces gained boundary hardening; process geometry, - SEMF/nuclear coefficients, source-plasma drive, medium intercomponent, - root-debt, import, CLI, and boundary index/smoke-pack coverage were added or - expanded. Focused parent pack passed `125 passed in 33.75s`; full pytest + medium-response source surfaces gained boundary hardening; we added or + expanded coverage for process geometry, SEMF/nuclear coefficients, + source-plasma drive, medium intercomponent, root-debt, import, CLI, and the + boundary index/smoke pack. Focused parent pack passed `125 passed in 33.75s`; full pytest passed `628 passed in 71.99s`; audit gate PASS reported 16 systems, 1517 variables, 24 constants, 959 equations, 619 root inputs, 253 leaves, 0 cycles, 0 hard failures, 0 large scope files, and 7 large project files; full verifier passed `4/4 gates passed in 73.38s`; read-only full verifier passed `4/4 gates passed in 75.17s`; final source-clean check reported `cache_dirs=0 pyc_files=0 pytest_cache_dirs=0 ruff_cache_dirs=0`. -* Finalized the scenario-audit selector/report ergonomics wave. +* Finished the scenario-audit selector/report ergonomics wave. `SCENARIO_TARGET_SETS` and `scenario_targets_for(...)` centralize advertised scenario targets; `scenario-audit --preset` selects packs; `scenario-audit --target [LABEL=]VARIABLE` overrides advertised targets; `ScenarioReport` @@ -175,7 +181,7 @@ As of April 18, 2026 the user asked for roughly five files per response. Keep th full verifier passed `4/4 gates passed in 72.95s`; read-only full verifier passed `4/4 gates passed in 80.75s`; final source-clean check reported `cache_dirs=0 pyc_files=0 pytest_cache_dirs=0`. -* Finalized the `scenario-audit` CLI wave over +* Finished the `scenario-audit` CLI wave over `scenarios.SOURCED_SCENARIO_PACKS`. It evaluates advertised target sets via `Preset.evaluate_targets(...)`, supports text and `--json` output, and `--fail-on-issues` returns nonzero when any sourced scenario target has @@ -189,7 +195,7 @@ As of April 18, 2026 the user asked for roughly five files per response. Keep th `4/4 gates passed in 70.64s`; read-only full verifier passed `4/4 gates passed in 76.89s`; final source-clean check reported `cache_dirs=0 pyc_files=0 pytest_cache_dirs=0`. -* Finalized the structured scenario-artifact wave: +* Finished the structured scenario-artifact wave: `Preset.evaluate_targets(...)` returns a structured `ScenarioReport` with one `ScenarioTargetReport` per requested target; `MissingFamilySummary` captures grouped missing-family summaries; and `scenario-report --json` emits the @@ -203,7 +209,7 @@ As of April 18, 2026 the user asked for roughly five files per response. Keep th `cache_dirs=0 pyc_files=0 pytest_cache_dirs=0`. SEMF numeric defaults remain blocked by source and semantics; cited scenario expansion and model expansion remain open. -* Finalized the diagnostics / resolve-family / provenance wave: +* Finished the diagnostics / resolve-family / provenance wave: user-visible surfaces include `root-debt --families`, `scenario-report --missing-families`, and `resolve --missing-families` missing-input grouping. Focused integration pack passed with @@ -228,9 +234,9 @@ As of April 18, 2026 the user asked for roughly five files per response. Keep th leaving unsourced fluence, pressure, temperature, focusing, heating, and efficiency roots open. SEMF factory tests now reject empty, nonnumeric, boolean, unknown, non-root, and derived-alias assignments without publishing - coefficient defaults. Root-debt CLI determinism, gas/thermal feasibility, - medium-response domain propagation, and preset export/discovery tests were - added. Focused integration pack passed with 118 tests. Full pytest passed + coefficient defaults. We added tests for root-debt CLI determinism, + gas/thermal feasibility, medium-response domain propagation, and preset + export/discovery. Focused integration pack passed with 118 tests. Full pytest passed with 488 tests, full verifier passed 4/4 in 57.47s, read-only full verifier passed 4/4 in 62.22s, and the final source-clean check reports `cache_dirs=0 pyc_files=0 pytest_cache_dirs=0`. @@ -263,8 +269,8 @@ As of April 18, 2026 the user asked for roughly five files per response. Keep th thermal-speed, and subluminal thermal-speed constraints; source valence quark roots are positive integer primitive boundaries; medium optical-response fractions/counts/resonance ratios gained structural constraints; shared SEMF - calibration roots gained focused boundary tests; and material preset - provenance scaffolding was extended without weakening composition-only + calibration roots gained focused boundary tests; and we extended material + preset provenance scaffolding without weakening composition-only caveats. Full pytest passed with 432 tests, full verifier passed 4/4 in 48.41s, read-only full verifier passed 4/4 in 49.69s, and the final source-clean check reports `cache_dirs=0 pyc_files=0`. @@ -791,8 +797,9 @@ As of April 18, 2026 the user asked for roughly five files per response. Keep th 0 collapsed approximation-validity predicates, and 233 tests. * Added resolver domain-constraint and structural approximation-validity hardening: - approximation validity predicates that collapsed under positive SymPy - assumptions are now recovered into structural domain checks; resolver + approximation-validity predicates that used to collapse to a trivial answer + under positive SymPy assumptions are now recovered as structural domain + checks; resolver constraints also report declared variable domains for assigned and derived scenario values; `gpu-stack audit` now reports `collapsed_approximation_validity` as a hard-failure signal; and the fast @@ -921,8 +928,8 @@ As of April 18, 2026 the user asked for roughly five files per response. Keep th and constraint helper evaluation respects selected variants. * Hardened expression-LHS constraints: relations such as `x + y <= z` now wire every registered LHS variable as a constraint owner, resolver diagnostics can - discover and evaluate them, and raw unregistered LHS symbols are surfaced by - audit instead of hiding outside RHS dependency scans. + discover and evaluate them, and audit now surfaces raw unregistered LHS + symbols instead of letting them hide outside RHS dependency scans. * Extended unit checking to expression-LHS relations: `check_units=True` now infers dimensional units on both sides when the LHS is not a bare registered variable, so constraints and equations like `x + y <= z` can be validated. @@ -962,7 +969,7 @@ Split `core.py` into a `core/` package: * `core/__init__.py`: re-exports so `from ..core import var, eq, System` still works in scope files. -Real bug caught immediately by the new cycle detection: thermal.dc.pue +The new cycle detection immediately caught a real bug: thermal.dc.pue and thermal.dc.total_power define each other. Noted for pass 18. ## Pass 2: constants.py (DONE) @@ -979,9 +986,10 @@ Expanded from 10 to 23 Constants. Added: * Math helpers (not Constants, but useful): LN_2, LN_10, PI, E_MATH, TWO_PI. -All Constants now carry sp_units for the dimensional checker. Organized -into labeled sections. Sources converted to structured Reference -conceptually (kept source string for backwards compat). +All Constants now carry sp_units for the dimensional checker, and the file +is organized into labeled sections. Sources are treated as structured +Reference objects conceptually, but we kept the source string for +backwards compatibility. ## Pass 3: scopes/__init__.py (DONE) diff --git a/DESIGN.md b/DESIGN.md index a5aadd1..e9e775e 100644 --- a/DESIGN.md +++ b/DESIGN.md @@ -69,9 +69,9 @@ components: **Creative North Star: "A project window inside CuperOS"** -The GitHub Pages site must feel like it belongs to Cuper's live portfolio, not like a separate AI-generated landing page. The surface is a retro desktop app: teal dotted desktop, zero-radius window chrome, indigo title bars, pixel icons, inset content panes, taskbar, and project-copy voice that sounds human, technical, and slightly amused. +The idea is simple. Cuper's portfolio site is styled as a retro operating system called CuperOS. The GitHub Pages site for this project must look like one window inside that OS, not like a separate AI-generated landing page. That means a teal dotted desktop, zero-radius window chrome, indigo title bars, pixel icons, inset content panes, a taskbar, and project copy that sounds human, technical, and slightly amused. -The README remains a GitHub-rendered article, but any browser page for the project should inherit the CuperOS design language. The page can be visual and interactive, but it should do that as an OS control surface, not as a full-bleed marketing hero. +The README stays a GitHub-rendered article. But any browser page for the project inherits the CuperOS design language. The page can be visual and interactive, as long as it behaves like an OS control surface and never like a full-bleed marketing hero. **Key Characteristics:** - Pixel OS chrome first, modern landing-page composition never. @@ -81,7 +81,7 @@ The README remains a GitHub-rendered article, but any browser page for the proje ## 2. Colors -The palette is inherited from the portfolio OS: teal desktop, gray chrome, indigo title bars, off-white document panes, and small utility accents. +The palette comes straight from the portfolio OS: teal desktop, gray chrome, indigo title bars, off-white document panes, and small utility accents. Nothing here is new. ### Primary - **CuperOS Indigo** (`title-bar`, `title-bar-end`): title bars, selected controls, primary project identity. @@ -97,17 +97,17 @@ The palette is inherited from the portfolio OS: teal desktop, gray chrome, indig ### Named Rules -**The No New Brand Rule.** Do not invent a separate gpu_stack palette. The project page is a child window inside CuperOS. +**The No New Brand Rule.** Do not invent a separate gpu_stack palette. The project page is a child window inside CuperOS, so it uses the parent's colors. -**The Small Accent Rule.** Gold, green, and red are status lights, not brand washes. Use them as signals, not backgrounds. +**The Small Accent Rule.** Gold, green, and red are status lights, not brand washes. Use them as signals, never as backgrounds. ## 3. Typography -**The Font Law (portfolio-wide, per Cuper):** every rendered glyph, regardless of size or role, comes from the approved pixel set: DotGothic16, Pixelify Sans, VT323, Handjet, or Silkscreen. No other typeface ever renders. No exceptions for paragraphs, tables, code, or fine print. +**The Font Law (portfolio-wide, per Cuper):** every rendered glyph, at every size and in every role, comes from the approved pixel set: DotGothic16, Pixelify Sans, VT323, Handjet, or Silkscreen. No other typeface ever renders. No exceptions for paragraphs, tables, code, or fine print. **Current mapping:** Pixelify Sans carries the interface and all prose (headings, buttons, labels, status lines, paragraphs). VT323, the terminal face, carries commands, identifiers, numeric values, intervals, tables, and console output. The other three approved faces are available but unused here. -**Legibility floor:** pixel faces break down under ~11px, so nothing renders smaller. Dense instrument fine print sits at 0.7rem minimum, chart ticks at 11px. +**Legibility floor:** pixel faces break down under about 11px, so nothing renders smaller. Dense instrument fine print sits at 0.7rem minimum, chart ticks at 11px. **Character:** The pixel face IS the voice of the OS, everywhere, at every size. @@ -126,7 +126,7 @@ The palette is inherited from the portfolio OS: teal desktop, gray chrome, indig ## 4. Elevation -Depth is not blur, glass, or soft shadow. It is the retro OS physical model: `2px outset` for buttons and frames, `2px inset` for content wells, and a crisp `2px 2px 0` shadow behind windows. +Depth here is not blur, glass, or soft shadow. It is the physical model of a retro OS: `2px outset` for buttons and frames, `2px inset` for content wells, and a crisp `2px 2px 0` shadow behind windows. A surface looks raised or sunken because its border says so. ### Shadow Vocabulary - **Pixel Window Shadow** (`2px 2px 0 oklch(0.08 0.004 250)`): top-level windows only. @@ -135,7 +135,7 @@ Depth is not blur, glass, or soft shadow. It is the retro OS physical model: `2p ### Named Rules -**The Chrome Is Structure Rule.** If an element needs hierarchy, give it a real OS affordance: title bar, inset pane, status light, or taskbar. Do not fake hierarchy with decorative cards. +**The Chrome Is Structure Rule.** If an element needs hierarchy, give it a real OS affordance: a title bar, an inset pane, a status light, or a taskbar entry. Do not fake hierarchy with decorative cards. ## 5. Components diff --git a/PRODUCT.md b/PRODUCT.md index f89a9c6..81940c7 100644 --- a/PRODUCT.md +++ b/PRODUCT.md @@ -10,10 +10,11 @@ brand ## Product Purpose -`gpu_stack` is a causal, uncertainty-aware virtual AI datacenter. It joins -learning progress, training and inference execution, communication, memory, -failures, power, cooling, grid behavior, and economics in one inspectable world -model. +`gpu_stack` is a causal, uncertainty-aware virtual AI datacenter. Causal means it +models what drives what, not just what correlates. Uncertainty-aware means it +says how sure it is. It joins learning progress, training and inference +execution, communication, memory, failures, power, cooling, grid behavior, and +economics in one inspectable world model. The same engine has three inseparable jobs: @@ -24,14 +25,15 @@ The same engine has three inseparable jobs: a large AI datacenter to screen. The recursive physical graph remains valuable, but graph depth is not the -objective. A deeper lithography or particle relation is research progress only -when it improves an externally evaluated prediction, reduces decision-relevant -uncertainty, explains a residual, or enables a falsifiable experiment. +objective. A deeper lithography or particle relation counts as research +progress only when it improves an externally evaluated prediction, reduces +decision-relevant uncertainty, explains a residual, or enables a falsifiable +experiment. Success means the engine transfers to held-out hardware and workloads, carries calibrated uncertainty, recommends interventions with low decision regret, and -makes the causal reason visible. A simulation result is a hypothesis, not -evidence about the real datacenter until measurements validate it. +makes the causal reason visible. A simulation result is a hypothesis. It +becomes evidence about the real datacenter only after measurements validate it. ## Brand Personality diff --git a/README.md b/README.md index f92e0ea..5b18f7c 100644 --- a/README.md +++ b/README.md @@ -13,27 +13,27 @@ The question was simple enough to be annoying: if frontier training is supposedl Not rhetorically. Physically. -A token passes through model architecture, kernels, collectives, memory bandwidth, transistor switching, lithography, materials, thermals, power delivery, and eventually a cost line item that someone has to pay. The stack is usually explained in slices. I wanted the uncomfortable version where the slices have to talk to each other. +A token passes through model architecture, kernels, collectives, memory bandwidth, transistor switching, lithography, materials, thermals, power delivery, and eventually a cost line item that someone has to pay. Each of those layers is usually explained on its own, in a slice. I wanted the version where the slices have to talk to each other. ## What This Is Now, And How It Got Here -The project grew in three stages, and knowing the stages makes everything else legible. +The project grew in three stages. Once you know the stages, everything else in this README makes sense. -First it was an equation graph: thousands of physics and engineering relations wired together so that a question like "what does one token cost" could be traced all the way down instead of stopping at a vendor slide. +First it was an equation graph: thousands of physics and engineering relations wired together, so that a question like "what does one token cost" could be traced all the way down instead of stopping at a vendor slide. -Then the graph learned to move. Events, failures, checkpoints, power draw, multi-site traffic. A static graph became a small virtual datacenter that can replay what a training run does over time. +Then the graph learned to move. Events, failures, checkpoints, power draw, multi-site traffic. The static graph became a small virtual datacenter that can replay what a training run does over time. -Now it is a lab. The virtual datacenter runs preregistered experiments. Preregistered means the pass/fail line is frozen before the run starts, so I cannot move the goalposts after seeing the result. Measurements calibrate the engine, the engine powers the explanation, the explanation exposes its own assumptions, and experiments produce new measurements. That loop is the whole point now. +Now it is a lab. The virtual datacenter runs preregistered experiments. Preregistered means the pass/fail line is frozen before the run starts, so I cannot move the goalposts after seeing the result. Measurements calibrate the engine. The engine powers the explanation. The explanation exposes its own assumptions. Experiments produce new measurements. That loop is the whole point now. So in one sentence: GPUSTACK is a virtual AI datacenter you can interrogate. It predicts what a training run does to time, power, and money, says how sure it is, and can show you what every one of its numbers is made of. -If that sounds like a weird amount of effort to understand GPU training, yes. That is more or less how the project happened. +That is a lot of machinery for one question about GPU training. It grew this way one honest step at a time, which is more or less how the project happened. ## The Shape Of The Stack ![Dependency cone from datacenter economics down through GPU systems, transistor physics, lithography, atoms, nucleons, quarks, and equations.](docs/assets/readme-equation-cone.svg) -`gpu_stack` treats the training stack like one inspectable dependency cone. Start from a single number at the top, collect everything it depends on, and the shape that falls out is a cone: one question at the tip, hundreds of assumptions at the base. +`gpu_stack` treats the training stack as one inspectable dependency cone. Pick a single number at the top, collect everything it depends on, and the shape that falls out is a cone: one question at the tip, hundreds of assumptions at the base. At the wide end are questions people actually ask: @@ -48,7 +48,7 @@ At the narrow end are the things the model refuses to pretend away: how the chip Most tooling stops at the first satisfying number. `gpu_stack` keeps asking: what is that number made of? -The answer can be an equation, a sourced scenario value, a universal constant, or a root input. A root input is a value the model needs but cannot yet derive, so it names it instead of hiding it. Root inputs are not a shame pile. They are visible modeling debt, which is much better than hidden modeling debt wearing a lab coat. +The answer can be an equation, a sourced scenario value, a universal constant, or a root input. A root input is a value the model needs but cannot yet derive, so it names the value instead of hiding it. A root input is not a failure. It is modeling debt made visible, and visible debt is much safer than hidden debt. ## The Central Idea @@ -103,7 +103,7 @@ The model spans: | Cluster and facility | nodes, racks, bisection, storage, reliability, power, cooling, PUE | | Economics | capex, opex, amortization, power cost, run cost, cost per token | -MFU means Model FLOPs Utilization. HBM means High Bandwidth Memory. PUE means Power Usage Effectiveness. The README should not assume the reader was born knowing datacenter abbreviations. Sadly, many datacenter docs do. If half the other words in that table are new to you, that is fine. The table is a map of where things live, not a quiz. +MFU means Model FLOPs Utilization. HBM means High Bandwidth Memory. PUE means Power Usage Effectiveness. You should not need to arrive already knowing datacenter abbreviations, so this README defines them. If half the other words in that table are new to you, that is fine. The table is a map of where things live, not a quiz. ## Try It Without Believing Me @@ -189,11 +189,11 @@ total_weight root_count family boundary_c Reading the columns: `total_weight` is how many downstream variables depend on the family's roots, `family` is the group of related roots, and `primitive_boundary` marks families sitting at the edge of what the model can currently derive. The live table also appends a `top_roots` column naming the heaviest individual roots per family, truncated here for line width. -This is one of the more useful commands because it prevents the project from drifting into "add equations wherever it feels cool." The graph can tell which unknowns are currently expensive. +This is one of the more useful commands, because it stops the project from adding equations wherever it feels interesting. The graph itself can tell you which unknowns are currently expensive. ## Scenario Reports -Presets can evaluate named targets and return structured artifacts. A preset is a saved bundle of scenario assignments, so a run is reproducible instead of vibes. +Presets can evaluate named targets and return structured artifacts. A preset is a saved bundle of scenario assignments, so a run is reproducible instead of a matter of memory. ```python from gpu_stack.presets import scenarios @@ -247,7 +247,7 @@ econ.cost.per_token = 3.000078e-06 That last line reads as three millionths of a dollar per token: for this synthetic scenario, a million tokens costs about three dollars of datacenter. -That fixture is synthetic. A fixture is a fixed test anchor: deterministic on purpose, not vendor truth, historical data, or a price recommendation. The distinction matters. Fake authority is how technical debt gets a haircut and calls itself strategy. +That fixture is synthetic. A fixture is a fixed test anchor: deterministic on purpose, not vendor truth, historical data, or a price recommendation. The distinction matters. A synthetic number wearing the costume of a measurement is exactly the kind of hidden assumption this project exists to avoid. ## Resolver Workflows @@ -298,9 +298,9 @@ In plain words, the six questions: A few experiment codes appear throughout the project: LC stands for learning calibration, PW for power waveform, SC for semantic consistency. E001-LC3 is just "the third learning-calibration run of experiment one." -Two results are worth telling as stories, because they are the project behaving the way it was designed to. +Two results are worth telling as stories, because they show the project behaving the way it was designed to. -The first: E001-SC1 stress-tested the adaptive controller across six failure patterns it had never seen. In three of them the controller recognized it was outside its calibrated experience, 104 times, and each time it recorded an abstention: a logged "I do not know" plus a fallback to the safe baseline, instead of a guess. The persisted conclusion is `abstain_without_policy_claim`. The system declined to claim a win it could not support. That refusal is the result, and it is the most honest thing in this repository. +The first: E001-SC1 stress-tested the adaptive controller across six failure patterns it had never seen. In three of them, the controller recognized it was outside its calibrated experience, 104 times, and each time it recorded an abstention: a logged "I do not know" plus a fallback to the safe baseline, instead of a guess. The persisted conclusion is `abstain_without_policy_claim`. The system declined to claim a win it could not support. That refusal is the result, and it is the most honest thing in this repository. The second: E002-PW1 completed all 32 runs and then invalidated itself, because its power meter turned out to sample 25 times slower than requested. The favorable-looking raw numbers were thrown out as inadmissible instead of being quietly kept. The rerun with a valid meter, PW2, is the result that counts. diff --git a/RELEASING.md b/RELEASING.md index 857e0e7..667a558 100644 --- a/RELEASING.md +++ b/RELEASING.md @@ -1,9 +1,11 @@ # Releasing gpu_stack -This document describes how to cut a release. +This document describes how to cut a release. The short version: you bump the version, tag the commit, and push the tag. CI does the building and publishing. Everything below is the detail behind that sentence. ## Prerequisites (one-time setup) +These steps happen once, before the first release. They let GitHub Actions publish to PyPI without anyone handling an API token. + 1. Create the project on PyPI at https://pypi.org/manage/projects/ using the name `gpu_stack`. @@ -55,6 +57,8 @@ sequence: ## Building locally (optional) +You do not need this for a normal release, but it is the fastest way to check that the package builds before tagging: + ``` pip install -e ".[release]" python -m build diff --git a/RESEARCH.md b/RESEARCH.md index 61f6596..605c8ee 100644 --- a/RESEARCH.md +++ b/RESEARCH.md @@ -14,16 +14,17 @@ GPUSTACK is one system with three expressions: - **Research lab:** an experimental environment for screening datacenter-scale ML systems hypotheses before asking for expensive real-world validation. -These are a feedback loop, not three product tracks. Measurements calibrate the -engine. The engine powers the visual explanation. The visual explanation makes -the hypothesis and its assumptions inspectable. Experiments produce new -measurements. +These three are not separate product tracks. They form a feedback loop. +Measurements calibrate the engine. The engine powers the visual explanation. +The visual explanation makes the hypothesis and its assumptions inspectable. +Experiments produce new measurements, and the loop repeats. ## Scientific Position Recent systems can already simulate operator timelines accurately, optimize individual serving mechanisms, switch parallelism online, or control facility -power. The open problem is their interaction. +power. Each does one of those things well. The open problem is how they +interact. GPUSTACK should answer questions of this form: @@ -33,7 +34,8 @@ GPUSTACK should answer questions of this form: > datacenter? The primary score is not equation count, root count, test count, or in-sample -fit. The primary scores are: +fit. In-sample fit means fitting the data you trained on, which any model can +do. The primary scores are: - held-out predictive error and uncertainty coverage; - configuration-ranking and intervention decision regret; @@ -70,7 +72,8 @@ symbolic registry: ## Visual Medium Contract The primary artifact is a causal observatory, not an equal-weight dashboard. -Each result uses semantic zoom: +Each result uses semantic zoom, meaning the reader can descend through five +levels of the same result: 1. **Question:** a plain-language claim and the immediate answer. 2. **Mechanism:** the few causal paths responsible for the result. @@ -187,25 +190,27 @@ The preregistered design is in ## Current Foundation And Next Research Order -The repository now contains the observation and split contracts, held-out -evaluation and replicated-panel aggregation, deterministic temporal and -multi-site mechanics, observable-only interventions, six scalar-plus-structured -protocols, E001 recovery mechanics, the artifact-driven causal observatory, -three successive measured learning questions through E001-LC3, the completed -E002-PW1 factorial preserved as measurement-invalid evidence, and E002-PW2's -valid cumulative-energy mechanism and salvage result. E002-PW3 now has a -frozen physical rack question, dependency-safe phase scheduler, distributed -two-rank-job runtime, direct multi-boundary telemetry, compact-plus-chunked -evidence artifact, and three-depth observatory projection. It does not yet have -a physical result. - -LC2 preserved two protocol failures without opening held-out evaluation. V1's -2,048-tick checkpoint was not late-stage: NLL improved `0.0862674` over its -final 256 ticks against a frozen `0.03` ceiling. V2's 8,192-tick checkpoint -passed that gate with `0.00453499` improvement and exact no-failure -equivalence, but its target was crossed at ticks 40 and 96 rather than the -frozen 192 to 288 window because late-stage NLL was non-monotonic. These -results invalidate the protocol instances, not the recovery candidate. +Here is what the repository already contains: the observation and split +contracts, held-out evaluation and replicated-panel aggregation, deterministic +temporal and multi-site mechanics, observable-only interventions, six +scalar-plus-structured protocols, E001 recovery mechanics, the artifact-driven +causal observatory, three successive measured learning questions through +E001-LC3, the completed E002-PW1 factorial preserved as measurement-invalid +evidence, and E002-PW2's valid cumulative-energy mechanism and salvage result. +E002-PW3 now has a frozen physical rack question, a dependency-safe phase +scheduler, a distributed two-rank-job runtime, direct multi-boundary telemetry, +a compact-plus-chunked evidence artifact, and a three-depth observatory +projection. It does not yet have a physical result. + +LC2 preserved two protocol failures without opening held-out evaluation. In +plain terms: both attempts disqualified themselves before touching the +evaluation data. V1's 2,048-tick checkpoint was not late-stage: NLL, the +negative log-likelihood training loss, improved `0.0862674` over its final 256 +ticks against a frozen `0.03` ceiling. V2's 8,192-tick checkpoint passed that +gate with `0.00453499` improvement and exact no-failure equivalence, but its +target was crossed at ticks 40 and 96 rather than the frozen 192 to 288 +window, because late-stage NLL was non-monotonic. These results invalidate the +protocol instances, not the recovery candidate. LC3 then compared fixed restart and adaptive continuation at the exact same 524,288-token canonical frontier across six untouched held-out schedules. @@ -213,19 +218,20 @@ Adaptive passed learning noninferiority: adaptive-minus-fixed NLL had median `0.0033385` and 90% interval `[0.00239279, 0.00850366]`, below the frozen `0.01` margin. It saved a median `3.0303%` attempted work and 40 opportunity ticks, and was earlier in all six schedules. The candidate was nevertheless -falsified solely on measured training-device energy: the adaptive/fixed ratio -had median `1.06839` and 90% interval `[1.001795, 1.134269]`, above the frozen -`1.05` upper bound. +falsified on a single measurement: training-device energy. The adaptive/fixed +energy ratio had median `1.06839` and 90% interval `[1.001795, 1.134269]`, +above the frozen `1.05` upper bound. -This is measured small-model learning and sampled board energy from one local -GPU. Opportunity ticks are simulated. Datacenter concurrency, WAN, storage, -host, cooling, and facility-energy behavior remain modeled or unmeasured. +Keep the scope of that result in view. It is measured small-model learning and +sampled board energy from one local GPU. Opportunity ticks are simulated. +Datacenter concurrency, WAN, storage, host, cooling, and facility-energy +behavior remain modeled or unmeasured. E002-PW1 executed that frozen 2x2 with exact LC3 warm binding and all 32 arms -complete. It did not produce an admissible attribution. Requested 20 ms NVML -polling yielded an effective 494.693 ms update period; the selected +250 ms lag -sat at the frozen boundary. The only active invalidators were -`insufficient_evaluation_power_updates` and +complete. It did not produce an admissible attribution, because its power meter +could not keep up. Requested 20 ms NVML polling yielded an effective 494.693 ms +update period; the selected +250 ms lag sat at the frozen boundary. The only +active invalidators were `insufficient_evaluation_power_updates` and `insufficient_pooled_cadence_phase_updates`, so the result is `measurement_invalid`. @@ -248,23 +254,23 @@ continuation passed all eight salvage gates with NLL upper bound `0.0085037`, ratio upper bound `1.00319`. The valid local conclusion is `checkpoint_cadence_attributed_sparse_continuation_survives`. -The frozen primary uses raw cumulative energy over the complete run window. -The idle-subtracted sensitivity was +One caveat stands. The frozen primary uses raw cumulative energy over the +complete run window. The idle-subtracted sensitivity was `3.9825e-6 [-8.0109e-6, 1.2479e-5] J/token` and crossed zero, so PW2 does not show that its local attribution is insensitive to estimated idle-baseline treatment. PW3 packages the rack-transfer question as an optional executable physical -calibration. At least -four independent two-rank jobs execute the same useful work and failure -windows under synchronized release, seeded legal jitter, storage-only pacing, -static cohorts, and visible-telemetry feedback. Only dependency-safe timing of -checkpoint capture/persist, state transfer, communicator rebuild, and rejoin -may move. UUID-bound GPU telemetry is aligned with direct rack-PDU, storage -activity, measured storage power, and cooling channels; missing boundaries -produce `measurement_invalid` instead of a modeled substitute. The current -machine cannot execute the claim because it has one GPU and no tenant-visible -rack meters. That boundary does not block the software research program. +calibration. At least four independent two-rank jobs execute the same useful +work and failure windows under synchronized release, seeded legal jitter, +storage-only pacing, static cohorts, and visible-telemetry feedback. Only +dependency-safe timing of checkpoint capture/persist, state transfer, +communicator rebuild, and rejoin may move. UUID-bound GPU telemetry is aligned +with direct rack-PDU, storage activity, measured storage power, and cooling +channels; missing boundaries produce `measurement_invalid` instead of a modeled +substitute. The current machine cannot execute the claim because it has one GPU +and no tenant-visible rack meters. That boundary does not block the software +research program. E001-SC1 has since completed the software-first semantic-consistency loop. Calibration selected `periodic_local`; adaptive switching failed its learning, diff --git a/ROADMAP.md b/ROADMAP.md index 628431b..e42aef9 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -10,10 +10,11 @@ The project now optimizes one closed loop: The June engine is a strong symbolic foundation. Its live `next-work` ranking now follows the measured E001 sequence; root-debt closure and graph -completeness remain diagnostic evidence, not the research roadmap. Recent work such as Charon -already reports operator-level training and inference simulation below 5.35% -overall timing error. `gpu_stack` will not compete by becoming another static -performance simulator with more equations. +completeness remain diagnostic evidence, not the research roadmap. The reason +is competitive honesty: recent work such as Charon already reports +operator-level training and inference simulation below 5.35% overall timing +error. `gpu_stack` will not compete by becoming another static performance +simulator with more equations. The research target is a causal, uncertainty-aware virtual datacenter that couples learning dynamics, temporal execution, failures, facility power, grid @@ -33,7 +34,7 @@ gates. ## Virtual Datacenter Foundation: Implemented -The first research substrate is no longer roadmap prose: +The first research substrate is no longer roadmap prose. It exists: - observations, calibration/evaluation splits, residual metrics, stratified coverage, Kendall tau-b ranking, decision regret, and repeated benchmark @@ -65,17 +66,19 @@ read-only full verifier `5/5` in `320.32s` with a 600-second gate ceiling. ## Recovery-Backed E001 Vertical Slice: Executed -The feature branch now carries one complete research path: scenario input, -transition-driven failure/recovery execution, four matched policies, a -content-addressed result, and a three-depth observatory projection. All four -runs reach durable frontier 8 and conserve attempted, retained, and lost work. +The feature branch now carries one complete research path from end to end: +scenario input, transition-driven failure/recovery execution, four matched +policies, a content-addressed result, and a three-depth observatory +projection. All four runs reach durable frontier 8 and conserve attempted, +retained, and lost work. -On this deterministic trace, adaptive recovery beats synchronous wait and -restore by 48 ms and 1.6 GB while losing 66.21 PFLOP less work. Fixed-local -restart is 20 ms faster and moves 3.2 GB fewer bytes than adaptive, but loses -96.07 PFLOP more work and uses 0.028 MJ more modeled energy. The oracle ties -adaptive on time and traffic and loses more work. The mechanics therefore -produce a Pareto split, not a validated controller hierarchy. +What did the runs show? On this deterministic trace, adaptive recovery beats +synchronous wait and restore by 48 ms and 1.6 GB while losing 66.21 PFLOP less +work. Fixed-local restart is 20 ms faster and moves 3.2 GB fewer bytes than +adaptive, but loses 96.07 PFLOP more work and uses 0.028 MJ more modeled +energy. The oracle ties adaptive on time and traffic and loses more work. So +each policy wins somewhere and loses somewhere else. The mechanics produce a +Pareto split, not a validated controller hierarchy. The observatory renders the same result at Freshman, Researcher, and Full trace depth, including the shared failure clock, restore/replay intervals, exact work @@ -89,15 +92,17 @@ observations and 30 untouched held-out evaluation observations across six strata. The frozen decision is `candidate_falsified_small_model_calibration`; `candidate_survives_lc1=False`. -Adaptive survivor continuation ended at better held-out loss: median final NLL -2.31465 versus fixed-local restart's 2.34115. The frozen finite-horizon -progress-per-FLOP objective nevertheless favored fixed restart because it -stopped with 458,752 attempted tokens rather than 524,288, exactly 12.5 percent -less work, while every policy crossed the 3.13759 target at the first 32-tick -observation. LC1 therefore published the candidate failure instead of -retuning the target or the six held-out schedules. The result artifact, -learning sidecar, and observatory projection preserve the measured curves, -paired interval, falsifier outcomes, device-energy boundary, and provenance. +The details explain why the candidate lost despite learning well. Adaptive +survivor continuation ended at better held-out loss: median final NLL 2.31465 +versus fixed-local restart's 2.34115. The frozen finite-horizon +progress-per-FLOP objective nevertheless favored fixed restart, because fixed +restart stopped with 458,752 attempted tokens rather than 524,288, exactly +12.5 percent less work, while every policy crossed the 3.13759 target at the +first 32-tick observation. LC1 therefore published the candidate failure +instead of retuning the target or the six held-out schedules. The result +artifact, learning sidecar, and observatory projection preserve the measured +curves, paired interval, falsifier outcomes, device-energy boundary, and +provenance. ## E001-LC2 Protocol Sequence: Preserved, Not Rewritten @@ -106,8 +111,9 @@ late-stage under the frozen gate. LC2 v2 established a valid 8,192-tick late-stage checkpoint and exact no-failure equivalence, then stopped because the calibration-only NLL target had already been crossed outside its frozen 192-to-288-tick validity window. Neither LC2 artifact is candidate evidence. -Together they exposed raw first crossing as an invalid late-stage endpoint and -selected LC3's equal-canonical-work question without smoothing or retuning. +That does not make them useless. Together they exposed raw first crossing as an +invalid late-stage endpoint and selected LC3's equal-canonical-work question +without smoothing or retuning. ## E001-LC3 Equal-Canonical-Work Result: Executed @@ -116,16 +122,17 @@ the exact same 524,288-canonical-token frontier. Adaptive was learning- noninferior, saved a median 3.03 percent attempted work, and saved a median 40 opportunity ticks. Those gates passed. Its sampled training-device-energy ratio had median 1.068 and paired 90 percent upper bound 1.134, above the frozen 1.05 -ceiling. The sole failed gate makes the conclusion +ceiling. One failed gate is enough: the conclusion is `candidate_falsified_equal_canonical_work`. The compact observatory sidecar is `docs/data/e001-equal-work-v1.json`. ## E002-PW1 Checkpoint-Power Result: Measurement Invalid PW1 completed all 32 frozen factorial runs with exact LC3 warm-state binding. -Its requested 20 ms logger observed effective NVML updates every 494.693 ms, -with selected +250 ms lag at the frozen boundary. The only active invalidators -are `insufficient_evaluation_power_updates` and +Then the measurement itself failed. Its requested 20 ms logger observed +effective NVML updates every 494.693 ms, with selected +250 ms lag at the +frozen boundary. The only active invalidators are +`insufficient_evaluation_power_updates` and `insufficient_pooled_cadence_phase_updates`. The preserved conclusion is `measurement_invalid`. @@ -330,6 +337,7 @@ and preset export/discovery tests. Build the research programs in `RESEARCH.md`: adaptive multi-datacenter training, power-waveform shaping, semantic fault tolerance, fluid inference topology, heterogeneous architecture co-design, and firm grid-responsive -inference. Every program must move through virtual screening, held-out -calibration, and counterfactual comparison. External physical validation is a -later adapter only for claims whose measurement boundary requires it. +inference. Every program must move through the same three steps: virtual +screening, held-out calibration, and counterfactual comparison. External +physical validation is a later adapter, used only for claims whose measurement +boundary requires it. diff --git a/archive/README.md b/archive/README.md index 6fc0ac8..7d270ba 100644 --- a/archive/README.md +++ b/archive/README.md @@ -1,10 +1,12 @@ # archive/ -This directory holds historical agent-session memory files moved from the +This directory is a historical archive. It holds agent-session memory files — +notes an agent wrote to itself while working — that were moved here from the repository root on 2026-06-10. The files are kept byte-identical for -provenance. They are not operational ledgers: they record the inner thread of -earlier agent sessions, coordination pseudo-git-log entries, break-room pauses, -and session start-here instructions from the first large pass of the project. +provenance: what you read here is exactly what was written then, unchanged. +They are not operational ledgers. They record the inner thread of earlier +agent sessions, coordination pseudo-git-log entries, break-room pauses, and +session start-here instructions from the first large pass of the project. ## What lives here @@ -19,12 +21,14 @@ and session start-here instructions from the first large pass of the project. ## Why they are here and not at root -The four canonical operational ledgers at root are: +The root of the repository is reserved for documents you need today. The four +canonical operational ledgers at root are: `CHANGELOG.md`, `SESSION_STATE.md`, `HANDOFF.md`, and `VISIBLE_BACKLOG.md`. Planning docs are `ROADMAP.md` and `IMPROVEMENT_MAP.md`. Reference docs are `README.md`, `PRODUCT.md`, and `DESIGN.md`. -The files archived here are per-session memory artifacts. They do not serve -day-to-day project navigation. Keeping them at root added noise to the root -inventory without adding operational value. They are preserved here in full -for historical reference. +The files in this archive are different. They are per-session memory +artifacts: a record of what one session thought and felt, useful for history +but not for finding your way around the project now. Keeping them at root +added noise to the root inventory without adding operational value. So they +live here instead, preserved in full for historical reference. diff --git a/docs/app.js b/docs/app.js index b01c0c1..9179aef 100644 --- a/docs/app.js +++ b/docs/app.js @@ -1,4 +1,5 @@ -// Published guide data. Keep these values aligned with the README summaries. +// The guide copy and numbers shown on the page. Keep them aligned with the +// README summaries, because readers will compare the two. const primerSteps = { target: { text: "Start with a human question, then follow the named dependencies upstream. Every hop should tell you whether you are looking at an equation, a scenario value, or an unresolved root input.", @@ -174,7 +175,8 @@ const traceMeterFoot = document.getElementById("trace-meter-foot"); const traceMeter = document.getElementById("trace-meter"); const clock = document.getElementById("clock"); -// Render helpers replace only panel content, preserving the surrounding markup. +// Each render helper swaps out only the content inside its panel. The +// surrounding markup never changes, so the page structure stays stable. function renderPrimer(key) { const step = primerSteps[key]; primerText.textContent = step.text; @@ -269,7 +271,8 @@ targetTabs.forEach((tab) => { tab.addEventListener("click", () => renderTrace(tab.dataset.target)); }); -// Guard each panel so one missing hook cannot take down the others. +// Each panel checks its own DOM hooks before rendering, so one missing +// element cannot take down the other panels. if (primerText && primerFacts && primerStatusTitle && primerStatusBody) { renderPrimer("target"); } @@ -308,9 +311,10 @@ if (clock) { }); })(); -// Scroll reveals. The hidden pre-reveal state only exists once html.js-anim -// is set, so readers without JS, without IntersectionObserver, or with -// reduced motion requested always get the fully visible page. +// Scroll reveals. Content is only hidden after html.js-anim is set, so a +// reader without JS, without IntersectionObserver, or with reduced motion +// requested always gets the fully visible page. Hiding first and revealing +// later would break for exactly those readers. (function initReveals() { const motionOk = !window.matchMedia("(prefers-reduced-motion: reduce)").matches; if (!("IntersectionObserver" in window) || !motionOk) { diff --git a/docs/cone-browser.js b/docs/cone-browser.js index 76dfb77..fb3b1c1 100644 --- a/docs/cone-browser.js +++ b/docs/cone-browser.js @@ -1,17 +1,19 @@ /** * cone-browser.js * --------------- - * Vanilla JS dependency-cone browser for the gpu_stack portfolio page. + * The dependency-cone browser on the gpu_stack portfolio page, in plain + * vanilla JS. A dependency cone is everything a target variable depends on, + * upstream hop by upstream hop. * - * Fetches docs/data/registry-cone.json (pre-generated by the - * `export-graph-json` CLI subcommand), then renders the upstream cone of - * the selected target as an expandable OS-styled tree. + * It fetches docs/data/registry-cone.json (pre-generated by the + * `export-graph-json` CLI subcommand) and renders the upstream cone of the + * selected target as an expandable OS-styled tree. * - * Degrades gracefully when the JSON fails to load (file:// or network - * error): shows an informative notice instead of a broken panel. + * If the JSON fails to load (file:// or a network error), the panel shows an + * informative notice instead of breaking. * - * Keyboard: all interactive elements are