From 7bcb2cc9da12b6b28a9a9a8ba37bd4f9bc7a5fe5 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 03:25:09 +0000 Subject: [PATCH 01/22] docs(topics): start finding-your-unknowns-integration Brief Interview contract for integrating the verified Finding-Your-Unknowns corpus (slice finding-your-unknowns-0f25bd45): round-1 decisions locked (vehicles, binding codification posture, conditional-verdict evidence pass, vertical order, targeted live-doc checks). Brief accumulates as rounds resolve; the Plan section stays empty for /planning:plan. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- .../finding-your-unknowns-integration/PLAN.md | 64 +++++++++++++++++++ 1 file changed, 64 insertions(+) create mode 100644 docs/topics/finding-your-unknowns-integration/PLAN.md diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md new file mode 100644 index 000000000..f21432264 --- /dev/null +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -0,0 +1,64 @@ +# finding-your-unknowns-integration + +## Brief + +### TLDR + +Integrate the verified "Finding Your Unknowns" corpus (Thariq Shihipar's field guide, its +X-Article draft, the 13 html-effectiveness pages, and the context-engineering companion; +slice `finding-your-unknowns-0f25bd45`) into this repo as judgment-preserving deltas on +existing skills plus a small set of citable reference docs, gated by a scripted +comparison-evidence pass over the ~20 named-skill collisions. + +### Goal + +Every corpus decision in +`.work/finding-your-unknowns/finding-your-unknowns-0f25bd45/corpus-inventory.md` reaches a +disposition (adopt / adapt / compare-then-adopt / cite / drop / treat-as-caution) that is +either executed as a repo change or recorded with its reason, with no silent drops. + +### Constraints + +1. Vehicles (interview Q1): (a) augmentation of existing skills and (b) citable + reference/convention docs are the primary vehicles; at most 2 genuinely new thin skills, + each only where the evidence pass confirms a real gap; CLAUDE.md/rules changes only if + the context-engineering material earns one on its own evidence. +2. Codification posture (Q2, BINDING): honor the source author's anti-premature-codification + warning — only judgment-preserving deltas (output contracts, checklists, conventions with + rationale); no generator-style "make me an X" skills; the warning itself is quoted in + whatever reference doc graduates. +3. Evidence discipline (Q3): adopt/adapt verdicts on named-skill collisions are CONDITIONAL + ("adopt X into S if S lacks it") until a read-only evidence pass grades each named skill + against its checkable claims; only surprises return to the human. +4. Vendor-claim discipline (Q5): corpus claims are vendor-blog anecdote unless the targeted + live-doc check (folded into the evidence pass) verifies them; the ~6 decision-relevant + harness claims (auto-memory, /doctor, ToolSearch deferred loading, artifacts-as-context, + 80%-claim scoping, rich references) get that check; nothing else does. +5. House style: all new prose obeys the repo's ai-slop/house-style rules + (.claude/rules/vendor-docs-are-not-style.md); citation shape follows + plugins/knowledge/reference/citation-shape.md (URL + retrieval date + content hash). + +### Acceptance criteria + +- (accumulating; interview in progress — see .work/finding-your-unknowns-integration/ + interview-checklist.md for the live decision ledger) + +### Captured assumptions + +- The published blog is the canonical citation source for wording; the X draft is citable + only for draft-only content (5 images, 3 links) — per corpus V7.1. +- Arm-B verification ran degraded (same-vendor adversarial refuter); accepted as sufficient + for this corpus. + +### Out-of-scope + +- Watching/transcribing the two linked videos (companion-classified; no ingest path). +- Re-digesting the 3 referenced-external related-posts articles. + +### Deferred questions + +- (populated as the interview defers items) + +## Plan + +(unfilled — /planning:plan territory) From 68baec8e08129f687f8f4731ba3512ea610cd4dd Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:06:54 +0000 Subject: [PATCH 02/22] docs(topics): lock the finding-your-unknowns integration contract at sign-off Operator sign-off (".confirm all", 2026-09-01) closes the interview: the Brief gains the seven acceptance criteria (sign-off Part G.5), the signed constraint set (evidence-gate classification, registry discipline, ctx-eng sequencing, schema stability), and the deferred-questions record with per-item triggers and arbiters. The signed decision sheet (rev 2, both final validators folded) is committed beside it as the durable input contract for /planning:plan, including the operator's delivery amendment: one session, one branch, one PR, waves as commit ordering. Register gates clean: registered=17 open=0 answered=17, brief=ok. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- .../finding-your-unknowns-integration/PLAN.md | 64 ++++- .../signoff-sheet.md | 219 ++++++++++++++++++ 2 files changed, 280 insertions(+), 3 deletions(-) create mode 100644 docs/topics/finding-your-unknowns-integration/signoff-sheet.md diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index f21432264..75d801abc 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -37,11 +37,53 @@ either executed as a repo change or recorded with its reason, with no silent dro 5. House style: all new prose obeys the repo's ai-slop/house-style rules (.claude/rules/vendor-docs-are-not-style.md); citation shape follows plugins/knowledge/reference/citation-shape.md (URL + retrieval date + content hash). +6. Execution contract (Q6): each conditional-batch unit is closed only when its verdict is + executed-or-recorded and the inventory row links the outcome; no silent drops. +7. Vehicle placements (Q7-Q10): the five-pass sequencing lands as a workflow section in the + central reference doc with cross-refs from planning:wayfind and session-flow:workflow (no + new orchestration skill); the reference doc owns the prompt-pattern catalog with one + canonical invocation line per owning skill; reply-affordance is a house convention + (default-with-judgment) owned by the doc with an artifact-design cross-ref. +8. Evidence-gate classification (sign-off S1): every delta is classified per-row under + PLUGIN-PHILOSOPHY's rubric (docs/PLUGIN-PHILOSOPHY.md:699-742). CONTRACT/POLICY/CONVENTION + rows land now as team conventions adopted by the sign-off; DOC rows are citation/doc + lines; BEHAVIORAL rows never land as standing instructions — they become reference-doc + heuristic lines plus candidate eval cases, awaiting observed-stumble evidence. +9. Registry discipline (sign-off S2): convention-registry rows for reply-affordance and + export-button only, each with an explicit conformance surface, landing owner-doc-first + (owner doc before a second adopting plugin, PLUGIN-PHILOSOPHY:594-595). +10. Sequencing with docs/topics/context-engineering-claude-5/ (sign-off S3): this effort + never edits that topic's audit-instructions criteria files; Wave 1 records F2's three + candidates in that topic's PLAN plus a Phase-10 re-inventory/rebase note; no freeze. +11. Schema stability (sign-off G.4 / D37): the `### Phase N` heading/tag vocabulary and any + parsed schema (check-open-questions.sh fields) are never renamed without a version bump + and changelog; block reordering is safe. + +The complete signed decision surface — classified delta roster (D/E/F/G rows), wave +assignments, per-finding dispositions — is ./signoff-sheet.md (rev 2, operator-confirmed +".confirm all" on 2026-09-01). ### Acceptance criteria -- (accumulating; interview in progress — see .work/finding-your-unknowns-integration/ - interview-checklist.md for the live decision ledger) +(Signed off with the sheet, 2026-09-01; sign-off sheet Part G.5.) + +1. Every decision in corpus-inventory.md has an executed-or-recorded disposition traceable + from ./signoff-sheet.md — verified by grep over the disposition lines, not asserted. +2. All waves green on the per-plugin gates: version bump + CHANGELOG entry; + check-changed-skills.sh (trigger-keyword preservation, listing cap, --require-evals); + listing budget respected; cheat-sheet regenerated once per PR series, series sequential; + scripts/affected-tests.sh --run green. +3. The F1 reference doc carries the quoted anti-premature-codification warning and the C6 + permission/quoting basis in its header. +4. Registry rows (reply-affordance, export-button) carry explicit conformance surfaces and + land owner-doc-first. +5. docs/topics/context-engineering-claude-5/PLAN.md carries the S3 sequencing rows. +6. Every Wave-2/Wave-3 contract delta lands with eval expectations in the same commit as + its contract lines (sign-off Part D eval-impact column). +7. /planning:plan consumes ./signoff-sheet.md + this Brief as its input contract. +8. Delivery (operator amendment at sign-off): all execution in one session on branch + claude/reading-feedback-j4sg96, delivered as ONE pull request; waves are commit + ordering, not separate PR series; deferred items may become filed issues. ### Captured assumptions @@ -57,7 +99,23 @@ either executed as a repo change or recorded with its reason, with no silent dro ### Deferred questions -- (populated as the interview defers items) +(No open-register rows were retired as deferred; the items below are sub-decisions the +sign-off explicitly deferred, each with its trigger and arbiter.) + +- E4-EXT | E4 skill extension (planning:prd durable-output pitch mode vs a design-handoff + layer) — deferred behind demand evidence; the doc-tier buy-in pattern ships now (S4). + Arbiter: human, at the recorded-candidate evidence point. +- D28-SCHEMA | Adding a mechanical scrutiny-flag field to the 5-field open-question register + schema — the resolution-field convention ships now, gate-invisible by design; the schema + change waits until a consumer needs mechanical reads (M4). Arbiter: that consumer's + change, via version bump + changelog per constraint 11. +- EVAL-CAND | Behavioral eval candidates (D5, D7, D11, D17b, D34-residue) — doc lines now; + promotion to standing skill instructions only on observed-stumble evidence per the + evidence-gated-additions rule (PLUGIN-PHILOSOPHY:699-711). Arbiter: that rule. +- Q11-FLIP | Flipping tweak-likelihood ordering from documented-default to requestable mode + in further consumers — deferred to own-usage evidence (interview Q11 residue). +- E1-REG | A convention-registry row for the deviation-log — retracted as premature (M3); + fires the moment a second plugin reads DEVIATIONS.md (recorded trigger, C5). ## Plan diff --git a/docs/topics/finding-your-unknowns-integration/signoff-sheet.md b/docs/topics/finding-your-unknowns-integration/signoff-sheet.md new file mode 100644 index 000000000..33e81e4b5 --- /dev/null +++ b/docs/topics/finding-your-unknowns-integration/signoff-sheet.md @@ -0,0 +1,219 @@ +# Final sign-off sheet — finding-your-unknowns integration (rev 2) + +2026-09-01. THE single review artifact for this topic, and (with the Brief in +./PLAN.md) the input contract /planning:plan consumes. Rev 2 folds both final +validators' challenges; every amendment is tagged [A]/[B]. + +**SIGNED: operator ".confirm all", 2026-09-01.** Every "Default: CONFIRM" row below +is confirmed; Parts A-G are adopted in full. + +Chain of custody: corpus (17 verified digest slices, slice +`finding-your-unknowns-0f25bd45`) -> interview locks -> evidence pass (7 graders + +targeted live-doc checks) -> validators A/B -> explore+research (fresh-verified) -> +blindspots -> brainstorm -> devils-advocate -> this sheet -> final validators A/B -> +rev 2 -> operator sign-off. The session working files behind each stage lived under +`.work/finding-your-unknowns-integration/` (unversioned by design); this committed +copy is the durable record. + +## Part A — structural fixes (confirmed) + +- S1 [DA-C1; refined per A+B] Evidence-gate compliance. Reading stated explicitly: + "standing instruction" = always-loaded contract text in a skill body. Every delta + is classified per-row under PLUGIN-PHILOSOPHY:699-742. CONTRACT/POLICY/CONVENTION + rows land now under the durable-tier carve-out as TEAM CONVENTIONS ADOPTED BY THIS + SIGN-OFF, citing corpus + industry grounding. DOC rows are citation/doc lines. + BEHAVIORAL rows never land as instructions: they become F1 doc lines + candidate + EVAL cases (evals outlive instructions), awaiting observed-stumble evidence. + Reclassified in rev 2 on validator evidence: D5 -> BEHAVIORAL [A+B], D34 -> + CORROBORATION + behavioral residue -> eval candidate (plan Steps 4.6/4.7 already + own the invariant) [A+B], D17 split into D17a contract / D17b behavioral [A]. +- S2 [DA-H1] Registry justification: owner-doc-before-a-SECOND-ADOPTING-PLUGIN rule + (PHILOSOPHY:594-595). Applies to reply-affordance + export-button only (see E-part; + E1's row is retracted per M3 [B]). +- S3 [DA-H2 + L4] Sequencing with docs/topics/context-engineering-claude-5/: + (a) this effort never edits the audit-instructions criteria files (zero overlap, + stated); (b) Wave 1 records F2's three candidates in that topic's PLAN; + (c) one line there: Phase 10's sweep re-inventories current state (its recorded + counts are stale [L4]) and rebases over landed waves. No freeze. +- S4 [DA-H3] E4 closure: five-section buy-in pattern + E5 objection-evidence + checklist land in the F1 DOC (durable, tracked; cites Rust RFC / Oxide RFD / + Amazon PR-FAQ + corpus). Skill extension (prd durable-output mode vs + design-handoff layer) deferred as a recorded candidate behind demand evidence. +- S5 [DA-H4] This sheet is the complete decision surface; rev 2 adds the rows both + validators found missing (F4, D18, D37, D3/D13 dispositions) and the M2 + eval-impact column. +- S6 [DA downgrade, verified] Em-dash reconciliation unnecessary: rule-em-dash is + disabled repo-wide (.claude/ai-slop.json, recorded rationale). F1 quotes stay + byte-verbatim; one F1 header line records the quoting posture [M1]. + +## Part B — interview locks (human-confirmed rounds 1-2; record) + +- Q1-Q5: vehicles; binding codification posture; conditional verdicts + evidence + pass; all seven verticals V2-first; targeted live-doc checks. +- Q6-Q12: batch execution contract; workflow section + cross-refs; no + brainstorm-first posture; central pattern catalog + one canonical line per skill; + reply-affordance convention (default-with-judgment); tweak-likelihood discharged + (plan already defaults to presentation ordering — reconciliation line in F1); + stop-and-wait port gate, token recommended-not-required. + +## Part C — open questions signed with this sheet (Q13-Q17 + amendments) + +- C1 [Q13] Merged round-3 fixes as amended by rev 2: D3 demoted; D13 dropped + + D14 scope-conditioned; D31 reconciliation line; D11 restored as doc-tier; + G-block adopted; plus rev-2 reclasses (S1). CONFIRMED. +- C2 [Q14/E4] Buy-in home per S4 (doc now, skill later on evidence). CONFIRMED. +- C3 [Q15/E6] Port gate in point-dont-copy via declared-step-deltas, scoped to + external-reference ports ("source of truth outside this repo's tree: vendored, + foreign-language, other-repo"); in-tree corrections stay do-it-now. + CONFIRMED with boundary sentence verbatim. +- C4 [Q16/F4] quiz-me answer-key fresh-context requirement — a classified row + (Part D, CONTRACT, W2) [A+B]. CONFIRMED. +- C5 [Q17/E1] Deviation-log ships as OPT-IN convention: F1 owner section + + implement-pipeline reference + RECORDED-TRIGGER NOTE in lieu of a registry row + ("the moment a second PLUGIN reads DEVIATIONS.md, the registry rule fires") — + rev 2 retracts the premature registry row per M3 [B; A concurs]. + CONFIRMED opt-in with trigger note. +- C6 [licensing, blindspot 4 restored in full [B]] F1's header records the + permission basis: short attributed verbatim excerpts from the named author's + public posts under fair-quotation practice; no license claimed; citation-shape + citations; bulk reproduction avoided; x.com lychee excludes. CONFIRMED. + +## Part D — classified delta roster (eval-impact column per M2) + +Waves: W1 = docs/registry/glossary/sequencing. W2 = CONTRACT/POLICY rows landing in +skill bodies (6 plugins). W3 = rows needing their own review moment (E6, E1+E2) — +the stated exception to the W2 mapping [B]. Format: +row | target | delta | CLASS | wave | eval impact. + +- D1 | discovery:blindspot | typed finding taxonomy | CONTRACT | W2 | extend +- D4 | discovery:blindspot | scan-scope disclosure line | POLICY | W2 | extend +- D5 | planning:brainstorm | observed-fact evidence bar | BEHAVIORAL -> F1 + eval + candidate [A+B reclass] | W1(doc) | eval-candidate +- D7 | planning:brainstorm | disconnected-work heuristic | BEHAVIORAL -> F1 + eval + candidate | W1(doc) | eval-candidate +- D9 | planning:brainstorm | brainstorm-practice citation line | DOC | W1 | skip +- D11 | improvement:find | disconnected-work recipe | BEHAVIORAL -> F1 + eval + candidate | W1(doc) | eval-candidate +- D12 | education:explain | vocabulary ladder | CONTRACT | W2 | extend +- D14 | education:explain | success condition (scoped to original-ask invocations) | + CONTRACT | W2 | extend +- D16 | education:quiz-me | source-anchor + on-miss routing | CONTRACT | W2 | extend +- D17a | education:quiz-me | diff-sourced question authoring | CONTRACT | W2 | extend +- D17b | education:quiz-me | non-obvious-behavior keying | BEHAVIORAL -> F1 + eval + candidate [A split] | W1(doc) | eval-candidate +- D18 | verification:confirm | quiz-layer cross-ref line (the REROUTE executed; + quiz-me untouched) [B restore] | DOC | W1 | skip +- D19 | verification:confirm | "existing behavior this leans on" callout | CONTRACT | + W2 | extend +- D20 | prototype:explore-directions | same-data control-variable rule | POLICY | + W2 | extend +- D21 | prototype:explore-directions | structured steal/graft capture | CONTRACT | + W2 | extend +- D22 | prototype:explore-directions | machine-legible reply template | CONTRACT | + W2 | extend +- D24 | prototype:pressure-test | validation-answer-set shape | CONTRACT | W2 | extend +- D25 | prototype:pressure-test | fake-data disclosure footnote | POLICY | W2 | extend +- D26 | prototype:pressure-test | per-option named costs | CONTRACT | W2 | extend +- D27 | prototype pair | mock-before-wire composition note | DOC | W1 | skip +- D28 | planning:interview | free-text scrutiny flag AS RESOLUTION-FIELD CONVENTION + (branch chosen on script evidence: 5-field enum, check-open-questions.sh:31-32, + :173; gate-invisible BY DESIGN — known limitation recorded; schema change deferred + until a consumer needs mechanical reads) [M4 resolved] | CONTRACT | W2 | extend +- D32 | planning:plan | switch condition on alternatives | CONTRACT | W2 | extend +- D33 | planning:plan | closing revision replies | CONTRACT | W2 | extend +- D34 | planning:plan | collapse self-check | CORROBORATION (Steps 4.6/4.7 own it) + - behavioral residue -> eval candidate [A+B reclass] | W1(doc) | eval-candidate +- D35 | planning:plan | lands-green forward-reference | DOC | W1 | skip +- D36 | planning:design | design gains the same tweak-likelihood presentation + ordering plan already documents (SKILL.md:223) [A reword] | CONTRACT | W2 | extend +- E7 | discipline:point-dont-copy | primitive->convention trap item | CONTRACT | + W2 | extend +- E8 | discipline:point-dont-copy | canonical invocation line | DOC | W1 | skip +- F4 | education:quiz-me | fresh-context answer-key requirement [A+B add] | + CONTRACT | W2 | extend +- E6 | discipline:point-dont-copy | semantics map + confirmation gate + (external-reference ports, boundary per C3) | CONTRACT | W3 | extend +- E1+E2 | implement pipeline | opt-in deviation-log convention + fold-back step + (per C5; trigger note, no registry row) | CONVENTION | W3 | extend +- E5 | F1 doc | objection-evidence checklist section | DOC | W1 | skip +- Recorded dispositions, no change: D2, D3 (demoted), D6, D8, D10, D13 (dropped), + D15, D23, D29, D30, D31 (+reconciliation line), E3, F3 [A coverage fix]. +- D37 | PLAN.md schema constraint -> Part G rule 4 [B restore]. + +Eval pricing [M2]: W2 touches 6 plugins; the planning-plugin PR concentrates 4 +contract deltas (D28, D32, D33, D36) + their expectations — budget it as its own +review; education PR carries 5 (D12, D14, D16, D17a, F4). "Extend" means the +skill's evals.json gains expectations for the new contract lines in the same PR; +authorship = the executing wave session; review = the PR reviewer. + +## Part E — doc/registry/glossary placements (Wave 1) + +- F1 doc at docs/ (name at plan time): taxonomy + lifecycle; pattern catalog; + reply-affordance convention + export-button rule; when-HTML taxonomy; buy-in + pattern + E5; codification warning quoted; HTML-scoping rule; C6 permission + basis; D5/D7/D11/D17b/D34-residue heuristics as doc lines; Q11/D31 + reconciliation line. +- Registry rows (S2): reply-affordance and export-button ONLY, each with explicit + conformance surface [M5]: "conformance = the template blocks in + prototype:explore-directions (D21/D22) and prototype:pressure-test (D24)", + landing owner-doc-first with adopters in W2. +- Glossary: curate-language dispositions for unknowns quadrants + D1's four finding + types; map/territory cite-only (G1). +- G1-G16: as merged in the round-3 audit (audit/merge-2026-09-01.md rows G1-G16). + CONFIRMED en bloc. +- Sequencing rows per S3 into the ctx-eng topic PLAN. + +## Part F — devils-advocate finding dispositions (per-finding [A]) + +- M1 em-dash: F1 header line only; no config change; addendum line 3 dropped. DONE + in rev 2 (S6). +- M2 eval obligation: eval-impact column + pricing added (Part D). DONE. +- M3 E1 registry row: retracted; recorded-trigger note (C5). DONE. +- M4 D28 branch: chosen on script evidence (Part D row). DONE. +- M5 registry conformance surfaces: stated per row (Part E). DONE. +- M6 Brief completeness: sign-off closed interview Steps 3-5; acceptance criteria + (Part G) written into the Brief at sign-off. DONE. +- L1 cap headroom: verified ample; no action. L2: fixed by F4/E5 rows. L3: + cheat-sheet contention — wave PR series stay sequential; regenerate once per + series (Part G). L4: folded into S3. + +## Part G — execution shape + definition of done + +1. Three waves; W2 = 6 plugins (discovery, planning, education, verification, + prototype, discipline) [B arithmetic fix]; W1 = docs + registry + glossary + + sequencing + the W1(doc) rows; W3 = E6, E1+E2. +2. Per wave and per plugin: version bump + CHANGELOG entry; + check-changed-skills.sh green (trigger-keyword preservation, listing cap, + --require-evals); check-listing-budget respected [B]; cheat-sheet regenerated + once per series, series sequential [L3]; scripts/affected-tests.sh --run green. +3. Eval expectations land in the same PR as their contract lines (Part D column). +4. [D37] The `### Phase N` heading/tag vocabulary and any parsed schema + (check-open-questions.sh fields) are never renamed without version bump + + changelog; block reordering is safe. +5. Acceptance criteria (written into the Brief at sign-off): (1) every inventory + decision has an executed-or-recorded disposition traceable from this sheet — + verified by grep over the disposition lines, not asserted; (2) all waves green + on rule-2 gates; (3) F1 carries the codification warning + C6 basis; (4) + registry rows carry conformance surfaces and land owner-doc-first; (5) ctx-eng + topic PLAN carries the S3 rows; (6) every W2/W3 contract delta has same-PR eval + expectations; (7) /planning:plan consumes this sheet + Brief as its input + contract. + +## Operator amendment at sign-off (2026-09-01) + +Delivery vehicle amended by the operator with the sign-off: ALL execution happens +in this session, on the single feature branch `claude/reading-feedback-j4sg96`, as +ONE pull request. Consequences, superseding the conflicting phrasing above: + +- The three waves survive as COMMIT ORDERING and review structure inside the one + branch (W1 docs commits, then per-plugin W2 commits, then W3), not as separate + PR series. Filing follow-up issues for deferred items is fine; landed work is not + split across PRs. +- Part G rule 2 gates run on the branch head: per-plugin version bump + CHANGELOG + entries all land in the same PR; cheat-sheet regenerated once at the end (L3's + sequential-series concern is moot with a single branch). +- Part D's eval pricing ("budget the planning PR as its own review") becomes + reviewer guidance for the corresponding commits within the single PR. +- "Same-PR eval expectations" (rule 3, criterion 6) is satisfied trivially by the + single PR, but the intent is kept stricter: eval expectations land in the SAME + COMMIT as their contract lines. From 3c95327339cbcf7204fac0ec240bce2e9335b9da Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:16:04 +0000 Subject: [PATCH 03/22] docs(topics): draft the finding-your-unknowns execution plan (pre-review) Ten phases sequencing the signed contract: F1 reference doc, governance placements (registry rows, glossary, ctx-eng sequencing bullets), six per-plugin contract-delta phases with same-commit eval expectations, the Wave-3 heavies (E6 port gate, E1/E2 deviation-log convention), and the close-out phase (issues, cheat-sheet check, acceptance verification, the single PR). Includes the durable G-block placement copy, the execution shape (sequential main-session), the gate-passed decisions table, and the Tier-C design-gate early-exit record. A fresh-context plan review is in flight; confirmed findings will amend this draft before execution. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- .../finding-your-unknowns-integration/PLAN.md | 257 +++++++++++++++++- .../design/design-resolution.md | 19 ++ 2 files changed, 275 insertions(+), 1 deletion(-) create mode 100644 docs/topics/finding-your-unknowns-integration/design/design-resolution.md diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index 75d801abc..39a4af629 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -119,4 +119,259 @@ sign-off explicitly deferred, each with its trigger and arbiter.) ## Plan -(unfilled — /planning:plan territory) +Written by /planning:plan on 2026-09-01, executing the signed contract (./signoff-sheet.md +rev 2 + this Brief). The delta roster, wave mapping, and definition of done are locked +decisions; these phases sequence their execution, they do not relitigate them. + +Standards grounding: no standards index exists in this repo (`.claude/` carries rules, not +an index); the plan is grounded directly in AGENTS.md (affected-tests contract), +docs/PLUGIN-PHILOSOPHY.md (convention registry :591-631, instruction economy :681-718, +two-lane posture :245-286), docs/MIGRATION-PLAYBOOK.md (per-plugin version bump + CHANGELOG +delivery), .claude/rules/vendor-docs-are-not-style.md, and +plugins/knowledge/reference/citation-shape.md — all read this session. + +### G-block placements (durable copy; source: round-3 audit merge, confirmed C1) + +All are doc-tier placements consumed by Phases 1-2. G1 map/territory: cite-only in F1, +never house vocabulary. G2 over/under-specify diagnostic: F1 doc line, skill edit deferred. +G3 disclose-starting-point primer: F1 pattern-catalog entry. G4 cost framing: F1 intro +rationale line. G5 long-horizon failure diagnostic: F1 doc line, debugging edit deferred. +G6 interactivity scope note: sequencing is chat-portable, artifact optional. G7 +fresh-session-per-phase: corroboration, F1 cite only. G8 stay-in-the-loop criterion: quoted +in F1 caution section. G9 density rubric: F1 when-HTML section. G10 ~100-line ceiling: F1, +labeled practitioner anecdote. G11 sharing argument: one F1 line (satisfied by the Artifact +tool). G12 throwaway-editor doctrine: F1 caution section beside the codification warning. +G13 HTML-diff noise: F1 scoping rule (HTML for ephemeral/published outputs, never +version-controlled instruction surfaces). G14 design-system/PR-explainer/report-HTML +options: F1 pattern entries, skill edits deferred. G15 workflow shape: covered by the +five-pass workflow section. G16 named-expert anecdotes: dropped from all graduated +artifacts. + +### Phase 1: F1 reference doc — docs/FINDING-YOUR-UNKNOWNS.md [TODO] + +Create the graduated reference doc following the top-level precedent shape (H1 → +`## Contents` anchor TOC → charter paragraph → H2 sections). Content contract (signed +Part E + G-block): + +- Header: provenance + the C6 permission basis (short attributed verbatim excerpts from + the named author's public posts under fair-quotation practice, no license claimed, bulk + reproduction avoided) + the quoting-posture line (quotes stay byte-verbatim; S6/M1) + + the citation shape stated locally (URL + ISO retrieval date + sha256 over snapshot + bytes) rather than path-citing the plugin-private shape doc (public-surface rule). +- Unknowns taxonomy + lifecycle: the quadrants, D1's four finding types, G4 cost-framing + rationale, G2/G5 diagnostic doc lines, G1 map/territory as citation only. +- Five-pass sequencing workflow section (Q7 home) + G6 interactivity note + G7 cite + + G15 coverage. +- Prompt-pattern catalog (Q9 home): G3 primer entry, G14 pattern entries, one canonical + invocation line per owning skill. +- Reply-affordance convention + export-button rule: the owner sections for both registry + rows, each naming its conformance surface verbatim ("conformance = the template blocks + in prototype:explore-directions (D21/D22) and prototype:pressure-test (D24)"). +- When-HTML taxonomy: G9 density rubric, G10 ceiling (labeled practitioner anecdote), + G11 sharing line, G13 scoping rule, the HTML-scoping rule. +- Buy-in pattern (S4) + E5 objection-evidence checklist, citing Rust RFC / Oxide RFD / + Amazon PR-FAQ + corpus. +- Caution section: the author's anti-premature-codification warning quoted verbatim with + citation stamp (Brief constraint 2), G8 stay-in-the-loop criterion, G12 + throwaway-editor doctrine. +- Behavioral heuristics as doc lines, each labeled as an eval candidate awaiting + observed-stumble evidence: D5 observed-fact evidence bar, D7/D11 disconnected-work + heuristic/recipe, D17b non-obvious-behavior keying, D34-residue collapse self-check. +- Q11/D31 reconciliation line (presentation ordering is already planning:plan's + documented default). +- If the doc cites x.com URLs, extend lychee.toml excludes for them (authorized by C6). + +**Sanity Check:** `test -f docs/FINDING-YOUR-UNKNOWNS.md`; `grep -c '^## ' docs/FINDING-YOUR-UNKNOWNS.md` ≥ 7; +`grep -q 'retrieved 2026' docs/FINDING-YOUR-UNKNOWNS.md` (citation stamps present); +`grep -qi 'fair-quotation\|fair quotation' docs/FINDING-YOUR-UNKNOWNS.md`; +`scripts/affected-tests.sh --run` green (hygiene lane: markdownlint, typos, lychee). + +### Phase 2: Governance placements — registry, glossary, ctx-eng sequencing [TODO] + +- docs/PLUGIN-PHILOSOPHY.md convention registry: add exactly 2 names-and-points rows + (reply-affordance → the F1 doc's owner section; export-button → same). No restating in + the registry; conformance surfaces live in the owner sections (Phase 1). +- docs/GLOSSARY.md via the /domain-driven-design:curate-language disposition procedure: + add the unknowns-quadrant vocabulary + D1's four finding types with "Avoid:" lines and + a dated Provenance entry; map/territory deliberately NOT added (G1 cite-only). +- docs/topics/context-engineering-claude-5/PLAN.md: add "Open, new" bullets under + `## Open questions` recording F2's three candidates (skill-quality genericness check; + audit-instructions I29 widening; claude-memory /doctor cross-ref against its existing + /doctor contract section) and one line noting Phase 10's sweep re-inventories current + state (recorded counts stale) and rebases over this topic's landed waves. NO phase + heading/tag renames in that file (D37 / Part G rule 4). + +**Sanity Check:** `grep -c 'FINDING-YOUR-UNKNOWNS' docs/PLUGIN-PHILOSOPHY.md` = 2; +`grep -q 'unknown' docs/GLOSSARY.md`; `grep -c 'Open, new' docs/topics/context-engineering-claude-5/PLAN.md` +increased by ≥3; `git diff docs/topics/context-engineering-claude-5/PLAN.md | grep -c '^[-].*### Phase'` = 0 +(no phase headings touched). + +### Phase 3: discovery plugin — blindspot contract deltas [TODO] + +plugins/discovery/skills/blindspot/SKILL.md: D1 typed finding taxonomy in the output +contract; D4 scan-scope disclosure line (POLICY). Extend evals/evals.json expectations for +both new contract lines in this same commit. plugins/discovery/.claude-plugin/plugin.json +version bump + CHANGELOG.md entry. + +**Sanity Check:** grep for the taxonomy + disclosure lines in blindspot SKILL.md; +`jq '.evals[].expectations | length' plugins/discovery/skills/blindspot/evals/evals.json` +shows growth vs HEAD; `git diff HEAD --name-only` includes discovery plugin.json + CHANGELOG; +`scripts/affected-tests.sh --run` green. + +### Phase 4: education plugin — explain + quiz-me contract deltas [TODO] + +explain SKILL.md: D12 vocabulary ladder; D14 success condition scoped to original-ask +invocations. quiz-me SKILL.md: D16 source-anchor + on-miss routing; D17a diff-sourced +question authoring; F4 fresh-context answer-key requirement. Same-commit evals extensions +for all five; education plugin version bump + CHANGELOG. + +**Sanity Check:** grep each of the five contract lines in its SKILL.md; evals.json +expectation growth in both skills; affected-tests green. + +### Phase 5: verification plugin — confirm deltas [TODO] + +confirm SKILL.md: D19 "existing behavior this leans on" callout (CONTRACT) + D18 +quiz-layer cross-ref doc line (the reroute executed; quiz-me untouched by D18). Same-commit +evals extension for D19; verification plugin version bump + CHANGELOG. + +**Sanity Check:** grep both lines in confirm SKILL.md; evals growth; affected-tests green. + +### Phase 6: prototype plugin — explore-directions + pressure-test deltas [TODO] + +explore-directions SKILL.md: D20 same-data control-variable rule (POLICY); D21 structured +steal/graft capture; D22 machine-legible reply template (shaped as the skill's OWN output +contract, never a consumer-repo format — lane-2 constraint). pressure-test SKILL.md: D24 +validation-answer-set shape; D25 fake-data disclosure footnote (POLICY); D26 per-option +named costs. D27 mock-before-wire composition note as a doc line in the pair. Same-commit +evals extensions; prototype plugin version bump + CHANGELOG. + +**Sanity Check:** grep the six contract/policy lines + D27 note; evals growth in both +skills; affected-tests green. + +### Phase 7: planning plugin — interview/plan/design/brainstorm deltas [TODO] + +interview SKILL.md: D28 free-text scrutiny flag as a RESOLUTION-FIELD convention (the +5-field register schema is untouched; gate-invisible by design, limitation recorded in the +skill text). plan SKILL.md: D32 switch condition on alternatives; D33 closing revision +replies; D35 lands-green forward-reference doc line. design SKILL.md: D36 the same +tweak-likelihood presentation ordering plan already documents. brainstorm SKILL.md: D9 +brainstorm-practice citation doc line. Same-commit evals extensions for the four contract +rows (D28, D32, D33, D36); planning plugin version bump + CHANGELOG. + +**Sanity Check:** grep each delta line; `bash plugins/planning/tests/`'s +check-open-questions suite still green via `scripts/affected-tests.sh --run` (D28 must not +break the register gate); evals growth for interview/plan/design. + +### Phase 8: discipline plugin — point-dont-copy W2 deltas [TODO] + +point-dont-copy SKILL.md: E7 primitive-to-convention trap item (CONTRACT); E8 canonical +invocation doc line. Same-commit evals extension for E7; discipline plugin version bump + +CHANGELOG. + +**Sanity Check:** grep E7/E8 lines; evals growth; affected-tests green. + +### Phase 9: Wave 3 — E6 port gate + E1/E2 deviation-log convention [TODO] + +Own review moment; lands after all W2 phases. + +- E6 (discipline:point-dont-copy): semantics map + stop-and-wait confirmation gate, + scoped to external-reference ports with the C3 boundary sentence verbatim ("source of + truth outside this repo's tree: vendored, foreign-language, other-repo"); in-tree + corrections stay do-it-now. Extends the discipline CHANGELOG entry from Phase 8 (one + version bump per plugin per PR). Same-commit evals extension. +- E1+E2 (implementation plugin: implement + implement-dispatch): opt-in deviation-log + convention + fold-back step; the F1 owner section (Phase 1) is the doc home; the + recorded-trigger note lands where the convention is stated ("the moment a second plugin + reads DEVIATIONS.md, the registry rule fires" — no registry row now, per C5/M3). + implementation plugin version bump + CHANGELOG; same-commit evals extension. + +**Sanity Check:** grep the boundary sentence verbatim in point-dont-copy SKILL.md; grep +the trigger note in the implementation plugin; both plugins' CHANGELOGs updated; +affected-tests green. + +### Phase 10: Close-out — issues, cheat-sheet, acceptance verification, PR [TODO] + +- File follow-up GitHub issues (authorized): one for the behavioral eval candidates + (D5, D7, D11, D17b, D34-residue), one for the E4 skill extension candidate, one for the + D28 schema-change deferral, one for Q11-FLIP + E1-REG recorded triggers (or fold small + ones into a single tracking issue — executor's judgment on granularity). +- `node scripts/generate-cheatsheet.mjs --check`; regenerate once if any frontmatter + changed (descriptions are NOT edited by any delta, so expected clean). +- Acceptance-criterion 1 verification by grep: every D/E/F/G row id from the sheet + resolves to a diff hunk or a recorded disposition line; write the sweep result into this + PLAN under a dated note. +- Full `scripts/affected-tests.sh --run` green on the branch head. +- Push and open the ONE pull request (explicitly requested), body mapping commits to + waves and citing ./signoff-sheet.md. + +**Sanity Check:** `node scripts/generate-cheatsheet.mjs --check` exit 0; +`scripts/affected-tests.sh --run` exit 0; issue URLs recorded in this PLAN; PR URL +recorded; the criterion-1 grep sweep output pasted as a dated note with zero unaccounted +rows. + +## Blast radius + +MEDIUM-HIGH. Six plugins' skill contracts change in one PR plus two governance docs and a +mid-flight sibling topic's PLAN. Mitigations: every delta is additive prose (no schema, +no executable surface); the parsed schemas in reach are explicitly frozen (Part G rule 4); +per-phase gates run before each commit; the sibling-topic edit is bullets-only. + +## Stress-test summary + +The execution shape this plan sequences was already adversarially tested this session +before sign-off: /planning:devils-advocate (1 CRITICAL / 4 HIGH / 6 MEDIUM / 4 LOW, all +folded), then two independent fresh-context Fable validators over the sign-off sheet +(amendments folded as rev 2). Step 3's fresh-context plan-reviewer ran over THIS phase +plan; confirmed findings fixed before execution (see dated note below if any). + +## Execution shape + +Fully sequential, all main-session, phases 1→10 in order (2-8 are order-independent among +themselves but run sequentially anyway; commits are serial on one branch). Basis for +main-session routing: the repo's on-demand convention surfaces (AGENTS.md table) do not +auto-load inside subagents; every phase is judgment-heavy house-style contract writing; +token budget is ample. Sub-agents are used only for fresh-context REVIEW (Step 3 reviewer; +any fresh-eyes checkpoint the executor requests), never for authoring. Sequential fallback +is the shape itself — no parallel orchestration to fall back from. + +| Phase | Surface | Basis | +|---|---|---| +| 1-10 | main session | convention surfaces + house style live in main context; serial commits on one branch | +| Step-3 reviewer | fresh sub-agent | mandatory fresh-context stress-test | + +## Open questions + +None blocking — every decision is signed (./signoff-sheet.md). Deferred items live in the +Brief's Deferred questions with triggers and arbiters. + +## Handoff to implementation + +### User-approval gates + +None remaining: the operator signed the full decision surface (".confirm all") and +explicitly directed all execution into this session, one branch +(claude/reading-feedback-j4sg96), one PR. Any mid-flight pivot that would change an +acceptance criterion still stops and asks. + +### Execution shape ([EXEC-SHAPE] tagged) + +Decisions made (gate-passed): + +| Decision | What it changes in the plan | Basis (evidence) | +|---|---|---| +| [EXEC-SHAPE] F1 doc path = `docs/FINDING-YOUR-UNKNOWNS.md` | Phase 1 target file | SCREAMING-KEBAB at docs/ root is the read precedent for durable cross-cutting references (docs/ inventory read this session); "name at plan time" was delegated by Part E | +| [EXEC-SHAPE] Per-plugin DOC rows (D9, D18, D27, D35, E8) ride their plugin's W2 commit instead of a separate W1 commit | Phases 3-8 contents | MIGRATION-PLAYBOOK: version bump + CHANGELOG is per plugin; folding avoids two bumps per plugin in one PR; wave intent (review structure) is preserved by commit ordering | +| [EXEC-SHAPE] Sequential main-session execution, no parallel fan-out | Execution shape | Convention surfaces don't auto-load in subagents (AGENTS.md); serial commits on one branch; judgment-heavy prose work | +| [EXEC-SHAPE] Registry rows point at the F1 doc as owner (form-2: repo file) | Phase 2 | Registry precedent allows repo-file owners; Part E assigns ownership of both conventions to the F1 doc | +| [EXEC-SHAPE] Deferred/eval-candidate tracking as GitHub issues, granularity at executor's judgment | Phase 10 | Operator: "Its OK if we file issues" | + +### Mechanical work + +- One commit per phase (Phases 1-2 may merge into one docs commit if small; never split a + plugin's bump across commits). Commit messages via `git commit -F - --cleanup=verbatim` + heredoc with the session's attribution footer. +- Per-commit verification: the phase's Sanity Check plus `scripts/affected-tests.sh --run`. +- Push with `git push -u origin claude/reading-feedback-j4sg96` (retry with backoff on + network failure only). +- PLAN.md phase tags advance `[TODO]` → `[DOING]` → `[DONE]` in the same commit as the + phase's changes. diff --git a/docs/topics/finding-your-unknowns-integration/design/design-resolution.md b/docs/topics/finding-your-unknowns-integration/design/design-resolution.md new file mode 100644 index 000000000..853cbe91f --- /dev/null +++ b/docs/topics/finding-your-unknowns-integration/design/design-resolution.md @@ -0,0 +1,19 @@ +# Design resolution — finding-your-unknowns-integration + +outcome: early-exit + +Tier C under /planning:plan's design-significance gate. This effort introduces no new +types, modules, package topology, or data models: every change is markdown contract text +in existing skill bodies, evals.json expectation strings, per-plugin CHANGELOG/version +metadata, and new reference documentation under docs/. The one structural artifact (the +F1 reference doc) follows an existing precedent shape (top-level SCREAMING-KEBAB doc with +a Contents TOC, per docs/PLUGIN-PHILOSOPHY.md), so no design exploration is warranted. + +Type sketch: none needed — no executable surface changes. The only parsed-schema surfaces +in reach (the `### Phase N` heading/tag vocabulary; check-open-questions.sh's 5-field +register rows) are explicitly frozen by the signed contract (signoff-sheet Part G rule 4 / +D37): the plan adds prose around them and never renames them. + +Design-tier decisions were resolved upstream by the signed decision chain: interview +rounds 1-3, evidence pass, dual validators, devils-advocate, final validators, operator +sign-off (../signoff-sheet.md, 2026-09-01). From eaee41b3a9bcaa138b8707a421984e83c24ee17c Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:23:02 +0000 Subject: [PATCH 04/22] docs(topics): fold the fresh-context plan review into the execution plan All 15 reviewer findings verified against the repo and folded: Q7/Q10 cross-refs gain phase homes (wayfind in Phase 7, a new session-flow Phase 9, artifact-design as a prose mention in F1), criterion-1 traceability flips to the corpus-to-sheet direction off a committed disposition ledger, the phase gates switch from affected-tests (which selects nothing for these paths) to check-changed-skills, direct markdownlint, the em-dash ratchet, and the direct register test, the criterion-4 reconciliation and criteria-2-7 walk are recorded, and the sanity-check commands become self-verifying before/after pairs. Waves renumber to 11 phases. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- .../finding-your-unknowns-integration/PLAN.md | 197 ++++++++++++------ 1 file changed, 138 insertions(+), 59 deletions(-) diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index 39a4af629..87a37be90 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -165,8 +165,12 @@ Part E + G-block): - Prompt-pattern catalog (Q9 home): G3 primer entry, G14 pattern entries, one canonical invocation line per owning skill. - Reply-affordance convention + export-button rule: the owner sections for both registry - rows, each naming its conformance surface verbatim ("conformance = the template blocks - in prototype:explore-directions (D21/D22) and prototype:pressure-test (D24)"). + rows, each following the owner-doc anatomy (the rule, who is bound, conformance) and + naming its conformance surface verbatim ("conformance = the template blocks in + prototype:explore-directions (D21/D22) and prototype:pressure-test (D24)"). The + reply-affordance section carries the one-line artifact-design cross-ref (Q10) as a prose + mention of the session built-in skill — the established idiom, since artifact-design has + no in-repo surface (verified: only prose mentions exist repo-wide). - When-HTML taxonomy: G9 density rubric, G10 ceiling (labeled practitioner anecdote), G11 sharing line, G13 scoping rule, the HTML-scoping rule. - Buy-in pattern (S4) + E5 objection-evidence checklist, citing Rust RFC / Oxide RFD / @@ -179,32 +183,46 @@ Part E + G-block): heuristic/recipe, D17b non-obvious-behavior keying, D34-residue collapse self-check. - Q11/D31 reconciliation line (presentation ordering is already planning:plan's documented default). -- If the doc cites x.com URLs, extend lychee.toml excludes for them (authorized by C6). +- The doc cites x.com URLs, so extend lychee.toml excludes (authorized by C6; verified + absent today — only `twitter.com` is excluded). -**Sanity Check:** `test -f docs/FINDING-YOUR-UNKNOWNS.md`; `grep -c '^## ' docs/FINDING-YOUR-UNKNOWNS.md` ≥ 7; +**Sanity Check:** `test -f docs/FINDING-YOUR-UNKNOWNS.md`; `test "$(grep -c '^## ' docs/FINDING-YOUR-UNKNOWNS.md)" -ge 7`; `grep -q 'retrieved 2026' docs/FINDING-YOUR-UNKNOWNS.md` (citation stamps present); -`grep -qi 'fair-quotation\|fair quotation' docs/FINDING-YOUR-UNKNOWNS.md`; -`scripts/affected-tests.sh --run` green (hygiene lane: markdownlint, typos, lychee). +`grep -qi 'fair.quotation' docs/FINDING-YOUR-UNKNOWNS.md`; `grep -q 'x\.com' lychee.toml`; +`npx markdownlint-cli2 docs/FINDING-YOUR-UNKNOWNS.md` exit 0 (affected-tests classes +docs/*.md as no-suite — the hygiene tools must be invoked directly). -### Phase 2: Governance placements — registry, glossary, ctx-eng sequencing [TODO] +### Phase 2: Governance placements — registry, glossary, ctx-eng sequencing, traceability ledger [TODO] - docs/PLUGIN-PHILOSOPHY.md convention registry: add exactly 2 names-and-points rows - (reply-affordance → the F1 doc's owner section; export-button → same). No restating in - the registry; conformance surfaces live in the owner sections (Phase 1). + (reply-affordance → the F1 doc's owner section; export-button → same). The registry + table is two-column and never restates, so acceptance criterion 4 ("rows carry + conformance surfaces") is discharged by the owner sections the rows point at — the + reconciliation is recorded here and verified in Phase 11. - docs/GLOSSARY.md via the /domain-driven-design:curate-language disposition procedure: add the unknowns-quadrant vocabulary + D1's four finding types with "Avoid:" lines and - a dated Provenance entry; map/territory deliberately NOT added (G1 cite-only). + a dated Provenance entry; map/territory lands as a `## Rejected terms` row pointing at + the F1 cite-only disposition (G1) so the next curate-language run sees the decision. - docs/topics/context-engineering-claude-5/PLAN.md: add "Open, new" bullets under `## Open questions` recording F2's three candidates (skill-quality genericness check; audit-instructions I29 widening; claude-memory /doctor cross-ref against its existing /doctor contract section) and one line noting Phase 10's sweep re-inventories current state (recorded counts stale) and rebases over this topic's landed waves. NO phase heading/tag renames in that file (D37 / Part G rule 4). - -**Sanity Check:** `grep -c 'FINDING-YOUR-UNKNOWNS' docs/PLUGIN-PHILOSOPHY.md` = 2; -`grep -q 'unknown' docs/GLOSSARY.md`; `grep -c 'Open, new' docs/topics/context-engineering-claude-5/PLAN.md` -increased by ≥3; `git diff docs/topics/context-engineering-claude-5/PLAN.md | grep -c '^[-].*### Phase'` = 0 -(no phase headings touched). +- Traceability ledger (acceptance criterion 1 runs corpus→sheet, not the reverse): commit + `docs/topics/finding-your-unknowns-integration/delta-resolution.md` (the full delta + wording record) and a `disposition-ledger.md` crosswalk keyed by the corpus-inventory + V-ids (V1.1-V7.n), each row naming its sheet/G row id and disposition, so criterion 1 + is verifiable from committed files and every corpus decision is accounted for in the + correct direction. + +**Sanity Check:** `test "$(grep -c 'FINDING-YOUR-UNKNOWNS' docs/PLUGIN-PHILOSOPHY.md)" -ge 2`; +`grep -q 'unknown' docs/GLOSSARY.md`; `grep -qi 'map.*territory' docs/GLOSSARY.md` (rejected-terms +row); `test "$(grep -c 'Open, new' docs/topics/context-engineering-claude-5/PLAN.md)" -ge 4`; +`! git diff HEAD~1 -- docs/topics/context-engineering-claude-5/PLAN.md | grep -q '^-.*### Phase'`; +`test -f docs/topics/finding-your-unknowns-integration/disposition-ledger.md` and every V-id in it +carries a non-empty disposition (`! grep -E '^\- V[0-9]' disposition-ledger.md | grep -q '\|\s*$'`); +`npx markdownlint-cli2` over the touched docs exit 0. ### Phase 3: discovery plugin — blindspot contract deltas [TODO] @@ -214,9 +232,11 @@ both new contract lines in this same commit. plugins/discovery/.claude-plugin/pl version bump + CHANGELOG.md entry. **Sanity Check:** grep for the taxonomy + disclosure lines in blindspot SKILL.md; -`jq '.evals[].expectations | length' plugins/discovery/skills/blindspot/evals/evals.json` -shows growth vs HEAD; `git diff HEAD --name-only` includes discovery plugin.json + CHANGELOG; -`scripts/affected-tests.sh --run` green. +`F=plugins/discovery/skills/blindspot/evals/evals.json; test "$(jq '[.evals[].expectations]|flatten|length' $F)" -gt "$(git show HEAD:$F | jq '[.evals[].expectations]|flatten|length')"`; +`git diff HEAD --name-only` includes discovery plugin.json + CHANGELOG; +`bash scripts/check-changed-skills.sh origin/main` green (affected-tests classes SKILL.md +as no-suite; this is the gate that runs trigger-keyword preservation, listing cap, +--require-evals per Part G rule 2). ### Phase 4: education plugin — explain + quiz-me contract deltas [TODO] @@ -226,15 +246,20 @@ question authoring; F4 fresh-context answer-key requirement. Same-commit evals e for all five; education plugin version bump + CHANGELOG. **Sanity Check:** grep each of the five contract lines in its SKILL.md; evals.json -expectation growth in both skills; affected-tests green. +expectation growth in both skills (before/after jq pair as in Phase 3); +`bash scripts/check-changed-skills.sh origin/main` green. ### Phase 5: verification plugin — confirm deltas [TODO] confirm SKILL.md: D19 "existing behavior this leans on" callout (CONTRACT) + D18 -quiz-layer cross-ref doc line (the reroute executed; quiz-me untouched by D18). Same-commit -evals extension for D19; verification plugin version bump + CHANGELOG. +quiz-layer cross-ref doc line (the reroute executed; quiz-me untouched by D18). D19 likely +also needs a row in the report template in context/outcome.md (read it first). Same-commit +evals extension for D19; verification plugin version bump + CHANGELOG. CONSTRAINT: the +em-dash ratchet (scripts/em-dash-purged-paths.txt) covers verification SKILL.md files — +no em dashes in any new line here. -**Sanity Check:** grep both lines in confirm SKILL.md; evals growth; affected-tests green. +**Sanity Check:** grep both lines in confirm SKILL.md; evals growth (before/after jq pair); +`bash scripts/check-changed-skills.sh origin/main` green; `bash scripts/check-purged-em-dashes.sh` green. ### Phase 6: prototype plugin — explore-directions + pressure-test deltas [TODO] @@ -242,11 +267,14 @@ explore-directions SKILL.md: D20 same-data control-variable rule (POLICY); D21 s steal/graft capture; D22 machine-legible reply template (shaped as the skill's OWN output contract, never a consumer-repo format — lane-2 constraint). pressure-test SKILL.md: D24 validation-answer-set shape; D25 fake-data disclosure footnote (POLICY); D26 per-option -named costs. D27 mock-before-wire composition note as a doc line in the pair. Same-commit -evals extensions; prototype plugin version bump + CHANGELOG. +named costs. D27 mock-before-wire composition note in the shared +plugins/prototype/context/discipline.md Composition table. Same-commit evals extensions; +prototype plugin version bump + CHANGELOG. CONSTRAINT: the em-dash ratchet covers +prototype SKILL.md files — no em dashes in any new SKILL.md line here. **Sanity Check:** grep the six contract/policy lines + D27 note; evals growth in both -skills; affected-tests green. +skills (before/after jq pairs); `bash scripts/check-changed-skills.sh origin/main` green; +`bash scripts/check-purged-em-dashes.sh` green. ### Phase 7: planning plugin — interview/plan/design/brainstorm deltas [TODO] @@ -255,59 +283,96 @@ interview SKILL.md: D28 free-text scrutiny flag as a RESOLUTION-FIELD convention skill text). plan SKILL.md: D32 switch condition on alternatives; D33 closing revision replies; D35 lands-green forward-reference doc line. design SKILL.md: D36 the same tweak-likelihood presentation ordering plan already documents. brainstorm SKILL.md: D9 -brainstorm-practice citation doc line. Same-commit evals extensions for the four contract -rows (D28, D32, D33, D36); planning plugin version bump + CHANGELOG. +brainstorm-practice citation doc line. wayfind SKILL.md: the five-pass workflow-section +cross-ref doc line (Brief constraint 7 / Q7 — points at the F1 workflow section as +marketplace-repo prose). Same-commit evals extensions for the four contract rows (D28, +D32, D33, D36); planning plugin version bump + CHANGELOG. -**Sanity Check:** grep each delta line; `bash plugins/planning/tests/`'s -check-open-questions suite still green via `scripts/affected-tests.sh --run` (D28 must not -break the register gate); evals growth for interview/plan/design. +**Sanity Check:** grep each delta line (incl. the wayfind cross-ref); +`bash plugins/planning/scripts/check-open-questions.test.sh` green (invoked directly — +affected-tests does not select it for SKILL.md/context edits; D28 must not break the +register gate); evals growth for interview/plan/design (before/after jq pairs); +`bash scripts/check-changed-skills.sh origin/main` green. ### Phase 8: discipline plugin — point-dont-copy W2 deltas [TODO] -point-dont-copy SKILL.md: E7 primitive-to-convention trap item (CONTRACT); E8 canonical -invocation doc line. Same-commit evals extension for E7; discipline plugin version bump + -CHANGELOG. +point-dont-copy SKILL.md: E7 primitive-to-convention trap item (CONTRACT) in the +"Audit. What to look for" list; E8 canonical invocation line (as `argument-hint` +frontmatter, not a description rewrite — trigger keywords stay intact; frontmatter change +means the cheat-sheet regenerates in Phase 11). Same-commit evals extension for E7; +discipline plugin version bump + CHANGELOG. The CHANGELOG entry is PROVISIONAL: Phase 10 +extends this same version's entry with the E6 line (one bump per plugin per PR; the +intermediate commit's entry knowingly under-describes and Phase 10 finalizes it). + +**Sanity Check:** grep E7 line + `argument-hint` in frontmatter; evals growth +(before/after jq pair); `bash scripts/check-changed-skills.sh origin/main` green. + +### Phase 9: session-flow plugin — workflow cross-ref [TODO] + +plugins/session-flow/skills/workflow/SKILL.md: the five-pass workflow-section cross-ref +doc line (Brief constraint 7 / Q7 — marketplace-repo prose pointing at the F1 workflow +section). session-flow plugin version bump + CHANGELOG (a doc-line-only bump; no eval +extension — no contract change). -**Sanity Check:** grep E7/E8 lines; evals growth; affected-tests green. +**Sanity Check:** grep the cross-ref line in workflow SKILL.md; session-flow plugin.json + +CHANGELOG in the diff; `bash scripts/check-changed-skills.sh origin/main` green. -### Phase 9: Wave 3 — E6 port gate + E1/E2 deviation-log convention [TODO] +### Phase 10: Wave 3 — E6 port gate + E1/E2 deviation-log convention [TODO] Own review moment; lands after all W2 phases. - E6 (discipline:point-dont-copy): semantics map + stop-and-wait confirmation gate, scoped to external-reference ports with the C3 boundary sentence verbatim ("source of truth outside this repo's tree: vendored, foreign-language, other-repo"); in-tree - corrections stay do-it-now. Extends the discipline CHANGELOG entry from Phase 8 (one - version bump per plugin per PR). Same-commit evals extension. + corrections stay do-it-now. Read plugins/discipline/context/re-anchor-audit-correct.md + FIRST and confirm the gate does not reverse its correct-forward doctrine (the gate lands + in point-dont-copy's own file, never the shared doc). Finalizes the discipline CHANGELOG + entry opened in Phase 8 (same version — one bump per plugin per PR). Same-commit evals + extension. - E1+E2 (implementation plugin: implement + implement-dispatch): opt-in deviation-log - convention + fold-back step; the F1 owner section (Phase 1) is the doc home; the - recorded-trigger note lands where the convention is stated ("the moment a second plugin - reads DEVIATIONS.md, the registry rule fires" — no registry row now, per C5/M3). - implementation plugin version bump + CHANGELOG; same-commit evals extension. + convention + fold-back step (implement Step 5 gains the read-DEVIATIONS.md item; the + schema/taxonomy extension lands in implement-dispatch's "Divergence in non-interactive + runs", which implement's interactive path then cites); the F1 owner section (Phase 1) + is the doc home; the recorded-trigger note lands where the convention is stated ("the + moment a second plugin reads DEVIATIONS.md, the registry rule fires" — no registry row + now, per C5/M3). implementation plugin version bump + CHANGELOG; same-commit evals + extension. CONSTRAINT: the em-dash ratchet covers implementation SKILL.md files — no + em dashes in any new line there. **Sanity Check:** grep the boundary sentence verbatim in point-dont-copy SKILL.md; grep the trigger note in the implementation plugin; both plugins' CHANGELOGs updated; -affected-tests green. +`bash scripts/check-changed-skills.sh origin/main` green; +`bash scripts/check-purged-em-dashes.sh` green. -### Phase 10: Close-out — issues, cheat-sheet, acceptance verification, PR [TODO] +### Phase 11: Close-out — issues, cheat-sheet, acceptance verification, PR [TODO] - File follow-up GitHub issues (authorized): one for the behavioral eval candidates (D5, D7, D11, D17b, D34-residue), one for the E4 skill extension candidate, one for the D28 schema-change deferral, one for Q11-FLIP + E1-REG recorded triggers (or fold small ones into a single tracking issue — executor's judgment on granularity). -- `node scripts/generate-cheatsheet.mjs --check`; regenerate once if any frontmatter - changed (descriptions are NOT edited by any delta, so expected clean). -- Acceptance-criterion 1 verification by grep: every D/E/F/G row id from the sheet - resolves to a diff hunk or a recorded disposition line; write the sweep result into this - PLAN under a dated note. -- Full `scripts/affected-tests.sh --run` green on the branch head. +- `node scripts/generate-cheatsheet.mjs --check`; regenerate once (E8's argument-hint is + a frontmatter change, so regeneration is expected). +- Acceptance-criterion 1 verification in the CORPUS→SHEET direction: iterate every V-id + row of the committed disposition-ledger.md (Phase 2) and confirm each names a + disposition + sheet row; then confirm every sheet CONTRACT/POLICY/CONVENTION/DOC row + resolves to a diff hunk. Write both sweep results into this PLAN under a dated note. +- Walk acceptance criteria 2-7 explicitly, one recorded line each (2: all phase gates + green on head; 3: F1 warning + C6 grep; 4: registry rows + owner-section conformance + text, per the Phase 2 reconciliation; 5: ctx-eng rows grep; 6: eval-expectation commits + co-located with contract commits via `git log --name-only`; 7: this PLAN + sheet were + the executed contract). +- /ai-slop:audit pass over the new F1 doc + a sample of edited SKILL.md hunks (Brief + constraint 5); fix findings before the PR. +- Full gate battery on the branch head: `bash scripts/check-changed-skills.sh origin/main`; + `bash scripts/check-purged-em-dashes.sh`; `npx markdownlint-cli2` on touched .md; + `scripts/affected-tests.sh --run` (covers any script/test surfaces touched). - Push and open the ONE pull request (explicitly requested), body mapping commits to waves and citing ./signoff-sheet.md. -**Sanity Check:** `node scripts/generate-cheatsheet.mjs --check` exit 0; -`scripts/affected-tests.sh --run` exit 0; issue URLs recorded in this PLAN; PR URL -recorded; the criterion-1 grep sweep output pasted as a dated note with zero unaccounted -rows. +**Sanity Check:** `node scripts/generate-cheatsheet.mjs --check` exit 0 after regeneration; +all five gate commands above exit 0; issue URLs recorded in this PLAN; PR URL recorded; +the criterion-1 double sweep pasted as a dated note with zero unaccounted rows; criteria +2-7 walk recorded. ## Blast radius @@ -321,12 +386,24 @@ per-phase gates run before each commit; the sibling-topic edit is bullets-only. The execution shape this plan sequences was already adversarially tested this session before sign-off: /planning:devils-advocate (1 CRITICAL / 4 HIGH / 6 MEDIUM / 4 LOW, all folded), then two independent fresh-context Fable validators over the sign-off sheet -(amendments folded as rev 2). Step 3's fresh-context plan-reviewer ran over THIS phase -plan; confirmed findings fixed before execution (see dated note below if any). +(amendments folded as rev 2). + +2026-09-01 — Step 3 fresh-context plan review over THIS phase plan returned 3 CRITICAL / +6 IMPORTANT / 6 SUGGESTION; all verified against the repo and folded: (1) Q7/Q10 +cross-refs gained homes (wayfind in Phase 7, session-flow as new Phase 9, artifact-design +resolved as prose mention in F1); (2) criterion-1 traceability now runs corpus→sheet off +a committed disposition ledger (Phase 2); (3) affected-tests was the wrong gate for +SKILL.md/docs edits — replaced with check-changed-skills.sh, direct markdownlint, the +em-dash ratchet, and the direct check-open-questions test; (4-6) Part G gates wired into +every phase, criterion-4 reconciliation recorded, criteria 2-7 walk added to Phase 11; +(7) F1-as-owner justification extended; (8) map/territory gets a Rejected-terms row; +(9) delta-resolution.md committed for compaction survivability; plus the six suggestions +(before/after eval counts, grep idioms, real test path, provisional-CHANGELOG note, +ai-slop audit pass, unconditional lychee check). ## Execution shape -Fully sequential, all main-session, phases 1→10 in order (2-8 are order-independent among +Fully sequential, all main-session, phases 1→11 in order (3-9 are order-independent among themselves but run sequentially anyway; commits are serial on one branch). Basis for main-session routing: the repo's on-demand convention surfaces (AGENTS.md table) do not auto-load inside subagents; every phase is judgment-heavy house-style contract writing; @@ -336,7 +413,7 @@ is the shape itself — no parallel orchestration to fall back from. | Phase | Surface | Basis | |---|---|---| -| 1-10 | main session | convention surfaces + house style live in main context; serial commits on one branch | +| 1-11 | main session | convention surfaces + house style live in main context; serial commits on one branch | | Step-3 reviewer | fresh sub-agent | mandatory fresh-context stress-test | ## Open questions @@ -362,8 +439,10 @@ Decisions made (gate-passed): | [EXEC-SHAPE] F1 doc path = `docs/FINDING-YOUR-UNKNOWNS.md` | Phase 1 target file | SCREAMING-KEBAB at docs/ root is the read precedent for durable cross-cutting references (docs/ inventory read this session); "name at plan time" was delegated by Part E | | [EXEC-SHAPE] Per-plugin DOC rows (D9, D18, D27, D35, E8) ride their plugin's W2 commit instead of a separate W1 commit | Phases 3-8 contents | MIGRATION-PLAYBOOK: version bump + CHANGELOG is per plugin; folding avoids two bumps per plugin in one PR; wave intent (review structure) is preserved by commit ordering | | [EXEC-SHAPE] Sequential main-session execution, no parallel fan-out | Execution shape | Convention surfaces don't auto-load in subagents (AGENTS.md); serial commits on one branch; judgment-heavy prose work | -| [EXEC-SHAPE] Registry rows point at the F1 doc as owner (form-2: repo file) | Phase 2 | Registry precedent allows repo-file owners; Part E assigns ownership of both conventions to the F1 doc | -| [EXEC-SHAPE] Deferred/eval-candidate tracking as GitHub issues, granularity at executor's judgment | Phase 10 | Operator: "Its OK if we file issues" | +| [EXEC-SHAPE] Registry rows point at the F1 doc as owner (form-2: repo file) | Phase 2 | Registry precedent allows repo-file owners (`lib/hook-utils.sh`, plugin surfaces); Part E assigns ownership of both conventions to the F1 doc; the owner sections carry the full owner-doc anatomy (rule, who is bound, conformance) and versioning rides git history exactly as it does for the existing non-directory owners | +| [EXEC-SHAPE] Deferred/eval-candidate tracking as GitHub issues, granularity at executor's judgment | Phase 11 | Operator: "Its OK if we file issues" | +| [EXEC-SHAPE] Q10's artifact-design cross-ref = a prose mention of the session built-in skill inside F1's reply-affordance section | Phase 1 | artifact-design has no in-repo surface (verified repo-wide); prose mention of built-ins is the established idiom | +| [EXEC-SHAPE] Q7's session-flow:workflow cross-ref is its own phase/commit with a doc-line-only version bump | Phase 9 | The cross-ref is a Part B obligation, not a Part D roster row; MIGRATION-PLAYBOOK requires the bump for any plugin content change | ### Mechanical work From 93461ef4e10e8e62a78668d446af898ac10fe72a Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:27:38 +0000 Subject: [PATCH 05/22] docs: add the Finding Your Unknowns methodology reference (Phase 1) The graduated reference doc the signed integration contract calls F1: unknowns taxonomy and lifecycle, the five-pass pre-implementation workflow mapped to owning skills, the prompt-pattern catalog, the reply-affordance and export-button owner sections (with conformance surfaces), the opt-in deviation-log convention with its recorded registry trigger, the when-HTML taxonomy and scoping rule, the buy-in pattern with industry grounding, the author's anti-premature-codification warning quoted byte-faithfully with citation stamps, and the behavioral heuristics recorded as eval candidates rather than standing instructions. Fair-quotation permission basis and local citation shape stated in the doc header. lychee already excludes x.com, so no config change rode along. Phase 1 of docs/topics/finding-your-unknowns-integration/PLAN.md; sanity checks green (markdownlint 0 issues, typos clean, section and stamp greps). Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- docs/FINDING-YOUR-UNKNOWNS.md | 288 ++++++++++++++++++ .../finding-your-unknowns-integration/PLAN.md | 9 +- 2 files changed, 293 insertions(+), 4 deletions(-) create mode 100644 docs/FINDING-YOUR-UNKNOWNS.md diff --git a/docs/FINDING-YOUR-UNKNOWNS.md b/docs/FINDING-YOUR-UNKNOWNS.md new file mode 100644 index 000000000..846e94ba7 --- /dev/null +++ b/docs/FINDING-YOUR-UNKNOWNS.md @@ -0,0 +1,288 @@ +# Finding your unknowns + +Graduated reference for the "Finding Your Unknowns" methodology: an artifact-first way of +working where, before and during an implementation, the agent produces small purpose-built +artifacts (explainers, brainstorms, interviews, mockups, plans) whose job is to surface +what you don't yet know while it is still cheap to find out. This doc owns the house +conventions the methodology graduated into this marketplace — the reply-affordance +convention, the export-button rule, and the opt-in deviation-log convention — plus the +pattern catalog and the boundaries (when HTML, when not; what deliberately stays +un-codified). Sibling docs: `PLUGIN-PHILOSOPHY.md` (governance), +`GLOSSARY.md` (vocabulary), `MIGRATION-PLAYBOOK.md` (delivery). + +**Sources and permission basis.** The material derives from public posts by their named +author (see [Sources](#sources-and-citation-shape)). This doc quotes short attributed +verbatim excerpts under fair-quotation practice; no license is claimed and bulk +reproduction is avoided. Quotes are reproduced exactly as published — punctuation +included — and are never edited to fit this repo's style rules. + +## Contents + +- [Why this exists](#why-this-exists) +- [The unknowns taxonomy](#the-unknowns-taxonomy) +- [The five-pass pre-implementation workflow](#the-five-pass-pre-implementation-workflow) +- [Prompt-pattern catalog](#prompt-pattern-catalog) +- [Reply-affordance convention](#reply-affordance-convention) +- [Export-button rule](#export-button-rule) +- [Deviation-log convention (opt-in)](#deviation-log-convention-opt-in) +- [When HTML, and when not](#when-html-and-when-not) +- [The buy-in pattern](#the-buy-in-pattern) +- [Cautions from the source author](#cautions-from-the-source-author) +- [Heuristics awaiting evidence](#heuristics-awaiting-evidence) +- [Sources and citation shape](#sources-and-citation-shape) + +## Why this exists + +The methodology's economic argument, in the author's words: "Every explainer, brainstorm, +interview, prototype, and reference is a cheap way to find out what you didn't know before +it gets expensive to fix." (Field guide, [Sources](#sources-and-citation-shape) S1.) Each +pass below trades a few minutes of artifact review for a class of rework. + +Caution on the framing: the author's stronger thesis — that output quality is now +bottlenecked by the human's ability to clarify the model's unknowns — is a single +practitioner's vendor-published claim and is treated here as direction, not doctrine. + +## The unknowns taxonomy + +Four quadrants, asked as "what are your unknowns?" before prompting: + +- **Known knowns** — what the prompt already states. +- **Known unknowns** — questions you know to ask but haven't answered yet. +- **Unknown knowns** — things you assume without realizing you're assuming them; the + agent can't see them until you disclose them. +- **Unknown unknowns** — the pothole you didn't know the road could have; only an + artifact that shows you the terrain surfaces these. + +The draft article's quadrant taglines ("questions you know to ask", "the pothole you +didn't know the road could have") appear only in the X draft (S4), which is the citable +source for draft-only content. + +Findings that surface during an unknowns pass fall into four types (adopted as +`discovery:blindspot`'s output taxonomy): **Landmine** (a change that will break +something non-obvious), **History** (a constraint that exists for a reason the code no +longer shows), **Convention** (an unwritten team rule the work must follow), and +**Missing concept** (a domain idea the prompt never named). + +Two diagnostics ride the taxonomy: + +- Over-specifying and under-specifying are the same failure seen from two sides: both + mean the split between what you locked and what you left open didn't match your actual + unknowns. +- When a long-horizon task comes back wrong, check the unknowns and the plan's + adaptability before blaming the model: the usual root cause is an unknown that was + never surfaced, not a capability gap. + +The lifecycle is a loop: what an artifact teaches you becomes the starting map for the +next round. The author frames this as matching the map to the territory (S1, "Matching +map and territory") — cited here as his metaphor, not adopted as house vocabulary (see +`GLOSSARY.md` rejected terms). + +## The five-pass pre-implementation workflow + +The corpus composes its pre-implementation demos into one ordered flow. This repo ships a +skill per pass; the composition itself is judgment, not a gate — run the passes whose +unknowns you actually have, in this order when you run several: + +1. **Blindspot pass** — `/discovery:blindspot`: surface unknown unknowns in the task's + blast radius. +2. **Brainstorm / prototype** — `/planning:brainstorm` for direction candidates; + `/prototype:explore-directions` or `/prototype:pressure-test` when the unknown is + visual or interactive. +3. **Interview** — `/planning:interview`: convert known unknowns into decisions on the + record. +4. **Reference port** — `/discipline:point-dont-copy` when the work leans on an external + reference whose semantics must survive the port. +5. **Plan** — `/planning:plan`: lock the approach with the unknowns now known. + +Notes: the sequencing is chat-portable — every pass works as plain conversation, the +artifact form is optional. Running later passes in a fresh session with the earlier +artifacts carried forward matches this repo's existing session-flow doctrine (the corpus +independently corroborates it; see `session-flow` plugin). + +## Prompt-pattern catalog + +Patterns the corpus demonstrated that have no owning skill; each entry is one canonical +prompt-line to adapt. Patterns with an owning skill are listed in the +[workflow](#the-five-pass-pre-implementation-workflow) above — invoke the skill instead. + +- **Disclose your starting point** (primer for any pass): "Before we start: my starting + point is X, my current thinking is Y, my experience level with this area is Z." +- **Teach me my unknowns** (explainer with a vocabulary ladder): served by + `/education:explain`; ask it to end with the terms you should now be using. +- **Design-system HTML file**: "Generate a single HTML page from this codebase's real + tokens and components — one section per component family — so future design + conversations can cite it as the reference." +- **PR explainer page**: "Make a single-file HTML explainer of this PR for reviewers: + annotated diff hunks, a module map of what talks to what, and the three questions a + reviewer should ask." +- **Report/audit HTML view**: for recurring documents (status, incident timeline), + ask the producing skill for an HTML rendering as an opt-in output, never the default. +- **Quiz me before I merge**: served by `/education:quiz-me`; the merge gate itself stays + with `/verification:confirm` (one mechanism per concern). + +Reconciliation note: the corpus's "tweakable plan" ordering (high-tweak decisions first, +mechanical work collapsed) is already `planning:plan`'s documented presentation default; +it needed no new mode here. + +## Reply-affordance convention + +**The rule.** A generated review artifact ends with a structured reply affordance: a +machine-legible way for the human's reaction to become the next prompt — steal/skip +choices, a chip-filled reply template, a decisions table, a confirmation token. Default +with judgment: apply it to artifacts that exist to collect a decision; skip it for purely +informational output. In session contexts that render artifacts (the `artifact-design` +built-in skill's territory), the affordance rides the artifact; in plain chat it is a +reply template in the closing message. + +**Who is bound.** Skills that generate decision-collecting artifacts cite this section +instead of restating it. + +**Conformance** = the template blocks in `prototype:explore-directions` (structured +steal/graft capture and the assembled-reply template) and `prototype:pressure-test` (the +validation answer set). Fleet audits check those surfaces against this section. + +## Export-button rule + +**The rule.** An interactive HTML artifact always ends with an export affordance that +turns UI state back into something the user can paste or commit. In the author's words: +"The trick is always to end with an export: a "copy as JSON" or "copy as prompt" button +that turns whatever I did in the UI back into something I can paste into Claude Code." +(S2, "Custom editing interfaces".) The doctrine recurs three times independently in the +corpus; it is what keeps a throwaway editor inside the agent loop instead of becoming a +dead end. + +**Who is bound.** Skills that emit interactive HTML artifacts cite this section. + +**Conformance** = the same template blocks named in the +[reply-affordance convention](#reply-affordance-convention); the export button is the +HTML-artifact form of the reply affordance. + +## Deviation-log convention (opt-in) + +**The rule (opt-in).** An implementation session MAY keep an append-only `DEVIATIONS.md` +beside `PLAN.md` recording, per entry: what the plan said, what was found, what was chosen, +and whether a human needs to revisit. Entry types: plan-confirmed / discovery / deviation / +human-decision. Default conservative: when in doubt, log. The convention's contract text +is owned by `implementation:implement-dispatch` ("Divergence in non-interactive runs"); +this section records the house posture: opt-in for interactive sessions, required only +where a skill's own contract says so. + +**Recorded trigger.** The moment a second plugin reads `DEVIATIONS.md` (rather than +writing its own), the convention-registry rule fires and this section graduates to a +registry row per `PLUGIN-PHILOSOPHY.md` "Convention registry". + +## When HTML, and when not + +The corpus's examples index (S3) organizes twenty demos into nine categories — +exploration and planning, code review and understanding, design, prototyping, +illustrations and diagrams, decks, research and learning, reports, custom editing +interfaces — which double as the "when is HTML worth it" taxonomy: reach for a rendered +page when the information is spatial (diffs, call graphs), comparative (side-by-side +directions), interactive (motion you can only feel), or recurring (reports that benefit +from structure and color). + +- **Density rubric**: HTML earns its cost through tables, CSS, SVG, interaction, and + spatial layout. Markdown pushed past its density limit produces the degraded + workarounds (ASCII diagrams, unicode color) that signal you wanted a page. +- **Reading ceiling**: the author's ~100-line markdown ceiling is a practitioner + anecdote, recorded as such — not a measured threshold. +- **Sharing**: the publish-and-share argument is satisfied in this environment by the + Artifact tool; nothing extra to build. +- **Scoping rule**: HTML artifacts are for ephemeral and published outputs. They never + replace version-controlled instruction surfaces — HTML diffs are noisy (the author's + own admission) and generation costs 2-4x the markdown equivalent, so plans, skills, + and docs stay markdown in git. + +## The buy-in pattern + +For work that needs stakeholder agreement, the corpus's buy-in document has five +sections: demo first; the pitch; pre-answered objections; spec at a glance; risk and +rollback with named per-person asks and a deadline. The pre-answered-objections element +is the industry-standard core: Amazon's PR/FAQ carries an internal FAQ anticipating hard +leadership questions (Bezos 2017 shareholder letter; Bryar & Carr's Working Backwards), +and every surveyed RFC process — Rust RFCs, Oxide RFDs, Google design docs, Uber-style +RFCs — requires drawbacks/alternatives-considered sections. In all of those orgs the +persuasion artifact and the decision record are one document with a lifecycle, which is +why this repo extends existing planning artifacts rather than minting a parallel one. + +**Objection-evidence checklist** (reusable in PR descriptions): for each objection you +expect, write the question, the factual answer, and the evidence citation — before +anyone asks. An objection you can't answer factually is an unknown; route it back +through the [workflow](#the-five-pass-pre-implementation-workflow). + +## Cautions from the source author + +The corpus carries its own warning against exactly the move a plugin marketplace is +tempted to make, and this repo treats it as binding (it is why the deltas that landed are +judgment-preserving contract lines and doc entries, never generator skills): + +> I’m a little bit afraid that people will read this article and turn it into a /html +> skill or something. While there might be some value in that, I want to emphasize that +> you don’t need to do much to get Claude to do this. You can just ask it to “make a HTML +> file” or “make a HTML artifact”. +> +> The trick is knowing what you want the artifact to do and how you might use it. You may +> over time make a skill, but for now I’d suggest just prompting from scratch to get a +> hang of how to use it in different cases. (S2, "How to Get Started".) + +Two companions to the warning: + +- **Stay in the loop** is the evaluation lens for any artifact tooling: "All of the above + is to say that I think the real reason I use HTML is that I feel much more in the loop + with Claude." (S2, "Stay in the Loop".) Tooling that produces artifacts the user never + forms judgment about fails this criterion even when it satisfies density, sharing, and + ease. +- **Throwaway-editor doctrine**: a custom editing interface is "not a product, or a + reusable tool" — it is built for the exact thing being worked on and discarded. The + marketplace instinct to generalize a good throwaway into a shipped generator is the + failure mode the warning names. + +## Heuristics awaiting evidence + +The following corpus heuristics are recorded here as doc lines and candidate eval cases, +not as standing skill instructions — per `PLUGIN-PHILOSOPHY.md` "Instruction economy", +they graduate into a skill body only on observed, repeated stumble evidence: + +- **Observed-fact evidence bar** (brainstorming): each candidate option cites an observed, + falsifiable fact about the codebase (a path plus a claim that could be wrong), not just + a plausible path. +- **Already-built-but-disconnected scan**: before proposing new work, scan for dead + imports, dark feature flags, and unread tables — the improvement may already exist, + disconnected. +- **Non-obvious-behavior keying** (quizzes): author questions against behaviors a reader + would skim past, not against what the diff makes obvious. +- **Collapse self-check** (plans): before collapsing a section as "mechanical, trust me", + re-check that nothing in it is actually a judgment call — the corpus's failure case is + a design decision hidden in a collapsed section. + +## Sources and citation shape + +Citations in this doc use: URL, ISO retrieval date, and `sha256:` over the raw +snapshot bytes captured at retrieval. Content drift produces a new citation, never an +in-place hash edit. + +- **S1** — "A field guide to Claude Fable 5: Finding your unknowns", Thariq Shihipar, + Anthropic blog, published 2026-07-06. + `https://claude.com/blog/a-field-guide-to-claude-fable-finding-your-unknowns` + (retrieved 2026-09-01, + `sha256:ac8229699555d38eb0dfe6c80dd2e85353f30471a7abff0894d342b5107aad26`) +- **S2** — "Using Claude Code: The Unreasonable Effectiveness of HTML", X article by the + same author. `https://x.com/trq212/status/2052809885763747935` (retrieved 2026-09-01, + `sha256:07dc71b1a7fabe264b9a80ee003edbcd1e74013895372a8ffe13ee4bb178e63c`) +- **S3** — HTML-effectiveness examples index (20 demos, 9 categories, plus the 11-demo + "Know your unknowns" sub-collection). `https://thariqs.github.io/html-effectiveness` + (retrieved 2026-08-31, + `sha256:7e6da98b6b447ec39efdc6deb34602204e4641dc59f4e311e3f05fb23d74f98e`) +- **S4** — X draft of the field guide (citable only for draft-only content: the quadrant + taglines, the lifecycle-loop image, and three links the published blog dropped). + `https://x.com/trq212/status/2073100352921215386` (retrieved 2026-09-01; snapshot + pinned in the corpus work slice) +- Buy-in grounding: Bezos 2017 shareholder letter + (`https://www.aboutamazon.com/news/company-news/2017-letter-to-shareholders`), the + Working Backwards PR/FAQ + (`https://workingbackwards.com/concepts/working-backwards-pr-faq-process/`), Rust RFCs + (`https://raw.githubusercontent.com/rust-lang/rfcs/master/README.md`), Oxide RFD 1 + (`https://rfd.shared.oxide.computer/rfd/0001`), Google design docs + (`https://www.industrialempathy.com/posts/design-docs-at-google/`), and Uber-style + RFCs (`https://blog.pragmaticengineer.com/scaling-engineering-teams-via-writing-things-down-rfcs/`), + all retrieved 2026-09-01. diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index 87a37be90..7b3f44c34 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -147,7 +147,7 @@ options: F1 pattern entries, skill edits deferred. G15 workflow shape: covered b five-pass workflow section. G16 named-expert anecdotes: dropped from all graduated artifacts. -### Phase 1: F1 reference doc — docs/FINDING-YOUR-UNKNOWNS.md [TODO] +### Phase 1: F1 reference doc — docs/FINDING-YOUR-UNKNOWNS.md [DONE] Create the graduated reference doc following the top-level precedent shape (H1 → `## Contents` anchor TOC → charter paragraph → H2 sections). Content contract (signed @@ -183,12 +183,13 @@ Part E + G-block): heuristic/recipe, D17b non-obvious-behavior keying, D34-residue collapse self-check. - Q11/D31 reconciliation line (presentation ordering is already planning:plan's documented default). -- The doc cites x.com URLs, so extend lychee.toml excludes (authorized by C6; verified - absent today — only `twitter.com` is excluded). +- The doc cites x.com URLs; lychee.toml already excludes `^https?://(www\.)?x\.com/` + (verified at execution — the review's "no x.com entry" finding was itself wrong), so no + config change was needed. **Sanity Check:** `test -f docs/FINDING-YOUR-UNKNOWNS.md`; `test "$(grep -c '^## ' docs/FINDING-YOUR-UNKNOWNS.md)" -ge 7`; `grep -q 'retrieved 2026' docs/FINDING-YOUR-UNKNOWNS.md` (citation stamps present); -`grep -qi 'fair.quotation' docs/FINDING-YOUR-UNKNOWNS.md`; `grep -q 'x\.com' lychee.toml`; +`grep -qi 'fair.quotation' docs/FINDING-YOUR-UNKNOWNS.md`; `grep -qF 'x\.com' lychee.toml`; `npx markdownlint-cli2 docs/FINDING-YOUR-UNKNOWNS.md` exit 0 (affected-tests classes docs/*.md as no-suite — the hygiene tools must be invoked directly). From 727d22e95e97ac06be5f5a58be8f1c9d5d8189dc Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:30:40 +0000 Subject: [PATCH 06/22] docs: land the unknowns governance placements (Phase 2) Two names-and-points convention-registry rows (reply affordance, export button) pointing at their owner sections in FINDING-YOUR-UNKNOWNS.md; glossary entries for the unknowns quadrants and blindspot finding types plus a map/territory rejected-terms row, curated per the curate-language entry discipline with a dated provenance note; three "Open, new" candidate bullets and a Phase-10 rebase note recorded in the context-engineering topic PLAN (no phase headings touched); and the acceptance-criterion-1 traceability spine committed into the topic dir - the 48-row V-id disposition ledger plus the delta wording record. Phase 2 of the integration plan; sanity checks green (markdownlint, typos, grep battery, no-phase-touch guard). Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- docs/GLOSSARY.md | 23 ++++- docs/PLUGIN-PHILOSOPHY.md | 2 + .../context-engineering-claude-5/PLAN.md | 13 +++ .../finding-your-unknowns-integration/PLAN.md | 2 +- .../delta-resolution.md | 88 +++++++++++++++++ .../disposition-ledger.md | 97 +++++++++++++++++++ 6 files changed, 223 insertions(+), 2 deletions(-) create mode 100644 docs/topics/finding-your-unknowns-integration/delta-resolution.md create mode 100644 docs/topics/finding-your-unknowns-integration/disposition-ledger.md diff --git a/docs/GLOSSARY.md b/docs/GLOSSARY.md index 3d7be3f41..37c962375 100644 --- a/docs/GLOSSARY.md +++ b/docs/GLOSSARY.md @@ -72,6 +72,21 @@ The healthiest of `context-guard`'s three context zones (`smart` / `acceptable` the band rather than any token figure — the band numbers are declared judgment defaults and tunable per consumer. +**unknowns quadrants** + +The four-way pre-prompt breakdown — known knowns, known unknowns, unknown knowns, unknown +unknowns — used to decide which unknown-finding pass a task needs. Owned by +[`FINDING-YOUR-UNKNOWNS.md`](FINDING-YOUR-UNKNOWNS.md); entries cite it rather than restating the +quadrants. + +**blindspot finding types** + +The typed taxonomy a blindspot pass reports its findings in: Landmine (breaks something +non-obvious), History (a constraint the code no longer shows), Convention (an unwritten team +rule), Missing concept (a domain idea the prompt never named). The output contract lives in +`discovery:blindspot`; the taxonomy's rationale in +[`FINDING-YOUR-UNKNOWNS.md`](FINDING-YOUR-UNKNOWNS.md). + ## Rejected terms Names considered for a concept this project already owns, recorded so they are not reintroduced. @@ -86,11 +101,17 @@ Each maps to the term or doctrine that owns the concept. | cache *(the doc-restating-environment sense)* | `docs-hygiene:audit-derivability`'s derivable-from-environment doctrine; the word is overloaded here (plugin cache, prompt cache) | | sediment | the `docs-hygiene` audit family's pruning doctrine; collides with the code-sense use in `playbooks:fable-5` | | sycophancy | nothing — a generic LLM-behavior term with no distinct project meaning. Free-prose use is unaffected; it is simply not project vocabulary | +| map / territory | the source author's metaphor, cited where it appears in [`FINDING-YOUR-UNKNOWNS.md`](FINDING-YOUR-UNKNOWNS.md) "The unknowns taxonomy"; never house vocabulary (metaphor-jargon risk) | ## Provenance -Every term above was graded and adopted in lane 6 of the AI Hero course vetting +Terms through "smart zone" were graded and adopted in lane 6 of the AI Hero course vetting (2026-08-18). The decision rows, including the basis for each verdict and the rejected-term mappings, are in [`upstream/aihero-course.md`](upstream/aihero-course.md) under "Term adoption". Materialization of this file was tracked as [#3000](https://github.com/melodic-software/claude-code-plugins/issues/3000). + +"unknowns quadrants", "blindspot finding types", and the map/territory rejected-terms row were +adopted at the finding-your-unknowns integration sign-off (2026-09-01); the decision rows are in +[`topics/finding-your-unknowns-integration/signoff-sheet.md`](topics/finding-your-unknowns-integration/signoff-sheet.md) +Part E. diff --git a/docs/PLUGIN-PHILOSOPHY.md b/docs/PLUGIN-PHILOSOPHY.md index 08df19e22..ac69f5475 100644 --- a/docs/PLUGIN-PHILOSOPHY.md +++ b/docs/PLUGIN-PHILOSOPHY.md @@ -629,6 +629,8 @@ doc before a second plugin adopts it. Fleet audits check conformance per row. | Always-on hook cost ceiling | [`docs/conventions/hook-budget/`](conventions/hook-budget/README.md) | | Tracker reference form inside a code comment | [`docs/conventions/tracker-reference-form/`](conventions/tracker-reference-form/README.md) | | Untrusted-content framing contract | [`docs/conventions/untrusted-content/`](conventions/untrusted-content/README.md) | +| Reply affordance on decision-collecting artifacts | [`docs/FINDING-YOUR-UNKNOWNS.md`](FINDING-YOUR-UNKNOWNS.md#reply-affordance-convention) | +| Export button on interactive HTML artifacts | [`docs/FINDING-YOUR-UNKNOWNS.md`](FINDING-YOUR-UNKNOWNS.md#export-button-rule) | ## Cross-platform contract diff --git a/docs/topics/context-engineering-claude-5/PLAN.md b/docs/topics/context-engineering-claude-5/PLAN.md index ac55023a7..ad80f04ad 100644 --- a/docs/topics/context-engineering-claude-5/PLAN.md +++ b/docs/topics/context-engineering-claude-5/PLAN.md @@ -1167,6 +1167,19 @@ them. - **Open, new** — whether `mcp-tools:audit` actually covers tool-search configuration. The gate defers the "deferred tool loading is unowned" remainder out of scope on that basis, which is a negative claim about a body nobody has read. Task #44. +- **Open, new** — a skill-body "genericness" check candidate for `skill-quality:check` (could this + skill body be pasted into any repo unchanged?), recorded as an input by the + finding-your-unknowns integration (2026-09-01, its topic's signoff-sheet, F2). Resolution + belongs to this topic's check-design phases, not that effort. +- **Open, new** — widening `claude-config:audit-instructions` check I29 (restatement families + I29-a/I29-b) with the companion article's repetition-myth duplication lens, recorded as an + input by the same F2 reroute. Owner: the phase that next touches the I-check catalog. +- **Open, new** — a `claude-memory:audit` cross-ref against this topic's `/doctor` prerequisite + contract (design/checks-and-sweep.md), recorded as an input by the same F2 reroute. +- **Note (2026-09-01):** Phase 10's sweep re-inventories current state before running — its + recorded surface counts predate the finding-your-unknowns waves landing on + `claude/reading-feedback-j4sg96` and are stale; the sweep rebases over whatever has landed. + That effort touches none of this topic's audit-instructions criteria files. ## Handoff to implementation diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index 7b3f44c34..cc0e1161a 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -193,7 +193,7 @@ Part E + G-block): `npx markdownlint-cli2 docs/FINDING-YOUR-UNKNOWNS.md` exit 0 (affected-tests classes docs/*.md as no-suite — the hygiene tools must be invoked directly). -### Phase 2: Governance placements — registry, glossary, ctx-eng sequencing, traceability ledger [TODO] +### Phase 2: Governance placements — registry, glossary, ctx-eng sequencing, traceability ledger [DONE] - docs/PLUGIN-PHILOSOPHY.md convention registry: add exactly 2 names-and-points rows (reply-affordance → the F1 doc's owner section; export-button → same). The registry diff --git a/docs/topics/finding-your-unknowns-integration/delta-resolution.md b/docs/topics/finding-your-unknowns-integration/delta-resolution.md new file mode 100644 index 000000000..24710e89c --- /dev/null +++ b/docs/topics/finding-your-unknowns-integration/delta-resolution.md @@ -0,0 +1,88 @@ +# Delta resolution — conditional verdicts resolved against grader evidence + +Committed copy of the session's delta wording record (originally a work-slice artifact); +the row wording here is what the wave phases implement. Row classifications and wave +assignments were finalized in ./signoff-sheet.md Part D, which supersedes the Status +column below where they differ (e.g. D5/D34 reclassed behavioral-tier, D13 dropped). + +2026-09-01. Inputs: six grader reports in evidence/ (file:line evidence there), the +round-1/2 interview locks, corpus-inventory.md. Every row cites its grader. Statuses: +DELTA (change to make), CORROBORATION (already present; optionally cite), NO-CHANGE +(present and stronger than corpus), REROUTE (corpus aimed at wrong home), RETURN +(genuine fork back to human/audit). All DELTA rows are subject to PLUGIN-PHILOSOPHY +"evidence-gated additions": each lands as a sourced, corpus-cited contract line with +the gate acknowledged, or waits for observed-stumble evidence per the audit's call. + +## D-block: V2 technique deltas + +| ID | Target | Change | Status | Grader | +|---|---|---|---|---| +| D1 | discovery:blindspot | typed finding taxonomy (Landmine/History/Convention/Missing-concept) as output contract | DELTA | blindspot-brainstorm A1 | +| D2 | discovery:blindspot | per-finding prompt-fix | CORROBORATION (present SKILL.md:59) | A2 | +| D3 | discovery:blindspot | fold-step: add explicit human confirm-before-final checkpoint | DELTA (partial today) | A3 | +| D4 | discovery:blindspot | output requires scan-scope disclosure line | DELTA (partial) | A4 | +| D5 | planning:brainstorm | tighten per-option evidence to observed-fact (falsifiable), not just path | DELTA (partial) | B5 | +| D6 | planning:brainstorm | cheapest-to-ambitious ordering | CORROBORATION | B6 | +| D7 | planning:brainstorm | named already-built-but-disconnected scan heuristic (dead imports, dark flags, unread tables) | DELTA (partial) | B7 | +| D8 | planning:brainstorm | structured closing pick | CORROBORATION | B8 | +| D9 | planning:brainstorm | session-start-brainstorm citable rationale line (Q8 lock: only-if-absent; absent) | DELTA (one line) | B9 | +| D10 | improvement:find | corpus S/M/L/XL axis | NO-CHANGE (WSJF+size stronger; adopting would regress) | C10 | +| D11 | improvement:find | disconnected-work recipe file (like hotspots.md) | DELTA (optional, small) | C11 | +| D12 | education:explain | vocabulary ladder (term + definition + modeled "say ->" sentence) | DELTA | edu A1 | +| D13 | education:explain | payoff-prompts closing (before/after prompt contrast) | DELTA | A2 | +| D14 | education:explain | success condition: "user's next prompt names what they mean" | DELTA | A4 | +| D15 | education:explain | three-tier restructure | NO-CHANGE (corpus itself marks hypothesis unvalidated; existing altitude tiers stay) | A3 | +| D16 | education:quiz-me | per-question source-anchor + on-miss routing to the skimmed section | DELTA | B6 | +| D17 | education:quiz-me | diff-sourced question authoring keyed to non-obvious behaviors | DELTA (partial+absent merged) | B7+B9 | +| D18 | quiz-as-merge-gate | REROUTE: quiz-me disclaims merge-gating twice by design; verification:confirm owns the PR gate (SKILL.md:119). Any merge-gate framing lands as a confirm cross-ref, never a second gate | REROUTE | edu B8xC11 | +| D19 | verification:confirm | explicit "existing behavior this leans on" out-of-diff coupling callout | DELTA (partial today) | C10 | +| D20 | prototype:explore-directions | same-data control-variable rule for mockup substrate | DELTA (partial) | proto A1 | +| D21 | prototype:explore-directions | structured steal/graft capture at single-decision granularity | DELTA (partial) | A3 | +| D22 | prototype:explore-directions | machine-legible assembled-reply template (direction/steal/skip/next-target) | DELTA | A4 | +| D23 | prototype:explore-directions | direction count | NO-CHANGE (default 3 cap 5 beats corpus's fixed 4) | A5 | +| D24 | prototype:pressure-test | validation-answer-set output shape (bounded fillable forced-choice answers) | DELTA | B6+B7 | +| D25 | prototype:pressure-test | fake-data/no-real-wiring disclosure footnote (non-dev audience risk) | DELTA (high value) | B8 | +| D26 | prototype:pressure-test | per-option named costs on forced-choice questions | DELTA | B9 | +| D27 | prototype pair | "mock before you wire" named ordering note in composition table | DELTA (small) | B11 | +| D28 | planning:interview | flag free-text answers for downstream scrutiny | DELTA (small) | plan-grader 2 | +| D29 | planning:interview | decisions-table / assembler | NO-CHANGE (register+arbiter cover it; Brief prose is by design) | 3+4 | +| D30 | planning:questionnaire | none | NO-CHANGE (intentional non-adopter: "interview the send, not the subject") | 5 | +| D31 | planning:plan | tweak-likelihood ordering | CORROBORATION (knob already at SKILL.md:223; Q11's "mode" is status quo) | 6+7 | +| D32 | planning:plan | alternatives carry a one-line switch condition | DELTA (partial) | 8 | +| D33 | planning:plan | closing pre-drafted revision replies tied to flagged decisions | DELTA | 9 | +| D34 | planning:plan | self-check before collapsing mechanical sections | DELTA (partial) | 10 | +| D35 | planning:plan | forward-reference implement's lands-green guarantee from Sanity Check | DELTA (one line) | 11 | +| D36 | planning:design | mirror plan's tweak-likelihood knob in Phase-5 discussion rounds | DELTA | 13 | +| D37 | PLAN.md schema | constraint: never rename `### Phase N` heading/tag vocabulary without version bump; block reordering is safe | CONSTRAINT (fact) | 12 | + +## E-block: V3/V4 deltas + +| ID | Target | Change | Status | Grader | +|---|---|---|---|---| +| E1 | implement pipeline | extend deviation-log convention (taxonomy: plan-confirmed/discovery/deviation/human-decision; plan-said/found/chosen fields; conservative default; blocking markers) from autonomous+Moderate to the interactive path | DELTA (the V3 headline) | port-impl 5+7 | +| E2 | implement Step 5 | required fold-back: read DEVIATIONS.md, emit plan-amendment bullets | DELTA | 6 | +| E3 | session-flow retro/handoff | live-vs-posthoc | NO-CHANGE (complementary by design; E1/E2 close the gap at the right home) | 8+9 | +| E4 | buy-in doc home | design-handoff has 0/4 persuasion elements; prd ships an HTML pitch view (closest precedent); visualize is static (demo routes via playwright/run) | RETURN (fork: extend prd pitch view vs. extend design-handoff vs. new thin skill slot) | 10+11+12 | +| E5 | buy-in components | standalone objection-evidence checklist (question+answer+evidence citation) | DELTA (home follows E4) | corpus f7e96d2a | +| E6 | discipline:point-dont-copy | externalized semantics-map artifact + confirmation gate for EXTERNAL-REFERENCE PORTS ONLY (in-tree corrector doctrine explicitly forbids stop-and-wait; scoping avoids the reversal) | DELTA scoped + RETURN flag (audit must confirm the scoping is clean) | 1+2 | +| E7 | discipline:point-dont-copy | trap check: source primitive with no target analogue -> name the carrying convention | DELTA | 3 | +| E8 | discipline:point-dont-copy | canonical invocation example line | DELTA (one line) | 4 | + +## F-block: V5/V6 + reference doc + +| ID | Target | Change | Status | Grader | +|---|---|---|---|---| +| F1 | graduated reference doc | new docs/ reference: unknowns taxonomy + lifecycle + pattern catalog + reply-affordance convention + export-button rule + 9-category when-HTML taxonomy + the skill-codification warning quoted; citations per plugins/knowledge/reference/citation-shape.md | DELTA (the Q7/Q9/Q10 vehicle; artifact-design/capabilities are session built-ins, not repo files, so conventions land here) | gov 10-12 | +| F2 | V6 routing | ALL context-engineering deltas route into docs/topics/context-engineering-claude-5/ open phases (article already decomposed, corroborated, gated; audit-instructions I6/I15 already cite it) | REROUTE (supersedes V6 entries 2,3; genericness check and I29-widening and /doctor cross-ref become candidate inputs to that topic, not this one) | gov 1-9 + surprise | +| F3 | claude-memory:audit | gotcha-vs-obvious heuristic | CORROBORATION (C2 Deletion Test + C5) | gov 6 | +| F4 | quiz-me fresh-eyes tension | quiz author self-grades in biased context vs. philosophy's fresh-eyes doctrine | RETURN (design question beyond corpus scope; surfaced to human) | edu surprise 2 | + +## Standing constraints carried into every delta + +1. PLUGIN-PHILOSOPHY evidence-gated additions (:699-711): corpus-anticipated is not + observed-stumble; each delta lands citing the corpus as source AND acknowledging the + gate, with the audit deciding per-delta whether the gate demands deferral. +2. One-mechanism-per-concern (:678-679): D18 reroute is the enforcement example. +3. MIGRATION-PLAYBOOK version pinning: D37; any parsed-schema touch needs changelog. +4. Two-lane convention posture: no hardcoded decision-record formats or literal + path/flag templates in conventions (proto grader note). diff --git a/docs/topics/finding-your-unknowns-integration/disposition-ledger.md b/docs/topics/finding-your-unknowns-integration/disposition-ledger.md new file mode 100644 index 000000000..f872bc261 --- /dev/null +++ b/docs/topics/finding-your-unknowns-integration/disposition-ledger.md @@ -0,0 +1,97 @@ +# Disposition ledger — corpus decision → executed-or-recorded outcome + +The acceptance-criterion-1 crosswalk, in the corpus→sheet direction: one row per decision +in the corpus inventory (V-ids; the inventory itself lives in the session work slice), +each naming its signed-sheet row(s) and disposition. Sheet: ./signoff-sheet.md. Wording +record: ./delta-resolution.md. No row may be blank. + +## V1 — framing and vocabulary + +- V1.1 quadrant taxonomy | adopt doc-tier | F1 taxonomy section + GLOSSARY "unknowns quadrants" +- V1.2 map/territory | cite-only | G1; GLOSSARY rejected-terms row +- V1.3 bottleneck thesis | treat-as-caution | F1 "Why this exists" caution line +- V1.4 over/under-specify diagnostic | doc line | G2 (F1 taxonomy section) +- V1.5 disclose-starting-point primer | adopt doc-tier | G3 (F1 pattern catalog) +- V1.6 cost framing | adopt doc-tier | G4 (F1 intro quote) +- V1.7 long-horizon diagnostic | doc line | G5 (F1 taxonomy section) + +## V2 — pre-implementation techniques + +- V2.1 blindspot block | adopt/corroborate | D1 (W2), D2 (corroboration), D3 (demoted to + corroboration), D4 (W2) +- V2.2 teach-me block | adopt scoped | D12 (W2), D13 (dropped — conflicts with explain's + one-line close), D14 (W2, scoped), D15 (no-change), G6 (F1 note) +- V2.3 brainstorm block | mixed | D5/D7 (behavioral → F1 + eval candidates), D6/D8 + (corroboration), D9 (W1 doc line), D10 (no-change), D11 (behavioral → F1) +- V2.4 design-directions block | adopt | D20-D22 (W2), D23 (no-change — repo default stronger) +- V2.5 mock-before-wire block | adopt | D24-D26 (W2), D27 (W1 doc note) +- V2.6 interview block | adopt small | D28 (W2 resolution-field convention), D29/D30 (no-change) +- V2.7 reference-port block | adopt scoped | E6 (W3), E7 (W2), E8 (W1 doc line) +- V2.8 tweakable-plan block | adopt/corroborate | D31 (corroboration + reconciliation line), + D32/D33 (W2), D34 (corroboration + eval candidate), D35 (W1 doc line), D36 (W2) +- V2.9 sequencing meta | adopt doc-tier | Q7: F1 workflow section + wayfind and + session-flow:workflow cross-refs; G15 + +## V3 — during implementation + +- V3.1 implementation-notes convention | adopt opt-in | E1+E2 (W3), E3 (no-change); + F1 deviation-log owner section + recorded registry trigger (C5) +- V3.2 fresh-session-per-phase | corroboration | G7 (F1 cite of session-flow doctrine) + +## V4 — post-implementation + +- V4.1 buy-in doc | adopt doc-tier | S4/C2: F1 buy-in section + E5 checklist; skill + extension deferred behind demand evidence (Brief deferred E4-EXT) +- V4.2 quiz-as-merge-gate | reroute + adopt | D16/D17a/F4 (W2), D17b (behavioral → F1), + D18 (reroute to verification:confirm cross-ref, W1), D19 (W2) + +## V5 — HTML-artifact methodology + +- V5.1 governing caution | BINDING constraint | Q2; quoted verbatim in F1 cautions +- V5.2 stay-in-the-loop | adopt as lens | G8 (F1 cautions quote) +- V5.3 nine-category taxonomy | adopt doc-tier | F1 when-HTML section +- V5.4 density rubric | doc line | G9 +- V5.5 ~100-line ceiling | recorded anecdote | G10 (labeled practitioner anecdote) +- V5.6 sharing argument | recorded | G11 (satisfied by the Artifact tool; one F1 line) +- V5.7 export-button doctrine | adopt convention | F1 owner section + registry row +- V5.8 throwaway-editor doctrine | adopt as caution | G12 (F1 cautions) +- V5.9 HTML-diff noise | adopt scoping rule | G13 (F1 when-HTML scoping rule) +- V5.10 design-system/PR-explainer/report patterns | doc-tier | G14 (F1 pattern entries; + skill edits deferred behind the evidence gate) +- V5.11 workflow shape | covered | G15 (F1 workflow section) + +## V6 — context-engineering companion (F2 reroute) + +All V6 decisions route to docs/topics/context-engineering-claude-5/ (F2); this effort +touches none of that topic's criteria files. + +- V6.1 80% system-prompt claim | treat-as-caution | recorded; live-doc-checked in the + evidence pass; never a pruning license here +- V6.2 conflicting-instructions example | corroboration | audit-instructions I6/I15 + already cite the article +- V6.3 S1→S2 rewrite pattern | rerouted | that topic's check-design phases own it +- V6.4 examples→interfaces | recorded, scoped | covers tool usage examples, not skill + trigger phrases; no action here +- V6.5 progressive-disclosure myth | corroboration | existing docs-hygiene audit + + AGENTS.md load-on-demand table; cite-only +- V6.6 repetition myth | candidate recorded | ctx-eng PLAN "Open, new" bullet + (I29 widening) +- V6.7 auto-memory myth | rerouted | claude-memory scoping falls under that topic's + /doctor + memory-audit work +- V6.8 rich-references myth | rerouted | that topic's check-design phases +- V6.9 genericness lens | candidate recorded | ctx-eng PLAN "Open, new" bullet + (skill-quality check) +- V6.10 claude doctor positioning | candidate recorded | ctx-eng PLAN "Open, new" bullet + (/doctor cross-ref) + +## V7 — corpus and process meta + +- V7.1 citation routing | executed | Brief captured assumptions; F1 sources S1/S4 split +- V7.2 companion-pair rule | executed | the two handoffs interviewed as one package +- V7.3 cc-applicable tags | executed | Q5: ~6 decision-relevant claims live-doc-checked; + rest treated as vendor anecdote +- V7.4 videos | out of scope | Brief out-of-scope list +- V7.5 named-expert anecdotes | dropped | G16 (absent from all graduated artifacts) +- V7.6 frontend-plugin pointer | recorded | verified absent from this repo; nothing to compare +- V7.7 verifier MINORs + pipeline lessons | recorded | correction records in the work + slice; candidate follow-up issue at close-out (Phase 11) From 89e6b590430269d9bf3a8c5b6ff0e522cbb76a75 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:33:40 +0000 Subject: [PATCH 07/22] feat(discovery): typed blindspot finding cards + scan-scope disclosure blindspot's output contract gains the four-way finding taxonomy (Landmine / History / Convention / Missing concept) on every card and a closing one-line scan-scope disclosure naming which lanes ran and what was and was not scanned. Adopted at the finding-your-unknowns integration sign-off as team-convention-tier contract lines (D1, D4); rationale and provenance in docs/FINDING-YOUR-UNKNOWNS.md and the topic's signoff-sheet. Evals gain expectations for both lines in this commit; plugin 0.16.18 -> 0.17.0. Phase 3 of docs/topics/finding-your-unknowns-integration/PLAN.md. check-changed-skills green (0 errors). Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- .../finding-your-unknowns-integration/PLAN.md | 2 +- plugins/discovery/.claude-plugin/plugin.json | 2 +- plugins/discovery/CHANGELOG.md | 14 ++++++++++++++ plugins/discovery/skills/blindspot/SKILL.md | 14 +++++++++++--- .../discovery/skills/blindspot/evals/evals.json | 2 ++ 5 files changed, 29 insertions(+), 5 deletions(-) diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index cc0e1161a..56d30b597 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -225,7 +225,7 @@ row); `test "$(grep -c 'Open, new' docs/topics/context-engineering-claude-5/PLAN carries a non-empty disposition (`! grep -E '^\- V[0-9]' disposition-ledger.md | grep -q '\|\s*$'`); `npx markdownlint-cli2` over the touched docs exit 0. -### Phase 3: discovery plugin — blindspot contract deltas [TODO] +### Phase 3: discovery plugin — blindspot contract deltas [DONE] plugins/discovery/skills/blindspot/SKILL.md: D1 typed finding taxonomy in the output contract; D4 scan-scope disclosure line (POLICY). Extend evals/evals.json expectations for diff --git a/plugins/discovery/.claude-plugin/plugin.json b/plugins/discovery/.claude-plugin/plugin.json index 63971a19d..ac22f752a 100644 --- a/plugins/discovery/.claude-plugin/plugin.json +++ b/plugins/discovery/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "discovery", - "version": "0.16.18", + "version": "0.17.0", "description": "Structured discovery before changes: explore the local codebase, run disciplined multi-source external research, and reconstruct why a past decision was made from evidence outside the code — each dispatching a purpose-built subagent by default so the reading stays out of the main conversation, with source tiers, falsification, recency gates, an intent-evidence tier, and a corpus-coverage ledger — persisting EXPLORE.md / RESEARCH.md / INTENT.md index-plus-sidecar handoff artifacts.", "author": { "name": "Melodic Software", diff --git a/plugins/discovery/CHANGELOG.md b/plugins/discovery/CHANGELOG.md index 2d4526c33..449a038c3 100644 --- a/plugins/discovery/CHANGELOG.md +++ b/plugins/discovery/CHANGELOG.md @@ -1,5 +1,19 @@ # Changelog — discovery plugin +## [0.17.0] + +### Added + +- **`blindspot`: typed finding cards and a scan-scope disclosure line.** Each blindspot card now + leads with a finding type from a four-way taxonomy — Landmine (breaks something non-obvious), + History (a constraint whose reason the code no longer shows), Convention (an unwritten team + rule), Missing concept (a domain idea the framing never named) — so repeated runs teach the user + which kinds of unknowns they tend to carry. The output also ends with a one-line scan-scope + disclosure naming which lane(s) ran and what was and was not scanned. Adopted from the + "Finding Your Unknowns" corpus at the integration sign-off (team-convention tier, evidence and + provenance in `docs/FINDING-YOUR-UNKNOWNS.md` in the marketplace repository); evals extended to + cover both contract lines. + ## [0.16.18] ### Fixed diff --git a/plugins/discovery/skills/blindspot/SKILL.md b/plugins/discovery/skills/blindspot/SKILL.md index 397abd83f..f3285e273 100644 --- a/plugins/discovery/skills/blindspot/SKILL.md +++ b/plugins/discovery/skills/blindspot/SKILL.md @@ -43,9 +43,9 @@ understanding rather than the agent's. would break. - **Domain lane**. Build a lightweight vocabulary ladder grounded in sources fetched this session (repo files, official docs), never bare training recall. -3. **Output. Blindspot cards.** One card per blindspot: the gap, why it matters here, and a copyable - prompt-fix line. Close by assembling the fixes into ONE improved implementation prompt the user can - run next. +3. **Output. Blindspot cards.** One card per blindspot, typed by the kind of gap it is (Landmine / + History / Convention / Missing concept): the gap, why it matters here, and a copyable prompt-fix + line. Close by assembling the fixes into ONE improved implementation prompt the user can run next. 4. **Escalate when depth warranted**, a domain too deep for a lightweight ladder gets a recommendation to run proper external research (`/discovery:research`) or whatever structured-learning capability the environment provides. @@ -54,6 +54,10 @@ understanding rather than the agent's. Present each blindspot as a card: +- **Type**, one of four: **Landmine** (the change would break something non-obvious), **History** + (a constraint whose reason the code no longer shows), **Convention** (an unwritten team rule the + work must follow), or **Missing concept** (a domain idea the user's framing never named). The + type tells the user which kind of unknown they were carrying, so repeated runs teach a pattern. - **Gap**, the specific thing the user's current framing did not account for. - **Why it matters here**, the concrete consequence in this codebase or domain, not a generic caution. - **Prompt-fix**, a single copyable line the user can drop into their prompt to close the gap. @@ -61,6 +65,10 @@ Present each blindspot as a card: Then assemble every prompt-fix into ONE improved implementation prompt, wrapped in clear copy-start / copy-end markers so the exact text to reuse is unambiguous. +End with one scan-scope disclosure line: which lane(s) ran and what was and was not scanned (areas +read, sources fetched), so the user knows what the cards do and do not cover. One line, not a +methodology dump. + This skill does NOT write `EXPLORE.md`. Its deliverable is the user's understanding plus the improved prompt. When the scan's findings also serve as stage-1 codebase exploration, offer to hand off to `/discovery:explore` to persist the `EXPLORE.md` artifact rather than diff --git a/plugins/discovery/skills/blindspot/evals/evals.json b/plugins/discovery/skills/blindspot/evals/evals.json index bcd50d44c..9f756491c 100644 --- a/plugins/discovery/skills/blindspot/evals/evals.json +++ b/plugins/discovery/skills/blindspot/evals/evals.json @@ -11,7 +11,9 @@ "Output builds the USER's understanding (blindspot cards + an improved prompt), not just the agent's internal findings", "The run asks one intake question about the user's starting point before scanning", "Each blindspot is expressed as a card naming the gap, why it matters here, and a copyable prompt-fix line", + "Each card carries a finding type from the four-way taxonomy: Landmine, History, Convention, or Missing concept", "Output closes by assembling the fixes into a single improved implementation prompt the user can run next", + "The output ends with a one-line scan-scope disclosure naming which lane(s) ran and what was and was not scanned", "The run does not write EXPLORE.md" ] }, From fcb99ef2f56a8dac07f1abb15eb545025aaff21b Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:37:37 +0000 Subject: [PATCH 08/22] feat(education): vocabulary ladders, anchored diff-sourced quizzes, fresh keys explain: rung-2 terms arrive as vocabulary-ladder entries (term, plain definition, a modeled "you can now say" sentence), and original-ask invocations gain a success condition judged by the user's next prompt (bare comprehension asks exempt by scope). quiz-me: questions are diff-sourced, every question anchors to the report section that teaches its answer with on-miss routing to that exact section, and the embedded answer key is fresh-context authored or verified. Adopted at the finding-your-unknowns integration sign-off (D12, D14, D16, D17a, F4); provenance in docs/FINDING-YOUR-UNKNOWNS.md and the topic signoff-sheet. Evals extended in this commit (including a new original-ask case); plugin 0.8.8 -> 0.9.0. Phase 4 of the integration plan. check-changed-skills green (0 errors). Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- .../finding-your-unknowns-integration/PLAN.md | 2 +- plugins/education/.claude-plugin/plugin.json | 2 +- plugins/education/CHANGELOG.md | 18 ++++++++++++++++++ plugins/education/skills/explain/SKILL.md | 10 +++++++++- .../education/skills/explain/evals/evals.json | 15 ++++++++++++++- plugins/education/skills/quiz-me/SKILL.md | 13 +++++++++++++ .../education/skills/quiz-me/evals/evals.json | 5 ++++- 7 files changed, 60 insertions(+), 5 deletions(-) diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index 56d30b597..29f1cb41c 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -239,7 +239,7 @@ version bump + CHANGELOG.md entry. as no-suite; this is the gate that runs trigger-keyword preservation, listing cap, --require-evals per Part G rule 2). -### Phase 4: education plugin — explain + quiz-me contract deltas [TODO] +### Phase 4: education plugin — explain + quiz-me contract deltas [DONE] explain SKILL.md: D12 vocabulary ladder; D14 success condition scoped to original-ask invocations. quiz-me SKILL.md: D16 source-anchor + on-miss routing; D17a diff-sourced diff --git a/plugins/education/.claude-plugin/plugin.json b/plugins/education/.claude-plugin/plugin.json index 4cc3cfd3c..f5683749b 100644 --- a/plugins/education/.claude-plugin/plugin.json +++ b/plugins/education/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "education", - "version": "0.8.8", + "version": "0.9.0", "description": "Interactive multi-session learning coach: teaches a general subject or a concept grounded in the consuming repo through the Knowledge-Skills-Wisdom progression, with persistent per-topic learning state. Also a single-session domain primer, a one-shot plain-language explainer that drops anything to genuinely plain words, and a post-work comprehension check that quizzes the human on a completed change.", "author": { "name": "Melodic Software", diff --git a/plugins/education/CHANGELOG.md b/plugins/education/CHANGELOG.md index e24611d85..40c2ad379 100644 --- a/plugins/education/CHANGELOG.md +++ b/plugins/education/CHANGELOG.md @@ -3,6 +3,24 @@ All notable changes to the `education` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [0.9.0] + +### Added + +- **`explain`: vocabulary-ladder entries and an original-ask success condition.** Rung-2 terms of + art now arrive as ladder entries — the term, an ordinary-words definition, and a modeled "you + can now say" sentence the user can reuse — and when the explanation serves a task the user was + stuck on, success is judged by whether their next prompt names what they mean (bare + comprehension asks are exempt by scope). Adopted from the "Finding Your Unknowns" corpus at the + integration sign-off (team-convention tier; provenance in `docs/FINDING-YOUR-UNKNOWNS.md` in the + marketplace repository). Evals extended, including a new original-ask case. +- **`quiz-me`: diff-sourced, anchored questions and a fresh-context answer key.** Quiz questions + are authored from the change's actual diff and the report's own sections; every question is + anchored to the section that teaches its answer and a miss routes the reader to that exact + section before any retry; the embedded answer key is produced or verified by a fresh-context + pass reading only the report and diff (re-derived from the artifact alone when no fresh surface + is available). Same adoption basis; evals extended for all three contract lines. + ## [0.8.8] ### Changed diff --git a/plugins/education/skills/explain/SKILL.md b/plugins/education/skills/explain/SKILL.md index 53113494e..918c194f5 100644 --- a/plugins/education/skills/explain/SKILL.md +++ b/plugins/education/skills/explain/SKILL.md @@ -62,7 +62,7 @@ actually know X"). Never front-load a higher rung. | Rung | Altitude | Move | |------|----------|------| | 1 (default) | **Plain / ELI5** | Concrete everyday analogy, zero jargon. The floor and the default landing. | -| 2 (on request) | **High-school** | Introduce one or two real terms of art, each defined as it appears. Keep the analogy as scaffolding. | +| 2 (on request) | **High-school** | Introduce one or two real terms of art as vocabulary-ladder entries: the term, its definition in ordinary words, and one modeled "you can now say: …" sentence showing the term doing work in the user's own next prompt. Keep the analogy as scaffolding. | | 3 (on request) | **Peer** | Full precision, jargon allowed, edge cases and tradeoffs, the explanation a colleague in the field would want. | Offer the next rung as a one-line invitation, not a wall of text: "That's the @@ -79,6 +79,14 @@ why, and ground harder (re-read the source, fetch the primary reference) before claiming to explain it. A confident-sounding restatement of jargon is the failure mode this check exists to catch. +## Success condition (original-ask invocations only) + +When the explanation serves a task the user was stuck on, the success test is their **next +prompt**: it names what they mean in the newly plain terms instead of re-gesturing at the +confusion. Judge the explanation by that, and shape rung-2 vocabulary entries so the user can +reuse them. A bare comprehension ask with no task behind it ("I don't get it", full stop) has no +next-prompt contrast; this check does not apply there. + ## Handoff to `education:teach` Close every explanation with a single lightweight line offering the multi-session diff --git a/plugins/education/skills/explain/evals/evals.json b/plugins/education/skills/explain/evals/evals.json index 714dd572d..c754f3290 100644 --- a/plugins/education/skills/explain/evals/evals.json +++ b/plugins/education/skills/explain/evals/evals.json @@ -34,7 +34,8 @@ "expectations": [ "Recognizes the explicit request for higher altitude and climbs the rung ladder rather than staying pinned at rung 1", "Treats altitude as request-driven — it does not default to dumping the peer-level explanation on an unqualified 'explain X'", - "Introduces real terms of art with definitions as altitude increases rather than assuming them silently" + "Introduces real terms of art with definitions as altitude increases rather than assuming them silently", + "Rung-2 terms arrive as vocabulary-ladder entries: the term, an ordinary-words definition, and a modeled 'you can now say' sentence the user can reuse in their next prompt" ] }, { @@ -84,6 +85,18 @@ "Asks 'What would you like explained?' (or equivalent) instead of proceeding blind or inventing a topic to explain", "Does not hallucinate a referent or produce a confident plain-language explanation of something never stated" ] + }, + { + "id": 8, + "name": "original-ask-success-condition", + "prompt": "I'm trying to write a Postgres migration that adds a partial index but it keeps rejecting my syntax, and I don't get what \"predicate\" means in the docs. Explain it so I can fix my migration.", + "expected_output": "Lands a rung-1 plain explanation of a partial-index predicate (concrete analogy, jargon defined inline), aimed at unblocking the migration the user is stuck on. Because an original ask exists, the close shapes the vocabulary so the user's NEXT prompt can name what they mean (using \"predicate\" correctly), treating that next prompt as the success test rather than the explanation itself.", + "files": [], + "expectations": [ + "Recognizes the original ask behind the comprehension request and aims the explanation at unblocking it", + "Treats the user's next prompt as the success test: the close equips them to name what they mean (reusing the term correctly) rather than re-gesturing at the confusion", + "Still lands at rung 1 first with a concrete analogy and inline definitions" + ] } ] } diff --git a/plugins/education/skills/quiz-me/SKILL.md b/plugins/education/skills/quiz-me/SKILL.md index 525d7e741..c229b3936 100644 --- a/plugins/education/skills/quiz-me/SKILL.md +++ b/plugins/education/skills/quiz-me/SKILL.md @@ -76,6 +76,19 @@ pattern ("a quiz at the bottom on the changes that I must pass"). Match each nar section's length to what the change needs: cover the substance, but do not pad with filler, redundant summaries, or boilerplate. +- **Questions are diff-sourced.** Author each quiz question from the change's actual diff + and the report sections that explain it, never from generic topic knowledge a reader + could answer without having followed this change. +- **Each question carries a source anchor, and a miss routes to it.** Anchor every + question to the report section that teaches its answer (a report-internal anchor, or a + durable pointer per the reference discipline below). On a missed question, send the + reader to that exact section — the skimmed material, quoted or linked — before any + retry; the miss's job is routing, not scoring. +- **The answer key is fresh-context authored.** Produce or verify the embedded key with a + fresh-context pass that reads only the report and the diff, not the authoring + conversation, so grading tests what the artifact actually supports rather than what its + author remembers meaning. Where no fresh sub-agent surface is available, re-derive the + key from the artifact alone before embedding it. - **Answer key persists with the artifact.** Embed the key in the report, a collapsed `
` block in HTML, an appendix section in the markdown fallback. Grade in-conversation in the same session; a later or compacted session grades by reading the diff --git a/plugins/education/skills/quiz-me/evals/evals.json b/plugins/education/skills/quiz-me/evals/evals.json index 07e37624b..2d07cb1f9 100644 --- a/plugins/education/skills/quiz-me/evals/evals.json +++ b/plugins/education/skills/quiz-me/evals/evals.json @@ -10,7 +10,10 @@ "expectations": [ "Produces a self-contained HTML report (inline CSS/JS, no remote fetch) with context/intuition/decisions/what-was-done sections and a quiz at the bottom", "Treats the human's comprehension of the completed work as the object under test — not whether the artifact is correct or works", - "Embeds the answer key in the artifact (collapsed details in HTML) rather than leaving grading impossible for a later session" + "Embeds the answer key in the artifact (collapsed details in HTML) rather than leaving grading impossible for a later session", + "Quiz questions are authored from the change's actual diff and the report's own sections, not from generic topic knowledge", + "Each question is anchored to the report section that teaches its answer, and a missed question routes the reader to that exact section before any retry", + "The answer key is produced or verified by a fresh-context pass reading only the report and diff, or re-derived from the artifact alone when no fresh surface is available" ] }, { From 173c6c678e7c7a0f6f68026222c733b6dea0862c Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:40:23 +0000 Subject: [PATCH 09/22] feat(verification): name the out-of-diff couplings in confirm's report Stage 2's intent match now requires naming the existing behavior a change leans on (unchanged code whose contract the diff depends on), and the outcome report template gains a dedicated couplings table with an evidence column. The PR-prep edge case records that a quiz-me comprehension layer may precede the gate while the merge gate stays in confirm, one mechanism per concern. Adopted at the finding-your-unknowns integration sign-off (D19 contract, D18 doc line); provenance in docs/FINDING-YOUR-UNKNOWNS.md. Evals extended; plugin 0.5.10 -> 0.6.0. Phase 5 of the integration plan. check-changed-skills and the em-dash ratchet both green. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- .../finding-your-unknowns-integration/PLAN.md | 2 +- plugins/verification/.claude-plugin/plugin.json | 2 +- plugins/verification/CHANGELOG.md | 13 +++++++++++++ plugins/verification/skills/confirm/SKILL.md | 4 ++-- .../verification/skills/confirm/context/outcome.md | 5 +++++ .../verification/skills/confirm/evals/evals.json | 1 + 6 files changed, 23 insertions(+), 4 deletions(-) diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index 29f1cb41c..780c72ca9 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -250,7 +250,7 @@ for all five; education plugin version bump + CHANGELOG. expectation growth in both skills (before/after jq pair as in Phase 3); `bash scripts/check-changed-skills.sh origin/main` green. -### Phase 5: verification plugin — confirm deltas [TODO] +### Phase 5: verification plugin — confirm deltas [DONE] confirm SKILL.md: D19 "existing behavior this leans on" callout (CONTRACT) + D18 quiz-layer cross-ref doc line (the reroute executed; quiz-me untouched by D18). D19 likely diff --git a/plugins/verification/.claude-plugin/plugin.json b/plugins/verification/.claude-plugin/plugin.json index 6136464a5..2fca01741 100644 --- a/plugins/verification/.claude-plugin/plugin.json +++ b/plugins/verification/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "verification", - "version": "0.5.10", + "version": "0.6.0", "description": "Outcome-verification stage: prove a change achieved its intended outcome (`/verification:confirm` — a mechanical build/test/lint prerequisite gate, then intent-match + evidence + verdict with the criterion auto-detected by change type), and verify measurable-improvement claims against a planning-time baseline (`/verification:measure`), never fabricating numbers.", "author": { "name": "Melodic Software", diff --git a/plugins/verification/CHANGELOG.md b/plugins/verification/CHANGELOG.md index b3860f028..8416da194 100644 --- a/plugins/verification/CHANGELOG.md +++ b/plugins/verification/CHANGELOG.md @@ -3,6 +3,19 @@ All notable changes to the `verification` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [0.6.0] + +### Added + +- **`confirm`: out-of-diff couplings get their own report table.** Stage 2's intent match now + names the existing behavior the change leans on: unchanged code whose contract the diff depends + on, carried in the outcome report as a dedicated table (coupling, where it lives, evidence it + still holds). A named coupling is checkable; an implied one is where regressions hide. The + PR-prep edge case also records that a comprehension layer (`education:quiz-me`, when installed) + may precede the gate while the merge gate itself stays here. Adopted from the "Finding Your + Unknowns" corpus at the integration sign-off (D19, D18; provenance in + `docs/FINDING-YOUR-UNKNOWNS.md` in the marketplace repository). Evals extended. + ## [0.5.10] ### Changed diff --git a/plugins/verification/skills/confirm/SKILL.md b/plugins/verification/skills/confirm/SKILL.md index c83b4e825..730c7cdd2 100644 --- a/plugins/verification/skills/confirm/SKILL.md +++ b/plugins/verification/skills/confirm/SKILL.md @@ -95,7 +95,7 @@ Read the criterion context file for the dispatched mode, then run the flow below 1. **Auto-trigger `/testing:run-e2e` (when runtime-affecting)**. Inspect changed files. If any match an `e2e-*` category from the Runtime-affecting paths above, or touch observability code paths verifiable end-to-end, or the user said "test the app", invoke `/testing:run-e2e` via the Skill tool when the `testing` plugin is installed. Otherwise drive the live app directly (Claude Code's bundled `/run`, or a manual orchestrator launch) and capture the same evidence. When present, it validates prerequisites, starts the app, exercises the changed flow, and captures evidence (screenshots, console, network, traces). Carry that into the evidence table. If not runtime-affecting (pure refactor, internal lib, doc-only): note "E2E not applicable" and skip. 2. **Intent retrieval**. Scan the conversation for the original request, the approved plan, refinements, and acceptance criteria. If none is clear, ask the user what the goal was. 3. **Implementation inventory**. Changed files, new capabilities, behavior changes, config/infra changes. -4. **Intent match**. Every requirement has implementation; every implementation traces to a requirement; flag scope additions and gaps (including implicit requirements. Error handling, edge cases, tests). +4. **Intent match**. Every requirement has implementation; every implementation traces to a requirement; flag scope additions and gaps (including implicit requirements. Error handling, edge cases, tests). Name the out-of-diff couplings: the existing behavior this change leans on, unchanged code whose contract the diff now depends on. A named coupling is checkable; an implied one is where regressions hide. The report carries them as their own table (see [context/outcome.md](context/outcome.md)). 5. **Evidence collection**. Stage-1 results, E2E results, test names + assertions proving the claimed behavior. For UI changes: the UI evidence artifacts per [context/outcome.md](context/outcome.md) (pre/action/post snapshot, console, network, behavior assertion. "screenshot looks fine" is NOT an assertion). When the plan states a measurable goal: the `/verification:measure` comparison table. 6. **Report + verdict**. Emit the outcome report (intent-match table, mechanical results, E2E + UI-evidence tables when triggered, evidence table, measurements when applicable) and a `CONFIRMED` / `NEEDS WORK` verdict. Report template and verdict criteria in [context/outcome.md](context/outcome.md). @@ -116,7 +116,7 @@ For "run the live app and watch it behave," beyond automated `/testing:run-e2e`, - **No git changes but user runs `/verification:confirm all`**: run Stage 1 across all ecosystems anyway (useful after a rebase or pull), then outcome verification if intent is in scope. - **Changed file outside any known ecosystem**: Stage 1 skips it with a note; Stage 2 still assesses intent match. - **Missing tools**: `/toolchain:check` / `/toolchain:lint` report `skip` with install hint, not failure. Except the core toolchain the project's own code requires. -- **Invoked from a PR-prep flow**: treat the verdict as a hard gate. Any FAIL or unresolved CRITICAL gap blocks PR creation. +- **Invoked from a PR-prep flow**: treat the verdict as a hard gate. Any FAIL or unresolved CRITICAL gap blocks PR creation. A comprehension layer (an `education:quiz-me` report, when that plugin is installed) may precede this gate and inform it; the merge gate itself lives here, one mechanism per concern. ## Skill chaining diff --git a/plugins/verification/skills/confirm/context/outcome.md b/plugins/verification/skills/confirm/context/outcome.md index fcd94d8f5..c4a3b7fdd 100644 --- a/plugins/verification/skills/confirm/context/outcome.md +++ b/plugins/verification/skills/confirm/context/outcome.md @@ -53,6 +53,11 @@ Justified additions are fine but should be noted. Unjustified additions should b |---|----------|-----------|-----------| | 1 | | Yes/No | | +### Existing behavior this leans on (out-of-diff couplings) +| # | Coupling | Where it lives | Evidence it still holds | +|---|----------|----------------|-------------------------| +| 1 | | | | + ### Assessment - Plan items: X/Y complete (Z%) - Deviations: N (all justified / N unjustified) diff --git a/plugins/verification/skills/confirm/evals/evals.json b/plugins/verification/skills/confirm/evals/evals.json index 92761bf42..3b68fb636 100644 --- a/plugins/verification/skills/confirm/evals/evals.json +++ b/plugins/verification/skills/confirm/evals/evals.json @@ -11,6 +11,7 @@ "Auto-detects the `outcome` criterion from the post-implementation context (no mode argument given)", "Stage 1 delegates to /toolchain:check and /toolchain:lint cross-cutting — does NOT reimplement build/test/lint or inline exec-bit/gitleaks bash", "Stage 2 retrieves intent (plan / conversation) and produces an intent-match table", + "The report names the out-of-diff couplings as their own table: existing behavior the change leans on, where it lives, and evidence it still holds", "Emits a CONFIRMED or NEEDS WORK verdict, not just a green build" ] }, From 33161611425fa7c70331b3b5f5414a61255c9168 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:44:04 +0000 Subject: [PATCH 10/22] feat(prototype): control-variable data, graft capture, answer sets, disclosure explore-directions: variants bind one identical data set (design is the only variable), the handover closes with a machine-legible direction/steal/skip/next-target reply template, and captures record steal/skip decisions at single-decision granularity so grafts compose. pressure-test: the HTML demo shell gains a validation answer set (forced choices whose options name their costs, free-text escape hatch, copy-out) and a visible fake-data disclosure footer; the capture step carries the filled answers into the durable record. Shared discipline gains the mock-before-you-wire ordering note. Adopted at the finding-your-unknowns integration sign-off (D20-D22, D24-D27); provenance in docs/FINDING-YOUR-UNKNOWNS.md. Evals extended; plugin 0.9.8 -> 0.10.0. Phase 6 of the integration plan. check-changed-skills and the em-dash ratchet both green. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- .../finding-your-unknowns-integration/PLAN.md | 2 +- plugins/prototype/.claude-plugin/plugin.json | 2 +- plugins/prototype/CHANGELOG.md | 21 ++++++++++++++++ plugins/prototype/context/discipline.md | 5 ++++ .../skills/explore-directions/SKILL.md | 25 ++++++++++++++++--- .../explore-directions/evals/evals.json | 5 +++- .../prototype/skills/pressure-test/SKILL.md | 17 ++++++++++--- .../skills/pressure-test/evals/evals.json | 3 ++- 8 files changed, 69 insertions(+), 11 deletions(-) diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index 780c72ca9..e5af0ee65 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -262,7 +262,7 @@ no em dashes in any new line here. **Sanity Check:** grep both lines in confirm SKILL.md; evals growth (before/after jq pair); `bash scripts/check-changed-skills.sh origin/main` green; `bash scripts/check-purged-em-dashes.sh` green. -### Phase 6: prototype plugin — explore-directions + pressure-test deltas [TODO] +### Phase 6: prototype plugin — explore-directions + pressure-test deltas [DONE] explore-directions SKILL.md: D20 same-data control-variable rule (POLICY); D21 structured steal/graft capture; D22 machine-legible reply template (shaped as the skill's OWN output diff --git a/plugins/prototype/.claude-plugin/plugin.json b/plugins/prototype/.claude-plugin/plugin.json index e32435c24..78f31422b 100644 --- a/plugins/prototype/.claude-plugin/plugin.json +++ b/plugins/prototype/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "prototype", - "version": "0.9.8", + "version": "0.10.0", "description": "Builds throwaway code to answer a design question before committing to architecture — a logic facet (an interactive terminal app over a portable state model) and a UI facet (radically different visual variants on one route).", "author": { "name": "Melodic Software", diff --git a/plugins/prototype/CHANGELOG.md b/plugins/prototype/CHANGELOG.md index f5be3658c..04a4e43aa 100644 --- a/plugins/prototype/CHANGELOG.md +++ b/plugins/prototype/CHANGELOG.md @@ -3,6 +3,27 @@ All notable changes to the `prototype` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [0.10.0] + +### Added + +- **`explore-directions`: same-data control variable, single-decision graft capture, and a + machine-legible reply template.** Variants now bind one identical data set (the data is the + control variable, so only the design differs); the handover closes with a fillable + direction/steal/skip/next-target reply template (what the mockup's copy-out terminator lifts); + and the capture step records steal/skip decisions per named piece so grafts compose across + variants. Adopted from the "Finding Your Unknowns" corpus at the integration sign-off (D20, + D21, D22; provenance in `docs/FINDING-YOUR-UNKNOWNS.md` in the marketplace repository). Evals + extended. +- **`pressure-test`: validation answer set and a fake-data disclosure footer on the HTML demo + shell.** The demo page carries the questions it exists to answer as forced choices whose + options each name their cost (plus a free-text escape hatch), with a copy-out control, and a + visible footer stating the page is synthetic end to end and where real wiring lives. The + capture step carries the filled answer set into the durable answer verbatim. Same adoption + basis (D24, D25, D26); evals extended. +- **Shared discipline: "mock before you wire" ordering note.** The composition table now states + that the throwaway mock runs before any real wiring when a change has both questions (D27). + ## [0.9.8] ### Added diff --git a/plugins/prototype/context/discipline.md b/plugins/prototype/context/discipline.md index 8579a124c..8af80cb5c 100644 --- a/plugins/prototype/context/discipline.md +++ b/plugins/prototype/context/discipline.md @@ -71,3 +71,8 @@ written here is gone, and the question gets re-litigated from scratch the next t | Architecture discovery surfaced a design question | `/architecture:improve` (when installed) | Improvement pass surfaces the opportunity → prototype validates the approach | | Prototype answered the question | `/planning:plan` (when installed) | Validated decision feeds the plan | | Logic module worth keeping | `/implementation:implement` (when installed) | Lift the pure module into production; delete the TUI shell | + +Ordering note — **mock before you wire**: when a change has both a "does the interaction work" +question and real integration work, run the throwaway mock (this plugin) before any wiring. A +mock that fails kills the wiring work for free; wiring first turns every design misfire into +rework of live code. diff --git a/plugins/prototype/skills/explore-directions/SKILL.md b/plugins/prototype/skills/explore-directions/SKILL.md index c103ecbb8..5546f5ef0 100644 --- a/plugins/prototype/skills/explore-directions/SKILL.md +++ b/plugins/prototype/skills/explore-directions/SKILL.md @@ -100,6 +100,10 @@ Constraints: - **Synthetic data only.** A throwaway prototype binds synthetic data, never real or captured values. +- **Same data across variants.** All variants bind one identical data set; the data is the + control variable, so the only thing that differs between variants is the design. On the real + stack, sub-shape A's shared fetching above the switcher already enforces this; on the mockup + substrate, define the synthetic set once and have every variant render it. - **No remote fetch by construction.** Vendor everything inline so the page opens straight from `file://`. No external scripts, fonts, or data fetches. Enforce this rather than trusting it: emit a restrictive CSP meta tag in the page `` so the browser blocks any remote resource: @@ -213,14 +217,27 @@ Requirements: ### 5. Hand it over Surface the URL and variant keys. Interesting feedback is usually "I want the header from B with -the sidebar from C". That's the actual design discovered. +the sidebar from C". That's the actual design discovered. Close the handover with a +machine-legible reply template the user fills, so their reaction comes back as the next prompt +rather than prose to re-parse (on the mockup substrate this is what the copy-out terminator +lifts): + +```text +direction: +steal: from (repeat per piece) +skip: because +next-target: +``` ### 6. Capture the answer and clean up Per the shared discipline. Record which variant won and why, and record the directions that lost -with their reasons. When the verdict is a graft rather than a single winner, say which piece came -from where **and what the discarded parts held that the graft deliberately left behind**. The -deletions below are irreversible: whatever is not written down now is gone. +with their reasons. Capture at single-decision granularity: each named piece (a header, a +hierarchy choice, a primary affordance) gets its own steal/skip/adapt entry, so grafts compose +across variants instead of collapsing into one "variant B, mostly" note. When the verdict is a +graft rather than a single winner, say which piece came from where **and what the discarded parts +held that the graft deliberately left behind**. The deletions below are irreversible: whatever is +not written down now is gone. - **Sub-shape A**. Delete losing variants and the switcher; fold the winner into the existing page. - **Sub-shape B**. Promote the winner to a real route; delete the throwaway route and switcher. diff --git a/plugins/prototype/skills/explore-directions/evals/evals.json b/plugins/prototype/skills/explore-directions/evals/evals.json index 4955a9430..5e5c81aa3 100644 --- a/plugins/prototype/skills/explore-directions/evals/evals.json +++ b/plugins/prototype/skills/explore-directions/evals/evals.json @@ -32,7 +32,9 @@ "files": [], "expectations": [ "The generated variants differ structurally (layout / information hierarchy / primary affordance), not only in color or copy", - "The response declines to make recolors the only difference between variants, while still letting each variant carry its own visual direction on top of its structure" + "The response declines to make recolors the only difference between variants, while still letting each variant carry its own visual direction on top of its structure", + "The handover closes with a machine-legible reply template (direction / steal / skip / next-target) for the user to fill", + "The capture records steal/skip decisions at single-decision granularity (per named piece), so grafts compose across variants" ] }, { @@ -54,6 +56,7 @@ "files": [], "expectations": [ "The mockup is a single self-contained `file://` HTML page using synthetic data only (no real/captured values)", + "All variants render one identical data set: the data is the control variable, so only the design differs between variants", "The page includes a restrictive CSP meta tag that blocks remote scripts/fonts/fetches", "The mockup file is generated into a temp or gitignored scratch location, not a tracked repo path" ] diff --git a/plugins/prototype/skills/pressure-test/SKILL.md b/plugins/prototype/skills/pressure-test/SKILL.md index 96f64d769..47add1201 100644 --- a/plugins/prototype/skills/pressure-test/SKILL.md +++ b/plugins/prototype/skills/pressure-test/SKILL.md @@ -73,6 +73,15 @@ business, not the reducer, because the driver is not reading code: case, an attempt at something that should be illegal. Each is a short plain-language description plus the ordered buttons to press; starting a walkthrough **resets to a known initial state** so the scenario runs the same way every time. +5. **Validation answer set**, the questions this demo exists to answer, each as a forced choice + with a small authored option set, every option naming its cost in plain language (what + picking it gives up), plus a free-text escape hatch for the answer the options missed. The + driver's picks are the demo's real output; pair the set with a copy-out control that lifts + the filled answers back out as text to paste into the session. +6. **Fake-data disclosure footer**, one visible line stating the page is synthetic end to end, + that nothing on it reads from or writes to the real app, and where the real wiring lives (or + will live) behind which flag. The driver is not reading code; the footer is what keeps a + convincing mock from being mistaken for the wired feature. Constraints (the same set as explore-directions' HTML mockup substrate): @@ -172,9 +181,11 @@ Prototypes evolve. ### 7. Capture the answer -When done, capture what the prototype taught (per the shared discipline). The logic module behind -the shell is often worth keeping; the shell, TUI or HTML page, is not: lift the validated -module into production and delete the shell. +When done, capture what the prototype taught (per the shared discipline). For the HTML demo +shell, carry the filled validation answer set into the durable answer verbatim: the chosen +option per question, with the cost the driver accepted, is the record of what was actually +decided. The logic module behind the shell is often worth keeping; the shell, TUI or HTML page, +is not: lift the validated module into production and delete the shell. ## Anti-patterns diff --git a/plugins/prototype/skills/pressure-test/evals/evals.json b/plugins/prototype/skills/pressure-test/evals/evals.json index 8a44ec4b7..0d53269be 100644 --- a/plugins/prototype/skills/pressure-test/evals/evals.json +++ b/plugins/prototype/skills/pressure-test/evals/evals.json @@ -11,7 +11,8 @@ "The prototype records the question being tested at the top of the file before building anything", "The scheduling logic lives in a pure module (e.g. a reducer or explicit state machine) separate from the TUI shell", "The TUI re-renders the whole frame on each action (replace, not append) and lists available actions", - "The prototype is runnable via a single command — wired into the project's existing task runner, or documented at the top of a prototype README when the project has no task runner" + "The prototype is runnable via a single command — wired into the project's existing task runner, or documented at the top of a prototype README when the project has no task runner", + "When the HTML demo shell is the substrate, the page carries a validation answer set (forced-choice questions whose options each name their cost, plus a free-text escape hatch) and a visible fake-data disclosure footer, and the capture step carries the filled answer set into the durable answer" ] }, { From 133ffe9e4eafcd044da4e0da6c6f09f4b03b79ec Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:48:08 +0000 Subject: [PATCH 11/22] feat(planning): switch conditions, revision replies, free-text flag, ordering plan: rejected alternatives carry a one-line switch condition; Step 5 closes with pre-drafted one-line revision replies per flagged decision; the sanity-check paragraph names the every-step-lands-green expectation those checks back. interview: free-text answers get a free-text: resolution-field flag (register schema untouched; gate-invisibility recorded as a known limitation in context/loop.md). design: Phase 5 discussion rounds present findings in tweak-likelihood order. brainstorm gains the session-start rationale line, wayfind the five-pass workflow cross-ref. Adopted at the finding-your-unknowns integration sign-off (D28, D32, D33, D35, D36, D9, Q7); provenance in docs/FINDING-YOUR-UNKNOWNS.md. Evals extended for the four contract rows; check-open-questions test suite green; plugin 0.34.15 -> 0.35.0. Phase 7 of the integration plan. check-changed-skills green (0 errors). Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- .../finding-your-unknowns-integration/PLAN.md | 2 +- plugins/planning/.claude-plugin/plugin.json | 2 +- plugins/planning/CHANGELOG.md | 25 +++++++++++++++++++ plugins/planning/skills/brainstorm/SKILL.md | 2 ++ plugins/planning/skills/design/SKILL.md | 2 +- .../planning/skills/design/evals/evals.json | 1 + .../planning/skills/interview/context/loop.md | 2 ++ .../skills/interview/evals/evals.json | 1 + plugins/planning/skills/plan/SKILL.md | 8 ++++-- plugins/planning/skills/plan/evals/evals.json | 2 ++ plugins/planning/skills/wayfind/SKILL.md | 4 +++ 11 files changed, 46 insertions(+), 5 deletions(-) diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index e5af0ee65..249651d68 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -277,7 +277,7 @@ prototype SKILL.md files — no em dashes in any new SKILL.md line here. skills (before/after jq pairs); `bash scripts/check-changed-skills.sh origin/main` green; `bash scripts/check-purged-em-dashes.sh` green. -### Phase 7: planning plugin — interview/plan/design/brainstorm deltas [TODO] +### Phase 7: planning plugin — interview/plan/design/brainstorm deltas [DONE] interview SKILL.md: D28 free-text scrutiny flag as a RESOLUTION-FIELD convention (the 5-field register schema is untouched; gate-invisible by design, limitation recorded in the diff --git a/plugins/planning/.claude-plugin/plugin.json b/plugins/planning/.claude-plugin/plugin.json index 67055ad46..8b8a1a528 100644 --- a/plugins/planning/.claude-plugin/plugin.json +++ b/plugins/planning/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "planning", - "version": "0.34.15", + "version": "0.35.0", "userConfig": { "use_ask_user_question": { "type": "boolean", diff --git a/plugins/planning/CHANGELOG.md b/plugins/planning/CHANGELOG.md index 47a43da00..62c0e541d 100644 --- a/plugins/planning/CHANGELOG.md +++ b/plugins/planning/CHANGELOG.md @@ -3,6 +3,31 @@ All notable changes to the `planning` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [0.35.0] + +### Added + +- **`plan`: switch conditions on alternatives and pre-drafted revision replies.** Every rejected + alternative in the plan template now carries a one-line switch condition (the observable fact + that would make it the better choice), and Step 5's presentation closes with 2-4 pre-drafted + one-line revision replies, one per flagged close call, so the user's cheapest reaction is + pasting one back. The sanity-check paragraph also names what those checks back downstream: + implementation's every-step-lands-green expectation. Adopted from the "Finding Your Unknowns" + corpus at the integration sign-off (D32, D33, D35; provenance in + `docs/FINDING-YOUR-UNKNOWNS.md` in the marketplace repository). Evals extended. +- **`interview`: free-text resolution-field flag.** An answer arriving as free text rather than + an authored option is recorded with a `free-text:` prefix in the register row's resolution + field so downstream passes scrutinize it. Deliberately a convention inside the free-form + field: `check-open-questions.sh` grades statuses, not resolutions, so the flag is + gate-invisible (limitation recorded in `context/loop.md`); the 5-field register schema is + unchanged (D28). Evals extended. +- **`design`: tweak-likelihood ordering in discussion rounds.** Phase 5 findings are presented + in the same tweak-likelihood order `plan` Step 5 documents: contracts, data shapes, and + user-facing surfaces lead; mechanical threads sit at the bottom (D36). Evals extended. +- **`brainstorm`: session-start rationale line; `wayfind`: five-pass workflow cross-ref.** One + citable line each (D9; Q7), both pointing at the marketplace repository's + `docs/FINDING-YOUR-UNKNOWNS.md`. + ## [0.34.15] ### Changed diff --git a/plugins/planning/skills/brainstorm/SKILL.md b/plugins/planning/skills/brainstorm/SKILL.md index f9cef000a..beb37f494 100644 --- a/plugins/planning/skills/brainstorm/SKILL.md +++ b/plugins/planning/skills/brainstorm/SKILL.md @@ -12,6 +12,8 @@ metadata: The divergence step before any scoping: unknown-knowns (criteria the user only recognizes when seen) surface cheapest at candidate-list time. Finding one mid-implementation costs a re-plan. A brainstorm round also calibrates scope: reacting to a cheapest→most-ambitious spread prevents locking a scope that is too narrow (missed the high-value approach) or too wide (ambition the problem doesn't need). +Opening a fresh session on a rough problem with a brainstorm is a citable practice, not a detour: the cheapest→ambitious spread is the cheapest artifact that surfaces criteria the user only recognizes when seen (rationale and sources: `docs/FINDING-YOUR-UNKNOWNS.md` in the marketplace repository). + Distinct neighbors: `/planning:design` Phase 1 decomposes the problem space WITHIN a design task already chosen; a proactive architecture-friction scan (e.g. `/architecture:improve`, if installed) hunts on its own lanes; a UI-variation prototyper (e.g. `/prototype:explore-directions`, if installed) builds visual variations of a chosen direction. This skill is the general, problem-shaped entry upstream of all three. Creative-domain ideation owned by a domain skill (e.g. songwriting brainstorms → `/songwriting:workflow`, if installed) stays with that skill. ## Task diff --git a/plugins/planning/skills/design/SKILL.md b/plugins/planning/skills/design/SKILL.md index 58541939d..e6fd6da2e 100644 --- a/plugins/planning/skills/design/SKILL.md +++ b/plugins/planning/skills/design/SKILL.md @@ -136,7 +136,7 @@ Systematic gap-finding. For each round: 1. Re-read all design artifacts 2. Identify underspecified types, missing contracts, boundary friction, pattern concerns, and design-default gaps (configurability, extension axes, observability, testability). Record these as design threads -3. Present findings to user for discussion +3. Present findings to user for discussion, ordered by tweak likelihood (the same presentation default `/planning:plan` Step 5 documents): the threads the user is most likely to redirect — public contracts, data shapes, user-facing surfaces — lead the round; settled-looking mechanical threads sit at the bottom. Presentation order only; thread dependencies still govern what can resolve when 4. When discussion surfaces project-wide principles, suggest codifying them immediately in the project's own rules Continue rounds until no new gaps surface. Then run the `handoff` action, which invokes `/planning:design-handoff` via the Skill tool for the binary gate and plan-ready summary. diff --git a/plugins/planning/skills/design/evals/evals.json b/plugins/planning/skills/design/evals/evals.json index ff93d8894..bfc0a04f6 100644 --- a/plugins/planning/skills/design/evals/evals.json +++ b/plugins/planning/skills/design/evals/evals.json @@ -10,6 +10,7 @@ "expectations": [ "Output explores the design collaboratively, asking the user rather than unilaterally deciding the types/boundaries", "Output tracks design threads with a resolution status (resolved / directional / deferred)", + "Discussion-round findings are presented in tweak-likelihood order: contracts, data shapes, and user-facing surfaces lead; settled-looking mechanical threads sit at the bottom", "Output builds design artifacts (e.g. capability-matrix / type-inventory / design-threads / topology) rather than jumping to an implementation plan" ] }, diff --git a/plugins/planning/skills/interview/context/loop.md b/plugins/planning/skills/interview/context/loop.md index 27f8905b5..31b8d66c9 100644 --- a/plugins/planning/skills/interview/context/loop.md +++ b/plugins/planning/skills/interview/context/loop.md @@ -206,6 +206,8 @@ Fields: `Q | status | round | question | resolution`. Statuses: `Q` matches the terminal numbering, runs continuously across rounds, and never has a gap — a gap means a row was dropped after it was written, and the gate refuses to grade a register with one. +**Free-text flag — a resolution-field convention.** When an answer arrives as free text rather than a pick from the authored options — the escape hatch, a partial answer, a "whatever you think is best" — lead the resolution field with `free-text:` before the answer. Downstream passes (answer audits, plan formulation) treat flagged rows as deserving scrutiny rather than as settled picks: a free-text answer is where a misread lands silently. This lives inside the free-form resolution field by design; `check-open-questions.sh` grades statuses, not resolutions, so the flag is gate-invisible (a known limitation, recorded here) — a consumer needing mechanical reads of it means a register-schema change, carried by a version bump per the plugin's changelog discipline. + ### Drift check — a reply that does not answer is not an answer **After every user reply, before doing anything else, check the reply against the register's `open` rows.** Any row the reply did not address stays `open`, and you restate it at the top of your next response — even when the reply changed the subject entirely, even when you are mid-answer to something else, and even when the reply reads as agreement. Conversational drift is never consent, and the user changing the subject is ordinary conversation, not a defect on their side. diff --git a/plugins/planning/skills/interview/evals/evals.json b/plugins/planning/skills/interview/evals/evals.json index c5982337c..fa9e31cbd 100644 --- a/plugins/planning/skills/interview/evals/evals.json +++ b/plugins/planning/skills/interview/evals/evals.json @@ -155,6 +155,7 @@ "expectations": [ "A register row per asked question exists at ask-time, before any reply arrives — not written only once an answer lands", "The unrelated reply does not resolve any open question; every unaddressed row stays `open`", + "An answer arriving as free text rather than an authored option is recorded with the free-text: resolution-field flag so downstream passes scrutinize it", "The next response restates the still-open questions in one line rather than continuing as though they were answered", "The skill does not lock the contract or hand off while a register row is still `open`" ] diff --git a/plugins/planning/skills/plan/SKILL.md b/plugins/planning/skills/plan/SKILL.md index 399e1797c..7e2955f01 100644 --- a/plugins/planning/skills/plan/SKILL.md +++ b/plugins/planning/skills/plan/SKILL.md @@ -95,7 +95,10 @@ Produce a structured plan using the template in [context/plan-template.md](conte - **Approach**: the specific steps, in order - **Test strategy**: how we'll verify the changes work. For which test type each kind of change needs (unit / integration / e2e / architecture / analyzer), `/testing:plan`'s file-type classification table is the SSOT **when the `testing` plugin is installed**; **invoke `/tdd:principles` via the Skill tool (if installed)** when formulating this section for authoritative guidance on what to test, which testing style fits, and when to mock; otherwise apply standard test-design judgment. TDD is the default approach. The test strategy should specify Red-Green-Refactor unless genuinely impractical. **Name the test boundaries**. The public interfaces the tests will drive, and for each whether it already exists or is being introduced (prefer driving an existing interface over introducing one for testability alone). Naming them is what lets Step 5's approval settle them, so implementation writes no test against a boundary the plan never named; on an unattended run, a boundary chosen during implementation that this section did not name is a deviation, logged for PR-time review (`DEVIATIONS.md` beside `PLAN.md` in the contract slice) rather than silently taken - **Files affected**: what gets created, modified, or deleted -- **Alternatives considered**: what was rejected and why +- **Alternatives considered**: what was rejected and why — and, per alternative, a one-line + switch condition: the observable fact that, if it turned up, would make this the better choice. + A rejection with no switch condition is not revisable; the condition is what lets a reviewer + (or a later phase) flip the decision without re-deriving the analysis - **Risks and mitigations**: what could go wrong Scale the plan to the task: @@ -117,7 +120,7 @@ Per-scale calibration examples live in [context/plan-template.md](context/plan-t **File inventory for large-scope plans**. When a plan or phase touches ≥10 files, emit a checkbox inventory table per phase (file, action, rationale). Checkboxes enforce verification discipline. The agent ticks each file as processed; the reviewer sees completeness at a glance. Include KEEP rows for files audited and deliberately left unchanged. Full format in [context/plan-template.md](context/plan-template.md) "File Inventory". -**Sanity-check verifiable-criterion enforcement**. Every phase ends with at least one `**Sanity Check:**` bullet. Criteria MUST be mechanically verifiable (a specific grep, file Read assertion, build exit code, test exit code, or runtime probe). Never vague (~~"documented appropriately"~~, ~~"behaves as expected"~~, ~~"all cases covered"~~). Rewrite vague criteria as exact commands a fresh session can execute without inferential judgement. Full format guide in [context/plan-template.md](context/plan-template.md) "Sanity-Check Format". +**Sanity-check verifiable-criterion enforcement**. Every phase ends with at least one `**Sanity Check:**` bullet. Criteria MUST be mechanically verifiable (a specific grep, file Read assertion, build exit code, test exit code, or runtime probe). Never vague (~~"documented appropriately"~~, ~~"behaves as expected"~~, ~~"all cases covered"~~). Rewrite vague criteria as exact commands a fresh session can execute without inferential judgement. Full format guide in [context/plan-template.md](context/plan-template.md) "Sanity-Check Format". Phase-scoped sanity checks are also what backs implementation's every-step-lands-green expectation: each phase can be verified green at its own boundary rather than only at the end. **Build-technique selection**. BEFORE ordering phases, pick the de-risking technique by the task's *uncertainty type*, not by habit. A design or viability unknown (*might abandon*) resolves **upstream**. `/planning:design` for design-significant questions, a throwaway `/prototype:pressure-test` spike (if installed) or research for raw feasibility. And plan **consumes that outcome** rather than re-deriving it inline. Plan's own call is the *kept* slice: a tracer bullet / walking skeleton when you are committed to ship and the risk is integration. When an upstream feasibility spike and ship-commitment both hold, sequence the kept tracer-bullet slice after the spike's outcome lands. Trivial / pure-horizontal work skips all techniques. @@ -219,6 +222,7 @@ Present the final plan to the user. The plan is a proposal, not a commitment. Th 4. **Execution shape** (from Step 4.5). Parallelism shape AND per-phase routing table. Skipped for single-phase plans 5. **Decisions made (gate-passed)** (from Step 4.6). TABLE per [context/tag-decisions.md](context/tag-decisions.md) "Presentation contract": `Decision | What it changes in the plan | Basis (evidence)`, one row per gate-passed `[EXEC-SHAPE]` / `[FALLBACK]` tag, written for a cold reader (no session shorthand). Below-bar decisions never appear here. They were interviewed before the plan locked. An empty section ("no unilateral decisions. Every PLAN item traces to brief") is also valid output 6. **Explicit approval request**: "Approve this plan to proceed to execution, or provide feedback to revise. Anything tagged `[EXEC-SHAPE]` or `[FALLBACK]` above is /planning:plan's discretion. Flag any you want changed." +7. **Highest-leverage replies**: close with 2-4 pre-drafted one-line revision replies, one per flagged close call or gate-passed decision, each a copyable sentence that flips exactly that decision (e.g. "Switch phase 2 to the queue-based alternative"). The user's cheapest possible reaction is pasting one back; a presentation whose flagged decisions have no pre-drafted flip line makes the user compose the revision themselves **Presentation order. Tweak-likelihood first.** Order the presentation by what the user is most likely to change on review: data-model/schema choices, type interfaces and public contracts, and user-facing surfaces LEAD (flag close calls with their alternatives); mechanical refactoring and low-judgment work sits at the bottom. Presentation order only. Phase EXECUTION order stays integration-first per Step 2. Optionally offer a self-contained HTML plan view (decisions-first layout, flagged choices with toggleable alternatives), rendered to the topic-docs **ephemeral tier**, never the contract slice beside `PLAN.md`, which stays the tracked record. Placement and rules: [`${CLAUDE_PLUGIN_ROOT}/reference/topic-docs.md`](${CLAUDE_PLUGIN_ROOT}/reference/topic-docs.md). diff --git a/plugins/planning/skills/plan/evals/evals.json b/plugins/planning/skills/plan/evals/evals.json index a896cc324..367187559 100644 --- a/plugins/planning/skills/plan/evals/evals.json +++ b/plugins/planning/skills/plan/evals/evals.json @@ -9,8 +9,10 @@ "files": [], "expectations": [ "Output produces a structured plan with at minimum a goal, an ordered approach, and a test strategy", + "Each rejected alternative carries a one-line switch condition naming the observable fact that would make it the better choice", "Output includes a blast-radius assessment line", "Output ends by requesting explicit user approval before execution rather than starting to implement", + "The presentation closes with pre-drafted one-line revision replies, one per flagged close call, each a copyable sentence that flips exactly that decision", "The run does not write or edit any source code file" ] }, diff --git a/plugins/planning/skills/wayfind/SKILL.md b/plugins/planning/skills/wayfind/SKILL.md index c7d801166..e80219761 100644 --- a/plugins/planning/skills/wayfind/SKILL.md +++ b/plugins/planning/skills/wayfind/SKILL.md @@ -179,6 +179,10 @@ owns the trigger's meaning (too-big + fog, both, not either alone). | External-evidence item | `/discovery:research` | `research`-typed items route here (autonomous) | | The plan itself | `/planning:plan` | Graduation target when the destination is a PLAN | +For pre-implementation efforts, the routed items above compose into a known five-pass order +(blindspot → brainstorm/prototype → interview → reference port → plan); the workflow section of +`docs/FINDING-YOUR-UNKNOWNS.md` in the marketplace repository states it with rationale. + ## What this skill does NOT do - **Does not do build work.** A map holds decisions; build items live on the ordinary tracker From 5c64a6540e1026656a5731c232dcefa9d40b38ee Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:50:47 +0000 Subject: [PATCH 12/22] feat(discipline): no-analogue port trap + canonical invocation hint point-dont-copy's audit list gains the cross-stack port trap: a source-side primitive with no target-side analogue whose invariant the port silently drops must name the convention now carrying it, or the finding stands. The skill gains an argument-hint showing the canonical invocation (frontmatter only; description and trigger keywords untouched). Adopted at the finding-your-unknowns integration sign-off (E7, E8); provenance in docs/FINDING-YOUR-UNKNOWNS.md. Evals extended; plugin 0.12.20 -> 0.13.0 (changelog entry provisional, finalized when the same version's E6 port gate lands in the Wave-3 commit). Phase 8 of the integration plan. check-changed-skills green (0 errors). Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- .../finding-your-unknowns-integration/PLAN.md | 2 +- plugins/discipline/.claude-plugin/plugin.json | 2 +- plugins/discipline/CHANGELOG.md | 14 ++++++++++++++ plugins/discipline/skills/point-dont-copy/SKILL.md | 8 +++++++- .../skills/point-dont-copy/evals/evals.json | 3 ++- 5 files changed, 25 insertions(+), 4 deletions(-) diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index 249651d68..f2b0f87f1 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -295,7 +295,7 @@ affected-tests does not select it for SKILL.md/context edits; D28 must not break register gate); evals growth for interview/plan/design (before/after jq pairs); `bash scripts/check-changed-skills.sh origin/main` green. -### Phase 8: discipline plugin — point-dont-copy W2 deltas [TODO] +### Phase 8: discipline plugin — point-dont-copy W2 deltas [DONE] point-dont-copy SKILL.md: E7 primitive-to-convention trap item (CONTRACT) in the "Audit. What to look for" list; E8 canonical invocation line (as `argument-hint` diff --git a/plugins/discipline/.claude-plugin/plugin.json b/plugins/discipline/.claude-plugin/plugin.json index f85625b3a..ef66f7884 100644 --- a/plugins/discipline/.claude-plugin/plugin.json +++ b/plugins/discipline/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "discipline", - "version": "0.12.20", + "version": "0.13.0", "description": "Discipline correctors that re-anchor a standing rule mid-session, then audit both the work in flight and the pre-existing state and choices it trusts, and correct what has drifted: do-your-research (research and no-assumptions discipline; sibling do-your-research-deep escalates to a typed full inventory of the session's claims — assumptions, asserted facts, concrete specifics, load-bearing premises — verified at a configurable depth and reported as a per-item ledger), follow-our-standards (alignment to the consuming org's engineering conventions), point-dont-copy (pointer-over-copy discipline — no copied content, internal-name coupling, or closed capability lists), reason-dont-recite (interrogate inherited content — precedent is evidence of what is, never self-justifying authority), tighten-your-output (terseness discipline — fewer words or lines with no loss of meaning or correctness), recheck-against-upstream (existing state is not evidence of its own correctness — audit config, code, and infra against current official upstream docs; sibling recheck-against-upstream-deep fans subagents doc-by-doc over a whole subsystem), pick-for-the-problem (tool, library, framework, and approach selection fitted to the problem, not reached for out of habit, availability, incumbency, or preconception), mind-your-maxims (cooperative-communication discipline per Grice plus the AI-augmented transparency maxim), script-the-deterministic-work (offload deterministic sub-work — counts, diffs, sorts, transforms, and scaffolds — to a script that runs, reserving model output for judgment over its real output; the audit runs both ways, also catching an existing script that over-reaches into judgement), use-your-skills (actually use the skills already in context — scan the listing, map the task, invoke the fitting skill instead of reinventing it, and name skills when delegating to a subagent), and reuse-or-replace (anti-fragmentation — new work reuses an established way of doing something or openly replaces it (migrate the old uses, record the decision), never silently stands up a second parallel way; divergence is allowed but owes a recorded reason proportional to blast radius), and scrutinize-dont-coast (adversarial self-scrutiny — stop coasting on your own recent output and re-examine whether it is sound, not merely confidently produced, through a fresh-context pass blind to the reasoning that made it, then remediate with the user; it stops the trajectory first and remediates collaboratively rather than autonomously). Plus further species that are not correctors (examples, not a fixed list — each skill's own description is authoritative), including setup, sweep-all, a posture-batch runbook that composes them — it fans out an audit-only subagent per in-scope corrector, then applies the corrections on the main thread in a fixed order, with batch membership and order set by each corrector's own colocated tier metadata and an optional userConfig overlay — and wait-what, a one-shot user-invoked-only communication repair: type /discipline:wait-what when the last message did not land and the model re-pitches it, backing up as far as needed, adding the missing context, in ASD-STE100 Simplified Technical English, using the project's ubiquitous language; never model-invoked and never in the batch. Firing a corrector is a re-anchor, not an accusation; the audit may return clean.", "author": { "name": "Melodic Software", diff --git a/plugins/discipline/CHANGELOG.md b/plugins/discipline/CHANGELOG.md index efcd06008..cca7a2289 100644 --- a/plugins/discipline/CHANGELOG.md +++ b/plugins/discipline/CHANGELOG.md @@ -5,6 +5,20 @@ All notable changes to the `discipline` plugin are documented here. Format follo Entries below `0.9.0` were released under the plugin's former name, `re-anchor`. +## [0.13.0] + +### Added + +- **`point-dont-copy`: the no-analogue trap check and a canonical invocation hint.** The audit + list gains the cross-stack port trap: a source-side primitive with no target-side analogue (a + language feature, a library guarantee, an implicit runtime behavior) whose invariant the port + silently drops — the port must name the convention now carrying that invariant, or the finding + stands. The skill also gains an `argument-hint` showing the canonical invocation. Adopted from + the "Finding Your Unknowns" corpus at the integration sign-off (E7, E8; provenance in + `docs/FINDING-YOUR-UNKNOWNS.md` in the marketplace repository). Evals extended. + (Provisional entry: the same version's external-reference port gate lands in a follow-up + commit on this branch and extends this entry.) + ## [0.12.20] ### Changed diff --git a/plugins/discipline/skills/point-dont-copy/SKILL.md b/plugins/discipline/skills/point-dont-copy/SKILL.md index e2e8e89f8..17f9730d1 100644 --- a/plugins/discipline/skills/point-dont-copy/SKILL.md +++ b/plugins/discipline/skills/point-dont-copy/SKILL.md @@ -1,5 +1,6 @@ --- description: "Re-anchor pointer-over-copy discipline, then audit the work in flight for copied content, internal-name coupling, and closed capability lists, and correct by pointing at the living source. Use when: 'point don't copy', 'you copied that', 'don't duplicate the docs', 'cite instead of paste', 'link don't restate', 'you enumerated the tools', 'that couples to internal names', 'this will drift', or at conversation start on documentation work." +argument-hint: "[target] (e.g., /discipline:point-dont-copy the new setup guide, or empty to audit the work in flight)" user-invocable: true disable-model-invocation: false metadata: @@ -81,7 +82,12 @@ Name concrete, located findings (per the method doc's step 2, self-audit): invocation contract would do; - a closed enumeration of duties or mechanisms that will drift as the surface evolves; -- the same passage, literal, or concept appearing in two or more places. +- the same passage, literal, or concept appearing in two or more places; +- in a port from another stack or language: a source-side primitive with no + target-side analogue (a language feature, a library guarantee, an + implicit runtime behavior) whose invariant the port silently drops. The + port must name the convention now carrying that invariant, or the + finding stands. Correct each forward now: replace the copy with a pointer to its owner, swap an internal-name reference for the public contract, and reopen a diff --git a/plugins/discipline/skills/point-dont-copy/evals/evals.json b/plugins/discipline/skills/point-dont-copy/evals/evals.json index 5dc589872..a24331e84 100644 --- a/plugins/discipline/skills/point-dont-copy/evals/evals.json +++ b/plugins/discipline/skills/point-dont-copy/evals/evals.json @@ -10,7 +10,8 @@ "expectations": [ "Flags both the verbatim-pasted reference and the reworded paraphrase as duplication findings", "States that a reworded paraphrase drifts the same as a verbatim copy", - "Corrects forward by replacing the copied content with a pointer/citation to the owning source" + "Corrects forward by replacing the copied content with a pointer/citation to the owning source", + "In a cross-stack port, a source-side primitive with no target analogue is flagged unless the port names the convention now carrying its invariant" ] }, { From 6221b960ed4ec05cf085a4cfd79304665f01ca61 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:53:03 +0000 Subject: [PATCH 13/22] docs(session-flow): five-pass pre-implementation cross-ref in workflow One doc line under the at-a-glance diagram: stages 0-3 expand, for unfamiliar territory, into the five-pass order the marketplace repo's docs/FINDING-YOUR-UNKNOWNS.md states with rationale (integration sign-off Q7). Doc-line-only bump 0.34.14 -> 0.34.15; no contract change, no eval delta. Phase 9 of the integration plan. check-changed-skills green. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- docs/topics/finding-your-unknowns-integration/PLAN.md | 2 +- plugins/session-flow/.claude-plugin/plugin.json | 2 +- plugins/session-flow/CHANGELOG.md | 11 +++++++++++ plugins/session-flow/skills/workflow/SKILL.md | 4 ++++ 4 files changed, 17 insertions(+), 2 deletions(-) diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index f2b0f87f1..75b86002d 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -308,7 +308,7 @@ intermediate commit's entry knowingly under-describes and Phase 10 finalizes it) **Sanity Check:** grep E7 line + `argument-hint` in frontmatter; evals growth (before/after jq pair); `bash scripts/check-changed-skills.sh origin/main` green. -### Phase 9: session-flow plugin — workflow cross-ref [TODO] +### Phase 9: session-flow plugin — workflow cross-ref [DONE] plugins/session-flow/skills/workflow/SKILL.md: the five-pass workflow-section cross-ref doc line (Brief constraint 7 / Q7 — marketplace-repo prose pointing at the F1 workflow diff --git a/plugins/session-flow/.claude-plugin/plugin.json b/plugins/session-flow/.claude-plugin/plugin.json index 184493729..c2c1d965f 100644 --- a/plugins/session-flow/.claude-plugin/plugin.json +++ b/plugins/session-flow/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "session-flow", - "version": "0.34.14", + "version": "0.34.15", "description": "Session-lifecycle toolkit of fourteen skills: workflow (navigate a staged dev workflow and suggest the next stage), handoff (write a save-point and resume prompt for /clear-and-resume), continue-in-background (delegate the task to a fresh background agent that continues it now \u2014 same save-point engine as handoff, delivered by launching a detached claude --bg session seeded with the resume prompt; launches only on explicit user request), keep-going (recover and continue after any interruption OR when live off-thread work looks stalled \u2014 inventory off-thread work, inspect its real output, act only on evidence, then continue; after a usage limit lifts it continues rather than summarizing-and-stalling), find-handoff (recover a lost handoff after /clear \u2014 when the resume prompt was written but never copied \u2014 via a read-only detection ladder: known-location glob of the handoffs dir, then a bounded, recency-ranked transcript scan for the handoff directive and dashed-rail markers, then a confirm-before-resume gate; surfaces only the resume prompt + metadata, never raw transcript content), clean-stop (get to a durable, linked stopping point before the machine may go away \u2014 sweep every repo/worktree for uncommitted, unpushed, or PR-less work, push it durable, put breadcrumbs in PR/issue bodies, then give a free-and-clear verdict), retro (structured end-of-session retrospective with transcript metrics and learning codification), running-retro (in-flight retrospective checkpoints that spawn a subagent to analyze the transcript so far and append classified findings to a cumulative running ledger \u2014 capture and route only, the live counterpart to retro; also owns a detached-observer substrate that can watch a session out-of-band and run the checkpoint autonomously after the session ends), orient (read-only session orientation \u2014 synthesize where we stand, what we are doing, and why, from durable + off-thread state the built-in /recap never sees: ledgers, handoffs, workflow checklists, running-retro ledgers, open PRs and work-items, and git), orchestrate (arm a session or worker with proactive-orchestration imperatives), reanchor (verify a session's working assumptions are still true against live reality \u2014 referenced PRs/issues/branches, base-branch drift, renamed/version-drifted surfaces, stale memory-tier files, and the goal a handoff records, compared across the chain so a re-derived goal reports as drift \u2014 before building on them), reconcile (retire finished off-thread work and reconcile this session's task ledger with reality \u2014 the prune-and-reconcile counterpart to keep-going's resume: inventory the work this session spawned, inspect its real state, retire the finished and close proven-done tasks, auto-settling the finished and gating any kill of still-running work; sibling sessions in the project are reported read-only), setup (check-centric verification of the observer's runtime prerequisites and configuration), and show-options (lay out which skills fit this moment as a ranked, nothing-hidden menu \u2014 a shortlist per bucket plus the complete remainder by name, resolved from the full installed catalog rather than the truncated in-context listing, so the human decides and no option is withheld for looking already-done).", "author": { "name": "Melodic Software", diff --git a/plugins/session-flow/CHANGELOG.md b/plugins/session-flow/CHANGELOG.md index 9a8853c62..823bd3be7 100644 --- a/plugins/session-flow/CHANGELOG.md +++ b/plugins/session-flow/CHANGELOG.md @@ -1,5 +1,16 @@ # Changelog — session-flow plugin +## [0.34.15] + +### Added + +- **`workflow`: five-pass pre-implementation cross-ref.** One doc line under the at-a-glance + diagram noting that stages 0-3 expand, for unfamiliar territory, into the known five-pass + order (blindspot → brainstorm/prototype → interview → reference port → plan), stated with + rationale in `docs/FINDING-YOUR-UNKNOWNS.md` in the marketplace repository. Recorded at the + finding-your-unknowns integration sign-off (Q7); no contract or trigger change, so no eval + delta. + ## [0.34.14] ### Changed diff --git a/plugins/session-flow/skills/workflow/SKILL.md b/plugins/session-flow/skills/workflow/SKILL.md index 0bf20f571..d6643398a 100644 --- a/plugins/session-flow/skills/workflow/SKILL.md +++ b/plugins/session-flow/skills/workflow/SKILL.md @@ -80,6 +80,10 @@ Parse the first argument to determine mode; when it is `continue`, parse the sec PR lifecycle: prep → create → monitor CI → merge (runs after step 7) ``` +Stages 0-3 expand, for unfamiliar territory, into a known five-pass pre-implementation order +(blindspot → brainstorm/prototype → interview → reference port → plan); the workflow section of +`docs/FINDING-YOUR-UNKNOWNS.md` in the marketplace repository states it with rationale. + ### 2. Detect current position Check conversation context for evidence of completed stages: From ba7a040aa8dba7ad59448d5ec3dc16313df3f54f Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 18:57:29 +0000 Subject: [PATCH 14/22] feat(discipline,implementation): port gate + deviation-log convention (Wave 3) point-dont-copy gains the external-reference port gate as a declared step delta on the shared re-anchor/audit/correct loop: a five-section semantics map (side-by-side pairs, preserved/changed/dropped ledger, edge-case parity, open questions) and a stop-and-wait confirmation gate, scoped to ports whose source of truth lives outside this repo's tree (vendored, foreign-language, other-repo); in-tree corrections stay do-it-now, and the no-analogue trap check feeds the dropped ledger. Verified against the shared method doc before landing: its declared-step-deltas allowance is the exact seam, no correct-forward reversal. Finalizes discipline 0.13.0's changelog entry; a new eval case covers the gate. The implementation plugin's deviation log gains typed entries (plan-confirmed / discovery / deviation / human-decision) with four deviation fields (plan said / found / chose / revisit) in implement-dispatch's owning contract, an interactive opt-in in implement Step 3, and a completion fold-back in Step 5 that reads DEVIATIONS.md and emits plan-amendment bullets. Recorded registry trigger: a second plugin reading the log graduates the contract to an owner doc. No registry row now, per the signed C5/M3. Evals extended in both skills; implementation 0.15.8 -> 0.16.0. Adopted at the finding-your-unknowns integration sign-off (E6 per C3 boundary, E1+E2 per C5); provenance in docs/FINDING-YOUR-UNKNOWNS.md. Phase 10 of the integration plan. check-changed-skills and the em-dash ratchet both green. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- .../finding-your-unknowns-integration/PLAN.md | 2 +- plugins/discipline/CHANGELOG.md | 10 ++++++++-- .../skills/point-dont-copy/SKILL.md | 20 +++++++++++++++++++ .../skills/point-dont-copy/evals/evals.json | 13 ++++++++++++ .../implementation/.claude-plugin/plugin.json | 2 +- plugins/implementation/CHANGELOG.md | 17 ++++++++++++++++ .../skills/implement-dispatch/SKILL.md | 3 +++ .../implement-dispatch/evals/evals.json | 1 + .../implementation/skills/implement/SKILL.md | 9 ++++++--- .../skills/implement/evals/evals.json | 3 ++- 10 files changed, 72 insertions(+), 8 deletions(-) diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index 75b86002d..d9c99e2ce 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -318,7 +318,7 @@ extension — no contract change). **Sanity Check:** grep the cross-ref line in workflow SKILL.md; session-flow plugin.json + CHANGELOG in the diff; `bash scripts/check-changed-skills.sh origin/main` green. -### Phase 10: Wave 3 — E6 port gate + E1/E2 deviation-log convention [TODO] +### Phase 10: Wave 3 — E6 port gate + E1/E2 deviation-log convention [DONE] Own review moment; lands after all W2 phases. diff --git a/plugins/discipline/CHANGELOG.md b/plugins/discipline/CHANGELOG.md index cca7a2289..70118c6b8 100644 --- a/plugins/discipline/CHANGELOG.md +++ b/plugins/discipline/CHANGELOG.md @@ -16,8 +16,14 @@ Entries below `0.9.0` were released under the plugin's former name, `re-anchor`. stands. The skill also gains an `argument-hint` showing the canonical invocation. Adopted from the "Finding Your Unknowns" corpus at the integration sign-off (E7, E8; provenance in `docs/FINDING-YOUR-UNKNOWNS.md` in the marketplace repository). Evals extended. - (Provisional entry: the same version's external-reference port gate lands in a follow-up - commit on this branch and extends this entry.) +- **`point-dont-copy`: semantics map + confirmation gate for external-reference ports.** A + declared step delta on the shared re-anchor/audit/correct loop, scoped to ports whose source + of truth lives outside this repo's tree (vendored, foreign-language, other-repo): between + audit and correct-forward, produce a five-section semantics map (what the source does, + side-by-side pairs, preserved/changed/dropped ledger, edge-case parity table, open questions) + and stop at a confirmation gate until the user confirms it — "semantics confirmed" recommended, + not required. In-tree corrections stay do-it-now; the no-analogue trap check feeds the dropped + ledger. Same adoption basis (E6, boundary per the signed C3); a new eval case covers the gate. ## [0.12.20] diff --git a/plugins/discipline/skills/point-dont-copy/SKILL.md b/plugins/discipline/skills/point-dont-copy/SKILL.md index 17f9730d1..b2bf0865b 100644 --- a/plugins/discipline/skills/point-dont-copy/SKILL.md +++ b/plugins/discipline/skills/point-dont-copy/SKILL.md @@ -95,6 +95,26 @@ closed enumeration into a general duty with marked examples. Where content is genuinely this project's own to hold (an adapted config, a self-pinned constraint, a dated research deliverable), say so and leave it. +## External-reference ports. Semantics map + confirmation gate + +**Scope.** This section governs external-reference ports only: work whose source of truth +lives outside this repo's tree: vendored, foreign-language, other-repo. In-tree +corrections stay on the method doc's do-it-now side; nothing here changes that. + +**Declared step delta** (per the method doc's "Declared step deltas" allowance): for an +external-reference port, insert between the loop's steps 2 and 3 a semantics map and a +confirmation gate, because a port that starts before the semantics are agreed bakes +misreads into working code, where the audit can no longer see them as findings: + +- **Semantics map**, externalized, five sections: what the source does; side-by-side + pairs (source construct against port construct); a preserved / changed / dropped + ledger; an edge-case parity table; the open questions the port cannot settle alone. + The no-analogue trap check above feeds the dropped ledger: every dropped source + primitive names the convention now carrying its invariant. +- **Confirmation gate.** Stop and wait for the user to confirm the map before port work + proceeds. A reply of "semantics confirmed" is the recommended token, not a required + one; any clear confirmation opens the gate. + ## What this skill does NOT do - **Does not strip legitimate local content.** An adapted config, a diff --git a/plugins/discipline/skills/point-dont-copy/evals/evals.json b/plugins/discipline/skills/point-dont-copy/evals/evals.json index a24331e84..42444ca4f 100644 --- a/plugins/discipline/skills/point-dont-copy/evals/evals.json +++ b/plugins/discipline/skills/point-dont-copy/evals/evals.json @@ -49,6 +49,19 @@ "Recommends consolidating to a single source and pointing at it", "Leaves room for a merits-based legitimate-divergence exception rather than an absolute rule" ] + }, + { + "id": 5, + "name": "external-reference-port-gate", + "prompt": "Port this Python rate-limiter module from the vendored library into our TypeScript services package. Keep the behavior identical.", + "expected_output": "Recognizes an external-reference port (source of truth outside this repo's tree: vendored, foreign-language, other-repo) and applies the declared step delta: produces the five-section semantics map (what the source does; side-by-side pairs; preserved/changed/dropped ledger; edge-case parity table; open questions), with the no-analogue trap check feeding the dropped ledger, then stops at the confirmation gate and waits for the user to confirm the map before any port code is written. A 'semantics confirmed' reply is recommended, not required.", + "files": [], + "expectations": [ + "Classifies the task as an external-reference port (source of truth outside this repo's tree) and applies the semantics-map step delta", + "Produces a semantics map with side-by-side pairs, a preserved/changed/dropped ledger, and an edge-case parity table before porting", + "Stops and waits for the user to confirm the map before port work proceeds, recommending but not requiring a 'semantics confirmed' reply", + "Does not extend the stop-and-wait gate to in-tree corrections, which stay do-it-now per the shared method doc" + ] } ] } diff --git a/plugins/implementation/.claude-plugin/plugin.json b/plugins/implementation/.claude-plugin/plugin.json index b005a5cee..aae0acca9 100644 --- a/plugins/implementation/.claude-plugin/plugin.json +++ b/plugins/implementation/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "implementation", - "version": "0.15.8", + "version": "0.16.0", "description": "Disciplined implementation stage: execute approved plans inline (`/implementation:implement`) or via orchestrated worker subagents (`/implementation:implement-dispatch`) with incremental validation, TDD-by-default cadence, green-checkpoint commits, scope-fence drift detection, and divergence detection that routes back to planning. Build/test/lint, testing, and outcome verification live in the companion `toolchain`, `testing`, and `verification` plugins, invoked when installed.", "author": { "name": "Melodic Software", diff --git a/plugins/implementation/CHANGELOG.md b/plugins/implementation/CHANGELOG.md index b40a95b59..9e2fe1f05 100644 --- a/plugins/implementation/CHANGELOG.md +++ b/plugins/implementation/CHANGELOG.md @@ -3,6 +3,23 @@ All notable changes to the `implementation` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [0.16.0] + +### Added + +- **Deviation-log convention: typed entries, interactive opt-in, and a completion fold-back.** + `implement-dispatch`'s "Divergence in non-interactive runs" contract gains entry types + (plan-confirmed / discovery / deviation / human-decision, the last marked blocking or + non-blocking) and four deviation fields (plan said / found / chose / revisit), stated as this + plugin's own output contract for its own log file, never a format imposed on consumer repos. + `implement` Step 3 gains the interactive opt-in (same log, same contract, for long or + contested sessions), and Step 5 gains a deviation fold-back item: read `DEVIATIONS.md` at + completion and emit one plan-amendment bullet per unresolved entry. Recorded trigger: when a + second plugin reads `DEVIATIONS.md`, the marketplace's convention-registry rule fires and the + contract graduates to an owner doc. Adopted from the "Finding Your Unknowns" corpus at the + integration sign-off (E1, E2; provenance in `docs/FINDING-YOUR-UNKNOWNS.md` in the + marketplace repository). Evals extended in both skills. + ## [0.15.8] ### Changed diff --git a/plugins/implementation/skills/implement-dispatch/SKILL.md b/plugins/implementation/skills/implement-dispatch/SKILL.md index fc78cdbf3..93e0a546f 100644 --- a/plugins/implementation/skills/implement-dispatch/SKILL.md +++ b/plugins/implementation/skills/implement-dispatch/SKILL.md @@ -68,6 +68,9 @@ In a session with no human to escalate to, stop-and-escalate on Moderate diverge - **Evidence is a pointer, not prose**. A commit SHA, a `file:line`, a test name, an artifact path. Prefer evidence a committed script produced over a hand-made one-off, so the reviewer can re-run it rather than believe it. - **Carry the outcome, not just the choice.** An entry whose result is still unknown says so (`unverified`) rather than reading as settled; an entry claiming a result names the check that produced it. State which work is unverified rather than omitting the distinction, the same grounding rule `work-items:work-loop` and `source-control:babysit-loop` apply to their cycle reports. - **One entry is one decision.** If it does not fit on a line or two, the decision is not crisp yet, split it, or say plainly that it is still open. +- **Entries are typed, and a deviation carries four fields.** Type each entry as one of: plan-confirmed (a load-bearing plan assumption checked out), discovery (something learned the plan never spoke to), deviation (the plan said X, the run did Y), or human-decision (a call only a person can make, marked blocking or non-blocking). A deviation entry answers: plan said / found / chose / revisit. This taxonomy is this plugin's own output contract for its own log file, never a format imposed on consumer repos. + +Interactive sessions may opt into this same log rather than leaving Moderate adjustments in scrollback (see `/implementation:implement` "Step 3: Divergence Detection"); the house posture and rationale live in `docs/FINDING-YOUR-UNKNOWNS.md` in the marketplace repository. Recorded trigger: the moment a second plugin READS `DEVIATIONS.md` rather than writing its own, the marketplace's convention-registry rule fires and this contract graduates to an owner doc with a registry row. An entry whose evidence does not resolve, or whose result was never verified, is the PR review catching a gap. That is the log working. Major divergence (fundamental assumption wrong) still STOPS even autonomously. Park the run with a handoff note rather than improvising a new design. Interactive sessions keep the `/implementation:implement` "Step 3: Divergence Detection" escalation ladder unchanged. diff --git a/plugins/implementation/skills/implement-dispatch/evals/evals.json b/plugins/implementation/skills/implement-dispatch/evals/evals.json index d5d4bab9e..28419c252 100644 --- a/plugins/implementation/skills/implement-dispatch/evals/evals.json +++ b/plugins/implementation/skills/implement-dispatch/evals/evals.json @@ -61,6 +61,7 @@ "Does NOT deadlock the autonomous run waiting for a human on a moderate divergence", "Picks the conservative option (closest to plan intent, smallest blast radius) for the moderate divergence", "Logs the deviation to a DEVIATIONS.md beside the plan artifact (what was planned, what was done, why, blast radius)", + "The logged entry is typed (plan-confirmed / discovery / deviation / human-decision), and a deviation entry answers plan said / found / chose / revisit", "Reserves a hard STOP for major divergence (a fundamental assumption wrong), not this moderate one" ] }, diff --git a/plugins/implementation/skills/implement/SKILL.md b/plugins/implementation/skills/implement/SKILL.md index 77d365a6e..ede901f1a 100644 --- a/plugins/implementation/skills/implement/SKILL.md +++ b/plugins/implementation/skills/implement/SKILL.md @@ -116,6 +116,8 @@ Most important discipline in execution. Plans are hypotheses, implementation is **Non-interactive fork (autonomous runs only):** see `/implementation:implement-dispatch` "Divergence in non-interactive runs". Moderate divergence takes the conservative option + a deviations log instead of deadlocking; Major still STOPS. Interactive sessions keep the escalation ladder above unchanged. +**Opt-in deviation log (interactive):** an interactive session may keep the same append-only `DEVIATIONS.md` beside the plan artifact, typing entries per the contract in `/implementation:implement-dispatch` "Divergence in non-interactive runs" (plan-confirmed / discovery / deviation / human-decision; a deviation answers plan said / found / chose / revisit). Worth opting into when the session is long, the plan is contested, or a handoff is likely: the Moderate rung's "document what changed and why" then has a durable home instead of scrollback, and Step 5's fold-back has something to read. + ## Step 3.5: Scope-fence drift detector (run at every decision boundary) **When**: at each phase boundary, at each worker-agent return, and BEFORE proposing any action not literally in the approved plan's work items. @@ -177,13 +179,14 @@ When all planned work is done: 1. **Final build check**. Invoke `/toolchain:check` via the Skill tool for all affected ecosystems when the `toolchain` plugin is installed; otherwise run the project's own build/test command 2. **Run all affected tests**. Include the tests you wrote and any tests your changes could impact -3. **Self-review (a floor, not the final verdict)**. The producing context converges on approval, so this catches slips but does not render the outcome verdict (step 5 hands to `/verification:confirm`, which renders it from outside the producing loop). Read through changes (`git diff HEAD~N`) looking for: +3. **Self-review (a floor, not the final verdict)**. The producing context converges on approval, so this catches slips but does not render the outcome verdict (step 6 hands to `/verification:confirm`, which renders it from outside the producing loop). Read through changes (`git diff HEAD~N`) looking for: - Consistency with existing patterns - No debugging artifacts left behind - No commented-out code - No TODO comments that should be actual work -4. **Rubber-duck advisor checkpoint (HIGH/CRITICAL only)**. For changes involving concurrency, security, cross-platform behavior, external API integration, or with significant divergence from the original plan, call the `advisor` tool (when available in the session) for a quick cross-model critique pass before the review gate. Skip for trivial changes -5. **Hand off to the pre-PR sequence**. Hand off, do not re-order: that sequence owns the step order (invoke `/session-flow:workflow pre-pr` via the Skill tool when the `session-flow` plugin is installed to read it; otherwise follow the consuming setup's own pre-PR checklist). Its order puts **review before outcome verification**, because the simplify pass sits between them and outcome verification must judge the code that ships. So: suggest the project's review flow first (`/review:quality-gate` when the `review` plugin is installed; otherwise the consuming setup's review step), then `/verification:confirm` for outcome verification once the diff is final (when the `verification` plugin is installed; otherwise self-verify the outcome against the plan/intent directly), then the PR (`/source-control:pull-request` when that plugin is installed; otherwise whatever the consuming setup provides. The user controls timing). Do not commit-and-push unilaterally, final staging and PR creation belong to that flow +4. **Deviation fold-back**. When a `DEVIATIONS.md` exists for this work (the non-interactive fork wrote one, or the session opted in per Step 3), read it now and emit one plan-amendment bullet per unresolved deviation or human-decision entry: what the plan should say next time, or what still needs a person. The log is the run's memory; a completion that never reads it back hands the PR reviewer deviations the author already knew about. Fold the bullets into the phase-boundary plan updates (Step 4 ritual) or the handoff summary +5. **Rubber-duck advisor checkpoint (HIGH/CRITICAL only)**. For changes involving concurrency, security, cross-platform behavior, external API integration, or with significant divergence from the original plan, call the `advisor` tool (when available in the session) for a quick cross-model critique pass before the review gate. Skip for trivial changes +6. **Hand off to the pre-PR sequence**. Hand off, do not re-order: that sequence owns the step order (invoke `/session-flow:workflow pre-pr` via the Skill tool when the `session-flow` plugin is installed to read it; otherwise follow the consuming setup's own pre-PR checklist). Its order puts **review before outcome verification**, because the simplify pass sits between them and outcome verification must judge the code that ships. So: suggest the project's review flow first (`/review:quality-gate` when the `review` plugin is installed; otherwise the consuming setup's review step), then `/verification:confirm` for outcome verification once the diff is final (when the `verification` plugin is installed; otherwise self-verify the outcome against the plan/intent directly), then the PR (`/source-control:pull-request` when that plugin is installed; otherwise whatever the consuming setup provides. The user controls timing). Do not commit-and-push unilaterally, final staging and PR creation belong to that flow ## Skill chaining during execution diff --git a/plugins/implementation/skills/implement/evals/evals.json b/plugins/implementation/skills/implement/evals/evals.json index beca6e706..155a37434 100644 --- a/plugins/implementation/skills/implement/evals/evals.json +++ b/plugins/implementation/skills/implement/evals/evals.json @@ -37,7 +37,8 @@ "Classifies the wrong SDK-capability assumption as major divergence, not a minor inline fixup", "STOPS writing code instead of pushing through with workarounds/hacks to force the original plan to fit", "Runs external research for alternatives before re-planning, then routes back to the planning skill for the user to approve the new direction", - "Does NOT silently expand scope or improvise a new design without surfacing it to the user" + "Does NOT silently expand scope or improvise a new design without surfacing it to the user", + "When a DEVIATIONS.md exists for the work, completion reads it back and emits a plan-amendment bullet per unresolved deviation or human-decision entry rather than handing the reviewer a log nobody folded back" ] }, { From ccb814d9da5721636f7dcd36914d853d0a831cf1 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 19:12:46 +0000 Subject: [PATCH 15/22] docs(topics): record the close-out verification results (Phase 11, partial) The dated close-out record: both criterion-1 sweeps (48/48 corpus rows, 45/45 sheet-row greps), the criteria 2-7 walk, and the three follow-up issues (#3589 eval candidates, #3590 E4 pitch-view deferral, #3591 recorded triggers). The phase tag stays open until the full affected-tests battery, still running, reports green; the PR follows it. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- .../finding-your-unknowns-integration/PLAN.md | 28 +++++++++++++++++++ 1 file changed, 28 insertions(+) diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index d9c99e2ce..dcd4789fd 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -375,6 +375,34 @@ all five gate commands above exit 0; issue URLs recorded in this PLAN; PR URL re the criterion-1 double sweep pasted as a dated note with zero unaccounted rows; criteria 2-7 walk recorded. +#### Close-out record (2026-09-01) + +- **Criterion 1 (traceability, both directions):** sweep A iterated all 48 V-id rows of + ./disposition-ledger.md — every row carries a disposition and target (zero blank). Sweep + B grep-verified all 45 executed sheet rows against their landed hunks — 45/45 PASS + (script: criterion1_sweep.sh, session scratchpad; output in the session record). Zero + unaccounted rows in either direction. +- **Criterion 2 (gates):** per-phase gates ran green before each commit; + check-changed-skills.sh origin/main = 53 skills, 0 failed; check-purged-em-dashes.sh + clean; markdownlint 0 issues and typos clean across all 34 changed markdown files; + `node scripts/generate-cheatsheet.mjs --check` exit 0 (frontmatter argument-hint does not + feed the sheet, so no regeneration was needed); ai-slop detector 0 findings across the + changed set, rubric pass on the F1 doc clean. +- **Criterion 3:** F1 carries the codification warning byte-faithfully with citation stamp + and the C6 fair-quotation basis (grep-verified in sweep B: C3-warn, C3-basis). +- **Criterion 4:** two names-and-points registry rows landed pointing at F1's owner + sections, which carry the conformance surfaces (the Phase-2 reconciliation: the + two-column registry never restates; criterion satisfied by owner-section text). +- **Criterion 5:** ctx-eng topic PLAN carries the three "Open, new" candidate bullets and + the Phase-10 rebase note; no phase headings touched (grep-verified). +- **Criterion 6:** every W2/W3 contract delta's eval expectations landed in the same + commit as its contract lines (verifiable via `git log --name-only` per phase commit). +- **Criterion 7:** this PLAN + ./signoff-sheet.md were the executed contract, phase by + phase. +- **Follow-up issues filed:** #3589 (behavioral eval candidates), #3590 (E4 prd pitch-view + extension), #3591 (recorded triggers: D28 schema, Q11 flip, E1 registry row, V7.7 + digest-pipeline lessons). + ## Blast radius MEDIUM-HIGH. Six plugins' skill contracts change in one PR plus two governance docs and a From 72a5929c04f69168a79a5ecf3022273d4e5baaf5 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 19:30:12 +0000 Subject: [PATCH 16/22] fix(planning): satisfy the interview-defenses ratchet for the D28 addition The digest suite caught two things the D28 change touched: the free-text-flag paragraph lands inside the digested open-question-register section of context/loop.md, and the new eval expectation sat in case 12, whose unrelated-reply scenario never has an answer arrive, making the criterion ungradeable there. Per the suite's own contract: both defenses re-read and confirmed intact (the flag adds scrutiny on answered rows and qualifies neither the ask-time write rule nor the gap/blocker register bindings), the expectation moved to case 1 where answers arrive, and the register-section and case-1 digests refreshed in the same change. Suite green (89/0); check-open-questions tests and check-changed-skills green; the battery's 13 other-ecosystem python suites run from their own lane, 785 passed + 330 subtests. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- docs/topics/finding-your-unknowns-integration/PLAN.md | 10 +++++++++- plugins/planning/CHANGELOG.md | 10 ++++++++++ plugins/planning/skills/interview/evals/evals.json | 4 ++-- plugins/planning/tests/interview-defenses.test.sh | 4 ++-- 4 files changed, 23 insertions(+), 5 deletions(-) diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index dcd4789fd..eb91935d7 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -345,7 +345,7 @@ the trigger note in the implementation plugin; both plugins' CHANGELOGs updated; `bash scripts/check-changed-skills.sh origin/main` green; `bash scripts/check-purged-em-dashes.sh` green. -### Phase 11: Close-out — issues, cheat-sheet, acceptance verification, PR [TODO] +### Phase 11: Close-out — issues, cheat-sheet, acceptance verification, PR [DOING] - File follow-up GitHub issues (authorized): one for the behavioral eval candidates (D5, D7, D11, D17b, D34-residue), one for the E4 skill extension candidate, one for the @@ -402,6 +402,14 @@ the criterion-1 double sweep pasted as a dated note with zero unaccounted rows; - **Follow-up issues filed:** #3589 (behavioral eval candidates), #3590 (E4 prd pitch-view extension), #3591 (recorded triggers: D28 schema, Q11 flip, E1 registry row, V7.7 digest-pipeline lessons). +- **Full battery:** `scripts/affected-tests.sh --run` selected the wide suite set (the + plugin.json edits fan out); 2569 assertions passed, one suite failed — the + `interview-defenses` digest ratchet caught D28's paragraph inside its digested register + section and the eval expectation in a case where it was ungradeable. Resolution per that + suite's own contract: defenses re-read and confirmed intact, the expectation moved to + case 1, digests refreshed in the same change; suite re-run green (89/0). The 13 + NOT-RUN-ecosystem suites were then run from their own lane: 785 passed + 330 subtests + (the one PowerShell suite is Windows-lane, CI covers it). ## Blast radius diff --git a/plugins/planning/CHANGELOG.md b/plugins/planning/CHANGELOG.md index 62c0e541d..671b53ff6 100644 --- a/plugins/planning/CHANGELOG.md +++ b/plugins/planning/CHANGELOG.md @@ -28,6 +28,16 @@ All notable changes to the `planning` plugin are documented here. Format follows citable line each (D9; Q7), both pointing at the marketplace repository's `docs/FINDING-YOUR-UNKNOWNS.md`. +### Changed + +- **`interview-defenses` digests refreshed for the D28 addition.** The free-text-flag paragraph + lands inside the digested open-question-register section of `context/loop.md`, and its eval + expectation moved from case 12 (no answer arrives in that scenario, so the criterion was + ungradeable there) to case 1 (answers arrive). Both defenses re-read and confirmed intact: + the flag adds scrutiny on answered rows and qualifies neither the ask-time write rule nor + the gap/blocker register bindings. Register-section and case-1 digests updated in the same + change, per the suite's own contract. + ## [0.34.15] ### Changed diff --git a/plugins/planning/skills/interview/evals/evals.json b/plugins/planning/skills/interview/evals/evals.json index fa9e31cbd..170c28b6e 100644 --- a/plugins/planning/skills/interview/evals/evals.json +++ b/plugins/planning/skills/interview/evals/evals.json @@ -11,7 +11,8 @@ "Output asks the round's independent frontier questions together as one numbered set, not spread across separate turns", "Each question leads with a recommended answer and a one-line basis", "The round is asked inline in prose, not via a side-by-side AskUserQuestion card", - "Where a question is answerable from the codebase, the skill resolves it by inspection instead of spending a question on it" + "Where a question is answerable from the codebase, the skill resolves it by inspection instead of spending a question on it", + "An answer arriving as free text rather than an authored option is recorded with the free-text: resolution-field flag so downstream passes scrutinize it" ] }, { @@ -155,7 +156,6 @@ "expectations": [ "A register row per asked question exists at ask-time, before any reply arrives — not written only once an answer lands", "The unrelated reply does not resolve any open question; every unaddressed row stays `open`", - "An answer arriving as free text rather than an authored option is recorded with the free-text: resolution-field flag so downstream passes scrutinize it", "The next response restates the still-open questions in one line rather than continuing as though they were answered", "The skill does not lock the contract or hand off while a register row is still `open`" ] diff --git a/plugins/planning/tests/interview-defenses.test.sh b/plugins/planning/tests/interview-defenses.test.sh index ac1efe470..db1137a1c 100755 --- a/plugins/planning/tests/interview-defenses.test.sh +++ b/plugins/planning/tests/interview-defenses.test.sh @@ -483,7 +483,7 @@ pin_section "loop.md open-question register section is unchanged (it binds gaps "$LOOP" \ "## The open-question register" \ "## Step 3 — Recognize the stop condition" \ - "66dd3352d4ad6152aa6237b92f1cc1abc31f2cdd5ad5a517adb859408f52f26e" + "6b3deb652af3048e961b5780e62d7d6fba8ca6edc015c6612b3931844d2cf380" # loop.md carries TWINS of two SKILL.md lines that are byte-pinned there: the # confirmation-gate exemption ("`lock` is exempt … its STOP-on-gap rule still applies") in # Step 3, and the `USER-RESERVED` arbiter guidance in Step 4. A twin with no pin is a @@ -577,7 +577,7 @@ pin_case_digest "case 12 still refuses to read drift as consent" \ "e45fe64a00cf78e6dd089837e81c52d8a0947ccf924be6472c6f47eeadcd5942" pin_case_digest "case 1 still resolves codebase-answerable questions without asking, and only those" \ "relentless-me-mode-frontier-rounds" \ - "a1d108f02272c18494b5ca0a0c108385b776ff63bc423cb075c6e052810abc4a" + "cddd012ee8334177839d116ae564a61a5158f3b9ee996c5903c4fc820b973a36" pin_file "case A fixture: the task context still plants the open decision" \ "$FIXTURES/lock-stop-on-gap/task-context.md" \ From 617d5ed82eeebfc127bab417045bcac5341b7058 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 19:31:15 +0000 Subject: [PATCH 17/22] docs(topics): close out the integration plan (Phase 11 done, PR recorded) All eleven phases DONE; PR #3592 recorded in the close-out record. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- docs/topics/finding-your-unknowns-integration/PLAN.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md index eb91935d7..db71a635c 100644 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ b/docs/topics/finding-your-unknowns-integration/PLAN.md @@ -345,7 +345,7 @@ the trigger note in the implementation plugin; both plugins' CHANGELOGs updated; `bash scripts/check-changed-skills.sh origin/main` green; `bash scripts/check-purged-em-dashes.sh` green. -### Phase 11: Close-out — issues, cheat-sheet, acceptance verification, PR [DOING] +### Phase 11: Close-out — issues, cheat-sheet, acceptance verification, PR [DONE] - File follow-up GitHub issues (authorized): one for the behavioral eval candidates (D5, D7, D11, D17b, D34-residue), one for the E4 skill extension candidate, one for the @@ -402,6 +402,8 @@ the criterion-1 double sweep pasted as a dated note with zero unaccounted rows; - **Follow-up issues filed:** #3589 (behavioral eval candidates), #3590 (E4 prd pitch-view extension), #3591 (recorded triggers: D28 schema, Q11 flip, E1 registry row, V7.7 digest-pipeline lessons). +- **PR:** (the single + PR the operator directed; commits map to phases 1-11). - **Full battery:** `scripts/affected-tests.sh --run` selected the wide suite set (the plugin.json edits fan out); 2569 assertions passed, one suite failed — the `interview-defenses` digest ratchet caught D28's paragraph inside its digested register From 770c8299be90cc641596682225e5ad796796323e Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 19:38:42 +0000 Subject: [PATCH 18/22] docs: graduate the integration record to ADR 0025 and prune the contract slice The topic-docs v3 wave completed the contract-tier burn-down on main while this branch was in flight: docs/topics/ is empty there, the prune gate now red-lines any slice left behind, and the context-engineering topic this branch had annotated graduated to ADR 0004. Close-out per the convention: the durable decision record becomes docs/adr/0025 (evidence chain, classification decision, consequences, recorded triggers), the glossary provenance points at the ADR, the branch's own slice (PLAN, signoff-sheet, disposition ledger, delta record, design resolution) prunes as working material preserved in this branch's history and pasted into the PR body, and the resurrected context-engineering PLAN.md is removed; its three candidate inputs re-home to a tracker issue, since acceptance criterion 5's named target no longer exists on main (deviation recorded in the PR). Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- docs/GLOSSARY.md | 6 +- ...-as-judgment-preserving-contract-deltas.md | 60 + .../context-engineering-claude-5/PLAN.md | 1217 ----------------- .../finding-your-unknowns-integration/PLAN.md | 495 ------- .../delta-resolution.md | 88 -- .../design/design-resolution.md | 19 - .../disposition-ledger.md | 97 -- .../signoff-sheet.md | 219 --- 8 files changed, 63 insertions(+), 2138 deletions(-) create mode 100644 docs/adr/0025-adopt-the-unknowns-corpus-as-judgment-preserving-contract-deltas.md delete mode 100644 docs/topics/context-engineering-claude-5/PLAN.md delete mode 100644 docs/topics/finding-your-unknowns-integration/PLAN.md delete mode 100644 docs/topics/finding-your-unknowns-integration/delta-resolution.md delete mode 100644 docs/topics/finding-your-unknowns-integration/design/design-resolution.md delete mode 100644 docs/topics/finding-your-unknowns-integration/disposition-ledger.md delete mode 100644 docs/topics/finding-your-unknowns-integration/signoff-sheet.md diff --git a/docs/GLOSSARY.md b/docs/GLOSSARY.md index 37c962375..c7b9514be 100644 --- a/docs/GLOSSARY.md +++ b/docs/GLOSSARY.md @@ -112,6 +112,6 @@ Materialization of this file was tracked as [#3000](https://github.com/melodic-software/claude-code-plugins/issues/3000). "unknowns quadrants", "blindspot finding types", and the map/territory rejected-terms row were -adopted at the finding-your-unknowns integration sign-off (2026-09-01); the decision rows are in -[`topics/finding-your-unknowns-integration/signoff-sheet.md`](topics/finding-your-unknowns-integration/signoff-sheet.md) -Part E. +adopted at the finding-your-unknowns integration sign-off (2026-09-01); the decision record is +[ADR 0025](adr/0025-adopt-the-unknowns-corpus-as-judgment-preserving-contract-deltas.md) and the +shipping PR carries the full decision sheet. diff --git a/docs/adr/0025-adopt-the-unknowns-corpus-as-judgment-preserving-contract-deltas.md b/docs/adr/0025-adopt-the-unknowns-corpus-as-judgment-preserving-contract-deltas.md new file mode 100644 index 000000000..58a3b4c4b --- /dev/null +++ b/docs/adr/0025-adopt-the-unknowns-corpus-as-judgment-preserving-contract-deltas.md @@ -0,0 +1,60 @@ +# Adopt the "Finding Your Unknowns" corpus as judgment-preserving contract deltas + +- Status: accepted +- Date: 2026-09-01 + +## Context + +A practitioner corpus on artifact-first development — Thariq Shihipar's "A field guide to +Claude Fable 5: Finding your unknowns" (Anthropic blog, 2026-07-06), its X-article +methodology substrate "The Unreasonable Effectiveness of HTML", and a 20-demo example +collection — was ingested as 17 verified digest slices (byte-exact quoting, dual +verification) and worked through a full decision chain: a relentless interview, a +read-only evidence pass grading every named-skill collision, dual fresh-context +validators, external research grounding seven practice areas in primary sources, +blindspot/brainstorm/devils-advocate passes, and a signed single-sheet decision surface. +The working material lived in the branch's contract slice and prunes with it per the +topic-docs convention; the shipping PR (#3592) carries the full plan and verification +record, and the corpus itself is the primary source a future auditor reads. + +The corpus's own author warns against exactly the move a plugin marketplace is tempted to +make — turning the material into generator skills — and the marketplace's instruction +economy separately requires observed, repeated stumble evidence before any standing +instruction lands. Genuine alternatives existed: adopt the techniques as new skills, +adopt them as standing instructions, or reject codification entirely. + +## Decision + +Absorb the corpus behind a per-row evidence-gate classification, with the author's +anti-premature-codification warning treated as a binding constraint: + +- CONTRACT / POLICY / CONVENTION rows land now as team conventions adopted at the + sign-off, as additive lines in the owning skills' bodies with same-commit eval + expectations, never as generator skills. +- BEHAVIORAL rows never land as standing instructions: they ship as doc lines in + `docs/FINDING-YOUR-UNKNOWNS.md` plus tracked eval candidates (#3589), awaiting + observed-stumble evidence. +- `docs/FINDING-YOUR-UNKNOWNS.md` is the graduated reference and the owner doc for the + reply-affordance and export-button conventions (registry rows point at it, + owner-doc-first); it quotes the warning byte-faithfully under a stated fair-quotation + basis. +- Quiz-as-merge-gate reroutes to `verification:confirm`'s existing gate (one mechanism + per concern); the external-reference port gate is scoped to sources of truth outside + the repo's tree via the corrector method's declared-step-delta seam; the deviation log + ships opt-in with a recorded registry trigger (a second plugin reading `DEVIATIONS.md` + graduates it to an owner doc). +- The corpus's context-engineering companion routes to the incumbent effort recorded in + [ADR 0004](0004-rightsize-instruction-surfaces-by-incumbent-first-arbitration.md) + rather than a parallel lane; its three candidate inputs are tracked on the issue + tracker since that effort's contract slice has graduated. + +## Consequences + +Eight plugins gained contract lines and minor version bumps (discovery, education, +verification, prototype, planning, discipline, session-flow, implementation), each with +evals extended in the same commit. Two conventions are in force with named conformance +surfaces. Deferred sub-decisions carry recorded triggers on the tracker (#3590 buy-in +skill extension behind demand evidence; #3591 register-schema flag, tweak-likelihood +flip, deviation-log registry row, digest-pipeline hardening). Reversal is possible but +priced: each convention names its conformance surfaces, and the eval expectations +outlive any instruction ablation, which is what makes a future deletion round provable. diff --git a/docs/topics/context-engineering-claude-5/PLAN.md b/docs/topics/context-engineering-claude-5/PLAN.md deleted file mode 100644 index ad80f04ad..000000000 --- a/docs/topics/context-engineering-claude-5/PLAN.md +++ /dev/null @@ -1,1217 +0,0 @@ -# Context engineering for Claude 5 — source absorption and rightsizing pass - -## Contents - -- [Brief](#brief) - - [TLDR](#tldr) - - [The design documents](#the-design-documents) - - [Goal](#goal) - - [Settled](#settled) - - [Coordination](#coordination) - - [Constraints](#constraints) - - [Shape — decided](#shape--decided) - - [Open](#open) - - [Acceptance criteria](#acceptance-criteria) -- [Plan](#plan) - - [Standards grounding](#standards-grounding) - - [Phase 1: Finish the fresh-docs sweep [DONE]](#phase-1-finish-the-fresh-docs-sweep-done) - - [Phase 2: Resolve the cross-plugin criteria seam [DONE]](#phase-2-resolve-the-cross-plugin-criteria-seam-done) - - [Phase 2.5: Proportionality gate — which detectors survive [DONE]](#phase-25-proportionality-gate--which-detectors-survive-done) - - [Phase 3: Specify the criteria edits [SPECIFIED — execution moves to Phase 8]](#phase-3-specify-the-criteria-edits-specified--execution-moves-to-phase-8) - - [Phase 4: Define the re-run contract [DONE]](#phase-4-define-the-re-run-contract-done) - - [Phase 5: Map each check to its owning plugin [DONE]](#phase-5-map-each-check-to-its-owning-plugin-done) - - [Phase 6: Design the detectors and the sweep [PARTIALLY SHIPPED]](#phase-6-design-the-detectors-and-the-sweep-partially-shipped--run-contract--suppression-landed-in-audit-pass-detector-catalog-still-open) - - [Gate — re-evaluate before implementation [TODO]](#gate--re-evaluate-before-implementation-todo) - - [Phase 7: Bring the setup-skill corpus to its owner doc [AUDITED]](#phase-7-bring-the-setup-skill-corpus-to-its-owner-doc-audited) - - [Phase 8: Implement the checks in their owning plugins [PARTIALLY SHIPPED]](#phase-8-implement-the-checks-in-their-owning-plugins-partially-shipped--audit-pass-run-contract-shipped-per-check-detectors-still-open) - - [Phase 9: Implement the sweep [PARTIALLY SHIPPED]](#phase-9-implement-the-sweep-partially-shipped--audit-pass-skill--run-contract-shipped-cross-plugin-sweep-wiring-still-open) - - [Phase 10: Reconcile, then run against this repository [TODO]](#phase-10-reconcile-then-run-against-this-repository-todo) - - [Phase 11: Acceptance gate and PR [TODO]](#phase-11-acceptance-gate-and-pr-todo) -- [Test strategy](#test-strategy) -- [Alternatives considered](#alternatives-considered) -- [Risks and mitigations](#risks-and-mitigations) -- [Blast radius](#blast-radius) -- [Open questions](#open-questions) -- [Handoff to implementation](#handoff-to-implementation) - - [User-approval gates](#user-approval-gates) - - [Execution shape](#execution-shape) - - [Mechanical work](#mechanical-work) - -## Brief - -Status: **in progress** — shape decided (a component, not a runbook), seam resolved, proportionality -gate closed. Phases 1, 2, 2.5, and 5 are done; design continues at Phase 3. - -### TLDR - -Absorb "The new rules of context engineering for Claude 5 models" into this marketplace and turn it -into something re-runnable against a git repository, starting with this repo. Most of the source is -already enforced — by `/doctor` and by `claude-config:audit-instructions`. The value is a **sweep -skill** — `/claude-config:audit-pass` — that applies every relevant existing skill in a fixed order, -plus detectors for the gaps nothing covers. - -### The design documents - -- [design/article-sections.md](design/article-sections.md) — the source decomposed into 15 sections, - every paragraph and claim, nothing dropped -- [design/official-corroboration.md](design/official-corroboration.md) — each claim checked against - documentation fetched 2026-07-24, marked confirmed / partly confirmed / `OPINION`-tier -- [design/coverage-matrix.md](design/coverage-matrix.md) — each rule against the incumbent that - already enforces it -- [design/skill-inventory.md](design/skill-inventory.md) — which plugins are instruments of the pass, - which are only targets, and the workload that inventory exposes. Counts are cited by command in - "Standards grounding" rather than transcribed here, because they drift -- [design/proportionality-gate.md](design/proportionality-gate.md) — Phase 2.5's record: the D1–D7 - dispositions with per-row evidence, the escalation and the operator's decision, the re-derived - deliverable shape, the homing map, the `OPINION`-tier policy, and both independent reviews -- [design/seam-resolution.md](design/seam-resolution.md) — Phase 2's record: no shared criteria - artifact, with shape 4 reversed on four verified findings - -### Goal - -A repeatable pass the operator points at a **git repository** that applies the source's rules — the -ones official documentation confirms — through the skills that already own each concern, plus new -detectors where nothing does. First target: this repository. - -**Narrowed from "a repo or folder" on 2026-07-24, deliberately.** Every mechanism the re-run contract -rests on is git-derived: target identity comes from the remote URL, two checkouts of one project are -distinguished by a worktree discriminator, and the headline proof that a run never wrote into its own -scan set is `git status --porcelain` being empty afterward. None of that is defined for a bare -folder. Supporting one would mean inventing a path-hash identity and a before/after file-hash -manifest — real work, unclear demand. A smaller true promise beats a larger false one, and the -mismatch is closed here rather than left standing in two places. - -### Settled - -- **Scope: everything.** Every section of the source is in scope; no rule is dropped for being - inconvenient. Rules with no official backing ship marked `OPINION`-tier rather than omitted. -- **Shape: a component — a sweep skill — not a runbook.** The concerns are already distributed - across `claude-config`, `claude-memory`, `docs-hygiene`, `skill-quality`, and `plugin-quality`, and - the operator's need is that *all of them get applied*, in order, repeatedly, against a named - target. - - **Corrected 2026-07-24. This read "Shape: a runbook", and the reason it gave did not support it.** - Distributed concerns plus "all of them get applied" is an argument for **delegation** — which a - component that delegates satisfies exactly as well as a runbook does. The proportionality gate - asked the question directly and answered it against the runbook framing on evidence: - [design/proportionality-gate.md](design/proportionality-gate.md), "Is the sweep a component or a - runbook a human could run by hand?". Its argument is that the checks are delegated but the *run - semantics* are not, and the run semantics are the product — the derived exclusion set, the - three-scope inventory, finding identity, suppression memory, resumability, an advisory lock, and - one human gate per run. None of that is reachable by invoking the incumbents by hand, which is the - operational definition of a runbook. The gate's answer survived cross-vendor review; the Brief's - line did not, so the Brief is what changes. -- **`discipline:sweep-all` is the structural precedent** — a router that fans out - audit-only lanes and applies corrections in a fixed order. That precedent is about **router - mechanics** and is unaffected by the component/runbook correction: the sweep either extends that - pattern or states why it does not. -- **`/doctor` is the incumbent for the CLAUDE.md half** and improves on Anthropic's cadence. The - sweep invokes or defers to it rather than reimplementing trim and migrate. -- **Conflict review is officially prescribed and only half automated.** The memory doc tells operators - to periodically review CLAUDE.md, nested CLAUDE.md, and `.claude/rules/` for conflicting - instructions. **Corrected 2026-07-25 — this previously read "Nothing does it. This is the strongest - gap", and that was false.** `claude-memory:audit` ships check C6 Consistency, which performs exactly - that review over the memory layer, citing the same official line. What no incumbent does is compare - *across* layers — a skill body against `CLAUDE.md`, which is the article's own headline example — - or reach agent definitions, prompt-type hooks, output styles, and the managed-policy tier. That - remainder is the gap, and it is narrower than this line claimed. -- **`@path` imports are a progressive-disclosure anti-fix.** Imported files load at launch. Only - skills and path-scoped rules defer load. Any "split this up" remediation must name the right - destination. -- **User-scope surfaces are routed, never edited.** `~/.claude/**` is chezmoi-managed; findings - there become recommendations backfilled through the dotfiles repo. - -### Coordination - -A parallel session is revising the `playbooks:fable-5` skill and may update other skills. This work -is deliberately downstream-tolerant: the sweep is what cleans up after that session, so it must -be re-runnable against a moved target rather than assume a frozen tree. - -### Constraints - -- Fresh-docs mandate governs. Every claim traces to a page fetched this session or a file read this - session; unverified claims are labeled. -- Nothing ships that `/doctor` already does. -- The plugin-acceptance gate applies: repo-agnostic, `userConfig`-configurable, plugin-form-safe, no - PII, semver'd, security-reviewed. -- The pass follows its own doctrine — a lightweight guide with progressive disclosure, not another - always-loaded rule wall. - -### Shape — decided - -**Two layers, both native to the Claude-configuration plane.** - -1. **Individual checks land where their concern already lives.** Each detector extends the plugin - that owns its surface — `claude-config`, `claude-memory`, `docs-hygiene`, `skill-quality` — and - becomes a new skill only where no owner exists. Every rule from the source is applied piecemeal - by whichever skill is the right home for it. -2. **A sweep skill in `claude-config` fans them all out against a named target.** Point it at a - repository and it runs the whole body of knowledge in one ordered pass. `claude-config` is the - native home because it already owns the Claude Code configuration plane and because - `audit-instructions` already builds the surface inventory the sweep needs. - -`discipline:sweep-all` is the structural precedent, not the home: it is session-posture -scoped (what discipline is Claude operating under right now), where this is target scoped (what does -this repository's instruction surface look like). Same router mechanics, different subject. - -**Fix-capable, not report-only.** The operating goal is repeated application to a repository, not a -findings report to read. `audit-instructions` stays report-only on its own; the sweep applies. The -human gate moves from per-finding to per-run. - -**Re-runnable by design.** Success is hammering the same repository repeatedly and watching the -finding set shrink. That requires the knowledge the sweep applies to be anchored in one versioned -source with explicit staleness triggers — not restated across the skills that consume it. - -### Open - -1. ~~Naming.~~ **Closed** — the sweep is `audit-pass`; see [design/checks-and-sweep.md](design/checks-and-sweep.md), "Naming — resolved". -2. ~~Which detectors extend an existing skill versus become new ones.~~ **Closed** — the homing map - is discharged in [design/proportionality-gate.md](design/proportionality-gate.md). Two host - plugins: `claude-config` and `claude-memory`. - -### Acceptance criteria - -- Pointed at this repository, the pass applies every check the inventory names in one ordered run, - with no step silently skipped and no surface silently excluded — in-repo worktrees excepted, and - named when dropped. -- Every source section maps to either a check that runs, an incumbent that already covers it, or an - explicit recorded exclusion — traceable back to - [design/article-sections.md](design/article-sections.md). -- A rerun over a **comparable** pair produces the same **derived-tier** result, never a different - one. Comparability is the design's own list and is cited rather than restated here, because it has - already grown twice; an unchanged tree is only its first and weakest input. A rerun after accepted - fixes drops every accepted finding **from whichever tier reported it** — not - necessarily a smaller derived set, because fixing a judged-only finding such as D1 legitimately - leaves the inventory, the exclusion set, shadowed definitions, and the raw candidate rows - identical, so a strict-shrink gate would fail the ordinary success case. The derived set grows - again only when a comparability input moved — a **detection-version** bump among them, which is - the catalog version, the host plugin's semver, and a - digest over the prompt text and scripts that actually decide detection, because the catalog version - alone does not cover `SKILL.md`'s Phase B and Phase C prompts. A skill authored between runs - legitimately grows it, and the tree is moving under a parallel session. That is the acceptance test - for "regular audit"; - [design/rerun-contract.md](design/rerun-contract.md) specifies it, including the finding-identity - function that makes "the same set" diffable. - - **The tier vocabulary changed, and this criterion changed with it.** It originally scoped the - diff-clean gate to the `mechanical` tier. Verified against the implementations, that tier cannot - carry it: `audit-instructions` says its deterministic pre-scan "is advisory… so the lane refines - every candidate rather than reporting it verbatim", and its Phase C re-judges *every* proposal, so - no dispatched check reaches the report without model judgment — including `mechanical`-tagged ones. - And `claude-memory`'s criteria carries 17 checks with **zero** occurrences of `mechanical` or - `behavioral`, so half the dispatched catalog was never in the vocabulary at all. The **derived** - tier — the three-scope inventory, the exclusion set, shadowed-definition findings, and raw script - candidate rows — is what is genuinely model-free, and it is where the exact-equality gate now sits. -- **Judged findings are held to a stability property, not to identity.** Detection there is a model - judgement, so two runs over an identical tree may legitimately differ; a finding-identity function - normalizes how a finding is *reported* and cannot make the *detection* deterministic. Judged - findings are reported separately, excluded from the diff-clean gate, and held instead to: none - contradicts an accepted suppression, and the **symmetric difference** between the two judged sets - stays within a stated tolerance over a comparable pair **whose judging configuration is also - equal** — **whose violation fails the run's self-check** rather than being absorbed by - recalibrating the tolerance. The metric counts removals as well as additions, because a check that - silently stops firing is the more damaging direction; and the judging configuration is a - precondition rather than an obligation, since a model swap moves the judged set legitimately. -- **A surface that silently leaves the inventory fails the gate.** The inventory and exclusion set are - *in* the derived tier, not scaffolding beneath it, so a scope regression is caught. This is the - property the original two-tier split could not express, and it matters more than a changed finding: - a shrinking scope looks like an improving report. - -## Plan - -**Scale:** large — cross-plugin, new contract surface, 60-plugin blast radius. -**Design tier:** A, sequenced into Phases 2–6 behind a gate, with a proportionality gate at Phase -2.5 that can re-derive the deliverable's shape — see -[design/design-resolution.md](design/design-resolution.md). - -### Standards grounding - -| Surface | Sections cited | Provenance | -|---|---|---| -| Plugin design | `docs/PLUGIN-PHILOSOPHY.md` — "Design boundary", "Naming", "Native-first", "Component stances", "Setup is explicit and repeatable", "Prerequisites and failure behavior", "Convention registry", "Cross-platform contract", "Fresh-eyes checkpoints" | Repo-owned | -| Catalog placement | `docs/CATALOG-TAXONOMY.md` — "Form rule", "Assignment principle", "Vocabulary" | Repo-owned | -| Repo operating rules | `CLAUDE.md` — fresh-docs mandate, design rules for plugins, branching and PRs | Repo-owned | -| Topic-doc lifecycle | `docs/conventions/topic-docs/` — tier placement and the prune-before-merge rule that governs these very documents | Repo-owned | -| Presence-gated fallbacks | `docs/conventions/seam-phrasing/` — the phrasing every optional cross-plugin invocation uses | Repo-owned | -| Consumer configuration | `docs/conventions/config-cascade/`, `docs/conventions/consumer-config-layering/` — load-bearing if Phase 2 chooses the consumer-artifact shape | Repo-owned, conditional | -| Acceptance gate | `docs/MIGRATION-PLAYBOOK.md` — per-plugin migration gate, plugin-acceptance security review | Repo-owned, read at Phase 11 | - -Four grounding findings reshape the plan and are carried as constraints throughout: - -- **Horizontal decoupling.** A plugin never imports a sibling's files and an installed plugin reaches - only inside `${CLAUDE_PLUGIN_ROOT}` — no `../`. A shared criteria catalog cannot simply be read - across plugin boundaries, and a repo-level doc is unreachable from an installed plugin's cache. - Phase 2 exists to resolve that seam. -- **Fixed verb meanings.** `audit` and `scan` are read-only, mutating only behind an explicit - override; `clean` / `tidy` / `fix` mutate. The fix-capable posture must be expressed through one of - those two shapes, not asserted. -- **The convention registry is a gate, not a courtesy.** "A new cross-plugin convention lands in an - owner doc before a second plugin adopts it. Fleet audits check conformance per row." This was - carried as a constraint because a shared catalog would have been such a convention. **It no longer - binds this work:** Phase 2 chose no shared artifact, so nothing here is a new cross-plugin - convention and no registry row is owed. The gate itself stands — and ground-truth verification - found it is currently unheld in practice, with 17 rows covering only 2 of the repository's 5 - materialization mechanisms, which is a finding about the repository rather than about this work. -- **Native-first is a gate on every customization surface.** `InstructionsLoaded`, `/context`, and - `claudeMdExcludes` are already recorded in `design/official-corroboration.md` as relevant official - mechanisms. Each is adopted, rejected with a reason, or deferred with a trigger — never bypassed by - building a filesystem walk instead. - -Counts are cited by command rather than transcribed, because they drift and the tree owns them: -`ls plugins/*/skills/*/SKILL.md | wc -l` (181 top-level skills at time of writing), -`find plugins -name SKILL.md | wc -l` (187, the extra six being vendored upstream materializations). - -### Phase 1: Finish the fresh-docs sweep [DONE] - -Task #18. Gates every subsequent phase: the remaining checks' premises are unverified until the -pages land. Fetch and record in `design/official-corroboration.md`, each with its URL: skills -(progressive disclosure inside a skill, subagent execution, dynamic context injection), plugins and -plugins-reference (manifest schema, `${CLAUDE_PLUGIN_ROOT}` / `${CLAUDE_PLUGIN_DATA}`, `userConfig`), -hooks (`InstructionsLoaded`, prompt-type hook text as an instruction surface), sub-agents (subagent -memory, startup loading), tools-reference (deferred tool loading via `ToolSearch` — the §S7 claim), -claude-directory (surface enumeration). - -**Four further pages are load-bearing and were missing from this list.** Each governs a surface the -pass audits or a native mechanism the Phase 6 gate must rule on: - -- `debug-your-config` — "diagnose why `CLAUDE.md` or settings aren't taking effect", per the memory - doc's Related resources. A native diagnostic aimed at this exact subject, so it is a native-first - candidate for the inventory alongside `InstructionsLoaded` and `/context`, not an optional read. -- `large-codebases` — the memory doc defers to it for "the full layout of root and per-directory - `CLAUDE.md` files and rules", which is the surface partition D1 compares across. -- `settings` — the settings layers `claudeMdExcludes` merges across, the `claudeMd` managed key, and - `disableBundledSkills` / `skillOverrides`, which decide whether `/doctor` is present at all. -- `context-window` — where `CLAUDE.md` sits in startup context, and what survives compaction. - -- **Sanity Check:** for each of the fourteen page slugs (`skills`, `plugins`, `plugins-reference`, - `hooks`, `sub-agents`, `tools-reference`, `claude-directory`, `debug-your-config`, - `large-codebases`, `settings`, `context-window`, `features-overview`, `output-styles`, `mcp`), - `rg -c "docs/en/>" design/official-corroboration.md` returns ≥ 1. Counting `https://` lines - does not work — a line may carry two URLs. The trailing `>` closing the autolink is required: a - bare `docs/en/plugins` also matches `docs/en/plugins-reference`, and `docs/en/mcp` also matches - `docs/en/mcp-quickstart`, so either slug would pass without its own page ever being fetched. -- **Sanity Check — the list itself is falsifiable.** The check above can only prove the slugs it - already names were fetched; it cannot notice a page nobody listed. Before closing the phase, walk - `https://code.claude.com/docs/llms.txt` and record every page whose subject is an instruction, - memory, or configuration surface as fetched or explicitly out of scope with a reason. - -**Outcome.** All fourteen slugs are recorded in `design/official-corroboration.md` and both sanity -checks pass. The list started at eleven; the `llms.txt` walk covered all 172 listed pages and -surfaced three more — `features-overview`, `output-styles`, and `mcp` — which were fetched rather -than deferred, and the check was widened to the fourteen it now names. Two findings change what -later phases are built against, and both are carried into Phase 2.5 rather than resolved here: - -- **`features-overview` already prescribes D7's routing.** Its "Compare similar features" section is - official guidance on choosing between `CLAUDE.md`, `.claude/rules/`, and skills, including the - 200-line rule and the enforcement boundary between an instruction and a hook. D7 must show what it - detects beyond restating that page. -- **Output styles are an unenumerated instruction surface.** They modify the system prompt directly, - default to *removing* Claude Code's built-in software-engineering instructions, ship from plugins in - an `output-styles/` directory, and can override the operator's selection via `force-for-plugin`. - D1's surface partition is incomplete without them. - -### Phase 2: Resolve the cross-plugin criteria seam [DONE] - -Review: architecture — satisfied by a cross-vendor review and a blind derivation, both recorded in -[design/proportionality-gate.md](design/proportionality-gate.md). - -**Ran after Phase 2.5, not before it.** [design/design-resolution.md](design/design-resolution.md) -says the Tier A classification and Phase 2's own existence rest on "a versioned criteria catalog -consumed by more than one plugin". Phase 2.5 decides whether there is one, so deciding the seam -first would have decided it against a premise the gate deletes. Information flows one way. - -**Outcome: no shared criteria artifact.** Each plugin owns its criteria outright; cooperation is a -presence-gated namespaced skill invocation with a documented standalone fallback — the seam these -two plugins already run in both directions. Shape 4 was the starting position and is **reversed**; -Shapes 1 and 2 are rejected with reasons. Full record, including the drift risk the choice accepts: -[design/seam-resolution.md](design/seam-resolution.md). - -Four verified findings reversed shape 4, and every one of them was unavailable when it was proposed: - -- **Its cited CI guarantee never fires for this artifact.** `check-cross-plugin-source-drift.sh` - clusters on the full path-within-plugin, and a criteria catalog lives at - `skills//reference/criteria.md` where the skill name differs by construction. Four - `criteria.md` files exist today at four distinct paths and form zero clusters; `--check` exits 0. - This plan's own claim that a byte-identical copy trips the check as an unregistered cluster is - **false** — the skip-list argument was necessary but not sufficient. -- **Relocation breaks a currently-green gate**, and then stops watching: `check-skill.sh` - existence-checks skill-internal refs, so the move emits `broken skill-internal ref`, while the - post-move `](../../reference/criteria.md)` form escapes its extractor entirely. -- **Adoption is six to seven registration points**, not one — and one of them, the cluster registry - entry, is unreachable because no cluster can form. -- **The catalog has no frontmatter to bump.** `Version: 1.0.0` is body prose under an H1; the - precedent's bump machinery reads a YAML key. - -The strongest argument is one no shape analysis had: **the corpus contains no instance of one plugin -reading another's `reference/criteria.md`**, and `docs-hygiene:audit-encapsulation` classifies -`reference/` as private surface. A shared catalog would be the first violation of the encapsulation -contract this repository enforces. - -- **Sanity Check — passes.** `rg -c "rejected" design/seam-resolution.md` returns 6 (≥ 3), and the - document names the chosen shape and cites `PLUGIN-PHILOSOPHY.md` "Design boundary" verbatim. -- **Sanity Check — satisfied by construction, re-asserted at Phase 8.** The design introduces no - cross-plugin file read, so there is no surface for `/docs-hygiene:audit-encapsulation` to flag. - Verified by inspection too: every criteria reference in the corpus is same-plugin and relative. - Running the skill against the implemented result remains a Phase 8 gate. - -### Phase 2.5: Proportionality gate — which detectors survive [DONE] - -Ran first, ahead of Phase 2. Full record with per-row evidence, the escalation, the operator's -decision, the re-derived shape, the homing map, the `OPINION`-tier policy, the D1 scope answers, and -both independent reviews: [design/proportionality-gate.md](design/proportionality-gate.md). - -**Outcome.** One officially-backed new check survives (D1, as I12). One further new check ships -`OPINION`-tier and default-off (D3). Everything else is an edit to a check that already exists. -The escalation condition fired and **the operator approved re-deriving the deliverable's shape**: -the cross-plugin catalog, its convention-registry owner doc, and the sync-script materialization are -dropped; the sweep, the re-run contract, and D1 survive. Two host plugins receive rules — -`claude-config` and `claude-memory` — not four. Nothing lands in `docs-hygiene` or `skill-quality`. - -Three findings from the gate that later phases inherit: - -- **The gate's own test had to be split in two.** "Does an incumbent already cover it" and "is it - officially backed" are orthogonal, and the first draft demoted D2 on the second while reporting it - as the first. Every disposition now names which test it fails. -- **`OPINION`-tier enablement inverts for suppressors.** D4 withholds findings rather than emitting - them, so defaulting it off deletes the mitigation for this plan's own High/High "detectors flag - correct constraint as over-constraint" risk and makes trimming strictly more aggressive. -- **D1's detection rule is narrower than "two instructions differ."** Where the official layering - rule already picks a winner — skills, subagents, and MCP servers override by name — differing - instructions are a resolved override, not a conflict. The comparison set is the `CLAUDE.md` - family, hooks, and output styles. - -The original phase text follows, retained because the dispositions were assigned against it. - -The seven detectors are not equal-weight, and the plan's own evidence says so. -`design/coverage-matrix.md` ranks the four gaps and finds only one justifies new surface: S3, -cross-surface instruction conflict, which is `ANTHROPIC-DOCS`-backed (the memory doc prescribes the -review under "Consistency" and repeats it in troubleshooting). **Corrected 2026-07-25: this said "which no incumbent performs".** `claude-memory:audit` C6 performs it over the memory layer; the unowned remainder is the cross-layer and non-memory case. The -matrix files the other three as *"authoring guidance, not an auditable defect"*, *"plausibly a check -added to `skill-quality:check`"*, and *"calibration refinements to incumbents, not standalone -surface"*. `design/official-corroboration.md` separately finds the interface half of S6, the S8 -placement rule, the S13 carve-out, and artifacts-as-references all `OPINION`-tier, unconfirmed by any -fetched page. - -**Corrected — the original argument here was an axis conflation and is struck.** It read "those two -documents were produced from different inputs and agree", and used that agreement as evidence. It is -not evidence. The matrix measures *who enforces a rule*; the corroboration document measures -*whether the rule is true*. A rule can be fully confirmed and still be the largest gap — S3 is -exactly that, and S7 is the mirror image — so the two axes run in opposite directions on the rows -that matter most. Nor are they independent inputs: both descend from `design/article-sections.md`. -The dispositions survive on the per-row evidence recorded in -[design/proportionality-gate.md](design/proportionality-gate.md), not on this argument. - -So D1 is a first-class new detector; D2–D5 are the weak gaps and the `OPINION` set at once. Assign -dispositions accordingly: - -- **D1 is a detector**, and is the deliverable's primary payload. -- **D2–D5 are calibration inputs to their incumbents by default** — a check added to - `skill-quality:check`, a rule fed to `docs-hygiene:extract-ssot`, a suppression input consulted by - the trimming detectors — not standalone components. Promoting one to a detector requires a written - reason that survives the matrix's own verdict on it. -- **D6 and D7 trace to `PARTIAL` remainders the matrix never ranked.** Give each the same - justification test as D2–D5 before it is treated as new surface. -- **`OPINION`-tier content is disabled by default and opt-in.** The existing catalog declares the - `OPINION` authority value but has never used it — all eleven seeds are `ANTHROPIC-DOCS` — so no - consumer has ever had to decide what an `OPINION` finding means. This work is the first to populate - the tier, and therefore owns defining its default enablement, its severity ceiling, and how a - consumer turns it on. Shipping `OPINION` rules enabled would let vendor blog advice that no - documentation confirms mutate a consumer's instruction corpus under the same banner as documented - doctrine. - -**If only D1 survives as a detector, stop and re-derive the deliverable's shape** before Phase 3 -builds a catalog. A versioned cross-plugin catalog, its convention-registry owner doc, a re-run -contract, and a sweep are machinery sized for a multi-detector program; carrying a single detector -plus a set of incumbent refinements is a different, smaller artifact. That re-derivation is a -user-approval gate, not an implementation detail. - -- **Sanity Check — passes, with the vocabulary corrected.** Every D1–D7 carries a disposition and a - reason. The three-value vocabulary proved incomplete: D4 emits nothing, so `suppression input` was - added rather than stretching a bad fit, and the gate says so instead of leaving the up-front claim - standing. The `OPINION` clause is amended for the same reason — no `OPINION` rule that *emits* is - enabled on bare invocation, while a rule that *withholds* must be. -- **Sanity Check — passes.** Exactly one officially-backed detector survived; the decision record - states that the operator approved re-deriving the shape rather than continuing with the full - machinery, on 2026-07-24. - -### Phase 3: Specify the criteria edits [SPECIFIED — execution moves to Phase 8] - -Task #34, **re-derived, and then re-sequenced.** There is no new catalog, no canonical relocation, -and no per-plugin materialization — Phase 2.5 deleted the multi-plugin premise and Phase 2 found the -mechanism would not have worked here anyway. What remains is edits to two files that already exist. - -**Those edits are implementation, and this phase cannot perform them.** The original Phase 3 built a -*design artifact*; the re-derived one edits `plugins/claude-config/**` and `plugins/claude-memory/**`, -which are live plugin files. Phase 8 is where the branch splits, and Phase 11 requires documentation -and implementation to land as separate PRs — so writing them here would put implementation on the -docs branch and break that rule. The re-derivation created this collision by changing what the phase -produces without changing where it sits. - -**Resolution: this phase is discharged as a specification.** The exact per-file edit list below is -the deliverable, and it is executed under Phase 8 on the implementation branch. Nothing is lost and -the phase count does not change; only the commit boundary moves. - -**`plugins/claude-config/skills/audit-instructions/reference/criteria.md`** - -- New check **I12** — D1, cross-surface instruction conflict, scoped per the gate's narrowing - (task #19). -- New locality check beside I3 — D3, on the definition-site axis rather than I3's load-timing axis, - `OPINION`-tier and default-off (task #23). -- **I9's Remediate line extended** with the interface destination — D2, `OPINION`-tier, default-off - (task #22). -- **I3's Remediate line gains a destination-qualifying test** ("a destination qualifies only if it - defers loading — `@path` imports do not") **and a move cost** (task #26, non-memory half). -- **A stopping condition on I6 and I8** — D4, default-**on**, per the suppressor inversion - (task #24). - -**`plugins/claude-memory/skills/audit/reference/criteria.md`** - -- **One consolidated C3 revision** — D7 and D6's memory half are the same rule. Adds an auto-memory - destination row, an `@path` non-deferring row, and a per-destination move cost drawn from the - compaction table the plugin already ships in `reference/official-guidance.md` and no check cites - (task #27). - -Every rule cites its source URL rather than restating doctrine, and carries the catalog's existing -recheck triggers so one staleness event fires all of them. The pre-flight consumer check is **done** -and its result is why the catalog stays put: three parse paths, all bare skill-relative markdown -links in `audit-instructions/SKILL.md`, no script or CI workflow reads the file, and a reverse -`[SKILL.md](../SKILL.md)` back-link in `criteria.md`'s "Output format" section constrains where it -could live at all. - -The convention-registry work item is **dropped**: the registry gates a new cross-plugin convention, -and extending one plugin's own reference file is not one. Task #43 decides separately whether the -version stops being body prose and becomes assertable frontmatter. - -- **Sanity Check — this phase.** Every `S` id in `design/article-sections.md` maps to a named edit - above or to a stated exclusion. That is assertable against the specification alone, and it is what - keeps traceability — an acceptance criterion — from being the thing the smaller shape loses. -- **Sanity Check — carried to Phase 8, where the edits land.** - `scripts/check-catalog-coverage.sh` (new, per the repo's `check-*.sh` idiom) exits 0; - `/skill-quality:check` passes for both modified skills; and - `scripts/check-cross-plugin-source-drift.sh --check` still exits 0, proving no edit accidentally - created a byte-identical cluster. - -### Phase 4: Define the re-run contract [DONE] - -Recorded in [design/rerun-contract.md](design/rerun-contract.md). Every work item below is answered -there as a numbered assertion rather than as prose intent, which the phase's own sanity check -demands and which the gate's argument for the sweep now depends on. - -The decisions that carry the most weight, because a later phase could plausibly have chosen -otherwise: - -- **The anchor is content-derived, never line-derived.** A line number shifts when anything above it - changes, so a line-anchored identity would churn the whole report on an unrelated edit and destroy - the property the contract exists to protect. -- **A run never writes into its own scan set**, and the whole property reduces to one command: - after a run against a clean worktree with no redirect argument, `git status --porcelain` is empty. -- **State is keyed by canonical repository identity plus a worktree discriminator**, so two - worktrees of one repository on different branches do not share a report — and the working - directory is never an input. -- **Read-only runs take no lock; applying runs refuse rather than queue**, because a sweep over a - large tree runs long and a silent queue looks like a hang. -- **Inline suppression is permitted only where the pass may write.** For a file it does not own, a - chezmoi-managed user-scope file, or a registered cluster copy, suppression is central and keyed by - finding id — an inline marker in a cluster copy would break the sync path the exclusion set exists - to protect. -- **The judged-tier tolerance is `max(2, ceil(0.10 × |J(R1)|))`**, with the floor there so a small - judged set cannot collapse the tolerance to zero and reintroduce identity by the back door. - It is a starting calibration, and Phase 10 is what tests it. -- **Two mechanisms inside the contract were proportionality-tested and moved** — - [design/proportionality-gate.md](design/proportionality-gate.md), "The run contract's own - machinery, proportionality-tested per mechanism". The state key's `repo-identity` half is - legibility rather than correctness, and per-lane input digests are deferred with a Phase 10 - trigger against a tree-wide refuse-to-resume check that closes the same hole more cheaply. - -The original work items follow. - -Task #35. Idempotence is the headline acceptance criterion, so it needs a machine-comparable -definition before any detector is designed against it. - -- **First work item — the finding-identity function.** "The same finding set" is undiffable while - findings are prose judgements. Define identity as a tuple (surface path, check id, anchor, - normalized claim) and emit findings to a machine-readable file so two runs can be diffed rather - than compared by reading. -- **Second work item — where the run report lives relative to the scan set.** If run 1 writes its - report into the tree, run 2's tree is not unchanged and the idempotence property is unfalsifiable. -- **Third work item — run-state keying and concurrency.** `${CLAUDE_PLUGIN_DATA}` is machine-global - rather than per-project, and this machine carries dozens of checkouts of this repository — see the - derived counts below, which correct an earlier transcribed figure that stood here too. State must - be keyed by canonical repository identity, not by - working directory, and the concurrent-run posture must be stated (a lock, or documented - last-write-wins). -- **Fourth work item — the suppression surface, per target class.** A deliberately-kept finding must - not resurface, but "its site" differs by class: a `SKILL.md` this pass does not own; a - chezmoi-managed `~/.claude/**` file the Brief says is routed and never edited; a registered - byte-identical cluster copy where an edit breaks the sync path. Decide each before Phase 6 designs - detectors that read the suppression record. -- **Fifth work item — mid-run resumability.** A run over 181 skills can be interrupted by compaction, - a rate limit, or a crash. Findings persist incrementally as collected, and an interrupted run - resumes from the last completed lane rather than restarting. - -Then the properties themselves: a comparable pair yields an identical finding set; accepted fixes -remove each accepted finding from the tier that reported it, which leaves the derived set unchanged -when the fix landed a judged-only finding; judged-tier findings carry the delete-and-watch -follow-through where their host check defines one, or the check's own stated loop where it does not -(D1); the set grows again only when a comparability input moved. - -- **Sanity Check:** `design/rerun-contract.md` states the identity tuple, the report location rule, the - state key, the concurrency posture, the per-class suppression surface, the checkpoint property, and - each idempotence property as a condition a test could assert — not as prose intent. - -### Phase 5: Map each check to its owning plugin [DONE] - -Task #33, discharged inside [design/proportionality-gate.md](design/proportionality-gate.md) rather -than as a separate artifact — it would have been a seven-row table restating that document's own. - -Two corrections the mapping forced, both from reading the incumbents' bodies instead of their -listing descriptions: - -- **The D4 carve-out needs no seam.** It was expected to be consulted by trimming rules in three - plugins. Every `docs-hygiene` trimmer already owns a stopping condition shaped to its own content - model — semantic-loss revert, always-admitted categories, fact ownership, reasoning-stays-inline — - and `skill-quality`'s skills remove no content. The gap is `claude-config`'s I6 and I8 alone. -- **`skill-quality:check` hosts nothing.** Its contract is "NO model invocation… reproducible in CI - or a pre-commit hook", and `argument-hint` is read by nothing in the plugin. A - representational-equivalence judgement would be the first non-reproducible check in a gate whose - value is that every check is reproducible. **Re-verified 2026-07-24 against open PR #1096, which - adds a twenty-first check** — the contract survives it, and the discriminator that keeps this - conclusion true is stated in [design/proportionality-gate.md](design/proportionality-gate.md), - "The determinism contract, re-verified against check 21". - -- **Sanity Check — passes.** One row per D1–D7 with its disposition; both named owning plugins exist - under `plugins/`; every row names a skill directory that already exists; the - `deferred-with-trigger` row names its trigger; no row reads "TBD". - -### Phase 6: Design the detectors and the sweep [PARTIALLY SHIPPED — run contract + suppression landed in `audit-pass`; detector catalog still open] - -Review: architecture - -Tasks #19 and #22–#27 (now two new checks plus four edits to existing ones, per Phase 2.5) and #28 -(the sweep). Each check needs a false-positive story before it ships. Naming resolves here via -`/naming:name-it-better`, constrained by the fixed verb meanings. - -**In progress — the design is in [design/checks-and-sweep.md](design/checks-and-sweep.md).** D1's -detection rule, its five must-not-flag cases, its remediation split by scope, the native-first -inventory ruling, the sweep's posture, its derived exclusion set, the `/doctor` prerequisite -contract, the dispatch order, and the `OPINION` discovery line are all written. **The sweep is named -`audit-pass`** — operator's choice from a 32-candidate tournament, with what -that name pays recorded at the naming site so it is not re-litigated. The suppression record is -resolved as two artifacts: an owner doc under `docs/conventions/` declaring the keys, and the -instance at `.claude/audit-pass.md`, keyed per finding id with the team layer winning conflicts. - -**The report schema and the lane decomposition stood open here and are now settled**, in -[design/checks-and-sweep.md](design/checks-and-sweep.md), "The report and the lanes": two files, an -append-only `findings.partial.jsonl` written per completed lane and a `findings.json` assembled at -the end whose sections are the three tiers plus `suppressed` and `skipped`; and lanes keyed as -**(check × surface class)**, which is the granularity the exclusion set and the three-scope inventory -already work at. - -**One correction that phase drafting forced back upstream.** The first statement of D1's scope -excluded skills, subagents, and MCP servers wholesale because they "override by name". That misreads -the rule — override-by-name resolves a collision between two entities *sharing a name*, where one is -simply inert. It says nothing about a skill body contradicting `CLAUDE.md`, which is the source -article's own headline example. Excluding skill bodies would have excluded the failure D1 exists to -detect. The exclusion is now narrow: a shadowed same-named definition is not a conflict; everything -else that holds instruction text is in the comparison set. - -**The sweep is the phase's centre of gravity now, not the detectors.** With one new -officially-backed check, the design work that carries risk is the run contract — the derived -exclusion set, the three-scope inventory, finding identity, suppression, resumability, and the apply -posture. That is also the argument for the sweep existing at all: the checks are delegated, the run -semantics are not, and invoking the incumbents by hand yields none of them. Recorded in -[design/proportionality-gate.md](design/proportionality-gate.md). - -**The exclusion set is derived, never hardcoded.** Three classes a fix-capable pass would corrupt, -all verified present: - -- **Registered byte-identical clusters** — `scripts/cross-plugin-source-registry.txt` registers - `hooks/hook-utils.sh` (13 live copies), `reference/artifact-protocol.md`, and - `reference/standards-contract.md`, each guarded by a dedicated CI drift check. A trim or compress on - any copy breaks the sync path and reds CI. -- **Vendored upstream materializations** — six `SKILL.md` files under `plugins/*/skills/*/vendor/`. - Hand-editing an upstream copy is a standing prohibition. -- **Worktrees** — derive from `git worktree list` plus gitignore-awareness; a git-tracked enumeration - excludes them for free where a filesystem walk does not. - - **The narrative claim here was wrong and is now checked rather than asserted.** It read "three - exist under `.claude/worktrees/`, which is gitignored at `.gitignore:15`". The `.gitignore:15` - half is correct; the count and the existence are not, and **existence is root-dependent, which is - why this correction itself needed correcting on 2026-07-25**: from a sibling task worktree - `.claude/worktrees/` does not exist, while the primary checkout holds exactly **one** nested - worktree there — `.claude/worktrees/ignition-rebind-note`, whose own copy of the tree accounts for - the entire `SKILL.md` remainder in - [design/skill-inventory.md](design/skill-inventory.md). `git worktree list` reports **50** entries, - the primary checkout plus 49 linked worktrees, nearly all in directories outside the repository. - The derivation mechanism is robust to every one of those variations — which is the point of - deriving rather than transcribing — but a plan - that audits other people's instruction files for stale unverified counts had a stale unverified - count of its own, twice, in the paragraph arguing for derivation. Counts here are cited by command - or not at all. - -**The inventory clears the native-first gate first.** Adopt, reject with a reason, or defer with a -trigger: the `InstructionsLoaded` hook ("log exactly which instruction files are loaded, when they -load, and why"), `/context` for confirming what actually loaded, and `claudeMdExcludes` as a -remediation option. Building a filesystem walk without recording that decision fails the gate. - -**`/doctor` gets a prerequisite contract, not a hand-wave.** Its trim requires Claude Code v2.1.206 or -later, and it "reports findings first and asks for confirmation before changing anything" — so it is -interactive and cannot be driven by an unattended sweep. State the version floor, classify absence -(required-for-correctness versus optional-feature), and make the handoff an operator instruction -rather than a dispatch. - -**`/doctor` is a bundled skill, and a bundled skill can be absent.** The skills doc: `/doctor` is -prompt-based rather than fixed logic, and "before v2.1.205, `/doctor` was a built-in command rather -than a bundled skill." `disableBundledSkills` spares it — but the doc's own escape hatch is "to hide it, set the -`DISABLE_DOCTOR_COMMAND` environment variable or a `skillOverrides` entry of `"doctor": "off"`". Two consequences the plan has to absorb, -because it deliberately builds nothing on `/doctor`'s half of the surface: - -- **Presence, not just version, is the prerequisite.** On a consumer machine where `/doctor` is - turned off, "nothing ships that `/doctor` already does" leaves the entire `CLAUDE.md` trim and - migrate half with no incumbent *and* no replacement. The design boundary's "report the missing - optional capability clearly" is the floor: the sweep names `/doctor` as the missing capability and - states what goes unchecked. Deciding whether to build a fallback beyond that is a Phase 6 call, and - either answer is acceptable if it is recorded. -- **A prompt-based delegate is not deterministic.** Any step that delegates to `/doctor` cannot - contribute to a diff-clean gate. Keep `/doctor`'s output in its own delegated tier, out of both the derived and judged finding sets. - -**Every detector asserts non-overlap with `/doctor`.** Per detector, against a freshly fetched commands -and memory doc — and the catalog carries a `/doctor` recheck trigger alongside its per-source-page -triggers, because `/doctor` moves on Anthropic's cadence and a one-time judgement decays. - -**Fresh-eyes delegation is designed into the artifacts, not left to the invoker.** The sweep's -apply-verify step and each detector's self-check are author-verifier arrangements, governed by -`docs/PLUGIN-PHILOSOPHY.md` "Fresh-eyes checkpoints" and "Delegation mechanics" — the latter arriving -with open PR #1096, whose vocabulary this work adopts rather than paralleling. Each checkpoint names -its fresh-context (non-fork) target, defaulting to the generic rung, and conforms to #1096's -declaration grammar. The deltas that bind sites here are enumerated in -[design/checks-and-sweep.md](design/checks-and-sweep.md), "Verification is designed in, not left to -the invoker". - -- **Sanity Check:** each of D1–D7 has a design section naming its detection rule, its remediation, and - at least one case it must NOT flag. -- **Sanity Check:** every split remediation names a load-deferring destination (a skill or a - path-scoped rule); no remediation anywhere proposes an `@path` import as a context saving — - `rg -c '@path'` over the detector designs returns hits only in the stated anti-fix warning. -- **Sanity Check:** the sweep design names its derived exclusion set, its dispatch order, its - `/doctor` version floor and absence classification, and its fresh-context verification checkpoint - in the "Delegation mechanics" vocabulary. - -### Gate — re-evaluate before implementation [TODO] - -Phases 2–6 are the design. Confirm the seam is chosen, Phase 2.5's dispositions held, every -surviving check has an owner, the re-run contract is testable, and the dispatch order is fixed. - -**The conditional `/planning:design` question is already answered: no.** Phase 2 chose the -no-shared-catalog shape, which fires this gate's condition, and the re-derived tier answers it — -that pass was warranted by a *new contract*, and there is no new contract. What remains is D1's -detection rule, the re-run contract, and the sweep's dispatch design, all already sequenced as -phases. Recorded in [design/proportionality-gate.md](design/proportionality-gate.md) under "The -design tier, actually re-derived", and this line closes the same question where -`design-resolution.md` also poses it. - -### Phase 7: Bring the setup-skill corpus to its owner doc [AUDITED] - -**The audit is done and recorded in [design/setup-corpus-audit.md](design/setup-corpus-audit.md); -the fixes are task #45.** 43 setup skills across 60 plugins: 41 conforming, 2 partial, 0 fully -non-conforming, 4 legitimately check-only with their premises settled from manifests rather than -from the skills' own prose. No setup skill reads a sibling plugin's files. - -The two most useful findings are not plugin defects. `context-guard` and `rate-limit-guard` share one -configuration shape — the user's own `settings.json` plus a plugin-owned machine file — resolved two -different ways, and the owner doc sanctions neither; and "non-trivial `userConfig`", which the -requirement gate turns on, is never defined, leaving three plugins' status unanswerable from the -doc's own text. Both are owner-doc fixes that dissolve plugin-level findings. - -One duplication decision remains open: a 61-line byte-identical block, about two thirds of each file, -shared by `discovery` and `verification`. It cannot ride the existing cluster registry, which takes -whole files only — extracting it to a per-plugin `reference/` file would make it registrable. - -The original phase text follows. - -Task #29, **reclassified.** The original framing — extract a new SSOT across the 43 `setup` skills — -is plugin-form-illegal: any single home is either a repo-level doc, unreachable from an installed -plugin's isolated cache, or a sibling-plugin file the design boundary forbids. Spot-checking confirms -the corpus is per-plugin by construction: `plugins/github/skills/setup/SKILL.md` references only -`${CLAUDE_PLUGIN_ROOT}/reference/*`. What is shared is the *shape*, and the shape already has an owner: -`docs/PLUGIN-PHILOSOPHY.md` "Setup is explicit and repeatable". - -So this phase audits conformance to that owner doc rather than inventing a second source. Where a -genuinely byte-identical fragment exists, it is a candidate for the existing cross-plugin cluster -registry — the mechanism the repo already uses for `hook-utils.sh` — not for a new extraction. - -Gated behind Phase 2 only if it turns out to need the same seam; otherwise file-disjoint from Phases -3–6 and parallel-safe. - -- **Sanity Check:** every `plugins/*/skills/setup/SKILL.md` either conforms to the owner doc's shape or - appears in a written exception list with a reason; any fragment promoted to a shared cluster appears - in `scripts/cross-plugin-source-registry.txt` and its drift check passes. - -### Phase 8: Implement the checks in their owning plugins [PARTIALLY SHIPPED — `audit-pass` run contract shipped; per-check detectors still open] - -Review: architecture - -Per the Phase 5 map. Each check ships with its criteria citation (never a restatement), its -false-positive carve-out, and evals where the owning plugin's conventions require them. - -**Phase 3's specified criteria edits execute here**, on the implementation branch, because they touch -live plugin files. Two further live-file corrections ride with them, both found during design and -neither belonging to the docs branch: the stale path-scoping claim in `claude-memory`'s -`reference/official-guidance.md` (task #47), and the setup-corpus fixes (task #45), whose first two -items are owner-doc changes to `docs/PLUGIN-PHILOSOPHY.md`. - -**Branch split happens here, before the first implementation commit.** Phase 11 requires documentation -and implementation to land as separate PRs; the fork point is the tip of -`docs/context-engineering-claude-5-topic` after Phase 7. - -**The conditional `setup`-skill work item is dropped.** It was contingent on Phase 2 choosing the -consumer-artifact shape, which it did not: no plugin gains a consumer-project configuration surface, -so neither `claude-memory` nor `docs-hygiene` needs a `setup` skill on this work's account. - -**Two plugins are modified, not four** — `claude-config` and `claude-memory`. `docs-hygiene` and -`skill-quality` are targets of the pass, never instruments of it. - -- **Sanity Check:** `/skill-quality:check` passes for every skill created or modified; no new skill - restates catalog content that Phase 2's seam makes citable. -- **Sanity Check — standalone usefulness.** Every new check, invoked with no catalog supplied, runs to - a documented reduced result and reports the missing optional capability by name. A check that only - works when the sweep calls it violates "Every plugin remains useful alone", which the - invocation-argument seam shape makes easy to breach silently. - -### Phase 9: Implement the sweep [PARTIALLY SHIPPED — `audit-pass` skill + run contract shipped; cross-plugin sweep wiring still open] - -Review: architecture, security - -Per Phase 6's design. Cross-plugin invocations are presence-gated with documented fallbacks — a bare -unguarded cross-plugin reference is a defect per the design boundary. Mutation sits behind the -explicit override the naming convention requires. - -- **Sanity Check:** every sibling-plugin invocation in the sweep body is inside a presence guard with - a stated fallback; bare invocation performs zero mutations. - -### Phase 10: Reconcile, then run against this repository [TODO] - -Tasks #31, #30, #20. Order matters and is not the task order. - -1. **Reconcile first.** Diff what the parallel `fable-5` session changed and re-check the affected rows - of `design/official-corroboration.md` — that skill is both a target of the pass and the doctrine - source its premises rest on. -2. **Inventory all three scopes before applying any side's fixes.** D1 detects cross-surface conflict; - it cannot see a repo↔user conflict from a repo-only inventory. `design/skill-inventory.md` names - `~/.claude/CLAUDE.md` against the 15-skill `discipline` plugin as the most likely conflict site. - Inventorying only the repo would apply fixes against half the picture. - - **A second example previously stood here and is struck as false.** It named "a `SessionStart` hook - injecting a persistent ruleset that no incumbent inventories as an always-loaded surface". No such - hook exists: the repo's only `SessionStart` arms a detached observer and emits no context, and - there is no `UserPromptSubmit` hook anywhere. Refuted and recorded in - [design/skill-inventory.md](design/skill-inventory.md), "Known conflict surface, already visible". - The three-scope requirement does not depend on it — the `~/.claude/CLAUDE.md` case carries it - alone. - - **The managed-policy scope is the third, and no design document names it.** The memory doc places - an organization-deployed `CLAUDE.md` at `/Library/Application Support/ClaudeCode/CLAUDE.md`, - `/etc/claude-code/CLAUDE.md`, and `C:\Program Files\ClaudeCode\CLAUDE.md`, plus a `claudeMd` key - honored in managed and policy settings only. It loads **before** user and project, and "managed - policy `CLAUDE.md` files cannot be excluded" — `claudeMdExcludes` does not reach it. A conflict - detector blind to that tier misses the highest-precedence surface, and worse, resolves a conflict - by proposing an edit to the lower surface when the authoritative side is org policy and the lower - side is the correct thing to keep. D1's surface partition must therefore carry a read-only, - never-remediated managed tier whose findings are reported as "conflicts with org policy at - ``", and the sweep must degrade cleanly when that path is unreadable — the common case on a - machine without one. This is in scope by default: the pass ships to organizations, not only to - this machine, where no managed policy file exists today. -3. **Then apply**, repo first: `claude-code-plugins` itself, recording hit rate, false positives, and - dispatch cost per check. -4. **Then route** the user-scope findings. Every one is a recommendation backfilled through - `melodic-software/dotfiles` — never an in-place edit. - -**This repository cannot exercise every rule, and the gap is measured rather than assumed.** Counted -at `cbf27e88a9`: **0** `@`-imports in the one root `CLAUDE.md` (63 lines) or in `AGENTS.md` -(28 lines); **0** nested `CLAUDE.md` files; **0** files under `.claude/rules/` — the directory does -not exist; **0** files carrying `paths:` frontmatter anywhere in the tree. So D6's two target -defects — `@path`-as-context-saving, and content in a compaction-losing destination that needs to -survive compaction — have **zero instances here**, because the destination class is empty. The repo -also already prohibits the practice in synced text: "Never `@import` (imports expand at launch, -defeating lazy load)" appears at line 211 of three files. - -Two consequences. A green dogfood run is not evidence that D6 works — it is evidence the repository -is clean, and the two must not be conflated in the Phase 10 report. And D6 needs a **synthetic -fixture** to be validated at all, which is a Phase 6 design obligation rather than a Phase 10 -discovery. The rules still ship: the pass targets organizations, not only this machine, and an -absent defect class here says nothing about a consumer's tree. - -#### Measured — every figure below carries the commit it was derived at - -**The design carried baseline numbers with no recorded commit and no script**, which is what let a -figure go stale unnoticed. These were re-derived by command at **`abe914eace`** and independently -reproduced in this lane. **The commit is part of each figure**, because the population itself moved: -this worktree's `plugins/` tree measures 181 skills and a 425-line maximum, while `abe914eace` -measures 182 and 497. Neither is wrong; a figure without its commit is. - -- **The skill listing budget — reproduces.** **131 listing-eligible skills, 80,220 description - characters.** Eligibility is *derived*, not assumed: listing-eligible ⟺ `disable-model-invocation` - is not `true`. **Zero of the 131 exceed the repo's own 1,536-character per-skill listing cap**, and - the longest is 1,161. So the fleet passes its per-skill gate 131 out of 131 while the **ungated - aggregate** stands at 80,220 — **the gap is a missing fleet-level cap, not a per-skill one**, which - is the finding that sharpens D6. Caveat that travels with the number: the repo fleet is not the - installed fleet, so 80,220 is this repository's listing cost, not any consumer's. -- **The line cap — reproduces, and it is the comparison that matters.** **0 of 182 skill bodies - exceed the 500-line hard cap; the maximum is 497** (`source-control/babysit-prs`) — a gate visibly - binding, a corpus written *to* its limit. Set against the re-attach cap, 182 of 182 pass the line - cap while 14 of 182 exceed truncation. **A conformance rate of 100% and a real defect population - of 14 coexist**, because the two caps measure different things. -- **The re-attach cap — reproduces with a correction**, and the eviction arithmetic **does not**. - Both are recorded in [design/article-sections.md](design/article-sections.md), row S1-c, with the - divisor and whole-file-versus-body choices stated alongside the count, because each moves the - answer more than the corpus change did. -- **Unreachable support files — did NOT reproduce. The old "~10 across 5 skills" must not ship.** - The verified count is **three**, each named, with the resolver-limitation methodology finding that - matters more than the count — [design/article-sections.md](design/article-sections.md), row S1-b. - -#### Coverage limits this run must state rather than imply - -A run that does not say what it could not reach reports its own blind spots as clean. - -- **Native-first is unsatisfiable when the runner cannot open a session inside its own target.** - `/context`, `/memory`, `/skills`, `/hooks`, `/status`, `InstructionsLoaded` logs, and - `--safe-mode` were all unavailable, so the pass degraded to exactly the filesystem-and-git walk the - design itself calls *"a candidate set, never the answer"*. Every native-first ruling in - [design/checks-and-sweep.md](design/checks-and-sweep.md) is therefore **designed but unexercised**. -- **The managed tier is entirely untested.** `C:\ProgramData\ClaudeCode` does not exist on this - machine, so every managed-tier scope rule — read-only, never remediated, highest precedence — is - unexercised rather than confirmed. -- **User scope is a fixed path, not a hierarchy ancestor.** A `CLAUDE.md` walk upward from the target - terminates at `D:\` and never reaches `C:\Users\\.claude\CLAUDE.md`. Any walk - implementation needs the fixed-path leg stated as its own step, or user scope is silently missing - from a run that looks complete. -- **A worktree lives inside the scan root** — `.claude/worktrees/ignition-rebind-note`, untracked. - A raw filesystem walk double-counts every surface beneath it; `git ls-files` does not. Concrete - evidence for the exclusion set's git-tracked-enumeration preference, which was argued from - principle before and is now argued from an instance. -- **The repository has no prompt-type instruction hook.** 15 `hooks.json` files, zero - `UserPromptSubmit`; the 13 plugins emitting `additionalContext` do so as `PostToolUse` - *diagnostics*, not standing instructions. Worth knowing before I12's surface set promises a - prompt-type hook surface it cannot exercise here. - -- **Sanity Check:** two consecutive runs over an unchanged tree emit machine-readable finding files - whose **derived-tier** sections diff clean, per the Phase 4 identity function — including the - three-scope inventory and the exclusion set, so a silent scope regression fails here. The - **judged-tier** section is held to Phase 4's stability tolerance instead, and exceeding it fails - the run's self-check rather than prompting a recalibration. `/doctor`'s output is the **delegated** - tier and is excluded from both — a prompt-based delegate cannot contribute to a determinism gate. -- **Sanity Check:** after an apply, the repo's cross-plugin drift checks still pass — - `scripts/check-cross-plugin-source-drift.sh`, `scripts/sync-hook-utils.sh --check`, - `scripts/sync-standards-contract.sh --check` — proving no registered cluster copy or vendored file - was mutated. - -### Phase 11: Acceptance gate and PR [TODO] - -Review: security - -Tasks #21, #32. `docs/MIGRATION-PLAYBOOK.md`'s per-plugin migration gate and plugin-acceptance security -review; repo-agnostic, `userConfig`-configurable, plugin-form-safe, no PII, explicit semver, catalog -assignment per `docs/CATALOG-TAXONOMY.md` — `claude-code`, since the subject *is* Claude Code and the -taxonomy's assignment principle gives subject priority over activity. - -#### The gate, walked against what PRs #1316 and #1318 actually ship - -Task #21, walked 2026-07-25 against the gate's own body — `docs/MIGRATION-PLAYBOOK.md` "Per-plugin -migration gate" and "Plugin-acceptance security review" — rather than against this plan's seven-word -summary of it, and against the two PRs' shipped files at `origin/feat/audit-pass-criteria` (#1316) -and `origin/feat/audit-pass-skill` (#1318). **One criterion fails, one cannot close, and two -divergences from this branch's own design are recorded.** A pass is recorded only where it was -substantiated. - -| Criterion | Verdict | Evidence | -|---|---|---| -| Repo-agnostic | **PASS** | `audit-pass/reference/exclusion-set.md` derives every exclusion class and names the empty case: "When the target documents no such registry, **this class is empty**." Zero hardcoded repo, marketplace, or machine identifiers across all four skill files | -| `userConfig`-configurable | **PASS**, by correct owner selection | No `userConfig` is the right answer, not an omission — `PLUGIN-PHILOSOPHY.md` "Configuration ownership and scope" owns a value by *kind*, and audit-pass has no personal-or-administrator scalar. Invocation choices are skill arguments, the suppression record is a consumer-project file, run state is `${CLAUDE_PLUGIN_DATA}` | -| Plugin-form-safe | **FAIL** | Three `../` reach-outs — see below | -| No PII or secrets | **PASS** today, with a named forward hazard | See below | -| Explicit semver | **PASS per PR, FAILS across the pair** | Each PR is internally consistent — `claude-config` `0.10.0` and `claude-memory` `0.5.0`, each with a matching changelog heading. But **both PRs bump `claude-config` from `0.9.2` to the same `0.10.0`**, so whichever merges second carries a version already taken. See below | -| Catalog assignment | **PASS** | `claude-config` is an existing entry already filed `claude-code`; `audit-pass` is a skill inside it, so no new marketplace entry is owed. `CATALOG-TAXONOMY.md` "Assignment principle" confirms subject-over-activity, and its `claude-code` scope requires "the subject to *be* Claude Code" — which it is | -| Plugin-acceptance security review | **CANNOT CLOSE** | See below | - -**The failure: three reach-outs to a repo-level document.** Security-review criterion 4 requires a -plugin reference only files inside itself and forbids `../` reach-outs; migration-gate step 4 says -the same positively. All three point at the same target, -`docs/conventions/finding-suppression/README.md`, which PR #1318 creates at the repository root: - -- `audit-pass/SKILL.md:163` — `../../../../docs/conventions/finding-suppression/README.md` -- `audit-pass/reference/run-contract.md:104` — `../../../../../docs/…` -- `audit-pass/reference/exclusion-set.md:72` — `../../../../../docs/…` - -Each resolves correctly in this repository and to **nothing** in an installed plugin, whose cache -root is the plugin directory. This is not a new hazard — it is precisely this plan's own -"Horizontal decoupling" grounding finding ("a repo-level doc is unreachable from an installed -plugin's cache") landing on the artifact that finding was written to protect, and the playbook lists -it first under "Plugin-form caveats (works in-repo, breaks as a plugin)". **The convention document -is the right home for the keys; the pointer to it is what has to change** — the plugin needs its own -bundled statement of the suppression shape, citing the convention by name rather than by relative -path. Owned by the implementation lane; not edited from this branch. - -**The semver collision is a cross-PR defect neither PR can see from inside itself.** Both PRs — the -criteria PR #1316 and the skill PR #1318 — modify -`plugins/claude-config/.claude-plugin/plugin.json`, and both take `0.9.2` → **`0.10.0`**. -Each is internally consistent, so each passes `changelog-parity-gate` on its own branch; the pair is -not. Whichever merges second lands a manifest version already claimed by the first, with a duplicate -`## [0.10.0]` changelog heading and no bump relative to the new `main`. **The second PR to merge must -rebase onto `0.11.0`** with its changelog heading renamed to match. `claude-memory` has no such -problem — only #1316 touches it (`0.4.0` → `0.5.0`). Recorded here because it is invisible from -either branch and surfaces only at the second merge, which is the expensive place to find it. - -**No PII passes today, and the hazard is in what is not yet built.** The shipped -`run-contract.md:17-20` scope-prefixes `surface` — "the file's *logical* path, never a -machine-absolute one" — so no report emits an operator's home directory. But the shipped contract -has **no liveness-basis concept at all**, and this branch's -[design/rerun-contract.md](design/rerun-contract.md) §6 requires a liveness-dependent finding to -record its basis *as report evidence*, including the launch directory and the effective merged -`claudeMdExcludes` — patterns that match **absolute paths**. When that lands it must take the same -scope-prefixed or redacted form `surface` already takes, or the report carries -`C:\Users\\…` into an artifact §2 explicitly permits redirecting into the target tree. Named -here so it is a designed-in constraint rather than a Phase 10 discovery. - -**The security review cannot be closed by this walk, and recording it as a pass would be false.** -Two of its own prerequisites are outstanding by this document's reckoning: the adversarial injection -fixture corpus (Phase 8) and verification that the exclusion-set derivation is not itself an -injection target (Phase 10), both owed by -[design/checks-and-sweep.md](design/checks-and-sweep.md), "What this section does not settle". The -review's surfaces 1, 2, 5, 6, and 7 are all clean — audit-pass ships no hooks, no MCP server, no -telemetry, no `bin/`, and no `settings.json` `agent` — so the residual is entirely the runtime -behavior the threat-model section covers. - -**Two divergences between this branch's design and what #1318 shipped, both since closed.** They are -kept rather than deleted because the record of what the walk found is the traceability this gate -rests on — but each is marked with what closed it, since a stale divergence claim on a durable -surface is the same defect this deliverable exists to detect. - -1. ~~**The shipped `run-contract.md:14` carries `identity = (surface, check, anchor, claim)`**~~ — the - superseded tuple, which [design/rerun-contract.md](design/rerun-contract.md) §1 replaced precisely - because it cannot express a pairwise finding, and D1 is the deliverable's entire payload. The - shipped file also carried no `sites` set, no `anchor/v1` version tag, and no `e:`/`s:` granularity - discriminator. **Closed:** #1318 now ships `identity = (check, claim, sites)` with `sites` a - canonically sorted set, the `anchor/v1` version tag, and both granularity prefixes. -2. ~~**Liveness is absent**~~ — Assertion 1.1 shipped unscoped and P1's exact equality shipped without - its liveness clause, both falsifiable in that form by correct behavior on a second machine. - **Closed:** Assertion 1.1 now reads "for a fixed tree **and a fixed live surface set**", and P1 is - conditioned on run comparability. - -**One sub-claim of divergence 1 outlived it and is now closed too: the missing harness-version -input.** #1318's P1 was conditioned on tree and live surface set alone, so a harness or -detection-version change — which legitimately moves the derived tier, since it reads a versioned -registry of harness behavior — read as a determinism failure. Closed by giving the shipped contract -a single **comparability** precondition, defined once and cited by every property, matching the -design's at the time: tree, live surface set, detection version, and harness version. - -**That parity has since lapsed in both directions, and reconciling it is Phase 8 work rather than a -design edit.** The design doc has since widened its first input to a **scan-baseline state digest** -(the shipped contract's own term, adopted from it) and added a **sweep version** the shipped -contract has no concept of; the shipped contract carries **behavior-affecting arguments** and an -**observable detection version** — a form it adopted after finding the design's catalog-version-plus- -prompt-digest triple unestablishable from outside a delegate plugin — that the design still lacks. -Neither list is wrong; they are two documents that moved at different times. The reconciliation is -recorded here rather than performed in passing, because changing either predicate changes what the -shipped pass reports. - -**Widening the cross-run input briefly left the within-run one behind, and that asymmetry is now -closed.** Assertion 3.4 recorded a *working-tree* content digest scoped to the repository while -comparability spanned every inventoried scope, so an edit to `~/.claude/CLAUDE.md` **between** runs -correctly rendered the pair non-comparable while the same edit **during** a run stayed invisible to -the mid-run integrity check. 3.4's endpoint capture is now the same all-scope state digest, matching -the shipped contract's assertion 6.1b. Within-run and cross-run integrity have to read the same -surfaces or neither is worth asserting — a rule worth carrying forward the next time either side -widens. - -None of these is a gate criterion; each would have made the shipped contract fail its own idempotence -claim, which is why they are recorded with the gate rather than after it. - -**The acceptance gate is not the whole security review, and assuming it was is a gap this plan -carried.** That gate checks distribution hygiene — properties of the artifact as a shipped plugin. -The sweep's exposure is *runtime behavior against hostile input*: it reads arbitrary instruction -text, which is text designed to steer a model, and then mutates files. No item on the acceptance -list surfaces that. [design/checks-and-sweep.md](design/checks-and-sweep.md), "Threat model — prompt -injection against the sweep", is an input to this phase: three threats grounded in the Claude Code -security page and OWASP LLM01:2025, with mitigations mapped to mechanisms in the design. Its two -deferred items — an adversarial injection fixture corpus (Phase 8) and verification of the -exclusion-set derivation (Phase 10) — are owed before this gate closes. - -**The traceability anchors before the prune.** `docs/conventions/topic-docs/` places -`docs/topics//` in the contract tier: committed on the task branch only, pruned before merge. But -acceptance criterion 2 requires traceability back to `design/article-sections.md`, and the catalog's -recheck triggers are anchored in `design/official-corroboration.md`. After the prune every such -reference dangles, and `link-check.yml` validates markdown links. - -**RESOLVED 2026-07-24 — task #37. Distil into the shipped catalog; do not graduate.** And it resolves -by *derivation* from the repo's own convention rather than by preference. - -`docs/conventions/topic-docs/README.md`, "Pointer discipline on durable surfaces", supplies a -surface-agnostic **rationale**: "The contract slice is deleted before merge and the memory slice -never leaves its checkout, so such pointers dangle by design. Cite the PR, the promoted location, or -distilled values instead." - -**The argument rests on that rationale and deliberately not on the enumeration.** The convention's -enumerated durable surfaces are "tickets, PR bodies, promoted docs" — a shipped plugin catalog is -**not** among them, and claiming it is would be exactly the transcribe-don't-derive failure this -effort keeps hitting. The rationale reaches it **a fortiori**: a shipped catalog outlives the prune -permanently, so the pointer dangles by precisely the mechanism the rule names, and more durably than -for any surface the rule does list. - -**The shape costs nothing new, because a sibling catalog already ships it.** -`plugins/claude-memory/skills/audit/reference/criteria.md` carries `Version: 1.2.0`, -`Last updated: 2026-07-11`, and a `Source:` line naming the official pages, plus per-check `**Why**` -lines quoting the official page. Verified this session: it greps clean for both `design/` and -`docs/topics` — **zero matches**. - -So: - -- **Every shipped catalog entry carries its source URL, its section id, and the recheck triggers - verbatim, distilled into the catalog.** No entry points at a `docs/topics/` path. -- **Traceability to `design/article-sections.md` is satisfied two ways** — at review time on the - branch, where the document still exists, and durably by **citing the PR that carried the design**, - never a `docs/topics/` path. -- **This is what unblocks the docs PR.** Nothing further is owed before the prune commit beyond - applying the shape, which task #34 does when it writes the catalog entries. - -- **Sanity Check — the repo's contract gate passes**, not a hand-picked subset. Sixteen `check-*.sh` - gates exist under `scripts/`; the load-bearing ones here are `check-changelog-parity.sh` (all 60 - plugins carry a `CHANGELOG.md`), `check-skill-leaf-names.sh` against `skill-leaf-name-registry.txt` - (a colliding leaf name must be registered as an owner set — registering a bare name is refused), - `check-skill-portability.sh`, `check-silent-skips.sh`, `check-orphaned-fixtures.sh`, plus - `validate-plugin-contracts.mjs` and `generate-catalog.mjs`. -- **Sanity Check:** `plugin.json` version bumped for every touched plugin; the marketplace entry carries - a category from the documented vocabulary; PR title matches `.github/workflows/pr-title.yml`; the PR - body carries a closing keyword and a non-empty `## Related` section per - `.github/workflows/pr-issue-linkage.yml` and `docs/conventions/pr-body-convention/`. -- **Sanity Check:** `rg -o 'design/[a-z-]+\.md'` across shipped plugin files and `docs/` returns no - reference to a pruned path. - -## Test strategy - -No runtime code — the deliverables are skills, a catalog, and their evals. Verification is therefore: - -- **Catalog fidelity** — every source section maps to a catalog entry or a recorded exclusion, - asserted by the Phase 3 sanity check against `design/article-sections.md`. -- **Detector precision** — each detector ships with at least one case it must NOT flag, drawn from - real files in this repo. The `audit-instructions` gotchas are the model: format-steering examples - are not scaffolding, bare-prohibition rewrites go positive before adding rationale. -- **Idempotence** — the Phase 4 contract asserted by running the sweep twice over an unchanged tree - and diffing the reports. -- **Structural gates** — `/skill-quality:check` on every new or modified skill; `/plugin-quality:audit` - on the resulting components. -- **Independence** — findings verified by fresh-context reviewers, never by the context that produced - them, per `PLUGIN-PHILOSOPHY.md` "Fresh-eyes checkpoints" and "Delegation mechanics"; conformance is - gated mechanically by `skill-quality:check` check 21. - -## Alternatives considered - -| Alternative | Why rejected | -|---|---| -| One new dedicated plugin holding every check | Duplicates the surface inventory `audit-instructions` already builds, forces a second install, and files the same subject under a second catalog entry | -| Extend `discipline:sweep-all` | Session-posture scoped, not target scoped. Its correctors audit the work in flight; this audits a repository at rest | -| Reimplement the CLAUDE.md trim | `/doctor` already trims, deduplicates, and migrates, on Anthropic's release cadence. The sweep hands off to it | -| Ship every rule the source states as an enforced check | Four rules have no official backing. They ship `OPINION`-tier so a consumer can weigh them, rather than being enforced as doctrine or silently dropped | - -## Risks and mitigations - -| Risk | Likelihood | Impact | Mitigation | -|---|---|---|---| -| Phase 2's seam forces the catalog into a shape that reintroduces drift | Medium | High | The gate after Phase 6 re-evaluates; shape 3 accepts drift explicitly rather than silently | -| Detectors flag correct constraint as over-constraint | High | High | D4's carve-out is a suppression input every trimming detector consults, designed before any of them ship | -| A fix-capable sweep mutates 181 skills on a bad rule | Low | Critical | Naming convention forces mutation behind an explicit override; bare invocation stays read-only; Phase 10 runs against this repo first | -| The parallel `fable-5` session and this branch diverge | Medium | Medium | Phase 10 reconciles before the first full run; PRs are required and squash-merged, so divergence surfaces at merge | -| Dispatch budget blows past what a session can hold | Medium | Medium | `audit-instructions` already gates near 20 dispatches; dynamic workflows are the documented mechanism above that ceiling — decided in Phase 6 | -| The pass becomes the thing it audits | Medium | Medium | It cites the catalog rather than restating it, and the source's own doctrine applies to it | -| An apply mutates a registered byte-identical cluster copy or a vendored upstream file, breaking a sync path and reddening CI | Medium | High | Phase 6 derives the exclusion set from `cross-plugin-source-registry.txt` + a `vendor/` rule + `git worktree list`; Phase 10 asserts the drift checks still pass after an apply | -| ~~The catalog is adopted by four plugins without an owner doc, violating the convention registry~~ | — | — | **Retired.** Phase 2 chose no shared artifact, so there is no cross-plugin catalog to adopt and no registry row is owed | -| Design documents are pruned at merge, dangling the catalog's traceability anchors | High | Medium | **Resolved (task #37).** Anchors are distilled into the shipped catalog — source URL, section id, recheck triggers verbatim — per `topic-docs`' pointer-discipline rationale; the PR, never a `docs/topics/` path, carries design traceability. Applied by task #34 | -| Injected instruction text in an audited surface steers a lane — suppressing a finding, or laundering an attacker's edit through an accepted fix | Medium | High | `design/checks-and-sweep.md` "Threat model — prompt injection against the sweep": examined text is denoted as data not instruction; suppression is never reachable by the apply step; least privilege on managed and user scope; no unattended applying run | -| A detector or the sweep self-grades its own output | Medium | High | Phase 6 requires each to name a fresh-context checkpoint conforming to `PLUGIN-PHILOSOPHY.md` "Delegation mechanics" (PR #1096), whose conformance gate is `skill-quality:check` check 21 | -| The machinery outweighs the payload: a versioned catalog, a convention-registry owner doc, a seam decision, a re-run contract, and a sweep built to carry one well-grounded detector | High | High | **Fired, and the mitigation worked.** Exactly one detector survived, the operator approved re-deriving the shape, and the catalog, the owner doc, the registry row, and the materialization are gone. What survives is the sweep and the re-run contract, whose justification was never detector count | -| `OPINION`-tier rules mutate a consumer's instruction corpus under the same banner as documented doctrine | Medium | High | The tier is populated for the first time by this work, so this work defines it: disabled on bare invocation, opt-in, with a severity ceiling — Phase 6 | -| Idempotence is asserted over detection that is a model judgement and cannot be deterministic | High | High | **Fired, and the first mitigation was itself wrong.** Scoping the gate to the `mechanical` tier did not work: no dispatched check reaches the report without model judgment, and half the dispatched catalog has no tier axis at all. Re-derived into three tiers — derived (model-free: inventory, exclusion set, shadowing, raw candidate rows) carries exact equality; judged carries a tolerance whose violation fails the run; delegated carries neither | -| `/doctor` is absent on a consumer machine — `DISABLE_DOCTOR_COMMAND` or `skillOverrides: {"doctor": "off"}` — leaving its half of the surface with no incumbent and nothing built to replace it | Medium | High | Phase 6 treats presence as a prerequisite alongside the version floor; the sweep names the missing capability and states what goes unchecked | -| D1 is blind to the managed-policy `CLAUDE.md` tier and proposes edits to a lower surface whose conflict is with unremovable org policy | Medium | High | Phase 10 inventories three scopes; the managed tier is read-only and never remediated, and its absence degrades cleanly | -| Phase 1's page list omits a load-bearing doc and its sanity check cannot notice | Medium | Medium | Four pages added; a second check walks `llms.txt` and records every instruction/memory/configuration page as fetched or explicitly out of scope | - -## Blast radius - -**Still HIGH, but for fewer reasons than before the gate.** Every plugin and skill in the marketplace -remains in scope as a **target** (counts cited by command in "Standards grounding" rather than -transcribed), and a fix-capable pass can mutate the marketplace's own instruction corpus — including, -if the exclusion set is wrong, 13 registered byte-identical cluster copies and six vendored upstream -materializations. That is what keeps the rating where it is. - -Two of the original four contributors are gone. **Two plugins are modified as instruments, not -four** — `claude-config` and `claude-memory`; `docs-hygiene` and `skill-quality` are targets only. -And **no new contract surface is consumed across plugin boundaries**: nothing crosses a boundary -except a presence-gated invocation. Recorded rather than silently re-rated, because a blast radius -that never moves is a blast radius nobody is reading. - -## Open questions - -Three of the five are closed. Kept with their answers rather than deleted, because later phases cite -them. - -- ~~Phase 2's seam choice.~~ **Closed:** no shared criteria artifact — - [design/seam-resolution.md](design/seam-resolution.md). -- ~~Whether D2–D7 survive as detectors.~~ **Closed:** they do not. One officially-backed new check - (D1), one `OPINION`-tier new check (D3), four edits to existing checks, one deferred — - [design/proportionality-gate.md](design/proportionality-gate.md). -- ~~What an `OPINION`-tier finding means to a consumer.~~ **Closed:** emitting rules default off with - an `info` severity ceiling, never fix-applied, and every run reports the tier's existence and the - argument that enables it; withholding rules default **on**. -- **Open** — whether the sweep's dispatch exceeds a session's ceiling and must become a dynamic - workflow. Phase 6. -- **Open** — naming, under the fixed verb meanings. Phase 6. -- **Open, new** — whether `mcp-tools:audit` actually covers tool-search configuration. The gate - defers the "deferred tool loading is unowned" remainder out of scope on that basis, which is a - negative claim about a body nobody has read. Task #44. -- **Open, new** — a skill-body "genericness" check candidate for `skill-quality:check` (could this - skill body be pasted into any repo unchanged?), recorded as an input by the - finding-your-unknowns integration (2026-09-01, its topic's signoff-sheet, F2). Resolution - belongs to this topic's check-design phases, not that effort. -- **Open, new** — widening `claude-config:audit-instructions` check I29 (restatement families - I29-a/I29-b) with the companion article's repetition-myth duplication lens, recorded as an - input by the same F2 reroute. Owner: the phase that next touches the I-check catalog. -- **Open, new** — a `claude-memory:audit` cross-ref against this topic's `/doctor` prerequisite - contract (design/checks-and-sweep.md), recorded as an input by the same F2 reroute. -- **Note (2026-09-01):** Phase 10's sweep re-inventories current state before running — its - recorded surface counts predate the finding-your-unknowns waves landing on - `claude/reading-feedback-j4sg96` and are stale; the sweep rebases over whatever has landed. - That effort touches none of this topic's audit-instructions criteria files. - -## Handoff to implementation - -### User-approval gates - -- The Phase 6 → Phase 8 gate: no implementation begins until the design phases land and the gate is - re-evaluated. -- Any mutation applied to a surface outside this repository, including every user-scope finding, - which is routed through `melodic-software/dotfiles` rather than applied. -- Phase 2's seam choice, if it lands on a shape that changes what the Brief promised. - -### Execution shape - -Phases 1 → 2 → 2.5 → 3 → 4 → 5 → 6 are sequential: each consumes the previous phase's output. Phase 7 is -file-disjoint from Phases 3–6 (it touches `plugins/*/skills/setup/`, they touch design documents) -and can run in parallel. Phases 8–11 are sequential and gated. - -| Phase | Surface | Basis | -|---|---|---| -| 1 | Main session | Doc fetches feeding a document the main thread owns | -| 2 | Main session | The load-bearing architectural decision | -| 2.5 | Main session | Scope gate with a user-approval escalation | -| 3–6 | Main session | Design work; judgment-heavy, tightly coupled | -| 7 | Sub-agent worker | Mechanical, file-disjoint, 30+ files of the same shape | -| 8 | Sub-agent workers, one per owning plugin | File-disjoint by construction once Phase 5 assigns owners | -| 9–11 | Main session | Integration, security posture, and the acceptance gate | - -Phase 7's worker is fenced to `plugins/*/skills/setup/**` and is forbidden PLAN.md, every design -document, and any plugin surface outside `skills/setup/`. - -### Mechanical work - -Commit at each phase boundary; PLAN.md phase tags advance in the same commit as that phase's output. -PLAN.md edits stay main-session only. Sequential fallback: if a parallel worker reports it cannot -complete or violates its fence, abort that worker and fold its phase back into the main session. diff --git a/docs/topics/finding-your-unknowns-integration/PLAN.md b/docs/topics/finding-your-unknowns-integration/PLAN.md deleted file mode 100644 index db71a635c..000000000 --- a/docs/topics/finding-your-unknowns-integration/PLAN.md +++ /dev/null @@ -1,495 +0,0 @@ -# finding-your-unknowns-integration - -## Brief - -### TLDR - -Integrate the verified "Finding Your Unknowns" corpus (Thariq Shihipar's field guide, its -X-Article draft, the 13 html-effectiveness pages, and the context-engineering companion; -slice `finding-your-unknowns-0f25bd45`) into this repo as judgment-preserving deltas on -existing skills plus a small set of citable reference docs, gated by a scripted -comparison-evidence pass over the ~20 named-skill collisions. - -### Goal - -Every corpus decision in -`.work/finding-your-unknowns/finding-your-unknowns-0f25bd45/corpus-inventory.md` reaches a -disposition (adopt / adapt / compare-then-adopt / cite / drop / treat-as-caution) that is -either executed as a repo change or recorded with its reason, with no silent drops. - -### Constraints - -1. Vehicles (interview Q1): (a) augmentation of existing skills and (b) citable - reference/convention docs are the primary vehicles; at most 2 genuinely new thin skills, - each only where the evidence pass confirms a real gap; CLAUDE.md/rules changes only if - the context-engineering material earns one on its own evidence. -2. Codification posture (Q2, BINDING): honor the source author's anti-premature-codification - warning — only judgment-preserving deltas (output contracts, checklists, conventions with - rationale); no generator-style "make me an X" skills; the warning itself is quoted in - whatever reference doc graduates. -3. Evidence discipline (Q3): adopt/adapt verdicts on named-skill collisions are CONDITIONAL - ("adopt X into S if S lacks it") until a read-only evidence pass grades each named skill - against its checkable claims; only surprises return to the human. -4. Vendor-claim discipline (Q5): corpus claims are vendor-blog anecdote unless the targeted - live-doc check (folded into the evidence pass) verifies them; the ~6 decision-relevant - harness claims (auto-memory, /doctor, ToolSearch deferred loading, artifacts-as-context, - 80%-claim scoping, rich references) get that check; nothing else does. -5. House style: all new prose obeys the repo's ai-slop/house-style rules - (.claude/rules/vendor-docs-are-not-style.md); citation shape follows - plugins/knowledge/reference/citation-shape.md (URL + retrieval date + content hash). -6. Execution contract (Q6): each conditional-batch unit is closed only when its verdict is - executed-or-recorded and the inventory row links the outcome; no silent drops. -7. Vehicle placements (Q7-Q10): the five-pass sequencing lands as a workflow section in the - central reference doc with cross-refs from planning:wayfind and session-flow:workflow (no - new orchestration skill); the reference doc owns the prompt-pattern catalog with one - canonical invocation line per owning skill; reply-affordance is a house convention - (default-with-judgment) owned by the doc with an artifact-design cross-ref. -8. Evidence-gate classification (sign-off S1): every delta is classified per-row under - PLUGIN-PHILOSOPHY's rubric (docs/PLUGIN-PHILOSOPHY.md:699-742). CONTRACT/POLICY/CONVENTION - rows land now as team conventions adopted by the sign-off; DOC rows are citation/doc - lines; BEHAVIORAL rows never land as standing instructions — they become reference-doc - heuristic lines plus candidate eval cases, awaiting observed-stumble evidence. -9. Registry discipline (sign-off S2): convention-registry rows for reply-affordance and - export-button only, each with an explicit conformance surface, landing owner-doc-first - (owner doc before a second adopting plugin, PLUGIN-PHILOSOPHY:594-595). -10. Sequencing with docs/topics/context-engineering-claude-5/ (sign-off S3): this effort - never edits that topic's audit-instructions criteria files; Wave 1 records F2's three - candidates in that topic's PLAN plus a Phase-10 re-inventory/rebase note; no freeze. -11. Schema stability (sign-off G.4 / D37): the `### Phase N` heading/tag vocabulary and any - parsed schema (check-open-questions.sh fields) are never renamed without a version bump - and changelog; block reordering is safe. - -The complete signed decision surface — classified delta roster (D/E/F/G rows), wave -assignments, per-finding dispositions — is ./signoff-sheet.md (rev 2, operator-confirmed -".confirm all" on 2026-09-01). - -### Acceptance criteria - -(Signed off with the sheet, 2026-09-01; sign-off sheet Part G.5.) - -1. Every decision in corpus-inventory.md has an executed-or-recorded disposition traceable - from ./signoff-sheet.md — verified by grep over the disposition lines, not asserted. -2. All waves green on the per-plugin gates: version bump + CHANGELOG entry; - check-changed-skills.sh (trigger-keyword preservation, listing cap, --require-evals); - listing budget respected; cheat-sheet regenerated once per PR series, series sequential; - scripts/affected-tests.sh --run green. -3. The F1 reference doc carries the quoted anti-premature-codification warning and the C6 - permission/quoting basis in its header. -4. Registry rows (reply-affordance, export-button) carry explicit conformance surfaces and - land owner-doc-first. -5. docs/topics/context-engineering-claude-5/PLAN.md carries the S3 sequencing rows. -6. Every Wave-2/Wave-3 contract delta lands with eval expectations in the same commit as - its contract lines (sign-off Part D eval-impact column). -7. /planning:plan consumes ./signoff-sheet.md + this Brief as its input contract. -8. Delivery (operator amendment at sign-off): all execution in one session on branch - claude/reading-feedback-j4sg96, delivered as ONE pull request; waves are commit - ordering, not separate PR series; deferred items may become filed issues. - -### Captured assumptions - -- The published blog is the canonical citation source for wording; the X draft is citable - only for draft-only content (5 images, 3 links) — per corpus V7.1. -- Arm-B verification ran degraded (same-vendor adversarial refuter); accepted as sufficient - for this corpus. - -### Out-of-scope - -- Watching/transcribing the two linked videos (companion-classified; no ingest path). -- Re-digesting the 3 referenced-external related-posts articles. - -### Deferred questions - -(No open-register rows were retired as deferred; the items below are sub-decisions the -sign-off explicitly deferred, each with its trigger and arbiter.) - -- E4-EXT | E4 skill extension (planning:prd durable-output pitch mode vs a design-handoff - layer) — deferred behind demand evidence; the doc-tier buy-in pattern ships now (S4). - Arbiter: human, at the recorded-candidate evidence point. -- D28-SCHEMA | Adding a mechanical scrutiny-flag field to the 5-field open-question register - schema — the resolution-field convention ships now, gate-invisible by design; the schema - change waits until a consumer needs mechanical reads (M4). Arbiter: that consumer's - change, via version bump + changelog per constraint 11. -- EVAL-CAND | Behavioral eval candidates (D5, D7, D11, D17b, D34-residue) — doc lines now; - promotion to standing skill instructions only on observed-stumble evidence per the - evidence-gated-additions rule (PLUGIN-PHILOSOPHY:699-711). Arbiter: that rule. -- Q11-FLIP | Flipping tweak-likelihood ordering from documented-default to requestable mode - in further consumers — deferred to own-usage evidence (interview Q11 residue). -- E1-REG | A convention-registry row for the deviation-log — retracted as premature (M3); - fires the moment a second plugin reads DEVIATIONS.md (recorded trigger, C5). - -## Plan - -Written by /planning:plan on 2026-09-01, executing the signed contract (./signoff-sheet.md -rev 2 + this Brief). The delta roster, wave mapping, and definition of done are locked -decisions; these phases sequence their execution, they do not relitigate them. - -Standards grounding: no standards index exists in this repo (`.claude/` carries rules, not -an index); the plan is grounded directly in AGENTS.md (affected-tests contract), -docs/PLUGIN-PHILOSOPHY.md (convention registry :591-631, instruction economy :681-718, -two-lane posture :245-286), docs/MIGRATION-PLAYBOOK.md (per-plugin version bump + CHANGELOG -delivery), .claude/rules/vendor-docs-are-not-style.md, and -plugins/knowledge/reference/citation-shape.md — all read this session. - -### G-block placements (durable copy; source: round-3 audit merge, confirmed C1) - -All are doc-tier placements consumed by Phases 1-2. G1 map/territory: cite-only in F1, -never house vocabulary. G2 over/under-specify diagnostic: F1 doc line, skill edit deferred. -G3 disclose-starting-point primer: F1 pattern-catalog entry. G4 cost framing: F1 intro -rationale line. G5 long-horizon failure diagnostic: F1 doc line, debugging edit deferred. -G6 interactivity scope note: sequencing is chat-portable, artifact optional. G7 -fresh-session-per-phase: corroboration, F1 cite only. G8 stay-in-the-loop criterion: quoted -in F1 caution section. G9 density rubric: F1 when-HTML section. G10 ~100-line ceiling: F1, -labeled practitioner anecdote. G11 sharing argument: one F1 line (satisfied by the Artifact -tool). G12 throwaway-editor doctrine: F1 caution section beside the codification warning. -G13 HTML-diff noise: F1 scoping rule (HTML for ephemeral/published outputs, never -version-controlled instruction surfaces). G14 design-system/PR-explainer/report-HTML -options: F1 pattern entries, skill edits deferred. G15 workflow shape: covered by the -five-pass workflow section. G16 named-expert anecdotes: dropped from all graduated -artifacts. - -### Phase 1: F1 reference doc — docs/FINDING-YOUR-UNKNOWNS.md [DONE] - -Create the graduated reference doc following the top-level precedent shape (H1 → -`## Contents` anchor TOC → charter paragraph → H2 sections). Content contract (signed -Part E + G-block): - -- Header: provenance + the C6 permission basis (short attributed verbatim excerpts from - the named author's public posts under fair-quotation practice, no license claimed, bulk - reproduction avoided) + the quoting-posture line (quotes stay byte-verbatim; S6/M1) + - the citation shape stated locally (URL + ISO retrieval date + sha256 over snapshot - bytes) rather than path-citing the plugin-private shape doc (public-surface rule). -- Unknowns taxonomy + lifecycle: the quadrants, D1's four finding types, G4 cost-framing - rationale, G2/G5 diagnostic doc lines, G1 map/territory as citation only. -- Five-pass sequencing workflow section (Q7 home) + G6 interactivity note + G7 cite + - G15 coverage. -- Prompt-pattern catalog (Q9 home): G3 primer entry, G14 pattern entries, one canonical - invocation line per owning skill. -- Reply-affordance convention + export-button rule: the owner sections for both registry - rows, each following the owner-doc anatomy (the rule, who is bound, conformance) and - naming its conformance surface verbatim ("conformance = the template blocks in - prototype:explore-directions (D21/D22) and prototype:pressure-test (D24)"). The - reply-affordance section carries the one-line artifact-design cross-ref (Q10) as a prose - mention of the session built-in skill — the established idiom, since artifact-design has - no in-repo surface (verified: only prose mentions exist repo-wide). -- When-HTML taxonomy: G9 density rubric, G10 ceiling (labeled practitioner anecdote), - G11 sharing line, G13 scoping rule, the HTML-scoping rule. -- Buy-in pattern (S4) + E5 objection-evidence checklist, citing Rust RFC / Oxide RFD / - Amazon PR-FAQ + corpus. -- Caution section: the author's anti-premature-codification warning quoted verbatim with - citation stamp (Brief constraint 2), G8 stay-in-the-loop criterion, G12 - throwaway-editor doctrine. -- Behavioral heuristics as doc lines, each labeled as an eval candidate awaiting - observed-stumble evidence: D5 observed-fact evidence bar, D7/D11 disconnected-work - heuristic/recipe, D17b non-obvious-behavior keying, D34-residue collapse self-check. -- Q11/D31 reconciliation line (presentation ordering is already planning:plan's - documented default). -- The doc cites x.com URLs; lychee.toml already excludes `^https?://(www\.)?x\.com/` - (verified at execution — the review's "no x.com entry" finding was itself wrong), so no - config change was needed. - -**Sanity Check:** `test -f docs/FINDING-YOUR-UNKNOWNS.md`; `test "$(grep -c '^## ' docs/FINDING-YOUR-UNKNOWNS.md)" -ge 7`; -`grep -q 'retrieved 2026' docs/FINDING-YOUR-UNKNOWNS.md` (citation stamps present); -`grep -qi 'fair.quotation' docs/FINDING-YOUR-UNKNOWNS.md`; `grep -qF 'x\.com' lychee.toml`; -`npx markdownlint-cli2 docs/FINDING-YOUR-UNKNOWNS.md` exit 0 (affected-tests classes -docs/*.md as no-suite — the hygiene tools must be invoked directly). - -### Phase 2: Governance placements — registry, glossary, ctx-eng sequencing, traceability ledger [DONE] - -- docs/PLUGIN-PHILOSOPHY.md convention registry: add exactly 2 names-and-points rows - (reply-affordance → the F1 doc's owner section; export-button → same). The registry - table is two-column and never restates, so acceptance criterion 4 ("rows carry - conformance surfaces") is discharged by the owner sections the rows point at — the - reconciliation is recorded here and verified in Phase 11. -- docs/GLOSSARY.md via the /domain-driven-design:curate-language disposition procedure: - add the unknowns-quadrant vocabulary + D1's four finding types with "Avoid:" lines and - a dated Provenance entry; map/territory lands as a `## Rejected terms` row pointing at - the F1 cite-only disposition (G1) so the next curate-language run sees the decision. -- docs/topics/context-engineering-claude-5/PLAN.md: add "Open, new" bullets under - `## Open questions` recording F2's three candidates (skill-quality genericness check; - audit-instructions I29 widening; claude-memory /doctor cross-ref against its existing - /doctor contract section) and one line noting Phase 10's sweep re-inventories current - state (recorded counts stale) and rebases over this topic's landed waves. NO phase - heading/tag renames in that file (D37 / Part G rule 4). -- Traceability ledger (acceptance criterion 1 runs corpus→sheet, not the reverse): commit - `docs/topics/finding-your-unknowns-integration/delta-resolution.md` (the full delta - wording record) and a `disposition-ledger.md` crosswalk keyed by the corpus-inventory - V-ids (V1.1-V7.n), each row naming its sheet/G row id and disposition, so criterion 1 - is verifiable from committed files and every corpus decision is accounted for in the - correct direction. - -**Sanity Check:** `test "$(grep -c 'FINDING-YOUR-UNKNOWNS' docs/PLUGIN-PHILOSOPHY.md)" -ge 2`; -`grep -q 'unknown' docs/GLOSSARY.md`; `grep -qi 'map.*territory' docs/GLOSSARY.md` (rejected-terms -row); `test "$(grep -c 'Open, new' docs/topics/context-engineering-claude-5/PLAN.md)" -ge 4`; -`! git diff HEAD~1 -- docs/topics/context-engineering-claude-5/PLAN.md | grep -q '^-.*### Phase'`; -`test -f docs/topics/finding-your-unknowns-integration/disposition-ledger.md` and every V-id in it -carries a non-empty disposition (`! grep -E '^\- V[0-9]' disposition-ledger.md | grep -q '\|\s*$'`); -`npx markdownlint-cli2` over the touched docs exit 0. - -### Phase 3: discovery plugin — blindspot contract deltas [DONE] - -plugins/discovery/skills/blindspot/SKILL.md: D1 typed finding taxonomy in the output -contract; D4 scan-scope disclosure line (POLICY). Extend evals/evals.json expectations for -both new contract lines in this same commit. plugins/discovery/.claude-plugin/plugin.json -version bump + CHANGELOG.md entry. - -**Sanity Check:** grep for the taxonomy + disclosure lines in blindspot SKILL.md; -`F=plugins/discovery/skills/blindspot/evals/evals.json; test "$(jq '[.evals[].expectations]|flatten|length' $F)" -gt "$(git show HEAD:$F | jq '[.evals[].expectations]|flatten|length')"`; -`git diff HEAD --name-only` includes discovery plugin.json + CHANGELOG; -`bash scripts/check-changed-skills.sh origin/main` green (affected-tests classes SKILL.md -as no-suite; this is the gate that runs trigger-keyword preservation, listing cap, ---require-evals per Part G rule 2). - -### Phase 4: education plugin — explain + quiz-me contract deltas [DONE] - -explain SKILL.md: D12 vocabulary ladder; D14 success condition scoped to original-ask -invocations. quiz-me SKILL.md: D16 source-anchor + on-miss routing; D17a diff-sourced -question authoring; F4 fresh-context answer-key requirement. Same-commit evals extensions -for all five; education plugin version bump + CHANGELOG. - -**Sanity Check:** grep each of the five contract lines in its SKILL.md; evals.json -expectation growth in both skills (before/after jq pair as in Phase 3); -`bash scripts/check-changed-skills.sh origin/main` green. - -### Phase 5: verification plugin — confirm deltas [DONE] - -confirm SKILL.md: D19 "existing behavior this leans on" callout (CONTRACT) + D18 -quiz-layer cross-ref doc line (the reroute executed; quiz-me untouched by D18). D19 likely -also needs a row in the report template in context/outcome.md (read it first). Same-commit -evals extension for D19; verification plugin version bump + CHANGELOG. CONSTRAINT: the -em-dash ratchet (scripts/em-dash-purged-paths.txt) covers verification SKILL.md files — -no em dashes in any new line here. - -**Sanity Check:** grep both lines in confirm SKILL.md; evals growth (before/after jq pair); -`bash scripts/check-changed-skills.sh origin/main` green; `bash scripts/check-purged-em-dashes.sh` green. - -### Phase 6: prototype plugin — explore-directions + pressure-test deltas [DONE] - -explore-directions SKILL.md: D20 same-data control-variable rule (POLICY); D21 structured -steal/graft capture; D22 machine-legible reply template (shaped as the skill's OWN output -contract, never a consumer-repo format — lane-2 constraint). pressure-test SKILL.md: D24 -validation-answer-set shape; D25 fake-data disclosure footnote (POLICY); D26 per-option -named costs. D27 mock-before-wire composition note in the shared -plugins/prototype/context/discipline.md Composition table. Same-commit evals extensions; -prototype plugin version bump + CHANGELOG. CONSTRAINT: the em-dash ratchet covers -prototype SKILL.md files — no em dashes in any new SKILL.md line here. - -**Sanity Check:** grep the six contract/policy lines + D27 note; evals growth in both -skills (before/after jq pairs); `bash scripts/check-changed-skills.sh origin/main` green; -`bash scripts/check-purged-em-dashes.sh` green. - -### Phase 7: planning plugin — interview/plan/design/brainstorm deltas [DONE] - -interview SKILL.md: D28 free-text scrutiny flag as a RESOLUTION-FIELD convention (the -5-field register schema is untouched; gate-invisible by design, limitation recorded in the -skill text). plan SKILL.md: D32 switch condition on alternatives; D33 closing revision -replies; D35 lands-green forward-reference doc line. design SKILL.md: D36 the same -tweak-likelihood presentation ordering plan already documents. brainstorm SKILL.md: D9 -brainstorm-practice citation doc line. wayfind SKILL.md: the five-pass workflow-section -cross-ref doc line (Brief constraint 7 / Q7 — points at the F1 workflow section as -marketplace-repo prose). Same-commit evals extensions for the four contract rows (D28, -D32, D33, D36); planning plugin version bump + CHANGELOG. - -**Sanity Check:** grep each delta line (incl. the wayfind cross-ref); -`bash plugins/planning/scripts/check-open-questions.test.sh` green (invoked directly — -affected-tests does not select it for SKILL.md/context edits; D28 must not break the -register gate); evals growth for interview/plan/design (before/after jq pairs); -`bash scripts/check-changed-skills.sh origin/main` green. - -### Phase 8: discipline plugin — point-dont-copy W2 deltas [DONE] - -point-dont-copy SKILL.md: E7 primitive-to-convention trap item (CONTRACT) in the -"Audit. What to look for" list; E8 canonical invocation line (as `argument-hint` -frontmatter, not a description rewrite — trigger keywords stay intact; frontmatter change -means the cheat-sheet regenerates in Phase 11). Same-commit evals extension for E7; -discipline plugin version bump + CHANGELOG. The CHANGELOG entry is PROVISIONAL: Phase 10 -extends this same version's entry with the E6 line (one bump per plugin per PR; the -intermediate commit's entry knowingly under-describes and Phase 10 finalizes it). - -**Sanity Check:** grep E7 line + `argument-hint` in frontmatter; evals growth -(before/after jq pair); `bash scripts/check-changed-skills.sh origin/main` green. - -### Phase 9: session-flow plugin — workflow cross-ref [DONE] - -plugins/session-flow/skills/workflow/SKILL.md: the five-pass workflow-section cross-ref -doc line (Brief constraint 7 / Q7 — marketplace-repo prose pointing at the F1 workflow -section). session-flow plugin version bump + CHANGELOG (a doc-line-only bump; no eval -extension — no contract change). - -**Sanity Check:** grep the cross-ref line in workflow SKILL.md; session-flow plugin.json + -CHANGELOG in the diff; `bash scripts/check-changed-skills.sh origin/main` green. - -### Phase 10: Wave 3 — E6 port gate + E1/E2 deviation-log convention [DONE] - -Own review moment; lands after all W2 phases. - -- E6 (discipline:point-dont-copy): semantics map + stop-and-wait confirmation gate, - scoped to external-reference ports with the C3 boundary sentence verbatim ("source of - truth outside this repo's tree: vendored, foreign-language, other-repo"); in-tree - corrections stay do-it-now. Read plugins/discipline/context/re-anchor-audit-correct.md - FIRST and confirm the gate does not reverse its correct-forward doctrine (the gate lands - in point-dont-copy's own file, never the shared doc). Finalizes the discipline CHANGELOG - entry opened in Phase 8 (same version — one bump per plugin per PR). Same-commit evals - extension. -- E1+E2 (implementation plugin: implement + implement-dispatch): opt-in deviation-log - convention + fold-back step (implement Step 5 gains the read-DEVIATIONS.md item; the - schema/taxonomy extension lands in implement-dispatch's "Divergence in non-interactive - runs", which implement's interactive path then cites); the F1 owner section (Phase 1) - is the doc home; the recorded-trigger note lands where the convention is stated ("the - moment a second plugin reads DEVIATIONS.md, the registry rule fires" — no registry row - now, per C5/M3). implementation plugin version bump + CHANGELOG; same-commit evals - extension. CONSTRAINT: the em-dash ratchet covers implementation SKILL.md files — no - em dashes in any new line there. - -**Sanity Check:** grep the boundary sentence verbatim in point-dont-copy SKILL.md; grep -the trigger note in the implementation plugin; both plugins' CHANGELOGs updated; -`bash scripts/check-changed-skills.sh origin/main` green; -`bash scripts/check-purged-em-dashes.sh` green. - -### Phase 11: Close-out — issues, cheat-sheet, acceptance verification, PR [DONE] - -- File follow-up GitHub issues (authorized): one for the behavioral eval candidates - (D5, D7, D11, D17b, D34-residue), one for the E4 skill extension candidate, one for the - D28 schema-change deferral, one for Q11-FLIP + E1-REG recorded triggers (or fold small - ones into a single tracking issue — executor's judgment on granularity). -- `node scripts/generate-cheatsheet.mjs --check`; regenerate once (E8's argument-hint is - a frontmatter change, so regeneration is expected). -- Acceptance-criterion 1 verification in the CORPUS→SHEET direction: iterate every V-id - row of the committed disposition-ledger.md (Phase 2) and confirm each names a - disposition + sheet row; then confirm every sheet CONTRACT/POLICY/CONVENTION/DOC row - resolves to a diff hunk. Write both sweep results into this PLAN under a dated note. -- Walk acceptance criteria 2-7 explicitly, one recorded line each (2: all phase gates - green on head; 3: F1 warning + C6 grep; 4: registry rows + owner-section conformance - text, per the Phase 2 reconciliation; 5: ctx-eng rows grep; 6: eval-expectation commits - co-located with contract commits via `git log --name-only`; 7: this PLAN + sheet were - the executed contract). -- /ai-slop:audit pass over the new F1 doc + a sample of edited SKILL.md hunks (Brief - constraint 5); fix findings before the PR. -- Full gate battery on the branch head: `bash scripts/check-changed-skills.sh origin/main`; - `bash scripts/check-purged-em-dashes.sh`; `npx markdownlint-cli2` on touched .md; - `scripts/affected-tests.sh --run` (covers any script/test surfaces touched). -- Push and open the ONE pull request (explicitly requested), body mapping commits to - waves and citing ./signoff-sheet.md. - -**Sanity Check:** `node scripts/generate-cheatsheet.mjs --check` exit 0 after regeneration; -all five gate commands above exit 0; issue URLs recorded in this PLAN; PR URL recorded; -the criterion-1 double sweep pasted as a dated note with zero unaccounted rows; criteria -2-7 walk recorded. - -#### Close-out record (2026-09-01) - -- **Criterion 1 (traceability, both directions):** sweep A iterated all 48 V-id rows of - ./disposition-ledger.md — every row carries a disposition and target (zero blank). Sweep - B grep-verified all 45 executed sheet rows against their landed hunks — 45/45 PASS - (script: criterion1_sweep.sh, session scratchpad; output in the session record). Zero - unaccounted rows in either direction. -- **Criterion 2 (gates):** per-phase gates ran green before each commit; - check-changed-skills.sh origin/main = 53 skills, 0 failed; check-purged-em-dashes.sh - clean; markdownlint 0 issues and typos clean across all 34 changed markdown files; - `node scripts/generate-cheatsheet.mjs --check` exit 0 (frontmatter argument-hint does not - feed the sheet, so no regeneration was needed); ai-slop detector 0 findings across the - changed set, rubric pass on the F1 doc clean. -- **Criterion 3:** F1 carries the codification warning byte-faithfully with citation stamp - and the C6 fair-quotation basis (grep-verified in sweep B: C3-warn, C3-basis). -- **Criterion 4:** two names-and-points registry rows landed pointing at F1's owner - sections, which carry the conformance surfaces (the Phase-2 reconciliation: the - two-column registry never restates; criterion satisfied by owner-section text). -- **Criterion 5:** ctx-eng topic PLAN carries the three "Open, new" candidate bullets and - the Phase-10 rebase note; no phase headings touched (grep-verified). -- **Criterion 6:** every W2/W3 contract delta's eval expectations landed in the same - commit as its contract lines (verifiable via `git log --name-only` per phase commit). -- **Criterion 7:** this PLAN + ./signoff-sheet.md were the executed contract, phase by - phase. -- **Follow-up issues filed:** #3589 (behavioral eval candidates), #3590 (E4 prd pitch-view - extension), #3591 (recorded triggers: D28 schema, Q11 flip, E1 registry row, V7.7 - digest-pipeline lessons). -- **PR:** (the single - PR the operator directed; commits map to phases 1-11). -- **Full battery:** `scripts/affected-tests.sh --run` selected the wide suite set (the - plugin.json edits fan out); 2569 assertions passed, one suite failed — the - `interview-defenses` digest ratchet caught D28's paragraph inside its digested register - section and the eval expectation in a case where it was ungradeable. Resolution per that - suite's own contract: defenses re-read and confirmed intact, the expectation moved to - case 1, digests refreshed in the same change; suite re-run green (89/0). The 13 - NOT-RUN-ecosystem suites were then run from their own lane: 785 passed + 330 subtests - (the one PowerShell suite is Windows-lane, CI covers it). - -## Blast radius - -MEDIUM-HIGH. Six plugins' skill contracts change in one PR plus two governance docs and a -mid-flight sibling topic's PLAN. Mitigations: every delta is additive prose (no schema, -no executable surface); the parsed schemas in reach are explicitly frozen (Part G rule 4); -per-phase gates run before each commit; the sibling-topic edit is bullets-only. - -## Stress-test summary - -The execution shape this plan sequences was already adversarially tested this session -before sign-off: /planning:devils-advocate (1 CRITICAL / 4 HIGH / 6 MEDIUM / 4 LOW, all -folded), then two independent fresh-context Fable validators over the sign-off sheet -(amendments folded as rev 2). - -2026-09-01 — Step 3 fresh-context plan review over THIS phase plan returned 3 CRITICAL / -6 IMPORTANT / 6 SUGGESTION; all verified against the repo and folded: (1) Q7/Q10 -cross-refs gained homes (wayfind in Phase 7, session-flow as new Phase 9, artifact-design -resolved as prose mention in F1); (2) criterion-1 traceability now runs corpus→sheet off -a committed disposition ledger (Phase 2); (3) affected-tests was the wrong gate for -SKILL.md/docs edits — replaced with check-changed-skills.sh, direct markdownlint, the -em-dash ratchet, and the direct check-open-questions test; (4-6) Part G gates wired into -every phase, criterion-4 reconciliation recorded, criteria 2-7 walk added to Phase 11; -(7) F1-as-owner justification extended; (8) map/territory gets a Rejected-terms row; -(9) delta-resolution.md committed for compaction survivability; plus the six suggestions -(before/after eval counts, grep idioms, real test path, provisional-CHANGELOG note, -ai-slop audit pass, unconditional lychee check). - -## Execution shape - -Fully sequential, all main-session, phases 1→11 in order (3-9 are order-independent among -themselves but run sequentially anyway; commits are serial on one branch). Basis for -main-session routing: the repo's on-demand convention surfaces (AGENTS.md table) do not -auto-load inside subagents; every phase is judgment-heavy house-style contract writing; -token budget is ample. Sub-agents are used only for fresh-context REVIEW (Step 3 reviewer; -any fresh-eyes checkpoint the executor requests), never for authoring. Sequential fallback -is the shape itself — no parallel orchestration to fall back from. - -| Phase | Surface | Basis | -|---|---|---| -| 1-11 | main session | convention surfaces + house style live in main context; serial commits on one branch | -| Step-3 reviewer | fresh sub-agent | mandatory fresh-context stress-test | - -## Open questions - -None blocking — every decision is signed (./signoff-sheet.md). Deferred items live in the -Brief's Deferred questions with triggers and arbiters. - -## Handoff to implementation - -### User-approval gates - -None remaining: the operator signed the full decision surface (".confirm all") and -explicitly directed all execution into this session, one branch -(claude/reading-feedback-j4sg96), one PR. Any mid-flight pivot that would change an -acceptance criterion still stops and asks. - -### Execution shape ([EXEC-SHAPE] tagged) - -Decisions made (gate-passed): - -| Decision | What it changes in the plan | Basis (evidence) | -|---|---|---| -| [EXEC-SHAPE] F1 doc path = `docs/FINDING-YOUR-UNKNOWNS.md` | Phase 1 target file | SCREAMING-KEBAB at docs/ root is the read precedent for durable cross-cutting references (docs/ inventory read this session); "name at plan time" was delegated by Part E | -| [EXEC-SHAPE] Per-plugin DOC rows (D9, D18, D27, D35, E8) ride their plugin's W2 commit instead of a separate W1 commit | Phases 3-8 contents | MIGRATION-PLAYBOOK: version bump + CHANGELOG is per plugin; folding avoids two bumps per plugin in one PR; wave intent (review structure) is preserved by commit ordering | -| [EXEC-SHAPE] Sequential main-session execution, no parallel fan-out | Execution shape | Convention surfaces don't auto-load in subagents (AGENTS.md); serial commits on one branch; judgment-heavy prose work | -| [EXEC-SHAPE] Registry rows point at the F1 doc as owner (form-2: repo file) | Phase 2 | Registry precedent allows repo-file owners (`lib/hook-utils.sh`, plugin surfaces); Part E assigns ownership of both conventions to the F1 doc; the owner sections carry the full owner-doc anatomy (rule, who is bound, conformance) and versioning rides git history exactly as it does for the existing non-directory owners | -| [EXEC-SHAPE] Deferred/eval-candidate tracking as GitHub issues, granularity at executor's judgment | Phase 11 | Operator: "Its OK if we file issues" | -| [EXEC-SHAPE] Q10's artifact-design cross-ref = a prose mention of the session built-in skill inside F1's reply-affordance section | Phase 1 | artifact-design has no in-repo surface (verified repo-wide); prose mention of built-ins is the established idiom | -| [EXEC-SHAPE] Q7's session-flow:workflow cross-ref is its own phase/commit with a doc-line-only version bump | Phase 9 | The cross-ref is a Part B obligation, not a Part D roster row; MIGRATION-PLAYBOOK requires the bump for any plugin content change | - -### Mechanical work - -- One commit per phase (Phases 1-2 may merge into one docs commit if small; never split a - plugin's bump across commits). Commit messages via `git commit -F - --cleanup=verbatim` - heredoc with the session's attribution footer. -- Per-commit verification: the phase's Sanity Check plus `scripts/affected-tests.sh --run`. -- Push with `git push -u origin claude/reading-feedback-j4sg96` (retry with backoff on - network failure only). -- PLAN.md phase tags advance `[TODO]` → `[DOING]` → `[DONE]` in the same commit as the - phase's changes. diff --git a/docs/topics/finding-your-unknowns-integration/delta-resolution.md b/docs/topics/finding-your-unknowns-integration/delta-resolution.md deleted file mode 100644 index 24710e89c..000000000 --- a/docs/topics/finding-your-unknowns-integration/delta-resolution.md +++ /dev/null @@ -1,88 +0,0 @@ -# Delta resolution — conditional verdicts resolved against grader evidence - -Committed copy of the session's delta wording record (originally a work-slice artifact); -the row wording here is what the wave phases implement. Row classifications and wave -assignments were finalized in ./signoff-sheet.md Part D, which supersedes the Status -column below where they differ (e.g. D5/D34 reclassed behavioral-tier, D13 dropped). - -2026-09-01. Inputs: six grader reports in evidence/ (file:line evidence there), the -round-1/2 interview locks, corpus-inventory.md. Every row cites its grader. Statuses: -DELTA (change to make), CORROBORATION (already present; optionally cite), NO-CHANGE -(present and stronger than corpus), REROUTE (corpus aimed at wrong home), RETURN -(genuine fork back to human/audit). All DELTA rows are subject to PLUGIN-PHILOSOPHY -"evidence-gated additions": each lands as a sourced, corpus-cited contract line with -the gate acknowledged, or waits for observed-stumble evidence per the audit's call. - -## D-block: V2 technique deltas - -| ID | Target | Change | Status | Grader | -|---|---|---|---|---| -| D1 | discovery:blindspot | typed finding taxonomy (Landmine/History/Convention/Missing-concept) as output contract | DELTA | blindspot-brainstorm A1 | -| D2 | discovery:blindspot | per-finding prompt-fix | CORROBORATION (present SKILL.md:59) | A2 | -| D3 | discovery:blindspot | fold-step: add explicit human confirm-before-final checkpoint | DELTA (partial today) | A3 | -| D4 | discovery:blindspot | output requires scan-scope disclosure line | DELTA (partial) | A4 | -| D5 | planning:brainstorm | tighten per-option evidence to observed-fact (falsifiable), not just path | DELTA (partial) | B5 | -| D6 | planning:brainstorm | cheapest-to-ambitious ordering | CORROBORATION | B6 | -| D7 | planning:brainstorm | named already-built-but-disconnected scan heuristic (dead imports, dark flags, unread tables) | DELTA (partial) | B7 | -| D8 | planning:brainstorm | structured closing pick | CORROBORATION | B8 | -| D9 | planning:brainstorm | session-start-brainstorm citable rationale line (Q8 lock: only-if-absent; absent) | DELTA (one line) | B9 | -| D10 | improvement:find | corpus S/M/L/XL axis | NO-CHANGE (WSJF+size stronger; adopting would regress) | C10 | -| D11 | improvement:find | disconnected-work recipe file (like hotspots.md) | DELTA (optional, small) | C11 | -| D12 | education:explain | vocabulary ladder (term + definition + modeled "say ->" sentence) | DELTA | edu A1 | -| D13 | education:explain | payoff-prompts closing (before/after prompt contrast) | DELTA | A2 | -| D14 | education:explain | success condition: "user's next prompt names what they mean" | DELTA | A4 | -| D15 | education:explain | three-tier restructure | NO-CHANGE (corpus itself marks hypothesis unvalidated; existing altitude tiers stay) | A3 | -| D16 | education:quiz-me | per-question source-anchor + on-miss routing to the skimmed section | DELTA | B6 | -| D17 | education:quiz-me | diff-sourced question authoring keyed to non-obvious behaviors | DELTA (partial+absent merged) | B7+B9 | -| D18 | quiz-as-merge-gate | REROUTE: quiz-me disclaims merge-gating twice by design; verification:confirm owns the PR gate (SKILL.md:119). Any merge-gate framing lands as a confirm cross-ref, never a second gate | REROUTE | edu B8xC11 | -| D19 | verification:confirm | explicit "existing behavior this leans on" out-of-diff coupling callout | DELTA (partial today) | C10 | -| D20 | prototype:explore-directions | same-data control-variable rule for mockup substrate | DELTA (partial) | proto A1 | -| D21 | prototype:explore-directions | structured steal/graft capture at single-decision granularity | DELTA (partial) | A3 | -| D22 | prototype:explore-directions | machine-legible assembled-reply template (direction/steal/skip/next-target) | DELTA | A4 | -| D23 | prototype:explore-directions | direction count | NO-CHANGE (default 3 cap 5 beats corpus's fixed 4) | A5 | -| D24 | prototype:pressure-test | validation-answer-set output shape (bounded fillable forced-choice answers) | DELTA | B6+B7 | -| D25 | prototype:pressure-test | fake-data/no-real-wiring disclosure footnote (non-dev audience risk) | DELTA (high value) | B8 | -| D26 | prototype:pressure-test | per-option named costs on forced-choice questions | DELTA | B9 | -| D27 | prototype pair | "mock before you wire" named ordering note in composition table | DELTA (small) | B11 | -| D28 | planning:interview | flag free-text answers for downstream scrutiny | DELTA (small) | plan-grader 2 | -| D29 | planning:interview | decisions-table / assembler | NO-CHANGE (register+arbiter cover it; Brief prose is by design) | 3+4 | -| D30 | planning:questionnaire | none | NO-CHANGE (intentional non-adopter: "interview the send, not the subject") | 5 | -| D31 | planning:plan | tweak-likelihood ordering | CORROBORATION (knob already at SKILL.md:223; Q11's "mode" is status quo) | 6+7 | -| D32 | planning:plan | alternatives carry a one-line switch condition | DELTA (partial) | 8 | -| D33 | planning:plan | closing pre-drafted revision replies tied to flagged decisions | DELTA | 9 | -| D34 | planning:plan | self-check before collapsing mechanical sections | DELTA (partial) | 10 | -| D35 | planning:plan | forward-reference implement's lands-green guarantee from Sanity Check | DELTA (one line) | 11 | -| D36 | planning:design | mirror plan's tweak-likelihood knob in Phase-5 discussion rounds | DELTA | 13 | -| D37 | PLAN.md schema | constraint: never rename `### Phase N` heading/tag vocabulary without version bump; block reordering is safe | CONSTRAINT (fact) | 12 | - -## E-block: V3/V4 deltas - -| ID | Target | Change | Status | Grader | -|---|---|---|---|---| -| E1 | implement pipeline | extend deviation-log convention (taxonomy: plan-confirmed/discovery/deviation/human-decision; plan-said/found/chosen fields; conservative default; blocking markers) from autonomous+Moderate to the interactive path | DELTA (the V3 headline) | port-impl 5+7 | -| E2 | implement Step 5 | required fold-back: read DEVIATIONS.md, emit plan-amendment bullets | DELTA | 6 | -| E3 | session-flow retro/handoff | live-vs-posthoc | NO-CHANGE (complementary by design; E1/E2 close the gap at the right home) | 8+9 | -| E4 | buy-in doc home | design-handoff has 0/4 persuasion elements; prd ships an HTML pitch view (closest precedent); visualize is static (demo routes via playwright/run) | RETURN (fork: extend prd pitch view vs. extend design-handoff vs. new thin skill slot) | 10+11+12 | -| E5 | buy-in components | standalone objection-evidence checklist (question+answer+evidence citation) | DELTA (home follows E4) | corpus f7e96d2a | -| E6 | discipline:point-dont-copy | externalized semantics-map artifact + confirmation gate for EXTERNAL-REFERENCE PORTS ONLY (in-tree corrector doctrine explicitly forbids stop-and-wait; scoping avoids the reversal) | DELTA scoped + RETURN flag (audit must confirm the scoping is clean) | 1+2 | -| E7 | discipline:point-dont-copy | trap check: source primitive with no target analogue -> name the carrying convention | DELTA | 3 | -| E8 | discipline:point-dont-copy | canonical invocation example line | DELTA (one line) | 4 | - -## F-block: V5/V6 + reference doc - -| ID | Target | Change | Status | Grader | -|---|---|---|---|---| -| F1 | graduated reference doc | new docs/ reference: unknowns taxonomy + lifecycle + pattern catalog + reply-affordance convention + export-button rule + 9-category when-HTML taxonomy + the skill-codification warning quoted; citations per plugins/knowledge/reference/citation-shape.md | DELTA (the Q7/Q9/Q10 vehicle; artifact-design/capabilities are session built-ins, not repo files, so conventions land here) | gov 10-12 | -| F2 | V6 routing | ALL context-engineering deltas route into docs/topics/context-engineering-claude-5/ open phases (article already decomposed, corroborated, gated; audit-instructions I6/I15 already cite it) | REROUTE (supersedes V6 entries 2,3; genericness check and I29-widening and /doctor cross-ref become candidate inputs to that topic, not this one) | gov 1-9 + surprise | -| F3 | claude-memory:audit | gotcha-vs-obvious heuristic | CORROBORATION (C2 Deletion Test + C5) | gov 6 | -| F4 | quiz-me fresh-eyes tension | quiz author self-grades in biased context vs. philosophy's fresh-eyes doctrine | RETURN (design question beyond corpus scope; surfaced to human) | edu surprise 2 | - -## Standing constraints carried into every delta - -1. PLUGIN-PHILOSOPHY evidence-gated additions (:699-711): corpus-anticipated is not - observed-stumble; each delta lands citing the corpus as source AND acknowledging the - gate, with the audit deciding per-delta whether the gate demands deferral. -2. One-mechanism-per-concern (:678-679): D18 reroute is the enforcement example. -3. MIGRATION-PLAYBOOK version pinning: D37; any parsed-schema touch needs changelog. -4. Two-lane convention posture: no hardcoded decision-record formats or literal - path/flag templates in conventions (proto grader note). diff --git a/docs/topics/finding-your-unknowns-integration/design/design-resolution.md b/docs/topics/finding-your-unknowns-integration/design/design-resolution.md deleted file mode 100644 index 853cbe91f..000000000 --- a/docs/topics/finding-your-unknowns-integration/design/design-resolution.md +++ /dev/null @@ -1,19 +0,0 @@ -# Design resolution — finding-your-unknowns-integration - -outcome: early-exit - -Tier C under /planning:plan's design-significance gate. This effort introduces no new -types, modules, package topology, or data models: every change is markdown contract text -in existing skill bodies, evals.json expectation strings, per-plugin CHANGELOG/version -metadata, and new reference documentation under docs/. The one structural artifact (the -F1 reference doc) follows an existing precedent shape (top-level SCREAMING-KEBAB doc with -a Contents TOC, per docs/PLUGIN-PHILOSOPHY.md), so no design exploration is warranted. - -Type sketch: none needed — no executable surface changes. The only parsed-schema surfaces -in reach (the `### Phase N` heading/tag vocabulary; check-open-questions.sh's 5-field -register rows) are explicitly frozen by the signed contract (signoff-sheet Part G rule 4 / -D37): the plan adds prose around them and never renames them. - -Design-tier decisions were resolved upstream by the signed decision chain: interview -rounds 1-3, evidence pass, dual validators, devils-advocate, final validators, operator -sign-off (../signoff-sheet.md, 2026-09-01). diff --git a/docs/topics/finding-your-unknowns-integration/disposition-ledger.md b/docs/topics/finding-your-unknowns-integration/disposition-ledger.md deleted file mode 100644 index f872bc261..000000000 --- a/docs/topics/finding-your-unknowns-integration/disposition-ledger.md +++ /dev/null @@ -1,97 +0,0 @@ -# Disposition ledger — corpus decision → executed-or-recorded outcome - -The acceptance-criterion-1 crosswalk, in the corpus→sheet direction: one row per decision -in the corpus inventory (V-ids; the inventory itself lives in the session work slice), -each naming its signed-sheet row(s) and disposition. Sheet: ./signoff-sheet.md. Wording -record: ./delta-resolution.md. No row may be blank. - -## V1 — framing and vocabulary - -- V1.1 quadrant taxonomy | adopt doc-tier | F1 taxonomy section + GLOSSARY "unknowns quadrants" -- V1.2 map/territory | cite-only | G1; GLOSSARY rejected-terms row -- V1.3 bottleneck thesis | treat-as-caution | F1 "Why this exists" caution line -- V1.4 over/under-specify diagnostic | doc line | G2 (F1 taxonomy section) -- V1.5 disclose-starting-point primer | adopt doc-tier | G3 (F1 pattern catalog) -- V1.6 cost framing | adopt doc-tier | G4 (F1 intro quote) -- V1.7 long-horizon diagnostic | doc line | G5 (F1 taxonomy section) - -## V2 — pre-implementation techniques - -- V2.1 blindspot block | adopt/corroborate | D1 (W2), D2 (corroboration), D3 (demoted to - corroboration), D4 (W2) -- V2.2 teach-me block | adopt scoped | D12 (W2), D13 (dropped — conflicts with explain's - one-line close), D14 (W2, scoped), D15 (no-change), G6 (F1 note) -- V2.3 brainstorm block | mixed | D5/D7 (behavioral → F1 + eval candidates), D6/D8 - (corroboration), D9 (W1 doc line), D10 (no-change), D11 (behavioral → F1) -- V2.4 design-directions block | adopt | D20-D22 (W2), D23 (no-change — repo default stronger) -- V2.5 mock-before-wire block | adopt | D24-D26 (W2), D27 (W1 doc note) -- V2.6 interview block | adopt small | D28 (W2 resolution-field convention), D29/D30 (no-change) -- V2.7 reference-port block | adopt scoped | E6 (W3), E7 (W2), E8 (W1 doc line) -- V2.8 tweakable-plan block | adopt/corroborate | D31 (corroboration + reconciliation line), - D32/D33 (W2), D34 (corroboration + eval candidate), D35 (W1 doc line), D36 (W2) -- V2.9 sequencing meta | adopt doc-tier | Q7: F1 workflow section + wayfind and - session-flow:workflow cross-refs; G15 - -## V3 — during implementation - -- V3.1 implementation-notes convention | adopt opt-in | E1+E2 (W3), E3 (no-change); - F1 deviation-log owner section + recorded registry trigger (C5) -- V3.2 fresh-session-per-phase | corroboration | G7 (F1 cite of session-flow doctrine) - -## V4 — post-implementation - -- V4.1 buy-in doc | adopt doc-tier | S4/C2: F1 buy-in section + E5 checklist; skill - extension deferred behind demand evidence (Brief deferred E4-EXT) -- V4.2 quiz-as-merge-gate | reroute + adopt | D16/D17a/F4 (W2), D17b (behavioral → F1), - D18 (reroute to verification:confirm cross-ref, W1), D19 (W2) - -## V5 — HTML-artifact methodology - -- V5.1 governing caution | BINDING constraint | Q2; quoted verbatim in F1 cautions -- V5.2 stay-in-the-loop | adopt as lens | G8 (F1 cautions quote) -- V5.3 nine-category taxonomy | adopt doc-tier | F1 when-HTML section -- V5.4 density rubric | doc line | G9 -- V5.5 ~100-line ceiling | recorded anecdote | G10 (labeled practitioner anecdote) -- V5.6 sharing argument | recorded | G11 (satisfied by the Artifact tool; one F1 line) -- V5.7 export-button doctrine | adopt convention | F1 owner section + registry row -- V5.8 throwaway-editor doctrine | adopt as caution | G12 (F1 cautions) -- V5.9 HTML-diff noise | adopt scoping rule | G13 (F1 when-HTML scoping rule) -- V5.10 design-system/PR-explainer/report patterns | doc-tier | G14 (F1 pattern entries; - skill edits deferred behind the evidence gate) -- V5.11 workflow shape | covered | G15 (F1 workflow section) - -## V6 — context-engineering companion (F2 reroute) - -All V6 decisions route to docs/topics/context-engineering-claude-5/ (F2); this effort -touches none of that topic's criteria files. - -- V6.1 80% system-prompt claim | treat-as-caution | recorded; live-doc-checked in the - evidence pass; never a pruning license here -- V6.2 conflicting-instructions example | corroboration | audit-instructions I6/I15 - already cite the article -- V6.3 S1→S2 rewrite pattern | rerouted | that topic's check-design phases own it -- V6.4 examples→interfaces | recorded, scoped | covers tool usage examples, not skill - trigger phrases; no action here -- V6.5 progressive-disclosure myth | corroboration | existing docs-hygiene audit + - AGENTS.md load-on-demand table; cite-only -- V6.6 repetition myth | candidate recorded | ctx-eng PLAN "Open, new" bullet - (I29 widening) -- V6.7 auto-memory myth | rerouted | claude-memory scoping falls under that topic's - /doctor + memory-audit work -- V6.8 rich-references myth | rerouted | that topic's check-design phases -- V6.9 genericness lens | candidate recorded | ctx-eng PLAN "Open, new" bullet - (skill-quality check) -- V6.10 claude doctor positioning | candidate recorded | ctx-eng PLAN "Open, new" bullet - (/doctor cross-ref) - -## V7 — corpus and process meta - -- V7.1 citation routing | executed | Brief captured assumptions; F1 sources S1/S4 split -- V7.2 companion-pair rule | executed | the two handoffs interviewed as one package -- V7.3 cc-applicable tags | executed | Q5: ~6 decision-relevant claims live-doc-checked; - rest treated as vendor anecdote -- V7.4 videos | out of scope | Brief out-of-scope list -- V7.5 named-expert anecdotes | dropped | G16 (absent from all graduated artifacts) -- V7.6 frontend-plugin pointer | recorded | verified absent from this repo; nothing to compare -- V7.7 verifier MINORs + pipeline lessons | recorded | correction records in the work - slice; candidate follow-up issue at close-out (Phase 11) diff --git a/docs/topics/finding-your-unknowns-integration/signoff-sheet.md b/docs/topics/finding-your-unknowns-integration/signoff-sheet.md deleted file mode 100644 index 33e81e4b5..000000000 --- a/docs/topics/finding-your-unknowns-integration/signoff-sheet.md +++ /dev/null @@ -1,219 +0,0 @@ -# Final sign-off sheet — finding-your-unknowns integration (rev 2) - -2026-09-01. THE single review artifact for this topic, and (with the Brief in -./PLAN.md) the input contract /planning:plan consumes. Rev 2 folds both final -validators' challenges; every amendment is tagged [A]/[B]. - -**SIGNED: operator ".confirm all", 2026-09-01.** Every "Default: CONFIRM" row below -is confirmed; Parts A-G are adopted in full. - -Chain of custody: corpus (17 verified digest slices, slice -`finding-your-unknowns-0f25bd45`) -> interview locks -> evidence pass (7 graders + -targeted live-doc checks) -> validators A/B -> explore+research (fresh-verified) -> -blindspots -> brainstorm -> devils-advocate -> this sheet -> final validators A/B -> -rev 2 -> operator sign-off. The session working files behind each stage lived under -`.work/finding-your-unknowns-integration/` (unversioned by design); this committed -copy is the durable record. - -## Part A — structural fixes (confirmed) - -- S1 [DA-C1; refined per A+B] Evidence-gate compliance. Reading stated explicitly: - "standing instruction" = always-loaded contract text in a skill body. Every delta - is classified per-row under PLUGIN-PHILOSOPHY:699-742. CONTRACT/POLICY/CONVENTION - rows land now under the durable-tier carve-out as TEAM CONVENTIONS ADOPTED BY THIS - SIGN-OFF, citing corpus + industry grounding. DOC rows are citation/doc lines. - BEHAVIORAL rows never land as instructions: they become F1 doc lines + candidate - EVAL cases (evals outlive instructions), awaiting observed-stumble evidence. - Reclassified in rev 2 on validator evidence: D5 -> BEHAVIORAL [A+B], D34 -> - CORROBORATION + behavioral residue -> eval candidate (plan Steps 4.6/4.7 already - own the invariant) [A+B], D17 split into D17a contract / D17b behavioral [A]. -- S2 [DA-H1] Registry justification: owner-doc-before-a-SECOND-ADOPTING-PLUGIN rule - (PHILOSOPHY:594-595). Applies to reply-affordance + export-button only (see E-part; - E1's row is retracted per M3 [B]). -- S3 [DA-H2 + L4] Sequencing with docs/topics/context-engineering-claude-5/: - (a) this effort never edits the audit-instructions criteria files (zero overlap, - stated); (b) Wave 1 records F2's three candidates in that topic's PLAN; - (c) one line there: Phase 10's sweep re-inventories current state (its recorded - counts are stale [L4]) and rebases over landed waves. No freeze. -- S4 [DA-H3] E4 closure: five-section buy-in pattern + E5 objection-evidence - checklist land in the F1 DOC (durable, tracked; cites Rust RFC / Oxide RFD / - Amazon PR-FAQ + corpus). Skill extension (prd durable-output mode vs - design-handoff layer) deferred as a recorded candidate behind demand evidence. -- S5 [DA-H4] This sheet is the complete decision surface; rev 2 adds the rows both - validators found missing (F4, D18, D37, D3/D13 dispositions) and the M2 - eval-impact column. -- S6 [DA downgrade, verified] Em-dash reconciliation unnecessary: rule-em-dash is - disabled repo-wide (.claude/ai-slop.json, recorded rationale). F1 quotes stay - byte-verbatim; one F1 header line records the quoting posture [M1]. - -## Part B — interview locks (human-confirmed rounds 1-2; record) - -- Q1-Q5: vehicles; binding codification posture; conditional verdicts + evidence - pass; all seven verticals V2-first; targeted live-doc checks. -- Q6-Q12: batch execution contract; workflow section + cross-refs; no - brainstorm-first posture; central pattern catalog + one canonical line per skill; - reply-affordance convention (default-with-judgment); tweak-likelihood discharged - (plan already defaults to presentation ordering — reconciliation line in F1); - stop-and-wait port gate, token recommended-not-required. - -## Part C — open questions signed with this sheet (Q13-Q17 + amendments) - -- C1 [Q13] Merged round-3 fixes as amended by rev 2: D3 demoted; D13 dropped + - D14 scope-conditioned; D31 reconciliation line; D11 restored as doc-tier; - G-block adopted; plus rev-2 reclasses (S1). CONFIRMED. -- C2 [Q14/E4] Buy-in home per S4 (doc now, skill later on evidence). CONFIRMED. -- C3 [Q15/E6] Port gate in point-dont-copy via declared-step-deltas, scoped to - external-reference ports ("source of truth outside this repo's tree: vendored, - foreign-language, other-repo"); in-tree corrections stay do-it-now. - CONFIRMED with boundary sentence verbatim. -- C4 [Q16/F4] quiz-me answer-key fresh-context requirement — a classified row - (Part D, CONTRACT, W2) [A+B]. CONFIRMED. -- C5 [Q17/E1] Deviation-log ships as OPT-IN convention: F1 owner section + - implement-pipeline reference + RECORDED-TRIGGER NOTE in lieu of a registry row - ("the moment a second PLUGIN reads DEVIATIONS.md, the registry rule fires") — - rev 2 retracts the premature registry row per M3 [B; A concurs]. - CONFIRMED opt-in with trigger note. -- C6 [licensing, blindspot 4 restored in full [B]] F1's header records the - permission basis: short attributed verbatim excerpts from the named author's - public posts under fair-quotation practice; no license claimed; citation-shape - citations; bulk reproduction avoided; x.com lychee excludes. CONFIRMED. - -## Part D — classified delta roster (eval-impact column per M2) - -Waves: W1 = docs/registry/glossary/sequencing. W2 = CONTRACT/POLICY rows landing in -skill bodies (6 plugins). W3 = rows needing their own review moment (E6, E1+E2) — -the stated exception to the W2 mapping [B]. Format: -row | target | delta | CLASS | wave | eval impact. - -- D1 | discovery:blindspot | typed finding taxonomy | CONTRACT | W2 | extend -- D4 | discovery:blindspot | scan-scope disclosure line | POLICY | W2 | extend -- D5 | planning:brainstorm | observed-fact evidence bar | BEHAVIORAL -> F1 + eval - candidate [A+B reclass] | W1(doc) | eval-candidate -- D7 | planning:brainstorm | disconnected-work heuristic | BEHAVIORAL -> F1 + eval - candidate | W1(doc) | eval-candidate -- D9 | planning:brainstorm | brainstorm-practice citation line | DOC | W1 | skip -- D11 | improvement:find | disconnected-work recipe | BEHAVIORAL -> F1 + eval - candidate | W1(doc) | eval-candidate -- D12 | education:explain | vocabulary ladder | CONTRACT | W2 | extend -- D14 | education:explain | success condition (scoped to original-ask invocations) | - CONTRACT | W2 | extend -- D16 | education:quiz-me | source-anchor + on-miss routing | CONTRACT | W2 | extend -- D17a | education:quiz-me | diff-sourced question authoring | CONTRACT | W2 | extend -- D17b | education:quiz-me | non-obvious-behavior keying | BEHAVIORAL -> F1 + eval - candidate [A split] | W1(doc) | eval-candidate -- D18 | verification:confirm | quiz-layer cross-ref line (the REROUTE executed; - quiz-me untouched) [B restore] | DOC | W1 | skip -- D19 | verification:confirm | "existing behavior this leans on" callout | CONTRACT | - W2 | extend -- D20 | prototype:explore-directions | same-data control-variable rule | POLICY | - W2 | extend -- D21 | prototype:explore-directions | structured steal/graft capture | CONTRACT | - W2 | extend -- D22 | prototype:explore-directions | machine-legible reply template | CONTRACT | - W2 | extend -- D24 | prototype:pressure-test | validation-answer-set shape | CONTRACT | W2 | extend -- D25 | prototype:pressure-test | fake-data disclosure footnote | POLICY | W2 | extend -- D26 | prototype:pressure-test | per-option named costs | CONTRACT | W2 | extend -- D27 | prototype pair | mock-before-wire composition note | DOC | W1 | skip -- D28 | planning:interview | free-text scrutiny flag AS RESOLUTION-FIELD CONVENTION - (branch chosen on script evidence: 5-field enum, check-open-questions.sh:31-32, - :173; gate-invisible BY DESIGN — known limitation recorded; schema change deferred - until a consumer needs mechanical reads) [M4 resolved] | CONTRACT | W2 | extend -- D32 | planning:plan | switch condition on alternatives | CONTRACT | W2 | extend -- D33 | planning:plan | closing revision replies | CONTRACT | W2 | extend -- D34 | planning:plan | collapse self-check | CORROBORATION (Steps 4.6/4.7 own it) - - behavioral residue -> eval candidate [A+B reclass] | W1(doc) | eval-candidate -- D35 | planning:plan | lands-green forward-reference | DOC | W1 | skip -- D36 | planning:design | design gains the same tweak-likelihood presentation - ordering plan already documents (SKILL.md:223) [A reword] | CONTRACT | W2 | extend -- E7 | discipline:point-dont-copy | primitive->convention trap item | CONTRACT | - W2 | extend -- E8 | discipline:point-dont-copy | canonical invocation line | DOC | W1 | skip -- F4 | education:quiz-me | fresh-context answer-key requirement [A+B add] | - CONTRACT | W2 | extend -- E6 | discipline:point-dont-copy | semantics map + confirmation gate - (external-reference ports, boundary per C3) | CONTRACT | W3 | extend -- E1+E2 | implement pipeline | opt-in deviation-log convention + fold-back step - (per C5; trigger note, no registry row) | CONVENTION | W3 | extend -- E5 | F1 doc | objection-evidence checklist section | DOC | W1 | skip -- Recorded dispositions, no change: D2, D3 (demoted), D6, D8, D10, D13 (dropped), - D15, D23, D29, D30, D31 (+reconciliation line), E3, F3 [A coverage fix]. -- D37 | PLAN.md schema constraint -> Part G rule 4 [B restore]. - -Eval pricing [M2]: W2 touches 6 plugins; the planning-plugin PR concentrates 4 -contract deltas (D28, D32, D33, D36) + their expectations — budget it as its own -review; education PR carries 5 (D12, D14, D16, D17a, F4). "Extend" means the -skill's evals.json gains expectations for the new contract lines in the same PR; -authorship = the executing wave session; review = the PR reviewer. - -## Part E — doc/registry/glossary placements (Wave 1) - -- F1 doc at docs/ (name at plan time): taxonomy + lifecycle; pattern catalog; - reply-affordance convention + export-button rule; when-HTML taxonomy; buy-in - pattern + E5; codification warning quoted; HTML-scoping rule; C6 permission - basis; D5/D7/D11/D17b/D34-residue heuristics as doc lines; Q11/D31 - reconciliation line. -- Registry rows (S2): reply-affordance and export-button ONLY, each with explicit - conformance surface [M5]: "conformance = the template blocks in - prototype:explore-directions (D21/D22) and prototype:pressure-test (D24)", - landing owner-doc-first with adopters in W2. -- Glossary: curate-language dispositions for unknowns quadrants + D1's four finding - types; map/territory cite-only (G1). -- G1-G16: as merged in the round-3 audit (audit/merge-2026-09-01.md rows G1-G16). - CONFIRMED en bloc. -- Sequencing rows per S3 into the ctx-eng topic PLAN. - -## Part F — devils-advocate finding dispositions (per-finding [A]) - -- M1 em-dash: F1 header line only; no config change; addendum line 3 dropped. DONE - in rev 2 (S6). -- M2 eval obligation: eval-impact column + pricing added (Part D). DONE. -- M3 E1 registry row: retracted; recorded-trigger note (C5). DONE. -- M4 D28 branch: chosen on script evidence (Part D row). DONE. -- M5 registry conformance surfaces: stated per row (Part E). DONE. -- M6 Brief completeness: sign-off closed interview Steps 3-5; acceptance criteria - (Part G) written into the Brief at sign-off. DONE. -- L1 cap headroom: verified ample; no action. L2: fixed by F4/E5 rows. L3: - cheat-sheet contention — wave PR series stay sequential; regenerate once per - series (Part G). L4: folded into S3. - -## Part G — execution shape + definition of done - -1. Three waves; W2 = 6 plugins (discovery, planning, education, verification, - prototype, discipline) [B arithmetic fix]; W1 = docs + registry + glossary + - sequencing + the W1(doc) rows; W3 = E6, E1+E2. -2. Per wave and per plugin: version bump + CHANGELOG entry; - check-changed-skills.sh green (trigger-keyword preservation, listing cap, - --require-evals); check-listing-budget respected [B]; cheat-sheet regenerated - once per series, series sequential [L3]; scripts/affected-tests.sh --run green. -3. Eval expectations land in the same PR as their contract lines (Part D column). -4. [D37] The `### Phase N` heading/tag vocabulary and any parsed schema - (check-open-questions.sh fields) are never renamed without version bump + - changelog; block reordering is safe. -5. Acceptance criteria (written into the Brief at sign-off): (1) every inventory - decision has an executed-or-recorded disposition traceable from this sheet — - verified by grep over the disposition lines, not asserted; (2) all waves green - on rule-2 gates; (3) F1 carries the codification warning + C6 basis; (4) - registry rows carry conformance surfaces and land owner-doc-first; (5) ctx-eng - topic PLAN carries the S3 rows; (6) every W2/W3 contract delta has same-PR eval - expectations; (7) /planning:plan consumes this sheet + Brief as its input - contract. - -## Operator amendment at sign-off (2026-09-01) - -Delivery vehicle amended by the operator with the sign-off: ALL execution happens -in this session, on the single feature branch `claude/reading-feedback-j4sg96`, as -ONE pull request. Consequences, superseding the conflicting phrasing above: - -- The three waves survive as COMMIT ORDERING and review structure inside the one - branch (W1 docs commits, then per-plugin W2 commits, then W3), not as separate - PR series. Filing follow-up issues for deferred items is fine; landed work is not - split across PRs. -- Part G rule 2 gates run on the branch head: per-plugin version bump + CHANGELOG - entries all land in the same PR; cheat-sheet regenerated once at the end (L3's - sequential-series concern is moot with a single branch). -- Part D's eval pricing ("budget the planning PR as its own review") becomes - reviewer guidance for the corresponding commits within the single PR. -- "Same-PR eval expectations" (rule 3, criterion 6) is satisfied trivially by the - single PR, but the intent is kept stricter: eval expectations land in the SAME - COMMIT as their contract lines. From c69c2e9a9425555cf18fbc8e1e7c94cf998667ca Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 19:42:08 +0000 Subject: [PATCH 19/22] fix(planning,prototype): scope the free-text flag and the disclosure footer Two Codex review findings, both verified real. The interview free-text flag now applies only to replies that RESOLVE their question (a complete free-text answer, or an explicit "you pick", which resolves to the recommendation); a partial or non-resolving reply keeps its row open under the drift check, so the flag can never launder a non-answer into a terminal answered row. The pressure-test disclosure footer no longer forces inventing production details: it states wiring location and flag when decided, and says "not decided / no flag planned" explicitly otherwise. The interview-defenses register-section digest is refreshed for the reworded paragraph (defenses re-read: the fix strengthens the no-silent-capture posture; suite green 89/0). Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- plugins/planning/skills/interview/context/loop.md | 2 +- plugins/planning/tests/interview-defenses.test.sh | 2 +- plugins/prototype/skills/pressure-test/SKILL.md | 8 +++++--- 3 files changed, 7 insertions(+), 5 deletions(-) diff --git a/plugins/planning/skills/interview/context/loop.md b/plugins/planning/skills/interview/context/loop.md index 31b8d66c9..a7858879a 100644 --- a/plugins/planning/skills/interview/context/loop.md +++ b/plugins/planning/skills/interview/context/loop.md @@ -206,7 +206,7 @@ Fields: `Q | status | round | question | resolution`. Statuses: `Q` matches the terminal numbering, runs continuously across rounds, and never has a gap — a gap means a row was dropped after it was written, and the gate refuses to grade a register with one. -**Free-text flag — a resolution-field convention.** When an answer arrives as free text rather than a pick from the authored options — the escape hatch, a partial answer, a "whatever you think is best" — lead the resolution field with `free-text:` before the answer. Downstream passes (answer audits, plan formulation) treat flagged rows as deserving scrutiny rather than as settled picks: a free-text answer is where a misread lands silently. This lives inside the free-form resolution field by design; `check-open-questions.sh` grades statuses, not resolutions, so the flag is gate-invisible (a known limitation, recorded here) — a consumer needing mechanical reads of it means a register-schema change, carried by a version bump per the plugin's changelog discipline. +**Free-text flag — a resolution-field convention.** When a reply RESOLVES its question but arrives as free text rather than a pick from the authored options — the escape hatch, a complete answer in the user's own words, an explicit "you pick" (which resolves to the recommendation) — lead the resolution field with `free-text:` before the answer. Downstream passes (answer audits, plan formulation) treat flagged rows as deserving scrutiny rather than as settled picks: a free-text answer is where a misread lands silently. The flag never launders a non-answer into `answered`: a partial or non-resolving reply keeps its row `open` under the drift check below, exactly as if the reply had changed the subject. This lives inside the free-form resolution field by design; `check-open-questions.sh` grades statuses, not resolutions, so the flag is gate-invisible (a known limitation, recorded here) — a consumer needing mechanical reads of it means a register-schema change, carried by a version bump per the plugin's changelog discipline. ### Drift check — a reply that does not answer is not an answer diff --git a/plugins/planning/tests/interview-defenses.test.sh b/plugins/planning/tests/interview-defenses.test.sh index db1137a1c..c500945fc 100755 --- a/plugins/planning/tests/interview-defenses.test.sh +++ b/plugins/planning/tests/interview-defenses.test.sh @@ -483,7 +483,7 @@ pin_section "loop.md open-question register section is unchanged (it binds gaps "$LOOP" \ "## The open-question register" \ "## Step 3 — Recognize the stop condition" \ - "6b3deb652af3048e961b5780e62d7d6fba8ca6edc015c6612b3931844d2cf380" + "1507ecb169de8ff22e11e6fec211906069344d6afa4fcd05e283e456bd6926da" # loop.md carries TWINS of two SKILL.md lines that are byte-pinned there: the # confirmation-gate exemption ("`lock` is exempt … its STOP-on-gap rule still applies") in # Step 3, and the `USER-RESERVED` arbiter guidance in Step 4. A twin with no pin is a diff --git a/plugins/prototype/skills/pressure-test/SKILL.md b/plugins/prototype/skills/pressure-test/SKILL.md index 47add1201..7bd78be53 100644 --- a/plugins/prototype/skills/pressure-test/SKILL.md +++ b/plugins/prototype/skills/pressure-test/SKILL.md @@ -79,9 +79,11 @@ business, not the reducer, because the driver is not reading code: driver's picks are the demo's real output; pair the set with a copy-out control that lifts the filled answers back out as text to paste into the session. 6. **Fake-data disclosure footer**, one visible line stating the page is synthetic end to end, - that nothing on it reads from or writes to the real app, and where the real wiring lives (or - will live) behind which flag. The driver is not reading code; the footer is what keeps a - convincing mock from being mistaken for the wired feature. + that nothing on it reads from or writes to the real app, and, when decided, where the real + wiring lives (or will live) and behind which flag; when integration is not yet decided, or no + flag is planned, the footer says so explicitly rather than inventing production details. The + driver is not reading code; the footer is what keeps a convincing mock from being mistaken + for the wired feature. Constraints (the same set as explore-directions' HTML mockup substrate): From 6e36f3aa76e0592ea600805fa0e329fa791246f9 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 19:48:09 +0000 Subject: [PATCH 20/22] docs(planning): sync the free-text-flag changelog bullet with the narrowed rule Review finding: the 0.35.0 bullet still described the flag as applying to any free-text answer, but c69c2e9a narrowed it to replies that resolve their question, with partial or non-resolving replies staying open under the drift check. The bullet now matches context/loop.md. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- plugins/planning/CHANGELOG.md | 9 ++++++--- 1 file changed, 6 insertions(+), 3 deletions(-) diff --git a/plugins/planning/CHANGELOG.md b/plugins/planning/CHANGELOG.md index 1d4a46603..edfeb81ce 100644 --- a/plugins/planning/CHANGELOG.md +++ b/plugins/planning/CHANGELOG.md @@ -15,9 +15,12 @@ All notable changes to the `planning` plugin are documented here. Format follows implementation's every-step-lands-green expectation. Adopted from the "Finding Your Unknowns" corpus at the integration sign-off (D32, D33, D35; provenance in `docs/FINDING-YOUR-UNKNOWNS.md` in the marketplace repository). Evals extended. -- **`interview`: free-text resolution-field flag.** An answer arriving as free text rather than - an authored option is recorded with a `free-text:` prefix in the register row's resolution - field so downstream passes scrutinize it. Deliberately a convention inside the free-form +- **`interview`: free-text resolution-field flag.** A reply that RESOLVES its question but + arrives as free text rather than an authored option (a complete answer in the user's own + words, or an explicit "you pick", which resolves to the recommendation) is recorded with a + `free-text:` prefix in the register row's resolution field so downstream passes scrutinize + it; a partial or non-resolving reply keeps its row `open` under the drift check, so the flag + never launders a non-answer into `answered`. Deliberately a convention inside the free-form field: `check-open-questions.sh` grades statuses, not resolutions, so the flag is gate-invisible (limitation recorded in `context/loop.md`); the 5-field register schema is unchanged (D28). Evals extended. From 84460c04124fee04c3238dab1683a09361823836 Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 19:48:36 +0000 Subject: [PATCH 21/22] docs(prototype): sync the disclosure-footer changelog bullet with the fix Same staleness the planning changelog had: the 0.10.0 bullet still described the pre-fix footer contract; it now records the explicit "not decided / no flag planned" fallback c69c2e9a added. Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- plugins/prototype/CHANGELOG.md | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/plugins/prototype/CHANGELOG.md b/plugins/prototype/CHANGELOG.md index 04a4e43aa..18e4f5148 100644 --- a/plugins/prototype/CHANGELOG.md +++ b/plugins/prototype/CHANGELOG.md @@ -18,9 +18,11 @@ All notable changes to the `prototype` plugin are documented here. Format follow - **`pressure-test`: validation answer set and a fake-data disclosure footer on the HTML demo shell.** The demo page carries the questions it exists to answer as forced choices whose options each name their cost (plus a free-text escape hatch), with a copy-out control, and a - visible footer stating the page is synthetic end to end and where real wiring lives. The - capture step carries the filled answer set into the durable answer verbatim. Same adoption - basis (D24, D25, D26); evals extended. + visible footer stating the page is synthetic end to end and, when decided, where real wiring + lives and behind which flag; when integration is not yet decided, or no flag is planned, the + footer says so explicitly rather than inventing production details. The capture step carries + the filled answer set into the durable answer verbatim. Same adoption basis (D24, D25, D26); + evals extended. - **Shared discipline: "mock before you wire" ordering note.** The composition table now states that the throwaway mock runs before any real wiring when a change has both questions (D27). From 5f614aebeed1efc2e15d2e6b10c92fa8d2e4a8dc Mon Sep 17 00:00:00 2001 From: Claude Date: Tue, 1 Sep 2026 19:51:33 +0000 Subject: [PATCH 22/22] fix(planning): sync the case-1 eval criterion with the narrowed free-text rule Review finding: the relentless-me-mode-frontier-rounds expectation still graded the pre-fix, over-broad rule (any free-text answer flagged). It now matches the narrowed loop.md contract: a resolving free-text reply gets the flag, a partial or non-resolving reply keeps its row open. The case digest in interview-defenses.test.sh is refreshed in the same change per the suite's contract; suite green (89/0). Co-Authored-By: Claude Fable 5 Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1 --- plugins/planning/skills/interview/evals/evals.json | 2 +- plugins/planning/tests/interview-defenses.test.sh | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/plugins/planning/skills/interview/evals/evals.json b/plugins/planning/skills/interview/evals/evals.json index 170c28b6e..6093739d6 100644 --- a/plugins/planning/skills/interview/evals/evals.json +++ b/plugins/planning/skills/interview/evals/evals.json @@ -12,7 +12,7 @@ "Each question leads with a recommended answer and a one-line basis", "The round is asked inline in prose, not via a side-by-side AskUserQuestion card", "Where a question is answerable from the codebase, the skill resolves it by inspection instead of spending a question on it", - "An answer arriving as free text rather than an authored option is recorded with the free-text: resolution-field flag so downstream passes scrutinize it" + "A reply that resolves its question but arrives as free text rather than an authored option is recorded with the free-text: resolution-field flag so downstream passes scrutinize it; a partial or non-resolving reply keeps its row open" ] }, { diff --git a/plugins/planning/tests/interview-defenses.test.sh b/plugins/planning/tests/interview-defenses.test.sh index c500945fc..c9bd1b1a5 100755 --- a/plugins/planning/tests/interview-defenses.test.sh +++ b/plugins/planning/tests/interview-defenses.test.sh @@ -577,7 +577,7 @@ pin_case_digest "case 12 still refuses to read drift as consent" \ "e45fe64a00cf78e6dd089837e81c52d8a0947ccf924be6472c6f47eeadcd5942" pin_case_digest "case 1 still resolves codebase-answerable questions without asking, and only those" \ "relentless-me-mode-frontier-rounds" \ - "cddd012ee8334177839d116ae564a61a5158f3b9ee996c5903c4fc820b973a36" + "5e38782253a410890320cddc3c5c44d71ab64666bb596bc954167f1447059bdd" pin_file "case A fixture: the task context still plants the open decision" \ "$FIXTURES/lock-stop-on-gap/task-context.md" \