Skip to content

feat: integrate the Finding Your Unknowns corpus — reference doc, conventions, and contract deltas across 8 plugins - #3592

Merged
kyle-sexton merged 23 commits into
mainfrom
claude/reading-feedback-j4sg96
Sep 1, 2026
Merged

feat: integrate the Finding Your Unknowns corpus — reference doc, conventions, and contract deltas across 8 plugins#3592
kyle-sexton merged 23 commits into
mainfrom
claude/reading-feedback-j4sg96

Conversation

@kyle-sexton

@kyle-sexton kyle-sexton commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

No linked issue

Summary

Lands the signed finding-your-unknowns integration in one PR (operator-directed delivery: one session, one branch): the corpus of Thariq Shihipar's "Finding Your Unknowns" field guide, its X-article methodology substrate, and the 20-demo HTML-effectiveness collection, absorbed as judgment-preserving contract deltas on existing skills plus a citable reference doc — never generator skills, per the source author's own warning, quoted byte-faithfully in the doc and treated as a binding constraint throughout. The decision record is ADR 0025 (docs/adr/0025-adopt-the-unknowns-corpus-as-judgment-preserving-contract-deltas.md); the branch's contract slice was graduated and pruned per the topic-docs convention (the prune gate), with the approved plan pasted below.

Fix

Three waves as commit ordering:

  • Wave 1 (docs/governance): docs/FINDING-YOUR-UNKNOWNS.md (taxonomy + lifecycle, five-pass workflow, pattern catalog, reply-affordance + export-button owner sections with conformance surfaces, opt-in deviation-log posture, when-HTML scoping, buy-in pattern with industry grounding, cautions, eval-candidate heuristics, fair-quotation basis and local citation shape); two names-and-points convention-registry rows; glossary entries (unknowns quadrants, blindspot finding types) + a map/territory rejected-terms row via the curate-language discipline; ADR 0025 as the graduated decision record.
  • Wave 2 (contract deltas, 7 plugins with same-commit eval extensions and minor version bumps): discovery (typed blindspot cards + scan-scope disclosure), education (vocabulary ladders, original-ask success condition, diff-sourced anchored quizzes, fresh-context answer keys), verification (out-of-diff couplings table + quiz-layer cross-ref), prototype (same-data control variable, single-decision graft capture, machine-legible reply template, validation answer sets, fake-data disclosure footer, mock-before-wire note), planning (switch conditions on alternatives, pre-drafted revision replies, free-text register flag, design tweak-likelihood ordering, brainstorm/wayfind doc lines), discipline (no-analogue port trap, argument-hint), session-flow (five-pass cross-ref, doc-line bump).
  • Wave 3 (own review moment): the external-reference port gate in discipline:point-dont-copy (semantics map + stop-and-wait confirmation, scoped to sources of truth outside this repo's tree; declared step delta on the shared corrector loop) and the typed deviation-log convention in the implementation plugin (interactive opt-in, completion fold-back, recorded registry trigger).

Behavioral-tier heuristics deliberately did NOT land as standing instructions (evidence-gated additions rule); they ship as doc lines + eval candidates tracked in #3589. Deferred items: #3590 (E4 prd pitch-view extension), #3591 (recorded triggers).

Recorded deviations after the base moved (topic-docs v3 wave merged mid-flight):

  • Acceptance criterion 5 named docs/topics/context-engineering-claude-5/PLAN.md as the home for three candidate inputs; that slice graduated to ADR 0004 on main, so the candidates re-homed to Context-engineering effort: three candidate inputs + re-inventory note from the unknowns integration #3593 and the branch's edit to that file was dropped rather than resurrecting a pruned slice.
  • The topic's own contract slice (PLAN, signoff-sheet, disposition ledger, delta record) was graduated (ADR 0025) and pruned in the final commit per docs/conventions/topic-docs/ and the contract-slice-prune-gate; the working material is preserved in this branch's history (last revision 72a5929c) and the approved PLAN.md is pasted below.
  • Merge with main re-picked versions so the merged releases supersede both sides (discovery 0.19.0, session-flow 0.34.16; planning 0.35.0 / verification 0.6.0 / implementation 0.16.0 stand above main's entries), with both sides' changelog entries stacked newest-first.

Verification

  • Per-phase gates before each commit; branch head: scripts/check-changed-skills.sh origin/main = 0 failed; scripts/check-purged-em-dashes.sh clean; markdownlint 0 issues + typos clean across all changed markdown; node scripts/generate-cheatsheet.mjs --check exit 0; ai-slop detector 0 findings over the changed set; contract-slice-prune-gate green after the graduation commit.
  • scripts/affected-tests.sh --run: 2569 assertions passed; the one failure was the interview-defenses digest ratchet correctly catching the D28 edit inside its pinned register section — defenses re-read and confirmed intact, the eval expectation moved to a gradeable case, digests refreshed per the suite's own contract, re-run green (89/0). The 13 other-ecosystem python suites run from their own lane: 785 passed + 330 subtests (the PowerShell suite is the Windows CI lane).
  • Acceptance-criterion traceability, both directions: all 48 corpus-inventory decisions carried dispositions in the ledger; all 45 executed sheet rows grep-verified against landed hunks (close-out record inside the pasted PLAN below).
Approved PLAN.md (contract slice, pruned per the topic-docs convention; last in-tree revision 72a5929)
# finding-your-unknowns-integration

## Brief

### TLDR

Integrate the verified "Finding Your Unknowns" corpus (Thariq Shihipar's field guide, its
X-Article draft, the 13 html-effectiveness pages, and the context-engineering companion;
slice `finding-your-unknowns-0f25bd45`) into this repo as judgment-preserving deltas on
existing skills plus a small set of citable reference docs, gated by a scripted
comparison-evidence pass over the ~20 named-skill collisions.

### Goal

Every corpus decision in the corpus inventory reaches a disposition (adopt / adapt /
compare-then-adopt / cite / drop / treat-as-caution) that is either executed as a repo
change or recorded with its reason, with no silent drops.

### Constraints

1. Vehicles (interview Q1): (a) augmentation of existing skills and (b) citable
   reference/convention docs are the primary vehicles; at most 2 genuinely new thin skills,
   each only where the evidence pass confirms a real gap; CLAUDE.md/rules changes only if
   the context-engineering material earns one on its own evidence.
2. Codification posture (Q2, BINDING): honor the source author's anti-premature-codification
   warning — only judgment-preserving deltas (output contracts, checklists, conventions with
   rationale); no generator-style "make me an X" skills; the warning itself is quoted in
   whatever reference doc graduates.
3. Evidence discipline (Q3): adopt/adapt verdicts on named-skill collisions are CONDITIONAL
   until a read-only evidence pass grades each named skill against its checkable claims;
   only surprises return to the human.
4. Vendor-claim discipline (Q5): corpus claims are vendor-blog anecdote unless the targeted
   live-doc check verifies them; the ~6 decision-relevant harness claims got that check.
5. House style: all new prose obeys the repo's ai-slop/house-style rules; citation shape is
   URL + retrieval date + content hash.
6. Execution contract (Q6): each conditional-batch unit is closed only when its verdict is
   executed-or-recorded and the inventory row links the outcome; no silent drops.
7. Vehicle placements (Q7-Q10): the five-pass sequencing lands as a workflow section in the
   central reference doc with cross-refs from planning:wayfind and session-flow:workflow (no
   new orchestration skill); the reference doc owns the prompt-pattern catalog with one
   canonical invocation line per owning skill; reply-affordance is a house convention
   (default-with-judgment) owned by the doc with an artifact-design cross-ref.
8. Evidence-gate classification (sign-off S1): every delta is classified per-row under
   PLUGIN-PHILOSOPHY's rubric. CONTRACT/POLICY/CONVENTION rows land now as team conventions
   adopted by the sign-off; DOC rows are citation/doc lines; BEHAVIORAL rows never land as
   standing instructions — they become reference-doc heuristic lines plus candidate eval
   cases, awaiting observed-stumble evidence.
9. Registry discipline (sign-off S2): convention-registry rows for reply-affordance and
   export-button only, each with an explicit conformance surface, landing owner-doc-first.
10. Sequencing with the context-engineering effort (sign-off S3): this effort never edits
    that effort's audit-instructions criteria files; its three candidate inputs are
    recorded as that effort's inputs (re-homed to #3593 after that slice graduated).
11. Schema stability (sign-off G.4 / D37): the `### Phase N` heading/tag vocabulary and any
    parsed schema (check-open-questions.sh fields) are never renamed without a version bump
    and changelog; block reordering is safe.

### Acceptance criteria (sign-off Part G.5)

1. Every corpus-inventory decision has an executed-or-recorded disposition — verified by
   grep over the disposition lines, not asserted (48/48 + 45/45, close-out record below).
2. All waves green on the per-plugin gates (version bump + CHANGELOG;
   check-changed-skills; listing budget; cheat-sheet; test battery).
3. The reference doc carries the quoted anti-premature-codification warning and the
   fair-quotation basis in its header.
4. Registry rows carry explicit conformance surfaces and land owner-doc-first.
5. The context-engineering effort carries the S3 sequencing rows (re-homed to #3593 —
   recorded deviation; that effort's slice graduated to ADR 0004 while this branch flew).
6. Every Wave-2/Wave-3 contract delta lands with eval expectations in the same commit.
7. /planning:plan consumed the signed sheet + this Brief as its input contract.
8. Delivery (operator amendment): one session, one branch, ONE pull request; waves are
   commit ordering; deferred items became filed issues.

### Deferred questions (each with trigger and arbiter)

- E4-EXT: prd pitch-view skill extension — behind demand evidence (#3590).
- D28-SCHEMA: mechanical scrutiny-flag field in the 5-field register — when a consumer
  needs mechanical reads (#3591).
- EVAL-CAND: behavioral eval candidates D5/D7/D11/D17b/D34-residue — observed-stumble
  evidence per the evidence-gated-additions rule (#3589).
- Q11-FLIP: tweak-likelihood requestable mode — own-usage evidence (#3591).
- E1-REG: deviation-log registry row — fires when a second plugin reads DEVIATIONS.md (#3591).

## Plan (all phases DONE)

- Phase 1: docs/FINDING-YOUR-UNKNOWNS.md — the graduated reference doc (taxonomy,
  five-pass workflow, pattern catalog, convention owner sections, when-HTML, buy-in
  pattern, cautions with byte-faithful quotes, eval-candidate heuristics, sources).
- Phase 2: governance placements — 2 registry rows, glossary via curate-language,
  traceability ledger (later graduated to ADR 0025 + this PR body).
- Phases 3-9: per-plugin contract deltas with same-commit evals and minor bumps
  (discovery D1/D4; education D12/D14/D16/D17a/F4; verification D19/D18; prototype
  D20-D27; planning D28/D32/D33/D35/D36/D9 + wayfind cross-ref; discipline E7/E8;
  session-flow cross-ref).
- Phase 10: Wave 3 — E6 external-reference port gate (declared step delta, C3 boundary
  verbatim, verified against the shared corrector method doc before landing) and E1/E2
  typed deviation-log convention (opt-in, fold-back, recorded registry trigger).
- Phase 11: close-out — issues #3589/#3590/#3591 (+#3593 post-merge re-home), cheat-sheet
  check, criterion sweeps, full battery, this PR; then the topic-docs graduation (ADR
  0025) and slice prune once the base's v3 wave made the prune gate binding.

## Close-out record (2026-09-01)

- Criterion 1: 48/48 corpus V-rows carried dispositions; 45/45 executed sheet rows
  grep-verified against landed hunks; zero unaccounted in either direction.
- Criterion 2: check-changed-skills 0 failed; em-dash ratchet clean; markdownlint +
  typos clean; cheat-sheet in sync; ai-slop detector 0 findings; affected-tests battery:
  2569 shell assertions + 785 python tests + 330 subtests green after the
  interview-defenses digest refresh (defenses confirmed intact per that suite's contract).
- Criteria 3-4: grep-verified (warning + basis in the doc; registry rows + owner-section
  conformance text).
- Criterion 5: satisfied via #3593 (recorded deviation, base moved).
- Criteria 6-8: eval expectations co-located per phase commit; the signed sheet + Brief
  were the executed contract; delivered as this single PR.

## Stress-test summary

Pre-sign-off: /planning:devils-advocate (1C/4H/6M/4L, all folded) + two independent
fresh-context validators over the sign-off sheet. Pre-execution: a fresh-context plan
review returned 3 CRITICAL / 6 IMPORTANT / 6 SUGGESTION, all verified against the repo
and folded (gate rewiring, traceability direction, cross-ref homes, criterion
reconciliations, self-verifying sanity checks).

## Execution shape

Fully sequential, all main-session, phases 1→11; sub-agents used only for fresh-context
review. Seven gate-passed [EXEC-SHAPE] decisions recorded (doc path, DOC-row folding,
sequential routing, F1-as-owner, issues granularity, artifact-design prose mention,
session-flow phase).

Related

Refs #3589, Refs #3590, Refs #3591, Refs #3593. Decision record: docs/adr/0025-adopt-the-unknowns-corpus-as-judgment-preserving-contract-deltas.md (this PR); incumbent context-engineering record: docs/adr/0004-rightsize-instruction-surfaces-by-incumbent-first-arbitration.md. Full working material (signoff-sheet, disposition ledger, delta record, verbatim PLAN): this branch's history at 72a5929c.

🤖 Generated with Claude Code

https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1

Interview contract for integrating the verified Finding-Your-Unknowns corpus
(slice finding-your-unknowns-0f25bd45): round-1 decisions locked (vehicles,
binding codification posture, conditional-verdict evidence pass, vertical
order, targeted live-doc checks). Brief accumulates as rounds resolve; the
Plan section stays empty for /planning:plan.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
…sign-off

Operator sign-off (".confirm all", 2026-09-01) closes the interview: the
Brief gains the seven acceptance criteria (sign-off Part G.5), the signed
constraint set (evidence-gate classification, registry discipline, ctx-eng
sequencing, schema stability), and the deferred-questions record with
per-item triggers and arbiters. The signed decision sheet (rev 2, both
final validators folded) is committed beside it as the durable input
contract for /planning:plan, including the operator's delivery amendment:
one session, one branch, one PR, waves as commit ordering.

Register gates clean: registered=17 open=0 answered=17, brief=ok.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
…iew)

Ten phases sequencing the signed contract: F1 reference doc, governance
placements (registry rows, glossary, ctx-eng sequencing bullets), six
per-plugin contract-delta phases with same-commit eval expectations, the
Wave-3 heavies (E6 port gate, E1/E2 deviation-log convention), and the
close-out phase (issues, cheat-sheet check, acceptance verification, the
single PR). Includes the durable G-block placement copy, the execution
shape (sequential main-session), the gate-passed decisions table, and the
Tier-C design-gate early-exit record. A fresh-context plan review is in
flight; confirmed findings will amend this draft before execution.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
All 15 reviewer findings verified against the repo and folded: Q7/Q10
cross-refs gain phase homes (wayfind in Phase 7, a new session-flow Phase 9,
artifact-design as a prose mention in F1), criterion-1 traceability flips to
the corpus-to-sheet direction off a committed disposition ledger, the phase
gates switch from affected-tests (which selects nothing for these paths) to
check-changed-skills, direct markdownlint, the em-dash ratchet, and the
direct register test, the criterion-4 reconciliation and criteria-2-7 walk
are recorded, and the sanity-check commands become self-verifying
before/after pairs. Waves renumber to 11 phases.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
The graduated reference doc the signed integration contract calls F1:
unknowns taxonomy and lifecycle, the five-pass pre-implementation workflow
mapped to owning skills, the prompt-pattern catalog, the reply-affordance
and export-button owner sections (with conformance surfaces), the opt-in
deviation-log convention with its recorded registry trigger, the when-HTML
taxonomy and scoping rule, the buy-in pattern with industry grounding, the
author's anti-premature-codification warning quoted byte-faithfully with
citation stamps, and the behavioral heuristics recorded as eval candidates
rather than standing instructions. Fair-quotation permission basis and
local citation shape stated in the doc header. lychee already excludes
x.com, so no config change rode along.

Phase 1 of docs/topics/finding-your-unknowns-integration/PLAN.md; sanity
checks green (markdownlint 0 issues, typos clean, section and stamp greps).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
Two names-and-points convention-registry rows (reply affordance, export
button) pointing at their owner sections in FINDING-YOUR-UNKNOWNS.md;
glossary entries for the unknowns quadrants and blindspot finding types
plus a map/territory rejected-terms row, curated per the curate-language
entry discipline with a dated provenance note; three "Open, new" candidate
bullets and a Phase-10 rebase note recorded in the context-engineering
topic PLAN (no phase headings touched); and the acceptance-criterion-1
traceability spine committed into the topic dir - the 48-row V-id
disposition ledger plus the delta wording record.

Phase 2 of the integration plan; sanity checks green (markdownlint, typos,
grep battery, no-phase-touch guard).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
blindspot's output contract gains the four-way finding taxonomy (Landmine /
History / Convention / Missing concept) on every card and a closing
one-line scan-scope disclosure naming which lanes ran and what was and was
not scanned. Adopted at the finding-your-unknowns integration sign-off as
team-convention-tier contract lines (D1, D4); rationale and provenance in
docs/FINDING-YOUR-UNKNOWNS.md and the topic's signoff-sheet. Evals gain
expectations for both lines in this commit; plugin 0.16.18 -> 0.17.0.

Phase 3 of docs/topics/finding-your-unknowns-integration/PLAN.md.
check-changed-skills green (0 errors).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
…resh keys

explain: rung-2 terms arrive as vocabulary-ladder entries (term, plain
definition, a modeled "you can now say" sentence), and original-ask
invocations gain a success condition judged by the user's next prompt
(bare comprehension asks exempt by scope). quiz-me: questions are
diff-sourced, every question anchors to the report section that teaches
its answer with on-miss routing to that exact section, and the embedded
answer key is fresh-context authored or verified. Adopted at the
finding-your-unknowns integration sign-off (D12, D14, D16, D17a, F4);
provenance in docs/FINDING-YOUR-UNKNOWNS.md and the topic signoff-sheet.
Evals extended in this commit (including a new original-ask case);
plugin 0.8.8 -> 0.9.0.

Phase 4 of the integration plan. check-changed-skills green (0 errors).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
Stage 2's intent match now requires naming the existing behavior a change
leans on (unchanged code whose contract the diff depends on), and the
outcome report template gains a dedicated couplings table with an
evidence column. The PR-prep edge case records that a quiz-me
comprehension layer may precede the gate while the merge gate stays in
confirm, one mechanism per concern. Adopted at the finding-your-unknowns
integration sign-off (D19 contract, D18 doc line); provenance in
docs/FINDING-YOUR-UNKNOWNS.md. Evals extended; plugin 0.5.10 -> 0.6.0.

Phase 5 of the integration plan. check-changed-skills and the em-dash
ratchet both green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
…isclosure

explore-directions: variants bind one identical data set (design is the
only variable), the handover closes with a machine-legible
direction/steal/skip/next-target reply template, and captures record
steal/skip decisions at single-decision granularity so grafts compose.
pressure-test: the HTML demo shell gains a validation answer set (forced
choices whose options name their costs, free-text escape hatch, copy-out)
and a visible fake-data disclosure footer; the capture step carries the
filled answers into the durable record. Shared discipline gains the
mock-before-you-wire ordering note. Adopted at the finding-your-unknowns
integration sign-off (D20-D22, D24-D27); provenance in
docs/FINDING-YOUR-UNKNOWNS.md. Evals extended; plugin 0.9.8 -> 0.10.0.

Phase 6 of the integration plan. check-changed-skills and the em-dash
ratchet both green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
…ordering

plan: rejected alternatives carry a one-line switch condition; Step 5
closes with pre-drafted one-line revision replies per flagged decision;
the sanity-check paragraph names the every-step-lands-green expectation
those checks back. interview: free-text answers get a free-text:
resolution-field flag (register schema untouched; gate-invisibility
recorded as a known limitation in context/loop.md). design: Phase 5
discussion rounds present findings in tweak-likelihood order. brainstorm
gains the session-start rationale line, wayfind the five-pass workflow
cross-ref. Adopted at the finding-your-unknowns integration sign-off
(D28, D32, D33, D35, D36, D9, Q7); provenance in
docs/FINDING-YOUR-UNKNOWNS.md. Evals extended for the four contract rows;
check-open-questions test suite green; plugin 0.34.15 -> 0.35.0.

Phase 7 of the integration plan. check-changed-skills green (0 errors).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
point-dont-copy's audit list gains the cross-stack port trap: a
source-side primitive with no target-side analogue whose invariant the
port silently drops must name the convention now carrying it, or the
finding stands. The skill gains an argument-hint showing the canonical
invocation (frontmatter only; description and trigger keywords untouched).
Adopted at the finding-your-unknowns integration sign-off (E7, E8);
provenance in docs/FINDING-YOUR-UNKNOWNS.md. Evals extended; plugin
0.12.20 -> 0.13.0 (changelog entry provisional, finalized when the same
version's E6 port gate lands in the Wave-3 commit).

Phase 8 of the integration plan. check-changed-skills green (0 errors).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
One doc line under the at-a-glance diagram: stages 0-3 expand, for
unfamiliar territory, into the five-pass order the marketplace repo's
docs/FINDING-YOUR-UNKNOWNS.md states with rationale (integration sign-off
Q7). Doc-line-only bump 0.34.14 -> 0.34.15; no contract change, no eval
delta.

Phase 9 of the integration plan. check-changed-skills green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
… (Wave 3)

point-dont-copy gains the external-reference port gate as a declared step
delta on the shared re-anchor/audit/correct loop: a five-section semantics
map (side-by-side pairs, preserved/changed/dropped ledger, edge-case
parity, open questions) and a stop-and-wait confirmation gate, scoped to
ports whose source of truth lives outside this repo's tree (vendored,
foreign-language, other-repo); in-tree corrections stay do-it-now, and the
no-analogue trap check feeds the dropped ledger. Verified against the
shared method doc before landing: its declared-step-deltas allowance is
the exact seam, no correct-forward reversal. Finalizes discipline 0.13.0's
changelog entry; a new eval case covers the gate.

The implementation plugin's deviation log gains typed entries
(plan-confirmed / discovery / deviation / human-decision) with four
deviation fields (plan said / found / chose / revisit) in
implement-dispatch's owning contract, an interactive opt-in in implement
Step 3, and a completion fold-back in Step 5 that reads DEVIATIONS.md and
emits plan-amendment bullets. Recorded registry trigger: a second plugin
reading the log graduates the contract to an owner doc. No registry row
now, per the signed C5/M3. Evals extended in both skills; implementation
0.15.8 -> 0.16.0.

Adopted at the finding-your-unknowns integration sign-off (E6 per C3
boundary, E1+E2 per C5); provenance in docs/FINDING-YOUR-UNKNOWNS.md.
Phase 10 of the integration plan. check-changed-skills and the em-dash
ratchet both green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
…rtial)

The dated close-out record: both criterion-1 sweeps (48/48 corpus rows,
45/45 sheet-row greps), the criteria 2-7 walk, and the three follow-up
issues (#3589 eval candidates, #3590 E4 pitch-view deferral, #3591
recorded triggers). The phase tag stays open until the full
affected-tests battery, still running, reports green; the PR follows it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
…ition

The digest suite caught two things the D28 change touched: the
free-text-flag paragraph lands inside the digested open-question-register
section of context/loop.md, and the new eval expectation sat in case 12,
whose unrelated-reply scenario never has an answer arrive, making the
criterion ungradeable there. Per the suite's own contract: both defenses
re-read and confirmed intact (the flag adds scrutiny on answered rows and
qualifies neither the ask-time write rule nor the gap/blocker register
bindings), the expectation moved to case 1 where answers arrive, and the
register-section and case-1 digests refreshed in the same change. Suite
green (89/0); check-open-questions tests and check-changed-skills green;
the battery's 13 other-ecosystem python suites run from their own lane,
785 passed + 330 subtests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 1, 2026

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review Completed 2026-09-01T19:37:55.683660Z 72a5929 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

…ded)

All eleven phases DONE; PR #3592 recorded in the close-out record.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
Main advanced with the topic-docs v3.0.0 wave while this branch was in
flight, bumping five of the same plugins. Resolution keeps both sides'
changelog entries stacked newest-first and re-picks versions so the merged
release supersedes both: discovery 0.17.0 -> 0.19.0 (main shipped its own
0.17.0 and 0.18.0), session-flow 0.34.15 -> 0.34.16 (identical-number
collision with main's release), planning 0.35.0, verification 0.6.0, and
implementation 0.16.0 stand above main's 0.34.16 / 0.5.11 / 0.15.9.
Version/changelog parity verified for all eight touched plugins;
check-changed-skills, interview-defenses, markdownlint, and the
cheat-sheet check green on the merged tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
@claude

claude Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 6m 34s —— View job


Code review complete

  • Gather context (PR body, changed files, skill instructions)
  • Check skip gate (open/draft/trivial/duplicate review) — proceeded, PR is open and non-trivial
  • Read docs/FINDING-YOUR-UNKNOWNS.md and topic docs
  • Review Wave 2 contract deltas (7 plugins: discovery, education, verification, prototype, planning, discipline, session-flow)
  • Review Wave 3 changes (point-dont-copy port gate, implementation deviation-log)
  • Review test/eval changes and interview-defenses.test.sh
  • Post inline findings + summary

Finding (posted inline)

docs/topics/context-engineering-claude-5/PLAN.md (1217 new lines) is an unrelated file swept into this PR by branch basis, not by this PR's own work. Diffed against origin/main it shows as wholly new, but it was already present in this branch's history — introduced by a separate, still-unmerged commit (7cd56c9, refactor: batch-simplify sweep, waves 4-12 + deferred-work pass (#3494)), confirmed not an ancestor of origin/main. This PR's own Wave-1 commit (727d22e) touched only 13 lines of it (the 3 "Open, new" bullets + the Phase-10 rebase note the PR body mentions). The file is also demonstrably stale (it claims .claude/rules/ "does not exist"; main has 5 files there today). Net effect: merging as-is lands an entire separate, partly-[TODO] topic plan into main, unreviewed and never mentioned in the Summary/Fix/Related sections — and it's ~44% of this PR's stated diff. See the inline comment for the full trail; recommend rebasing this branch onto current origin/main (or landing #3494 first) so the diff isolates the actual finding-your-unknowns work.

Everything else checked out

Spot-checked cross-file consistency across the wave-2/wave-3 contract deltas and the accompanying eval/CHANGELOG/version-bump machinery (discovery:blindspot, education:explain/quiz-me, verification:confirm, prototype:explore-directions/pressure-test, planning:brainstorm/design/interview/plan/wayfind, discipline:point-dont-copy, implementation:implement/implement-dispatch, session-flow:workflow):

  • Plugin.json version bumps match their CHANGELOG headers exactly (0.13.0, 0.19.0, 0.9.0, 0.16.0, 0.35.0, 0.10.0, 0.6.0; session-flow's patch-only 0.34.16 is explicitly justified as a doc-only, no-eval-delta line).
  • SKILL.md step renumbering in implementation:implement (Step 4→6 insertion) correctly updates its own internal forward-reference ("step 5 hands to..." → "step 6 hands to...").
  • The interview-defenses.test.sh digest updates correctly track the actual byte ranges/eval-case bodies that changed (context/loop.md's new free-text-flag paragraph, and the relentless-me-mode-frontier-rounds eval case's new expectation line) — no stale or over-broad re-pins.
  • Cross-referenced doc anchors (GLOSSARY.md new terms ↔ FINDING-YOUR-UNKNOWNS.md sections, PLUGIN-PHILOSOPHY.md convention-registry rows ↔ the doc's reply-affordance/export-button sections) resolve correctly.
  • Spot-checked the delta-resolution/disposition-ledger process docs against a few landed changes (D1/D4 blindspot, D12/D14 explain, D16/D17 quiz-me, D28 interview, D32/D33/D36 plan/design, E1/E2/E6/E7 discipline/implementation) — all match what's actually in the shipped SKILL.md/evals diffs, and the one traced non-obvious gap (D3, blindspot's confirm-before-final checkpoint) is recorded as an explicit "demoted" disposition rather than a silent drop.

No other correctness issues found in the reviewed diff.
· branch claude/reading-feedback-j4sg96

@kyle-sexton kyle-sexton changed the title Integrate the "Finding Your Unknowns" corpus: reference doc, conventions, and contract deltas across 8 plugins feat: integrate the Finding Your Unknowns corpus — reference doc, conventions, and contract deltas across 8 plugins Sep 1, 2026
@claude

claude Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 4s —— View job


Security review — /review:security-review

Skip gate: stopping early — no security-relevant surface in this diff.

This PR is entirely documentation and skill-instruction content: 4 new docs under docs/topics/finding-your-unknowns-integration/ and docs/FINDING-YOUR-UNKNOWNS.md, SKILL.md/CHANGELOG.md/evals.json edits across 8 plugins, and one hash-digest update in plugins/planning/tests/interview-defenses.test.sh. Confirmed via git diff origin/main...HEAD:

  • No .github/workflows/** changes (no permissions, triggers, or Actions logic touched).
  • No shell/script logic changed beyond two pinned SHA-256 digest literals in the test file (re-pinning an existing content hash, not new logic).
  • No auth/access-control, injection surface (command/SQL/path/template), secret/token/credential handling, or network calls introduced anywhere in the diff.
  • The new "external-reference port gate" in plugins/discipline/skills/point-dont-copy/SKILL.md (added stop-and-wait confirmation step for cross-repo ports) is agent-workflow governance prose, not executable code — no trust-boundary or exploitable path.

No security issues found. This lane defers style/quality/coverage commentary to /review:code-review.

@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown

Last security-reviewed head: 5f614aebeed1efc2e15d2e6b10c92fa8d2e4a8dc. On the next push, the relevance gate compares only the commits since this SHA; delete this comment to force a full re-review.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 72a5929c04

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/planning/skills/interview/context/loop.md Outdated
Comment thread plugins/prototype/skills/pressure-test/SKILL.md Outdated
…act slice

The topic-docs v3 wave completed the contract-tier burn-down on main while
this branch was in flight: docs/topics/ is empty there, the prune gate now
red-lines any slice left behind, and the context-engineering topic this
branch had annotated graduated to ADR 0004. Close-out per the convention:
the durable decision record becomes docs/adr/0025 (evidence chain,
classification decision, consequences, recorded triggers), the glossary
provenance points at the ADR, the branch's own slice (PLAN, signoff-sheet,
disposition ledger, delta record, design resolution) prunes as working
material preserved in this branch's history and pasted into the PR body,
and the resurrected context-engineering PLAN.md is removed; its three
candidate inputs re-home to a tracker issue, since acceptance criterion 5's
named target no longer exists on main (deviation recorded in the PR).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
…footer

Two Codex review findings, both verified real. The interview free-text
flag now applies only to replies that RESOLVE their question (a complete
free-text answer, or an explicit "you pick", which resolves to the
recommendation); a partial or non-resolving reply keeps its row open
under the drift check, so the flag can never launder a non-answer into a
terminal answered row. The pressure-test disclosure footer no longer
forces inventing production details: it states wiring location and flag
when decided, and says "not decided / no flag planned" explicitly
otherwise. The interview-defenses register-section digest is refreshed
for the reworded paragraph (defenses re-read: the fix strengthens the
no-silent-capture posture; suite green 89/0).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
Comment thread docs/topics/context-engineering-claude-5/PLAN.md Outdated
@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown

Claude has reviewed this PR 5 times. The lane skips further automatic reviews after 5; deleting this comment resets the count.

@claude

claude Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 57s —— View job


Security review — /review:security-review

Skip gate: stopping early — no security-relevant surface in this diff.

Re-reviewed the full PR diff (git diff origin/main...HEAD, 50 files, +718/-38) against the prior pass posted at head 8d3544c. The two commits added since then (770c829, c69c2e9) don't change the conclusion:

  • 770c829 removes docs/topics/context-engineering-claude-5/PLAN.md (the 1217-line accidental branch-basis artifact flagged in the earlier review comment) along with the pruned finding-your-unknowns-integration contract-slice docs — net effect is fewer files in scope, not more.
  • c69c2e9 only touches plugins/planning/skills/interview/context/loop.md and plugins/prototype/skills/pressure-test/SKILL.md (both agent-instruction prose) plus a re-pinned SHA-256 digest literal in plugins/planning/tests/interview-defenses.test.sh — a content-hash fixture update, not new logic.

Full-diff file list confirms: no .github/workflows/** changes (no permissions, triggers, or Actions logic touched), no other executable code (.py/.js/.ts) anywhere in the changeset — every changed file is .md, .json (eval fixtures / plugin-version manifests), or the one digest-only .sh line noted above. No auth/access-control, injection surface (command/SQL/path/template), secret/token/credential handling, or network calls anywhere in the diff.

No security issues found. This lane defers style/quality/coverage/scope commentary (e.g. the accidental-file concern from the earlier review, now resolved by 770c829) to /review:code-review.

@claude

claude Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 4m 39s —— View job


Code review — /review:code-review

Scope: this is the CODE-REVIEW lane (correctness + maintainability; security is out of scope here — see the separate security-review comment above, which already found no security-relevant surface). Reviewed the full diff against origin/main, all 50 changed files.

Prior review findings already resolved. Both Codex findings (free-text flag treating incomplete replies as terminal; pressure-test footer forcing invented wiring details) and the earlier Claude finding (an unrelated docs/topics/context-engineering-claude-5/PLAN.md file appearing in the diff) are gone at this head — the first two were fixed in c69c2e9, and the third file simply isn't part of this diff anymore (confirmed via git diff origin/main...HEAD, 50 files / 718+/-38, matching the PR's stated file list).

New findings (2, both posted inline):

  1. plugins/planning/CHANGELOG.md (0.35.0, "interview: free-text resolution-field flag") — the bullet still describes the pre-fix, over-broad rule ("any free-text answer gets flagged"). Commit c69c2e9 narrowed the actual rule to "only replies that resolve their question" but never touched this CHANGELOG entry.
  2. plugins/prototype/CHANGELOG.md (0.10.0, "pressure-test: ... fake-data disclosure footer") — same pattern: still says the footer states "where real wiring lives," omitting the "not decided / no flag planned" fallback c69c2e9 added.

Both are cases where a fix commit corrected the actual skill contract but left the CHANGELOG describing the old, already-flagged-as-buggy behavior — worth fixing since these CHANGELOGs are this repo's authoritative record of what a delta does.

Also checked, no issues found:

  • All 8 plugin.json version bumps match their CHANGELOG head entries.
  • All 12 touched evals.json files are valid JSON with no duplicate eval IDs introduced by this PR (one pre-existing ID gap in verification/skills/confirm/evals.json, already present on origin/main, untouched by this PR).
  • Cross-file consistency: the renumbered steps in implementation:implement SKILL.md (Step 5Step 6 for the pre-PR handoff cross-reference) were updated correctly; external references to that skill's "Step 5: Completion and Handoff" section header still resolve since the header name didn't change.
  • The new discipline:point-dont-copy external-reference port gate correctly follows its own method doc's "Declared step deltas" allowance (states the reason and the insertion point between the loop's steps 2 and 3, per plugins/discipline/context/re-anchor-audit-correct.md).
  • ADR 0025 and the GLOSSARY.md new terms/provenance note are internally consistent with the rest of the diff.
  • Citation hashes in docs/FINDING-YOUR-UNKNOWNS.md are well-formed 64-hex-char sha256 digests.
    · branch claude/reading-feedback-j4sg96

Comment thread plugins/planning/CHANGELOG.md Outdated
Comment thread plugins/prototype/CHANGELOG.md Outdated
…rowed rule

Review finding: the 0.35.0 bullet still described the flag as applying to
any free-text answer, but c69c2e9 narrowed it to replies that resolve
their question, with partial or non-resolving replies staying open under
the drift check. The bullet now matches context/loop.md.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
… fix

Same staleness the planning changelog had: the 0.10.0 bullet still
described the pre-fix footer contract; it now records the explicit
"not decided / no flag planned" fallback c69c2e9 added.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
@claude

claude Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 2m 4s —— View job


Code review — /review:code-review

Note: the trigger's supplied HEAD SHA was 6e36f3a, but two commits later this branch already had a newer head (84460c0) at review time — that commit itself fixes one of the two findings from the prior review pass, so I reviewed the actual current state.

Prior findings — status

Both CHANGELOG staleness findings from the previous code-review pass are now fixed:

  • plugins/planning/CHANGELOG.md (free-text flag bullet) — fixed in 6e36f3a.
  • plugins/prototype/CHANGELOG.md (disclosure-footer bullet) — fixed in 84460c0.

Both now correctly describe the narrowed rule from c69c2e9 (flag only applies to free-text replies that resolve their question; footer allows an explicit "not decided / no flag planned" fallback). Confirmed by diffing each bullet against the current text of plugins/planning/skills/interview/context/loop.md and plugins/prototype/skills/pressure-test/SKILL.md.

Also confirmed docs/topics/context-engineering-claude-5/PLAN.md (the unrelated branch-basis artifact flagged in an earlier review round) is no longer part of this diff — git diff origin/main...HEAD --stat now matches the PR's stated 50-file list exactly.

New finding (posted inline)

plugins/planning/skills/interview/evals/evals.json:15 — same staleness pattern as the two CHANGELOG bullets, but unfixed. The eval criterion for case relentless-me-mode-frontier-rounds still reads "an answer arriving as free text ... is recorded with the free-text: flag," with no "only when it resolves the question" qualifier. c69c2e9 narrowed the actual rule in context/loop.md but never touched this grading criterion, so a judge grading strictly against this line would still expect the flag on any free-text reply — including a non-resolving one — which is exactly the over-broad behavior the fix removed. The pin_case_digest refresh in the same commit only re-pinned this case's existing (stale) text rather than catching the mismatch, since that digest detects drift, not correctness against the current rule.

Everything else checked

Diffed the full PR against origin/main (50 files, +723/-38, matches the PR's stated count) and spot-checked the remaining wave-2/wave-3 files for regressions since the last full pass — no other issues found. Plugin version bumps still match CHANGELOG headers; eval JSON is well-formed with no new duplicate IDs; cross-file anchors (glossary ↔ reference doc, registry rows ↔ owner sections) still resolve.
· branch claude/reading-feedback-j4sg96

Comment thread plugins/planning/skills/interview/evals/evals.json Outdated
@claude

claude Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 2m 38s —— View job


Code review — /review:code-review

…text rule

Review finding: the relentless-me-mode-frontier-rounds expectation still
graded the pre-fix, over-broad rule (any free-text answer flagged). It
now matches the narrowed loop.md contract: a resolving free-text reply
gets the flag, a partial or non-resolving reply keeps its row open. The
case digest in interview-defenses.test.sh is refreshed in the same change
per the suite's contract; suite green (89/0).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GE7YPqWwqGSNYfVWj8DdF1
@claude

claude Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 51s —— View job


Security review — /review:security-review

Skip gate: stopping early — no security-relevant surface in this diff.

Re-verified the full PR diff (git diff origin/main...HEAD, 50 files, +723/-38 — matches the PR's stated stats exactly) at the current head (5f614ae). The two commits added since the last security-review pass (6e36f3a, 5f614ae) don't change the conclusion:

  • 6e36f3a and 5f614ae only edit plugins/planning/CHANGELOG.md, plugins/prototype/CHANGELOG.md, and plugins/planning/skills/interview/evals/evals.json — all prose/eval-criterion text syncing stale wording to the c69c2e9 fix already reviewed. No logic changes.
  • No .github/workflows/** changes anywhere in the branch (git diff origin/main...HEAD -- '.github/workflows/**' is empty) — no permissions, triggers, or Actions logic touched.
  • The only non-.md/.json file touched anywhere in the diff is plugins/planning/tests/interview-defenses.test.sh, and its only change is two re-pinned SHA-256 digest literals (content-hash fixture updates tracking prose changes elsewhere in the same commits, not new logic).
  • Scanned the full diff for URLs, exec/subprocess patterns, and credential-shaped strings: every URL hit is a citation reference inside documentation prose (blog posts, RFCs, json.schemastore.org schema refs already used repo-wide) — no new network calls, no secrets/tokens/credentials introduced.
  • No auth/access-control, injection surface (command/SQL/path/template), or trust-boundary code anywhere in the changeset — every changed file is .md, .json (eval fixtures / plugin-version manifests), or the one digest-only .sh diff noted above.

No security issues found. This lane defers style/quality/coverage/scope commentary to /review:code-review.
· branch claude/reading-feedback-j4sg96

@claude

claude Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

Claude finished @kyle-sexton's task in 1m 53s —— View job


Code review — /review:code-review

Skip gate: not applicable — PR is open, non-draft, and non-trivial (50 files, +723/-38).

Prior finding — now fixed

The one outstanding finding from the previous code-review pass (posted at head 84460c0) is resolved by this branch's only new commit, 5f614ae:

plugins/planning/skills/interview/evals/evals.json:15 — the relentless-me-mode-frontier-rounds eval criterion previously described the pre-fix, over-broad free-text-flag rule. It now reads "A reply that resolves its question but arrives as free text ... is recorded with the free-text: ... flag ...; a partial or non-resolving reply keeps its row open," which matches the narrowed rule in plugins/planning/skills/interview/context/loop.md:209.

I independently recomputed the pinned digest rather than trusting the commit message: jq -S -c over the relentless-me-mode-frontier-rounds eval case, sha256-hashed, produces 5e38782253a410890320cddc3c5c44d71ab64666bb596bc954167f1447059bdd — matching the updated pin in plugins/planning/tests/interview-defenses.test.sh:580 exactly. (I could not execute the full interview-defenses.test.sh suite in this lane — its sandbox only permits git/gh/Read/Grep/Glob-class commands, not arbitrary script execution — so the "89/0 green" claim in the commit message is author-claimed and unverified beyond this one digest I recomputed directly.)

Full diff re-check

Confirmed the previously-flagged unrelated docs/topics/context-engineering-claude-5/PLAN.md artifact remains absent (git diff origin/main...HEAD --stat -- 'docs/topics/*' is empty), and git diff origin/main...HEAD --stat matches the PR's stated 50 files / +723/-38 exactly. Diffed all evals.json changes across the PR against their corresponding SKILL.md/context prose; all read as consistent with the delivered contract text.

No new findings. This PR looks ready from the code-review lane's perspective.
· branch claude/reading-feedback-j4sg96

@kyle-sexton
kyle-sexton merged commit f43bd07 into main Sep 1, 2026
62 checks passed
@kyle-sexton
kyle-sexton deleted the claude/reading-feedback-j4sg96 branch September 1, 2026 21:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants