Skip to content

redgate: point hooks at the handlers that exist; sort runs in C locale - #136

Merged
JRichlen merged 3 commits into
mainfrom
claude/redgate-hooks-path-fix
Sep 15, 2026
Merged

JRichlen merged 3 commits into
mainfrom
claude/redgate-hooks-path-fix

Conversation

@JRichlen

Copy link
Copy Markdown
Owner

Two fixes lifted out of #115 so they do not wait on 76k lines. Three changed lines, plus a cheap-tier section that would have caught the first one.

The shipped defect

plugins/redgate/hooks/hooks.json referenced ${CLAUDE_PLUGIN_ROOT}/hooks-handlers/<handler>.sh. Both handlers live under hooks/hooks-handlers/, and ${CLAUDE_PLUGIN_ROOT} is the plugin directory on an installed plugin — agent-compiler's own hooks.json resolves the same way. There is no plugins/redgate/hooks-handlers/ directory.

So on every installed copy of redgate, the PreToolUse write guard — the hook that stops a write reaching a ratified run's contract while its round is in MIDDLE — and the SessionStart announcer silently never ran. The cheap tier stayed green throughout, because nothing resolved handler paths against the filesystem.

-  "command": "bash \"${CLAUDE_PLUGIN_ROOT}/hooks-handlers/guard-redgate-paths.sh\""
+  "command": "bash \"${CLAUDE_PLUGIN_ROOT}/hooks/hooks-handlers/guard-redgate-paths.sh\""
-  "command": "bash \"${CLAUDE_PLUGIN_ROOT}/hooks-handlers/session-start.sh\""
+  "command": "bash \"${CLAUDE_PLUGIN_ROOT}/hooks/hooks-handlers/session-start.sh\""

The locale sort

criteria-index.sh sorted run directories with the host locale. Its own header promises deterministic output and --check compares against it, so an en_US host can order slugs differently from CI and report drift that is not there.

-  for run in $(ls -d "$RG"/*/ 2>/dev/null | sort); do
+  for run in $(ls -d "$RG"/*/ 2>/dev/null | LC_ALL=C sort); do

The guard

New cheap-tier section 2b resolves every ${CLAUDE_PLUGIN_ROOT}/<rel> in every plugins/*/hooks/hooks.json to a file under plugins/<p>/<rel>, one PASS/FAIL per handler, and fails closed if the walk finds nothing. It is resolved against the filesystem rather than pattern-matched on purpose: voice's hooks.json uses the identical-looking hooks-handlers/ path and is correct, because voice keeps its handler directly under the plugin root. A pattern rule would have flagged the healthy plugin and could not have told the two apart.

Verification

Check Result
cheap tier, this branch 1295 passed, 0 failed (1290 on main + 5 handler paths)
section 2b against main's pre-fix hooks.json both redgate hooks reported missing — 1293 / 2 failed
mutation: agent-compiler handler renamed away red (1294 / 1)
mutation: walk pointed at an empty glob red — "resolved ZERO handler paths" (1290 / 1)
all five current handler paths resolve: agent-compiler ×2, redgate ×2, voice ×1

Not verified here, stated rather than implied:

  • The locale ordering difference. No en_US locale is installed in this container; LC_ALL=en_US.UTF-8 falls back to C and sorts identically. The one-line fix rests on sort(1) semantics and feat(evals): marketplace-wide agentic test framework (T01–T52) with red-team lane #115's report of a red cheap tier on en_US hosts, not on a reproduction.
  • The deep tier. Its path filter matches criteria-index.sh (plugins/*/skills/**/scripts/**), but redgate ships no pier pack, so deep tier (pier) will report green because its leg is skipped, not because anything ran. That is the existing gate design, named here so a green is not read as proof.

No SKILL.md, command, or skill references/ touched, so demonstration discipline does not apply. No eval tier, workflow job, or eval pack added, removed, renamed, or re-scoped, so docs/testing.md is unchanged and section 20 stays green.

Independent of the OpenRouter credit outage blocking #131 and #133 — nothing here calls a model.

🤖 Generated with Claude Code

https://claude.ai/code/session_01XHYtoYrtWhTGy59GC9uJGd


Generated by Claude Code

Two fixes lifted out of #115 so they do not wait on 76k lines.

hooks.json referenced ${CLAUDE_PLUGIN_ROOT}/hooks-handlers/<h>.sh, but both
handlers live under hooks/hooks-handlers/, and ${CLAUDE_PLUGIN_ROOT} is the
plugin dir on an installed plugin (agent-compiler's own hooks.json resolves
the same way). No plugins/redgate/hooks-handlers/ directory exists. So on
every installed copy the PreToolUse write guard — the hook that stops a write
reaching a ratified run's contract while its round is in MIDDLE — and the
SessionStart announcer silently never ran. The cheap tier was green the whole
time because nothing resolved handler paths against the filesystem.

criteria-index.sh sorted run dirs with the host locale. Its own header
promises deterministic output and --check compares against it, so an en_US
host can order slugs differently from CI and report drift that is not there.
LC_ALL=C pins it.

New cheap-tier section 2b resolves every ${CLAUDE_PLUGIN_ROOT}/<rel> in every
plugins/*/hooks/hooks.json to a file under plugins/<p>/<rel>, and fails
closed if the walk finds nothing. It is resolved against the filesystem, not
pattern-matched: voice's hooks.json uses the identical-looking
hooks-handlers/ path and is CORRECT, because voice keeps its handler directly
under the plugin root.

Verified: cheap tier 1295 passed / 0 failed (1290 on main + 5 handler paths).
Against main's pre-fix hooks.json the new section reports both redgate hooks
as missing (1293/2). Two further mutations — an agent-compiler handler
renamed away, and the walk pointed at an empty glob — each go red.

Not verified here: the locale ordering difference. No en_US locale is
installed in this container; LC_ALL=en_US.UTF-8 falls back to C and sorts
identically, so the one-line fix rests on sort(1) semantics and #115's report
of a red cheap tier on en_US hosts, not on a reproduction. The deep tier's
path filter matches criteria-index.sh, but redgate ships no pier pack, so
that check will report green because its leg is skipped, not because it ran.

No SKILL.md, command, or skill references/ touched; demonstration discipline
does not apply. No eval tier, job, or pack added, removed, or re-scoped, so
docs/testing.md is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHYtoYrtWhTGy59GC9uJGd
Copilot AI lite review requested due to automatic review settings September 15, 2026 02:14
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 15, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-15T02:20:52.997686Z a9145b8 PR opened
🔒 Security Review ✅ Completed 2026-09-15T02:32:38.805357Z a9145b8 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The handler-path validation needs canonical containment checks and reliable walker error propagation.

Get a fresh assessment by requesting another Copilot review.

Pull request overview

This PR fixes Redgate hook paths, makes criteria indexing locale-independent, and adds cheap-tier handler validation.

Changes:

  • Corrects both Redgate hook handler paths.
  • Forces criteria sorting with LC_ALL=C.
  • Adds filesystem-based hook path checks.
File summaries
File Summary
plugins/redgate/skills/criteria-contract/scripts/criteria-index.sh Makes run ordering deterministic across locales.
plugins/redgate/hooks/hooks.json Corrects handler locations.
evals/cheap/run.sh Validates resolved hook handler paths.
Review details

Suppressed comments (1)

evals/cheap/run.sh:82

  • The Python walker runs in process substitution, so its non-zero exit status is not propagated by set -o pipefail and the loop only checks whether it saw any output. If an earlier plugin emits valid paths and a later syntactically valid but structurally malformed hooks.json raises inside this walker, hook_paths_seen remains nonzero and section 2b reports no failure instead of failing closed. Capture and check the walker's exit status (or run it as a normal command) before accepting the results.
)
  • Files reviewed: 3/3 changed files
  • Comments generated: 1
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread evals/cheap/run.sh Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a9145b8edd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread evals/cheap/run.sh Outdated
… is nothing-to-check

Two corrections to the section added in a9145b8, both found by running it
against something other than the tree it was written on.

1. The counterfeit corpus's synthetic marketplace ships no hooks.json at all,
   and the section's "zero paths resolved" fail-closed turned that root red:
   the corpus's own calibration guard reported "baseline plugin is NOT green —
   corpus is miscalibrated, every rejection below is meaningless" and the
   required counterfeit tier failed on CI. Reproduced locally (24 passed / 1
   failed) before touching anything. The section now counts hooks.json FILES
   separately from extracted PATHS: no files is a legitimate "nothing to
   check", the same PASS shape every other section uses for an absent surface;
   files present but zero paths extracted still fails closed, because that is
   the walk breaking, not the plugins being clean.

2. The resolver checked isfile() only, while the section's name promises
   containment. `${CLAUDE_PLUGIN_ROOT}/../shared/x.sh`, or a symlink out of the
   plugin, would report OK when the target exists even though the installed
   command resolves outside the plugin root. Copilot's finding on #136. Both
   paths are now realpath'd and the target must share the plugin dir as its
   commonpath — the same idiom the relative-link resolver further down already
   uses.

Verified, in order:
  counterfeit harness on the fixed file: 25 passed / 0 failed, baseline green,
    all 18 counterfeits rejected by their expected gate (was 24 / 1).
  cheap tier on the repo: 1295 passed / 0 failed, unchanged.
  Five mutations against the final section, each restored after:
    main's pre-fix redgate hooks.json      -> 2 FAIL (1293 / 2)
    agent-compiler handler renamed away    -> 1 FAIL (1294 / 1)
    ../ escape to a REAL file in redgate   -> 1 FAIL (1294 / 1)  isfile alone said OK
    extraction regex broken, 3 files kept  -> 1 FAIL "3 file(s) present but the
                                              walk extracted ZERO handler paths"
    glob pointed at a root with no hooks   -> PASS "nothing to check" (1291 / 0)
  The regex mutation's first attempt did not apply (a sed pattern that missed
  the heredoc's escaping left the file untouched and the tier at 1295 / 0);
  that was a null test, not evidence, and is recorded here so the second,
  applied attempt is the one that counts.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHYtoYrtWhTGy59GC9uJGd

Copy link
Copy Markdown
Owner Author

Pushed 2af3101. Two corrections to section 2b, one of them mine to own.

Counterfeit tier was red, and it was this PR's fault. The corpus's synthetic marketplace ships no hooks.json at all, and the section's "zero paths resolved" fail-closed turned that root red — the corpus's own calibration guard said "baseline plugin is NOT green — corpus is miscalibrated, every rejection below is meaningless". Reproduced locally at 24 / 1 before touching anything. The section now counts hooks.json files separately from extracted paths: no files is a legitimate "nothing to check" (the same PASS shape every other section uses for an absent surface); files present but zero paths extracted still fails closed, because that is the walk breaking, not the plugins being clean. Counterfeit harness now 25 / 0, baseline green, all 18 counterfeits rejected by their expected gate.

Copilot's containment finding was correct and is fixed — both paths are realpath'd and the target must share the plugin dir as its commonpath, the same idiom the link resolver further down already uses. The ../ case is now a mutation in the verification set: pointed at a real file in redgate from agent-compiler's hooks.json, isfile alone says OK and the section correctly fails.

Five mutations against the final section, cheap tier 1295 / 0 unchanged. One null result recorded honestly: the regex-break mutation's first attempt didn't apply (a sed pattern missed the heredoc's escaping and left the file untouched at 1295 / 0) — that was no evidence at all, so it was redone with an applied edit, which fires "3 file(s) present but the walk extracted ZERO handler paths".

Two things this run revealed that are outside this diff:

  • The behavioral tier — promptfoo (redgate) leg ran for real (by design — plugins/<p>/hooks/** is behavior surface) and passed 9 / 9 with 0 errors on the live subject. That is the first paid-tier evidence in six days that the OpenRouter account is funded again; evals: confirm the SUBJECT model resolves, and repin it to qwen/qwen3.8-flash #131 and Portal/shunt delegation research, and three skills it sharpens #133 are no longer blocked on credit.
  • That same leg logs evals/paid/capture-example.sh: Permission denied — the script isn't executable, so the gallery capture step is a silent no-op in CI (guarded by || true; the 4 KB "snapshot" artifact is the checked-in docs/examples/data/redgate.json, not a fresh capture). Pre-existing on main, untouched here; noting it rather than widening this PR.

Generated by Claude Code

Codex's finding on #136, confirmed by reproduction before touching anything:
the handler-path scanner ran in process substitution, so its exit status was
discarded. A manifest that is valid JSON but not the shape the walk expects —
`"command": null` — raised TypeError, the remaining manifests were never
scanned, and because an EARLIER manifest had already emitted records the
zero-paths guard stayed quiet. On the unmodified file, nulling voice's command
made the `voice hooks.json SessionStart` line simply vanish; section 2b
reported no failure and the required tier exited 0 (1294 / 0), while stderr
carried the traceback nobody reads. Nulling agent-compiler's instead went red
only incidentally, via the zero-paths guard, and misreported "1 file(s)" when
three exist — the scan had been silently truncated.

Three layers now, each verified on its own:
  * Every shape the walk touches is checked, not assumed. Nonsense in a
    valid-JSON manifest emits a MALFORMED record naming the plugin and event
    instead of raising, and the walk continues to the next manifest.
  * The scanner writes to a temp file and its exit status is captured; a
    nonzero one is a recorded FAIL saying an unknown number of manifests were
    never checked.
  * The temp file itself fails closed if it cannot be created, and the
    "nothing to check" PASS moved inside the scanner-ran branch, so no
    silent-scanner path can produce that green line.

Verified against the final file:
  null command, last manifest   -> FAIL "MALFORMED hook entry (a hook has no
                                   string "command")"          1294 / 1, exit 1
  null command, first manifest  -> same FAIL, and redgate + voice still
                                   scanned after it              1294 / 1
  crash on every manifest       -> FAIL "scanner exited 1" + zero-paths FAIL
  crash on the LAST manifest    -> FAIL "scanner exited 1"      1294 / 1
    (the exact earlier-records shape that used to pass)
  "hooks" not an object / event not a list / entry's hooks not a list
                                -> one MALFORMED FAIL each
  TMPDIR unwritable             -> FAIL, and no "nothing to check" green
  The five prior mutations unchanged: main's pre-fix redgate (2 FAIL), handler
  renamed (1), ../ escape to a real file (1), regex broken (now reports
  "3 file(s)", not 1 — the walk no longer truncates), no-hooks root (PASS).
  cheap tier 1295 / 0 unchanged; counterfeit harness 25 / 0, baseline green.

The pre-fix reproduction and the full mutation sweep were run by a delegated
agent; the fixed-file cases above, both roots, and the hooks.json files being
byte-identical to HEAD were re-verified independently before this commit.

Survey, not fixed here: sections 10 (per-plugin safety-pack discovery) and 12
(install-smoke coverage) feed `while read` from the same `< <(python3 …)`
shape and neither asserts the enumerated count against marketplace.json, so a
scanner crash after N plugins would silently run only N packs and stay green.
Section 10 is the one this repo's safety story rests on. Follow-up.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XHYtoYrtWhTGy59GC9uJGd
@JRichlen
JRichlen merged commit 9ace33a into main Sep 15, 2026
56 checks passed
JRichlen pushed a commit that referenced this pull request Sep 18, 2026
The merge brought in #136, which added checks. The figure this branch
introduced (~19s / 1290 checks, measured 2026-09-14) was already stale
against its own instruction to re-measure rather than trust it.

Measured on this checkout at 2026-09-18: 13.6s, 1296 checks, 25 plugins.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb
JRichlen added a commit that referenced this pull request Sep 24, 2026
… fits instead (#134)

* Research: where the Harness Knowledge Graph lands in this marketplace

Placement analysis for the Harness Software Delivery Knowledge Graph,
diffed against every shipped skill and against this repo's own prior
verdict on graph memory in docs/research/agentic-patterns-corpus.md.

Verdict: not a new plugin. Harness's payoff needs heterogeneous,
non-git-native estate data with no single strongly-consistent query
surface; a repo fleet already has one. But the corpus rejected graph
memory on evidence quality (one 198-document paper), and Harness
retires that reason — the rejection is re-based on domain fit, which
is stronger and survives.

Lands as evidence inside two shipped skills:

- eval-ladder: their published eval methodology is a production eval
  ladder for a retrieval system, and names a rung we don't —
  "a registered relationship is not necessarily a usable relationship"
  (declared != populated != fresh). Their product-backed validation is
  independent, citable confirmation of "grade the surface closest to
  the harm".

- fleet-playbook-curator: three of Harness's four theses were arrived
  at independently (canonical identity on node_id, freshness as an
  independent clock, index-not-CMDB). The fourth is a verified gap —
  the fleet manifest is a flat entity table with no edge field, so
  every relationship exists only as uncited prose.

Records what the research could not establish, including that the
15-25x token claim is an unreplicated vendor benchmark and that the
docs page was read through a search-index fetch because the origin is
egress-blocked here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

* Correct the fleet-playbook relationship claim: not uncited, unmodeled

Codex flagged the note as asserting that relationship prose in a fleet
playbook is "uncited by construction". Checked against source, and the
finding is right:

- SKILL.md:98 names "cross-repo interactions" as exactly the kind of
  thing that belongs in a playbook.
- The per-claim rule covers EVERY substantive claim, relationships
  included: repo@sha:path plus an as-of stamp, or omitted/flagged STALE.
- validate-citations.sh fails the build on a claim citing a repo not
  read this pass, or a path absent from that repo's gathered tree.

So a relationship claim is cited, and its citation is machine-checked
for traceability. Saying otherwise understated a safety mechanism the
plugin actually has, and rested the edge argument on a false premise.

The gap is narrower and survives restating. Relationships are not
MODELED, with two consequences:

- No edge is diffable. diff-fleet.sh cascades over membership and
  pushed_at; the staleness clock stamps head_sha per member and nothing
  per edge, so a relationship that stops holding raises no signal of
  its own.
- Traceable is not supported. validate-citations.sh says so in its own
  comments: semantic support is the behavioral layer's job. For a
  single-repo claim the cited file usually is the evidence; for an edge
  the evidence is the join, and a claim can cite two real, genuinely
  read paths while asserting an edge neither supports.

Cheap tier: 1290 passed, 0 failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

* Drop "strongly-consistent" for "authoritative" — the argument needs the weaker claim

Copilot flagged that the note asserts there is "no single strongly-consistent
query surface" for a DevOps estate and then names `gh api` as one for a repo
fleet. GitHub publishes no consistency guarantee for REST list endpoints, so
the phrasing reads as a guarantee the note cannot substantiate.

The finding is right, and the fix costs the argument nothing: what actually
carries it is that ONE surface is the system of record and its join key
(node_id) is stable — not any ordering or read-your-writes property. Both
occurrences now say "authoritative".

Added a parenthetical noting that fleet-playbook-curator's own prose calls
`gh api orgs/<owner>/repos` "strongly-consistent", and that the defensible
contrast it reaches for is with the Search API, which documents its indexing
lag. Left that SKILL.md wording alone — out of scope for a docs note, and
worth its own look.

Cheap tier: 1290 passed, 0 failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

* Fix six review findings in the Harness research note

A full review of the note against this repo's own source turned up six
claims that do not hold. All six are corrected here; the note's argument
and verdict are unchanged.

- diff-fleet.sh does not key on pushed_at. It joins on node_id and its
  content-drift bucket keys on head_sha (diff-fleet.sh:28-30) — head_sha
  is the diff's own key, not a separate clock beside it. The "no edge is
  diffable" point survives; the mechanism given for it was wrong. This
  error was introduced by c1d3d9a, the commit that fixed the last one.
- eval-ladder already has the rung. Rung 1 (discriminating corpus) sits
  exactly between structural and code assertion, so what is missing is
  the retrieval-layer vocabulary, not a rung.
- The fleet-playbook table counts now agree: four rows, four theses, and
  the gap named as a fifth item absent from the table.
- Provenance said five sources; the list has six. The uncounted product
  page is load-bearing further down.
- "CMDB" is Harness's vocabulary, not ours. SKILL.md:45 says "not a
  runbook" and says nothing about a CMDB.
- The corpus now back-links here. agentic-patterns-corpus.md and its
  .json twin carried the retired evidence-quality rejection with no
  pointer, so a reader arriving via red-gate-protocol.md:745 read a dead
  verdict as live. The original verdicts are left intact and annotated:
  the evidence half is marked superseded, the domain-fit half upheld and
  now scoped to git-native repo data.

Cheap tier: 1290 passed, 0 failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

* Add companion research note on LLM-generated wikis and repo maps

Asks what structure is right for the fleet-playbook edge gap, which the
Harness note named but never answered. Surveys the LLM-wiki family
(DeepWiki, Google Code Wiki, DeepWiki-Open, Auto Wiki), repo maps,
SCIP/LSP indexes, embeddings, llms.txt, AGENTS.md, graphify and doctests.

Verdict: nothing in the family closes the gap, and PR #134's verdict
stands. The sharpening is worth more than the answer — on a DeepWiki page
pinned to a commit that is still HEAD (zero drift), three of five rows in
the architecture table are wrong, one naming a function absent from the
file. Every citation resolves. Traceability was never the problem, which
is the gap validate-citations.sh names in its own comments, reproduced at
industrial scale.

Derives an entry condition for any future edge: an edge is only worth
adding if it comes with a command that re-derives it and a comparison that
can go red. That disqualifies one of the three candidate edges in the
companion note.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

* Correct two claims that verification found false

Both were caught by dives verifying this repo's own assertions against
primary sources, and both were load-bearing where they sat.

harness-knowledge-graph.md called node_id "stable". GitHub does not say
that. The GraphQL global-node-ID guide promises only that "it's best
practice to persist the global node ID so you can easily reference objects
across API versions", and the migration guide states that "The legacy
format will be closing down and replaced with a new format" — so the
identifier's string is on the record as changing, with no shutdown date.
The verdict's domain-fit argument rests on this claim, so it now says what
is actually supported: node_id is opaque, unambiguous and independent of
the mutable full_name within a pass. It does not license joining a manifest
captured before the format migration against one captured after, which is
what diff-fleet.sh does across passes — a fleet straddling that boundary
sees every member as removed plus added rather than renamed, and nothing
in the plugin would say why.

AGENTS.md:63 and evals/cheap/run.sh:3 both claimed the cheap tier runs
"under a second". Measured on this checkout: 18.63s for 1290 checks across
25 plugins. Off by roughly 19x, in the file that tells every contributor
and every agent what the tier costs. Replaced with the measured figure, its
date, and the fact that it scales with plugin count — a number that will
go stale again is worth less than the instruction to re-measure.

Cheap tier green after the change: 1290 passed, 0 failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

* Add knowledge-management corpus: 140 patterns, 7 dives, 57 corrections

Ten scouts swept ten source clusters — retrieval and indexing, code graphs,
context files, agent memory, provenance and freshness, developer portals,
semantic layers and lineage, enterprise search and MCP, evaluating knowledge
systems, and knowledge-management prior art outside software. Seven dives
then re-verified the load-bearing claims against primary sources, organised
by cross-domain convergence rather than by scout.

Verdict: nobody machine-checks a knowledge claim. The field's answer is to
make claims that do not need checking — derive them from a parser or
compiler, execute them, or accept them as unchecked and label how they were
produced. Where a semantic check is attempted the measured ceiling is ~77%
balanced accuracy on prose (0.55 informedness against a 50% chance
baseline), ~61 on the hardest prose split, ~58 on contested cases, and 0.17
span-F1 for a prose-trained checker on code as evidence.

Three findings the survey did not expect:

Fail-open is the norm and nobody says so. Every expiry and drift mechanism
examined degrades to green rather than red — an env var that skips every
check, a swallowed network error, a doc preprocessor silent on a missing
anchor for nearly seven years, an ownership gate that goes vacuous exactly
where ownership is broken. The rule "the comparison must be able to go red"
is insufficient; it must also go red when it cannot be made.

Citation is not transcription, proven on this corpus itself. 57 of ~140
scout claims needed correction — not for missing citations but for misread
ones. The cleanest instance is external: "84% of KM programmes fail" traces
through three hops of valid, resolvable citations to a 1997 article saying
the rate is one third, and that the number is an estimate rather than a
study. Pinning where is not pinning what, and repo@sha:path pins only where.

This marketplace is further ahead than its own scouts believed and its best
evidence is buried: verify-before-claim ran a negative control three times
across six scenarios, got a null every time, and ships without a calibration
case with the reasoning recorded — an in-house reproduction of the year's
most-cited context-file null, reached before the paper it matches was
revised. It lives in a YAML comment no tier or doc references.

Two claims recorded but deliberately not acted on, because both are
plugins/** changes that trigger the behavioral tier and the demonstration
review gate: SKILL.md:65 asserts GitHub's "stable node_id", which GitHub
declines to guarantee; and index.schema.json requires repo ("owner/name")
with additionalProperties false and no node_id field, while
validate-citations.sh matches on full_name, so a rename leaves every prior
claim carrying a dead name that can resolve to the wrong repository.

Cheap tier: 1291 passed, 0 failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

* Re-measure the cheap tier after merging main

The merge brought in #136, which added checks. The figure this branch
introduced (~19s / 1290 checks, measured 2026-09-14) was already stale
against its own instruction to re-measure rather than trust it.

Measured on this checkout at 2026-09-18: 13.6s, 1296 checks, 25 plugins.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

* fleet-playbook-curator: key citations on node_id, not the mutable repo name

The plugin's thesis is that membership is "joined on node_id, never
full_name". That held for the manifest and the diff and was abandoned at
the claim ledger: gather-context.sh reduced each member to full_name and
dropped node_id one line before it was needed, so validate-citations.sh
could only match on the mutable name, and index.schema.json had no
node_id field with additionalProperties:false forbidding one.

Measured, not assumed. A plain rename already failed CLOSED (the old name
is absent from context.json, so the claim reads as fabricated). The case
that actually bites is a repo name later REUSED by a different repository:
the stale citation matched the impostor's gathered surface and resolved to
the wrong repository, silently, exit 0. GitHub's own rename doc warns that
reusing an old name retargets the redirect.

Threaded node_id through the seam: gather-context.sh carries it into every
context entry, the schema admits it as an OPTIONAL claim field, and
validate-citations.sh prefers it when present. A rename now keeps the
citation resolving (with a note); a reused name fails closed. Ledgers
written before the field validate unchanged on the old full_name match,
and the validator now says so on every run instead of leaving it implicit.

Two prose claims corrected in the same pass, both verified unsupported
against GitHub's own documentation:
- "GitHub's stable node_id" -> opaque. GitHub promises only that it is the
  value to persist "across API versions" and states the legacy format
  "will be closing down and replaced with a new format".
- "strongly-consistent rate bucket" -> authoritative. GitHub publishes no
  consistency guarantee for REST list endpoints. This is the same claim
  Copilot flagged in the research note on this branch; the note's copy was
  fixed and the plugin's original was left behind.

Regression coverage that can go red: a cite-index-hijack.json fixture plus
three checks. Verified by reverting the node_id preference and watching the
hijack check fail, then restoring it. Cheap tier 1300 passed, 0 failed.

Touches plugins/*/skills/**/scripts/** — a deep_safety_path — so the gated
pier run fires on this PR by design.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

* Move the fleet-playbook-curator fix to its own PR

Reverts b4a4dc4 on this branch. The change is not being dropped: it now
lives on claude/fleet-playbook-node-id-citations as its own pull request,
cherry-picked onto main.

This PR was opened as docs-only, and its body says so: no plugins/**, no
SKILL.md, so no behavioral tier and no demonstration comment. b4a4dc4 made
that statement false by adding a safety-path change to
plugins/*/skills/**/scripts/**. Splitting makes the body true again, lets
the research notes merge without waiting on a pier run they don't need,
and gives the plugin change its own review, demonstration comment and deep
tier run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants