Skip to content

fleet-playbook-curator: key citations on node_id, not the mutable repo name - #138

Merged
JRichlen merged 7 commits into
mainfrom
claude/fleet-playbook-node-id-citations
Sep 24, 2026
Merged

JRichlen merged 7 commits into
mainfrom
claude/fleet-playbook-node-id-citations

Conversation

@JRichlen

@JRichlen JRichlen commented Sep 23, 2026 •

Copy link
Copy Markdown
Owner

Split out of #134, which is docs-only again. The first commit is a cherry-pick onto main. The two follow-ups address review findings on this PR.

The defect

The plugin's thesis is "membership is joined on node_id, never full_name". That holds for the manifest and for the diff. It breaks at the claim ledger.

Stage Keys on
list-fleet-members.sh node_id ✅
diff-fleet.sh node_id — every bucket carries it ✅
gather-context.sh:31 reduces each member to full_name and drops node_id ❌
validate-citations.sh can therefore only match full_name ❌
index.schema.json no node_id field, and additionalProperties: false forbids adding one ❌
PROMPT.md:42-43 (the only instructions the deployed curator reads) lists repo, path, sha, curated_at as the ledger fields ❌

What that does in practice — measured against the real validator

A plain rename already failed closed. The old name is missing from context.json, so the claim reads as fabricated and the validator exits 1. The finding as I first recorded it in #134 said the ledger "cannot survive a rename". That was wrong. It breaks safely.

What actually goes wrong is a reused name. The original repo leaves the fleet and a different repository takes its old name. The stale citation then matches the impostor's gathered tree and resolves to the wrong repository with exit 0. Nothing in the output says so. GitHub's rename documentation warns that reusing an old name retargets the redirect.

The fix

  • gather-context.sh writes node_id into every context entry.
  • index.schema.json accepts node_id as an optional claim field.
  • validate-citations.sh: a claim that carries node_id is matched on node_id only. A rename resolves with a note. A reused name fails closed, including against a context that has no node_ids (3e8d366, from a Copilot finding). A claim without node_id falls back to full_name, and the validator says so on every run.
  • PROMPT.md requires node_id on every ledger entry, copied verbatim from context.json (ceb325d, from a Codex finding). Without this, the deployed curator would never have written a keyed claim.

Prose claims corrected

Both were checked against GitHub's own documentation and neither is supported. They are now fixed in SKILL.md, templates/fleet.example.yaml and scripts/list-fleet-members.sh:

  • "GitHub's stable node_id" → opaque. GitHub documents it only as the value to persist "across API versions", and states that the legacy global-ID format "will be closing down and replaced with a new format".
  • "strongly-consistent" → authoritative. GitHub publishes no consistency guarantee for REST list endpoints.

Gates

  • Cheap tier: 1302 passed, 0 failed. New checks cover: the hijack is rejected, including against a legacy context; gather-context carries node_id; the schema admits it; PROMPT.md's field list includes it. Each check was confirmed to go red with its fix reverted.
  • Deep tier (pier): the seed context carries node_id, the oracle writes keyed claims, and the verifier fails any file citation without one. After the ANTHROPIC_API_KEY account was fixed, the re-run on 3e8d366 passed with a real agent: claude-code OK (reward=1), oracle OK (reward=1), nop OK (reward=0). So the agent read the new PROMPT.md and wrote keyed claims unprompted.
  • Behavioral tier: the same re-run passed all six scenarios 3/3 under the Sonnet grader. That includes the new one, where the user asks for "just repo/path/sha" and the subject still copies node_id.
  • Demonstration comments: posted below for each skill change.

What this does not fix

  • Existing ledgers, and a curator that ignores PROMPT.md. node_id stays optional for backward compatibility, so validate-citations.sh only warns on an unkeyed claim. Only the pier verifier fails one. A strict mode for new claims would need a way to tell them from old ones. It has not been built.
  • The legacy global-ID format migration. diff-fleet.sh joins manifests across passes on node_id. If one manifest was captured before GitHub's format change and the other after, every member looks like removed + added instead of renamed. It is item 6 on the knowledge-management corpus roadmap in Research: where the Harness Knowledge Graph lands, and what structure fits instead #134.

🤖 Generated with Claude Code

https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

…o name

The plugin's thesis is that membership is "joined on node_id, never
full_name". That held for the manifest and the diff and was abandoned at
the claim ledger: gather-context.sh reduced each member to full_name and
dropped node_id one line before it was needed, so validate-citations.sh
could only match on the mutable name, and index.schema.json had no
node_id field with additionalProperties:false forbidding one.

Measured, not assumed. A plain rename already failed CLOSED (the old name
is absent from context.json, so the claim reads as fabricated). The case
that actually bites is a repo name later REUSED by a different repository:
the stale citation matched the impostor's gathered surface and resolved to
the wrong repository, silently, exit 0. GitHub's own rename doc warns that
reusing an old name retargets the redirect.

Threaded node_id through the seam: gather-context.sh carries it into every
context entry, the schema admits it as an OPTIONAL claim field, and
validate-citations.sh prefers it when present. A rename now keeps the
citation resolving (with a note); a reused name fails closed. Ledgers
written before the field validate unchanged on the old full_name match,
and the validator now says so on every run instead of leaving it implicit.

Two prose claims corrected in the same pass, both verified unsupported
against GitHub's own documentation:
- "GitHub's stable node_id" -> opaque. GitHub promises only that it is the
  value to persist "across API versions" and states the legacy format
  "will be closing down and replaced with a new format".
- "strongly-consistent rate bucket" -> authoritative. GitHub publishes no
  consistency guarantee for REST list endpoints. This is the same claim
  Copilot flagged in the research note on this branch; the note's copy was
  fixed and the plugin's original was left behind.

Regression coverage that can go red: a cite-index-hijack.json fixture plus
three checks. Verified by reverting the node_id preference and watching the
hijack check fail, then restoring it. Cheap tier 1300 passed, 0 failed.

Touches plugins/*/skills/**/scripts/** — a deep_safety_path — so the gated
pier run fires on this PR by design.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb
Copilot AI lite review requested due to automatic review settings September 23, 2026 02:06
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-23T02:09:02.124146Z 853c817 PR opened
🔒 Security Review ✅ Completed 2026-09-23T02:20:30.005949Z 853c817 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Copy link
Copy Markdown
Owner Author

Demonstration — fleet-playbook-curator on real input

This PR edits a SKILL.md and two files under a skill's scripts/, so the repo's demonstration discipline applies. Everything below was run on this PR's branch (853c817), and the before comes from main (9ace33a), this PR's actual base. I re-ran all of it here instead of relying on the earlier run on #134, because a cherry-pick is a new commit.

The skill's deterministic core is list-fleet-members.sh → diff-fleet.sh → gather-context.sh → validate-citations.sh. I exercised the last stage with context bundles shaped exactly like gather-context.sh output, modelling two behaviours GitHub documents: a rename, and an old name reused by a different repository.

Input: a claim curated before anything moved, citing acme/billing-svc@aaaaaaa:README.md.

The case that bites: the old name taken by a different repository

The original repo left the fleet, and a new repo (different node_id) now holds acme/billing-svc.

Before (main):

validate-citations: all 1 citations traceable to a read surface or the manifest.
-> exit=0

The check is green, but the citation resolved against the impostor's tree. Anyone following it lands on an unrelated repository and nothing tells them. That happens even though the claim carries a node_id, because main's validator never reads that field.

After (this PR):

FABRICATED CITATION: claim cites acme/billing-svc@…:README.md (node_id R_kgDO_ORIGINAL)
but that repository was not read this pass (no such node_id in context.json).
A renamed member keeps its node_id; a REUSED name does not.
-> exit=1

The rename case, which was never the danger

The same repo, renamed to acme/payments-svc and still in the fleet.

Claim form Before (main) After
legacy (no node_id) exit 1, fails closed exit 1, unchanged
node_id-keyed exit 1, fails closed exit 0 with note: …renamed to acme/payments-svc…the citation still resolves because it is keyed on node_id

A plain rename was always fail-closed. The difference is that a keyed citation now follows the rename instead of being thrown away as fabricated.

The rule that produced each change

gather-context.sh:31 reduced each member to full_name and dropped node_id, one line before the validator needed it. The manifest and the diff are keyed on node_id and the ledger was not. There was exactly one seam.

Proof the new check can go red

I disabled the node_id preference in the validator and re-ran the cheap tier:

[FAIL] validate-citations ACCEPTED a node_id-keyed citation whose repo left the fleet
       while its NAME was reused — the citation resolved to the wrong repository

With the preference restored, the cheap tier on this branch is 1299 passed, 0 failed.

The misses

  • My first attempt at the reused-name fixture was wrong (on Research: where the Harness Knowledge Graph lands, and what structure fits instead #134, before the split). I left the renamed original in the fleet next to the impostor, predicted FAIL, and got a pass. The code was correct: the node_id followed the rename. My fixture just didn't model the threat. I rebuilt it before drawing any conclusion.
  • Legacy ledgers are flagged, not fixed. node_id is optional for backward compatibility, so any claim curated before this change still matches on the mutable name and still has the hole. You can see this in the legacy + reused name row above: exit 0, now with a warning printed, but not a failure.
  • This is unit-level, not end-to-end. The context bundles are synthetic. I never called the real gh api, so this proves the validator's logic, not that a live fleet pass fills in node_id on every code path.
  • Neither prose correction is demonstrated here. Changing "stable node_id" to "opaque" and "strongly-consistent" to "authoritative" fixes unsupported claims, checked against GitHub's docs. No run here exercises them; the behavioral tier covers prose.
  • The cross-format-migration case is not covered. If manifests straddle GitHub's legacy→new global-ID migration, every member reads as removed+added. That is out of scope and listed in the PR body.

Generated by Claude Code

Copy link
Copy Markdown
Owner Author

CI red — not this PR's. confirm grader model resolves fails, and behavioral tier (promptfoo) fails only because of it (GRADER: failure, RUN: skipped). One root cause:

##[error]grader model 'claude-sonnet-5' did not resolve (HTTP 400).

Why it isn't this PR's

What's actually wrong can't be read from the log. claude-sonnet-5 is a current, valid model ID. An unknown model returns 404; this is 400 (invalid_request_error). So the check's advice to "fix the slug" is very likely wrong. My best guess is account-level: Anthropic returns 400 when the credit balance is too low. I can't confirm that, though, because the check runs curl -o /dev/null and throws the API's error body away.

Fix: there's nothing I can port in-repo. Whoever holds the ANTHROPIC_API_KEY secret needs to look at the account (credit balance, key status, model access).

Proposed patch, not pushed here, so the next failure explains itself (.github/workflows/evals.yml, grader ping loop):

-            code=$(curl -sS -o /dev/null -w '%{http_code}' https://api.anthropic.com/v1/messages \
+            resp=$(mktemp)
+            code=$(curl -sS -o "$resp" -w '%{http_code}' https://api.anthropic.com/v1/messages \
               ...
             if [ "$code" = "200" ]; then
               echo "grader model '$grader' resolved OK."
             else
-              echo "::error::grader model '$grader' did not resolve (HTTP $code). Fix the slug in that plugin's promptfooconfig.yaml."
+              why=$(jq -r '"\(.error.type // "unknown"): \(.error.message // "")"' "$resp" 2>/dev/null || head -c 300 "$resp")
+              case "$code" in
+                404)     hint="Fix the slug in that plugin's promptfooconfig.yaml." ;;
+                401|403) hint="The ANTHROPIC_API_KEY secret is invalid or lacks access." ;;
+                *)       hint="Not a slug problem - read the API error above." ;;
+              esac
+              echo "::error::grader model '$grader' failed (HTTP $code) - $why. $hint"
               fail=1
             fi
+            rm -f "$resp"

⚠ Specific to this PR: because the grader never resolved, the promptfoo run was skipped, not passed. This PR changes SKILL.md prose, which is exactly what the behavioral tier exists to check, so that change currently has no behavioral coverage. The cheap tier (1299) and the pier run still hold. But this shouldn't merge on the strength of a behavioral tier that never executed. Once the account is fixed, it needs a real run.

Re-running the failed jobs once to confirm it reproduces.


Generated by Claude Code

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 853c817f57

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

One or more issues must be addressed before approval.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 High severity · 1 Medium severity · 1 Low severity

Open (3)
What changed in this PR

This PR hardens fleet citation identity by carrying GitHub node_id through context gathering and validation, while preserving legacy ledger compatibility.

Changes:

  • Adds optional node_id claim support and identity-aware validation.
  • Updates skill prose and adds reused-name hijack coverage.
  • Retains legacy full_name fallback with diagnostic output.
File Description
plugins/​fleet-playbook-curator/​skills/​fleet-playbook-curator/​templates/​fleet-playbook/​index.schema.json Updated as part of this pull request.
plugins/​fleet-playbook-curator/​skills/​fleet-playbook-curator/​SKILL.md Updated as part of this pull request.
plugins/​fleet-playbook-curator/​skills/​fleet-playbook-curator/​scripts/​validate-citations.sh Updated as part of this pull request.
plugins/​fleet-playbook-curator/​skills/​fleet-playbook-curator/​scripts/​gather-context.sh Updated as part of this pull request.
plugins/​fleet-playbook-curator/​evals/​cheap/​fixtures/​cite-index-hijack.json Updated as part of this pull request.
plugins/​fleet-playbook-curator/​evals/​cheap/​fixtures/​cite-context.json Updated as part of this pull request.
plugins/​fleet-playbook-curator/​evals/​cheap/​checks.sh Updated as part of this pull request.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

The deployed curate step pipes only PROMPT.md and context.json to the agent
(fleet-sync.yml), and PROMPT.md still listed repo/path/sha/curated_at as the
ledger fields. The validator matches on node_id only when a claim carries it,
so freshly curated claims stayed keyed on the mutable repo name -- the exact
reused-name misattribution the previous commit set out to close.

- PROMPT.md: node_id is a required ledger field, copied verbatim from
  context.json, never invented.
- pier: the seed context carries node_id (as gather-context now emits it), the
  oracle writes keyed claims, and the verifier fails a file citation with no
  node_id. Checked locally: oracle ledger -> reward 1, same ledger with node_id
  stripped -> reward 0.
- promptfoo: a scenario where the user asks for a short repo/path/sha entry and
  the rubric requires node_id copied exactly.
- cheap: PROMPT.md's field list must include node_id (verified to FAIL with it
  removed).
- fleet.example.yaml: "stable node_id" -> "opaque", matching SKILL.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

Copy link
Copy Markdown
Owner Author

Demonstration: ceb325d (PROMPT.md requires node_id)

Input. The PR's existing name-reuse fixture, evals/cheap/fixtures/cite-index-hijack.json. One claim cites o/new-svc:README.md and was curated when that name belonged to repository R_ORIGINAL. This pass, cite-context.json shows the name o/new-svc now belongs to R_NEW.

Before. This is the ledger the old PROMPT.md:42-43 field list produced (repo, path, sha, curated_at, no node_id):

validate-citations: all 1 citations traceable to a read surface or the manifest.
validate-citations: 1 claim(s) carried no node_id and were matched on the MUTABLE full_name. …
exit=0

The citation resolves to the wrong repository and CI passes.

After. The same claim written under the new field list (node_id copied from context):

FABRICATED CITATION: claim cites o/new-svc@…:README.md (node_id R_ORIGINAL) but that repository was not read this pass (no such node_id in context.json). A renamed member keeps its node_id; a REUSED name does not.
validate-citations: 1 non-traceable citation(s) — failing.
exit=1

Rule that produced the change: PROMPT.md step 2: "every claim has an entry with node_id, repo, path, sha, curated_at … Copy node_id verbatim from the repo's entry in context.json."

Pier verifier, run locally against the task's real seed. Docker was not used. The paths were rewritten to a scratch dir.

oracle ledger (keyed)            … PASS every file citation carries node_id      reward=1
same ledger, node_id stripped    … FAIL 1 file citation(s) carry no node_id      reward=0

Misses

  • No model has run the new prose yet. Both runs above are deterministic: they show what the rule enforces once a model follows it, not whether a model does. The behavioral scenario and the claude-code pier agent could show that. Both are currently blocked by the ANTHROPIC_API_KEY account problem already reported on this PR. This container has no key either.
  • The production gate still accepts unkeyed claims. validate-citations.sh only warns on them, so that older ledgers keep validating. A curator that ignores PROMPT.md still passes fleet-sync.yml CI. Only the pier verifier fails it. One option is a strict mode (fail on unkeyed new claims), which would need a way to tell new claims from old ones. I have not built it.
  • fleet.example.yaml still said "stable node_id" after the last commit. My earlier sweep for "stable" missed it. It is fixed here.

Generated by Claude Code

…xt format

validate-citations matched on node_id only when BOTH the claim and context.json
carried node_ids. Against a context with none, a keyed claim fell back to
full_name, so the reused-name hijack the previous commits close was accepted
again -- and the summary misreported it as a claim that "carried no node_id".
A claim that carries node_id is now matched on node_id only; the full_name
fallback is reserved for claims that omit it (legacy ledgers).

cheap: new fixture cite-context-legacy.json (cite-context.json with context
node_ids stripped) and a check that the keyed hijack still fails against it.
Verified the check FAILs with the validator change reverted.

list-fleet-members.sh: header comments still called node_id STABLE and the
repo listings strongly consistent; aligned with SKILL.md (opaque, authoritative).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

Copy link
Copy Markdown
Owner Author

Demonstration: 3e8d366 (a keyed claim fails closed regardless of context format)

Input. cite-index-hijack.json, a claim keyed R_ORIGINAL that cites o/new-svc:README.md. It is checked against cite-context.json with the node_id removed from every context[] entry. That is the context shape from before gather-context emitted node_ids, and it is now committed as cite-context-legacy.json.

Before (the validator at ceb325d):

validate-citations: all 1 citations traceable to a read surface or the manifest.
validate-citations: 1 claim(s) carried no node_id and were matched on the MUTABLE full_name. …
exit=0

It accepts the reused name, and the summary is also wrong: the claim did carry a node_id.

After:

FABRICATED CITATION: claim cites o/new-svc@…:README.md (node_id R_ORIGINAL) but that repository was not read this pass (no such node_id in context.json). …
exit=1

Rule: if [ -n "$node_id" ]. The branch now depends on the claim alone, and an empty read_node_ids fails the lookup instead of routing to the full_name fallback. cite-index-good.json still passes against the normal context (exit 0, 3 citations).

Misses

  • The new cheap fixture strips node_ids from context[] only. removed[] still carries R_GONE, as it did before this PR, because the diff always had node_ids. I did not test a context with no removed[] node_ids, since the validator never reads removed[].
  • This makes a keyed claim strictly less forgiving. If someone runs a new ledger against an old gather-context.sh, every keyed file claim now fails. I think that is correct: it is the hijack case. But it is a behavior change for a mixed-version install, and no test here covers an upgrade sequence.

Generated by Claude Code

Copy link
Copy Markdown
Owner Author

deep tier — pier run (fleet-playbook-curator) is red on 3e8d366, and the cause is outside this PR. The part of the check that tests this PR's change passes.

claude-code  ERROR (ERR:NonZeroAgentExitCodeError)
oracle       OK (reward=1, expected=1)
nop          OK (reward=0, expected=0)
  • Oracle and nop both pass. This is the first CI run with the new verifier check ("every file citation carries node_id"). The oracle's keyed ledger scores 1 and nop scores 0. So the seed, the oracle and the verifier all behave correctly with this PR's change.
  • claude-code fails. The agent exits nonzero before it produces any output, and it did the same on 853c817. That run uses the ANTHROPIC_API_KEY secret, the same key whose grader ping returns HTTP 400 on this PR and on Research: where the Harness Knowledge Graph lands, and what structure fits instead #134. Research: where the Harness Knowledge Graph lands, and what structure fits instead #134 does not touch fleet-playbook-curator and fails identically on the grader. No fix exists that I can port, because the fix is at the account level.
  • Re-run: I am holding this PR's one re-run until the key is fixed. A re-run now would only reproduce the same failure. Once the key works, that re-run is the real test of whether a model follows the new PROMPT.md.

The deep tier (pier) failure on ceb325d is not a signal at all. That run was cancelled, because 3e8d366 superseded it.


Generated by Claude Code

Brings in #137 (paid tiers gated behind the paid-evals label).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

Copy link
Copy Markdown
Owner Author

Resolved: the key is fixed, and the held re-run on 3e8d366 passed all 52 jobs. This also closes the first miss in the ceb325d demonstration ("No model has run the new prose yet"):

  • Behavioral (Sonnet grader, 3 repeats per scenario): all six scenarios passed 3/3. That includes the new one, keys a new ledger entry on node_id, not only the mutable repo name. The subject model copied R_kgDOM0n1 into the entry even though the user asked for "just repo/path/sha".
  • Deep (pier): claude-code OK (reward=1), oracle OK (reward=1), nop OK (reward=0). The claude-code agent read the new PROMPT.md and produced a ledger that passes verifier check 4, "every file citation carries node_id". So a real harness wrote keyed claims unprompted.

Since then I've merged main in (f4cea99), which brings #137. On this head the paid legs skip unless the PR carries paid-evals. That merge changed only .github/workflows/evals.yml, and the cheap tier passes 1302/1302.


Generated by Claude Code

Brings in #131 (subject model repinned to qwen/qwen3.8-flash, subject-model preflight, pack rework).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb
Conflict in fleet-playbook-curator's cheap checks: kept both the node_id
identity checks (this branch) and the pier seed workflow/tree check (#140).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb
@JRichlen JRichlen added the paid-evals label Sep 24, 2026 — with Claude
…r entry

The behavioral tier on fc2294c failed "keys a new ledger entry on node_id" at
1/3. One subject answer kept node_id but wrote `<head_sha>` for the sha and a
truncated repo name. SKILL.md's own example ("as of `<sha>` it held the
inventory") shows the placeholder form, and only PROMPT.md, which the
behavioral tier does not inject, said a placeholder sha is invalid.

SKILL.md now states that a ledger entry holds real values copied verbatim
(node_id and the full repo name from context.json, sha from the manifest
head_sha), and that the `<sha>` in its examples is never a value to write.
The rubric is unchanged. A cheap check pins the sentence, and it fails with
the sentence removed (1321/1322) and passes with it (1322/1322).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DghMFX2SuCASz3fjHoYwjb

JRichlen commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner Author

Demonstration: the no-placeholder-sha rule in SKILL.md (15708d7)

Input. The behavioral scenario "keys a new ledger entry on node_id". The user supplies context.json for acme/ansible-homelab-monitoring (node_id R_kgDOM0n1), a manifest head_sha of mmm2222, and asks for a short ledger entry. SKILL.md is injected. The subject is qwen/qwen3.8-flash and the grader is Haiku. Both are real CI runs.

Before: fc2294c, 1/3, below the 0.6 floor

One failing row, verbatim from the job log:

{
  "node_id": "R_kgDOM0n1",
  "repo": "acme/ansible-homelab-monitor",
  "claim": "check `site.yml` — as of <head_sha> it is the monitoring playbook entry point",
  "citation": "acme/ansible-homelab-monitor@<head_sha>:site.yml",
  ...
}

The subject kept node_id, but wrote the placeholder <head_sha> instead of mmm2222 and truncated the repo name. The grader failed it on the sha, correctly.

Why. SKILL.md's own directive example is "check hosts.yml — as of <sha> it held the inventory." The rule that a placeholder sha is invalid lived only in PROMPT.md. The deployed curator reads that file, but this tier doesn't inject it.

The rule added

A ledger entry holds real values, copied verbatim. node_id and repo (the full_name, unabridged) come from the repo's context.json entry; sha is that repo's manifest head_sha. The <sha> in the examples here marks where the real sha goes — it is never a value to write. An entry with a placeholder or empty sha is invalid; if you do not have the sha, you do not have a citation.

The rubric is unchanged. A cheap check pins the sentence: it fails with the sentence removed (1321/1322) and passes with it in (1322/1322).

After: 15708d7, 3/3, all six scenarios 3/3

Two of the three rows, from the job log's results table (the table truncates each cell):

{ "node_id": "R_kgDOM0n1", "repo": "acme/ansible-homelab-monitoring", "sha": "mmm2222",
  "path": "site.yml", "claim": "site.yml installs prometheus/grafana", "as_of": "2026-07-11" }
"repo": "acme/ansible-homelab-monitoring", "sha": "mmm2222", "path": "site.yml", "as_of": "2026-07-11",
"claim": "check site.yml — as of mmm2222 it configures prometheus/grafana"

The third row opens by flagging that the subject has the structural pieces but has not read site.yml's contents this pass. That's the "traceable ≠ supported" distinction the skill draws. All three show the real sha and the full repo name. Deep tier on the same commit: pier fleet-playbook-curator passed with Haiku as the agent against the verifier's "every file citation carries node_id" check.

Misses and limits

  • The other before-failure is not something this rule addresses. One fc2294c row never wrote an entry. It explained an invented {"node_id": "R_kgDOM0n1", "full_name": "acme/homelab"} object instead. That is a subject task-following failure, and no sentence in SKILL.md targets it. It didn't recur in these 3 samples, but 3 samples is no evidence it's gone.
  • n = 3 on each side. This shows the placeholder failure went from present to absent across one run each. It is not a measured rate.
  • The truncated table hides one row's node_id. For the first after-row, the node_id line is above the visible cells in the log. The grader passed it, and the rubric fails any entry that omits or alters node_id.
  • A placeholder sha is forbidden but not enforced at runtime. index.schema.json already rejects one: sha has pattern ^[0-9a-f]{7,40}$, described as "no placeholders". But nothing in fleet-sync.yml validates the ledger against the schema, and validate-citations.sh checks only that the repo was read and that the path is in its tree, not the sha. A deployed curator that wrote <head_sha> would pass today's pipeline. This rule makes that less likely; it does not make it impossible. Wiring schema validation into fleet-sync.yml is a separate change, not made here.

Generated by Claude Code

@JRichlen
JRichlen merged commit 25c6e10 into main Sep 24, 2026
55 checks passed
JRichlen pushed a commit that referenced this pull request Sep 24, 2026
JRichlen pushed a commit that referenced this pull request Sep 25, 2026
main's #138 changed validate-citations.sh, so the pos-01 card's pinned
sha256 no longer matched and redteam/bin/generate.py --check crashed with
"task input hash mismatch" (test_redteam_design drift tests, CI run
36078213798). Update the pin and regenerate the redteam configs, which
also picks up placebo drift from #133's skill edits.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
JRichlen added a commit that referenced this pull request Sep 26, 2026
…ed-team lane (#115)

* feat: add jori coordination plugin

* docs(jori): add sourced model routing guidance

* Apply batched suggestions from code review

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

* fix(jori): clarify GitHub routing controls

* docs(examples): correct marketplace coverage counts

* feat(evals): stage GLM subject and price monitor

* redgate: make criteria-index run ordering locale-stable (LC_ALL=C sort)

The committed .redgate/INDEX.md was generated under a C-locale sort; under
en_US.UTF-8 the same corpus sorts slice2-reconcile before slice2-reconcile-r2
and the cheap-tier drift gate reported a phantom drift. Pin the sort so the
index is byte-identical regardless of the invoking shell's locale.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN

* redgate: point hooks.json at hooks/hooks-handlers/ where the handlers actually live

hooks.json wired both handlers as ${CLAUDE_PLUGIN_ROOT}/hooks-handlers/<script>.sh,
but the scripts are checked in one level deeper at hooks/hooks-handlers/. On an
installed plugin the PreToolUse write guard and the SessionStart announcement
therefore failed with 'No such file' and the guard was silently absent. Found by
the agentic protocol lane's real hook-subprocess tests (T20-T22), which resolve
hook commands from hooks.json instead of a hardcoded list.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN

* evals/agentic: core, measurement and protocol lanes (wave 1, T01-T10, T19-T24, T32-T41)

Stdlib-only unittest framework under evals/agentic/: frozen vocabulary
(contract.py, io.py with a fail-closed JSON-Schema subset validator), terminal-
state classifier and controls/detectors (classify.py, controls.py), attempt
accounting/analysis/reporting per the benchmark spec (accounting.py,
analysis.py, reporting.py, usage/judgement schemas), and real hook-subprocess +
MCP stdio protocol fixtures (protocols.py). 202 tests, all offline. Catalog
fragments for core/measurement/protocol; counterfeit fixtures 21, 22, 24 (inert
until the integration lane wires cheap section 22 and counterfeit staging).

Shared-file edits (integration-owned): package skeleton, docs/testing.md gains
'eval-dir: evals/agentic', and check-testing-doc.sh skips __pycache__/ which the
unittest tiers create under evals/.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN

* evals/agentic + evals/redteam: registry/corpus, native adapters, red-team lanes (wave 2, T11-T18, T25-T31, T42-T48)

Registry: live 25-plugin roster, catalog loader/resolver, corpus validator with
two-sided verifiers, vacuity, holdout/leakage/strata; 75 cards (positive,
negative, near-miss for every plugin) with executable outcome AND adoption
verifiers; pairing.py exposure parity as matched substitution, four estimand
arms, guidance-only degeneracy recorded not zeroed, version estimand targets
found by git survey (agent-compiler, fleet-playbook-curator, redgate, voice).

Adapters: driver configs loaded from fixtures and checked against the installed
claude/codex --help on every run; spawn refused without an approval token
(proven by a Popen trap); append-only HMAC hash-chained host ledger whose
witness is fixed at construction; replay/worker sessions can only ever yield
SIMULATED/REAL_FIXTURE evidence; no stream grammar shipped (none captured).
T26-T29 exist only as *__offline_form and stay BLOCKED pending approval.

Red team: pinned promptfoo 0.122.0 by path (never npx), fail-closed version
check, documented custom-provider interface, frozen sha256 corpus, safe/
vulnerable/refusenik scripted controls, 31 generated offline configs for the
clean/adversarial x baseline/placebo/treatment design with parity digests,
protected-effect assertions that outrank rubric prose, verdict.py as sole
judge calling the agentic native-proof gate, docker-only netproof. T45/T46
stay paid-required.

321 offline tests green; counterfeit fixtures 23, 25-31 added (inert until
the integration lane wires cheap section 22). test_protocols.py updated for the
fixed redgate hook path; docs/testing.md gains 'eval-dir: evals/redteam'.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN

* evals/agentic: integration lane — public CLI, catalog index, lifecycle paths, cheap/counterfeit wiring, docs (wave 3, T49-T52)

run.py exposes --offline/--gate/--catalog/--id/--lane, coverage --json and
driver --dry-run; driver --spawn raises ApprovalRequired without a token in a
manifest's approvals. --catalog resolves all 52 IDs to executing tests with
assertion counting, negative-control siblings that must fail under the control
fixture, and prints the frozen summary (46 executed, 6 BLOCKED pending
approval). --gate is the root-portable subset (T52 sweeps the catalog, with
requires_real_marketplace / reentrant_unsafe / heavy_external exclusions) and
runs in ~5 s; a T52 self-recursion and a T51 whole-corpus recursion that made
the gate take 30+ minutes were removed. Cheap tier section 22 runs both new
suites fail-closed; counterfeit staging covers both trees, the corpus is 31
fixtures (all fire), and COUNTERFEIT_ONLY selects one fixture for the bounded
T51 test. Counterfeit runs now shim npx/npm out of PATH after fixture 26's
first mutation executed a real npx call and upgraded the host's shared cache;
promptfoo is pinned to a separate verified 0.122.0 install. verdict.py
classifies provider faults as FAULT before the VACUOUS check. Lifecycle tests
cover terminal, correction, approval, compaction and cancellation paths with
negative controls. Docs carry measured costs; testing-plan L3/L4 read
'framework live (offline forms); native runs approval-gated'.

Verified by the coordinator: 357 tests OK; cheap 1,740/0 (22 s); counterfeits
38/0 (5.5 min); agentic --offline PASS (4.3 min); --catalog PASS (2 min);
redteam --offline PASS (1.7 min); redteam --gate ~1 s; doc guard 84 entries;
git diff --check clean.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN

* evals/agentic + evals/redteam: independent adversarial review and repair (wave 4/4B)

Six independent review lenses (native trust, statistics, causal validity and
corpus, red-team effects, catalog vacuity with mutation testing, process safety
and docs honesty) reported 83 findings, 53 blocker/major; lane-scoped repair
agents applied the confirmed ones and a second pass closed the cross-lane
residuals. Highlights:

Trust: HostLedger keys are host capabilities, never raw caller bytes; a lying
reader with zero host-observed entries can no longer satisfy assert_native_
claims; verdict.py lost its raw-key flag and its native path is now a live,
tested gate instead of dead code; attempt.schema binds native-proven to a
native adapter; claims_native counts event_ids; driver --dry-run validates
flags against the installed help; attempt_from_session is the production seam
from a live session to a contract.Attempt (native branch untestable offline,
recorded in known-gaps.md).

Statistics: matched_pairs/difference_interval apply the scoring-valid filter;
tri-state verdicts are never coerced; per-card rates over each arm's own
trials; 2x2 exclusions printed in every format; min_valid/min_clusters/margin
declared in the manifest, no code defaults; DEFF floor; FAULT-starved red-team
tranches are INCOMPLETE under a declared fault ceiling; clustered intervals on
every red-team cell and interaction delta.

Corpus: verifiers are card-bound (a copied fixture no longer passes another
card); a correct direct baseline can pass the outcome verifier; adoption
verifiers are per-plugin; deterministic arm ids and config hashes; evidence
manifests generated for all 150 fixtures with forged/stale detection; stale
README and DEFECT text corrected; run.py's parity probe covers all 25 arms;
holdout paraphrases selected per attempt from the run seed; planned_n and
SPAWN/EXIT accounting in the manifest.

Repo hygiene: 78 generated red-team logs untracked; cheap tier now fails on
tracked-but-ignored files; runners proven to write only git-ignored paths.

Verified by the coordinator after all repairs: 540 tests OK (182 s); cheap
1,903/0 (30 s); counterfeits 38/0 with all 31 fixtures firing (544 s);
agentic --offline PASS (326 s); --catalog 46 executed / 6 BLOCKED (146 s);
redteam --offline PASS (103 s); gates 10 s / 1 s; doc guard 84 entries;
git diff --check clean; no leaked processes or containers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN

* ci: install the pinned tooling the agentic and red-team gates verify against

The cheap and counterfeit jobs now set up Node 22 and install exactly
promptfoo@0.122.0, @anthropic-ai/claude-code@2.1.263 and @openai/codex@0.153.4
into the runner temp dir (never npx, never @latest), exporting PROMPTFOO_HOME
and the .bin PATH. The red-team gate compares the pin by package.json and
dist-manifest digest, and the adapter lane reads the installed CLIs' --help to
refuse unsupported driver flags; neither logs in or calls a model. Without
this the always-on gate was host-bound (known-gaps.md R11) and red in CI.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN

* evals/agentic: deepen graveyard, redgate and egress-gate to 8 cards each (90 cards, 3 plugins measurable)

Five new distinct cards per plugin (2 positive, 2 must-not-fire, 1 near-miss),
each with real fixtures, card-bound two-sided verifiers, hidden pass/fail/
near-fail workspaces, holdout paraphrases and generated evidence manifests.
Coverage now reports total_cards=90, measured_plugins=3 at min_clusters=8, so
per-plugin effect intervals for these three plugins are no longer
'unavailable' by construction. Leakage scan: 0 overlaps against every
plugin's SKILL.md and commands.

Finalize audit finding (not fixed here, recorded in known-gaps.md): 26 of 31
negative cards' outcome check is trivially satisfied on an empty workspace
because the deliverable is byte-identical across a negative card's own
pass/fail fixtures; the adoption verifier still discriminates, so the 2x2
remains informative, but the outcome axis of negative cards is weak
corpus-wide.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN

* ci/adapters: surface the CLI help text when flag conformance cannot read it

On the x64 runner the npm-installed codex wrapper answers 'exec --help' with
49 characters and the T25 assertion hid what they were. Print the captured
text in the assertion message and add post-install --version/--help
diagnostics to the CI tooling step.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN

* ci: expose node on the driver's allowlisted PATH so the npm codex wrapper can run

The adapter driver executes each CLI under an env allowlist with
PATH=/usr/bin:/bin (contract 10.1). On the runner the npm-installed codex is
a '#!/usr/bin/env node' wrapper and node lives in the setup-node tool cache,
so 'codex exec --help' printed only 'env: node: No such file or directory'.
Symlink the setup-node binary into /usr/bin in the tooling step.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN

* adapters: resolve a driver's binary by its declared basename, not the config name

Control configs (invented-flag, dangerous-flag) declare the real claude binary
under a different config name. On a host where the committed absolute path is
gone (the CI runner) the loader fell back to shutil.which(<config name>),
found nothing, and T25's negative sibling could not even load -- reported as
'negative control executed 0 assertions'. Fall back by the declared binary's
basename and add a foreign-host test with its own negative.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN

* evals: first native evidence (T26-T29 live on Claude Code 2.1.263) and first red-team tranche with the real CLI as subject

Native (10 approved sessions, token user-approved-2026-09-07-native):
CliDriver.spawn is wired to a NativeSession over a HOST_OBSERVED HostLedger;
a real stream-json capture and the grammar derived from it are committed
(codex still has none and load_grammar keeps raising for it). T26 session and
turn acks, T27 resume/fork/fresh isolation (workspace diff empty, probe
answered UNKNOWN), T28 mid-turn cancel with process-group teardown, and T29
in-process ledger verification each executed native-proven under
run.py --id <T> --approval-token; without a token they stay BLOCKED and the
frozen catalog line is unchanged. Driver config deviations forced by the
installed CLI: --session-id moved to the fresh mode, fork = --resume +
--fork-session, --verbose required by stream-json, CLAUDE_CONFIG_DIR slot.
Observed: the CLI echoes a caller-minted --session-id on system/init, so the
echo is the ack; the fork id is the one harness-minted id observed.

Red-team tranche 2026-09-07-first (declared before running; token
user-approved-2026-09-07-redteam-tranche): graveyard, redgate, egress-gate,
6 cells x 8 items x repeat 1 = 144 attempts, 148/150 model calls, 0 faults,
1h25m. bin/tranche.py brokers rows over an AF_UNIX socket so one
HOST_OBSERVED ledger per plugin lives in the process that judges, and the
native gate passed for the first time (432 host-observed events per plugin,
0 caller-asserted). No safety qualification is granted: every plugin
produced protected-effect failures. egress-gate is the only nonzero safety
interaction (+0.75 [0.18, 1.32] vs placebo) and it comes with a clean-task
utility loss and a clean-condition safety loss, reported as a trade. The
injected detector fired 10 times with 0 true positives against a real
subject; corpus injections did not work and the instrument recorded ten
successes. Textual effects only; no second grader existed.

549 tests OK (4 live forms skip without a token); cheap 2,012/0; gates
11 s / 1 s; tranche validate OK.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Gx6juLZRTy4fVdS666GJAN

* fix(evals): verify task artifacts and preserve honest test outcomes

Replace claim-based task grading with isolated artifact checks, strengthen mutation and forgery controls, preserve paired uncertainty and evidence scope, and correct fixture and host portability defects.

Validation: 645 unittests (641 passed, 4 existing native-required skips); cheap gate 2036/0; all 31 counterfeit fixtures rejected with 7 controls passing. Model thresholds and historical failures remain unchanged.

* test(evals): separate plan structure from live roster coverage

* fix(evals): bound paid CI concurrency and clarify grading contracts

* evals: repin fleet-playbook-curator task input after main's #138

main's #138 changed validate-citations.sh, so the pos-01 card's pinned
sha256 no longer matched and redteam/bin/generate.py --check crashed with
"task input hash mismatch" (test_redteam_design drift tests, CI run
36078213798). Update the pin and regenerate the redteam configs, which
also picks up placebo drift from #133's skill edits.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C

* jori evals: grade calibration controls as quoted element checklists

Three jori negative controls read below floor on run 36078738817
(GitHub routing 0/3, rough-equivalence 0/3, bounded-work 1/3). The rows
show the Haiku grader inverting the "PASS only if the answer fails to
give..." double negative rather than the stub producing Jori's rules:
the rough-equivalence stub accepted the parity table and was graded as
rejecting it; the HyDRA stub never named HyDRA and was failed for not
treating HydraFusion as real.

Rewrite those three rubrics in the shape the authority control already
uses (3/3 on the same run): distinct elements, a quotation per element,
FAIL iff every element is present, and explicit notes on what does not
count. Bounded-work now keys on Jori's per-assignment dispatch contract
(model and effort, permitted actions, per-assignment stop condition,
distinct assigned/running/completed/verified states), which the real
skill's answer on that run states and a generic plan does not.

No real-skill rubric, floor, repeat count, or model changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C

---------

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Jordan Richlen <9574264+JRichlen@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants