You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Testing handoff: document the existing eval architecture, plan the agentic/composition test strategy, and add standing orders that keep both current #89
Handoff brief from the PR #82 / RQ-003 session. Three deliverables, sequenced so each lands independently. Context that motivated this: #82 reframed redgate as a specialist-first verification envelope, but no tier tests the two things redgate (and Agent OS, #83) actually claim to be — a multi-round protocol and a composition contract. Those claims are currently defended only by prose greps.
1. Document the existing eval architecture (do this first — it is the baseline the plan diffs against)
evals/README.md predates several tiers. Produce one authoritative doc (expand evals/README.md, or docs/testing.md linked from it) covering every tier that actually runs today, and for each: what it proves, what it structurally cannot prove, when it fires (per-PR / path-gated / release), cost, and how to run it locally.
Current inventory to document (verify against .github/workflows/ at time of writing — do not trust this list blindly):
routing tier (evals/routing/): roster-level trigger routing, deterministic regex verdicts, repeat: 5, evals/paid/pass-rate.sh as the arbiter (FAULTs excluded, 0.8 floor over valid samples, fail-closed on starvation, PROMPTFOO_RETRY_5XX for transient 5xx — see the wiring rationale in the bc07de3 commit on Make Redgate the default interactive workflow #82).
scale (evals/scale, e.g. redgate lifecycle stress, agent-compiler kernel stress): randomized offline stress of the same gates.
deep tier (pier): sandboxed cross-harness end-to-end, path-gated to the safety surface; the gate-switch caveat in the root CLAUDE.md WARNING.
counterfeit tier: the corpus run.
demonstration discipline: the human review gate (root CLAUDE.md) — not machine-enforced, and why.
The statistical spine shared by all LLM tiers: repeat:, k-of-N floor, FAULT/verdict separation, fail-closed starvation.
2. Plan (then implement incrementally) the agentic/composition testing strategy
Full rationale in this PR #82 discussion and the session that produced it. The layered model, cheapest first — each layer catches what the one below cannot:
L1 — decision-point probes (extends existing promptfoo machinery): freeze a trajectory prefix (a fabricated mid-run transcript) as the prompt and assert on the single next move. Families redgate lacks: gate classification (PATCH auto-pass vs MAJOR stop under blanket approval), T0 calibration pass-through, re-pin refusal mid-conversation, JUDGE independence without subagents. Every must-fire gets a must-not-fire twin — over-activation is the signature failure mode of an envelope skill.
L2 — plan-audit tier (sibling of evals/routing/): give the model a realistic problem + the installed roster, ask for its execution plan before any work, and grade the plan for explicit application of dependencies. Labeled corpus with expected composition per problem (e.g. flaky prod bug → diagnosing-bugs inside a redgate run with a MAJOR gate; trivial rename → specialist only, no envelope; an Agent OS curation case → recipe/control plane consuming, not restating, the interaction contract). Two-stage grading: deterministic assertions first (names a roster specialist; declares criteria/verifier when invoking redgate; places a gate class on the irreversible step), then an anchored LLM-judge rubric per plugin-factory's judge-calibration reference, with a stub-skill negative control. Reuses the whole statistical spine.
L3 — trajectory runs with artifact audit (extends pier): scripted gate-responder plays the human; the reward is a deterministic post-hoc audit of .redgate/<slug>/ (criteria pinned before TRACE evidence, gates.log matches actions, no green except via the pinned verifier, JUDGE independent). No LLM judge in the loop. Path-gated like the current deep tier.
L4 — cross-plugin composition runs: redgate + a specialist + Agent OS installed together, one scenario per boundary claim, release-cadence only.
Guiding principles to carry into the plan doc: test decisions not scripts; force protocol state onto disk and test the artifacts; a negative on both sides of every boundary; plans are the cheap proxy for trajectories; FAULT/verdict separation at every LLM layer; cross-harness runs only for invariants.
Coordinate with #83 (Agent OS) and the RQ-002 composition/trajectory redesign — L2's corpus encodes the boundary those define, so the corpus labels should be reviewed against whatever contract they land.
3. Standing orders: keep the testing docs continuously up to date
Documentation that isn't enforced drifts — this repo already knows that (docs-hygiene, the branch-protection drift guard). Two mechanisms, both required:
Standing order in the instruction files: add to the root CLAUDE.md/AGENTS.md eval-discipline section (and the testing doc itself): any PR that adds, removes, renames, or re-scopes an eval tier, workflow job, or per-plugin pack MUST update the testing doc in the same PR — same-PR, not follow-up, so the doc can never describe a tier that no longer exists.
Cheap-tier drift check that makes the order bite: a deterministic check in evals/cheap/run.sh that derives the live tier inventory (job names from .github/workflows/*.yml, pack directories from plugins/*/evals/*/, evals/*/) and asserts each appears in the testing doc — and, reverse direction, that the doc names no tier that no longer exists. Same pattern as the roster-sync and branch-protection guards: the doc becomes generated-or-verified, never trusted. A rename then fails cheap until the doc moves with it, which is exactly the standing order, machine-enforced.
Acceptance criteria
Testing doc exists, covers every live tier with proves/cannot-prove/trigger/cost/local-run, and is linked from evals/README.md and the root CLAUDE.md.
Standing order added to CLAUDE.md/AGENTS.md; cheap-tier doc-drift check red when a tier and the doc disagree (prove it red first on a deliberate mismatch, per repo discipline).
Handoff brief from the PR #82 / RQ-003 session. Three deliverables, sequenced so each lands independently. Context that motivated this: #82 reframed redgate as a specialist-first verification envelope, but no tier tests the two things redgate (and Agent OS, #83) actually claim to be — a multi-round protocol and a composition contract. Those claims are currently defended only by prose greps.
1. Document the existing eval architecture (do this first — it is the baseline the plan diffs against)
evals/README.mdpredates several tiers. Produce one authoritative doc (expandevals/README.md, ordocs/testing.mdlinked from it) covering every tier that actually runs today, and for each: what it proves, what it structurally cannot prove, when it fires (per-PR / path-gated / release), cost, and how to run it locally.Current inventory to document (verify against
.github/workflows/at time of writing — do not trust this list blindly):evals/cheap/run.sh): deterministic, offline; executed script gates + load-bearing-sentence greps; manifest/marketplace/roster sync; secret gate; branch-protection drift guard (ci/check_branch_protection.py+ci/required-checks.json).repeat:+ pass-rate floor.evals/routing/): roster-level trigger routing, deterministic regex verdicts,repeat: 5,evals/paid/pass-rate.shas the arbiter (FAULTs excluded, 0.8 floor over valid samples, fail-closed on starvation,PROMPTFOO_RETRY_5XXfor transient 5xx — see the wiring rationale in thebc07de3commit on Make Redgate the default interactive workflow #82).evals/scale, e.g. redgate lifecycle stress, agent-compiler kernel stress): randomized offline stress of the same gates.CLAUDE.mdWARNING.CLAUDE.md) — not machine-enforced, and why.repeat:, k-of-N floor, FAULT/verdict separation, fail-closed starvation.2. Plan (then implement incrementally) the agentic/composition testing strategy
Full rationale in this PR #82 discussion and the session that produced it. The layered model, cheapest first — each layer catches what the one below cannot:
evals/routing/): give the model a realistic problem + the installed roster, ask for its execution plan before any work, and grade the plan for explicit application of dependencies. Labeled corpus with expected composition per problem (e.g. flaky prod bug →diagnosing-bugsinside a redgate run with a MAJOR gate; trivial rename → specialist only, no envelope; an Agent OS curation case → recipe/control plane consuming, not restating, the interaction contract). Two-stage grading: deterministic assertions first (names a roster specialist; declares criteria/verifier when invoking redgate; places a gate class on the irreversible step), then an anchored LLM-judge rubric perplugin-factory's judge-calibration reference, with a stub-skill negative control. Reuses the whole statistical spine..redgate/<slug>/(criteria pinned before TRACE evidence, gates.log matches actions, no green except via the pinned verifier, JUDGE independent). No LLM judge in the loop. Path-gated like the current deep tier.Guiding principles to carry into the plan doc: test decisions not scripts; force protocol state onto disk and test the artifacts; a negative on both sides of every boundary; plans are the cheap proxy for trajectories; FAULT/verdict separation at every LLM layer; cross-harness runs only for invariants.
Coordinate with #83 (Agent OS) and the RQ-002 composition/trajectory redesign — L2's corpus encodes the boundary those define, so the corpus labels should be reviewed against whatever contract they land.
3. Standing orders: keep the testing docs continuously up to date
Documentation that isn't enforced drifts — this repo already knows that (
docs-hygiene, the branch-protection drift guard). Two mechanisms, both required:CLAUDE.md/AGENTS.mdeval-discipline section (and the testing doc itself): any PR that adds, removes, renames, or re-scopes an eval tier, workflow job, or per-plugin pack MUST update the testing doc in the same PR — same-PR, not follow-up, so the doc can never describe a tier that no longer exists.evals/cheap/run.shthat derives the live tier inventory (job names from.github/workflows/*.yml, pack directories fromplugins/*/evals/*/,evals/*/) and asserts each appears in the testing doc — and, reverse direction, that the doc names no tier that no longer exists. Same pattern as the roster-sync and branch-protection guards: the doc becomes generated-or-verified, never trusted. A rename then fails cheap until the doc moves with it, which is exactly the standing order, machine-enforced.Acceptance criteria
evals/README.mdand the rootCLAUDE.md.CLAUDE.md/AGENTS.md; cheap-tier doc-drift check red when a tier and the doc disagree (prove it red first on a deliberate mismatch, per repo discipline).