Skip to content

Testing handoff: document the existing eval architecture, plan the agentic/composition test strategy, and add standing orders that keep both current #89

Description

@JRichlen

Handoff brief from the PR #82 / RQ-003 session. Three deliverables, sequenced so each lands independently. Context that motivated this: #82 reframed redgate as a specialist-first verification envelope, but no tier tests the two things redgate (and Agent OS, #83) actually claim to be — a multi-round protocol and a composition contract. Those claims are currently defended only by prose greps.

1. Document the existing eval architecture (do this first — it is the baseline the plan diffs against)

evals/README.md predates several tiers. Produce one authoritative doc (expand evals/README.md, or docs/testing.md linked from it) covering every tier that actually runs today, and for each: what it proves, what it structurally cannot prove, when it fires (per-PR / path-gated / release), cost, and how to run it locally.

Current inventory to document (verify against .github/workflows/ at time of writing — do not trust this list blindly):

  • cheap (evals/cheap/run.sh): deterministic, offline; executed script gates + load-bearing-sentence greps; manifest/marketplace/roster sync; secret gate; branch-protection drift guard (ci/check_branch_protection.py + ci/required-checks.json).
  • install tier: marketplace install-smoke + per-plugin cheap packs.
  • behavioral (per-plugin promptfoo packs): single-round pressured scenarios, LLM-judged, calibration-stub negative controls, repeat: + pass-rate floor.
  • routing tier (evals/routing/): roster-level trigger routing, deterministic regex verdicts, repeat: 5, evals/paid/pass-rate.sh as the arbiter (FAULTs excluded, 0.8 floor over valid samples, fail-closed on starvation, PROMPTFOO_RETRY_5XX for transient 5xx — see the wiring rationale in the bc07de3 commit on Make Redgate the default interactive workflow #82).
  • scale (evals/scale, e.g. redgate lifecycle stress, agent-compiler kernel stress): randomized offline stress of the same gates.
  • deep tier (pier): sandboxed cross-harness end-to-end, path-gated to the safety surface; the gate-switch caveat in the root CLAUDE.md WARNING.
  • counterfeit tier: the corpus run.
  • demonstration discipline: the human review gate (root CLAUDE.md) — not machine-enforced, and why.
  • The statistical spine shared by all LLM tiers: repeat:, k-of-N floor, FAULT/verdict separation, fail-closed starvation.

2. Plan (then implement incrementally) the agentic/composition testing strategy

Full rationale in this PR #82 discussion and the session that produced it. The layered model, cheapest first — each layer catches what the one below cannot:

  • L1 — decision-point probes (extends existing promptfoo machinery): freeze a trajectory prefix (a fabricated mid-run transcript) as the prompt and assert on the single next move. Families redgate lacks: gate classification (PATCH auto-pass vs MAJOR stop under blanket approval), T0 calibration pass-through, re-pin refusal mid-conversation, JUDGE independence without subagents. Every must-fire gets a must-not-fire twin — over-activation is the signature failure mode of an envelope skill.
  • L2 — plan-audit tier (sibling of evals/routing/): give the model a realistic problem + the installed roster, ask for its execution plan before any work, and grade the plan for explicit application of dependencies. Labeled corpus with expected composition per problem (e.g. flaky prod bug → diagnosing-bugs inside a redgate run with a MAJOR gate; trivial rename → specialist only, no envelope; an Agent OS curation case → recipe/control plane consuming, not restating, the interaction contract). Two-stage grading: deterministic assertions first (names a roster specialist; declares criteria/verifier when invoking redgate; places a gate class on the irreversible step), then an anchored LLM-judge rubric per plugin-factory's judge-calibration reference, with a stub-skill negative control. Reuses the whole statistical spine.
  • L3 — trajectory runs with artifact audit (extends pier): scripted gate-responder plays the human; the reward is a deterministic post-hoc audit of .redgate/<slug>/ (criteria pinned before TRACE evidence, gates.log matches actions, no green except via the pinned verifier, JUDGE independent). No LLM judge in the loop. Path-gated like the current deep tier.
  • L4 — cross-plugin composition runs: redgate + a specialist + Agent OS installed together, one scenario per boundary claim, release-cadence only.

Guiding principles to carry into the plan doc: test decisions not scripts; force protocol state onto disk and test the artifacts; a negative on both sides of every boundary; plans are the cheap proxy for trajectories; FAULT/verdict separation at every LLM layer; cross-harness runs only for invariants.

Coordinate with #83 (Agent OS) and the RQ-002 composition/trajectory redesign — L2's corpus encodes the boundary those define, so the corpus labels should be reviewed against whatever contract they land.

3. Standing orders: keep the testing docs continuously up to date

Documentation that isn't enforced drifts — this repo already knows that (docs-hygiene, the branch-protection drift guard). Two mechanisms, both required:

  1. Standing order in the instruction files: add to the root CLAUDE.md/AGENTS.md eval-discipline section (and the testing doc itself): any PR that adds, removes, renames, or re-scopes an eval tier, workflow job, or per-plugin pack MUST update the testing doc in the same PR — same-PR, not follow-up, so the doc can never describe a tier that no longer exists.
  2. Cheap-tier drift check that makes the order bite: a deterministic check in evals/cheap/run.sh that derives the live tier inventory (job names from .github/workflows/*.yml, pack directories from plugins/*/evals/*/, evals/*/) and asserts each appears in the testing doc — and, reverse direction, that the doc names no tier that no longer exists. Same pattern as the roster-sync and branch-protection guards: the doc becomes generated-or-verified, never trusted. A rename then fails cheap until the doc moves with it, which is exactly the standing order, machine-enforced.

Acceptance criteria

  • Testing doc exists, covers every live tier with proves/cannot-prove/trigger/cost/local-run, and is linked from evals/README.md and the root CLAUDE.md.
  • Standing order added to CLAUDE.md/AGENTS.md; cheap-tier doc-drift check red when a tier and the doc disagree (prove it red first on a deliberate mismatch, per repo discipline).
  • Strategy plan reviewed (L1–L4 scoped, costed, sequenced against experiment: Agent OS design-evidence harness (paid workflows) [do not merge — #85] #83/RQ-002); L1 + L2 implemented as the first increment; L3/L4 ticketed separately if not built here.
  • Every new LLM tier inherits the statistical spine (repeat, floor, FAULT exclusion, fail-closed starvation).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions