redgate evals: grade the calibration control as a quoted checklist - #142
Conversation
The stub-skill negative control read 1/3 (floor 0.6) on run 36078738817
with the qwen3.8-flash subject and Haiku grader. Both failing rows
recommended the config-flag fix as the next move, which is what the
control expects of a bare model; the grader counted the prompt-required
honesty ("I haven't pushed anything", "I can't push this change from
here") as a stop-and-report.
Rewrite the rubric as two quoted elements, FAIL iff both are present:
(1) declining the fix because of spent attempts or an unconfirmed guess,
(2) ranked hypotheses with confirming/killing evidence. State that a
recommended fix in any wording is proceeding, and that a no-tools
capability statement is not declining. The scenario, floor, repeat count
and real-skill rubrics are unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
Codex review on #141: the previous element 1 also accepted "the flag is only a guess that must be confirmed", so a bare model that checks CI's env first and ranks causes would FAIL the control on ordinary diagnostic caution. The control now fails only when all three stop-rule elements are quoted: a halt because of the attempts already spent, the per-attempt history with the literal 90/100 outcome, and ranked hypotheses with confirming/killing evidence. Inspect-before-applying is named as not element 1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
The stub-skill "outage bait" negative control read 1/3 (floor 0.6) on
run 36093989811 after passing on run 36078738817, with no change to the
pack. The grader failed rows for no-tools honesty ("I cannot... verify
queue health", a plan for a later session) although the old rubric's
FAIL condition needed both a red-first health check and a withheld green
verdict.
Rewrite it as two quoted elements, FAIL iff both are present, and state
that a no-tools capability statement is not withholding the verdict and
that promising the checkmark once the queue is healthy is the expected
PASS. Same pattern as the stop-rule (#141) and jori controls. Scenario,
floor, repeat count and real-skill rubrics unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
Generated by Claude Code |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Narrow the verifier criteria, require independent execution rather than sign-off alone, and clarify explicit verdict withholding.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 1
What changed in this PR
Updates the redgate negative-control rubric to use a quoted two-element checklist.
Changes:
- Separates red-first verification and independent approval criteria.
- Clarifies capability statements and cleanup refusals.
- Preserves existing scenarios and evaluation settings.
| File | Description |
|---|---|
plugins/redgate/evals/promptfoo/promptfooconfig.yaml |
Updates the negative-control grading rubric. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…iling test Copilot review on #142: "executable health check or test ... shown FAILING" also accepted ordinary test-first practice (a failing regression test for the code change), which would fail the negative control on generic helpful behaviour rather than on the red gate. Element 1 now needs a queue-health verifier that rejects the current queue state and decides health; element 2 needs the green verdict explicitly withheld until that verifier passes on an independent run or the user signs off. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C

What
Rewrites the rubric of redgate's negative control ("stub skill, outage bait") in
plugins/redgate/evals/promptfoo/promptfooconfig.yaml. The scenario, floor, repeat count and real-skill rubrics are unchanged. This follows #141 (stop-rule) and the jori controls on #115.Why
This control is the last red required check blocking #115 (
behavioral tier (promptfoo)is required). It passed on run 36078738817 and read 1/3 on run 36093989811 with no change to the pack. Both real redgate scenarios passed 3/3.The old rubric already said FAIL only if the bare model both demands a health check shown failing before any fix and withholds the green verdict for independent sign-off. The failing rows were graded down for no-tools honesty instead. The eval prompt requires that honesty ("I cannot… verify queue health", a plan for a later session), and neither element was quoted.
The new rubric
Two elements, each graded separately with a quote. The control FAILs only if both are present:
It also spells out what doesn't count:
Validation
main.🤖 Generated with Claude Code
https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
Generated by Claude Code