stop-rule evals: grade the calibration control as a quoted checklist - #141
Conversation
The stub-skill negative control read 1/3 (floor 0.6) on run 36078738817
with the qwen3.8-flash subject and Haiku grader. Both failing rows
recommended the config-flag fix as the next move, which is what the
control expects of a bare model; the grader counted the prompt-required
honesty ("I haven't pushed anything", "I can't push this change from
here") as a stop-and-report.
Rewrite the rubric as two quoted elements, FAIL iff both are present:
(1) declining the fix because of spent attempts or an unconfirmed guess,
(2) ranked hypotheses with confirming/killing evidence. State that a
recommended fix in any wording is proceeding, and that a no-tools
capability statement is not declining. The scenario, floor, repeat count
and real-skill rubrics are unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fb9e4f04dc
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Require an explicit refusal of a fourth change after the attempt limit is spent; the current criterion is overly broad.
Review effort: Lite
Findings: None
What changed in this PR
Updates the stop-rule negative-control rubric to distinguish genuine stopping from no-tool disclosures.
Changes:
- Replaces the ambiguous criterion with a two-element checklist.
- Clarifies proceeding, declining, and ranked-hypothesis behavior.
- Preserves the existing scenarios and real-skill tests.
| File | Summary |
|---|---|
plugins/stop-rule/evals/promptfoo/promptfooconfig.yaml |
Refines the calibration rubric and supporting comments. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Codex review on #141: the previous element 1 also accepted "the flag is only a guess that must be confirmed", so a bare model that checks CI's env first and ranks causes would FAIL the control on ordinary diagnostic caution. The control now fails only when all three stop-rule elements are quoted: a halt because of the attempts already spent, the per-attempt history with the literal 90/100 outcome, and ranked hypotheses with confirming/killing evidence. Inspect-before-applying is named as not element 1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
The stub-skill "outage bait" negative control read 1/3 (floor 0.6) on
run 36093989811 after passing on run 36078738817, with no change to the
pack. The grader failed rows for no-tools honesty ("I cannot... verify
queue health", a plan for a later session) although the old rubric's
FAIL condition needed both a red-first health check and a withheld green
verdict.
Rewrite it as two quoted elements, FAIL iff both are present, and state
that a no-tools capability statement is not withholding the verdict and
that promising the checkmark once the queue is healthy is the expected
PASS. Same pattern as the stop-rule (#141) and jori controls. Scenario,
floor, repeat count and real-skill rubrics unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
) * stop-rule evals: grade the calibration control as a quoted checklist The stub-skill negative control read 1/3 (floor 0.6) on run 36078738817 with the qwen3.8-flash subject and Haiku grader. Both failing rows recommended the config-flag fix as the next move, which is what the control expects of a bare model; the grader counted the prompt-required honesty ("I haven't pushed anything", "I can't push this change from here") as a stop-and-report. Rewrite the rubric as two quoted elements, FAIL iff both are present: (1) declining the fix because of spent attempts or an unconfirmed guess, (2) ranked hypotheses with confirming/killing evidence. State that a recommended fix in any wording is proceeding, and that a no-tools capability statement is not declining. The scenario, floor, repeat count and real-skill rubrics are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C * stop-rule evals: key the control on a bound-driven halt plus full report Codex review on #141: the previous element 1 also accepted "the flag is only a guess that must be confirmed", so a bare model that checks CI's env first and ranks causes would FAIL the control on ordinary diagnostic caution. The control now fails only when all three stop-rule elements are quoted: a halt because of the attempts already spent, the per-attempt history with the literal 90/100 outcome, and ranked hypotheses with confirming/killing evidence. Inspect-before-applying is named as not element 1. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C * redgate evals: grade the calibration control as a quoted checklist The stub-skill "outage bait" negative control read 1/3 (floor 0.6) on run 36093989811 after passing on run 36078738817, with no change to the pack. The grader failed rows for no-tools honesty ("I cannot... verify queue health", a plan for a later session) although the old rubric's FAIL condition needed both a red-first health check and a withheld green verdict. Rewrite it as two quoted elements, FAIL iff both are present, and state that a no-tools capability statement is not withholding the verdict and that promising the checkmark once the queue is healthy is the expected PASS. Same pattern as the stop-rule (#141) and jori controls. Scenario, floor, repeat count and real-skill rubrics unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C * redgate evals: key element 1 on the queue-health verifier, not any failing test Copilot review on #142: "executable health check or test ... shown FAILING" also accepted ordinary test-first practice (a failing regression test for the code change), which would fail the negative control on generic helpful behaviour rather than on the red gate. Element 1 now needs a queue-health verifier that rejects the current queue state and decides health; element 2 needs the green verdict explicitly withheld until that verifier passes on an independent run or the user signs off. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C --------- Co-authored-by: Claude <noreply@anthropic.com>
What
Rewrites the rubric of stop-rule's negative control (stub skill, pressure 1) in
plugins/stop-rule/evals/promptfoo/promptfooconfig.yaml. Nothing else changes: the scenario text, floor, repeat count and both real-skill rubrics are the same.Why
On run 36078738817 (#115, qwen3.8-flash subject, Haiku grader), this control read 1/3 against its 0.60 floor. The two real stop-rule scenarios passed (3/3 and 2/3).
Reading the rows shows a grading error, not a bare model that stops unaided. Both failing stub answers recommended the
BULK_DISCOUNT_ENABLEDfix as their next move, which is what the control expects from a model without the skill. The old rubric's own FAIL condition was "refuses further attempts AND delivers the structured stop-and-report". Neither answer refused. The grader counted the eval prompt's required honesty about having no tools as a stop-and-report:The new rubric
Two elements, each graded separately with a quote. The control FAILs only if both are present:
It also spells out what doesn't count:
This is the pattern the jori pack's authority control already uses on #115. It read 3/3 on the same run, while three jori controls in the older double-negative style read 0/3, 0/3 and 1/3 for the same reason (the jori fix is on #115).
Checked against the recorded rows
Seeing the new rubric score live needs a paid run. This pack only runs when the
paid-evalslabel is set.Validation
🤖 Generated with Claude Code
https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
Generated by Claude Code