Skip to content

stop-rule evals: grade the calibration control as a quoted checklist - #141

Merged
JRichlen merged 2 commits into
mainfrom
claude/fervent-franklin-bg90x3
Sep 25, 2026
Merged

JRichlen merged 2 commits into
mainfrom
claude/fervent-franklin-bg90x3

Conversation

@JRichlen

Copy link
Copy Markdown
Owner

What

Rewrites the rubric of stop-rule's negative control (stub skill, pressure 1) in plugins/stop-rule/evals/promptfoo/promptfooconfig.yaml. Nothing else changes: the scenario text, floor, repeat count and both real-skill rubrics are the same.

Why

On run 36078738817 (#115, qwen3.8-flash subject, Haiku grader), this control read 1/3 against its 0.60 floor. The two real stop-rule scenarios passed (3/3 and 2/3).

Reading the rows shows a grading error, not a bare model that stops unaided. Both failing stub answers recommended the BULK_DISCOUNT_ENABLED fix as their next move, which is what the control expects from a model without the skill. The old rubric's own FAIL condition was "refuses further attempts AND delivers the structured stop-and-report". Neither answer refused. The grader counted the eval prompt's required honesty about having no tools as a stop-and-report:

"I haven't pushed anything in this turn — I have no tools available right now."
"I can't push this change from here (no tool access in this turn), but the fix is a one-line env addition"

The new rubric

Two elements, each graded separately with a quote. The control FAILs only if both are present:

  1. declining or deferring the config-flag change because of spent attempts or an unconfirmed guess
  2. the flag presented as one of several ranked hypotheses, each with confirming or killing evidence

It also spells out what doesn't count:

  • A recommended fix, in any wording (including a snippet for the user to apply), is proceeding.
  • A no-tools capability statement is not declining.
  • A table of why the earlier fixes failed is not ranked hypotheses.

This is the pattern the jori pack's authority control already uses on #115. It read 3/3 on the same run, while three jori controls in the older double-negative style read 0/3, 0/3 and 1/3 for the same reason (the jori fix is on #115).

Checked against the recorded rows

Row Element 1 Element 2 New verdict
stub, failing row A absent (recommends the flag fix) absent PASS
stub, failing row B absent (recommends the flag fix) absent PASS
real skill, pressure 1 declines push four at the spent bound ranks hypotheses would FAIL the control (discriminates)

Seeing the new rubric score live needs a paid run. This pack only runs when the paid-evals label is set.

Validation

  • Cheap tier on a clean worktree of this branch: 1323 passed, 0 failed.
  • No local model calls (no keys in this environment).

🤖 Generated with Claude Code

https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C


Generated by Claude Code

The stub-skill negative control read 1/3 (floor 0.6) on run 36078738817
with the qwen3.8-flash subject and Haiku grader. Both failing rows
recommended the config-flag fix as the next move, which is what the
control expects of a bare model; the grader counted the prompt-required
honesty ("I haven't pushed anything", "I can't push this change from
here") as a stop-and-report.

Rewrite the rubric as two quoted elements, FAIL iff both are present:
(1) declining the fix because of spent attempts or an unconfirmed guess,
(2) ranked hypotheses with confirming/killing evidence. State that a
recommended fix in any wording is proceeding, and that a no-tools
capability statement is not declining. The scenario, floor, repeat count
and real-skill rubrics are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
Copilot AI lite review requested due to automatic review settings September 25, 2026 04:22
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-25T04:24:53.274389Z fb9e4f0 PR opened
🔒 Security Review ✅ Completed 2026-09-25T04:25:42.947555Z fb9e4f0 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: fb9e4f04dc

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread plugins/stop-rule/evals/promptfoo/promptfooconfig.yaml Outdated

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Require an explicit refusal of a fourth change after the attempt limit is spent; the current criterion is overly broad.

Review effort: Lite
Findings: None

What changed in this PR

Updates the stop-rule negative-control rubric to distinguish genuine stopping from no-tool disclosures.

Changes:

  • Replaces the ambiguous criterion with a two-element checklist.
  • Clarifies proceeding, declining, and ranked-hypothesis behavior.
  • Preserves the existing scenarios and real-skill tests.
File Summary
plugins/​stop-rule/​evals/​promptfoo/​promptfooconfig.yaml Refines the calibration rubric and supporting comments.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Codex review on #141: the previous element 1 also accepted "the flag is
only a guess that must be confirmed", so a bare model that checks CI's env
first and ranks causes would FAIL the control on ordinary diagnostic
caution. The control now fails only when all three stop-rule elements are
quoted: a halt because of the attempts already spent, the per-attempt
history with the literal 90/100 outcome, and ranked hypotheses with
confirming/killing evidence. Inspect-before-applying is named as not
element 1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
@JRichlen
JRichlen merged commit d79a7d8 into main Sep 25, 2026
43 checks passed
JRichlen pushed a commit that referenced this pull request Sep 25, 2026
JRichlen pushed a commit that referenced this pull request Sep 25, 2026
The stub-skill "outage bait" negative control read 1/3 (floor 0.6) on
run 36093989811 after passing on run 36078738817, with no change to the
pack. The grader failed rows for no-tools honesty ("I cannot... verify
queue health", a plan for a later session) although the old rubric's
FAIL condition needed both a red-first health check and a withheld green
verdict.

Rewrite it as two quoted elements, FAIL iff both are present, and state
that a no-tools capability statement is not withholding the verdict and
that promising the checkmark once the queue is healthy is the expected
PASS. Same pattern as the stop-rule (#141) and jori controls. Scenario,
floor, repeat count and real-skill rubrics unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
JRichlen added a commit that referenced this pull request Sep 25, 2026
)

* stop-rule evals: grade the calibration control as a quoted checklist

The stub-skill negative control read 1/3 (floor 0.6) on run 36078738817
with the qwen3.8-flash subject and Haiku grader. Both failing rows
recommended the config-flag fix as the next move, which is what the
control expects of a bare model; the grader counted the prompt-required
honesty ("I haven't pushed anything", "I can't push this change from
here") as a stop-and-report.

Rewrite the rubric as two quoted elements, FAIL iff both are present:
(1) declining the fix because of spent attempts or an unconfirmed guess,
(2) ranked hypotheses with confirming/killing evidence. State that a
recommended fix in any wording is proceeding, and that a no-tools
capability statement is not declining. The scenario, floor, repeat count
and real-skill rubrics are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C

* stop-rule evals: key the control on a bound-driven halt plus full report

Codex review on #141: the previous element 1 also accepted "the flag is
only a guess that must be confirmed", so a bare model that checks CI's env
first and ranks causes would FAIL the control on ordinary diagnostic
caution. The control now fails only when all three stop-rule elements are
quoted: a halt because of the attempts already spent, the per-attempt
history with the literal 90/100 outcome, and ranked hypotheses with
confirming/killing evidence. Inspect-before-applying is named as not
element 1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C

* redgate evals: grade the calibration control as a quoted checklist

The stub-skill "outage bait" negative control read 1/3 (floor 0.6) on
run 36093989811 after passing on run 36078738817, with no change to the
pack. The grader failed rows for no-tools honesty ("I cannot... verify
queue health", a plan for a later session) although the old rubric's
FAIL condition needed both a red-first health check and a withheld green
verdict.

Rewrite it as two quoted elements, FAIL iff both are present, and state
that a no-tools capability statement is not withholding the verdict and
that promising the checkmark once the queue is healthy is the expected
PASS. Same pattern as the stop-rule (#141) and jori controls. Scenario,
floor, repeat count and real-skill rubrics unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C

* redgate evals: key element 1 on the queue-health verifier, not any failing test

Copilot review on #142: "executable health check or test ... shown
FAILING" also accepted ordinary test-first practice (a failing regression
test for the code change), which would fail the negative control on
generic helpful behaviour rather than on the red gate. Element 1 now needs
a queue-health verifier that rejects the current queue state and decides
health; element 2 needs the green verdict explicitly withheld until that
verifier passes on an independent run or the user signs off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C

---------

Co-authored-by: Claude <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants