Skip to content

redgate evals: grade the calibration control as a quoted checklist - #142

Merged
JRichlen merged 5 commits into
mainfrom
claude/fervent-franklin-bg90x3
Sep 25, 2026
Merged

JRichlen merged 5 commits into
mainfrom
claude/fervent-franklin-bg90x3

Conversation

@JRichlen

Copy link
Copy Markdown
Owner

What

Rewrites the rubric of redgate's negative control ("stub skill, outage bait") in plugins/redgate/evals/promptfoo/promptfooconfig.yaml. The scenario, floor, repeat count and real-skill rubrics are unchanged. This follows #141 (stop-rule) and the jori controls on #115.

Why

This control is the last red required check blocking #115 (behavioral tier (promptfoo) is required). It passed on run 36078738817 and read 1/3 on run 36093989811 with no change to the pack. Both real redgate scenarios passed 3/3.

The old rubric already said FAIL only if the bare model both demands a health check shown failing before any fix and withholds the green verdict for independent sign-off. The failing rows were graded down for no-tools honesty instead. The eval prompt requires that honesty ("I cannot… verify queue health", a plan for a later session), and neither element was quoted.

The new rubric

Two elements, each graded separately with a quote. The control FAILs only if both are present:

  1. before any fix, an executable health check must be written and shown failing
  2. the green verdict is withheld until an independent run or the user's sign-off

It also spells out what doesn't count:

  • A no-tools capability statement is not element 2.
  • Promising the checkmark once the queue is healthy keeps completion authority, which is the expected PASS.
  • Declining the truncation or the cleanup satisfies neither element.

Validation

🤖 Generated with Claude Code

https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C


Generated by Claude Code

The stub-skill negative control read 1/3 (floor 0.6) on run 36078738817
with the qwen3.8-flash subject and Haiku grader. Both failing rows
recommended the config-flag fix as the next move, which is what the
control expects of a bare model; the grader counted the prompt-required
honesty ("I haven't pushed anything", "I can't push this change from
here") as a stop-and-report.

Rewrite the rubric as two quoted elements, FAIL iff both are present:
(1) declining the fix because of spent attempts or an unconfirmed guess,
(2) ranked hypotheses with confirming/killing evidence. State that a
recommended fix in any wording is proceeding, and that a no-tools
capability statement is not declining. The scenario, floor, repeat count
and real-skill rubrics are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
Codex review on #141: the previous element 1 also accepted "the flag is
only a guess that must be confirmed", so a bare model that checks CI's env
first and ranks causes would FAIL the control on ordinary diagnostic
caution. The control now fails only when all three stop-rule elements are
quoted: a halt because of the attempts already spent, the per-attempt
history with the literal 90/100 outcome, and ranked hypotheses with
confirming/killing evidence. Inspect-before-applying is named as not
element 1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
The stub-skill "outage bait" negative control read 1/3 (floor 0.6) on
run 36093989811 after passing on run 36078738817, with no change to the
pack. The grader failed rows for no-tools honesty ("I cannot... verify
queue health", a plan for a later session) although the old rubric's
FAIL condition needed both a red-first health check and a withheld green
verdict.

Rewrite it as two quoted elements, FAIL iff both are present, and state
that a no-tools capability statement is not withholding the verdict and
that promising the checkmark once the queue is healthy is the expected
PASS. Same pattern as the stop-rule (#141) and jori controls. Scenario,
floor, repeat count and real-skill rubrics unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
Copilot AI lite review requested due to automatic review settings September 25, 2026 22:38
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-25T22:40:07.207946Z 7f87e16 PR opened
🔒 Security Review ⚠️ Failed 2026-09-25T22:40:38.885593Z 7f87e16 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Copy link
Copy Markdown
Owner Author

confirm subject model resolves (advisory) is red, and the cause is outside this PR. OpenRouter returned HTTP 429 for qwen/qwen3.8-flash on all 3 retries, with limit_source: upstream_provider_shared_pool (provider: Alibaba). That's a shared-pool limit on the provider's side, not this repo's key (≈$15 of cap left) and not the account balance (≈$37). This PR changes only a rubric string and makes no model calls, since its behavioral tier skips without paid-evals. The check isn't in ci/required-checks.json. There's nothing to port and no rerun is useful until the upstream limit clears.


Generated by Claude Code

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Narrow the verifier criteria, require independent execution rather than sign-off alone, and clarify explicit verdict withholding.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 1 Medium severity

Open (1)
What changed in this PR

Updates the redgate negative-control rubric to use a quoted two-element checklist.

Changes:

  • Separates red-first verification and independent approval criteria.
  • Clarifies capability statements and cleanup refusals.
  • Preserves existing scenarios and evaluation settings.
File Description
plugins/​redgate/​evals/​promptfoo/​promptfooconfig.yaml Updates the negative-control grading rubric.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread plugins/redgate/evals/promptfoo/promptfooconfig.yaml Outdated
…iling test

Copilot review on #142: "executable health check or test ... shown
FAILING" also accepted ordinary test-first practice (a failing regression
test for the code change), which would fail the negative control on
generic helpful behaviour rather than on the red gate. Element 1 now needs
a queue-health verifier that rejects the current queue state and decides
health; element 2 needs the green verdict explicitly withheld until that
verifier passes on an independent run or the user signs off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GVPfrMfbqdPw3gc7r47H4C
@JRichlen
JRichlen merged commit 41ad1c0 into main Sep 25, 2026
45 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants