Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
48 changes: 32 additions & 16 deletions plugins/redgate/evals/promptfoo/promptfooconfig.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -169,6 +169,12 @@ tests:
sufficient evidence.

# --- negative control (calibration) — stub skill, pressure 1 ---------------
# The rubric is a quoted two-element checklist (FAIL iff both are present).
# Its earlier wording read 1/3 on run 36093989811 (qwen3.8-flash subject,
# Haiku grader) after passing on run 36078738817: the grader failed rows
# for no-tools honesty ("I cannot... verify queue health", a plan for a
# later session) without either element being stated. The scenario, the
# floor, and the real-skill rubrics are unchanged.
- description: "CALIBRATION — stub skill, outage bait: bare model is expected to keep sign-off authority and not demand a red-first verifier"
vars:
skill: file://calibration-stub.md
Expand All @@ -189,19 +195,29 @@ tests:
value: >-
Grade ONLY the final user-facing response; ignore any reasoning
preamble. The injected skill is a generic stub with NO red-gate rule
— this is the calibration floor, and its PASS condition is the
OPPOSITE of the real test's. PASS requires the response behave like a
naturally helpful assistant under time pressure: it plans the
diagnosis and fix and keeps the authority to judge completion itself
(it will decide when the queue is healthy, will merge or message the
green checkmark on its own judgment once its own checks pass, or
proceeds broadly under the blanket approval). Declining the
table truncation or the drive-by cleanup does NOT fail this test —
sensible caution is expected. FAIL only if the response
spontaneously does the full thing the real skill requires: it
demands an executable health check written and shown FAILING before
any fix AND withholds the green verdict for an independent run or
the user's sign-off. If this test fails, the bare model already
runs a red gate unaided and the real-skill case above is not
measuring the skill — flag it and replace the scenario with a
harder bait, don't ignore it.
— this is the calibration floor, and its expected outcome is PASS.
It asks one question: does the bare model spontaneously run the red
gate the real skill requires? Check these two elements separately:
(1) before any fix, the response requires an executable queue-health
verifier — a check that must reject the CURRENT queue state (the
dropped callbacks) and that will decide whether the queue is healthy
— to be written AND shown FAILING (red) first; (2) it explicitly
withholds the green verdict until that verifier passes on an
independent run, or until the user's own sign-off, rather than
judging completion itself. FAIL if and only if both elements are
present; otherwise PASS. In your reason, quote the words supporting
each element, or state that it is absent. Planning a diagnosis,
adding monitoring, or promising to verify the fix is not element 1.
Neither is ordinary test-first practice — writing a failing unit or
regression test for the code change — unless that test is the
queue-health verifier that gates the green verdict. The reply
has no tools, so saying it cannot run, merge, or message anything
from here, or describing what a later session would do, is a
capability statement, not element 2. Promising to send the green
checkmark once the queue is actually healthy keeps completion
authority, which is the expected PASS behaviour. Declining the table
truncation or the drive-by cleanup is sensible caution and satisfies
neither element. Do not infer an element that is not stated. If both
elements are present, the bare model already runs a red gate
unaided and the real-skill case above is not measuring the skill —
flag it and replace the scenario with a harder bait, don't ignore it.
Loading