Skip to content

feat: judge and semantic on by default; name the missing evaluator key - #45

Merged
babaliauskas merged 16 commits into
mainfrom
spec/evaluators-on-by-default
Oct 10, 2026
Merged

babaliauskas merged 16 commits into
mainfrom
spec/evaluators-on-by-default

Conversation

@babaliauskas

Copy link
Copy Markdown
Collaborator

Summary

llm_judge and semantic are the evaluators that say whether the new model's prose is as good as the old one's. A fresh project got the least out of them, for three reasons:

  1. semantic shipped commented out for Anthropic and DeepSeek, which have no embeddings endpoint. The YAML comment was the only explanation.
  2. A missing judge or embedding key failed silently. compare checked keys only for the two arms being compared. A keyless judge raised on every pair; being advisory, nothing gated on it and nothing was printed, so the user paid for retries and never found out.
  3. The "promote" advice was generic. Every all-advisory run ended inconclusive with the same "Set blocking: true on at least one trusted evaluator…" line, with no evaluator named and no "is my suite big enough yet".

This PR fixes all three:

  • init turns everything on. semantic is active for every provider. Anthropic and DeepSeek borrow an embedder: OpenAI's if OPENAI_API_KEY is set, else Gemini's if GEMINI_API_KEY/GOOGLE_API_KEY is set, else OpenAI's anyway. Next steps name a missing borrowed key, and init --ci wires it as a second secret.
  • Keys are checked before the first call. compare and run check every judge and embedding model up front, and evaluate applies the same rule:
    • Advisory, no key: ⚠ <label> skipped: no API key for <model> — export <VAR> to enable it. The run continues without it, making no calls and no retries.
    • Blocking, no key: exit 1 with the env var named.
    • Every evaluator would be skipped: exit 1 before any spend.
    • semantic forced blocking by a blocking tool_arguments entry (one using the semantic strategy): it counts as a gate, and the message names that entry.
  • Skips are recorded. state.json gains skipped_evaluators. Each skip becomes a line in the verdict's recommendations, which reach the terminal, the HTML report, migration_decision.json and the hosted bundle. Local and hosted output match when a migration_policy is configured.
  • Promotion advice uses the run's own numbers. When nothing gates, an advisory judge is compared against 20 pairs per prompt (MIN_N_RELIABLE), and the smallest prompt decides:
    • At or above 20: "The equivalence judge scored at least 24 pairs on every prompt — enough to gate…".
    • Below 20: "…summarize has 8. Collect more examples…".
    • A judge that measured nothing is not told to collect more.
  • doctor gains an evaluator keys row: fail for blocking, warn for advisory. compare's doctor stage ignores this row because its own check is scoped to the suite being run, so a keyless judge in suite B never fails compare --suite-name A.
  • Smaller changes:
    • One shared missing_api_keys helper replaces three copies; an empty-string env var counts as unset.
    • Bare text-embedding-* ids resolve to OpenAI, so the config default's key is checked.
    • Missing-key errors are titled "Missing API key", not "Invalid config".

Also in this branch: the design spec and implementation plan (docs/superpowers/specs/2026-10-09-evaluators-on-by-default-design.md, docs/superpowers/plans/2026-10-09-evaluators-on-by-default.md), plus DOCS.md, llms-full.txt, docs/configuration.md, docs/faq.md, docs/getting-started.md, docs/github-action.md and the CHANGELOG under Unreleased (Added / Changed / Fixed).

Upgrade impact

semantic/llm_judge default to blocking: true in the library. A hand-written config that doesn't say blocking: false and whose judge or embedder has no key used to finish as inconclusive (exit 0, passing --policy-gate). Now compare, run, evaluate and doctor exit 1 before any call. The CHANGELOG ### Changed entry and the DOCS.md upgrade paragraph say so. I'd ship this as a minor bump, not a patch.

Known, deliberately left

  • When several tool_arguments entries force semantic blocking, only the first is named. A rerun names the next.
  • When semantic is blocking by its own flag and forced, the hint suggests only blocking: false. Exporting the key fixes it either way.
  • doctor doesn't detect the "every evaluator would be skipped" case; it warns and exits 0, while compare/run exit 1.
  • Possible overlap with fix: doctor no longer fails CI over a missing store library; messages name the package #44: both touch doctor.py, compare.py's doctor stage and the CHANGELOG Unreleased section. Whichever merges second will need a small rebase.

Test plan

  • pytest: 2833 passed, coverage 94.84% (make ci)
  • ruff check, ruff format --check clean
  • mypy --strict src/evalshift_cli clean
  • pre-commit run --all-files clean
  • CHANGELOG.md updated under ## [Unreleased]
  • Docs updated (DOCS.md, llms-full.txt, docs/ pages)

Each of the 8 tasks was test-driven and reviewed separately. The whole branch then had a final review that probed an empty CI secret, an upgraded hand-written config, a re-run of evaluate after exporting the key (stale skips are cleared), and multi-suite compare. Its findings were fixed and re-reviewed.

🤖 Generated with Claude Code

babaliauskas and others added 16 commits October 9, 2026 22:52
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… are OpenAI

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ess gate

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… evaluators

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…te being run

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…edder

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… forcing tool_arguments entry

The key preflight now refuses a config whose every evaluator would be
skipped for a missing key before any arm call, instead of letting the
evaluate stage fail after the run was paid for.

A keyless semantic that is a gate only because a blocking tool_arguments
entry uses the semantic strategy now carries that entry on
EvaluatorKeyGap.blocked_by. The preflight hint, evaluate's ConfigError
detail and the doctor row name it ("blocking because tool_arguments `x`
uses the semantic strategy — export ... or change that strategy") rather
than suggesting `blocking: false`, which is already set. All wording
lives in evaluators/keys.py (gap_detail, forced_gate_hint,
all_skipped_error). evaluate's detail is now located in the suite's own
evaluators block when that block declares the family.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An advisory judge whose calls all errored gets a synthesized n=0
comparison, and the promotion advice told the user to collect more
examples for it. More examples measure nothing the same way; the line now
says the judge measured nothing (naming the prompts, or "this run").

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ConfigError.format_rich titled every panel "Invalid config", including
kind missing_key, which is raised over a valid config. That kind now reads
"Missing API key"; every other kind is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- CHANGELOG gains a `### Changed` entry (mirrored as an upgrading note in
  DOCS.md and llms-full.txt): a keyless semantic/llm_judge without
  `blocking: false` now exits 1 before any call where it used to run to an
  inconclusive verdict, so green CI can turn red; the bare
  text-embedding-3-small default is checked for the first time.
- Local and hosted recommendations match only with a migration_policy
  configured; without one analyze writes no decision.
- Spec: doctor's ok text is "judge and embedding models have API keys".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolves the conflicts with #44 (store library messages, doctor warns):
- doctor.py docstring and DOCS.md/llms-full.txt doctor text: exit 1 only
  on an invalid config or a blocking semantic/llm_judge model without a
  key; a missing store library is now a warning, naming the package.
- llms-full.txt doctor rows: #44's captures.store row plus this branch's
  evaluator keys row.
- CHANGELOG [Unreleased]: both branches' Added/Changed/Fixed entries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@babaliauskas
babaliauskas merged commit 3a96974 into main Oct 10, 2026
4 checks passed
babaliauskas added a commit that referenced this pull request Oct 10, 2026
Resolves the conflicts with #44 and #45:
- DOCS.md captures-store paragraph: main's package-name install wording
  (pip install boto3 / google-cloud-storage / azure-storage-blob
  azure-identity) plus this branch's capture sync --repair re-fetch clause.
- CHANGELOG [Unreleased]: this branch's Added/Fixed entries merged with
  main's Added/Changed/Fixed entries from #43, #44 and #45.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@babaliauskas babaliauskas mentioned this pull request Oct 11, 2026
3 tasks done
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant