Repository navigation
feat: judge and semantic on by default; name the missing evaluator key - #45
Merged
Merged
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… are OpenAI Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ess gate Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… evaluators Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…te being run Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…edder Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… forcing tool_arguments entry
The key preflight now refuses a config whose every evaluator would be
skipped for a missing key before any arm call, instead of letting the
evaluate stage fail after the run was paid for.
A keyless semantic that is a gate only because a blocking tool_arguments
entry uses the semantic strategy now carries that entry on
EvaluatorKeyGap.blocked_by. The preflight hint, evaluate's ConfigError
detail and the doctor row name it ("blocking because tool_arguments `x`
uses the semantic strategy — export ... or change that strategy") rather
than suggesting `blocking: false`, which is already set. All wording
lives in evaluators/keys.py (gap_detail, forced_gate_hint,
all_skipped_error). evaluate's detail is now located in the suite's own
evaluators block when that block declares the family.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
An advisory judge whose calls all errored gets a synthesized n=0 comparison, and the promotion advice told the user to collect more examples for it. More examples measure nothing the same way; the line now says the judge measured nothing (naming the prompts, or "this run"). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ConfigError.format_rich titled every panel "Invalid config", including kind missing_key, which is raised over a valid config. That kind now reads "Missing API key"; every other kind is unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- CHANGELOG gains a `### Changed` entry (mirrored as an upgrading note in DOCS.md and llms-full.txt): a keyless semantic/llm_judge without `blocking: false` now exits 1 before any call where it used to run to an inconclusive verdict, so green CI can turn red; the bare text-embedding-3-small default is checked for the first time. - Local and hosted recommendations match only with a migration_policy configured; without one analyze writes no decision. - Spec: doctor's ok text is "judge and embedding models have API keys". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolves the conflicts with #44 (store library messages, doctor warns): - doctor.py docstring and DOCS.md/llms-full.txt doctor text: exit 1 only on an invalid config or a blocking semantic/llm_judge model without a key; a missing store library is now a warning, naming the package. - llms-full.txt doctor rows: #44's captures.store row plus this branch's evaluator keys row. - CHANGELOG [Unreleased]: both branches' Added/Changed/Fixed entries. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
babaliauskas
added a commit
that referenced
this pull request
Oct 10, 2026
Resolves the conflicts with #44 and #45: - DOCS.md captures-store paragraph: main's package-name install wording (pip install boto3 / google-cloud-storage / azure-storage-blob azure-identity) plus this branch's capture sync --repair re-fetch clause. - CHANGELOG [Unreleased]: this branch's Added/Fixed entries merged with main's Added/Changed/Fixed entries from #43, #44 and #45. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
llm_judgeandsemanticare the evaluators that say whether the new model's prose is as good as the old one's. A fresh project got the least out of them, for three reasons:semanticshipped commented out for Anthropic and DeepSeek, which have no embeddings endpoint. The YAML comment was the only explanation.comparechecked keys only for the two arms being compared. A keyless judge raised on every pair; being advisory, nothing gated on it and nothing was printed, so the user paid for retries and never found out.inconclusivewith the same "Set blocking: true on at least one trusted evaluator…" line, with no evaluator named and no "is my suite big enough yet".This PR fixes all three:
initturns everything on.semanticis active for every provider. Anthropic and DeepSeek borrow an embedder: OpenAI's ifOPENAI_API_KEYis set, else Gemini's ifGEMINI_API_KEY/GOOGLE_API_KEYis set, else OpenAI's anyway. Next steps name a missing borrowed key, andinit --ciwires it as a second secret.compareandruncheck every judge and embedding model up front, andevaluateapplies the same rule:⚠ <label> skipped: no API key for <model> — export <VAR> to enable it.The run continues without it, making no calls and no retries.semanticforced blocking by a blockingtool_argumentsentry (one using thesemanticstrategy): it counts as a gate, and the message names that entry.state.jsongainsskipped_evaluators. Each skip becomes a line in the verdict'srecommendations, which reach the terminal, the HTML report,migration_decision.jsonand the hosted bundle. Local and hosted output match when amigration_policyis configured.MIN_N_RELIABLE), and the smallest prompt decides:doctorgains anevaluator keysrow:failfor blocking,warnfor advisory.compare's doctor stage ignores this row because its own check is scoped to the suite being run, so a keyless judge in suite B never failscompare --suite-name A.missing_api_keyshelper replaces three copies; an empty-string env var counts as unset.text-embedding-*ids resolve to OpenAI, so the config default's key is checked.Also in this branch: the design spec and implementation plan (
docs/superpowers/specs/2026-10-09-evaluators-on-by-default-design.md,docs/superpowers/plans/2026-10-09-evaluators-on-by-default.md), plus DOCS.md, llms-full.txt, docs/configuration.md, docs/faq.md, docs/getting-started.md, docs/github-action.md and the CHANGELOG underUnreleased(Added / Changed / Fixed).Upgrade impact
semantic/llm_judgedefault toblocking: truein the library. A hand-written config that doesn't sayblocking: falseand whose judge or embedder has no key used to finish asinconclusive(exit 0, passing--policy-gate). Nowcompare,run,evaluateanddoctorexit 1 before any call. The CHANGELOG### Changedentry and the DOCS.md upgrade paragraph say so. I'd ship this as a minor bump, not a patch.Known, deliberately left
tool_argumentsentries forcesemanticblocking, only the first is named. A rerun names the next.semanticis blocking by its own flag and forced, the hint suggests onlyblocking: false. Exporting the key fixes it either way.doctordoesn't detect the "every evaluator would be skipped" case; it warns and exits 0, whilecompare/runexit 1.doctor.py,compare.py's doctor stage and theCHANGELOGUnreleasedsection. Whichever merges second will need a small rebase.Test plan
pytest: 2833 passed, coverage 94.84% (make ci)ruff check,ruff format --checkcleanmypy --strict src/evalshift_clicleanpre-commit run --all-filesclean## [Unreleased]Each of the 8 tasks was test-driven and reviewed separately. The whole branch then had a final review that probed an empty CI secret, an upgraded hand-written config, a re-run of
evaluateafter exporting the key (stale skips are cleared), and multi-suitecompare. Its findings were fixed and re-reviewed.🤖 Generated with Claude Code