Repository navigation
docs: cover trace invariants on the summary surfaces; fix advisory-tag scope and hosted-gate wording - #37
Merged
Conversation
…g scope and hosted-gate wording Cross-repo docs audit after #35/#36. The reference sections (DOCS.md Trace invariants, docs/configuration.md, llms-full.txt) were right; the summary surfaces around them were not updated: - Profile tables (DOCS.md, llms-full.txt) gain the trace-rule-violations column (0 in every profile). Report-layout lists (DOCS.md, llms-full.txt, README) gain the "Trace rules broken" panel and say the executive summary and the strip's mean score delta come from blocking evaluators. - README agent section and docs/agents.md ("four new evaluators", the top-level placement carve-out, "Reading a tool run's report") mention trace_invariants; the "tool evaluators are not written at top level" comments in DOCS.md and llms-full.txt carry the same carve-out. - Hosted wording (docs/hosted.md, DOCS.md, llms-full.txt, methodology.md, CHANGELOG): the server accepts and stores max_invariant_violations since 2026-10-08 and never re-evaluates it; a run that carries its own policy is answered with the CLI's decision, only policy-less runs get the six-budget re-check. The stale "waits on that server deploy" sentence is gone. - Inconclusive FAQs (docs/faq.md, DOCS.md, llms-full.txt): a breached max_invariant_violations is never inconclusive. Resume FAQs: a CLI upgrade that added config fields moves config_hash. - Corrections: the "Overall, by evaluator" advisory tag applies to every blocking: false evaluator, not only trace-rule entries (CHANGELOG, DOCS.md, llms-full.txt, configuration.md); call_count has min_calls/max_calls, not an "exact" key (CHANGELOG, evaluators.md); "they all run on every pair" and the blocking qualifier in evaluators.md/traces.md; DOCS.md applies_to and the four other families' applies_to rows in configuration.md now say it is accepted but enforced only by trace_invariants; round numbering in llms-full.txt uses the replay's 1-based convention; CHANGELOG mentions the analyze recommendation for source-shared breaks. - push.py docstring: six of ten budgets, other four. - Re-vendored schemas/bundle_manifest.schema.json from evalshift-server PR #30 (one description string). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Cross-repo documentation audit after #35/#36 (siblings: evalshift-server #30, evalshift-client PR to follow). The reference sections — DOCS.md Trace invariants,
docs/configuration.md#evaluatorstrace_invariants, thellms-full.txtevaluator block — were already right and consistent. The surfaces that summarise them were not updated. Prose only, plus one docstring and the re-vendored schema.Missing → added
DOCS.md,llms-full.txt): a trace-rule violations ≤ column,0in every profile (whatinitwrites).DOCS.md,llms-full.txt,README.md): the "Trace rules broken" panel; the executive summary and the strip's mean score Δ are drawn from blocking evaluators (advisory ones only on a prompt nothing else scored).docs/agents.md("four new evaluators", the top-level placement carve-out, a paragraph in "Reading a tool run's report"), and the "tool evaluators are not written at top level" comments inDOCS.md/llms-full.txtnow carry thetrace_invariantsexception.docs/hosted.md,DOCS.md,llms-full.txt,docs/methodology.md, CHANGELOG): the server accepts and storesmax_invariant_violationssince 2026-10-08 and never re-evaluates it; a run that carries its own policy is answered with the CLI's decision, only policy-less runs get the six-budget re-check. The CHANGELOG's "this CLI release waits on that server deploy" is gone.max_invariant_violationsis neverinconclusive(docs/faq.md,DOCS.md,llms-full.txt); a CLI upgrade that added config fields movesconfig_hash(resume FAQs in the same three files). CHANGELOG mentions theanalyzerecommendation for source-shared breaks.Corrections
blocking: falseevaluator, not only trace-rule entries — CHANGELOG,DOCS.md,llms-full.txt,docs/configuration.mdsaid the latter.call_counthasmin_calls/max_calls; there is noexactkey (CHANGELOG,docs/evaluators.md).docs/evaluators.md: "they all run on every pair" qualified fortrace_invariants; the budget gates blocking entries (alsodocs/traces.md, which now also notes the imported-tracesevaluateerror andapplies_to).applies_to:DOCS.mdsaid "where supported"; the four other families' rows indocs/configuration.mdsaid only "Glob list". All now say accepted, enforced only bytrace_invariants.llms-full.txtround numbering uses the replay's 1-based convention everywhere (violationround_indexnoted as 0-based).hosted/push.pydocstring: six of ten budgets, other four.Schema
src/evalshift_cli/hosted/bundle_manifest.schema.jsonre-vendored from evalshift-server #30 (one description string onBudgetResult.denominator). Merge server #30 first: until then the local-onlytest_bundle_shape.py::test_the_vendored_schema_matches_the_server_exportdiffers from servermainby exactly that string (CI skips that test; it needs the sibling checkout).Checks
ruff check/ruff format --check/mypy --strict: green.pytest: 2581 passed, 1 failed — the vendored-schema test above, for the stated reason only (verified: 0 differences against the #30 schema).tests/unit/test_docs_currency.pygreen.🤖 Generated with Claude Code