Skip to content

docs: cover trace invariants on the summary surfaces; fix advisory-tag scope and hosted-gate wording - #37

Merged
babaliauskas merged 1 commit into
mainfrom
docs/trace-invariants-coverage
Oct 8, 2026
Merged

babaliauskas merged 1 commit into
mainfrom
docs/trace-invariants-coverage

Conversation

@babaliauskas

Copy link
Copy Markdown
Collaborator

Cross-repo documentation audit after #35/#36 (siblings: evalshift-server #30, evalshift-client PR to follow). The reference sections — DOCS.md Trace invariants, docs/configuration.md#evaluatorstrace_invariants, the llms-full.txt evaluator block — were already right and consistent. The surfaces that summarise them were not updated. Prose only, plus one docstring and the re-vendored schema.

Missing → added

  • Profile tables (DOCS.md, llms-full.txt): a trace-rule violations ≤ column, 0 in every profile (what init writes).
  • Report layout (DOCS.md, llms-full.txt, README.md): the "Trace rules broken" panel; the executive summary and the strip's mean score Δ are drawn from blocking evaluators (advisory ones only on a prompt nothing else scored).
  • README agent section, docs/agents.md ("four new evaluators", the top-level placement carve-out, a paragraph in "Reading a tool run's report"), and the "tool evaluators are not written at top level" comments in DOCS.md / llms-full.txt now carry the trace_invariants exception.
  • Hosted wording (docs/hosted.md, DOCS.md, llms-full.txt, docs/methodology.md, CHANGELOG): the server accepts and stores max_invariant_violations since 2026-10-08 and never re-evaluates it; a run that carries its own policy is answered with the CLI's decision, only policy-less runs get the six-budget re-check. The CHANGELOG's "this CLI release waits on that server deploy" is gone.
  • FAQs: a breached max_invariant_violations is never inconclusive (docs/faq.md, DOCS.md, llms-full.txt); a CLI upgrade that added config fields moves config_hash (resume FAQs in the same three files). CHANGELOG mentions the analyze recommendation for source-shared breaks.

Corrections

  • The "Overall, by evaluator" advisory tag (from fix(report): advisory evaluators out of the summary; trace-rule causes counted as the budget does #36) applies to every blocking: false evaluator, not only trace-rule entries — CHANGELOG, DOCS.md, llms-full.txt, docs/configuration.md said the latter.
  • call_count has min_calls/max_calls; there is no exact key (CHANGELOG, docs/evaluators.md).
  • docs/evaluators.md: "they all run on every pair" qualified for trace_invariants; the budget gates blocking entries (also docs/traces.md, which now also notes the imported-traces evaluate error and applies_to).
  • applies_to: DOCS.md said "where supported"; the four other families' rows in docs/configuration.md said only "Glob list". All now say accepted, enforced only by trace_invariants.
  • llms-full.txt round numbering uses the replay's 1-based convention everywhere (violation round_index noted as 0-based).
  • hosted/push.py docstring: six of ten budgets, other four.

Schema

src/evalshift_cli/hosted/bundle_manifest.schema.json re-vendored from evalshift-server #30 (one description string on BudgetResult.denominator). Merge server #30 first: until then the local-only test_bundle_shape.py::test_the_vendored_schema_matches_the_server_export differs from server main by exactly that string (CI skips that test; it needs the sibling checkout).

Checks

ruff check / ruff format --check / mypy --strict: green. pytest: 2581 passed, 1 failed — the vendored-schema test above, for the stated reason only (verified: 0 differences against the #30 schema). tests/unit/test_docs_currency.py green.

🤖 Generated with Claude Code

…g scope and hosted-gate wording

Cross-repo docs audit after #35/#36. The reference sections (DOCS.md Trace
invariants, docs/configuration.md, llms-full.txt) were right; the summary
surfaces around them were not updated:

- Profile tables (DOCS.md, llms-full.txt) gain the trace-rule-violations
  column (0 in every profile). Report-layout lists (DOCS.md, llms-full.txt,
  README) gain the "Trace rules broken" panel and say the executive summary
  and the strip's mean score delta come from blocking evaluators.
- README agent section and docs/agents.md ("four new evaluators", the
  top-level placement carve-out, "Reading a tool run's report") mention
  trace_invariants; the "tool evaluators are not written at top level"
  comments in DOCS.md and llms-full.txt carry the same carve-out.
- Hosted wording (docs/hosted.md, DOCS.md, llms-full.txt, methodology.md,
  CHANGELOG): the server accepts and stores max_invariant_violations since
  2026-10-08 and never re-evaluates it; a run that carries its own policy is
  answered with the CLI's decision, only policy-less runs get the six-budget
  re-check. The stale "waits on that server deploy" sentence is gone.
- Inconclusive FAQs (docs/faq.md, DOCS.md, llms-full.txt): a breached
  max_invariant_violations is never inconclusive. Resume FAQs: a CLI upgrade
  that added config fields moves config_hash.
- Corrections: the "Overall, by evaluator" advisory tag applies to every
  blocking: false evaluator, not only trace-rule entries (CHANGELOG, DOCS.md,
  llms-full.txt, configuration.md); call_count has min_calls/max_calls, not
  an "exact" key (CHANGELOG, evaluators.md); "they all run on every pair" and
  the blocking qualifier in evaluators.md/traces.md; DOCS.md applies_to and
  the four other families' applies_to rows in configuration.md now say it is
  accepted but enforced only by trace_invariants; round numbering in
  llms-full.txt uses the replay's 1-based convention; CHANGELOG mentions the
  analyze recommendation for source-shared breaks.
- push.py docstring: six of ten budgets, other four.
- Re-vendored schemas/bundle_manifest.schema.json from evalshift-server PR #30
  (one description string).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@babaliauskas
babaliauskas merged commit 6dac5a0 into main Oct 8, 2026
4 checks passed
@babaliauskas
babaliauskas deleted the docs/trace-invariants-coverage branch October 9, 2026 19:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant