Skip to content

Add audit sampling and routing evidence (#24) - #43

Merged
NingyuSUN merged 1 commit into
mainfrom
audit-sampling
Oct 6, 2026
Merged

NingyuSUN merged 1 commit into
mainfrom
audit-sampling

Conversation

@NingyuSUN

Copy link
Copy Markdown
Owner

Implements #24: audit sampling of auto-admitted records, and evidence for which routing signals predict errors.

What it adds

bioevidence review audit-sample: a seeded random sample of one method's auto-admitted records for one use, written as a blank annotation sheet.

  • --target 0.01 sizes the sample so that, if no audited record is wrong, the auto-admitted error rate is below 1% at 95% confidence. That is 299 records; a 2% target needs 149, 5% needs 59 and 10% needs 29.
  • --controls mixes in records from the other routes in shuffled order, so auditors cannot tell which route a record took.
  • The manifest records the seed, the size of each route and the sampled hashes. It stays with the maintainer.

bioevidence review audit-score: reports, for each route:

  • its share of records;
  • the audited error rate, with a Wilson interval and an exact one-sided (Clopper–Pearson) upper bound;
  • the most errors the route can hold;
  • its share of expert time, estimated from minutes_spent.

It refuses labels for records that were not sampled, or that were made on a different record or profile.

Routing evidence (evaluation/llm_benchmark/risk_signals.py): for each signal that sends a record to a person, how often first answers were wrong when it fired and when it did not.

Signal Default Wrong when flagged vs otherwise Verdict
BEV004 contradicting evidence on held-out 82% vs 33%; external 70% vs 38%; literature pilot 91% vs 16% shown
BEV022 semantic cue opt-in 4% vs 9% not shown
BEV025 ASCT+B conflict opt-in 57% vs 40% not shown
BEV026 definition contradicted opt-in 47% vs 41% not shown

BEV004 is the only risk signal on by default, and it holds on every split. I set the criterion after the runs: the risk ratio's 95% interval must lie above 1 on a held-out or external split.

Audit dry run (evaluation/llm_benchmark/audit_celltype.py), on the single-cell held-out results. No expert has audited these records. The dry run uses the dataset authors' labels as a stand-in auditor:

  • The sample was 59 auto-admitted records and 10 controls. The packet shows each record without its route, its model or the authors' label.
  • The stand-in found 14 errors in 59: 23.7%, Wilson 14.7–36.0%, upper bound 34.6%.
  • The true rate over all 255 auto-admitted records is 19.6% (50/255). It falls inside the interval.
  • So the route misses a 5% target. An audit is how a deployment would learn that before relying on it.

Done-when in #24

  • Routing uses only risk signals shown to predict errors, each with its evidence.
  • bioevidence review audit-sample draws a seeded random sample of auto-admitted records.
  • An upper bound on the auto-admitted error rate is reported (0 errors in 299 audited records gives a bound below 1% at 95% confidence).
  • Per-route statistics: share of records, audited error rate, and share of expert time. Expert time stays empty in the dry run because the stand-in records none.

The expert audit itself still needs experts (#23).

Checks

  • ruff, mypy and pytest pass, with 98.5% coverage. tests/test_audit.py checks the bounds against the closed form and against the binomial CDF.
  • mkdocs build --strict passes.
  • tools/reproduce.py now has 20 steps, and all of them reproduce identically, including the 2 new ones.

- audit: seeded random sample of auto-admitted records, sized from a target
  error bound, with blind controls from the other routes; per-route error
  rates with Wilson and exact one-sided (Clopper-Pearson) bounds, and share
  of expert time from minutes_spent.
- bioevidence review audit-sample / audit-score.
- risk_signals.py: whether each signal that routes a record to a person
  predicts a wrong answer. BEV004 is shown on every split; the opt-in
  signals BEV022, BEV025 and BEV026 are not.
- audit_celltype.py: dry run on the single-cell held-out results with the
  authors' labels as a stand-in auditor; the census rate lies inside the
  audit's interval.
- reproduce.py gains both steps (20 in all); docs, dossier and changelog.
@NingyuSUN
NingyuSUN merged commit ccd8857 into main Oct 6, 2026
11 checks passed
@NingyuSUN
NingyuSUN deleted the audit-sampling branch October 6, 2026 06:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant