Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,17 @@

## Unreleased

- Add audit sampling (`audit`; `bioevidence review audit-sample` and `audit-score`, #24). A seeded random
sample of auto-admitted records, sized from a target (`--target 0.01` audits 299 records: no error found
bounds the rate below 1% at 95%), with blind controls from the other routes and a manifest kept from the
auditors. Scoring reports each route's share of records, error rate with Wilson and exact one-sided
(Clopper–Pearson) bounds, and share of expert time from `minutes_spent`.
- Add the routing evidence (`evaluation/llm_benchmark/risk_signals.py`): whether each signal that sends a
record to a person predicts a wrong answer. BEV004, the only one on by default, is shown on every split
(82% vs 33% wrong on the single-cell held-out split); the opt-in signals BEV022, BEV025 and BEV026 are not.
- Add an audit dry run on the single-cell held-out results with the authors' labels as a stand-in auditor:
14 errors in 59 sampled auto-admitted records, below 34.6% at 95%; the census rate, 19.6%, lies inside.
`tools/reproduce.py` now runs 20 steps.
- Add the validation dossier (`docs/VALIDATION_DOSSIER.md`), structured after the seven steps of FDA's
draft AI credibility framework: question of interest, context of use, model risk, credibility plan,
execution, results with deviations, and adequacy. Identifier and citation integrity of admitted records
Expand Down
9 changes: 6 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,9 +109,10 @@ Across all 276 held-out annotations, stage by stage:
| PubMed rate limits, server errors | `Library.fetch` | Exponential backoff on 429 and 5xx, a response cache, a 100 MB download cap |
| Wrong paper, retracted paper, quote not in the paper | `LiteratureGrounder` (BEV016, BEV017, BEV019) | The reason goes back to the model; at most three submissions |
| ID of another term, obsolete term, gene alias | `OntologyGrounder`, `GeneGrounder` (BEV017, BEV023, BEV024) | The reason goes back, naming the correct identifier |
| The model's own evidence contradicts its answer | BEV004 | Sent to a person; not fed back, so the model is never asked to hide it |
| The model's own evidence contradicts its answer | BEV004 | Sent to a person; not fed back, so the model is never asked to hide it. Flagged answers were wrong 82% of the time vs 33% ([routing evidence](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/evaluation/llm_benchmark/results/risk-signals/summary.md)) |
| A revision drops evidence that was verified | `feedback.carry` | The verified evidence is carried into the revision |
| Still not admitted after the last round | `feedback.revise` | Sent to a person |
| Admitted but wrong | A seeded random audit (`bioevidence review audit-sample`, `audit-score`) | Each route's error rate with an exact upper bound; a [dry run](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/evaluation/llm_benchmark/results/celltype-audit/summary.md) shows the held-out route misses a 5% target |
| A long batch is interrupted | One file per episode | A rerun skips finished episodes (used when a 486-episode run hit a time limit) |

## How it works
Expand Down Expand Up @@ -251,10 +252,11 @@ uv run --frozen --with matplotlib==3.11.2 python tools/reproduce.py
```

One command regenerates every committed benchmark table, summary and figure, offline, in about three minutes,
and compares each with the repository byte for byte (text files with line endings normalised). It runs 18 steps,
and compares each with the repository byte for byte (text files with line endings normalised). It runs 20 steps,
also run in CI on every change:
- the three real-data cases;
- ten LLM-benchmark result sets;
- the routing evidence and the audit dry run;
- the task files and the error taxonomy;
- the figures.

Expand All @@ -281,7 +283,8 @@ context of use, model risk, credibility evidence, adequacy): the

Tracked in the [AI validation roadmap](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/AI_VALIDATION_ROADMAP.md) (#26):
- expert review of the benchmark cases (#23);
- calibrated triage and audit sampling of admitted records, to measure what still gets through (#24);
- an expert audit of admitted records with `bioevidence review audit-sample` and `audit-score`, which bound
the error rate of what still gets through. The tool and a dry run are done (#24); the audit needs experts.
- the [validation dossier](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/VALIDATION_DOSSIER.md),
structured after FDA's draft AI credibility framework, states what is established and what is not (#25).

Expand Down
4 changes: 2 additions & 2 deletions docs/AI_VALIDATION_ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -82,8 +82,8 @@ source → structured extraction → provenance → deterministic validation
| Now | [#31](https://github.com/NingyuSUN/bioai-evidence-validator/issues/31) | Semantic support check: does a verbatim quote support the claim's direction? | #20 |
| Next | [#21](https://github.com/NingyuSUN/bioai-evidence-validator/issues/21) | Real AI extraction experiment on openly licensed sources | #18, #20 |
| Next | [#22](https://github.com/NingyuSUN/bioai-evidence-validator/issues/22) | Model and agent evaluation against expert labels | #17, #23 |
| Then | [#24](https://github.com/NingyuSUN/bioai-evidence-validator/issues/24) | Calibrated triage and audit of auto-admitted records | #23 |
| Last | [#25](https://github.com/NingyuSUN/bioai-evidence-validator/issues/25) | Reproducible benchmark, funnel figure and validation dossier | all |
| Done | [#24](https://github.com/NingyuSUN/bioai-evidence-validator/issues/24) | Calibrated triage and audit of auto-admitted records (tooling, routing evidence and a dry run; the expert audit needs #23) | #23 |
| Done | [#25](https://github.com/NingyuSUN/bioai-evidence-validator/issues/25) | Reproducible benchmark, funnel figure and validation dossier | all |

## Definition of done for 0.8.0

Expand Down
19 changes: 19 additions & 0 deletions docs/ENGINEERING.md
Original file line number Diff line number Diff line change
Expand Up @@ -244,6 +244,25 @@ but cannot withdraw verified evidence against its answer. The literature benchma
(`evaluation/llm_benchmark/agent_loop.py`) showed why: before that rule, an agent whose record held a
misquote and a verified quote against its decision dropped the latter and was admitted.

### Routing evidence and audits

A record reaches a person because it cannot be verified (BEV006, BEV015, BEV018, BEV020), because the use requires
a person (BEV008–BEV013, BEV021), or because a risk signal flags it. The first two are admission policy. A risk
signal routes by default only when it is shown to predict a wrong answer.
`evaluation/llm_benchmark/risk_signals.py` measures this on the first answers of the committed runs:

| Signal | Default | Wrong when flagged vs otherwise | Verdict |
|---|---|---|---|
| BEV004 contradicting evidence line | on | held-out 82% vs 33%; external 70% vs 38%; literature pilot 91% vs 16% | shown |
| BEV022 semantic cue | opt-in | literature pilot 4% vs 9% | not shown |
| BEV025 reference conflict (ASCT+B) | opt-in | single-cell pilot 57% vs 40% | not shown |
| BEV026 definition contradicted | opt-in, experimental | external 47% vs 41% | not shown |

What routing lets through is measured by an audit (`audit`, `bioevidence review audit-sample` and `audit-score`):
a seeded random sample of auto-admitted records, optionally mixed with blind controls from the other routes, and
each route's error rate with an exact one-sided (Clopper–Pearson) upper bound. The procedure is in
[the gold-standard workflow](GOLD_STANDARD.md#auditing-what-is-admitted-automatically).

## Semantic checks

Grounding proves that a quote is real; it cannot tell whether the quote supports the claim.
Expand Down
26 changes: 26 additions & 0 deletions docs/GOLD_STANDARD.md
Original file line number Diff line number Diff line change
Expand Up @@ -113,6 +113,32 @@ rather than discarding difficult cases silently. Report reviewer agreement befor
adjudication (`bioevidence review agreement`). Any statistical intervals should respect concept-level dependence; do not
treat multiple aliases or mutations of one concept as independent samples.

## Auditing what is admitted automatically

A reference set measures a method once. In use, the question is how many auto-admitted records are wrong.
An audit answers it on a seeded random sample:

1. Export one prediction per record and use, in the format `bioevidence review score` reads, with the route the
method took: `admitted`, `review_required` or `rejected`.
2. `bioevidence review audit-sample` draws the sample. `--target 0.01` sizes it so that, if no audited record
is wrong, the auto-admitted error rate is below 1% at 95% confidence: 299 records (149 for 2%, 59 for 5%,
29 for 10%). `--controls` mixes in records from the other routes, in shuffled order, so auditors cannot tell
a record's route. The manifest records the seed, each route's size and the sampled hashes; keep it from the
auditors.
3. Auditors label the sheet as in this protocol, blind to the route, the model and each other, and record
`minutes_spent`.
4. `bioevidence review audit-score` reports each route's error rate with an exact one-sided upper bound. An
auto-admitted record is an error when the auditor would not admit it for the use or its claim is incorrect;
a record on another route, when the auditor would have admitted it as it was.

Report the bound with the rate: "2 errors in 300 audited, below 2.1% at 95% confidence". Audit again when the
model, prompts, profile or references change. One auditor per record is the default for audits
(`--min-reviewers 1`); disagreements between several still need adjudication.

Records reach a person for one of three reasons: they cannot be verified, the use requires a person, or a risk
signal flags them. A risk signal belongs in routing only when it is shown to predict errors; the evidence for
each is in [`results/risk-signals`](../evaluation/llm_benchmark/results/risk-signals/summary.md).

An improvement against this reference can support a claim about the defined curation
task. It does not establish breed-genotype membership, universal biological truth, or
clinical validity.
12 changes: 8 additions & 4 deletions docs/VALIDATION_DOSSIER.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,8 @@ State: release 0.8.0, October 2026.
|---|---|---|
| Admitted records carry no invalid identifier or citation | 0 of 721 admitted AI answers across four experiments (upper bound 0.5%); models alone produced such errors in 10–76% of answers | **Established** |
| The feedback loop corrects fixable errors without hiding evidence | Single-cell held-out: 57 of 63 answers with identifier or marker errors corrected and admitted; verified evidence cannot be withdrawn | **Established** |
| Admitted records are semantically correct | 19.6% [15.2, 24.9] of admitted single-cell annotations disagree with the authors' term. Part of this is label noise. No expert audit yet | **Not established** |
| Admitted records are semantically correct | 19.6% [15.2, 24.9] of admitted single-cell annotations disagree with the authors' term. Part of this is label noise. The audit tool is ready and was checked in a dry run; no expert audit yet | **Not established** |
| Default routing rests on signals that predict errors | BEV004 answers were wrong 82% vs 33% (held-out) and 70% vs 38% (external); the opt-in signals (semantic cues, ASCT+B, definitions) showed no such effect and stay off | **Established** for BEV004 |
| Routing load is workable | 7.6% [5.0, 11.3] of single-cell annotations sent to a person (held-out) | **Plausible**, one run |
| A new semantic check (Cell Ontology definitions) helps | No better than chance on external data, and harmful when fed back | **Refuted**; kept experimental |

Expand Down Expand Up @@ -75,7 +76,9 @@ consequence**, the impact of a wrong decision.
- protocols fixed before results are seen;
- negative controls;
- full reproducibility;
- a measure of what still gets through. That last one needs an expert audit, which has not been done.
- a measure of what still gets through. That last one needs an expert audit, which has not been done. The
procedure and its statistics are ready (`bioevidence review audit-sample`, `audit-score`), and a dry run
with the authors' labels as a stand-in recovered the census rate (19.6%) inside its interval.

## 4. Credibility assessment plan

Expand Down Expand Up @@ -173,7 +176,7 @@ Fed back in the loop, its findings turned 15 correct answers into wrong ones. Co
- **One run was resumed.** The external run hit a time limit at 321 of 486 episodes and was resumed. Episodes are
independent, and none was repeated.
- **Two planned evaluations have not been done.** The literature held-out set was built but not run, and no expert
review or audit of admitted records has been done (#23, #24).
review or audit of admitted records has been done (#23). The audit tooling is in place (#24).

## 7. Adequacy for the context of use

Expand All @@ -199,7 +202,8 @@ Proposed acceptance criteria for the next evaluation, to be fixed before it is r
protocol, and compare per model. The integrity result should not move; the semantic result can.
- **Reference change.** A new Cell Ontology, HGNC or ClinVar release changes the pinned snapshots. Re-run
`tools/reproduce.py` and the affected splits, and keep the earlier snapshots for comparison.
- **Monitoring in use.** Audit a random sample of admitted records on a schedule (#24), and report the audited
- **Monitoring in use.** Audit a random sample of admitted records on a schedule with
`bioevidence review audit-sample` and `audit-score` (#24), and report the audited
error rate with its interval next to the figures above.

## Sources
Expand Down
3 changes: 2 additions & 1 deletion docs/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,9 +105,10 @@ Across all 276 held-out annotations, stage by stage:
| PubMed rate limits, server errors | `Library.fetch` | Exponential backoff on 429 and 5xx, a response cache, a 100 MB download cap |
| Wrong paper, retracted paper, quote not in the paper | `LiteratureGrounder` (BEV016, BEV017, BEV019) | The reason goes back to the model; at most three submissions |
| ID of another term, obsolete term, gene alias | `OntologyGrounder`, `GeneGrounder` (BEV017, BEV023, BEV024) | The reason goes back, naming the correct identifier |
| The model's own evidence contradicts its answer | BEV004 | Sent to a person; not fed back, so the model is never asked to hide it |
| The model's own evidence contradicts its answer | BEV004 | Sent to a person; not fed back, so the model is never asked to hide it. Flagged answers were wrong 82% of the time vs 33% ([routing evidence](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/evaluation/llm_benchmark/results/risk-signals/summary.md)) |
| A revision drops evidence that was verified | `feedback.carry` | The verified evidence is carried into the revision |
| Still not admitted after the last round | `feedback.revise` | Sent to a person |
| Admitted but wrong | A seeded random audit (`bioevidence review audit-sample`, `audit-score`) | Each route's error rate with an exact upper bound; a [dry run](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/evaluation/llm_benchmark/results/celltype-audit/summary.md) shows the held-out route misses a 5% target |
| A long batch is interrupted | One file per episode | A rerun skips finished episodes |

How the checks work, rule by rule: [engineering contract](ENGINEERING.md).
Expand Down
2 changes: 2 additions & 0 deletions evaluation/gold_standard/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,8 @@ zero reviewed cases. An editor changing its status does not complete human revie
| `bioevidence review adjudication-sheet annotations.csv --output adjudications.csv` | Blank rows for every disagreement; never overwrites an existing file |
| `bioevidence review score --annotations … --adjudications … --predictions …` | Resolves final labels and scores predictions on the test split; rejects missing or hash-mismatched rows |
| `bioevidence review freeze --manifest … --annotations … --adjudications …` | Refuses unresolved cases; records counts, reference type and file hashes |
| `bioevidence review audit-sample --predictions … --method … --use … --target 0.01 --seed … --profile-id … --output-dir …` | A seeded random sample of auto-admitted records as a blank annotation sheet, sized so that no error found bounds their error rate below the target; `--controls` mixes in records from the other routes. The manifest of routes stays with you |
| `bioevidence review audit-score --manifest … --annotations …` | Per route: share of records, audited error rate with Wilson and exact one-sided bounds, and share of expert time from `minutes_spent` |

A worked, domain-specific kit is in [`../clinvar_review/`](../clinvar_review/README.md).
The VBO replay command still evaluates only its source-derived reference set and controlled faults.
34 changes: 34 additions & 0 deletions evaluation/llm_benchmark/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -448,6 +448,40 @@ Which need experts: evidence that is mixed, depends on the endpoint, or where th
every model disagree; these are exactly the cases the expert review (#23) is for. Five errors on two
tasks are too few for rates; the held-out test set would measure them with intervals.

## Routing evidence and the audit (#24)

**Which signals should send a record to a person?** Only those shown to predict a wrong answer.
[`results/risk-signals`](results/risk-signals/summary.md) compares, on first answers, how often an answer was wrong
when a signal fired and when it did not:

| Signal | Default | Split | Wrong when flagged | Wrong otherwise | Risk ratio (95% CI) |
|---|---|---|---:|---:|---:|
| BEV004 contradicting evidence | on | single-cell held-out | 14/17 (82%) | 86/259 (33%) | 2.48 (1.87–3.28) |
| | | single-cell external | 38/54 (70%) | 163/425 (38%) | 1.83 (1.49–2.27) |
| | | literature agent pilot | 10/11 (91%) | 11/69 (16%) | 5.7 (3.21–10.12) |
| BEV022 semantic cue | opt-in | literature pilot | 1/24 (4%) | 4/44 (9%) | 0.46 (0.05–3.87) |
| BEV025 ASCT+B conflict | opt-in | single-cell pilot | 13/23 (57%) | 46/115 (40%) | 1.41 (0.93–2.16) |
| BEV026 definition contradicted | opt-in | single-cell external | 41/88 (47%) | 160/391 (41%) | 1.14 (0.88–1.47) |

The only risk signal on by default, BEV004, is shown on every split. The opt-in ones are not, and stay off by
default. The criterion was set after the runs: the risk ratio's interval lies above 1 on a held-out or external
split.

**What gets through?** An audit of a seeded random sample of auto-admitted records answers that with an upper
bound. No expert has audited these records yet, so
[`results/celltype-audit`](results/celltype-audit/summary.md) is a dry run on the held-out split. It runs the
whole procedure with a stand-in auditor, the dataset authors' labels:
1. export the loop's routes;
2. `bioevidence review audit-sample --target 0.05 --controls 10` draws 59 auto-admitted records and 10 others;
3. a [packet](results/celltype-audit/audit_packet.md) shows each record without its route, model or the
authors' label;
4. `bioevidence review audit-score` scores it.

The stand-in found 14 errors in 59, so the auto-admitted error rate is below 34.6% at 95% confidence (23.7%,
Wilson 14.7–36.0%). Because the authors' labels cover every record, the census can be checked: 50 of 255
(19.6%), inside the interval. The route does not meet a 5% target. A deployment would learn from such an audit
that these annotations need review, or a better model, before research summaries rely on them.

## Run

```bash
Expand Down
Loading
Loading