diff --git a/CHANGELOG.md b/CHANGELOG.md index ac10990..810518a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,10 @@ ## Unreleased +- Add the validation dossier (`docs/VALIDATION_DOSSIER.md`), structured after the seven steps of FDA's + draft AI credibility framework: question of interest, context of use, model risk, credibility plan, + execution, results with deviations, and adequacy. Identifier and citation integrity of admitted records + is established (0 of 721); semantic correctness is not, pending an expert audit. - Add `tools/reproduce.py`: one command regenerates every committed benchmark table, summary and figure offline (18 steps, about three minutes) and compares each with the repository; CI runs it on every change. It caught the ClinVar and VBO figures still captioned v0.7.0 after the release, now fixed. diff --git a/README.md b/README.md index 9258e3c..275700a 100644 --- a/README.md +++ b/README.md @@ -273,13 +273,17 @@ committed files and `--list` shows the steps. was run once. - **Not for clinical use.** +What is and is not established, in the structure of FDA's draft AI credibility framework (question of interest, +context of use, model risk, credibility evidence, adequacy): the +[validation dossier](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/VALIDATION_DOSSIER.md). + ## Roadmap Tracked in the [AI validation roadmap](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/AI_VALIDATION_ROADMAP.md) (#26): - expert review of the benchmark cases (#23); - calibrated triage and audit sampling of admitted records, to measure what still gets through (#24); -- a validation dossier structured like FDA's draft AI credibility framework (#25). One-command reproduction - and the funnel figure are done. +- the [validation dossier](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/VALIDATION_DOSSIER.md), + structured after FDA's draft AI credibility framework, states what is established and what is not (#25). ## Contributing and citing diff --git a/docs/VALIDATION_DOSSIER.md b/docs/VALIDATION_DOSSIER.md new file mode 100644 index 0000000..cd62f71 --- /dev/null +++ b/docs/VALIDATION_DOSSIER.md @@ -0,0 +1,211 @@ +# Validation dossier: AI-assisted curation behind bioevidence + +This dossier follows the seven steps of the risk-based credibility assessment framework in FDA's draft +guidance *Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug +and Biological Products* (January 2025). It borrows the structure only. It is **not a regulatory submission**, +it has not been reviewed by FDA, and the system is **not for clinical use**. + +Every number comes from the committed results and can be regenerated offline with +`uv run --frozen --with matplotlib==3.11.2 python tools/reproduce.py`. Intervals are Wilson 95%. +State: release 0.8.0, October 2026. + +## Summary + +| Claim | Evidence | Status for the context of use | +|---|---|---| +| Admitted records carry no invalid identifier or citation | 0 of 721 admitted AI answers across four experiments (upper bound 0.5%); models alone produced such errors in 10–76% of answers | **Established** | +| The feedback loop corrects fixable errors without hiding evidence | Single-cell held-out: 57 of 63 answers with identifier or marker errors corrected and admitted; verified evidence cannot be withdrawn | **Established** | +| Admitted records are semantically correct | 19.6% [15.2, 24.9] of admitted single-cell annotations disagree with the authors' term. Part of this is label noise. No expert audit yet | **Not established** | +| Routing load is workable | 7.6% [5.0, 11.3] of single-cell annotations sent to a person (held-out) | **Plausible**, one run | +| A new semantic check (Cell Ontology definitions) helps | No better than chance on external data, and harmful when fed back | **Refuted**; kept experimental | + +## 1. Question of interest + +Can records produced by general-purpose LLMs and agents be admitted to a research knowledge base without an +expert reviewing each one, and which records must an expert see? Two kinds of records are in scope: +- cell-type annotations of single-cell clusters; +- literature evidence for cancer-variant claims. + +## 2. Context of use + +- **The system assessed** is a model and bioevidence together, because the question concerns the records that + are admitted. + - **Proposer:** one of six general-purpose models, or an agent that searches PubMed, proposes a record. A + record is a claim with the evidence it rests on: marker genes of a cluster, or quotes from papers. + - **Validator:** bioevidence checks the record against pinned sources and reference releases, under a + profile for the intended use. + - **Feedback:** fixable findings go back to the proposer, for up to three answers. + - **Admission:** an admitted record enters the knowledge base. A record with conflicting evidence, or still + failing after the last round, goes to a curator. +- **Intended use:** research summaries and a research knowledge base. + + Out of scope: + - clinical interpretation; + - patient care; + - regulatory decisions; + - training data, unless a human accepts each record. The profiles require that. +- **Other evidence in the decision:** curators decide every routed record. Admitted records are not reviewed by + anyone today; that is the residual risk this dossier measures. +- **Models:** + + | Key | Model | Tier | Runs through | + |---|---|---|---| + | claude-opus | `claude-opus-5-5` | frontier | Claude Code 2.1.283 | + | gpt-astra | `gpt-6-astra` | frontier | codex-cli 0.153.4 | + | gemini-pro | `gemini-3.1-pro-high` | frontier | Antigravity 1.2.13 | + | claude-haiku | `claude-haiku-4-5-20251001` | fast | Claude Code 2.1.283 | + | gpt-luna | `gpt-5.6-luna` | fast | codex-cli 0.153.4 | + | gemini-flash | `gemini-3.8-flash-medium` | fast | Antigravity 1.2.13 | + + Every reply must match a JSON Schema, and the harness, not the model's CLI, executes the tools. Bioevidence + must not depend on which model proposes, so all six are evaluated with one pipeline. + +## 3. Model risk + +The guidance combines **model influence**, how much the output drives the decision, with **decision +consequence**, the impact of a wrong decision. + +- **Influence is high for admitted records:** nothing else checks them. It is low for routed records, where a + curator decides. +- **Consequence is moderate in this context.** A wrong cell type or citation in a research knowledge base + spreads to analyses, summaries and possibly to training sets, but it does not reach patients. Outside the + context (clinical use) it would be high. +- **Risk is therefore medium.** It calls for: + - evaluation on data not used to design the system; + - protocols fixed before results are seen; + - negative controls; + - full reproducibility; + - a measure of what still gets through. That last one needs an expert audit, which has not been done. + +## 4. Credibility assessment plan + +| Element | What was planned and done | +|---|---| +| Sources | Pinned by SHA-256, with licences and attribution: | +| Reference standards | No independent expert labels yet (#23) | +| Splits | | +| Protocol freezes | Each protocol was committed before the split it was run on: Protocol 1 was fixed before the pilot and committed with its results | +| Metrics | Defined in the scoring code before the runs: Each is reported per model and pooled | +| Negative controls | | +| Acceptance criteria | **Not set numerically in advance.** This is a gap; section 7 proposes criteria for the next evaluation | +| Reproducibility | `tools/reproduce.py` regenerates all 18 result sets and figures. CI runs it on every change | + +## 5. Execution + +| Run | Scope | Records | Protocol | +|---|---|---|---| +| Literature pilots | 6 models × 18 CIViC claims: with the source, plain questions (1a, 1b), semantic review | 108 answers each | fixed before each run | +| Literature agent | 6 models × 18 claims, PubMed tools, feedback; with and without stances | 108 episodes each | fixed before each run | +| Single-cell pilot | 6 models × 24 clusters | 144 episodes | protocol 1 | +| Single-cell held-out | 6 models × 46 clusters | 276 episodes | protocol 2 (`d134d1e`) | +| Single-cell external | 6 models × 81 clusters, six new studies | 486 episodes | protocol 3 (`c86d5d4`) | +| ClinVar, VBO, CIViC cases | Deterministic replays from pinned sources | 5,026 + 72 + controls | engine and grounders | + +## 6. Results + +### 6.1 Identifier and citation integrity + +| Experiment | Model's answers with an invalid identifier or citation | Admitted with one | +|---|---|---| +| Single-cell held-out | 63/276, 22.8% [18.3, 28.1] | 0/255 [0, 1.5] | +| Single-cell external | 70/479, 14.6% [11.7, 18.1] | 0/390 [0, 1.0] | +| Literature, claim only, no source | 38/50, 76.0% [62.6, 85.7] | 0/3 [0, 56.1] | +| Literature agent with PubMed | 8/77, 10.4% [5.4, 19.2] | 0/73 [0, 5.0] | +| Pooled | | **0/721 [0, 0.5]** | + +In single-cell annotation, the typical error is an ID that belongs to another term. For example, "pancreatic +delta cell" was given the ID of a brown preadipocyte. Error rates differ by model: on the held-out split, +Claude Haiku 4.5 had 23 of 46 answers with an identifier or marker error, GPT-5.6-Luna 16, and Claude Opus 5.5 4. +The outcome after bioevidence does not differ by model. + +Deterministic controls: + +| Control | Schema only | Validator | Validator with grounding | +|---|---|---|---| +| ClinVar controlled faults | 160/160 admitted | 0/160 [0, 2.3] | 0/160 | +| ClinVar trust-boundary forgeries | 32/32 | 32/32 | 0/32 [0, 10.7] | +| VBO controlled faults / forgeries | 160/160 / 48/48 | 0/160 / 48/48 | 0/160 / 0/48 | + +### 6.2 Agreement with the reference (single-cell held-out, protocol 2) + +![Funnel for the single-cell held-out clusters, stage by stage](assets/ai_validation_funnel.svg) + +| Configuration | Compatible with the authors' term | Wrong or invalid | Sent to a person | +|---|---|---|---| +| Model alone | 176/276, 63.8% [57.9, 69.2] | 100/276, 36.2% [30.8, 42.1] | 0 | +| Gate | 161 admitted | 37/198 admitted, 18.7% [13.9, 24.7] | 78 | +| Feedback loop | 205/276, 74.3% [68.8, 79.1] | 50/255 admitted, **19.6% [15.2, 24.9]** | 21, 7.6% [5.0, 11.3] | + +What the remaining disagreements are: +- 7 call a duodenal cluster that the authors label "B cell" a plasma cell. Its top markers are JCHAIN and MZB1, + so this is probably label noise. +- 11 call a "mesenchymal cell" a fibroblast or stellate cell. +- The rest are fine-grained confusions: NK T against gamma-delta T cells, memory against naive CD4 T cells. + +On the external split the base rates are higher. Wrong or invalid annotations fell from 42.0% [37.6, 46.4] of +the models' answers to 37.9% [33.3, 42.9] of those admitted. That run used protocol 3, whose definition check is +discussed in 6.3. + +In the literature agent (pilot), stances and a "conflicting" outcome reduced wrong decisions that were admitted +from 15/108, 13.9% [8.6, 21.7], to 10/108, 9.3% [5.1, 16.2]. Another 10 records went to an expert as +conflicting. + +### 6.3 Negative result: the Cell Ontology definition check + +The check compares the markers a claimed term is defined to have or lack with the cluster's measurements. Its +thresholds were set on the development datasets, where 70% of the answers it flagged were wrong. + +On the external split, the answers it flagged were wrong 46.6% [36.5, 56.9] of the time, against a base rate of +42.0%: no better than chance. The cause is protein definitions that do not hold for transcripts: +- mast cell CCR3 and neutrophil CEACAM8 are not detected; +- NK cells are defined as lacking CD3 epsilon, but their CD3E transcripts are detected. + +Fed back in the loop, its findings turned 15 correct answers into wrong ones. Compatible answers in the loop +(242 of 486) fell below the model alone (278). + +### 6.4 Deviations from the plan + +- **Protocols changed between splits, never within one.** After the pilot, the ASCT+B cross-check was removed, + because its biomarker lists are not specificity statements, and the request for contradicting markers was + narrowed. Protocol 2 was then frozen. +- **One protocol detail changed before any results.** Gemini models first tried to use their own tools. The + prompt gained "answer from what you know" before the pilot. +- **One run was resumed.** The external run hit a time limit at 321 of 486 episodes and was resumed. Episodes are + independent, and none was repeated. +- **Two planned evaluations have not been done.** The literature held-out set was built but not run, and no expert + review or audit of admitted records has been done (#23, #24). + +## 7. Adequacy for the context of use + +- **Identifier and citation integrity: adequate.** No invalid identifier or citation was admitted in 721 admitted + AI answers, for every model. The check is deterministic and independent of the model that proposes, so the + result transfers to new models as long as the pinned references cover the records. +- **Semantic correctness: not established.** About one admitted single-cell annotation in five disagrees with the + authors' term. Part of that is label noise, but how much cannot be known without experts. Until an audit + measures it, admitted records are fit for research summaries that state this error rate, and not for uses that + depend on fine cell-type distinctions. +- **Literature direction: indicative only.** It rests on an 18-claim pilot. +- **Clinical or regulatory use: not adequate,** and outside the context of use. + +Proposed acceptance criteria for the next evaluation, to be fixed before it is run: +1. Admitted records with an invalid identifier or citation: upper 95% bound below 1%. Met today. +2. Semantic error of admitted records, measured by an expert audit of a random sample: upper 95% bound below a + threshold set per use, for example 10% for research summaries. +3. Records sent to a person: at most 20%. + +## 8. Life-cycle maintenance + +- **Model change.** When a model or its version changes, re-run the held-out and external splits with the frozen + protocol, and compare per model. The integrity result should not move; the semantic result can. +- **Reference change.** A new Cell Ontology, HGNC or ClinVar release changes the pinned snapshots. Re-run + `tools/reproduce.py` and the affected splits, and keep the earlier snapshots for comparison. +- **Monitoring in use.** Audit a random sample of admitted records on a schedule (#24), and report the audited + error rate with its interval next to the figures above. + +## Sources + +- [LLM benchmark](benchmarks/llm-benchmark.md): literature pilots, scenarios 1–2b, semantic checks, scenario 3. +- [Single-cell case](benchmarks/singlecell.md), [ClinVar case](benchmarks/clinvar.md), + [VBO case](benchmarks/vbo-canine.md), [CIViC literature case](benchmarks/civic-literature.md). +- [Engineering contract](ENGINEERING.md): rule codes, grounders, the feedback loop, reproduction. +- [AI validation roadmap](AI_VALIDATION_ROADMAP.md) and [error taxonomy](ERROR_TAXONOMY.md). diff --git a/docs/index.md b/docs/index.md index d9b96b7..e8335bc 100644 --- a/docs/index.md +++ b/docs/index.md @@ -131,6 +131,7 @@ result rest on that alone (`BEV008`). The manually curated version of the same c | I want to… | Read | |---|---| | See the AI benchmarks | [LLM benchmark](benchmarks/llm-benchmark.md), [single-cell annotation](benchmarks/singlecell.md) | +| Judge whether it is fit for a use | [Validation dossier](VALIDATION_DOSSIER.md) | | Write records quickly or from LLM output | [Drafts](DRAFTS.md) | | Encode my domain's admission policy | [Profiles](PROFILES.md), [community profiles](community-profiles.md) | | Check records in pull requests | [GitHub Action](https://github.com/NingyuSUN/bioai-evidence-validator#quickstart) | diff --git a/mkdocs.yml b/mkdocs.yml index 655d697..5481b2f 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -55,6 +55,7 @@ nav: - Reference: - Engineering contract: ENGINEERING.md - Standards alignment: STANDARDS.md + - Validation dossier: VALIDATION_DOSSIER.md - AI validation roadmap: AI_VALIDATION_ROADMAP.md - Error taxonomy: ERROR_TAXONOMY.md - Design: