diff --git a/CHANGELOG.md b/CHANGELOG.md
index ac10990..810518a 100644
--- a/CHANGELOG.md
+++ b/CHANGELOG.md
@@ -2,6 +2,10 @@
## Unreleased
+- Add the validation dossier (`docs/VALIDATION_DOSSIER.md`), structured after the seven steps of FDA's
+ draft AI credibility framework: question of interest, context of use, model risk, credibility plan,
+ execution, results with deviations, and adequacy. Identifier and citation integrity of admitted records
+ is established (0 of 721); semantic correctness is not, pending an expert audit.
- Add `tools/reproduce.py`: one command regenerates every committed benchmark table, summary and figure
offline (18 steps, about three minutes) and compares each with the repository; CI runs it on every
change. It caught the ClinVar and VBO figures still captioned v0.7.0 after the release, now fixed.
diff --git a/README.md b/README.md
index 9258e3c..275700a 100644
--- a/README.md
+++ b/README.md
@@ -273,13 +273,17 @@ committed files and `--list` shows the steps.
was run once.
- **Not for clinical use.**
+What is and is not established, in the structure of FDA's draft AI credibility framework (question of interest,
+context of use, model risk, credibility evidence, adequacy): the
+[validation dossier](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/VALIDATION_DOSSIER.md).
+
## Roadmap
Tracked in the [AI validation roadmap](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/AI_VALIDATION_ROADMAP.md) (#26):
- expert review of the benchmark cases (#23);
- calibrated triage and audit sampling of admitted records, to measure what still gets through (#24);
-- a validation dossier structured like FDA's draft AI credibility framework (#25). One-command reproduction
- and the funnel figure are done.
+- the [validation dossier](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/VALIDATION_DOSSIER.md),
+ structured after FDA's draft AI credibility framework, states what is established and what is not (#25).
## Contributing and citing
diff --git a/docs/VALIDATION_DOSSIER.md b/docs/VALIDATION_DOSSIER.md
new file mode 100644
index 0000000..cd62f71
--- /dev/null
+++ b/docs/VALIDATION_DOSSIER.md
@@ -0,0 +1,211 @@
+# Validation dossier: AI-assisted curation behind bioevidence
+
+This dossier follows the seven steps of the risk-based credibility assessment framework in FDA's draft
+guidance *Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug
+and Biological Products* (January 2025). It borrows the structure only. It is **not a regulatory submission**,
+it has not been reviewed by FDA, and the system is **not for clinical use**.
+
+Every number comes from the committed results and can be regenerated offline with
+`uv run --frozen --with matplotlib==3.11.2 python tools/reproduce.py`. Intervals are Wilson 95%.
+State: release 0.8.0, October 2026.
+
+## Summary
+
+| Claim | Evidence | Status for the context of use |
+|---|---|---|
+| Admitted records carry no invalid identifier or citation | 0 of 721 admitted AI answers across four experiments (upper bound 0.5%); models alone produced such errors in 10–76% of answers | **Established** |
+| The feedback loop corrects fixable errors without hiding evidence | Single-cell held-out: 57 of 63 answers with identifier or marker errors corrected and admitted; verified evidence cannot be withdrawn | **Established** |
+| Admitted records are semantically correct | 19.6% [15.2, 24.9] of admitted single-cell annotations disagree with the authors' term. Part of this is label noise. No expert audit yet | **Not established** |
+| Routing load is workable | 7.6% [5.0, 11.3] of single-cell annotations sent to a person (held-out) | **Plausible**, one run |
+| A new semantic check (Cell Ontology definitions) helps | No better than chance on external data, and harmful when fed back | **Refuted**; kept experimental |
+
+## 1. Question of interest
+
+Can records produced by general-purpose LLMs and agents be admitted to a research knowledge base without an
+expert reviewing each one, and which records must an expert see? Two kinds of records are in scope:
+- cell-type annotations of single-cell clusters;
+- literature evidence for cancer-variant claims.
+
+## 2. Context of use
+
+- **The system assessed** is a model and bioevidence together, because the question concerns the records that
+ are admitted.
+ - **Proposer:** one of six general-purpose models, or an agent that searches PubMed, proposes a record. A
+ record is a claim with the evidence it rests on: marker genes of a cluster, or quotes from papers.
+ - **Validator:** bioevidence checks the record against pinned sources and reference releases, under a
+ profile for the intended use.
+ - **Feedback:** fixable findings go back to the proposer, for up to three answers.
+ - **Admission:** an admitted record enters the knowledge base. A record with conflicting evidence, or still
+ failing after the last round, goes to a curator.
+- **Intended use:** research summaries and a research knowledge base.
+
+ Out of scope:
+ - clinical interpretation;
+ - patient care;
+ - regulatory decisions;
+ - training data, unless a human accepts each record. The profiles require that.
+- **Other evidence in the decision:** curators decide every routed record. Admitted records are not reviewed by
+ anyone today; that is the residual risk this dossier measures.
+- **Models:**
+
+ | Key | Model | Tier | Runs through |
+ |---|---|---|---|
+ | claude-opus | `claude-opus-5-5` | frontier | Claude Code 2.1.283 |
+ | gpt-astra | `gpt-6-astra` | frontier | codex-cli 0.153.4 |
+ | gemini-pro | `gemini-3.1-pro-high` | frontier | Antigravity 1.2.13 |
+ | claude-haiku | `claude-haiku-4-5-20251001` | fast | Claude Code 2.1.283 |
+ | gpt-luna | `gpt-5.6-luna` | fast | codex-cli 0.153.4 |
+ | gemini-flash | `gemini-3.8-flash-medium` | fast | Antigravity 1.2.13 |
+
+ Every reply must match a JSON Schema, and the harness, not the model's CLI, executes the tools. Bioevidence
+ must not depend on which model proposes, so all six are evaluated with one pipeline.
+
+## 3. Model risk
+
+The guidance combines **model influence**, how much the output drives the decision, with **decision
+consequence**, the impact of a wrong decision.
+
+- **Influence is high for admitted records:** nothing else checks them. It is low for routed records, where a
+ curator decides.
+- **Consequence is moderate in this context.** A wrong cell type or citation in a research knowledge base
+ spreads to analyses, summaries and possibly to training sets, but it does not reach patients. Outside the
+ context (clinical use) it would be high.
+- **Risk is therefore medium.** It calls for:
+ - evaluation on data not used to design the system;
+ - protocols fixed before results are seen;
+ - negative controls;
+ - full reproducibility;
+ - a measure of what still gets through. That last one needs an expert audit, which has not been done.
+
+## 4. Credibility assessment plan
+
+| Element | What was planned and done |
+|---|---|
+| Sources | Pinned by SHA-256, with licences and attribution:
- 12 CELLxGENE datasets (CC BY 4.0)
- the Cell Ontology 2026-06-08 (CC BY 4.0)
- HGNC 2026-09-30 (CC0)
- HuBMAP ASCT+B (CC BY 4.0)
- CIViC evidence (CC0)
- PMC open-access full texts (CC BY / CC0)
- a ClinVar 2023-09 sample, followed to 2026-09
|
+| Reference standards | - Single-cell: each study's own author annotations, compared through the ontology (exact, coarser, finer or wrong)
- Literature: the direction CIViC records for the claim
- ClinVar: NCBI's review status, and what happened to each classification three years later
No independent expert labels yet (#23) |
+| Splits | - Single-cell: a pilot (24 clusters) for design; a held-out test (46 clusters, same datasets); an external split (81 clusters from six other studies)
- Literature: pilot only (18 claims); its held-out set (92 claims) is built but not run
|
+| Protocol freezes | Each protocol was committed before the split it was run on: - protocol 2 at `d134d1e`, before the held-out test
- protocol 3 at `c86d5d4`, before the external split
Protocol 1 was fixed before the pilot and committed with its results |
+| Metrics | Defined in the scoring code before the runs: - identifier or citation errors in the model's answers, and in what is admitted
- agreement with the reference
- wrong answers admitted
- records sent to a person
Each is reported per model and pooled |
+| Negative controls | - ClinVar: 160 controlled faults and 32 trust-boundary forgeries
- VBO: 160 faults and 48 forgeries
- CIViC: retracted papers, PMIDs that do not exist, misattributed quotes
|
+| Acceptance criteria | **Not set numerically in advance.** This is a gap; section 7 proposes criteria for the next evaluation |
+| Reproducibility | `tools/reproduce.py` regenerates all 18 result sets and figures. CI runs it on every change |
+
+## 5. Execution
+
+| Run | Scope | Records | Protocol |
+|---|---|---|---|
+| Literature pilots | 6 models × 18 CIViC claims: with the source, plain questions (1a, 1b), semantic review | 108 answers each | fixed before each run |
+| Literature agent | 6 models × 18 claims, PubMed tools, feedback; with and without stances | 108 episodes each | fixed before each run |
+| Single-cell pilot | 6 models × 24 clusters | 144 episodes | protocol 1 |
+| Single-cell held-out | 6 models × 46 clusters | 276 episodes | protocol 2 (`d134d1e`) |
+| Single-cell external | 6 models × 81 clusters, six new studies | 486 episodes | protocol 3 (`c86d5d4`) |
+| ClinVar, VBO, CIViC cases | Deterministic replays from pinned sources | 5,026 + 72 + controls | engine and grounders |
+
+## 6. Results
+
+### 6.1 Identifier and citation integrity
+
+| Experiment | Model's answers with an invalid identifier or citation | Admitted with one |
+|---|---|---|
+| Single-cell held-out | 63/276, 22.8% [18.3, 28.1] | 0/255 [0, 1.5] |
+| Single-cell external | 70/479, 14.6% [11.7, 18.1] | 0/390 [0, 1.0] |
+| Literature, claim only, no source | 38/50, 76.0% [62.6, 85.7] | 0/3 [0, 56.1] |
+| Literature agent with PubMed | 8/77, 10.4% [5.4, 19.2] | 0/73 [0, 5.0] |
+| Pooled | | **0/721 [0, 0.5]** |
+
+In single-cell annotation, the typical error is an ID that belongs to another term. For example, "pancreatic
+delta cell" was given the ID of a brown preadipocyte. Error rates differ by model: on the held-out split,
+Claude Haiku 4.5 had 23 of 46 answers with an identifier or marker error, GPT-5.6-Luna 16, and Claude Opus 5.5 4.
+The outcome after bioevidence does not differ by model.
+
+Deterministic controls:
+
+| Control | Schema only | Validator | Validator with grounding |
+|---|---|---|---|
+| ClinVar controlled faults | 160/160 admitted | 0/160 [0, 2.3] | 0/160 |
+| ClinVar trust-boundary forgeries | 32/32 | 32/32 | 0/32 [0, 10.7] |
+| VBO controlled faults / forgeries | 160/160 / 48/48 | 0/160 / 48/48 | 0/160 / 0/48 |
+
+### 6.2 Agreement with the reference (single-cell held-out, protocol 2)
+
+
+
+| Configuration | Compatible with the authors' term | Wrong or invalid | Sent to a person |
+|---|---|---|---|
+| Model alone | 176/276, 63.8% [57.9, 69.2] | 100/276, 36.2% [30.8, 42.1] | 0 |
+| Gate | 161 admitted | 37/198 admitted, 18.7% [13.9, 24.7] | 78 |
+| Feedback loop | 205/276, 74.3% [68.8, 79.1] | 50/255 admitted, **19.6% [15.2, 24.9]** | 21, 7.6% [5.0, 11.3] |
+
+What the remaining disagreements are:
+- 7 call a duodenal cluster that the authors label "B cell" a plasma cell. Its top markers are JCHAIN and MZB1,
+ so this is probably label noise.
+- 11 call a "mesenchymal cell" a fibroblast or stellate cell.
+- The rest are fine-grained confusions: NK T against gamma-delta T cells, memory against naive CD4 T cells.
+
+On the external split the base rates are higher. Wrong or invalid annotations fell from 42.0% [37.6, 46.4] of
+the models' answers to 37.9% [33.3, 42.9] of those admitted. That run used protocol 3, whose definition check is
+discussed in 6.3.
+
+In the literature agent (pilot), stances and a "conflicting" outcome reduced wrong decisions that were admitted
+from 15/108, 13.9% [8.6, 21.7], to 10/108, 9.3% [5.1, 16.2]. Another 10 records went to an expert as
+conflicting.
+
+### 6.3 Negative result: the Cell Ontology definition check
+
+The check compares the markers a claimed term is defined to have or lack with the cluster's measurements. Its
+thresholds were set on the development datasets, where 70% of the answers it flagged were wrong.
+
+On the external split, the answers it flagged were wrong 46.6% [36.5, 56.9] of the time, against a base rate of
+42.0%: no better than chance. The cause is protein definitions that do not hold for transcripts:
+- mast cell CCR3 and neutrophil CEACAM8 are not detected;
+- NK cells are defined as lacking CD3 epsilon, but their CD3E transcripts are detected.
+
+Fed back in the loop, its findings turned 15 correct answers into wrong ones. Compatible answers in the loop
+(242 of 486) fell below the model alone (278).
+
+### 6.4 Deviations from the plan
+
+- **Protocols changed between splits, never within one.** After the pilot, the ASCT+B cross-check was removed,
+ because its biomarker lists are not specificity statements, and the request for contradicting markers was
+ narrowed. Protocol 2 was then frozen.
+- **One protocol detail changed before any results.** Gemini models first tried to use their own tools. The
+ prompt gained "answer from what you know" before the pilot.
+- **One run was resumed.** The external run hit a time limit at 321 of 486 episodes and was resumed. Episodes are
+ independent, and none was repeated.
+- **Two planned evaluations have not been done.** The literature held-out set was built but not run, and no expert
+ review or audit of admitted records has been done (#23, #24).
+
+## 7. Adequacy for the context of use
+
+- **Identifier and citation integrity: adequate.** No invalid identifier or citation was admitted in 721 admitted
+ AI answers, for every model. The check is deterministic and independent of the model that proposes, so the
+ result transfers to new models as long as the pinned references cover the records.
+- **Semantic correctness: not established.** About one admitted single-cell annotation in five disagrees with the
+ authors' term. Part of that is label noise, but how much cannot be known without experts. Until an audit
+ measures it, admitted records are fit for research summaries that state this error rate, and not for uses that
+ depend on fine cell-type distinctions.
+- **Literature direction: indicative only.** It rests on an 18-claim pilot.
+- **Clinical or regulatory use: not adequate,** and outside the context of use.
+
+Proposed acceptance criteria for the next evaluation, to be fixed before it is run:
+1. Admitted records with an invalid identifier or citation: upper 95% bound below 1%. Met today.
+2. Semantic error of admitted records, measured by an expert audit of a random sample: upper 95% bound below a
+ threshold set per use, for example 10% for research summaries.
+3. Records sent to a person: at most 20%.
+
+## 8. Life-cycle maintenance
+
+- **Model change.** When a model or its version changes, re-run the held-out and external splits with the frozen
+ protocol, and compare per model. The integrity result should not move; the semantic result can.
+- **Reference change.** A new Cell Ontology, HGNC or ClinVar release changes the pinned snapshots. Re-run
+ `tools/reproduce.py` and the affected splits, and keep the earlier snapshots for comparison.
+- **Monitoring in use.** Audit a random sample of admitted records on a schedule (#24), and report the audited
+ error rate with its interval next to the figures above.
+
+## Sources
+
+- [LLM benchmark](benchmarks/llm-benchmark.md): literature pilots, scenarios 1–2b, semantic checks, scenario 3.
+- [Single-cell case](benchmarks/singlecell.md), [ClinVar case](benchmarks/clinvar.md),
+ [VBO case](benchmarks/vbo-canine.md), [CIViC literature case](benchmarks/civic-literature.md).
+- [Engineering contract](ENGINEERING.md): rule codes, grounders, the feedback loop, reproduction.
+- [AI validation roadmap](AI_VALIDATION_ROADMAP.md) and [error taxonomy](ERROR_TAXONOMY.md).
diff --git a/docs/index.md b/docs/index.md
index d9b96b7..e8335bc 100644
--- a/docs/index.md
+++ b/docs/index.md
@@ -131,6 +131,7 @@ result rest on that alone (`BEV008`). The manually curated version of the same c
| I want to… | Read |
|---|---|
| See the AI benchmarks | [LLM benchmark](benchmarks/llm-benchmark.md), [single-cell annotation](benchmarks/singlecell.md) |
+| Judge whether it is fit for a use | [Validation dossier](VALIDATION_DOSSIER.md) |
| Write records quickly or from LLM output | [Drafts](DRAFTS.md) |
| Encode my domain's admission policy | [Profiles](PROFILES.md), [community profiles](community-profiles.md) |
| Check records in pull requests | [GitHub Action](https://github.com/NingyuSUN/bioai-evidence-validator#quickstart) |
diff --git a/mkdocs.yml b/mkdocs.yml
index 655d697..5481b2f 100644
--- a/mkdocs.yml
+++ b/mkdocs.yml
@@ -55,6 +55,7 @@ nav:
- Reference:
- Engineering contract: ENGINEERING.md
- Standards alignment: STANDARDS.md
+ - Validation dossier: VALIDATION_DOSSIER.md
- AI validation roadmap: AI_VALIDATION_ROADMAP.md
- Error taxonomy: ERROR_TAXONOMY.md
- Design: