Skip to content

Validation dossier (#25) - #42

Merged
NingyuSUN merged 1 commit into
mainfrom
validation-dossier
Oct 6, 2026
Merged

NingyuSUN merged 1 commit into
mainfrom
validation-dossier

Conversation

@NingyuSUN

Copy link
Copy Markdown
Owner

The last item of #25: docs/VALIDATION_DOSSIER.md, structured after the seven steps of FDA's draft guidance Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products (January 2025). It borrows the structure only: it is not a regulatory submission and not for clinical use.

The seven steps:

  1. Question of interest: can LLM- and agent-produced curation records (single-cell cell types, literature evidence) be admitted to a research knowledge base without expert review, and which must an expert see?

  2. Context of use: the model and bioevidence together; research summaries and a research knowledge base. Clinical and regulatory use are excluded.

  3. Model risk: influence is high for admitted records and consequence is moderate, so the risk is medium.

  4. Credibility plan: pinned sources, reference standards, splits, protocol freezes with their commit hashes, metrics, negative controls and reproducibility. It records honestly that numerical acceptance criteria were not set in advance.

  5. Execution: the runs.

  6. Results, with Wilson intervals and deviations:

    • integrity: 0/721 admitted with an invalid identifier or citation (upper bound 0.5%);
    • agreement: 19.6% [15.2, 24.9] of admitted single-cell annotations disagree with the authors' term;
    • the refuted definition check;
    • the funnel figure.
  7. Adequacy:

    • integrity is established;
    • semantic correctness is not;
    • literature direction is indicative;
    • clinical use is not adequate.

    It also proposes acceptance criteria for the next evaluation.

Life-cycle maintenance covers model changes, reference updates and audit sampling.

Linked from the README, the docs homepage and the site navigation. mkdocs build --strict passes, and every number was checked against the committed summaries.

Structured after the seven steps of FDA's draft AI credibility
framework (January 2025): question of interest, context of use, model
risk, credibility assessment plan, execution, results with deviations,
and adequacy for the context of use, plus life-cycle maintenance. Every
number comes from the committed results, with Wilson intervals.

Identifier and citation integrity of admitted records is established
(0 of 721); semantic correctness is not, pending an expert audit; the
definition check is reported as refuted.
@NingyuSUN
NingyuSUN merged commit 5da9f04 into main Oct 6, 2026
11 checks passed
@NingyuSUN
NingyuSUN deleted the validation-dossier branch October 6, 2026 04:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant