Skip to content

Single-cell cell-type annotation case and benchmark scenario 3 - #36

Merged
NingyuSUN merged 3 commits into
mainfrom
singlecell-celltype
Oct 5, 2026
Merged

NingyuSUN merged 3 commits into
mainfrom
singlecell-celltype

Conversation

@NingyuSUN

@NingyuSUN NingyuSUN commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Stacked on #35. A second real AI experiment outside the literature: single-cell cell-type annotation, using the generic grounders and feedback loop from #35.

Case (examples/singlecell_celltype)

  • Data: six normal human CELLxGENE datasets (CC BY 4.0), each annotated by its authors with Cell Ontology terms: liver, kidney, adult duodenum, pancreas, skin, and fetal lung T/NK/ILC.
  • Marker tables: each author cell type with ≥ 30 cells is a cluster. The table keeps the top 50 genes by log fold change from raw counts, named by current HGNC symbols. That gives 70 clusters: 24 pilot and 46 held out.
  • Pinned references: the full Cell Ontology 2026-06-08 (it keeps the disjointness axioms), HGNC 2026-09-30, and HuBMAP ASCT+B for five organs (1,027 gene marker_of cell type assertions).
  • Committed: 6.5 MB of content-addressed snapshots. The expression matrices stay out; their SHA-256 are in manifest.json.

Scenario 3 (evaluation/llm_benchmark/celltype_loop.py)

Each of the six models annotates a cluster from its top 20 markers with a CL term and markers that support it or argue against it. feedback.revise runs up to three answers:

  • fixable findings go back to the model;
  • a record with the model's own contradicting markers (BEV004), or an ASCT+B conflict (BEV025), goes to a person.

The scorer compares answers with the author term through the ontology, independently of the grounders.

Pilot, 24 clusters × 6 models:

Configuration Annotated Exact Coarser Finer Wrong or invalid Identifier or marker error Routed to a person
Model alone 138/144 41 21 17 59 38/138 0
+ bioevidence gate 56/144 32 13 1 10 0/56 82
+ feedback loop 67/144 35 13 4 15 0/67 71
  • Identifier hallucination: 35 of 138 first answers gave a CL ID that belongs to another term, e.g. "Paneth cell" under CL:0000147 (pigment cell). By model: Haiku and Luna 12/23 each, Flash 7, Gemini Pro 4, Astra 2, Opus 1. None was admitted. The loop fixed and admitted 11; 26 went to a person because the revised answer still cited markers against itself.
  • Routing concentrates wrong answers:
    • Wrong or invalid: 43% of first answers, against 18% of those admitted behind the gate and 22% in the loop.
    • Model-declared conflict is the strongest signal: wrong 63% of the time (39/62) when present, 26% without.
  • ASCT+B as a conflict source is noise. It fired on 23 first answers, mostly for spurious reasons (CD44 is listed only for a lung T cell, FCER1G only for granulocytes). Its biomarker lists are not specificity statements. The grounder stays generic, but this protocol should not use it that way.
  • Cost: half of the annotations go to a person. Models cited markers against their own answer in 64 of 144 first answers, because the prompt asks for "any that argue against it".

Held-out test split (protocol 2)

Two lessons from the pilot became protocol 2, committed in d134d1e before the 46 held-out clusters were run:

  • contradicting markers are asked for only when some clearly point to another cell type (a mixed cluster or doublets);
  • the ASCT+B cross-check is off.

Nothing was changed after seeing the results.

Configuration Annotated Exact Coarser Finer Wrong or invalid Identifier or marker error Routed to a person
Model alone 276/276 117 33 26 100 63/276 0
+ bioevidence gate 198/276 113 27 21 37 0/198 78
+ feedback loop 255/276 138 31 36 50 0/255 21
  • The loop makes annotations better, not only safer:
    • Answers compatible with the authors' term rose from 176 to 205, and exact matches from 117 to 138.
    • Wrong or invalid fell from 36% of answers to 20% of those admitted.
    • 8% went to a person, against 49% in the pilot.
  • The gain is mostly identifiers:
    • 62 first answers had an ID that does not match its label. By model: Haiku 23/46, Luna 16, Gemini Pro 10, Flash 9, Opus 4, Astra 1.
    • The finding names the term the label belongs to; 56 of the 62 were fixed and admitted.
    • None of these errors was admitted.
  • Model-declared conflict is now rare and sharp: it appeared in 17 first answers, 82% of them wrong (33% for the rest).
  • What remains is semantic, and much of it is the reference:

Limits

  • Author labels are the reference; some are coarse or debatable.
  • One run per split. The 46 held-out clusters come from the same six datasets as the pilot.

Checks

  • ruff, mypy, pytest (98.81% coverage);
  • 15 case tests, including byte-for-byte replays of the pilot and test scores;
  • mkdocs build --strict.

Six CELLxGENE datasets annotated by their authors become 70 clusters
with pinned marker tables, the Cell Ontology, HGNC and ASCT+B. Six
models annotate the 24 pilot clusters, alone and in bioevidence's
feedback loop, validated by the generic table, ontology, gene and
reference grounders.

35 of 138 first answers gave a Cell Ontology ID that belongs to another
term; none was admitted. Wrong or invalid annotations fell from 43% of
answers to 18% of those admitted behind the gate, with half of the
annotations routed to a person. The ASCT+B cross-check proved to be
noise as a conflict source.
From the pilot: ask for contradicting markers only when some clearly
point to another cell type, and leave out the ASCT+B cross-check, whose
biomarker lists are not specificity statements. Protocol 1 stays
available and the pilot replays unchanged. Committed before the test
split is run.
46 clusters, six models, run with the protocol frozen in d134d1e.
Identifier errors in 62 of 276 first answers, none admitted. In the
feedback loop, answers compatible with the authors' term rose from 176
to 205, wrong or invalid fell from 36% of answers to 20% of those
admitted, and 21 of 276 went to a person.
@NingyuSUN NingyuSUN changed the title Single-cell cell-type annotation case and benchmark scenario 3 (pilot) Single-cell cell-type annotation case and benchmark scenario 3 Oct 1, 2026
@NingyuSUN
NingyuSUN changed the base branch from generic-grounders to main October 5, 2026 05:33
@NingyuSUN
NingyuSUN merged commit b42f436 into main Oct 5, 2026
10 checks passed
@NingyuSUN
NingyuSUN deleted the singlecell-celltype branch October 6, 2026 02:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant