Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,16 @@

## Unreleased

- Add the Cell Ontology definition check (`definitions`): the presence and absence marker axioms of a
claimed term, own and inherited, mapped to genes and compared with the subject's measurements; a
contradiction sends the record to review (BEV026). `Ontology` now reads logical-definition relations
and gene names of protein terms. The single-cell case adds per-cluster definition panels and an
external split of six datasets from other studies (81 clusters). Run there with thresholds frozen on
the first six datasets (protocol 3), the check did not transfer. It flagged wrong annotations no better
than chance (47% against 42%), because several protein definitions (mast cell CCR3, neutrophil
CEACAM8, NK cells lacking CD3 epsilon) do not hold for transcripts. Fed back in the loop, its findings
turned 15 correct answers into wrong ones. Identifier errors (70 of 479 answers) were still all caught.
The check is marked experimental; its findings still go back to the proposer in the feedback loop.
- Add the single-cell cell-type annotation case (`examples/singlecell_celltype`): marker tables derived
from six CELLxGENE datasets annotated by their authors, with the Cell Ontology, HGNC and ASCT+B pinned.
Benchmark scenario 3 runs six models on it, alone and in the feedback loop. In the pilot, 35 of 138 first
Expand Down
42 changes: 42 additions & 0 deletions docs/ENGINEERING.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@ evidence items), and duplicate source/item/line/adjudication IDs cannot be admit
| BEV023 | Grounding: a cited identifier is obsolete, withdrawn, or not its current name (a previous gene symbol or alias) | Reject |
| BEV024 | Grounding: a cited identifier is malformed, or of the wrong kind for its place (e.g. not a cell type) | Reject |
| BEV025 | Cross-check: a pinned reference resource contradicts the claim or its supporting evidence | Review |
| BEV026 | Definition check: measurements contradict a marker the ontology defines the claimed term to have or lack | Review |

A human acceptance does not erase contradictions, missing evidence, source mismatches,
or other human rejection/deferral. Low-strength extraction permissions are explicit
Expand Down Expand Up @@ -181,6 +182,47 @@ bioevidence validate record.json --ontology cl-basic.obo --term-root cell_type=C
`TableGrounder` and `ReferenceGrounder` need their configuration (evidence types, keys, relation) and are
used from Python.

### Ontology definitions as checks

Many Cell Ontology terms are defined by marker proteins. A CD8-positive, alpha-beta T cell `has plasma membrane
part` the CD8 co-receptor and `lacks plasma membrane part` CD4; a natural killer cell lacks CD3 epsilon. These
axioms belong to a pinned release, not to a judgment about one dataset, so they can be checked against
measurements without knowing the right answer.

- `MarkerDefinitions` (`definitions`) collects each term's presence and absence axioms, its own and inherited
through `is_a`. It maps each protein to HGNC genes:
- by its PRO short label or gene-based synonym;
- as a family (`Fcgr3` is FCGR3A and FCGR3B);
- through the components of a complex (the CD8 co-receptor is CD8A and CD8B);
- or through the protein a modified form belongs to.
- It leaves some axioms out:
- isoform-specific markers (CD45RA), which gene-level counts cannot see;
- axioms about relative amounts, which compare with another cell type rather than with the rest of a
dataset.

It needs the full `cl.obo`, which keeps the logical definitions.
- `DefinitionGrounder` compares the claimed term with `measure(record)`, which gives each gene's detection rate
and log fold change for the record's subject. It raises BEV026 in two cases:
- **presence:** every gene of a defining protein is detected in fewer than 10% of the cells;
- **absence:** a gene of an excluded protein is detected in at least half of them and more than elsewhere.

The thresholds were set on the single-cell case's first six datasets, before it was run on six others.

BEV026 sends a record to review and does not reject it, because a transcript is not a protein and some
definitions are written for one species. In the feedback loop it goes back to the proposer together with the
measurement. The proposer cannot make the finding go away by leaving evidence out, because the check reads the
data, not the evidence the proposer cites. The check is **experimental**.

**It did not transfer to new data.** On the single-cell case's external split, the thresholds set on the
first six datasets flagged wrong annotations no better than chance (47% against 42%). The cause was definitions
that do not hold for transcripts:
- mast cells defined by CCR3, and neutrophils by CEACAM8, whose transcripts are not detected;
- NK cells defined as lacking CD3 epsilon, although their CD3E transcripts are.

Fed back to models, these findings turned correct answers into wrong ones (see the benchmark README). Use it
on transcript data only with markers validated at the mRNA level, since the loop passes its findings to the
proposer as they are.

### The feedback loop

`feedback.revise(propose, validator)` runs the loop around any proposer: a model call, an agent or a person.
Expand Down
53 changes: 53 additions & 0 deletions evaluation/llm_benchmark/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -345,6 +345,59 @@ Limits:
"wrong" is partly label noise; exact plus coarser plus finer is the fairer measure of agreement.
- One run per split. The test split has 46 clusters from the same six datasets as the pilot.

### External split (protocol 3): the definition check did not transfer

Protocol 3 adds the Cell Ontology definition check (`bioevidence_validator.definitions`, BEV026) to
protocol 2. The check compares the markers a claimed term is defined to have or lack with the cluster's
measurements:
- a defining marker detected in under 10% of cells is a contradiction;
- an excluded marker detected in half or more, and higher than elsewhere, is a contradiction.

The thresholds were set on the 70 clusters of the first six datasets. There, the 20 first answers the check
flagged were wrong 70% of the time, against 38% overall. Protocol 3 was committed before it was run on six
datasets from other studies, 81 clusters, mostly immune ([results](results/celltype-external/summary.md)).
The gate is reported both with and without the check, on the same first answers.

| Configuration | Annotated | Compatible (exact + coarser + finer) | Exact | Wrong or invalid | Routed to a person |
|---|---:|---:|---:|---:|---:|
| Model alone | 479/486 | 278 | 169 | 201 | 0 |
| Model + gate without the definition check | 362/486 | 242 | 154 | 120 | 117 |
| Model + gate with it | 302/486 | 202 | 114 | 100 | 177 |
| Model + feedback loop with it | 390/486 | 242 | 125 | 148 | 89 |

**What this shows: a negative result.**

- **On new data the check flagged answers no better than chance.** It flagged 88 first answers: 41 wrong,
43 exactly the authors' term and 4 finer. That is 47% wrong, against 42% wrong overall. Behind the gate it
stopped 20 more wrong answers and 40 more correct ones, and left the error rate of what was admitted
unchanged (33% with and without it).
- **It pushed models away from correct answers.** In the loop, the check's findings go back to the model,
and models believed them. Of 40 answers that matched the authors' term exactly and were revised:
- 15 were changed to a wrong term, 8 of which were then admitted;
- 12 became coarser or finer;
- 13 stayed exact.

As a result, compatible answers in the loop (242) fell below the model alone (278).
- **The cause is transcripts against proteins.** The Cell Ontology defines cell types by surface proteins,
and several definitions do not hold at the mRNA level:
- a mast cell is defined by CCR3, detected in 0% of all four mast cell clusters;
- a neutrophil is defined by CEACAM8 (CD66b), also undetected;
- a natural killer cell is defined as lacking CD3 epsilon, yet CD3E transcripts are detected in 50–68% of
NK and ILC3 cells.

These rules fire on the authors' own labels: 11 of 81 external clusters. The development datasets had
almost no mast cells or neutrophils, so the problem did not show before.
- **The identifier layer held.** 70 of 479 first answers had an identifier error (GPT-5.6-Luna 33,
Claude Haiku 4.5 23), and none was admitted.

Two lessons:
- **Feed back only findings that are facts.** An identifier that belongs to another term is a fact. A
definition that may not hold for transcripts is not, and a model in a feedback loop treats any finding as
an instruction. A finding of uncertain precision belongs with a person, not in the model's next prompt.
- **Validate before transferring.** A protein-level definition can become a transcript-level check only for
markers shown to be reliable at the mRNA level, for example against CITE-seq data that measures both.
That needs its own validation on data not used to choose the markers.

## Semantic checks (#31)

Grounding cannot tell whether a verbatim quote supports the claim, so two semantic layers were
Expand Down
54 changes: 45 additions & 9 deletions evaluation/llm_benchmark/celltype_loop.py
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,10 @@
against the answer and cross-checked supporting markers with ASCT+B. Protocol 2 (the held-out test split), from
what the pilot showed, asks for contradicting markers only when some clearly point to another cell type (a mixed
cluster, doublets), and leaves out the ASCT+B cross-check, whose biomarker lists are not specificity statements.
Protocol 3 (the external split, six datasets from other studies) keeps protocol 2 and adds the Cell Ontology
definition check (`bioevidence_validator.definitions`): a claimed term whose defining markers the cluster's
measurements contradict (a CD8 T cell in which CD4 is detected in most cells) goes back to the model with the
measurement, as a finding to review (BEV026). Its thresholds were set on the first six datasets.

One run gives three views: the model alone (its first answer as given), behind a bioevidence gate (its first
answer, validated) and in the feedback loop (its last answer, validated). Scoring is independent of the
Expand All @@ -27,6 +31,7 @@
from __future__ import annotations

import argparse
import collections
import concurrent.futures
import datetime as dt
import hashlib
Expand All @@ -40,6 +45,7 @@

from bioevidence_validator import feedback
from bioevidence_validator.crosscheck import ReferenceGrounder
from bioevidence_validator.definitions import DefinitionGrounder, MarkerDefinitions
from bioevidence_validator.engine import RecordValidator
from bioevidence_validator.grounding import SnapshotStore, SourceBytesGrounder
from bioevidence_validator.identifiers import GeneGrounder, Genes, Ontology, OntologyGrounder, _label_key
Expand Down Expand Up @@ -91,7 +97,7 @@
different cell type, for example in a mixed cluster or doublets, list those with stance "contradicts". If the \
markers do not identify a cell type, set "decision" to "uncertain". Answer from what you know: do not run \
commands, search the web or read files."""
PROMPTS = {1: PROMPT_V1, 2: PROMPT}
PROMPTS = {1: PROMPT_V1, 2: PROMPT, 3: PROMPT}
REVISION = """{prompt}

Your previous answer:
Expand Down Expand Up @@ -121,6 +127,16 @@ def __init__(self) -> None:
self.tasks = {t["task_id"]: t for t in load_jsonl(SOURCES / "tasks.jsonl")}
self.versions = {d["markers_sha256"]: d["dataset_version_id"] for d in self.manifest["datasets"]}
self.tables = {sha: parse_table(self.store.get(sha) or b"")[1] for sha in self.versions}
self.panels: dict[str, dict[str, dict[str, tuple[float, float]]]] = {}
for dataset in self.manifest["datasets"]:
for row in parse_table(self.store.get(dataset["panel_sha256"]) or b"")[1]:
cluster = self.panels.setdefault(dataset["name"], {}).setdefault(row["cluster"], {})
cluster[row["gene"]] = (float(row["pct_in"]), float(row["logfc"]))

def measure(self, record: dict) -> dict[str, tuple[float, float]] | None:
"""The definition panel of the record's cluster (`cluster:<dataset>/<cluster>`)."""
dataset, _, cluster = record["statement"]["subject"]["id"].removeprefix("cluster:").partition("/")
return self.panels.get(dataset, {}).get(cluster)

def validator(self, protocol: int = 2) -> RecordValidator:
grounders: list = [SourceBytesGrounder(self.store), TableGrounder(self.store, ["marker_gene"]),
Expand All @@ -129,6 +145,8 @@ def validator(self, protocol: int = 2) -> RecordValidator:
grounders.append(ReferenceGrounder.from_table(self.reference, label="ASCT+B", evidence_key="gene",
relation="marker_of", ontology=self.ontology,
conflict="disjoint"))
if protocol >= 3:
grounders.append(DefinitionGrounder(MarkerDefinitions(self.ontology, self.genes), self.measure))
return RecordValidator(profile=CASE / "profile.yaml", grounders=grounders)

def markers(self, task: dict) -> set[str]:
Expand Down Expand Up @@ -294,6 +312,12 @@ def problems(case: Case, task: dict, answer: dict) -> list[str]:

VIEWS = [("model", "Model alone (first answer)"), ("gate", "Model + bioevidence gate (first answer)"),
("loop", "Model + bioevidence feedback loop (last answer)")]
# Protocol 3 only: the gate as protocol 2 would have it, on the same first answers (the prompts are the same).
WITHOUT_DEFINITIONS = ("gate_without_definitions", "Model + gate without the definition check (first answer)")


def views_for(protocol: int) -> list[tuple[str, str]]:
return VIEWS[:2] + [WITHOUT_DEFINITIONS] + VIEWS[2:] if protocol >= 3 else VIEWS


def view(row: dict, name: str) -> dict:
Expand All @@ -305,10 +329,11 @@ def view(row: dict, name: str) -> dict:
return {"answer": first if first and first["decision"] == "annotate" else None, "routed": False, "expert": False}
if not attempts:
return {"answer": None, "routed": False, "expert": False}
chosen = attempts[0] if name == "gate" else attempts[-1]
index = 0 if name == "gate" else len(attempts) - 1
first = name in ("gate", WITHOUT_DEFINITIONS[0])
chosen = attempts[0] if first else attempts[-1]
index = 0 if first else len(attempts) - 1
annotating = [a for a in answered if a["decision"] == "annotate" and a["cell_type_id"].strip()]
if chosen["status"] == "admitted":
if chosen["status"] == "admitted" or (name == WITHOUT_DEFINITIONS[0] and set(chosen["codes"]) == {"BEV026"}):
return {"answer": annotating[index], "routed": False, "expert": False}
return {"answer": None, "routed": True, "expert": chosen["to_expert"]}

Expand Down Expand Up @@ -337,15 +362,23 @@ def metrics(mine: list[dict], name: str) -> dict:

split = case.tasks[rows[0]["task_id"]]["split"] if rows else "pilot"
protocol = rows[0].get("protocol", 1) if rows else 2
names = views_for(protocol)
summary = {"benchmark": f"singlecell-celltype-v{protocol}", "split": split,
"results": {m: {n: metrics([r for r in rows if r["model"] == m], n) for n, _ in VIEWS} for m in models},
"pooled": {n: metrics(rows, n) for n, _ in VIEWS}, "episodes": len(rows),
"results": {m: {n: metrics([r for r in rows if r["model"] == m], n) for n, _ in names} for m in models},
"pooled": {n: metrics(rows, n) for n, _ in names}, "episodes": len(rows),
"calls": sum(len(r["calls"]) for r in rows), "failed_calls": sum(c["error"] is not None
for r in rows for c in r["calls"]),
"revised": sum(len(r["attempts"]) > 1 for r in rows),
"codes": {c: sum(c in a["codes"] for r in rows for a in r["attempts"])
for c in sorted({c for r in rows for a in r["attempts"] for c in a["codes"]})},
"episodes_sha256": hashlib.sha256((results / "episodes.jsonl").read_bytes()).hexdigest()}
if protocol >= 3: # first answers the definition check flagged, by how they compare with the authors' term
flagged = collections.Counter()
for r in rows:
got = [c["answer"] for c in r["calls"] if c["answer"] and c["answer"]["decision"] == "annotate"]
if r["attempts"] and "BEV026" in r["attempts"][0]["codes"]:
flagged[outcome(case, truth[r["task_id"]]["term"], got[0])] += 1
summary["definition_flags_on_first_answers"] = dict(sorted(flagged.items()))
(results / "summary.json").write_text(json.dumps(summary, indent=2, sort_keys=True) + "\n", encoding="utf-8",
newline="\n")
(results / "summary.md").write_text(render(summary), encoding="utf-8", newline="\n")
Expand All @@ -365,7 +398,7 @@ def render(summary: dict) -> str:
"or marker error | Routed to a person |", "|---|---|---:|---:|---:|---:|---:|---:|---:|---:|"]
blocks = [(score_claims.MODEL_NAMES[m], v) for m, v in summary["results"].items()] + [("All models", summary["pooled"])]
for name, views in blocks:
for key, label in VIEWS:
for key, label in views_for(protocol):
m = views[key]
lines.append(f"| {name} | {label} | {f(m['answered'])} | {f(m['exact'])} | {f(m['coarser'])} | "
f"{f(m['finer'])} | {f(m['wrong'])} | {f(m['invalid'])} | "
Expand All @@ -377,6 +410,9 @@ def render(summary: dict) -> str:
"", f"{summary['episodes']} episodes, {summary['calls']} model calls ({summary['failed_calls']} failed), "
f"{summary['revised']} revised after feedback. Rule codes raised across all records: "
+ ", ".join(f"{c} {n}" for c, n in summary["codes"].items()) + ".", ""]
if "definition_flags_on_first_answers" in summary:
lines[-1:] = ["", "First answers the definition check flagged, by comparison with the authors' term: "
+ ", ".join(f"{k} {n}" for k, n in summary["definition_flags_on_first_answers"].items()) + ".", ""]
return "\n".join(lines)


Expand All @@ -385,11 +421,11 @@ def main(argv: list[str] | None = None) -> int:
steps = parser.add_subparsers(dest="step", required=True)
r = steps.add_parser("run")
r.add_argument("--output", type=Path, required=True)
r.add_argument("--split", choices=["pilot", "test"], default="pilot")
r.add_argument("--split", choices=["pilot", "test", "external"], default="pilot")
r.add_argument("--models", nargs="+", choices=list(run_models.MODELS), default=list(run_models.MODELS))
r.add_argument("--workers", type=int, default=2)
r.add_argument("--limit", type=int)
r.add_argument("--protocol", type=int, choices=sorted(PROMPTS), default=2)
r.add_argument("--protocol", type=int, choices=sorted(PROMPTS), default=3)
c = steps.add_parser("collect")
c.add_argument("--output", type=Path, required=True)
c.add_argument("--results", type=Path, required=True)
Expand Down
Loading
Loading