diff --git a/CHANGELOG.md b/CHANGELOG.md index 810518a..e5fa7a8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -2,6 +2,17 @@ ## Unreleased +- Add audit sampling (`audit`; `bioevidence review audit-sample` and `audit-score`, #24). A seeded random + sample of auto-admitted records, sized from a target (`--target 0.01` audits 299 records: no error found + bounds the rate below 1% at 95%), with blind controls from the other routes and a manifest kept from the + auditors. Scoring reports each route's share of records, error rate with Wilson and exact one-sided + (Clopper–Pearson) bounds, and share of expert time from `minutes_spent`. +- Add the routing evidence (`evaluation/llm_benchmark/risk_signals.py`): whether each signal that sends a + record to a person predicts a wrong answer. BEV004, the only one on by default, is shown on every split + (82% vs 33% wrong on the single-cell held-out split); the opt-in signals BEV022, BEV025 and BEV026 are not. +- Add an audit dry run on the single-cell held-out results with the authors' labels as a stand-in auditor: + 14 errors in 59 sampled auto-admitted records, below 34.6% at 95%; the census rate, 19.6%, lies inside. + `tools/reproduce.py` now runs 20 steps. - Add the validation dossier (`docs/VALIDATION_DOSSIER.md`), structured after the seven steps of FDA's draft AI credibility framework: question of interest, context of use, model risk, credibility plan, execution, results with deviations, and adequacy. Identifier and citation integrity of admitted records diff --git a/README.md b/README.md index 275700a..2da45e6 100644 --- a/README.md +++ b/README.md @@ -109,9 +109,10 @@ Across all 276 held-out annotations, stage by stage: | PubMed rate limits, server errors | `Library.fetch` | Exponential backoff on 429 and 5xx, a response cache, a 100 MB download cap | | Wrong paper, retracted paper, quote not in the paper | `LiteratureGrounder` (BEV016, BEV017, BEV019) | The reason goes back to the model; at most three submissions | | ID of another term, obsolete term, gene alias | `OntologyGrounder`, `GeneGrounder` (BEV017, BEV023, BEV024) | The reason goes back, naming the correct identifier | -| The model's own evidence contradicts its answer | BEV004 | Sent to a person; not fed back, so the model is never asked to hide it | +| The model's own evidence contradicts its answer | BEV004 | Sent to a person; not fed back, so the model is never asked to hide it. Flagged answers were wrong 82% of the time vs 33% ([routing evidence](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/evaluation/llm_benchmark/results/risk-signals/summary.md)) | | A revision drops evidence that was verified | `feedback.carry` | The verified evidence is carried into the revision | | Still not admitted after the last round | `feedback.revise` | Sent to a person | +| Admitted but wrong | A seeded random audit (`bioevidence review audit-sample`, `audit-score`) | Each route's error rate with an exact upper bound; a [dry run](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/evaluation/llm_benchmark/results/celltype-audit/summary.md) shows the held-out route misses a 5% target | | A long batch is interrupted | One file per episode | A rerun skips finished episodes (used when a 486-episode run hit a time limit) | ## How it works @@ -251,10 +252,11 @@ uv run --frozen --with matplotlib==3.11.2 python tools/reproduce.py ``` One command regenerates every committed benchmark table, summary and figure, offline, in about three minutes, -and compares each with the repository byte for byte (text files with line endings normalised). It runs 18 steps, +and compares each with the repository byte for byte (text files with line endings normalised). It runs 20 steps, also run in CI on every change: - the three real-data cases; - ten LLM-benchmark result sets; +- the routing evidence and the audit dry run; - the task files and the error taxonomy; - the figures. @@ -281,7 +283,8 @@ context of use, model risk, credibility evidence, adequacy): the Tracked in the [AI validation roadmap](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/AI_VALIDATION_ROADMAP.md) (#26): - expert review of the benchmark cases (#23); -- calibrated triage and audit sampling of admitted records, to measure what still gets through (#24); +- an expert audit of admitted records with `bioevidence review audit-sample` and `audit-score`, which bound + the error rate of what still gets through. The tool and a dry run are done (#24); the audit needs experts. - the [validation dossier](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/docs/VALIDATION_DOSSIER.md), structured after FDA's draft AI credibility framework, states what is established and what is not (#25). diff --git a/docs/AI_VALIDATION_ROADMAP.md b/docs/AI_VALIDATION_ROADMAP.md index 5206851..ca22d93 100644 --- a/docs/AI_VALIDATION_ROADMAP.md +++ b/docs/AI_VALIDATION_ROADMAP.md @@ -82,8 +82,8 @@ source → structured extraction → provenance → deterministic validation | Now | [#31](https://github.com/NingyuSUN/bioai-evidence-validator/issues/31) | Semantic support check: does a verbatim quote support the claim's direction? | #20 | | Next | [#21](https://github.com/NingyuSUN/bioai-evidence-validator/issues/21) | Real AI extraction experiment on openly licensed sources | #18, #20 | | Next | [#22](https://github.com/NingyuSUN/bioai-evidence-validator/issues/22) | Model and agent evaluation against expert labels | #17, #23 | -| Then | [#24](https://github.com/NingyuSUN/bioai-evidence-validator/issues/24) | Calibrated triage and audit of auto-admitted records | #23 | -| Last | [#25](https://github.com/NingyuSUN/bioai-evidence-validator/issues/25) | Reproducible benchmark, funnel figure and validation dossier | all | +| Done | [#24](https://github.com/NingyuSUN/bioai-evidence-validator/issues/24) | Calibrated triage and audit of auto-admitted records (tooling, routing evidence and a dry run; the expert audit needs #23) | #23 | +| Done | [#25](https://github.com/NingyuSUN/bioai-evidence-validator/issues/25) | Reproducible benchmark, funnel figure and validation dossier | all | ## Definition of done for 0.8.0 diff --git a/docs/ENGINEERING.md b/docs/ENGINEERING.md index 119e3e9..e448d3b 100644 --- a/docs/ENGINEERING.md +++ b/docs/ENGINEERING.md @@ -244,6 +244,25 @@ but cannot withdraw verified evidence against its answer. The literature benchma (`evaluation/llm_benchmark/agent_loop.py`) showed why: before that rule, an agent whose record held a misquote and a verified quote against its decision dropped the latter and was admitted. +### Routing evidence and audits + +A record reaches a person because it cannot be verified (BEV006, BEV015, BEV018, BEV020), because the use requires +a person (BEV008–BEV013, BEV021), or because a risk signal flags it. The first two are admission policy. A risk +signal routes by default only when it is shown to predict a wrong answer. +`evaluation/llm_benchmark/risk_signals.py` measures this on the first answers of the committed runs: + +| Signal | Default | Wrong when flagged vs otherwise | Verdict | +|---|---|---|---| +| BEV004 contradicting evidence line | on | held-out 82% vs 33%; external 70% vs 38%; literature pilot 91% vs 16% | shown | +| BEV022 semantic cue | opt-in | literature pilot 4% vs 9% | not shown | +| BEV025 reference conflict (ASCT+B) | opt-in | single-cell pilot 57% vs 40% | not shown | +| BEV026 definition contradicted | opt-in, experimental | external 47% vs 41% | not shown | + +What routing lets through is measured by an audit (`audit`, `bioevidence review audit-sample` and `audit-score`): +a seeded random sample of auto-admitted records, optionally mixed with blind controls from the other routes, and +each route's error rate with an exact one-sided (Clopper–Pearson) upper bound. The procedure is in +[the gold-standard workflow](GOLD_STANDARD.md#auditing-what-is-admitted-automatically). + ## Semantic checks Grounding proves that a quote is real; it cannot tell whether the quote supports the claim. diff --git a/docs/GOLD_STANDARD.md b/docs/GOLD_STANDARD.md index 5f4f66e..3c6b477 100644 --- a/docs/GOLD_STANDARD.md +++ b/docs/GOLD_STANDARD.md @@ -113,6 +113,32 @@ rather than discarding difficult cases silently. Report reviewer agreement befor adjudication (`bioevidence review agreement`). Any statistical intervals should respect concept-level dependence; do not treat multiple aliases or mutations of one concept as independent samples. +## Auditing what is admitted automatically + +A reference set measures a method once. In use, the question is how many auto-admitted records are wrong. +An audit answers it on a seeded random sample: + +1. Export one prediction per record and use, in the format `bioevidence review score` reads, with the route the + method took: `admitted`, `review_required` or `rejected`. +2. `bioevidence review audit-sample` draws the sample. `--target 0.01` sizes it so that, if no audited record + is wrong, the auto-admitted error rate is below 1% at 95% confidence: 299 records (149 for 2%, 59 for 5%, + 29 for 10%). `--controls` mixes in records from the other routes, in shuffled order, so auditors cannot tell + a record's route. The manifest records the seed, each route's size and the sampled hashes; keep it from the + auditors. +3. Auditors label the sheet as in this protocol, blind to the route, the model and each other, and record + `minutes_spent`. +4. `bioevidence review audit-score` reports each route's error rate with an exact one-sided upper bound. An + auto-admitted record is an error when the auditor would not admit it for the use or its claim is incorrect; + a record on another route, when the auditor would have admitted it as it was. + +Report the bound with the rate: "2 errors in 300 audited, below 2.1% at 95% confidence". Audit again when the +model, prompts, profile or references change. One auditor per record is the default for audits +(`--min-reviewers 1`); disagreements between several still need adjudication. + +Records reach a person for one of three reasons: they cannot be verified, the use requires a person, or a risk +signal flags them. A risk signal belongs in routing only when it is shown to predict errors; the evidence for +each is in [`results/risk-signals`](../evaluation/llm_benchmark/results/risk-signals/summary.md). + An improvement against this reference can support a claim about the defined curation task. It does not establish breed-genotype membership, universal biological truth, or clinical validity. diff --git a/docs/VALIDATION_DOSSIER.md b/docs/VALIDATION_DOSSIER.md index cd62f71..fdb9ada 100644 --- a/docs/VALIDATION_DOSSIER.md +++ b/docs/VALIDATION_DOSSIER.md @@ -15,7 +15,8 @@ State: release 0.8.0, October 2026. |---|---|---| | Admitted records carry no invalid identifier or citation | 0 of 721 admitted AI answers across four experiments (upper bound 0.5%); models alone produced such errors in 10–76% of answers | **Established** | | The feedback loop corrects fixable errors without hiding evidence | Single-cell held-out: 57 of 63 answers with identifier or marker errors corrected and admitted; verified evidence cannot be withdrawn | **Established** | -| Admitted records are semantically correct | 19.6% [15.2, 24.9] of admitted single-cell annotations disagree with the authors' term. Part of this is label noise. No expert audit yet | **Not established** | +| Admitted records are semantically correct | 19.6% [15.2, 24.9] of admitted single-cell annotations disagree with the authors' term. Part of this is label noise. The audit tool is ready and was checked in a dry run; no expert audit yet | **Not established** | +| Default routing rests on signals that predict errors | BEV004 answers were wrong 82% vs 33% (held-out) and 70% vs 38% (external); the opt-in signals (semantic cues, ASCT+B, definitions) showed no such effect and stay off | **Established** for BEV004 | | Routing load is workable | 7.6% [5.0, 11.3] of single-cell annotations sent to a person (held-out) | **Plausible**, one run | | A new semantic check (Cell Ontology definitions) helps | No better than chance on external data, and harmful when fed back | **Refuted**; kept experimental | @@ -75,7 +76,9 @@ consequence**, the impact of a wrong decision. - protocols fixed before results are seen; - negative controls; - full reproducibility; - - a measure of what still gets through. That last one needs an expert audit, which has not been done. + - a measure of what still gets through. That last one needs an expert audit, which has not been done. The + procedure and its statistics are ready (`bioevidence review audit-sample`, `audit-score`), and a dry run + with the authors' labels as a stand-in recovered the census rate (19.6%) inside its interval. ## 4. Credibility assessment plan @@ -173,7 +176,7 @@ Fed back in the loop, its findings turned 15 correct answers into wrong ones. Co - **One run was resumed.** The external run hit a time limit at 321 of 486 episodes and was resumed. Episodes are independent, and none was repeated. - **Two planned evaluations have not been done.** The literature held-out set was built but not run, and no expert - review or audit of admitted records has been done (#23, #24). + review or audit of admitted records has been done (#23). The audit tooling is in place (#24). ## 7. Adequacy for the context of use @@ -199,7 +202,8 @@ Proposed acceptance criteria for the next evaluation, to be fixed before it is r protocol, and compare per model. The integrity result should not move; the semantic result can. - **Reference change.** A new Cell Ontology, HGNC or ClinVar release changes the pinned snapshots. Re-run `tools/reproduce.py` and the affected splits, and keep the earlier snapshots for comparison. -- **Monitoring in use.** Audit a random sample of admitted records on a schedule (#24), and report the audited +- **Monitoring in use.** Audit a random sample of admitted records on a schedule with + `bioevidence review audit-sample` and `audit-score` (#24), and report the audited error rate with its interval next to the figures above. ## Sources diff --git a/docs/index.md b/docs/index.md index e8335bc..e36b138 100644 --- a/docs/index.md +++ b/docs/index.md @@ -105,9 +105,10 @@ Across all 276 held-out annotations, stage by stage: | PubMed rate limits, server errors | `Library.fetch` | Exponential backoff on 429 and 5xx, a response cache, a 100 MB download cap | | Wrong paper, retracted paper, quote not in the paper | `LiteratureGrounder` (BEV016, BEV017, BEV019) | The reason goes back to the model; at most three submissions | | ID of another term, obsolete term, gene alias | `OntologyGrounder`, `GeneGrounder` (BEV017, BEV023, BEV024) | The reason goes back, naming the correct identifier | -| The model's own evidence contradicts its answer | BEV004 | Sent to a person; not fed back, so the model is never asked to hide it | +| The model's own evidence contradicts its answer | BEV004 | Sent to a person; not fed back, so the model is never asked to hide it. Flagged answers were wrong 82% of the time vs 33% ([routing evidence](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/evaluation/llm_benchmark/results/risk-signals/summary.md)) | | A revision drops evidence that was verified | `feedback.carry` | The verified evidence is carried into the revision | | Still not admitted after the last round | `feedback.revise` | Sent to a person | +| Admitted but wrong | A seeded random audit (`bioevidence review audit-sample`, `audit-score`) | Each route's error rate with an exact upper bound; a [dry run](https://github.com/NingyuSUN/bioai-evidence-validator/blob/main/evaluation/llm_benchmark/results/celltype-audit/summary.md) shows the held-out route misses a 5% target | | A long batch is interrupted | One file per episode | A rerun skips finished episodes | How the checks work, rule by rule: [engineering contract](ENGINEERING.md). diff --git a/evaluation/gold_standard/README.md b/evaluation/gold_standard/README.md index df0f681..bf3a558 100644 --- a/evaluation/gold_standard/README.md +++ b/evaluation/gold_standard/README.md @@ -32,6 +32,8 @@ zero reviewed cases. An editor changing its status does not complete human revie | `bioevidence review adjudication-sheet annotations.csv --output adjudications.csv` | Blank rows for every disagreement; never overwrites an existing file | | `bioevidence review score --annotations … --adjudications … --predictions …` | Resolves final labels and scores predictions on the test split; rejects missing or hash-mismatched rows | | `bioevidence review freeze --manifest … --annotations … --adjudications …` | Refuses unresolved cases; records counts, reference type and file hashes | +| `bioevidence review audit-sample --predictions … --method … --use … --target 0.01 --seed … --profile-id … --output-dir …` | A seeded random sample of auto-admitted records as a blank annotation sheet, sized so that no error found bounds their error rate below the target; `--controls` mixes in records from the other routes. The manifest of routes stays with you | +| `bioevidence review audit-score --manifest … --annotations …` | Per route: share of records, audited error rate with Wilson and exact one-sided bounds, and share of expert time from `minutes_spent` | A worked, domain-specific kit is in [`../clinvar_review/`](../clinvar_review/README.md). The VBO replay command still evaluates only its source-derived reference set and controlled faults. diff --git a/evaluation/llm_benchmark/README.md b/evaluation/llm_benchmark/README.md index 7548ec8..c8141da 100644 --- a/evaluation/llm_benchmark/README.md +++ b/evaluation/llm_benchmark/README.md @@ -448,6 +448,40 @@ Which need experts: evidence that is mixed, depends on the endpoint, or where th every model disagree; these are exactly the cases the expert review (#23) is for. Five errors on two tasks are too few for rates; the held-out test set would measure them with intervals. +## Routing evidence and the audit (#24) + +**Which signals should send a record to a person?** Only those shown to predict a wrong answer. +[`results/risk-signals`](results/risk-signals/summary.md) compares, on first answers, how often an answer was wrong +when a signal fired and when it did not: + +| Signal | Default | Split | Wrong when flagged | Wrong otherwise | Risk ratio (95% CI) | +|---|---|---|---:|---:|---:| +| BEV004 contradicting evidence | on | single-cell held-out | 14/17 (82%) | 86/259 (33%) | 2.48 (1.87–3.28) | +| | | single-cell external | 38/54 (70%) | 163/425 (38%) | 1.83 (1.49–2.27) | +| | | literature agent pilot | 10/11 (91%) | 11/69 (16%) | 5.7 (3.21–10.12) | +| BEV022 semantic cue | opt-in | literature pilot | 1/24 (4%) | 4/44 (9%) | 0.46 (0.05–3.87) | +| BEV025 ASCT+B conflict | opt-in | single-cell pilot | 13/23 (57%) | 46/115 (40%) | 1.41 (0.93–2.16) | +| BEV026 definition contradicted | opt-in | single-cell external | 41/88 (47%) | 160/391 (41%) | 1.14 (0.88–1.47) | + +The only risk signal on by default, BEV004, is shown on every split. The opt-in ones are not, and stay off by +default. The criterion was set after the runs: the risk ratio's interval lies above 1 on a held-out or external +split. + +**What gets through?** An audit of a seeded random sample of auto-admitted records answers that with an upper +bound. No expert has audited these records yet, so +[`results/celltype-audit`](results/celltype-audit/summary.md) is a dry run on the held-out split. It runs the +whole procedure with a stand-in auditor, the dataset authors' labels: +1. export the loop's routes; +2. `bioevidence review audit-sample --target 0.05 --controls 10` draws 59 auto-admitted records and 10 others; +3. a [packet](results/celltype-audit/audit_packet.md) shows each record without its route, model or the + authors' label; +4. `bioevidence review audit-score` scores it. + +The stand-in found 14 errors in 59, so the auto-admitted error rate is below 34.6% at 95% confidence (23.7%, +Wilson 14.7–36.0%). Because the authors' labels cover every record, the census can be checked: 50 of 255 +(19.6%), inside the interval. The route does not meet a 5% target. A deployment would learn from such an audit +that these annotations need review, or a better model, before research summaries rely on them. + ## Run ```bash diff --git a/evaluation/llm_benchmark/audit_celltype.py b/evaluation/llm_benchmark/audit_celltype.py new file mode 100644 index 0000000..09b2ef8 --- /dev/null +++ b/evaluation/llm_benchmark/audit_celltype.py @@ -0,0 +1,191 @@ +"""Audit dry run (#24): what an audit of the auto-admitted single-cell annotations would report. + +No expert has audited these records. This dry run runs the audit procedure end to end on the held-out results +(protocol 2: 6 models x 46 clusters) and checks its statistics against a census: + +1. export each episode's final route (auto-admitted, to a person) as a predictions file, one record per model and + cluster, under an opaque case id; +2. draw a seeded random sample with `bioevidence review audit-sample`: enough auto-admitted records to bound their + error rate below 5% if none is wrong, and 10 records from the other route mixed in; +3. write the auditors' packet: each record's cluster, marker table and claim, in shuffled order, without its route, + model or the authors' label; +4. fill the audit sheet with a stand-in auditor, the dataset authors' own cluster labels: a claim is correct when it is + the authors' term, a coarser or a finer one, and admissible when it is correct and passes the identifier checks; +5. score it with `bioevidence review audit-score`, and compare with the census of all auto-admitted records, which + the authors' labels allow here and a real audit does not. + +The stand-in is not an expert, and it records no time, so the expert-time columns stay empty. + + python evaluation/llm_benchmark/audit_celltype.py --output +""" +from __future__ import annotations + +import argparse +import hashlib +import json +import sys +from collections import Counter +from pathlib import Path + +ROOT = Path(__file__).resolve().parent +sys.path.insert(0, str(ROOT)) +import celltype_loop as cl # noqa: E402 + +from bioevidence_validator import audit, cli, review # noqa: E402 + +RESULTS = ROOT / "results" +METHOD, USE, PROFILE = "llm-feedback-loop", "research_summary", "singlecell-celltype" +SEED, TARGET, CONTROLS = 20261006, 0.05, 10 +STAND_IN = "stand-in: the dataset authors' cluster labels (not an expert)" + + +def sha256_text(text: str) -> str: + return hashlib.sha256(text.encode("utf-8")).hexdigest() + + +def export(case: cl.Case) -> dict[str, dict]: + """case id -> the episode's final record, its route and what the stand-in needs.""" + profile = sha256_text((cl.CASE / "profile.yaml").read_text(encoding="utf-8").replace("\r\n", "\n")) + cases = {} + for row in cl.load_jsonl(RESULTS / "celltype-test" / "episodes.jsonl"): + if not row["attempts"]: + continue # no annotation, no record + task = case.tasks[row["task_id"]] + got = [c["answer"] for c in row["calls"] if c["answer"] and c["answer"]["decision"] == "annotate"] + answer = got[len(row["attempts"]) - 1] + record = cl.record(case, task, answer) + case_id = "sc-" + sha256_text(f"{row['task_id']}|{row['model']}")[:12] + cases[case_id] = {"task": task, "answer": answer, "route": row["attempts"][-1]["status"], "profile": profile, + "record_sha256": sha256_text(json.dumps(record, sort_keys=True, separators=(",", ":")))} + return cases + + +def packet(case: cl.Case, cases: dict[str, dict], sheet: list[dict[str, str]]) -> str: + lines = ["# Audit packet: single-cell annotations for research summaries", "", + "For each record, decide from the cluster's markers:", "", + "- `mapping_label`: is the claimed cell type right for this cluster? `correct`, `incorrect` or `uncertain`. " + "A correct but coarser type is `correct`.", + "- `admission_label`: should the record be admitted for a research summary as it stands? `admitted`, " + "`review_required` or `rejected`.", "", + "Write both, with a short rationale, your reviewer id and qualification, the time and `minutes_spent`, in " + "`audit_sheet.csv`. Records are in random order; their route, model and the authors' label are withheld.", + ""] + for number, row in enumerate(sheet, start=1): + c = cases[row["case_id"]] + task, answer = c["task"], c["answer"] + markers = ", ".join(f"{m['gene']} ({float(m['logfc']):.1f}; {float(m['pct_in']):.0%} vs {float(m['pct_out']):.0%})" + for m in task["markers"]) + cited = ", ".join(f"{m['gene']} ({m['stance'].replace('_', ' ')})" for m in answer["markers"]) or "none" + lines += [f"## {number}. `{row['case_id']}`", "", + f"- Cluster `{task['cluster']}` of a {task['species']} {task['tissue']} dataset ({task['assay']})", + f"- Top markers (log fold change; share of cells in vs out of the cluster): {markers}", + f"- **Claim:** {answer['cell_type_id']} *{answer['cell_type_label']}*", + f"- Cited markers: {cited}", ""] + return "\n".join(lines) + + +def stand_in(case: cl.Case, truth: dict[str, dict], c: dict) -> dict[str, str]: + """The labels the authors' term implies for one record.""" + kind = cl.outcome(case, truth[c["task"]["task_id"]]["term"], c["answer"]) + issues = cl.problems(case, c["task"], c["answer"]) + correct = kind in ("exact", "coarser", "finer") + t = truth[c["task"]["task_id"]] + return {"mapping_label": "correct" if correct else "incorrect", + "mapping_rationale": f"authors' term {t['term']} ({t['label']}); the claim is {kind}", + "admission_label": "admitted" if correct and not issues else "rejected", + "admission_rationale": "; ".join(issues) or ("passes the identifier checks" if correct else "wrong cell type")} + + +def main(argv: list[str] | None = None) -> int: + options = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + options.add_argument("--output", type=Path, required=True, help="A new directory") + args = options.parse_args(argv) + out: Path = args.output + if out.exists() and any(out.iterdir()): + raise SystemExit(f"{out} is not empty; refusing to overwrite an audit") + out.mkdir(parents=True, exist_ok=True) + case = cl.Case() + truth = {t["task_id"]: t for t in cl.load_jsonl(cl.SOURCES / "truth.jsonl")} + cases = export(case) + + predictions = [{"case_id": k, "requested_use": USE, "record_sha256": c["record_sha256"], + "profile_sha256": c["profile"], "method": METHOD, "predicted_status": c["route"]} + for k, c in sorted(cases.items())] + review.write_csv(out / "predictions.csv", predictions, review.PREDICTION_COLUMNS) + sampled = out / "sample" + status = cli.main(["review", "audit-sample", "--predictions", str(out / "predictions.csv"), "--method", METHOD, + "--use", USE, "--target", str(TARGET), "--controls", str(CONTROLS), "--seed", str(SEED), + "--profile-id", PROFILE, "--output-dir", str(sampled)]) + if status: + return status + sheet = list(review._read_csv(sampled / "audit_sheet.csv", review.ANNOTATION_COLUMNS, + review.OPTIONAL_ANNOTATION_COLUMNS)) + (out / "audit_packet.md").write_text(packet(case, cases, sheet), encoding="utf-8", newline="\n") + + filled = [{**row, "annotation_id": f"stand-in:{row['case_id']}", "reviewer_id": "stand-in-authors-labels", + "reviewer_qualification": STAND_IN, "annotated_at": "2026-10-06T00:00:00Z", + "evidence_refs_json": json.dumps([f"authors-label:{cases[row['case_id']]['task']['task_id']}"]), + **stand_in(case, truth, cases[row["case_id"]])} for row in sheet] + review.write_csv(out / "stand_in_annotations.csv", filled, + review.ANNOTATION_COLUMNS + review.OPTIONAL_ANNOTATION_COLUMNS) + status = cli.main(["review", "audit-score", "--manifest", str(sampled / "audit_manifest.json"), "--annotations", + str(out / "stand_in_annotations.csv"), "--output", str(out / "audit_report.json")]) + if status: + return status + + report = json.loads((out / "audit_report.json").read_text(encoding="utf-8")) + census: dict[str, Counter] = {} + for c in cases.values(): + labels = stand_in(case, truth, c) + census.setdefault(c["route"], Counter())["errors" if audit._is_error(c["route"], labels) else "fine"] += 1 + checks = {} + for route, s in report["routes"].items(): + errors, n = census[route]["errors"], sum(census[route].values()) + rate = errors / n + checks[route] = {"census_errors": errors, "census_records": n, "census_error_rate": round(rate, 4), + "audit_error_rate": s["error_rate"], "audit_wilson_95": s["error_rate_wilson_95"], + "audit_upper_bound": s["error_rate_upper_bound"], + "census_within_wilson": s["error_rate_wilson_95"][0] <= rate <= s["error_rate_wilson_95"][1], + "census_below_upper_bound": rate <= s["error_rate_upper_bound"]} + summary = {"dry_run": "audit of the single-cell held-out results with a stand-in auditor (#24)", + "stand_in": STAND_IN, "seed": SEED, "target": TARGET, "controls": CONTROLS, "audit": report, + "census": checks, "episodes_sha256": hashlib.sha256((RESULTS / "celltype-test" / "episodes.jsonl") + .read_bytes().replace(b"\r\n", b"\n")).hexdigest()} + (out / "summary.json").write_text(json.dumps(summary, indent=2, sort_keys=True) + "\n", encoding="utf-8", + newline="\n") + (out / "summary.md").write_text(render(summary), encoding="utf-8", newline="\n") + print(render(summary)) + return 0 + + +def render(summary: dict) -> str: + report = summary["audit"] + admitted = report["routes"]["admitted"] + check = summary["census"]["admitted"] + lines = ["# Audit dry run: single-cell held-out annotations", "", + f"**Not an expert audit.** The auditor is a stand-in: {summary['stand_in'].removeprefix('stand-in: ')}. " + "It shows what the audit procedure reports, and checks its statistics against a census.", "", + audit.render_audit(report).split("\n", 2)[2].rstrip(), "", + "## Against the census", "", + "The authors' labels cover every record, so here the error rate of the whole route is known:", "", + "| Route | Census error rate | Audit estimate | Wilson 95% | Upper bound (95%) | Census inside |", + "|---|---:|---:|---|---:|---|"] + for route, c in summary["census"].items(): + low, high = c["audit_wilson_95"] + inside = "yes" if c["census_within_wilson"] and c["census_below_upper_bound"] else "no" + lines.append(f"| {audit.ROUTES[route]} | {c['census_errors']}/{c['census_records']} " + f"({100 * c['census_error_rate']:.1f}%) | {100 * c['audit_error_rate']:.1f}% | " + f"{100 * low:.1f}%–{100 * high:.1f}% | {100 * c['audit_upper_bound']:.1f}% | {inside} |") + verdict = "meets" if report["admitted_error_upper_bound"] < summary["target"] else "does not meet" + lines += ["", f"The audit was sized to show an auto-admitted error rate below {100 * summary['target']:.0f}% if " + f"none of {admitted['audited']} sampled records was wrong. It found {admitted['errors']}, so this route " + f"{verdict} a {100 * summary['target']:.0f}% target: the census rate is " + f"{100 * check['census_error_rate']:.1f}%. An audit is how a deployment would learn that these " + "annotations need review, or a better model, before research summaries rely on them.", "", + f"Seed {summary['seed']}; `sample/audit_manifest.json` holds the routes and stays with the maintainer. " + "Expert time is not measured: the stand-in records none."] + return "\n".join(lines) + "\n" + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/evaluation/llm_benchmark/results/celltype-audit/audit_packet.md b/evaluation/llm_benchmark/results/celltype-audit/audit_packet.md new file mode 100644 index 0000000..0f9af2e --- /dev/null +++ b/evaluation/llm_benchmark/results/celltype-audit/audit_packet.md @@ -0,0 +1,491 @@ +# Audit packet: single-cell annotations for research summaries + +For each record, decide from the cluster's markers: + +- `mapping_label`: is the claimed cell type right for this cluster? `correct`, `incorrect` or `uncertain`. A correct but coarser type is `correct`. +- `admission_label`: should the record be admitted for a research summary as it stands? `admitted`, `review_required` or `rejected`. + +Write both, with a short rationale, your reviewer id and qualification, the time and `minutes_spent`, in `audit_sheet.csv`. Records are in random order; their route, model and the authors' label are withheld. + +## 1. `sc-ba0d303f9838` + +- Cluster `c08` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): CCL2 (2.8; 90% vs 26%), RGS5 (2.6; 93% vs 8%), C11orf96 (2.5; 98% vs 15%), IGFBP5 (2.5; 98% vs 35%), IFI27 (2.4; 100% vs 31%), TM4SF1 (2.4; 100% vs 28%), JUNB (2.3; 100% vs 64%), TMSB4X (2.3; 100% vs 75%), SOCS3 (2.2; 98% vs 37%), CEBPD (2.1; 98% vs 51%), IGFBP7 (2.1; 100% vs 68%), HSPA1A (2.0; 98% vs 68%), TIMP3 (1.9; 98% vs 40%), MT1A (1.9; 85% vs 16%), DNAJB1 (1.9; 98% vs 62%), CDKN1A (1.9; 98% vs 46%), IRF1 (1.8; 95% vs 37%), NEAT1 (1.8; 100% vs 70%), JUN (1.8; 100% vs 58%), APOLD1 (1.7; 93% vs 15%) +- **Claim:** CL:0000669 *pericyte* +- Cited markers: RGS5 (supports), IGFBP5 (supports), IGFBP7 (supports), TM4SF1 (supports), TIMP3 (supports) + +## 2. `sc-2138b2030213` + +- Cluster `c06` of a Homo sapiens pancreas dataset (CEL-seq2) +- Top markers (log fold change; share of cells in vs out of the cluster): PPY (6.4; 100% vs 54%), SCG2 (2.0; 100% vs 90%), PEG10 (2.0; 100% vs 88%), ID2 (1.8; 98% vs 73%), PAX6 (1.7; 100% vs 79%), ETV1 (1.6; 99% vs 41%), AQP3 (1.6; 90% vs 46%), MEIS2 (1.5; 98% vs 75%), ITM2C (1.3; 99% vs 92%), PCSK2 (1.2; 100% vs 94%), CHGB (1.2; 100% vs 93%), THSD7A (1.2; 86% vs 14%), PAM (1.2; 100% vs 94%), GAD2 (1.2; 98% vs 70%), NEUROD1 (1.2; 99% vs 80%), SERTM1 (1.1; 84% vs 9%), ABCC9 (1.1; 96% vs 71%), UCHL1 (1.1; 97% vs 72%), ABCC8 (1.1; 99% vs 73%), SLC6A4 (1.1; 91% vs 56%) +- **Claim:** CL:0002275 *pancreatic PP cell* +- Cited markers: PPY (supports), PAX6 (supports), PCSK2 (supports), SCG2 (supports), CHGB (supports), PAM (supports), NEUROD1 (supports), SERTM1 (supports) + +## 3. `sc-8d801d5b870e` + +- Cluster `c13` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): DEFB1 (3.3; 100% vs 38%), SLC12A3 (3.2; 98% vs 4%), TMEM52B (2.7; 100% vs 13%), KNG1 (2.4; 99% vs 16%), MT-ND1 (2.3; 100% vs 94%), WNK1 (2.1; 98% vs 28%), ATP1B1 (2.0; 100% vs 59%), SPP1 (2.0; 98% vs 57%), MT-CO1 (2.0; 100% vs 99%), MT-ND2 (1.9; 100% vs 97%), MT-CO3 (1.9; 100% vs 99%), MT-ATP6 (1.8; 100% vs 97%), MT-ND4 (1.8; 100% vs 97%), MT-CYB (1.8; 100% vs 96%), CA12 (1.7; 98% vs 30%), MALAT1 (1.7; 100% vs 84%), MT-CO2 (1.7; 100% vs 98%), MT-ND5 (1.7; 100% vs 87%), ATP1A1 (1.6; 99% vs 55%), MT-ND3 (1.6; 100% vs 97%) +- **Claim:** CL:1000849 *kidney distal convoluted tubule epithelial cell* +- Cited markers: SLC12A3 (supports), DEFB1 (supports), TMEM52B (supports), KNG1 (supports), WNK1 (supports), CA12 (supports), ATP1B1 (supports), ATP1A1 (supports) + +## 4. `sc-2c1b4befc878` + +- Cluster `c08` of a Homo sapiens duodenum dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): CXCL14 (4.5; 90% vs 7%), ADAMDEC1 (4.4; 90% vs 7%), LUM (3.7; 90% vs 4%), DCN (3.7; 90% vs 4%), IGFBP7 (3.7; 99% vs 4%), CFD (3.5; 90% vs 11%), GPX3 (3.3; 94% vs 3%), RARRES2 (3.0; 91% vs 4%), APOE (2.9; 84% vs 5%), IFITM3 (2.8; 100% vs 25%), CALD1 (2.8; 99% vs 3%), VIM (2.6; 100% vs 34%), COL3A1 (2.6; 90% vs 3%), TCF21 (2.4; 90% vs 1%), COL1A2 (2.4; 90% vs 2%), CXCL6 (2.3; 88% vs 1%), MFAP4 (2.3; 89% vs 1%), C1S (2.1; 89% vs 1%), ADH1B (2.1; 90% vs 9%), CTSC (2.1; 88% vs 32%) +- **Claim:** CL:0000057 *fibroblast* +- Cited markers: CXCL14 (supports), LUM (supports), DCN (supports), CFD (supports), RARRES2 (supports), CALD1 (supports), COL3A1 (supports), TCF21 (supports), COL1A2 (supports), MFAP4 (supports), C1S (supports), ADAMDEC1 (supports) + +## 5. `sc-2eff7bb75fad` + +- Cluster `c02` of a Homo sapiens caudate lobe of liver dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): HBB (7.7; 100% vs 31%), HBA1 (7.0; 100% vs 14%), HBA2 (6.8; 100% vs 18%), HBD (4.2; 92% vs 3%), AHSP (3.2; 98% vs 2%), CA1 (3.0; 91% vs 2%), SLC25A37 (2.9; 96% vs 13%), HBM (2.8; 89% vs 1%), ALAS2 (2.4; 96% vs 1%), SLC4A1 (1.9; 85% vs 1%), BLVRB (1.8; 92% vs 43%), SNCA (1.8; 91% vs 2%), GYPA (1.7; 79% vs 0%), SLC25A39 (1.6; 91% vs 17%), BNIP3L (1.6; 89% vs 18%), BSG (1.5; 94% vs 29%), YBX3 (1.5; 87% vs 16%), GLRX5 (1.4; 87% vs 22%), GYPC (1.4; 92% vs 25%), HMBS (1.4; 81% vs 6%) +- **Claim:** CL:0000232 *erythrocyte* +- Cited markers: HBB (supports), HBA1 (supports), HBA2 (supports), HBD (supports), AHSP (supports), CA1 (supports), SLC25A37 (supports), HBM (supports), ALAS2 (supports), SLC4A1 (supports), BLVRB (supports), SNCA (supports), GYPA (supports), SLC25A39 (supports), BNIP3L (supports), BSG (supports), YBX3 (supports), GLRX5 (supports), GYPC (supports), HMBS (supports) + +## 6. `sc-47865daed309` + +- Cluster `c03` of a Homo sapiens duodenum dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): JCHAIN (5.6; 100% vs 99%), IGLL5 (3.0; 100% vs 89%), MZB1 (3.0; 100% vs 19%), SSR4 (2.7; 100% vs 90%), DERL3 (2.6; 99% vs 15%), HERPUD1 (2.3; 100% vs 66%), FKBP11 (2.3; 99% vs 34%), CD79A (2.2; 98% vs 6%), XBP1 (2.2; 99% vs 74%), TNFRSF17 (2.1; 98% vs 6%), SEC11C (2.0; 99% vs 57%), SPCS2 (1.9; 100% vs 72%), DNAJB9 (1.8; 96% vs 30%), ITM2C (1.8; 92% vs 25%), CYBA (1.8; 100% vs 82%), UBE2J1 (1.7; 99% vs 54%), SRGN (1.7; 96% vs 15%), LGALS1 (1.6; 94% vs 17%), JUN (1.6; 99% vs 94%), SPCS1 (1.5; 100% vs 79%) +- **Claim:** CL:0000786 *plasma cell* +- Cited markers: MZB1 (supports), TNFRSF17 (supports), JCHAIN (supports), XBP1 (supports), DERL3 (supports), FKBP11 (supports), CD79A (supports), IGLL5 (supports) + +## 7. `sc-43489d3f91fb` + +- Cluster `c09` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): WFDC2 (2.9; 97% vs 21%), KRT7 (2.8; 100% vs 13%), IGFBP5 (2.0; 97% vs 35%), CA12 (1.9; 97% vs 33%), TMEM213 (1.9; 97% vs 21%), CD9 (1.8; 100% vs 50%), ATP6V1G3 (1.8; 90% vs 10%), MALAT1 (1.8; 100% vs 85%), ATP6V0B (1.7; 100% vs 52%), LGALS3 (1.7; 100% vs 39%), ATP6V0D2 (1.6; 94% vs 13%), TSPAN8 (1.6; 87% vs 8%), ATP6V1B1 (1.6; 94% vs 20%), KRT8 (1.4; 94% vs 38%), HEPACAM2 (1.4; 90% vs 8%), ATP6AP2 (1.4; 97% vs 44%), S100A2 (1.3; 81% vs 23%), ID3 (1.3; 87% vs 33%), IL18 (1.3; 90% vs 12%), SLC26A4 (1.3; 77% vs 1%) +- **Claim:** CL:0002306 *epithelial cell of proximal tubule* +- Cited markers: WFDC2 (supports), KRT7 (supports), KRT8 (supports), CA12 (supports), SLC26A4 (supports), ATP6V1G3 (supports), ATP6V0B (supports), ATP6V0D2 (supports), ATP6V1B1 (supports), ATP6AP2 (supports), IGFBP5 (supports), TSPAN8 (supports), HEPACAM2 (supports) + +## 8. `sc-457686ee8397` + +- Cluster `c07` of a Homo sapiens zone of skin dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): CCL21 (4.2; 97% vs 20%), TFF3 (3.6; 100% vs 8%), CLDN5 (2.2; 92% vs 13%), TFPI (2.0; 94% vs 19%), GNG11 (1.9; 93% vs 30%), MMRN1 (1.8; 85% vs 2%), IGFBP7 (1.7; 99% vs 65%), TM4SF1 (1.6; 91% vs 35%), ADIRF (1.6; 99% vs 74%), SNCG (1.5; 79% vs 13%), ECSCR (1.4; 79% vs 10%), RAMP2 (1.4; 81% vs 17%), FABP4 (1.3; 65% vs 5%), PROX1 (1.3; 72% vs 3%), HLA-E (1.3; 95% vs 62%), LYVE1 (1.3; 67% vs 1%), PPFIBP1 (1.2; 72% vs 10%), SOX4 (1.2; 74% vs 34%), CAVIN2 (1.2; 68% vs 4%), ARL4A (1.1; 78% vs 28%) +- **Claim:** CL:0002138 *endothelial cell of lymphatic vessel* +- Cited markers: CCL21 (supports), TFF3 (supports), PROX1 (supports), LYVE1 (supports), MMRN1 (supports), CLDN5 (supports), TFPI (supports), GNG11 (supports), ECSCR (supports), RAMP2 (supports), CAVIN2 (supports), PPFIBP1 (supports), SNCG (supports), TM4SF1 (supports), FABP4 (supports) + +## 9. `sc-1bb81802ec74` + +- Cluster `c08` of a Homo sapiens zone of skin dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): TM4SF1 (3.0; 98% vs 29%), IFI27 (2.6; 98% vs 29%), SELE (2.1; 72% vs 10%), ACKR1 (2.1; 81% vs 14%), SPARCL1 (2.0; 91% vs 20%), CD74 (1.9; 99% vs 61%), TSC22D1 (1.9; 90% vs 40%), SPRY1 (1.7; 81% vs 20%), AQP1 (1.5; 86% vs 28%), SAT1 (1.5; 97% vs 78%), HLA-E (1.4; 94% vs 60%), ADIRF (1.4; 98% vs 72%), IFITM3 (1.4; 98% vs 66%), GNG11 (1.4; 81% vs 26%), RCAN1 (1.3; 71% vs 17%), SOCS3 (1.3; 83% vs 41%), PCAT19 (1.3; 73% vs 5%), CLDN5 (1.3; 69% vs 9%), RAMP2 (1.2; 72% vs 13%), MCTP1 (1.2; 70% vs 6%) +- **Claim:** CL:0002543 *vein endothelial cell* +- Cited markers: ACKR1 (supports), SELE (supports), CLDN5 (supports), RAMP2 (supports), GNG11 (supports), AQP1 (supports), SPARCL1 (supports), TM4SF1 (supports) + +## 10. `sc-01ab2950a674` + +- Cluster `c07` of a Homo sapiens caudate lobe of liver dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): GNLY (2.0; 57% vs 7%), NKG7 (1.8; 65% vs 23%), TMSB10 (1.3; 97% vs 76%), IGKC (1.2; 60% vs 34%), GZMB (1.2; 53% vs 4%), TUBA1B (1.2; 63% vs 31%), TMSB4X (1.2; 98% vs 89%), PTMA (1.1; 99% vs 91%), MALAT1 (1.1; 99% vs 97%), CORO1A (1.1; 75% vs 29%), CCL5 (1.1; 54% vs 20%), FGFBP2 (1.0; 45% vs 2%), KLRD1 (1.0; 55% vs 14%), PFN1 (1.0; 95% vs 79%), H4C3 (1.0; 64% vs 48%), HMGB2 (1.0; 57% vs 23%), HMGN2 (1.0; 69% vs 40%), CYBA (1.0; 84% vs 44%), CCL4 (1.0; 57% vs 22%), HLA-B (1.0; 96% vs 86%) +- **Claim:** CL:0000623 *natural killer cell* +- Cited markers: GNLY (supports), NKG7 (supports), IGKC (contradicts), GZMB (supports), CCL5 (supports), FGFBP2 (supports), KLRD1 (supports), CCL4 (supports) + +## 11. `sc-9bd026197028` + +- Cluster `c11` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): FXYD4 (2.8; 83% vs 8%), AQP2 (2.7; 80% vs 4%), AQP3 (2.1; 80% vs 12%), MALAT1 (2.1; 100% vs 84%), MT-CO1 (1.6; 100% vs 99%), CLU (1.6; 84% vs 39%), WFDC2 (1.6; 78% vs 20%), TACSTD2 (1.6; 78% vs 21%), GDF15 (1.5; 80% vs 29%), MT-CO2 (1.4; 100% vs 98%), S100A6 (1.4; 92% vs 67%), MT-CO3 (1.4; 100% vs 99%), CDH16 (1.3; 86% vs 32%), KRT19 (1.3; 68% vs 10%), MT-ND4 (1.3; 100% vs 98%), KRT18 (1.3; 83% vs 38%), MT-CYB (1.3; 100% vs 97%), ELF3 (1.2; 76% vs 25%), MT-ATP6 (1.2; 100% vs 97%), NEAT1 (1.2; 97% vs 69%) +- **Claim:** CL:1001431 *kidney collecting duct principal cell* +- Cited markers: AQP2 (supports), AQP3 (supports), FXYD4 (supports), CDH16 (supports), KRT18 (supports), KRT19 (supports), TACSTD2 (supports), WFDC2 (supports) + +## 12. `sc-66d221735c3f` + +- Cluster `c03` of a Homo sapiens caudate lobe of liver dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): C1QA (3.6; 100% vs 14%), C1QB (3.5; 99% vs 11%), CD163 (2.9; 99% vs 10%), HLA-DRA (2.7; 98% vs 22%), CTSB (2.6; 99% vs 38%), MARCO (2.6; 86% vs 4%), FTL (2.5; 100% vs 100%), SLC40A1 (2.5; 94% vs 32%), C1QC (2.5; 95% vs 8%), MS4A6A (2.4; 95% vs 17%), CD74 (2.3; 98% vs 40%), MS4A7 (2.3; 91% vs 8%), CTSS (2.3; 96% vs 24%), AIF1 (2.2; 95% vs 15%), TYROBP (2.2; 97% vs 28%), NPC2 (2.2; 97% vs 42%), SAT1 (2.2; 100% vs 72%), GPX1 (2.1; 99% vs 66%), HMOX1 (2.1; 85% vs 9%), FCER1G (2.1; 94% vs 24%) +- **Claim:** CL:0000091 *Kupffer cell* +- Cited markers: C1QA (supports), C1QB (supports), C1QC (supports), CD163 (supports), MARCO (supports), CD74 (supports), HLA-DRA (supports), MS4A6A (supports), MS4A7 (supports), AIF1 (supports), CTSB (supports), CTSS (supports), TYROBP (supports), FCER1G (supports), HMOX1 (supports) + +## 13. `sc-81eace82fee7` + +- Cluster `c05` of a Homo sapiens duodenum dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): LYZ (2.9; 98% vs 25%), MT1G (2.4; 98% vs 77%), GUCA2B (2.2; 61% vs 25%), BEST4 (2.2; 74% vs 1%), SPIB (2.0; 84% vs 4%), MT2A (1.9; 96% vs 82%), MT1E (1.8; 99% vs 70%), MT1H (1.8; 82% vs 52%), CFTR (1.7; 92% vs 28%), MT1X (1.6; 96% vs 60%), GSN (1.5; 82% vs 34%), CPA2 (1.4; 75% vs 3%), CA7 (1.4; 68% vs 1%), KRT20 (1.4; 83% vs 44%), CTSE (1.2; 90% vs 29%), GUCA2A (1.2; 50% vs 2%), FXYD3 (1.2; 94% vs 49%), S100A6 (1.2; 100% vs 95%), MT1M (1.1; 74% vs 27%), STARD10 (1.1; 94% vs 32%) +- **Claim:** CL:0000584 *enterocyte* +- Cited markers: BEST4 (supports), SPIB (supports), CFTR (supports), CA7 (supports), CPA2 (supports), GUCA2B (supports), GUCA2A (supports), LYZ (supports) + +## 14. `sc-c218ebe7c140` + +- Cluster `c05` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): TAGLN (3.7; 95% vs 8%), MYL9 (3.1; 97% vs 16%), RGS5 (3.1; 84% vs 6%), ACTA2 (2.9; 89% vs 5%), TPM2 (2.8; 94% vs 8%), CALD1 (2.8; 99% vs 34%), ADIRF (2.7; 95% vs 56%), C11orf96 (2.6; 93% vs 13%), MGP (2.4; 96% vs 20%), LGALS1 (2.3; 97% vs 33%), IGFBP7 (2.3; 97% vs 68%), VIM (2.2; 94% vs 34%), MALAT1 (2.1; 100% vs 84%), SPARCL1 (2.1; 87% vs 13%), MYH11 (1.9; 72% vs 2%), JUNB (1.9; 94% vs 63%), BGN (1.9; 84% vs 5%), SOD3 (1.8; 81% vs 19%), FLNA (1.8; 86% vs 24%), MT2A (1.7; 99% vs 90%) +- **Claim:** CL:0000192 *smooth muscle cell* +- Cited markers: TAGLN (supports), MYL9 (supports), RGS5 (supports), ACTA2 (supports), TPM2 (supports), CALD1 (supports), MYH11 (supports) + +## 15. `sc-4937458f58f8` + +- Cluster `c08` of a Homo sapiens caudate lobe of liver dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): S100A9 (2.9; 89% vs 13%), TYROBP (2.7; 99% vs 23%), LYZ (2.7; 90% vs 8%), S100A8 (2.4; 68% vs 11%), AIF1 (2.3; 92% vs 10%), HLA-DRA (2.2; 86% vs 18%), S100A6 (2.1; 87% vs 31%), CTSS (2.1; 92% vs 20%), CD74 (2.1; 92% vs 37%), S100A4 (2.1; 86% vs 28%), TMSB10 (2.0; 100% vs 75%), FCER1G (1.9; 92% vs 19%), VIM (1.9; 92% vs 33%), GPX1 (1.9; 98% vs 64%), C1QA (1.8; 63% vs 12%), SAT1 (1.8; 97% vs 71%), S100A11 (1.8; 90% vs 23%), LST1 (1.7; 83% vs 8%), CST3 (1.7; 98% vs 63%), TMSB4X (1.7; 100% vs 88%) +- **Claim:** CL:0000235 *macrophage* +- Cited markers: S100A9 (supports), TYROBP (supports), LYZ (supports), S100A8 (supports), AIF1 (supports), HLA-DRA (supports), CTSS (supports), CD74 (supports), C1QA (supports) + +## 16. `sc-f6b33d5ab782` + +- Cluster `c05` of a Homo sapiens zone of skin dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): KRT1 (3.3; 97% vs 30%), DMKN (3.2; 98% vs 32%), KRT10 (3.1; 98% vs 49%), SFN (3.1; 98% vs 45%), LGALS7B (2.9; 97% vs 14%), PERP (2.7; 98% vs 42%), KRTDAP (2.7; 89% vs 12%), LY6D (2.6; 96% vs 23%), S100A14 (2.5; 98% vs 29%), LYPD3 (2.1; 91% vs 18%), AQP3 (2.1; 94% vs 24%), DSP (2.1; 94% vs 19%), TACSTD2 (1.7; 91% vs 24%), SERPINB5 (1.7; 90% vs 16%), KRT14 (1.6; 95% vs 67%), CCL27 (1.5; 86% vs 12%), RNF144B (1.5; 79% vs 19%), MIR205HG (1.4; 88% vs 21%), CLDN1 (1.4; 81% vs 11%), FXYD3 (1.4; 86% vs 17%) +- **Claim:** CL:0000312 *keratinocyte* +- Cited markers: KRT1 (supports), KRT10 (supports), KRTDAP (supports), DMKN (supports), SFN (supports), LGALS7B (supports), PERP (supports), LY6D (supports), S100A14 (supports), DSP (supports), KRT14 (supports), CCL27 (supports), AQP3 (supports), SERPINB5 (supports), MIR205HG (supports), CLDN1 (supports) + +## 17. `sc-7c5b8a9c5536` + +- Cluster `c03` of a Homo sapiens zone of skin dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): HBB (7.2; 91% vs 27%), HBA2 (6.2; 89% vs 12%), HBA1 (5.5; 87% vs 5%), HBD (2.2; 79% vs 0%), SLC25A37 (1.0; 60% vs 25%), SLC25A39 (1.0; 53% vs 11%), HBM (0.6; 28% vs 0%), GYPC (0.6; 49% vs 26%), FBXO7 (0.6; 37% vs 14%), FKBP8 (0.4; 46% vs 33%), SNCA (0.4; 25% vs 8%), UBBP4 (0.1; 28% vs 28%) +- **Claim:** CL:0000232 *erythrocyte* +- Cited markers: HBB (supports), HBA2 (supports), HBA1 (supports), HBD (supports), HBM (supports), SLC25A37 (supports), SLC25A39 (supports), GYPC (supports), FBXO7 (supports), SNCA (supports) + +## 18. `sc-dc6252fb95cb` + +- Cluster `c12` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): TMEM213 (2.3; 98% vs 16%), SLC4A1 (2.0; 96% vs 4%), SPINK1 (1.9; 72% vs 4%), ATP6V0D2 (1.9; 98% vs 8%), ATP6V1G3 (1.9; 96% vs 4%), ATP6AP2 (1.8; 98% vs 41%), ADGRF5 (1.8; 96% vs 20%), C12orf75 (1.7; 97% vs 36%), LGALS3 (1.7; 96% vs 35%), SLC26A7 (1.7; 95% vs 4%), MT-CO1 (1.6; 100% vs 99%), ATP6V0B (1.6; 97% vs 49%), RHCG (1.6; 92% vs 9%), MALAT1 (1.6; 100% vs 84%), ATP6V1B1 (1.6; 96% vs 15%), BSG (1.6; 99% vs 55%), MT-ND1 (1.6; 100% vs 94%), CKB (1.6; 96% vs 34%), MT-CO3 (1.5; 100% vs 99%), CA12 (1.5; 98% vs 28%) +- **Claim:** CL:0005011 *renal alpha-intercalated cell* +- Cited markers: SLC4A1 (supports), SLC26A7 (supports), ATP6V0D2 (supports), ATP6V1G3 (supports), ATP6AP2 (supports), ATP6V0B (supports), ATP6V1B1 (supports), RHCG (supports), CA12 (supports) + +## 19. `sc-fef6c4941874` + +- Cluster `c04` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): CD74 (2.3; 100% vs 62%), CXCR4 (2.2; 85% vs 16%), CD52 (2.1; 85% vs 7%), IGKC (2.0; 59% vs 4%), MALAT1 (2.0; 100% vs 85%), RPS29 (1.9; 100% vs 97%), RPS27 (1.9; 100% vs 99%), HLA-DRA (1.9; 91% vs 43%), RPS2 (1.7; 100% vs 91%), RPLP2 (1.7; 100% vs 98%), RPS21 (1.6; 97% vs 89%), RPS8 (1.6; 100% vs 95%), RPS15A (1.6; 100% vs 96%), RPL18A (1.5; 100% vs 96%), RPS19 (1.5; 100% vs 96%), CD79A (1.5; 59% vs 0%), LTB (1.5; 62% vs 4%), TMSB4X (1.5; 97% vs 75%), RPS10 (1.5; 94% vs 76%), NBEAL1 (1.5; 85% vs 58%) +- **Claim:** CL:0000236 *B cell* +- Cited markers: CD79A (supports), IGKC (supports), CD74 (supports), HLA-DRA (supports), CD52 (supports), CXCR4 (supports), LTB (supports) + +## 20. `sc-ae30f1136126` + +- Cluster `c12` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): LTB (2.2; 99% vs 57%), IL7R (2.1; 98% vs 35%), IL4I1 (1.6; 84% vs 5%), LST1 (1.5; 86% vs 17%), CCR6 (1.4; 86% vs 4%), VIM (1.4; 96% vs 66%), S100A4 (1.4; 92% vs 62%), RORC (1.4; 82% vs 6%), S100A6 (1.4; 89% vs 44%), TNFAIP3 (1.3; 96% vs 60%), AQP3 (1.3; 88% vs 30%), RGS1 (1.2; 92% vs 54%), ITM2C (1.1; 85% vs 34%), JAML (1.0; 76% vs 17%), DDIT4 (1.0; 90% vs 60%), CXCR4 (1.0; 86% vs 50%), KIT (1.0; 66% vs 4%), TMIGD2 (0.9; 84% vs 46%), IL23R (0.9; 70% vs 8%), ZFP36L1 (0.9; 98% vs 88%) +- **Claim:** CL:0000899 *T-helper 17 cell* +- Cited markers: RORC (supports), CCR6 (supports), IL23R (supports), IL7R (supports), LTB (supports), LST1 (supports), TNFAIP3 (supports), RGS1 (supports), ITM2C (supports), JAML (supports) + +## 21. `sc-3440a71c316e` + +- Cluster `c04` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): LTB (1.7; 97% vs 56%), IL32 (1.7; 98% vs 61%), CD52 (1.3; 97% vs 67%), CD27 (1.3; 88% vs 36%), S100A4 (1.1; 88% vs 61%), CD2 (1.0; 84% vs 34%), FOXP3 (1.0; 59% vs 0%), ID3 (1.0; 76% vs 31%), CD3E (1.0; 98% vs 66%), CTLA4 (1.0; 61% vs 1%), SELL (1.0; 78% vs 34%), TNFRSF1B (1.0; 75% vs 25%), PIM2 (0.9; 80% vs 37%), SIRPG (0.9; 72% vs 21%), FXYD5 (0.9; 92% vs 68%), TRAC (0.8; 81% vs 38%), DGKA (0.8; 75% vs 31%), ITM2A (0.8; 79% vs 44%), CD5 (0.7; 67% vs 20%), LEF1 (0.7; 72% vs 35%) +- **Claim:** CL:0000815 *regulatory T cell* +- Cited markers: FOXP3 (supports), CTLA4 (supports), CD27 (supports), CD3E (supports), TRAC (supports), CD2 (supports), CD52 (supports), SELL (supports), LEF1 (supports), IL32 (supports) + +## 22. `sc-bffa8ceffd0e` + +- Cluster `c10` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): NKG7 (2.1; 100% vs 67%), GNLY (2.0; 80% vs 23%), GZMB (1.9; 94% vs 23%), SPON2 (1.8; 92% vs 21%), FCGR3A (1.7; 94% vs 26%), PRF1 (1.7; 99% vs 54%), FGFBP2 (1.6; 85% vs 14%), GZMH (1.6; 88% vs 23%), CST7 (1.6; 99% vs 53%), CCL5 (1.6; 88% vs 47%), CLIC3 (1.5; 96% vs 47%), TYROBP (1.5; 99% vs 50%), FCER1G (1.4; 99% vs 57%), CCL4 (1.4; 95% vs 48%), GZMA (1.3; 100% vs 65%), ADGRG1 (1.3; 85% vs 22%), CX3CR1 (1.3; 80% vs 14%), TBX21 (1.3; 86% vs 29%), CTSW (1.2; 99% vs 72%), MYOM2 (1.2; 58% vs 11%) +- **Claim:** CL:0000623 *natural killer cell* +- Cited markers: NKG7 (supports), GNLY (supports), GZMB (supports), PRF1 (supports), FGFBP2 (supports), GZMH (supports), FCGR3A (supports), TYROBP (supports), FCER1G (supports), CX3CR1 (supports), TBX21 (supports), CCL5 (supports) + +## 23. `sc-b855c9140617` + +- Cluster `c07` of a Homo sapiens caudate lobe of liver dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): GNLY (2.0; 57% vs 7%), NKG7 (1.8; 65% vs 23%), TMSB10 (1.3; 97% vs 76%), IGKC (1.2; 60% vs 34%), GZMB (1.2; 53% vs 4%), TUBA1B (1.2; 63% vs 31%), TMSB4X (1.2; 98% vs 89%), PTMA (1.1; 99% vs 91%), MALAT1 (1.1; 99% vs 97%), CORO1A (1.1; 75% vs 29%), CCL5 (1.1; 54% vs 20%), FGFBP2 (1.0; 45% vs 2%), KLRD1 (1.0; 55% vs 14%), PFN1 (1.0; 95% vs 79%), H4C3 (1.0; 64% vs 48%), HMGB2 (1.0; 57% vs 23%), HMGN2 (1.0; 69% vs 40%), CYBA (1.0; 84% vs 44%), CCL4 (1.0; 57% vs 22%), HLA-B (1.0; 96% vs 86%) +- **Claim:** CL:0000623 *natural killer cell* +- Cited markers: GNLY (supports), NKG7 (supports), GZMB (supports), FGFBP2 (supports), KLRD1 (supports), CCL5 (supports), CCL4 (supports), CORO1A (supports), IGKC (contradicts) + +## 24. `sc-00a40a840007` + +- Cluster `c07` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): LTB (1.3; 97% vs 57%), CXCR4 (1.3; 89% vs 50%), IL32 (1.3; 93% vs 62%), GZMK (1.2; 77% vs 28%), IL7R (1.1; 85% vs 36%), TRDV2 (1.1; 44% vs 4%), S100A4 (1.1; 91% vs 62%), TNFAIP3 (1.0; 86% vs 60%), DUSP2 (0.9; 93% vs 80%), SPOCK2 (0.9; 86% vs 44%), KLRB1 (0.9; 94% vs 73%), NFKBIA (0.9; 94% vs 81%), TC2N (0.9; 74% vs 28%), S100A6 (0.9; 81% vs 45%), CD3E (0.9; 92% vs 67%), IL23R (0.8; 57% vs 9%), CD69 (0.8; 98% vs 92%), FOSB (0.7; 90% vs 74%), DUSP1 (0.7; 99% vs 93%), SLC4A10 (0.7; 50% vs 5%) +- **Claim:** CL:0000940 *mucosal invariant T cell* +- Cited markers: SLC4A10 (supports), KLRB1 (supports), IL23R (supports), GZMK (supports), IL7R (supports), CD3E (supports), LTB (supports), CXCR4 (supports), IL32 (supports), S100A4 (supports), CD69 (supports), TRDV2 (contradicts) + +## 25. `sc-4d6f03e0ceaf` + +- Cluster `c08` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): CCL2 (2.8; 90% vs 26%), RGS5 (2.6; 93% vs 8%), C11orf96 (2.5; 98% vs 15%), IGFBP5 (2.5; 98% vs 35%), IFI27 (2.4; 100% vs 31%), TM4SF1 (2.4; 100% vs 28%), JUNB (2.3; 100% vs 64%), TMSB4X (2.3; 100% vs 75%), SOCS3 (2.2; 98% vs 37%), CEBPD (2.1; 98% vs 51%), IGFBP7 (2.1; 100% vs 68%), HSPA1A (2.0; 98% vs 68%), TIMP3 (1.9; 98% vs 40%), MT1A (1.9; 85% vs 16%), DNAJB1 (1.9; 98% vs 62%), CDKN1A (1.9; 98% vs 46%), IRF1 (1.8; 95% vs 37%), NEAT1 (1.8; 100% vs 70%), JUN (1.8; 100% vs 58%), APOLD1 (1.7; 93% vs 15%) +- **Claim:** CL:0000669 *pericyte* +- Cited markers: RGS5 (supports), IGFBP5 (supports), IGFBP7 (supports), TIMP3 (supports) + +## 26. `sc-c09fd2a58e63` + +- Cluster `c05` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): TAGLN (3.7; 95% vs 8%), MYL9 (3.1; 97% vs 16%), RGS5 (3.1; 84% vs 6%), ACTA2 (2.9; 89% vs 5%), TPM2 (2.8; 94% vs 8%), CALD1 (2.8; 99% vs 34%), ADIRF (2.7; 95% vs 56%), C11orf96 (2.6; 93% vs 13%), MGP (2.4; 96% vs 20%), LGALS1 (2.3; 97% vs 33%), IGFBP7 (2.3; 97% vs 68%), VIM (2.2; 94% vs 34%), MALAT1 (2.1; 100% vs 84%), SPARCL1 (2.1; 87% vs 13%), MYH11 (1.9; 72% vs 2%), JUNB (1.9; 94% vs 63%), BGN (1.9; 84% vs 5%), SOD3 (1.8; 81% vs 19%), FLNA (1.8; 86% vs 24%), MT2A (1.7; 99% vs 90%) +- **Claim:** CL:0000359 *vascular smooth muscle cell* +- Cited markers: ACTA2 (supports), MYH11 (supports), TAGLN (supports), MYL9 (supports), TPM2 (supports), CALD1 (supports), RGS5 (supports), MGP (supports), SPARCL1 (supports), BGN (supports), SOD3 (supports) + +## 27. `sc-7c94fc7a0196` + +- Cluster `c07` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): GNLY (4.0; 95% vs 2%), NKG7 (3.7; 96% vs 2%), GZMB (2.6; 84% vs 2%), CCL5 (2.4; 80% vs 4%), CCL4 (2.4; 77% vs 14%), TMSB4X (2.2; 100% vs 74%), MALAT1 (2.1; 100% vs 84%), CST7 (2.1; 79% vs 2%), B2M (1.8; 100% vs 96%), AREG (1.7; 64% vs 15%), HLA-C (1.7; 96% vs 71%), TMSB10 (1.7; 100% vs 88%), FGFBP2 (1.6; 62% vs 0%), SRGN (1.6; 80% vs 33%), KLRD1 (1.6; 59% vs 2%), CREM (1.5; 72% vs 36%), HLA-A (1.5; 99% vs 77%), HCST (1.5; 66% vs 8%), CYBA (1.5; 86% vs 55%), GZMA (1.5; 59% vs 1%) +- **Claim:** CL:0000623 *natural killer cell* +- Cited markers: GNLY (supports), NKG7 (supports), GZMB (supports), GZMA (supports), CST7 (supports), FGFBP2 (supports), KLRD1 (supports), CCL5 (supports), CCL4 (supports), HCST (supports), SRGN (supports) + +## 28. `sc-6cf84b7db7a2` + +- Cluster `c03` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): CD8B (1.6; 89% vs 8%), LTB (1.4; 100% vs 55%), LEF1 (1.3; 97% vs 32%), CD52 (1.2; 100% vs 66%), CCR7 (1.2; 90% vs 23%), TCF7 (1.1; 90% vs 33%), CD8A (1.1; 88% vs 29%), SELL (1.1; 86% vs 32%), RGS10 (1.1; 94% vs 35%), LINC02446 (1.1; 63% vs 3%), LDHB (1.0; 99% vs 73%), AIF1 (1.0; 93% vs 35%), NOSIP (1.0; 94% vs 52%), IL7R (1.0; 84% vs 33%), CD3E (1.0; 99% vs 65%), LRRN3 (0.9; 72% vs 13%), CD27 (0.9; 91% vs 34%), RPS8 (0.8; 100% vs 100%), S100B (0.8; 52% vs 12%), RPLP0 (0.8; 100% vs 98%) +- **Claim:** CL:0000900 *naive thymus-derived CD8-positive, alpha-beta T cell* +- Cited markers: CD8B (supports), CD8A (supports), CCR7 (supports), LEF1 (supports), TCF7 (supports), SELL (supports), IL7R (supports), CD27 (supports), CD3E (supports), LTB (supports) + +## 29. `sc-92d87108bf26` + +- Cluster `c03` of a Homo sapiens caudate lobe of liver dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): C1QA (3.6; 100% vs 14%), C1QB (3.5; 99% vs 11%), CD163 (2.9; 99% vs 10%), HLA-DRA (2.7; 98% vs 22%), CTSB (2.6; 99% vs 38%), MARCO (2.6; 86% vs 4%), FTL (2.5; 100% vs 100%), SLC40A1 (2.5; 94% vs 32%), C1QC (2.5; 95% vs 8%), MS4A6A (2.4; 95% vs 17%), CD74 (2.3; 98% vs 40%), MS4A7 (2.3; 91% vs 8%), CTSS (2.3; 96% vs 24%), AIF1 (2.2; 95% vs 15%), TYROBP (2.2; 97% vs 28%), NPC2 (2.2; 97% vs 42%), SAT1 (2.2; 100% vs 72%), GPX1 (2.1; 99% vs 66%), HMOX1 (2.1; 85% vs 9%), FCER1G (2.1; 94% vs 24%) +- **Claim:** CL:0000091 *Kupffer cell* +- Cited markers: C1QA (supports), C1QB (supports), C1QC (supports), CD163 (supports), MARCO (supports), SLC40A1 (supports), MS4A6A (supports), MS4A7 (supports), FCER1G (supports), TYROBP (supports), HLA-DRA (supports), CD74 (supports) + +## 30. `sc-26e57da30416` + +- Cluster `c07` of a Homo sapiens duodenum dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): TMSB4X (2.9; 100% vs 100%), CD7 (2.7; 92% vs 11%), KLRB1 (2.4; 92% vs 10%), CD52 (2.3; 97% vs 18%), CORO1A (2.1; 96% vs 20%), GZMA (2.0; 59% vs 7%), CD3E (2.0; 97% vs 7%), EVL (1.9; 97% vs 21%), PTPRCAP (1.9; 97% vs 28%), ACTB (1.9; 100% vs 99%), HCST (1.8; 93% vs 12%), ARHGDIB (1.7; 95% vs 23%), NKG7 (1.6; 80% vs 5%), CD247 (1.6; 76% vs 2%), ALOX5AP (1.6; 89% vs 11%), VIM (1.5; 92% vs 34%), COTL1 (1.5; 91% vs 44%), RAC2 (1.5; 87% vs 21%), ACAP1 (1.5; 88% vs 14%), LCK (1.5; 86% vs 5%) +- **Claim:** CL:0000084 *T cell* +- Cited markers: CD3E (supports), CD247 (supports), LCK (supports), CD7 (supports), KLRB1 (supports), GZMA (supports), NKG7 (supports), CD52 (supports), CORO1A (supports), PTPRCAP (supports), HCST (supports), ACAP1 (supports) + +## 31. `sc-2258691ddec9` + +- Cluster `c06` of a Homo sapiens pancreas dataset (CEL-seq2) +- Top markers (log fold change; share of cells in vs out of the cluster): PPY (6.4; 100% vs 54%), SCG2 (2.0; 100% vs 90%), PEG10 (2.0; 100% vs 88%), ID2 (1.8; 98% vs 73%), PAX6 (1.7; 100% vs 79%), ETV1 (1.6; 99% vs 41%), AQP3 (1.6; 90% vs 46%), MEIS2 (1.5; 98% vs 75%), ITM2C (1.3; 99% vs 92%), PCSK2 (1.2; 100% vs 94%), CHGB (1.2; 100% vs 93%), THSD7A (1.2; 86% vs 14%), PAM (1.2; 100% vs 94%), GAD2 (1.2; 98% vs 70%), NEUROD1 (1.2; 99% vs 80%), SERTM1 (1.1; 84% vs 9%), ABCC9 (1.1; 96% vs 71%), UCHL1 (1.1; 97% vs 72%), ABCC8 (1.1; 99% vs 73%), SLC6A4 (1.1; 91% vs 56%) +- **Claim:** CL:0002275 *pancreatic PP cell* +- Cited markers: PPY (supports), PAX6 (supports), NEUROD1 (supports), SCG2 (supports), CHGB (supports), PCSK2 (supports), PAM (supports) + +## 32. `sc-00ad508b5979` + +- Cluster `c07` of a Homo sapiens zone of skin dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): CCL21 (4.2; 97% vs 20%), TFF3 (3.6; 100% vs 8%), CLDN5 (2.2; 92% vs 13%), TFPI (2.0; 94% vs 19%), GNG11 (1.9; 93% vs 30%), MMRN1 (1.8; 85% vs 2%), IGFBP7 (1.7; 99% vs 65%), TM4SF1 (1.6; 91% vs 35%), ADIRF (1.6; 99% vs 74%), SNCG (1.5; 79% vs 13%), ECSCR (1.4; 79% vs 10%), RAMP2 (1.4; 81% vs 17%), FABP4 (1.3; 65% vs 5%), PROX1 (1.3; 72% vs 3%), HLA-E (1.3; 95% vs 62%), LYVE1 (1.3; 67% vs 1%), PPFIBP1 (1.2; 72% vs 10%), SOX4 (1.2; 74% vs 34%), CAVIN2 (1.2; 68% vs 4%), ARL4A (1.1; 78% vs 28%) +- **Claim:** CL:0002138 *lymphatic endothelial cell* +- Cited markers: CCL21 (supports), CLDN5 (supports), MMRN1 (supports), PROX1 (supports), LYVE1 (supports), CAVIN2 (supports), RAMP2 (supports), ECSCR (supports), GNG11 (supports), TFPI (supports), TM4SF1 (supports), FABP4 (supports) + +## 33. `sc-a271a8ebacae` + +- Cluster `c11` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): FXYD4 (2.8; 83% vs 8%), AQP2 (2.7; 80% vs 4%), AQP3 (2.1; 80% vs 12%), MALAT1 (2.1; 100% vs 84%), MT-CO1 (1.6; 100% vs 99%), CLU (1.6; 84% vs 39%), WFDC2 (1.6; 78% vs 20%), TACSTD2 (1.6; 78% vs 21%), GDF15 (1.5; 80% vs 29%), MT-CO2 (1.4; 100% vs 98%), S100A6 (1.4; 92% vs 67%), MT-CO3 (1.4; 100% vs 99%), CDH16 (1.3; 86% vs 32%), KRT19 (1.3; 68% vs 10%), MT-ND4 (1.3; 100% vs 98%), KRT18 (1.3; 83% vs 38%), MT-CYB (1.3; 100% vs 97%), ELF3 (1.2; 76% vs 25%), MT-ATP6 (1.2; 100% vs 97%), NEAT1 (1.2; 97% vs 69%) +- **Claim:** CL:1001431 *kidney collecting duct principal cell* +- Cited markers: AQP2 (supports), FXYD4 (supports), AQP3 (supports), CDH16 (supports), KRT19 (supports), KRT18 (supports) + +## 34. `sc-f0f257d1b718` + +- Cluster `c05` of a Homo sapiens zone of skin dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): KRT1 (3.3; 97% vs 30%), DMKN (3.2; 98% vs 32%), KRT10 (3.1; 98% vs 49%), SFN (3.1; 98% vs 45%), LGALS7B (2.9; 97% vs 14%), PERP (2.7; 98% vs 42%), KRTDAP (2.7; 89% vs 12%), LY6D (2.6; 96% vs 23%), S100A14 (2.5; 98% vs 29%), LYPD3 (2.1; 91% vs 18%), AQP3 (2.1; 94% vs 24%), DSP (2.1; 94% vs 19%), TACSTD2 (1.7; 91% vs 24%), SERPINB5 (1.7; 90% vs 16%), KRT14 (1.6; 95% vs 67%), CCL27 (1.5; 86% vs 12%), RNF144B (1.5; 79% vs 19%), MIR205HG (1.4; 88% vs 21%), CLDN1 (1.4; 81% vs 11%), FXYD3 (1.4; 86% vs 17%) +- **Claim:** CL:0000312 *keratinocyte* +- Cited markers: KRT1 (supports), DMKN (supports), KRT10 (supports), SFN (supports), LGALS7B (supports), PERP (supports), KRTDAP (supports), LY6D (supports), S100A14 (supports), LYPD3 (supports), AQP3 (supports), DSP (supports), TACSTD2 (supports), SERPINB5 (supports), CCL27 (supports), CLDN1 (supports) + +## 35. `sc-2d06b761bbe6` + +- Cluster `c03` of a Homo sapiens caudate lobe of liver dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): C1QA (3.6; 100% vs 14%), C1QB (3.5; 99% vs 11%), CD163 (2.9; 99% vs 10%), HLA-DRA (2.7; 98% vs 22%), CTSB (2.6; 99% vs 38%), MARCO (2.6; 86% vs 4%), FTL (2.5; 100% vs 100%), SLC40A1 (2.5; 94% vs 32%), C1QC (2.5; 95% vs 8%), MS4A6A (2.4; 95% vs 17%), CD74 (2.3; 98% vs 40%), MS4A7 (2.3; 91% vs 8%), CTSS (2.3; 96% vs 24%), AIF1 (2.2; 95% vs 15%), TYROBP (2.2; 97% vs 28%), NPC2 (2.2; 97% vs 42%), SAT1 (2.2; 100% vs 72%), GPX1 (2.1; 99% vs 66%), HMOX1 (2.1; 85% vs 9%), FCER1G (2.1; 94% vs 24%) +- **Claim:** CL:0000091 *Kupffer cell* +- Cited markers: C1QA (supports), C1QB (supports), C1QC (supports), CD163 (supports), MARCO (supports), SLC40A1 (supports), HMOX1 (supports), MS4A7 (supports), MS4A6A (supports), HLA-DRA (supports), CD74 (supports), AIF1 (supports), CTSS (supports), TYROBP (supports), FCER1G (supports), FTL (supports) + +## 36. `sc-74db4609c7a9` + +- Cluster `c07` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): LTB (1.3; 97% vs 57%), CXCR4 (1.3; 89% vs 50%), IL32 (1.3; 93% vs 62%), GZMK (1.2; 77% vs 28%), IL7R (1.1; 85% vs 36%), TRDV2 (1.1; 44% vs 4%), S100A4 (1.1; 91% vs 62%), TNFAIP3 (1.0; 86% vs 60%), DUSP2 (0.9; 93% vs 80%), SPOCK2 (0.9; 86% vs 44%), KLRB1 (0.9; 94% vs 73%), NFKBIA (0.9; 94% vs 81%), TC2N (0.9; 74% vs 28%), S100A6 (0.9; 81% vs 45%), CD3E (0.9; 92% vs 67%), IL23R (0.8; 57% vs 9%), CD69 (0.8; 98% vs 92%), FOSB (0.7; 90% vs 74%), DUSP1 (0.7; 99% vs 93%), SLC4A10 (0.7; 50% vs 5%) +- **Claim:** CL:0000798 *gamma-delta T cell* +- Cited markers: TRDV2 (supports), CD3E (supports), IL7R (supports), KLRB1 (supports), IL23R (supports), SLC4A10 (supports), GZMK (supports), IL32 (supports), TC2N (supports) + +## 37. `sc-554e5e45515f` + +- Cluster `c15` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): UMOD (3.4; 85% vs 18%), SLC12A1 (3.0; 96% vs 5%), DEFB1 (2.4; 99% vs 34%), KNG1 (1.9; 83% vs 12%), ATP1B1 (1.8; 99% vs 56%), MT-ND1 (1.7; 100% vs 94%), ATP1A1 (1.7; 98% vs 52%), MALAT1 (1.7; 100% vs 83%), CD24 (1.6; 98% vs 41%), S100A6 (1.5; 100% vs 64%), MT-ND2 (1.5; 100% vs 96%), MT-CYB (1.5; 100% vs 96%), MT-CO1 (1.5; 100% vs 99%), MT-CO3 (1.5; 100% vs 99%), MT-ND4 (1.5; 100% vs 97%), MT-ATP6 (1.4; 100% vs 97%), MT-ND3 (1.4; 100% vs 97%), MT-CO2 (1.3; 100% vs 98%), CA12 (1.3; 92% vs 26%), S100A2 (1.2; 70% vs 18%) +- **Claim:** CL:0002306 *thick ascending limb cell* +- Cited markers: UMOD (supports), SLC12A1 (supports), ATP1A1 (supports), ATP1B1 (supports), CA12 (supports) + +## 38. `sc-c7c4c2041f88` + +- Cluster `c01` of a Homo sapiens zone of skin dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): CD52 (2.2; 87% vs 16%), IL32 (2.1; 86% vs 28%), SRGN (1.9; 94% vs 43%), CD3D (1.6; 71% vs 5%), ARHGDIB (1.6; 80% vs 23%), CXCR4 (1.5; 68% vs 16%), CD69 (1.4; 63% vs 6%), CREM (1.4; 76% vs 41%), PTPRCAP (1.4; 65% vs 4%), SARAF (1.3; 77% vs 36%), DUSP2 (1.3; 67% vs 21%), HCST (1.3; 67% vs 15%), PTPRC (1.2; 63% vs 9%), LTB (1.2; 54% vs 6%), SAMSN1 (1.2; 60% vs 10%), FXYD5 (1.1; 76% vs 39%), ALOX5AP (1.1; 59% vs 12%), STK17B (1.1; 63% vs 17%), S100A4 (1.1; 97% vs 86%), RGCC (1.1; 67% vs 34%) +- **Claim:** CL:0000084 *T cell* +- Cited markers: CD3D (supports), IL32 (supports), CD52 (supports), LTB (supports), PTPRC (supports), PTPRCAP (supports), CD69 (supports), CXCR4 (supports), HCST (supports), ARHGDIB (supports), SRGN (supports), DUSP2 (supports), STK17B (supports), SAMSN1 (supports) + +## 39. `sc-b93c03bf06d2` + +- Cluster `c06` of a Homo sapiens duodenum dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): HLA-DPB1 (4.1; 100% vs 22%), HLA-DPA1 (4.0; 100% vs 30%), C1QA (3.6; 89% vs 13%), C1QC (3.5; 87% vs 8%), C1QB (3.5; 87% vs 10%), TYROBP (3.4; 95% vs 9%), HLA-DRA (3.2; 100% vs 70%), TMSB4X (3.1; 100% vs 100%), HLA-DQA1 (3.0; 99% vs 4%), CD74 (3.0; 100% vs 95%), HLA-DRB1 (3.0; 100% vs 64%), CST3 (2.7; 100% vs 84%), AIF1 (2.7; 100% vs 5%), LYZ (2.6; 97% vs 23%), MS4A6A (2.6; 96% vs 3%), HLA-DQB1 (2.5; 99% vs 24%), SELENOP (2.3; 81% vs 64%), NPC2 (2.0; 100% vs 76%), FCER1G (2.0; 93% vs 5%), HLA-DMA (2.0; 96% vs 37%) +- **Claim:** CL:0000235 *macrophage* +- Cited markers: HLA-DPB1 (supports), HLA-DPA1 (supports), C1QA (supports), C1QC (supports), C1QB (supports), TYROBP (supports), HLA-DRA (supports), TMSB4X (supports), HLA-DQA1 (supports), CD74 (supports), HLA-DRB1 (supports), CST3 (supports), AIF1 (supports), LYZ (supports), MS4A6A (supports), HLA-DQB1 (supports), SELENOP (supports), NPC2 (supports), FCER1G (supports), HLA-DMA (supports) + +## 40. `sc-0ed3dae30cb0` + +- Cluster `c10` of a Homo sapiens zone of skin dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): KRT14 (3.7; 100% vs 67%), S100A2 (3.2; 99% vs 46%), KRT5 (3.0; 99% vs 34%), SFN (2.9; 100% vs 46%), DST (2.3; 97% vs 30%), MIR205HG (2.2; 96% vs 21%), SERPINB2 (2.2; 89% vs 18%), PERP (2.1; 98% vs 43%), AQP3 (1.9; 92% vs 25%), S100A14 (1.8; 92% vs 30%), TACSTD2 (1.8; 92% vs 25%), LGALS7B (1.8; 90% vs 16%), ERRFI1 (1.6; 87% vs 27%), SERPINB5 (1.6; 88% vs 18%), AREG (1.5; 68% vs 8%), RND3 (1.5; 88% vs 37%), FGFBP1 (1.5; 79% vs 17%), CCL27 (1.4; 82% vs 13%), KRT15 (1.4; 70% vs 11%), ACTG1 (1.3; 100% vs 86%) +- **Claim:** CL:0000312 *keratinocyte* +- Cited markers: KRT14 (supports), S100A2 (supports), KRT5 (supports), SFN (supports), DST (supports), MIR205HG (supports), SERPINB2 (supports), PERP (supports), AQP3 (supports), S100A14 (supports), TACSTD2 (supports), LGALS7B (supports), ERRFI1 (supports), SERPINB5 (supports), AREG (supports), RND3 (supports), FGFBP1 (supports), CCL27 (supports), KRT15 (supports) + +## 41. `sc-835f4e01b78e` + +- Cluster `c03` of a Homo sapiens pancreas dataset (CEL-seq2) +- Top markers (log fold change; share of cells in vs out of the cluster): SST (5.7; 100% vs 100%), RBP4 (2.6; 100% vs 68%), PCSK1 (1.2; 100% vs 53%), PRG4 (1.2; 80% vs 14%), BCHE (0.7; 76% vs 4%), LEPR (0.7; 93% vs 18%), SEC11C (0.6; 100% vs 96%), RGS2 (0.6; 81% vs 49%), AQP3 (0.5; 86% vs 44%), ISL1 (0.5; 100% vs 70%), TPPP3 (0.5; 89% vs 56%), CASR (0.4; 97% vs 37%), HHEX (0.4; 89% vs 12%), UCHL1 (0.4; 100% vs 71%), GABRG2 (0.4; 58% vs 8%), HADH (0.3; 97% vs 64%), UNC5B (0.3; 91% vs 31%), PCP4 (0.3; 92% vs 48%), DIRAS3 (0.3; 83% vs 36%), TENM3 (0.3; 96% vs 41%) +- **Claim:** CL:0000173 *pancreatic D cell* +- Cited markers: SST (supports), RBP4 (supports), PCSK1 (supports), ISL1 (supports), HHEX (supports), UCHL1 (supports) + +## 42. `sc-dd56d93c4db8` + +- Cluster `c07` of a Homo sapiens caudate lobe of liver dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): GNLY (2.0; 57% vs 7%), NKG7 (1.8; 65% vs 23%), TMSB10 (1.3; 97% vs 76%), IGKC (1.2; 60% vs 34%), GZMB (1.2; 53% vs 4%), TUBA1B (1.2; 63% vs 31%), TMSB4X (1.2; 98% vs 89%), PTMA (1.1; 99% vs 91%), MALAT1 (1.1; 99% vs 97%), CORO1A (1.1; 75% vs 29%), CCL5 (1.1; 54% vs 20%), FGFBP2 (1.0; 45% vs 2%), KLRD1 (1.0; 55% vs 14%), PFN1 (1.0; 95% vs 79%), H4C3 (1.0; 64% vs 48%), HMGB2 (1.0; 57% vs 23%), HMGN2 (1.0; 69% vs 40%), CYBA (1.0; 84% vs 44%), CCL4 (1.0; 57% vs 22%), HLA-B (1.0; 96% vs 86%) +- **Claim:** CL:0000623 *natural killer cell* +- Cited markers: GNLY (supports), NKG7 (supports), GZMB (supports), FGFBP2 (supports), KLRD1 (supports), CCL5 (supports), CCL4 (supports), IGKC (contradicts) + +## 43. `sc-1ecb98228b96` + +- Cluster `c13` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): DEFB1 (3.3; 100% vs 38%), SLC12A3 (3.2; 98% vs 4%), TMEM52B (2.7; 100% vs 13%), KNG1 (2.4; 99% vs 16%), MT-ND1 (2.3; 100% vs 94%), WNK1 (2.1; 98% vs 28%), ATP1B1 (2.0; 100% vs 59%), SPP1 (2.0; 98% vs 57%), MT-CO1 (2.0; 100% vs 99%), MT-ND2 (1.9; 100% vs 97%), MT-CO3 (1.9; 100% vs 99%), MT-ATP6 (1.8; 100% vs 97%), MT-ND4 (1.8; 100% vs 97%), MT-CYB (1.8; 100% vs 96%), CA12 (1.7; 98% vs 30%), MALAT1 (1.7; 100% vs 84%), MT-CO2 (1.7; 100% vs 98%), MT-ND5 (1.7; 100% vs 87%), ATP1A1 (1.6; 99% vs 55%), MT-ND3 (1.6; 100% vs 97%) +- **Claim:** CL:1000849 *kidney distal convoluted tubule epithelial cell* +- Cited markers: SLC12A3 (supports), TMEM52B (supports), KNG1 (supports), WNK1 (supports), DEFB1 (supports), CA12 (supports), ATP1B1 (supports), ATP1A1 (supports), SPP1 (supports) + +## 44. `sc-734a08d5281e` + +- Cluster `c05` of a Homo sapiens zone of skin dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): KRT1 (3.3; 97% vs 30%), DMKN (3.2; 98% vs 32%), KRT10 (3.1; 98% vs 49%), SFN (3.1; 98% vs 45%), LGALS7B (2.9; 97% vs 14%), PERP (2.7; 98% vs 42%), KRTDAP (2.7; 89% vs 12%), LY6D (2.6; 96% vs 23%), S100A14 (2.5; 98% vs 29%), LYPD3 (2.1; 91% vs 18%), AQP3 (2.1; 94% vs 24%), DSP (2.1; 94% vs 19%), TACSTD2 (1.7; 91% vs 24%), SERPINB5 (1.7; 90% vs 16%), KRT14 (1.6; 95% vs 67%), CCL27 (1.5; 86% vs 12%), RNF144B (1.5; 79% vs 19%), MIR205HG (1.4; 88% vs 21%), CLDN1 (1.4; 81% vs 11%), FXYD3 (1.4; 86% vs 17%) +- **Claim:** CL:0000312 *keratinocyte* +- Cited markers: KRT1 (supports), DMKN (supports), KRT10 (supports), SFN (supports), LGALS7B (supports), PERP (supports), KRTDAP (supports), LY6D (supports), S100A14 (supports), LYPD3 (supports), AQP3 (supports), DSP (supports), TACSTD2 (supports), SERPINB5 (supports), KRT14 (supports), CCL27 (supports), RNF144B (supports), MIR205HG (supports), CLDN1 (supports), FXYD3 (supports) + +## 45. `sc-c5207ef0df45` + +- Cluster `c07` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): LTB (1.3; 97% vs 57%), CXCR4 (1.3; 89% vs 50%), IL32 (1.3; 93% vs 62%), GZMK (1.2; 77% vs 28%), IL7R (1.1; 85% vs 36%), TRDV2 (1.1; 44% vs 4%), S100A4 (1.1; 91% vs 62%), TNFAIP3 (1.0; 86% vs 60%), DUSP2 (0.9; 93% vs 80%), SPOCK2 (0.9; 86% vs 44%), KLRB1 (0.9; 94% vs 73%), NFKBIA (0.9; 94% vs 81%), TC2N (0.9; 74% vs 28%), S100A6 (0.9; 81% vs 45%), CD3E (0.9; 92% vs 67%), IL23R (0.8; 57% vs 9%), CD69 (0.8; 98% vs 92%), FOSB (0.7; 90% vs 74%), DUSP1 (0.7; 99% vs 93%), SLC4A10 (0.7; 50% vs 5%) +- **Claim:** CL:0000940 *mucosal invariant T cell* +- Cited markers: LTB (supports), CXCR4 (supports), IL32 (supports), GZMK (supports), IL7R (supports), TRDV2 (contradicts), S100A4 (supports), TNFAIP3 (supports), DUSP2 (supports), SPOCK2 (supports), KLRB1 (supports), NFKBIA (supports), TC2N (supports), S100A6 (supports), CD3E (supports), IL23R (supports), CD69 (supports), FOSB (supports), DUSP1 (supports), SLC4A10 (supports) + +## 46. `sc-c92402aae05f` + +- Cluster `c07` of a Homo sapiens pancreas dataset (CEL-seq2) +- Top markers (log fold change; share of cells in vs out of the cluster): COL1A1 (5.5; 100% vs 57%), SPARC (4.5; 100% vs 27%), COL3A1 (4.4; 100% vs 25%), COL1A2 (4.3; 100% vs 32%), FN1 (3.7; 100% vs 19%), COL6A1 (3.5; 100% vs 37%), COL6A3 (3.5; 100% vs 10%), COL4A2 (3.3; 100% vs 24%), COL5A1 (3.2; 100% vs 12%), COL15A1 (3.1; 96% vs 9%), COL5A2 (3.0; 100% vs 11%), TIMP1 (3.0; 100% vs 84%), MICAL2 (2.9; 100% vs 29%), PXDN (2.8; 100% vs 14%), CALD1 (2.8; 100% vs 43%), SERPINE1 (2.7; 97% vs 13%), IGFBP7 (2.6; 99% vs 87%), VIM (2.6; 100% vs 48%), CYGB (2.6; 97% vs 8%), SFRP2 (2.6; 96% vs 6%) +- **Claim:** CL:0000057 *fibroblast* +- Cited markers: COL1A1 (supports), COL3A1 (supports), COL1A2 (supports), COL6A1 (supports), COL6A3 (supports), COL4A2 (supports), COL5A1 (supports), COL5A2 (supports), COL15A1 (supports), SPARC (supports), FN1 (supports), SFRP2 (supports), CYGB (supports) + +## 47. `sc-42f3008ea7cd` + +- Cluster `c10` of a Homo sapiens duodenum dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): TFF3 (5.2; 100% vs 38%), SPINK4 (4.1; 93% vs 20%), CLCA1 (3.6; 98% vs 7%), ZG16 (2.9; 95% vs 52%), MUC2 (2.7; 94% vs 7%), REG4 (2.6; 89% vs 11%), ITLN1 (2.1; 89% vs 4%), AGR2 (2.0; 94% vs 55%), RNASE1 (1.8; 94% vs 14%), SH3BGRL3 (1.7; 100% vs 82%), SPINK1 (1.5; 100% vs 69%), GSN (1.5; 89% vs 34%), LGALS4 (1.5; 100% vs 82%), STARD10 (1.4; 96% vs 33%), KRT18 (1.3; 100% vs 74%), FXYD3 (1.3; 94% vs 49%), CDC42EP5 (1.2; 95% vs 45%), KLK1 (1.2; 84% vs 14%), ST6GALNAC1 (1.2; 84% vs 28%), LRRC26 (1.2; 75% vs 10%) +- **Claim:** CL:0000160 *goblet cell* +- Cited markers: MUC2 (supports), TFF3 (supports), CLCA1 (supports), ITLN1 (supports), AGR2 (supports), SPINK4 (supports), REG4 (supports), ZG16 (supports) + +## 48. `sc-8202e57eff36` + +- Cluster `c03` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): HLA-DRA (3.4; 100% vs 41%), CD74 (3.2; 100% vs 61%), HLA-DPA1 (3.2; 98% vs 25%), HLA-DPB1 (3.1; 98% vs 24%), SRGN (3.0; 98% vs 31%), HLA-DRB1 (2.9; 99% vs 38%), TYROBP (2.7; 98% vs 6%), CXCL8 (2.4; 72% vs 20%), HLA-DQB1 (2.4; 93% vs 18%), IL1B (2.3; 72% vs 8%), GPR183 (2.2; 86% vs 5%), TMSB4X (2.2; 100% vs 74%), HLA-DQA1 (2.2; 93% vs 14%), PLAUR (2.1; 88% vs 15%), RGS1 (2.1; 82% vs 6%), NFKBIA (2.1; 94% vs 56%), RGS2 (2.0; 88% vs 22%), AIF1 (2.0; 92% vs 4%), CCL3 (1.9; 58% vs 12%), CXCL2 (1.8; 72% vs 26%) +- **Claim:** CL:0000451 *dendritic cell* +- Cited markers: HLA-DRA (supports), CD74 (supports), HLA-DPA1 (supports), HLA-DPB1 (supports), HLA-DRB1 (supports), HLA-DQB1 (supports), HLA-DQA1 (supports), GPR183 (supports), TYROBP (supports), AIF1 (supports), SRGN (supports), CXCL8 (supports), IL1B (supports), CCL3 (supports) + +## 49. `sc-919d16f1e63f` + +- Cluster `c08` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): IL32 (1.2; 93% vs 62%), GZMK (1.1; 71% vs 28%), CD3E (1.0; 98% vs 67%), ITM2C (0.9; 74% vs 35%), CD27 (0.9; 85% vs 37%), TRDV2 (0.8; 35% vs 4%), AIF1 (0.7; 82% vs 38%), CNN2 (0.7; 92% vs 64%), CXCR3 (0.7; 71% vs 20%), ZNF683 (0.7; 37% vs 4%), CD8A (0.7; 61% vs 33%), CD52 (0.7; 95% vs 68%), COTL1 (0.7; 80% vs 59%), SIRPG (0.6; 65% vs 23%), CD3D (0.6; 97% vs 69%), ID3 (0.6; 64% vs 32%), TCF7 (0.6; 73% vs 37%), CXCR4 (0.5; 80% vs 50%), LTB (0.5; 87% vs 57%), XCL1 (0.5; 74% vs 47%) +- **Claim:** CL:0000625 *CD8-positive, alpha-beta T cell* +- Cited markers: CD3E (supports), CD3D (supports), CD8A (supports), GZMK (supports), CD27 (supports), CXCR3 (supports), TCF7 (supports), ZNF683 (supports), IL32 (supports), CD52 (supports) + +## 50. `sc-3e8815615c78` + +- Cluster `c07` of a Homo sapiens pancreas dataset (CEL-seq2) +- Top markers (log fold change; share of cells in vs out of the cluster): COL1A1 (5.5; 100% vs 57%), SPARC (4.5; 100% vs 27%), COL3A1 (4.4; 100% vs 25%), COL1A2 (4.3; 100% vs 32%), FN1 (3.7; 100% vs 19%), COL6A1 (3.5; 100% vs 37%), COL6A3 (3.5; 100% vs 10%), COL4A2 (3.3; 100% vs 24%), COL5A1 (3.2; 100% vs 12%), COL15A1 (3.1; 96% vs 9%), COL5A2 (3.0; 100% vs 11%), TIMP1 (3.0; 100% vs 84%), MICAL2 (2.9; 100% vs 29%), PXDN (2.8; 100% vs 14%), CALD1 (2.8; 100% vs 43%), SERPINE1 (2.7; 97% vs 13%), IGFBP7 (2.6; 99% vs 87%), VIM (2.6; 100% vs 48%), CYGB (2.6; 97% vs 8%), SFRP2 (2.6; 96% vs 6%) +- **Claim:** CL:0002410 *pancreatic stellate cell* +- Cited markers: COL1A1 (supports), COL1A2 (supports), COL3A1 (supports), SPARC (supports), FN1 (supports), COL6A1 (supports), COL6A3 (supports), CYGB (supports), SFRP2 (supports) + +## 51. `sc-28afeb3a02f1` + +- Cluster `c03` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): HLA-DRA (3.4; 100% vs 41%), CD74 (3.2; 100% vs 61%), HLA-DPA1 (3.2; 98% vs 25%), HLA-DPB1 (3.1; 98% vs 24%), SRGN (3.0; 98% vs 31%), HLA-DRB1 (2.9; 99% vs 38%), TYROBP (2.7; 98% vs 6%), CXCL8 (2.4; 72% vs 20%), HLA-DQB1 (2.4; 93% vs 18%), IL1B (2.3; 72% vs 8%), GPR183 (2.2; 86% vs 5%), TMSB4X (2.2; 100% vs 74%), HLA-DQA1 (2.2; 93% vs 14%), PLAUR (2.1; 88% vs 15%), RGS1 (2.1; 82% vs 6%), NFKBIA (2.1; 94% vs 56%), RGS2 (2.0; 88% vs 22%), AIF1 (2.0; 92% vs 4%), CCL3 (1.9; 58% vs 12%), CXCL2 (1.8; 72% vs 26%) +- **Claim:** CL:0000235 *macrophage* +- Cited markers: HLA-DRA (supports), CD74 (supports), HLA-DPA1 (supports), HLA-DPB1 (supports), HLA-DRB1 (supports), TYROBP (supports), AIF1 (supports), IL1B (supports), CCL3 (supports) + +## 52. `sc-856e7a60cdcd` + +- Cluster `c10` of a Homo sapiens caudate lobe of liver dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): DNASE1L3 (2.8; 96% vs 13%), PRSS23 (2.7; 92% vs 10%), TIMP1 (2.5; 98% vs 33%), RAMP3 (2.4; 94% vs 8%), IGFBP7 (2.3; 96% vs 16%), ID1 (2.3; 91% vs 20%), LIFR (2.3; 91% vs 10%), INMT (2.2; 88% vs 5%), TIMP3 (2.2; 92% vs 17%), ID3 (2.1; 87% vs 12%), RNASE1 (2.1; 90% vs 11%), TAGLN (2.1; 73% vs 4%), ENG (2.1; 90% vs 10%), IFI27 (1.9; 86% vs 10%), HSPG2 (1.9; 87% vs 8%), C7 (1.9; 79% vs 7%), HLA-E (1.8; 97% vs 59%), PTPRB (1.8; 86% vs 8%), FCN3 (1.7; 66% vs 14%), PLAC8 (1.7; 87% vs 24%) +- **Claim:** CL:0000632 *hepatic stellate cell* +- Cited markers: TIMP1 (supports), TIMP3 (supports), IGFBP7 (supports), TAGLN (supports), HSPG2 (supports), RAMP3 (supports), LIFR (supports), ID1 (supports), ID3 (supports), PRSS23 (supports), ENG (supports) + +## 53. `sc-6a41569ec051` + +- Cluster `c12` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): LTB (2.2; 99% vs 57%), IL7R (2.1; 98% vs 35%), IL4I1 (1.6; 84% vs 5%), LST1 (1.5; 86% vs 17%), CCR6 (1.4; 86% vs 4%), VIM (1.4; 96% vs 66%), S100A4 (1.4; 92% vs 62%), RORC (1.4; 82% vs 6%), S100A6 (1.4; 89% vs 44%), TNFAIP3 (1.3; 96% vs 60%), AQP3 (1.3; 88% vs 30%), RGS1 (1.2; 92% vs 54%), ITM2C (1.1; 85% vs 34%), JAML (1.0; 76% vs 17%), DDIT4 (1.0; 90% vs 60%), CXCR4 (1.0; 86% vs 50%), KIT (1.0; 66% vs 4%), TMIGD2 (0.9; 84% vs 46%), IL23R (0.9; 70% vs 8%), ZFP36L1 (0.9; 98% vs 88%) +- **Claim:** CL:0001071 *group 3 innate lymphoid cell* +- Cited markers: RORC (supports), IL7R (supports), KIT (supports), IL23R (supports), CCR6 (supports), LTB (supports) + +## 54. `sc-c3eb60920f61` + +- Cluster `c03` of a Homo sapiens caudate lobe of liver dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): C1QA (3.6; 100% vs 14%), C1QB (3.5; 99% vs 11%), CD163 (2.9; 99% vs 10%), HLA-DRA (2.7; 98% vs 22%), CTSB (2.6; 99% vs 38%), MARCO (2.6; 86% vs 4%), FTL (2.5; 100% vs 100%), SLC40A1 (2.5; 94% vs 32%), C1QC (2.5; 95% vs 8%), MS4A6A (2.4; 95% vs 17%), CD74 (2.3; 98% vs 40%), MS4A7 (2.3; 91% vs 8%), CTSS (2.3; 96% vs 24%), AIF1 (2.2; 95% vs 15%), TYROBP (2.2; 97% vs 28%), NPC2 (2.2; 97% vs 42%), SAT1 (2.2; 100% vs 72%), GPX1 (2.1; 99% vs 66%), HMOX1 (2.1; 85% vs 9%), FCER1G (2.1; 94% vs 24%) +- **Claim:** CL:0000091 *Kupffer cell* +- Cited markers: MARCO (supports), CD163 (supports), C1QA (supports), C1QB (supports), C1QC (supports), SLC40A1 (supports), HMOX1 (supports), MS4A7 (supports) + +## 55. `sc-080a4ad54b3b` + +- Cluster `c11` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): FXYD4 (2.8; 83% vs 8%), AQP2 (2.7; 80% vs 4%), AQP3 (2.1; 80% vs 12%), MALAT1 (2.1; 100% vs 84%), MT-CO1 (1.6; 100% vs 99%), CLU (1.6; 84% vs 39%), WFDC2 (1.6; 78% vs 20%), TACSTD2 (1.6; 78% vs 21%), GDF15 (1.5; 80% vs 29%), MT-CO2 (1.4; 100% vs 98%), S100A6 (1.4; 92% vs 67%), MT-CO3 (1.4; 100% vs 99%), CDH16 (1.3; 86% vs 32%), KRT19 (1.3; 68% vs 10%), MT-ND4 (1.3; 100% vs 98%), KRT18 (1.3; 83% vs 38%), MT-CYB (1.3; 100% vs 97%), ELF3 (1.2; 76% vs 25%), MT-ATP6 (1.2; 100% vs 97%), NEAT1 (1.2; 97% vs 69%) +- **Claim:** CL:1001431 *kidney collecting duct principal cell* +- Cited markers: FXYD4 (supports), AQP2 (supports), AQP3 (supports), CDH16 (supports) + +## 56. `sc-b27ceddaabda` + +- Cluster `c09` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): CCL3 (2.0; 87% vs 34%), CCL4 (1.8; 96% vs 47%), NKG7 (1.7; 100% vs 66%), TYROBP (1.7; 98% vs 49%), XCL1 (1.5; 90% vs 40%), IL2RB (1.5; 94% vs 50%), GZMK (1.5; 82% vs 19%), CCL5 (1.5; 80% vs 48%), GZMA (1.2; 97% vs 65%), TRDC (1.2; 84% vs 39%), KLRB1 (1.2; 99% vs 69%), XCL2 (1.2; 85% vs 41%), ALOX5AP (1.2; 91% vs 58%), CMC1 (1.2; 87% vs 57%), ITM2C (1.1; 66% vs 30%), KLRC1 (1.1; 83% vs 39%), GSTP1 (1.0; 88% vs 57%), CST7 (1.0; 96% vs 53%), IER2 (1.0; 99% vs 94%), SRGN (1.0; 99% vs 90%) +- **Claim:** CL:0000623 *natural killer cell* +- Cited markers: NKG7 (supports), TYROBP (supports), IL2RB (supports), KLRC1 (supports), XCL1 (supports), XCL2 (supports), CST7 (supports), GZMA (supports), GZMK (supports), CCL5 (supports) + +## 57. `sc-3e4827f1cf8e` + +- Cluster `c03` of a Homo sapiens pancreas dataset (CEL-seq2) +- Top markers (log fold change; share of cells in vs out of the cluster): SST (5.7; 100% vs 100%), RBP4 (2.6; 100% vs 68%), PCSK1 (1.2; 100% vs 53%), PRG4 (1.2; 80% vs 14%), BCHE (0.7; 76% vs 4%), LEPR (0.7; 93% vs 18%), SEC11C (0.6; 100% vs 96%), RGS2 (0.6; 81% vs 49%), AQP3 (0.5; 86% vs 44%), ISL1 (0.5; 100% vs 70%), TPPP3 (0.5; 89% vs 56%), CASR (0.4; 97% vs 37%), HHEX (0.4; 89% vs 12%), UCHL1 (0.4; 100% vs 71%), GABRG2 (0.4; 58% vs 8%), HADH (0.3; 97% vs 64%), UNC5B (0.3; 91% vs 31%), PCP4 (0.3; 92% vs 48%), DIRAS3 (0.3; 83% vs 36%), TENM3 (0.3; 96% vs 41%) +- **Claim:** CL:0000173 *pancreatic D cell* +- Cited markers: SST (supports), HHEX (supports), RBP4 (supports), LEPR (supports), PCSK1 (supports), ISL1 (supports) + +## 58. `sc-fbc3fa6e37a7` + +- Cluster `c07` of a Homo sapiens caudate lobe of liver dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): GNLY (2.0; 57% vs 7%), NKG7 (1.8; 65% vs 23%), TMSB10 (1.3; 97% vs 76%), IGKC (1.2; 60% vs 34%), GZMB (1.2; 53% vs 4%), TUBA1B (1.2; 63% vs 31%), TMSB4X (1.2; 98% vs 89%), PTMA (1.1; 99% vs 91%), MALAT1 (1.1; 99% vs 97%), CORO1A (1.1; 75% vs 29%), CCL5 (1.1; 54% vs 20%), FGFBP2 (1.0; 45% vs 2%), KLRD1 (1.0; 55% vs 14%), PFN1 (1.0; 95% vs 79%), H4C3 (1.0; 64% vs 48%), HMGB2 (1.0; 57% vs 23%), HMGN2 (1.0; 69% vs 40%), CYBA (1.0; 84% vs 44%), CCL4 (1.0; 57% vs 22%), HLA-B (1.0; 96% vs 86%) +- **Claim:** CL:0000623 *natural killer cell* +- Cited markers: NKG7 (supports), GNLY (supports), GZMB (supports), KLRD1 (supports), CCL5 (supports), CCL4 (supports), CORO1A (supports), IGKC (contradicts) + +## 59. `sc-eab9c0d47ac4` + +- Cluster `c05` of a Homo sapiens duodenum dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): LYZ (2.9; 98% vs 25%), MT1G (2.4; 98% vs 77%), GUCA2B (2.2; 61% vs 25%), BEST4 (2.2; 74% vs 1%), SPIB (2.0; 84% vs 4%), MT2A (1.9; 96% vs 82%), MT1E (1.8; 99% vs 70%), MT1H (1.8; 82% vs 52%), CFTR (1.7; 92% vs 28%), MT1X (1.6; 96% vs 60%), GSN (1.5; 82% vs 34%), CPA2 (1.4; 75% vs 3%), CA7 (1.4; 68% vs 1%), KRT20 (1.4; 83% vs 44%), CTSE (1.2; 90% vs 29%), GUCA2A (1.2; 50% vs 2%), FXYD3 (1.2; 94% vs 49%), S100A6 (1.2; 100% vs 95%), MT1M (1.1; 74% vs 27%), STARD10 (1.1; 94% vs 32%) +- **Claim:** CL:0002071 *enterocyte of epithelium of small intestine* +- Cited markers: BEST4 (supports), CA7 (supports), SPIB (supports), CFTR (supports), GUCA2B (supports), GUCA2A (supports), CPA2 (supports), CTSE (supports), KRT20 (supports) + +## 60. `sc-c11395392128` + +- Cluster `c12` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): LTB (2.2; 99% vs 57%), IL7R (2.1; 98% vs 35%), IL4I1 (1.6; 84% vs 5%), LST1 (1.5; 86% vs 17%), CCR6 (1.4; 86% vs 4%), VIM (1.4; 96% vs 66%), S100A4 (1.4; 92% vs 62%), RORC (1.4; 82% vs 6%), S100A6 (1.4; 89% vs 44%), TNFAIP3 (1.3; 96% vs 60%), AQP3 (1.3; 88% vs 30%), RGS1 (1.2; 92% vs 54%), ITM2C (1.1; 85% vs 34%), JAML (1.0; 76% vs 17%), DDIT4 (1.0; 90% vs 60%), CXCR4 (1.0; 86% vs 50%), KIT (1.0; 66% vs 4%), TMIGD2 (0.9; 84% vs 46%), IL23R (0.9; 70% vs 8%), ZFP36L1 (0.9; 98% vs 88%) +- **Claim:** CL:0001071 *group 3 innate lymphoid cell* +- Cited markers: IL7R (supports), CCR6 (supports), RORC (supports), IL23R (supports), KIT (supports), LTB (supports), JAML (supports), LST1 (contradicts) + +## 61. `sc-1bd2021fc890` + +- Cluster `c05` of a Homo sapiens duodenum dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): LYZ (2.9; 98% vs 25%), MT1G (2.4; 98% vs 77%), GUCA2B (2.2; 61% vs 25%), BEST4 (2.2; 74% vs 1%), SPIB (2.0; 84% vs 4%), MT2A (1.9; 96% vs 82%), MT1E (1.8; 99% vs 70%), MT1H (1.8; 82% vs 52%), CFTR (1.7; 92% vs 28%), MT1X (1.6; 96% vs 60%), GSN (1.5; 82% vs 34%), CPA2 (1.4; 75% vs 3%), CA7 (1.4; 68% vs 1%), KRT20 (1.4; 83% vs 44%), CTSE (1.2; 90% vs 29%), GUCA2A (1.2; 50% vs 2%), FXYD3 (1.2; 94% vs 49%), S100A6 (1.2; 100% vs 95%), MT1M (1.1; 74% vs 27%), STARD10 (1.1; 94% vs 32%) +- **Claim:** CL:0000584 *enterocyte* +- Cited markers: BEST4 (supports), CA7 (supports), CFTR (supports), GUCA2A (supports), GUCA2B (supports), KRT20 (supports) + +## 62. `sc-37758b97b56f` + +- Cluster `c12` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): LTB (2.2; 99% vs 57%), IL7R (2.1; 98% vs 35%), IL4I1 (1.6; 84% vs 5%), LST1 (1.5; 86% vs 17%), CCR6 (1.4; 86% vs 4%), VIM (1.4; 96% vs 66%), S100A4 (1.4; 92% vs 62%), RORC (1.4; 82% vs 6%), S100A6 (1.4; 89% vs 44%), TNFAIP3 (1.3; 96% vs 60%), AQP3 (1.3; 88% vs 30%), RGS1 (1.2; 92% vs 54%), ITM2C (1.1; 85% vs 34%), JAML (1.0; 76% vs 17%), DDIT4 (1.0; 90% vs 60%), CXCR4 (1.0; 86% vs 50%), KIT (1.0; 66% vs 4%), TMIGD2 (0.9; 84% vs 46%), IL23R (0.9; 70% vs 8%), ZFP36L1 (0.9; 98% vs 88%) +- **Claim:** CL:0001071 *group 3 innate lymphoid cell* +- Cited markers: IL7R (supports), RORC (supports), KIT (supports), CCR6 (supports), IL23R (supports), LTB (supports) + +## 63. `sc-0e9dd71166b2` + +- Cluster `c06` of a Homo sapiens pancreas dataset (CEL-seq2) +- Top markers (log fold change; share of cells in vs out of the cluster): PPY (6.4; 100% vs 54%), SCG2 (2.0; 100% vs 90%), PEG10 (2.0; 100% vs 88%), ID2 (1.8; 98% vs 73%), PAX6 (1.7; 100% vs 79%), ETV1 (1.6; 99% vs 41%), AQP3 (1.6; 90% vs 46%), MEIS2 (1.5; 98% vs 75%), ITM2C (1.3; 99% vs 92%), PCSK2 (1.2; 100% vs 94%), CHGB (1.2; 100% vs 93%), THSD7A (1.2; 86% vs 14%), PAM (1.2; 100% vs 94%), GAD2 (1.2; 98% vs 70%), NEUROD1 (1.2; 99% vs 80%), SERTM1 (1.1; 84% vs 9%), ABCC9 (1.1; 96% vs 71%), UCHL1 (1.1; 97% vs 72%), ABCC8 (1.1; 99% vs 73%), SLC6A4 (1.1; 91% vs 56%) +- **Claim:** CL:0002275 *pancreatic PP cell* +- Cited markers: PPY (supports), SCG2 (supports), PAX6 (supports), PCSK2 (supports), CHGB (supports), NEUROD1 (supports) + +## 64. `sc-06c333c4660a` + +- Cluster `c02` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): CXCR4 (2.2; 83% vs 14%), MALAT1 (2.1; 100% vs 84%), RPS29 (2.0; 100% vs 97%), RPS27 (1.9; 100% vs 99%), TMSB4X (1.9; 99% vs 74%), BTG1 (1.8; 94% vs 67%), SRGN (1.6; 80% vs 32%), CD52 (1.6; 65% vs 5%), RPL17 (1.5; 89% vs 70%), RPS3 (1.5; 100% vs 92%), RPS15A (1.5; 100% vs 96%), IL7R (1.5; 57% vs 2%), RPS19 (1.5; 99% vs 96%), CD2 (1.5; 59% vs 1%), RPLP2 (1.4; 100% vs 98%), RPS2 (1.4; 99% vs 91%), PTPRC (1.4; 61% vs 7%), RPL41 (1.4; 100% vs 99%), B2M (1.4; 100% vs 96%), RPS21 (1.4; 97% vs 89%) +- **Claim:** CL:0000084 *T cell* +- Cited markers: CD2 (supports), IL7R (supports), CD52 (supports), PTPRC (supports) + +## 65. `sc-7f4823b086ad` + +- Cluster `c03` of a Homo sapiens pancreas dataset (CEL-seq2) +- Top markers (log fold change; share of cells in vs out of the cluster): SST (5.7; 100% vs 100%), RBP4 (2.6; 100% vs 68%), PCSK1 (1.2; 100% vs 53%), PRG4 (1.2; 80% vs 14%), BCHE (0.7; 76% vs 4%), LEPR (0.7; 93% vs 18%), SEC11C (0.6; 100% vs 96%), RGS2 (0.6; 81% vs 49%), AQP3 (0.5; 86% vs 44%), ISL1 (0.5; 100% vs 70%), TPPP3 (0.5; 89% vs 56%), CASR (0.4; 97% vs 37%), HHEX (0.4; 89% vs 12%), UCHL1 (0.4; 100% vs 71%), GABRG2 (0.4; 58% vs 8%), HADH (0.3; 97% vs 64%), UNC5B (0.3; 91% vs 31%), PCP4 (0.3; 92% vs 48%), DIRAS3 (0.3; 83% vs 36%), TENM3 (0.3; 96% vs 41%) +- **Claim:** CL:0000173 *pancreatic D cell* +- Cited markers: SST (supports), HHEX (supports), RBP4 (supports), LEPR (supports), BCHE (supports), PRG4 (supports), CASR (supports), ISL1 (supports), PCSK1 (supports) + +## 66. `sc-fd438ca2c872` + +- Cluster `c05` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): LTB (1.6; 98% vs 52%), CD52 (1.4; 98% vs 64%), TCF7 (1.3; 91% vs 30%), LEF1 (1.3; 92% vs 29%), SELL (1.3; 88% vs 29%), IL7R (1.2; 84% vs 30%), LDHB (1.2; 97% vs 72%), CCR7 (1.1; 84% vs 20%), RGS10 (1.0; 90% vs 33%), NOSIP (1.0; 90% vs 51%), LRRN3 (0.9; 71% vs 10%), CD3E (0.9; 97% vs 63%), VIM (0.9; 95% vs 62%), RPS8 (0.9; 100% vs 100%), AIF1 (0.8; 82% vs 33%), CAMK4 (0.8; 79% vs 25%), RPL36A (0.8; 94% vs 78%), CD27 (0.8; 84% vs 32%), CORO1B (0.8; 76% vs 31%), CD40LG (0.8; 65% vs 7%) +- **Claim:** CL:0000895 *naive thymus-derived CD4-positive, alpha-beta T cell* +- Cited markers: CCR7 (supports), SELL (supports), LEF1 (supports), TCF7 (supports), LRRN3 (supports), IL7R (supports), CAMK4 (supports), CD27 (supports), LTB (supports), CD3E (supports), CD40LG (supports), NOSIP (supports) + +## 67. `sc-01004e474ad5` + +- Cluster `c09` of a Homo sapiens kidney dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): WFDC2 (2.9; 97% vs 21%), KRT7 (2.8; 100% vs 13%), IGFBP5 (2.0; 97% vs 35%), CA12 (1.9; 97% vs 33%), TMEM213 (1.9; 97% vs 21%), CD9 (1.8; 100% vs 50%), ATP6V1G3 (1.8; 90% vs 10%), MALAT1 (1.8; 100% vs 85%), ATP6V0B (1.7; 100% vs 52%), LGALS3 (1.7; 100% vs 39%), ATP6V0D2 (1.6; 94% vs 13%), TSPAN8 (1.6; 87% vs 8%), ATP6V1B1 (1.6; 94% vs 20%), KRT8 (1.4; 94% vs 38%), HEPACAM2 (1.4; 90% vs 8%), ATP6AP2 (1.4; 97% vs 44%), S100A2 (1.3; 81% vs 23%), ID3 (1.3; 87% vs 33%), IL18 (1.3; 90% vs 12%), SLC26A4 (1.3; 77% vs 1%) +- **Claim:** CL:0002201 *renal beta-intercalated cell* +- Cited markers: SLC26A4 (supports), ATP6V1B1 (supports), ATP6V0D2 (supports), ATP6V1G3 (supports), ATP6V0B (supports), ATP6AP2 (supports), TMEM213 (supports), HEPACAM2 (supports), CA12 (supports), WFDC2 (supports), KRT7 (supports), KRT8 (supports), TSPAN8 (supports) + +## 68. `sc-8328256db783` + +- Cluster `c08` of a Homo sapiens lung dataset (10x 5' v1) +- Top markers (log fold change; share of cells in vs out of the cluster): IL32 (1.2; 93% vs 62%), GZMK (1.1; 71% vs 28%), CD3E (1.0; 98% vs 67%), ITM2C (0.9; 74% vs 35%), CD27 (0.9; 85% vs 37%), TRDV2 (0.8; 35% vs 4%), AIF1 (0.7; 82% vs 38%), CNN2 (0.7; 92% vs 64%), CXCR3 (0.7; 71% vs 20%), ZNF683 (0.7; 37% vs 4%), CD8A (0.7; 61% vs 33%), CD52 (0.7; 95% vs 68%), COTL1 (0.7; 80% vs 59%), SIRPG (0.6; 65% vs 23%), CD3D (0.6; 97% vs 69%), ID3 (0.6; 64% vs 32%), TCF7 (0.6; 73% vs 37%), CXCR4 (0.5; 80% vs 50%), LTB (0.5; 87% vs 57%), XCL1 (0.5; 74% vs 47%) +- **Claim:** CL:0000625 *CD8-positive, alpha-beta T cell* +- Cited markers: CD3E (supports), CD3D (supports), CD8A (supports), GZMK (supports), CXCR3 (supports), TRDV2 (contradicts) + +## 69. `sc-264c388c143f` + +- Cluster `c06` of a Homo sapiens duodenum dataset (10x 3' v2) +- Top markers (log fold change; share of cells in vs out of the cluster): HLA-DPB1 (4.1; 100% vs 22%), HLA-DPA1 (4.0; 100% vs 30%), C1QA (3.6; 89% vs 13%), C1QC (3.5; 87% vs 8%), C1QB (3.5; 87% vs 10%), TYROBP (3.4; 95% vs 9%), HLA-DRA (3.2; 100% vs 70%), TMSB4X (3.1; 100% vs 100%), HLA-DQA1 (3.0; 99% vs 4%), CD74 (3.0; 100% vs 95%), HLA-DRB1 (3.0; 100% vs 64%), CST3 (2.7; 100% vs 84%), AIF1 (2.7; 100% vs 5%), LYZ (2.6; 97% vs 23%), MS4A6A (2.6; 96% vs 3%), HLA-DQB1 (2.5; 99% vs 24%), SELENOP (2.3; 81% vs 64%), NPC2 (2.0; 100% vs 76%), FCER1G (2.0; 93% vs 5%), HLA-DMA (2.0; 96% vs 37%) +- **Claim:** CL:0000235 *macrophage* +- Cited markers: C1QA (supports), C1QB (supports), C1QC (supports), AIF1 (supports), MS4A6A (supports), TYROBP (supports), FCER1G (supports), LYZ (supports), CD74 (supports), HLA-DRA (supports), HLA-DPB1 (supports), HLA-DPA1 (supports), HLA-DQA1 (supports), HLA-DRB1 (supports), HLA-DQB1 (supports), HLA-DMA (supports), CST3 (supports), SELENOP (supports), NPC2 (supports) diff --git a/evaluation/llm_benchmark/results/celltype-audit/audit_report.json b/evaluation/llm_benchmark/results/celltype-audit/audit_report.json new file mode 100644 index 0000000..fca5404 --- /dev/null +++ b/evaluation/llm_benchmark/results/celltype-audit/audit_report.json @@ -0,0 +1,70 @@ +{ + "method": "llm-feedback-loop", + "requested_use": "research_summary", + "confidence": 0.95, + "seed": 20261006, + "population": { + "admitted": 255, + "rejected": 3, + "review_required": 18 + }, + "routes": { + "admitted": { + "audited": 59, + "errors": 14, + "uncertain": 0, + "route_name": "auto-admitted", + "population": 255, + "share_of_records": 0.9239, + "error_rate": 0.2373, + "error_rate_wilson_95": [ + 0.146946, + 0.359749 + ], + "error_rate_upper_bound": 0.345827, + "errors_in_route_at_most": 89, + "audit_minutes": null, + "expert_minutes_estimated": null, + "share_of_expert_time": null + }, + "rejected": { + "audited": 2, + "errors": 0, + "uncertain": 0, + "route_name": "rejected", + "population": 3, + "share_of_records": 0.0109, + "error_rate": 0.0, + "error_rate_wilson_95": [ + 0.0, + 0.65762 + ], + "error_rate_upper_bound": 0.776394, + "errors_in_route_at_most": 3, + "audit_minutes": null, + "expert_minutes_estimated": null, + "share_of_expert_time": null + }, + "review_required": { + "audited": 8, + "errors": 1, + "uncertain": 0, + "route_name": "expert review", + "population": 18, + "share_of_records": 0.0652, + "error_rate": 0.125, + "error_rate_wilson_95": [ + 0.022417, + 0.470888 + ], + "error_rate_upper_bound": 0.47068, + "errors_in_route_at_most": 9, + "audit_minutes": null, + "expert_minutes_estimated": null, + "share_of_expert_time": null + } + }, + "not_yet_audited": [], + "admitted_error_upper_bound": 0.345827, + "expert_minutes_estimated": null +} diff --git a/evaluation/llm_benchmark/results/celltype-audit/predictions.csv b/evaluation/llm_benchmark/results/celltype-audit/predictions.csv new file mode 100644 index 0000000..7fba3e1 --- /dev/null +++ b/evaluation/llm_benchmark/results/celltype-audit/predictions.csv @@ -0,0 +1,277 @@ +case_id,requested_use,record_sha256,profile_sha256,method,predicted_status +sc-00277595642d,research_summary,be465df4041d9a6af260ad0297030b664f53b11ff079c62961d0a0b4bd27566a,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-00a40a840007,research_summary,0611f7cec9d37b055231413531e1befb17272f7a2f6ead20946e3cd081161842,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-00ad508b5979,research_summary,7c1d935bda9f312817ca35e84513fc48551e186242fdb8c0cc04f0a58d9d5bf0,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-01004e474ad5,research_summary,3cfb6d6aa30d20d951385b016a1348a5daa34144f47ecd49ff83b86d0d331fc5,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-0130e09b3cf9,research_summary,e44b69015215c4b0067c3fb2c8bc97380d01879d3dc2369b7a81d331fbf92c35,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-01ab2950a674,research_summary,ad96d5f842fd994e6b36ad3a8b750803dc0ee9198b20004fcdff3550f0144d6b,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-026f0e31de34,research_summary,44cb8b6fdc2784b7067473984f9736a919846aa603b863a95ab86ff6840c97e0,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-0367242cb046,research_summary,b83c06c2011c1f5a2c133e00c7f2cefa14211d7969103b5478d25b9f276ee857,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-06841a5648e9,research_summary,0c46a8a8b8313de70c97dc74e2ce3f33e53ee8ae0e78b789af67f6a0898ab133,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-06c333c4660a,research_summary,0ceb3c532e8dff710feca6d264986161ad274f1aabe0caeee26a93cdd6f705cc,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-080a4ad54b3b,research_summary,51533adf0d1ec9ee8affa7fbfdab0d08da75d6cf2a554e421d16649b91c759f0,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-0890d15729ee,research_summary,033e4af2493c43033d8917d721f4c2784ce8480902b40bd9bc94fd6c67fcfa92,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-08a2410a9405,research_summary,a21506fdf2e5abc556b09ca818d0a4aabe40ae8e34c28f5004d77a0160db2333,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-0adfe6138a72,research_summary,dc18689afc2eed13b95d14f713679f75deea5bb3f7cdb962960ffdac031b6a9e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-0c6b4f84f573,research_summary,9faf4523b950ba815cdb2a48adab08413dbf4d13c1a3ad7036f2031d094aa0bf,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-0cf6ae2934a9,research_summary,003dce60ca8a24a2dc4f90a58ad11e48646642321932e30880d405d4f1afbe9c,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-0d6bf5bd2973,research_summary,ac82b9366d4bef990108ce5f8bad4955b47c583394e2960daeb21a34e980185e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-0e6517e09cc2,research_summary,30ffa9f3014316e971d25e446051d837119a108b331bb982da99073e6594b9a7,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-0e9dd71166b2,research_summary,aa9554ff83c256ac3fffd2f5a5c2e88b1cc5f943eda5ec75063d3df90a695eca,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-0ed3dae30cb0,research_summary,1bc41867170d8bf55c13476c5bfd294fbd88bf0fb3d26f7545aafe63a1cdafdd,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-106b535c7840,research_summary,423208b37c59d49c65a4b384738775e6ab74c55fd57e17d9ea11eb325676d0bb,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-107213b1b6bf,research_summary,f0b29398c287f16159cf54c5a2f558509ca3bb4ef3377d5996b23abe871ba101,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-1a02c2605533,research_summary,28ca11d950fb093d860f2f9e8c2cf277a383b9f52b38b62c3e326f1afb2c60aa,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-1a6f4557c4dc,research_summary,b72f308dac185ffd75017c51b715e8262fb506bc54d03689ed120fd0559e3618,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-1bb81802ec74,research_summary,0e93d1c9adf4dc389eb871d7e3982d6b8d9bac2521a60bf67185a8319910e9fe,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-1bd2021fc890,research_summary,72220776a8460cad0b99a0990e4e409ffb88f140ab8504c739c77eb33dacb9fa,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-1ce9c117d5e3,research_summary,b2f6e12425b3fe5ac5525d8a4ec9fb1fa7f1f2a83b2ee015e2f82ca4c0eb1699,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-1d83a5741b1f,research_summary,5a67daeac01ee91753894f0f03f6cf04b83d4821e822efdb6639b9bc9b4d59cc,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-1ecb98228b96,research_summary,9322381342fad881e0e73ad6d35ab969fc7b9cfde109e501bc8725cc0418d8c2,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-1f0addd61a6b,research_summary,0bca0ca0d90efe133513ce70d2406bb0f2cb0da45829cf8722b06d49e38b999e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-1f9da4b8e364,research_summary,4f4557d0be7c864cc3ceecf67d04c8306fe2f2cd85fa7abf3d3ba224b88f5108,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-2138b2030213,research_summary,ae99cf1ac3504977ea7502d5f5c4b140bdd4f137b896b47a6ad1491979a1efa9,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-22547df41f7f,research_summary,1a12331dcaec1a76c107ffdddbba1d0e47deed1dc974d88a902dbf1ff132c9e3,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-2258691ddec9,research_summary,5a4779afa2cd62890d0c29bb83e05415ce40b4c915496eccad12f599178d2054,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-23f6a4c1c8be,research_summary,c1e78010e5c91daa1db40e958a265938ce319bf9810e87e5931e62573b12332d,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-264c388c143f,research_summary,6781f580a06fb4ecc2dc3af1d831c135e931dd10184e39febfe36ec8b814fa44,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-26e57da30416,research_summary,e921db952a6e408cdead4ab59e363ebd2e06573d490b434ff5307f4fa761bc87,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-285c5fc2eccc,research_summary,83ef5e49ae5a615d254467a63448d10c7925f4e814a8e8e8a998385cd4714704,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-28afeb3a02f1,research_summary,f320aa538194d0f478775cd69d49a5460d6e5c213a817f1529dd5b1d03a13d00,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-29278e133334,research_summary,d1b5f5c1e21342d0b3431d472a5ffd6af4298fa983f4ebb1c4d85fe6ec444f8e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-2a4ac9d42c4b,research_summary,e602c258e76bae8066331502db71d070d6c4c54f31f3366da8524cf1e5695dc7,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-2a67ac7a3771,research_summary,239198f34a400233c02a721a42311e5792a686c2d402c0c999898a89bb8dc163,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-2a709b047cd1,research_summary,1d531377c6c7153f1b033023565408aa04ab42e03c237c6b278b53148bcf5d75,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-2b181f6a406b,research_summary,f0cf52a28493ae8f5736d8b46d26313e17e0d6fba8b02299d1728ebd6f370382,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-2bee4a3eb6bb,research_summary,4e2c28a4414b8e073504c86af4b093b49682e53d13b4621cdc506ea44315ceca,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-2c1b4befc878,research_summary,0a58922afef83897a5ae960e2bae3c5ab30c5ecd3ed7d908dbd67872ca3fdf8f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-2d06b761bbe6,research_summary,9a81cceac78844c5ef5156cf0666940402f94941c3bd805788cc4c432884425a,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-2eff7bb75fad,research_summary,621d4e81f8cea24ee1b7bf5ec16642aeb14042281998979c6d08fb1b20bd0d07,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-304cf26fb2ea,research_summary,415ae20c938ec3145e2cbff50aa23ce53a302e8329c1be6c42af064766907231,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-309790b7d1f9,research_summary,de8cde69d07ff44431a42524e29cb424c0471b0440f67f8c0f36089178509927,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-335da704ab02,research_summary,68b7930f43ff8621bfecc0cbda2b19cb1423e6ad2eb82a522b0838e1748d3333,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-3440a71c316e,research_summary,43de0c39755de5852f9048867ed093b1969abe7dbd3555272c5f14177651e4cf,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-34992e2bf4c0,research_summary,2f678a63efec9ac2537b90d18c1f19a80356b4af283916b16f2c0dd140285aa3,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-34cb72e5b6d5,research_summary,c6045322823bd42660532e0086e57fc5465fda1636c4e7bc212dd815de15b163,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-3511ba96d3a0,research_summary,85f2943f41b938996ba1c88ee3c88b299e44a0b535bac62699c3ec10f78a8c9a,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-37758b97b56f,research_summary,4b5faca597cf488ea5ba0987388a6c050f8f9a2134b476b8dad227e21c37e099,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-3a5b3533a81d,research_summary,d9b1fb6a9438e2803807b251803199e909b1ebfd4f763abacc71a4724f363ad4,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-3b7c8e4937bb,research_summary,391b53098b6cb98381b221224a257cf63370a69d15be9dd07a7ab8ed739fd546,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-3d81f1a882b4,research_summary,f3f38b25f3b8a8a48b99ef043fb8572a21a95eb6160a8701b3fef1fbf906e7b3,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-3e4827f1cf8e,research_summary,46c6d2ac0be5e4ef8f09ff50dea04d791baff09d454ea9cdd2b2b1033419b8d0,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-3e8815615c78,research_summary,204fbd911d94f7e79c007890be794cb0b7e4de2149c5a65d628e5204a6d3c15e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-3fd4070a536c,research_summary,662548dd829f6ed2c599dfb9f08ce75a73472c671615272a1bcd042541fb15ba,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-400c4bba9e95,research_summary,f069beda82b42b6319364299bf01bfc2ec920659efeebf0716febe53ad47c995,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-42f3008ea7cd,research_summary,e6fa8ef66a64479a65e796ff86e32713aa8df0f81a75c460d6e99cc40003aace,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-43489d3f91fb,research_summary,ea8d77cbff3b8717daabadc19de23b8c44105e1dfd225760713b9777e7ebfdae,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-457686ee8397,research_summary,9b58d12cf37dc801382c820a1bdb4cd14a42f126f3997a3c5dd565d6cda58955,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-45fef2cc0b08,research_summary,9eb0b971b53bac0793019c8bdec71032e878966a932e0fc4ff9a110d81ae45f1,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-46206c0096e6,research_summary,b9e39724b58195b5edcbf9535a3513203d3c9d5c1940367a3602e94940f04994,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-46891659601b,research_summary,de5c1506328486319fbdfa1d7c1f40583c37386c64ba4ee3089832536af1a488,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-46db31b5e7be,research_summary,b5fd6dd3c5cbb04a2bec2d0c287821b11a04d3a97e390373fce4c838c274703b,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-47865daed309,research_summary,07e51b6bdecbc8c001d641151be60aef83f85f8dca46744aef17a777790779fc,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-47a38a347209,research_summary,8df4d7eccfcba81e2c3da20acc245def98a3e6d7ecd4077f57a55c28a9d3da3b,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-47ebf1f674cd,research_summary,38be7b8ee4638cb5b5eb6144c0a46d51ff7dfc7d05fbc993f066ada4a91ce9d6,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-48954d06cb44,research_summary,c0c48c586867e75b85a22bed5a90214f0a97b0576f0e35e6997826e9e8c5d79c,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-4937458f58f8,research_summary,4f9898ad50bb99d5365ec89e9d01290b0eb826a02b04293418b27ac944a1e748,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-4983f82802c0,research_summary,e9d111260b965654177983d37126f8167d658486dd86bb8cd6dc305fde5554b9,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-4b2b578b0c50,research_summary,9408167ff4e67bf89f12b461964245cf2fee1bc37f2d2d62bea072c6aecdbe72,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-4c0924ec42b7,research_summary,e3257e43784af581aa6b074cb2d4110f6347c8472e94b4ebf00e738873f6ee29,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-4d6f03e0ceaf,research_summary,9e0a1068a7ba05e9f107663f7bcbaf083ae1b5b9db0715a5af6b8e65e0367430,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-4d871ed1c820,research_summary,2b83633a07c88f0dfd10323dd7e405eb375b1dc956109b5bb02541c18cee72d7,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-4e18192d32a2,research_summary,fb9f635f8cfc56295e31f1cab4a6e0af2a6f2847ba8649df5372c6e21f106d16,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-4e4ba6e493eb,research_summary,d4c2672a737449c7cb3116987dda0ba43376969fcc91dc4480cdeaaec0f87f26,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-4fd4a0288258,research_summary,c877e237a4a8996372ee2d8e0bd45717680c02f3861108b630110fb775473280,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-50747b367f5d,research_summary,779de4b7676426fb7b200d56a400bebbd19d6c0f1efb939828277c2c1a765247,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-5293ffb3e55e,research_summary,01ddc93b5b51902130a88fac9e234d5434f6cf7f4fc2f81b8219bae5931e98b5,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-5343ef0e0a90,research_summary,9ccb8e7ce5f879412c1f37e08b01ffd7f8f049a568a6f2f04d361bcf540f7226,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-554e5e45515f,research_summary,affa29ae0d37010f82da64b640bf9c24b4a399d7eef972563c9a35707bd41249,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,rejected +sc-573469fad85a,research_summary,f92830e4955adb97d7789688281cbe79151b85ddeb4f689e1b87bf0c662a20aa,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-5815e4f8e247,research_summary,5f00742ab41542851ea7c6bb378043787eef72f286d933d8deabb3e047572244,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-58ea25123f9d,research_summary,dee6f91c0cf7ac455909ecfdf5ed6a7af12fe6a1786b0fa247069525ce379dc2,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-59338be2cc93,research_summary,6e6fc3975238c943f8383791bf15d41467e2bcb69c55ea074b920055d4433068,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-5e87ea97e1ab,research_summary,1c08c4f0a1e99856bda699d5e00c954896946423bf543b809b18618d9145c606,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-5eb36404cead,research_summary,2b7a76cbc367c08330b993bdbeec0f22e6b7b40cac39dc8902cb471a1cae8d88,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-5eecde57a7e9,research_summary,41cdc71f0009e7c3f177a286cafde6d0831d6e39eb986a9b225705f28337a702,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-619d47c2b2f6,research_summary,3531ab9a3f9be4395991ace60d558f571f30d798faf160d357b60617b30c3825,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-633108b86a26,research_summary,c5dd776309ad7fb7aab0bb43755c06ab3f6ae7766e0cf28a363538c4ec8b4285,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-668653b22fa8,research_summary,36fe7a40324155bc2e8531652f52d731880dc8f5b4a266d66a933fe4613c3fc0,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-6690ea870e56,research_summary,861c7dd30faf8f304424f132d2f107c86196542eba6bb6a9d6a917972d26ad5d,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-66d221735c3f,research_summary,b9d375a69d9952f7fd8d0efcb1e59c1314b7494d667cf022485037f41e8db0ff,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-670c0bddd691,research_summary,3d1a22f88be2914b385cfcf7d65b148081fe6cbc367acb65e8d782457448aaab,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-6932d3513da1,research_summary,3a2d5e73aafe6bd2814854f75fae9d79cc3e3085c883aa5f75df145425966e76,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-6a41569ec051,research_summary,5b5c66ad8898d38b773591f52bbb1aca59cebfb5a7ad74598631aacb8a241f9a,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-6ad63dcdc29d,research_summary,128977bb30bb638326f871ee62df9132831e9d9d0216f1b1f7feb75ad06c5bc0,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-6bb906ad7218,research_summary,bf6812dd7507daf9fac480973705b942fb95f8d7bea800ec671fdf17760eb115,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-6c2a8c2fa19e,research_summary,771736587ce25f978e0172bdc58ad15db7e2bee155b55eaa04f62e1a5ae94cff,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-6cd02c4fc1c1,research_summary,b5b9cdebce4d2c3992c292a2b7ce9e0dc1f6a5531132efb9ff8805dfe280fccf,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-6cf84b7db7a2,research_summary,31c4cb60d9e999ef0e782e52b033b8c774e52207fbdc359c23a7dbbbbf319758,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-6d6d69069a7f,research_summary,ed46d68de452056936f5ae24a7a88c8e709d625ff9838827a0d7278c1ec8beba,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-6e1d75f47b55,research_summary,a15db8d8df67e2dff800939f833299490416ca3bcd231d50dbd1ed8803100bbc,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-6f3294041cc1,research_summary,79e7a7259cf7e36ba23f8ffcbc1f0c0a669c2300f740ad9550b7071338533db5,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-6fc0ad2e9d97,research_summary,1404380a7c8f7e15f9aee3bf240ae1d0272c295c8c2eec88fea141de859993a8,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-70a0d47430d6,research_summary,f9a8388ce8b46a4144467d9130763e712cfc2579de10ba169b320994ba2c180f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-71214e23d402,research_summary,0c781aa8cb379d392cacd1517d5fd605ac09f6494be926e06ef92f7f9331315e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-728dfb136a8a,research_summary,3f79e661e531de605e57c7eae1ed0d28712fdf5070a3303a51cd11b98c45a112,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-72d42e62cdc6,research_summary,3267c89e8386f89236777351c7f541cadb45a57e2b67b74a694c8c64530a0988,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-730cd6d51bd1,research_summary,779de4b7676426fb7b200d56a400bebbd19d6c0f1efb939828277c2c1a765247,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-734a08d5281e,research_summary,210fdd7bec320b34952c12f2e19a5b6cb9a0d3f44946573e7669b4f98204ec50,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-74cc3c4a595c,research_summary,70528347ad467df8f03197f78d2173ff7af74493d6b57188615da0bbcd8e5c44,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-74db4609c7a9,research_summary,6650925da933f2ffc3f9dab8f9266218a308a8a1fb34450da087adee70b47540,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-75f47dcf8d9c,research_summary,4bf8d39a1e38875116d2fac2b8b79c7b5ef127ef2be423ddc6029bd2c6d81d25,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-77c00dc710e2,research_summary,7883420f682cefe570709bf3f3ae5b1b9d2e7c89a292533c3b2cc187f3e9871a,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-7887c53ee718,research_summary,81739c46a1d9568f45e67f04d0d6ca6038eaaa386e791970c9ca22ddcd616398,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-79de44a00fee,research_summary,16e1ff90de5f3f98a4733e63c67beb6503592ee0b9361f2a5c4e66ae7ee21ed7,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-7adf05153b42,research_summary,2a29b5c22d83e4c9ac415b8b5663cafedabdb53c89d7e8372bbdb557235b2701,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-7c5b8a9c5536,research_summary,173d72b8c2a89e460fb9f905fee37bce6e61b4dcc654c395fdee8e13a22de672,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-7c94fc7a0196,research_summary,c2ade4a742b1fef9655982d865c7a215ee8f69e1e0662ce4e4d8c4b984aa4051,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-7c9d343877a6,research_summary,1ceac066fac6ce403155a4eb625286f4b11d68edc15c98c184e781f9c07cb743,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-7d1dd5e6fcbd,research_summary,c719f42d749801ce90cedcd7d0473abb6895a8bfe03da6ab590d7469d037bf20,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-7d8c8b16e7fc,research_summary,b7730b368dae93bfa0e327cf62fe9bfc45b6d38571b49654fc54b8933c02feb8,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-7d9bc5f5e66c,research_summary,f065d1919d2112e1ae8a542e9857e3a22f706d5c159184714f54cd76cc178160,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-7eb2a0a289b9,research_summary,bbfd01b4ccbc1dec9967c4277ba568c14ba50a5463bb9387882d079969f25b83,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-7f4823b086ad,research_summary,18a1561edc78b3029d85aef6235b7038c00ea420a52dacbcffff93eef9b29d1b,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-7f6c4ef22f85,research_summary,0b3aadf9601fb01b25d1de0d527ab86afbbfe559f2068dc1714624b76df08cd7,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-81eace82fee7,research_summary,b92dd1b7b8cee3736c0fa11225615f5e83772836dd679aaf003bba8cc7ccbe7f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-8202e57eff36,research_summary,e9e17c38bede211d3c4c8dc08e73b996470340b3b45e307a8a8adffd6abd1b4f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-8252343129c2,research_summary,9a483a7a2bdaacbe30af89170f222aea55ffa86524f90d997ac5a9cc9e26a80d,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-832530c66057,research_summary,1a0a6825e8b1f9a0310877e534ad25252d0905be8fa6981c846aec265e97a707,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-8328256db783,research_summary,684488da7dda60d59cc529a2dcf5ac267e7a6e4c0f3be87c6a7fd0d6d363d374,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-835f4e01b78e,research_summary,8ef9f110b4598c9a630032519e5fee3011887f1562fadc691f5c406c841d5ae3,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-8487eef049db,research_summary,dfaa9d7fea88195a2f901f93890ac2236368e58aae9be0b4be78d91b370ede2f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-8560e3933c36,research_summary,7b9f2cd14ce10d6372bef6ec0f7aba507b9b42805942024a767da2f3b131e86f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-856abf20234e,research_summary,9c6b7800ab987d02ffc45c16578b336ff79c3e8e7dc4801946bcc9b6826cee0d,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-856e7a60cdcd,research_summary,e72fdc21fb141684fb603e456d20e2801bd728c28c4a29d5d79a083805a073f6,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-85eb1335aede,research_summary,8416b794d93f2a9430411c226b3e5cf9f9450240d050a9f6eb9ee7898b2203fa,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-86ca16f9bb46,research_summary,8c06d24ce068d3082c704b6f7a25416870e701b717e31fa223ceaf664278c0b1,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-879dbbbcac93,research_summary,effda7127299156c40d7c553b2480f2f0c9443014b8d41a8eaefda20b1f5c082,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-8bcf5b176a81,research_summary,f51820bd1c2caf2157d5697fbdadad88452372100d1935d8e0d7c588bd72b05d,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-8c2c9d757e48,research_summary,f31ae376b79807690fd1c8c959bb730c9cca06cbc5987069d0016c733a526195,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-8d801d5b870e,research_summary,211e269bd267768441c540a7f958adbd19b3f3f6b100bdeb1e96c2c657384952,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-8ea02b185c19,research_summary,2c2113370ddc61137091a19a241204ebcd4b83974979329c46c6c2437daef1eb,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-90ac6d618b7b,research_summary,845c598e91675be1009fdcd029f66803da534b3d081d3da8db3de20c0bea248e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-90d7dddbac41,research_summary,b154142ccffed32c83ef74c07c766e977958d2f01d3fece6a74521bafcd6ba4f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-919d16f1e63f,research_summary,af0d10fbf0dca77d29b06d8bd8aaf5482d3cfade4c56c0754834ae3c6ad963fc,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-91e1ebfb2171,research_summary,7ce784f868f5222af4ac94a49ae453c9c1cf07c9f4998befb88313f987c42a85,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-92bac05bf45c,research_summary,ac82b9366d4bef990108ce5f8bad4955b47c583394e2960daeb21a34e980185e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-92cba9842842,research_summary,6a413235ec3182b776ae88c0183b311a6a205987f481ead25d12dc53c746e9c7,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-92d87108bf26,research_summary,f847bc220d329fb0fd55f4cdbc16d417fdac2d28e44153db544c7ca636f05088,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-945e3181de3a,research_summary,391b53098b6cb98381b221224a257cf63370a69d15be9dd07a7ab8ed739fd546,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-9645b7ad4203,research_summary,42cc6a857bfd432a03b37f2bec5f3d9003ed6275c517f1e1baa547303d392fad,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-96b1aea1b618,research_summary,4df70100a17927981d24983c470f6cba8738655e7d1f5ecab475f71839778a8c,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-97e581083a26,research_summary,391b53098b6cb98381b221224a257cf63370a69d15be9dd07a7ab8ed739fd546,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-9855612538a0,research_summary,20d0451281232792e387fede5a1b3cf14a5c952926c2c09a9080cec2e50b7e64,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-9906b31850f1,research_summary,fadaa957551037fedff5bb6142367fed360683b54be97f9a3f7621282f107bb0,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-991cd5a94091,research_summary,b7cf218ea36fc9b24b2370d85bdd17f7edc91e75fa264afda2b1b4289d1d6b0a,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-991d099d7935,research_summary,996eb0cec79bf133210cea8b0e02a5788cbb722bb201b4f97cc8548139753383,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-99c42e420e3f,research_summary,7e5c52648e72f79cb9e346caec0c0efd817522b9308bf5e7f6cf0caa750beed9,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-9ac73104c904,research_summary,af69b94c0e33c863cba2cebcb20a86b648fd0dc29409f45a4c2a5915baa593d8,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-9b1c2c058448,research_summary,0613b24bff4f153f2d3895f8397a181bd0c6df66d8404acd9476ab044a22ff19,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-9b801242a21b,research_summary,53a82205beae7f756a662f97f92560f3e10a8c161e5d79001797b173a62917e0,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-9bd026197028,research_summary,2d0c6f562bead3ff26e683bd36e6f8074fd1cab8039702dc9f433e69b2c10429,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-9c570887850c,research_summary,a19b9e6b788e5501713d7d8296da8f810a6ded91139a48d1a5fa7d2fde265b44,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-9dfe0b22bad6,research_summary,374dc2da7592623333c430f76d01523780c7732019b43e56c6ff104cdd0ad33b,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-9ed9c4d428ab,research_summary,f3482cdd7ffc57137ae1ea66a56e21233629b9130ed46890921a8baae00969e9,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-a13988e07332,research_summary,68b7930f43ff8621bfecc0cbda2b19cb1423e6ad2eb82a522b0838e1748d3333,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-a271a8ebacae,research_summary,eef620b565898efe760dc90ce5a8f8024a70fa1417a81d70c316bfa8f4400c8e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-a3b0b993f709,research_summary,2c6eec729728f458497cdffccd9981099225da86ada964ee93fec16a2986cd92,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-a546f7a4467a,research_summary,f99f7ff74692762b1280fc7fa3f20e302ab7a8915c78d5c18305486ff358805d,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-a65245502801,research_summary,bd55c7ba76241f16c27cdf4578b19ce35c1dffb4aa3a9a3cfc738ec03931e97a,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,rejected +sc-a6c1c56e8c0e,research_summary,38be7b8ee4638cb5b5eb6144c0a46d51ff7dfc7d05fbc993f066ada4a91ce9d6,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-a8495317b783,research_summary,0d444aeec3d145c41e62e4f2fe62a5f92388f5802bf88f9afd7e1ba46ba0bb6a,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-aaab2f69a5c1,research_summary,4903c48d250fae39f0c430930901f07b8ff51539a133ad22302e2f32cf35fd67,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-ad1e5e17edc0,research_summary,3685ab39e3b73350c07d15e0b770542054c62b3f1f3f686e608c1cd90806b104,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-ae30f1136126,research_summary,53a15d7ce9a564d71ccec7c6fc589dc320a4ed988bf9408cf00de1c20d048bb5,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-b12683a80490,research_summary,8b187aaf963fa8ce6a638d03cad4b21269688b9cb120b41866c20ef9153ca8a1,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-b27ceddaabda,research_summary,7f4ec8baa52d6c4b542153cbf2de0882d59491a02c0f17ff9d2a3522626d071f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-b2fe664a11c9,research_summary,023d1e3141c9e80825a2a890949efd48d8c0917de09bbb295fa17526269370a9,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-b339b3b4f6fc,research_summary,31455d79da87285a4588a799923b390e3689bb54c1026aa3dd00c35771b2dcf4,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-b37666b02883,research_summary,e83957cacd33adc30203d7406baeccd362776cb8f6c4e4bbbbad8ede6b74817f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-b46e14e7de53,research_summary,141ed895853d40bb212e8f5e25058d97952bf3131236d88ff5cefdf24397f7d2,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-b48d958ac385,research_summary,24d305b202e118a9dcb1d18609ca6d7363f9a713841960bc3e86839332524f60,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-b4c3e5bd3dc8,research_summary,acddbb25d92438c46f25a1ad3412e9eff55cd04563bc3c5110caabe5bb195651,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-b4f532b115a4,research_summary,b28978ba4228049462dfbd727e02181b1aed4bca074e458ba1df89480691402e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-b52479efff6a,research_summary,be7bd5d48a76c37ae2e2e02e4b0743134ab491335db79b8bf3aab10f52f6eb3c,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-b5a736bbbf3c,research_summary,c0a5f95b3cb4ab95818e7efb27a6b4bb117db0bc647a98a6aa1e5f58a7bbbd52,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-b855c9140617,research_summary,8ee8037f386c19bbce0d1680ebb598b161a564c31c2361127c3a2b83d7c70624,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-b93c03bf06d2,research_summary,dac5734267097f38712035da4d2406243143d216a7164a987571f8ba642058df,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-ba0d303f9838,research_summary,bbf506f7e62bf6433855f3f7730fde829c6abda531dbe029785534be2887654e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-ba818ec5daf1,research_summary,b73204480b37d6d3824a39fe70da359789149791f64d11336e3ba43b652bd005,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-bcd46cf18076,research_summary,751cc19174ab65fbc5b91b43bfe8d4f45daac6124b9658162ca4f4caa2b98061,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-bcd7110c536d,research_summary,f766f75409f6651a5ea4fb5cd370bbe1acb6af0e8ff2b5cadeac6ccef8f8806c,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-bfaad52755f6,research_summary,765553f722991d141d2e0ec5b91b73d50a577f1a7fb9eaf6d88d6748cd8f965b,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-bffa8ceffd0e,research_summary,b38c5e9c1f080d11bd6b2bc7cc528e9c0356fb5c33d19c873529a98c94d17bbe,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c0676f67077f,research_summary,4b73124fcc4f6c8834b992a1e5d0e606d974780f1cb76a03ffb2f56a10e44cb0,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c09fd2a58e63,research_summary,ae5501aa3d92dcca86c86e52ec1390bb3339d447270f9ec3c1c4d629ef8995cd,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c108bfa40b03,research_summary,634e443f448984dcdf4aed647a26858eda25d2f7099b01fd33fb4c4e4f8fe97f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c10efb1ce3d5,research_summary,2565f4ed3438e884c3deac41a42d97aee245915677256cc4cf3d3cec8582d0ab,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c11395392128,research_summary,0c4b04e1f9f2942dfaf6a2a12d20ca885222d36bc8f0bbfe1090e42cb9bd029a,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-c13dd2c571dc,research_summary,e632e8f1a81b296628905caa5ce2495383a666659f928a9e7398e58eaccc7c4e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c218ebe7c140,research_summary,e71e711ad73a8cce2789e2cf472a0525aad6febd1f4fc56a4045a174271a9573,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c3eb60920f61,research_summary,1c8893e9185c72da9617cb8133fe91a9261ca252f1f52917283f250f6cb0ba50,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c4a67e0cf79a,research_summary,d5d653864f5810564bdc8ebcd42f762c8e81b4fe83ef39659c82bb0578b8d88f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c4d582e32978,research_summary,52ff1c18be20e1214a57c32d2b14d3bc3164c3e4aacf2b36796fa00bfb6d86a7,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c5207ef0df45,research_summary,4dcf8369ea88de2bb3139aabb81bb7e5dd61e4d97b3197c0b20b0085c141c889,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-c607e8acc18f,research_summary,4870dea94575e3b9edf9c39f729e7993146faefca1664f7e557c4f7f5fabb1a4,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c6fd68ad3aa1,research_summary,ab11fe505a7b1794b582f8d65b1f97313c3b42c7f11be6fcfcc98e744d618e9e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c7c4c2041f88,research_summary,c6bce172487a6914c6190ac736f076cc15b49589e8b09a49d965cc150f757d7a,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c92402aae05f,research_summary,c01c8d13fe8c4539efa4a98242ded967bbf444b0d0cb63985efdbfb3dbdd5309,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c94150f71f1f,research_summary,c0a5f95b3cb4ab95818e7efb27a6b4bb117db0bc647a98a6aa1e5f58a7bbbd52,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-c9e3071825a4,research_summary,71f309707644e4a27ff8ddd9681b5bd88732b03c27e5f6cbf2f77c1736720d53,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-ca174a50bf27,research_summary,5e339c4f02d5b91bd492304290f524d74c10534e8c6ad509355829b436049b45,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-caae0e7fdd2d,research_summary,17359813ae2243308f77277cc7eae90c5dde5cf427998d0a3f551acc64fae33f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-cac273f692ab,research_summary,c0a5f95b3cb4ab95818e7efb27a6b4bb117db0bc647a98a6aa1e5f58a7bbbd52,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-cb03d6eb4359,research_summary,34e6412c047048c7c20817f2fa0239ef886097fbb78d873649e4390e7ff4879f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-cbb7f4d54c1c,research_summary,822b5c35747f51dffdfdd00c93eff03dc4a0feb088a667074b606bc10a178e53,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-cc8bd6470ece,research_summary,f348774822007b6093e6c6dfd32ddb935b9a3b308291072ebc08ed628450cae7,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-cd8bbf36394e,research_summary,9d1792949b1563c34e1ca210501759274fe687a7355ef0c23ebf946cd87b6ce3,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-ce1afb129beb,research_summary,4f5536594c564a8aa2b8f7ec0527ebd2bd26a24ca8d06bc5d018c6247defb8a4,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-cf1bb309315d,research_summary,f5f85abf319a118f5fba437cf5dde31276df8bbc0143dd5df27859a5f5ff552b,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-d2450cacf9b6,research_summary,afcd77e372c26d267ada490ba0db9a7ce779ef46e1e7f2746973586c08727b87,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-d35e99ad4363,research_summary,18f8ffff30a5a4cc3b9d25863f2bb55305e8763655cee729f247ba5d5d471b7c,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-d521c7cdaa97,research_summary,b7730b368dae93bfa0e327cf62fe9bfc45b6d38571b49654fc54b8933c02feb8,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-d77117eedc40,research_summary,60ebb2928cfc0d2fb76474c27865763bacb4f044c45baf5751cb1e0a9c65d3c4,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-d78c5b1dc1d0,research_summary,073586d4df89db0a0294d5401275ddc5d5de7f516a639caebb510a3ceb4596c7,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-d94cb2f71ce3,research_summary,3dd2bd9bc2b48af4238b5a142fda7d6d3f05616ecf149c7319d944736d2bf3c3,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-daf63d7f1fff,research_summary,3272eb45d1b12bfc8b20d8d7426ddd0373bb939ff88dce6b8b0423a45cd6fe78,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-dbd8a4047bed,research_summary,637b80b04731fb65966e788e747749a1d480812412e321e67f804eab55858828,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-dc6252fb95cb,research_summary,43c24acf44503dde0b0cf37c8c3d0b7e27e89d3008eb6cff2b67ddabb6c2dca0,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-dcca2c7c2821,research_summary,48ec861b46307a56e2e18d3ac91642841a76c4df36c00d5c2ce40dba8b63dcef,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-dd56d93c4db8,research_summary,3685ab39e3b73350c07d15e0b770542054c62b3f1f3f686e608c1cd90806b104,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-de3e7c0a8784,research_summary,c5056699dff22194c8182e7eb18e2244bc2ded1682d89f68789849d0c827d99f,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-de765dfd1bab,research_summary,6f2aef306df6a78bf07aa437753e7ef9abd4a3606525ee64d0865215e8869417,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-de8b5c7e9f23,research_summary,62956b28721b1738b4ef80362abb7e19b0aa70a221d8aede37f61c6797a3940a,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-df903e5a3d01,research_summary,699f71677f2441200fa633c239946fb22fe57e212acc6d0f62bfc77e74ca8e70,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-e518d430ec7f,research_summary,d9d87544b4ba7d4751b5ddf324ac6af67cb0178e80b2c72ca1f0fa4e54c87aed,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-e557881820bb,research_summary,b48c08bc3baba2bd2ac4089891a1faaf972de578d91601a63ffe634d4ecb99fe,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-e8bb5de922cc,research_summary,fb19df53b1692576da96b73ec4b936e37846e83579f2634b4d504b094914a145,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-eab9c0d47ac4,research_summary,37dbb4ea3eeae3bff9fe9f5b0d2204528e6ebcbfe87ef7a2b1c00abb06411ffe,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,rejected +sc-eba81ca4c4b6,research_summary,ff746b154f520b502d9228bb432d2583469544c2aa3a6d1a97247b4f580dcef7,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-ec0b520f715d,research_summary,ad3bfdae8ae08ac26afcb05859d3443fec97d7d7df94edec5b002e147b40609d,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-ed68892668d7,research_summary,668bf7e9fccf887d7477f555272e68d70ae62f2ddfbc3a9ad62524e06b17ca01,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-ee8c3534bc49,research_summary,3579ecb66700ca125cedae6723ab9b76e47668d39602389a13373cc8f1b486de,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-effff2d7367c,research_summary,6bd49357238d637207e17ee5cb3a30ac1b3ce631db4ae0670d6c2c0f29cbf957,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-f0f257d1b718,research_summary,57332ab94937875bd53490d09060e2d5fa2a462ce097fbdbca20b24145e9d6f7,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-f18b747a87e5,research_summary,393ce4eacb8614408c41e11a57c10753453c13f2518ec184f6c583e4a3c18848,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-f24286d38554,research_summary,bafaa9de51864a663262098a96fda2672797971e0d9b5bf4f3f89c8bfd89b2c9,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-f274b94ae60e,research_summary,f918196a101565541e3bf05122381d33d647f6448ca0f0b376af9af62fe29a5a,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-f2bbdbe5e5de,research_summary,1819088a09ef6e9cf2ef5ad8a77704914391a036bd062572faada577a05c9834,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-f2fa7bec2b2f,research_summary,37d7f9f1ca7cacca0c45f35a72092de2f078a8b0c6a70b46fb5d3c66ad9373ed,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-f3246beee305,research_summary,a095b72543f773fb8231e92c41ca1bb3628ae33d1f83f0d490ae70fb38060488,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-f39228266fae,research_summary,e1622895c7ac59a20eb7d0a8982e62caf589dc67c3ed6d16c67bf5804b852de0,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-f4633f5391ff,research_summary,902146033d7c641ce91f85a44ddc51cde41e8c0efe837a92dc921c7aa47c7611,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-f52b70678883,research_summary,b651c77a6530f9587af3466828b6d40a92d10627b40d66e08a3466497fa30273,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-f6b33d5ab782,research_summary,0e6776d237e290d5169fbe779372888018cdf67d358293b2c4e97b4cbda8d3de,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-f73fcb28995b,research_summary,e0ecf55f7d8d7b13d3945b1047886f8c0c1574bd8016fda853901f38a151b08e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-f7cdead36c0d,research_summary,96fbd0b0187e4e7a7532775b967692d51d365e6fbd7832eb06528a9ff82254a3,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-fac311ffaf6d,research_summary,d67bd3304d08d0e735fbb3d3eb073bb61209c5b4458c912b9995cf3f65c9d2ed,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-fb23b494abe8,research_summary,1698c7b130f542e32644717e6a3d3c8137475a8fb71ed4a2984cdcc148607547,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-fb3520534757,research_summary,993c7f0cca2809cd94a56fd3ba21a32d219211ea4202ccf1dd50a6dd7bdf1af1,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-fbc3fa6e37a7,research_summary,b5b8df27b9b4691bc225d0b0dbd26f4ebcccf977435ef8bafccf119ea2b1c5f9,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,review_required +sc-fc09d78ea413,research_summary,55f4caab382d0a05f1fd48d458713c4d97a7dd086ab2f10fbc621f8737fce09d,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-fcdfc9c9b6d4,research_summary,f7268d618e4e1d001b2976709f84196b25c6ec2a2931143c18d0eab2b99d7daa,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-fd3b2b2b3a3a,research_summary,6b700c7063ed6a0b62ed7f3a5418e0b78c81bfff593ae338e45a97a68ea6641e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-fd438ca2c872,research_summary,e2df13cf28cb0dca80590ff755dad89fe8a5788f2e75714b0fe1a1b17a9bf1f1,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-fe6cf5b8c028,research_summary,74c48bb12074111d8d63fa89b661e3a9348ac1fc8cf7efb0b2bdf53531ee0e1e,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-fef6c4941874,research_summary,b7730b368dae93bfa0e327cf62fe9bfc45b6d38571b49654fc54b8933c02feb8,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted +sc-ff6089326323,research_summary,3135f36fe62fea6207089a222cf097955dd8a7a154c1bf13d79a442989d8c350,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,llm-feedback-loop,admitted diff --git a/evaluation/llm_benchmark/results/celltype-audit/sample/audit_manifest.json b/evaluation/llm_benchmark/results/celltype-audit/sample/audit_manifest.json new file mode 100644 index 0000000..91176e5 --- /dev/null +++ b/evaluation/llm_benchmark/results/celltype-audit/sample/audit_manifest.json @@ -0,0 +1,367 @@ +{ + "audit": "bioevidence audit sample: keep from the auditors", + "validator_version": "0.8.0", + "method": "llm-feedback-loop", + "requested_use": "research_summary", + "profile_id": "singlecell-celltype", + "seed": 20261006, + "predictions_sha256": "25a597594ab7a9c8a0517e90abaf6732573f00b3406a7a895b881f9026e2af68", + "population": { + "admitted": 255, + "rejected": 3, + "review_required": 18 + }, + "sample": { + "admitted": 59, + "controls": 10 + }, + "confidence": 0.95, + "if_no_errors_admitted_error_below": 0.049508, + "sampled": { + "sc-00a40a840007": { + "route": "review_required", + "record_sha256": "0611f7cec9d37b055231413531e1befb17272f7a2f6ead20946e3cd081161842", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-00ad508b5979": { + "route": "admitted", + "record_sha256": "7c1d935bda9f312817ca35e84513fc48551e186242fdb8c0cc04f0a58d9d5bf0", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-01004e474ad5": { + "route": "admitted", + "record_sha256": "3cfb6d6aa30d20d951385b016a1348a5daa34144f47ecd49ff83b86d0d331fc5", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-01ab2950a674": { + "route": "review_required", + "record_sha256": "ad96d5f842fd994e6b36ad3a8b750803dc0ee9198b20004fcdff3550f0144d6b", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-06c333c4660a": { + "route": "admitted", + "record_sha256": "0ceb3c532e8dff710feca6d264986161ad274f1aabe0caeee26a93cdd6f705cc", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-080a4ad54b3b": { + "route": "admitted", + "record_sha256": "51533adf0d1ec9ee8affa7fbfdab0d08da75d6cf2a554e421d16649b91c759f0", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-0e9dd71166b2": { + "route": "admitted", + "record_sha256": "aa9554ff83c256ac3fffd2f5a5c2e88b1cc5f943eda5ec75063d3df90a695eca", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-0ed3dae30cb0": { + "route": "admitted", + "record_sha256": "1bc41867170d8bf55c13476c5bfd294fbd88bf0fb3d26f7545aafe63a1cdafdd", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-1bb81802ec74": { + "route": "admitted", + "record_sha256": "0e93d1c9adf4dc389eb871d7e3982d6b8d9bac2521a60bf67185a8319910e9fe", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-1bd2021fc890": { + "route": "admitted", + "record_sha256": "72220776a8460cad0b99a0990e4e409ffb88f140ab8504c739c77eb33dacb9fa", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-1ecb98228b96": { + "route": "admitted", + "record_sha256": "9322381342fad881e0e73ad6d35ab969fc7b9cfde109e501bc8725cc0418d8c2", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-2138b2030213": { + "route": "admitted", + "record_sha256": "ae99cf1ac3504977ea7502d5f5c4b140bdd4f137b896b47a6ad1491979a1efa9", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-2258691ddec9": { + "route": "admitted", + "record_sha256": "5a4779afa2cd62890d0c29bb83e05415ce40b4c915496eccad12f599178d2054", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-264c388c143f": { + "route": "admitted", + "record_sha256": "6781f580a06fb4ecc2dc3af1d831c135e931dd10184e39febfe36ec8b814fa44", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-26e57da30416": { + "route": "admitted", + "record_sha256": "e921db952a6e408cdead4ab59e363ebd2e06573d490b434ff5307f4fa761bc87", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-28afeb3a02f1": { + "route": "admitted", + "record_sha256": "f320aa538194d0f478775cd69d49a5460d6e5c213a817f1529dd5b1d03a13d00", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-2c1b4befc878": { + "route": "admitted", + "record_sha256": "0a58922afef83897a5ae960e2bae3c5ab30c5ecd3ed7d908dbd67872ca3fdf8f", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-2d06b761bbe6": { + "route": "admitted", + "record_sha256": "9a81cceac78844c5ef5156cf0666940402f94941c3bd805788cc4c432884425a", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-2eff7bb75fad": { + "route": "admitted", + "record_sha256": "621d4e81f8cea24ee1b7bf5ec16642aeb14042281998979c6d08fb1b20bd0d07", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-3440a71c316e": { + "route": "admitted", + "record_sha256": "43de0c39755de5852f9048867ed093b1969abe7dbd3555272c5f14177651e4cf", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-37758b97b56f": { + "route": "admitted", + "record_sha256": "4b5faca597cf488ea5ba0987388a6c050f8f9a2134b476b8dad227e21c37e099", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-3e4827f1cf8e": { + "route": "admitted", + "record_sha256": "46c6d2ac0be5e4ef8f09ff50dea04d791baff09d454ea9cdd2b2b1033419b8d0", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-3e8815615c78": { + "route": "admitted", + "record_sha256": "204fbd911d94f7e79c007890be794cb0b7e4de2149c5a65d628e5204a6d3c15e", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-42f3008ea7cd": { + "route": "admitted", + "record_sha256": "e6fa8ef66a64479a65e796ff86e32713aa8df0f81a75c460d6e99cc40003aace", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-43489d3f91fb": { + "route": "admitted", + "record_sha256": "ea8d77cbff3b8717daabadc19de23b8c44105e1dfd225760713b9777e7ebfdae", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-457686ee8397": { + "route": "admitted", + "record_sha256": "9b58d12cf37dc801382c820a1bdb4cd14a42f126f3997a3c5dd565d6cda58955", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-47865daed309": { + "route": "admitted", + "record_sha256": "07e51b6bdecbc8c001d641151be60aef83f85f8dca46744aef17a777790779fc", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-4937458f58f8": { + "route": "admitted", + "record_sha256": "4f9898ad50bb99d5365ec89e9d01290b0eb826a02b04293418b27ac944a1e748", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-4d6f03e0ceaf": { + "route": "admitted", + "record_sha256": "9e0a1068a7ba05e9f107663f7bcbaf083ae1b5b9db0715a5af6b8e65e0367430", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-554e5e45515f": { + "route": "rejected", + "record_sha256": "affa29ae0d37010f82da64b640bf9c24b4a399d7eef972563c9a35707bd41249", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-66d221735c3f": { + "route": "admitted", + "record_sha256": "b9d375a69d9952f7fd8d0efcb1e59c1314b7494d667cf022485037f41e8db0ff", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-6a41569ec051": { + "route": "admitted", + "record_sha256": "5b5c66ad8898d38b773591f52bbb1aca59cebfb5a7ad74598631aacb8a241f9a", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-6cf84b7db7a2": { + "route": "admitted", + "record_sha256": "31c4cb60d9e999ef0e782e52b033b8c774e52207fbdc359c23a7dbbbbf319758", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-734a08d5281e": { + "route": "admitted", + "record_sha256": "210fdd7bec320b34952c12f2e19a5b6cb9a0d3f44946573e7669b4f98204ec50", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-74db4609c7a9": { + "route": "admitted", + "record_sha256": "6650925da933f2ffc3f9dab8f9266218a308a8a1fb34450da087adee70b47540", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-7c5b8a9c5536": { + "route": "admitted", + "record_sha256": "173d72b8c2a89e460fb9f905fee37bce6e61b4dcc654c395fdee8e13a22de672", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-7c94fc7a0196": { + "route": "admitted", + "record_sha256": "c2ade4a742b1fef9655982d865c7a215ee8f69e1e0662ce4e4d8c4b984aa4051", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-7f4823b086ad": { + "route": "admitted", + "record_sha256": "18a1561edc78b3029d85aef6235b7038c00ea420a52dacbcffff93eef9b29d1b", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-81eace82fee7": { + "route": "admitted", + "record_sha256": "b92dd1b7b8cee3736c0fa11225615f5e83772836dd679aaf003bba8cc7ccbe7f", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-8202e57eff36": { + "route": "admitted", + "record_sha256": "e9e17c38bede211d3c4c8dc08e73b996470340b3b45e307a8a8adffd6abd1b4f", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-8328256db783": { + "route": "review_required", + "record_sha256": "684488da7dda60d59cc529a2dcf5ac267e7a6e4c0f3be87c6a7fd0d6d363d374", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-835f4e01b78e": { + "route": "admitted", + "record_sha256": "8ef9f110b4598c9a630032519e5fee3011887f1562fadc691f5c406c841d5ae3", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-856e7a60cdcd": { + "route": "admitted", + "record_sha256": "e72fdc21fb141684fb603e456d20e2801bd728c28c4a29d5d79a083805a073f6", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-8d801d5b870e": { + "route": "admitted", + "record_sha256": "211e269bd267768441c540a7f958adbd19b3f3f6b100bdeb1e96c2c657384952", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-919d16f1e63f": { + "route": "admitted", + "record_sha256": "af0d10fbf0dca77d29b06d8bd8aaf5482d3cfade4c56c0754834ae3c6ad963fc", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-92d87108bf26": { + "route": "admitted", + "record_sha256": "f847bc220d329fb0fd55f4cdbc16d417fdac2d28e44153db544c7ca636f05088", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-9bd026197028": { + "route": "admitted", + "record_sha256": "2d0c6f562bead3ff26e683bd36e6f8074fd1cab8039702dc9f433e69b2c10429", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-a271a8ebacae": { + "route": "admitted", + "record_sha256": "eef620b565898efe760dc90ce5a8f8024a70fa1417a81d70c316bfa8f4400c8e", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-ae30f1136126": { + "route": "admitted", + "record_sha256": "53a15d7ce9a564d71ccec7c6fc589dc320a4ed988bf9408cf00de1c20d048bb5", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-b27ceddaabda": { + "route": "admitted", + "record_sha256": "7f4ec8baa52d6c4b542153cbf2de0882d59491a02c0f17ff9d2a3522626d071f", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-b855c9140617": { + "route": "review_required", + "record_sha256": "8ee8037f386c19bbce0d1680ebb598b161a564c31c2361127c3a2b83d7c70624", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-b93c03bf06d2": { + "route": "admitted", + "record_sha256": "dac5734267097f38712035da4d2406243143d216a7164a987571f8ba642058df", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-ba0d303f9838": { + "route": "admitted", + "record_sha256": "bbf506f7e62bf6433855f3f7730fde829c6abda531dbe029785534be2887654e", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-bffa8ceffd0e": { + "route": "admitted", + "record_sha256": "b38c5e9c1f080d11bd6b2bc7cc528e9c0356fb5c33d19c873529a98c94d17bbe", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-c09fd2a58e63": { + "route": "admitted", + "record_sha256": "ae5501aa3d92dcca86c86e52ec1390bb3339d447270f9ec3c1c4d629ef8995cd", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-c11395392128": { + "route": "review_required", + "record_sha256": "0c4b04e1f9f2942dfaf6a2a12d20ca885222d36bc8f0bbfe1090e42cb9bd029a", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-c218ebe7c140": { + "route": "admitted", + "record_sha256": "e71e711ad73a8cce2789e2cf472a0525aad6febd1f4fc56a4045a174271a9573", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-c3eb60920f61": { + "route": "admitted", + "record_sha256": "1c8893e9185c72da9617cb8133fe91a9261ca252f1f52917283f250f6cb0ba50", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-c5207ef0df45": { + "route": "review_required", + "record_sha256": "4dcf8369ea88de2bb3139aabb81bb7e5dd61e4d97b3197c0b20b0085c141c889", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-c7c4c2041f88": { + "route": "admitted", + "record_sha256": "c6bce172487a6914c6190ac736f076cc15b49589e8b09a49d965cc150f757d7a", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-c92402aae05f": { + "route": "admitted", + "record_sha256": "c01c8d13fe8c4539efa4a98242ded967bbf444b0d0cb63985efdbfb3dbdd5309", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-dc6252fb95cb": { + "route": "admitted", + "record_sha256": "43c24acf44503dde0b0cf37c8c3d0b7e27e89d3008eb6cff2b67ddabb6c2dca0", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-dd56d93c4db8": { + "route": "review_required", + "record_sha256": "3685ab39e3b73350c07d15e0b770542054c62b3f1f3f686e608c1cd90806b104", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-eab9c0d47ac4": { + "route": "rejected", + "record_sha256": "37dbb4ea3eeae3bff9fe9f5b0d2204528e6ebcbfe87ef7a2b1c00abb06411ffe", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-f0f257d1b718": { + "route": "admitted", + "record_sha256": "57332ab94937875bd53490d09060e2d5fa2a462ce097fbdbca20b24145e9d6f7", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-f6b33d5ab782": { + "route": "admitted", + "record_sha256": "0e6776d237e290d5169fbe779372888018cdf67d358293b2c4e97b4cbda8d3de", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-fbc3fa6e37a7": { + "route": "review_required", + "record_sha256": "b5b8df27b9b4691bc225d0b0dbd26f4ebcccf977435ef8bafccf119ea2b1c5f9", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-fd438ca2c872": { + "route": "admitted", + "record_sha256": "e2df13cf28cb0dca80590ff755dad89fe8a5788f2e75714b0fe1a1b17a9bf1f1", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + }, + "sc-fef6c4941874": { + "route": "admitted", + "record_sha256": "b7730b368dae93bfa0e327cf62fe9bfc45b6d38571b49654fc54b8933c02feb8", + "profile_sha256": "a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d" + } + } +} diff --git a/evaluation/llm_benchmark/results/celltype-audit/sample/audit_sheet.csv b/evaluation/llm_benchmark/results/celltype-audit/sample/audit_sheet.csv new file mode 100644 index 0000000..1ddbb98 --- /dev/null +++ b/evaluation/llm_benchmark/results/celltype-audit/sample/audit_sheet.csv @@ -0,0 +1,70 @@ +annotation_id,case_id,record_sha256,profile_id,profile_sha256,requested_use,group_id,split,reviewer_id,reviewer_qualification,annotated_at,mapping_label,mapping_rationale,admission_label,admission_rationale,evidence_refs_json,later_information_seen,minutes_spent +,sc-ba0d303f9838,bbf506f7e62bf6433855f3f7730fde829c6abda531dbe029785534be2887654e,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-ba0d303f9838,test,,,,,,,,[],, +,sc-2138b2030213,ae99cf1ac3504977ea7502d5f5c4b140bdd4f137b896b47a6ad1491979a1efa9,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-2138b2030213,test,,,,,,,,[],, +,sc-8d801d5b870e,211e269bd267768441c540a7f958adbd19b3f3f6b100bdeb1e96c2c657384952,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-8d801d5b870e,test,,,,,,,,[],, +,sc-2c1b4befc878,0a58922afef83897a5ae960e2bae3c5ab30c5ecd3ed7d908dbd67872ca3fdf8f,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-2c1b4befc878,test,,,,,,,,[],, +,sc-2eff7bb75fad,621d4e81f8cea24ee1b7bf5ec16642aeb14042281998979c6d08fb1b20bd0d07,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-2eff7bb75fad,test,,,,,,,,[],, +,sc-47865daed309,07e51b6bdecbc8c001d641151be60aef83f85f8dca46744aef17a777790779fc,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-47865daed309,test,,,,,,,,[],, +,sc-43489d3f91fb,ea8d77cbff3b8717daabadc19de23b8c44105e1dfd225760713b9777e7ebfdae,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-43489d3f91fb,test,,,,,,,,[],, +,sc-457686ee8397,9b58d12cf37dc801382c820a1bdb4cd14a42f126f3997a3c5dd565d6cda58955,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-457686ee8397,test,,,,,,,,[],, +,sc-1bb81802ec74,0e93d1c9adf4dc389eb871d7e3982d6b8d9bac2521a60bf67185a8319910e9fe,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-1bb81802ec74,test,,,,,,,,[],, +,sc-01ab2950a674,ad96d5f842fd994e6b36ad3a8b750803dc0ee9198b20004fcdff3550f0144d6b,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-01ab2950a674,test,,,,,,,,[],, +,sc-9bd026197028,2d0c6f562bead3ff26e683bd36e6f8074fd1cab8039702dc9f433e69b2c10429,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-9bd026197028,test,,,,,,,,[],, +,sc-66d221735c3f,b9d375a69d9952f7fd8d0efcb1e59c1314b7494d667cf022485037f41e8db0ff,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-66d221735c3f,test,,,,,,,,[],, +,sc-81eace82fee7,b92dd1b7b8cee3736c0fa11225615f5e83772836dd679aaf003bba8cc7ccbe7f,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-81eace82fee7,test,,,,,,,,[],, +,sc-c218ebe7c140,e71e711ad73a8cce2789e2cf472a0525aad6febd1f4fc56a4045a174271a9573,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c218ebe7c140,test,,,,,,,,[],, +,sc-4937458f58f8,4f9898ad50bb99d5365ec89e9d01290b0eb826a02b04293418b27ac944a1e748,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-4937458f58f8,test,,,,,,,,[],, +,sc-f6b33d5ab782,0e6776d237e290d5169fbe779372888018cdf67d358293b2c4e97b4cbda8d3de,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-f6b33d5ab782,test,,,,,,,,[],, +,sc-7c5b8a9c5536,173d72b8c2a89e460fb9f905fee37bce6e61b4dcc654c395fdee8e13a22de672,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-7c5b8a9c5536,test,,,,,,,,[],, +,sc-dc6252fb95cb,43c24acf44503dde0b0cf37c8c3d0b7e27e89d3008eb6cff2b67ddabb6c2dca0,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-dc6252fb95cb,test,,,,,,,,[],, +,sc-fef6c4941874,b7730b368dae93bfa0e327cf62fe9bfc45b6d38571b49654fc54b8933c02feb8,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-fef6c4941874,test,,,,,,,,[],, +,sc-ae30f1136126,53a15d7ce9a564d71ccec7c6fc589dc320a4ed988bf9408cf00de1c20d048bb5,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-ae30f1136126,test,,,,,,,,[],, +,sc-3440a71c316e,43de0c39755de5852f9048867ed093b1969abe7dbd3555272c5f14177651e4cf,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-3440a71c316e,test,,,,,,,,[],, +,sc-bffa8ceffd0e,b38c5e9c1f080d11bd6b2bc7cc528e9c0356fb5c33d19c873529a98c94d17bbe,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-bffa8ceffd0e,test,,,,,,,,[],, +,sc-b855c9140617,8ee8037f386c19bbce0d1680ebb598b161a564c31c2361127c3a2b83d7c70624,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-b855c9140617,test,,,,,,,,[],, +,sc-00a40a840007,0611f7cec9d37b055231413531e1befb17272f7a2f6ead20946e3cd081161842,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-00a40a840007,test,,,,,,,,[],, +,sc-4d6f03e0ceaf,9e0a1068a7ba05e9f107663f7bcbaf083ae1b5b9db0715a5af6b8e65e0367430,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-4d6f03e0ceaf,test,,,,,,,,[],, +,sc-c09fd2a58e63,ae5501aa3d92dcca86c86e52ec1390bb3339d447270f9ec3c1c4d629ef8995cd,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c09fd2a58e63,test,,,,,,,,[],, +,sc-7c94fc7a0196,c2ade4a742b1fef9655982d865c7a215ee8f69e1e0662ce4e4d8c4b984aa4051,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-7c94fc7a0196,test,,,,,,,,[],, +,sc-6cf84b7db7a2,31c4cb60d9e999ef0e782e52b033b8c774e52207fbdc359c23a7dbbbbf319758,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-6cf84b7db7a2,test,,,,,,,,[],, +,sc-92d87108bf26,f847bc220d329fb0fd55f4cdbc16d417fdac2d28e44153db544c7ca636f05088,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-92d87108bf26,test,,,,,,,,[],, +,sc-26e57da30416,e921db952a6e408cdead4ab59e363ebd2e06573d490b434ff5307f4fa761bc87,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-26e57da30416,test,,,,,,,,[],, +,sc-2258691ddec9,5a4779afa2cd62890d0c29bb83e05415ce40b4c915496eccad12f599178d2054,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-2258691ddec9,test,,,,,,,,[],, +,sc-00ad508b5979,7c1d935bda9f312817ca35e84513fc48551e186242fdb8c0cc04f0a58d9d5bf0,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-00ad508b5979,test,,,,,,,,[],, +,sc-a271a8ebacae,eef620b565898efe760dc90ce5a8f8024a70fa1417a81d70c316bfa8f4400c8e,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-a271a8ebacae,test,,,,,,,,[],, +,sc-f0f257d1b718,57332ab94937875bd53490d09060e2d5fa2a462ce097fbdbca20b24145e9d6f7,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-f0f257d1b718,test,,,,,,,,[],, +,sc-2d06b761bbe6,9a81cceac78844c5ef5156cf0666940402f94941c3bd805788cc4c432884425a,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-2d06b761bbe6,test,,,,,,,,[],, +,sc-74db4609c7a9,6650925da933f2ffc3f9dab8f9266218a308a8a1fb34450da087adee70b47540,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-74db4609c7a9,test,,,,,,,,[],, +,sc-554e5e45515f,affa29ae0d37010f82da64b640bf9c24b4a399d7eef972563c9a35707bd41249,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-554e5e45515f,test,,,,,,,,[],, +,sc-c7c4c2041f88,c6bce172487a6914c6190ac736f076cc15b49589e8b09a49d965cc150f757d7a,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c7c4c2041f88,test,,,,,,,,[],, +,sc-b93c03bf06d2,dac5734267097f38712035da4d2406243143d216a7164a987571f8ba642058df,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-b93c03bf06d2,test,,,,,,,,[],, +,sc-0ed3dae30cb0,1bc41867170d8bf55c13476c5bfd294fbd88bf0fb3d26f7545aafe63a1cdafdd,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-0ed3dae30cb0,test,,,,,,,,[],, +,sc-835f4e01b78e,8ef9f110b4598c9a630032519e5fee3011887f1562fadc691f5c406c841d5ae3,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-835f4e01b78e,test,,,,,,,,[],, +,sc-dd56d93c4db8,3685ab39e3b73350c07d15e0b770542054c62b3f1f3f686e608c1cd90806b104,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-dd56d93c4db8,test,,,,,,,,[],, +,sc-1ecb98228b96,9322381342fad881e0e73ad6d35ab969fc7b9cfde109e501bc8725cc0418d8c2,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-1ecb98228b96,test,,,,,,,,[],, +,sc-734a08d5281e,210fdd7bec320b34952c12f2e19a5b6cb9a0d3f44946573e7669b4f98204ec50,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-734a08d5281e,test,,,,,,,,[],, +,sc-c5207ef0df45,4dcf8369ea88de2bb3139aabb81bb7e5dd61e4d97b3197c0b20b0085c141c889,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c5207ef0df45,test,,,,,,,,[],, +,sc-c92402aae05f,c01c8d13fe8c4539efa4a98242ded967bbf444b0d0cb63985efdbfb3dbdd5309,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c92402aae05f,test,,,,,,,,[],, +,sc-42f3008ea7cd,e6fa8ef66a64479a65e796ff86e32713aa8df0f81a75c460d6e99cc40003aace,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-42f3008ea7cd,test,,,,,,,,[],, +,sc-8202e57eff36,e9e17c38bede211d3c4c8dc08e73b996470340b3b45e307a8a8adffd6abd1b4f,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-8202e57eff36,test,,,,,,,,[],, +,sc-919d16f1e63f,af0d10fbf0dca77d29b06d8bd8aaf5482d3cfade4c56c0754834ae3c6ad963fc,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-919d16f1e63f,test,,,,,,,,[],, +,sc-3e8815615c78,204fbd911d94f7e79c007890be794cb0b7e4de2149c5a65d628e5204a6d3c15e,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-3e8815615c78,test,,,,,,,,[],, +,sc-28afeb3a02f1,f320aa538194d0f478775cd69d49a5460d6e5c213a817f1529dd5b1d03a13d00,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-28afeb3a02f1,test,,,,,,,,[],, +,sc-856e7a60cdcd,e72fdc21fb141684fb603e456d20e2801bd728c28c4a29d5d79a083805a073f6,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-856e7a60cdcd,test,,,,,,,,[],, +,sc-6a41569ec051,5b5c66ad8898d38b773591f52bbb1aca59cebfb5a7ad74598631aacb8a241f9a,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-6a41569ec051,test,,,,,,,,[],, +,sc-c3eb60920f61,1c8893e9185c72da9617cb8133fe91a9261ca252f1f52917283f250f6cb0ba50,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c3eb60920f61,test,,,,,,,,[],, +,sc-080a4ad54b3b,51533adf0d1ec9ee8affa7fbfdab0d08da75d6cf2a554e421d16649b91c759f0,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-080a4ad54b3b,test,,,,,,,,[],, +,sc-b27ceddaabda,7f4ec8baa52d6c4b542153cbf2de0882d59491a02c0f17ff9d2a3522626d071f,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-b27ceddaabda,test,,,,,,,,[],, +,sc-3e4827f1cf8e,46c6d2ac0be5e4ef8f09ff50dea04d791baff09d454ea9cdd2b2b1033419b8d0,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-3e4827f1cf8e,test,,,,,,,,[],, +,sc-fbc3fa6e37a7,b5b8df27b9b4691bc225d0b0dbd26f4ebcccf977435ef8bafccf119ea2b1c5f9,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-fbc3fa6e37a7,test,,,,,,,,[],, +,sc-eab9c0d47ac4,37dbb4ea3eeae3bff9fe9f5b0d2204528e6ebcbfe87ef7a2b1c00abb06411ffe,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-eab9c0d47ac4,test,,,,,,,,[],, +,sc-c11395392128,0c4b04e1f9f2942dfaf6a2a12d20ca885222d36bc8f0bbfe1090e42cb9bd029a,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c11395392128,test,,,,,,,,[],, +,sc-1bd2021fc890,72220776a8460cad0b99a0990e4e409ffb88f140ab8504c739c77eb33dacb9fa,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-1bd2021fc890,test,,,,,,,,[],, +,sc-37758b97b56f,4b5faca597cf488ea5ba0987388a6c050f8f9a2134b476b8dad227e21c37e099,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-37758b97b56f,test,,,,,,,,[],, +,sc-0e9dd71166b2,aa9554ff83c256ac3fffd2f5a5c2e88b1cc5f943eda5ec75063d3df90a695eca,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-0e9dd71166b2,test,,,,,,,,[],, +,sc-06c333c4660a,0ceb3c532e8dff710feca6d264986161ad274f1aabe0caeee26a93cdd6f705cc,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-06c333c4660a,test,,,,,,,,[],, +,sc-7f4823b086ad,18a1561edc78b3029d85aef6235b7038c00ea420a52dacbcffff93eef9b29d1b,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-7f4823b086ad,test,,,,,,,,[],, +,sc-fd438ca2c872,e2df13cf28cb0dca80590ff755dad89fe8a5788f2e75714b0fe1a1b17a9bf1f1,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-fd438ca2c872,test,,,,,,,,[],, +,sc-01004e474ad5,3cfb6d6aa30d20d951385b016a1348a5daa34144f47ecd49ff83b86d0d331fc5,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-01004e474ad5,test,,,,,,,,[],, +,sc-8328256db783,684488da7dda60d59cc529a2dcf5ac267e7a6e4c0f3be87c6a7fd0d6d363d374,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-8328256db783,test,,,,,,,,[],, +,sc-264c388c143f,6781f580a06fb4ecc2dc3af1d831c135e931dd10184e39febfe36ec8b814fa44,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-264c388c143f,test,,,,,,,,[],, diff --git a/evaluation/llm_benchmark/results/celltype-audit/stand_in_annotations.csv b/evaluation/llm_benchmark/results/celltype-audit/stand_in_annotations.csv new file mode 100644 index 0000000..ec9fe64 --- /dev/null +++ b/evaluation/llm_benchmark/results/celltype-audit/stand_in_annotations.csv @@ -0,0 +1,70 @@ +annotation_id,case_id,record_sha256,profile_id,profile_sha256,requested_use,group_id,split,reviewer_id,reviewer_qualification,annotated_at,mapping_label,mapping_rationale,admission_label,admission_rationale,evidence_refs_json,later_information_seen,minutes_spent +stand-in:sc-ba0d303f9838,sc-ba0d303f9838,bbf506f7e62bf6433855f3f7730fde829c6abda531dbe029785534be2887654e,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-ba0d303f9838,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000669 (pericyte); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-d9bc81cfff""]",, +stand-in:sc-2138b2030213,sc-2138b2030213,ae99cf1ac3504977ea7502d5f5c4b140bdd4f137b896b47a6ad1491979a1efa9,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-2138b2030213,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0002275 (pancreatic PP cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-a26635e6cd""]",, +stand-in:sc-8d801d5b870e,sc-8d801d5b870e,211e269bd267768441c540a7f958adbd19b3f3f6b100bdeb1e96c2c657384952,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-8d801d5b870e,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:1000849 (kidney distal convoluted tubule epithelial cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-c7858eb46a""]",, +stand-in:sc-2c1b4befc878,sc-2c1b4befc878,0a58922afef83897a5ae960e2bae3c5ab30c5ecd3ed7d908dbd67872ca3fdf8f,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-2c1b4befc878,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0008019 (mesenchymal cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-f7c13b08bb""]",, +stand-in:sc-2eff7bb75fad,sc-2eff7bb75fad,621d4e81f8cea24ee1b7bf5ec16642aeb14042281998979c6d08fb1b20bd0d07,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-2eff7bb75fad,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000232 (erythrocyte); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-73d8984fbc""]",, +stand-in:sc-47865daed309,sc-47865daed309,07e51b6bdecbc8c001d641151be60aef83f85f8dca46744aef17a777790779fc,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-47865daed309,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000236 (B cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-560d5829d2""]",, +stand-in:sc-43489d3f91fb,sc-43489d3f91fb,ea8d77cbff3b8717daabadc19de23b8c44105e1dfd225760713b9777e7ebfdae,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-43489d3f91fb,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0002201 (renal beta-intercalated cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-6807a1005a""]",, +stand-in:sc-457686ee8397,sc-457686ee8397,9b58d12cf37dc801382c820a1bdb4cd14a42f126f3997a3c5dd565d6cda58955,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-457686ee8397,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0002138 (endothelial cell of lymphatic vessel); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-cc1cfcedd3""]",, +stand-in:sc-1bb81802ec74,sc-1bb81802ec74,0e93d1c9adf4dc389eb871d7e3982d6b8d9bac2521a60bf67185a8319910e9fe,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-1bb81802ec74,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0002139 (endothelial cell of vascular tree); the claim is finer,admitted,passes the identifier checks,"[""authors-label:ct-3875326824""]",, +stand-in:sc-01ab2950a674,sc-01ab2950a674,ad96d5f842fd994e6b36ad3a8b750803dc0ee9198b20004fcdff3550f0144d6b,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-01ab2950a674,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000798 (gamma-delta T cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-20d7fe4cf5""]",, +stand-in:sc-9bd026197028,sc-9bd026197028,2d0c6f562bead3ff26e683bd36e6f8074fd1cab8039702dc9f433e69b2c10429,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-9bd026197028,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0005009 (renal principal cell); the claim is finer,admitted,passes the identifier checks,"[""authors-label:ct-54741cc370""]",, +stand-in:sc-66d221735c3f,sc-66d221735c3f,b9d375a69d9952f7fd8d0efcb1e59c1314b7494d667cf022485037f41e8db0ff,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-66d221735c3f,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000235 (macrophage); the claim is finer,admitted,passes the identifier checks,"[""authors-label:ct-c12bd3b011""]",, +stand-in:sc-81eace82fee7,sc-81eace82fee7,b92dd1b7b8cee3736c0fa11225615f5e83772836dd679aaf003bba8cc7ccbe7f,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-81eace82fee7,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000682 (M cell of gut); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-daa29fa7c3""]",, +stand-in:sc-c218ebe7c140,sc-c218ebe7c140,e71e711ad73a8cce2789e2cf472a0525aad6febd1f4fc56a4045a174271a9573,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c218ebe7c140,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000359 (vascular associated smooth muscle cell); the claim is coarser,admitted,passes the identifier checks,"[""authors-label:ct-884f62d6f1""]",, +stand-in:sc-4937458f58f8,sc-4937458f58f8,4f9898ad50bb99d5365ec89e9d01290b0eb826a02b04293418b27ac944a1e748,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-4937458f58f8,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000863 (M1 macrophage); the claim is coarser,admitted,passes the identifier checks,"[""authors-label:ct-6846a2fbd5""]",, +stand-in:sc-f6b33d5ab782,sc-f6b33d5ab782,0e6776d237e290d5169fbe779372888018cdf67d358293b2c4e97b4cbda8d3de,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-f6b33d5ab782,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000312 (keratinocyte); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-506f19166c""]",, +stand-in:sc-7c5b8a9c5536,sc-7c5b8a9c5536,173d72b8c2a89e460fb9f905fee37bce6e61b4dcc654c395fdee8e13a22de672,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-7c5b8a9c5536,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000232 (erythrocyte); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-838e09fefd""]",, +stand-in:sc-dc6252fb95cb,sc-dc6252fb95cb,43c24acf44503dde0b0cf37c8c3d0b7e27e89d3008eb6cff2b67ddabb6c2dca0,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-dc6252fb95cb,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0005011 (renal alpha-intercalated cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-90023ca6ff""]",, +stand-in:sc-fef6c4941874,sc-fef6c4941874,b7730b368dae93bfa0e327cf62fe9bfc45b6d38571b49654fc54b8933c02feb8,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-fef6c4941874,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000236 (B cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-18ce9fb955""]",, +stand-in:sc-ae30f1136126,sc-ae30f1136126,53a15d7ce9a564d71ccec7c6fc589dc320a4ed988bf9408cf00de1c20d048bb5,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-ae30f1136126,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0001071 (group 3 innate lymphoid cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-3ec5acbb2d""]",, +stand-in:sc-3440a71c316e,sc-3440a71c316e,43de0c39755de5852f9048867ed093b1969abe7dbd3555272c5f14177651e4cf,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-3440a71c316e,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000815 (regulatory T cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-f3161764a6""]",, +stand-in:sc-bffa8ceffd0e,sc-bffa8ceffd0e,b38c5e9c1f080d11bd6b2bc7cc528e9c0356fb5c33d19c873529a98c94d17bbe,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-bffa8ceffd0e,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,"authors' term CL:0000939 (CD16-positive, CD56-dim natural killer cell, human); the claim is coarser",admitted,passes the identifier checks,"[""authors-label:ct-2e16bde736""]",, +stand-in:sc-b855c9140617,sc-b855c9140617,8ee8037f386c19bbce0d1680ebb598b161a564c31c2361127c3a2b83d7c70624,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-b855c9140617,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000798 (gamma-delta T cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-20d7fe4cf5""]",, +stand-in:sc-00a40a840007,sc-00a40a840007,0611f7cec9d37b055231413531e1befb17272f7a2f6ead20946e3cd081161842,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-00a40a840007,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000921 (type I NK T cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-1161617d72""]",, +stand-in:sc-4d6f03e0ceaf,sc-4d6f03e0ceaf,9e0a1068a7ba05e9f107663f7bcbaf083ae1b5b9db0715a5af6b8e65e0367430,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-4d6f03e0ceaf,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000669 (pericyte); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-d9bc81cfff""]",, +stand-in:sc-c09fd2a58e63,sc-c09fd2a58e63,ae5501aa3d92dcca86c86e52ec1390bb3339d447270f9ec3c1c4d629ef8995cd,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c09fd2a58e63,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000359 (vascular associated smooth muscle cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-884f62d6f1""]",, +stand-in:sc-7c94fc7a0196,sc-7c94fc7a0196,c2ade4a742b1fef9655982d865c7a215ee8f69e1e0662ce4e4d8c4b984aa4051,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-7c94fc7a0196,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000623 (natural killer cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-def3067099""]",, +stand-in:sc-6cf84b7db7a2,sc-6cf84b7db7a2,31c4cb60d9e999ef0e782e52b033b8c774e52207fbdc359c23a7dbbbbf319758,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-6cf84b7db7a2,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,"authors' term CL:0000625 (CD8-positive, alpha-beta T cell); the claim is finer",admitted,passes the identifier checks,"[""authors-label:ct-6b9f55ce2e""]",, +stand-in:sc-92d87108bf26,sc-92d87108bf26,f847bc220d329fb0fd55f4cdbc16d417fdac2d28e44153db544c7ca636f05088,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-92d87108bf26,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000235 (macrophage); the claim is finer,admitted,passes the identifier checks,"[""authors-label:ct-c12bd3b011""]",, +stand-in:sc-26e57da30416,sc-26e57da30416,e921db952a6e408cdead4ab59e363ebd2e06573d490b434ff5307f4fa761bc87,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-26e57da30416,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000814 (mature NK T cell); the claim is coarser,admitted,passes the identifier checks,"[""authors-label:ct-13d48ee1d6""]",, +stand-in:sc-2258691ddec9,sc-2258691ddec9,5a4779afa2cd62890d0c29bb83e05415ce40b4c915496eccad12f599178d2054,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-2258691ddec9,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0002275 (pancreatic PP cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-a26635e6cd""]",, +stand-in:sc-00ad508b5979,sc-00ad508b5979,7c1d935bda9f312817ca35e84513fc48551e186242fdb8c0cc04f0a58d9d5bf0,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-00ad508b5979,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0002138 (endothelial cell of lymphatic vessel); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-cc1cfcedd3""]",, +stand-in:sc-a271a8ebacae,sc-a271a8ebacae,eef620b565898efe760dc90ce5a8f8024a70fa1417a81d70c316bfa8f4400c8e,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-a271a8ebacae,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0005009 (renal principal cell); the claim is finer,admitted,passes the identifier checks,"[""authors-label:ct-54741cc370""]",, +stand-in:sc-f0f257d1b718,sc-f0f257d1b718,57332ab94937875bd53490d09060e2d5fa2a462ce097fbdbca20b24145e9d6f7,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-f0f257d1b718,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000312 (keratinocyte); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-506f19166c""]",, +stand-in:sc-2d06b761bbe6,sc-2d06b761bbe6,9a81cceac78844c5ef5156cf0666940402f94941c3bd805788cc4c432884425a,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-2d06b761bbe6,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000235 (macrophage); the claim is finer,admitted,passes the identifier checks,"[""authors-label:ct-c12bd3b011""]",, +stand-in:sc-74db4609c7a9,sc-74db4609c7a9,6650925da933f2ffc3f9dab8f9266218a308a8a1fb34450da087adee70b47540,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-74db4609c7a9,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000921 (type I NK T cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-1161617d72""]",, +stand-in:sc-554e5e45515f,sc-554e5e45515f,affa29ae0d37010f82da64b640bf9c24b4a399d7eef972563c9a35707bd41249,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-554e5e45515f,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:1001106 (kidney loop of Henle thick ascending limb epithelial cell); the claim is wrong,rejected,label_mismatch,"[""authors-label:ct-7954bd765f""]",, +stand-in:sc-c7c4c2041f88,sc-c7c4c2041f88,c6bce172487a6914c6190ac736f076cc15b49589e8b09a49d965cc150f757d7a,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c7c4c2041f88,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000084 (T cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-e76b10ee43""]",, +stand-in:sc-b93c03bf06d2,sc-b93c03bf06d2,dac5734267097f38712035da4d2406243143d216a7164a987571f8ba642058df,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-b93c03bf06d2,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000766 (myeloid leukocyte); the claim is finer,admitted,passes the identifier checks,"[""authors-label:ct-1fcb853772""]",, +stand-in:sc-0ed3dae30cb0,sc-0ed3dae30cb0,1bc41867170d8bf55c13476c5bfd294fbd88bf0fb3d26f7545aafe63a1cdafdd,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-0ed3dae30cb0,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:1000428 (stem cell of epidermis); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-5633d2ce11""]",, +stand-in:sc-835f4e01b78e,sc-835f4e01b78e,8ef9f110b4598c9a630032519e5fee3011887f1562fadc691f5c406c841d5ae3,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-835f4e01b78e,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000173 (pancreatic D cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-50ec79996f""]",, +stand-in:sc-dd56d93c4db8,sc-dd56d93c4db8,3685ab39e3b73350c07d15e0b770542054c62b3f1f3f686e608c1cd90806b104,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-dd56d93c4db8,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000798 (gamma-delta T cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-20d7fe4cf5""]",, +stand-in:sc-1ecb98228b96,sc-1ecb98228b96,9322381342fad881e0e73ad6d35ab969fc7b9cfde109e501bc8725cc0418d8c2,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-1ecb98228b96,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:1000849 (kidney distal convoluted tubule epithelial cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-c7858eb46a""]",, +stand-in:sc-734a08d5281e,sc-734a08d5281e,210fdd7bec320b34952c12f2e19a5b6cb9a0d3f44946573e7669b4f98204ec50,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-734a08d5281e,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000312 (keratinocyte); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-506f19166c""]",, +stand-in:sc-c5207ef0df45,sc-c5207ef0df45,4dcf8369ea88de2bb3139aabb81bb7e5dd61e4d97b3197c0b20b0085c141c889,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c5207ef0df45,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000921 (type I NK T cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-1161617d72""]",, +stand-in:sc-c92402aae05f,sc-c92402aae05f,c01c8d13fe8c4539efa4a98242ded967bbf444b0d0cb63985efdbfb3dbdd5309,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c92402aae05f,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0008019 (mesenchymal cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-75516baf75""]",, +stand-in:sc-42f3008ea7cd,sc-42f3008ea7cd,e6fa8ef66a64479a65e796ff86e32713aa8df0f81a75c460d6e99cc40003aace,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-42f3008ea7cd,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0019031 (intestine goblet cell); the claim is coarser,admitted,passes the identifier checks,"[""authors-label:ct-9ce2902282""]",, +stand-in:sc-8202e57eff36,sc-8202e57eff36,e9e17c38bede211d3c4c8dc08e73b996470340b3b45e307a8a8adffd6abd1b4f,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-8202e57eff36,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000235 (macrophage); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-0c19dd5456""]",, +stand-in:sc-919d16f1e63f,sc-919d16f1e63f,af0d10fbf0dca77d29b06d8bd8aaf5482d3cfade4c56c0754834ae3c6ad963fc,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-919d16f1e63f,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000922 (type II NK T cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-52a80b5c6b""]",, +stand-in:sc-3e8815615c78,sc-3e8815615c78,204fbd911d94f7e79c007890be794cb0b7e4de2149c5a65d628e5204a6d3c15e,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-3e8815615c78,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0008019 (mesenchymal cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-75516baf75""]",, +stand-in:sc-28afeb3a02f1,sc-28afeb3a02f1,f320aa538194d0f478775cd69d49a5460d6e5c213a817f1529dd5b1d03a13d00,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-28afeb3a02f1,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000235 (macrophage); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-0c19dd5456""]",, +stand-in:sc-856e7a60cdcd,sc-856e7a60cdcd,e72fdc21fb141684fb603e456d20e2801bd728c28c4a29d5d79a083805a073f6,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-856e7a60cdcd,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0019021 (endothelial cell of periportal hepatic sinusoid); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-4d4de5c13f""]",, +stand-in:sc-6a41569ec051,sc-6a41569ec051,5b5c66ad8898d38b773591f52bbb1aca59cebfb5a7ad74598631aacb8a241f9a,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-6a41569ec051,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0001071 (group 3 innate lymphoid cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-3ec5acbb2d""]",, +stand-in:sc-c3eb60920f61,sc-c3eb60920f61,1c8893e9185c72da9617cb8133fe91a9261ca252f1f52917283f250f6cb0ba50,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c3eb60920f61,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000235 (macrophage); the claim is finer,admitted,passes the identifier checks,"[""authors-label:ct-c12bd3b011""]",, +stand-in:sc-080a4ad54b3b,sc-080a4ad54b3b,51533adf0d1ec9ee8affa7fbfdab0d08da75d6cf2a554e421d16649b91c759f0,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-080a4ad54b3b,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0005009 (renal principal cell); the claim is finer,admitted,passes the identifier checks,"[""authors-label:ct-54741cc370""]",, +stand-in:sc-b27ceddaabda,sc-b27ceddaabda,7f4ec8baa52d6c4b542153cbf2de0882d59491a02c0f17ff9d2a3522626d071f,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-b27ceddaabda,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,"authors' term CL:0000938 (CD16-negative, CD56-bright natural killer cell, human); the claim is coarser",admitted,passes the identifier checks,"[""authors-label:ct-bb928b569a""]",, +stand-in:sc-3e4827f1cf8e,sc-3e4827f1cf8e,46c6d2ac0be5e4ef8f09ff50dea04d791baff09d454ea9cdd2b2b1033419b8d0,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-3e4827f1cf8e,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000173 (pancreatic D cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-50ec79996f""]",, +stand-in:sc-fbc3fa6e37a7,sc-fbc3fa6e37a7,b5b8df27b9b4691bc225d0b0dbd26f4ebcccf977435ef8bafccf119ea2b1c5f9,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-fbc3fa6e37a7,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000798 (gamma-delta T cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-20d7fe4cf5""]",, +stand-in:sc-eab9c0d47ac4,sc-eab9c0d47ac4,37dbb4ea3eeae3bff9fe9f5b0d2204528e6ebcbfe87ef7a2b1c00abb06411ffe,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-eab9c0d47ac4,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000682 (M cell of gut); the claim is wrong,rejected,label_mismatch,"[""authors-label:ct-daa29fa7c3""]",, +stand-in:sc-c11395392128,sc-c11395392128,0c4b04e1f9f2942dfaf6a2a12d20ca885222d36bc8f0bbfe1090e42cb9bd029a,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-c11395392128,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0001071 (group 3 innate lymphoid cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-3ec5acbb2d""]",, +stand-in:sc-1bd2021fc890,sc-1bd2021fc890,72220776a8460cad0b99a0990e4e409ffb88f140ab8504c739c77eb33dacb9fa,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-1bd2021fc890,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000682 (M cell of gut); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-daa29fa7c3""]",, +stand-in:sc-37758b97b56f,sc-37758b97b56f,4b5faca597cf488ea5ba0987388a6c050f8f9a2134b476b8dad227e21c37e099,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-37758b97b56f,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0001071 (group 3 innate lymphoid cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-3ec5acbb2d""]",, +stand-in:sc-0e9dd71166b2,sc-0e9dd71166b2,aa9554ff83c256ac3fffd2f5a5c2e88b1cc5f943eda5ec75063d3df90a695eca,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-0e9dd71166b2,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0002275 (pancreatic PP cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-a26635e6cd""]",, +stand-in:sc-06c333c4660a,sc-06c333c4660a,0ceb3c532e8dff710feca6d264986161ad274f1aabe0caeee26a93cdd6f705cc,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-06c333c4660a,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000084 (T cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-e0c7aa5d8b""]",, +stand-in:sc-7f4823b086ad,sc-7f4823b086ad,18a1561edc78b3029d85aef6235b7038c00ea420a52dacbcffff93eef9b29d1b,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-7f4823b086ad,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000173 (pancreatic D cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-50ec79996f""]",, +stand-in:sc-fd438ca2c872,sc-fd438ca2c872,e2df13cf28cb0dca80590ff755dad89fe8a5788f2e75714b0fe1a1b17a9bf1f1,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-fd438ca2c872,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,"authors' term CL:0000897 (CD4-positive, alpha-beta memory T cell); the claim is wrong",rejected,wrong cell type,"[""authors-label:ct-f5c625b34c""]",, +stand-in:sc-01004e474ad5,sc-01004e474ad5,3cfb6d6aa30d20d951385b016a1348a5daa34144f47ecd49ff83b86d0d331fc5,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-01004e474ad5,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0002201 (renal beta-intercalated cell); the claim is exact,admitted,passes the identifier checks,"[""authors-label:ct-6807a1005a""]",, +stand-in:sc-8328256db783,sc-8328256db783,684488da7dda60d59cc529a2dcf5ac267e7a6e4c0f3be87c6a7fd0d6d363d374,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-8328256db783,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,incorrect,authors' term CL:0000922 (type II NK T cell); the claim is wrong,rejected,wrong cell type,"[""authors-label:ct-52a80b5c6b""]",, +stand-in:sc-264c388c143f,sc-264c388c143f,6781f580a06fb4ecc2dc3af1d831c135e931dd10184e39febfe36ec8b814fa44,singlecell-celltype,a87416b47d0ff42c2d0d6df91760a51103ac423380e89d591ad391bfa93fc11d,research_summary,sc-264c388c143f,test,stand-in-authors-labels,stand-in: the dataset authors' cluster labels (not an expert),2026-10-06T00:00:00Z,correct,authors' term CL:0000766 (myeloid leukocyte); the claim is finer,admitted,passes the identifier checks,"[""authors-label:ct-1fcb853772""]",, diff --git a/evaluation/llm_benchmark/results/celltype-audit/summary.json b/evaluation/llm_benchmark/results/celltype-audit/summary.json new file mode 100644 index 0000000..d2f6ebf --- /dev/null +++ b/evaluation/llm_benchmark/results/celltype-audit/summary.json @@ -0,0 +1,119 @@ +{ + "audit": { + "admitted_error_upper_bound": 0.345827, + "confidence": 0.95, + "expert_minutes_estimated": null, + "method": "llm-feedback-loop", + "not_yet_audited": [], + "population": { + "admitted": 255, + "rejected": 3, + "review_required": 18 + }, + "requested_use": "research_summary", + "routes": { + "admitted": { + "audit_minutes": null, + "audited": 59, + "error_rate": 0.2373, + "error_rate_upper_bound": 0.345827, + "error_rate_wilson_95": [ + 0.146946, + 0.359749 + ], + "errors": 14, + "errors_in_route_at_most": 89, + "expert_minutes_estimated": null, + "population": 255, + "route_name": "auto-admitted", + "share_of_expert_time": null, + "share_of_records": 0.9239, + "uncertain": 0 + }, + "rejected": { + "audit_minutes": null, + "audited": 2, + "error_rate": 0.0, + "error_rate_upper_bound": 0.776394, + "error_rate_wilson_95": [ + 0.0, + 0.65762 + ], + "errors": 0, + "errors_in_route_at_most": 3, + "expert_minutes_estimated": null, + "population": 3, + "route_name": "rejected", + "share_of_expert_time": null, + "share_of_records": 0.0109, + "uncertain": 0 + }, + "review_required": { + "audit_minutes": null, + "audited": 8, + "error_rate": 0.125, + "error_rate_upper_bound": 0.47068, + "error_rate_wilson_95": [ + 0.022417, + 0.470888 + ], + "errors": 1, + "errors_in_route_at_most": 9, + "expert_minutes_estimated": null, + "population": 18, + "route_name": "expert review", + "share_of_expert_time": null, + "share_of_records": 0.0652, + "uncertain": 0 + } + }, + "seed": 20261006 + }, + "census": { + "admitted": { + "audit_error_rate": 0.2373, + "audit_upper_bound": 0.345827, + "audit_wilson_95": [ + 0.146946, + 0.359749 + ], + "census_below_upper_bound": true, + "census_error_rate": 0.1961, + "census_errors": 50, + "census_records": 255, + "census_within_wilson": true + }, + "rejected": { + "audit_error_rate": 0.0, + "audit_upper_bound": 0.776394, + "audit_wilson_95": [ + 0.0, + 0.65762 + ], + "census_below_upper_bound": true, + "census_error_rate": 0.0, + "census_errors": 0, + "census_records": 3, + "census_within_wilson": true + }, + "review_required": { + "audit_error_rate": 0.125, + "audit_upper_bound": 0.47068, + "audit_wilson_95": [ + 0.022417, + 0.470888 + ], + "census_below_upper_bound": true, + "census_error_rate": 0.2778, + "census_errors": 5, + "census_records": 18, + "census_within_wilson": true + } + }, + "controls": 10, + "dry_run": "audit of the single-cell held-out results with a stand-in auditor (#24)", + "episodes_sha256": "1826196d4e1b645b90612b674708c850332ca35f38e141fe47e253cdd10b3fa0", + "seed": 20261006, + "stand_in": "stand-in: the dataset authors' cluster labels (not an expert)", + "target": 0.05 +} diff --git a/evaluation/llm_benchmark/results/celltype-audit/summary.md b/evaluation/llm_benchmark/results/celltype-audit/summary.md new file mode 100644 index 0000000..826ee2c --- /dev/null +++ b/evaluation/llm_benchmark/results/celltype-audit/summary.md @@ -0,0 +1,27 @@ +# Audit dry run: single-cell held-out annotations + +**Not an expert audit.** The auditor is a stand-in: the dataset authors' cluster labels (not an expert). It shows what the audit procedure reports, and checks its statistics against a census. + +| Route | Records | Share of records | Audited | Errors | Error rate | Wilson 95% | Upper bound (95%, one-sided) | Share of expert time | +|---|---:|---:|---:|---:|---:|---|---:|---:| +| auto-admitted | 255 | 92.4% | 59 | 14 | 23.7% | 14.7%–36.0% | 34.6% | – | +| rejected | 3 | 1.09% | 2 | 0 | 0.00% | 0.00%–65.8% | 77.6% | – | +| expert review | 18 | 6.52% | 8 | 1 | 12.5% | 2.24%–47.1% | 47.1% | – | + +An error on the auto-admitted route is a record the reference says should not have been admitted, or whose claim is wrong. On another route, it is a record that could have been admitted as it was. + +**14 error(s) in 59 audited auto-admitted records: the auto-admitted error rate is below 34.6% at 95% confidence**, at most 89 of 255 records. + +## Against the census + +The authors' labels cover every record, so here the error rate of the whole route is known: + +| Route | Census error rate | Audit estimate | Wilson 95% | Upper bound (95%) | Census inside | +|---|---:|---:|---|---:|---| +| auto-admitted | 50/255 (19.6%) | 23.7% | 14.7%–36.0% | 34.6% | yes | +| rejected | 0/3 (0.0%) | 0.0% | 0.0%–65.8% | 77.6% | yes | +| expert review | 5/18 (27.8%) | 12.5% | 2.2%–47.1% | 47.1% | yes | + +The audit was sized to show an auto-admitted error rate below 5% if none of 59 sampled records was wrong. It found 14, so this route does not meet a 5% target: the census rate is 19.6%. An audit is how a deployment would learn that these annotations need review, or a better model, before research summaries rely on them. + +Seed 20261006; `sample/audit_manifest.json` holds the routes and stays with the maintainer. Expert time is not measured: the stand-in records none. diff --git a/evaluation/llm_benchmark/results/risk-signals/summary.json b/evaluation/llm_benchmark/results/risk-signals/summary.json new file mode 100644 index 0000000..f3417f9 --- /dev/null +++ b/evaluation/llm_benchmark/results/risk-signals/summary.json @@ -0,0 +1,491 @@ +{ + "analysis": "risk signals against wrong first answers", + "first_answers": { + "agent-stance-pilot": 80, + "celltype-external": 479, + "celltype-pilot": 138, + "celltype-test": 276, + "semantic-pilot": 68 + }, + "signals": { + "BEV004": { + "evidence": [ + { + "flagged": 62, + "flagged_wrong": 39, + "flagged_wrong_rate": 0.629, + "flagged_wrong_wilson_95": [ + 0.5046, + 0.7384 + ], + "kind": "pilot", + "results": "celltype-pilot", + "risk_ratio": { + "ci_95": [ + 1.57, + 3.65 + ], + "ratio": 2.39 + }, + "unflagged": 76, + "unflagged_wrong": 20, + "unflagged_wrong_rate": 0.2632, + "unflagged_wrong_wilson_95": [ + 0.1773, + 0.3718 + ], + "what": "Single-cell annotation, pilot (protocol 1, with ASCT+B)" + }, + { + "flagged": 17, + "flagged_wrong": 14, + "flagged_wrong_rate": 0.8235, + "flagged_wrong_wilson_95": [ + 0.5897, + 0.9381 + ], + "kind": "held-out", + "results": "celltype-test", + "risk_ratio": { + "ci_95": [ + 1.87, + 3.28 + ], + "ratio": 2.48 + }, + "unflagged": 259, + "unflagged_wrong": 86, + "unflagged_wrong_rate": 0.332, + "unflagged_wrong_wilson_95": [ + 0.2775, + 0.3915 + ], + "what": "Single-cell annotation, held-out clusters (protocol 2)" + }, + { + "flagged": 54, + "flagged_wrong": 38, + "flagged_wrong_rate": 0.7037, + "flagged_wrong_wilson_95": [ + 0.5717, + 0.8086 + ], + "kind": "external", + "results": "celltype-external", + "risk_ratio": { + "ci_95": [ + 1.49, + 2.27 + ], + "ratio": 1.83 + }, + "unflagged": 425, + "unflagged_wrong": 163, + "unflagged_wrong_rate": 0.3835, + "unflagged_wrong_wilson_95": [ + 0.3385, + 0.4306 + ], + "what": "Single-cell annotation, external studies (protocol 3, with definitions)" + }, + { + "flagged": 11, + "flagged_wrong": 10, + "flagged_wrong_rate": 0.9091, + "flagged_wrong_wilson_95": [ + 0.6226, + 0.9838 + ], + "kind": "pilot", + "results": "agent-stance-pilot", + "risk_ratio": { + "ci_95": [ + 3.21, + 10.12 + ], + "ratio": 5.7 + }, + "unflagged": 69, + "unflagged_wrong": 11, + "unflagged_wrong_rate": 0.1594, + "unflagged_wrong_wilson_95": [ + 0.0914, + 0.2633 + ], + "what": "Literature claims, agent with stances (scenario 2b)" + } + ], + "role": "risk", + "routing": "on: engine rule", + "signal": "Contradicting evidence line (the answer's own evidence disagrees with it)", + "verdict": "shown" + }, + "BEV016": { + "evidence": [ + { + "flagged": 4, + "flagged_wrong": 3, + "flagged_wrong_rate": 0.75, + "flagged_wrong_wilson_95": [ + 0.3006, + 0.9544 + ], + "kind": "pilot", + "results": "celltype-pilot", + "risk_ratio": { + "ci_95": [ + 0.98, + 3.27 + ], + "ratio": 1.79 + }, + "unflagged": 134, + "unflagged_wrong": 56, + "unflagged_wrong_rate": 0.4179, + "unflagged_wrong_wilson_95": [ + 0.3378, + 0.5026 + ], + "what": "Single-cell annotation, pilot (protocol 1, with ASCT+B)" + }, + { + "flagged": 3, + "flagged_wrong": 2, + "flagged_wrong_rate": 0.6667, + "flagged_wrong_wilson_95": [ + 0.2077, + 0.9385 + ], + "kind": "held-out", + "results": "celltype-test", + "risk_ratio": { + "ci_95": [ + 0.82, + 4.2 + ], + "ratio": 1.86 + }, + "unflagged": 273, + "unflagged_wrong": 98, + "unflagged_wrong_rate": 0.359, + "unflagged_wrong_wilson_95": [ + 0.3044, + 0.4175 + ], + "what": "Single-cell annotation, held-out clusters (protocol 2)" + } + ], + "role": "fixable", + "routing": "fed back to the model", + "signal": "A cited identifier or marker does not exist in the source", + "verdict": "not shown" + }, + "BEV017": { + "evidence": [ + { + "flagged": 35, + "flagged_wrong": 26, + "flagged_wrong_rate": 0.7429, + "flagged_wrong_wilson_95": [ + 0.5793, + 0.8584 + ], + "kind": "pilot", + "results": "celltype-pilot", + "risk_ratio": { + "ci_95": [ + 1.65, + 3.26 + ], + "ratio": 2.32 + }, + "unflagged": 103, + "unflagged_wrong": 33, + "unflagged_wrong_rate": 0.3204, + "unflagged_wrong_wilson_95": [ + 0.2381, + 0.4156 + ], + "what": "Single-cell annotation, pilot (protocol 1, with ASCT+B)" + }, + { + "flagged": 59, + "flagged_wrong": 48, + "flagged_wrong_rate": 0.8136, + "flagged_wrong_wilson_95": [ + 0.6962, + 0.8926 + ], + "kind": "held-out", + "results": "celltype-test", + "risk_ratio": { + "ci_95": [ + 2.6, + 4.43 + ], + "ratio": 3.4 + }, + "unflagged": 217, + "unflagged_wrong": 52, + "unflagged_wrong_rate": 0.2396, + "unflagged_wrong_wilson_95": [ + 0.1877, + 0.3006 + ], + "what": "Single-cell annotation, held-out clusters (protocol 2)" + }, + { + "flagged": 69, + "flagged_wrong": 49, + "flagged_wrong_rate": 0.7101, + "flagged_wrong_wilson_95": [ + 0.5943, + 0.8038 + ], + "kind": "external", + "results": "celltype-external", + "risk_ratio": { + "ci_95": [ + 1.57, + 2.33 + ], + "ratio": 1.92 + }, + "unflagged": 410, + "unflagged_wrong": 152, + "unflagged_wrong_rate": 0.3707, + "unflagged_wrong_wilson_95": [ + 0.3254, + 0.4185 + ], + "what": "Single-cell annotation, external studies (protocol 3, with definitions)" + }, + { + "flagged": 10, + "flagged_wrong": 4, + "flagged_wrong_rate": 0.4, + "flagged_wrong_wilson_95": [ + 0.1682, + 0.6873 + ], + "kind": "pilot", + "results": "agent-stance-pilot", + "risk_ratio": { + "ci_95": [ + 0.69, + 3.91 + ], + "ratio": 1.65 + }, + "unflagged": 70, + "unflagged_wrong": 17, + "unflagged_wrong_rate": 0.2429, + "unflagged_wrong_wilson_95": [ + 0.1575, + 0.355 + ], + "what": "Literature claims, agent with stances (scenario 2b)" + } + ], + "role": "fixable", + "routing": "fed back to the model", + "signal": "The record disagrees with the source (e.g. a label that is not the term's name)", + "verdict": "shown" + }, + "BEV022": { + "evidence": [ + { + "flagged": 24, + "flagged_wrong": 1, + "flagged_wrong_rate": 0.0417, + "flagged_wrong_wilson_95": [ + 0.0074, + 0.2024 + ], + "kind": "pilot", + "results": "semantic-pilot", + "risk_ratio": { + "ci_95": [ + 0.05, + 3.87 + ], + "ratio": 0.46 + }, + "unflagged": 44, + "unflagged_wrong": 4, + "unflagged_wrong_rate": 0.0909, + "unflagged_wrong_wilson_95": [ + 0.0359, + 0.2116 + ], + "what": "Literature claims, answers with the paper (semantic cues)" + } + ], + "role": "risk", + "routing": "opt-in: `semantic` cues", + "signal": "Semantic cue: negated, hedged or non-human quote", + "verdict": "not shown" + }, + "BEV023": { + "evidence": [ + { + "flagged": 2, + "flagged_wrong": 2, + "flagged_wrong_rate": 1.0, + "flagged_wrong_wilson_95": [ + 0.3424, + 1.0 + ], + "kind": "pilot", + "results": "celltype-pilot", + "risk_ratio": { + "ci_95": [ + 1.15, + 3.42 + ], + "ratio": 1.99 + }, + "unflagged": 136, + "unflagged_wrong": 57, + "unflagged_wrong_rate": 0.4191, + "unflagged_wrong_wilson_95": [ + 0.3395, + 0.5031 + ], + "what": "Single-cell annotation, pilot (protocol 1, with ASCT+B)" + }, + { + "flagged": 1, + "flagged_wrong": 1, + "flagged_wrong_rate": 1.0, + "flagged_wrong_wilson_95": [ + 0.2065, + 1.0 + ], + "kind": "held-out", + "results": "celltype-test", + "risk_ratio": { + "ci_95": [ + 0.92, + 4.7 + ], + "ratio": 2.08 + }, + "unflagged": 275, + "unflagged_wrong": 99, + "unflagged_wrong_rate": 0.36, + "unflagged_wrong_wilson_95": [ + 0.3056, + 0.4183 + ], + "what": "Single-cell annotation, held-out clusters (protocol 2)" + }, + { + "flagged": 1, + "flagged_wrong": 1, + "flagged_wrong_rate": 1.0, + "flagged_wrong_wilson_95": [ + 0.2065, + 1.0 + ], + "kind": "external", + "results": "celltype-external", + "risk_ratio": { + "ci_95": [ + 0.8, + 4.02 + ], + "ratio": 1.79 + }, + "unflagged": 478, + "unflagged_wrong": 200, + "unflagged_wrong_rate": 0.4184, + "unflagged_wrong_wilson_95": [ + 0.375, + 0.4631 + ], + "what": "Single-cell annotation, external studies (protocol 3, with definitions)" + } + ], + "role": "fixable", + "routing": "fed back to the model", + "signal": "An obsolete identifier or a previous gene symbol", + "verdict": "indicative" + }, + "BEV025": { + "evidence": [ + { + "flagged": 23, + "flagged_wrong": 13, + "flagged_wrong_rate": 0.5652, + "flagged_wrong_wilson_95": [ + 0.3681, + 0.7437 + ], + "kind": "pilot", + "results": "celltype-pilot", + "risk_ratio": { + "ci_95": [ + 0.93, + 2.16 + ], + "ratio": 1.41 + }, + "unflagged": 115, + "unflagged_wrong": 46, + "unflagged_wrong_rate": 0.4, + "unflagged_wrong_wilson_95": [ + 0.3151, + 0.4914 + ], + "what": "Single-cell annotation, pilot (protocol 1, with ASCT+B)" + } + ], + "role": "risk", + "routing": "opt-in: `ReferenceGrounder`; dropped from the cell-type validator after the pilot", + "signal": "A pinned reference resource contradicts the claim (ASCT+B)", + "verdict": "not shown" + }, + "BEV026": { + "evidence": [ + { + "flagged": 88, + "flagged_wrong": 41, + "flagged_wrong_rate": 0.4659, + "flagged_wrong_wilson_95": [ + 0.3653, + 0.5694 + ], + "kind": "external", + "results": "celltype-external", + "risk_ratio": { + "ci_95": [ + 0.88, + 1.47 + ], + "ratio": 1.14 + }, + "unflagged": 391, + "unflagged_wrong": 160, + "unflagged_wrong_rate": 0.4092, + "unflagged_wrong_wilson_95": [ + 0.3616, + 0.4586 + ], + "what": "Single-cell annotation, external studies (protocol 3, with definitions)" + } + ], + "role": "risk", + "routing": "opt-in, experimental: `DefinitionGrounder`; fed back to the model", + "signal": "Measurements contradict the ontology's definition", + "verdict": "not shown" + } + }, + "wrong_first_answers": { + "agent-stance-pilot": 21, + "celltype-external": 201, + "celltype-pilot": 59, + "celltype-test": 100, + "semantic-pilot": 5 + } +} diff --git a/evaluation/llm_benchmark/results/risk-signals/summary.md b/evaluation/llm_benchmark/results/risk-signals/summary.md new file mode 100644 index 0000000..40eaeb3 --- /dev/null +++ b/evaluation/llm_benchmark/results/risk-signals/summary.md @@ -0,0 +1,26 @@ +# Risk signals: which routing signals predict a wrong answer + +First answers only, before any feedback. *Shown*: the risk ratio's 95% interval lies above 1 on a held-out or external split, and no such split points the other way. *Indicative*: that holds only on a pilot. The criterion was set after the runs. + +| Signal | Role | Default routing | Split | Wrong when flagged | Wrong otherwise | Risk ratio (95% CI) | Verdict | +|---|---|---|---|---:|---:|---:|---| +| **BEV004** Contradicting evidence line (the answer's own evidence disagrees with it) | risk | on: engine rule | celltype-pilot (pilot) | 39/62 (63%) | 20/76 (26%) | 2.39 (1.57–3.65) | shown | +| | | | celltype-test (held-out) | 14/17 (82%) | 86/259 (33%) | 2.48 (1.87–3.28) | | +| | | | celltype-external (external) | 38/54 (70%) | 163/425 (38%) | 1.83 (1.49–2.27) | | +| | | | agent-stance-pilot (pilot) | 10/11 (91%) | 11/69 (16%) | 5.7 (3.21–10.12) | | +| **BEV022** Semantic cue: negated, hedged or non-human quote | risk | opt-in: `semantic` cues | semantic-pilot (pilot) | 1/24 (4%) | 4/44 (9%) | 0.46 (0.05–3.87) | not shown | +| **BEV025** A pinned reference resource contradicts the claim (ASCT+B) | risk | opt-in: `ReferenceGrounder`; dropped from the cell-type validator after the pilot | celltype-pilot (pilot) | 13/23 (57%) | 46/115 (40%) | 1.41 (0.93–2.16) | not shown | +| **BEV026** Measurements contradict the ontology's definition | risk | opt-in, experimental: `DefinitionGrounder`; fed back to the model | celltype-external (external) | 41/88 (47%) | 160/391 (41%) | 1.14 (0.88–1.47) | not shown | +| **BEV016** A cited identifier or marker does not exist in the source | fixable | fed back to the model | celltype-pilot (pilot) | 3/4 (75%) | 56/134 (42%) | 1.79 (0.98–3.27) | not shown | +| | | | celltype-test (held-out) | 2/3 (67%) | 98/273 (36%) | 1.86 (0.82–4.2) | | +| **BEV017** The record disagrees with the source (e.g. a label that is not the term's name) | fixable | fed back to the model | celltype-pilot (pilot) | 26/35 (74%) | 33/103 (32%) | 2.32 (1.65–3.26) | shown | +| | | | celltype-test (held-out) | 48/59 (81%) | 52/217 (24%) | 3.4 (2.6–4.43) | | +| | | | celltype-external (external) | 49/69 (71%) | 152/410 (37%) | 1.92 (1.57–2.33) | | +| | | | agent-stance-pilot (pilot) | 4/10 (40%) | 17/70 (24%) | 1.65 (0.69–3.91) | | +| **BEV023** An obsolete identifier or a previous gene symbol | fixable | fed back to the model | celltype-pilot (pilot) | 2/2 (100%) | 57/136 (42%) | 1.99 (1.15–3.42) | indicative | +| | | | celltype-test (held-out) | 1/1 (100%) | 99/275 (36%) | 2.08 (0.92–4.7) | | +| | | | celltype-external (external) | 1/1 (100%) | 200/478 (42%) | 1.79 (0.8–4.02) | | + +Risk signals that route by default: BEV004 (shown). Opt-in: BEV022 (not shown), BEV025 (not shown), BEV026 (not shown); a signal that is not shown stays off by default. BEV026 is fed back to the model first; a record it still flags after feedback goes to a person, which is why it is experimental. Fixable findings predict wrong answers too, partly by construction (an unknown identifier is an invalid answer); they go back to the model. + +Rules that route a record because it cannot be verified (BEV006, BEV015, BEV018, BEV020) or because the use requires a person (BEV008–BEV013, BEV021) are admission policy, not predictions, and are not listed. diff --git a/evaluation/llm_benchmark/risk_signals.py b/evaluation/llm_benchmark/risk_signals.py new file mode 100644 index 0000000..18f42c4 --- /dev/null +++ b/evaluation/llm_benchmark/risk_signals.py @@ -0,0 +1,194 @@ +"""Which signals that send a record to a person actually predict a wrong answer (#24). + +Routing should rest on evidence. For every first answer in the committed runs, this compares how often the answer +was wrong when a signal fired and when it did not, with Wilson 95% intervals and a risk ratio (Katz 95% interval). +Wrong means: for single-cell annotation, a cell type the authors' term rules out (wrong or invalid); for the +literature claims, a decision other than the expected one. The criterion was set in this analysis, after the runs +(each run's protocol was frozen before it): a risk signal is *shown* when the risk ratio's interval lies above 1 on +a held-out or external split and no such split points the other way; *indicative* when that holds only on a pilot; +otherwise *not shown*. + + python evaluation/llm_benchmark/risk_signals.py --output evaluation/llm_benchmark/results/risk-signals +""" +from __future__ import annotations + +import argparse +import json +import math +import sys +from pathlib import Path + +ROOT = Path(__file__).resolve().parent +sys.path.insert(0, str(ROOT)) +import celltype_loop # noqa: E402 +import score_claims # noqa: E402 + +RESULTS = ROOT / "results" +Z = 1.959963984540054 + +# Signals that send a record to a person because it may be wrong. Rules that route because a record cannot be +# verified (BEV006, BEV015, BEV018, BEV020) or because the use requires a person (BEV008-BEV013, BEV021) need no +# such evidence: they are admission policy, not predictions. +RISK = { + "BEV004": ("Contradicting evidence line (the answer's own evidence disagrees with it)", "on: engine rule"), + "BEV022": ("Semantic cue: negated, hedged or non-human quote", "opt-in: `semantic` cues"), + "BEV025": ("A pinned reference resource contradicts the claim (ASCT+B)", + "opt-in: `ReferenceGrounder`; dropped from the cell-type validator after the pilot"), + "BEV026": ("Measurements contradict the ontology's definition", + "opt-in, experimental: `DefinitionGrounder`; fed back to the model"), +} +# Fixable findings, fed back to the model rather than routed; shown for comparison. +FIXABLE = { + "BEV016": "A cited identifier or marker does not exist in the source", + "BEV017": "The record disagrees with the source (e.g. a label that is not the term's name)", + "BEV023": "An obsolete identifier or a previous gene symbol", +} +SPLITS = [ # (results, kind, what) + ("celltype-pilot", "pilot", "Single-cell annotation, pilot (protocol 1, with ASCT+B)"), + ("celltype-test", "held-out", "Single-cell annotation, held-out clusters (protocol 2)"), + ("celltype-external", "external", "Single-cell annotation, external studies (protocol 3, with definitions)"), + ("agent-stance-pilot", "pilot", "Literature claims, agent with stances (scenario 2b)"), + ("semantic-pilot", "pilot", "Literature claims, answers with the paper (semantic cues)"), +] + + +def wilson(events: int, n: int) -> list[float] | None: + if n == 0: + return None + p = events / n + centre = (p + Z * Z / (2 * n)) / (1 + Z * Z / n) + half = Z * math.sqrt(p * (1 - p) / n + Z * Z / (4 * n * n)) / (1 + Z * Z / n) + return [round(max(0.0, centre - half), 4), round(min(1.0, centre + half), 4)] + + +def risk_ratio(a: int, n1: int, c: int, n2: int) -> dict: + """Wrong among flagged (a of n1) over wrong among the rest (c of n2), Katz interval; 0.5 added to every cell + when one is empty.""" + if not n1 or not n2: + return {"ratio": None, "ci_95": None} + b, d = n1 - a, n2 - c + if 0 in (a, b, c, d): + a, b, c, d = a + 0.5, b + 0.5, c + 0.5, d + 0.5 + ratio = (a / (a + b)) / (c / (c + d)) + se = math.sqrt(1 / a - 1 / (a + b) + 1 / c - 1 / (c + d)) + return {"ratio": round(ratio, 2), "ci_95": [round(ratio * math.exp(-Z * se), 2), round(ratio * math.exp(Z * se), 2)]} + + +def first_answers(name: str) -> list[tuple[set[str], bool]]: + """(codes on the first record, wrong?) for every first answer the checks saw.""" + out = [] + if name.startswith("celltype"): + case = celltype_loop.Case() + truth = {t["task_id"]: t["term"] for t in celltype_loop.load_jsonl(celltype_loop.SOURCES / "truth.jsonl")} + for row in celltype_loop.load_jsonl(RESULTS / name / "episodes.jsonl"): + if row["attempts"]: + got = [c["answer"] for c in row["calls"] if c["answer"] and c["answer"]["decision"] == "annotate"] + kind = celltype_loop.outcome(case, truth[row["task_id"]], got[0]) + out.append((set(row["attempts"][0]["codes"]), kind in ("wrong", "invalid"))) + elif name == "agent-stance-pilot": + expected = score_claims.expected_directions() + for row in score_claims.load_jsonl(RESULTS / name / "episodes.jsonl"): + first = row["submissions"][0] if row["submissions"] else None + if first and first["decision"] != "stop": + out.append((set(first["codes"]), first["decision"] != expected[row["task_id"]])) + else: # semantic cues: what the cues stopped beyond grounding, from the committed summary + pipeline = json.loads((RESULTS / name / "summary.json").read_text(encoding="utf-8"))["pipeline"] + cues, grounded = pipeline["cues"], pipeline["grounded"] + flagged_right = cues["correct_answers_stopped"]["events"] - grounded["correct_answers_stopped"]["events"] + flagged_wrong = cues["wrong_answers_stopped"]["events"] - grounded["wrong_answers_stopped"]["events"] + right, wrong = cues["correct_answers_stopped"]["n"], cues["wrong_answers_stopped"]["n"] + out += [({"BEV022"}, False)] * flagged_right + [({"BEV022"}, True)] * flagged_wrong + out += [(set(), False)] * (right - flagged_right) + [(set(), True)] * (wrong - flagged_wrong) + return out + + +def table(answers: list[tuple[set[str], bool]], code: str) -> dict: + flagged = [wrong for codes, wrong in answers if code in codes] + rest = [wrong for codes, wrong in answers if code not in codes] + a, c = sum(flagged), sum(rest) + return {"flagged": len(flagged), "flagged_wrong": a, "flagged_wrong_rate": round(a / len(flagged), 4) if flagged else None, + "flagged_wrong_wilson_95": wilson(a, len(flagged)), "unflagged": len(rest), "unflagged_wrong": c, + "unflagged_wrong_rate": round(c / len(rest), 4) if rest else None, + "unflagged_wrong_wilson_95": wilson(c, len(rest)), "risk_ratio": risk_ratio(a, len(flagged), c, len(rest))} + + +def verdict(rows: list[dict]) -> str: + def above(r: dict) -> bool: + ci = r["risk_ratio"]["ci_95"] + return bool(ci) and ci[0] > 1 + + def below(r: dict) -> bool: + ci = r["risk_ratio"]["ci_95"] + return bool(ci) and ci[1] < 1 + + held = [r for r in rows if r["kind"] != "pilot"] + if any(above(r) for r in held) and not any(below(r) for r in held): + return "shown" + if any(above(r) for r in rows) and not any(below(r) for r in rows): + return "indicative" + return "not shown" + + +def analyse() -> dict: + answers = {name: first_answers(name) for name, _, _ in SPLITS} + signals = {} + for code, label in [*((c, v[0]) for c, v in RISK.items()), *FIXABLE.items()]: + rows = [{"results": name, "kind": kind, "what": what, **table(answers[name], code)} + for name, kind, what in SPLITS if any(code in codes for codes, _ in answers[name])] + signals[code] = {"signal": label, "role": "risk" if code in RISK else "fixable", + "routing": RISK[code][1] if code in RISK else "fed back to the model", + "verdict": verdict(rows), "evidence": rows} + return {"analysis": "risk signals against wrong first answers", + "first_answers": {name: len(a) for name, a in answers.items()}, + "wrong_first_answers": {name: sum(w for _, w in a) for name, a in answers.items()}, "signals": signals} + + +def render(summary: dict) -> str: + def pct(events: int, n: int) -> str: + return f"{events}/{n} ({100 * events / n:.0f}%)" if n else "–" + + lines = ["# Risk signals: which routing signals predict a wrong answer", "", + "First answers only, before any feedback. *Shown*: the risk ratio's 95% interval lies above 1 on a " + "held-out or external split, and no such split points the other way. *Indicative*: that holds only on a " + "pilot. The criterion was set after the runs.", "", + "| Signal | Role | Default routing | Split | Wrong when flagged | Wrong otherwise | Risk ratio (95% CI) | " + "Verdict |", "|---|---|---|---|---:|---:|---:|---|"] + for code, s in summary["signals"].items(): + for i, r in enumerate(s["evidence"]): + rr = r["risk_ratio"] + ratio = "–" if rr["ratio"] is None else f"{rr['ratio']} ({rr['ci_95'][0]}–{rr['ci_95'][1]})" + head = (f"**{code}** {s['signal']}", s["role"], s["routing"]) if i == 0 else ("", "", "") + lines.append(f"| {head[0]} | {head[1]} | {head[2]} | {r['results']} ({r['kind']}) | " + f"{pct(r['flagged_wrong'], r['flagged'])} | {pct(r['unflagged_wrong'], r['unflagged'])} | " + f"{ratio} | {s['verdict'] if i == 0 else ''} |") + risk = {c: s for c, s in summary["signals"].items() if s["role"] == "risk"} + + def listed(codes: list[str]) -> str: + return ", ".join(f"{c} ({risk[c]['verdict']})" for c in codes) or "none" + + lines += ["", f"Risk signals that route by default: {listed([c for c, s in risk.items() if s['routing'].startswith('on')])}. " + f"Opt-in: {listed([c for c, s in risk.items() if not s['routing'].startswith('on')])}; a signal that is " + "not shown stays off by default. BEV026 is fed back to the model first; a record it still flags after " + "feedback goes to a person, which is why it is experimental. Fixable findings predict wrong answers too, " + "partly by construction (an unknown identifier is an invalid answer); they go back to the model.", + "", "Rules that route a record because it cannot be verified (BEV006, BEV015, BEV018, BEV020) or because " + "the use requires a person (BEV008–BEV013, BEV021) are admission policy, not predictions, and are not " + "listed."] + return "\n".join(lines) + "\n" + + +def main(argv: list[str] | None = None) -> int: + options = argparse.ArgumentParser(description=__doc__.splitlines()[0]) + options.add_argument("--output", type=Path, required=True) + args = options.parse_args(argv) + summary = analyse() + args.output.mkdir(parents=True, exist_ok=True) + (args.output / "summary.json").write_text(json.dumps(summary, indent=2, sort_keys=True) + "\n", encoding="utf-8", + newline="\n") + (args.output / "summary.md").write_text(render(summary), encoding="utf-8", newline="\n") + print(render(summary)) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/src/bioevidence_validator/audit.py b/src/bioevidence_validator/audit.py new file mode 100644 index 0000000..ff0a945 --- /dev/null +++ b/src/bioevidence_validator/audit.py @@ -0,0 +1,216 @@ +"""Audit sampling: measure the error of what an automated route lets through. + +Routing sends each record one way: admitted automatically, to an expert, or rejected. Experts decide every +routed record, so the error nobody sees is the error of what was admitted. An audit measures it: experts label a +seeded random sample of admitted records with the usual annotation protocol (docs/GOLD_STANDARD.md). A few +records from the other routes can be mixed in, in shuffled order, so that auditors cannot tell which route a record +took. Each route's error rate is then reported with an exact one-sided upper bound: 0 errors in 299 audited +admitted records bounds the admitted error rate below 1% at 95% confidence. + +`audit_sample` returns the auditors' sheet (blank annotation rows) and a manifest. The manifest holds the seed, +each route's population and the route each sampled record took. **Keep the manifest from the auditors:** it +unblinds them. `audit_score` reads the manifest and the resolved labels. + +An admitted record is an error when the reference says it should not have been admitted for the use +(`admission_label` review_required or rejected) or that its claim is wrong (`mapping_label` incorrect). A record +from another route is an error when the reference says it could have been admitted as it was (admitted and +correct): expert time spent for nothing, or a good record blocked. +""" +from __future__ import annotations + +import math +import random +from collections import Counter, defaultdict +from typing import Any + +from . import __version__ +from .review import ANNOTATION_COLUMNS, Unit, wilson + +ROUTES = {"admitted": "auto-admitted", "review_required": "expert review", "rejected": "rejected", + "not_admitted": "not admitted"} +EXPERT_ROUTES = {"review_required"} # every record on these routes is decided by an expert + + +def _log_binomial_cdf(errors: int, n: int, p: float) -> float: + """log P(X <= errors) for X ~ Binomial(n, p), 0 < p < 1.""" + terms = [math.lgamma(n + 1) - math.lgamma(k + 1) - math.lgamma(n - k + 1) + k * math.log(p) + + (n - k) * math.log1p(-p) for k in range(errors + 1)] + top = max(terms) + return top + math.log(sum(math.exp(t - top) for t in terms)) + + +def upper_bound(errors: int, n: int, confidence: float = 0.95) -> float | None: + """Exact (Clopper-Pearson) one-sided upper bound on the error rate after `errors` errors in `n` audited records. + + Sampling without replacement from a finite route makes the true bound a little lower; this one is conservative. + """ + if n <= 0: + return None + if not 0 <= errors <= n or not 0 < confidence < 1: + raise ValueError("need 0 <= errors <= n and 0 < confidence < 1") + if errors == n: + return 1.0 + low, high, alpha = errors / n, 1.0, math.log(1 - confidence) + for _ in range(60): # bisection: P(X <= errors) falls as the rate rises + middle = (low + high) / 2 + if _log_binomial_cdf(errors, n, middle) > alpha: + low = middle + else: + high = middle + return math.ceil(high * 1e6) / 1e6 # rounded up, never below the exact bound + + +def sample_size(target: float, confidence: float = 0.95, errors: int = 0) -> int: + """The smallest audit that bounds the error rate below `target` if it finds at most `errors` errors.""" + if not 0 < target < 1: + raise ValueError("target must be between 0 and 1") + + def enough(n: int) -> bool: + return n > errors and (upper_bound(errors, n, confidence) or 1.0) < target + + high = errors + 1 + while not enough(high): + high *= 2 + if high > 100_000_000: + raise ValueError("target too small to audit") + low = high // 2 + while high - low > 1: + middle = (low + high) // 2 + if enough(middle): + high = middle + else: + low = middle + return high + + +def audit_sample(predictions: list[dict[str, str]], *, method: str, use: str, size: int, seed: int, + profile_id: str, controls: int = 0, split: str = "test", confidence: float = 0.95, + predictions_sha256: str = "") -> tuple[list[dict[str, str]], dict[str, Any]]: + """(auditors' sheet, manifest): a seeded random sample of one method's admitted records for one use, with + `controls` records from its other routes mixed in.""" + rows = sorted((p for p in predictions if p["method"] == method and p["requested_use"] == use), + key=lambda p: p["case_id"]) + if not rows: + raise ValueError(f"No predictions for method {method!r} and use {use!r}") + admitted = [p for p in rows if p["predicted_status"] == "admitted"] + others = [p for p in rows if p["predicted_status"] != "admitted"] + if size <= 0 or not admitted: + raise ValueError("Nothing to audit: the sample size must be positive and some records admitted") + rng = random.Random(seed) + chosen = rng.sample(admitted, min(size, len(admitted))) + rng.sample(others, min(max(controls, 0), len(others))) + rng.shuffle(chosen) + sheet = [{**dict.fromkeys(ANNOTATION_COLUMNS, ""), "case_id": p["case_id"], "record_sha256": p["record_sha256"], + "profile_id": profile_id, "profile_sha256": p["profile_sha256"], "requested_use": use, + "group_id": p["case_id"], "split": split, "evidence_refs_json": "[]"} for p in chosen] + audited = min(size, len(admitted)) + manifest = { + "audit": "bioevidence audit sample: keep from the auditors", "validator_version": __version__, + "method": method, "requested_use": use, "profile_id": profile_id, "seed": seed, + "predictions_sha256": predictions_sha256, + "population": dict(sorted(Counter(p["predicted_status"] for p in rows).items())), + "sample": {"admitted": audited, "controls": len(chosen) - audited}, + "confidence": confidence, "if_no_errors_admitted_error_below": upper_bound(0, audited, confidence), + "sampled": {p["case_id"]: {"route": p["predicted_status"], "record_sha256": p["record_sha256"], + "profile_sha256": p["profile_sha256"]} for p in sorted(chosen, key=lambda p: p["case_id"])}, + } + return sheet, manifest + + +def _is_error(route: str, label: dict[str, str]) -> bool: + if route == "admitted": + return label["admission_label"] != "admitted" or label["mapping_label"] == "incorrect" + return label["admission_label"] == "admitted" and label["mapping_label"] == "correct" + + +def audit_score(manifest: dict[str, Any], final: dict[Unit, dict[str, str]], + annotations: list[dict[str, str]] | None = None, *, confidence: float = 0.95) -> dict[str, Any]: + """Per route: share of records, audited records, errors, error rate with Wilson and exact one-sided bounds, and + expert time when the annotations record `minutes_spent`. + + Expert time per route is estimated as: on expert routes, the mean audited minutes per record times the route's + records (an expert decides each one); on other routes, the audit's own minutes (experts see only the sample). + """ + use, sampled, population = manifest["requested_use"], manifest["sampled"], manifest["population"] + stray = sorted({u for u in final if u[1] == use and u[0] not in sampled}) + if stray: + raise ValueError(f"{len(stray)} labelled record(s) were not in the audit sample, e.g. {stray[0]}") + total = sum(population.values()) + minutes: defaultdict[str, float] = defaultdict(float) + timed: Counter[str] = Counter() + for row in annotations or []: + case = sampled.get(row["case_id"]) + recorded = (row.get("minutes_spent") or "").strip() + if case and row["requested_use"] == use and recorded: + try: + minutes[case["route"]] += float(recorded) + except ValueError: + raise ValueError(f"{row['case_id']}: minutes_spent must be a number, not {recorded!r}") from None + timed[case["route"]] += 1 + routes: dict[str, dict[str, Any]] = {} + pending = [] + for case_id, case in sorted(sampled.items()): + label = final.get((case_id, use)) + if label is None: + pending.append(case_id) + continue + if (label["record_sha256"], label["profile_sha256"]) != (case["record_sha256"], case["profile_sha256"]): + raise ValueError(f"{case_id}: the audited record or profile differs from the sampled one") + stats = routes.setdefault(case["route"], {"audited": 0, "errors": 0, "uncertain": 0}) + stats["audited"] += 1 + stats["errors"] += _is_error(case["route"], label) + stats["uncertain"] += label["mapping_label"] == "uncertain" + workload: dict[str, float | None] = {} + for route, stats in routes.items(): + n, errors, size = stats["audited"], stats["errors"], population.get(route, 0) + bound = upper_bound(errors, n, confidence) + stats.update(route_name=ROUTES.get(route, route), population=size, + share_of_records=round(size / total, 4) if total else None, + error_rate=round(errors / n, 4), error_rate_wilson_95=wilson(errors, n), + error_rate_upper_bound=bound, + errors_in_route_at_most=math.ceil(round(bound * size, 6)) if bound is not None else None, + audit_minutes=round(minutes[route], 1) if timed[route] else None) + if route in EXPERT_ROUTES: + workload[route] = minutes[route] / timed[route] * size if timed[route] else None + else: + workload[route] = minutes[route] if minutes else None + known = [w for w in workload.values() if w is not None] + spent = sum(known) if known and len(known) == len(workload) else None + for route, stats in routes.items(): + estimate = workload[route] + stats["expert_minutes_estimated"] = round(estimate, 1) if estimate is not None else None + stats["share_of_expert_time"] = round(estimate / spent, 4) if spent and estimate is not None else None + admitted = routes.get("admitted", {}) + return {"method": manifest["method"], "requested_use": use, "confidence": confidence, "seed": manifest["seed"], + "population": population, "routes": dict(sorted(routes.items())), "not_yet_audited": pending, + "admitted_error_upper_bound": admitted.get("error_rate_upper_bound"), + "expert_minutes_estimated": round(spent, 1) if spent is not None else None} + + +def render_audit(report: dict[str, Any]) -> str: + def pct(value: float | None) -> str: # two decimals under 10%, where a bound such as 0.997% matters + return "–" if value is None else f"{100 * value:.{2 if value < 0.1 else 1}f}%" + + level = f"{100 * report['confidence']:g}%" + lines = [f"# Audit of `{report['method']}` for {report['requested_use']}", "", + f"| Route | Records | Share of records | Audited | Errors | Error rate | Wilson 95% | Upper bound " + f"({level}, one-sided) | Share of expert time |", + "|---|---:|---:|---:|---:|---:|---|---:|---:|"] + for s in report["routes"].values(): + low_high = s["error_rate_wilson_95"] + interval = "–" if low_high is None else f"{pct(low_high[0])}–{pct(low_high[1])}" + lines.append(f"| {s['route_name']} | {s['population']} | {pct(s['share_of_records'])} | {s['audited']} | " + f"{s['errors']} | {pct(s['error_rate'])} | {interval} | {pct(s['error_rate_upper_bound'])} | " + f"{pct(s['share_of_expert_time'])} |") + lines += ["", "An error on the auto-admitted route is a record the reference says should not have been admitted, " + "or whose claim is wrong. On another route, it is a record that could have been admitted as it was."] + if report["admitted_error_upper_bound"] is not None: + s = report["routes"]["admitted"] + lines += ["", f"**{s['errors']} error(s) in {s['audited']} audited auto-admitted records: the auto-admitted " + f"error rate is below {pct(report['admitted_error_upper_bound'])} at {level} confidence**, at most " + f"{s['errors_in_route_at_most']} of {s['population']} records."] + if report["not_yet_audited"]: + lines += ["", f"{len(report['not_yet_audited'])} sampled record(s) are not yet audited and are left out."] + if report["expert_minutes_estimated"] is not None: + lines += ["", f"Expert time, estimated from the recorded minutes: {report['expert_minutes_estimated']} minutes. " + "Expert routes count every record at the audited mean; other routes count only the audit."] + return "\n".join(lines) + "\n" diff --git a/src/bioevidence_validator/cli.py b/src/bioevidence_validator/cli.py index edde33a..aa89dc6 100644 --- a/src/bioevidence_validator/cli.py +++ b/src/bioevidence_validator/cli.py @@ -9,6 +9,7 @@ import yaml +from . import audit as audit_module from . import review as review_module from .draft import build_record, draft_json_schema, load_draft from .engine import default_schema_path, generate_json_schema, list_profiles, profile_path, validate_record @@ -85,6 +86,29 @@ def parser() -> argparse.ArgumentParser: frozen.add_argument("--frozen-at", required=True, help="ISO 8601 time of the freeze") frozen.add_argument("--min-reviewers", type=int, default=2) frozen.add_argument("--output", type=Path, required=True) + sample = steps.add_parser("audit-sample", help="Seeded random sample of auto-admitted records for an expert audit") + sample.add_argument("--predictions", type=Path, required=True) + sample.add_argument("--method", required=True, help="The method whose routes are audited") + sample.add_argument("--use", required=True, help="The requested use whose admissions are audited") + size = sample.add_mutually_exclusive_group(required=True) + size.add_argument("--size", type=int, help="Admitted records to audit") + size.add_argument("--target", type=float, + help="Audit enough admitted records to bound their error rate below this if none is wrong") + sample.add_argument("--confidence", type=float, default=0.95) + sample.add_argument("--controls", type=int, default=0, + help="Records from the other routes to mix in, so auditors cannot tell the routes apart") + sample.add_argument("--seed", type=int, required=True) + sample.add_argument("--profile-id", required=True) + sample.add_argument("--split", default="test", choices=sorted(review_module.SPLITS)) + sample.add_argument("--output-dir", type=Path, required=True, + help="Writes audit_sheet.csv for the auditors and audit_manifest.json, which stays with you") + audited = steps.add_parser("audit-score", help="Error rate of each route, with an exact upper bound") + audited.add_argument("--manifest", type=Path, required=True) + audited.add_argument("--annotations", type=Path, required=True) + audited.add_argument("--adjudications", type=Path) + audited.add_argument("--min-reviewers", type=int, default=1) + audited.add_argument("--confidence", type=float, default=0.95) + audited.add_argument("--output", type=Path) return result @@ -210,8 +234,31 @@ def _run(args) -> int: return 0 +def _audit_sample(args) -> int: + rv, au = review_module, audit_module + size = args.size if args.size is not None else au.sample_size(args.target, args.confidence) + sheet_path, manifest_path = args.output_dir / "audit_sheet.csv", args.output_dir / "audit_manifest.json" + if sheet_path.exists() or manifest_path.exists(): + raise ValueError(f"{args.output_dir} already holds an audit; refusing to overwrite it") + sheet, manifest = au.audit_sample(rv.load_predictions(args.predictions), method=args.method, use=args.use, + size=size, seed=args.seed, profile_id=args.profile_id, controls=args.controls, + split=args.split, confidence=args.confidence, + predictions_sha256=rv.sha256_file(args.predictions)) + args.output_dir.mkdir(parents=True, exist_ok=True) + rv.write_csv(sheet_path, sheet, rv.ANNOTATION_COLUMNS + rv.OPTIONAL_ANNOTATION_COLUMNS) + _write_json(manifest_path, manifest) + picked = manifest["sample"] + print(f"{len(sheet)} records to audit ({picked['admitted']} auto-admitted, {picked['controls']} controls) in " + f"{sheet_path}. If none of the auto-admitted is wrong, their error rate is below " + f"{100 * manifest['if_no_errors_admitted_error_below']:.2f}% at {100 * args.confidence:g}% " + f"confidence. Keep {manifest_path.name} from the auditors.") + return 0 + + def _review(args) -> int: rv = review_module + if args.step == "audit-sample": + return _audit_sample(args) annotations = rv.load_annotations(args.annotations) if args.step == "check": adjudications = rv.load_adjudications(args.adjudications, annotations) if args.adjudications else [] @@ -232,8 +279,15 @@ def _review(args) -> int: rv.write_csv(args.output, rows, rv.ADJUDICATION_COLUMNS) print(f"{len(rows)} disagreement(s) to adjudicate; {len(incomplete)} unit(s) with too few reviews") return 0 - adjudications = rv.load_adjudications(args.adjudications, annotations) + adjudications = rv.load_adjudications(args.adjudications, annotations) if args.adjudications else [] final = rv.resolve(annotations, adjudications, min_reviewers=args.min_reviewers) + if args.step == "audit-score": + manifest = json.loads(args.manifest.read_text(encoding="utf-8")) + report = audit_module.audit_score(manifest, final, annotations, confidence=args.confidence) + if args.output: + _write_json(args.output, report) + print(audit_module.render_audit(report), end="") + return 0 if args.step == "score": report = rv.score(final, rv.load_predictions(args.predictions), split=args.split) if args.output: diff --git a/tests/test_audit.py b/tests/test_audit.py new file mode 100644 index 0000000..23b1508 --- /dev/null +++ b/tests/test_audit.py @@ -0,0 +1,154 @@ +"""Audit sampling: exact bounds, a seeded blind sample, and per-route error rates (#24).""" +import csv +import json +import math + +import pytest + +from bioevidence_validator import audit, review +from bioevidence_validator.cli import main + +H2 = "b" * 64 + + +def binomial_cdf(errors, n, p): + return sum(math.comb(n, k) * p ** k * (1 - p) ** (n - k) for k in range(errors + 1)) + + +def test_zero_error_bound_has_the_closed_form(): + for n in (1, 10, 59, 299, 300, 1000): + assert audit.upper_bound(0, n) == pytest.approx(1 - 0.05 ** (1 / n), abs=2e-6) + assert audit.upper_bound(0, 299) < 0.01 <= audit.upper_bound(0, 298) + assert audit.upper_bound(0, 0) is None and audit.upper_bound(5, 5) == 1.0 + + +@pytest.mark.parametrize("errors,n", [(1, 300), (3, 59), (14, 59), (50, 255)]) +def test_bound_with_errors_is_exact_and_conservative(errors, n): + bound = audit.upper_bound(errors, n) + assert binomial_cdf(errors, n, bound) <= 0.05 < binomial_cdf(errors, n, bound - 1e-5) + assert bound > errors / n + assert audit.upper_bound(errors, n, 0.99) > bound + + +def test_bound_rejects_impossible_counts(): + for errors, n, confidence in ((3, 2, 0.95), (-1, 5, 0.95), (1, 5, 1.0)): + with pytest.raises(ValueError): + audit.upper_bound(errors, n, confidence) + + +def test_sample_size(): + assert audit.sample_size(0.01) == 299 and audit.sample_size(0.05) == 59 and audit.sample_size(0.10) == 29 + n = audit.sample_size(0.01, errors=1) + assert audit.upper_bound(1, n) < 0.01 <= audit.upper_bound(1, n - 1) + assert audit.sample_size(0.01, confidence=0.99) > 299 + with pytest.raises(ValueError): + audit.sample_size(0) + + +def prediction(case, status, method="loop", use="u"): + digest = f"{int(case[1:]):064x}" + return {"case_id": case, "requested_use": use, "record_sha256": digest, "profile_sha256": H2, + "method": method, "predicted_status": status} + + +PREDICTIONS = ([prediction(f"c{i:03d}", "admitted") for i in range(100)] + + [prediction(f"c{i:03d}", "review_required") for i in range(100, 120)] + + [prediction(f"c{i:03d}", "rejected") for i in range(120, 125)] + + [prediction("c000", "admitted", method="other")]) + + +def test_sample_is_seeded_blind_and_only_of_the_method_and_use(): + sheet, manifest = audit.audit_sample(PREDICTIONS, method="loop", use="u", size=30, seed=7, controls=5, + profile_id="p") + again, _ = audit.audit_sample(list(reversed(PREDICTIONS)), method="loop", use="u", size=30, seed=7, controls=5, + profile_id="p") + other, _ = audit.audit_sample(PREDICTIONS, method="loop", use="u", size=30, seed=8, controls=5, profile_id="p") + assert sheet == again and sheet != other + assert manifest["population"] == {"admitted": 100, "rejected": 5, "review_required": 20} + assert manifest["sample"] == {"admitted": 30, "controls": 5} and len(sheet) == 35 + routes = [manifest["sampled"][row["case_id"]]["route"] for row in sheet] + assert routes.count("admitted") == 30 + assert routes != sorted(routes, key=lambda r: r != "admitted") # controls are shuffled in, not appended + assert set(sheet[0]) == set(review.ANNOTATION_COLUMNS) # nothing on the sheet names the route + assert all(row["mapping_label"] == "" and row["profile_id"] == "p" for row in sheet) + assert manifest["if_no_errors_admitted_error_below"] == audit.upper_bound(0, 30) + + +def test_sample_is_capped_and_refuses_nothing_to_audit(): + _, manifest = audit.audit_sample(PREDICTIONS, method="loop", use="u", size=500, seed=1, controls=500, + profile_id="p") + assert manifest["sample"] == {"admitted": 100, "controls": 25} + with pytest.raises(ValueError, match="No predictions"): + audit.audit_sample(PREDICTIONS, method="loop", use="other", size=5, seed=1, profile_id="p") + with pytest.raises(ValueError, match="Nothing to audit"): + audit.audit_sample([prediction("c1", "rejected")], method="loop", use="u", size=5, seed=1, profile_id="p") + + +def label(row, admission, mapping, minutes=""): + return {**row, "annotation_id": f"a:{row['case_id']}", "reviewer_id": "R1", "reviewer_qualification": "curator", + "annotated_at": "2026-10-06", "admission_label": admission, "mapping_label": mapping, + "evidence_refs_json": '["packet"]', "minutes_spent": minutes} + + +def test_score_counts_errors_per_route_with_bounds_and_expert_time(): + sheet, manifest = audit.audit_sample(PREDICTIONS, method="loop", use="u", size=40, seed=3, controls=10, + profile_id="p") + route = {row["case_id"]: manifest["sampled"][row["case_id"]]["route"] for row in sheet} + rows, wrong_admitted = [], 0 + for row in sheet[:-1]: # the last record is not yet audited + if route[row["case_id"]] == "admitted": + wrong = wrong_admitted < 2 + wrong_admitted += wrong + rows.append(label(row, "rejected" if wrong else "admitted", "incorrect" if wrong else "correct", "2")) + else: # a routed record that could have been admitted as it was is a wasted review + rows.append(label(row, "admitted", "correct", "6")) + final = review.resolve(rows, [], min_reviewers=1) + report = audit.audit_score(manifest, final, rows) + admitted = report["routes"]["admitted"] + assert len(report["not_yet_audited"]) == 1 + assert admitted["errors"] == 2 and admitted["error_rate_upper_bound"] == audit.upper_bound(2, admitted["audited"]) + assert report["admitted_error_upper_bound"] == admitted["error_rate_upper_bound"] + assert admitted["share_of_records"] == 0.8 + assert admitted["errors_in_route_at_most"] == math.ceil(admitted["error_rate_upper_bound"] * 100) + others = [s for r, s in report["routes"].items() if r != "admitted"] + assert all(s["errors"] == s["audited"] for s in others) + # Expert time: the audit's own minutes on the admitted route; every record at the audited mean on review. + assert admitted["expert_minutes_estimated"] == 2 * admitted["audited"] + assert report["routes"]["review_required"]["expert_minutes_estimated"] == 6 * 20 + assert sum(s["share_of_expert_time"] for s in report["routes"].values()) == pytest.approx(1, abs=1e-3) + text = audit.render_audit(report) + assert "auto-admitted error rate is below" in text and "not yet audited" in text + + +def test_score_refuses_a_different_record_or_a_record_not_sampled(): + sheet, manifest = audit.audit_sample(PREDICTIONS, method="loop", use="u", size=5, seed=3, profile_id="p") + changed = [label({**sheet[0], "record_sha256": "f" * 64}, "admitted", "correct")] + with pytest.raises(ValueError, match="differs"): + audit.audit_score(manifest, review.resolve(changed, [], min_reviewers=1)) + stray = [label({**sheet[0], "case_id": "c124"}, "admitted", "correct")] + with pytest.raises(ValueError, match="not in the audit sample"): + audit.audit_score(manifest, review.resolve(stray, [], min_reviewers=1)) + + +def test_cli_audit_sample_and_score(tmp_path, capsys): + predictions = tmp_path / "predictions.csv" + review.write_csv(predictions, PREDICTIONS, review.PREDICTION_COLUMNS) + out = tmp_path / "audit" + args = ["review", "audit-sample", "--predictions", str(predictions), "--method", "loop", "--use", "u", + "--target", "0.05", "--controls", "4", "--seed", "11", "--profile-id", "p", "--output-dir", str(out)] + assert main(args) == 0 + assert "59 auto-admitted, 4 controls" in capsys.readouterr().out + assert main(args) == 3 # never overwrites an audit (input error) + assert "refusing to overwrite" in capsys.readouterr().err + manifest = json.loads((out / "audit_manifest.json").read_text(encoding="utf-8")) + assert manifest["predictions_sha256"] == review.sha256_file(predictions) + with (out / "audit_sheet.csv").open(encoding="utf-8", newline="") as handle: + sheet = list(csv.DictReader(handle)) + filled = [label(row, "admitted", "correct", "3") for row in sheet] + review.write_csv(tmp_path / "labels.csv", filled, review.ANNOTATION_COLUMNS + review.OPTIONAL_ANNOTATION_COLUMNS) + assert main(["review", "audit-score", "--manifest", str(out / "audit_manifest.json"), "--annotations", + str(tmp_path / "labels.csv"), "--output", str(tmp_path / "report.json")]) == 0 + printed = capsys.readouterr().out + assert "0 error(s) in 59 audited auto-admitted records" in printed and "below 4.95%" in printed + report = json.loads((tmp_path / "report.json").read_text(encoding="utf-8")) + assert report["admitted_error_upper_bound"] < 0.05 diff --git a/tools/reproduce.py b/tools/reproduce.py index d6ecd33..4700ab3 100644 --- a/tools/reproduce.py +++ b/tools/reproduce.py @@ -96,6 +96,14 @@ def case(name: str, what: str) -> Step: ["score", "--results", "{out}"], ["episodes.jsonl"], ["summary.json", "summary.md"]), replay("celltype-external", "Scenario 3: single-cell annotation, external studies (protocol 3)", "celltype_loop.py", ["score", "--results", "{out}"], ["episodes.jsonl"], ["summary.json", "summary.md"]), + Step("risk-signals", "Which routing signals predict a wrong answer, from the committed runs", + [py(BENCH / "risk_signals.py", "--output", "{out}")], + outputs={RESULTS / "risk-signals" / n: n for n in ("summary.json", "summary.md")}), + Step("celltype-audit", "Audit dry run of the auto-admitted single-cell annotations, with a stand-in auditor", + [py(BENCH / "audit_celltype.py", "--output", "{out}/run")], + outputs={RESULTS / "celltype-audit" / n: f"run/{n}" for n in ( + "predictions.csv", "sample/audit_sheet.csv", "sample/audit_manifest.json", "audit_packet.md", + "stand_in_annotations.csv", "audit_report.json", "summary.json", "summary.md")}), Step("figures", "Figures in docs/assets, from the committed summaries", [py(REPO / "examples" / "clinvar_germline" / "plot.py", "--output", "{out}/clinvar_germline_benchmark.svg"), py(REPO / "examples" / "vbo_canine" / "plot.py", "--output", "{out}/vbo_canine_benchmark.svg"),