Repository navigation
Add audit sampling and routing evidence (#24) - #43
Merged
Merged
Conversation
- audit: seeded random sample of auto-admitted records, sized from a target error bound, with blind controls from the other routes; per-route error rates with Wilson and exact one-sided (Clopper-Pearson) bounds, and share of expert time from minutes_spent. - bioevidence review audit-sample / audit-score. - risk_signals.py: whether each signal that routes a record to a person predicts a wrong answer. BEV004 is shown on every split; the opt-in signals BEV022, BEV025 and BEV026 are not. - audit_celltype.py: dry run on the single-cell held-out results with the authors' labels as a stand-in auditor; the census rate lies inside the audit's interval. - reproduce.py gains both steps (20 in all); docs, dossier and changelog.
4 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements #24: audit sampling of auto-admitted records, and evidence for which routing signals predict errors.
What it adds
bioevidence review audit-sample: a seeded random sample of one method's auto-admitted records for one use, written as a blank annotation sheet.--target 0.01sizes the sample so that, if no audited record is wrong, the auto-admitted error rate is below 1% at 95% confidence. That is 299 records; a 2% target needs 149, 5% needs 59 and 10% needs 29.--controlsmixes in records from the other routes in shuffled order, so auditors cannot tell which route a record took.bioevidence review audit-score: reports, for each route:minutes_spent.It refuses labels for records that were not sampled, or that were made on a different record or profile.
Routing evidence (
evaluation/llm_benchmark/risk_signals.py): for each signal that sends a record to a person, how often first answers were wrong when it fired and when it did not.BEV004 is the only risk signal on by default, and it holds on every split. I set the criterion after the runs: the risk ratio's 95% interval must lie above 1 on a held-out or external split.
Audit dry run (
evaluation/llm_benchmark/audit_celltype.py), on the single-cell held-out results. No expert has audited these records. The dry run uses the dataset authors' labels as a stand-in auditor:Done-when in #24
bioevidence review audit-sampledraws a seeded random sample of auto-admitted records.The expert audit itself still needs experts (#23).
Checks
tests/test_audit.pychecks the bounds against the closed form and against the binomial CDF.mkdocs build --strictpasses.tools/reproduce.pynow has 20 steps, and all of them reproduce identically, including the 2 new ones.