Resource
What differs from patch/test benchmarks
The real-software lane does not award credit for a patch or passing tests alone. One unchanged frozen agent must enter at least four outsider-created situations containing misleading evidence and either causally achieve the independently checked outcome or return a truth-checked non-win naming the smallest recovery condition. Admission requires at least 3/4 correct dispositions, zero unauthorized actions, fixed budgets, complete deterministic replay, and comparison with no-exploration, matched-nonlearning, corrupted-information, oracle, briefed, strongest direct, and incumbent arms.
A separate synthetic lane requires outside game authors and an independent custodian, split learner/executor processes, non-enumeration proofs, semantic-corruption falsification, oracle headroom, and independent rendering.
Scope question
Would this fit the collection as an evaluation methodology now, or only after the first independently authored cases and results exist? I am not asking that preparation be listed as a completed benchmark. Mismatch-first protocol criticism and independent author/custodian/adjudicator interest are especially welcome. The exact public recruitment and review entry points are https://github.com/Parslee-ai/errata-external-evaluation/issues .
No hidden case material, evaluator logic, winning actions, secrets, or private repository content should be posted publicly.
Resource
What differs from patch/test benchmarks
The real-software lane does not award credit for a patch or passing tests alone. One unchanged frozen agent must enter at least four outsider-created situations containing misleading evidence and either causally achieve the independently checked outcome or return a truth-checked non-win naming the smallest recovery condition. Admission requires at least 3/4 correct dispositions, zero unauthorized actions, fixed budgets, complete deterministic replay, and comparison with no-exploration, matched-nonlearning, corrupted-information, oracle, briefed, strongest direct, and incumbent arms.
A separate synthetic lane requires outside game authors and an independent custodian, split learner/executor processes, non-enumeration proofs, semantic-corruption falsification, oracle headroom, and independent rendering.
Scope question
Would this fit the collection as an evaluation methodology now, or only after the first independently authored cases and results exist? I am not asking that preparation be listed as a completed benchmark. Mismatch-first protocol criticism and independent author/custodian/adjudicator interest are especially welcome. The exact public recruitment and review entry points are https://github.com/Parslee-ai/errata-external-evaluation/issues .
No hidden case material, evaluator logic, winning actions, secrets, or private repository content should be posted publicly.