Skip to content

Scope check: OutcomeBench / Errata for outcome-level software-agent evaluation #10

Description

@mliotta

Resource

What differs from patch/test benchmarks

The real-software lane does not award credit for a patch or passing tests alone. One unchanged frozen agent must enter at least four outsider-created situations containing misleading evidence and either causally achieve the independently checked outcome or return a truth-checked non-win naming the smallest recovery condition. Admission requires at least 3/4 correct dispositions, zero unauthorized actions, fixed budgets, complete deterministic replay, and comparison with no-exploration, matched-nonlearning, corrupted-information, oracle, briefed, strongest direct, and incumbent arms.

A separate synthetic lane requires outside game authors and an independent custodian, split learner/executor processes, non-enumeration proofs, semantic-corruption falsification, oracle headroom, and independent rendering.

Scope question

Would this fit the collection as an evaluation methodology now, or only after the first independently authored cases and results exist? I am not asking that preparation be listed as a completed benchmark. Mismatch-first protocol criticism and independent author/custodian/adjudicator interest are especially welcome. The exact public recruitment and review entry points are https://github.com/Parslee-ai/errata-external-evaluation/issues .

No hidden case material, evaluator logic, winning actions, secrets, or private repository content should be posted publicly.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions