Skip to content

injection_guard: a labelled set from InjecAgent and AgentDojo #3

Description

@Yunaik

evals/labelled/injection.jsonl holds 20 items and the injection_guard rail scores 20 of 20 on it. A set that size says little about precision and recall on real attacks.

Build a larger labelled set from the public benchmarks in the same JSONL shape (state.tool, state.text, label, note): InjecAgent (1,054 tool-output injections) and AgentDojo (97 tasks, 629 security cases). Then run the rail on it and add the precision and recall table to docs/benchmarks.md:

uv run s1a run injection_guard --slot jev --labelled-set evals/labelled/injection-public.jsonl

docs/roadmap.md § "Prompt-injection guard" has the plan and the two sources. Needs a Jev key (TYPESAFE_API_KEY or OPENROUTER_API_KEY).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions