Framework-native, reproducible prompt-injection benchmark, CI gate, ImpactTwin procurement tests, and attested ProofRun evidence for AI agents.
-
Updated
Aug 13, 2026 - Python
Framework-native, reproducible prompt-injection benchmark, CI gate, ImpactTwin procurement tests, and attested ProofRun evidence for AI agents.
Action-graded severity scoring (L0-L6) for tool-using AI agents, computed from red-team execution traces.
Benchmarking schema-valid false tool observations and defense baselines for tool-using LLM agents.
A local-first research scaffold for evaluating models, agent harnesses, and complete agent products on realistic, stateful tasks.
LangChain-native AgentDojo benchmark: utility + ASR evaluation across banking, slack, travel, and workspace suites.
Evaluating provenance-gated tool calls as a prompt injection defense on AgentDojo. Reproducible runs, per-case analysis, and published results.
Security audit of LLM-based multi-agent systems with indirect prompt-injection PoCs and mitigations.
Personal research project — solo, unaffiliated. Inspect AI evaluation framework for LLM agent security: ASR, benign utility, and Transparency Rate across prompt injection, tool poisoning, and psych attacks.
Reproducible evaluation of deterministic sequence rules for AI-agent tool-call traces.
Add a description, image, and links to the agentdojo topic page so that developers can more easily learn about it.
To associate your repository with the agentdojo topic, visit your repo's landing page and select "manage topics."