Open-source, 100% reproducible AI Agent Runtime Security Benchmark & Sandbox Environment (RFC-010 Draft Protocol).
-
Updated
Oct 8, 2026 - Python
Open-source, 100% reproducible AI Agent Runtime Security Benchmark & Sandbox Environment (RFC-010 Draft Protocol).
Adversarial security benchmark for agent authorization: does a compromised agent's policy-violating proposal become an unauthorized external effect? 61 trials, nine families, an independent oracle, per-mechanism ablation, confidence intervals. 0 unauthorized effects in 61 attack trials (95% CI [0.0%, 5.9%]). Reproduction is partial.
Open, privacy-bounded assurance for AI agents: containment provenance, identity passports, authorization twins, OTel evidence, CI gates, and OSCAL.
Agent Memory Integrity: a conformance test suite that measures whether AI agent memory and checkpoint stores notice tampering. Sixteen measured rows across LangGraph (SQLite, Postgres, Redis), OpenAI Agents SDK, LlamaIndex, Letta, Mem0 and six integrity tools. IETF draft-khandelwal-bmwg-agent-memory-integrity.
Local-first workbench to run, inspect, compare, report, and gate OpenAI Codex Security scans.
Open deterministic security tests for unsafe multi-agent handoffs and authority escalation.
Deterministic security benchmark for tool-using AI agents
Vendor-neutral benchmark measuring how MCP security proxies/gateways DEFEND against 22+ attack vectors — crosswalked to NIST AI RMF & OWASP LLM/Agentic Top 10. CI-gated, reproducible, DOI-cited. Submit your tool to the leaderboard.
21 passive security evaluation inputs: conventional appsec and exploratory hardware, firmware, PLC, RTOS, database and policy domains. No scans started.
Production-grade microservices security benchmark featuring OWASP Top 10 logic exploits, automated remediation, custom Semgrep SAST rules, and CI/CD DevSecOps gates.
Security evaluation input | Database-engine source, SQL parser, storage, transaction and extension boundaries | production-control-no-vulnerability-claim
Internal PyPI SCA precision and recall benchmark corpus
FreightSkillBench is a reproducible benchmark for evaluating document-to-transaction integrity, prompt-injection risk, and security controls in AI-enabled shipping and logistics workflows.
Open AI-for-security validation benchmark: non-LLM scorer + a SOTA-validation loop. Labeled positive corpus withheld pending coordinated disclosure.
TURNCOAT: an independent, reproducible prompt-injection benchmark for AI coding agents. Open corpus + methodology, responsible disclosure.
ReplayBench-IoT: reproducible IoT replay-defense benchmark with Monte Carlo sweeps, CI, static demo, and hardware-validation artifacts.
Reproducible intentionally vulnerable JavaScript/TypeScript SAST accuracy benchmark
Security evaluation input | Labeled Python injection, deserialization, crypto and trust-boundary cases | labeled-positive-negative
Security evaluation input | GraphQL authorization, queries, mutations, subscriptions, injection and query complexity and resource-exhaustion scenarios | intentional-vulnerabilities
Security evaluation input | Cisco ACLs, zone firewall policy, iptables, topology, routing and rule changes | labeled-examples-in-library
To associate your repository with the security-benchmark topic, visit your repo's landing page and select "manage topics."