Agent Reliability Engineering: applying SRE principles to AI agent systems. Evals, imp@k metrics, self-improvement, config versioning, transfer experiments.
-
Updated
Mar 26, 2026 - Shell
Agent Reliability Engineering: applying SRE principles to AI agent systems. Evals, imp@k metrics, self-improvement, config versioning, transfer experiments.
A systematic framework for reliable LLM Agent skills in production. Based on 6 months of real-world deployment with OpenClaw.
The ARE Incident Database: an OWASP-ASI-indexed registry of real agent failures, with honest coverage boundaries.
A billing agent that's allowed to move money and mostly decides not to. Gmail in; Stripe, Linear and Slack out.
Evidence discipline for AI-agent systems: an evidence ladder, canaries and counterexamples, and seven verified failure patterns from our own operating cycle. A control is worth nothing until it has been seen to fail. Roadmap items are stated as roadmap, never as implemented.
Measure, monitor, and improve AI agent reliability with SRE-style tools for production agent systems
To associate your repository with the agent-reliability-engineering topic, visit your repo's landing page and select "manage topics."