Agents ship the code
the harness steers the current—
governance, the lock.
demos and case studies: wxl3.com
AI agent that turns regulated enterprise conversations into decision-ready briefs and audited actions. Permission-aware RAG with grounded citations, deterministic policy gates, HITL approval loop, …
TypeScript
An eval and observability cockpit for coding agents. It runs policy-controlled coding agents in sandboxed toy repos, tool-use traces, MCP tools, compares harness policies, scores recovery and safet…
Python
Adversarial Testing Lab for Agentic Safeguards (ATLAS). A synthetic multi-agent eval environment for adversarial fraud decisioning inspired by Anthropic's Project Deal. Measures how model quality, …
Python 1
A regulated-agent deployment kit for turning traces, evals, regressions, and approval gates into launch/no-launch decisions
Python
Can you eval an art form? Canon is a continuity linter for serialized TV, YouTube and micro-drama fiction. Canon plays the role of whats currently the scriptwriting coordinator, verifies your story…
Python
A voice agent demo and prompt evaluation harness for insurance first notice of loss claims
TypeScript