I study how to evaluate and supervise AI agents, and build tools that make experiments easier to inspect and reproduce.
- Terminal-Bench: hardness and verifier audit — distinguishing genuine task difficulty from task defects and evaluation failures. Poster accepted at the NeurIPS 2026 Verify-Agents workshop; submitted to AAMAS 2027, not accepted. Related note.
- Factor(U,T) — testing what a trusted monitor can detect from an untrusted model’s plan. Accepted at the AAAI 2026 TrustAgent workshop. Paper · Digest.
- Eval Evidence — evaluation records with configuration, artifacts, and file hashes. Reproducible walkthrough.
Website · All research & projects · CV · Email · LinkedIn · X




