Comparing sequential forecasters via confidence sequences & e-processes
-
Updated
Oct 24, 2023 - Jupyter Notebook
Comparing sequential forecasters via confidence sequences & e-processes
让 prompt 从"感觉好了"变成"真好了" · Where "feels better" becomes "measurably better" for your AI prompts.
Taming False Positives in Out-of-Distribution Detection with Human Feedback (AISTATS '24)
Compute confidence Intervals (CLT, Chebyshev, Hoeffding) and Confidence Sequences (in a sequential setting if needed) in Julia
Sharp review-count and cost bounds with exact certified-uncertainty policies for AI-assisted financial audits
A GxP-purpose-built LLM evaluation framework: anytime-valid validation, judge qualification, ALCOA+ audit, human plus AI workflow validation, drift monitoring, listener hooks, and a high-level facade on Inspect AI. Enables responsible AI use under risk-based assurance.
Always-valid sequential A/B testing engine: mSPRT confidence sequences make peeking safe by construction, CUPED cuts variance up to 49%, SRM gates bad data. Built-in adversarial peeking harness proves the claim: naive daily peeking hit 27.5% false positives in 2,000 simulations; this engine held 1.7%. All numbers reproducible from committed seeds.
Official implementation of On the Tightness and Computational Tractability of Higher-Dimensional Confidence Sequences
Runtime-verified fidelity for long-context LLM serving: detect, label-free, when sparse-attention KV compression silently degrades output, and bound it with anytime-valid confidence sequences. Custom HF attention backend + elastic probe scheduler. H1/H4 confirmed at 7B across Qwen2.5 & Mistral on 2x H100 · 88 tests · pre-registered hypotheses.
Add a description, image, and links to the confidence-sequences topic page so that developers can more easily learn about it.
To associate your repository with the confidence-sequences topic, visit your repo's landing page and select "manage topics."