Skip to content

Commit 5ad201f

Browse files
mstrathmanclaude
andauthored
feat: agent skills for usage, receipts, distillation, and backtesting (#11)
* feat: agent skills for usage, receipts, distillation, and backtesting Four agentskills.io-format skills in skills/: sqlite-predict (core usage: operation selection, the aggregate convention, statuses vs errors, defaults), prediction-receipts (agent-owned provenance: a _predict_receipts table convention, hashing the result document, replay verification; possible because serving is deterministic and models are content-hashed), distill-lifecycle (verify holdout before serving, drift and re-distill cadence, license discipline), and interpret-backtest (MASE against the naive floor, coverage honesty, conformal judgment, when to narrow auto's pool). README gains an agent skills pointer. Skills carry judgment, not wrappers: the SQL surface is already the API. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: spec-polish the skill frontmatter All four validate with the agentskills reference validator (skills-ref). Add the optional license field and quote metadata values per the spec's string-map recommendation. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: receipts reference script and CI conformance checks skills/prediction-receipts/scripts/receipt.py is the interoperability anchor: one stdlib-only implementation of the convention (record, verify with exit-code semantics, list), with canonicalization defined in one place (aggregate documents hashed verbatim; row results as compact JSON with shortest round-trip floats). The skill now points at it as the preferred path, keeping the hand-rolled steps as the spec. tests/test_skills.py wires both into the normal pytest run: every SKILL.md is held to the agentskills spec (name/dir match, name charset, description bounds and when-to-use, line budget, house dash rule), and the receipt script is proven end to end: record, replay-match, re-record determinism, tamper detection via exit code 2, and registry content_hash pinning for bundled models and distilled students. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: predict_sha256 makes the receipts workflow pure SQL Python was the receipt script's only dependency, and the environments this extension targets (phone, embedded, wasm) may not have it. The lightest dependency is the extension itself: expose the vendored SHA-256 (the hash that already pins model weights) as a deterministic, innocuous predict_sha256(x) scalar (TEXT/BLOB, NULL passthrough), and rewrite the prediction-receipts skill so record and replay-verify run in pure SQL. The Python script stays as optional convenience for row-shaped results (canonical float serialization) and mismatch diagnostics. CI: hash vectors against hashlib, plus a pure-SQL record/verify/tamper round trip; functions reference documents the new scalar. Also gitignores autogluon's local model dumps, which must never be committed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: address all nine CodeRabbit findings on the skills PR Security (critical): receipt.py locks extension loading immediately after loading predict0, so a tampered receipt's replayed input_sql can no longer call load_extension() on an arbitrary library; a regression test replays exactly that attack and asserts sqlite refuses it. Interoperability (major): canonicalization now follows the recorded operation instead of the result shape, so a one-text-cell backtest or predict result hashes as {"columns","rows"} rather than masquerading as an aggregate document; verify() reads the stored operation. Claims discipline (the reviewer enforcing our own path instructions): size/latency claims in distill-lifecycle now cite the benchmarks and state measured ranges; the soft-label rescue is scoped and qualified; the licensing section says obligations 'may' apply, tells users to record the exact license with each student, and states that accept_license is an acknowledgment, not compliance; the 0.57 coverage figure carries its model/dataset scope and source; the conformal fragment became a complete backtest() call; and conformal coverage is 'substantially improves, reached nominal in our benchmarks, no finite-sample guarantee' instead of a promise. Tests: adversarial cases for the one-cell TVF canonicalization, unknown option keys asserting the exact PREDICT_ERR_OPTIONS code, missing receipt ids, and the load_extension replay attack. Eight pass. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: drop the vacuous or-clause from the replay-attack assertion Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: address round-two CodeRabbit findings predict_sha256 rejects INTEGER/REAL with PREDICT_ERR_SCHEMA instead of hashing SQLite's 15-digit number-to-text coercion, which could collapse distinct doubles into one hash; the functions reference documents the contract. The receipt CLI gains stable RECEIPT_ERR_* failure codes so agents branch on outcomes the way they do on PREDICT_ERR_* (the test asserts the code, not prose). The receipts skill now distinguishes result hashing from weight pinning explicitly, ships an executable no-registry recording variant, labels the pure-SQL verify as a spot check of one query's hash and routes full replay to the reference script, and tags the command block for MD040. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * chore: keep the scale-study harness out of the skills PR It lands on main with the scale results. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: verify re-derives the serving model; claims narrowed to behavior The skill claimed the reference script re-derives everything from stored fields; it only replayed input_sql and compared hashes, so a tampered model_id or options column still verified. Now verify() re-derives the serving model from the replayed result and fails the verification (exit 2, id_matches false) when it contradicts the recorded model_id; row-form receipts recorded via --model-id report id_matches null since no model is derivable. The options column is declared informational in both the report (options_verified: false) and the skill text: its authoritative copy is the options text inside input_sql, which the replay executes verbatim, and the receipt row itself is self-attested. Adversarial tests cover the detected forgery (model_id) and pin the documented limit (options). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: replay treats stored SQL as untrusted Locking extension loading closed one door and left the write door open: a tampered receipt's input_sql could DELETE, DROP, ATTACH, or flip pragmas during verification. verify() and list now open the database read-only at the file level and install an authorizer that permits only reads and function calls; every legitimate replay (aggregates, predict, backtest) is pure and passes, every write shape is denied at prepare time. Adversarial tests replay DELETE, DROP, ATTACH, and PRAGMA receipts, assert refusal, and assert the verifier's data is untouched. The skill states the replay posture explicitly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: complete the untrusted-replay audit in one pass Rather than waiting for the next review round to find the next adjacent gap, this closes the full remaining surface found by an adversarial self-audit of the receipt tool: a VM-step budget aborts runaway replays (recursive CTE bombs) instead of hanging the verifier; BLOB cells canonicalize deterministically as their sha256 instead of crashing json serialization; --options must parse as JSON before anything executes; database-open and SQL failures exit with stable RECEIPT_ERR_* codes instead of tracebacks. Every path has a test: the bomb aborts under a small step budget, blob receipts round-trip deterministically, malformed options and missing databases fail clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
1 parent a53939b commit 5ad201f

12 files changed

Lines changed: 1135 additions & 0 deletions

File tree

‎.gitignore‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -57,3 +57,6 @@ bindings/node/node_modules/
5757
bindings/rust/csrc/
5858
bindings/rust/target/
5959
bindings/rust/Cargo.lock
60+
61+
# autogluon working dirs (Mitra benchmark runs dump model copies here)
62+
benchmarks/AutogluonModels/

‎README.md‎

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -142,6 +142,17 @@ permissively licensed (MIT/Apache-2.0), so you ship it inside your product
142142
rather than rent it. And it composes with the rest of the in-database AI
143143
toolbox: sqlite-vec gave SQLite vector search; this gives it prediction.
144144

145+
### Agent skills
146+
147+
For agents, the SQL surface is the whole API, so what ships in
148+
[`skills/`](skills/) is judgment, not wrappers: four
149+
[agentskills.io](https://agentskills.io)-format skills covering core
150+
usage, provenance receipts (record what was asked, which model answered,
151+
and a hash that replays; serving determinism makes this verifiable),
152+
the distillation lifecycle (verify before serving, re-distill on drift,
153+
license discipline), and backtest interpretation. Install them into any
154+
skills-capable agent, or read them as the condensed operator's manual.
155+
145156
## Operations
146157

147158
| Function | Question | Returns |
Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,13 @@
1+
{"dataset": "Diabetes130US", "n_train": 1500, "n_test": 500, "d": 40, "xgboost": 0.906, "xgb_s": 0.1, "tabicl": 0.916, "tabicl_s": 2.2, "tabicl_device": "mps", "leak_gap": 0.0007, "ic_maxp": 0.9251, "collapsed": 1, "gbt<-tabicl soft": 0.916, "oof_s": 7.6, "oof_gap": 0.0007, "gbt<-tabicl soft-oof": 0.916}
2+
{"dataset": "Diabetes130US", "n_train": 9999, "n_test": 3334, "d": 40, "xgboost": 0.9070185962807439, "xgb_s": 0.1, "tabicl": 0.9130173965206959, "tabicl_s": 101.9, "tabicl_device": "mps", "leak_gap": -0.0, "ic_maxp": 0.922, "collapsed": 1, "gbt<-tabicl soft": 0.9130173965206959, "oof_s": 116.4, "oof_gap": -0.0, "gbt<-tabicl soft-oof": 0.9130173965206959}
3+
{"dataset": "airline_satisfaction", "n_train": 1500, "n_test": 500, "d": 21, "xgboost": 0.91, "xgb_s": 0.0, "tabicl": 0.928, "tabicl_s": 1.6, "tabicl_device": "mps", "leak_gap": 0.0693, "ic_maxp": 0.8787, "gbt<-tabicl": 0.908, "distill_s": 0.5, "blob_kb": 103.4, "gbt<-tabicl soft": 0.904, "oof_s": 5.5, "oof_gap": -0.0073, "gbt<-tabicl soft-oof": 0.906}
4+
{"dataset": "GiveMeSomeCredit", "n_train": 1500, "n_test": 500, "d": 10, "xgboost": 0.922, "xgb_s": 0.0, "tabicl": 0.936, "tabicl_s": 1.2, "tabicl_device": "mps", "leak_gap": 0.0067, "ic_maxp": 0.9407, "gbt<-tabicl": 0.932, "distill_s": 0.3, "blob_kb": 97.1, "gbt<-tabicl soft": 0.926, "oof_s": 4.1, "oof_gap": -0.0113, "gbt<-tabicl soft-oof": 0.93}
5+
{"dataset": "GiveMeSomeCredit", "n_train": 9999, "n_test": 3334, "d": 10, "xgboost": 0.9289142171565686, "xgb_s": 0.1, "tabicl": 0.9322135572885423, "tabicl_s": 8.4, "tabicl_device": "mps", "leak_gap": 0.021, "ic_maxp": 0.9452, "gbt<-tabicl": 0.9334133173365327, "distill_s": 2.6, "blob_kb": 104.3, "gbt<-tabicl soft": 0.9334133173365327, "oof_s": 29.5, "oof_gap": 0.0009, "gbt<-tabicl soft-oof": 0.9322135572885423}
6+
{"dataset": "APSFailure", "n_train": 1500, "n_test": 500, "d": 40, "xgboost": 0.994, "xgb_s": 0.1, "tabicl": 0.996, "tabicl_s": 2.6, "tabicl_device": "mps", "leak_gap": -0.004, "ic_maxp": 0.9902, "gbt<-tabicl": 0.99, "distill_s": 1.4, "blob_kb": 76.6, "gbt<-tabicl soft": 0.99, "oof_s": 8.2, "oof_gap": -0.0093, "gbt<-tabicl soft-oof": 0.996}
7+
{"dataset": "SDSS17", "n_train": 1500, "n_test": 500, "d": 11, "xgboost": 0.97, "xgb_s": 0.1, "tabicl": 0.972, "tabicl_s": 1.4, "tabicl_device": "mps", "leak_gap": 0.0167, "ic_maxp": 0.9892, "gbt<-tabicl": 0.968, "distill_s": 1.1, "blob_kb": 145.8, "gbt<-tabicl soft": 0.968, "oof_s": 4.3, "oof_gap": 0.0007, "gbt<-tabicl soft-oof": 0.964}
8+
{"dataset": "SDSS17", "n_train": 9999, "n_test": 3334, "d": 11, "xgboost": 0.9679064187162567, "xgb_s": 0.3, "tabicl": 0.9718056388722256, "tabicl_s": 8.7, "tabicl_device": "mps", "leak_gap": 0.0205, "ic_maxp": 0.9906, "gbt<-tabicl": 0.9658068386322736, "distill_s": 9.2, "blob_kb": 159.2, "gbt<-tabicl soft": 0.9661067786442712, "oof_s": 30.3, "oof_gap": 0.0052, "gbt<-tabicl soft-oof": 0.9676064787042592}
9+
{"dataset": "taiwanese_bankruptcy", "n_train": 1500, "n_test": 500, "d": 40, "xgboost": 0.966, "xgb_s": 0.1, "tabicl": 0.968, "tabicl_s": 2.4, "tabicl_device": "mps", "leak_gap": 0.004, "ic_maxp": 0.9721, "gbt<-tabicl": 0.968, "distill_s": 2.9, "blob_kb": 93.8, "gbt<-tabicl soft": 0.966, "oof_s": 8.2, "oof_gap": 0.0007, "gbt<-tabicl soft-oof": 0.968}
10+
{"dataset": "online_shoppers_intention", "n_train": 1500, "n_test": 500, "d": 17, "xgboost": 0.882, "xgb_s": 0.1, "tabicl": 0.89, "tabicl_s": 1.5, "tabicl_device": "mps", "leak_gap": 0.0507, "ic_maxp": 0.8863, "gbt<-tabicl": 0.89, "distill_s": 0.4, "blob_kb": 100.9, "gbt<-tabicl soft": 0.888, "oof_s": 4.9, "oof_gap": 0.0153, "gbt<-tabicl soft-oof": 0.886}
11+
{"dataset": "credit_card_default", "n_train": 1500, "n_test": 500, "d": 23, "xgboost": 0.81, "xgb_s": 0.1, "tabicl": 0.836, "tabicl_s": 1.7, "tabicl_device": "mps", "leak_gap": -0.0247, "ic_maxp": 0.8101, "gbt<-tabicl": 0.798, "distill_s": 0.9, "blob_kb": 93.5, "gbt<-tabicl soft": 0.81, "oof_s": 5.8, "oof_gap": -0.0153, "gbt<-tabicl soft-oof": 0.838}
12+
{"dataset": "airline_satisfaction", "n_train": 9999, "n_test": 3334, "d": 21, "xgboost": 0.9448110377924415, "xgb_s": 0.3, "tabicl": 0.9511097780443911, "tabicl_s": 139.6, "tabicl_device": "cpu", "leak_gap": 0.0465, "ic_maxp": 0.9842, "gbt<-tabicl": 0.9337132573485303, "distill_s": 3.5, "blob_kb": 105.7, "gbt<-tabicl soft": 0.9340131973605279, "oof_s": 486.7, "oof_gap": 0.0043, "gbt<-tabicl soft-oof": 0.9337132573485303}
13+
{"dataset": "credit_card_default", "n_train": 9999, "n_test": 3334, "d": 23, "xgboost": 0.8047390521895621, "xgb_s": 0.2, "tabicl": 0.8191361727654469, "tabicl_s": 167.5, "tabicl_device": "cpu", "leak_gap": 0.0205, "ic_maxp": 0.8387, "gbt<-tabicl": 0.8197360527894421, "distill_s": 7.4, "blob_kb": 96.8, "gbt<-tabicl soft": 0.8182363527294542, "oof_s": 571.6, "oof_gap": -0.001, "gbt<-tabicl soft-oof": 0.8194361127774445}

‎benchmarks/results/scale.md‎

Lines changed: 25 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,25 @@
1+
# Scale study: distillation beyond 1500 rows
2+
3+
TabICL v2 as the teacher on the naturally large TabArena
4+
datasets, at increasing training sizes. Questions: does student
5+
retention hold, where does tuned xgboost catch the zero-shot
6+
teacher, and does the in-context label leak grow with context
7+
(the gap column; see permissive-teachers.md).
8+
9+
Device: mps with cpu fallback (per-cell `dev`).
10+
11+
| dataset | n_train | xgboost | TabICL | gbt<-TabICL | soft | soft-oof | leak gap | oof gap | teacher s | distill s | blob KB |
12+
|---|---|---|---|---|---|---|---|---|---|---|---|
13+
| APSFailure | 1500 | 0.994 | 0.996 | 0.990 | 0.990 | 0.996 | -0.0040 | -0.0093 | 2.6 | 1.4 | 76.6 |
14+
| Diabetes130US | 1500 | 0.906 | 0.916 | - | 0.916 | 0.916 | 0.0007 | 0.0007 | 2.2 | - | - |
15+
| Diabetes130US | 9999 | 0.907 | 0.913 | - | 0.913 | 0.913 | -0.0000 | -0.0000 | 101.9 | - | - |
16+
| GiveMeSomeCredit | 1500 | 0.922 | 0.936 | 0.932 | 0.926 | 0.930 | 0.0067 | -0.0113 | 1.2 | 0.3 | 97.1 |
17+
| GiveMeSomeCredit | 9999 | 0.929 | 0.932 | 0.933 | 0.933 | 0.932 | 0.0210 | 0.0009 | 8.4 | 2.6 | 104.3 |
18+
| SDSS17 | 1500 | 0.970 | 0.972 | 0.968 | 0.968 | 0.964 | 0.0167 | 0.0007 | 1.4 | 1.1 | 145.8 |
19+
| SDSS17 | 9999 | 0.968 | 0.972 | 0.966 | 0.966 | 0.968 | 0.0205 | 0.0052 | 8.7 | 9.2 | 159.2 |
20+
| SDSS17 | 49999 | 0.975 | 0.610 | - | 0.610 | - | -0.0000 | - | 118.9 | - | - |
21+
| airline_satisfaction | 1500 | 0.910 | 0.928 | 0.908 | 0.904 | 0.906 | 0.0693 | -0.0073 | 1.6 | 0.5 | 103.4 |
22+
| credit_card_default | 1500 | 0.810 | 0.836 | 0.798 | 0.810 | 0.838 | -0.0247 | -0.0153 | 1.7 | 0.9 | 93.5 |
23+
| credit_card_default | 9999 | 0.805 | 0.771 | 0.776 | 0.776 | 0.776 | 0.0019 | 0.0054 | 11.7 | 7.4 | 96.2 |
24+
| online_shoppers_intention | 1500 | 0.882 | 0.890 | 0.890 | 0.888 | 0.886 | 0.0507 | 0.0153 | 1.5 | 0.4 | 100.9 |
25+
| taiwanese_bankruptcy | 1500 | 0.966 | 0.968 | 0.968 | 0.966 | 0.968 | 0.0040 | 0.0007 | 2.4 | 2.9 | 93.8 |

‎skills/distill-lifecycle/SKILL.md‎

Lines changed: 84 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,84 @@
1+
---
2+
name: distill-lifecycle
3+
description: >-
4+
Distill a teacher model into a native sqlite-predict student and manage
5+
it over time: verify quality before serving, re-distill on drift, and
6+
respect the teacher's license. Use when per-call serving needs to be
7+
instant and self-contained, when a prediction runs at volume, or when a
8+
model must travel inside the database file.
9+
license: MIT OR Apache-2.0
10+
metadata:
11+
version: "0.1.0"
12+
---
13+
14+
# The distillation lifecycle
15+
16+
Distillation compresses a teacher (your own labels, your existing model's
17+
predictions, or a licensed foundation model) into a native student the
18+
serving core executes with no runtime. Measured on the repo's benchmark
19+
suite: student blobs run tens to hundreds of kilobytes, single-row
20+
serving is sub-millisecond on CPU, and fits complete in seconds at
21+
benchmark scale (see `benchmarks/results/` in the sqlite-predict repo).
22+
The lifecycle judgment is yours.
23+
24+
## Distill
25+
26+
```sql
27+
-- tabular: target column holds labels or your model's predictions
28+
SELECT model_id, holdout_metric FROM distill_predict(
29+
'SELECT f1, f2, label FROM training',
30+
'{"target":"label","student_id":"churn-v1","student_kind":"gbt"}');
31+
32+
-- time series: from windows, or with a registered onnx teacher
33+
SELECT model_id FROM distill_forecast('SELECT series_key, value FROM obs',
34+
'{"context":96,"horizon":24,"student_id":"traffic-v1"}');
35+
```
36+
37+
Students register in `_predict_models` as content-hashed rows; they
38+
snapshot, fork, and sync with the database. Serve by name:
39+
`predict(NULL, apply_sql, '{"model":"churn-v1"}')` or
40+
`forecast(ts, value, 24, '{"model":"traffic-v1"}')`.
41+
42+
## Verify before serving
43+
44+
Never serve a student on faith:
45+
46+
1. Read `holdout_metric` from the distill call. It is measured on a
47+
holdout of your data. Compare it to a floor you trust (the majority
48+
class, last-value carry-forward, or your current model).
49+
2. For forecast students, run `backtest` on the same series and compare
50+
the student against `theta-classic` and the naive floor (see the
51+
interpret-backtest skill).
52+
3. Know the measured shape of distillation loss: on our benchmark suite
53+
classification students give up a median half point of accuracy
54+
against their teacher, but regression tails are worse. Check
55+
regression students more skeptically.
56+
4. Soft-label distillation (`proba`/`classes`) preserves the teacher's
57+
calibration when the teacher emits probabilities, and usually rescues
58+
imbalanced datasets where hard labels collapse to one class (the
59+
distiller refuses loudly on collapse rather than fitting a
60+
constant).
61+
62+
## Watch for drift, re-distill cheaply
63+
64+
A student is frozen; the world is not. When input distributions move,
65+
quality decays silently, so schedule verification rather than assuming:
66+
67+
- Re-run `backtest` (forecast) or score a fresh labeled sample (tabular)
68+
on a cadence proportional to how fast the data changes.
69+
- Re-distilling costs seconds. Register the new student under a
70+
versioned id (`churn-v2`), verify, then switch the serving call. Keep
71+
the old row until the new one is trusted; retire it after.
72+
73+
## Licensing is part of the lifecycle
74+
75+
A student derives from its teacher, and the teacher's license may
76+
impose obligations on what you distill and distribute; what applies
77+
depends on the exact weights, license version, and how the student is
78+
used. Record the teacher's license alongside each student you register.
79+
Your own labels and models carry no such limits. Restrictively licensed
80+
teachers require an explicit `accept_license` opt-in before they will
81+
run, but that gate is an acknowledgment, not compliance: read the
82+
license of the exact weights you use before distilling for anything
83+
beyond evaluation. The project documentation's license notes are
84+
orientation, not legal advice.

‎skills/interpret-backtest/SKILL.md‎

Lines changed: 79 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,79 @@
1+
---
2+
name: interpret-backtest
3+
description: >-
4+
Read sqlite-predict backtest() output to choose models, trust or
5+
distrust prediction intervals, and decide when auto-selection needs
6+
narrowing. Use before serving forecasts that feed decisions, when
7+
intervals matter, or when choosing between models on a specific series.
8+
license: MIT OR Apache-2.0
9+
metadata:
10+
version: "0.1.0"
11+
---
12+
13+
# Interpreting backtest()
14+
15+
`backtest` answers "how would this model have done on my series" with
16+
rolling-origin evaluation: it repeatedly hides the tail of the series,
17+
forecasts it, and scores against what actually happened.
18+
19+
```sql
20+
SELECT * FROM backtest('SELECT ts, value FROM readings', 24,
21+
'{"model":"theta-classic","folds":5}');
22+
```
23+
24+
## The metrics that matter
25+
26+
- **MASE** is the headline: error relative to the seasonal-naive floor.
27+
Below 1.0 beats naive; above 1.0 means the model is losing to
28+
last-season carry-forward and you should not serve it on this series.
29+
Compare models by MASE on the same series and folds.
30+
- **Coverage** is the fraction of actuals that landed inside the
31+
prediction interval. Compare it to the confidence level you asked for:
32+
a 0.90 band with 0.60 measured coverage is lying to you.
33+
- Per-fold rows expose stability. A model that wins on average but
34+
swings wildly across folds is riskier than a slightly worse, steady
35+
one.
36+
37+
## Interval judgment
38+
39+
The default Gaussian band can be overconfident on smooth series: the
40+
sqlite-predict benchmarks measured 0.57 coverage at a nominal 0.90 for
41+
theta-class models on smooth gluonts series (`benchmarks/results/` in
42+
the repo). When interval truth matters, request conformal intervals in
43+
the same call:
44+
45+
```sql
46+
SELECT * FROM backtest('SELECT ts, value FROM readings', 24,
47+
'{"model":"theta-classic",
48+
"interval_method":"conformal"}');
49+
```
50+
51+
Conformal intervals calibrate to measured out-of-sample residuals and
52+
substantially improve empirical coverage (they reached the nominal level
53+
in our benchmarks, but finite folds carry no guarantee). They need
54+
enough history to calibrate and apply to the statistical models only; a
55+
short series makes the call refuse rather than fabricate. Always verify
56+
with `backtest` coverage on your own series rather than trusting either
57+
method's reputation, ours included.
58+
59+
## Auto-selection judgment
60+
61+
Bare `forecast(ts, value, h)` lets `auto` pick per series by rolling-
62+
origin MASE. Use `backtest` when you want to see what auto sees:
63+
64+
- If one model wins consistently across your series, pin it
65+
(`'{"model":"theta-classic"}'`) and save the selection cost.
66+
- If the pool is polluted (a distilled student trained for a different
67+
regime keeps winning on stale patterns), narrow it:
68+
`'{"candidates":["theta-classic","tsb"]}'`.
69+
- Intermittent series (many zeros, sporadic demand) are `tsb` territory;
70+
if auto is not picking it, check whether the series reaches it and
71+
consider pinning.
72+
73+
## Reading degraded outcomes
74+
75+
`backtest` needs enough history for its folds: expect loud errors or
76+
degraded statuses on short series rather than fabricated confidence.
77+
That is the tool working. A series too short to backtest is a series too
78+
short to trust a model on; fall back to wider intervals and say so in
79+
whatever the forecast feeds.

0 commit comments

Comments
 (0)