Repository navigation
Commit 5ad201f
feat: agent skills for usage, receipts, distillation, and backtesting (#11)
* feat: agent skills for usage, receipts, distillation, and backtesting
Four agentskills.io-format skills in skills/: sqlite-predict (core
usage: operation selection, the aggregate convention, statuses vs
errors, defaults), prediction-receipts (agent-owned provenance:
a _predict_receipts table convention, hashing the result document,
replay verification; possible because serving is deterministic and
models are content-hashed), distill-lifecycle (verify holdout before
serving, drift and re-distill cadence, license discipline), and
interpret-backtest (MASE against the naive floor, coverage honesty,
conformal judgment, when to narrow auto's pool). README gains an agent
skills pointer. Skills carry judgment, not wrappers: the SQL surface is
already the API.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: spec-polish the skill frontmatter
All four validate with the agentskills reference validator
(skills-ref). Add the optional license field and quote metadata values
per the spec's string-map recommendation.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: receipts reference script and CI conformance checks
skills/prediction-receipts/scripts/receipt.py is the interoperability
anchor: one stdlib-only implementation of the convention (record,
verify with exit-code semantics, list), with canonicalization defined
in one place (aggregate documents hashed verbatim; row results as
compact JSON with shortest round-trip floats). The skill now points at
it as the preferred path, keeping the hand-rolled steps as the spec.
tests/test_skills.py wires both into the normal pytest run: every
SKILL.md is held to the agentskills spec (name/dir match, name charset,
description bounds and when-to-use, line budget, house dash rule), and
the receipt script is proven end to end: record, replay-match,
re-record determinism, tamper detection via exit code 2, and registry
content_hash pinning for bundled models and distilled students.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* feat: predict_sha256 makes the receipts workflow pure SQL
Python was the receipt script's only dependency, and the environments
this extension targets (phone, embedded, wasm) may not have it. The
lightest dependency is the extension itself: expose the vendored
SHA-256 (the hash that already pins model weights) as a deterministic,
innocuous predict_sha256(x) scalar (TEXT/BLOB, NULL passthrough), and
rewrite the prediction-receipts skill so record and replay-verify run
in pure SQL. The Python script stays as optional convenience for
row-shaped results (canonical float serialization) and mismatch
diagnostics. CI: hash vectors against hashlib, plus a pure-SQL
record/verify/tamper round trip; functions reference documents the new
scalar. Also gitignores autogluon's local model dumps, which must never
be committed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: address all nine CodeRabbit findings on the skills PR
Security (critical): receipt.py locks extension loading immediately
after loading predict0, so a tampered receipt's replayed input_sql can
no longer call load_extension() on an arbitrary library; a regression
test replays exactly that attack and asserts sqlite refuses it.
Interoperability (major): canonicalization now follows the recorded
operation instead of the result shape, so a one-text-cell backtest or
predict result hashes as {"columns","rows"} rather than masquerading
as an aggregate document; verify() reads the stored operation.
Claims discipline (the reviewer enforcing our own path instructions):
size/latency claims in distill-lifecycle now cite the benchmarks and
state measured ranges; the soft-label rescue is scoped and qualified;
the licensing section says obligations 'may' apply, tells users to
record the exact license with each student, and states that
accept_license is an acknowledgment, not compliance; the 0.57 coverage
figure carries its model/dataset scope and source; the conformal
fragment became a complete backtest() call; and conformal coverage is
'substantially improves, reached nominal in our benchmarks, no
finite-sample guarantee' instead of a promise.
Tests: adversarial cases for the one-cell TVF canonicalization, unknown
option keys asserting the exact PREDICT_ERR_OPTIONS code, missing
receipt ids, and the load_extension replay attack. Eight pass.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* test: drop the vacuous or-clause from the replay-attack assertion
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: address round-two CodeRabbit findings
predict_sha256 rejects INTEGER/REAL with PREDICT_ERR_SCHEMA instead of
hashing SQLite's 15-digit number-to-text coercion, which could collapse
distinct doubles into one hash; the functions reference documents the
contract. The receipt CLI gains stable RECEIPT_ERR_* failure codes so
agents branch on outcomes the way they do on PREDICT_ERR_* (the test
asserts the code, not prose). The receipts skill now distinguishes
result hashing from weight pinning explicitly, ships an executable
no-registry recording variant, labels the pure-SQL verify as a spot
check of one query's hash and routes full replay to the reference
script, and tags the command block for MD040.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* chore: keep the scale-study harness out of the skills PR
It lands on main with the scale results.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: verify re-derives the serving model; claims narrowed to behavior
The skill claimed the reference script re-derives everything from
stored fields; it only replayed input_sql and compared hashes, so a
tampered model_id or options column still verified. Now verify()
re-derives the serving model from the replayed result and fails the
verification (exit 2, id_matches false) when it contradicts the
recorded model_id; row-form receipts recorded via --model-id report
id_matches null since no model is derivable. The options column is
declared informational in both the report (options_verified: false)
and the skill text: its authoritative copy is the options text inside
input_sql, which the replay executes verbatim, and the receipt row
itself is self-attested. Adversarial tests cover the detected forgery
(model_id) and pin the documented limit (options).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: replay treats stored SQL as untrusted
Locking extension loading closed one door and left the write door
open: a tampered receipt's input_sql could DELETE, DROP, ATTACH, or
flip pragmas during verification. verify() and list now open the
database read-only at the file level and install an authorizer that
permits only reads and function calls; every legitimate replay
(aggregates, predict, backtest) is pure and passes, every write shape
is denied at prepare time. Adversarial tests replay DELETE, DROP,
ATTACH, and PRAGMA receipts, assert refusal, and assert the verifier's
data is untouched. The skill states the replay posture explicitly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: complete the untrusted-replay audit in one pass
Rather than waiting for the next review round to find the next
adjacent gap, this closes the full remaining surface found by an
adversarial self-audit of the receipt tool: a VM-step budget aborts
runaway replays (recursive CTE bombs) instead of hanging the verifier;
BLOB cells canonicalize deterministically as their sha256 instead of
crashing json serialization; --options must parse as JSON before
anything executes; database-open and SQL failures exit with stable
RECEIPT_ERR_* codes instead of tracebacks. Every path has a test:
the bomb aborts under a small step budget, blob receipts round-trip
deterministically, malformed options and missing databases fail clean.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>1 parent a53939b commit 5ad201f
12 files changed
Lines changed: 1135 additions & 0 deletions
File tree
- benchmarks/results
- skills
- distill-lifecycle
- interpret-backtest
- prediction-receipts
- scripts
- sqlite-predict
- tests
- website/src/content/docs/reference
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
57 | 57 | | |
58 | 58 | | |
59 | 59 | | |
| 60 | + | |
| 61 | + | |
| 62 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
142 | 142 | | |
143 | 143 | | |
144 | 144 | | |
| 145 | + | |
| 146 | + | |
| 147 | + | |
| 148 | + | |
| 149 | + | |
| 150 | + | |
| 151 | + | |
| 152 | + | |
| 153 | + | |
| 154 | + | |
| 155 | + | |
145 | 156 | | |
146 | 157 | | |
147 | 158 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
| 1 | + | |
| 2 | + | |
| 3 | + | |
| 4 | + | |
| 5 | + | |
| 6 | + | |
| 7 | + | |
| 8 | + | |
| 9 | + | |
| 10 | + | |
| 11 | + | |
| 12 | + | |
| 13 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
| 1 | + | |
| 2 | + | |
| 3 | + | |
| 4 | + | |
| 5 | + | |
| 6 | + | |
| 7 | + | |
| 8 | + | |
| 9 | + | |
| 10 | + | |
| 11 | + | |
| 12 | + | |
| 13 | + | |
| 14 | + | |
| 15 | + | |
| 16 | + | |
| 17 | + | |
| 18 | + | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
| 1 | + | |
| 2 | + | |
| 3 | + | |
| 4 | + | |
| 5 | + | |
| 6 | + | |
| 7 | + | |
| 8 | + | |
| 9 | + | |
| 10 | + | |
| 11 | + | |
| 12 | + | |
| 13 | + | |
| 14 | + | |
| 15 | + | |
| 16 | + | |
| 17 | + | |
| 18 | + | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| 26 | + | |
| 27 | + | |
| 28 | + | |
| 29 | + | |
| 30 | + | |
| 31 | + | |
| 32 | + | |
| 33 | + | |
| 34 | + | |
| 35 | + | |
| 36 | + | |
| 37 | + | |
| 38 | + | |
| 39 | + | |
| 40 | + | |
| 41 | + | |
| 42 | + | |
| 43 | + | |
| 44 | + | |
| 45 | + | |
| 46 | + | |
| 47 | + | |
| 48 | + | |
| 49 | + | |
| 50 | + | |
| 51 | + | |
| 52 | + | |
| 53 | + | |
| 54 | + | |
| 55 | + | |
| 56 | + | |
| 57 | + | |
| 58 | + | |
| 59 | + | |
| 60 | + | |
| 61 | + | |
| 62 | + | |
| 63 | + | |
| 64 | + | |
| 65 | + | |
| 66 | + | |
| 67 | + | |
| 68 | + | |
| 69 | + | |
| 70 | + | |
| 71 | + | |
| 72 | + | |
| 73 | + | |
| 74 | + | |
| 75 | + | |
| 76 | + | |
| 77 | + | |
| 78 | + | |
| 79 | + | |
| 80 | + | |
| 81 | + | |
| 82 | + | |
| 83 | + | |
| 84 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
| 1 | + | |
| 2 | + | |
| 3 | + | |
| 4 | + | |
| 5 | + | |
| 6 | + | |
| 7 | + | |
| 8 | + | |
| 9 | + | |
| 10 | + | |
| 11 | + | |
| 12 | + | |
| 13 | + | |
| 14 | + | |
| 15 | + | |
| 16 | + | |
| 17 | + | |
| 18 | + | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| 26 | + | |
| 27 | + | |
| 28 | + | |
| 29 | + | |
| 30 | + | |
| 31 | + | |
| 32 | + | |
| 33 | + | |
| 34 | + | |
| 35 | + | |
| 36 | + | |
| 37 | + | |
| 38 | + | |
| 39 | + | |
| 40 | + | |
| 41 | + | |
| 42 | + | |
| 43 | + | |
| 44 | + | |
| 45 | + | |
| 46 | + | |
| 47 | + | |
| 48 | + | |
| 49 | + | |
| 50 | + | |
| 51 | + | |
| 52 | + | |
| 53 | + | |
| 54 | + | |
| 55 | + | |
| 56 | + | |
| 57 | + | |
| 58 | + | |
| 59 | + | |
| 60 | + | |
| 61 | + | |
| 62 | + | |
| 63 | + | |
| 64 | + | |
| 65 | + | |
| 66 | + | |
| 67 | + | |
| 68 | + | |
| 69 | + | |
| 70 | + | |
| 71 | + | |
| 72 | + | |
| 73 | + | |
| 74 | + | |
| 75 | + | |
| 76 | + | |
| 77 | + | |
| 78 | + | |
| 79 | + | |
0 commit comments