English · 繁體中文 · 简体中文 │ Portfolio · CV (PDF) · Email
Each terminal replays a real run against the commit it names; the output is copied, not written.
Three works from October 2026 where the film, the music and the cut are all source code. Each repository documents its pipeline and its limits.
- 卜 ORACLE — A 4:30 short film: at 3 a.m. someone asks an AI "Will she get better?" Three.js rendered CPU-only in headless Chromium; score and sound design synthesised in Python. The README cites sources for the film's historical details and marks what was reconstructed.
Three.js r169 · SwiftShader (CPU) · numpy/scipy score· ▶ Watch - 病名為AI · The Disease Called AI — An original song and hand-painted watercolour music video, 3:35, 67 shots. The score is a Python program; DiffSinger vocals render bit-exactly from fixed seeds; Whisper transcription is used as a diction check.
p5.js + p5.brush · DiffSinger · Kokoro · Whisper QA· ▶ Watch - world.execute(me); · Claude Code — Mili's world.execute(me); staged as a Claude Code session and played live in the terminal. Pure Node, no dependencies; every frame is a function of song time. Unofficial fan work after MisakaZentai's DeepSeek Harness version; the song is not included.
Node 20, zero dependencies · 24-bit ANSI · braille/sextant canvases· ▶ Watch
Results that did not go the way I hoped stay public, with the same frozen artifacts as the ones that did.
- ✗ RuleShift — Simple retrieval matched the more complex memory strategies; the complexity did not pay for itself.
- ✗ RuleShift-Web — The no-LLM controller beat both models on the frozen held-out matrix.
- ✗ RuleDiff negative result — 0.99 on development, 0.67 held out; the full-paper follow-up was stopped under its preregistered rule and frozen as this report.
- ✗ Charlie Alpha 4B — No improvement on P-Bench or StatQA; only the simulator benchmark moved.
01 Frame. I write the question, the boundary, and what would count as done: which tests, which replay, which hash.
02 Execute. Claude Code and Codex (including Codex Cloud) write most of the code, tests and docs, in branches I review. Most lines in these repositories were typed by an agent; every claim is mine.
03 Decide. Evidence decides, not confidence: tests, independent replay checkers, content hashes. Negative results stay published with the same care as positive ones.
This site and the GitHub profile README are generated from one catalog file; the build fails if they drift.
- Shipped three works made entirely as code: ORACLE, The Disease Called AI, and world.execute(me).
- Froze the RuleDiff negative result as a technical report; the RuleShift-Web manuscript is pre-submission.
- Two upstream fixes merged: DeepSeek Harness Desktop and dsh-engram.
- Learning Lean 4 / Mathlib through ProofWeave and SAIR.
- dsh-tauri/deepseek-harness-desktop#740 — Normalise symlinked worktree paths so re-creating a worktree cannot misjudge and delete uncommitted changes. 2026-09-26
- kenz1117/dsh-engram#4 — Trace legacy database claims and add a conservative migration. 2026-09-21
- EmiyaKatuz/Codex-Dream-Skin-Needy-Girl-Overdose#10 — Keep Windows verification failures distinguishable: narrower native-window fallback, standalone helper loading. 2026-07-28 · case study
Everything else — 10 more records
Tools
- Verified Search — DeepSeek Harness search plugin that retains citation excerpts and keeps evidence gaps visible. Only verified_search is stable; four extensions are experimental. 250 tests; no independent validation. v0.1.1 · stable search / experimental extensions
- Second Agent Kit — Patches for DeepSeek Harness on macOS: Seatbelt confinement for shell processes, input-call limits, and per-project memory isolation. Gaps are documented; it is not a universal firewall. v0.1.4 · macOS only · experimental parts
- DSH Architecture Lab — Development-preview toolkit for isolated memory/planning experiments on DeepSeek Harness (Lima VM or Seatbelt), with external judging and cost metering. Results so far come from one small repair task. 85 tests. v0.1.0-dev.9 · development preview
Research
- ProofWeave Core — Checks author-written structured Markdown proofs against pinned Lean 4 / Mathlib, and reports certification separately from human-confirmed statement alignment. No natural-language translation. experimental · Core 2
- RuleShift — Deterministic local testbed for how agent memory strategies cope with changing rules. In this pilot, simple retrieval matched more complex strategies (153 / 160) at lower cost. research pilot · 800 paired tasks
- RuleShift-Web — Web-policy memory audit workbench: 3,200 model runs and 1,600 controls. The no-LLM controller was stronger than either model. Draft manuscript, not reviewed. pre-submission research snapshot
- Charlie Alpha 4B — Experimental Qwen3.5-4B MLX fine-tune that picks statistical procedures locally. Improves on a simulator benchmark (DGP-Regret −34%) but not on P-Bench or StatQA. experimental v0.3.0 · mixed results
Learning
- TokenScope — Bilingual in-browser lab: a hand-set one-head 5×5 attention toy, sampling controls (temperature, top-k, top-p), a step-by-step BPE merge demo, and exportable numbers. educational browser lab
- MiniHarness — Python workshop where learners build a small agent harness in eight steps, offline with a scripted mock model, backed by a 38-module Traditional Chinese curriculum. Not production. 8-step workshop · zh-TW curriculum
Other
- DeepSeek Girl — One animation atlas, two unofficial host packages: a 16-direction animated pet for Codex Desktop, and a DeepSeek Harness plugin that reacts to session state, offline. Codex v0.1.0 · Harness v0.2.0 · unofficial
Generated from projects.json by scripts/render-profile.mjs; edits by hand are overwritten. Motion respects prefers-reduced-motion.


