Skip to content
View f0909172434's full-sized avatar

Highlights

  • Pro

Block or report f0909172434

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
f0909172434/README.md

Chih-Kai Wang — a terminal session that prints a short profile: Taipei; Python and TypeScript tools for inspectable AI and mathematical research; open to software and AI internships.

English · 繁體中文 · 简体中文  │  Portfolio · CV (PDF) · Email

Selected work  ls -l --pinned

HonestCI — Green CI should mean the tests you expected actually ran. RigorGraph — Ties each claim to the bytes of its evidence; one edited byte fails the audit. Finite Witness — Exhaustive counterexample search on small graphs, with certificates anyone can replay. SAIR Proof Press — Lean-checked proofs or finite countermodels; 1,669 of 1,669 released inputs accepted. 卜 ORACLE — A 4:30 short film rendered entirely from code: at 3 a.m., someone asks an AI a question. RuleDiff negative result — 0.99 on development, 0.67 held out. It didn't hold, so it was frozen and published.

Each terminal replays a real run against the commit it names; the output is copied, not written.

Films made as code  ckw play --loop *

Three works from October 2026 where the film, the music and the cut are all source code. Each repository documents its pipeline and its limits.

卜 ORACLE 病名為AI · The Disease Called AI world.execute(me); · Claude Code

  • 卜 ORACLE — A 4:30 short film: at 3 a.m. someone asks an AI "Will she get better?" Three.js rendered CPU-only in headless Chromium; score and sound design synthesised in Python. The README cites sources for the film's historical details and marks what was reconstructed. Three.js r169 · SwiftShader (CPU) · numpy/scipy score · ▶ Watch
  • 病名為AI · The Disease Called AI — An original song and hand-painted watercolour music video, 3:35, 67 shots. The score is a Python program; DiffSinger vocals render bit-exactly from fixed seeds; Whisper transcription is used as a diction check. p5.js + p5.brush · DiffSinger · Kokoro · Whisper QA · ▶ Watch
  • world.execute(me); · Claude Code — Mili's world.execute(me); staged as a Claude Code session and played live in the terminal. Pure Node, no dependencies; every frame is a function of song time. Unofficial fan work after MisakaZentai's DeepSeek Harness version; the song is not included. Node 20, zero dependencies · 24-bit ANSI · braille/sextant canvases · ▶ Watch

Negative results, kept  ckw verify --keep-negatives

Results that did not go the way I hoped stay public, with the same frozen artifacts as the ones that did.

  • ✗ RuleShift — Simple retrieval matched the more complex memory strategies; the complexity did not pay for itself.
  • ✗ RuleShift-Web — The no-LLM controller beat both models on the frozen held-out matrix.
  • ✗ RuleDiff negative result — 0.99 on development, 0.67 held out; the full-paper follow-up was stopped under its preregistered rule and frozen as this report.
  • ✗ Charlie Alpha 4B — No improvement on P-Bench or StatQA; only the simulator benchmark moved.

How I work  git log --graph

01 Frame. I write the question, the boundary, and what would count as done: which tests, which replay, which hash.

02 Execute. Claude Code and Codex (including Codex Cloud) write most of the code, tests and docs, in branches I review. Most lines in these repositories were typed by an agent; every claim is mine.

03 Decide. Evidence decides, not confidence: tests, independent replay checkers, content hashes. Negative results stay published with the same care as positive ones.

This site and the GitHub profile README are generated from one catalog file; the build fails if they drift.

Now · Oct 2026  ckw log --now

  • Shipped three works made entirely as code: ORACLE, The Disease Called AI, and world.execute(me).
  • Froze the RuleDiff negative result as a technical report; the RuleShift-Web manuscript is pre-submission.
  • Two upstream fixes merged: DeepSeek Harness Desktop and dsh-engram.
  • Learning Lean 4 / Mathlib through ProofWeave and SAIR.

Merged upstream  gh pr list --state merged

Everything else — 10 more records

Tools

  • Verified Search — DeepSeek Harness search plugin that retains citation excerpts and keeps evidence gaps visible. Only verified_search is stable; four extensions are experimental. 250 tests; no independent validation. v0.1.1 · stable search / experimental extensions
  • Second Agent Kit — Patches for DeepSeek Harness on macOS: Seatbelt confinement for shell processes, input-call limits, and per-project memory isolation. Gaps are documented; it is not a universal firewall. v0.1.4 · macOS only · experimental parts
  • DSH Architecture Lab — Development-preview toolkit for isolated memory/planning experiments on DeepSeek Harness (Lima VM or Seatbelt), with external judging and cost metering. Results so far come from one small repair task. 85 tests. v0.1.0-dev.9 · development preview

Research

  • ProofWeave Core — Checks author-written structured Markdown proofs against pinned Lean 4 / Mathlib, and reports certification separately from human-confirmed statement alignment. No natural-language translation. experimental · Core 2
  • RuleShift — Deterministic local testbed for how agent memory strategies cope with changing rules. In this pilot, simple retrieval matched more complex strategies (153 / 160) at lower cost. research pilot · 800 paired tasks
  • RuleShift-Web — Web-policy memory audit workbench: 3,200 model runs and 1,600 controls. The no-LLM controller was stronger than either model. Draft manuscript, not reviewed. pre-submission research snapshot
  • Charlie Alpha 4B — Experimental Qwen3.5-4B MLX fine-tune that picks statistical procedures locally. Improves on a simulator benchmark (DGP-Regret −34%) but not on P-Bench or StatQA. experimental v0.3.0 · mixed results

Learning

  • TokenScope — Bilingual in-browser lab: a hand-set one-head 5×5 attention toy, sampling controls (temperature, top-k, top-p), a step-by-step BPE merge demo, and exportable numbers. educational browser lab
  • MiniHarness — Python workshop where learners build a small agent harness in eight steps, offline with a scripted mock model, backed by a 38-module Traditional Chinese curriculum. Not production. 8-step workshop · zh-TW curriculum

Other

  • DeepSeek Girl — One animation atlas, two unofficial host packages: a 16-direction animated pet for Codex Desktop, and a DeepSeek Harness plugin that reacts to session state, offline. Codex v0.1.0 · Harness v0.2.0 · unofficial

exit

Generated from projects.json by scripts/render-profile.mjs; edits by hand are overwritten. Motion respects prefers-reduced-motion.

Pinned Loading

  1. honest-ci honest-ci Public

    Make green CI mean the tests you expected actually ran.

    TypeScript 3

  2. finite-witness-webmcp finite-witness-webmcp Public

    Finite-graph counterexample search with inspectable certificates, independent Python replay, and eight shared WebMCP tools.

    JavaScript

  3. rigorgraph rigorgraph Public

    Local-first claim-evidence graphs and deterministic audit reports for AI-assisted research.

    Python 1

  4. sair-stage2-proof-press sair-stage2-proof-press Public

    Public companion to Lean-checked equational implication solvers: frozen artifacts, released-input evaluations, and an English research paper.

    TypeScript 1

  5. ORACLE ORACLE Public

    卜 ORACLE — 凌晨三點,有人問 AI:「她會好起來嗎?」A short film where every frame and every note is generated by code.

    JavaScript

  6. rulediff-negative-result rulediff-negative-result Public

    Public technical report and reproducibility artifact for the RuleDiff confirmatory negative result.

    TeX