diff --git a/TODO.md b/TODO.md index 952d0c3..6bb5d2b 100644 --- a/TODO.md +++ b/TODO.md @@ -274,6 +274,25 @@ root `AGENTS.md` mock-policy reconciled; `docs/REFACTOR_READINESS.md` marked a h --- +## Round-2 agent-ergonomics pass — 2026-08-31 (lane E) + +- [x] **scripts/README.md covered 6 of 24 scripts by name; 4 on-disk scripts + entirely undocumented** (`build_pages_site.py`, `derive_captions_from_json.py`, + `sync_youtube_metadata.py`, `txt_to_segments.py`). Added a pipeline map + (ingest → transcribe → captions → speakers → journal v2 maintenance → + pages/translation surfaces) at the top plus per-script sections for the four + missing ones. Acceptance: a mechanical name-scan of `scripts/*.py` vs + `scripts/README.md` reports zero undocumented scripts; all referenced paths + exist on disk. ✓ this pass +- [x] **Verified Makefile target references and docs script references resolve** + (`-m` module paths and `scripts/*.py` args in `Makefile`, plus every + `scripts/…py` mention in `docs/*.md`). All resolve; no stale commands found + this pass. Acceptance: zero missing refs in the same mechanical scan. ✓ +- [ ] **Documentation test-count snapshot drift** — `tests/README.md` / + `tests/AGENTS.md` counts are dated snapshots by policy; refresh from a live + `uv run pytest tests/ -q` whenever a pass touches tests. Deferred: this + round changed docs only; no test surface moved. + ## Notes for the next reviewer - Test/corpus counts are live snapshots — run `uv run pytest tests/ -q`; do not hardcode totals diff --git a/scripts/README.md b/scripts/README.md index bdc0b60..865eea8 100644 --- a/scripts/README.md +++ b/scripts/README.md @@ -2,6 +2,32 @@ CLI utilities for data processing, transcription, and course scaffolding. +Every script here is a thin CLI orchestrator over `src/journal_utilities/` or the +sibling `ActiveInferenceJournal` checkout. Run any of them as +`uv run python scripts/.py --help` for exact flags. + +## Pipeline map (what exists and how it chains) + +```text +YouTube ingest download_channel.py --> data/output/channel_videos.json + sync_youtube_metadata.py (enrich titles/tags/chapters from local metadata) +Transcription transcribe_missing.py / transcribe_worklist.py (mlx-whisper, Mac) + transcription_status.py (read-only coverage report) +Caption/subtitle layer derive_captions_from_json.py (transcript.json -> captions/*.srt) + txt_to_segments.py (plain .txt -> pseudo-timed segments) +Speaker attribution speaker_cues.py -> apply_speaker_names.py (repair + name mapping) +Recovery/repair repair_split_transcripts.py, recover_whisperx.py, + patch_whisperx.py, fix_scheduled_dates.py, fetch_chapters.py +Journal v2 maintenance enrich_metadata.py -> repair_split_transcripts.py + -> generate_journal_indexes.py -> validate_journal.py (read-only gate) +Publication surfaces build_pages_site.py (journal -> static GitHub Pages bundle) +Translation translate_subtitles_openrouter.py (docs/translation.md) +Curriculum scaffold_youtube_courses.py +``` + +The canonical maintenance chain is documented in `docs/JOURNAL_SCHEMA.md`; the +read-only `validate_journal.py` is the acceptance gate for every step above. + ## `download_channel.py` Primary CLI for enumerate-and-download of YouTube channel content. @@ -255,3 +281,28 @@ Uses OpenRouter API (requires `OPENROUTER_API_KEY` in `.env` or environment): python scripts/translate_subtitles_openrouter.py --journal ../ActiveInferenceJournal --series "Livestream" ``` + +## `build_pages_site.py` + +Compile the sibling `ActiveInferenceJournal` checkout into a static GitHub Pages +bundle (HTML pages + index) under the journal's own output directory. Pass +`--help` for the journal-path and output flags; read-only with respect to this repo. + +## `derive_captions_from_json.py` + +Derive base English `captions/*.srt` files from each item's diarized +`transcript.json` for journal items that currently lack caption SRTs. Dry-run by +default; add `--apply` to write the `.srt` files to disk. + +## `sync_youtube_metadata.py` + +Closed-loop metadata synchronizer: merges locally cached channel data +(`data/output/channel_videos.json`, `data/input/video_chapters.json`) and the +InstituteOS source tree to derive video tags and refresh transcript-adjacent +metadata. Uses repo-relative paths automatically. + +## `txt_to_segments.py` + +Convert a timestamp-free plain-text transcript +(`data/output/transcripts/.txt`) into pseudo-timed ~30-second segments +(~80 words each) so downstream diarization/segment tooling has a uniform input.