All-in-one subtitle generation tool integrating English transcription + subtitle segmentation + translation. All processing supports fully offline operation.
| Feature | Approach | Used in |
|---|---|---|
| ASR | WhisperX (faster-whisper-large-v3) + Silero VAD | Transcription (scripts/transcribe.py) |
| Alignment | Wav2Vec2 phoneme-level forced alignment | Transcription (scripts/transcribe.py) |
| Segment | Local LLM via llama-server or online API | Transcription (scripts/transcribe.py) |
| Translate | Local LLM via llama-server or online API | Translation (scripts/translate.py) |
| File | Purpose |
|---|---|
| main.py | Orchestrator entry point. Calls scripts/transcribe.py for transcription and/or scripts/translate.py for translation in a single run. |
| batch.py | Batch processing. Iterates all files in an input folder, spawning one main.py subprocess per file (each gets a clean CUDA context). |
| translate_config.json | Translation config. Language, punctuation, glossary, merge limits & split tightness (merge_fix_preset) — edited by the user, read by scripts/translate.py at startup. |
| llm_config.json | LLM backend & model config. API keys, base URLs, default models, local model names, layer counts, chat templates, reasoning flags. Loaded by tools/llm_call.py at startup — edit to add/change backends or models. |
| README_CN.md | Chinese (中文) documentation — setup guide, usage, and file reference. |
| README.md | This file — English documentation. |
| Script | Purpose |
|---|---|
| scripts/transcribe.py | Transcription pipeline. WhisperX ASR → Wav2Vec2 alignment → LLM segmentation → SRT + TXT export. Can run standalone. |
| scripts/translate.py | Translation. Translates existing SRT (preserving timecodes) or TXT files using a local or API-based LLM. Can run standalone. |
| Module | Purpose |
|---|---|
| tools/__init__.py | Package marker — module docstring only. |
| tools/config.py | System configuration. Sets HTTP_PROXY / NO_PROXY (bypasses proxy for PyTorch/HuggingFace CDNs) and CUDA_VISIBLE_DEVICES. Runs at import time, before any model download. |
| tools/env_check.py | Environment verification. check_ffmpeg() — verifies FFmpeg is on PATH. check_cuda() — GPU detection + real kernel-launch test, returns True/False. |
| tools/llm_call.py | Unified LLM call layer. create_llm_call() factory with three backend classes (LLMCallLocal, LLMCallOpenAI, LLMCallAnthropic). Shared by both segmentation and translation pipelines. Loads llm_config.json at startup. |
| tools/music_detect.py | YAMNet music/speech analysis. YamNetScan — one YAMNet pass over the waveform exposing high-confidence music regions (chunk boundaries) and per-interval speech peaks. filter_hallucinations() — Layer 2 confidence filter dropping low-confidence segments only when no concurrent speech is detected. |
| tools/yamnet_frontend.py | YAMNet numpy log-mel frontend. log_mel_spectrogram(), to_patches(), classify() — compute log-mel features and classify patches for the split yamnet_clf.onnx classifier (used by tools/music_detect.py). |
| tools/extract.py | Word extraction, dedup & retry resolution. extract_words_from_result() — pulls (start, end, text) word tuples from WhisperX result dicts. dedup_overlap_words() / dedup_seam_words() — drop words decoded twice across overlapping / seamed decode windows. resolve_retry_segments() — island lead-in recovery. fix_word_timestamps() — post-processing to correct inflated timestamps. |
| tools/cache.py | Word-level caching. save_words_cache() / load_words_cache() — persist/restore extracted word data as JSON for re-split debugging. |
| tools/format.py | SRT time formatting. format_srt_time() — converts float seconds to HH:MM:SS,mmm SRT format. |
| tools/export.py | Subtitle export. export_srt() — writes SRT with index + timestamps. export_txt() — writes plain text or SRT-format text. export_word_level() — debug-only word-level SRT/TXT (currently commented out). |
| tools/download_models.py | Model pre-downloader. One-shot script to download Silero VAD, Wav2Vec2, NLTK, and YAMNet data to local cache so the pipeline runs fully offline. |
| tools/segment.py | LLM segmentation engine. segment_words() function — runs the full 10-phase segmentation algorithm (see docs/segmentation_pipeline.md for full algorithm chart). |
| tools/seg_rules.py | Rule-based segmentation helpers. _comma_split(), _conjunction_split(), and _find_ambiguous_conjunctions() — pure rule logic extracted from tools/segment.py (no LLM dependency). |
| tools/seg_diff.py | Character-diff for LLM punctuation. _build_char_to_word(), _find_new_breaks(), _find_new_commas() — maps LLM-added punctuation back to input word indices using difflib.SequenceMatcher. |
| tools/non_split_bigrams.py | Phrase protection. NON_SPLIT_BIGRAMS set (~5700 entries) of fixed expressions, phrasal verbs, and collocations that must not be split across subtitle segments. would_break_phrase() look-up used by all segmentation phases. |
| tools/vad_patch.py | Gap- & music-aware VAD merge patch. Replaces whisperx Vad.merge_chunks with a data-driven cutter (adaptive gap, music boundary, music-interior, cap-with-overlap). Emits overlap regions and retry twin windows for music-masked islands — consumed by tools/extract.py dedup. |
| tools/llama/ | llama.cpp binaries. Pre-compiled llama-server.exe + CUDA 12 DLLs for local LLM inference. Excluded from git — see Setup → llama.cpp Binary. |
| File | Purpose |
|---|---|
| docs/segmentation_pipeline.md | 10-phase LLM segmentation algorithm doc (English). Full algorithm flow chart and detailed description of all 10 phases, guard mechanisms, and LLM call summary. |
| docs/segmentation_pipeline_CN.md | 10-phase LLM segmentation algorithm doc (Chinese). Same content as above in Chinese. |
| docs/transcription_n_translation.md | Transcription & translation technical implementation (English). Design rationale, parameter choices, and code walkthrough for the ASR-to-alignment-to-segmentation pipeline and the translation factory backend. |
| docs/transcription_n_translation_CN.md | Transcription & translation technical implementation (Chinese). Same content as above in Chinese. |
| docs/models_n_performance.md | Models & performance (English). Hardware requirements, LLM selection guide, and inference benchmarks. |
| docs/models_n_performance_CN.md | Models & performance (Chinese). Same content as above in Chinese. |
| Category | Examples | Install method |
|---|---|---|
| System tools | NVIDIA driver (≥ CUDA 12.8), FFmpeg, Python venv | Manual (one-time) |
| llama.cpp binary | llama-server.exe |
Download zip → extract to tools/llama/ (see variant notes below) |
| Pre-downloaded models | faster-whisper-large-v3, Phi-4 GGUF | Manual download to models/ |
| Auto-downloaded models | Silero VAD, Wav2Vec2, NLTK, YAMNet | python tools/download_models.py |
GPU vs CPU: The entire pipeline (ASR + alignment + LLM) auto-detects CUDA and falls back to CPU if no GPU is available. No special flags needed — just choose the right llama.cpp variant below.
This project depends on PyTorch CUDA 12.8 (see torch==2.8.0+cu128 in requirements.txt), which requires a compatible NVIDIA driver for GPU acceleration.
| Requirement | Notes |
|---|---|
| NVIDIA GPU | Architecture ≥ Maxwell (GTX 900 series or newer), minimum 4 GB VRAM (8 GB+ recommended) |
| NVIDIA Driver | Version ≥ R570 (Final driver support for CUDA 12.8 was released July 2026; update to the latest Game Ready / Studio Driver) |
| CUDA Toolkit | Not required. PyTorch ships its own CUDA runtime — torch==2.8.0+cu128 includes the necessary cudart and cublas libraries. A sufficiently recent NVIDIA driver is all you need. |
| No GPU / CPU mode | The pipeline auto-detects CUDA and falls back to CPU when unavailable (slower). No additional configuration needed. |
To verify, run nvidia-smi and check the CUDA Version line at the top. If it's below 12.8 or nvidia-smi produces no output, update your driver from NVIDIA Driver Downloads.
Required by Whisper for audio extraction.
winget install FFmpeg
# or: choco install ffmpeg
# or download from https://ffmpeg.org/download.html and add to PATH
# CN mirror: https://mirrors.tuna.tsinghua.edu.cn/ffmpeg/Verify: ffmpeg -version
# Create .venv (Python 3.10)
python -m venv .venv
# Activate
.venv\Scripts\activate
# Install dependencies
pip install -r requirements.txtllama.cpp provides the local LLM inference engine. Choose the variant matching your hardware:
| Your setup | Download zip |
|---|---|
| NVIDIA GPU (CUDA 12.x) | llama-bNNNN-bin-win-cuda12.4-x64.zip (e.g. b9888) |
| CPU only / no GPU | llama-bNNNN-bin-win-cpu-x64.zip (same release page) |
- Go to the llama.cpp releases page
- CN accelerator proxy:
https://ghproxy.com/https://github.com/ggml-org/llama.cpp/releases
- CN accelerator proxy:
- Download the zip for your setup
- Extract all files into
tools/llama/
Note: The release page also lists a "CUDA 12.4 DLLs" package (
cudart-llama-bin-win-cuda-12.4-x64.zip, ~750 MB). This is an optional supplement — it provides NVIDIA's official cuBLAS libraries which can give slightly better GPU performance thanggml-cuda.dll's built-in implementation. You don't need it; the standard CUDA build works fine on its own. If you do download it, extract the 3 DLLs into the sametools/llama/folder.
Place these in the models/ directory:
models/
├── faster-whisper-large-v3/ # ~3 GB, from HuggingFace
└── *.gguf # LLM models (Phi-4, etc.)
Download links for the LLM GGUF models: docs/models_n_performance.md.
faster-whisper-large-v3:
pip install huggingface-hub
# Method A — default download
hf download Systran/faster-whisper-large-v3 --local-dir models/faster-whisper-large-v3
# Method B — CN mirror (hf-mirror.com), faster in China
# Windows PowerShell:
$env:HF_ENDPOINT = "https://hf-mirror.com"
hf download Systran/faster-whisper-large-v3 --local-dir models/faster-whisper-large-v3
# Or manually download via mirror direct link:
# https://hf-mirror.com/Systran/faster-whisper-large-v3/tree/mainYAMNet (music detection): auto-downloaded by download_models.py → models/yamnet/ — see Auto-download.
python tools/download_models.pyIf download is slow, set the mirror before running:
$env:HF_ENDPOINT = "https://hf-mirror.com"
Downloads to models/hub/ (Silero VAD, Wav2Vec2, NLTK) and models/yamnet/ (YAMNet):
| Model | Size | Used by | Description |
|---|---|---|---|
| Silero VAD | ~35 MB | whisperx (VAD) | Voice Activity Detection — finds speech segments in audio |
| Wav2Vec2 (base 960h) | ~360 MB | WhisperX align | Phoneme-level word timestamp alignment |
| NLTK punkt_tab | ~64 MB | WhisperX alignment (internal) | Sentence boundary detection for phoneme-level forced alignment |
| YAMNet | ~16 MB | music_detect (YAMNet scan) | Music-region detection + concurrent-speech evidence for the hallucination filter |
YAMNet (music detection):
The pipeline runs one YAMNet pass over the audio to locate music (used to cut chunks at real content boundaries) and concurrent-speech evidence (used to keep real speech spoken over music from the hallucination filter). download_models.py installs it into models/yamnet/ as a split model: the Qualcomm yamnet_clf.onnx classifier (log-mel input) plus the numpy log-mel frontend params in yamnet_frontend.npz (window + mel, extracted once from the original monolithic model). If the split assets are absent (e.g. the classifier download failed), the pipeline falls back to the monolithic yamnet.onnx.
The packaged exe already bundles YAMNet (under models/yamnet/), so this step is only needed for source installs.
# 1. Transcribe a single file (accepts mp4/mkv/avi/mov/wav/mp3/m4a/flac)
python scripts/transcribe.py -i input/input.mp4
# 2. Translate existing subtitles (auto-detects SRT or TXT)
python scripts/translate.py -i output/input.srt
# 3. Transcribe + translate in one command
python main.py -i input/input.mp4 -translate true
# 4. Transcribe + translate with different backends
python main.py -i input/input.mp4 -translate true -seg_backend local -transl_backend deepseek
# 5. Process all videos in a folder
python batch.py
# 6. Batch with translation
python batch.py -translate trueBoth the segmentation pipeline (scripts/transcribe.py) and the translation engine (scripts/translate.py) share a unified LLM call layer. Available backends:
| Backend | API format | Default model | Auth |
|---|---|---|---|
local |
llama-server subprocess | phi-4 |
— |
deepseek |
Auto-detect: OpenAI /chat/completions or Anthropic Messages |
deepseek-v4-flash |
openai_api_key |
openai |
OpenAI-compatible | gpt-5.6-terra |
openai_api_key |
qwen |
OpenAI-compatible | qwen3.7-max |
openai_api_key |
gemini |
OpenAI-compatible | gemini-3.5-flash |
openai_api_key |
anthropic |
Anthropic Messages | claude-opus-4-8 |
anthropic_api_key |
Backend URLs, default models, and local model definitions (GGUF filename, layer count, chat template, reasoning flag) are configured in llm_config.json — see transcription_n_translation.md - "1.6 llm_config.json Configuration Reference" for the full field reference.
API keys and custom base URL sit at the top of llm_config.json (structure sketch, // comments not present in the real file):
The API key fields (openai_api_key, anthropic_api_key, api_base_url) are shared — they apply to both segmentation backend and translation backend. Translation-specific fields (target_lang, target_lang_code, etc.) are configured in translate_config.json and detailed in transcription_n_translation.md - "3.1 Translation Configuration".
Or use the .env file for API keys (place at project root as .env): only takes effect when the corresponding key fields in llm_config.json are left empty
# ── LLM backend API keys (shared by segmentation & translation) ──
OPENAI_API_KEY=sk-your-key-here
# Anthropic-specific (only needed when backend is anthropic)
ANTHROPIC_API_KEY=sk-ant-your-key-here
python scripts/transcribe.py -i <input> [options]Transcribes audio/video to aligned word-level timestamps via WhisperX + Wav2Vec2, then segments the words into subtitle blocks via LLM. Produces SRT + TXT.
On re-run loads a word-level JSON cache (cache/<stem>_words.json) to skip the transcription + alignment step. For more details see transcription_n_translation.md - "2. Transcription Pipeline".
| Argument | Default | Description |
|---|---|---|
-i, --input <path> |
input/input.mp4 |
Input video or audio file (.mp4/.mkv/.avi/.mov/.wav/.mp3/.m4a/.flac) |
-o, --output <path> |
output/<stem>.srt |
Output SRT file path |
-seg_backend <name> |
local |
Segmentation backend. See 1. LLM Backend & Shared Configuration for available options. |
-seg_model <name> |
per-backend | Model for segmentation (e.g. phi-4, gpt-5.6-terra; default per backend) |
-gpu_layers <N> |
auto-detect | Number of local LLM layers to offload to GPU (0 = CPU only) |
-no_cache |
— | Disable word-level .json cache (caching is enabled by default) |
Input video
│
├─ 1. Audio decoding + YAMNet music scan
├─ 2. WhisperX ASR (faster-whisper-large-v3 + Silero VAD, gap- & music-aware chunk merge)
├─ 3. Island lead-in recovery + speech-evidence hallucination filter
├─ 4. Wav2Vec2 phoneme-level forced alignment
├─ 5. Word dedup (overlap-window + window-seam) & timestamp fixing
├─ 6. LLM punctuation & segmentation (configured via -seg_backend / -seg_model)
│ Phase 1-4: LLM fill missing punctuation (chunked + context)
│ Phase 5: Build segments from LLM-inferred sentence boundaries
│ Phase 6: Comma-based force-split of overlength segments
│ Phase 7: Conjunction split (LLM classifier for ambiguous "and")
│ Phase 8: Run-on re-punctuation (recursive LLM split)
│ Phase 9: Conjunction fragment merge (LLM classifier)
│ Phase 10: Emergency character-count split
└─ 7. Export SRT + TXT
| File | Description |
|---|---|
output/<stem>.srt |
Subtitle file with timestamps |
output/<stem>.txt |
Plain text (with timestamp markers) |
cache/<stem>_words.json |
Cached ASR word data (reused on re-run) |
# Default
python scripts/transcribe.py -i input/input.mp4
# Custom output path
python scripts/transcribe.py -i input/input.mp4 -o D:/output/lecture.srt
# GPU-free LLM segmentation (CPU only)
python scripts/transcribe.py -i input/input.mp4 -gpu_layers 0
# Use an online LLM for segmentation
python scripts/transcribe.py -i input/input.mp4 -seg_backend openai -seg_model gpt-5.6-terra
# No caching
python scripts/transcribe.py -i input/input.mp4 -no_cacheTranslates SRT (preserving timecodes) or plain TXT files using a local LLM (via llama-server) or any of 5 online API backends. Auto-detects input format by content — SRT files keep their timecodes, TXT files are treated as plain text lines.
Two translation modes are available:
accurate(default): Recommended for local models.flexible: Recommended for online API backends.allow_flexible_word_orderandallow_simplify_wordingare only effective in this mode.
For detailed mechanics and window parameters see transcription_n_translation.md - "3.3.1 Translation Mode".
python scripts/translate.py -i <input> [options]| Argument | Default | Description |
|---|---|---|
-i, --input <path> |
(required) | Input .srt or .txt file (auto-detects format by content) |
-o, --output <path> |
auto | Output file path |
-transl_backend <name> |
local |
Translation backend. Same options as scripts/transcribe.py. |
-transl_model <name> |
per-backend | Model for the selected backend. Same options as scripts/transcribe.py. |
-src_lang <name> |
English |
Source language name (only English supported for now). Overrides source_lang in translate_config.json. |
-tgt_lang <name> |
Chinese |
Target language name. Overrides target_lang in translate_config.json. Auto-completes -tgt_lang_code if omitted. |
-tgt_lang_code <code> |
CN |
Language code for filename suffix. Overrides target_lang_code in translate_config.json. Auto-completes -tgt_lang if omitted. |
-mode <name> |
accurate |
Translation mode: accurate (2 lines/batch) or flexible (4 lines/batch, with timecodes) |
-max_merge_chars <N> |
no limit | Maximum characters in a merged translation output (CJK characters count as 1 each). Overly long merges are rejected and lines re-translated individually. |
-max_merge_gap <N> |
no limit | Maximum time gap (seconds) across merged subtitles. Merges spanning a wider gap are rejected. Timecoded input only. |
-gpu_layers <N> |
auto-detect | Number of local LLM layers to offload to GPU (0 = CPU only). Same as scripts/transcribe.py |
# Local Phi-4 — Translate TXT to Chinese (default)
python scripts/translate.py -i output/input.txt
# → output/input_CN.txt
# Local Phi-4 — Translate SRT (preserves timecodes)
python scripts/translate.py -i output/input.srt
# → output/input_CN.srt + output/input_CN.txt
# Online backends
python scripts/translate.py -i output/input.srt -transl_backend deepseek
python scripts/translate.py -i output/input.srt -transl_backend openai -transl_model gpt-5.6-terra
python scripts/translate.py -i output/input.srt -transl_backend anthropic -transl_model claude-opus-4-8
# Translate to Japanese using DeepSeek (flexible mode)
python scripts/translate.py -i output/input.txt -transl_backend deepseek -tgt_lang Japanese -tgt_lang_code JP -mode flexible
# Custom output path (TXT input → .txt output; -o sets dir + filename stem, extension auto-selected)
python scripts/translate.py -i output/input.txt -o D:/output/lecture_JP
# Override source language
python scripts/translate.py -i output/input.txt -src_lang English -tgt_lang Chinesepython main.py -i <input> [options]main.py is the orchestrator that imports and calls scripts.transcribe.transcribe_file() and/or scripts.translate.translate_file() in a single run.
It does not re-implement the pipelines — it calls the same functions the standalone scripts expose.
-transcribe |
-translate |
Behaviour |
|---|---|---|
true (default) |
false (default) |
Transcribe only (ASR + segment → SRT/TXT) |
true |
true |
Transcribe, then translate the SRT |
false |
true |
Translate-only (input must be .srt or .txt) |
false |
false |
Error: at least one must be enabled |
| Argument | Default | Description |
|---|---|---|
-i, --input <path> |
input/input.mp4 |
Input file (video/audio for transcribe, .srt/.txt for translate-only) |
-o, --output <path> |
output/<stem>.srt |
Same as scripts/transcribe.py |
-transcribe |
true |
Enable transcription (true/false) |
-translate |
false |
Enable translation (true/false) |
-seg_backend <name> |
local |
Same as scripts/transcribe.py |
-seg_model <name> |
per-backend | Same as scripts/transcribe.py |
-transl_backend <name> |
local |
Same as scripts/translate.py |
-transl_model <name> |
per-backend | Same as scripts/translate.py |
-src_lang <name> |
English |
Same as scripts/translate.py |
-tgt_lang <name> |
Chinese |
Same as scripts/translate.py |
-tgt_lang_code <code> |
CN |
Same as scripts/translate.py |
-mode <name> |
accurate |
Same as scripts/translate.py |
-max_merge_chars <N> |
no limit | Same as scripts/translate.py |
-max_merge_gap <N> |
no limit | Same as scripts/translate.py |
-gpu_layers <N> |
auto-detect | Same as scripts/transcribe.py and scripts/translate.py |
-no_cache |
— | Same as scripts/transcribe.py |
Note:
-moderequires-translate true. Passing it without translation prints an error and exits.
# Transcribe only (default)
python main.py -i input/input.mp4
# Transcribe + translate to Chinese (local Phi-4)
python main.py -i input/input.mp4 -translate true
# Transcribe + translate with online API (flexible mode)
python main.py -i input/input.mp4 -translate true -transl_backend deepseek -mode flexible
# Translate-only (existing subtitles)
python main.py -transcribe false -translate true -i output/lecture.srt
# Different models for segmentation vs translation
python main.py -i input/input.mp4 -translate true -seg_backend local -seg_model phi-4 -transl_backend openai -transl_model gpt-5.6-terraProcess all files in a folder, one at a time in isolated subprocesses (each file gets its own CUDA context — no VRAM leak between files).
Supports both video (.mp4, .mkv, .avi, .mov) and audio (.wav, .mp3, .m4a, .flac) inputs via the -ext argument.
python batch.py [options]| Argument | Default | Description |
|---|---|---|
-i, --input <dir> |
input/ |
Input folder |
-o, --output <dir> |
output/ |
Output folder |
-ext <ext> |
.mp4 / .srt / .txt |
File extension to search for (e.g. .mp3, .wav, .mkv). Defaults to .mp4 when transcribe is on, .srt & .txt in translate-only mode. |
-transcribe |
true |
Same as main.py |
-translate |
false |
Same as main.py |
-seg_backend <name> |
local |
Same as scripts/transcribe.py |
-seg_model <name> |
per-backend | Same as scripts/transcribe.py |
-transl_backend <name> |
local |
Same as scripts/translate.py |
-transl_model <name> |
per-backend | Same as scripts/translate.py |
-src_lang <name> |
English |
Same as scripts/translate.py |
-tgt_lang <name> |
Chinese |
Same as scripts/translate.py |
-tgt_lang_code <code> |
CN |
Same as scripts/translate.py |
-mode <name> |
accurate |
Same as scripts/translate.py |
-max_merge_chars <N> |
no limit | Same as scripts/translate.py |
-max_merge_gap <N> |
no limit | Same as scripts/translate.py |
-gpu_layers <N> |
auto-detect | Same as scripts/transcribe.py and scripts/translate.py |
-no_cache |
— | Same as scripts/transcribe.py |
# Process all MP4s in input/
python batch.py
# Custom folders
python batch.py -i D:/videos -o D:/output
# Handle MKV files
python batch.py -ext .mkv
# Process audio files (MP3, WAV)
python batch.py -ext .mp3
python batch.py -ext .wav
# Batch + translate all files (local Phi-4)
python batch.py -translate true
# Batch + translate with online API (flexible mode)
python batch.py -translate true -transl_backend deepseek -transl_model deepseek-v4-flash -mode flexible
# Batch translate-only (existing subtitles)
python batch.py -transcribe false -translate true -i D:/subtitles/ -o D:/subtitles_translated/ -transl_backend deepseek
# Online segmentation + local translation (not recommended)
python batch.py -seg_backend openai -seg_model gpt-5.6-terra -translate true.mp4, .mkv, .avi, .mov, .wav, .mp3, .m4a, .flac
M2L3/
├── main.py # Orchestrator — transcribe + translate in one command
├── batch.py # Batch processing entry
├── translate_config.json # Translation user configuration
├── llm_config.json # LLM backend & model configuration
├── README.md # English documentation (This file)
├── README_CN.md # Chinese documentation
├── requirements.txt
├── scripts/
│ ├── transcribe.py # Transcription pipeline (ASR → align → segment → export)
│ └── translate.py # Translation pipeline (6 backends, accurate/flexible modes)
├── models/
│ ├── hub/ # Local model cache
│ │ ├── checkpoints/ # Wav2Vec2 alignment model (~360 MB)
│ │ ├── nltk_data/ # NLTK punkt tokenizers (~64 MB)
│ │ └── snakers4_silero-vad_master/ # Silero VAD (~35 MB)
│ ├── yamnet/ # YAMNet split model (~16 MB, from download_models.py)
│ ├── faster-whisper-large-v3/ # ASR model (~3 GB)
│ └── *.gguf # LLM models for segmentation & translation (Phi-4, etc.)
├── docs/
│ ├── segmentation_pipeline.md # 10-phase LLM segmentation algorithm (English)
│ ├── segmentation_pipeline_CN.md # 10 阶段 LLM 分割算法文档(中文)
│ ├── transcription_n_translation.md # Transcription & translation technical implementation (English)
│ ├── transcription_n_translation_CN.md # 转录与翻译技术实现文档(中文)
│ ├── models_n_performance.md # Models & performance (English)
│ └── models_n_performance_CN.md # 模型与性能(中文)
├── tools/
│ ├── llama/
│ │ └── llama-server.exe # llama.cpp inference server (CUDA)
│ ├── __init__.py # Package marker
│ ├── config.py # Proxy & GPU environment config
│ ├── env_check.py # FFmpeg & CUDA verification
│ ├── llm_call.py # Unified LLM call layer (all backends)
│ ├── extract.py # Word extraction + dedup (overlap/seam/retry)
│ ├── cache.py # Word-level JSON cache
│ ├── format.py # SRT time formatting
│ ├── export.py # SRT / TXT export
│ ├── download_models.py # Auto-download script
│ ├── segment.py # LLM segmentation engine (10-phase)
│ ├── seg_rules.py # Rule-based segmentation helpers
│ ├── seg_diff.py # Character-diff for LLM punctuation
│ ├── non_split_bigrams.py # Phrase protection bigrams
│ ├── music_detect.py # YAMNet music scan + hallucination filter
│ ├── yamnet_frontend.py # YAMNet numpy log-mel frontend (split model)
│ ├── vad_patch.py # Gap-/music-aware VAD merge patch
│ └── _patch_transformers.py # Transformers offline patch (internal)
├── input/ # Default input folder
├── output/ # Default output folder
└── cache/ # Word-level ASR cache
This project is licensed under the MIT License.
{ "openai_api_key": "", "anthropic_api_key": "", "api_base_url": "", "backends": { // online API backends — see table above }, "local_models": { // local GGUF models } }