Skip to content

fix(cli): annotate speaker turns in tts --multi-speaker - #71

Merged
copybara-service[bot] merged 1 commit into
mainfrom
copybara/994564718
Oct 6, 2026
Merged

copybara-service[bot] merged 1 commit into
mainfrom
copybara/994564718

Conversation

@copybara-service

@copybara-service copybara-service Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

fix(cli): annotate speaker turns in tts --multi-speaker

The default TTS model moved to gemini-3.8-flash-tts, which rejects a multi-speaker script sent as one text block ("Multi-speaker interactions must specify a speaker for each text turn.", HTTP 400). Every --multi-speaker call on the default model failed, including the tts --help example.

Split the script into one text block per turn, each annotated with speech_metadata naming its speaker, and strip the <Speaker>: prefix from the turn text (Gemini 3.8 TTS treats text as a verbatim transcript and takes the speaker via speech_metadata). A turn starts at a <Speaker>: label for a declared --multi-speaker name: at the start of a line (optionally after a list or quote marker: -, •, >, “, 1.) or mid-line after punctuation ("Alice: hi. Bob: yo.", "Alice:Bob:"), never inside a longer word ("MaryAnn:", "श्रीराम:") or after a word on the same line ("I told Bob: wait."), with any Unicode space around the name. Words before the first label, no label at all, or only empty turns are usage errors (exit 2) instead of the API's opaque 400.

Legacy gemini-2.5-* and gemini-3.1-* TTS models reject speech_metadata annotations, so they keep the single plain block.

The "speaker does not appear in the text" warning now follows the parsed turns on annotating models and keeps the word match on the legacy ones. -f / --stdin text drops a leading UTF-8 byte-order mark, so Windows-saved scripts parse.

@copybara-service
copybara-service Bot requested a review from a team October 6, 2026 21:15
@copybara-service
copybara-service Bot force-pushed the copybara/994564718 branch 5 times, most recently from c88d434 to 9b91412 Compare October 6, 2026 22:35
The default TTS model moved to `gemini-3.8-flash-tts`, which rejects a multi-speaker script sent as one text block ("Multi-speaker interactions must specify a speaker for each text turn.", HTTP 400). Every `--multi-speaker` call on the default model failed, including the `tts --help` example.

Split the script into one text block per turn, each annotated with `speech_metadata` naming its speaker, and strip the `<Speaker>:` prefix from the turn text (Gemini 3.8 TTS treats `text` as a verbatim transcript and takes the speaker via `speech_metadata`). A turn starts at a `<Speaker>:` label for a declared `--multi-speaker` name: at the start of a line (optionally after a list or quote marker: `-`, `•`, `>`, `“`, `1.`) or mid-line after punctuation (`"Alice: hi. Bob: yo."`, `"Alice:Bob:"`), never inside a longer word (`"MaryAnn:"`, `"श्रीराम:"`) or after a word on the same line (`"I told Bob: wait."`), with any Unicode space around the name. Words before the first label, no label at all, or only empty turns are usage errors (exit 2) instead of the API's opaque 400.

Legacy `gemini-2.5-*` and `gemini-3.1-*` TTS models reject `speech_metadata` annotations, so they keep the single plain block.

The "speaker does not appear in the text" warning now follows the parsed turns on annotating models and keeps the word match on the legacy ones. `-f` / `--stdin` text drops a leading UTF-8 byte-order mark, so Windows-saved scripts parse.

PiperOrigin-RevId: 994669788
@copybara-service
copybara-service Bot merged commit 1f78c8a into main Oct 6, 2026
2 checks passed
@copybara-service
copybara-service Bot deleted the copybara/994564718 branch October 6, 2026 22:38
2ynn added a commit that referenced this pull request Oct 6, 2026
Bounding lineMarkerStart at the previous label also made "only
whitespace since the previous label" count as the start of a line, so
a short turn after a label was read as the next label's list marker
and dropped with it. #71 fixed that for numbers (`Alice: 42. Bob:`), but
punctuation-only turns were still lost: `Alice: ... Bob: Yes?`,
`Alice: — Bob: Sorry, go on.` and `Alice: ?! Bob: What?` each sent only
Bob's turn.

Judge the line start on the whole text for every marker. The scan still
stops at the previous label, so emoji names (`😀: 🤖:`) don't panic.
The live API accepts annotated turns of only `...`, `—` or `-` (HTTP
200).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants