Repository navigation
fix(cli): annotate speaker turns in tts --multi-speaker - #71
Merged
Merged
Conversation
copybara-service
Bot
force-pushed
the
copybara/994564718
branch
5 times, most recently
from
October 6, 2026 22:35
c88d434 to
9b91412
Compare
The default TTS model moved to `gemini-3.8-flash-tts`, which rejects a multi-speaker script sent as one text block ("Multi-speaker interactions must specify a speaker for each text turn.", HTTP 400). Every `--multi-speaker` call on the default model failed, including the `tts --help` example.
Split the script into one text block per turn, each annotated with `speech_metadata` naming its speaker, and strip the `<Speaker>:` prefix from the turn text (Gemini 3.8 TTS treats `text` as a verbatim transcript and takes the speaker via `speech_metadata`). A turn starts at a `<Speaker>:` label for a declared `--multi-speaker` name: at the start of a line (optionally after a list or quote marker: `-`, `•`, `>`, `“`, `1.`) or mid-line after punctuation (`"Alice: hi. Bob: yo."`, `"Alice:Bob:"`), never inside a longer word (`"MaryAnn:"`, `"श्रीराम:"`) or after a word on the same line (`"I told Bob: wait."`), with any Unicode space around the name. Words before the first label, no label at all, or only empty turns are usage errors (exit 2) instead of the API's opaque 400.
Legacy `gemini-2.5-*` and `gemini-3.1-*` TTS models reject `speech_metadata` annotations, so they keep the single plain block.
The "speaker does not appear in the text" warning now follows the parsed turns on annotating models and keeps the word match on the legacy ones. `-f` / `--stdin` text drops a leading UTF-8 byte-order mark, so Windows-saved scripts parse.
PiperOrigin-RevId: 994669788
copybara-service
Bot
force-pushed
the
copybara/994564718
branch
from
October 6, 2026 22:38
9b91412 to
1f78c8a
Compare
2ynn
added a commit
that referenced
this pull request
Oct 6, 2026
Bounding lineMarkerStart at the previous label also made "only whitespace since the previous label" count as the start of a line, so a short turn after a label was read as the next label's list marker and dropped with it. #71 fixed that for numbers (`Alice: 42. Bob:`), but punctuation-only turns were still lost: `Alice: ... Bob: Yes?`, `Alice: — Bob: Sorry, go on.` and `Alice: ?! Bob: What?` each sent only Bob's turn. Judge the line start on the whole text for every marker. The scan still stops at the previous label, so emoji names (`😀: 🤖:`) don't panic. The live API accepts annotated turns of only `...`, `—` or `-` (HTTP 200).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
fix(cli): annotate speaker turns in
tts --multi-speakerThe default TTS model moved to
gemini-3.8-flash-tts, which rejects a multi-speaker script sent as one text block ("Multi-speaker interactions must specify a speaker for each text turn.", HTTP 400). Every--multi-speakercall on the default model failed, including thetts --helpexample.Split the script into one text block per turn, each annotated with
speech_metadatanaming its speaker, and strip the<Speaker>:prefix from the turn text (Gemini 3.8 TTS treatstextas a verbatim transcript and takes the speaker viaspeech_metadata). A turn starts at a<Speaker>:label for a declared--multi-speakername: at the start of a line (optionally after a list or quote marker:-,•,>,“,1.) or mid-line after punctuation ("Alice: hi. Bob: yo.","Alice:Bob:"), never inside a longer word ("MaryAnn:","श्रीराम:") or after a word on the same line ("I told Bob: wait."), with any Unicode space around the name. Words before the first label, no label at all, or only empty turns are usage errors (exit 2) instead of the API's opaque 400.Legacy
gemini-2.5-*andgemini-3.1-*TTS models rejectspeech_metadataannotations, so they keep the single plain block.The "speaker does not appear in the text" warning now follows the parsed turns on annotating models and keeps the word match on the legacy ones.
-f/--stdintext drops a leading UTF-8 byte-order mark, so Windows-saved scripts parse.