feat(speechify): add Speechify TTS plugin - #2434
Open
luke-speechify wants to merge 5 commits into
Open
Conversation
🦋 Changeset detectedLatest commit: faef238 The changes in this PR will be included in the next version bump. This PR includes changesets to release 39 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
…s, word spacing, sanitized errors
…lay) and fix per-segment metrics
…failure after first frame instead of retrying
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds
@livekit/agents-plugin-speechify— a Speechify text-to-speech plugin with streaming synthesis and word-level timestamps, built on Speechify's official@speechify/apiSDK. It follows the existing plugin conventions and mirrors thecartesiapackage layout.Features
/v1/audio/stream/with-timestampsendpoint (Server-Sent Events) viaclient.audio.streamWithTimestamps(), so audio and aligned speech marks arrive incrementally while the audio is still being generated. Declarescapabilities: { streaming: true, alignedTranscript: true }.SynthesizeStreamchunks streamed input into sentences and streams one request per sentence;ChunkedStreamhandles one-shotsynthesize(). Both deferfinal: trueto the last frame of the stream and attach each request's word timestamps to a frame (speech-mark times are absolute milliseconds, converted to seconds).pcm_24000).TTSOptions. Defaults to thedominic_32voice and thesimba-3.2model.apiKeyoption or theSPEECHIFY_API_KEYenvironment variable.Custom headers
Every request carries two custom headers so Speechify can attribute API usage to this integration per release:
Speechify-Caller: livekit-typescriptSpeechify-Caller-Version: <plugin version>They are set on the SDK client and merged into every request; the SDK sets no caller of its own, and per-request options do not clobber them.
Tests
@livekit/agents-plugins-test) runs whenSPEECHIFY_API_KEYandOPENAI_API_KEYare set.pnpm build(ESM + CJS + type declarations) andapi-extractorpass.Notes
turbo.jsonenv allowlist (SPEECHIFY_API_KEY).Usage