Official Python & JavaScript clients for the Gandr TTS API, text to speech built for voice agents.
- Audio streams back as it is generated, so playback can start before the whole clip has rendered. First audio byte in 146 ms over the open internet, 116 ms p50 first audio, server side warm
- WER 1.982% on a 1,088-line set, the human recordings score 2.171% on the same scorer
- $10 a month for one million tokens (one token is one character), or unlimited, unmetered stream plans from $150/mo (annual)
- Every render watermarked (imperceptible, detectable)
- Numbers, dates, addresses, and order IDs read back correctly
pip install gandr # Python 3.9+, zero dependencies
npm install gandr # Node 18+, zero dependenciesPython
from gandr import Gandr
g = Gandr("gnd_...") # your API key, https://gandr.ai
audio = g.say("Your table for two is confirmed for Thursday at seven.")
open("confirmation.wav", "wb").write(audio)JavaScript
import { Gandr } from "gandr";
const g = new Gandr("gnd_...");
const wav = await g.say("Your table for two is confirmed for Thursday at seven.");
// wav is a Uint8Array of a WAV fileg.say(
"Order number 4-2-7-1 ships on March 3rd.",
voice="gandr-jenny", # ava, dane, jenny, leo, lewis, mia
sample_rate=8000, # 8000-48000, resampled server-side (telephony: 8000)
temperature=0.9, # 0.1-1.2, pitch range / melody
cfg_weight=0.4, # 0.2-1.0, pacing (lower = more spacious)
speed=1.1, # 0.6-1.5
pronunciation=[{"text": "Nguyen", "pronunciation": "win"}],
)Omit a dial and you get the tuned default, per-voice temperature tuning included.
gandr-ava |
gandr-leo |
gandr-dane |
gandr-lewis |
gandr-jenny |
gandr-mia |
g.voices() returns the live catalog.
The client moves to the next endpoint when one is unreachable. A real answer, including an error, is never retried elsewhere, so you always see the response you actually got.
from gandr import Gandr, GandrError
try:
g.say("...")
except GandrError as e:
e.status # 401 invalid key · 402 quota spent · 400 bad input · 429 concurrency
e.payload # the API's JSON, with a hint field- One request carries up to 2,000 characters, split longer text at sentence boundaries.
- The streaming WebSocket lane (
wss://tts.gandr.ai/ws) is what voice agents should use. This SDK'ssay()returns the complete file, which suits batch and IVR work.
- Product: https://gandr.ai · Docs: https://gandr.ai/docs
- API spec: https://gandr.ai/openapi.yml · Status: https://gandr.ai/status
- Security disclosures: https://gandr.ai/.well-known/security.txt