stt-tts mode (talk.speak calls this
same synthesis path). Provider-native realtime Talk sessions synthesize
speech inside the realtime provider instead. transcription sessions never
synthesize an assistant voice reply.
This page is an index. Text-to-speech is documented on seven pages, one per
reader job. Open the page that matches your task.
Where each section moved
Every section heading from the previous single-page version keeps its anchor here, so an existing link such as/tools/tts#per-agent-voice-overrides still
resolves. Each entry points at the page that now holds the content.
- Quick start
- Supported providers
- Configuration
- Local Speech Swift and speech-core
- Per-agent voice overrides
- Personas
- Minimal persona
- Full persona (provider-specific shaping)
- Persona resolution
- Custom persona shaping
- Fallback policy
- Model-driven directives
- Slash commands
- Per-user preferences
- Output formats
- Auto-TTS behavior
- Field reference
- Inworld primary
- Agent tool
- Gateway RPC
- Full persona (provider-specific shaping)
Component anchors
The previous single-page version also minted an anchor for every step, tab, accordion, and field. Those anchors are preserved here so that any deep link into the old page still resolves. Nine accordion anchors lost a-1 suffix
when the tab that shared their slug moved to a different page. The stub below
keeps the old id and points at the new one.
Quickstart
Configuration
- Azure Speech
- ElevenLabs
- Fish Audio
- Google Gemini
- Gradium
- Inworld
- Local CLI
- Microsoft (no key)
- MiniMax
- OpenAI + ElevenLabs
- OpenRouter
- Volcengine
- xAI
- Xiaomi MiMo
- macOS HTTP
- macOS CLI
- Linux CLI
- Windows CLI
- Top-level tts.*
- Top-level tts.* →
auto - Top-level tts.* →
enabled - Top-level tts.* →
mode - Top-level tts.* →
provider - Top-level tts.* →
persona - Top-level tts.* →
personas.<id> - Top-level tts.* →
summaryModel - Top-level tts.* →
modelOverrides - Top-level tts.* →
providers.<id> - Top-level tts.* →
maxTextLength - Top-level tts.* →
timeoutMs - Azure Speech
- Azure Speech →
apiKey - Azure Speech →
region - Azure Speech →
endpoint - Azure Speech →
speakerVoice - Azure Speech →
lang - Azure Speech →
outputFormat - Azure Speech →
voiceNoteOutputFormat - ElevenLabs
- ElevenLabs →
apiKey - ElevenLabs →
model - ElevenLabs →
speakerVoiceId - ElevenLabs →
voiceSettings - ElevenLabs →
applyTextNormalization - ElevenLabs →
languageCode - ElevenLabs →
seed - ElevenLabs →
baseUrl - Google Gemini
- Google Gemini →
apiKey - Google Gemini →
model - Google Gemini →
speakerVoice - Google Gemini →
audioProfile - Google Gemini →
speakerName - Google Gemini →
promptTemplate - Google Gemini →
personaPrompt - Google Gemini →
baseUrl - Gradium
- Gradium →
apiKey - Gradium →
baseUrl - Gradium →
speakerVoiceId - Inworld
- Inworld →
apiKey - Inworld →
baseUrl - Inworld →
modelId - Inworld →
speakerVoiceId - Inworld →
temperature - Local CLI (tts-local-cli)
- Local CLI (tts-local-cli) →
command - Local CLI (tts-local-cli) →
args - Local CLI (tts-local-cli) →
outputFormat - Local CLI (tts-local-cli) →
timeoutMs - Local CLI (tts-local-cli) →
cwd - Local CLI (tts-local-cli) →
env - Microsoft (no API key)
- Microsoft (no API key) →
enabled - Microsoft (no API key) →
speakerVoice - Microsoft (no API key) →
lang - Microsoft (no API key) →
outputFormat - Microsoft (no API key) →
rate / pitch / volume - Microsoft (no API key) →
saveSubtitles - Microsoft (no API key) →
proxy - Microsoft (no API key) →
timeoutMs - Microsoft (no API key) →
edge.* - MiniMax
- MiniMax →
apiKey - MiniMax →
baseUrl - MiniMax →
model - MiniMax →
speakerVoiceId - MiniMax →
speed - MiniMax →
vol - MiniMax →
pitch - OpenAI
- OpenAI →
apiKey - OpenAI →
model - OpenAI →
speakerVoice - OpenAI →
instructions - OpenAI →
responseFormat - OpenAI →
extraBody / extra_body - OpenAI →
baseUrl - OpenRouter
- OpenRouter →
apiKey - OpenRouter →
baseUrl - OpenRouter →
model - OpenRouter →
speakerVoice - OpenRouter →
responseFormat - OpenRouter →
speed - Volcengine (BytePlus Seed Speech)
- Volcengine (BytePlus Seed Speech) →
apiKey - Volcengine (BytePlus Seed Speech) →
resourceId - Volcengine (BytePlus Seed Speech) →
appKey - Volcengine (BytePlus Seed Speech) →
baseUrl - Volcengine (BytePlus Seed Speech) →
speakerVoice - Volcengine (BytePlus Seed Speech) →
speedRatio - Volcengine (BytePlus Seed Speech) →
emotion - Volcengine (BytePlus Seed Speech) →
appId / token / cluster - xAI
- xAI →
apiKey - xAI →
baseUrl - xAI →
speakerVoiceId - xAI →
language - xAI →
responseFormat - xAI →
speed - Xiaomi MiMo
- Xiaomi MiMo →
apiKey - Xiaomi MiMo →
baseUrl - Xiaomi MiMo →
model - Xiaomi MiMo →
speakerVoice - Xiaomi MiMo →
format - Xiaomi MiMo →
style
Service links
- Azure Speech provider
- Azure Speech REST text-to-speech
- ElevenLabs provider
- ElevenLabs Authentication
- ElevenLabs Text to Speech
- Fish Audio provider
- Gradium
- Inworld provider
- Inworld TTS API
- Microsoft Speech output formats
- MiniMax provider
- MiniMax T2A v2 API
- node-edge-tts
- OpenAI provider
- OpenAI Audio API reference
- OpenAI text-to-speech guide
- speech-core
- Speech Swift
- Volcengine TTS HTTP API
- xAI provider
- xAI text to speech
- Xiaomi MiMo speech synthesis