- Native macOS/iOS/Android Talk: native speech recognition, Gateway chat, and
talk.speakTTS. Apple Speech recognition on macOS/iOS may use network services. Android behavior depends on the installed speech service. Nodes advertise thetalkcapability and declare whichtalk.*commands they support. - iOS Talk (realtime): client-owned WebRTC for OpenAI realtime configs that select
webrtctransport or omit transport. This path includes framed and frameless transcript/audio events. Explicitgateway-relay,provider-websocket, and non-OpenAI realtime configs stay on the Gateway-owned relay. Non-realtime configs use the native speech loop. - Apple Watch standalone Talk: native WebRTC/Opus over UDP with Gateway-owned call control (
gateway-control-v1). The Watch uses the Gateway’s configured realtime provider. It keeps tools and transcript ownership on the Gateway. Unsupported configurations fail visibly without a relay fallback. - Browser Talk:
talk.client.createfor client-ownedwebrtc/provider-websocketsessions, ortalk.session.createfor Gateway-ownedgateway-relaysessions.managed-roomis reserved for Gateway handoff and walkie-talkie rooms. - Android Talk (realtime): Android uses Gateway-owned relay realtime when
talk.catalogreports the realtime group ready. For supported GPT-Live models, the Gateway advertises relay capability throughtalk.config, so Android startstalk.session.createinstead of native speech recognition andtalk.speak. Android never opens a client-owned WebRTC session. Older Gateways without the capability hint and unlisted GPT-Live routes retain native Talk with a visible explanation. Explicitstt-ttsmode also stays native. Relay capability is not authentication readiness: catalog and session creation still validate the selected account. Subscription-authenticated Android live audio verification remains pending. - Transcription-only clients:
talk.session.create({ mode: "transcription", transport: "gateway-relay", brain: "none" }), thentalk.session.appendAudioandtalk.session.closefor captions/dictation without an assistant voice response. One-shot uploaded voice notes still use the media understanding audio path.
talk.speak).
For replies dominated by fenced code, talk.speak uses a short spoken message directing the listener to the screen. Inline code and ordinary prose remain part of the spoken reply.
Apple Watch also retains Talk to Claw, the separate one-turn companion flow. That flow uses native dictation, text relayed through the iPhone, and system-voice readback. Talk on Watch is the realtime path included in normal Watch setup. See standalone voice setup.
Talk documentation pages
Talk mode is documented on this page and four child pages, one per reader job. This page keeps the voice directives and thetalk configuration reference.
Open the child page that matches your task.
Where each section moved
Every section heading from the previous single-page version keeps its anchor here, so an existing link such as/nodes/talk#session-ownership still
resolves. Each entry points at the page that now holds the content.
- Choose a Talk voice from chat
- Session ownership
- Behavior (macOS)
- Behavior (macOS)
- Realtime Talk over the Gateway relay (macOS)
- Realtime Talk over the Gateway relay (macOS)
- When realtime cannot start
- macOS UI
- Apple Watch UI
- Android UI
Voice directives in replies
The assistant can prefix a reply with a single JSON line to control voice:- First non-empty line only. The JSON line is stripped before TTS playback.
- Unknown keys are ignored.
once: trueapplies to the current reply only. Without it, the voice becomes the new Talk mode default.
voice / voice_id / voiceId, model / model_id / modelId, speed, rate (WPM), stability, similarity, style, speakerBoost, seed, normalize, lang, output_format, latency_tier, once.
Config (~/.openclaw/openclaw.json)
Angle-bracket values such as <elevenlabs-api-key> are placeholders. Replace them
with your own values.
gpt-live-1 for the public API or gpt-live-1-codex for the Codex route in
Settings → Talk. The public API requires a Platform key; the Codex route
prefers an OpenClaw ChatGPT OAuth profile and falls back to Platform API-key
authentication. Browser Talk uses client WebRTC with Gateway-owned control.
Gateway relay uses direct Platform-key WebSockets for gpt-live-1 and
Gateway-owned WebRTC for gpt-live-1-codex. Discord uses these same Gateway
bridges when configured with the corresponding Live model.
Account-issued, unlisted routes can be set in talk.realtime.model, but are
not published through catalogs or diagnostics. They never use OAuth and require
a Platform key. GPT-Live browser Talk also requires the bundled openai plugin
registered in full mode. A restrictive plugins.allow list fails session
creation with “OpenAI GPT-Live
browser session broker is unavailable”.
Runtime bounds: 8 concurrent sessions per Gateway and a 30-minute session TTL.
Browser sessions also use 60-second single-use offer tokens.
The Codex route uses arbor, breeze, cove, ember, juniper, maple,
sol, spruce, and vale, with cove as the default. The public API defaults
to marin and has its own voice list.
Choose the same model and supported voice in Talk and Discord to use the same
voice; cove with the public gpt-live-1 model falls back to marin.
Unlisted routes use their account-issued voice contract; the current unlisted
Platform profile accepts marin and cedar. A rejected session does not
identify the cause by itself. Check the selected account, model, and voice.
GPT-Live handles interruption natively and produces continuous audio without
requiring completed-response events. Discord preserves each speaker’s identity
through delegated OpenClaw work. Its voice-model connections remain separate
per speaker; shared room context belongs to the OpenClaw agent conversation.
See GPT-Live in Discord
for configuration and the limits on host-enforced meeting participation.
These rows describe implemented transport paths, not account entitlement or a
successful live call on every device. iOS implements frameless transcripts and
the Gateway offer exchange. Android follows the Gateway relay capability hint and retains its legacy GPT-Live model gate only when that hint is absent.
For model capability limits, see Discord voice policies
and Voice Call tools.
The Gateway-owned WebRTC route keeps OAuth and Platform credentials away from
relay clients. Backend WebSocket paths keep the Platform key on the Gateway.
OpenClaw converts telephony G.711 u-law audio to and from GPT-Live’s 24 kHz PCM
contract.
For GA
gpt-realtime-2.1, gpt-realtime-2.1-mini, and gpt-realtime-2
browser sessions, Platform credentials remain preferred in this order: the
configured realtime API key, an openai API-key profile, then
OPENAI_API_KEY. With none configured, browser Talk falls back to an OpenClaw
ChatGPT OAuth profile and exchanges SDP through the Gateway’s single-use offer
broker, so the OAuth token never reaches the browser. A configured Platform
credential that cannot be resolved fails closed instead of silently falling
through to OAuth.
iOS client-owned WebRTC and GA Gateway relay, including Android GA realtime,
remain Platform-key-only. GA browser Talk keeps the existing client-owned data channel
and talk.client.toolCall loop. Only the credential owner and SDP exchange path
change under OAuth. The Codex GPT-Live route remains OAuth-first with
Platform fallback for browser and Gateway-owned WebRTC, including Android Talk
and Discord. Public GPT-Live, direct backend sockets, and unlisted GPT-Live routes remain
Platform-key-only.
talk.catalog exposes canonical provider ids and registry aliases. It exposes each provider’s valid modes/transports/brain strategies/realtime audio formats/capability flags. It exposes the runtime-selected readiness result. First-party Talk clients should read that catalog instead of maintaining provider aliases locally. Treat an older Gateway that omits group readiness as unverified rather than definitively unconfigured. Streaming transcription providers are discovered through talk.catalog.transcription. The current Gateway relay uses the Voice Call streaming provider config until a dedicated Talk transcription config surface ships.
Notes
- Native speech recognition requires the platform’s speech and microphone access. Standalone Watch realtime requires microphone access, not local speech recognition.
- Native Talk uses the active Gateway session and only falls back to history polling when response events are unavailable.
- The gateway resolves Talk playback through
talk.speakusing the active Talk provider. Android falls back to local system TTS only when that RPC is unavailable. - macOS local MLX playback uses the bundled
openclaw-mlx-ttshelper when present, or an executable onPATH. SetOPENCLAW_MLX_TTS_BINto point at a custom helper binary during development. The helper streams PCM, keeps one selected model resident, and supports Fish S2 Pro reference audio throughproviders.mlx.referenceAudioPathplusreferenceText. - Voice directive value ranges (ElevenLabs): the
stability,similarity, andstylekeys accept0..1. Thespeedkey accepts0.5..2. Thelatency_tierkey accepts0..4.
Related
- Voice wake
- Audio and voice notes
- Media understanding
- Google Meet plugin
- Media overview — how the media tools fit together