Skip to main content
Talk mode covers these runtime shapes:
  • Native macOS/iOS/Android Talk: native speech recognition, Gateway chat, and talk.speak TTS. Apple Speech recognition on macOS/iOS may use network services. Android behavior depends on the installed speech service. Nodes advertise the talk capability and declare which talk.* commands they support.
  • iOS Talk (realtime): client-owned WebRTC for OpenAI realtime configs that select webrtc transport or omit transport. This path includes framed and frameless transcript/audio events. Explicit gateway-relay, provider-websocket, and non-OpenAI realtime configs stay on the Gateway-owned relay. Non-realtime configs use the native speech loop.
  • Apple Watch standalone Talk: native WebRTC/Opus over UDP with Gateway-owned call control (gateway-control-v1). The Watch uses the Gateway’s configured realtime provider. It keeps tools and transcript ownership on the Gateway. Unsupported configurations fail visibly without a relay fallback.
  • Browser Talk: talk.client.create for client-owned webrtc/provider-websocket sessions, or talk.session.create for Gateway-owned gateway-relay sessions. managed-room is reserved for Gateway handoff and walkie-talkie rooms.
  • Android Talk (realtime): Android uses Gateway-owned relay realtime when talk.catalog reports the realtime group ready. For supported GPT-Live models, the Gateway advertises relay capability through talk.config, so Android starts talk.session.create instead of native speech recognition and talk.speak. Android never opens a client-owned WebRTC session. Older Gateways without the capability hint and unlisted GPT-Live routes retain native Talk with a visible explanation. Explicit stt-tts mode also stays native. Relay capability is not authentication readiness: catalog and session creation still validate the selected account. Subscription-authenticated Android live audio verification remains pending.
  • Transcription-only clients: talk.session.create({ mode: "transcription", transport: "gateway-relay", brain: "none" }), then talk.session.appendAudio and talk.session.close for captions/dictation without an assistant voice response. One-shot uploaded voice notes still use the media understanding audio path.
Native Talk is a continuous loop. It listens for speech. It sends the transcript to the model through the active session. It waits for the response. It then speaks the response through the configured Talk provider (talk.speak). For replies dominated by fenced code, talk.speak uses a short spoken message directing the listener to the screen. Inline code and ordinary prose remain part of the spoken reply. Apple Watch also retains Talk to Claw, the separate one-turn companion flow. That flow uses native dictation, text relayed through the iPhone, and system-voice readback. Talk on Watch is the realtime path included in normal Watch setup. See standalone voice setup.

Talk documentation pages

Talk mode is documented on this page and four child pages, one per reader job. This page keeps the voice directives and the talk configuration reference. Open the child page that matches your task.

Where each section moved

Every section heading from the previous single-page version keeps its anchor here, so an existing link such as /nodes/talk#session-ownership still resolves. Each entry points at the page that now holds the content.

Voice directives in replies

The assistant can prefix a reply with a single JSON line to control voice:
Rules:
  • First non-empty line only. The JSON line is stripped before TTS playback.
  • Unknown keys are ignored.
  • once: true applies to the current reply only. Without it, the voice becomes the new Talk mode default.
Supported keys: voice / voice_id / voiceId, model / model_id / modelId, speed, rate (WPM), stability, similarity, style, speakerBoost, seed, normalize, lang, output_format, latency_tier, once.

Config (~/.openclaw/openclaw.json)

Angle-bracket values such as <elevenlabs-api-key> are placeholders. Replace them with your own values.
OpenAI browser WebRTC and Gateway-relay Talk support native GPT-Live. Select gpt-live-1 for the public API or gpt-live-1-codex for the Codex route in Settings → Talk. The public API requires a Platform key; the Codex route prefers an OpenClaw ChatGPT OAuth profile and falls back to Platform API-key authentication. Browser Talk uses client WebRTC with Gateway-owned control. Gateway relay uses direct Platform-key WebSockets for gpt-live-1 and Gateway-owned WebRTC for gpt-live-1-codex. Discord uses these same Gateway bridges when configured with the corresponding Live model. Account-issued, unlisted routes can be set in talk.realtime.model, but are not published through catalogs or diagnostics. They never use OAuth and require a Platform key. GPT-Live browser Talk also requires the bundled openai plugin registered in full mode. A restrictive plugins.allow list fails session creation with “OpenAI GPT-Live browser session broker is unavailable”. Runtime bounds: 8 concurrent sessions per Gateway and a 30-minute session TTL. Browser sessions also use 60-second single-use offer tokens. The Codex route uses arbor, breeze, cove, ember, juniper, maple, sol, spruce, and vale, with cove as the default. The public API defaults to marin and has its own voice list. Choose the same model and supported voice in Talk and Discord to use the same voice; cove with the public gpt-live-1 model falls back to marin. Unlisted routes use their account-issued voice contract; the current unlisted Platform profile accepts marin and cedar. A rejected session does not identify the cause by itself. Check the selected account, model, and voice. GPT-Live handles interruption natively and produces continuous audio without requiring completed-response events. Discord preserves each speaker’s identity through delegated OpenClaw work. Its voice-model connections remain separate per speaker; shared room context belongs to the OpenClaw agent conversation. See GPT-Live in Discord for configuration and the limits on host-enforced meeting participation. These rows describe implemented transport paths, not account entitlement or a successful live call on every device. iOS implements frameless transcripts and the Gateway offer exchange. Android follows the Gateway relay capability hint and retains its legacy GPT-Live model gate only when that hint is absent. For model capability limits, see Discord voice policies and Voice Call tools. The Gateway-owned WebRTC route keeps OAuth and Platform credentials away from relay clients. Backend WebSocket paths keep the Platform key on the Gateway. OpenClaw converts telephony G.711 u-law audio to and from GPT-Live’s 24 kHz PCM contract. For GA gpt-realtime-2.1, gpt-realtime-2.1-mini, and gpt-realtime-2 browser sessions, Platform credentials remain preferred in this order: the configured realtime API key, an openai API-key profile, then OPENAI_API_KEY. With none configured, browser Talk falls back to an OpenClaw ChatGPT OAuth profile and exchanges SDP through the Gateway’s single-use offer broker, so the OAuth token never reaches the browser. A configured Platform credential that cannot be resolved fails closed instead of silently falling through to OAuth. iOS client-owned WebRTC and GA Gateway relay, including Android GA realtime, remain Platform-key-only. GA browser Talk keeps the existing client-owned data channel and talk.client.toolCall loop. Only the credential owner and SDP exchange path change under OAuth. The Codex GPT-Live route remains OAuth-first with Platform fallback for browser and Gateway-owned WebRTC, including Android Talk and Discord. Public GPT-Live, direct backend sockets, and unlisted GPT-Live routes remain Platform-key-only. talk.catalog exposes canonical provider ids and registry aliases. It exposes each provider’s valid modes/transports/brain strategies/realtime audio formats/capability flags. It exposes the runtime-selected readiness result. First-party Talk clients should read that catalog instead of maintaining provider aliases locally. Treat an older Gateway that omits group readiness as unverified rather than definitively unconfigured. Streaming transcription providers are discovered through talk.catalog.transcription. The current Gateway relay uses the Voice Call streaming provider config until a dedicated Talk transcription config surface ships.

Notes

  • Native speech recognition requires the platform’s speech and microphone access. Standalone Watch realtime requires microphone access, not local speech recognition.
  • Native Talk uses the active Gateway session and only falls back to history polling when response events are unavailable.
  • The gateway resolves Talk playback through talk.speak using the active Talk provider. Android falls back to local system TTS only when that RPC is unavailable.
  • macOS local MLX playback uses the bundled openclaw-mlx-tts helper when present, or an executable on PATH. Set OPENCLAW_MLX_TTS_BIN to point at a custom helper binary during development. The helper streams PCM, keeps one selected model resident, and supports Fish S2 Pro reference audio through providers.mlx.referenceAudioPath plus referenceText.
  • Voice directive value ranges (ElevenLabs): the stability, similarity, and style keys accept 0..1. The speed key accepts 0.5..2. The latency_tier key accepts 0..4.