Voice
Discord has two distinct voice surfaces: realtime voice channels (continuous conversations) and voice message attachments (the waveform preview format). The gateway supports both.Voice channels
Setup checklist:- Enable Message Content Intent in the Discord Developer Portal.
- Enable Server Members Intent when role/user allowlists are used.
- Invite the bot with
botandapplications.commandsscopes. - Grant Connect, Speak, Send Messages, and Read Message History in the target voice channel.
- Enable native commands (
commands.nativeorchannels.discord.commands.native). - Configure
channels.discord.voice.
/vc join|leave|status to control sessions. The command uses the account default agent and follows the same allowlist and group policy rules as other Discord commands. Type these as slash commands in a Discord message box, not in a shell:
GPT-Live in Discord
Discord can use the same GPT-Live model and voice as Talk, Google Meet, and Voice Call. These surfaces share provider resolution, native delegation, and interruption policy while retaining their own audio transports. Unpinned Discord configurations keep the provider’s existing default; select GPT-Live explicitly. For the Codex GPT-Live route withcove, sign in with
openclaw models auth login --provider openai, then configure:
gpt-live-1 and a
supported public voice such as marin; Discord uses the Gateway’s direct
Platform-key WebSocket bridge. The model and voice must belong to the same
route: cove is a Codex GPT-Live voice, and selecting it with gpt-live-1
falls back to marin. See OpenAI voice and speech.
GPT-Live owns response timing and interruption. Discord plays its continuous
audio without waiting for a completed-response event, and does not add local
speaker-start interruption. Microphone audio remains admitted during playback
so GPT-Live can hear and handle interruptions itself. Delegated tasks use the routed OpenClaw agent with
the originating speaker’s Discord identity and tool permissions. Room controls
clear buffered local playback before requesting their spoken result; they keep
the continuous provider stream open so it can deliver that result.
The Gateway paces microphone input continuously, including silence between
speaker captures. Playback preserves quiet PCM within an active stream, including
pauses delivered after earlier speech has already played. When no unheard speech
remains for the player’s two-second idle grace, accepted audio drains before the
room player is released. Speech arriving during that retirement queues under a
fresh resource so it is not discarded with the old stream.
Continuous playback builds a 120 ms startup buffer to absorb brief delivery
gaps. A 120 ms startup deadline keeps short replies from waiting for a completed
response event.
GPT-Live does not support host-enforced wake-name gating or a forced consult
before every reply. Its defaults leave those policies off. Explicit
requireWakeName: true or consultPolicy: "always" fails at startup with a
configuration error; remove those settings or choose gpt-realtime-2.1 when
you need host-controlled turns in a shared meeting. Each speaker still has a
separate voice-model connection; the OpenClaw agent conversation provides
shared room context.
Notes:
- The OpenAI
agent-proxyresponse and wake-name policies below require a GA realtime model, such asgpt-realtime-2.1. GPT-Live follows its provider-owned response and delegation flow described above. - Discord voice is opt-in for text-only configs; set
channels.discord.voice.enabled=true(or keep an existingchannels.discord.voiceblock) to enable/vccommands, the voice runtime, and theGuildVoiceStatesgateway intent.channels.discord.intents.voiceStatescan explicitly override the intent subscription; leave it unset to follow effective voice enablement. voice.modecontrols the conversation path. The default isagent-proxy: a realtime voice front end handles turn timing, interruption, and playback, delegates substantive work to the routed OpenClaw agent throughopenclaw_agent_consult, and treats the result like a typed Discord prompt from that speaker.stt-ttskeeps the older batch STT plus TTS flow.bidilets the realtime model converse directly while exposingopenclaw_agent_consultfor the OpenClaw brain.- Realtime voice keeps each speaker’s audio in a separate provider connection, so delayed transcripts and tool calls retain that speaker’s Discord identity. Everyone still uses the same routed OpenClaw agent conversation and one room playback queue. Direct
bidiconversation history belongs to each speaker’s realtime connection; use the OpenClaw agent consult for shared room history. Multiple speakers can consume more provider connections, and provider account limits still apply. The room retains at most eight speaker connections and reclaims idle connections once their captures, requests, and playback have finished. voice.agentSessioncontrols which OpenClaw conversation receives voice turns. Leave it unset for the voice channel’s own session, or set{ mode: "target", target: "channel:<text-channel-id>" }to make the voice channel act as the microphone/speaker extension of an existing Discord text channel session such as#maintainers.- Steering a running agent from Discord voice requires a backend that can recheck the speaker’s permissions at dispatch. Copilot currently lacks this guarded input capability; status and cancellation remain available, and you can start a new request.
voice.modeloverrides the OpenClaw agent brain for Discord voice responses and realtime consults. Leave it unset to inherit the routed agent model. It is separate fromvoice.realtime.model.voice.followUserslets the bot join, move, and leave Discord voice with selected users. See Follow users in voice.agent-proxyroutes speech throughdiscord-voice, which preserves normal owner/tool authorization for the speaker and target session but hides the agentttstool because Discord voice owns playback. By default,agent-proxygives the consult full owner-equivalent tool access for owner speakers (voice.realtime.toolPolicy: "owner"). Models with host-controlled turns strongly prefer consulting the OpenClaw agent before substantive answers (voice.realtime.consultPolicy: "always"); GPT-Live uses provider-owned delegation withconsultPolicy: "auto". Inalwaysmode, the realtime layer does not auto-speak filler before the consult answer; it captures and transcribes speech, then speaks the routed OpenClaw answer. If multiple forced consult answers finish while Discord is still playing the first answer, later exact-speech answers are queued until playback idles instead of replacing speech mid-sentence.- Realtime voice buffers generated audio when Discord playback temporarily falls behind and tolerates brief provider or network gaps. Each provider response keeps its own buffered audio, including native tool continuations. Normal backpressure does not cancel the response, and queued answers wait until Discord finishes playing the previous answer, even if its provider response or audio encoder has already finished.
- GA OpenAI Realtime and xAI interruptions truncate each retained native audio item at the amount Discord consumed. Queued items are discarded at zero, and completed replies that have finished playing are left intact. Playback progress survives temporary gaps in the same response; GA OpenAI Realtime’s echo guard uses the combined consumed duration of retained items. GPT-Live handles interruption natively and does not use item truncation.
- If a speaker’s realtime connection fails, other speakers stay connected. Check the
realtime speaker failedlog and try speaking again to open a new connection. If the initial provider connection fails during/vc join, joining fails while an already-connected recorder keeps recording; check therealtime session failed terminallylog and retry/vc join. Temporary provider reconnects do not end the Discord voice session. - In
stt-ttsmode, STT usestools.media.audio;voice.modeldoes not affect transcription. stt-ttsreplies remain active until Discord finishes playing them; long responses are not cut off by a fixed one-minute playback deadline.- In realtime modes,
voice.realtime.provider,voice.realtime.model, andvoice.realtime.speakerVoiceconfigure the realtime audio session. For OpenAI Realtime 2.1 plus the Codex brain, usevoice.realtime.model: "gpt-realtime-2.1"andvoice.model: "openai/gpt-6-astra". - Realtime voice modes include small
IDENTITY.md,USER.md, andSOUL.mdprofile files in the realtime provider instructions by default so fast direct turns keep the same identity, user grounding, and persona as the routed OpenClaw agent. Setvoice.realtime.bootstrapContextFilesto a subset to customize this, or[]to disable it. Only those profile files are supported;AGENTS.mdstays in the normal agent context. The injected profile context does not replaceopenclaw_agent_consultfor workspace work, current facts, memory lookup, or tool-backed actions. - In OpenAI
agent-proxyrealtime mode, wake-name gating adapts to the room by default: one human can talk naturally without a wake name, while two or more humans must start or end a turn with one. Other bots do not count as people. Setvoice.realtime.requireWakeName: trueto always require a wake name orfalseto never require one. Configured wake names must be one or two words. Ifvoice.realtime.wakeNamesis unset, OpenClaw uses the routed agentnameplusOpenClaw, falling back to the agent id plusOpenClaw. An active wake-name gate disables realtime provider auto-response, routes accepted turns through the OpenClaw agent consult path, and gives a short spoken acknowledgement when an exact leading wake name is recognized from partial transcription before the final transcript arrives. Fuzzy name matching waits for the final transcript, so an unfinished ordinary word does not trigger an acknowledgement. The policy follows live joins and leaves without reconnecting voice. - The OpenAI realtime provider accepts current Realtime 2 event names and legacy Codex-compatible aliases for output audio and transcript events, so compatible provider snapshots can drift without dropping assistant audio.
- For response-based models,
voice.realtime.bargeIncontrols whether audible microphone input interrupts active realtime playback. Silent packets and speaker-start notifications alone do not interrupt it. If unset, it follows the realtime provider’s input-audio interruption setting. GPT-Live ignores this setting because it owns interruption. voice.realtime.minBargeInAudioEndMscontrols the minimum assistant playback duration before a GA OpenAI Realtime barge-in truncates audio. Default:250. Set0for immediate interruption in low-echo rooms, or raise it for echo-heavy speaker setups. It does not apply to GPT-Live.voice.ttsoverridesttsforstt-ttsvoice playback only; realtime modes usevoice.realtime.speakerVoiceinstead. For an OpenAI voice on Discord playback, setvoice.tts.provider: "openai"and choose a Text-to-speech voice undervoice.tts.providers.openai.speakerVoice.cedaris a good masculine-sounding choice on the current OpenAI TTS model.- Per-channel Discord
systemPromptoverrides apply to voice transcript turns for that voice channel. - When OpenClaw joins a voice channel, the routed agent session receives a silent system event with the current participant roster. Later participant joins and leaves update that session without triggering an unsolicited spoken reply; Discord display names are treated as untrusted labels. Authorized voice turns also receive a fresh roster snapshot.
- Voice transcript turns and
/vccommands use Discord entries incommands.ownerAllowFromfor owner status. When no Discord command owner is configured, the selected Discord account’sallowFrom(or legacydm.allowFrom) can still authorize voice access without granting owner status. Agent tool visibility follows the configured tool policy for the routed session. - Authorized speakers can ask the agent to list or change the current realtime call’s voice with
talk_voice, without command-owner configuration. The change applies to the shared room and keeps saved voice defaults unchanged. Voice control remains bound to the admitted turn and stops when the call ends, the turn finishes or is canceled, or its access policy changes; other owner-only tools retain their usual permissions. - Canceled or revoked voice turns are not retried as new agent requests. Leaving prevents pending batch turns from starting agent work or speech synthesis. Work already running can finish; transcript recording is managed separately.
- If
voice.autoJoinhas multiple entries for the same guild, OpenClaw joins the last configured channel for that guild. voice.autoJoin[].whenOccupieddefaults tofalse. Set it totruefor an auto-managed room that should contain the bot only while at least one human is present. OpenClaw joins on the first human arrival and leaves after the last human departs; the OpenClaw bot and other bots do not count. Startup, fresh gateway sessions, and resumed gateway sessions reconcile from Discord’s voice-state roster.- Occupancy management owns only sessions that it joined. A manual
/vc join, standalone transcript-only session, follow-user session, active session in another channel, or other ad-hoc join is not moved or disconnected when the configured room empties. Attaching transcript capture to an occupancy-managed session preserves that ownership. voice.allowedChannelsis an optional residency allowlist. Leave it unset to allow/vc joininto any authorized Discord voice channel. When set,/vc join, startup auto-join, and bot voice-state moves are restricted to the listed{ guildId, channelId }entries. Set it to an empty array to deny all Discord voice joins. If Discord moves the bot outside the allowlist, OpenClaw leaves that channel and rejoins the configured auto-join target when one is available.voice.daveEncryptionandvoice.decryptionFailureTolerancepass through to@discordjs/voicejoin options; the upstream defaults aredaveEncryption=trueanddecryptionFailureTolerance=24.- OpenClaw uses the bundled
libopus-wasmcodec for Discord voice receive and realtime raw PCM playback. It ships a pinned libopus WebAssembly build and does not require native opus addons. Discord voice sockets, codecs, playback conversion, and packet pacing run in a worker thread. GPT-Live continuous output travels directly from its media worker to the Discord playback worker rather than relaying every audio chunk through the Gateway event loop. Speaker admission, agent work, and transcripts remain on the Gateway; a busy Gateway can still delay those control operations. voice.connectTimeoutMscontrols the initial@discordjs/voiceReady wait for/vc joinand auto-join attempts. Default:30000.voice.reconnectGraceMscontrols how long OpenClaw waits for a disconnected voice session to begin reconnecting before destroying it. Default:15000.- In
stt-ttsmode, voice playback does not stop just because another user starts speaking. To avoid feedback loops, OpenClaw does not admit new conversational turns while TTS is playing; an explicitly started capture still records that speech. Speak after playback finishes for the next conversational turn. Response-based realtime models receive audible authorized microphone input as barge-in signals when interruption is enabled; GPT-Live handles incoming audio itself. - In GA OpenAI Realtime, echo from speakers into an open mic can look like barge-in and interrupt playback. For echo-heavy Discord rooms, set
voice.realtime.providers.openai.interruptResponseOnInputAudio: falseto keep the provider from auto-interrupting on input audio. Addvoice.realtime.bargeIn: trueif you still want audible Discord microphone input to interrupt active playback. The GA OpenAI realtime bridge ignores playback truncations shorter thanvoice.realtime.minBargeInAudioEndMsas likely echo/noise and logs them as skipped instead of clearing Discord playback. These controls do not override GPT-Live’s native interruption. voice.captureSilenceGraceMscontrols how long OpenClaw waits after Discord reports a speaker has stopped before finalizing that audio segment for STT. Default:2000; raise it if Discord splits normal pauses into choppy partial transcripts.- When ElevenLabs is the selected TTS provider, Discord voice playback uses streaming TTS and starts from the provider response stream. Providers without streaming support fall back to the synthesized temp-file path.
- OpenClaw watches receive decrypt failures and auto-recovers by leaving/rejoining the voice channel after repeated failures in a short window.
- If receive logs repeatedly show
DecryptionFailed(UnencryptedWhenPassthroughDisabled)after updating, collect a dependency report and logs. The bundled@discordjs/voiceline includes the upstream padding fix from discord.js PR #11449, which closed discord.js issue #11419. The operation was abortedreceive events are expected when OpenClaw finalizes a captured speaker segment; they are verbose diagnostics, not warnings.- Verbose Discord voice logs include a bounded one-line STT transcript preview for each accepted speaker segment, so debugging shows both the user side and the agent reply side without dumping unbounded transcript text.
- In
agent-proxymode, forced consult fallback skips likely incomplete transcript fragments such as text ending in...or a trailing connector like “and”, plus complete non-actionable closings like “I’ll be right back” or “bye”. Requests that mention a closing, such as “write a goodbye email”, still reach the agent. Closing detection uses a bounded English-language heuristic; unrecognized wording continues to the agent. Logs showforced agent consult skipped reason=...when this prevents a stale queued answer.
Voice messages
Discord voice messages show a waveform preview and require OGG/Opus audio. OpenClaw generates the waveform automatically, but needsffmpeg and ffprobe on the gateway host to inspect and convert.
- Provide a local file path (URLs are rejected).
- Omit text content (Discord rejects text + voice message in the same payload).
- Any audio format is accepted; OpenClaw converts to OGG/Opus as needed.
message tool, not from a shell: