Config
The common Chrome agent path only needs the plugin enabled, BlackHole, SoX, a realtime provider key, and a configured OpenClaw TTS provider:Defaults
chrome.audioBridgeCommand and chrome.audioBridgeHealthCommand let an external bridge own the whole local audio path instead of chrome.audioInputCommand/chrome.audioOutputCommand; see Notes for the constraint on which mode can use them.
An openclaw doctor --fix migration exists for the legacy realtime.provider: "google" shape: it moves that intent to realtime.voiceProvider: "google" plus realtime.transcriptionProvider: "openai" when those fields are not already set.
GPT-Live with Cove
Selectbidi mode for GPT-Live speech. agent mode continues to use realtime
transcription and regular OpenClaw TTS; changing the realtime voice model does
not change that mode’s voice.
Sign in on the Gateway host with openclaw models auth login --provider openai,
then configure the existing model and provider fields:
chrome.audioFormat at its 24 kHz PCM16 default. Existing unpinned meeting
configurations keep the provider’s default model; choose Live explicitly.
Remove any explicit chrome.audioInputCommand to use Live’s managed isolated
browser input. Existing custom input commands retain their provider-input
behavior for other models; Live rejects that path because its isolation cannot
be verified. Selecting Live does not silently discard a configured input source.
For the public Platform API route, use model: "gpt-live-1" and a supported
voice such as voice: "marin", with an OpenAI API key on the Gateway host.
cove belongs to the Codex route. See
OpenAI voice and speech for authentication
and model/voice compatibility.
Optional overrides
tts.providers.elevenlabs.speakerVoiceId. Agent replies can also use per-reply [[tts:speakerVoiceId=... model=eleven_v3]] directives when TTS model overrides are enabled, but config is the deterministic default for meetings. On join, logs show transcriptionProvider=elevenlabs, and each spoken reply logs provider=elevenlabs model=eleven_v3 speakerVoiceId=<voiceId>.
Twilio-only config:
voiceCall.enabled: true (the default) and Twilio transport, Voice Call places the DTMF sequence before opening the realtime media stream, then uses the saved intro text as the initial realtime greeting. If voice-call is not enabled, Google Meet can still validate and record the dial plan but cannot place the Twilio call.
Leave voiceCall.gatewayUrl unset to use the local trusted Gateway runtime, which preserves the
invoking agent for the full call. A configured Gateway URL remains an explicit WebSocket target and
cannot authenticate plugin provenance; non-default agent joins fail closed instead of silently
using another agent. Run Google Meet and Voice Call in the same Gateway process when per-agent
routing is required.