Tool
Agents use thegoogle_meet tool:
test_speech always forces mode: "agent" or "bidi" and fails if asked to run in mode: "transcribe", because observe-only sessions cannot emit speech. speechOutputVerified requires both fresh realtime output bytes and a matching non-silent waveform on the native virtual microphone capture path during that output. With generated input commands, participant input is captured separately from browser playback; explicit input commands retain their configured capture path. A reused session’s older output or loopback signal does not count, and sink-byte growth alone does not report verified speech. This verifies local microphone injection; a second participant is needed to confirm remote audibility.
For Chrome transports, leave keeps a reused user-owned tab open after clicking Meet’s Leave call button. Tabs opened by OpenClaw are closed after departure.
Use transport: "chrome" when Chrome runs on the Gateway host, transport: "chrome-node" when it runs on a paired node. In both cases the model providers and openclaw_agent_consult run on the Gateway host, so model credentials stay there. Agent-mode logs include the resolved transcription provider/model at bridge startup and the TTS provider/model/voice/output format/sample rate after each synthesized reply. Raw mode: "realtime" is still accepted as a legacy compatibility alias for mode: "agent", but it is no longer advertised in the tool’s mode enum.
create with an API-backed room and explicit access policy:
status includes Chrome health when available:
Agent and bidi modes
agent mode: the realtime transcription provider hears meeting audio, final participant transcripts route through the configured OpenClaw agent, and the answer is spoken through regular OpenClaw TTS. Nearby final-transcript fragments are coalesced before the consult so one spoken turn does not produce several stale partial answers. Isolated browser input remains available during TTS; only legacy transports with mixed loopback input retain playback and transcript echo suppression.
bidi mode: the realtime voice model answers directly and delegates deeper reasoning, current information, or normal OpenClaw tools to the configured agent. Providers with function tools use openclaw_agent_consult; GPT-Live uses native delegation through the same Gateway voice runtime as Discord and Talk. Both paths preserve meeting context and realtime.toolPolicy. The resulting answer is spoken by the realtime voice model.
For GPT-Live with Cove, use the Live configuration. Live owns response timing and interruption: incoming participant audio stays open during playback, and OpenClaw does not add a local barge-in detector or wait for a response-completion event to play short replies. Stopping the meeting or replacing the provider session cancels active delegation and ignores late results.
By default consults run against the main agent; set realtime.agentId to point a Meet lane at a dedicated agent workspace, model defaults, tool policy, memory, and session history. Agent-mode consults use a per-meeting agent:<id>:subagent:google-meet:<session> session key so follow-up questions keep meeting context while inheriting normal agent policy. When an agent calls google_meet in agent mode, the consultant session forks the caller’s current transcript before answering participant speech; the Meet session stays separate so meeting follow-ups do not mutate the caller transcript directly.
realtime.toolPolicy controls the consult run:
The consult session key is scoped per Meet session, so follow-up consult calls reuse prior consult context during the same meeting.
Force a spoken readiness check after Chrome has fully joined: