Skip to main content
OpenClaw has separate plugins for Google Meet, Microsoft Teams meetings, and Zoom. All three can join through Chrome, use the same participation modes, and run Chrome either on the Gateway host or on a paired node. Their platform URLs, installation model, and extra capabilities differ. These plugins participate in meetings. They are separate from messaging channels such as the Microsoft Teams channel and from the Voice call plugin.

Choose a plugin

Choose Google Meet when you need meeting creation, Google API artifacts, or a Twilio phone path. Choose Teams or Zoom for direct browser guest participation on those platforms. The Teams and Zoom plugins do not create meetings, dial in, call the vendor API, or capture audio/video recordings.

Choose a mode

The three plugins share the same modes: Use transcribe when the agent only needs meeting text. Use agent for normal OpenClaw reasoning and tools. Use bidi when low-latency direct voice is more important than routing each turn through the regular agent. Google Meet, Teams, and Zoom use one shared meeting audio engine. In bidi mode they support GPT-Live through the same provider selection, native agent delegation, and interruption policy used by Discord and Talk. GPT-Live owns interruptions; participant audio stays open while the model speaks. In bidi mode, recoverable provider diagnostics are logged without stopping the audio bridge. Providers that support reconnecting own that recovery. Terminal provider closure, exhausted recovery, failed initial setup, and local audio-transport failures still stop the bridge; leaving the meeting also stops it. Stopping a meeting requests cancellation of any active agent consult. In agent mode, OpenClaw finishes active output and turn events before closing the session, then ignores late speech synthesis and audio delivery results. The bounded live transcript remains available only in transcribe mode. In all three modes, browser joins also persist completed caption rows and meeting notes to the shared state database. Notes update about every five minutes when new speech is saved, using the meeting agent’s utility model with primary-model and heuristic fallbacks. Leaving the meeting finalizes visible captions and writes the final summary; use openclaw transcripts to list, inspect, or export it. This durable notes path does not change the live agent-consult transcript or create an audio/video recording. Automatic notes are on by default. Set transcripts.enabled: false to disable durable notes globally. An explicitly selected transcribe session retains its bounded live-caption tail without writing durable rows. Caption availability still depends on the meeting platform, account, language, and host policy.

Configure Teams or Zoom

The Teams and Zoom plugins share the same configuration shape for their common meeting runtime. Add an entry only when you need to override a default. This example selects the normal agent path, changes the guest display name, and runs Chrome on a paired node:
Use "zoom-meetings" as the entry id for Zoom. Omit chromeNode to run Chrome on the Gateway host. For GPT-Live with Cove, set defaultMode: "bidi", realtime.voiceProvider: "openai", realtime.model: "gpt-live-1-codex", and realtime.providers.openai.voice: "cove". Sign in with openclaw models auth login --provider openai on the Gateway host. The Google Meet configuration example uses the same fields; substitute teams-meetings or zoom-meetings for the plugin entry. Unpinned configurations keep their provider’s default model. agent mode continues to use regular OpenClaw TTS.

Prepare Chrome and audio

Chrome can run on the Gateway host or on a paired node. A remote Chrome node must allow browser.proxy plus the platform command: For agent or bidi mode through Chrome, install the native audio dependencies on the host that runs Chrome. On macOS:
On a Linux desktop with PipeWire-Pulse, install the PulseAudio command-line tools. OpenClaw creates and reuses an OpenClaw Meeting Audio null sink and matching source in the desktop user’s audio session:
Run the Gateway or paired node as the same desktop user that runs Chrome. A root service or headless service without that user’s XDG_RUNTIME_DIR cannot access the PipeWire-Pulse socket and fails setup with an actionable error. The Gateway host still owns the OpenClaw agent and model credentials when Chrome runs on a paired node. Configure a realtime transcription provider and OpenClaw TTS for agent mode, or a realtime voice provider for bidi mode. The platform guides contain the provider and audio-command options. Chrome sessions without an explicit chrome.audioInputCommand capture participant audio directly from browser playback and keep that playback off the virtual microphone. The native output command injects assistant speech into the microphone; the native input command verifies that injection. Isolated participant input remains open during both realtime speech and TTS. A capture failure stops the bridge with an error; mixed loopback audio is not used as a fallback for Live. Explicit input commands retain their existing provider-input behavior on both local and paired-node Chrome, including custom capture, filters, and mixers. They keep the existing echo protection. To select GPT-Live, remove the input override and use managed isolated capture; Live rejects a custom command path whose isolation cannot be verified instead of silently replacing it.

Install or disable plugins

Install the meeting plugins you need. Each is enabled by default after installation:
Disable any meeting plugin you do not use:
These changes apply to a running Gateway automatically. If it is offline, start it before joining. Check the application result, then run the platform setup check below; see Apply changes and inspect.

Verify and join

Treat any failed setup check as a blocker for that transport and mode. For an observe-only smoke test, select transcribe mode and confirm that status reports an in-call session before expecting caption text. For talk-back smoke tests, verified speech requires more than bytes accepted by the playback command. The shared command-pair bridge correlates a bounded waveform fingerprint from the current output generation with audio returning on the selected virtual microphone capture path; Google Meet, Teams, and Zoom do not report speechOutputVerified: true when only the output-byte counter advances or unrelated participant audio is present. That verifies local microphone injection. Use a controlled second participant to prove remote audibility and interruption during an actual meeting.

Handle platform policy prompts

Browser automation handles the normal guest-name, prejoin camera and microphone, join, in-call, and leave controls. It does not bypass platform or organizer policy.
  • Google Meet may require Google sign-in, host admission, or a browser permission decision.
  • Microsoft Teams may require tenant sign-in, email verification, or organizer admission.
  • Zoom may require authentication, email verification, a passcode, CAPTCHA completion, or host admission; an account can also disable browser join.
When a join or status result includes manualAction, complete its reported step in the same OpenClaw Chrome profile before retrying. Repeatedly opening new tabs does not resolve an account, tenant, lobby, or CAPTCHA gate. Only join meetings where the operator is authorized to add an agent. Tell participants when local policy or consent rules require disclosure of automated participation, transcription, or synthesized speech.

Discord voice chat

Discord voice channels provide native, audio-only realtime conversation without browser meeting automation. OpenClaw can join a voice channel, listen, route turns through an OpenClaw agent or realtime voice model, and speak replies. It does not send or receive camera video or screen sharing, even when people use video in the same Discord channel, so Discord voice is a related live-conversation surface rather than a fourth browser meeting plugin.

Platform guides