Skip to main content
Deepgram is a speech-to-text API. OpenClaw uses it for inbound audio/voice-note transcription through tools.media.audio and for Voice Call streaming STT through plugins.entries.voice-call.config.streaming. Batch transcription uploads the complete audio file to Deepgram. Flux models use Deepgram’s one-shot WebSocket instead. Both paths inject the transcript into the reply pipeline ({{Transcript}} + [Audio] block). Voice Call streaming forwards live G.711 u-law frames over Deepgram’s WebSocket listen endpoint and emits partial/final transcripts as Deepgram returns them.

Getting started

1

Set your API key

2

Enable the audio provider

3

Send a voice note

Send an audio message through any connected channel. OpenClaw transcribes it via Deepgram and injects the transcript into the reply pipeline.

Configuration options

providerOptions.deepgram merges extra query params directly into the Deepgram /listen request, so any Deepgram-supported param name works (for example detect_language, punctuate, smart_format):

Flux models

Use flux-general-en or flux-general-multi for Deepgram Flux. OpenClaw converts the voice note to 16 kHz mono linear16 audio with ffmpeg, then sends it to Deepgram’s /v2/listen WebSocket endpoint. Set tools.media.models[].model to either Flux model in the getting-started configuration above. Voice-note conversion processes the complete audio file, including notes longer than 20 minutes. Decoded audio uses a private temporary file, which is removed after the transcription attempt. The configured input-size and request-timeout limits still apply; failed conversion or upload does not return a partial transcript. Flux supports eager_eot_threshold, eot_threshold, eot_timeout_ms, keyterm, language_hint, mip_opt_out, numerals, profanity_filter, redact, and tag in providerOptions.deepgram. OpenClaw ignores batch-only options such as detect_language, punctuate, and smart_format on Flux. Language hints apply only to flux-general-multi: OpenClaw maps the model entry’s language setting to language_hint, with an explicit provider option taking precedence. Both settings are ignored for the English-only flux-general-en model.

Voice Call streaming STT

The bundled deepgram plugin also registers a realtime transcription provider for the Voice Call plugin.
For a Deepgram custom endpoint, set baseUrl to the endpoint root, including any base path but not /listen. Realtime endpoints accept http://, https://, ws://, and wss://. HTTP maps to WS, HTTPS maps to WSS, and explicit WebSocket schemes stay unchanged. Malformed URLs and other schemes fail during session setup.
Voice Call receives telephony audio as 8 kHz G.711 u-law. The Deepgram streaming provider defaults to encoding: "mulaw" and sampleRate: 8000, so Twilio media frames can be forwarded directly.

Notes

Authentication follows the standard provider auth order. DEEPGRAM_API_KEY is the simplest path.
Override endpoints or headers on the Deepgram tools.media.models[] entry when using a proxy.
Output follows the same audio rules as other providers (size caps, timeouts, transcript injection).
Install ffmpeg with the gateway host’s package manager before selecting a Flux model.

Media tools

Audio, image, and video processing pipeline overview.

Configuration

Full config reference including media tool settings.

Troubleshooting

Common issues and debugging steps.

FAQ

Frequently asked questions about OpenClaw setup.