Skip to main content
llmman pulls models as OCI artifacts from any registry or Hugging Face and serves them through unmodified upstream engines: llama-server for GGUF, vllm or mlx-lm for safetensors. One daemon exposes Ollama-, OpenAI-, and Anthropic-compatible APIs and can also forward requests to hosted providers. OpenClaw talks to it through the generic openai-completions adapter on /v1. Three modes are supported:
llmman serve has no authentication and no TLS. Keep the default loopback bind unless a trusted network boundary restricts access, and never expose it on a public interface.
Version scope: this page is verified against llmman v0.1.334, commit 22b6d73. Model names, hybrid routing, vision, and tool calling on this page were exercised against that build with qwen3.8.

Auth rules

llmman serve never checks credentials. OpenClaw still needs a non-empty apiKey on the provider entry so the provider counts as configured. This page uses LLMMAN_API_KEY=llmman-local with apiKey: "${LLMMAN_API_KEY}", mirroring the OLLAMA_API_KEY=ollama-local convention. A literal apiKey: "llmman-local" works too.
For llmman.hybrid/... and llmman.provider/... refs, llmman forwards the bearer OpenClaw presents to the hosted provider as that provider’s API key. The one exception is the literal placeholder llmman, which tells the daemon to use its own key. See Hybrid inference for both patterns. A marker such as llmman-local would be sent to the hosted provider and rejected there.
A daemon started with LLMMAN_HOST=0.0.0.0 in its environment binds every interface with no auth. Point OpenClaw at a LAN host only inside a network you trust; there is no credential OpenClaw can send that llmman would enforce. A daemon reachable off loopback also refuses to spend its own hosted-provider key for callers that presented none.
${LLMMAN_API_KEY} resolves from the Gateway process environment or ~/.openclaw/.env. If the variable is missing, OpenClaw logs a config warning and treats the provider as unavailable, so put the export where the Gateway can read it. See Environment.

Getting started

1

Install llmman

Windows: irm https://llmmanorg.github.io/install.ps1 | iex or winget install llmmanorg.llmman. Other options: cargo binstall llmman.
2

Pull a model and start the server

Bare names such as qwen3.8 resolve to Docker Hub’s curated docker.io/ai/<name>:latest. owner/repo names resolve to hf.co/owner/repo; a full reference (ghcr.io/..., hf.co/...) is used as-is.llmman serve requires no arguments (an optional model argument preloads that model): it listens on 127.0.0.1:17434 and loads any pulled model on the first request that names it, then unloads it after five idle minutes. GPU acceleration (CUDA, ROCm, Vulkan, Metal) is auto-detected and the matching llama-server is downloaded if none is on PATH. Everything else is tuned through the daemon’s environment: LLMMAN_HOST for the bind address, LLMMAN_CONTEXT_LENGTH for the server context, LLMMAN_KEEP_ALIVE for the idle unload timer (see Advanced configuration).By default llmman uses up to 262,144 tokens, capped to the model’s trained context (262,144 for qwen3.8) and, on out-of-memory, retries with the context halved down to a 16,384 floor. llmman ps shows the context a loaded model actually got; keep the OpenClaw model’s contextWindow at or below that value.
3

Verify the server

There is no /health route; use /api/version or /v1/models as the readiness probe. /v1/models lists fully qualified ids such as docker.io/ai/qwen3.8:latest; requests may use either that form or the short name.
4

Set the marker credential

Add to ~/.openclaw/.env (or export in the Gateway’s shell):
5

Add the provider and select the model

Add the config below, then:
llmman launch openclaw --model qwen3.8 can bootstrap a first-run OpenClaw install by running non-interactive onboarding against the llmman endpoint. It only applies when no openclaw.json exists yet; for an existing install use the explicit config on this page.

Full config example

Qwen3.8 on a local llmman server:
models.mode: "merge" keeps hosted providers available as fallbacks. timeoutSeconds gives cold model loads and long generations room before the model request timeout fires.

Model discovery

llmman is not a bundled OpenClaw plugin, so there is no implicit discovery. List every model you want under models.providers.llmman.models with a provider-local id (no llmman/ prefix).
To add a model, pull it and add a matching entry:

Smoke tests

A narrow text probe that skips the full agent tool surface:
Add --file with an image for a lean vision-model probe (PNG/JPEG/WebP; non-image files are rejected before llmman is called; use openclaw infer audio transcribe for audio):
Neither path loads chat tools, memory, or session context. If a probe succeeds while normal agent replies fail, the issue is usually tool-schema handling or context pressure in the backend, not the endpoint; see Troubleshooting. A full agent turn with tool calling is the real test:

Hybrid inference

llmman can pair a local model with a hosted one under a single model name and choose a side per request. OpenClaw configures the pair once as an ordinary model id and gets local-first inference with hosted overflow, without an agent-level fallback switch. The reference is llmman.hybrid/<local>,<provider>/<model>. With qwen3.8 as the local half and OpenAI’s gpt-5.6-luna as the hosted half:
Which side serves a request:
  1. x-llmman-route: local or cloud request header wins. Any other value is a 400.
  2. Otherwise, size. A request body larger than the local context can hold goes to the hosted model. The budget is 4 bytes per token of LLMMAN_CONTEXT_LENGTH; LLMMAN_HYBRID_LOCAL_BYTES sets it directly and 0 disables the size rule.
  3. Otherwise, local.
If a request llmman kept local is then refused by the local backend as larger than its context, llmman resends it to the hosted half before anything reaches OpenClaw. A local pin is never overridden this way. Every routed request is logged with the side and the reason. The raw completion response may report the backend model or GGUF path; OpenClaw’s result envelope retains the configured pair ref.

Hosted-provider key

The hosted half authenticates like any llmman --provider request. Pick one:

Hybrid config

The daemon computes one hybrid byte budget at startup: 4 bytes per token of LLMMAN_CONTEXT_LENGTH, or 262144 × 4 = 1048576 bytes when unset. This budget does not follow a model’s trained-context cap or a later out-of-memory reduction. Set LLMMAN_CONTEXT_LENGTH=65536 in the daemon’s environment to match the example, or set LLMMAN_HYBRID_LOCAL_BYTES to pick the byte budget directly. An explicit LLMMAN_CONTEXT_LENGTH=0 disables size-based routing unless a positive LLMMAN_HYBRID_LOCAL_BYTES supplies a budget; setting the byte override to 0 also disables that rule. Local context-refusal fallback still applies unless the request is pinned local. Set contextWindow on the pair to the local model’s usable token context. OpenClaw then compacts around the local model’s limit, so most turns stay local; llmman still overflows to gpt-5.6-luna when a request exceeds it. Set it to the hosted model’s window instead if you prefer fewer compactions and more hosted traffic.

Pinning a side

OpenClaw sends provider-level headers on every request, so a second provider entry on the same base URL can force one side of the pair:
Switch with /model llmman-cloud/... for a turn that should go hosted, or use headers: { "x-llmman-route": "local" } for an entry that must never leave the machine.

Hybrid versus OpenClaw fallbacks

They compose. A common shape is a hybrid pair as primary with a direct hosted model as a fallback for when llmman itself is down:

Hosted models through llmman

llmman.provider/<provider>/<model> forwards to a hosted provider with no local half. Use it when you want every model, local or hosted, behind one endpoint and one key-handling story:
The provider catalog comes from models.dev and is cached by the daemon. llmman providers shows which providers have a key; llmman list --provider openai lists that provider’s models and prices. For direct hosted access without llmman in the path, configure the OpenAI provider instead.

Vision and image description

Models that ship a companion mmproj projector are vision-capable; qwen3.8 and gemma4:e4b are. llmman show <model> logs found companion mmproj file and /api/show reports a vision capability. Mark those models input: ["text", "image"] so image attachments are injected into agent turns.
--model must be a full <provider/model> ref. Use infer image describe for OpenClaw’s image-understanding flow and configured imageModel; use infer model run --file for a raw multimodal probe with a custom prompt. To make a llmman model the default image-understanding provider for inbound media:
OpenClaw rejects image-description requests for models not marked image-capable. Slow local vision models can need a longer image-understanding timeout than hosted models; models.providers.llmman.timeoutSeconds still governs the underlying HTTP request for normal model calls.

Configuration

Common recipes

Replace model ids with names from llmman list or openclaw models list --provider llmman.
Use the Full config example for the provider entry, and set LLMMAN_CONTEXT_LENGTH=65536 in the daemon’s environment if you want the server context to match it exactly.
The Hybrid config above: llmman.hybrid/qwen3.8,openai/gpt-5.6-luna as primary, LLMMAN_API_KEY set to the OpenAI key, LLMMAN_CONTEXT_LENGTH matching the pair’s contextWindow. Add openai/gpt-5.6-luna to fallbacks so a stopped daemon does not block replies.
Local models served through a custom openai-completions provider use structured Tool Search automatically when unset. The explicit setting below pins that surface, keeping optional capabilities available while loading their schemas only when needed. Cap the context to what the host can run with LLMMAN_CONTEXT_LENGTH=32768 in the daemon’s environment:
Use compat.supportsTools: false only when the model or server reliably fails on tool schemas; it disables tool use entirely. For a deliberately narrower agent, prefer tools.profile or a per-agent tool policy.
Custom provider ids when running more than one daemon; each gets its own host, models, and timeout:
llmman can also pool several daemons itself: llmman config set aggregation.peers <host>,<host> (or LLMMAN_PEERS) makes one daemon forward to peers, so OpenClaw sees a single endpoint whose model list spans the group.
Model ids are whatever llmman resolves. Pull with the full reference and use the same string as the OpenClaw model id:
The agent ref is then llmman/hf.co/unsloth/Qwen3.5-0.8B-GGUF. Run llmman login <registry> first for private registries.

Model selection

For slow local models, prefer provider-scoped tuning before raising the whole agent runtime timeout: models.providers.llmman.timeoutSeconds covers connection setup, headers, body streaming, and the total guarded-fetch abort for that provider’s model requests only.

Quick verification

For remote hosts, replace 127.0.0.1 with the baseUrl host. If curl works but OpenClaw does not, check whether the Gateway runs on a different machine, container, or service account.

Advanced configuration

LLMMAN_CONTEXT_LENGTH is the server-side context (there is no flag). Semantics by backend:
  • llama-server (GGUF): set, it is passed as --ctx-size for generation models; 0 means the trained context. Unset, llmman uses 262144 or the model’s trained context if smaller, and on out-of-memory retries with the context halved down to a 16,384 floor.
  • vLLM (safetensors): a positive value becomes --max-model-len; unset uses the vLLM default.
  • mlx_lm.server: not forwarded.
LLMMAN_NUM_PARALLEL scales --ctx-size up by that factor so each slot keeps the full context.On the OpenClaw side, contextWindow declares the model’s window and contextTokens caps active input. Keep contextWindow at or below the server value; OpenClaw derives compaction and preflight thresholds from it. OpenClaw’s contextWindow does not change llmman’s hybrid byte budget; configure the daemon separately as described in Hybrid config.
Qwen3.8 thinks by default; llmman returns the reasoning as reasoning_content, which OpenClaw’s openai-completions adapter separates from the final text. Requests are proxied to llama-server, so chat_template_kwargs passes through. To turn thinking off for agent turns with a local Qwen model:
For per-run or session control, declare the local Qwen model’s thinking format:
With this declaration, openclaw agent --model llmman/qwen3.8 --thinking off, /think off, and openclaw infer model run --local --model llmman/qwen3.8 --thinking off --prompt "Reply with exactly: pong" --json map the thinking setting to chat_template_kwargs.enable_thinking. Without it, the generic proxy defaults do not send this control or reasoning_effort.The lean infer model run path does not read the agent-level params recipe above; use the compatibility declaration and --thinking off for that probe. Do not combine a fixed enable_thinking agent param with per-run control, since the fixed param overrides the generated value. Apply Qwen-specific controls to a hybrid ref only if both its local and hosted backends accept them.
Models load on their first request, each in its own backend subprocess, and unload after five idle minutes. Tune with:llmman ps shows loaded models with their context and expiry; llmman stop <model> unloads one now. A first request after startup or an idle unload pays the load cost, so set timeoutSeconds on the provider and raise LLMMAN_KEEP_ALIVE to keep the daemon warm for chat surfaces.
llmman probes CUDA, ROCm, Vulkan (Linux/Windows) or Metal (macOS) and downloads a matching llama-server release if none is on PATH. Override with LLMMAN_LLM_LIBRARY: cpu, cuda, cuda13, rocm, vulkan, or metal. Other knobs: LLMMAN_FLASH_ATTENTION (on/off/auto), LLMMAN_KV_CACHE_TYPE (f16, q8_0, q4_0), LLMMAN_SCHED_SPREAD for multi-GPU layer splitting, LLMMAN_IGPU_ENABLE to count integrated GPUs. On Linux, llmman serve --ociman docker|podman runs llama-server from the ghcr.io/ggml-org/llama.cpp images instead of a local binary. LLMMAN_DEBUG=1 prints the probe result.
llmman serves /v1/embeddings for GGUF embedding models, so memory search can use it through the generic openai-compatible embedding provider:
Embedding models are capped to their trained context regardless of LLMMAN_CONTEXT_LENGTH. See Memory config for the remaining fields.
llmman also implements Ollama’s native /api/chat, /api/tags, /api/show, and /api/ps, and OLLAMA_HOST=127.0.0.1:17434 ollama run <model> works against it. Prefer api: "openai-completions" from OpenClaw anyway: llmman’s /api/show reports only completion and vision capabilities, so the bundled Ollama plugin’s discovery would mark every llmman model compat.supportsTools: false. If you do point the Ollama plugin at http://127.0.0.1:17434 (no /v1), list models explicitly instead of relying on discovery.
llmman forwards message content and tool schemas to the backend without normalizing them, so compatibility depends on the selected engine and model. Structured content parts (text + image) and OpenClaw’s full tool schema work with qwen3.8 on llama-server. If a different backend or model rejects them:
  • messages[].content: invalid type: sequence, expected a string → set compat.requiresStringContent: true on the model entry. OpenClaw then flattens pure text content parts into plain strings.
  • 400 JSON schema conversion failed → llama-server could not compile a tool schema into its grammar subset. Update OpenClaw first; if a third-party tool or MCP server contributes the offending schema, disable it for that agent, and use compat.supportsTools: false only as a last resort.
Because llmman is a non-native openai-completions endpoint, OpenClaw treats it as a proxy route: no service_tier, no Responses store, no prompt-cache hints, no OpenAI reasoning-compat payload shaping, no hidden OpenClaw attribution headers, and compat.supportsDeveloperRole is forced to false. Vendor-specific fields can be merged into the request body with agents.defaults.models["llmman/<model>"].params.extra_body.
Local llmman models are free, so set all costs to 0. The hosted half of a hybrid or llmman.provider/... ref is billed by that provider; llmman list --provider openai shows its per-million-token prices.

Troubleshooting

llmman serve is not running or is not reachable at the configured address. The default is 127.0.0.1:17434; if you set LLMMAN_HOST, update the OpenClaw baseUrl and healthUrl to match.
models.providers.llmman.apiKey: *** env var "LLMMAN_API_KEY" in the config warnings means the substitution found no value. Add LLMMAN_API_KEY=llmman-local to ~/.openclaw/.env, or replace "${LLMMAN_API_KEY}" with the literal "llmman-local".
The model is not pulled, or the id in config does not match what llmman resolves. Compare against llmman list; a short name maps to docker.io/ai/<name>:latest, and an owner/repo name maps to hf.co/owner/repo.
Large models can take minutes to load, especially on the first request after an idle unload. Raise models.providers.llmman.timeoutSeconds, warm the model with a first openclaw infer model run, and consider a longer LLMMAN_KEEP_ALIVE on the daemon.
llmman routed the request to the hosted half and found no usable key. Either the bearer OpenClaw sent was a marker (llmman-local) rather than a real key, or you used the llmman placeholder without giving the daemon its own OPENAI_API_KEY, or the daemon is bound off loopback and refuses to spend its own key. See Hosted-provider key.
Check the daemon log; every routed request records the side and the reason. Routing is by request body size against LLMMAN_CONTEXT_LENGTH (4 bytes per token) unless x-llmman-route is set. Lower LLMMAN_HYBRID_LOCAL_BYTES to overflow sooner, raise it to stay local longer, or pin a side with a provider headers entry as in Pinning a side.
Both probes are tool-free, so compat.supportsTools cannot change this failure. Check the configured base URL, model id, and LLMMAN_API_KEY, inspect the daemon and backend logs, and compare the two request payloads.
The agent turn adds a larger prompt and tool schemas. A 400 JSON schema conversion failed is llama-server rejecting a tool schema; update OpenClaw and check third-party tools or MCP servers. Otherwise enable Tool Search to defer schemas, confirm the server’s actual context allocation, and use compat.supportsTools: false only as a last resort. See Smaller or stricter backends.
If schema errors are gone but the spawned llama-server still crashes on larger turns, treat it as an upstream llama.cpp or model limitation. Lower LLMMAN_CONTEXT_LENGTH, set LLMMAN_KV_CACHE_TYPE=q8_0 to reduce memory, or switch the backend or model.
Confirm the model’s chat template supports tool calling and that the request reached /v1/chat/completions (not the Ollama or Anthropic surfaces via a proxy). If the model only calls tools when forced, set params.extra_body.tool_choice: "required" on that model ref as described in Local models.
More help: Troubleshooting and FAQ.

Local models

Running OpenClaw against local model servers.

Local model services

Starting local model servers on demand for configured providers.

OpenAI

Direct hosted access to the models used as the hybrid overflow half.

Inference CLI

openclaw infer model run and the other one-shot probes used on this page.

Model providers

Overview of all providers, model refs, and failover behavior.

Gateway troubleshooting

Debugging local OpenAI-compatible backends that pass probes but fail agent runs.