llama-server for GGUF, vllm or mlx-lm for safetensors. One
daemon exposes Ollama-, OpenAI-, and Anthropic-compatible APIs and can also
forward requests to hosted providers. OpenClaw talks to it through the generic
openai-completions adapter on /v1.
Three modes are supported:
22b6d73. Model names, hybrid routing, vision, and tool calling on this page were exercised against that build with qwen3.8.Auth rules
Local llmman has no auth
Local llmman has no auth
llmman serve never checks credentials. OpenClaw still needs a non-empty apiKey on the provider entry so the provider counts as configured. This page uses LLMMAN_API_KEY=llmman-local with apiKey: "${LLMMAN_API_KEY}", mirroring the OLLAMA_API_KEY=ollama-local convention. A literal apiKey: "llmman-local" works too.The bearer matters for hybrid and hosted refs
The bearer matters for hybrid and hosted refs
llmman.hybrid/... and llmman.provider/... refs, llmman forwards the bearer OpenClaw presents to the hosted provider as that provider’s API key. The one exception is the literal placeholder llmman, which tells the daemon to use its own key. See Hybrid inference for both patterns. A marker such as llmman-local would be sent to the hosted provider and rejected there.Remote llmman hosts
Remote llmman hosts
LLMMAN_HOST=0.0.0.0 in its environment binds every interface with no auth. Point OpenClaw at a LAN host only inside a network you trust; there is no credential OpenClaw can send that llmman would enforce. A daemon reachable off loopback also refuses to spend its own hosted-provider key for callers that presented none.Env var substitution
Env var substitution
${LLMMAN_API_KEY} resolves from the Gateway process environment or ~/.openclaw/.env. If the variable is missing, OpenClaw logs a config warning and treats the provider as unavailable, so put the export where the Gateway can read it. See Environment.Getting started
Install llmman
irm https://llmmanorg.github.io/install.ps1 | iex or winget install llmmanorg.llmman. Other options: cargo binstall llmman.Pull a model and start the server
qwen3.8 resolve to Docker Hub’s curated docker.io/ai/<name>:latest. owner/repo names resolve to hf.co/owner/repo; a full reference (ghcr.io/..., hf.co/...) is used as-is.llmman serve requires no arguments (an optional model argument preloads that model): it listens on 127.0.0.1:17434 and loads any pulled model on the first request that names it, then unloads it after five idle minutes. GPU acceleration (CUDA, ROCm, Vulkan, Metal) is auto-detected and the matching llama-server is downloaded if none is on PATH. Everything else is tuned through the daemon’s environment: LLMMAN_HOST for the bind address, LLMMAN_CONTEXT_LENGTH for the server context, LLMMAN_KEEP_ALIVE for the idle unload timer (see Advanced configuration).By default llmman uses up to 262,144 tokens, capped to the model’s trained context (262,144 for qwen3.8) and, on out-of-memory, retries with the context halved down to a 16,384 floor. llmman ps shows the context a loaded model actually got; keep the OpenClaw model’s contextWindow at or below that value.Verify the server
/health route; use /api/version or /v1/models as the readiness probe. /v1/models lists fully qualified ids such as docker.io/ai/qwen3.8:latest; requests may use either that form or the short name.Set the marker credential
~/.openclaw/.env (or export in the Gateway’s shell):Add the provider and select the model
llmman launch openclaw --model qwen3.8 can bootstrap a first-run OpenClaw install by running non-interactive onboarding against the llmman endpoint. It only applies when no openclaw.json exists yet; for an existing install use the explicit config on this page.Full config example
Qwen3.8 on a local llmman server:models.mode: "merge" keeps hosted providers available as fallbacks. timeoutSeconds gives cold model loads and long generations room before the model request timeout fires.
Model discovery
llmman is not a bundled OpenClaw plugin, so there is no implicit discovery. List every model you want undermodels.providers.llmman.models with a provider-local id (no llmman/ prefix).
Smoke tests
A narrow text probe that skips the full agent tool surface:--file with an image for a lean vision-model probe (PNG/JPEG/WebP;
non-image files are rejected before llmman is called; use
openclaw infer audio transcribe for audio):
Hybrid inference
llmman can pair a local model with a hosted one under a single model name and choose a side per request. OpenClaw configures the pair once as an ordinary model id and gets local-first inference with hosted overflow, without an agent-level fallback switch. The reference isllmman.hybrid/<local>,<provider>/<model>. With qwen3.8
as the local half and OpenAI’s gpt-5.6-luna as the hosted half:
x-llmman-route: localorcloudrequest header wins. Any other value is a400.- Otherwise, size. A request body larger than the local context can hold goes to the hosted model. The budget is 4 bytes per token of
LLMMAN_CONTEXT_LENGTH;LLMMAN_HYBRID_LOCAL_BYTESsets it directly and0disables the size rule. - Otherwise, local.
local pin is never overridden this way. Every routed request is
logged with the side and the reason. The raw completion response may report
the backend model or GGUF path; OpenClaw’s result envelope retains the
configured pair ref.
Hosted-provider key
The hosted half authenticates like any llmman--provider request. Pick one:
- OpenClaw presents the key (recommended)
- llmman holds the key
apiKey to the hosted provider’s key. llmman forwards it per request and never persists it. Local-only models on the same provider entry ignore the value.apiKey: "${LLMMAN_API_KEY}" in the config below.Hybrid config
LLMMAN_CONTEXT_LENGTH, or 262144 × 4 = 1048576 bytes when unset. This budget
does not follow a model’s trained-context cap or a later out-of-memory
reduction. Set LLMMAN_CONTEXT_LENGTH=65536 in the daemon’s environment to
match the example, or set LLMMAN_HYBRID_LOCAL_BYTES to pick the byte budget
directly. An explicit LLMMAN_CONTEXT_LENGTH=0 disables size-based routing
unless a positive LLMMAN_HYBRID_LOCAL_BYTES supplies a budget; setting the
byte override to 0 also disables that rule. Local context-refusal fallback
still applies unless the request is pinned local.
Set contextWindow on the pair to the local model’s usable token context. OpenClaw then compacts
around the local model’s limit, so most turns stay local; llmman still
overflows to gpt-5.6-luna when a request exceeds it. Set it to the hosted
model’s window instead if you prefer fewer compactions and more hosted
traffic.
Pinning a side
OpenClaw sends provider-levelheaders on every request, so a second provider
entry on the same base URL can force one side of the pair:
/model llmman-cloud/... for a turn that should go hosted, or use
headers: { "x-llmman-route": "local" } for an entry that must never leave
the machine.
Hybrid versus OpenClaw fallbacks
primary with a direct hosted
model as a fallback for when llmman itself is down:
Hosted models through llmman
llmman.provider/<provider>/<model> forwards to a hosted provider with no
local half. Use it when you want every model, local or hosted, behind one
endpoint and one key-handling story:
llmman providers shows which providers have a key;
llmman list --provider openai lists that provider’s models and prices. For
direct hosted access without llmman in the path, configure the
OpenAI provider instead.
Vision and image description
Models that ship a companionmmproj projector are vision-capable; qwen3.8
and gemma4:e4b are. llmman show <model> logs found companion mmproj file
and /api/show reports a vision capability. Mark those models
input: ["text", "image"] so image attachments are injected into agent turns.
--model must be a full <provider/model> ref. Use infer image describe
for OpenClaw’s image-understanding flow and configured imageModel; use
infer model run --file for a raw multimodal probe with a custom prompt.
To make a llmman model the default image-understanding provider for inbound
media:
models.providers.llmman.timeoutSeconds still
governs the underlying HTTP request for normal model calls.
Configuration
- Local only
- LAN llmman host
- On-demand startup
Common recipes
Replace model ids with names fromllmman list or
openclaw models list --provider llmman.
Local Qwen3.8 as the default agent model
Local Qwen3.8 as the default agent model
LLMMAN_CONTEXT_LENGTH=65536 in the daemon’s environment if you want the server context to match it exactly.Local first, hosted overflow
Local first, hosted overflow
llmman.hybrid/qwen3.8,openai/gpt-5.6-luna as primary, LLMMAN_API_KEY set to the OpenAI key, LLMMAN_CONTEXT_LENGTH matching the pair’s contextWindow. Add openai/gpt-5.6-luna to fallbacks so a stopped daemon does not block replies.Small local profile
Small local profile
openai-completions provider use structured Tool Search automatically when unset. The explicit setting below pins that surface, keeping optional capabilities available while loading their schemas only when needed. Cap the context to what the host can run with LLMMAN_CONTEXT_LENGTH=32768 in the daemon’s environment:compat.supportsTools: false only when the model or server reliably fails on tool schemas; it disables tool use entirely. For a deliberately narrower agent, prefer tools.profile or a per-agent tool policy.Multiple llmman hosts
Multiple llmman hosts
llmman config set aggregation.peers <host>,<host> (or LLMMAN_PEERS) makes one daemon forward to peers, so OpenClaw sees a single endpoint whose model list spans the group.Hugging Face or private-registry models
Hugging Face or private-registry models
llmman/hf.co/unsloth/Qwen3.5-0.8B-GGUF. Run llmman login <registry> first for private registries.Model selection
models.providers.llmman.timeoutSeconds covers
connection setup, headers, body streaming, and the total guarded-fetch abort
for that provider’s model requests only.
Quick verification
127.0.0.1 with the baseUrl host. If curl works
but OpenClaw does not, check whether the Gateway runs on a different machine,
container, or service account.
Advanced configuration
Context windows
Context windows
LLMMAN_CONTEXT_LENGTH is the server-side context (there is no flag). Semantics by backend:- llama-server (GGUF): set, it is passed as
--ctx-sizefor generation models;0means the trained context. Unset, llmman uses262144or the model’s trained context if smaller, and on out-of-memory retries with the context halved down to a 16,384 floor. - vLLM (safetensors): a positive value becomes
--max-model-len; unset uses the vLLM default. - mlx_lm.server: not forwarded.
LLMMAN_NUM_PARALLEL scales --ctx-size up by that factor so each slot keeps the full context.On the OpenClaw side, contextWindow declares the model’s window and contextTokens caps active input. Keep contextWindow at or below the server value; OpenClaw derives compaction and preflight thresholds from it. OpenClaw’s contextWindow does not change llmman’s hybrid byte budget; configure the daemon separately as described in Hybrid config.Thinking control
Thinking control
reasoning_content, which OpenClaw’s openai-completions adapter separates from the final text. Requests are proxied to llama-server, so chat_template_kwargs passes through. To turn thinking off for agent turns with a local Qwen model:openclaw agent --model llmman/qwen3.8 --thinking off, /think off, and openclaw infer model run --local --model llmman/qwen3.8 --thinking off --prompt "Reply with exactly: pong" --json map the thinking setting to chat_template_kwargs.enable_thinking. Without it, the generic proxy defaults do not send this control or reasoning_effort.The lean infer model run path does not read the agent-level params recipe above; use the compatibility declaration and --thinking off for that probe. Do not combine a fixed enable_thinking agent param with per-run control, since the fixed param overrides the generated value. Apply Qwen-specific controls to a hybrid ref only if both its local and hosted backends accept them.Model lifecycle and keep-alive
Model lifecycle and keep-alive
llmman ps shows loaded models with their context and expiry; llmman stop <model> unloads one now. A first request after startup or an idle unload pays the load cost, so set timeoutSeconds on the provider and raise LLMMAN_KEEP_ALIVE to keep the daemon warm for chat surfaces.GPU and backend selection
GPU and backend selection
llama-server release if none is on PATH. Override with LLMMAN_LLM_LIBRARY: cpu, cuda, cuda13, rocm, vulkan, or metal. Other knobs: LLMMAN_FLASH_ATTENTION (on/off/auto), LLMMAN_KV_CACHE_TYPE (f16, q8_0, q4_0), LLMMAN_SCHED_SPREAD for multi-GPU layer splitting, LLMMAN_IGPU_ENABLE to count integrated GPUs. On Linux, llmman serve --ociman docker|podman runs llama-server from the ghcr.io/ggml-org/llama.cpp images instead of a local binary. LLMMAN_DEBUG=1 prints the probe result.Memory embeddings
Memory embeddings
/v1/embeddings for GGUF embedding models, so memory search can use it through the generic openai-compatible embedding provider:LLMMAN_CONTEXT_LENGTH. See Memory config for the remaining fields.Ollama-compatible API
Ollama-compatible API
/api/chat, /api/tags, /api/show, and /api/ps, and OLLAMA_HOST=127.0.0.1:17434 ollama run <model> works against it. Prefer api: "openai-completions" from OpenClaw anyway: llmman’s /api/show reports only completion and vision capabilities, so the bundled Ollama plugin’s discovery would mark every llmman model compat.supportsTools: false. If you do point the Ollama plugin at http://127.0.0.1:17434 (no /v1), list models explicitly instead of relying on discovery.Compat flags
Compat flags
qwen3.8 on llama-server. If a different backend or model rejects them:messages[].content: invalid type: sequence, expected a string→ setcompat.requiresStringContent: trueon the model entry. OpenClaw then flattens pure text content parts into plain strings.400 JSON schema conversion failed→llama-servercould not compile a tool schema into its grammar subset. Update OpenClaw first; if a third-party tool or MCP server contributes the offending schema, disable it for that agent, and usecompat.supportsTools: falseonly as a last resort.
Proxy-style behavior
Proxy-style behavior
openai-completions endpoint, OpenClaw treats it as a proxy route: no service_tier, no Responses store, no prompt-cache hints, no OpenAI reasoning-compat payload shaping, no hidden OpenClaw attribution headers, and compat.supportsDeveloperRole is forced to false. Vendor-specific fields can be merged into the request body with agents.defaults.models["llmman/<model>"].params.extra_body.Model costs
Model costs
0. The hosted half of a hybrid or llmman.provider/... ref is billed by that provider; llmman list --provider openai shows its per-million-token prices.Troubleshooting
curl /v1/models fails
curl /v1/models fails
llmman serve is not running or is not reachable at the configured address. The default is 127.0.0.1:17434; if you set LLMMAN_HOST, update the OpenClaw baseUrl and healthUrl to match.Unknown model
Unknown model
llmman list; a short name maps to docker.io/ai/<name>:latest, and an owner/repo name maps to hf.co/owner/repo.Cold model times out
Cold model times out
models.providers.llmman.timeoutSeconds, warm the model with a first openclaw infer model run, and consider a longer LLMMAN_KEEP_ALIVE on the daemon.Hybrid request fails with no API key for provider
Hybrid request fails with no API key for provider
llmman-local) rather than a real key, or you used the llmman placeholder without giving the daemon its own OPENAI_API_KEY, or the daemon is bound off loopback and refuses to spend its own key. See Hosted-provider key.Hybrid requests never go hosted (or always do)
Hybrid requests never go hosted (or always do)
LLMMAN_CONTEXT_LENGTH (4 bytes per token) unless x-llmman-route is set. Lower LLMMAN_HYBRID_LOCAL_BYTES to overflow sooner, raise it to stay local longer, or pin a side with a provider headers entry as in Pinning a side.Direct /v1/chat/completions calls pass but openclaw infer model run fails
Direct /v1/chat/completions calls pass but openclaw infer model run fails
compat.supportsTools cannot change this failure. Check the configured base URL, model id, and LLMMAN_API_KEY, inspect the daemon and backend logs, and compare the two request payloads.Model run passes but a normal agent turn fails
Model run passes but a normal agent turn fails
400 JSON schema conversion failed is llama-server rejecting a tool schema; update OpenClaw and check third-party tools or MCP servers. Otherwise enable Tool Search to defer schemas, confirm the server’s actual context allocation, and use compat.supportsTools: false only as a last resort. See Smaller or stricter backends.llama-server crashes on larger agent turns
llama-server crashes on larger agent turns
llama-server still crashes on larger turns, treat it as an upstream llama.cpp or model limitation. Lower LLMMAN_CONTEXT_LENGTH, set LLMMAN_KV_CACHE_TYPE=q8_0 to reduce memory, or switch the backend or model.Model outputs tool JSON as text
Model outputs tool JSON as text
/v1/chat/completions (not the Ollama or Anthropic surfaces via a proxy). If the model only calls tools when forced, set params.extra_body.tool_choice: "required" on that model ref as described in Local models.Related
Local models
Local model services
OpenAI
Inference CLI
openclaw infer model run and the other one-shot probes used on this page.