LM Warden

Smaller things you’ll find once it’s running.

God mode
Watch one key’s prompts and responses as they happen, from a dock on its page. Held in memory only, streamed only while the dock is open, and off unless you set VW_GODMODE_ENABLED.
Admin tokens
vwa_ tokens let scripts, CI and agents drive the control API without the admin password, a cookie jar or a CSRF header. Shown once, stored as a SHA-256 hash.
A HuggingFace cache manager
What is on disk and which model owns it. Garbage-collect orphans, and export or import the whole cache as a tarball.
One port
The UI, the control API and /v1 share :8080. Your own TLS terminator, ingress or SSO goes in front of it unchanged.
No route out needed
Three image tarballs and an optional model-cache tarball are the whole transport for an air-gapped host. HF_HUB_OFFLINE=1 stops even the weight downloads.
Router stats
Every decision on /router/stats: local, passed through, fell back, refused, with the reason, per rule and per target, and the breaker state. Per process, never a header or a body.
A setup page per client
Connect writes the setup for Claude Code, Codex CLI, OpenCode, Aider, Grok CLI, the Anthropic and OpenAI SDKs and three editors, with this warden’s address, a key and a loaded model, and sends a test request.
Codex CLI
Run against a warden on 5 October 2026 with Codex 0.160: shell tool calls, a file edit, the session kept on one replica. Thinking is on by default; model_reasoning_effort = "none" turns it off. Local models only: router rules do not apply.
A chat playground
Talk to a loaded model in the browser, images included, before you hand anyone a key.

Two things we learned the hard way.

The engine died four times in one day, and the proxy kept forwarding to it.

The vLLM engine core crashed while the vllm serve wrapper around it stayed up, so the process looked alive and requests went on being sent to it. Now a watchdog asks the engine’s own /health endpoint instead of the process table. When the engine stops answering, it keeps the evidence, including a bounded tail of the log, and reloads the model.

A model’s live log panel under load: vLLM API server lines answering GET /metrics, POST /v1/chat/completions from a client, engine lines reporting 10 requests running and 0 waiting, and one GET /health line, all returning 200 OK The watchdog’s probe, answered by the engine
Fig. 15The live log on a model’s page during the same load: scrapes of /metrics, chat completions coming in, the engine’s own count of 10 requests running and none waiting, and one /health probe between them. Scroll the image sideways to read the full lines.

A 16 GiB card holds less than the arithmetic says.

The weights of a bf16 8B model are about 15 GiB on their own, and it runs out of memory while loading. We tried each row below on 16 GiB RTX A4000 and Quadro RTX 5000 cards and judged it by reading the model’s answers in the chat playground, not by trusting a health check.

ModelWhat worksWhat does not
openai/gpt-oss-20bTwo cards, tensor parallel 2, max_model_len 32000, gpu_memory_utilization 0.7 to 0.9One card: the 20B MoE doesn’t fit
Llama 3.1 8B, Mistral 7B classOne card, with an AWQ-INT4 checkpoint, --enforce-eager and gpu_memory_utilization 0.9The bf16 repo: about 15 GiB of weights, out of memory during load
Two models at onceOne model per card, each fitting its own cardTwo models sharing one card

Checked against vLLM 0.20.0, before the engine moved to 0.26.0, the version this release ships and the replica measurements above used. The rules have held across those upgrades; the exact numbers are worth rechecking on your engine. The full table and the reasons behind it are in HAZARDS.md. Not on NVIDIA? Other GPUs are on the roadmap.