What people ask before they install it.
What is LM Warden?
A control plane around two inference engines, vLLM and llama.cpp. You pull a model from HuggingFace, load it onto your own NVIDIA GPUs, mint an API key and call it through an OpenAI-compatible /v1 endpoint. A browser UI covers loading and unloading models, live engine logs, a chat playground, GPU and token statistics, and a page per API key. It does not make models faster and does not modify the engines. It adds the parts a bare engine does not ship.
Which GPUs and operating systems does it run on?
NVIDIA GPUs on a Linux x86_64 host. The floor is Turing (sm_75, the RTX 20-series generation): the base image ships CUDA 13, which dropped Maxwell, Pascal and Volta, so a GTX 1080 Ti, a Titan X or a V100 will not work. Turing through Blackwell is covered, including Ada, Hopper and the RTX 50-series. There is no ROCm, Apple Silicon, Intel GPU or CPU path today; Apple silicon, AMD Radeon and Intel GPUs are on the immediate roadmap, and which comes first depends on the feedback we get in the project’s GitHub Discussions. Several GPUs in one box are supported; one model split across several machines is not.
Will my existing OpenAI client work with it?
Yes. The gateway at /v1/* is OpenAI-compatible, so the openai SDK, LangChain, OpenWebUI and your own agents work unchanged once you change the base_url and the key. GET /v1/models also reports each model’s max_model_len, so a client does not have to guess the context window. Codex CLI, which speaks only the OpenAI Responses API, works too: POST /v1/responses is translated to Chat Completions for a model you loaded (stateless, and local only: router rules do not apply).
Can I use it with Claude Code and keep my Anthropic account?
Yes. Point ANTHROPIC_BASE_URL at the warden, keep your normal claude login, and put the warden key in ANTHROPIC_CUSTOM_HEADERS as X-LMWarden-Key. A rule maps a model name Claude Code asks for, such as claude-haiku*, to a model loaded on your GPUs; every other model, and every other /v1 path, is passed to Anthropic byte for byte on your own login. It needs Claude Code 2.1.227 or newer, it is off by default, and a key must be flagged “May relay to Anthropic” before anything is forwarded. Older clients can still use local-only mode.
Does Anthropic see my warden key, or the warden my Anthropic login?
Anthropic never sees the warden key: it travels in its own header and is removed before a request is passed on. The warden does handle your Anthropic login, because it has to forward it, but only to Anthropic. It is never sent to an engine and never logged, and nothing from a passed-through request is stored or shown in god mode.
Which inference engines does it use?
Mainline vLLM and llama.cpp, chosen per model and run unmodified: no patched image, no fork, no monkeypatching. Picking a .gguf file selects llama.cpp, which fits larger models onto smaller cards with GGUF quantisation and keeps older GPUs useful. TensorRT-LLM, SGLang, MLX and Ollama’s runtime are not shipped.
Does anything leave my machine?
Nothing leaves the host unless you turn on the Claude Code router and flag a key “May relay to Anthropic”. Then the requests from that key for Claude models no rule sends to your GPUs go to Anthropic, on the user’s own Anthropic login, and nothing else does. There is no account, licence check or analytics. Otherwise the stack only calls out when you ask it to: huggingface.co to pull weights (HF_HUB_OFFLINE=1 stops even that), Docker Hub when you open the engine-version picker, and the registry for the release images, which docker load replaces in an offline or air-gapped install.
Are prompts and responses stored?
Not by default. The request history behind the charts and per-key usage holds metadata only and has no column for prompt or completion text. Two diagnostic features can capture content, and both are off unless you turn them on: god mode (VW_GODMODE_ENABLED), a bounded in-memory view of one key’s traffic, and the content log (VW_CONTENT_LOG_ENABLED), which writes to disk only for the token ids you list.
How do API keys work for a team?
Every app or person gets its own vw_ key with its own page. Set a priority from P0 to P9 (a waiting P9 request always goes before a P8 one), pause a key so its clients get 403 token paused while running requests finish, or rotate it with a grace window in which the old secret keeps working. Requests, prompt tokens and completion tokens are counted against the key and the client IP, with queue wait and latency per key. Scripts, CI and agents that drive the control API use separate vwa_ admin tokens instead of the admin password.
Can I run several models on one GPU?
No. One loaded model per GPU is a rule enforced before an engine starts. You can load different models on different cards of the same host, and split one model across several cards with vLLM’s tensor parallelism.
What does installing it take?
A Linux host with Docker, Docker Compose v2.24 or newer, an NVIDIA GPU, the NVIDIA Container Toolkit and about 40 GB free where Docker keeps its images. The installer checks the host, lets you pick GPUs, generates the secrets, pulls the release images and offers to start the stack. A first-run wizard in the browser then asks for the GPUs to use, a HuggingFace token and your admin account.
What does it cost?
Nothing. LM Warden is licensed under Apache-2.0. It is software you run, not a hosted API: you supply the hardware, the driver, the disk and the electricity, and once the hardware is paid for a request costs electricity, not a per-token fee.