LM Warden

Claude Code on your own GPUs.

Keep your Claude login and run the models your rules name here, or run it fully local.

Protocol
Anthropic Messages
Support
Full support
Verified
Run against a live warden on 2026-10-04 with Claude Code 2.1.289. Local-only and router modes run live. Router: claude-haiku-4-5 was answered here, including a file read, and the default model went to Anthropic on the Claude login. Local only: a normal reply came back from a served model.

The Claude models your rules name run here; every other model and /v1 path goes to Anthropic on your own login. Never put the warden key in ANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN in router mode: that would replace your login. Local only runs everything on one served model, including the small background tasks, and nothing goes to Anthropic.

Requirements

  • Claude Code 2.1.227 or newer (ANTHROPIC_CUSTOM_HEADERS).
  • Routing must be on, with at least one enabled rule.
  • The key must be allowed to relay to Anthropic (May relay).
  • Needs at least 49,152 tokens of context: the first turn alone is about 38,000 tokens before any file is read.
  • 131,072 tokens or more for real sessions; less works for short ones.
  • Claude Code may warn that the model name is not in its catalogue. That is informational: it assumes a 200k window and auto-compacts accordingly. Set CLAUDE_CODE_MAX_CONTEXT_TOKENS to the model’s real context window.

Setup

In the console, Connect writes these files with your warden’s address, your key and the model you picked. Here they are with placeholders.

Router: keep your Claude login

Your Claude login keeps flowing to Anthropic; the warden key travels in its own header, and the models your rules name are answered here.

Shell
export ANTHROPIC_BASE_URL=https://your-warden
export ANTHROPIC_CUSTOM_HEADERS="X-LMWarden-Key: vw_YOUR_KEY"
claude
settings.json (~/.claude/settings.json)
{
  "env": {
    "ANTHROPIC_BASE_URL": "https://your-warden",
    "ANTHROPIC_CUSTOM_HEADERS": "X-LMWarden-Key: vw_YOUR_KEY"
  }
}

Local only (no Anthropic account)

Every model slot is the picked served model and the warden key is the auth token. Nothing goes to Anthropic.

Shell
export ANTHROPIC_BASE_URL=https://your-warden
export ANTHROPIC_AUTH_TOKEN=vw_YOUR_KEY
export ANTHROPIC_MODEL=your-served-model-name
export ANTHROPIC_DEFAULT_OPUS_MODEL=your-served-model-name
export ANTHROPIC_DEFAULT_SONNET_MODEL=your-served-model-name
export ANTHROPIC_DEFAULT_HAIKU_MODEL=your-served-model-name
export CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1
export CLAUDE_CODE_MAX_CONTEXT_TOKENS=32768
# The model's context window is unknown: 32768 is a guess. Set it to the model's real window.
settings.json (~/.claude/settings.json)
{
  "env": {
    "ANTHROPIC_BASE_URL": "https://your-warden",
    "ANTHROPIC_AUTH_TOKEN": "vw_YOUR_KEY",
    "ANTHROPIC_MODEL": "your-served-model-name",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "your-served-model-name",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "your-served-model-name",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "your-served-model-name",
    "CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS": "1",
    "CLAUDE_CODE_MAX_CONTEXT_TOKENS": "32768"
  }
}

Official documentation: https://code.claude.com/docs/en/llm-gateway (read 2026-10-04).

One ANTHROPIC_BASE_URL, two places an answer can come from.

Every request Claude Code sends to Anthropic counts against your plan’s usage limits, or is billed per token on an API key, and takes its prompt, your code included, off your machine. Many of those requests are small. Answer them on cards you already own and they cost electricity, they don’t count toward your limits, and their prompts stay on your host.

Claude itself never runs here. A rule names the Claude models this box answers for, claude-haiku* for example, and a model you loaded answers in their place, under the name Claude Code asked for. Anything else, Opus included, is passed to Anthropic byte for byte on your own login, streaming and all. Your warden key rides in its own header, so the login is never touched, and nothing is relayed for a key without the flag “May relay to Anthropic”.

Is a small model good enough? We have not measured answer quality, so this is judgement, not a benchmark. Claude Code uses its Haiku-class model for lighter background work, and a small local model is a reasonable stand-in there and for simple, well-specified edits. For changes across many files, debugging and long plans, keep Sonnet and Opus on Anthropic, which is what the catch-all does unless you add a rule. A bigger model on bigger cards moves that line; our Claude Code and replica measurements below used a 4B.

Size the context for Claude Code, not for chat: its first turn in our session was 38,167 tokens, so the local model needs well over 40k tokens of context.

Where a request from Claude Code goes Claude Code sends a request to LM Warden with two headers: Authorization, its own login, and X-LMWarden-Key, the warden key. LM Warden tries its rules in order and the first match wins. A request for claude-haiku or claude-sonnet matches a rule, is answered by your local model under the Claude name it asked for, and goes through translate, the scheduler and the replica router to one of four vLLM replicas, the conversation staying on its home replica. A request that matches no rule, such as Opus, or any other /v1 path, is passed to api.anthropic.com with the login forwarded and the warden key removed. If the local leg fails before the first byte, a dashed path sends the request to Anthropic, or the rule can refuse with a 529 that Claude Code retries. Claude Code ANTHROPIC_BASE_URL points at the warden Authorization your own login X-LMWarden-Key vw_…, the warden key POST /v1/messages LM Warden rules, first match wins claude-haiku* your local model claude-sonnet* your local model anything else pass through reads the warden key, never forwards it match Local leg translate to the engine scheduler, key priority replica router breaker: 3 failures, open 60 s session id picks the replica vLLM replica 0 vLLM replica 1 vLLM replica 2 vLLM replica 3 a conversation stays on its replica no match: Opus, or any other /v1 path passed on byte for byte, streaming included only for a key with “May relay to Anthropic” local leg failed before the first byte: the rule falls back to Anthropic, or refuses with a 529 that Claude Code retries api.anthropic.com your login forwarded, the warden key removed Off by default. The content of a passed-through request is never logged, stored or shown.
Where a request from Claude Code goes The same flow as the wide drawing, top to bottom. Claude Code sends a request to LM Warden with its own login in Authorization and the warden key in X-LMWarden-Key. LM Warden tries its rules in order; a request for claude-haiku or claude-sonnet is answered by your local model and goes down through the local leg (translate, scheduler, replica router) to one of four vLLM replicas, where the conversation stays. A request that matches no rule goes down the right-hand line to api.anthropic.com, with the login forwarded and the warden key removed. A dashed line from the local leg to Anthropic is the fallback when the local leg fails before the first byte; the rule can instead refuse with a 529 that Claude Code retries. Claude Code Authorization your own login X-LMWarden-Key the warden key ANTHROPIC_BASE_URL = the warden POST /v1/messages LM Warden rules, first match wins claude-haiku* your local model claude-sonnet* your local model anything else pass through reads the warden key, never forwards it match Local leg translate to the engine scheduler, key priority replica router breaker: 3 failures, open 60 s session id picks the replica vLLM replica 0 vLLM replica 1 vLLM replica 2 vLLM replica 3 a conversation stays on its replica no match: Opus, any other /v1 path dashed: the local leg failed before its first byte, so the rule falls back, or refuses with a 529 that Claude Code retries api.anthropic.com your login forwarded, the warden key removed Off by default. Only a key with “May relay to Anthropic” is passed through.
Fig. 2Where a request goes. Rules are tried in order and the first match wins. A local answer passes the same scheduler and replica router as any other request; a replica is one full copy of the model, here one per GPU. A local failure before the first byte goes to Anthropic, or comes back as a 529 that Claude Code retries, and after 3 failures in a row the model’s breaker opens: for 60 s nothing is sent to it.

In words: Claude Code sends the request to the warden with its own Anthropic login in Authorization and the warden key in X-LMWarden-Key. The warden tries its rules in order. A request for a model a rule names is translated, queued by the scheduler and sent by the replica router to one vLLM replica. Any other request, such as Opus, is passed on to api.anthropic.com with the login forwarded and the warden key removed. If the local leg fails before its first byte, the request goes to Anthropic, or the rule refuses it with a 529 that Claude Code retries.

export ANTHROPIC_BASE_URL=https://warden.example
export ANTHROPIC_CUSTOM_HEADERS="X-LMWarden-Key: vw_…"
claude

Claude Code 2.1.227 or newer. Keep your normal claude login; the key goes in its own header, not in ANTHROPIC_API_KEY.

The Router page of the console. A green banner reads “Routing normally” and “96% answered locally since 13:38”. Under “Where each Claude model goes. First match wins”, rule 1 sends claude-haiku* to the loaded model qwen3-4b-dp4, with “If it fails: Anthropic” and “Thinking off”. The last row is the catch-all: every other Claude model and /v1 path goes to Anthropic on the user’s own login and needs a key with “May relay to Anthropic”. The Traffic panel counts 107 local, 2 Anthropic pass-through, 0 fell back to Anthropic and 2 refused or error, and lists two reasons, no_upstream_credential and relay_not_allowed, one each. First match wins; every other model goes to Anthropic Anthropic is reached on the user’s own login
Fig. 3The Router page after the warden had been running since 13:38: “Routing normally”, with 96% of requests answered locally since the restart. One rule sends claude-haiku* to a local model and the catch-all passes everything else to Anthropic; rules are numbered in the order they are tried. The Traffic panel counts 107 local, 2 passed through, 0 fell back and 2 refused. Scroll the image sideways to see the whole page.

Expect the first turn to be the slow one. In an earlier session, Claude Code’s first turn was 38,167 tokens and the local model was loaded with a 32,768-token context, so the engine refused it (status_400) and the request went to Anthropic instead. With the context raised to 65,536 tokens and an fp8 KV cache (the attention cache kept at 8 bits, half the memory), the 4B model on 16 GiB cards answered that same turn locally in 28 s, with the whole 38k-token prompt read from cold and nothing in the cache yet. Later turns of a conversation find most of their prompt already in the engine’s cache: the median time to first byte for the claude-haiku* rule in fig. 4 is 290 ms.

The Router activity page. Four counters: local 107, Anthropic pass-through 2, fell back 0, and refused · errors 2 · 0. A table “By rule” gives the rule claude-haiku* to qwen3-4b-dp4 with its local, fell-back and refused counts, latency and time to first byte at p50 and p95, and its breaker state, closed; the pass-through row gives the same timings for the requests passed to Anthropic. A table of local models shows qwen3-4b-dp4 with its breaker closed, and a panel “Fallback reasons” lists no_upstream_credential and relay_not_allowed, one each. The breaker for this model: closed Why a rule went to Anthropic or was refused, by reason
Fig. 4Router activity for the same period: 107 requests answered locally, 2 passed through to Anthropic, 0 fell back, and 2 refused with 0 errors. Below, latency and time to first byte per rule at p50 and p95, the breaker per local model, and the reasons a rule went to Anthropic or refused. Counters are per process; a restart zeroes them. Recorded on 4 October 2026. Scroll the image sideways to see the whole page.