Claude Code on your own GPUs.
Keep your Claude login and run the models your rules name here, or run it fully local.
- Protocol
- Anthropic Messages
- Support
- Full support
- Verified
- Run against a live warden on 2026-10-04 with Claude Code 2.1.289. Local-only and router modes run live. Router: claude-haiku-4-5 was answered here, including a file read, and the default model went to Anthropic on the Claude login. Local only: a normal reply came back from a served model.
The Claude models your rules name run here; every other model and /v1 path goes to Anthropic on your own login. Never put the warden key in ANTHROPIC_API_KEY or ANTHROPIC_AUTH_TOKEN in router mode: that would replace your login. Local only runs everything on one served model, including the small background tasks, and nothing goes to Anthropic.
Requirements
- Claude Code 2.1.227 or newer (ANTHROPIC_CUSTOM_HEADERS).
- Routing must be on, with at least one enabled rule.
- The key must be allowed to relay to Anthropic (May relay).
- Needs at least 49,152 tokens of context: the first turn alone is about 38,000 tokens before any file is read.
- 131,072 tokens or more for real sessions; less works for short ones.
- Claude Code may warn that the model name is not in its catalogue. That is informational: it assumes a 200k window and auto-compacts accordingly. Set CLAUDE_CODE_MAX_CONTEXT_TOKENS to the model’s real context window.
Setup
In the console, Connect writes these files with your warden’s address, your key and the model you picked. Here they are with placeholders.
Router: keep your Claude login
Your Claude login keeps flowing to Anthropic; the warden key travels in its own header, and the models your rules name are answered here.
export ANTHROPIC_BASE_URL=https://your-warden export ANTHROPIC_CUSTOM_HEADERS="X-LMWarden-Key: vw_YOUR_KEY" claude
~/.claude/settings.json){
"env": {
"ANTHROPIC_BASE_URL": "https://your-warden",
"ANTHROPIC_CUSTOM_HEADERS": "X-LMWarden-Key: vw_YOUR_KEY"
}
}Local only (no Anthropic account)
Every model slot is the picked served model and the warden key is the auth token. Nothing goes to Anthropic.
export ANTHROPIC_BASE_URL=https://your-warden export ANTHROPIC_AUTH_TOKEN=vw_YOUR_KEY export ANTHROPIC_MODEL=your-served-model-name export ANTHROPIC_DEFAULT_OPUS_MODEL=your-served-model-name export ANTHROPIC_DEFAULT_SONNET_MODEL=your-served-model-name export ANTHROPIC_DEFAULT_HAIKU_MODEL=your-served-model-name export CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1 export CLAUDE_CODE_MAX_CONTEXT_TOKENS=32768 # The model's context window is unknown: 32768 is a guess. Set it to the model's real window.
~/.claude/settings.json){
"env": {
"ANTHROPIC_BASE_URL": "https://your-warden",
"ANTHROPIC_AUTH_TOKEN": "vw_YOUR_KEY",
"ANTHROPIC_MODEL": "your-served-model-name",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "your-served-model-name",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "your-served-model-name",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "your-served-model-name",
"CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS": "1",
"CLAUDE_CODE_MAX_CONTEXT_TOKENS": "32768"
}
}Official documentation: https://code.claude.com/docs/en/llm-gateway (read 2026-10-04).
One ANTHROPIC_BASE_URL, two places an answer can come from.
Every request Claude Code sends to Anthropic counts against your plan’s usage limits, or is billed per token on an API key, and takes its prompt, your code included, off your machine. Many of those requests are small. Answer them on cards you already own and they cost electricity, they don’t count toward your limits, and their prompts stay on your host.
Claude itself never runs here. A rule names the Claude models this box answers for,
claude-haiku* for example, and a model you loaded answers in their place, under
the name Claude Code asked for. Anything else, Opus included, is passed to Anthropic byte
for byte on your own login, streaming and all. Your warden key rides in its own header, so
the login is never touched, and nothing is relayed for a key without the flag
“May relay to Anthropic”.
Is a small model good enough? We have not measured answer quality, so this is judgement, not a benchmark. Claude Code uses its Haiku-class model for lighter background work, and a small local model is a reasonable stand-in there and for simple, well-specified edits. For changes across many files, debugging and long plans, keep Sonnet and Opus on Anthropic, which is what the catch-all does unless you add a rule. A bigger model on bigger cards moves that line; our Claude Code and replica measurements below used a 4B.
Size the context for Claude Code, not for chat: its first turn in our session was 38,167 tokens, so the local model needs well over 40k tokens of context.
In words: Claude Code sends the request to the warden with its own Anthropic login in Authorization and the warden key in X-LMWarden-Key. The warden tries its rules in order. A request for a model a rule names is translated, queued by the scheduler and sent by the replica router to one vLLM replica. Any other request, such as Opus, is passed on to api.anthropic.com with the login forwarded and the warden key removed. If the local leg fails before its first byte, the request goes to Anthropic, or the rule refuses it with a 529 that Claude Code retries.
export ANTHROPIC_BASE_URL=https:// warden.example export ANTHROPIC_CUSTOM_HEADERS= "X-LMWarden-Key: vw_…" claude
Claude Code 2.1.227 or newer. Keep your normal claude login; the key goes in its own header, not in ANTHROPIC_API_KEY.
First match wins; every other model goes to Anthropic
Anthropic is reached on the user’s own login
claude-haiku* to a local model and the catch-all passes everything else to Anthropic; rules are numbered in the order they are tried. The Traffic panel counts 107 local, 2 passed through, 0 fell back and 2 refused. Scroll the image sideways to see the whole page.
Expect the first turn to be the slow one. In an earlier session, Claude Code’s first turn
was 38,167 tokens and the local model was loaded with a 32,768-token context, so the engine
refused it (status_400) and the request went to Anthropic instead. With the
context raised to 65,536 tokens and an fp8 KV cache (the attention cache kept at 8 bits,
half the memory), the 4B model on 16 GiB cards answered that same turn locally in
28 s, with the whole 38k-token prompt read from cold and nothing in the cache yet. Later turns
of a conversation find most of their prompt already in the engine’s cache: the median time
to first byte for the claude-haiku* rule in fig. 4 is 290 ms.
The breaker for this model: closed
Why a rule went to Anthropic or was refused, by reason