LM Warden

One key per app, and a page for every key.

Give each colleague, laptop, CI job and agent its own key. A key’s page shows what it spent and how long its requests waited in the queue, median and p95, over any window you drag or type. Rotate a key with a grace window, and set its priority from P0 to P9: a waiting P9 request always goes before a P8 one.

Pause it, and new requests are refused while the ones already running finish:

$ curl -i http://gpu-box:8080/v1/chat/completions \
    -H "Authorization: Bearer vw_…" -d '{…}'
HTTP/1.1 403 Forbidden
content-type: application/json

{"detail":"token paused"}

From a local instance on 23 September 2026. Host and key replaced, other headers cut.

Fig. 13The page of demo-agent, the key that ran the load in fig. 1: 37 requests and 43k generated tokens in the hour on screen. Then its god-mode dock opens (an opt-in live view of one key’s prompts and replies, off by default) and its replies stream in as they are written. The prompts are ours.
The Requests chart from the Stats page over the last 7 days, 16 to 23 September: one mark per request, coloured by API key, sized by generated tokens and shaped by finish reason, against duration on a log scale from 100 ms to 10,000 s. Most marks sit between 3 and 100 seconds; one key shows two dense columns reaching from under a second to several hundred seconds.
Fig. 14Every key’s requests on one chart, from the same Stats page: the last 7 days, 28,906 requests, one in fifteen drawn. A mark is one request, coloured by its key, sized by the tokens it generated and shaped by how it finished, at the height of its duration on a log scale. The steady amber is opencode-laptop, the two dense pink columns are research-notebook. Recorded on 23 September 2026; the key names are replaced. Scroll the image sideways to see the whole week.

Leaving is a base_url change.

It speaks /v1/chat/completions and /v1/models, and /v1/models reports each model’s max_model_len. It also speaks the Anthropic Messages API at /v1/messages, so Claude Code points ANTHROPIC_BASE_URL at a warden’s root URL (Claude Code adds the /v1/messages path itself), and the OpenAI Responses API at /v1/responses, so Codex CLI can use a model you loaded. The weights sit in an ordinary HuggingFace cache that make export-hf-cache hands you as a tarball, and the exact argv each engine was started with is one GET away, so a working setup can be rebuilt without us. Read the diff below backwards and that is how you leave.

  import os
  from openai import OpenAI

  client = OpenAI(
-     base_url="https://api.openai.com/v1",
-     api_key=os.environ["OPENAI_API_KEY"],
+     base_url="http://gpu-box:8080/v1",
+     api_key=os.environ["LM_WARDEN_KEY"],   # a vw_ key from API tokens
  )

  reply = client.chat.completions.create(
-     model="gpt-4o-mini",
+     model="gpt-oss-20b",   # the served name of a model you loaded
      messages=[{"role": "user", "content": "What is a KV cache?"}],
  )