One key per app, and a page for every key.
Give each colleague, laptop, CI job and agent its own key. A key’s page shows what it spent and how long its requests waited in the queue, median and p95, over any window you drag or type. Rotate a key with a grace window, and set its priority from P0 to P9: a waiting P9 request always goes before a P8 one.
Pause it, and new requests are refused while the ones already running finish:
$ curl -i http://gpu-box:8080/v1/chat/completions \
-H "Authorization: Bearer vw_…" -d '{…}'
HTTP/1.1 403 Forbidden
content-type: application/json
{"detail":"token paused"}
From a local instance on 23 September 2026. Host and key replaced, other headers cut.
demo-agent, the key that ran the load in fig. 1: 37 requests and 43k generated tokens in the hour on screen. Then its god-mode dock opens (an opt-in live view of one key’s prompts and replies, off by default) and its replies stream in as they are written. The prompts are ours.
opencode-laptop, the two dense pink columns are research-notebook. Recorded on 23 September 2026; the key names are replaced. Scroll the image sideways to see the whole week.Leaving is a base_url change.
It speaks /v1/chat/completions and /v1/models, and
/v1/models reports each model’s max_model_len. It also
speaks the Anthropic Messages API at /v1/messages, so Claude Code
points ANTHROPIC_BASE_URL at a warden’s root URL (Claude Code adds the
/v1/messages path itself), and the OpenAI Responses API at
/v1/responses, so Codex CLI can use a model you loaded. The weights sit
in an ordinary HuggingFace cache that make export-hf-cache hands you as a
tarball, and the exact argv each engine was started with is one GET away, so a
working setup can be rebuilt without us. Read the diff below backwards and
that is how you leave.
import os from openai import OpenAI client = OpenAI( - base_url="https://api.openai.com/v1", - api_key=os.environ["OPENAI_API_KEY"], + base_url="http://gpu-box:8080/v1", + api_key=os.environ["LM_WARDEN_KEY"], # a vw_ key from API tokens ) reply = client.chat.completions.create( - model="gpt-4o-mini", + model="gpt-oss-20b", # the served name of a model you loaded messages=[{"role": "user", "content": "What is a KV cache?"}], )