LM Warden

GPU monitoring: the box behind these numbers.

Unless a caption says otherwise, the readings on this page come from one workstation: four RTX A4000s, with a single model loaded across all of them, busy with our own test traffic. These are the numbers the console showed while we recorded figs. 1 and 5. The fit tests further down also used a Quadro RTX 5000, and say so.

GPUs
4 × NVIDIA RTX A4000, 16 GiB each (Ampere, compute 8.6)
VRAM in use
59.4 of 64.0 GiB, one model across all four cards
PCIe
Gen 4 x16 per card
Driver
610.57.04, CUDA 13.3
GPU load, power and tokens, read off fig. 1
ReadingAt capture24 h busy median24 h peak
GPU utilisation, busiest card (%)100100100
GPU power, four cards summed (W)548530559
Tokens per minutenot shown190k880k

Your numbers will be different. That is the point of measuring them.

A number it can’t measure reads “not reported”.

Most GPU dashboards average four cards into one line and draw a missing value as 0. This one gives each card its own panel: temperature against the throttle point its driver reports, power against the card’s own cap, clocks, fan, the PCIe link as it is negotiated right now, and ECC and NVLink when the card has them. Two cards bought eighteen months apart stay two cards, and a value the driver doesn’t give is labelled as missing instead of drawn as zero.

Four GPU panels from the Stats page under load, one per RTX A4000: 99 to 100 percent utilisation in power state P2, temperatures from 79 to 94 degrees against a 100 degree throttle mark, a thermal slowdown warning on three cards, power at 134 to 138 of a 140 watt cap, clocks, driver and CUDA version, PCIe link, NVLink and ECC Throttle point from the driver, not a guess Why the clock is down, as the driver reports it Power against its own 140 W cap NVLink and ECC, as the card reports them
Fig. 5The four GPU panels from the same Stats page, at full size, taken during the same load. Every card is at 99 or 100% in power state P2, and three of them, at 90 to 94 °C, are held by the driver to between 76 and 84% of their top clock. Scroll the image sideways to see all four.
What is read where
Per cardTemperature and throttle point, power and its cap, core and memory clocks, fan, PCIe generation and width, ECC, NVLink, and the engine process using the card
Per hostVRAM in use, GPU utilisation (the busiest card), power (the sum of the cards) and tokens per minute
Per requestTime to first token, inter-token latency and duration, measured at the proxy the same way for both engines

What each request is waiting for.

The Stats page lists every request in flight with the client’s session and where the request is in its life: queued at the warden, prefill (the engine reading the prompt), thinking, writing a tool call, or answering. A badge gives the share of the prompt likely to be in the replica’s cache. A prefill that runs longer than this box’s learned prefill rate predicts says why it probably waits: behind other prompts on the same replica, a cache that was probably evicted, or slow for no reason the warden can see. The session comes from the id each tool already sends: Claude Code, Codex, OpenCode and pi out of the box, Hermes and Aider with one line of setup.

The In flight table with 14 active requests, all from the key agent-farm at one client address. Columns: token, client IP, session, model, phase, context window and elapsed. Six rows are in tool call or answering and eight in prefill, most with a badge estimating 94 to 100% of the prompt cached; three prefills carry an evict? badge and two an hourglass badge, 1·r0 and 1·r4. Below the table: prefill model 439 tok/s learned from 1132 requests, cache estimate right 57% of the time, and the hit rate by idle gap, 68% after up to 30 seconds falling to 18% after 5 to 15 minutes. No token yet, cache probably evicted One prompt ahead of it on replica 4
Fig. 12Fourteen requests in flight on the seven-GPU box. Under the table is what the warden has learned from finished requests: this box prefills about 439 tokens a second per request, its cache estimate was right 57% of the time, and a conversation idle for more than two minutes usually finds its cache gone (20% hit after 2 to 5 minutes). The key name is a stand-in and the address a documentation address. Below 768 pixels wide the table becomes a list of cards. Recorded on 5 October 2026. Scroll the image sideways to see every column.