The engine died four times in one day, and the proxy kept forwarding to it.
The vLLM engine core crashed while the vllm serve wrapper around it
stayed up, so the process looked alive and requests went on being sent to it.
Now a watchdog asks the engine’s own /health endpoint instead of the
process table. When the engine stops answering, it keeps the evidence, including
a bounded tail of the log, and reloads the model.
The watchdog’s probe, answered by the engine
/metrics, chat completions coming in, the engine’s own count of 10 requests running and none waiting, and one /health probe between them. Scroll the image sideways to read the full lines.