LM Warden

OpenAI SDK on your own GPUs.

Any OpenAI SDK, or anything with an OpenAI-compatible base URL setting.

Protocol
OpenAI Chat Completions
Support
Local models only
Verified
Run against a live warden on 2026-10-04 with openai 3.24.0 (Python), openai 7.28.0 (TypeScript), curl. Python, TypeScript and curl, one chat completion each.

Local models only. There is no pass-through to OpenAI: a request for gpt-* gets 404. Chat Completions and Completions are served; /v1/responses is a stateless translation to Chat Completions (see Codex CLI).

Setup

In the console, Connect writes these files with your warden’s address, your key and the model you picked. Here they are with placeholders.

Local models

base_url is this warden’s /v1 and the warden key is the API key.

Python
from openai import OpenAI

client = OpenAI(base_url="https://your-warden/v1", api_key="vw_YOUR_KEY")
response = client.chat.completions.create(
    model="your-served-model-name",
    messages=[{"role": "user", "content": "Reply with the single word OK."}],
)
print(response.choices[0].message.content)
TypeScript
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://your-warden/v1", apiKey: "vw_YOUR_KEY" });
const response = await client.chat.completions.create({
  model: "your-served-model-name",
  messages: [{ role: "user", content: "Reply with the single word OK." }],
});
console.log(response.choices[0].message.content);
curl
curl https://your-warden/v1/chat/completions \
  -H "Authorization: Bearer vw_YOUR_KEY" \
  -H "content-type: application/json" \
  -d '{"model": "your-served-model-name", "stream": false, "messages": [{"role": "user", "content": "Reply with the single word OK."}]}'

Official documentation: https://github.com/openai/openai-python#readme (read 2026-10-04).