Codex CLI on your own GPUs.
LM Warden as a custom model provider for Codex CLI, one served model.
- Protocol
- OpenAI Responses
- Support
- Local models only
- Verified
- Run against a live warden on 2026-10-05 with codex-cli 0.160.0. Two turns with a shell tool call, a workspace-write file edit, the session pinned by session-id and cached tokens measured on turn 2. Thinking on by default and off with model_reasoning_effort = “none”, both run live.
Local models only. Codex speaks only the Responses API; the warden translates POST /v1/responses to Chat Completions for the served model, so the reply, tool calls and usage reach Codex as Responses events. It is stateless: Codex sends store:false and the whole conversation each turn, and previous_response_id is refused. The thinking is on by default (Codex asks for reasoning summaries and the model’s reasoning is shown as a summary; to turn it off, set model_reasoning_effort = “none” in config.toml), hosted web search is dropped (set web_search = “disabled”), and router rules do not apply. Set model_context_window to the model’s real window: Codex assumes 272k for names it does not know. The engine needs tool calling enabled (vLLM: --enable-auto-tool-choice and a --tool-call-parser). A request for a model that is not loaded gets 404 and Codex retries for several seconds before it gives up.
Requirements
- Needs a model with tool calling (agentic edits).
- Works best with at least 32,768 tokens of context.
Setup
In the console, Connect writes these files with your warden’s address, your key and the model you picked. Here they are with placeholders.
Local models
Add LM Warden as a model provider in ~/.codex/config.toml.
~/.codex/config.toml)model = "your-served-model-name" model_provider = "lmwarden" # Codex assumes a 272k window for models it does not know; set the real one. model_context_window = 32768 [model_providers.lmwarden] name = "LM Warden" base_url = "https://your-warden/v1" env_key = "LMWARDEN_KEY" wire_api = "responses" # The model's context window is unknown: 32768 is a guess. Set it to the model's real window.
export LMWARDEN_KEY=vw_YOUR_KEY codex exec "Reply with the single word OK."
Official documentation: https://learn.chatgpt.com/docs/config-file/config-reference (read 2026-10-04).