Continue on your own GPUs.
Continue’s OpenAI provider pointed at this warden, one served model.
- Protocol
- OpenAI Chat Completions
- Support
- Configured from its docs
- Verified
- Checked against its docs on 2026-10-04. Config checked against the docs; the streaming Chat Completions request it sends, with a tool call, passed with curl.
Local models only, over Chat Completions. The config follows Continue’s docs and the request it sends was tested from this page; the editor itself was not run. Router mode would need custom headers Continue does not document.
Requirements
- Works best with at least 32,768 tokens of context.
Setup
In the console, Connect writes these files with your warden’s address, your key and the model you picked. Here they are with placeholders.
Local models
A model entry in ~/.continue/config.yaml.
~/.continue/config.yaml)name: LM Warden
version: 0.0.1
schema: v1
models:
- name: your-served-model-name
provider: openai
model: your-served-model-name
apiBase: https://your-warden/v1
apiKey: vw_YOUR_KEY
roles: [chat, edit, apply]
capabilities: [tool_use]
useResponsesApi: false # names like o*/gpt-5* would otherwise use /responses
defaultCompletionOptions:
contextLength: 32768
maxTokens: 8192
# The model's context window is unknown: 32768 is a guess. Set it to the model's real window.~/.continue/config.yaml)name: LM Warden
version: 0.0.1
schema: v1
models:
- name: your-served-model-name
provider: openai
model: your-served-model-name
apiBase: https://your-warden/v1
apiKey: vw_YOUR_KEY
roles: [chat, edit, apply]
useResponsesApi: false # names like o*/gpt-5* would otherwise use /responses
defaultCompletionOptions:
contextLength: 32768
maxTokens: 8192
# The model's context window is unknown: 32768 is a guess. Set it to the model's real window.Official documentation: https://docs.continue.dev/reference (read 2026-10-04).