LM Warden

Continue on your own GPUs.

Continue’s OpenAI provider pointed at this warden, one served model.

Protocol
OpenAI Chat Completions
Support
Configured from its docs
Verified
Checked against its docs on 2026-10-04. Config checked against the docs; the streaming Chat Completions request it sends, with a tool call, passed with curl.

Local models only, over Chat Completions. The config follows Continue’s docs and the request it sends was tested from this page; the editor itself was not run. Router mode would need custom headers Continue does not document.

Requirements

  • Works best with at least 32,768 tokens of context.

Setup

In the console, Connect writes these files with your warden’s address, your key and the model you picked. Here they are with placeholders.

Local models

A model entry in ~/.continue/config.yaml.

config.yaml (~/.continue/config.yaml)
name: LM Warden
version: 0.0.1
schema: v1
models:
  - name: your-served-model-name
    provider: openai
    model: your-served-model-name
    apiBase: https://your-warden/v1
    apiKey: vw_YOUR_KEY
    roles: [chat, edit, apply]
    capabilities: [tool_use]
    useResponsesApi: false   # names like o*/gpt-5* would otherwise use /responses
    defaultCompletionOptions:
      contextLength: 32768
      maxTokens: 8192
# The model's context window is unknown: 32768 is a guess. Set it to the model's real window.
config.yaml (~/.continue/config.yaml)
name: LM Warden
version: 0.0.1
schema: v1
models:
  - name: your-served-model-name
    provider: openai
    model: your-served-model-name
    apiBase: https://your-warden/v1
    apiKey: vw_YOUR_KEY
    roles: [chat, edit, apply]
    useResponsesApi: false   # names like o*/gpt-5* would otherwise use /responses
    defaultCompletionOptions:
      contextLength: 32768
      maxTokens: 8192
# The model's context window is unknown: 32768 is a guess. Set it to the model's real window.

Official documentation: https://docs.continue.dev/reference (read 2026-10-04).