How to use Warp with Subconscious

Let your AI do the setup. Paste this into Claude Code, Cursor, ChatGPT or any assistant.

Set up Warp to use Subconscious as its model provider.

Subconscious is an inference API for agents. It serves open models through OpenAI- and Anthropic-compatible endpoints, billed per token.
- OpenAI-compatible base URL: https://api.subconscious.dev/v1
- Anthropic-compatible base URL: https://api.subconscious.dev
- Model: subconscious/glm-5.3-marathon
- API key: read it from the SUBCONSCIOUS_API_KEY environment variable. Never commit the key or paste it into project files. If it is not set, ask me to create one at https://platform.subconscious.dev.

Follow these steps. Run commands and edit files yourself. For any step in a settings screen, tell me exactly what to click and what to enter.

1. Create a Subconscious API key
Sign in at https://platform.subconscious.dev and create an API key. You will paste it into the settings screen below.

2. Open custom inference settings
In Warp, open Settings and search for "inference endpoint".

3. Add the Subconscious endpoint
Add a new endpoint with these values. You can give the model a short alias.
   - Endpoint URL: https://api.subconscious.dev/v1
   - API key: Your Subconscious API key
   - Model identifier: subconscious/glm-5.3-marathon

4. Select the model
Save, then pick the model in the model picker or with /model followed by its alias.

Full guide: https://www.subconscious.dev/agents/terminal-and-ide/warp.md

When you are done, confirm the tool is using subconscious/glm-5.3-marathon, then summarize what you changed.

To use Warp with Subconscious, open Settings, search for "inference endpoint", and add a custom endpoint with URL https://api.subconscious.dev/v1, your Subconscious API key and model subconscious/glm-5.3-marathon. Select it in the model picker and Warp's agent runs on Subconscious, billed per token.

Last verified against the Warp docs on

Before you start

  • The Warp terminal app
  • A Subconscious account and API key

Warp docs

Set up Warp

  1. Step 1: Create a Subconscious API key

    Sign in at https://platform.subconscious.dev and create an API key. You will paste it into the settings screen below.

  2. Step 2: Open custom inference settings

    In Warp, open Settings and search for "inference endpoint".

  3. Step 3: Add the Subconscious endpoint

    Add a new endpoint with these values. You can give the model a short alias.

    Endpoint URL
    https://api.subconscious.dev/v1
    API key
    Your Subconscious API key
    Model identifier
    subconscious/glm-5.3-marathon
  4. Step 4: Select the model

    Save, then pick the model in the model picker or with /model followed by its alias.

Why run Warp on Subconscious

Long sessions stay fast
2x faster task completion on long tasks than standard open-model inference, especially past 200k tokens of context.
Context that keeps going
The OrangeLine runtime compresses the parts of context that stopped mattering, on the GPU, while the run continues. That gives a 5M+ token context window with no stop-and-compact pause.
Lower cost on long tasks
50% to 80% lower cost on long tasks, because the runtime processes far fewer tokens. Cached input bills at a steep discount, and agentic coding sees cache hit rates above 95%.
Your data stays yours
Prompts, inputs and outputs are not recorded, and nothing you send is used to train a model.

Connection details

For any client that takes a custom OpenAI- or Anthropic-compatible endpoint.

OpenAI-compatible base URL
https://api.subconscious.dev/v1
Anthropic-compatible base URL
https://api.subconscious.dev
API key
From subc login or the Subconscious dashboard
Model
subconscious/glm-5.3-marathon
Multimodal, high-throughput model
subconscious/deepseek-v4.1-flash-marathon

Full reference in the Subconscious API docs.

Questions

Which Warp plans support custom inference endpoints?
Warp's docs list custom endpoints on Free and paid individual plans and for organizations with 10 or fewer employees; larger teams need Business or Enterprise. Requests route through Warp's servers, and custom endpoints are not used by Cloud Agents or Auto models.
Which Subconscious model should I use with Warp?
Start with subconscious/glm-5.3-marathon, an open frontier coding model built for long agent runs. subconscious/deepseek-v4.1-flash-marathon is a cheaper, high-throughput alternative that also takes image input. Use the full model ID, including the subconscious/ prefix.
Is Warp on Subconscious OpenAI- or Anthropic-compatible?
This guide connects Warp through the OpenAI-compatible Subconscious API. Subconscious serves both formats: OpenAI Chat Completions at https://api.subconscious.dev/v1 and Anthropic Messages at https://api.subconscious.dev.
How is Warp usage on Subconscious billed?
Per token by default, with no commitment. Heavy users can switch to a monthly token plan: a fixed price for a large daily token allowance shared across your team.

More terminal and IDE agents

See all