How to use Warp with Subconscious
Let your AI do the setup. Paste this into Claude Code, Cursor, ChatGPT or any assistant.
Set up Warp to use Subconscious as its model provider. Subconscious is an inference API for agents. It serves open models through OpenAI- and Anthropic-compatible endpoints, billed per token. - OpenAI-compatible base URL: https://api.subconscious.dev/v1 - Anthropic-compatible base URL: https://api.subconscious.dev - Model: subconscious/glm-5.3-marathon - API key: read it from the SUBCONSCIOUS_API_KEY environment variable. Never commit the key or paste it into project files. If it is not set, ask me to create one at https://platform.subconscious.dev. Follow these steps. Run commands and edit files yourself. For any step in a settings screen, tell me exactly what to click and what to enter. 1. Create a Subconscious API key Sign in at https://platform.subconscious.dev and create an API key. You will paste it into the settings screen below. 2. Open custom inference settings In Warp, open Settings and search for "inference endpoint". 3. Add the Subconscious endpoint Add a new endpoint with these values. You can give the model a short alias. - Endpoint URL: https://api.subconscious.dev/v1 - API key: Your Subconscious API key - Model identifier: subconscious/glm-5.3-marathon 4. Select the model Save, then pick the model in the model picker or with /model followed by its alias. Full guide: https://www.subconscious.dev/agents/terminal-and-ide/warp.md When you are done, confirm the tool is using subconscious/glm-5.3-marathon, then summarize what you changed.
To use Warp with Subconscious, open Settings, search for "inference endpoint", and add a custom endpoint with URL https://api.subconscious.dev/v1, your Subconscious API key and model subconscious/glm-5.3-marathon. Select it in the model picker and Warp's agent runs on Subconscious, billed per token.
Last verified against the Warp docs on
Before you start
- The Warp terminal app
- A Subconscious account and API key
Set up Warp
Step 1: Create a Subconscious API key
Sign in at https://platform.subconscious.dev and create an API key. You will paste it into the settings screen below.
Step 2: Open custom inference settings
In Warp, open Settings and search for "inference endpoint".
Step 3: Add the Subconscious endpoint
Add a new endpoint with these values. You can give the model a short alias.
- Endpoint URL
- https://api.subconscious.dev/v1
- API key
- Your Subconscious API key
- Model identifier
- subconscious/glm-5.3-marathon
Step 4: Select the model
Save, then pick the model in the model picker or with /model followed by its alias.
Why run Warp on Subconscious
- Long sessions stay fast
- 2x faster task completion on long tasks than standard open-model inference, especially past 200k tokens of context.
- Context that keeps going
- The OrangeLine runtime compresses the parts of context that stopped mattering, on the GPU, while the run continues. That gives a 5M+ token context window with no stop-and-compact pause.
- Lower cost on long tasks
- 50% to 80% lower cost on long tasks, because the runtime processes far fewer tokens. Cached input bills at a steep discount, and agentic coding sees cache hit rates above 95%.
- Your data stays yours
- Prompts, inputs and outputs are not recorded, and nothing you send is used to train a model.
Connection details
For any client that takes a custom OpenAI- or Anthropic-compatible endpoint.
- OpenAI-compatible base URL
- https://api.subconscious.dev/v1
- Anthropic-compatible base URL
- https://api.subconscious.dev
- API key
- From subc login or the Subconscious dashboard
- Model
- subconscious/glm-5.3-marathon
- Multimodal, high-throughput model
- subconscious/deepseek-v4.1-flash-marathon
Full reference in the Subconscious API docs.
Questions
- Which Warp plans support custom inference endpoints?
- Warp's docs list custom endpoints on Free and paid individual plans and for organizations with 10 or fewer employees; larger teams need Business or Enterprise. Requests route through Warp's servers, and custom endpoints are not used by Cloud Agents or Auto models.
- Which Subconscious model should I use with Warp?
- Start with subconscious/glm-5.3-marathon, an open frontier coding model built for long agent runs. subconscious/deepseek-v4.1-flash-marathon is a cheaper, high-throughput alternative that also takes image input. Use the full model ID, including the subconscious/ prefix.
- Is Warp on Subconscious OpenAI- or Anthropic-compatible?
- This guide connects Warp through the OpenAI-compatible Subconscious API. Subconscious serves both formats: OpenAI Chat Completions at https://api.subconscious.dev/v1 and Anthropic Messages at https://api.subconscious.dev.
- How is Warp usage on Subconscious billed?
- Per token by default, with no commitment. Heavy users can switch to a monthly token plan: a fixed price for a large daily token allowance shared across your team.
More terminal and IDE agents
See allMarathon
subcSubconscious's own terminal coding agent, written in Rust.
Claude Code
subcAnthropic's terminal coding agent.
Codex
subcOpenAI's open-source terminal coding agent.
OpenCode
subcOpen-source terminal coding agent.
Pi
subcMinimal, extensible terminal coding agent.
DeepSeek Harness
subcDeepSeek's coding agent with a local web UI.