How to use Hermes Agent with Subconscious
Let your AI do the setup. Paste this into Claude Code, Cursor, ChatGPT or any assistant.
Set up Hermes Agent to use Subconscious as its model provider.
Subconscious is an inference API for agents. It serves open models through OpenAI- and Anthropic-compatible endpoints, billed per token.
- OpenAI-compatible base URL: https://api.subconscious.dev/v1
- Anthropic-compatible base URL: https://api.subconscious.dev
- Model: subconscious/glm-5.3-marathon
- API key: read it from the SUBCONSCIOUS_API_KEY environment variable. Never commit the key or paste it into project files. If it is not set, ask me to create one at https://platform.subconscious.dev.
Follow these steps. Run commands and edit files yourself. For any step in a settings screen, tell me exactly what to click and what to enter.
1. Install Hermes Agent
Reload your shell afterwards (source ~/.zshrc or ~/.bashrc) so the hermes command is on your PATH.
```bash
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
```
2. Save your Subconscious API key
Create a key in the Subconscious platform and add it to the Hermes env file.
```bash
SUBCONSCIOUS_API_KEY=your_key
```
3. Add Subconscious as a provider
Hermes needs at least 64k tokens of context and reads it from the model list when the API does not report one, so set context_length explicitly. 5,000,000 matches what the Subconscious CLI uses.
```yaml
providers:
subconscious:
api: https://api.subconscious.dev/v1
key_env: SUBCONSCIOUS_API_KEY
transport: chat_completions
models:
subconscious/glm-5.3-marathon:
context_length: 5000000
model:
default: subconscious/glm-5.3-marathon
provider: custom:subconscious
```
4. Start Hermes
Run hermes. Inside a session, /model switches between configured providers.
```bash
hermes
```
Full guide: https://www.subconscious.dev/agents/agent-harnesses/hermes.md
When you are done, confirm the tool is using subconscious/glm-5.3-marathon, then summarize what you changed.To use Hermes Agent with Subconscious, put your key in ~/.hermes/.env, add a subconscious provider to ~/.hermes/config.yaml with API https://api.subconscious.dev/v1 and the chat_completions transport, and set the default model to subconscious/glm-5.3-marathon on custom:subconscious. Run hermes and it uses Subconscious.
Last verified against the Hermes Agent docs on
Before you start
- macOS, Linux or WSL
- A Subconscious account and API key
Set up Hermes Agent
Step 1: Install Hermes Agent
Reload your shell afterwards (source ~/.zshrc or ~/.bashrc) so the hermes command is on your PATH.
Terminalcurl -fsSL https://hermes-agent.nousresearch.com/install.sh | bashStep 2: Save your Subconscious API key
Create a key in the Subconscious platform and add it to the Hermes env file.
~/.hermes/.envSUBCONSCIOUS_API_KEY=your_keyStep 3: Add Subconscious as a provider
Hermes needs at least 64k tokens of context and reads it from the model list when the API does not report one, so set context_length explicitly. 5,000,000 matches what the Subconscious CLI uses.
~/.hermes/config.yamlproviders: subconscious: api: https://api.subconscious.dev/v1 key_env: SUBCONSCIOUS_API_KEY transport: chat_completions models: subconscious/glm-5.3-marathon: context_length: 5000000 model: default: subconscious/glm-5.3-marathon provider: custom:subconsciousStep 4: Start Hermes
Run hermes. Inside a session, /model switches between configured providers.
Terminalhermes
Why run Hermes Agent on Subconscious
- Context that keeps going
- The OrangeLine runtime compresses the parts of context that stopped mattering, on the GPU, while the run continues. That gives a 5M+ token context window with no stop-and-compact pause.
- Long sessions stay fast
- 2x faster task completion on long tasks than standard open-model inference, especially past 200k tokens of context.
- Lower cost on long tasks
- 50% to 80% lower cost on long tasks, because the runtime processes far fewer tokens. Cached input bills at a steep discount, and agentic coding sees cache hit rates above 95%.
- Your data stays yours
- Prompts, inputs and outputs are not recorded, and nothing you send is used to train a model.
Connection details
For any client that takes a custom OpenAI- or Anthropic-compatible endpoint.
- OpenAI-compatible base URL
- https://api.subconscious.dev/v1
- Anthropic-compatible base URL
- https://api.subconscious.dev
- API key
- From subc login or the Subconscious dashboard
- Model
- subconscious/glm-5.3-marathon
- Multimodal, high-throughput model
- subconscious/deepseek-v4.1-flash-marathon
Full reference in the Subconscious API docs.
Questions
- Can I set up Hermes Agent without editing config.yaml?
- Yes. Run hermes model, choose Custom endpoint, and enter the Subconscious base URL, your API key and the model name when prompted.
- Can Hermes Agent use the Anthropic-compatible Subconscious API?
- Yes. Set api to https://api.subconscious.dev and transport to anthropic_messages on the subconscious provider.
- Which Subconscious model should I use with Hermes Agent?
- Start with subconscious/glm-5.3-marathon, an open frontier coding model built for long agent runs. subconscious/deepseek-v4.1-flash-marathon is a cheaper, high-throughput alternative that also takes image input. Use the full model ID, including the subconscious/ prefix.
- Is Hermes Agent on Subconscious OpenAI- or Anthropic-compatible?
- This guide connects Hermes Agent through the OpenAI-compatible Subconscious API. Subconscious serves both formats: OpenAI Chat Completions at https://api.subconscious.dev/v1 and Anthropic Messages at https://api.subconscious.dev.
- How is Hermes Agent usage on Subconscious billed?
- Per token by default, with no commitment. Heavy users can switch to a monthly token plan: a fixed price for a large daily token allowance shared across your team.