How to use Hermes Agent with Subconscious

Let your AI do the setup. Paste this into Claude Code, Cursor, ChatGPT or any assistant.

Set up Hermes Agent to use Subconscious as its model provider.

Subconscious is an inference API for agents. It serves open models through OpenAI- and Anthropic-compatible endpoints, billed per token.
- OpenAI-compatible base URL: https://api.subconscious.dev/v1
- Anthropic-compatible base URL: https://api.subconscious.dev
- Model: subconscious/glm-5.3-marathon
- API key: read it from the SUBCONSCIOUS_API_KEY environment variable. Never commit the key or paste it into project files. If it is not set, ask me to create one at https://platform.subconscious.dev.

Follow these steps. Run commands and edit files yourself. For any step in a settings screen, tell me exactly what to click and what to enter.

1. Install Hermes Agent
Reload your shell afterwards (source ~/.zshrc or ~/.bashrc) so the hermes command is on your PATH.
```bash
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
```

2. Save your Subconscious API key
Create a key in the Subconscious platform and add it to the Hermes env file.
```bash
SUBCONSCIOUS_API_KEY=your_key
```

3. Add Subconscious as a provider
Hermes needs at least 64k tokens of context and reads it from the model list when the API does not report one, so set context_length explicitly. 5,000,000 matches what the Subconscious CLI uses.
```yaml
providers:
  subconscious:
    api: https://api.subconscious.dev/v1
    key_env: SUBCONSCIOUS_API_KEY
    transport: chat_completions
    models:
      subconscious/glm-5.3-marathon:
        context_length: 5000000
model:
  default: subconscious/glm-5.3-marathon
  provider: custom:subconscious
```

4. Start Hermes
Run hermes. Inside a session, /model switches between configured providers.
```bash
hermes
```

Full guide: https://www.subconscious.dev/agents/agent-harnesses/hermes.md

When you are done, confirm the tool is using subconscious/glm-5.3-marathon, then summarize what you changed.

To use Hermes Agent with Subconscious, put your key in ~/.hermes/.env, add a subconscious provider to ~/.hermes/config.yaml with API https://api.subconscious.dev/v1 and the chat_completions transport, and set the default model to subconscious/glm-5.3-marathon on custom:subconscious. Run hermes and it uses Subconscious.

Last verified against the Hermes Agent docs on

Before you start

  • macOS, Linux or WSL
  • A Subconscious account and API key

Hermes Agent docs · Source on GitHub

Set up Hermes Agent

  1. Step 1: Install Hermes Agent

    Reload your shell afterwards (source ~/.zshrc or ~/.bashrc) so the hermes command is on your PATH.

    Terminal
    curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
  2. Step 2: Save your Subconscious API key

    Create a key in the Subconscious platform and add it to the Hermes env file.

    ~/.hermes/.env
    SUBCONSCIOUS_API_KEY=your_key
  3. Step 3: Add Subconscious as a provider

    Hermes needs at least 64k tokens of context and reads it from the model list when the API does not report one, so set context_length explicitly. 5,000,000 matches what the Subconscious CLI uses.

    ~/.hermes/config.yaml
    providers:
      subconscious:
        api: https://api.subconscious.dev/v1
        key_env: SUBCONSCIOUS_API_KEY
        transport: chat_completions
        models:
          subconscious/glm-5.3-marathon:
            context_length: 5000000
    model:
      default: subconscious/glm-5.3-marathon
      provider: custom:subconscious
  4. Step 4: Start Hermes

    Run hermes. Inside a session, /model switches between configured providers.

    Terminal
    hermes

Why run Hermes Agent on Subconscious

Context that keeps going
The OrangeLine runtime compresses the parts of context that stopped mattering, on the GPU, while the run continues. That gives a 5M+ token context window with no stop-and-compact pause.
Long sessions stay fast
2x faster task completion on long tasks than standard open-model inference, especially past 200k tokens of context.
Lower cost on long tasks
50% to 80% lower cost on long tasks, because the runtime processes far fewer tokens. Cached input bills at a steep discount, and agentic coding sees cache hit rates above 95%.
Your data stays yours
Prompts, inputs and outputs are not recorded, and nothing you send is used to train a model.

Connection details

For any client that takes a custom OpenAI- or Anthropic-compatible endpoint.

OpenAI-compatible base URL
https://api.subconscious.dev/v1
Anthropic-compatible base URL
https://api.subconscious.dev
API key
From subc login or the Subconscious dashboard
Model
subconscious/glm-5.3-marathon
Multimodal, high-throughput model
subconscious/deepseek-v4.1-flash-marathon

Full reference in the Subconscious API docs.

Questions

Can I set up Hermes Agent without editing config.yaml?
Yes. Run hermes model, choose Custom endpoint, and enter the Subconscious base URL, your API key and the model name when prompted.
Can Hermes Agent use the Anthropic-compatible Subconscious API?
Yes. Set api to https://api.subconscious.dev and transport to anthropic_messages on the subconscious provider.
Which Subconscious model should I use with Hermes Agent?
Start with subconscious/glm-5.3-marathon, an open frontier coding model built for long agent runs. subconscious/deepseek-v4.1-flash-marathon is a cheaper, high-throughput alternative that also takes image input. Use the full model ID, including the subconscious/ prefix.
Is Hermes Agent on Subconscious OpenAI- or Anthropic-compatible?
This guide connects Hermes Agent through the OpenAI-compatible Subconscious API. Subconscious serves both formats: OpenAI Chat Completions at https://api.subconscious.dev/v1 and Anthropic Messages at https://api.subconscious.dev.
How is Hermes Agent usage on Subconscious billed?
Per token by default, with no commitment. Heavy users can switch to a monthly token plan: a fixed price for a large daily token allowance shared across your team.

More agent harnesses

See all