How to use Temporal with Subconscious
Let your AI do the setup. Paste this into Claude Code, Cursor, ChatGPT or any assistant.
Make Subconscious the model provider in my Temporal project.
Subconscious is an inference API for agents. It serves open models through OpenAI- and Anthropic-compatible endpoints, billed per token.
- OpenAI-compatible base URL: https://api.subconscious.dev/v1
- Anthropic-compatible base URL: https://api.subconscious.dev
- Model: subconscious/glm-5.3-marathon
- API key: read it from the SUBCONSCIOUS_API_KEY environment variable. Never commit the key or paste it into project files. If it is not set, ask me to create one at https://platform.subconscious.dev.
Follow these steps. Run commands and edit files yourself. If I already have agent code, change its model setup to match the example instead of adding a new file.
1. Install Temporal
Install the framework and its OpenAI-compatible client.
```bash
pip install temporalio openai
```
2. Create a Subconscious API key
Sign in at https://platform.subconscious.dev, create an API key, and export it so the tool can read it.
```bash
export SUBCONSCIOUS_API_KEY="your_key"
```
3. Point it at Subconscious
Model calls belong in Activities, never in workflow code. max_retries=0 leaves retries to Temporal, and openai is imported through the workflow sandbox.
```python
import asyncio, os
from datetime import timedelta
from temporalio import activity, workflow
from temporalio.client import Client
from temporalio.worker import Worker
with workflow.unsafe.imports_passed_through():
from openai import AsyncOpenAI
@activity.defn
async def ask_llm(prompt: str) -> str:
client = AsyncOpenAI(base_url="https://api.subconscious.dev/v1",
api_key=os.environ["SUBCONSCIOUS_API_KEY"], max_retries=0)
resp = await client.chat.completions.create(
model="subconscious/glm-5.3-marathon",
messages=[{"role": "user", "content": prompt}])
return resp.choices[0].message.content or ""
@workflow.defn
class AgentWorkflow:
@workflow.run
async def run(self, prompt: str) -> str:
return await workflow.execute_activity(
ask_llm, prompt, start_to_close_timeout=timedelta(minutes=10))
async def main():
client = await Client.connect("localhost:7233")
async with Worker(client, task_queue="subconscious",
workflows=[AgentWorkflow], activities=[ask_llm]):
print(await client.execute_workflow(AgentWorkflow.run, "Hello!",
id="subconscious-demo", task_queue="subconscious"))
if __name__ == "__main__":
asyncio.run(main())
```
4. Run it
Start a local Temporal server, then run the worker and workflow.
```bash
temporal server start-dev
# in a second terminal
python agent.py
```
Full guide: https://www.subconscious.dev/agents/frameworks/temporal.md
When you are done, confirm the tool is using subconscious/glm-5.3-marathon, then summarize what you changed.To use Temporal with Subconscious, call the Subconscious API from an Activity using the OpenAI SDK with base_url="https://api.subconscious.dev/v1" and model="subconscious/glm-5.3-marathon", and run that Activity from your Workflow. Temporal then retries and resumes model calls durably, which suits long agent runs.
Last verified against the Temporal docs on
Before you start
- Python 3.10+ and the Temporal CLI
- A Subconscious account and API key
Set up Temporal
Step 1: Install Temporal
Install the framework and its OpenAI-compatible client.
Terminalpip install temporalio openaiStep 2: Create a Subconscious API key
Sign in at https://platform.subconscious.dev, create an API key, and export it so the tool can read it.
Terminalexport SUBCONSCIOUS_API_KEY="your_key"Step 3: Point it at Subconscious
Model calls belong in Activities, never in workflow code. max_retries=0 leaves retries to Temporal, and openai is imported through the workflow sandbox.
agent.pyimport asyncio, os from datetime import timedelta from temporalio import activity, workflow from temporalio.client import Client from temporalio.worker import Worker with workflow.unsafe.imports_passed_through(): from openai import AsyncOpenAI @activity.defn async def ask_llm(prompt: str) -> str: client = AsyncOpenAI(base_url="https://api.subconscious.dev/v1", api_key=os.environ["SUBCONSCIOUS_API_KEY"], max_retries=0) resp = await client.chat.completions.create( model="subconscious/glm-5.3-marathon", messages=[{"role": "user", "content": prompt}]) return resp.choices[0].message.content or "" @workflow.defn class AgentWorkflow: @workflow.run async def run(self, prompt: str) -> str: return await workflow.execute_activity( ask_llm, prompt, start_to_close_timeout=timedelta(minutes=10)) async def main(): client = await Client.connect("localhost:7233") async with Worker(client, task_queue="subconscious", workflows=[AgentWorkflow], activities=[ask_llm]): print(await client.execute_workflow(AgentWorkflow.run, "Hello!", id="subconscious-demo", task_queue="subconscious")) if __name__ == "__main__": asyncio.run(main())Step 4: Run it
Start a local Temporal server, then run the worker and workflow.
Terminaltemporal server start-dev # in a second terminal python agent.py
Why run Temporal on Subconscious
- No new SDK
- Subconscious speaks OpenAI Chat Completions and Anthropic Messages, so the client you already use works with a base URL change.
- Pay per token
- No plan and no commitment: sign up, get an API key and pay for the tokens your agents use.
- Long sessions stay fast
- 2x faster task completion on long tasks than standard open-model inference, especially past 200k tokens of context.
- Context that keeps going
- The OrangeLine runtime compresses the parts of context that stopped mattering, on the GPU, while the run continues. That gives a 5M+ token context window with no stop-and-compact pause.
Connection details
For any client that takes a custom OpenAI- or Anthropic-compatible endpoint.
- OpenAI-compatible base URL
- https://api.subconscious.dev/v1
- Anthropic-compatible base URL
- https://api.subconscious.dev
- API key
- From subc login or the Subconscious dashboard
- Model
- subconscious/glm-5.3-marathon
- Multimodal, high-throughput model
- subconscious/deepseek-v4.1-flash-marathon
Full reference in the Subconscious API docs.
Questions
- How long should the Temporal activity timeout be for agent runs?
- Long enough for your slowest model call. Long agent runs can take minutes, so set start_to_close_timeout generously and consider heartbeating from the activity.
- Which Subconscious model should I use with Temporal?
- Start with subconscious/glm-5.3-marathon, an open frontier coding model built for long agent runs. subconscious/deepseek-v4.1-flash-marathon is a cheaper, high-throughput alternative that also takes image input. Use the full model ID, including the subconscious/ prefix.
- Is Temporal on Subconscious OpenAI- or Anthropic-compatible?
- This guide connects Temporal through the OpenAI-compatible Subconscious API. Subconscious serves both formats: OpenAI Chat Completions at https://api.subconscious.dev/v1 and Anthropic Messages at https://api.subconscious.dev.
- How is Temporal usage on Subconscious billed?
- Per token by default, with no commitment. Heavy users can switch to a monthly token plan: a fixed price for a large daily token allowance shared across your team.
More frameworks
See allLangChain
Open-source framework for building LLM apps and agents.
LangGraph
Low-level orchestration for stateful, long-running agent graphs.
Deep Agents
LangChain agent harness with planning, a filesystem and subagents built in.
eve
Vercel's open-source, filesystem-first TypeScript framework for durable agents.
Mastra
TypeScript framework for agents, workflows, RAG, memory and evals.
AI SDK
Vercel's TypeScript toolkit for building AI apps across model providers.