# Together AI vs Modal

> Modal sells serverless GPU containers for your own code and weights. Together sells hosted open models by the token, plus fine-tuning and clusters.

Canonical: https://www.subconscious.dev/compare/together-ai-vs-modal · By The Subconscious Team · Updated September 30, 2026

## How they compare

The core difference is who owns the serving stack. Modal has no model catalog and no per-token price. A developer decorates a Python function with a GPU type, and Modal builds, schedules and autoscales the container, billing per second from $0.59 an hour for a T4 to $3.95 for an H100 at list. Together hosts thirty-plus open text models plus media and embeddings behind an OpenAI-compatible endpoint, so there is nothing to package. Together also runs managed fine-tuning and reserved clusters, with H100s from $3.19 an hour reserved.

Modal suits custom or non-LLM work: OCR, embeddings, transcription, fine-tunes with bespoke serving code, and spiky jobs that scale to zero. Its free Starter plan renews $30 in credits monthly, where Together has no free tier. Watch Modal's production math, though: non-preemptible US capacity runs about 3.75x list, near $14.81 for an H100, and keeping containers warm erases the serverless savings. For standard open LLMs at steady volume, Together's per-token pricing is simpler. For bring-your-own code, Modal is the better fit.

## What each one does

### Together AI

Together AI is the broadest open-model platform in the category. One bill covers per-token serverless inference, batch at up to 50% off, provisioned throughput with a 99% SLA, dedicated deployments, raw GPU clusters, managed fine-tuning and code sandboxes for agents. The text catalog runs past thirty open models, including DeepSeek V4, Kimi K3, GLM 5.2, Qwen 3.8 and MiniMax M3, plus image, video, speech and embedding models. Token prices sit at parity with Fireworks and Baseten.

### Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

## Which is best, and when

### Choose Together AI for

- Popular open LLMs with no serving code to write
- Steady production traffic billed per token
- Managed SFT without writing training infrastructure

### Choose Modal for

- Custom models and pipelines in plain Python
- Bursty GPU jobs that should scale to zero
- Free monthly credits for early experiments

## At a glance

| Attribute | Together AI | Modal |
|---|---|---|
| Model access | Open weights | Bring your own weights |
| Flagship models | Kimi K3, DeepSeek V4, GLM 5.2, Qwen 3.8 | None hosted |
| Speed | 0.99s TTFT on DeepSeek V4 Pro | ~1s container boot |
| Price | Parity with Fireworks and Baseten | Per second; H100 $3.95/hr list |
| Customization | LoRA and full SFT; RL in beta | Run any training code |
| Deployment | Serverless, dedicated, GPU clusters | Serverless GPU containers |
| Long context | 512K on DeepSeek V4 Pro | Depends on the model you deploy |

## FAQ

### What is the difference between Together AI and Modal?

Modal sells serverless GPU containers for your own code and weights. Together sells hosted open models by the token, plus fine-tuning and clusters.

### When should I choose Together AI over Modal?

Popular open LLMs with no serving code to write; Steady production traffic billed per token; Managed SFT without writing training infrastructure.

### When should I choose Modal over Together AI?

Custom models and pipelines in plain Python; Bursty GPU jobs that should scale to zero; Free monthly credits for early experiments.

### Is Together AI or Modal cheaper?

Together AI: Parity with Fireworks and Baseten. Modal: Per second; H100 $3.95/hr list. The cheaper choice depends on the model and workload.

### Which has more context, Together AI or Modal?

Together AI: 512K on DeepSeek V4 Pro. Modal: Depends on the model you deploy.

## Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious: [Subconscious vs Together AI](https://www.subconscious.dev/compare/subconscious-vs-together-ai.md), [Subconscious vs Modal](https://www.subconscious.dev/compare/subconscious-vs-modal.md).

Full profiles: [Together AI](https://www.subconscious.dev/providers/together-ai.md), [Modal](https://www.subconscious.dev/providers/modal.md).
