vs

Together AI vs xAI

xAI sells closed Grok models with native X data. Together hosts open weights you can fine-tune and move. Different bets on control versus a single lab's roadmap.

By The Subconscious Team · Updated

Together AI vs xAI: key differences

xAI is a closed lab with one API. Grok 4.6 is the flagship at $2 in and $6 out with a 500K window, and older Grok 4.20 and 4.3 keep 1M context at $1.25 in and $2.50 out. Its unique asset is live data: server-side Web Search and X Search pull posts and current events straight from X. Together has no first-party models. It hosts open weights like Kimi K3, DeepSeek V4 and GLM 5.2, which you can fine-tune with LoRA or full SFT and run on dedicated hardware or reserved clusters. Grok weights stay with xAI.

Long prompts are a real cost factor on xAI: once a prompt reaches 200K tokens, the whole request bills at double. On Together, batch cuts up to 50% for work that can wait. xAI also has a thinner enterprise footprint and fewer cloud marketplace options. Choose xAI for social listening, news and market agents where fresh X data is the product. Choose Together when owning the weights, training on proprietary data or avoiding a single vendor matters more.

What Together AI and xAI do

Together AI

Together AI is the broadest open-model platform in the category. One bill covers per-token serverless inference, batch at up to 50% off, provisioned throughput with a 99% SLA, dedicated deployments, raw GPU clusters, managed fine-tuning and code sandboxes for agents. The text catalog runs past thirty open models, including DeepSeek V4, Kimi K3, GLM 5.2, Qwen 3.8 and MiniMax M3, plus image, video, speech and embedding models. Token prices sit at parity with Fireworks and Baseten.

Example models: Kimi K3, DeepSeek V4 Pro

Full Together AI profile

xAI

xAI sells the Grok models through its own API. Grok 4.6 is the current flagship and xAI tells developers to use it for everything outside audio, image and video, code included. It has a 500K context window and costs $2 in and $6 out per million tokens under 200K prompt tokens. Older Grok 4.20 and 4.3 models keep a 1M window at $1.25 in and $2.50 out, which is aggressive for that capability class.

Example models: Grok 4.6, Grok 4.20

Full xAI profile

Should you choose Together AI or xAI?

Together AI

Choose Together AI for

  • Fine-tuning open weights on proprietary data
  • Avoiding lock-in to one lab's closed models
  • Batch jobs at up to 50% off

xAI

Choose xAI for

  • Agents that need live posts and events from X
  • Cheap output tokens on a closed reasoning model
  • 1M context on Grok 4.20 at $1.25 in

Together AI vs xAI at a glance

AttributeTogether AIxAI
Model accessOpen weightsClosed
Flagship modelsKimi K3, DeepSeek V4, GLM 5.2, Qwen 3.8Grok 4.6, Grok 4.20, grok-build
Speed0.99s TTFT on DeepSeek V4 Pro~54 tok/s on Grok 4.6
PriceParity with Fireworks and Baseten$2 in, $6 out (Grok 4.6); 2x past 200K
CustomizationLoRA and full SFT; RL in betaUnknown
DeploymentServerless, dedicated, GPU clustersFirst-party API
Long context512K on DeepSeek V4 Pro500K (4.6), 1M (4.20, 4.3)

Frequently asked questions

What is the difference between Together AI and xAI?

xAI sells closed Grok models with native X data. Together hosts open weights you can fine-tune and move. Different bets on control versus a single lab's roadmap.

When should I choose Together AI over xAI?

Fine-tuning open weights on proprietary data; Avoiding lock-in to one lab's closed models; Batch jobs at up to 50% off.

When should I choose xAI over Together AI?

Agents that need live posts and events from X; Cheap output tokens on a closed reasoning model; 1M context on Grok 4.20 at $1.25 in.

Is Together AI or xAI cheaper?

Together AI: Parity with Fireworks and Baseten. xAI: $2 in, $6 out (Grok 4.6); 2x past 200K. The cheaper choice depends on the model and workload.

Which has more context, Together AI or xAI?

Together AI: 512K on DeepSeek V4 Pro. xAI: 500K (4.6), 1M (4.20, 4.3).

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.