vs

Modal vs xAI

Modal rents serverless GPUs by the second for models you bring. xAI sells closed Grok models with live X data through its own API. Compute platform versus model vendor.

By The Subconscious Team · Updated

Modal vs xAI: key differences

Modal and xAI sell different things. xAI is a closed lab: Grok 4.6 is its flagship at $2 in and $6 out per million tokens with a 500K window, and older Grok 4.20 and 4.3 keep a 1M window at $1.25 in and $2.50 out. Its distinct feature is server-side Web Search and X Search, which pull live posts from X. Modal has no models and no per-token price. A developer decorates a Python function with the GPU it needs, and Modal builds, schedules, autoscales and scales the container to zero, billing per second from $0.59 an hour for a T4 to $3.95 for an H100 at list.

So the question is whether you want a finished model or a place to run your own. Grok fits news, market and social-sentiment agents that need fresh data, and cost-sensitive reasoning where cheap output tokens help. Watch the 200K line, past which the whole request bills at double. Modal fits custom models, fine-tunes, embeddings, OCR and batch jobs. Its catches are cold starts from loading weights and non-preemptible US production at about 3.75x list, near $14.81 an hour for an H100.

What Modal and xAI do

Modal

Modal is serverless compute with GPUs attached. A developer decorates a Python function with the hardware it needs, such as gpu="H100", and Modal builds the container, schedules it, autoscales it and scales it back to zero. Billing runs per second with no minimum increment, from $0.59 an hour for a T4 to $3.95 for an H100 at list. Containers can hold up to 8 GPUs across T4 through B300.

Example models: none hosted by default; teams deploy their own, such as open LLMs on vLLM or Whisper

Full Modal profile

xAI

xAI sells the Grok models through its own API. Grok 4.6 is the current flagship and xAI tells developers to use it for everything outside audio, image and video, code included. It has a 500K context window and costs $2 in and $6 out per million tokens under 200K prompt tokens. Older Grok 4.20 and 4.3 models keep a 1M window at $1.25 in and $2.50 out, which is aggressive for that capability class.

Example models: Grok 4.6, Grok 4.20

Full xAI profile

Should you choose Modal or xAI?

Modal

Choose Modal for

  • Serving your own fine-tuned or custom model.
  • Bursty GPU jobs billed by the second.
  • Python teams shipping a GPU service quickly.

xAI

Choose xAI for

  • Agents that need live data from X.
  • Cheap output tokens on a closed reasoning model.
  • Up to 1M context on Grok 4.20 and 4.3.

Modal vs xAI at a glance

AttributeModalxAI
Model accessBring your own weightsClosed
Flagship modelsNone hostedGrok 4.6, Grok 4.20, grok-build
Speed~1s container boot~54 tok/s on Grok 4.6
PricePer second; H100 $3.95/hr list$2 in, $6 out (Grok 4.6); 2x past 200K
CustomizationRun any training codeUnknown
DeploymentServerless GPU containersFirst-party API
Long contextDepends on the model you deploy500K (4.6), 1M (4.20, 4.3)

Frequently asked questions

What is the difference between Modal and xAI?

Modal rents serverless GPUs by the second for models you bring. xAI sells closed Grok models with live X data through its own API. Compute platform versus model vendor.

When should I choose Modal over xAI?

Serving your own fine-tuned or custom model; Bursty GPU jobs billed by the second; Python teams shipping a GPU service quickly.

When should I choose xAI over Modal?

Agents that need live data from X; Cheap output tokens on a closed reasoning model; Up to 1M context on Grok 4.20 and 4.3.

Is Modal or xAI cheaper?

Modal: Per second; H100 $3.95/hr list. xAI: $2 in, $6 out (Grok 4.6); 2x past 200K. The cheaper choice depends on the model and workload.

Which has more context, Modal or xAI?

Modal: Depends on the model you deploy. xAI: 500K (4.6), 1M (4.20, 4.3).

Related comparisons

Running long-horizon agents?

If your agents run past 200K tokens, compare both against Subconscious. Our inference stack treats a long-horizon trace as the primary workload, so speed, cost, and accuracy hold up deep into the trace.