Batch Inference | Together AI
Batch Inference
Process massive workloads asynchronously
Scale to 30 billion tokens per model with any serverless model or private deployment.
Start building now Explore the Docs
Why Batch Inference with Together AI?
Lower cost, higher limits, and predictable processing
Up to 50% cost savings
Run batch jobs at up to half the cost of our real-time API for most serverless models. Process millions of requests economically without sacrificing quality or speed.
- 30B enqueued tokens per model
Run massive batch jobs that scale to 30 billion enqueued tokens per model per user. Need more? We'll customize limits for your specific use case.
- <24h processing time SLA
Jobs consistently finish well under 24 hours — often within just hours. Submit and forget while we handle the scale.
Key capabilities, purpose built for AI natives
Run massive asynchronous inference jobs against a serverless model or dedicated model inference.
- Universal model access
- Any serverless model
- Dedicated deployments
Run batch jobs across any serverless model or private deployment — no limitations on model choice.
- Up and running in minutes
- Launch in three steps
- No DevOps required
- No orchestration required
Launch massive inference jobs simply by uploading a JSONL file. Start processing batches with just a few clicks. No orchestration or monitoring setup required.
- 50% off many top models
- DeepSeek, Llama, Qwen & more
- No minimum volume
Run batch jobs at up to half the cost of our real-time API for most serverless models. Process millions of requests without sacrificing quality or speed.
Research-optimized, best-in-class performance
We achieved up to 2x faster serverless inference for the most demanding LLMs, including GPT-OSS, Qwen, Kimi, and DeepSeek.
- GPT-OSS-20B
- Qwen3 235B 2507
- Kimi K2 0905
- DeepSeek V3.1
- DeepSeek R1 0528
Together AI vs other providers
- Nearly 2x faster serverless inference performance for GPT-OSS-20B.
- 2.75x faster for Qwen3 235B 2507.
- 65% faster for Kimi K2 0905.
- 10% faster for DeepSeek V3.1.
- 13% faster for DeepSeek R1 0528.
Learn more about fastest inference for the top open-source models.
Deployment options
Run models using different deployment options depending on latency needs, traffic patterns, and infrastructure control.
- Serverless Inference: A fully managed real-time or batch inference API with access to dozens of the most popular AI models. Best for variable or unpredictable traffic.
- Provisioned Throughput: Reserved token capacity with SLA guarantees. Best for production workloads.
- Dedicated Model Inference: An inference endpoint backed by reserved, isolated compute resources. Best for predictable or steady traffic.
- Dedicated Container Inference: Run inference with your own engine and model on fully-managed, scalable infrastructure. Best for generative media models.
Production-grade security and data privacy
We take security and compliance seriously, with strict data privacy controls to keep your information protected.
Customers Running Inference in Production
- 30B Enqueued tokens
- 24h SLA
"We rely on the Batch Inference API to process very large amounts of requests. The high rate limits—up to 30B enqueued tokens—let us run massive experiments without bottlenecks, and jobs consistently finish well under the 24-hour SLA, often within just hours. It’s transformed the pace at which we can test and iterate."
Volodymyr Kuleshov
Co-Founder, Inception Labs