Serverless Inference | Together AI
Serverless Inference
The fastest way to run open ‑ source models on demand
High-performance inference, powered by our in-house research. No infrastructure to manage, no long-term commitments.
Serverless inference on Together AI
Access all the top open-source models in one place.
- Up to 2.75x faster inference
- Every modality, one API
- Built on cutting-edge systems research
Build with leading models
Explore top-performing models across text, image, video, code, and voice.
Browse models Deploy own model
| Model Name | Description |
|---|---|
| DeepSeek V4 Pro | Chat Model |
| MiniMax M3 | Chat Model |
| GLM-5.2 | Chat Model |
| ByteDance Seedance 2.0 | Video Model |
| Kimi K3 | Chat Model |
| Inkling Small | Chat Model |
Key capabilities, purpose built for AI natives
Call any leading open-weight model through a single API, with nothing to deploy or manage, and pay only for the tokens you use.
- Adaptive speculative decoding: Reduce end-to-end latency by predicting and validating multiple tokens per step.
- OpenAI-compatible API: Same API, better models. No code changes required.
- Quantization without compromise: Run quantized models at full quality, improving speed without sacrificing output accuracy.
Deployment options
Run models using different deployment options depending on latency needs, traffic patterns, and infrastructure control.
- Serverless Inference: A fully managed real-time or batch inference API.
- Provisioned Throughput: Reserved token capacity with SLA guarantees.
- Dedicated Model Inference: An inference endpoint backed by reserved, isolated compute resources.
- Dedicated Container Inference: Run inference with your own engine and model on fully-managed, scalable infrastructure.
Production-grade security and data privacy
We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.
Customers running inference in production
Vercept
"Together has helped us deploy VyUI, our state-of-the-art computer AI model."
The Washington Post
"Together AI offers optimized performance at scale, and at a lower cost than closed-source providers."
Salesforce AI Research
"We’ve been thoroughly impressed with Together. They delivered a 2x reduction in latency and cut our costs by approximately a third."
Run inference with Serverless Endpoints
High-performance inference, powered by our in-house research.