Customer Stories | Together AI
The teams shipping AI at production scale
Together AI is the end-to-end platform trusted for reliability, leading price economics, and research-backed performance. Hear from the teams building on the AI Native Cloud.
Customer Success Stories
How Cursor partnered with Together AI to deliver real-time, low-latency inference at scale
- Category: Inference, GPU clusters, RESEARCH • Enterprise
How Decagon Engineered Sub-Second Voice AI with Together AI
- Category: Inference, GPU clusters, RESEARCH
- Cost Reduction: 6x per turn vs. GPT-5 mini
Faster Inference with Vercept
- Factor: 11x faster inference
.svg)
- Time-to-first-token: <500ms
- Story: How Deep Cogito trained and deployed frontier reasoning models on Together AI
Yutori's Browser-Use AI Agents
- Speed Factor: 2x faster inference
Cartesia's Real-Time Voice AI
- Model Latency: 90ms
XY.AI Labs' EOB Parsers
- Accuracy: 87%
Decagon's Voice AI
- Cost Per Turn: 6× reduction
Cursor's Inference Impact
- GPU Count: 72 GPUs in NVL72 topology
- Story: Real-time, low-latency inference at scale
Scaled Cognition's APT-1 Training
- Time Saved: ~3 months
Runware’s API Scaling
- Performance: 5-10× vs. competitors in generative video & image APIs
Latent Health's Clinical AI Development
- Training Cost Reduction: 7×
The Washington Post's AI Independence
- Response Time: 2 seconds
Slingshot AI's Mental Health AI
- Training Frequency Improvement: 3x
HeroUI's Speed to Launch
- Launch Speed: 10× faster
Hedra's Cost Savings
- Savings: 60%
Arcee AI's Inference Flexibility
- TTFT: 95% faster
LegionEdge's Prototyping Speed
- Launch Time: 3 months faster
Thai Language Models
- Cloud Savings: 50%
Zomato's AI Customer Support Bot
- CSAT Score Improvement: 2x
Dippy AI's Token Performance
- Median TTFT: 0.4s
Testimonials
"Together AI offers optimized performance at scale, and at a lower cost than closed-source providers – all while maintaining strict privacy standards."
— Vineet Khosla, CTO, The Washington Post
"We’ve been thoroughly impressed with Together. They delivered a 2x reduction in latency and cut our costs by approximately a third."
— Caiming Xiong, VP, Salesforce AI Research
"Together GPU Clusters provided a combination of amazing training performance, expert support, and the ability to scale to meet our rapid growth to help us serve our growing community of AI creators."
— Demi Guo, CEO, Pika
"Together AI provides the performance and reliability we need for real-time, high-quality image and video generation at scale. We value that Together AI is much more than an infrastructure provider — they're a true innovation partner, enabling us to push creative boundaries without compromise."
— Victor Perez, Co-Founder, Krea
"Together AI’s infrastructure has the capacity to soak up our viral moments without breaking a sweat. During major traffic surges, Dedicated Container Inference scales seamlessly while maintaining performance."
— Terrance Wang, Founding ML Engineer, Hedra
Why Together AI?
Cost Reduction
- Factor: 6x with Together Inference
Workload at Scale
- Configuration: 72 GPUs on Together Managed Clusters
Lower Latency
- Factor: 95× with Together Inference