# The teams shipping AI at production scale

Together AI is the end-to-end platform trusted for reliability, leading price economics, and research-backed performance. Hear from the teams building on the AI Native Cloud.

## Customer Success Stories

### How Cursor partnered with Together AI to deliver real-time, low-latency inference at scale
- **Category**: Inference, GPU clusters, RESEARCH  •  Enterprise

### How Decagon Engineered Sub-Second Voice AI with Together AI
- **Category**: Inference, GPU clusters, RESEARCH
- **Cost Reduction**: 6x per turn vs. GPT-5 mini

### Faster Inference with Vercept
- **Factor**: 11x faster inference

.svg)
- **Time-to-first-token**: <500ms
- **Story**: How Deep Cogito trained and deployed frontier reasoning models on Together AI

### Yutori's Browser-Use AI Agents
- **Speed Factor**: 2x faster inference

### Cartesia's Real-Time Voice AI
- **Model Latency**: 90ms

### XY.AI Labs' EOB Parsers
- **Accuracy**: 87%

### Decagon's Voice AI
- **Cost Per Turn**: 6× reduction

### Cursor's Inference Impact
- **GPU Count**: 72 GPUs in NVL72 topology
- **Story**: Real-time, low-latency inference at scale

### Scaled Cognition's APT-1 Training
- **Time Saved**: ~3 months

### Runware’s API Scaling
- **Performance**: 5-10× vs. competitors in generative video & image APIs

### Latent Health's Clinical AI Development
- **Training Cost Reduction**: 7×

### The Washington Post's AI Independence
- **Response Time**: 2 seconds

### Slingshot AI's Mental Health AI
- **Training Frequency Improvement**: 3x

### HeroUI's Speed to Launch
- **Launch Speed**: 10× faster

### Hedra's Cost Savings
- **Savings**: 60%

### Arcee AI's Inference Flexibility
- **TTFT**: 95% faster

### LegionEdge's Prototyping Speed
- **Launch Time**: 3 months faster

### Thai Language Models
- **Cloud Savings**: 50%

### Zomato's AI Customer Support Bot
- **CSAT Score Improvement**: 2x

### Dippy AI's Token Performance
- **Median TTFT**: 0.4s

## Testimonials

"Together AI offers optimized performance at scale, and at a lower cost than closed-source providers – all while maintaining strict privacy standards."  
— Vineet Khosla, CTO, The Washington Post

"We’ve been thoroughly impressed with Together. They delivered a 2x reduction in latency and cut our costs by approximately a third."  
— Caiming Xiong, VP, Salesforce AI Research

"Together GPU Clusters provided a combination of amazing training performance, expert support, and the ability to scale to meet our rapid growth to help us serve our growing community of AI creators."  
— Demi Guo, CEO, Pika

"Together AI provides the performance and reliability we need for real-time, high-quality image and video generation at scale. We value that Together AI is much more than an infrastructure provider — they're a true innovation partner, enabling us to push creative boundaries without compromise."  
— Victor Perez, Co-Founder, Krea

"Together AI’s infrastructure has the capacity to soak up our viral moments without breaking a sweat. During major traffic surges, Dedicated Container Inference scales seamlessly while maintaining performance."  
— Terrance Wang, Founding ML Engineer, Hedra

## Why Together AI?

### Cost Reduction
- **Factor**: 6x with Together Inference

### Workload at Scale
- **Configuration**: 72 GPUs on Together Managed Clusters

### Lower Latency
- **Factor**: 95× with Together Inference
