Voice | Together AI
Deploy real-time voice agents for every use case
Build voice agents that sound natural. Combine the best STT, LLM, and TTS models on co-located infrastructure for ultra-low latency and production-scale reliability.
Why Together AI for Voice Agents
The complete voice stack, built for real-time production use.
One platform for every voice use case
Deploy fast, expressive, multilingual, or cloned models for any use case. Access MiniMax, Rime, Deepgram, OpenAI, Cartesia through a single API. Swap configurations and switch models without rebuilding integrations.
Ultra-low latency conversations
Sub-second STT-to-TTS latency, built into the infrastructure. The entire pipeline runs co-located, keeping end-to-end latency under 500ms for conversations that feel instant.
Scales without breaking
Autoscale dynamically to thousands of concurrent calls across 25+ global regions. Dedicated GPU endpoints with a 99.9% uptime SLA keep traffic spikes running on pre-warmed capacity, every time.
The complete voice model library
Open-source and proprietary models across the full voice pipeline, on one platform. Switch between models optimized for emotion, pronunciation, code-switching, or cloning — with minimal code changes.
Models
- Cartesia Sonic 3.5
- NVIDIA Nemotron 3 ASR Streaming 0.6B
- MiniMax Speech 2.8
- NVIDIA Parakeet TDT 0.6B v3
- Mist v3 Omni
- MiniMax Speech 2.6 Turbo
- NVIDIA Nemotron 3.5 ASR
- Deepgram Nova-3 Multilingual
- Deepgram Flux
- Whisper Large v3
- Deepgram Aura-2
- Arcana V3 Turbo
Have your own model?
Deploy custom containers on Together’s managed GPU infrastructure with automatic scaling, job queues, and built-in observability.
Trusted by teams building voice at scale
- 6× cost reduction
- <400ms p95 model latency
- Weekly model deployments
"Low latency is especially important for voice because there’s a much higher UX bar. Together helped us push latency down by optimizing our models with techniques like speculative decoding, and they’ve been a reliable production partner — proactive about risks and fast when issues come up."
Max Lu
Head of Research, Decagon