NVIDIA GPU Clusters: H100, H200, B200, GB200 | Together AI
Together GPU Clusters
Reliable self-serve, AI-ready GPU clusters at scale
Go from zero to production in minutes. Bare-metal performance, InfiniBand networking, and managed orchestration — with flexible pricing for both on-demand and reserved capacity.
Why Together GPU Clusters
Infrastructure that keeps long-running jobs on track — with automated recovery, elastic scale, and zero DevOps overhead.
Research & experimentation
Spin up a cluster in minutes, test your hypothesis, shut it down when you're done. On-demand pricing means scratchpad experiments don't carry the cost of production workloads — so your team can move fast without burning budget.
Distributed training
Reserve from 8 to 4,000+ GPUs without rethinking your architecture. Our InfiniBand interconnect keeps gradient synchronization fast and communication overhead low — so your training runs finish faster, not just bigger.
Inference & serving
Go from trained model to production endpoint without switching providers. Low-latency networking, Kubernetes-native deployment, and flexible ingress configuration make it easy to serve at scale — and scale back down when traffic drops.
Everything you need to train at scale
Managed infrastructure with built-in observability, orchestration flexibility, and research-grade performance.
- Sustained performance
- Multi-week stability
- Reduced stragglers
- Predictable latency
Maintain high utilization across multi-week training runs and model serving. Kernel, hardware, and storage acceleration reduce stragglers and keep latencies predictable.
Frontier research-powered training performance
The Together Kernel Collection, built by our Chief Scientist Tri Dao (creator of FlashAttention), delivers improved training and inference performance.
Together Kernel Collection
- AI Training Performance: NVIDIA Hopper to Blackwell, with TKC
TKC vs SOTA Approaches
90% faster training
Training a 70B parameter Llama-architecture model (BF16) with an optimized TorchTitan + Together Kernel Collection (TKC) reached 15,264 tokens/second/GPU on NVIDIA HGX B200, up from 8,080 tokens/second on NVIDIA HGX H100—a 90% jump in training speed.
Fully managed, high-performance shared filesystems for faster training and innovation cycle
Provision and attach shared storage volumes for your GPU clusters to store and persist your training data, model weights — ensure your GPUs do not starve for data.
WEKA excels at high IOPS workloads.
With strong metadata performance. It scored 826.86 on IO500 and delivers sub 200 microsecond latency. Heavy small file operations and metadata intensive tasks like checkpoint discovery across hundreds of training ranks.
VAST simplifies operations.
VAST's disaggregated architecture separates compute from storage for straightforward capacity expansion and a unified namespace. Built for enterprise environments where operational simplicity and broad feature coverage matter most.
Flexible pricing models
Both options are fully self-serve. Choose based on your capacity requirements and commitment level.
On-Demand
Standard hourly rate
- Commitment: None—pay hourly, terminate anytime
- Best for: Starting with on-demand for flexibility
- Capacity: Based on real-time availability
- Scale: Up to 256 GPUs
Reserved
Lower hourly rate
- Commitment: Up to 6 months, pay upfront
- Best for: Guaranteed access with better economics
- Capacity: Locked in for your duration
- Scale: Up to 4,000+ GPUs
Orchestration flexibility for your AI workloads
Self-serve GPUs with hourly pricing.
- Managed Kubernetes for training and inference.
- Slurm on Kubernetes for training workloads.
Regions and availability zones
Launch close to your users and data across 25+ cities.
USA
2GW+ in the portfolio with 600MW of near-term capacity in the US.
Europe
150 MW+ available in Europe: France, Netherlands, Sweden, and Romania.
Asia & Middle East
Options available based on the scale of the projects.
Production-grade security.
We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.
As an NVIDIA Cloud Partner, Together builds and operates clusters on NVIDIA NCP reference architectures for predictable performance and faster time to production. Your data and models remain under your control with strict privacy safeguards and SOC 2–compliant security practices.
Customers running inference in production
"Together GPU Clusters provided a combination of amazing training performance, expert support, and the ability to scale to meet our rapid growth to help us serve our growing community of AI creators."
Demi Guo, CEO, Pika
"Together AI provides the performance and reliability we need for real-time, high-quality image and video generation at scale. We value that Together AI is much more than an infrastructure provider — they're a true innovation partner, enabling us to push creative boundaries without compromise."
Victor Perez, Co-Founder, Krea