NVIDIA HGX H200 Cluster Pricing & Specs | Rent HGX H200 GPUs | Together AI
NVIDIA H200
Rent NVIDIA H200 GPUs on demand or reserve a cluster, with 141GB HBM3e and 4.8TB/s for memory-bound LLM inference and training.
NVIDIA H200 Pricing & Specs
The NVIDIA H200 is the memory-upgraded Hopper GPU, the first with HBM3e: 141GB at 4.8TB/s on the same compute as the H100. That capacity suits memory-bound work like long-context inference and large-batch serving. Rent it on Together on demand or as a dedicated cluster.
Performance
| Metric | Value |
|---|---|
| Faster training | 2x vs. H100 |
| Faster inference | 110x higher |
| Better efficiency | 1.4x vs. H100 |
Pricing
Rent HGX H200 on demand, or reserve dedicated capacity at a lower rate. All prices are per GPU per hour.
| Plan | Price |
|---|---|
| On-demand (pay as you go) | $5.99 |
| Reserved 7-30 days | $4.99 |
| Reserved 31-90 days | $4.15 |
| Reserved 91-180 days | $3.99 |
| Reserved 181+ days | Contact us |
Prices as of August 2026.
Technical specification
- Model data: NVIDIA H200 (SXM)
- Architecture: NVIDIA Hopper
- GPU memory (VRAM): 141GB HBM3e
- Memory bandwidth: 4.8TB/s
- FP64: 34 TFLOPS
- FP64 Tensor Core: 67 TFLOPS
- FP32: 67 TFLOPS
- TF32 Tensor Core: 989 TFLOPS
- FP16 / BF16 Tensor Core: 1,979 TFLOPS
- FP8 Tensor Core: 3,958 TFLOPS
- INT8 Tensor Core: 3,958 TOPS
- NVLink bandwidth: 900GB/s
- Interconnect: NVLink 900GB/s, PCIe Gen5 128GB/s
- Multi-Instance GPU (MIG): Up to 7 instances
- Max thermal design power (TDP): Up to 700W (configurable)
- Form factor: HGX H200, 4 or 8 GPUs
Why Rent NVIDIA H200 on Together
The world’s most powerful AI infrastructure. Delivered faster. Tuned smarter.
Key Benefits
- 2x performance over H100: Each H200 GPU cluster offers double the inference throughput compared to H100, ideal for deploying LLMs at unprecedented scale.
- Enhanced memory bandwidth: With 141GB HBM3e GPU memory and 4.8TB/s bandwidth, H200 significantly accelerates memory-intensive generative AI workloads and HPC applications.
- Maximum efficiency and TCO savings: Achieve higher performance within the same power profile as previous-gen GPUs, drastically reducing energy consumption and total cost of ownership.
- Run by researchers who train models: Our research team actively runs and tunes training workloads on NVIDIA H200 systems for edge-of-possibility expertise.
FAQ
What is the NVIDIA H200?
The NVIDIA H200 is a Hopper-architecture data center GPU with 141GB of HBM3e memory and 4.8TB/s of bandwidth, the first GPU with HBM3e, built for memory-bound LLM inference and large-scale training.
How much does it cost to rent an NVIDIA H200?
On Together GPU Clusters, H200 GPUs are $5.99 per GPU per hour on demand, with reserved rates from $3.99 per GPU per hour for longer commitments. See the pricing section above or contact sales for volume pricing.
Can I rent H200 GPUs by the hour?
Yes. Launch an on-demand H200 cluster in minutes and pay per GPU, or reserve dedicated capacity for a defined training window.
How much memory does the H200 have?
Each H200 has 141GB of HBM3e memory, up from 80GB on the H100.
What is the H200 memory bandwidth?
The H200 delivers 4.8TB/s of memory bandwidth, roughly 1.4x the H100.
What is the power consumption of the H200?
The H200 SXM has a configurable thermal design power of up to 700W.
Is the H200 available on demand or only reserved?
Both. Together offers on-demand H200 capacity with real-time availability and reserved clusters for longer terms.
Which regions are H200 clusters available in?
Across the US, Europe (UK, Spain, France, Portugal, Iceland), and select Asia and Middle East regions, supporting data-residency and compliance needs.
What is the NVIDIA H200 used for?
The H200 is used for long-context and large-batch LLM inference, serving models that exceed 80GB at full precision, and large-scale training where memory capacity and bandwidth are the binding constraint.