Pricing | Together AI
💰 Announcing our Series C. Intelligence should be abundant, not expensive →
🤝 Together AI & Y Combinator announce partnership to deliver the first dedicated YC GPU cluster →
⚡ On-demand B200s now available on Together GPU Clusters →
🚀 Now serving MiniMax-M3 for efficient inference →
INFERENCE
Compute
Model Shaping
Need help choosing?
Our team can help you find the best fit for your needs.
Pricing
Pricing
Serverless Inference
Most teams start with serverless inference and move to dedicated endpoints at scale.
Chat Vision Image Audio Video Transcribe Embeddings Rerank Moderation
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Price per 1M tokens
Batch API price
| Model | Input | output |
|---|---|---|
| <br><br>DeepSeek V4 Pro](/content/models/deepseek-v4-pro/index.html) | $1.74 $0.20 (cached) |
$3.48 |
| <br><br>MiniMax M3](/content/models/minimax-m3/index.html) | $0.30 $0.06 (cached) |
$1.20 |
| <br><br>GLM-5.2](/content/models/glm-52/index.html) | $1.40 $0.26 (cached) |
$4.40 |
| <br><br>Kimi K3](/content/models/kimi-k3/index.html) | $3.00 $0.30 (cached) |
$15.00 |
| <br><br>Inkling Small](/content/models/inkling-small/index.html) | $0.50 $0.10 (cached) |
$1.20 |
| <br><br>DeepSeek V4 Flash 0731](/content/models/deepseek-v4-flash-0731/index.html) | $0.14 $0.03 (cached) |
$0.28 |
| <br><br>Gemma 4 31B](/content/models/gemma-4-31b/index.html) | $0.39 | $0.97 |
| <br><br>NVIDIA Nemotron 3 Ultra](/content/models/nvidia-nemotron-3-ultra/index.html) | $0.60 $0.20 (cached) |
$3.60 |
| <br><br>Kimi K2.7 Code](/content/models/kimi-k27-code/index.html) | $0.95 $0.19 (cached) |
$4.00 |
| <br><br>Qwen3.7-Plus](/content/models/qwen37-plus/index.html) | $0.32 | $1.28 |
| <br><br>PrismMLTernary Bonsai 27B](/content/models/prism-ml-ternary-bonsai-27b/index.html) | 0.00 | |
| <br><br>Inkling](/content/models/inkling/index.html) | $1.00 $0.17 (cached) |
$4.05 |
| <br><br>LFM2.5-8B-A1B](/content/models/liquid-lfm2-5-8b-a1b/index.html) | $0.03 | $0.12 |
| <br><br>Kimi K2.6](/content/models/kimi-k26/index.html) | $1.20 $0.20 (cached) |
$4.50 |
| <br><br>Qwen3.7-Max](/content/models/qwen37-max/index.html) | $1.25 $0.13 (cached) |
$3.75 |
| <br><br>gpt-oss-120B](/content/models/gpt-oss-120b/index.html) | $0.15 | $0.60 |
| <br><br>Qwen3.5-397B-A17B](/content/models/qwen3-5-397b-a17b/index.html) | $0.60 $0.35 (cached) |
$3.60 |
| <br><br>Qwen3.5 9B](/content/models/qwen3-5-9b/index.html) | $0.17 | $0.25 |
| <br><br>Gemma-4-31B-it-Pearl](/content/models/gemma-4-31b-it-pearl/index.html) | $0.28 | $0.86 |
| <br><br>Cogito v2.1 671B](/content/models/cogito-v2-1-671b/index.html) | $1.25 | $1.25 |
| <br><br>Rnj-1 Instruct](/content/models/rnj-1-instruct/index.html) | $0.15 | $0.15 |
| <br><br>Llama 3.3 70B](/content/models/llama-3-3-70b/index.html) | $1.04 | $1.04 |
| <br><br>Gemma 3n E4B Instruct](/content/models/gemma-3n-e4b-it/index.html) | $0.06 | $0.12 |
| <br><br>gpt-oss-20B](/content/models/gpt-oss-20b/index.html) | $0.05 | $0.20 |
| <br><br>GLM-5.1](/content/models/glm-51/index.html) | $1.40 $0.26 (cached) |
$4.40 |
| <br><br>MiniMax M2.7](/content/models/minimax-m2-7/index.html) | $0.30 $0.06 (cached) |
$1.20 |
| <br><br>Qwen3.6-Plus](/content/models/qwen36-plus/index.html) | $0.50 | $3.00 |
| <br><br>Qwen2.5 7B Instruct Turbo](/content/models/qwen2-5-7b-instruct-turbo/index.html) | $0.30 | $0.30 |
| <br><br>Llama 3 8B Instruct Lite](/content/models/llama-3-8b-instruct-lite/index.html) | $0.14 | $0.14 |
| <br><br>Qwen3 235B A22B Instruct 2507 FP8 Throughput](/content/models/qwen3-235b-a22b-instruct-2507-fp8/index.html) | $0.20 | $0.60 |
Displayed prices refer to the lowest resolution/duration settings. Actual prices might vary.
Price per 1M tokens
| Model | Input | output |
|---|---|---|
| <br><br>MiniMax M3](/content/models/minimax-m3/index.html) | $0.30 | $1.20 |
| <br><br>Kimi K3](/content/models/kimi-k3/index.html) | $3.00 | $15.00 |
| <br><br>Inkling Small](/content/models/inkling-small/index.html) | $0.50 | $1.20 |
| <br><br>Gemma 4 31B](/content/models/gemma-4-31b/index.html) | $0.39 | $0.97 |
| <br><br>Kimi K2.7 Code](/content/models/kimi-k27-code/index.html) | $0.95 | $4.00 |
| <br><br>Qwen3.7-Plus](/content/models/qwen37-plus/index.html) | $0.32 | $1.28 |
| <br><br>PrismMLTernary Bonsai 27B](/content/models/prism-ml-ternary-bonsai-27b/index.html) | 0.00 | |
| <br><br>Inkling](/content/models/inkling/index.html) | $1.00 | $4.05 |
| <br><br>Kimi K2.6](/content/models/kimi-k26/index.html) | $1.20 | $4.50 |
| <br><br>Qwen3.5 9B](/content/models/qwen3-5-9b/index.html) | $0.17 | $0.25 |
| <br><br>Gemma 3n E4B Instruct](/content/models/gemma-3n-e4b-it/index.html) | $0.06 | $0.12 |
| <br><br>Qwen3.6-Plus](/content/models/qwen36-plus/index.html) | $0.50 | $3.00 |
Displayed prices refer to the lowest resolution/duration settings. Actual prices might vary.
| Model | Price per mp | Price per iMAGE | Default steps |
|---|---|---|---|
| <br><br>Qwen3.7-Plus](/content/models/qwen37-plus/index.html) | - | - | - |
| <br><br>GPT Image 2](/content/models/gpt-image-2/index.html) | - | $0.053 | - |
| <br><br>Wan 2.6 Image](/content/models/wan-2-6-image/index.html) | - | $0.03 | - |
| <br><br>Nano Banana Pro (Gemini 3 Pro Image)](/content/models/nano-banana-pro/index.html) | - | $0.134 | - |
| <br><br>FLUX.2 [pro]](/content/models/flux-2-pro/index.html) | - | $0.03 | - |
| <br><br>Ideogram 4.0](/content/models/ideogram-40/index.html) | - | $0.06 | - |
| <br><br>Gemini 3.1 Flash Image (Nano Banana 2)](/content/models/gemini-31-flash-image/index.html) | - | $0.05 | - |
| <br><br>Qwen Image 2.0 Pro](/content/models/qwen-image-20-pro/index.html) | - | $0.08 | - |
| <br><br>Qwen Image 2.0](/content/models/qwen-image-20/index.html) | - | $0.04 | - |
| <br><br>FLUX.2 [dev]](/content/models/flux-2-dev/index.html) | - | $0.0154 | - |
| <br><br>FLUX.2 [flex]](/content/models/flux-2-flex/index.html) | - | $0.03 | - |
| <br><br>FLUX.2 [max]](/content/models/flux-2-max/index.html) | $0.070 | - | 50 |
| <br><br>FLUX.1 Kontext [pro]](/content/models/flux-1-kontext-pro/index.html) | $0.04 | - | 28 |
| <br><br>FLUX1.1 [pro]](/content/models/flux1-1-pro/index.html) | $0.04 | - | - |
| <br><br>Juggernaut Pro Flux](/content/models/juggernaut-pro-flux/index.html) | $0.0049 | - | - |
| <br><br>GPT Image 1.5](/content/models/gpt-image-1-5/index.html) | - | $0.034 | - |
| <br><br>FLUX.1 Kontext [max]](/content/models/flux-1-kontext-max/index.html) | $0.08 | - | 28 |
| <br><br>FLUX.1 [schnell]](/content/models/flux-1-schnell-2/index.html) | $0.0027 | - | 4 |
| <br><br>SD XL](/content/models/sd-xl/index.html) | $0.0019 | - | - |
| <br><br>Ideogram 3.0](/content/models/ideogram-3-0/index.html) | $0.06 | - | - |
| <br><br>HiDream-I1-Full](/content/models/hidream-i1-full/index.html) | $0.009 | - | - |
| <br><br>Juggernaut Lightning Flux](/content/models/juggernaut-lightning-flux/index.html) | $0.0017 | - | - |
| <br><br>Qwen Image](/content/models/qwen-image/index.html) | $0.0058 | - | - |
| <br><br>Google Imagen 4.0 Fast](/content/models/google-imagen-4-0-fast/index.html) | $0.02 | - | - |
| <br><br>ByteDance Seedream 4.0](/content/models/bytedance-seedream-4-0/index.html) | $0.03 | - | - |
| <br><br>Google Imagen 4.0 Preview](/content/models/google-imagen-4-0-preview/index.html) | $0.04 | - | - |
| <br><br>Gemini Flash Image 2.5 (Nano Banana)](/content/models/gemini-flash-image-2-5/index.html) | - | $0.039 | - |
| <br><br>Google Imagen 4.0 Ultra](/content/models/google-imagen-4-0-ultra/index.html) | $0.06 | - | - |
| <br><br>ByteDance Seedream 3.0](/content/models/bytedance-seedream-3-0/index.html) | $0.018 | - | - |
Prices include default steps shown above. Additional costs apply only when exceeding default steps. See full pricing details →
Price per 1M Characters
| Model | Price |
|---|---|
| <br><br>Inkling Small](/content/models/inkling-small/index.html) | $0.50 |
| <br><br>Inkling](/content/models/inkling/index.html) | $1.00 |
| <br><br>NVIDIA Parakeet TDT 0.6B V3 Realtime](/content/models/parakeet-tdt-0-6b-v3-realtime/index.html) | $0.0035 |
| <br><br>NVIDIA Nemotron 3 ASR Streaming 0.6B](/content/models/nemotron-3-asr-streaming-0-6b/index.html) | $0.0015 |
| <br><br>Cartesia Sonic-3](/content/models/cartesia-sonic-3/index.html) | $65.00 |
| <br><br>Orpheus TTS](/content/models/orpheus-tts/index.html) | $15 |
| <br><br>Kokoro-82M TTS](/content/models/kokoro-82m/index.html) | $4.00 |
| <br><br>Cartesia Sonic-2](/content/models/cartesia-sonic/index.html) | $65.00 |
Price per video
| Model | Price |
|---|---|
| <br><br>ByteDance Seedance 2.0](/content/models/seedance-2-0/index.html) | $0.16 |
| <br><br>ByteDance Seedance 2.5](/content/models/seedance-2-5/index.html) | $0.115 |
| <br><br>FLUX 3](/content/models/flux-3/index.html) | $0.17 |
| <br><br>Google Veo 3.0](/content/models/google-veo-3-0/index.html) | $1.60 |
| <br><br>Kling 1.6 Standard](/content/models/kling-1-6-standard/index.html) | $0.19 |
| <br><br>Kling 2.1 Master](/content/models/kling-2-1-master/index.html) | $0.92 |
| <br><br>Kling 2.1 Pro](/content/models/kling-2-1-pro/index.html) | $0.32 |
| <br><br>Kling 2.1 Standard](/content/models/kling-2-1-standard/index.html) | $0.18 |
| <br><br>Vidu 2.0](/content/models/vidu-2-0/index.html) | $0.28 |
| <br><br>Vidu Q1](/content/models/vidu-q1/index.html) | $0.22 |
| <br><br>Wan 2.2 I2V](/content/models/wan-2-2-i2v/index.html) | $0.31 |
| <br><br>Wan 2.2 T2V](/content/models/wan-2-2-t2v/index.html) | $0.66 |
| <br><br>Qwen3.6-Plus](/content/models/qwen36-plus/index.html) | $0.50 |
| <br><br>Sora 2](/content/models/sora-2/index.html) | $0.80 |
| <br><br>PixVerse v5](/content/models/pixverse-v5/index.html) | $0.30 |
| <br><br>ByteDance Seedance 1.0 Lite](/content/models/bytedance-seedance-1-0-lite/index.html) | $0.14 |
| <br><br>ByteDance Seedance 1.0 Pro](/content/models/bytedance-seedance-1-0-pro/index.html) | $0.57 |
| <br><br>Google Veo 3.0 Fast + Audio](/content/models/google-veo-3-0-fast-audio/index.html) | $1.20 |
| <br><br>Google Veo 3.0 Fast](/content/models/google-veo-3-0-fast/index.html) | $0.80 |
| <br><br>Google Veo 3.0 + Audio](/content/models/google-veo-3-0-audio/index.html) | $3.20 |
| <br><br>Google Veo 2.0](/content/models/google-veo-2-0/index.html) | $2.50 |
| <br><br>MiniMax Hailuo 02](/content/models/minimax-hailuo-02/index.html) | $0.49 |
| <br><br>MiniMax 01 Director](/content/models/minimax-01-director/index.html) | $0.28 |
Price per audio minute
Batch API price
| Model | Price |
|---|---|
| <br><br>NVIDIA Parakeet TDT 0.6B v3](/content/models/parakeet-tdt-0-6b-v3/index.html) | $0.0015 |
| <br><br>NVIDIA Nemotron 3.5 ASR](/content/models/nvidia-nemotron-35-asr/index.html) | $0.0045/min |
| <br><br>Whisper Large v3](/content/models/openai-whisper-large-v3/index.html) | $0.0015 |
| <br><br>Whisper Large v3 (Streaming)](/content/models/whisper-large-v3-streaming/index.html) | $0.0035/min |
Price per 1M tokens
| Model | Price |
|---|---|
| <br><br>Multilingual e5 large instruct](/content/models/multilingual-e5-large-instruct/index.html) | $0.02 |
Price per 1M tokens
| Model | Price |
|---|
No matching models
Price per 1M tokens
| Model | Price |
|---|---|
| <br><br>Llama Guard 4 12B](/content/models/llama-guard-4-12b/index.html) | $0.20 |
Provisioned Throughput
Reserve dedicated capacity in throughput units (PTUs). Each PTU represents fixed capacity. The tokens-per-minute it delivers depends on the model and the token type.
Estimate your PTUs & cost
Model
MiniMax M3GLM-5.2Kimi K3
Compare vs.
GPT-5.6 SolGPT-5.6 TerraGPT-5.6 LunaGPT 5.5GPT 5.4GPT 5.2GPT 5.4 miniClaude Fable 5Claude Opus 4.8Claude Sonnet 5Claude Haiku 4.5Gemini 3.5 FlashGemini 3 Flash
Peak requests / Sec
Cache hit rate
%
Input tokens / request
output tokens / request
PTUs required
353
Est. monthly cost
$773,070
Est. monthly savings
$1,802,37070% LOWER
| Compute costs | Input TPM/PTU |
Cached TPM/PTU |
Output TPM/PTU |
Price PTU/MIN |
|---|---|---|---|---|
| Kimi K3 | 16,667 | 166,667 | 3,333 | $0.05 |
| MiniMax M3 | 138,840 | 694,200 | 23,140 | $0.05 |
| GLM-5.2 | 35,731 | 192,400 | 9,620 | $0.05 |
Savings compare Together PTU cost against the selected commercial model's published list price ($/1M tokens) on the same traffic profile. Estimates assume continuous 24/7 provisioning (~43,800 min/mo).
Dedicated Inference
Deploy models on custom hardware with guaranteed performance and full control.
Single-tenant GPU instances with:
Guaranteed performance (no sharing)
Support for custom models
Autoscaling & traffic spike handling
| Hardware All prices per gpu per hour |
On-demand Pay as you go |
Reserved |
|---|---|---|
| NVIDIA HGX H100 | $5.49 | Contact sales |
| NVIDIA HGX H200 | Contact us | |
| NVIDIA HGX B200 | $8.99 | |
| NVIDIA HGX B300 | Contact us | |
| NVIDIA GB200 NVL72 | Contact us | |
| NVIDIA GB300 NVL72 | Contact us |
GPU Clusters
On-demand
Pay as you go GPU capacity on an hourly basis.
| Hardware | Hourly |
|---|---|
| NVIDIAGB200 NVL72 | — |
| NVIDIAGB300 NVL72 | — |
| NVIDIAHGX B200 | $8.19 |
| NVIDIAHGX B300 | — |
| NVIDIAHGX H100 | $3.99 |
| NVIDIAHGX H200 | $5.99 |
On-demand hourly rates and reserved capacity
All prices are per GPU per hour.
| Hardware | ON-Demand | Reserved |
|---|---|---|
| Pay as you go | 7-30 days | 31-90 days |
| --- | --- | --- |
| NVIDIAHGX H100 | $3.99 | $3.69 |
| NVIDIAHGX H200 | $5.99 | $4.99 |
| NVIDIAHGX B200 | $8.19 | $7.99 |
| NVIDIAGB200 NVL72 | — | Contact us |
| NVIDIAGB300 NVL72 | — | Contact us |
| NVIDIAHGX B300 | — | Contact us |
Sandbox
Code Sandbox
Customize a deployment of VM sandboxes for large development environments.
| Compute costs | Price/Hour |
|---|---|
| Per vCPU | $0.0446 |
| Per GiB RAM | $0.0149 |
Code Interpreter
Execute LLM-generated code securely using our API.
| Duration? | Price/Session |
|---|---|
| Session (60 minutes) | $0.03 |
Storage
High-bandwidth, parallel filesystem colocated with your compute.
| Compute costs | Price | Unit |
|---|---|---|
| Shared Filesystem | $0.16 | GiB/month |
Fine-Tuning
Train open-source models for real production use.
Per 1M tokens
| Supervised Fine-Tuning | Direct Preference Optimization | |
|---|---|---|
| Size | LoRA | Full Fine-Tuning |
| --- | --- | --- |
| Up to 16B | $0.48 | $0.54 |
| 17B-69B | $1.50 | $1.65 |
| 70-100B | $2.90 | $3.20 |
Price is based on the sum of tokens processed in the fine-tuning training dataset (training dataset size * number of epochs) plus any tokens in the optional evaluation dataset (validation dataset size * number of evaluations). Each job is subject to a minimum charge of $4.00.
| Size | Supervised Fine-Tuning (LoRA) |
Direct Preference Optimization (LoRA) |
Minimum charge |
|---|---|---|---|
| DeepSeek-R1 DeepSeek-R1-0528 DeepSeek-V3 DeepSeek-V3-0324 DeepSeek-V3.1 DeepSeek-V3.1-Base |
$10.00 | $25.00 | $20.00 |
| GLM-4.6 GLM-4.7 |
$9.00 | $22.50 | $27.00 |
| GLM-5 GLM-5.1 |
$40 | $100 | $60 |
| gpt-oss-120B | $5.00 | $12.50 | $6.00 |
| Kimi K2 Thinking Kimi K2 Instruct-0905 Kimi K2 Instruct Kimi K2 Base Kimi K2.5 Kimi K2.6 Kimi K2.7-Code |
$15.00 | $37.50 | $60.00 |
| Llama 4 Maverick Llama 4 Maverick Instruct |
$8.00 | $20.00 | $16.00 |
| Llama 4 Scout Llama 4 Scout |
$3.00 | $7.50 | $6.00 |
| Qwen3-Coder-480B-A35B-Instruct | $9.00 | $22.50 | $18.00 |
| Qwen3-235B-A22B Qwen3-235B-A22B-Instruct-2507 |
$6.00 | $15.00 | No min. price |
| Qwen3.5-122B-A10B | $6.00 | $15.00 | $10.00 |
| Qwen3.5-397B-A17B | $8.00 | $20.00 | $22.00 |
Price is based on the sum of tokens processed in the fine-tuning training dataset (training dataset size * number of epochs) plus any tokens in the optional evaluation dataset (validation dataset size * number of evaluations).