Together AI | The AI Native Cloud

Build what's next on the AI Native Cloud

Full-stack AI platform, powered by cutting-edge research.

Start building Contact Sales

Trusted by

The Together AI Platform

Accelerate inference, model shaping and pre-training on a research-optimized platform.

Faster inference

2x

powered by cutting-edge research.

Learn how

Lower cost

60%

with workload-specific optimization.

Learn how

Faster pre-training

90%

with Together Kernel Collection.

Learn how

Full-stack cloud

Powering every step of the AI development journey
—from experimentation to massive scale.

Serverless Inference

The fastest way to run open-source models on demand. Powered by cutting-edge inference research. No infrastructure to manage, no long-term commitments.

Learn more

Batch Inference

Cost-effectively process massive workloads asynchronously. Scale to 30 billion tokens per model with any serverless model or private deployment.

Learn more

Provisioned Throughput

Committed inference capacity with token-based pricing, reserved throughput, and a 99% uptime SLA. Drop-in API compatibility for production workloads with no infrastructure to manage.

Learn more

Dedicated Model Inference

Deploy models on dedicated infrastructure. Purpose-built for teams who need speed, control, and the best economics in the market.

Learn more

Dedicated Container Inference

GPU infrastructure purpose-built for generative media workloads. Deploy video, audio, and image models with performance acceleration powered by Together Research.

Learn more

Accelerated Compute

Scale from self-serve instant clusters to thousands of GPUs, all optimized for better performance with Together Kernel Collection.

Learn more

Sandbox

Use fast, secure code sandboxes at scale to set up full-scale development environments for AI apps and agents.

Learn more


Managed Storage

High-performance managed storage for AI-native workloads. Object storage and parallel filesystems optimized for AI, with zero egress fees.

Learn more


Fine-Tuning

Fine-tune open-source models for production workloads, using the latest research techniques. Improve accuracy, reduce hallucinations, and control behavior — without managing training infrastructure.

Learn more

Grounded in cutting-edge research

Foundational systems research for production AI.

Agents

ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale

Hao Kang, Ziyang Li, Weili Xu, Xinyu Yang et al.

Read More

Together AI at ICML 2026: frontier research across the full stack

Together Research

Read More

Kernels

ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)

Willy Chan, Nathan Paek, Simon Guo et al.

Read More

Agents

Violin: An open-source video translation skill that breaks language barriers

Shang Zhu, Kevin Qinghong Lin (Oxford) et al.

Read More

Inference

Accelerate RL rollouts by up to 50% with distribution-aware speculative decoding

Zelei Shao, Vikranth Srivatsa et al.

Read More

Architecture

Parcae: Doing more with fewer parameters using stable looped models

Hayden Prairie, Zachary Novack et al.

Read More

Agents

EinsteinArena: Harnessing the collective intelligence of agents in the wild to advance science

Federico Bianchi,* Yongchan Kwon et al.

Read More

Agents

AI for Systems: Using LLMs to Optimize Database Query Execution

Mehmet Hamza Erol, Xiangpeng Hao et al.

Read More

Kernels

Inside the Together AI kernels team

Will Van Eaton

Read More

Inference

Aurora

Junxiong Wang, Fengxiang Bie, Jisen Li et al.

Read More

Agents

Plan, divide, and conquer: How weak models excel at long context tasks

Zhen Xu, Shang Zhu, Jue Wang, Junlin Wang et al.

Read More

Architecture

Mamba-3

Aakash Lahoti* (CMU), Kevin Y. Li* (CMU) et al.

Read More

Kernels

Key research and product announcements at the AI Native Conf

Read More

Kernels

FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling

Ted Zadouri (Princeton University et al.

Read More

Inference

Cache-aware prefill–decode disaggregation (CPD) for up to 40% faster long-context LLM serving

Jiejing Zhang, Yubo Wang, Yinghui Liu et al.

Read More

Agents

CoderForge-Preview: SOTA open dataset for training efficient coding agents

By Alpay Ariyak*, Junda Zhang et al.

Read More

Agents

How speech models fail where it matters the most and what to do about it

Kaitlyn Zhou, Martijn Bartelds et al.

Read More

Inference

Consistency diffusion language models: Up to 14x faster inference without sacrificing quality

Minseo Kim, Chenfeng Xu et al.

Read More

Agents

What do LLMs think when you don't tell them what to think about?

Yongchan Kwon and James Zou

Read More

Agents

DSGym: A holistic framework for evaluating and training data science agents

Fan Nie, Junlin Wang, Harper Hua et al.

Read More

Kernels

Research POV: Yes, AGI Can Happen – A Computational Perspective

Together AI

Read More

Model Shaping

How to run TorchForge reinforcement learning pipelines in the Together AI Native Cloud

Together AI Training and Research et al.

Read More

Model Shaping

Introducing AutoJudge: Streamlined inference acceleration via automated dataset curation

Roman Garipov, Fedor Velikonivtsev et al.

Read More

Agents

Large Reasoning Models Fail to Follow Instructions During Reasoning: A Benchmark Study

Yongchan Kwon, Shang Zhu, Federico Bianchi et al.

Read More

Inference

AdapTive-LeArning Speculator System (ATLAS): A New Paradigm in LLM Inference via Runtime-Learning Accelerators

Junxiong Wang, Shirley Wu, Zelei Shao et al.

Read More

Agents

How Together AI Uses AI Agents to Automate Complex Engineering Tasks: Lessons from Developing Efficient LLM Inference Systems

Shang Zhu, Federico Bianchi, Wai Tong Chung et al.

Read More

Agents

Back to The Future: Evaluating AI Agents on Predicting Future Events

Federico Bianchi, Junlin Wang, Zain Hasan et al.

Read More

Inference

DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL

Michael Luo*, Naman Jain*, Jaskirat Singh* et al.

Read More

Agents

From Zero to One: Building An Autonomous and Open Data Scientist Agent from Scratch

Federico Bianchi, Shang Zhu, Zain Hasan et al.

Read More

Inference

Model-Preserving Adaptive Rounding with YAQA

Albert Tseng, Zhaofeng Sun, and Chris De Sa

Read More

Agents

Mixture-of-Agents Alignment: Harnessing the Collective Intelligence of Open-Source LLMs to Improve Post-Training

Junlin Wang, Roy Xie, Shang Zhu, Jue Wang et al.

Read More

Inference

Boosting DeepSeek-R1’s Speed with Customized Speculative Decoding

Wai Tong Chung, Dan Waters, Avner May et al.

Read More

Kernels

Chipmunk: Training-Free Acceleration of Diffusion Transformers with Dynamic Column-Sparse Deltas

Austin Silveria, Soham Govande, Dan Fu

Read More

Model Shaping

Direct Preference Optimization: A Technical Deep Dive

Ivan Provilkov, Zain Hasan, Max Ryabinin

Read More

Model Shaping

Continued Fine-tuning of LLMs: A Technical Deep Dive

Artem Chumachenko, Zain Hasan et al.

Read More

Agents

Open Deep Research

Together AI

Read More

Inference

DeepCoder: A Fully Open-Source 14B Coder at O3-mini Level

Michael Luo*, Sijun Tan*, Roy Huang* et al.

Read More

Kernels

ThunderKittens Now Optimized for NVIDIA Blackwell GPUs

Benjamin Spector, Aaryan Singhal, Dan Fu et al.

Read More

Inference

Minions: embracing small LMs, shifting compute on-device, and cutting cloud costs in the process

Avanika Narayan*, Dan Biderman* et al.

Read More

Model Shaping

Long Context Fine-Tuning: A Technical Deep Dive

George Grigorev, Zain Hasan, Max Ryabinin

Read More

Model Shaping

Fine-Tuning LLMs for Multi-Turn Conversations: A Technical Deep Dive

Artem Chumachenko, Zain Hasan et al.

Read More

Inference

Even Better, Even Faster Quantized LLMs with QTIP

Albert Tseng, Qingyao Sun, David Hou et al.

Read More

Architecture

Linearizing LLMs with LoLCATs

Michael Zhang, Simran Arora et al.

Read More

Applications

Multimodal Document RAG with Llama 3.2 Vision and ColQwen2

Zain Hasan

Read More

Architecture

The Mamba in the Llama: Distilling and Accelerating Hybrid Models

Junxiong Wang, Daniele Paliotta, Avner May et al.

Read More

Inference

Speculative decoding for high-throughput long-context inference

Jian Chen, Vashisth Tiwari, Ranajoy Sadhukhan et al.

Read More

Inference

TEAL: Training-Free Activation Sparsity in Large Language Models

James Liu, Pragaash Ponnusamy, Tianle Cai et al.

Read More

Kernels

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Jay Shah (Colfax Research) et al.

Read More

Applications

Building a personalized code assistant with open-source LLMs using RAG Fine-tuning

Kezhen Chen, Linda He, Ben Athiwaratkun et al.

Read More

Inference

SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices

Ruslan Svirschevski, Avner May et al.

Read More

Agents

Together MoA — collective intelligence of open-source models pushing the frontier of LLM capabilities

Junlin Wang, Jue Wang, Ben Athiwaratkun et al.

Read More

Architecture

Dragonfly: A large vision-language model with multi-resolution zoom

Kezhen Chen, Rahul Thapa, Rahul Chalamala et al.

Read More

Kernels

ThunderKittens: A Simple Embedded DSL for AI kernels

Benjamin Spector, Aaryan Singhal et al.

Read More

Inference

FAQ: Building LLMs with RedPajama-v2, a 30 trillion token web dataset

Together AI

Read More

Inference

Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

Zhuoming Chen, Avner May et al.

Read More

Architecture

BASED: Simple linear attention language models balance the recall-throughput tradeoff

Simran Arora, Sabri Eyuboglu, Michael Zhang et al.

Read More

Architecture

Evo: Long-context modeling from molecular to genome scale

Eric Nguyen, Michael Poli, Matthew Durrant et al.

Read More

Inference

BitDelta: Your Fine-Tune May Only Be Worth One Bit

James Liu, Guangxuan Xiao, Kai Li et al.

Read More

Inference

Long context retrieval models with Monarch Mixer

Jon Saad-Falcon, Dan Fu, Simran Arora

Read More

Architecture

Mamba-3B-SlimPJ: State-space models rivaling the best Transformer architecture

Tri Dao, Albert Gu

Read More

Architecture

Paving the way to efficient architectures: StripedHyena-7B, open source models offering a glimpse into a world beyond Transformers

Together

Read More

Kernels

FlashFFTConv: Efficient Convolutions for Long Sequences with Tensor Cores

Dan Fu, Hermann Kumbong, Eric Nguyen et al.

Read More

Inference

RedPajama-Data-v2: An open dataset with 30 trillion tokens for training large language models

Together

Read More

Kernels

Flash-Decoding for long-context inference

Tri Dao, Daniel Haziza, Francisco Massa et al.

Read More

Inference

Medusa: Simple framework for accelerating LLM generation with multiple decoding heads

Tianle Cai*, Yuhong Li*, Zhengyang Geng et al.

Read More

Model Shaping

Llama-2-7B-32K-Instruct — and fine-tuning for Llama-2 models with Together API

Together

Read More

Inference

Faster inference enables up to 5x price reduction on Together API

Together

Read More

Inference

Preparing for the era of 32K context: Early learnings and explorations

Together

Read More

Architecture

Monarch Mixer: A new model architecture for increased efficiency

Dan Fu, Simran Arora, Chris Ré

Read More

Model Shaping

Fine-tuning language models over slow networks using activation compression with guarantees

Jue Wang, Binhang Yuan, Luka Rimanic et al.

Read More

Model Shaping

Decentralized training of foundation models in heterogeneous environments

Binhang Yuan, Yongjun He et al.

Read More

Kernels

FlashAttention: Fast and memory-efficient exact attention with IO-Awareness

Tri Dao, Daniel Y. Fu, Stefano Ermon et al.

Read More

Model Shaping

CocktailSGD: Fine-tuning foundation models over 500Mbps networks

Jue Wang, Binhang Yuan, Luka Rimanic et al.

Read More

Inference

FlexGen: High-throughput generative inference of large language models with a single GPU

Ying Sheng, Lianmin Zheng, Binhang Yuan et al.

Read More

Architecture

Hyena Hierarchy: Towards larger convolutional language models

Michael Poli, Stefano Massaroli, Eric Nguyen et al.

Read More

Kernels

FlashConv: Speeding up state space models

Dan Fu and Tri Dao

Read More

Architecture

Hungry Hungry Hippos: Towards language modeling with state space models

Daniel Y. Fu, Tri Dao, Khaled K. Saab et al.

Read More

Model Shaping

NeurIPS 2022: Overcoming communication bottlenecks for decentralized training (2/2)

Together

Read More

Model Shaping

NeurIPS 2022: Overcoming communication bottlenecks for decentralized training (1/2)

Together

Read More

Inference

HELM: benchmarking large language models on the Together Research Computer

Together

Read More

recognized by

-

-

- -

AI natives build on Together AI

See how Together AI powers customers building the next generation of AI products.

\ \ How Cursor partnered with Together AI to deliver real-time, low-latency inference at scale\ \ Inference, GPU clusters, RESEARCH  •  Enterprise\ \ \ \ How Decagon Engineered Sub-Second Voice AI with Together AI\ \ Inference, GPU clusters, RESEARCH\ \ 6x\ \ Cost reduction per turn vs. gpt-5 mini\ \ \ \ 11x\ \ Faster Inference\ \

View All Stories

What’s new at Together AI

All blog posts

\ \ Company\ \ Announcing our $800M Series C to accelerate the shift to open-source AI\ \ We raised $800M to accelerate the shift to open-source AI. Here's why the economics of closed models don't scale, and what we're building next.](/content/blog/announcing-our-series-c/index.html)

\ \ GPU Clusters\ \ Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community \ \ No more two-year compute contracts. Together AI and YC just gave YC startups a faster way to get GPUs.](/content/blog/together-yc-gpu-cluster/index.html)

\ \ Model Library\ \ DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding\ \ We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.](/content/blog/deepseek-v4-flash-0731-vs-gpt-5-6-luna-on-deepswe-cost-and-coding/index.html)

\ \ Research\ \ Together AI at ICML 2026: frontier research across the full stack\ \ Nine papers at ICML 2026 across the full stack. The research that becomes the Together platform. Find us at booth B714 in Seoul.](/content/blog/icml-2026/index.html)

All blog posts