Together AI | The AI Native Cloud
Build what's next on the AI Native Cloud
Full-stack AI platform, powered by cutting-edge research.
Trusted by
The Together AI Platform
Accelerate inference, model shaping and pre-training on a research-optimized platform.
Faster inference
2x
powered by cutting-edge research.
Lower cost
60%
with workload-specific optimization.
Faster pre-training
90%
with Together Kernel Collection.
Full-stack cloud
Powering every step of the AI development journey —from experimentation to massive scale.
Inference
Compute
Model shaping
Serverless Inference
The fastest way to run open-source models on demand. Powered by cutting-edge inference research. No infrastructure to manage, no long-term commitments.
Batch Inference
Cost-effectively process massive workloads asynchronously. Scale to 30 billion tokens per model with any serverless model or private deployment.
Provisioned Throughput
Committed inference capacity with token-based pricing, reserved throughput, and a 99% uptime SLA. Drop-in API compatibility for production workloads with no infrastructure to manage.
Dedicated Model Inference
Deploy models on dedicated infrastructure. Purpose-built for teams who need speed, control, and the best economics in the market.
Dedicated Container Inference
GPU infrastructure purpose-built for generative media workloads. Deploy video, audio, and image models with performance acceleration powered by Together Research.
Accelerated Compute
Scale from self-serve instant clusters to thousands of GPUs, all optimized for better performance with Together Kernel Collection.
Sandbox
Use fast, secure code sandboxes at scale to set up full-scale development environments for AI apps and agents.
Managed Storage
High-performance managed storage for AI-native workloads. Object storage and parallel filesystems optimized for AI, with zero egress fees.
Fine-Tuning
Fine-tune open-source models for production workloads, using the latest research techniques. Improve accuracy, reduce hallucinations, and control behavior — without managing training infrastructure.
Grounded in cutting-edge research
Foundational systems research for production AI.
Agents
ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale
Hao Kang, Ziyang Li, Weili Xu, Xinyu Yang et al.
Together AI at ICML 2026: frontier research across the full stack
Together Research
Kernels
ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet)
Willy Chan, Nathan Paek, Simon Guo et al.
Agents
Violin: An open-source video translation skill that breaks language barriers
Shang Zhu, Kevin Qinghong Lin (Oxford) et al.
Inference
Accelerate RL rollouts by up to 50% with distribution-aware speculative decoding
Zelei Shao, Vikranth Srivatsa et al.
Architecture
Parcae: Doing more with fewer parameters using stable looped models
Hayden Prairie, Zachary Novack et al.
Agents
EinsteinArena: Harnessing the collective intelligence of agents in the wild to advance science
Federico Bianchi,* Yongchan Kwon et al.
Agents
AI for Systems: Using LLMs to Optimize Database Query Execution
Mehmet Hamza Erol, Xiangpeng Hao et al.
Kernels
Inside the Together AI kernels team
Will Van Eaton
Inference
Aurora
Junxiong Wang, Fengxiang Bie, Jisen Li et al.
Agents
Plan, divide, and conquer: How weak models excel at long context tasks
Zhen Xu, Shang Zhu, Jue Wang, Junlin Wang et al.
Architecture
Mamba-3
Aakash Lahoti* (CMU), Kevin Y. Li* (CMU) et al.
Kernels
Key research and product announcements at the AI Native Conf
Kernels
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
Ted Zadouri (Princeton University et al.
Inference
Cache-aware prefill–decode disaggregation (CPD) for up to 40% faster long-context LLM serving
Jiejing Zhang, Yubo Wang, Yinghui Liu et al.
Agents
CoderForge-Preview: SOTA open dataset for training efficient coding agents
By Alpay Ariyak*, Junda Zhang et al.
Agents
How speech models fail where it matters the most and what to do about it
Kaitlyn Zhou, Martijn Bartelds et al.
Inference
Consistency diffusion language models: Up to 14x faster inference without sacrificing quality
Minseo Kim, Chenfeng Xu et al.
Agents
What do LLMs think when you don't tell them what to think about?
Yongchan Kwon and James Zou
Agents
DSGym: A holistic framework for evaluating and training data science agents
Fan Nie, Junlin Wang, Harper Hua et al.
Kernels
Research POV: Yes, AGI Can Happen – A Computational Perspective
Together AI
Model Shaping
How to run TorchForge reinforcement learning pipelines in the Together AI Native Cloud
Together AI Training and Research et al.
Model Shaping
Introducing AutoJudge: Streamlined inference acceleration via automated dataset curation
Roman Garipov, Fedor Velikonivtsev et al.
Agents
Large Reasoning Models Fail to Follow Instructions During Reasoning: A Benchmark Study
Yongchan Kwon, Shang Zhu, Federico Bianchi et al.
Inference
AdapTive-LeArning Speculator System (ATLAS): A New Paradigm in LLM Inference via Runtime-Learning Accelerators
Junxiong Wang, Shirley Wu, Zelei Shao et al.
Agents
How Together AI Uses AI Agents to Automate Complex Engineering Tasks: Lessons from Developing Efficient LLM Inference Systems
Shang Zhu, Federico Bianchi, Wai Tong Chung et al.
Agents
Back to The Future: Evaluating AI Agents on Predicting Future Events
Federico Bianchi, Junlin Wang, Zain Hasan et al.
Inference
DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL
Michael Luo*, Naman Jain*, Jaskirat Singh* et al.
Agents
From Zero to One: Building An Autonomous and Open Data Scientist Agent from Scratch
Federico Bianchi, Shang Zhu, Zain Hasan et al.
Inference
Model-Preserving Adaptive Rounding with YAQA
Albert Tseng, Zhaofeng Sun, and Chris De Sa
Agents
Mixture-of-Agents Alignment: Harnessing the Collective Intelligence of Open-Source LLMs to Improve Post-Training
Junlin Wang, Roy Xie, Shang Zhu, Jue Wang et al.
Inference
Boosting DeepSeek-R1’s Speed with Customized Speculative Decoding
Wai Tong Chung, Dan Waters, Avner May et al.
Kernels
Chipmunk: Training-Free Acceleration of Diffusion Transformers with Dynamic Column-Sparse Deltas
Austin Silveria, Soham Govande, Dan Fu
Model Shaping
Direct Preference Optimization: A Technical Deep Dive
Ivan Provilkov, Zain Hasan, Max Ryabinin
Model Shaping
Continued Fine-tuning of LLMs: A Technical Deep Dive
Artem Chumachenko, Zain Hasan et al.
Agents
Open Deep Research
Together AI
Inference
DeepCoder: A Fully Open-Source 14B Coder at O3-mini Level
Michael Luo*, Sijun Tan*, Roy Huang* et al.
Kernels
ThunderKittens Now Optimized for NVIDIA Blackwell GPUs
Benjamin Spector, Aaryan Singhal, Dan Fu et al.
Inference
Minions: embracing small LMs, shifting compute on-device, and cutting cloud costs in the process
Avanika Narayan*, Dan Biderman* et al.
Model Shaping
Long Context Fine-Tuning: A Technical Deep Dive
George Grigorev, Zain Hasan, Max Ryabinin
Model Shaping
Fine-Tuning LLMs for Multi-Turn Conversations: A Technical Deep Dive
Artem Chumachenko, Zain Hasan et al.
Inference
Even Better, Even Faster Quantized LLMs with QTIP
Albert Tseng, Qingyao Sun, David Hou et al.
Architecture
Linearizing LLMs with LoLCATs
Michael Zhang, Simran Arora et al.
Applications
Multimodal Document RAG with Llama 3.2 Vision and ColQwen2
Zain Hasan
Architecture
The Mamba in the Llama: Distilling and Accelerating Hybrid Models
Junxiong Wang, Daniele Paliotta, Avner May et al.
Inference
Speculative decoding for high-throughput long-context inference
Jian Chen, Vashisth Tiwari, Ranajoy Sadhukhan et al.
Inference
TEAL: Training-Free Activation Sparsity in Large Language Models
James Liu, Pragaash Ponnusamy, Tianle Cai et al.
Kernels
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Jay Shah (Colfax Research) et al.
Applications
Building a personalized code assistant with open-source LLMs using RAG Fine-tuning
Kezhen Chen, Linda He, Ben Athiwaratkun et al.
Inference
SpecExec: Massively Parallel Speculative Decoding for Interactive LLM Inference on Consumer Devices
Ruslan Svirschevski, Avner May et al.
Agents
Together MoA — collective intelligence of open-source models pushing the frontier of LLM capabilities
Junlin Wang, Jue Wang, Ben Athiwaratkun et al.
Architecture
Dragonfly: A large vision-language model with multi-resolution zoom
Kezhen Chen, Rahul Thapa, Rahul Chalamala et al.
Kernels
ThunderKittens: A Simple Embedded DSL for AI kernels
Benjamin Spector, Aaryan Singhal et al.
Inference
FAQ: Building LLMs with RedPajama-v2, a 30 trillion token web dataset
Together AI
Inference
Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
Zhuoming Chen, Avner May et al.
Architecture
BASED: Simple linear attention language models balance the recall-throughput tradeoff
Simran Arora, Sabri Eyuboglu, Michael Zhang et al.
Architecture
Evo: Long-context modeling from molecular to genome scale
Eric Nguyen, Michael Poli, Matthew Durrant et al.
Inference
BitDelta: Your Fine-Tune May Only Be Worth One Bit
James Liu, Guangxuan Xiao, Kai Li et al.
Inference
Long context retrieval models with Monarch Mixer
Jon Saad-Falcon, Dan Fu, Simran Arora
Architecture
Mamba-3B-SlimPJ: State-space models rivaling the best Transformer architecture
Tri Dao, Albert Gu
Architecture
Paving the way to efficient architectures: StripedHyena-7B, open source models offering a glimpse into a world beyond Transformers
Together
Kernels
FlashFFTConv: Efficient Convolutions for Long Sequences with Tensor Cores
Dan Fu, Hermann Kumbong, Eric Nguyen et al.
Inference
RedPajama-Data-v2: An open dataset with 30 trillion tokens for training large language models
Together
Kernels
Flash-Decoding for long-context inference
Tri Dao, Daniel Haziza, Francisco Massa et al.
Inference
Medusa: Simple framework for accelerating LLM generation with multiple decoding heads
Tianle Cai*, Yuhong Li*, Zhengyang Geng et al.
Model Shaping
Llama-2-7B-32K-Instruct — and fine-tuning for Llama-2 models with Together API
Together
Inference
Faster inference enables up to 5x price reduction on Together API
Together
Inference
Preparing for the era of 32K context: Early learnings and explorations
Together
Architecture
Monarch Mixer: A new model architecture for increased efficiency
Dan Fu, Simran Arora, Chris Ré
Model Shaping
Fine-tuning language models over slow networks using activation compression with guarantees
Jue Wang, Binhang Yuan, Luka Rimanic et al.
Model Shaping
Decentralized training of foundation models in heterogeneous environments
Binhang Yuan, Yongjun He et al.
Kernels
FlashAttention: Fast and memory-efficient exact attention with IO-Awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon et al.
Model Shaping
CocktailSGD: Fine-tuning foundation models over 500Mbps networks
Jue Wang, Binhang Yuan, Luka Rimanic et al.
Inference
FlexGen: High-throughput generative inference of large language models with a single GPU
Ying Sheng, Lianmin Zheng, Binhang Yuan et al.
Architecture
Hyena Hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen et al.
Kernels
FlashConv: Speeding up state space models
Dan Fu and Tri Dao
Architecture
Hungry Hungry Hippos: Towards language modeling with state space models
Daniel Y. Fu, Tri Dao, Khaled K. Saab et al.
Model Shaping
NeurIPS 2022: Overcoming communication bottlenecks for decentralized training (2/2)
Together
Model Shaping
NeurIPS 2022: Overcoming communication bottlenecks for decentralized training (1/2)
Together
Inference
HELM: benchmarking large language models on the Together Research Computer
Together
recognized by
-
-
- -
AI natives build on Together AI
See how Together AI powers customers building the next generation of AI products.
\ \ How Cursor partnered with Together AI to deliver real-time, low-latency inference at scale\ \ Inference, GPU clusters, RESEARCH • Enterprise\ \ \ \ How Decagon Engineered Sub-Second Voice AI with Together AI\ \ Inference, GPU clusters, RESEARCH\ \ 6x\ \ Cost reduction per turn vs. gpt-5 mini\ \ \ \ 11x\ \ Faster Inference\ \
What’s new at Together AI
\ \ Company\ \ Announcing our $800M Series C to accelerate the shift to open-source AI\ \ We raised $800M to accelerate the shift to open-source AI. Here's why the economics of closed models don't scale, and what we're building next.](/content/blog/announcing-our-series-c/index.html)
\ \ GPU Clusters\ \ Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community \ \ No more two-year compute contracts. Together AI and YC just gave YC startups a faster way to get GPUs.](/content/blog/together-yc-gpu-cluster/index.html)
\ \ Model Library\ \ DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding\ \ We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.](/content/blog/deepseek-v4-flash-0731-vs-gpt-5-6-luna-on-deepswe-cost-and-coding/index.html)
\ \ Research\ \ Together AI at ICML 2026: frontier research across the full stack\ \ Nine papers at ICML 2026 across the full stack. The research that becomes the Together platform. Find us at booth B714 in Seoul.](/content/blog/icml-2026/index.html)