Research | Together AI
Foundational research for production AI
Our research areas
Inference
Design and optimization of production inference systems, spanning scheduling, batching, and hardware–software co-design for reliable high throughput.
Read papersKernels
Development of high-performance GPU kernels for training and inference, optimizing memory, attention, and custom operators at production scale.
Read papersModel Shaping
Advancement of post-training methods like fine-tuning, distillation, and quantization to shape efficient, controllable model behavior.
Read papersAgents
Studies of long-horizon reasoning and decision-making, focusing on tool use, multi-step planning, and reinforcement learning for reliable agentic systems.
Read papers
Recognized research
Papers accepted at top conferences
ICML
DSGym: A Holistic Framework for Evaluating and Training Data Science Agents
Fan Nie, Junlin Wang, Harper Hua, Federico Bianchi et al.
Read MoreThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System
Hao Kang, Ziyang Li, Weili Xu, Xinyu Yang, Yinfang Chen et al.
Read MoreLearning to Discover at Test Time (TTT-Discover)
Mert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi et al.
Read MoreEscaping the Verifier: Learning to Reason via Demonstrations (RARO)
Locke Cai, Max Ryabinin, Ivan Provilkov
Read MoreV1: Unifying Generation and Self-Verification for Parallel Reasoners
Harman Singh, Xiuyu Li, Kusha Sareen et al.
Read MoreWhen RL Meets Adaptive Speculative Training: A Unified Training-Serving System (Aurora)
Junxiong Wang, Fengxiang Bie, Jisen Li, Zhongzhu Zhou et al.
Read MoreUntied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking
Ravi Ghadia, Maksim Abraham, Sergei Vorobyov, Max Ryabinin
Read MoreOpportunistic Expert Activation: Batch-Aware Expert Routing for Faster Decode Without Retraining (OEA)
Costin-Andrei Oncescu, Qingyang Wu, Wai Tong Chung et al.
Read MoreParallelKernelBench: Benchmarking LLMs on Multi-GPU Kernel Generation
Willy Chan, Nathan Paek, Simon Guo, Simran Arora, Daniel Y. Fu.
Read More
Spotlight · ICLR
- ThunderKittens: Simple, Fast, and Adorable AI Kernels
Benjamin F. Spector, Simran Arora, Aaryan Singhal, Daniel Y. Fu et al.
Read More
Outstanding Paper · COLM
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Albert Gu, Tri Dao
Read More
Best Paper · ICML HAET
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, Christopher Ré
Read More
Oral · ICLR
- Mamba-3: Improved Sequence Modeling using State Space Principles
Aakash Lahoti, Kevin Y. Li, Berlin Chen, Caitlin Wang, Aviv Bick et al.
Read More
Key open-source projects
FlashAttention
IO-aware exact attention, universally adopted
Read MoreFlash Decoding
8× faster long-context token generation
Read MoreMixture of Agents
Open models, working together, beat GPT-4o
Read MoreDragonfly
Tiny 8B model beats Med-Gemini on every benchmark
Read MoreRed Pajama Datasets
100T+ tokens powering 500+ models
Read MoreDeepCoder
First open model to match o3-mini on code
Read MoreOpen Deep Research
Open-source multi-model deep research agent
Read MoreOpen Data Scientist Agent
Autonomous agent tops Adyen's real-world benchmark
Read More
Featured talks and conference presentations by our researchers
- At Slush 2025, Together AI VP of Kernels Dan Fu dives into building, using, and managing AI agents.
(Video available)
Research team
Researchers and engineers pushing the boundaries of AI
- Ce Zhang - Founder & CTO
- Chris Ré - Founder
- Tri Dao - Founder & Chief Scientist
- Percy Liang - Founder
- Ben Athiwaratkun - Core ML
- Dan Fu - Kernels
- James Zou - Frontier Agents
- Max Ryabinin - Model Shaping
- Simran Arora - Kernels
- Yineng Zhang - Inference