Custom Training: RL and SFT for Open Models | Together AI
Reinforcement Learning - now in beta
Post-training, all the way to production
RL and SFT, full-weight and LoRA, on frontier open models. The highest-quality custom models through fast, large-scale experimentation.
Why Together Custom Training
One platform that takes a model from first experiment to production, without leaving the stack.
- State-of-the-art quality
- Cost-efficient at scale
- End-to-end platform
Choose your training method
Both run through a granular Python SDK, with high configurability over every training job.
Reinforcement learning
Optimize your model against a reward.- GRPO and custom losses
- Reward and KL tracking per run
- Frontier open models, full-weight or LoRA
Supervised fine-tuning
Teach the model from your own demonstrations.- Your data, your formats
- Instruction and conversation tuning
- Same SDK, same deployment path
Everything you need to reach production
Your code defines the run. Full configurability, dedicated capacity, and one path from training to serving.
- High-scale experimentation
- Parallel adapters
- One deployment
- Fast test loop
Run many LoRA adapter experiments in parallel inside one training deployment. Converge on your best model in one fast test loop.
Training precision
We match the computations between training and inference and minimize any discrepancies, keeping training stable even for the largest runs. Reach the strongest version of your model.Dedicated capacity
Your experiments run on capacity that is completely yours. Predictable performance and predictable cost, with your code defining the run.End-to-end platform
Iterate on experiments, deploy multiple versions, and promote the best to production, with a dashboard tracking every run. One stack carries the model straight to serving.
Backed by frontier research
Every run sits on Together's own training and inference research.
- Raro
- Upipe
- FFT Optimizer
Accuracy (%) on DeepMath reasoning benchmark
- Throughput (TPS)
UPipe vs other SOTA Approaches
82.5% less memory
FFT Optimizer results
25% less memory
security and data privacy
We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.
preferred partner
- SOC 2 Type II
- ISO 27001:2022