Autoscaling endpoints for LLM inference
Autoscaling endpoints for LLM inference
Choosing scaling metrics, tuning windows, and budgeting for cold starts on dedicated inference.
Authors: Zain Hasan, SoYoung Park, Nikitha Suryadevara, Ted Cui
Table of contents
Summary
With Dedicated Model Inference on the Together AI platform, you can get your deployments to autoscale on metrics the inference engine understands. You can set replica bounds, pick a metric and target, and then tune two windows that control how eagerly it scales up and how patiently it scales down. Understanding and choosing the right metric is important because it impacts the latency your users will see. Below we'll cover how to choose the right metric to autoscale on.
Over- and under-provisioning are both expensive
With dedicated inference, you pay per replica-minute, making capacity planning a balance between two failure modes:
- Over-provision: you're paying for GPUs to sit at 15% utilization just to handle peak traffic when it arrives. Almost never viable.
- Under-provision: your p95 degrades sharply when traffic exceeds what your replicas can batch.
Just autoscale it