Autoscaling endpoints for LLM inference

Autoscaling endpoints for LLM inference

Choosing scaling metrics, tuning windows, and budgeting for cold starts on dedicated inference.

Summary

With Dedicated Model Inference on the Together AI platform, you can get your deployments to autoscale on metrics the inference engine understands. You can set replica bounds, pick a metric and target, and then tune two windows that control how eagerly it scales up and how patiently it scales down. Understanding and choosing the right metric is important because it impacts the latency your users will see. Below we'll cover how to choose the right metric to autoscale on.

Over- and under-provisioning are both expensive

With dedicated inference, you pay per replica-minute, making capacity planning a balance between two failure modes:

Just autoscale it