# Autoscaling endpoints for LLM inference

Choosing scaling metrics, tuning windows, and budgeting for cold starts on dedicated inference.

- Authors: Zain Hasan, SoYoung Park, Nikitha Suryadevara, Ted Cui

- Table of contents
  - [Over- and under-provisioning are both expensive](/content/blog/autoscaling-endpoints-for-llm-inference#over-and-under-provisioning-are-both-expensive/index.html)
  - [How it works](/content/blog/autoscaling-endpoints-for-llm-inference#how-it-works/index.html)
  - [Under the hood: metrics to autoscale on](/content/blog/autoscaling-endpoints-for-llm-inference#under-the-hood-metrics-to-autoscale-on/index.html)
  - [Idle shutdown and cold starts](/content/blog/autoscaling-endpoints-for-llm-inference#idle-shutdown-and-cold-starts/index.html)
  - [Edge cases](/content/blog/autoscaling-endpoints-for-llm-inference#edge-cases/index.html)
  - [Autoscaling the same load with various policies](/content/blog/autoscaling-endpoints-for-llm-inference#autoscaling-the-same-load-with-various-policies/index.html)
  - [Try it yourself!](/content/blog/autoscaling-endpoints-for-llm-inference#try-it-yourself/index.html)

## Summary

With Dedicated Model Inference on the Together AI platform, you can get your deployments to autoscale on metrics the inference engine understands. You can set replica bounds, pick a metric and target, and then tune two windows that control how eagerly it scales up and how patiently it scales down. Understanding and choosing the right metric is important because it impacts the latency your users will see. Below we'll cover how to choose the right metric to autoscale on.

## Over- and under-provisioning are both expensive

With dedicated inference, you pay per replica-minute, making capacity planning a balance between two failure modes:

- **Over-provision:** you're paying for GPUs to sit at 15% utilization just to handle peak traffic when it arrives. Almost never viable.
- **Under-provision:** your p95 degrades sharply when traffic exceeds what your replicas can batch.

Just autoscale it
