Serverless Inference | Together AI

Serverless Inference

The fastest way to run open source models on demand

High-performance inference, powered by our in-house research. No infrastructure to manage, no long-term commitments.

Get started Explore demos

Serverless inference on Together AI

Access all the top open-source models in one place.

Build with leading models

Explore top-performing models across text, image, video, code, and voice.

Browse models Deploy own model

Model Name Description
DeepSeek V4 Pro Chat Model
MiniMax M3 Chat Model
GLM-5.2 Chat Model
ByteDance Seedance 2.0 Video Model
Kimi K3 Chat Model
Inkling Small Chat Model

Key capabilities, purpose built for AI natives

Call any leading open-weight model through a single API, with nothing to deploy or manage, and pay only for the tokens you use.

Deployment options

Run models using different deployment options depending on latency needs, traffic patterns, and infrastructure control.

Production-grade security and data privacy

We take security and compliance seriously, with strict data privacy controls to keep your information protected. Your data and models remain fully under your ownership, safeguarded by robust security measures.

Customers running inference in production

Vercept

"Together has helped us deploy VyUI, our state-of-the-art computer AI model."

The Washington Post

"Together AI offers optimized performance at scale, and at a lower cost than closed-source providers."

Salesforce AI Research

"We’ve been thoroughly impressed with Together. They delivered a 2x reduction in latency and cut our costs by approximately a third."

View All Stories

Run inference with Serverless Endpoints

High-performance inference, powered by our in-house research.

Explore models Contact sales