FLUX 3 API | Together AI

FLUX 3

Multi-shot video generation with native audio from text, images, or keyframes

Try now Read docs

About model

FLUX 3 is Black Forest Labs' video generation model, built on a unified multimodal architecture that generates video and synchronized audio together. It produces clips of up to 20 seconds with multiple scenes and camera angles in a single take, starting from text, an image, or ordered keyframes. Audio is optional and native: multilingual speech with strong lipsync and accurate accents, plus sound effects and ambience generated with the frames. The model spans styles well beyond cinematic output, renders text and typography that feels native to the scene, and supports video continuation, extending an existing clip from its final frame with momentum, framing, and scene logic carried forward. Built on Self-Flow, Black Forest Labs' approach to aligning multimodal generation and understanding within one architecture. Available on Together AI.

Max Clip Length

Multiple Scenes and Camera Angles

Audio Generation

Frame-Level Control

Model key capabilities

API usage




Model card

Architecture Overview:

Training Methodology:

Performance Characteristics:

Prompting

Together AI API Access:

Applications & use cases

Filmmaking & Storyboarding:

Advertising & Brand Content:

Explainers & Social Content:

Model details

Run in Playground
Quickstart docs
Deploy model