Announcing Together Inference Engine 2.0 with new Turbo and Lite endpoints

Announcing Together Inference Engine 2.0 with new Turbo and Lite endpoints

Today we are announcing a new inference stack, which provides decoding throughput 4x faster than open-source vLLM, and outperforms commercial solutions including Amazon Bedrock, Azure AI, Fireworks, and Octo AI by 1.3x to 2.5x. The Together Inference Engine achieves over 400 tokens per second on Meta Llama 3 8B by building on the latest advances from Together AI, including FlashAttention-3, faster GEMM & MHA kernels, innovations in quality-preserving quantization, and speculative decoding.

We are also introducing new Together Turbo and Together Lite endpoints, starting today with Meta Llama 3, and expanding to other models soon. Together Turbo and Together Lite enable performance, quality, and price flexibility so enterprises do not have to compromise. They offer the most accurate quantization available on the market; with Together Turbo closely matching the quality of full-precision FP16 models. These advances make Together Inference the fastest engine for Nvidia GPUs, and the most accurate and cost-efficient solution to build with Generative AI at production scale.

Today, over 100,000 developers and companies like Zomato, DuckDuckGo, and the Washington Post build and run their Generative AI applications on the Together Inference Engine.

Today’s release includes:

Together Turbo endpoints

Together Turbo, our new flagship endpoints, provide the best combination of performance, quality, and cost-efficiency. Together Turbo endpoints leverage the most accurate quantization techniques and proprietary innovations to provide leading performance without compromising quality.

Quality

The below figures provide a comprehensive quality analysis of the Together Turbo endpoints on three leading quality benchmarks: HELM classic, WikiText/C4, and AlpacaEval 2.0. We compare Together Turbo models against our full precision reference implementation, as well as two commercial model providers, Fireworks and Groq, that provide quantized models by default.

Performance

Together Turbo provides up to 4.5x performance improvement over vLLM (version 0.5.1) on Llama-3-8B-Instruct and Llama-3-70B-Instruct. Across common inference regimes with different batch sizes and context lengths, Together Turbo consistently outperforms vLLM.

Cost efficiency

Together Turbo endpoints are available through the Together API as serverless endpoints at more than 10x lower cost than GPT-4o, and provide significant efficiency benefits for customers hosting their own dedicated endpoints on the Together Cloud.

Together Lite endpoints

Together Lite endpoints are designed for applications demanding fast performance and high capacity at the lowest cost.

In summary, Together Lite endpoints provide a highly economical solution with a modest compromise in quality.

Together Reference endpoints

Quality has always been at the heart of the Together Inference solution. Together Reference endpoints provide full precision FP16 quality consistent with model providers’ base implementations and reference architectures.

This combination of full reference quality, with high performance, meets the needs of the most demanding enterprises.

Together Inference Engine technical advances

The Together inference solution is a dynamic system that continuously incorporates cutting-edge innovations from both the community and our in-house research. These advancements span various areas, which include kernels (e.g., FlashAttention-3 and FlashDecoding), models and architectures (e.g., Mamba and StripedHyena), quality-preserving quantization, speculative decoding, and other innovative runtime and compiler techniques.

As a research-focused company, we will continue to push the envelope of AI acceleration. The Together Inference Engine is built for extensibility and rapid iteration, enabling us to quickly add support for new models, techniques, and kernels.

Together, we hope these innovations give you the flexibility to scale your applications with the performance, quality, and cost-efficiency your business demands.