Announcing Together Inference Engine 2.0 with new Turbo and Lite endpoints
Announcing Together Inference Engine 2.0 with new Turbo and Lite endpoints
Today we are announcing a new inference stack, which provides decoding throughput 4x faster than open-source vLLM, and outperforms commercial solutions including Amazon Bedrock, Azure AI, Fireworks, and Octo AI by 1.3x to 2.5x. The Together Inference Engine achieves over 400 tokens per second on Meta Llama 3 8B by building on the latest advances from Together AI, including FlashAttention-3, faster GEMM & MHA kernels, innovations in quality-preserving quantization, and speculative decoding.
We are also introducing new Together Turbo and Together Lite endpoints, starting today with Meta Llama 3, and expanding to other models soon. Together Turbo and Together Lite enable performance, quality, and price flexibility so enterprises do not have to compromise. They offer the most accurate quantization available on the market; with Together Turbo closely matching the quality of full-precision FP16 models. These advances make Together Inference the fastest engine for Nvidia GPUs, and the most accurate and cost-efficient solution to build with Generative AI at production scale.
Today, over 100,000 developers and companies like Zomato, DuckDuckGo, and the Washington Post build and run their Generative AI applications on the Together Inference Engine.
Today’s release includes:
- Together Turbo endpoints, providing fast FP8 performance while maintaining quality, closely matching FP16 reference models and exceeding other FP8 solutions on AlpacaEval 2.0 by up to 1.9 points (length-corrected win rate) and up to 2.5 points (win rate) – making them the most accurate, cost-efficient, and performant models available. Together Turbo endpoints are available at $0.18 for 8B and $0.88 for 70B, 17x lower cost than GPT-4o.
- Together Lite endpoints, which leverage several optimizations including INT4 quantization, provide the most cost-efficient and scalable Llama 3 models available anywhere, while maintaining excellent quality relative to full precision reference implementations. With Together Lite high-quality AI models are now more affordable than ever, with Llama 3 8B Lite priced at $0.10 per million tokens, 6x lower cost than GPT-4o-mini.
- Together Reference endpoints, the fastest full-precision FP16 support for Meta Llama 3 models at up to 4x faster performance than vLLM.
Together Turbo endpoints
Together Turbo, our new flagship endpoints, provide the best combination of performance, quality, and cost-efficiency. Together Turbo endpoints leverage the most accurate quantization techniques and proprietary innovations to provide leading performance without compromising quality.
Quality
The below figures provide a comprehensive quality analysis of the Together Turbo endpoints on three leading quality benchmarks: HELM classic, WikiText/C4, and AlpacaEval 2.0. We compare Together Turbo models against our full precision reference implementation, as well as two commercial model providers, Fireworks and Groq, that provide quantized models by default.
Performance
Together Turbo provides up to 4.5x performance improvement over vLLM (version 0.5.1) on Llama-3-8B-Instruct and Llama-3-70B-Instruct. Across common inference regimes with different batch sizes and context lengths, Together Turbo consistently outperforms vLLM.
Cost efficiency
Together Turbo endpoints are available through the Together API as serverless endpoints at more than 10x lower cost than GPT-4o, and provide significant efficiency benefits for customers hosting their own dedicated endpoints on the Together Cloud.
Together Lite endpoints
Together Lite endpoints are designed for applications demanding fast performance and high capacity at the lowest cost.
In summary, Together Lite endpoints provide a highly economical solution with a modest compromise in quality.
Together Reference endpoints
Quality has always been at the heart of the Together Inference solution. Together Reference endpoints provide full precision FP16 quality consistent with model providers’ base implementations and reference architectures.
This combination of full reference quality, with high performance, meets the needs of the most demanding enterprises.
Together Inference Engine technical advances
The Together inference solution is a dynamic system that continuously incorporates cutting-edge innovations from both the community and our in-house research. These advancements span various areas, which include kernels (e.g., FlashAttention-3 and FlashDecoding), models and architectures (e.g., Mamba and StripedHyena), quality-preserving quantization, speculative decoding, and other innovative runtime and compiler techniques.
As a research-focused company, we will continue to push the envelope of AI acceleration. The Together Inference Engine is built for extensibility and rapid iteration, enabling us to quickly add support for new models, techniques, and kernels.
Together, we hope these innovations give you the flexibility to scale your applications with the performance, quality, and cost-efficiency your business demands.