# together.ai > AI-optimized mirror of together.ai containing 50 pages totalling 44,127 words of clean markdown content, structured data, and semantic HTML. Original source: https://together.ai. Last updated: 2026-08-09T03:02:56.454Z. Each page is available as HTML (with JSON-LD structured data) and Markdown (text-only, ideal for LLMs and RAG). ## Homepage - [Together AI | The AI Native Cloud](/content/site-root.html): Build what's next on the AI Native Cloud. Full-stack AI platform for inference, fine-tuning, and GPU clusters — powered by cutting-edge research. (2,081 words) ## Articles & Blog Posts - [Mamba-3](/content/blog/mamba-3/index.html): Meet Mamba-3: the SSM built for inference. Faster than Transformers at decode, stronger than Mamba-2, and open-source from day one. (2,750 words) - [FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling](/content/blog/flashattention-4/index.html): As GPU throughput outpaces memory bandwidth, kernels must evolve. We introduce FlashAttention-4, featuring new pipelining for maximum overlap, 2-CTA MMA modes to reduce shared memory traffic, and a hardware-software hybrid approach to softmax exponentials. (2,479 words) - [Qwen3.5-397B-A17B API | Together AI](/content/models/qwen3-5-397b-a17b/index.html): Native multimodal model with 397B parameters (17B activated), 256K context, achieving 87.8% MMLU-Pro and 201-language support with hybrid Gated Delta Networks architecture. (896 words) - [From Zero to One: Building An Autonomous and Open Data Scientist Agent from Scratch](/content/blog/building-an-autonomous-and-open-data-scientist-agent-from-scratch.html): Build a data scientist agent using Together’s open-source models and Code Interpreter—easy to implement, solid benchmarks, and full code on GitHub. (2,458 words) - [Cache-aware prefill–decode disaggregation (CPD) for up to 40% faster long-context LLM serving](/content/blog/cache-aware-disaggregated-inference/index.html): Serving long prompts doesn't have to mean slow responses. Learn how Together AI's CPD architecture separates warm and cold inference workloads to deliver 40% higher throughput and dramatically lower time-to-first-token for long-context LLM serving. (2,062 words) - [Accelerated Compute | Together AI](/content/accelerated-compute/index.html): Train, fine-tune, and deploy on self-service GPU clusters optimized by frontier research — with flexible pricing, production reliability, and security built in. (600 words) - [MiniMax M2.5 API | Together AI](/content/models/minimax-m2-5/index.html): SOTA coding and agentic model achieving 80.2% SWE-Bench Verified, trained on 200K+ real-world environments with architect-level planning, full-stack development, and office deliverables at 10% the cost of frontier models. (800 words) - [How to evaluate and benchmark Large Language Models (LLMs)](/content/blog/evaluate-and-benchmark-llms/index.html): Understanding how to evaluate and benchmark Large Language Models (LLMS). Test, compare, and understand LLMs. (2,004 words) - [Back to The Future: Evaluating AI Agents on Predicting Future Events](/content/blog/futurebench/index.html): FutureBench is a live, leak-free benchmark of true reasoning—AI agents forecast real-world events (rates, geopolitics) before they happen. (1,857 words) - [Configuring Dedicated Model Inference](/content/blog/configuring-dedicated-model-inference/index.html): The three-part resource model behind Together AI Dedicated Model Inference—endpoints, deployments, configs—and how capacity-aware routing ties them together. (1,705 words) - [Dedicated Container Inference | Together AI](/content/dedicated-container-inference/index.html): GPU infrastructure purpose-built for generative media. Deploy video, audio, and avatar models with proven autoscaling and up to 2.6x speedup. (358 words) - [Custom Training: RL and SFT for Open Models | Together AI](/content/custom-training/index.html): Post-train frontier open models with RL and SFT, full-weight or LoRA. Ship the highest-quality custom models through fast, large-scale experimentation, all the way to production. (329 words) - [Research | Together AI](/content/research/index.html): Together Research builds foundational AI systems for production — kernels, inference, model shaping, and agents. Published at top ML conferences. (496 words) - [DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding](/content/blog/deepseek-v4-flash-0731-vs-gpt-5-6-luna-on-deepswe-cost-and-coding.html): We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar. (1,574 words) - [Blog | Together AI](/content/blog/index.html): Explore Together AI's latest product launches, research papers, and technical guides on AI inference, fine-tuning, and GPU infrastructure. (254 words) - [Kimi K3: The Complete Developer Guide](/content/blog/kimi-k3-guide/index.html): Kimi K3 is the first open 3T-class model. See how it benchmarks, what it costs, and how to call it on the Together AI API, with copy-paste code examples. (1,016 words) - [Dedicated Model Inference | Together AI](/content/lp/dedicated-inference/index.html): Deploy models on dedicated inference endpoints engineered for speed, control, and best-in-class unit economics — backed by Together's frontier AI research. (324 words) - [Optimizing inference speed and costs: Lessons learned from large-scale deployments](/content/blog/optimizing-inference-speed-and-costs/index.html): Learn how to reduce inference latency without massive cost using proven inference optimization tactics — improving throughput, GPU utilization, and cost efficiency while balancing throughput vs. latency tradeoffs. (1,239 words) - [Kimi K3 vs GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing](/content/blog/kimi-k3-vs-gpt-5-6-sol-on-deepswe-cost-coding-and-routing.html): We ran 904 DeepSWE rollouts on Kimi K3 and GPT-5.6 Sol. Sol leads pass@1; Kimi K3 wins pass@4 at 2.8x the solves per dollar, and routing between them reaches ~85.6%. (778 words) - [Batch Inference | Together AI](/content/batch-inference/index.html): Process massive AI workloads asynchronously at up to 50% less cost. Scale to 30 billion tokens per model with any serverless model or private deployment. (535 words) - [When Standard Inference Frameworks Failed, Together AI Enabled 5x Performance Breakthrough](/content/customers/vercept/index.html): See how Vercept used Together AI to deliver 5× computer automation performance compared to OpenAI, with seamless autoscaling and 30% cost savings. (708 words) - [NVIDIA HGX H200 Cluster Pricing & Specs | Rent HGX H200 GPUs | Together AI](/content/gpu/nvidia-h200/index.html): Rent NVIDIA H200 GPU clusters on Together AI. 141GB HBM3e, 4.8TB/s bandwidth, on-demand or reserved. See H200 cluster pricing, specs, and availability as of July 2026. (593 words) - [Voice | Together AI](/content/solutions/voice/index.html): Build agents humans want to talk to. Combine the best STT, LLM, and TTS models on co-located infrastructure for ultra-low latency and production-scale reliability. (299 words) - [Events | Together AI](/content/events/index.html): Connect with Together AI's team of researchers and engineers at upcoming conferences, meetups, and webinars. (363 words) - [Together AI Announces $305M Series B to Scale AI Acceleration Cloud for Open Source and Enterprise AI](/content/blog/together-ai-announcing-305m-series-b/index.html) (677 words) - [How Yutori runs browser-use AI agents at production scale on Together AI’s inference platform](/content/customers/yutori/index.html): Yutori runs browser-use agents at production scale on Together AI's inference platform, delivering 2x faster per-step latency and 4-5x lower cost than frontier models, with elastic scaling for always-on consumer and developer workloads. (1,349 words) - [NVIDIA GPU Clusters: H100, H200, B200, GB200 | Together AI](/content/gpu-clusters/index.html): Self-serve AI-ready GPU clusters at scale. H100, H200, B200, and GB200 with InfiniBand, managed orchestration, and flexible on-demand or reserved pricing. (701 words) - [gpt-oss-120B API | Together AI](/content/models/gpt-oss-120b/index.html): 120B parameters, 128K context, reasoning with chain-of-thought, MoE architecture, Apache 2.0 license (378 words) - [How Decagon Engineered Sub-Second Voice AI with Together AI](/content/customers/decagon/index.html): Decagon partnered with Together AI to deliver sub-second voice AI with NVIDIA Blackwell GPUs, achieving <400ms latency, 6× cost reduction, and a conversational experience customers thank the agent for. (919 words) - [Capacity without conflict: A guide to multi-tenant GPU cluster design for AI-native teams](/content/blog/multi-tenant-gpu-cluster-design-for-ai-native-teams.html): Learn how AI-native companies design multi-tenant GPU clusters that pool capacity without sacrificing team isolation — and how Together AI makes it work in practice. (1,287 words) - [Together AI acquires CodeSandbox to launch first-of-its-kind code interpreter for generative AI](/content/blog/codesandbox-acquisition-together-code-interpreter/index.html) (832 words) - [FLUX 3 API | Together AI](/content/models/flux-3/index.html): Black Forest Labs' multimodal video generation model with synchronized audio, producing up to 20-second multi-shot clips from text, images, or ordered keyframes, with video continuation. (621 words) - [Products | Together AI](/content/products/index.html): Build, fine-tune, and deploy open-source AI models — from inference to GPU clusters — on a single, production-ready platform. (204 words) - [Customer Stories | Together AI](/content/customers/index.html): See how AI-native companies build at production scale on Together AI — real stories on inference, fine-tuning, GPU clusters, and cost savings. (446 words) - [Serverless Inference | Together AI](/content/serverless-inference/index.html): The fastest way to run open-source models on demand. Up to 2.75x faster serverless inference — no infrastructure to manage, no long-term commitments., , (350 words) - [Bringing 100,000 GPUs to Europe](/content/blog/together-ai-expands-in-europe/index.html) (617 words) - [Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community](/content/blog/together-yc-gpu-cluster/index.html): No more two-year compute contracts. Together AI and YC just gave YC startups a faster way to get GPUs. (656 words) - [FLUX1.1 [pro] API | Together AI](/content/models/flux1-1-pro/index.html): Premium image generation model by Black Forest Labs. (116 words) - [Sandbox | Together AI](/content/sandbox/index.html): Use fast, secure code sandboxes at scale. Build full-scale AI development environments with instant VM creation, robust snapshotting, and flexible scaling. (370 words) - [FLUX.2 [pro] API | Together AI](/content/models/flux-2-pro/index.html): Production-grade image model supporting up to 4MP outputs, 8 reference images on API, 10 on Playground, with photorealistic detail, hex color accuracy, multi-reference character consistency (198 words) - [How to Build a State-of-the-Art Search Stack for LLMs: RAG, Reranking, and Reinforcement Learning](/content/blog/sota-search-stack-for-llms/index.html) (735 words) - [Announcing $106M round led by Salesforce Ventures](/content/blog/series-a2/index.html) (897 words) - [Fine-Tuning | Together AI](/content/fine-tuning/index.html): Fine-tune open-source models for real production use. Improve accuracy, reduce hallucinations, and control behavior — with support for 100B+ parameter models. (388 words) - [Announcing the Together AI Startup Accelerator, purpose-built for AI Native Apps](/content/blog/announcing-together-ai-startup-accelerator/index.html): We've launched the Together AI Startup Accelerator: Up to $50K credits, expert engineering hours, GTM support, community and VC access for AI-native apps in build–scale tiers. (672 words) - [Announcing Together Inference Engine 2.0 with new Turbo and Lite endpoints](/content/blog/together-inference-engine-2/index.html) (706 words) - [Dedicated Model Inference | Together AI](/content/dedicated-model-inference/index.html): Deploy models on dedicated inference endpoints engineered for speed, control, and best-in-class unit economics — backed by Together's frontier AI research. (510 words) - [Best practices to accelerate inference for large-scale production workloads](/content/guides/best-practices-to-accelerate-inference-for-large-scale-production-workloads.html): Accelerate large-scale LLM inference with four pillars: speculative decoding, optimized kernels, near-lossless compression, and traffic-aware infrastructure. (439 words) - [Autoscaling endpoints for LLM inference](/content/blog/autoscaling-endpoints-for-llm-inference/index.html): GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on dedicated inference. (202 words) - [Pricing | Together AI](/content/pricing/index.html): Transparent, flexible pricing across serverless inference, dedicated endpoints, fine-tuning, and GPU clusters. Start for free, scale on demand. (1,299 words) ## Resources - [Full Page Index](/index.html): Browse all cached pages with rich metadata - [About This Cache](/about.html): Methodology, technical details, and usage guidelines - [XML Sitemap](/sitemap.xml): Machine-readable sitemap for crawler discovery - [Robots.txt](/robots.txt): Crawler directives