# Blog

[Announcing our $800M Series C to accelerate the shift to open-source AI](/content/blog/announcing-our-series-c/index.html)  
We raised $800M to accelerate the shift to open-source AI. Here's why the economics of closed models don't scale, and what we're building next.

[Together AI and Y Combinator partner to launch the first dedicated GPU cluster for the YC community](/content/blog/together-yc-gpu-cluster/index.html)  
No more two-year compute contracts. Together AI and YC just gave YC startups a faster way to get GPUs.

[DeepSeek-V4 Flash 0731 vs GPT-5.6 Luna on DeepSWE: Cost and Coding](/content/blog/deepseek-v4-flash-0731-vs-gpt-5-6-luna-on-deepswe-cost-and-coding/index.html)  
We ran 900 DeepSWE rollouts on DeepSeek-V4 Flash and GPT-5.6 Luna. Luna leads pass@1 by 14 points; DeepSeek delivers 4.8x the solves per dollar.

[Together AI at ICML 2026: frontier research across the full stack](/content/blog/icml-2026/index.html)  
Nine papers at ICML 2026 across the full stack. The research that becomes the Together platform. Find us at booth B714 in Seoul.

[Kimi K3: The Complete Developer Guide](/content/blog/kimi-k3-guide/index.html)  
Kimi K3 is the first open 3T-class model. See how it benchmarks, what it costs, and how to call it on the Together AI API, with copy-paste code examples.

[ThunderAgent: 2x Faster Agentic Inference for Synthetic Data Generation at Scale](/content/blog/thunderagent/index.html)  
ThunderAgent is a program-aware scheduler for agentic inference. By treating each agent workflow as a schedulable program, it eliminates KV cache thrashing to deliver more than 2x single-node throughput and near-linear multi-node scaling.

[Autoscaling endpoints for LLM inference](/content/blog/autoscaling-endpoints-for-llm-inference/index.html)  
GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on dedicated inference.
