gpt-oss-120B API | Together AI
gpt-oss-120B
Advanced open reasoning model with enterprise-grade capabilities.
About model
Enterprise-Ready Open Reasoning:
gpt-oss-120B delivers sophisticated chain-of-thought reasoning capabilities in a fully open model. Built with community feedback and released under Apache 2.0, this 120B parameter model provides transparency, customization, and deployment flexibility for organizations requiring complete data security & privacy control.
Quickstart guides
Performance benchmarks
| Model | FrontierMath Tier 4 | GPQA Diamond | HLE | SciCode | GDPval-AA | Terminal-Bench 2.1 | Agent Arena | FrontierCode | DeepSWE |
|---|---|---|---|---|---|---|---|---|---|
gpt-oss-120B |
78.2% | 18% | 39% | 15% | 26% | ||||
Claude Fable 5 |
87.8% | 92.6% | 53% | 60% | 62% | 85% | 53.5% | 70% | |
Claude Opus 5 |
73.2% | 93.2% | 53% | 56% | 68% | 89% | +16.4pp | 53.4% | 74% |
GPT-5.6 Sol |
82.9% | 94.1% | 47% | 56% | 61% | 88% | 47.5% | 73% | |
Grok 4.5 |
24.4% | 93.1% | 40% | 54% | 51% | 82% | +4.2pp | 42.4% | 54% |
GPT-5.6 Luna |
61.0% | 91.1% | 37% | 53% | 54% | 81% | 39.8% | 67% |
API usage
- cURL
- Python
- Typescript
Model card
Architecture Overview:
- Mixture-of-Experts (MoE) architecture with SwiGLU activations
- Alternating attention layers between full context and sliding 128-token window
- Learned attention sink per-head for enhanced performance
Training Methodology:
- Comprehensive safety training and evaluation protocols
- Community feedback integration from global listening sessions
- Rigorous testing under Preparedness Framework
- Standard GPT-4o tokenizer with additional Harmony format tokens
Performance Characteristics:
- Native FP4 quantization for efficient inference
- 128K context window with RoPE positional encoding
- Chain-of-thought reasoning with adjustable effort levels
Applications & use cases
Enterprise Applications:
- Complex reasoning and analysis tasks
- Research and development support
- Technical documentation generation
- Strategic planning and decision support
Developer Use Cases:
- Code generation and review
- API development and integration
- System architecture design
- Technical troubleshooting and debugging
Industry Solutions:
- Healthcare: Clinical decision support and medical research
- Finance: Risk analysis and regulatory compliance
- Legal: Contract analysis and legal research
- Education: Curriculum development and tutoring systems
Deployment Scenarios:
- On-premises infrastructure for data sovereignty
- Private cloud deployments for security compliance
- Custom fine-tuning for domain-specific applications
- Multi-modal integration with existing systems
Model specifications
- Provider: OpenAI
- Type: Reasoning, Chat
- Main use cases: Chat, Small & Fast, Medium General Purpose
- Features: JSON Mode
- Speed: High
- Intelligence: High
- Deployment: Serverless, Dedicated
- Endpoint: openai/gpt-oss-120b
- Parameters: 120B
- Context length: 128K
- Input price: $0.15 / 1M tokens
- Output price: $0.60 / 1M tokens
- Input modalities: Text
- Output modalities: Text
- Released: August 4, 2025
- Last updated: August 18, 2025
- External link: HuggingFace
- Category: Chat