# gpt-oss-120B

Advanced open reasoning model with enterprise-grade capabilities.

## About model

**Enterprise-Ready Open Reasoning:**

gpt-oss-120B delivers sophisticated chain-of-thought reasoning capabilities in a fully open model. Built with community feedback and released under Apache 2.0, this 120B parameter model provides transparency, customization, and deployment flexibility for organizations requiring complete data security & privacy control.

### Quickstart guides

- [Building a RAG Workflow](https://docs.together.ai/docs/building-a-rag-workflow)
- [Agent Workflows](https://docs.together.ai/docs/workflows)
- [Next.js Chat Quickstart](https://docs.together.ai/docs/nextjs-chat-quickstart)

### Performance benchmarks

| Model | FrontierMath Tier 4 | GPQA Diamond | HLE | SciCode | GDPval-AA | Terminal-Bench 2.1 | Agent Arena | FrontierCode | DeepSWE |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| <br>gpt-oss-120B |  | 78.2% | 18% | 39% | 15% | 26% |  |  |  |
| <br>Claude Fable 5 | 87.8% | 92.6% | 53% | 60% | 62% | 85% |  | 53.5% | 70% |
| <br>Claude Opus 5 | 73.2% | 93.2% | 53% | 56% | 68% | 89% | +16.4pp | 53.4% | 74% |
| <br>GPT-5.6 Sol | 82.9% | 94.1% | 47% | 56% | 61% | 88% |  | 47.5% | 73% |
| <br>Grok 4.5 | 24.4% | 93.1% | 40% | 54% | 51% | 82% | +4.2pp | 42.4% | 54% |
| <br>GPT-5.6 Luna | 61.0% | 91.1% | 37% | 53% | 54% | 81% |  | 39.8% | 67% |

### API usage

- cURL

```shell
  
  ```

- Python

```python
  
  ```

- Typescript

```typescript
  
  ```

### Model card

**Architecture Overview:**

- Mixture-of-Experts (MoE) architecture with SwiGLU activations
- Alternating attention layers between full context and sliding 128-token window
- Learned attention sink per-head for enhanced performance

**Training Methodology:**

- Comprehensive safety training and evaluation protocols
- Community feedback integration from global listening sessions
- Rigorous testing under Preparedness Framework
- Standard GPT-4o tokenizer with additional Harmony format tokens

**Performance Characteristics:**

- Native FP4 quantization for efficient inference
- 128K context window with RoPE positional encoding
- Chain-of-thought reasoning with adjustable effort levels

### Applications & use cases

**Enterprise Applications:**

- Complex reasoning and analysis tasks
- Research and development support
- Technical documentation generation
- Strategic planning and decision support

**Developer Use Cases:**

- Code generation and review
- API development and integration
- System architecture design
- Technical troubleshooting and debugging

**Industry Solutions:**

- Healthcare: Clinical decision support and medical research
- Finance: Risk analysis and regulatory compliance
- Legal: Contract analysis and legal research
- Education: Curriculum development and tutoring systems

**Deployment Scenarios:**

- On-premises infrastructure for data sovereignty
- Private cloud deployments for security compliance
- Custom fine-tuning for domain-specific applications
- Multi-modal integration with existing systems

### Model specifications

- **Provider:** OpenAI  
- **Type:** Reasoning, Chat  
- **Main use cases:** Chat, Small & Fast, Medium General Purpose  
- **Features:** JSON Mode  
- **Speed:** High  
- **Intelligence:** High  
- **Deployment:** Serverless, Dedicated  
- **Endpoint:** openai/gpt-oss-120b  
- **Parameters:** 120B  
- **Context length:** 128K  
- **Input price:** $0.15 / 1M tokens  
- **Output price:** $0.60 / 1M tokens  
- **Input modalities:** Text  
- **Output modalities:** Text  
- **Released:** August 4, 2025  
- **Last updated:** August 18, 2025  
- **External link:** [HuggingFace](https://huggingface.co/openai/gpt-oss-120b)  
- **Category:** Chat
