gpt-oss-120b

gpt-oss-120b API in Europe

The production workhorse - built for agents, not editors.

OpenAI's open-weight model, measured at 713 tokens per second on our EU infrastructure in Munich. Reliable production performance without flagship costs.

OpenAI Quality, Open-Weight Freedom

gpt-oss-120b is OpenAI's first open-weight model - Apache 2.0 licensed, designed for production agentic workloads. It's not the flashiest model, but it's the one you can rely on day after day.

Built-in Reasoning

Chain-of-thought reasoning with adjustable effort levels - optimize for speed or accuracy per task.

Production-Ready

Matches GPT-4o on most tasks. Beats it on reasoning-heavy benchmarks.

Best Value

Best price-to-intelligence ratio per Artificial Analysis.

Efficient by Design

Total Parameters117B
Active Parameters5.1B per forward pass
ArchitectureMixture of Experts (MoE)
Experts128 experts, Top-4 routing per token
Layers36
Context Length128K tokens
LicenseApache 2.0
EU HostedFastest Model

Measured on EU Infrastructure

Output Throughput
713tok/s
Time to First Token
388ms
End-to-End Latency
1.789s
Context Length
128Ktokens

10K input / 1K output, 1 concurrent, 10 requests

Up to 772 tok/s on shorter prompts. Last measured: July 2026.

See full benchmarks and methodology

Why It's So Fast

The MoE architecture means you get 117B model quality while only running 5.1B parameters per request - that's why it's so fast.

  • 22x fewer active parameters per inference
  • Lower memory bandwidth requirements
  • Expert routing optimized for each token
  • Same quality, fraction of the compute
The architecture behind 713 tok/s →

Not for Developers. For Agents.

"If you're building a public-facing AI agent, gpt-oss is your best bet - it's the best privately hostable model that functions on a single high-end GPU in production."

- Tigris

Reasoning Control

Adjust thinking effort (low/medium/high) per task

Function Calling

Native tool use for agentic workflows

Structured Outputs

JSON mode for reliable parsing

Web Browsing

Built-in capability for research agents

Navigate websites, extract data, and perform multi-step research tasks autonomously.

Code Execution

Python execution for data analysis agents

Run Python in a sandboxed environment for data processing, calculations, and analysis.

The Right Model for the Right Task

Not every request needs your most expensive model. Smart teams use gpt-oss-120b as part of a multi-model strategy.

"The technical quality is undeniable, and the chain-of-thought reasoning system is genuinely innovative in the open-weight space."

- Apatero (2026 Review)

Balanced Mode

In balanced mode: Matches GPT-4o on most tasks

Deep Mode

In deep mode: Beats GPT-4o on reasoning (MATH, HumanEval)

Cost Efficiency

At a fraction of the cost of proprietary models

ScenarioModel Choice
Complex reasoninggpt-oss-120b (high effort)
Standard tasksgpt-oss-120b (medium effort)
Simple queriesgpt-oss-120b (low effort)
Premium tasksMiniMax M2.7 Ultraspeed

"We optimized workflows twice: once for accuracy + latency, and once for accuracy + cost-capturing the tradeoffs that matter most in real-world deployments."

- DataRobot

OpenAI Open-Weight on EU Infrastructure

Run OpenAI's open-weight model without sending data to the US:

  • Hosted in Germany on Infercom-owned infrastructure
  • Full GDPR compliance with EU-based DPA
  • No US CLOUD Act exposure
  • ISO 27001 certified
  • Apache 2.0 license - full freedom to deploy
ISO 27001 Certified
GDPR Compliant
German Datacenter
Apache 2.0 Licensed

Start Building in Minutes

quickstart.py
from openai import OpenAI

client = OpenAI(
    api_key="your-infercom-key",
    base_url="https://api.infercom.ai/v1"
)

response = client.chat.completions.create(
    model="gpt-oss-120b",
    messages=[{"role": "user", "content": "Your prompt here"}],
    max_tokens=4096
)

print(response.choices[0].message.content)

OpenAI-compatible API. Drop-in replacement for your existing code.

Pay only for the tokens you use.

gpt-oss-120b: Frequently Asked Questions

How fast is gpt-oss-120b on Infercom?

713 tokens per second output throughput and 388 ms time to first token, both server-side p50 measured at 10K input / 1K output on our production API in Munich. Shorter prompts peak at up to 772 tok/s. That is up to 10x faster than GPU-based alternatives, and you can reproduce every figure with our open-source benchmark tool.

See the full benchmarks
Is gpt-oss-120b open-weight?

Yes. OpenAI released gpt-oss-120b under the Apache 2.0 licence, so the weights are published and you are free to deploy, fine-tune and run the model commercially without a separate agreement. Running it through Infercom simply means we host and operate it for you on EU infrastructure.

What open-weight actually means
Can I run gpt-oss-120b GDPR-compliant in the EU?

Yes. Infercom serves gpt-oss-120b from Infercom-owned hardware in a Tier III+ datacenter in Munich, Germany. Requests are processed inside EU jurisdiction with no US CLOUD Act exposure, the infrastructure is ISO 27001 certified, and an EU data processing agreement is available on request. We do not train on your data.

How EU sovereignty works here
Is the API OpenAI-compatible?

Yes. Point the OpenAI SDK at https://api.infercom.ai/v1, use your Infercom API key and pass gpt-oss-120b as the model name. Code that already speaks the OpenAI chat completions API works without changes, including streaming, function calling and structured output.

Read the API documentation
What is the context window of gpt-oss-120b?

128K tokens (131,072), shared between your prompt and the model's response. That is room for large codebases, long agent traces and multi-document prompts in a single request.

What a context window is
What does gpt-oss-120b cost?

Pay-per-token with no minimums and no monthly commitment: EUR 0.22 per million input tokens and EUR 0.59 per million output tokens (excl. VAT). You are billed only for what you use, and live usage and spend are visible in the cloud portal.

See full pricing

Ready to Build with Enterprise-Grade AI?

Start with a pilot, scale to production. Record-breaking performance with dedicated enterprise support.