MiniMax M2.7 API in Europe
229B frontier reasoning at 428 tokens per second. Fully EU sovereign.
The latest 229B parameter frontier model from MiniMax, running on Infercom's fully EU sovereign infrastructure in Germany. No US hyperscalers. No CLOUD Act exposure. Native multi-agent support, 30% coding improvement over the previous generation.
New in M2.7
M2.7 adds built-in self-critique for better first-attempt results, plus native multi-agent orchestration as a core capability.
Built-in Self-Critique
New in M2.7: The model automatically reviews and refines its outputs before responding. Better first-attempt results on complex coding and reasoning tasks.
Native Agent Teams
New in M2.7: Multi-agent collaboration as a core model capability - not just prompting. Stable role boundaries, adversarial reasoning, and behavioral differentiation built in.
Improved: Software Engineering
56.22% on SWE-Pro (matching GPT-5.3-Codex), 55.6% on VIBE-Pro for full project delivery, 57.0% on Terminal Bench 2 for complex system understanding.
Improved: Professional Skills
97% skill compliance across 40+ complex professional skills. Enhanced Excel, PowerPoint, and Word editing with multi-round revision support.
Measured on Infercom EU Infrastructure
Output Throughput
428tok/s
Time to First Token
690ms
Context Window
192Ktokens
Server-side p50, 10K input / 1K output, single request
Up to 444 tok/s on shorter prompts. Last measured: July 2026.
See full benchmarks and methodologyDeep dives in our glossary
Frontier Benchmark Performance
M2.7 achieves top-tier scores across coding, engineering, and ML benchmarks - matching or exceeding GPT-5.3-Codex on SWE-Pro.
SWE-Pro
56.22%
Matches GPT-5.3-Codex
SWE Multilingual
76.5%
Real-world engineering
VIBE-Pro
55.6%
Near Opus 4.6 level
Terminal Bench 2
57.0%
Deep system understanding
MLE Bench Lite
66.6%
Medal rate (9 gold, 5 silver)
Multi SWE Bench
52.7%
Complex codebases
Built for Agentic Workflows
M2.7 excels at long-horizon agent tasks that require autonomous decision-making, tool use, and multi-step reasoning across complex professional domains.
Financial Workflows
End-to-end research, Excel modeling, report generation
SRE & DevOps
Log analysis, incident response, production debugging
Document Processing
Word, Excel, PowerPoint with multi-round editing
ML Competitions
66.6% medal rate on MLE Bench Lite (9 gold, 5 silver)
Agentic Coding
56.22% on SWE-Pro
Matches GPT-5.3-Codex on the most demanding software engineering benchmark. Works with Aider, OpenCode, Cline, Cursor, Continue, Goose, Windsurf, and Claude Code.
Set up agentic codingMulti-Agent Teams
Native Agent Orchestration
Internalized multi-agent collaboration as a native capability. Stable identity across roles, enhanced emotional intelligence, and robust causal reasoning for production decisions.
Frontier Performance, Fraction of the Cost
Run a 229B parameter frontier model at open-source pricing.
| Model | Input (per 1M) | Output (per 1M) | Relative Cost |
|---|---|---|---|
| Claude Opus 4.6 | $15.00 | $75.00 | ~30x more |
| Claude Sonnet 4.6 | $3.00 | $15.00 | ~6x more |
| MiniMax M2.7 Ultraspeed | €0.60 | €2.40 | Baseline |
Pricing as of May 2026.
Why EU Hosting Matters
Running MiniMax through Infercom means full EU sovereignty - the fastest inference in Europe with zero foreign jurisdiction exposure:
- Your data is processed exclusively in Germany - never leaves EU jurisdiction
- No US CLOUD Act exposure - no American hyperscaler involvement
- Full GDPR compliance with EU-based Data Processing Agreement
- ISO 27001 certified infrastructure owned and operated by Infercom
- Zero data retention - we never train on your data
Start Building in Minutes
from openai import OpenAI
client = OpenAI(
api_key="your-infercom-key",
base_url="https://api.infercom.ai/v1"
)
response = client.chat.completions.create(
model="MiniMax-M2.7",
messages=[{"role": "user", "content": "Your prompt here"}],
max_tokens=4096
)
print(response.choices[0].message.content)OpenAI-compatible API. Drop-in replacement for your existing code. Use model name MiniMax M2.7 Ultraspeed in your API calls.
Pay only for the tokens you use.
MiniMax M2.7: Frequently Asked Questions
428 tokens per second output throughput and 690 ms time to first token, both server-side p50 measured at 10K input / 1K output on our production API in Munich. Shorter prompts peak at up to 444 tok/s. For a 229B frontier reasoning model that is up to 10x faster than GPU-based alternatives, and every figure is reproducible with our open-source benchmark tool.
See the full benchmarks192K tokens (196,608), shared between your prompt and the model's response. That is designed for long agentic runs, multi-file code changes and document sets that would otherwise need to be chunked across several requests.
What a context window isYes. MiniMax publishes the M2.7 weights openly on Hugging Face. You can put the model into production through our API straight away, and because the weights are open you are never locked into a single provider.
What open-weight actually meansYes. Infercom serves M2.7 from Infercom-owned hardware in a Tier III+ datacenter in Munich, Germany. Requests are processed inside EU jurisdiction with no US CLOUD Act exposure, the infrastructure is ISO 27001 certified, and an EU data processing agreement is available on request. We do not train on your data.
How EU sovereignty works herePay-per-token with no minimums and no monthly commitment: EUR 0.60 per million input tokens and EUR 2.40 per million output tokens. You are billed only for what you use, and live usage and spend are visible in the cloud portal.
See full pricingLearn More
API Setup Guide
Base URL, model IDs, and code examples
Performance Benchmarks
See how MiniMax M2.7 Ultraspeed performs on our EU infrastructure
Agentic Coding Guide
Set up Aider, OpenCode, Cursor with MiniMax
gpt-oss-120b
When raw throughput matters more than frontier reasoning
EU Sovereign AI
Where your data lives and who can reach it
Pricing
Transparent per-token pricing
