MiniMax M2.7 API in Europe

229B frontier reasoning at 428 tokens per second. Fully EU sovereign.

The latest 229B parameter frontier model from MiniMax, running on Infercom's fully EU sovereign infrastructure in Germany. No US hyperscalers. No CLOUD Act exposure. Native multi-agent support, 30% coding improvement over the previous generation.

New in M2.7

M2.7 adds built-in self-critique for better first-attempt results, plus native multi-agent orchestration as a core capability.

Built-in Self-Critique

New in M2.7: The model automatically reviews and refines its outputs before responding. Better first-attempt results on complex coding and reasoning tasks.

Native Agent Teams

New in M2.7: Multi-agent collaboration as a core model capability - not just prompting. Stable role boundaries, adversarial reasoning, and behavioral differentiation built in.

Improved: Software Engineering

56.22% on SWE-Pro (matching GPT-5.3-Codex), 55.6% on VIBE-Pro for full project delivery, 57.0% on Terminal Bench 2 for complex system understanding.

Improved: Professional Skills

97% skill compliance across 40+ complex professional skills. Enhanced Excel, PowerPoint, and Word editing with multi-round revision support.

EU Sovereign Munich, Germany

Measured on Infercom EU Infrastructure

Output Throughput

428tok/s

Time to First Token

690ms

Context Window

192Ktokens

Server-side p50, 10K input / 1K output, single request

Up to 444 tok/s on shorter prompts. Last measured: July 2026.

See full benchmarks and methodology

30%

Coding performance improvement over the previous generation

Why we call it Ultraspeed

Frontier Benchmark Performance

M2.7 achieves top-tier scores across coding, engineering, and ML benchmarks - matching or exceeding GPT-5.3-Codex on SWE-Pro.

SWE-Pro

56.22%

Matches GPT-5.3-Codex

SWE Multilingual

76.5%

Real-world engineering

VIBE-Pro

55.6%

Near Opus 4.6 level

Terminal Bench 2

57.0%

Deep system understanding

MLE Bench Lite

66.6%

Medal rate (9 gold, 5 silver)

Multi SWE Bench

52.7%

Complex codebases

Built for Agentic Workflows

M2.7 excels at long-horizon agent tasks that require autonomous decision-making, tool use, and multi-step reasoning across complex professional domains.

Financial Workflows

End-to-end research, Excel modeling, report generation

SRE & DevOps

Log analysis, incident response, production debugging

Document Processing

Word, Excel, PowerPoint with multi-round editing

ML Competitions

66.6% medal rate on MLE Bench Lite (9 gold, 5 silver)

Agentic Coding

56.22% on SWE-Pro

Matches GPT-5.3-Codex on the most demanding software engineering benchmark. Works with Aider, OpenCode, Cline, Cursor, Continue, Goose, Windsurf, and Claude Code.

Set up agentic coding

Multi-Agent Teams

Native Agent Orchestration

Internalized multi-agent collaboration as a native capability. Stable identity across roles, enhanced emotional intelligence, and robust causal reasoning for production decisions.

Pricing

Run a 229B parameter frontier model with transparent, usage-based pricing.

ModelInput (per 1M)Output (per 1M)Context
MiniMax M2.7 Ultraspeed€0.60€2.40192K

Prices in EUR excl. VAT. EU sovereign deployment with full GDPR compliance.

Why EU Hosting Matters

Running MiniMax through Infercom keeps your requests under EU jurisdiction:

  • Your data is processed exclusively in Germany - never leaves EU jurisdiction
  • No US CLOUD Act exposure - no American hyperscaler involvement
  • Full GDPR compliance with EU-based Data Processing Agreement
  • ISO 27001 certified infrastructure owned and operated by Infercom
  • Zero data retention - we never train on your data
ISO 27001 Certified
GDPR Compliant
German Datacenter
Infercom-Owned Hardware

Start Building in Minutes

quickstart.py
from openai import OpenAI

client = OpenAI(
    api_key="your-infercom-key",
    base_url="https://api.infercom.ai/v1"
)

response = client.chat.completions.create(
    model="MiniMax-M2.7",
    messages=[{"role": "user", "content": "Your prompt here"}],
    max_tokens=4096
)

print(response.choices[0].message.content)

OpenAI-compatible API. Drop-in replacement for your existing code. Use model name MiniMax M2.7 Ultraspeed in your API calls.

Pay only for the tokens you use.

MiniMax M2.7: Frequently Asked Questions

How fast is MiniMax M2.7 on Infercom?

428 tokens per second output throughput and 690 ms time to first token, both server-side p50 measured at 10K input / 1K output on our production API in Munich. Shorter prompts peak at up to 444 tok/s. For a 229B frontier reasoning model that is up to 10x faster than GPU-based alternatives, and every figure is reproducible with our open-source benchmark tool.

See the full benchmarks
What is the context window of MiniMax M2.7?

192K tokens (196,608), shared between your prompt and the model's response. That is designed for long agentic runs, multi-file code changes and document sets that would otherwise need to be chunked across several requests.

What a context window is
Is MiniMax M2.7 open-weight?

Yes. MiniMax publishes the M2.7 weights openly on Hugging Face. You can put the model into production through our API straight away, and because the weights are open you are never locked into a single provider.

What open-weight actually means
Can I run MiniMax M2.7 in the EU?

Yes. Infercom serves M2.7 from Infercom-owned hardware in a Tier III+ datacenter in Munich, Germany. Requests are processed inside EU jurisdiction with no US CLOUD Act exposure, the infrastructure is ISO 27001 certified, and an EU data processing agreement is available on request. We do not train on your data.

How EU sovereignty works here
What does MiniMax M2.7 cost?

Pay-per-token with no minimums and no monthly commitment: EUR 0.60 per million input tokens and EUR 2.40 per million output tokens (excl. VAT). You are billed only for what you use, and live usage and spend are visible in the cloud portal.

See full pricing

Ready to Build with Enterprise-Grade AI?

Start with a pilot, scale to production. Record-breaking performance with dedicated enterprise support.