Performance

Europe's Fastest LLM Inference

Don't take our word for it - verify it yourself.

Real benchmark results from our EU infrastructure. Independently verified technology. Open-source benchmarks you can run yourself.

Tokens per Second, Measured

These numbers are from our production API with our servers in the EU. Not vendor marketing - real measurements you can reproduce.

Last measured: July 2026

EU SovereignFastest Throughput

gpt-oss-120b

Output Throughput
713tok/s
Time to First Token
388ms
End-to-End Latency
1.789s

10K input / 1K output, single request

Up to 772 tok/s on shorter prompts

EU SovereignBest Reasoning

MiniMax M2.7 Ultraspeed

Output Throughput
428tok/s
Time to First Token
690ms
End-to-End Latency
3.023s

10K input / 1K output, single request

Up to 444 tok/s on shorter prompts

EU SovereignNative Multimodal

Gemma 4 31B

Output Throughput
199tok/s
Time to First Token
1189ms
End-to-End Latency
6.206s

10K input / 1K output, single request

Native multimodal understanding for images and text

Server-side metrics (p50). Measured using our open-source benchmark tool. Your client-side results will vary based on network location and conditions.

Fastest LLM Inference, by Workload

Which number decides your speed depends on what you are building. Same measured runs, cut three ways.

Chat and voice

Time to First Token

388ms

gpt-oss-120b - p50

Conversational and voice interfaces are judged on how fast the first word appears, not on how fast the rest arrives.

Time to First Token

Agents and coding

Output Throughput

713tok/s

gpt-oss-120b - p50

Agent loops and coding assistants wait for the whole response, so tokens per second sets the pace. For frontier reasoning across long runs, MiniMax M2.7 Ultraspeed measures 428 tok/s with a 192K context window.

Output Throughput

Batch and RAG

End-to-End Latency

1.789s

gpt-oss-120b - p50

Pipelines care about the total round trip for a complete response. Gemma 4 31B takes image and text in the same request for document and vision workloads.

End-to-End Latency

Fast is half of it. Every number above was measured on hardware we own in Munich - see what EU sovereign actually means.

EU-Hosted Inference Infrastructure

All benchmarks measured on Infercom's production infrastructure in Germany. Your data never leaves European jurisdiction.

Location

Munich, Germany

Hardware

SambaNova SN40L Dataflow Architecture

Certification

ISO 27001 Certified

Ownership

Infercom-Owned Hardware

Your requests terminate on infrastructure we own - not rented cloud capacity from hyperscalers.

Data Residency

100% EU - No CLOUD Act Exposure

Data never leaves European jurisdiction. No US CLOUD Act exposure, no third-country transfers.

What We Measure

Three metrics that matter for production AI inference

Time to First Token (TTFT)

How quickly the model starts responding after your request. Critical for interactive applications and chat interfaces.

Output Throughput

Tokens generated per second after the first token. Determines how fast a complete response is delivered to the user.

End-to-End Latency

Total time from request to complete response. Includes TTFT plus full generation time. The number that matters for batch workloads.

Open Source

Run Your Own Benchmark

Our benchmark tool is fully open source. Run it against our API with your API key, or clone the repository and run it locally. Same code, same methodology, your results.

Why is it this fast? The architecture behind our speed →

Synthetic Performance

Fixed input/output token counts for controlled comparisons across models

Real Workload Simulation

Variable request rates mimicking production traffic patterns

Custom Dataset

Upload your own prompts and measure performance on your actual workload

Interactive Chat

Per-response metrics in a live chat interface - see TTFT and throughput on every reply

Performance Without the Power Bill

Up to 5x more energy efficient than GPU-based inference. Speed doesn't have to come at the planet's expense.

10 kW

Per Rack

vs. 40-50 kW+ for equivalent GPU infrastructure. Dramatic reduction in power consumption and cooling requirements.

Air Cooled

No Liquid Cooling

Standard air cooling simplifies deployment, eliminates water usage, and reduces operational complexity and cost.

Up to 5x

More Efficient

More intelligence per joule of energy consumed. Validated by Stanford Hazy Research methodology.

LLM Inference Speed: Frequently Asked Questions

Who has the fastest LLM inference in Europe?

Infercom serves open-weight models such as gpt-oss-120b on SambaNova dataflow hardware in Munich. Our highest measured output is 713 tokens per second, with peaks up to 772 tok/s on shorter prompts - up to 10x faster than GPU-based alternatives. Every figure is a server-side p50 from our production API, not a vendor estimate, and you can reproduce it with our open-source benchmark tool.

Run the benchmark yourself
Which model is fastest, and for what?

gpt-oss-120b leads output throughput at 713 tok/s and also has our lowest time to first token at 388 ms, which makes it the default for chat, voice and high-volume agent loops. MiniMax M2.7 Ultraspeed is our 229B frontier reasoning model at 428 tok/s with a 192K context window, for long agentic runs and hard coding tasks. Gemma 4 31B measures 199 tok/s and adds native image and text input. Chat and voice depend on time to first token; agents and coding depend on throughput.

Compare the models
How is this measured - can I reproduce it?

Yes. Every number is a server-side p50 at 10K input and 1K output tokens, single request, measured against our production API with our open-source benchmark tool. Run it online against our API with your own key, or clone the repository and run it locally against your own prompts. Your client-side results will differ depending on network location and conditions.

View the benchmark tool source
How do you keep latency low under concurrency?

Dataflow hardware keeps its compute pipeline busy at low batch sizes, so per-request speed does not depend on stacking many requests together the way GPU serving does. Batching still raises total system throughput, but it is not what buys the per-request numbers above. The figures published here are single-request p50s; if concurrency behaviour matters for your workload, the benchmark tool has a real-workload mode that replays variable request rates against our API.

How the dataflow architecture works
Is the speed independently verified?

The dataflow architecture Infercom runs on is continuously benchmarked by Artificial Analysis and has been covered by VentureBeat and TechRadar, with the underlying energy-efficiency methodology validated by Stanford Hazy Research. Those sources cover the technology; the per-model numbers on this page are our own measurements on our own Munich infrastructure, which is why we publish the benchmark tool so you can check them.

See the independent benchmarks
Why is dataflow faster than GPUs?

A GPU repeatedly moves model weights and activations between compute and memory, and during token-by-token decoding it spends much of its time waiting on that memory traffic. A dataflow architecture maps the model onto the chip and streams data through it, so far more of the silicon does useful work per token. That is where the up to 10x speed advantage over GPU-based inference comes from, alongside up to 5x better energy efficiency.

Dataflow vs. GPU architecture, explained

Ready to Build with Enterprise-Grade AI?

Start with a pilot, scale to production. Record-breaking performance with dedicated enterprise support.