Skip to content

Benchmarks

Speed test

P50 = median response time for a standard 500-token output. Measured from EU (Amsterdam). P95 = tail latency — 95% of requests complete within this time. Three runs per model per test cycle; values are medians across cycles.

Tier S< 200 ms
Tier A< 500 ms
Tier B< 1000 ms
Tier C> 1000 ms
P50 (median)P95 (tail)
01FLUX.1 Kontext [max] — Multi-Image Fusion
Tier S0 ms
P95: 0 ms
02FLUX.1 Kontext [pro] — Multi-Image Fusion
Tier S0 ms
P95: 0 ms
03gpt-5.6-terra
Tier S0 ms
P95: 0 ms
04NVIDIA Nemotron Super 49B v1.5
Tier S16 ms
P95: 17 ms
05Qwen3-Coder-30B-A3B-Instruct
Tier S81 ms
P95: 83 ms
06Mistral-Nemo-Instruct-2407
Tier S90 ms
P95: 104 ms
07SDXL 1.0
Tier S101 ms
P95: 101 ms
08gpt-5.2-chat-latest
Tier S121 ms
P95: 128 ms
09Mistral-Small-3.2-24B-Instruct-2506
Tier S121 ms
P95: 147 ms
10gpt-5.3-chat-latest
Tier S126 ms
P95: 128 ms
How we measure: Each model receives an identical prompt targeting a ~500-token output. We run 3 sequential calls per test cycle and compute P50/P95 across the distribution. Tests run 4× per day from a single EU endpoint. Network overhead is included.