Skip to content

Benchmarks

Leaderboard

Every active text model we can reach, with how long it takes to answer and how often it wins a head-to-head. Nothing is hidden: a model we have not measured still gets a row, and the row says why.

Typical
How long a normal answer takes. Half of the calls came back faster than this, half slower — so it is what you should expect on an ordinary request.
Slow case
How bad it gets on an off day. Roughly one call in twenty is slower than this. A model with a low typical time but a high slow case is fast until it is not.
Win rate
Out of 100 head-to-head comparisons against an average model on this board, how many this one wins. 50 is average, higher is better. It is worked out from every judged comparison, not from a single score.

104 of 109 models fully measured · 0 timed once · 0 awaiting a first run · 5 we cannot measure.

Filter:
#
1Qwen3-Coder-30B-A3B-InstructB73 ms
2Meta-Llama-3_3-70B-InstructB127 ms
3Mistral-Small-3.2-24B-Instruct-2506B157 ms
4NVIDIA Nemotron Super 49B v1.5A173 ms
5Qwen2.5-VL-72B-InstructB187 ms
6Qwen3.5-397B-A17BA229 ms
7Mistral-7B-Instruct-v0.3C230 ms
8gpt-oss-20bC239 ms
9Mistral-Nemo-Instruct-2407C288 ms
10Nous Hermes 3 70BA362 ms
11Llama 4 ScoutA408 ms
12gpt-4o-miniC424 ms
13Qwen 2.5 VL 72B InstructA430 ms
14gpt-5.4-miniA431 ms
15gpt-3.5-turbo-0125C451 ms
16Gemini Flash-Lite LatestC462 ms
17gpt-4.1-mini-2025-04-14C470 ms
18gpt-oss-120bC489 ms
19Gemini 3.1 Flash LiteB494 ms
20Qwen3-32BB495 ms
21Llama 3.3 70B InstructA499 ms
22gpt-4o-mini-2024-07-18C501 ms
23Llama 4 MaverickA501 ms
24gpt-3.5-turbo-16kC502 ms
25Gemini 2.5 Flash-LiteB512 ms
26gpt-3.5-turboC518 ms
27gpt-4.1-nanoC520 ms
28gpt-4.1-2025-04-14C543 ms
29gpt-4o-2024-05-13C554 ms
30Cohere Command-AA558 ms
31o3-miniC561 ms
32Qwen3.5-9BB566 ms
33o1-2024-12-17C567 ms
34gpt-4.1-nano-2025-04-14C568 ms
35gpt-4o-2024-11-20C576 ms
36gpt-4o-2024-08-06C578 ms
37Gemini 3.5 FlashA581 ms
38o1C587 ms
39Gemini 2.5 FlashA589 ms
40gpt-4.1-miniC592 ms
41Gemini 3 Flash PreviewC594 ms
42gpt-5.1-2025-11-13B597 ms
43gpt-4.1B600 ms
44o4-miniC612 ms
45gpt-4oC614 ms
46gpt-5.4-mini-2026-03-17A619 ms
47o4-mini-2025-04-16B662 ms
48gpt-5.1B664 ms
49Claude Haiku 4.5A674 ms
50o3C674 ms
51o3-2025-04-16B676 ms
52gpt-5.4-2026-03-05B701 ms
53Qwen 3.6 PlusA706 ms
54gpt-5.4-nanoC712 ms
55gpt-5-nanoC738 ms
56gpt-5-nano-2025-08-07B746 ms
57MiniMax M2.5A761 ms
58gpt-5C763 ms
59gpt-5.2-2025-12-11B779 ms
60gpt-5.2B814 ms
61Gemini Flash LatestB820 ms
62gpt-5.4-nano-2026-03-17A829 ms
63gpt-5-mini-2025-08-07B842 ms
64gpt-4C867 ms
65gpt-5.4A871 ms
66o3-mini-2025-01-31C892 ms
67DeepSeek v3.2A893 ms
68Claude Opus 4.5B902 ms
69gpt-5-miniC929 ms
70Claude Sonnet 4.5B963 ms
71gpt-5-2025-08-07B983 ms
72Qwen 3.7 MaxA992 ms
73gpt-4-0613C1.0 s
74Gemini 2.5 ProA1.1 s
75Claude Opus 4.8A1.1 s
76gpt-4-turboC1.1 s
77gpt-5.5C1.2 s
78Qwen3.7 PlusB1.2 s
79Claude Sonnet 5A1.2 s
80Gemini 3.1 Pro Preview Custom ToolsC1.3 s
81gpt-5.5-2026-04-23A1.4 s
82gpt-5-search-apiC1.4 s
83gpt-4-turbo-2024-04-09C1.5 s
84gpt-5-search-api-2025-10-14B1.6 s
85Claude Opus 4.7B1.6 s
86GLM-4.6V (vision)A1.9 s
87Claude Opus 5A1.9 s
88Gemini 3.1 Pro PreviewC2.1 s
89DeepSeek v4 ProA2.2 s
90Gemini Pro LatestC2.3 s
91gpt-3.5-turbo-1106C2.5 s
92GLM-4.5V (vision)A2.5 s
93Qwen3.7 MaxA2.9 s
94GLM-5A2.9 s
95GLM-5.2A3.3 s
96Claude Fable 5A3.6 s
97Claude Sonnet 4.6A4.1 s
98GLM-4.5 AirB4.5 s
99Claude Opus 4.6B8.0 s
100GLM-5 TurboB8.6 s
101GLM-5.1A8.8 s
102GLM-4.7A9.0 s
103GLM-4.6A9.3 s
104GLM-4.5A15.6 s
·Deep Research Max Preview (Apr-21-2026)BCannot be measured — answers on a different API
·Deep Research Preview (Apr-21-2026)BCannot be measured — answers on a different API
·Deep Research Pro Preview (Dec-12-2025)BCannot be measured — answers on a different API
·Gemini 2.5 Computer Use Preview 10-2025BCannot be measured — answers on a different API
·gpt-5.6-terraCannot be measured — no probe for this provider

109 of 109 models · click a column to sort

“Cannot be measured” means the model does not answer on the chat API we time, or its provider uses a protocol our prober does not speak yet. Waiting will not fill these in.

Fast (under 500 ms)
Medium (500–1000 ms)
Slow (over 1000 ms)
Measured four times a day (02:00, 08:00, 14:00, 20:00 UTC) from Amsterdam.