Skip to content

Benchmarks

Leaderboard

Every active text model we can reach, with how long it takes to answer and how often it wins a head-to-head. Nothing is hidden: a model we have not measured still gets a row, and the row says why.

Typical
How long a normal answer takes. Half of the calls came back faster than this, half slower — so it is what you should expect on an ordinary request.
Slow case
How bad it gets on an off day. Roughly one call in twenty is slower than this. A model with a low typical time but a high slow case is fast until it is not.
Win rate
Out of 100 head-to-head comparisons against an average model on this board, how many this one wins. 50 is average, higher is better. It is worked out from every judged comparison, not from a single score.

104 of 109 models fully measured · 0 timed once · 0 awaiting a first run · 5 we cannot measure.

Filter:
#
1Qwen3-Coder-30B-A3B-InstructB82 ms
2Meta-Llama-3_3-70B-InstructB136 ms
3Qwen2.5-VL-72B-InstructB145 ms
4Mistral-Nemo-Instruct-2407C158 ms
5NVIDIA Nemotron Super 49B v1.5A173 ms
6Mistral-Small-3.2-24B-Instruct-2506B174 ms
7MiniMax M2.5A175 ms
8Llama 4 ScoutA202 ms
9gpt-oss-120bC209 ms
10Nous Hermes 3 70BA212 ms
11Mistral-7B-Instruct-v0.3C228 ms
12gpt-oss-20bC273 ms
13Qwen3.5-397B-A17BA320 ms
14Cohere Command-AA402 ms
15gpt-4.1-mini-2025-04-14C446 ms
16gpt-4.1-nanoC488 ms
17gpt-5.4-mini-2026-03-17A516 ms
18Gemini 2.5 FlashA518 ms
19Gemini 2.5 Flash-LiteB532 ms
20Qwen3.5-9BB538 ms
21Gemini 3.1 Flash LiteB542 ms
22o1-2024-12-17C543 ms
23gpt-4.1B544 ms
24gpt-4o-2024-05-13C544 ms
25gpt-5.4-miniA548 ms
26Llama 4 MaverickA548 ms
27Qwen3-32BB548 ms
28gpt-4.1-nano-2025-04-14C550 ms
29gpt-4o-mini-2024-07-18C554 ms
30Claude Haiku 4.5A561 ms
31Gemini Flash-Lite LatestC562 ms
32Llama 3.3 70B InstructA567 ms
33gpt-4o-miniC570 ms
34gpt-4.1-miniC572 ms
35o3-miniC579 ms
36Qwen 2.5 VL 72B InstructA582 ms
37gpt-3.5-turbo-0125C590 ms
38gpt-5.1B613 ms
39gpt-4o-2024-11-20C617 ms
40Qwen 3.6 PlusA622 ms
41gpt-5.4A646 ms
42gpt-5.4-nanoC668 ms
43gpt-4.1-2025-04-14C669 ms
44gpt-4o-2024-08-06C691 ms
45Qwen 3.7 MaxA742 ms
46gpt-5-nanoC755 ms
47o3-2025-04-16B765 ms
48gpt-5.4-nano-2026-03-17A776 ms
49Claude Sonnet 4.5B782 ms
50o3C822 ms
51Gemini 3 Flash PreviewC834 ms
52Claude Opus 4.5B845 ms
53Claude Opus 4.7B846 ms
54gpt-4-turboC849 ms
55gpt-5.4-2026-03-05B850 ms
56gpt-5.2-2025-12-11B851 ms
57gpt-5.1-2025-11-13B853 ms
58gpt-3.5-turbo-16kC884 ms
59gpt-5-mini-2025-08-07B893 ms
60gpt-5-2025-08-07B919 ms
61Gemini 3.5 FlashA925 ms
62gpt-4-0613C925 ms
63o4-miniC934 ms
64o4-mini-2025-04-16B960 ms
65gpt-4-turbo-2024-04-09C972 ms
66DeepSeek v4 ProA973 ms
67Claude Opus 4.8A987 ms
68o3-mini-2025-01-31C988 ms
69gpt-4C993 ms
70o1C1.0 s
71Claude Sonnet 5A1.1 s
72Qwen3.7 MaxA1.1 s
73Gemini Flash LatestB1.1 s
74Qwen3.7 PlusB1.1 s
75gpt-5C1.2 s
76gpt-5.5C1.2 s
77gpt-5-miniC1.2 s
78Gemini 3.1 Pro Preview Custom ToolsC1.3 s
79Claude Opus 4.6B1.3 s
80gpt-4oC1.3 s
81Gemini 2.5 ProA1.4 s
82DeepSeek v3.2A1.4 s
83gpt-3.5-turbo-1106C1.4 s
84gpt-5.2B1.4 s
85GLM-4.5V (vision)A1.5 s
86gpt-5-search-apiC1.5 s
87gpt-5-search-api-2025-10-14B1.5 s
88gpt-3.5-turboC1.6 s
89Claude Opus 5A1.9 s
90Gemini 3.1 Pro PreviewC2.0 s
91GLM-4.5 AirB2.0 s
92GLM-4.6V (vision)A2.0 s
93Gemini Pro LatestC2.1 s
94gpt-5.5-2026-04-23A2.1 s
95GLM-4.5A2.2 s
96Claude Sonnet 4.6A2.2 s
97gpt-5-nano-2025-08-07B2.5 s
98GLM-5A2.7 s
99Claude Fable 5A2.9 s
100GLM-5.2A3.1 s
101GLM-4.6A4.2 s
102GLM-5.1A4.3 s
103GLM-4.7A4.4 s
104GLM-5 TurboB5.8 s
·Deep Research Max Preview (Apr-21-2026)BCannot be measured — answers on a different API
·Deep Research Preview (Apr-21-2026)BCannot be measured — answers on a different API
·Deep Research Pro Preview (Dec-12-2025)BCannot be measured — answers on a different API
·Gemini 2.5 Computer Use Preview 10-2025BCannot be measured — answers on a different API
·gpt-5.6-terraCannot be measured — no probe for this provider

109 of 109 models · click a column to sort

“Cannot be measured” means the model does not answer on the chat API we time, or its provider uses a protocol our prober does not speak yet. Waiting will not fill these in.

Fast (under 500 ms)
Medium (500–1000 ms)
Slow (over 1000 ms)
Measured four times a day (02:00, 08:00, 14:00, 20:00 UTC) from Amsterdam.