Skip to content

Benchmarks

Leaderboard

Every active text model we can reach, with how long it takes to answer and how often it wins a head-to-head. Nothing is hidden: a model we have not measured still gets a row, and the row says why.

Typical
How long a normal answer takes. Half of the calls came back faster than this, half slower — so it is what you should expect on an ordinary request.
Slow case
How bad it gets on an off day. Roughly one call in twenty is slower than this. A model with a low typical time but a high slow case is fast until it is not.
Win rate
Out of 100 head-to-head comparisons against an average model on this board, how many this one wins. 50 is average, higher is better. It is worked out from every judged comparison, not from a single score.

103 of 108 models fully measured · 0 timed once · 0 awaiting a first run · 5 we cannot measure.

Filter:
#
1Qwen3-Coder-30B-A3B-InstructB81 ms
2Meta-Llama-3_3-70B-InstructB129 ms
3Mistral-Nemo-Instruct-2407C132 ms
4Qwen2.5-VL-72B-InstructB141 ms
5Mistral-7B-Instruct-v0.3C171 ms
6NVIDIA Nemotron Super 49B v1.5A173 ms
7Llama 4 ScoutA186 ms
8Nous Hermes 3 70BA188 ms
9Llama 3.3 70B InstructA210 ms
10gpt-oss-20bC244 ms
11Qwen3.5-397B-A17BA279 ms
12Qwen 2.5 VL 72B InstructA329 ms
13gpt-4o-miniC459 ms
14gpt-5.4-mini-2026-03-17A493 ms
15gpt-4.1-miniC511 ms
16Cohere Command-AA539 ms
17gpt-4-0613C547 ms
18gpt-5.4-miniA552 ms
19gpt-4.1-mini-2025-04-14C557 ms
20gpt-3.5-turbo-16kC559 ms
21Gemini 2.5 Flash-LiteB560 ms
22gpt-4o-mini-2024-07-18C564 ms
23gpt-4o-2024-05-13C572 ms
24DeepSeek v4 ProA581 ms
25Qwen3.5-9BB583 ms
26gpt-5-nanoC584 ms
27Llama 4 MaverickA588 ms
28gpt-5.4-nanoC592 ms
29Claude Haiku 4.5A593 ms
30gpt-4C593 ms
31gpt-4.1-nano-2025-04-14C600 ms
32o3-mini-2025-01-31C604 ms
33Gemini 2.5 FlashA608 ms
34gpt-5-nano-2025-08-07B608 ms
35gpt-5-2025-08-07B625 ms
36gpt-3.5-turbo-1106C631 ms
37gpt-5.4-nano-2026-03-17A642 ms
38o3-miniC644 ms
39o3C646 ms
40gpt-5C687 ms
41gpt-4oC694 ms
42gpt-4o-2024-11-20C707 ms
43gpt-4o-2024-08-06C710 ms
44o4-mini-2025-04-16B718 ms
45gpt-5.1-2025-11-13B723 ms
46o4-miniC727 ms
47o3-2025-04-16B736 ms
48gpt-5-mini-2025-08-07B737 ms
49gpt-oss-120bC741 ms
50gpt-5.1B755 ms
51gpt-5.4A766 ms
52gpt-5.4-2026-03-05B766 ms
53gpt-4.1B773 ms
54o1-2024-12-17C782 ms
55Gemini 3.5 FlashA784 ms
56Gemini 3.1 Flash LiteB797 ms
57Claude Opus 4.7B817 ms
58gpt-4.1-nanoC832 ms
59gpt-5.2-2025-12-11B842 ms
60Claude Sonnet 4.5B844 ms
61Claude Opus 4.5B857 ms
62Gemini Flash-Lite LatestC860 ms
63gpt-5.2B873 ms
64o1C875 ms
65Qwen 3.7 MaxA893 ms
66MiniMax M2.5A907 ms
67Claude Sonnet 4.6A918 ms
68Claude Sonnet 5A924 ms
69gpt-4-turbo-2024-04-09C934 ms
70gpt-5.5C939 ms
71DeepSeek v3.2A963 ms
72Gemini Flash LatestB981 ms
73Qwen3.7 PlusB981 ms
74Gemini 3 Flash PreviewC989 ms
75gpt-5-miniC992 ms
76Claude Opus 4.8A1.0 s
77Qwen 3.6 PlusA1.1 s
78gpt-4.1-2025-04-14C1.1 s
79Qwen3.7 MaxA1.1 s
80gpt-4-turboC1.1 s
81gpt-3.5-turbo-0125C1.2 s
82gpt-3.5-turboC1.3 s
83Gemini 3.1 Pro Preview Custom ToolsC1.3 s
84gpt-5.5-2026-04-23A1.3 s
85GLM-4.5V (vision)A1.4 s
86gpt-5-search-api-2025-10-14B1.5 s
87Claude Opus 5A1.6 s
88Gemini 2.5 ProA1.7 s
89GLM-4.6V (vision)A1.8 s
90gpt-5-search-apiC1.8 s
91Mistral-Small-3.2-24B-Instruct-2506B1.8 s
92Claude Opus 4.6B1.8 s
93Gemini Pro LatestC2.0 s
94Gemini 3.1 Pro PreviewC2.0 s
95GLM-4.5 AirB2.4 s
96Claude Fable 5A2.5 s
97GLM-5A2.7 s
98GLM-5.2A3.1 s
99GLM-5.1A5.1 s
100GLM-5 TurboB5.3 s
101GLM-4.5A6.6 s
102GLM-4.6A6.7 s
103GLM-4.7A7.1 s
·Deep Research Max Preview (Apr-21-2026)BCannot be measured — answers on a different API
·Deep Research Preview (Apr-21-2026)BCannot be measured — answers on a different API
·Deep Research Pro Preview (Dec-12-2025)BCannot be measured — answers on a different API
·Gemini 2.5 Computer Use Preview 10-2025BCannot be measured — answers on a different API
·gpt-5.6-terraCannot be measured — no probe for this provider

108 of 108 models · click a column to sort

“Cannot be measured” means the model does not answer on the chat API we time, or its provider uses a protocol our prober does not speak yet. Waiting will not fill these in.

Fast (under 500 ms)
Medium (500–1000 ms)
Slow (over 1000 ms)
Measured four times a day (02:00, 08:00, 14:00, 20:00 UTC) from Amsterdam.