Skip to content
Tier B — Production
Runs in:FranceMade in:China
OVH AI Endpoints (GRA)

Qwen2.5-VL-72B-Instruct

Tier B — Production

Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency100 runs
923102611291211213108-2109-14ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

56%
Coding
judge mean 97
65%
Creative
judge mean 90
33%
Factual
judge mean 70
52%
Multilingual
judge mean 97
74%
Reasoning
judge mean 99

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing

What you pay per million tokens when you use this model on Tokonomix, plus an estimate for a typical conversation.

💰
API rates — Qwen2.5-VL-72B-Instruct
$1.78 per 1M input tokens
$1.78 per 1M output tokens
≈ $0.0014 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$1.78
per 1M output tokens$1.78
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)1379 / avg 1314
213619

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

visionownedBy: Qwen
Section 06

Availability

Availability

No measurements yet

We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 1 runs
1 correct0 partial0 wrong100% accuracy
claude-sonnet-4-592/100 · 82 runs
71 correct9 partial2 wrong87% accuracy
2026-09-13

Vision capability confirmed, stable quality and latency maintained

Qwen2.5-VL-72B-Instruct through OVH AI Endpoints continues to demonstrate stable performance in its second benchmark window with vision capabilities. The model maintains its quality score of 93.3, showing consistency in response accuracy and helpfulness. Average latency remains steady at 15.3 seconds, indicating reliable performance characteristics for this multimodal model. The vision capability, newly detected in the previous window, is now confirmed as an established feature of this endpoint. Users can expect dependable performance for both text and visual understanding tasks. The model serves as a capable option for applications requiring vision-language understanding with predictable response times. While no improvements are observed in this window, the lack of degradation in either quality or speed suggests stable infrastructure and model serving. Organizations evaluating this endpoint can rely on the consistent metrics for capacity planning. The 93.3 quality score positions this model competitively for production workloads requiring multimodal capabilities, though users should consider the 15.3-second latency when designing time-sensitive applications.

Quality

Latency p50

Test runs

0

Quality stable at 93.3 Latency consistent at 15.3s Vision capability confirmed
Last automated test
Sep 14, 2026 · 20:02 UTC · Speed benchmark
P50 latency
145 ms
P95 latency
394 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·September 14, 2026