Skip to content
Tier B — Production
Runs in:FranceMade in:China
OVH AI Endpoints (GRA)

Qwen2.5-VL-72B-Instruct

Tier B — Production

Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency105 runs
923102611291211213108-1009-05ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

56%
Coding
judge mean 99
67%
Creative
judge mean 93
34%
Factual
judge mean 73
55%
Multilingual
judge mean 98
74%
Reasoning
judge mean 99

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — Qwen2.5-VL-72B-Instruct
$0.9100 per 1M input tokens
$0.9100 per 1M output tokens
≈ $0.0007 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$0.9100
per 1M output tokens$0.9100

Pricing over time

Input & output per 1M tokens · step-line = price changes

$0.9100

input / 1M

— stable

$0.9100

output / 1M

— stable

2026-06-142026-07-262026-08-30
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)1389 / avg 1290
213619

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

visionownedBy: Qwen
Section 06

Availability

Availability

No measurements yet

We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 1 runs
1 correct0 partial0 wrong100% accuracy
claude-sonnet-4-593/100 · 72 runs
63 correct8 partial1 wrong88% accuracy
2026-08-30

Vision capability added, quality stable at 93.3, latency steady at 15.3s

Qwen2.5-VL-72B-Instruct maintains consistent performance in this benchmark window with quality holding steady at 93.3 and latency unchanged at 15.3 seconds. The model continues to offer vision capabilities that were introduced in the previous period, expanding its utility beyond text-only tasks. Performance metrics show no degradation despite the multimodal functionality, indicating stable infrastructure and model serving. The latency of 15.3 seconds remains elevated compared to text-only models in its class, which users should account for in latency-sensitive applications. Quality scores in the mid-90s position this model as a reliable option for vision-language tasks requiring strong reasoning capabilities. Organizations evaluating this endpoint should note the consistency across benchmark windows, suggesting predictable performance for production deployments. The absence of fluctuation in key metrics indicates mature model deployment, though the latency profile may warrant consideration for real-time use cases. Overall, this window demonstrates operational stability with no notable regressions or improvements in core performance indicators.

Quality

Latency p50

Test runs

0

Quality stable at 93.3 Latency consistent at 15.3s Vision capabilities maintained
Last automated test
Sep 5, 2026 · 08:01 UTC · Speed benchmark
P50 latency
144 ms
P95 latency
3196 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·September 5, 2026