Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.
Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.
Pricing history
Direct provider rates per million tokens, plus a typical-conversation cost estimate.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$0.9100
input / 1M
— stable
$0.9100
output / 1M
— stable
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
No measurements yet
We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.
Tokonomix benchmark verdicts
Vision capability added, quality stable at 93.3, latency steady at 15.3s
Qwen2.5-VL-72B-Instruct maintains consistent performance in this benchmark window with quality holding steady at 93.3 and latency unchanged at 15.3 seconds. The model continues to offer vision capabilities that were introduced in the previous period, expanding its utility beyond text-only tasks. Performance metrics show no degradation despite the multimodal functionality, indicating stable infrastructure and model serving. The latency of 15.3 seconds remains elevated compared to text-only models in its class, which users should account for in latency-sensitive applications. Quality scores in the mid-90s position this model as a reliable option for vision-language tasks requiring strong reasoning capabilities. Organizations evaluating this endpoint should note the consistency across benchmark windows, suggesting predictable performance for production deployments. The absence of fluctuation in key metrics indicates mature model deployment, though the latency profile may warrant consideration for real-time use cases. Overall, this window demonstrates operational stability with no notable regressions or improvements in core performance indicators.
Quality
—
Latency p50
—
Test runs
0
Qwen2.5-VL-72B-Instruct
by OVH AI Endpoints (GRA)
- Context window
- — tokens
- Input price
- $0.9100 / 1M
- Output price
- $0.9100 / 1M
- Tier
- Tier B — Production
- Modality
- Text + vision
- API type
- REST · streaming
- Benchmark runs
- 474
More from OVH AI Endpoints (GRA)