Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.
Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.
Pricing history
Direct provider rates per million tokens, plus a typical-conversation cost estimate.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$0.0800
input / 1M
— stable
$0.2300
output / 1M
— stable
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
No measurements yet
We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.
Tokonomix benchmark verdicts
Qwen3-32B quality continues decline to 57.4, down 9 points from 66.5
Qwen3-32B at OVH AI Endpoints shows continued performance degradation, with overall quality dropping from 66.5 to 57.4, marking a 9.1-point decline in this benchmark window. This represents the second consecutive period of quality regression for this model. Performance across categories is mixed but concerning. Coding ability decreased from 71 to 68, maintaining a downward trend from previous drops. Factual understanding fell from 53 to 45, representing a significant 8-point decline. Creative tasks now score 45, a new category measured this window that shows relatively weak performance. The bright spot is reasoning, which scores 72, though no previous comparison exists. Multilingual capability, previously at 76, was not measured in this window. Latency showed meaningful improvement, with p50 dropping from 18761ms to 14446ms, a 23% reduction that brings response times closer to usability thresholds. Users should be aware that despite faster responses, the model's output quality has declined substantially across most measured dimensions, particularly in factual accuracy and creative tasks.
Quality
57.4
Latency p50
14,446 ms
Test runs
5
Qwen3-32B
by OVH AI Endpoints (GRA)
- Context window
- — tokens
- Input price
- $0.0800 / 1M
- Output price
- $0.2300 / 1M
- Tier
- Tier B — Production
- Modality
- Text
- API type
- REST · streaming
- Benchmark runs
- 474
More from OVH AI Endpoints (GRA)