Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.
Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.
Pricing history
Direct provider rates per million tokens, plus a typical-conversation cost estimate.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$0.1200
input / 1M
— stable
$0.1800
output / 1M
— stable
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
No measurements yet
We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.
Tokonomix benchmark verdicts
Qwen3.5-9B surges to 92.0 with major coding improvement, latency increases
Qwen3.5-9B has demonstrated substantial improvement in this benchmark window, climbing from 75.5 to 92.0 overall quality. The most dramatic change comes in coding performance, which jumped from 59 to 92, representing a 33-point gain and signaling significant capability enhancement in technical tasks. This coding advancement appears to be the primary driver of the overall quality improvement of 16.5 points. However, the current window only includes a single test run compared to three in the previous period, which may limit the statistical robustness of these results. Latency has increased from 15791ms to 18139ms at the median, representing approximately a 15% slowdown that users should factor into latency-sensitive applications. The previous window showed strong multilingual performance at 92, but no multilingual score is available in the current window for comparison. The dramatic coding improvement suggests either model updates or optimization changes that have substantially enhanced technical reasoning capabilities, though users should monitor whether these gains persist across additional test runs.
Quality
92.0
Latency p50
18,139 ms
Test runs
1
Qwen3.5-9B
by OVH AI Endpoints (GRA)
- Context window
- — tokens
- Input price
- $0.1200 / 1M
- Output price
- $0.1800 / 1M
- Tier
- Tier B — Production
- Modality
- Text
- API type
- REST · streaming
- Benchmark runs
- 474
More from OVH AI Endpoints (GRA)