Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.
Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.
Pricing history
Direct provider rates per million tokens, plus a typical-conversation cost estimate.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$0.7100
input / 1M
— stable
$4.25
output / 1M
— stable
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
No measurements yet
We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.
Tokonomix benchmark verdicts
Qwen3.5-397B-A17B reaches perfect 100/100 score with faster response times
Qwen3.5-397B-A17B has achieved a perfect quality score of 100 out of 100 in the current benchmark window, marking a substantial 15.5-point improvement from the previous period's 84.5 score. This advancement represents continued momentum following earlier quality gains. Response latency has improved significantly, with the median response time dropping from 5064ms to 2719ms, a 46% reduction that enhances the model's practical usability. The coding category maintained its strong performance at 100, up from 96 in the previous window. However, it should be noted that the current window contains only a single test run compared to three runs previously, which means these results may be less statistically robust. The absence of multilingual category data in the current window, which previously scored 73, leaves questions about performance stability across diverse language tasks. Users should consider that while the perfect score and improved latency are encouraging signs, additional test runs would provide greater confidence in the consistency of these results. The model appears to be trending positively in both quality and speed metrics.
Quality
100.0
Latency p50
2,719 ms
Test runs
1
Qwen3.5-397B-A17B
by OVH AI Endpoints (GRA)
- Context window
- — tokens
- Input price
- $0.7100 / 1M
- Output price
- $4.25 / 1M
- Tier
- Tier A — Frontier
- Modality
- Text
- API type
- REST · streaming
- Benchmark runs
- 474
More from OVH AI Endpoints (GRA)