Skip to content
Tier A — Frontier
Runs in:FranceMade in:China
OVH AI Endpoints (GRA)

Qwen3.5-397B-A17B

Tier A — Frontier

Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency105 runs
149210140536004795608-1009-05ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

50%
Coding
judge mean 90
6%
Creative
judge mean 43
6%
Factual
judge mean 14
28%
Multilingual
judge mean 38
4%
Reasoning
judge mean 1

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — Qwen3.5-397B-A17B
$0.7100 per 1M input tokens
$4.25 per 1M output tokens
≈ $0.0013 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$0.7100
per 1M output tokens$4.25

Pricing over time

Input & output per 1M tokens · step-line = price changes

$0.7100

input / 1M

— stable

$4.25

output / 1M

— stable

2026-06-142026-07-192026-08-30
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)1058 / avg 844
132650

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

ownedBy: Qwen
Section 06

Availability

Availability

No measurements yet

We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 1 runs
1 correct0 partial0 wrong100% accuracy
claude-sonnet-4-542/100 · 66 runs
22 correct2 partial42 wrong33% accuracy
2026-08-30

Qwen3.5-397B-A17B reaches perfect 100/100 score with faster response times

Qwen3.5-397B-A17B has achieved a perfect quality score of 100 out of 100 in the current benchmark window, marking a substantial 15.5-point improvement from the previous period's 84.5 score. This advancement represents continued momentum following earlier quality gains. Response latency has improved significantly, with the median response time dropping from 5064ms to 2719ms, a 46% reduction that enhances the model's practical usability. The coding category maintained its strong performance at 100, up from 96 in the previous window. However, it should be noted that the current window contains only a single test run compared to three runs previously, which means these results may be less statistically robust. The absence of multilingual category data in the current window, which previously scored 73, leaves questions about performance stability across diverse language tasks. Users should consider that while the perfect score and improved latency are encouraging signs, additional test runs would provide greater confidence in the consistency of these results. The model appears to be trending positively in both quality and speed metrics.

Quality

100.0

Latency p50

2,719 ms

Test runs

1

Perfect 100/100 quality score 46% faster response times Coding improved to 100 Only 1 test run completed
Last automated test
Sep 5, 2026 · 08:00 UTC · Speed benchmark
P50 latency
189 ms
P95 latency
729 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·September 5, 2026