Skip to content
Runs in:SGMade in:China
Alibaba Cloud Qwen (DashScope Intl)

Qwen3.7 Max

Tier A — Frontier · 1M tokens

Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency120 runs
768593911109162802145008-1209-11ms
Section 02

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — Qwen3.7 Max
$2.50 per 1M input tokens
$7.50 per 1M output tokens
≈ $0.0030 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$2.50
per 1M output tokens$7.50

Pricing over time

Input & output per 1M tokens · step-line = price changes

$2.50

input / 1M

▲ +100% since first

$7.50

output / 1M

▲ +100% since first

2026-08-022026-08-092026-08-09
Input
Output
Price change
⟳ synced weekly
Section 03

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)157 / avg 169
25850

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 04

Availability

Availability

No measurements yet

We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.

Section 05

Tokonomix benchmark verdicts

2026-08-09

Qwen3.7 Max shows slight performance decline across key benchmarks

Qwen3.7 Max has experienced measurable performance decreases across multiple evaluation categories in the latest benchmark window. The model's overall score dropped from 68.1 to 66.3, reflecting a broader pattern of declining capabilities. Coding performance, previously a standout strength, fell from 72.9 to 70.5, while instruction following decreased from 77.7 to 75.8. Mathematical reasoning also declined from 71.9 to 70.3, and creative writing saw a reduction from 53.8 to 51.9. The model maintained its position in hard prompts with a modest improvement from 59.8 to 60.2, suggesting some stability in handling complex tasks. Style control remained relatively steady at 74.4 compared to the previous 75.0. Speed metrics show the model continues to operate quickly with a time to first token of 0.58 seconds and total time of 10.79 seconds. Despite these performance regressions, Qwen3.7 Max remains a capable model particularly for coding applications, though users should be aware of the downward trajectory in several key competency areas. The changes suggest potential model updates or infrastructure adjustments that have affected output quality.

Quality

Latency p50

Test runs

0

Overall score declined to 66.3 Coding performance dropped to 70.5 Instruction following decreased Hard prompts score improved slightly
Last automated test
Sep 11, 2026 · 08:02 UTC · Speed benchmark
P50 latency
1272 ms
P95 latency
1292 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·September 11, 2026