Skip to content
Tier B — Production
Runs in:FranceMade in:China
OVH AI Endpoints (GRA)

Qwen3-32B

Tier B — Production

Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency100 runs
367139924323464449608-2109-14ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

18%
Coding
judge mean 79
14%
Creative
judge mean 66
21%
Factual
judge mean 53
27%
Multilingual
judge mean 88
24%
Reasoning
judge mean 86

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing

What you pay per million tokens when you use this model on Tokonomix, plus an estimate for a typical conversation.

💰
API rates — Qwen3-32B
$0.2100 per 1M input tokens
$0.6000 per 1M output tokens
≈ $0.0002 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$0.2100
per 1M output tokens$0.6000
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)365 / avg 397
540169

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

ownedBy: Qwen
Section 06

Availability

Availability

How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.

Last 7 days

100.0%

n=6

Last 30 days

100.0%

n=6

Median response time

50,588ms

n=6

Based on 391 measurements over the last 30 days.

Technical details

Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.

Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.

Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.

Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.

Total calls (30d)

6

OK responses (30d)

6

Total calls (7d)

6

OK responses (7d)

6

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a95/100 · 1 runs
1 correct0 partial0 wrong100% accuracy
claude-sonnet-4-576/100 · 82 runs
51 correct16 partial15 wrong62% accuracy
2026-09-13

Qwen3-32B drops to 56.8 quality, latency improves 17% to 18.0s

Qwen3-32B shows a mixed performance shift in this benchmark window, with overall quality declining slightly from 58.0 to 56.8 while latency improved meaningfully from 21.7s to 18.0s. The 17% latency improvement brings response times closer to acceptable levels, though they remain notably high for a 32B parameter model. Category performance reveals dramatic swings: reasoning capability jumped impressively to 89 from a previous coding score of 52, and coding maintained strong performance at 72. However, factual accuracy collapsed to just 10, a concerning regression that suggests significant reliability issues with knowledge-based queries. The previous window's multilingual strength at 91 and creative capabilities are not represented in current testing categories, making direct comparison incomplete. This benchmarking period captures only 5 test runs, matching the previous window's sample size. Users should exercise caution with factual queries given the severe accuracy drop, while those prioritizing reasoning tasks may find improved utility. The latency gains are welcome but the model still requires substantial patience per request. Overall trajectory remains uncertain given the quality decline despite operational improvements.

Quality

56.8

Latency p50

17,953 ms

Test runs

5

Latency improved 17% Quality dropped to 56.8 Factual accuracy collapsed to 10 Reasoning strong at 89
Last automated test
Sep 14, 2026 · 20:02 UTC · Speed benchmark
P50 latency
548 ms
P95 latency
622 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·September 14, 2026