Naar inhoud
Tier B — Productie
Draait in:FranceGemaakt in:China
OVH AI Endpoints (GRA)

Qwen3-Coder-30B-A3B-Instruct

Tier B — Productie

Tokonomix-redactie·Gecontroleerd door Mes Kalkan··
Sectie 01

Snelheidsanalyse

Latency gemeten over alle benchmark-runs. P50 (mediaan) en P95 (95e percentiel) geven een realistisch beeld van de responssnelheid onder normale en piekbelasting.

P50 latency (mediaan)P95 latency100 runs
60792015780236403150008-2109-14ms
Sectie 02

Kwaliteitsscores

Hoe dit model zich per promptcategorie verhoudt tot de rest van het veld, uit een paarsgewijze berekening over dezelfde prompts. Het ruwe jurycijfer staat onder elk getal.

55%
Code generatie
jurygemiddelde 92
36%
Creatief
jurygemiddelde 78
57%
Feitelijk
jurygemiddelde 77
53%
Meertaligheid
jurygemiddelde 98
64%
Redeneren
jurygemiddelde 93

Winstpercentage per categorie: hoe vaak dit model een gemiddeld model verslaat op een prompt uit die categorie. 50% is gemiddeld, geen onvoldoende. Het is geen percentage goede antwoorden.

Sectie 03

Prijzen

Wat je per miljoen tokens betaalt als je dit model via Tokonomix gebruikt, plus een schatting voor een gemiddeld gesprek.

💰
API-tarieven — Qwen3-Coder-30B-A3B-Instruct
$0.1900 per 1M input-tokens
$0.6800 per 1M output-tokens
≈ $0.0003 per typisch gesprek (800 tokens)
Input vs output prijs (per 1M tokens)
per 1M input-tokens$0.1900
per 1M output-tokens$0.6800
Sectie 04

Tokens per seconde

Doorvoersnelheid in tokens per seconde, afgeleid uit gemeten P50-latency. Hogere waarden zijn beter; fluctuaties weerspiegelen serverbelasting bij de provider.

Doorvoer (tokens / s)2439 / avg 1626
328420

Geschat uit P50-latency × 200 output-tokens — het absolute getal hangt af van deze aanname; de trend is wat telt.

Sectie 05

Mogelijkheden

ownedBy: Qwen
Sectie 06

Beschikbaarheid

Beschikbaarheid

Nog geen meetdata

Er zijn nog niet genoeg API-aanroepen geregistreerd om beschikbaarheidsstatistieken voor dit model te tonen. Data verschijnt zodra het model live verkeer ontvangt.

Sectie 07

Tokonomix benchmark-oordelen

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 1 runs
1 correct0 partial0 wrong100% accuracy
claude-sonnet-4-589/100 · 82 runs
68 correct4 partial10 wrong83% accuracy
2026-09-13

Quality falls 9.5 points to 75.2 with near-doubled latency

Qwen3-Coder-30B-A3B-Instruct demonstrates significant performance degradation in this benchmark window. Overall quality dropped from 84.7 to 75.2, representing a 9.5 point decline that follows a previous 6.5 point decrease. This marks a concerning downward trend across consecutive windows. Latency deteriorated substantially, with p50 response times increasing 93% from 2009ms to 3886ms, nearly doubling the wait time for users. Category performance shows mixed results with sharp variations. Reasoning achieved a perfect 100 score, indicating strong logical capabilities. However, coding performance plummeted from 94 to 69, a 25 point drop that undermines the model's core positioning as a coding specialist. Factual accuracy scored 57, though this represents a new category without direct comparison. The previous window's multilingual and creative categories were not tested in the current period. The combination of declining quality metrics and significantly increased latency suggests potential infrastructure or model configuration issues. Users should expect notably slower responses and reduced coding performance compared to the previous benchmark period. The perfect reasoning score provides limited consolation given the substantial regression in the model's primary coding capabilities.

Kwaliteit

75.2

Latency p50

3,886 ms

Testruns

5

Quality dropped 9.5 points Latency increased 93% Coding score fell from 94 to 69 Perfect reasoning score achieved
Laatste automatische test
14 sep 2026 · 20:02 UTC · Snelheidstest
P50 latency
82 ms
P95 latency
84 ms
Fouten
0 / 6 runs
Laatst beoordeeld door Tokonomix-team·14 september 2026