Snelheidsanalyse
Latency gemeten over alle benchmark-runs. P50 (mediaan) en P95 (95e percentiel) geven een realistisch beeld van de responssnelheid onder normale en piekbelasting.
Kwaliteitsscores
Hoe dit model zich per promptcategorie verhoudt tot de rest van het veld, uit een paarsgewijze berekening over dezelfde prompts. Het ruwe jurycijfer staat onder elk getal.
Winstpercentage per categorie: hoe vaak dit model een gemiddeld model verslaat op een prompt uit die categorie. 50% is gemiddeld, geen onvoldoende. Het is geen percentage goede antwoorden.
Prijzen
Wat je per miljoen tokens betaalt als je dit model via Tokonomix gebruikt, plus een schatting voor een gemiddeld gesprek.
Tokens per seconde
Doorvoersnelheid in tokens per seconde, afgeleid uit gemeten P50-latency. Hogere waarden zijn beter; fluctuaties weerspiegelen serverbelasting bij de provider.
Geschat uit P50-latency × 200 output-tokens — het absolute getal hangt af van deze aanname; de trend is wat telt.
Mogelijkheden
Beschikbaarheid
Beschikbaarheid
Nog geen meetdata
Er zijn nog niet genoeg API-aanroepen geregistreerd om beschikbaarheidsstatistieken voor dit model te tonen. Data verschijnt zodra het model live verkeer ontvangt.
Tokonomix benchmark-oordelen
Quality falls 9.5 points to 75.2 with near-doubled latency
Qwen3-Coder-30B-A3B-Instruct demonstrates significant performance degradation in this benchmark window. Overall quality dropped from 84.7 to 75.2, representing a 9.5 point decline that follows a previous 6.5 point decrease. This marks a concerning downward trend across consecutive windows. Latency deteriorated substantially, with p50 response times increasing 93% from 2009ms to 3886ms, nearly doubling the wait time for users. Category performance shows mixed results with sharp variations. Reasoning achieved a perfect 100 score, indicating strong logical capabilities. However, coding performance plummeted from 94 to 69, a 25 point drop that undermines the model's core positioning as a coding specialist. Factual accuracy scored 57, though this represents a new category without direct comparison. The previous window's multilingual and creative categories were not tested in the current period. The combination of declining quality metrics and significantly increased latency suggests potential infrastructure or model configuration issues. Users should expect notably slower responses and reduced coding performance compared to the previous benchmark period. The perfect reasoning score provides limited consolation given the substantial regression in the model's primary coding capabilities.
Kwaliteit
75.2
Latency p50
3,886 ms
Testruns
5
Qwen3-Coder-30B-A3B-Instruct
door OVH AI Endpoints (GRA)
- Contextvenster
- — tokens
- Inputprijs
- $0.1900 / 1M
- Outputprijs
- $0.6800 / 1M
- Tier
- Tier B — Productie
- Modaliteit
- Tekst
- API-type
- REST · streaming
- Benchmark-runs
- 522
Meer van OVH AI Endpoints (GRA)