Zum Inhalt
Tier B — Produktion
Läuft in:FranceErstellt in:China
OVH AI Endpoints (GRA)

Qwen3-Coder-30B-A3B-Instruct

Tier B — Produktion

Tokonomix-Redaktionsteam·Geprüft von Mes Kalkan··
Abschnitt 01

Geschwindigkeitsanalyse

Latenz über alle Benchmark-Läufe gemessen. P50 (Median) und P95 (95. Perzentil) zeigen ein realistisches Bild der Antwortgeschwindigkeit bei normaler und Spitzenlast.

P50-Latenz (Median)P95-Latenz100 runs
60792015780236403150008-2109-14ms
Abschnitt 02

Qualitätswerte

Wie sich dieses Modell je Prompt-Kategorie zum übrigen Feld verhält, aus einem paarweisen Fit über dieselben Prompts. Die rohe Jury-Wertung steht unter jeder Zahl.

55%
Codegenerierung
Jury-Mittel 92
36%
Kreativ
Jury-Mittel 78
57%
Faktisch
Jury-Mittel 77
53%
Mehrsprachig
Jury-Mittel 98
64%
Schlussfolgern
Jury-Mittel 93

Gewinnrate je Kategorie: wie oft dieses Modell ein durchschnittliches Modell bei einem Prompt dieser Kategorie schlägt. 50% ist der Durchschnitt, keine schlechte Note. Es ist kein Prozentsatz richtiger Antworten.

Abschnitt 03

Preise

Was Sie pro Million Tokens zahlen, wenn Sie dieses Modell über Tokonomix nutzen, plus eine Schätzung für ein typisches Gespräch.

💰
API-Tarife — Qwen3-Coder-30B-A3B-Instruct
$0.1900 pro 1M Input-Tokens
$0.6800 pro 1M Output-Tokens
≈ $0.0003 pro typischem Gespräch (800 Tokens)
Input- vs. Output-Preis (pro 1M Tokens)
pro 1M Input-Tokens$0.1900
pro 1M Output-Tokens$0.6800
Abschnitt 04

Tokens pro Sekunde

Durchsatz in Tokens pro Sekunde, abgeleitet aus gemessener P50-Latenz. Höhere Werte sind besser; Schwankungen spiegeln die Provider-seitige Last wider.

Durchsatz (Tokens / s)2439 / avg 1626
328420

Geschätzt aus P50-Latenz × 200 Output-Tokens — die absolute Zahl hängt von dieser Annahme ab; entscheidend ist der Trend.

Abschnitt 05

Fähigkeiten

ownedBy: Qwen
Abschnitt 06

Verfügbarkeit

Verfügbarkeit

Noch keine Messdaten

Es wurden noch nicht genug API-Aufrufe aufgezeichnet, um Verfügbarkeitsstatistiken für dieses Modell anzuzeigen. Daten erscheinen, sobald das Modell Live-Traffic erhält.

Abschnitt 07

Tokonomix-Benchmark-Urteile

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 1 runs
1 correct0 partial0 wrong100% accuracy
claude-sonnet-4-589/100 · 82 runs
68 correct4 partial10 wrong83% accuracy
2026-09-13

Quality falls 9.5 points to 75.2 with near-doubled latency

Qwen3-Coder-30B-A3B-Instruct demonstrates significant performance degradation in this benchmark window. Overall quality dropped from 84.7 to 75.2, representing a 9.5 point decline that follows a previous 6.5 point decrease. This marks a concerning downward trend across consecutive windows. Latency deteriorated substantially, with p50 response times increasing 93% from 2009ms to 3886ms, nearly doubling the wait time for users. Category performance shows mixed results with sharp variations. Reasoning achieved a perfect 100 score, indicating strong logical capabilities. However, coding performance plummeted from 94 to 69, a 25 point drop that undermines the model's core positioning as a coding specialist. Factual accuracy scored 57, though this represents a new category without direct comparison. The previous window's multilingual and creative categories were not tested in the current period. The combination of declining quality metrics and significantly increased latency suggests potential infrastructure or model configuration issues. Users should expect notably slower responses and reduced coding performance compared to the previous benchmark period. The perfect reasoning score provides limited consolation given the substantial regression in the model's primary coding capabilities.

Qualität

75.2

Latenz p50

3,886 ms

Testläufe

5

Quality dropped 9.5 points Latency increased 93% Coding score fell from 94 to 69 Perfect reasoning score achieved
Letzter automatisierter Test
14. Sept. 2026 · 20:02 UTC · Geschwindigkeits-Benchmark
P50-Latenz
82 ms
P95-Latenz
84 ms
Fehler
0 / 6 Läufe
Zuletzt geprüft von Tokonomix-Team·14. September 2026