Zum Inhalt
Tier B — Produktion
Läuft in:FranceErstellt in:United States
OVH AI Endpoints (GRA)

Meta-Llama-3_3-70B-Instruct

Tier B — Produktion

Tokonomix-Redaktionsteam·Geprüft von Mes Kalkan··
Abschnitt 01

Geschwindigkeitsanalyse

Latenz über alle Benchmark-Läufe gemessen. P50 (Median) und P95 (95. Perzentil) zeigen ein realistisches Bild der Antwortgeschwindigkeit bei normaler und Spitzenlast.

P50-Latenz (Median)P95-Latenz105 runs
90172033514981661108-1009-05ms
Abschnitt 02

Qualitätswerte

Wie sich dieses Modell je Prompt-Kategorie zum übrigen Feld verhält, aus einem paarweisen Fit über dieselben Prompts. Die rohe Jury-Wertung steht unter jeder Zahl.

45%
Codegenerierung
Jury-Mittel 96
63%
Kreativ
Jury-Mittel 91
63%
Faktisch
Jury-Mittel 79
53%
Mehrsprachig
Jury-Mittel 98
74%
Schlussfolgern
Jury-Mittel 99

Gewinnrate je Kategorie: wie oft dieses Modell ein durchschnittliches Modell bei einem Prompt dieser Kategorie schlägt. 50% ist der Durchschnitt, keine schlechte Note. Es ist kein Prozentsatz richtiger Antworten.

Abschnitt 03

Preisverlauf

Direkte Provider-Tarife pro Million Tokens, plus eine typische Gesprächskostenschätzung.

💰
API-Tarife — Meta-Llama-3_3-70B-Instruct
$0.6700 pro 1M Input-Tokens
$0.6700 pro 1M Output-Tokens
≈ $0.0005 pro typischem Gespräch (800 Tokens)
Input- vs. Output-Preis (pro 1M Tokens)
pro 1M Input-Tokens$0.6700
pro 1M Output-Tokens$0.6700

Pricing over time

Input & output per 1M tokens · step-line = price changes

$0.6700

input / 1M

— stable

$0.6700

output / 1M

— stable

2026-06-142026-07-192026-08-30
Input
Output
Price change
⟳ synced weekly
Abschnitt 04

Tokens pro Sekunde

Durchsatz in Tokens pro Sekunde, abgeleitet aus gemessener P50-Latenz. Höhere Werte sind besser; Schwankungen spiegeln die Provider-seitige Last wider.

Durchsatz (Tokens / s)1639 / avg 1439
2200231

Geschätzt aus P50-Latenz × 200 Output-Tokens — die absolute Zahl hängt von dieser Annahme ab; entscheidend ist der Trend.

Abschnitt 05

Fähigkeiten

ownedBy: meta-llama
Abschnitt 06

Verfügbarkeit

Verfügbarkeit

Noch keine Messdaten

Es wurden noch nicht genug API-Aufrufe aufgezeichnet, um Verfügbarkeitsstatistiken für dieses Modell anzuzeigen. Daten erscheinen, sobald das Modell Live-Traffic erhält.

Abschnitt 07

Tokonomix-Benchmark-Urteile

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 1 runs
1 correct0 partial0 wrong100% accuracy
claude-sonnet-4-593/100 · 73 runs
66 correct3 partial4 wrong90% accuracy
2026-08-30

Quality rebounds 9.8 points to 85.3 with strong factual recovery

Meta-Llama-3.3-70B-Instruct demonstrates significant recovery in this benchmark window, climbing from 75.5 to 85.3 in overall quality. The most dramatic improvement comes in factual performance, which surged from 54 to a perfect 100, completely reversing the previous period's sharp decline. Coding capability also advanced from 81 to 92, showing solid progression. Reasoning performance debuts at an impressive 96, indicating strong logical processing capabilities. However, creative output presents a concerning new weakness at just 53, representing a notable gap in the model's capabilities. Multilingual support, previously measured at 92, was not evaluated in the current window. Latency improved modestly from 8442ms to 7416ms at the median, though response times remain relatively slow for real-time applications. The recovery suggests the previous window's issues may have been temporary, possibly related to infrastructure or configuration problems that have since been resolved. Users can expect reliable performance for technical, factual, and reasoning-heavy tasks, but should be aware of limitations in creative applications. The model appears well-suited for analytical workloads where accuracy and logical processing matter most.

Qualität

85.3

Latenz p50

7,416 ms

Testläufe

5

Factual score jumped to 100 Quality recovered 9.8 points Creative performance only 53 Latency improved by 1026ms
Letzter automatisierter Test
5. Sept. 2026 · 08:01 UTC · Geschwindigkeits-Benchmark
P50-Latenz
122 ms
P95-Latenz
123 ms
Fehler
0 / 6 Läufe
Zuletzt geprüft von Tokonomix-Team·5. September 2026