Archivé
Ce modèle a été retiré par le fournisseur. Les données historiques sont conservées.
Plus disponible depuis le 28 juin 2026.
Llama-3.1-8B-Instruct
Historique des tarifs
Tarifs directs du fournisseur par million de tokens, plus une estimation du coût d'une conversation typique.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$0.1000
input / 1M
— stable
$0.1000
output / 1M
— stable
Capacités
Disponibilité
Disponibilité
Pas encore de données
Nous n'avons pas encore enregistré suffisamment d'appels API pour afficher les statistiques de disponibilité de ce modèle. Les données apparaîtront dès que le modèle reçoit du trafic en direct.
Verdicts benchmark Tokonomix
Quality drops 29 points as performance degrades across all categories
Llama-3.1-8B-Instruct by OVH AI Endpoints has experienced a significant decline in performance this benchmark window. The overall quality score plummeted from 99.0 to 70.3, representing a 28.7-point drop that affects the model's competitive standing. The degradation is evident across all measured categories, with factual accuracy scoring just 57, reasoning at 74, and multilingual capabilities at 80. This contrasts sharply with the previous window where coding achieved 100, multilingual scored 97, and reasoning reached 100. The current window shows a different category composition, making direct comparisons complex, but the overall trend is unmistakably negative. On a positive note, latency has improved slightly from 9119ms to 7942ms at the median, offering users marginally faster response times. However, this speed gain is overshadowed by the substantial quality regression. Testing consistency remains stable with five runs in both windows. Users relying on this endpoint should be aware of the current performance limitations, particularly for fact-dependent tasks where the model now scores below 60. The cause of this regression warrants investigation to determine whether it stems from infrastructure changes, model configuration, or other factors.
Quality
70.3
Latency p50
7,942 ms
Test runs
5
Archivé
Ce modèle a été retiré par le fournisseur. Les données historiques sont conservées.
Plus disponible depuis le 28 juin 2026.
Llama-3.1-8B-Instruct
par OVH AI Endpoints (GRA)
- Fenêtre de contexte
- — tokens
- Prix d'entrée
- $0.1000 / 1M
- Prix de sortie
- $0.1000 / 1M
- Tier
- —
- Modalité
- Texte
- Type d'API
- REST · streaming
- Exécutions benchmark
- 152
Plus de OVH AI Endpoints (GRA)