Archived
This model has been discontinued by the provider. Historical data is preserved.
No longer available since May 7, 2027.
Speed analysis
Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.
Quality scores
How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.
Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.
Pricing history
Direct provider rates per million tokens, plus a typical-conversation cost estimate.
Pricing over time
Input & output per 1M tokens · step-line = price changes
$0.2500
input / 1M
— stable
$1.50
output / 1M
— stable
Tokens per second
Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.
Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.
Capabilities
Availability
Availability
No measurements yet
We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.
Tokonomix benchmark verdicts
Significant quality gains across coding and reasoning with faster response times
Gemini 3.1 Flash Lite demonstrates substantial improvements in this benchmark window, with overall quality rising from 71.0 to 81.8 points. The model achieved particularly strong performance in reasoning (96) and factual tasks (94), while coding capabilities jumped from 73 to 92. Response latency improved by 39 percent, dropping from 1902ms to 1167ms at the median, making interactions noticeably more responsive. However, creative performance emerged as a relative weakness at 45 points, representing the lowest category score in the current window. The previous window's strong multilingual performance (92) was not measured in the current testing cycle, so continuity in that capability cannot be confirmed. With five test runs completed in each window, the results suggest meaningful progress in core technical tasks and computational reasoning, though the model appears optimized for analytical rather than generative creative work. Users requiring fast, accurate responses for coding assistance, factual queries, and logical reasoning will find substantial value, while those prioritizing creative writing or content generation may encounter limitations.
Quality
81.8
Latency p50
1,167 ms
Test runs
5
Archived
This model has been discontinued by the provider. Historical data is preserved.
No longer available since May 7, 2027.
Gemini 3.1 Flash Lite
by Google Gemini
- Context window
- 1.048576M tokens
- Input price
- $0.2500 / 1M
- Output price
- $1.50 / 1M
- Tier
- Tier B — Production
- Modality
- Text + vision
- API type
- REST · streaming
- Benchmark runs
- 144
More from Google Gemini
Similar models