Skip to content
Tier B — Production
Runs in:USMade in:United States

Archived

This model has been discontinued by the provider. Historical data is preserved.

No longer available since May 7, 2027.

Google Gemini

Gemini 3.1 Flash Lite

Tier B — Production · 1.048576M tokens

Tokonomix Editorial Team·Reviewed by Mes Kalkan··
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency58 runs
325548771993121608-2209-05ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

53%
Coding
judge mean 94
44%
Creative
judge mean 86
60%
Factual
judge mean 88
62%
Multilingual
judge mean 99
74%
Reasoning
judge mean 99

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — Gemini 3.1 Flash Lite
$0.2500 per 1M input tokens
$1.50 per 1M output tokens
≈ $0.0004 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$0.2500
per 1M output tokens$1.50

Pricing over time

Input & output per 1M tokens · step-line = price changes

$0.2500

input / 1M

— stable

$1.50

output / 1M

— stable

2026-06-142026-07-192026-08-30
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)409 / avg 375
608208

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

toolssource: litellmvisionjson modepdf inputreasoningaudio inputjson schemaparallel toolsprompt cachingoutputTokenLimit: 65536max output tokens: 65536
Section 06

Availability

Availability

No measurements yet

We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 1 runs
1 correct0 partial0 wrong100% accuracy
claude-sonnet-4-594/100 · 72 runs
64 correct5 partial3 wrong89% accuracy
2026-08-30

Significant quality gains across coding and reasoning with faster response times

Gemini 3.1 Flash Lite demonstrates substantial improvements in this benchmark window, with overall quality rising from 71.0 to 81.8 points. The model achieved particularly strong performance in reasoning (96) and factual tasks (94), while coding capabilities jumped from 73 to 92. Response latency improved by 39 percent, dropping from 1902ms to 1167ms at the median, making interactions noticeably more responsive. However, creative performance emerged as a relative weakness at 45 points, representing the lowest category score in the current window. The previous window's strong multilingual performance (92) was not measured in the current testing cycle, so continuity in that capability cannot be confirmed. With five test runs completed in each window, the results suggest meaningful progress in core technical tasks and computational reasoning, though the model appears optimized for analytical rather than generative creative work. Users requiring fast, accurate responses for coding assistance, factual queries, and logical reasoning will find substantial value, while those prioritizing creative writing or content generation may encounter limitations.

Quality

81.8

Latency p50

1,167 ms

Test runs

5

Quality improved 10.8 points Latency reduced by 39% Coding score jumped to 92 Creative performance at 45
Last automated test
Sep 5, 2026 · 08:01 UTC · Speed benchmark
P50 latency
489 ms
P95 latency
498 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·September 5, 2026