Skip to content
Tier B — Production
Runs in:USMade in:United States

Retirement date

The provider lists May 7, 2027 as a tentative retirement date. The model is available as usual today.

Google Gemini

Gemini 3.1 Flash Lite

Tier B — Production · 1.048576M tokens

Tokonomix Editorial Team·Reviewed by Mes Kalkan·
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency96 runs
288520752984121608-2209-14ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

55%
Coding
judge mean 93
42%
Creative
judge mean 82
70%
Factual
judge mean 85
59%
Multilingual
judge mean 99
74%
Reasoning
judge mean 99

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing

What you pay per million tokens when you use this model on Tokonomix, plus an estimate for a typical conversation.

💰
API rates — Gemini 3.1 Flash Lite
$0.4900 per 1M input tokens
$2.93 per 1M output tokens
≈ $0.0009 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$0.4900
per 1M output tokens$2.93
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)369 / avg 389
688208

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

toolssource: litellmvisionjson modepdf inputreasoningaudio inputjson schemaparallel toolsprompt cachingoutputTokenLimit: 65536max output tokens: 65536
Section 06

Availability

Availability

How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.

Last 7 days

100.0%

n=8

Last 30 days

100.0%

n=8

Median response time

5,288ms

n=8

Based on 321 measurements over the last 30 days.

Technical details

Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.

Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.

Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.

Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.

Total calls (30d)

8

OK responses (30d)

8

Total calls (7d)

8

OK responses (7d)

8

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 1 runs
1 correct0 partial0 wrong100% accuracy
claude-sonnet-4-593/100 · 82 runs
71 correct6 partial5 wrong87% accuracy
2026-09-13

Quality gains of 6.8 points offset by 20% latency regression

Gemini 3.1 Flash Lite demonstrates meaningful quality improvements in this benchmark window, climbing from 77.2 to 84.0 overall. The most striking change appears in reasoning performance, which reached a perfect 100 score compared to its absence in previous testing. Factual response quality measured at 65, representing a new category in this evaluation period. Coding performance decreased from 95 to 87, indicating some regression in code generation capabilities despite the overall quality gains. Latency characteristics shifted notably, with p50 response times increasing from 1472ms to 1762ms, representing a 20% slowdown. This timing regression may impact user experience in latency-sensitive applications, though the model remains within reasonable response parameters for most use cases. The model previously excelled at multilingual tasks with a score of 92, though this category was not evaluated in the current window. Creative output had scored 45 previously, suggesting room for improvement in generative tasks. The current benchmark window shows a model that has strengthened its reasoning capabilities and maintained solid coding performance, though users should be aware of the latency tradeoff and the noted decline in code generation quality from previous levels.

Quality

84.0

Latency p50

1,762 ms

Test runs

5

Quality improved 6.8 points Reasoning reached perfect score Latency increased 20% Coding dropped from 95 to 87
Section 08

Full model profile

Gemini 3.1 Flash-Lite: Google's high-volume workhorse, built for speed over depth

Gemini 3.1 Flash-Lite is the lightest model in Google's Gemini 3 line, built on the Gemini 3 Pro foundation and positioned by Google as the fastest, most cost-efficient model in that generation. A preview version shipped March 3, 2026, and the generally available version sold here followed; Google has not published a specific GA date. Google measures its speed gains against Gemini 2.5 Flash, and positions the model for workloads that run at scale and need an answer fast rather than for the hardest reasoning tasks.

Where it shines

Google recommends this tier for translation, content moderation, generating user interfaces and dashboards, running simulations, and general instruction-following at volume — tasks where per-request latency and throughput matter more than squeezing out the last point of reasoning quality. According to Google, the model delivers 2.5x faster time-to-first-token and 45% higher output speed than Gemini 2.5 Flash, which is the practical reason to reach for it: a pipeline sending thousands of requests a day benefits far more from consistent low latency than from marginal accuracy gains on any single call.

Despite the "lite" label, Google reports respectable scores on general knowledge and multimodal reasoning: 86.9% on GPQA Diamond and 76.8% on MMMU-Pro, according to Google's own published figures. It is genuinely multimodal on the input side, accepting text, images, audio, and video, with a 1,000,000-token context window for input.

Under the hood

Output is capped at 65,536 tokens and is text-only — this tier does not generate images or audio, only reasons over them. Tool calling, structured JSON output against a schema, and prompt caching are all supported, so it fits into pipelines that need reliable structured extraction rather than free-form chat.

Where it falls short

The one clearly documented weak point is long-context retrieval. On Google's own MRCR v2 long-context benchmark, the model scores 60.1% at a 128k-token context and drops to 12.3% at the full 1,000,000-token window. In practical terms: the context window is large enough to accept a very long document, but the model's ability to accurately pull a specific detail back out of that document degrades substantially as the input approaches the ceiling. For workloads that genuinely need reliable recall near the top of a million-token window, this is not the right tier — reach for a larger Gemini model instead. It is also, by design, not the model to reach for on the hardest reasoning or agentic-coding tasks; Google built it for volume and speed, not frontier capability.

When to pick it

Pick Flash-Lite for high-volume, latency-sensitive work: bulk translation, moderation queues, classification at scale, UI or dashboard generation from structured input, or any pipeline where the same kind of request runs thousands of times a day. It is a poor fit for tasks that need either frontier reasoning or reliable recall from very long documents.

Alternatives worth comparing

Within Google's own lineup, Gemini 3 Pro is the base model this tier is distilled from, and is the better choice when a task needs deeper reasoning than the Lite tier can reliably deliver. Google has also since shipped a newer Gemini 3.5 Flash-Lite, which is worth checking against this generation on any workload sensitive to the latest quality improvements at the same speed-focused tier.

A note on naming

Tokonomix sells the generally available gemini-3.1-flash-lite, not the earlier preview build. The preview was a separate, time-boxed release Google used to gather feedback before GA; behavior and quality between preview and GA can differ, so results measured against the preview should be re-verified against the GA model before being relied on in production.

Last automated test
Sep 14, 2026 · 20:03 UTC · Speed benchmark
P50 latency
542 ms
P95 latency
671 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·September 14, 2026