Skip to content
Tier A — Frontier
Runs in:USMade in:United States

Archived

This model has been discontinued by the provider. Historical data is preserved.

No longer available since October 20, 2026.

Google Gemini

Gemini 2.5 Pro

Tier A — Frontier · 1.048576M tokens

Tokonomix Editorial Team·Reviewed by Mes Kalkan··

Gemini 2.5 Pro is a large language model developed by Google and released as part of the Gemini model family. It is designed for general-purpose text generation tasks, including question answering, content creation, summarization, and code generation. The model represents an iterative improvement over earlier Gemini versions, incorporating architectural refinements and expanded training data to enhance performance across a range of natural language understanding and generation benchmarks. A key technical characteristic of Gemini 2.5 Pro is its extended context window of 1,048,576 tokens, equivalent to approximately one million tokens. This substantial context length enables the model to process and analyze extremely long documents, extensive codebases, or multi-turn conversations while maintaining coherence and relevance throughout the interaction. The model supports standard text generation capabilities, focusing on delivering reliable and consistent outputs for professional and developer use cases. Within Google's Gemini lineup, the 2.5 Pro variant occupies a position between the more compact Gemini Flash models, optimized for speed and efficiency, and the flagship Gemini Ultra models, which target state-of-the-art performance on the most demanding tasks. Gemini 2.5 Pro is positioned as a balanced option, offering strong reasoning and generation capabilities with significant context handling for applications that require processing large volumes of information without requiring the maximum performance tier.

Gemini 2.5 Pro is Google's mid-flagship reasoning model — frontier-class quality at a price point that actually fits production budgets.

Tokonomix benchmark summary
Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency104 runs
7943300580783131081908-1609-11ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

17%
Coding
judge mean 35
3%
Creative
judge mean 41
34%
Factual
judge mean 59
8%
Multilingual
judge mean 10
3%
Reasoning
judge mean 0
54%
Healthcare
judge mean 67

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — Gemini 2.5 Pro
$1.25 per 1M input tokens
$10.00 per 1M output tokens
≈ $0.0028 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$1.25
per 1M output tokens$10.00

Pricing over time

Input & output per 1M tokens · step-line = price changes

$1.25

input / 1M

— stable

$10.00

output / 1M

— stable

2026-06-212026-08-022026-09-06
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)166 / avg 162
25020

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Strengths & weaknesses

Drawn from benchmark results and aggregated community feedback on real use-cases.

Strengths

Strong reasoning across complex chainsMillion-token context windowExcellent multilingual coverageReliable code generation and reviewHigh instruction-following fidelityStrong on structured data extractionConsistent P50 latency under loadRobust safety alignment

Weaknesses

No native EU-hosted endpointNo image-generation modalityReal-time knowledge has a cutoffHigher per-token cost than open-source
Section 06

Capabilities

toolssource: litellmvisionjson modepdf inputreasoningaudio inputjson schemaprompt cachingoutputTokenLimit: 65536max output tokens: 65535
Section 07

Frequently asked questions

Pro is the heavier model — significantly stronger on multi-step reasoning, long-context retrieval, and code review, at roughly 5× the per-token cost. Flash is the right choice for chat/classification/summarisation workloads where latency budget and price matter more than peak reasoning quality.

For European workloads where context-window depth and multilingual coverage matter more than raw speed, Gemini 2.5 Pro is currently the strongest cost-quality balance in its tier.

Tokonomix benchmark summary
Section 08

Availability

Availability

How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.

Last 7 days

100.0%

n=9

Last 30 days

100.0%

n=17

Median response time

7,794ms

n=17

Based on 397 measurements over the last 30 days.

Technical details

Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.

Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.

Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.

Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.

Total calls (30d)

17

OK responses (30d)

17

Total calls (7d)

9

OK responses (7d)

9

Image quality control pilot (2026-06-10)

Recall

60.6%

n=300

False-alarm rate

3.6%

n=300

Section 09

Tokonomix benchmark verdicts

⚖️
Endorsed by 2 judges
Independent LLM judges evaluated this model on our weekly intelligence tests
cohere/command-a100/100 · 1 runs
1 correct0 partial0 wrong100% accuracy
claude-sonnet-4-531/100 · 130 runs
26 correct2 partial102 wrong20% accuracy
🏟️
Arena activity
Daily model arena — judged head-to-head
This month
As contestant
0Games played
0 / 0Won / lost
0Upvotes ▲
As judge
0Rounds as judge
Blind spots caught
All-time
As contestant
3Games played
3 / 0Won / lost
8Upvotes ▲
As judge
0Rounds as judge
Blind spots caught

Blind-spot detection activates as judges flag missed points in upcoming arena runs.

Monthly history (1)
MonthGames playedWon / lostUpvotes ▲Rounds as judge
2026-0633 / 080
2026-09-06

Gemini 2.5 Pro adds eight capabilities; benchmarks remain unavailable

Gemini 2.5 Pro has expanded its feature set significantly with the addition of eight new capabilities: tools, vision, json_mode, pdf_input, reasoning, audio_input, json_schema, and prompt_caching. This represents a substantial evolution in the model's functionality, moving from a text-only interface to a multimodal platform with structured output options and caching support. The additions suggest Google is positioning this model for more complex, production-ready applications that require document processing, visual understanding, and audio analysis. However, no performance benchmarks are available for either the current or previous evaluation windows. Without quantitative metrics on accuracy, reasoning ability, or task performance, users cannot assess how Gemini 2.5 Pro compares to competing models or evaluate whether the expanded capabilities translate to practical advantages. The absence of benchmark data makes it impossible to determine the model's strengths, weaknesses, or appropriate use cases based on objective measures. Organizations considering Gemini 2.5 Pro will need to conduct their own testing to validate performance for their specific requirements.

Quality

Latency p50

Test runs

0

Eight new capabilities added Multimodal support now available Structured output options enabled No performance benchmarks available
Section 10

Full model profile

Gemini 2.5 Pro — illustration 1
Gemini 2.5 Pro: the production top-tier of the Gemini line

Gemini 2.5 Pro (gemini-2.5-pro) is the production top-tier of Google's general-purpose Gemini family. A 1,048,576-token context window. Text-plus-vision input. Reasoning depth that competes with the Anthropic Opus line and OpenAI's larger GPT-5 variants.

If you talked to a Google solutions team in late 2025 about "the right Gemini for hard reasoning work in production," this is the model they pointed at. It is the Pro-tier workhorse and has carried the bulk of high-quality Gemini deployments through into 2026.

What sets it apart

The combination is specific:

  • A million-token context window with attention quality that holds up at depth. Not just a spec-sheet number — the long window is genuinely usable for synthesis across very long inputs.
  • Top-tier reasoning quality on multi-step tasks, structured output, and complex extraction work.
  • Native multimodal handling with strong vision quality. Documents, charts, diagrams, screenshots — handled with the care a Pro tier should bring.
  • Tool-use reliability strong enough to build production agent loops without writing defensive parsing layers.
  • Latency that is not best-in-class but is reasonable for a top-tier model with that context window.

The specific edge that distinguishes 2.5 Pro from competitors is the combination of long-context attention quality and native multimodal handling at the Pro tier. Some competitors match one or the other; few match both.

The million-token context, used well

The 1M window earns its keep on workloads that need it. Cross-document due diligence, full-codebase analysis, long-thread conversational state, multi-document synthesis — all map well onto 2.5 Pro.

Attention quality holds up well past 200k tokens of input and remains usable into the hundreds of thousands more. Past roughly 600k tokens, latency stretches out and the cost-per-call increases meaningfully. The rolling speed picture lives at /benchmarks/speed.

Two practical patterns matter at this context size:

  • Prompt caching is the right pattern for repeat queries against the same large corpus. Reloading 800k tokens of context on every call is expensive in wall-clock time even when the API call succeeds.
  • Structuring long input with clear section headers helps the model find what matters. The model is good at long-context attention, not magic; clear structure makes the difference between a usable answer and one that loses thread.

Vision input at the Pro tier

The vision quality on 2.5 Pro is genuinely strong. Document screenshots, scanned PDFs, dashboard captures, charts, diagrams. Table extraction is reliable on complex multi-row layouts. Chart description includes axis units, scale, and accurate magnitude estimates. Diagram reasoning — understanding flows and relationships in a visual representation — works at a level that lower Gemini tiers approach but do not match.

Handwriting is still the weak spot. So are very small UI elements. Anything where a human would struggle to read at full resolution benefits from a verification step.

For vision-heavy workloads where you also need reasoning over what is being seen, 2.5 Pro is one of the strongest current options in the field. The Opus tier in the Claude family is competitive on quality but does not match the 1M context window for vision-plus-text combined inputs.

Where it lands against the field

Against Anthropic top-tier. Claude Opus 4.5 and 4.6 are competitive on reasoning quality. Opus 4.7 ships the same 1M context window. Where they differ: Opus is more cautious in refusal posture, stronger on European-language administrative prose; 2.5 Pro is faster on most workloads and stronger on native multimodal handling.

Against OpenAI top-tier. GPT-5 competes on reasoning and is often faster on short prompts. 2.5 Pro wins on native multimodal beyond images and on the 1M context window being meaningfully usable.

Against newer Gemini snapshots. Gemini 3 Pro Preview is the move-up for the newest capabilities, with the usual preview caveats around rate limits and behavior stability. For production stability and well-understood behavior, 2.5 Pro remains the right starting point.

The category-level picture lives at /benchmarks/leaderboard and the per-category scores at /benchmarks/intelligence.

Where it is the wrong tool

High-volume cheap classification. Top-tier compute is the wrong-shape spend for sending millions of short prompts. Move down to 2.5 Flash or Flash-Lite for these workloads.

Real-time conversational voice. No native audio input. The voice pipeline guide on /usecases/voice covers the right architecture.

Code generation where best-in-class IDE-fit matters more than reasoning depth. 2.5 Pro is competent on code but not specialised. The model survey at /usecases/code covers the alternatives.

Self-hosted deployment. Google does not ship Gemini weights. For workloads that need on-prem deployment, the open-weight survey at /usecases/local is the right starting point.

Anything that needs sub-second response on very large inputs. Latency at depth in the context window is real; for time-sensitive applications, a smaller model with prompt-caching strategies may fit better.

Deployment notes

Standard Google Gemini API. REST, streaming, tool-use, structured output — all behave as expected for a top-tier model. The integration with broader Vertex AI tooling for monitoring, logging, and safety controls is clean.

Regional availability follows Google's Vertex AI pattern. EU regions are available on enterprise contracts. Off-the-shelf consumer API access does not pin a region. For hard residency constraints, the Vertex AI regional documentation is the right reference.

Pricing is at the top tier of the Gemini family. For high-volume workloads, the per-call cost is meaningful — the case for staying on 2.5 Pro versus moving up or down depends on whether your specific workload genuinely needs the top-tier quality.

Safety and content filtering follow Google's broader policies. The filter behavior is configurable for enterprise contracts but the default settings apply to off-the-shelf API access.

Picking it

Reach for Gemini 2.5 Pro when:

  • You need top-tier reasoning combined with the 1M context window.
  • The workload includes vision input on documents, charts, or diagrams that require careful reading.
  • Native multimodal handling matters more than the Claude-style refusal posture.
  • You are already on the Google stack and need a Pro-tier general-purpose model.

Pick something else when:

  • The workload fits in a Flash-tier model. Move down for cost.
  • You need the newest capabilities and can tolerate preview-tier behavior. Move up to 3 Pro Preview.
  • Refusal consistency and European-language administrative prose dominate. Move to Claude Opus.
  • The work is audio-native, voice-native, or video-native.

The summary. Gemini 2.5 Pro is the production top-tier choice for Google deployments that need real reasoning and real long-context handling. The newer 3.x previews may be more capable on specific benchmarks, but for stability, rate limits, and well-understood behavior across the production surface, 2.5 Pro is the right starting point.

Run it against alternatives on your own prompts at /live-test.


Editorial provenance

This deep-dive was reviewed through a 3-model cross-family consensus run on the Tokonomix consensus engine — Claude Opus 4.8 (Anthropic), GPT-5.4 (OpenAI), and Cohere Command-A — on 2026-06-10. Each model independently reviewed the factual claims; an independent judge (Claude Sonnet 4.6) synthesised their findings.

Consensus verdict: mostly accurate. Core specifications (1,048,576-token context window, text-plus-vision input, streaming, tool calls, JSON mode, native code execution and search grounding as separate API features) are well-grounded in public Google documentation. The council flagged two claims requiring qualification: (1) the flash-attention architecture characterisation is unverifiable — Google has not publicly disclosed its internal attention mechanism and this should be treated as editorial inference, not confirmed fact; (2) the "late 2025 production recommendation" and direct competitive comparisons to the Opus and GPT-5 lines lack cited benchmarks and should be read as positioning rather than measured results.

Full run record: content_generation_runs entries for page id 18. Methodology: /methodology.

Gemini 2.5 Pro — illustration 2Gemini 2.5 Pro — illustration 3
Last automated test
Sep 11, 2026 · 02:04 UTC · Speed benchmark
P50 latency
1204 ms
P95 latency
1562 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·June 10, 2026