Skip to content
Tier A — Frontier
Runs in:CNMade in:China
Z.ai (GLM / Zhipu)

GLM-4.6V (vision)

Tier A — Frontier · 205K tokens

Tokonomix Editorial Team·Reviewed by Mes Kalkan··

GLM-4.6V is the vision-capable model of the GLM-4.6 line: it accepts images alongside text and reasons over both. It pairs multimodal understanding with the GLM-4.6 line’s large context window, at a notably low price.

Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency108 runs
80140057209104121361608-0909-05ms
Section 02

Quality scores

How this model compares to the rest of the field on each prompt category, from a pairwise fit over the same prompts. The raw judge score sits underneath each number.

44%
Coding
judge mean 65
4%
Creative
judge mean 40
27%
Factual
judge mean 49
5%
Reasoning
judge mean 16

Win rate per category: how often this model beats a field-average model on a prompt from that category. 50% is average, not a failing grade. It is not a percentage of correct answers.

Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — GLM-4.6V (vision)
$0.3000 per 1M input tokens
$0.9000 per 1M output tokens
≈ $0.0004 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$0.3000
per 1M output tokens$0.9000

Pricing over time

Input & output per 1M tokens · step-line = price changes

$0.3000

input / 1M

— stable

$0.9000

output / 1M

— stable

2026-07-122026-08-092026-08-30
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)170 / avg 133
24845

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

jsonnotes: GLM emits a non-standard reasoning_content field beside content; read content for the answer.toolsvision
Section 06

Availability

Availability

No measurements yet

We haven't recorded enough API calls to show availability stats for this model. Data appears once the model starts receiving live traffic.

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 1 judge
Independent LLM judges evaluated this model on our weekly intelligence tests
claude-sonnet-4-546/100 · 26 runs
9 correct0 partial17 wrong35% accuracy
2026-08-30

GLM-4.6V debuts with vision, tools, and JSON mode support

GLM-4.6V enters the benchmark as a new multimodal model from Zhipu AI, bringing vision capabilities alongside tools and structured JSON output support. The model shows competent performance across vision tasks, though detailed benchmark scores are not yet available for comparative analysis. As a fresh entry, GLM-4.6V represents Zhipu's expansion into multimodal AI, following their text-focused GLM series. The addition of vision processing means users can now submit both text and image inputs for analysis, generation, and reasoning tasks. Tool calling functionality enables the model to interact with external functions and APIs, while JSON mode ensures structured outputs for applications requiring consistent data formats. Without historical performance data, it's premature to assess stability or trajectory, but the model's feature set positions it as a general-purpose multimodal assistant. Users evaluating GLM-4.6V should test it against their specific use cases, particularly for vision-language tasks, structured data extraction, and function calling workflows. The model joins an increasingly competitive space where vision capabilities are becoming standard rather than exceptional. Early adopters should monitor future benchmark windows to track performance evolution and capability refinements.

Quality

Latency p50

Test runs

0

Vision capabilities added Tool calling support enabled JSON mode now available New multimodal model entry
Section 08

Full model profile

GLM-4.6V: long-context vision from Zhipu

GLM-4.6V is the vision-capable model of the GLM-4.6 line: it accepts images alongside text and reasons over both. It pairs multimodal understanding with the GLM-4.6 line’s large context window, at a notably low price.

z.ai publishes GLM-4.6V at $0.30 per 1M input tokens and $0.90 per 1M output tokens — an aggressively low price for a vision-capable model.

It advertises a large ~200K-token context window, so images can be combined with substantial text in one call.

Architecture & training signals

GLM-4.6V is the vision variant of Zhipu AI’s GLM-4.6. It extends the text model with image input, producing text output that reasons over the supplied images. It exposes tool-calling and JSON over an OpenAI-compatible endpoint. Its output modality is text (it describes/reasons about images, it does not generate them). Like the rest of the GLM line, it returns a non-standard reasoning_content field alongside content in its OpenAI-compatible responses; integrations should read content for the final answer and treat reasoning_content as an optional trace.

Where it shines

  • Image understanding — describing, comparing, extracting information from pictures and screenshots.
  • Combining images with long text context in a single call.
  • Very low price for a multimodal model.

Where it falls short

  • It reads images, it does not generate them — for image generation use z.ai’s GLM-Image / CogView models.
  • No Tokonomix benchmark data yet; verify vision quality on your own images.
  • Non-EU hosting.

Real-world use cases

  • Document, chart and screenshot understanding.
  • Visual QA and multimodal analysis at scale.
  • A cheap vision proposer in a multimodal consensus panel.

Tokonomix benchmark snapshot

GLM-4.6V is newly registered on Tokonomix and not yet activated, so we have not run it through our weekly intelligence test or speed benchmark. There are no Tokonomix scores to report yet — and we will not invent any.

When it goes live, it enters the same weekly harness as every other model: identical prompts, an independent cross-family judge, and reproducible latency and cost measurements. Until then, treat the pricing and capability notes on this page as the vendor-published starting point, not as measured Tokonomix results.

EU privacy & data residency

GLM-4.6V is built by Zhipu AI (z.ai), a China-headquartered lab, and is served from non-EU infrastructure. This is important to state plainly: routing a prompt to this model is not an EU-data-residency or GDPR-sovereign choice, and Tokonomix will never tag it as one.

If your use case requires data to stay within the EU, pick a model whose provider is EU-hosted (for example our OVH or Azure-EU routes) rather than a GLM model. Tokonomix keeps z.ai out of every EU-only / sovereign routing set by design. Use GLM where its capability or price is the priority and cross-border processing is acceptable for that workload.

Verdict & alternatives

GLM-4.6V is a strong-value vision model: multimodal understanding plus long context at a low price. For a free vision option, GLM-4.6V Flash; for image generation (not understanding), the GLM-Image / CogView models.

Last automated test
Sep 5, 2026 · 08:02 UTC · Speed benchmark
P50 latency
1177 ms
P95 latency
1294 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·July 8, 2026