GLM-4.6V is the vision-capable model of the GLM-4.6 line: it accepts images alongside text and reasons over both. It pairs multimodal understanding with the GLM-4.6 line’s large context window, at a notably low price.
z.ai publishes GLM-4.6V at $0.30 per 1M input tokens and $0.90 per 1M output tokens — an aggressively low price for a vision-capable model.
It advertises a large ~200K-token context window, so images can be combined with substantial text in one call.
Architecture & training signals
GLM-4.6V is the vision variant of Zhipu AI’s GLM-4.6. It extends the text model with image input, producing text output that reasons over the supplied images. It exposes tool-calling and JSON over an OpenAI-compatible endpoint. Its output modality is text (it describes/reasons about images, it does not generate them). Like the rest of the GLM line, it returns a non-standard reasoning_content field alongside content in its OpenAI-compatible responses; integrations should read content for the final answer and treat reasoning_content as an optional trace.
Where it shines
- Image understanding — describing, comparing, extracting information from pictures and screenshots.
- Combining images with long text context in a single call.
- Very low price for a multimodal model.
Where it falls short
- It reads images, it does not generate them — for image generation use z.ai’s GLM-Image / CogView models.
- No Tokonomix benchmark data yet; verify vision quality on your own images.
- Non-EU hosting.
Real-world use cases
- Document, chart and screenshot understanding.
- Visual QA and multimodal analysis at scale.
- A cheap vision proposer in a multimodal consensus panel.
Tokonomix benchmark snapshot
GLM-4.6V is newly registered on Tokonomix and not yet activated, so we have not run it through our weekly intelligence test or speed benchmark. There are no Tokonomix scores to report yet — and we will not invent any.
When it goes live, it enters the same weekly harness as every other model: identical prompts, an independent cross-family judge, and reproducible latency and cost measurements. Until then, treat the pricing and capability notes on this page as the vendor-published starting point, not as measured Tokonomix results.
EU privacy & data residency
GLM-4.6V is built by Zhipu AI (z.ai), a China-headquartered lab, and is served from non-EU infrastructure. This is important to state plainly: routing a prompt to this model is not an EU-data-residency or GDPR-sovereign choice, and Tokonomix will never tag it as one.
If your use case requires data to stay within the EU, pick a model whose provider is EU-hosted (for example our OVH or Azure-EU routes) rather than a GLM model. Tokonomix keeps z.ai out of every EU-only / sovereign routing set by design. Use GLM where its capability or price is the priority and cross-border processing is acceptable for that workload.
Verdict & alternatives
GLM-4.6V is a strong-value vision model: multimodal understanding plus long context at a low price. For a free vision option, GLM-4.6V Flash; for image generation (not understanding), the GLM-Image / CogView models.