GLM-4.6V-Flash

Z.ai·Pending·Released 2025-12-08textvisionvideo
Context128k
Max output32k
Input / 1M$0 ↗
Output / 1M$0 ↗
Cached in$0
Benchmarks15

Benchmarks

Cited public sources · rank is exact-benchmark, apples-to-apples

AndroidWorld#4/542.7% ↗
ChartQAPro#3/462.6% ↗
CharXiv_Val-Reasoning#2/259.6% ↗
Design2Code#4/569.8% ↗
MathVista#4/682.7% ↗
MMBench V1.1#2/286.9% ↗
MMBrowseComp#2/27.1% ↗
MMLongBench-Doc#2/353% ↗
MMMU (Val)#2/271.1% ↗
MMMU_Pro#4/460.6% ↗
MMStar#3/574.7% ↗
OCRBench#3/684.7% ↗
OSWorld#7/821.1% ↗
VideoMMMU#4/570.1% ↗
WebVoyager#3/371.8% ↗

GLM-4.6V-Flash FAQ

What is GLM-4.6V-Flash?
GLM-4.6V-Flash is a large language model from Z.ai, released on December 8, 2025. It accepts text, image and video input.
When was GLM-4.6V-Flash released?
Z.ai released GLM-4.6V-Flash on December 8, 2025.
How much does GLM-4.6V-Flash cost?
GLM-4.6V-Flash costs $0 per 1M input tokens and $0 per 1M output tokens through the Z.ai API. Cached input costs $0 per 1M tokens. Prices are from Z.ai's official pricing as of October 5, 2026.
What is GLM-4.6V-Flash's context window?
GLM-4.6V-Flash has a 128K-token context window and can generate up to 32K output tokens.
How does GLM-4.6V-Flash score on benchmarks?
GLM-4.6V-Flash has 15 cited benchmark scores, including AndroidWorld 42.7%, ChartQAPro 62.6% and CharXiv_Val-Reasoning 59.6%. Each score links to its source.

Access GLM-4.6V-Flash and every other model through one endpoint with automatic failover: Respan gateway.