GLM-4V
Vision
Vision GLM-4. Chinese-language multimodal model with document and chart understanding.
GLM-4V is a multimodal AI model from Zhipu AI. It costs $7.000 per million input tokens and $7.000 per million output tokens (blended $7.000/M), with a 2,000-token context window.
INPUT
$7.000/M
per million input tokens
OUTPUT
$7.000/M
per million output tokens
CONTEXT
2,000
tokens
What it is good at
- Vision-capable
- Strong Chinese OCR
- 128K context
Typical use cases
- Chinese document parsing
- Vision QA in Chinese
Benchmarks
vs. best public score
Scores inherited from GLM-4 Plus — this is a hosted variant of the same open-weights model, so the underlying benchmark scores are identical.
Hand-curated from each provider's published reports and public leaderboards. Methodology varies across sources, treat as directional rather than authoritative.
More from Zhipu AI
See all 5 →Frequently asked questions
How much does GLM-4V cost?
GLM-4V costs $7.000 per million input tokens and $7.000 per million output tokens, for a blended reference rate of $7.000 per million tokens.
What is GLM-4V's context window?
GLM-4V supports up to 2,000 tokens of context in a single request.
What is GLM-4V best for?
GLM-4V is well suited to Vision-capable, Strong Chinese OCR and 128K context.
Who makes GLM-4V?
GLM-4V is developed and served by Zhipu AI.