MiMo-V2.5
text+image+audio+video->text
MiMo-V2.5 is a general-purpose chat-tuned model. Available via xiaomi. 1,050,000-token context. Budget pricing.
MiMo-V2.5 is a multimodal AI model from xiaomi. It costs $0.140 per million input tokens and $0.280 per million output tokens (blended $0.182/M), with a 1,050,000-token context window.
INPUT
$0.140/M
per million input tokens
OUTPUT
$0.280/M
per million output tokens
CONTEXT
1,050,000
tokens
What it is good at
- Solid general chat performance
- Reasonable instruction following
- Multimodal: handles images alongside text
Typical use cases
- General chat
- Drafting and rewriting
- Personal productivity
Benchmarks
No published benchmark scores tracked for this model yet. Frontier and reasoning models from major providers have scores; smaller models and inference-host variants typically inherit the underlying open-weights score.
More from xiaomi
See all 1 →Frequently asked questions
How much does MiMo-V2.5 cost?
MiMo-V2.5 costs $0.140 per million input tokens and $0.280 per million output tokens, for a blended reference rate of $0.182 per million tokens.
What is MiMo-V2.5's context window?
MiMo-V2.5 supports up to 1,050,000 tokens of context in a single request.
What is MiMo-V2.5 best for?
MiMo-V2.5 is well suited to Solid general chat performance, Reasonable instruction following and Multimodal: handles images alongside text.
Who makes MiMo-V2.5?
MiMo-V2.5 is developed and served by xiaomi.