MultimodalAUTO PROFILELIVE PRICING

MiMo-V2.5

text+image+audio+video->text

MiMo-V2.5 is a general-purpose chat-tuned model. Available via xiaomi. 1,050,000-token context. Budget pricing.

MiMo-V2.5 is a multimodal AI model from xiaomi. It costs $0.140 per million input tokens and $0.280 per million output tokens (blended $0.182/M), with a 1,050,000-token context window.

INPUT
$0.140/M
per million input tokens
OUTPUT
$0.280/M
per million output tokens
BLENDED 70/30
$0.182/M
unchanged since 3 May
CONTEXT
1,050,000
tokens
What it is good at
  • Solid general chat performance
  • Reasonable instruction following
  • Multimodal: handles images alongside text
Typical use cases
  • General chat
  • Drafting and rewriting
  • Personal productivity

Benchmarks

No published benchmark scores tracked for this model yet. Frontier and reasoning models from major providers have scores; smaller models and inference-host variants typically inherit the underlying open-weights score.

More from xiaomi

See all 4 →

Frequently asked questions

How much does MiMo-V2.5 cost?

MiMo-V2.5 costs $0.140 per million input tokens and $0.280 per million output tokens, for a blended reference rate of $0.182 per million tokens.

What is MiMo-V2.5's context window?

MiMo-V2.5 supports up to 1,050,000 tokens of context in a single request.

What is MiMo-V2.5 best for?

MiMo-V2.5 is well suited to Solid general chat performance, Reasonable instruction following and Multimodal: handles images alongside text.

Who makes MiMo-V2.5?

MiMo-V2.5 is developed and served by xiaomi.

Terms used on this page