Xiaomi’s MiMo V2.6 Model Debuts With Unprecedented Efficiency Gains


TL;DR

  • Open Models: Xiaomi’s MiMo V2.6 Pro and Flash models give developers downloadable AI and hosted access for agents that use tools.
  • Flash Pricing: Flash’s listed input and output rates are about one-third of Pro’s, with a smaller discount on cached input.
  • Benchmark Lead: Pro leads open-weight models on Artificial Analysis’ composite index, while Xiaomi’s own tests show performance varies by task.
  • Training Rewards: Xiaomi says its reinforcement learning rewards cleaner code and penalizes shortcuts that copy existing fixes instead of solving assigned problems.

Xiaomi has released its MiMo V2.6 AI models for software agents that use tools to carry out tasks. Developers can pay for hosted access or run downloaded models themselves, taking responsibility for the computing infrastructure. Through OpenRouter, the smaller Flash model offers output at $0.28 per million tokens.

The flagship MiMo-V2.6-Pro and smaller MiMo-V2.6-Flash variants both accept text, images, audio and video and return text. Their million-token context windows let them consider long documents, code and conversation histories within a request. Tokens are the chunks of information a model processes, rather than a fixed number of words.

Both repositories label the downloadable models with the MIT license. The weights, the learned parameters that make the models work, can be used and modified under that permissive license. Developers can also access the models through Xiaomi’s API or through OpenRouter, a service that connects requests to model-hosting providers.

Flash Cuts the Price of Repeated Agent Calls

Xiaomi positions Pro for complex, extended tasks and Flash for frequent calls and large workloads. That distinction matters for agents: a single assignment can involve repeated model calls as software reads files, uses tools and checks its work.

OpenRouter’s Xiaomi-provider listings on September 22 show three separate charges: new input, reused input served from a cache, and generated output. The rates below are in US dollars per million tokens.