Back to Blog
Artificial Intelligence

Xiaomi Opens MiMo-V2.6 Under MIT, and Its Light Model Reads a Million Tokens for 14 Cents

Origami TeamEditorial Team
7 min read
Xiaomi Opens MiMo-V2.6 Under MIT, and Its Light Model Reads a Million Tokens for 14 Cents
Like what we publish? Pin Origami as a preferred source on Google.Add as a preferred source on Google

Xiaomi Opens MiMo-V2.6 Under MIT, and Its Light Model Reads a Million Tokens for 14 Cents

The short answer first: on September 22, 2026, Xiaomi launched the MiMo-V2.6 model family and published the weights under the MIT license, which allows commercial use and modification with almost no strings attached. The family has two models: Pro, with 1.02 trillion parameters, and Flash, with 309 billion. Both accept text, images, video and audio in a single model, with a context window of up to one million tokens. On OpenRouter, Flash starts at $0.14 per million input tokens. For your business that means two things: a self-hosting option with a comfortable license, and a very cheap API option for tasks that mix voice, images and text.

What Xiaomi actually released

Both models use a mixture-of-experts (MoE) architecture, meaning only a small share of the parameters is active for each token. That is what keeps running costs far below what the total size suggests:

  • MiMo-V2.6-Pro: 1.02 trillion parameters in total, 42 billion active per step, and a 1M-token context.
  • MiMo-V2.6-Flash: 309 billion parameters in total, 15 billion active, with the same context.
  • Pro-UltraSpeed: a version of Pro that Xiaomi says runs inference up to 20 times faster, offered through the API at ten times the price.

The company did not stop at the weights. It also published more than 7,000 reinforcement-learning training environments covering software engineering, vulnerability reproduction, knowledge work and web design, along with the training framework itself, the technical report, and a small 9-billion-parameter distilled model built on Qwen. The weights are published in FP8 on Hugging Face and can be downloaded without requesting access.

Where its performance really stands

Artificial Analysis gives Pro a score of 46.32 on its Intelligence Index. Xiaomi says that score puts it ahead of Kimi K3 and Qwen3.8 Max, making it the most capable open-source model at launch, while acknowledging that it still trails Claude Fable 5.1 and GPT-6 Astra.

The comparison table on the model card, however, measures it against the previous generation of closed models, not the current one. Here are its numbers against Claude Opus 5:

  • AutomationBench (task automation): 53.1 vs 50.3.
  • Terminal Bench 2.1 (command-line work): 89.9 vs 89.1.
  • DeepSWE (fixing real software bugs): 71.9 vs 74.0.
  • OSWorld-Verified (computer use): 82.0 vs 83.4.
  • The harder Terminal Bench 4.0: 34.9 vs 49.0.

The honest summary: a model close to last generation's top tier on agent and coding tasks, and clearly behind on the hardest ones. Do not read the announcement as an equal replacement for the latest closed models.

Price: the number that changes the operating math

On OpenRouter, each million input tokens and each million output tokens cost:

  • Flash: $0.14 input and $0.28 output.
  • Pro: $0.435 input and $0.87 output.
  • Pro-UltraSpeed: $4.35 input and $8.70 output.

An illustrative example, using our own assumptions: a customer-service assistant handling 1,000 conversations a day, each with about 3,000 input tokens and 500 output tokens. That adds up to 3 million input tokens and half a million output tokens a day. On Flash, the cost is about $0.56 a day, roughly SAR 63 a month. On Pro, about $1.74 a day, roughly SAR 196 a month. Real figures will move with the length of your conversations and with context caching, but the order of magnitude is what matters: cost is no longer the barrier to trying it. Xiaomi says Pro costs between one-twentieth and one-sixtieth of foreign models at the same intelligence level; that is the company's claim, not an independent measurement.

One model that hears, sees and reads

The part worth paying attention to is that one model understands all four input types. Today, if you want a system that analyzes customer call recordings, you usually need a speech-to-text service first, then a language model to analyze the transcript. The same goes for a photo of an invoice taken by a sales rep on a phone: one service to extract the text from the image, then a model to understand it. Every link in that chain is an extra cost and another point of failure.

A multimodal model shortens the chain: the call recording goes in as it is, the receipt photo as it is, the video clip from the warehouse camera as it is. That simplifies building systems such as call-quality review, automatic document entry into your accounting system, or checking photos of goods on receipt. Keep in mind, though, that the sources say nothing specific about its performance in Arabic or in Saudi dialects, and that is exactly what you should test on your own data before deciding anything.

Running it on your own servers: possible, with a budget

The MIT license means you can run the model inside your own infrastructure, modify it and build it into a commercial product. That matters for anyone processing personal data, because Saudi Arabia's Personal Data Protection Law places controls on transferring data outside the Kingdom, and self-hosting keeps it inside your environment.

The size is real, though. The Pro model card shows a reference setup that splits the model across 8 GPUs with vLLM, or across two nodes with SGLang. Flash needs a split across 4 GPUs. The card does not name the GPU type, but a model this size realistically needs top-tier AI servers. That is why Flash, with its 15 billion active parameters, is the realistic self-hosting candidate for most businesses. We covered the hardware math in detail in how to run a large AI model on your own machine.

What to watch before you rely on it

  • Security: Pro lags well behind on vulnerability-exploitation tests, scoring 47.9 on ExploitBench against 70.0 for Opus 5. If your task is security review of code, it is not your first choice.
  • Data path: using it through an API means your customers' data passes through an outside provider's servers. For sensitive data, self-hosting or anonymizing before sending is a requirement, not an option.
  • Vendor numbers: most benchmark results were published by Xiaomi itself. Test it on a task from your own work, with questions drawn from your own data.

How we approach it at Origami

We do not tie the systems we build to a single model. We put a middle layer in place so the model can be swapped without rebuilding the application, then measure each candidate on a real task from the client's work: accuracy, cost per operation and response time. MiMo-V2.6 joins that shortlist now, especially for tasks that combine voice, images and text, and wherever the monthly bill is what keeps an AI project from getting started. If you have a task like that and want to know its real cost before you commit, get in touch.

Sources

#MiMo-V2.6#Open-Source AI Models#Multimodal AI#AI Costs

Frequently asked questions

What is Xiaomi's MiMo-V2.6?+

A family of AI models Xiaomi launched on September 22, 2026, with weights published under the MIT license. It includes Pro, with 1.02 trillion parameters (42 billion active), and Flash, with 309 billion (15 billion active). Both understand text, images, video and audio with a one-million-token context.

How much does MiMo-V2.6 cost through an API?+

On OpenRouter, Flash costs $0.14 per million input tokens and $0.28 for output, while Pro costs $0.435 for input and $0.87 for output. An assistant handling a thousand average conversations a day could cost under SAR 70 a month on Flash, depending on conversation length.

Can I run it on my company's own servers?+

Yes. The MIT license allows running, modifying and commercial use. However, the reference setups on the Pro model card split it across 8 GPUs, while Flash needs a split across 4, which makes Flash the realistic self-hosting choice for most businesses.

Is it better than ChatGPT and Claude?+

No. It comes close to previous-generation models such as Claude Opus 5 on coding and automation tasks and falls behind on the hardest ones, and Xiaomi itself acknowledges it trails Claude Fable 5.1 and GPT-6 Astra. Its real advantages are price, the open license, and understanding voice and images in a single model.

Follow Origami in Google

Pin Origami as a preferred source and our articles will surface first for you in Google Search and Top Stories.

Add as a preferred source on Google

Related articles

Have a project in mind?

We build custom systems, apps and websites for your business. Tell us your idea and we will give you a straight answer on it.

One session. Twenty minutes. No commitments.