Xiaomi Opens MiMo-V2.6 Under MIT, and Its Light Model Reads a Million Tokens for 14 Cents

Xiaomi Opens MiMo-V2.6 Under MIT, and Its Light Model Reads a Million Tokens for 14 Cents
The short answer first: on September 22, 2026, Xiaomi launched the MiMo-V2.6 model family and published the weights under the MIT license, which allows commercial use and modification with almost no strings attached. The family has two models: Pro, with 1.02 trillion parameters, and Flash, with 309 billion. Both accept text, images, video and audio in a single model, with a context window of up to one million tokens. On OpenRouter, Flash starts at $0.14 per million input tokens. For your business that means two things: a self-hosting option with a comfortable license, and a very cheap API option for tasks that mix voice, images and text.
What Xiaomi actually released
Both models use a mixture-of-experts (MoE) architecture, meaning only a small share of the parameters is active for each token. That is what keeps running costs far below what the total size suggests:
- MiMo-V2.6-Pro: 1.02 trillion parameters in total, 42 billion active per step, and a 1M-token context.
- MiMo-V2.6-Flash: 309 billion parameters in total, 15 billion active, with the same context.
- Pro-UltraSpeed: a version of Pro that Xiaomi says runs inference up to 20 times faster, offered through the API at ten times the price.
The company did not stop at the weights. It also published more than 7,000 reinforcement-learning training environments covering software engineering, vulnerability reproduction, knowledge work and web design, along with the training framework itself, the technical report, and a small 9-billion-parameter distilled model built on Qwen. The weights are published in FP8 on Hugging Face and can be downloaded without requesting access.
Where its performance really stands
Artificial Analysis gives Pro a score of 46.32 on its Intelligence Index. Xiaomi says that score puts it ahead of Kimi K3 and Qwen3.8 Max, making it the most capable open-source model at launch, while acknowledging that it still trails Claude Fable 5.1 and GPT-6 Astra.
The comparison table on the model card, however, measures it against the previous generation of closed models, not the current one. Here are its numbers against Claude Opus 5:
- AutomationBench (task automation): 53.1 vs 50.3.
- Terminal Bench 2.1 (command-line work): 89.9 vs 89.1.
- DeepSWE (fixing real software bugs): 71.9 vs 74.0.
- OSWorld-Verified (computer use): 82.0 vs 83.4.
- The harder Terminal Bench 4.0: 34.9 vs 49.0.
The honest summary: a model close to last generation's top tier on agent and coding tasks, and clearly behind on the hardest ones. Do not read the announcement as an equal replacement for the latest closed models.
Price: the number that changes the operating math
On OpenRouter, each million input tokens and each million output tokens cost:
- Flash: $0.14 input and $0.28 output.
- Pro: $0.435 input and $0.87 output.
- Pro-UltraSpeed: $4.35 input and $8.70 output.
An illustrative example, using our own assumptions: a customer-service assistant handling 1,000 conversations a day, each with about 3,000 input tokens and 500 output tokens. That adds up to 3 million input tokens and half a million output tokens a day. On Flash, the cost is about $0.56 a day, roughly SAR 63 a month. On Pro, about $1.74 a day, roughly SAR 196 a month. Real figures will move with the length of your conversations and with context caching, but the order of magnitude is what matters: cost is no longer the barrier to trying it. Xiaomi says Pro costs between one-twentieth and one-sixtieth of foreign models at the same intelligence level; that is the company's claim, not an independent measurement.
One model that hears, sees and reads
The part worth paying attention to is that one model understands all four input types. Today, if you want a system that analyzes customer call recordings, you usually need a speech-to-text service first, then a language model to analyze the transcript. The same goes for a photo of an invoice taken by a sales rep on a phone: one service to extract the text from the image, then a model to understand it. Every link in that chain is an extra cost and another point of failure.
A multimodal model shortens the chain: the call recording goes in as it is, the receipt photo as it is, the video clip from the warehouse camera as it is. That simplifies building systems such as call-quality review, automatic document entry into your accounting system, or checking photos of goods on receipt. Keep in mind, though, that the sources say nothing specific about its performance in Arabic or in Saudi dialects, and that is exactly what you should test on your own data before deciding anything.
Running it on your own servers: possible, with a budget
The MIT license means you can run the model inside your own infrastructure, modify it and build it into a commercial product. That matters for anyone processing personal data, because Saudi Arabia's Personal Data Protection Law places controls on transferring data outside the Kingdom, and self-hosting keeps it inside your environment.
The size is real, though. The Pro model card shows a reference setup that splits the model across 8 GPUs with vLLM, or across two nodes with SGLang. Flash needs a split across 4 GPUs. The card does not name the GPU type, but a model this size realistically needs top-tier AI servers. That is why Flash, with its 15 billion active parameters, is the realistic self-hosting candidate for most businesses. We covered the hardware math in detail in how to run a large AI model on your own machine.
What to watch before you rely on it
- Security: Pro lags well behind on vulnerability-exploitation tests, scoring 47.9 on ExploitBench against 70.0 for Opus 5. If your task is security review of code, it is not your first choice.
- Data path: using it through an API means your customers' data passes through an outside provider's servers. For sensitive data, self-hosting or anonymizing before sending is a requirement, not an option.
- Vendor numbers: most benchmark results were published by Xiaomi itself. Test it on a task from your own work, with questions drawn from your own data.
How we approach it at Origami
We do not tie the systems we build to a single model. We put a middle layer in place so the model can be swapped without rebuilding the application, then measure each candidate on a real task from the client's work: accuracy, cost per operation and response time. MiMo-V2.6 joins that shortlist now, especially for tasks that combine voice, images and text, and wherever the monthly bill is what keeps an AI project from getting started. If you have a task like that and want to know its real cost before you commit, get in touch.
Sources
- Xiaomi — official MiMo-V2.6 announcement (September 22, 2026).
- MiMo-V2.6-Pro model card on Hugging Face (architecture, benchmarks and deployment settings).
- MiMo-V2.6-Flash model card on Hugging Face.
- Artificial Analysis — MiMo-V2.6-Pro evaluation page.
- OpenRouter — MiMo-V2.6-Flash pricing and MiMo-V2.6-Pro.
Frequently asked questions
What is Xiaomi's MiMo-V2.6?+
A family of AI models Xiaomi launched on September 22, 2026, with weights published under the MIT license. It includes Pro, with 1.02 trillion parameters (42 billion active), and Flash, with 309 billion (15 billion active). Both understand text, images, video and audio with a one-million-token context.
How much does MiMo-V2.6 cost through an API?+
On OpenRouter, Flash costs $0.14 per million input tokens and $0.28 for output, while Pro costs $0.435 for input and $0.87 for output. An assistant handling a thousand average conversations a day could cost under SAR 70 a month on Flash, depending on conversation length.
Can I run it on my company's own servers?+
Yes. The MIT license allows running, modifying and commercial use. However, the reference setups on the Pro model card split it across 8 GPUs, while Flash needs a split across 4, which makes Flash the realistic self-hosting choice for most businesses.
Is it better than ChatGPT and Claude?+
No. It comes close to previous-generation models such as Claude Opus 5 on coding and automation tasks and falls behind on the hardest ones, and Xiaomi itself acknowledges it trails Claude Fable 5.1 and GPT-6 Astra. Its real advantages are price, the open license, and understanding voice and images in a single model.
Follow Origami in Google
Pin Origami as a preferred source and our articles will surface first for you in Google Search and Top Stories.

Related articles
- Artificial IntelligenceAn Open Agent Built for Days of Work, Not Minutes: Atria Dawn and Its Real Running BillShanghai AI Lab released Atria Dawn under MIT: 744 billion parameters aimed at long multi-step tasks. What it is actually good for, and what running it in-house costs.
- Artificial IntelligenceThe Model Race Taps the Brakes: Outside Evaluators Get a Badge and a Desk Inside AnthropicOn September 12, 2026 Anthropic's CEO called for slowing AI capability gains and committed to letting independent evaluators inside. Here is what actually changes for your tech plan.
- Artificial IntelligenceSWE-2 Matches Frontier Coding at a Third of the Cost, Then Fails One TestCognition released SWE-2 on September 10, 2026: 50.0% on FrontierCode versus Fable 5.1's 50.9% at 64% lower cost, yet 28 points behind on Terminal-Bench 4.
- Artificial IntelligenceOpenAI's Two New Image Models Edit One Part of Your Product Photo Without a ReshootOpenAI shipped GPT Image 2.5 in two versions on September 8, 2026: Sunburst for editing precision, Flare for speed. What changed for your store, what it costs, where it helps.
- Artificial IntelligenceDeepSeek V4.1 Flash: 77% Cheaper, and Peak Hours Hit Your MorningDeepSeek ships V4.1 Flash on 10 September and routes V4 Pro requests to it at the cheaper rate: 77% off input, 70% off output. Its peak hours sit inside your working morning. The numbers, and the largest saving nobody notices.
- Artificial IntelligenceFrom Months to Hours: MHS Connects Your Factory and Lab Devices to One AI AgentAnthropic opened a research preview of MHS, a standard that lets one AI agent operate lab and factory instruments together, cutting integration from weeks to hours.
Have a project in mind?
We build custom systems, apps and websites for your business. Tell us your idea and we will give you a straight answer on it.
