Back to Blog
Artificial Intelligence

428 Billion Parameters That Speak Arabic — and You Still Cannot Run Them

Origami TeamEditorial Team
6 min read
428 Billion Parameters That Speak Arabic — and You Still Cannot Run Them
Like what we publish? Pin Origami as a preferred source on Google.Add as a preferred source on Google

428 Billion Parameters That Speak Arabic — and You Still Cannot Run Them

On September 3, 2026, at the LEAP conference in Riyadh, HUMAIN unveiled humain-m3: a frontier Arabic-focused language model with 428 billion parameters in a mixture-of-experts architecture, of which 23 billion activate per token. The company says it was trained on more than one trillion tokens of Arabic-native content.

This deserves attention, because weak Arabic in frontier models has for years been the ready-made excuse behind every stalled AI project in the Saudi market. But before you build a decision on it, three details inside the announcement itself change what you can actually do with it today.

What was actually announced

Per the official release and the model page:

  • 428 billion parameters in a mixture-of-experts architecture built on the MiniMax-M3 lineage, with 23 billion active per token.
  • Natively multimodal — trained jointly on text, image and video from inception, not bolted on afterwards.
  • Built for agents — tool use and screen operation for long-horizon tasks in Arabic and English.
  • Available now as a research and evaluation preview only, through the HUMAIN Node platform — not as a production service.

One point worth recording: the company itself frames the motivation as Arabic being spoken by hundreds of millions of people while remaining significantly underrepresented at the AI frontier. That diagnosis is correct, and it is the reason this news exists at all.

The numbers, and who published them

HUMAIN reports a 89.37 percent average across seven public Arabic benchmarks, broken down as: AraTrust for truthfulness at 97.53, MadinahQA for language proficiency at 95.44, ALRAGE for retrieval-augmented generation at 94.63, Translated MMLU at 93.20, ArabicMMLU for native knowledge at 90.70, AlGhafa for core understanding at 86.45, and Arabic EXAMS for academic reasoning at 67.67. It reports beating competing models that scored 87.34 and 87.30 on the same set.

Read the gap commercially. The average leads by under two percentage points, and the hardest academic benchmark drops to 67.67 — far below the rest. More importantly, these are results published by the party that owns the model, not by an independent third party. That does not make them wrong. It makes them a claim to verify, not a settled result.

The detail most coverage skipped

Calling the model open-weights today is premature. The official release states that the weights are expected to be published under the MiniMax Community License once safety training and alignment are complete, currently targeted for next month. What exists today is a promised timeline, not a file you download and run.

Second detail: the model was developed by MiniMax, a Chinese lab, commissioned by HUMAIN. That is not a technical objection, but it means the word sovereign here describes ownership, operation and hosting — not the full build chain. If your decision depends on where the technology originates, that is information you need.

The Origami view

We read this announcement as a market signal, not a finished product. The signal is that Arabic has finally entered the frontier race as a first-class target rather than an added language, and the practical effect on you is immediate: it is no longer acceptable for a vendor to excuse weak Arabic output by claiming the technology itself is weak in Arabic. Raise the bar on the proposals that reach you.

The decision itself, though, does not change today. We still build client systems so the model layer stays replaceable: one interface in front of the AI provider, and evaluation on your own data before any switch. A business built that way can trial humain-m3 within days of the weights landing. A business that wired its operating logic to a single provider will need a full project to do the same thing, and will pay the difference twice — once now, and once at the next model.

What to do this month

  • Do not stall a live project waiting for it. The model is in research preview; waiting costs you months against an unconfirmed gain.
  • Build an evaluation set from your own data. Fifty to a hundred real cases from your correspondence, contracts and customer questions, each with the correct answer. This is worth more to you than any public benchmark, because it measures your dialect and your terminology.
  • Isolate the model layer in your architecture if it is not already isolated, so changing providers is a setting rather than a rebuild.
  • Track the weights release and its license. The commercial-use terms are what decide whether running it inside your own infrastructure is even a legal option for you.
  • Ask for proof, not a claim. When a vendor pitches a model on its Arabic superiority, ask them to run it on your cases in a single session in front of you.

Conclusion

humain-m3 is good news for Arabic, and its numbers deserve verification rather than acceptance. Today it is a research preview with weights promised next month. Preparing for it properly does not start with the model — it starts with your architecture: a replaceable layer, and an evaluation set in your language and your data. Whoever has both benefits from whatever model lands next. Whoever has neither will keep reading each announcement as news that passes by, not an opportunity to take.

Sources

#HUMAIN#Language Models#Arabic AI#LEAP 2026#Vendor Selection

Frequently asked questions

Can I use humain-m3 in my company today?+

Not as a production system. The model is currently available as a research and evaluation preview through the HUMAIN Node platform, which suits testing and measurement rather than running a real customer service on it.

Is the model actually open-weights?+

Not yet. The official release states the weights are expected to be published under the MiniMax Community License once safety training and alignment are complete, currently targeted for next month. Until then you cannot download it and run it on your own servers.

Do the benchmark wins mean it is the best model for my case?+

Not necessarily. The scores were published by the party that owns the model, not an independent third party, and the average leads by under two percentage points. Public benchmarks measure Arabic in general; what matters to you is performance on your terminology, your customers dialect and your documents. The only measurement that counts is running it on fifty to a hundred cases from your own data.

What is the best practical step right now?+

Isolate the model layer in your system so switching providers is a setting rather than a rebuild, and prepare an evaluation set from your real cases. With both in place you can trial any new model in days instead of running a full project.

Follow Origami in Google

Pin Origami as a preferred source and our articles will surface first for you in Google Search and Top Stories.

Add as a preferred source on Google

Related articles

Weekly newsletter

The latest articles that matter to business owners, once a week. Just your email.

Have a project in mind?

We build custom systems, apps and websites for your business. Tell us your idea and we will give you a straight answer on it.

One session. Twenty minutes. No commitments.