Back to Blog
Artificial Intelligence

Reflection's Beam: 501 Billion Parameters, 23 Billion Doing the Work, Weights Not Out Yet

Origami TeamEditorial Team
7 min read
Reflection's Beam: 501 Billion Parameters, 23 Billion Doing the Work, Weights Not Out Yet
Like what we publish? Pin Origami as a preferred source on Google.Add as a preferred source on Google

Reflection Announces Beam: A 501-Billion-Parameter Open-Weight Model That Runs Only 23 Billion at a Time

The short answer: On October 5, 2026, the US company Reflection AI announced Beam, its first open-weight model. It has 501 billion parameters in total, but it is built as a mixture-of-experts model, so only 23 billion parameters are active for each token, and it handles an effective context of up to one million tokens. The weights will be released under the Apache 2.0 license, which allows commercial use, "later this month", together with a technical report, a model card and developer tooling. As of today you cannot download it; the only option is to sign up for the early access program.

This matters to any business considering running AI on its own servers instead of sending its data to a cloud service outside the Kingdom. But the published numbers deserve a calm reading before you build a plan on them.

What Reflection Actually Announced

  • Size: 501 billion parameters in total, 23 billion active per token, using a sparse mixture-of-experts architecture.
  • Training: 23.8 trillion tokens of web data and licensed datasets, on 6,144 NVIDIA GB300 GPUs in under four weeks.
  • Reinforcement learning: four more weeks on about 10,500 GB300 GPUs, with more than 100 million training rollouts.
  • Target use: coding, reasoning and agentic work, where the model carries out a sequence of steps on its own.
  • Reasoning control: a setting that decides how much the model "thinks" before answering, so you can choose between a short fast reply and a longer analysis.
  • Efficiency: the company says Beam achieves scores comparable to GLM-5.2 on advanced reasoning benchmarks, slightly below it in its own table, while using an estimated three to four times less inference compute. That figure is calculated from active parameters and generated tokens, not measured serving cost.

The Numbers as Published, and Who Beats It

Reflection published a comparison table against seven models, including GLM 5.2, Qwen 3.8 Max and Kimi K3. These are the main results for those three, as stated in the announcement:

  • Terminal Bench v2.1 (command-line tasks): Beam 80.1, against 81.0 for GLM 5.2, 86.6 for Qwen 3.8 Max and 88.3 for Kimi K3.
  • MCP Atlas (tool use over the MCP protocol): Beam 78.7, against 77.8, 84.5 and 82.3.
  • GPQA Diamond (expert science questions): Beam 90.5, against 91.2, 92.6 and 93.5.
  • DeepSWE v1.1 (fixing real software issues): Beam 44.4, against 44.0, 51.0 and 68.0.
  • AIME 2026 (mathematics): Beam 97.8 against 99.2 for GLM 5.2.

The honest reading of this table: Beam is not the strongest open model. Kimi K3 is ahead of it on 13 of the 14 benchmarks where both have scores (the one exception is tau3 banking, 38.0 for Beam against 37.1), and by a wide margin on code repair (68.0 against 44.4 on DeepSWE). What Reflection is selling is efficiency: performance close to GLM-5.2 at a lower running cost. Beam's own scores are reported by Reflection itself, while the company says it took the competitor scores from Artificial Analysis and DataCurve. Independent groups cannot re-test Beam until the weights and the technical report are published.

23 Billion Active Does Not Mean a Small Server

This is the most misunderstood point about mixture-of-experts models. The active parameter count sets the compute per token, meaning the speed and cost of each answer. But memory still has to hold the whole model. By simple arithmetic, 501 billion parameters need roughly 500 GB of GPU memory at 8 bits each, or about 1 TB at 16 bits, for the weights alone, before the context cache and concurrent users. Reflection has not said what precision it will release the weights in. You are talking about a server with several high-end GPUs, not a desktop machine. Reflection has not yet published official serving requirements or compressed versions, so wait for them before you estimate a budget.

What This Means for Saudi Businesses

The real value of any open-weight model is where it runs. When the model runs on your servers or in a data center inside the Kingdom, customer data, contracts and invoices stay under your control. That makes it easier to comply with the Personal Data Protection Law supervised by the Saudi Data and AI Authority (SDAIA), especially its rules on transferring data outside the Kingdom.

  • Internal coding agents: a tech team that wants a coding assistant without sending its source code to an outside service.
  • Process automation: an agent that reads requests and calls internal system tools, where the MCP Atlas score matters more than the math benchmarks.
  • Long-document analysis: a one-million-token context lets you read whole contracts or logs in one pass, provided you verify the quality of understanding on your own files.

One important caveat: the announcement does not name the supported languages, and it does not mention Arabic. The only multilingual result is SWE-Bench Multilingual at 78.0, and that refers to programming languages, not human ones. Do not assume Arabic quality before you test it.

How to Approach It, Step by Step

  • Do not change your plan today: the model is not available to download, and the numbers have not been independently reviewed.
  • Prepare a test set from your own work: 50 to 100 real tasks, such as Arabic customer messages, requests and documents, with the correct answer for each.
  • Compare when the weights land: Beam against the open model you are already considering, such as Kimi K3, on accuracy, latency and server cost.
  • Calculate the full cost: server, power, operations and maintenance, against the cost of a cloud API at the same request volume.
  • Build a switching layer: have your system call the model through a single interface, so you can move between models without rebuilding the application.

This is how we work at Origami when we build AI systems for our clients: we test the model on the business's own data, and we design the system so that the model is a replaceable part, so each new announcement does not turn into a rebuild project. If your data is sensitive, read our guide to the Personal Data Protection Law before choosing where to run it.

Sources

#Reflection Beam#Open-Weight Models#Self-Hosted AI#Mixture of Experts

Frequently asked questions

Can I download and run Beam now?+

Not yet. Reflection AI announced Beam on October 5, 2026, and said the weights and technical report will be released later in October under the Apache 2.0 license. For now, the only option is to sign up for the early access program on the company's platform.

Can my business use Beam commercially?+

Yes, according to the announcement: the weights will be released under Apache 2.0, which allows commercial use and modification. Check the license text and the model card when they are published to confirm there are no extra terms.

What kind of server does a 501-billion-parameter model need?+

Even though only 23 billion parameters are active per token, memory must hold the whole model. As a rough estimate, the weights alone need about 500 GB of GPU memory at 8 bits, or about 1 TB at 16 bits, which means a server with several high-end GPUs. Reflection has not yet published official requirements.

Does Beam support Arabic?+

The official announcement does not list supported languages or mention Arabic. Test it on a sample of your own Arabic customer messages and documents before relying on it in any customer-facing system.

Follow Origami in Google

Pin Origami as a preferred source and our articles will surface first for you in Google Search and Top Stories.

Add as a preferred source on Google

Related articles

Have a project in mind?

We build custom systems, apps and websites for your business. Tell us your idea and we will give you a straight answer on it.

One session. Twenty minutes. No commitments.