Back to Blog
Artificial Intelligence

The Model Race Taps the Brakes: Outside Evaluators Get a Badge and a Desk Inside Anthropic

Origami TeamEditorial Team
8 min read
The Model Race Taps the Brakes: Outside Evaluators Get a Badge and a Desk Inside Anthropic
Like what we publish? Pin Origami as a preferred source on Google.Add as a preferred source on Google

Three Steps to Slow the AI Race: What Anthropic Actually Committed To

The direct answer first: on September 12, 2026, Anthropic CEO Dario Amodei published an essay titled "We Must Pace the Frontier," arguing plainly that the AI industry must slow the pace at which it improves model capabilities, and laying out a three-step plan of escalating difficulty. Anthropic unilaterally committed to the first step: placing independent outside evaluators inside its own offices with near-employee access — a desk, a badge, a company laptop, and permissions broadly comparable to internal risk-assessment teams — along with the right to publish what they find without Anthropic holding editorial control over it. Within hours, Sam Altman said OpenAI agrees and will match the same commitment, and Elon Musk posted that "Dario is right."

For a business owner in Saudi Arabia, this is not a philosophical story about the future of humanity. It is an early signal that the tools your operations are being built on are entering a slower, better-documented, more third-party-reviewed phase — and that what is happening inside these labs today will, within a year or two, become the expected shape of any organization running AI in front of its customers, including yours.

The Three Steps as the Essay Lays Them Out

  • Embedded evaluators inside the lab: independent assessment organizations such as METR get permanent, employee-level access to verify that stated safety commitments are actually applied and that incidents are reported as they should be. This is the step Anthropic committed to alone, without waiting for anyone else.
  • Coordination among labs in democratic countries: common safety standards and agreed limits on the rate of unchecked progress, with government backing to address the antitrust concerns such coordination would raise. Here Amodei floats a "checkpoint" idea: if a model reaches a given capability, it should not ship before specific certifications establish particular alignment properties — possibly alongside limits on inputs themselves, such as training compute or the internal use of models to build other models.
  • Global coordination: an attempt to reach understandings with non-democratic governments, across four escalating levels running from narrow prohibitions on specific applications up to full development pauses, with an explicit acknowledgement of how hard compliance would be to verify.

Amodei gave no slowdown percentage and no timetable, but he did say that buying "an extra year or two" before models reach critical capability levels would greatly reduce the risk. At the same time he stressed that progress will still seem fast, and that the point is to use the time gained wisely rather than to stop.

Why the Competitors Agreed So Fast

The speed is the real story. An industry built on "whoever gets there first wins" found its four biggest names lined up within a day. The background is that Demis Hassabis of Google DeepMind had proposed on July 14 an industry-funded, federally overseen standards body modeled on the financial-industry regulator FINRA, which would test frontier-class models for up to thirty days before release — voluntarily at first, then as a condition of US deployment.

And on September 14, just two days later, Microsoft published a draft code of conduct for its in-house models, which Mustafa Suleyman, who heads the company's AI arm, described as a kind of constitution. Its central rule is that a model must never resist correction or shutdown, must communicate in terms people can understand, and must treat any breach of the code as an outright failure rather than an edge case. The company opened a six-week public comment period, after which it plans to train its next generation of models on the finished code.

What Actually Changes Inside Your Business

Three concrete shifts, none of which requires you to wait for legislation:

  • Do not build a plan on a capability that has not shipped. A great many stalled AI projects rest implicitly on the sentence "the next version will solve this." If the lab founders themselves are talking about easing the pace, the sound assumption is that what you build on next year is roughly what you can see today. Design against capabilities that actually exist, and measure them yourself.
  • The independent-evaluator principle scales down to your size. The core of Anthropic's commitment is simple: the party that builds the system is not the party that certifies it safe. In your business that means whoever developed the AI assistant should not also be the one who approves running it in front of customers. Have a different person or team test it against real cases, with genuine authority to say no.
  • The no-resisting-shutdown rule translates into a real kill switch. If you have an AI agent that sends messages, edits orders, or issues invoices, ask one question: how many seconds does a non-technical employee need to stop it completely? If the answer involves "we call the developer," you do not have a kill switch — you have a phone number.

Four Decisions Worth One Quarter of Your Attention

  • Write on a single page everything your AI system can do without human approval, then strike out every irreversible action: payments, deletions, outbound messages to external customers, inventory changes.
  • Put an approval gate on exactly those actions, and let the rest run freely so you do not kill the benefit along with the risk.
  • Turn on an audit log that shows what the system did, when, and on what input — because the first question after any mistake will be "why did it do that?"
  • Decide in advance who stops the system and who gets notified, and test it at least once before you need it.

This direction lines up with where the Kingdom is already heading. The AI Ethics Principles and the Personal Data Protection Law issued by the Saudi Data and AI Authority ask for the same things in different language: clear human oversight, explainability, and defined accountability for automated decisions. What you do today to prepare for this will not go to waste.

And What Does Not Change

These commitments should not be read as larger than they are. They are entirely voluntary, with no binding verification mechanism and no announced timetable. Competition from open-weight models coming out of China is subject to none of it and will keep pushing prices down and capabilities up. And Amodei's essay is a proposal, not a regulation. The practical message is not that AI will stop improving — it is that the layer separating a powerful model from your business, namely governance, testing, and oversight, has officially moved from a luxury to part of the product.

At Origami we build business AI systems on exactly this logic from day one: approval gates on irreversible actions, a complete audit log, and a clean handoff to a human when the conversation leaves the system's scope. Not because a regulator asked for it, but because it is the difference between a tool you trust in front of your customer and one you watch nervously.

Sources

  • Full text of Dario Amodei's "We Must Pace the Frontier," September 12, 2026: darioamodei.com
  • TechCrunch coverage of the plan, Anthropic's commitment, and Sam Altman's response: techcrunch.com
  • CNBC coverage of Microsoft's draft AI code of conduct: cnbc.com
  • Saudi Data and AI Authority — AI Ethics Principles and the Personal Data Protection Law: sdaia.gov.sa
#AI Governance#Anthropic#Technology Risk#AI Adoption

Frequently asked questions

Does this mean AI models will stop getting better?+

No. Amodei himself said progress will still seem fast, and that the goal is to buy an extra year or two before models reach critical capability levels. What changes is the cadence of large jumps and the amount of review before release, not the direction. In practice it means planning around what is available today rather than a promised version.

Are these commitments legally binding?+

No. They are entirely voluntary so far, with no binding verification mechanism and no timetable. Anthropic unilaterally committed to embedding independent evaluators, and OpenAI said it will do the same. The second and third steps require industry and then government coordination, and neither has happened yet.

How does this affect an AI project I am planning for my business?+

In three ways: design against capabilities that exist and that you have measured rather than promised ones; make sure the party approving the system for customer-facing use is not the party that built it; and ensure a non-technical employee can stop it within seconds. All three raise the project's odds of success regardless of what the labs decide.

What does this have to do with Saudi regulations like the PDPL?+

The direction is the same even if the language differs. The AI Ethics Principles and the Personal Data Protection Law from the Saudi Data and AI Authority require clear human oversight, explainability, and defined accountability for automated decisions. Preparing for that today serves you both in compliance terms and operationally.

Follow Origami in Google

Pin Origami as a preferred source and our articles will surface first for you in Google Search and Top Stories.

Add as a preferred source on Google

Related articles

Have a project in mind?

We build custom systems, apps and websites for your business. Tell us your idea and we will give you a straight answer on it.

One session. Twenty minutes. No commitments.