Back to Blog
AI for Business

Claude Leads 26% of Anthropic's R&D: The Scale That Measures AI's Real Share of Work

Origami TeamAI for Business
8 min read
Claude Leads 26% of Anthropic's R&D: The Scale That Measures AI's Real Share of Work
Like what we publish? Pin Origami as a preferred source on Google.Add as a preferred source on Google

Claude Leads 26% of Anthropic's R&D: The Scale That Measures AI's Real Share of Work

The direct answer: on 17 September 2026 Anthropic published its R&D Automation Index, the first detailed public measurement of how much of the work inside a frontier model lab is actually performed by AI. The headline result is that Claude "leads" 26% of the company's model research and development work as of August 2026, up from under 1% in February of the same year, with more than 90% of that work happening at the "collaborates" level or above, and no measured category reaching full autonomy. What matters most for your business is not the number. It is the method behind it, because you can reproduce it on your own operation with tools you already have.

What Anthropic actually published

Most conversations about "how much of our work is automated" run on impressions. A manager says half the team's work is now automatic; a developer on the same team says the tool only writes boilerplate. What Anthropic did was turn that question into a repeatable measurement with a fixed methodology instead of an impression.

The idea is simple at its core: catalogue every kind of work actually being done inside the model R&D departments, rate each kind on a scale that captures how involved the model is, then aggregate those ratings with a weight that reflects how much each kind matters. The company describes the index as a prototype rather than a final measurement, and frames it as part of a wider effort to narrow the gap between what frontier labs know and what the public knows about how fast this technology is moving.

A six-step scale: AL0 to AL5

The scale comes from Epoch AI and is called the Automation Level, running from zero to five:

  • AL0: no AI involvement at all.
  • AL1 and AL2: minimal involvement, then assistance — the human does the work while the model suggests or speeds up small steps.
  • AL3, "collaborates": the model completes large chunks of the task under close, continuous human direction.
  • AL4, "leads": the model finishes most of the task end to end from a high-level prompt, with human supervision.
  • AL5: full autonomy with no human in the loop — a level not reached in any measured category.

When Anthropic says "26%", it means the categories that reached AL4 specifically. The gap between AL3 and AL4 is the one that matters operationally: in the first, a human stays inside every step; in the second, the human's role shifts to setting the goal and reviewing the output.

The most important number in the report is not 26%

Buried in the methodology is a figure that deserves more attention than the headline. When the model's automation rating was compared against the rating given by the employee doing the work, the two matched exactly 59% of the time, and stayed within one level of each other 97% of the time. When one employee's rating was compared against a colleague's rating of the same work, they matched only 35% of the time.

Put plainly: humans disagree with each other about how automated their own work is more than they disagree with the model. That explains the meetings at your company that end without a conclusion. If two colleagues on the same team cannot agree on a description of their own work, then any AI investment decision built on collective impression is standing on soft ground. The fix is not more discussion. It is a written, fixed measure re-applied to the same thing at intervals.

How to build your own version of the index in a week

The methodology Anthropic described scales down to a mid-sized company. The steps, in short:

  • Catalogue the real work, not the job description: Anthropic sampled a random 20% of staff in the relevant departments during July 2026 and reviewed their actual work from internal conversations and documentation, producing roughly 15,000 granular tasks. For you, a far smaller sample of one real working week for one or two teams is enough.
  • Arrange the tasks into a tree: those 15,000 tasks were grouped into a tree of 542 nodes, 378 of them leaf categories. A mid-sized company reaches a useful picture with thirty to fifty leaf categories.
  • Weight by time, not by feel: they used the person-time dedicated to each task as a proxy for its importance. That stops flashy, low-impact tasks from dominating the picture.
  • Rate each category on the AL0–AL5 scale: and write a one-line reason for the rating so it can be reviewed later.
  • Freeze the basket and re-measure: the essential point is that the category list is frozen once built, so each re-measurement runs against the same thing rather than a fresh list. Without freezing, you are measuring changes in your list, not progress in automation.

The output of this exercise is not a number to boast about. It is a map telling you where AI actually stands in your operation: which categories have reached "collaborates" and deserve deeper automation, and which are still at zero despite consuming many hours — and those are your nearest opportunities.

What the "leads" level means in practice

When a category of work in your company reaches AL4, the shape of management changes, not just the size of the team. Direction becomes high-level and review moves to the final output, which means output quality rests on two things: clarity of the written instructions, and a review point before any irreversible effect — something sent to a customer, a change to real data, or a financial commitment.

In the projects we build for clients, this is the difference between automation you rely on and automation that worries you. A model can lead an entire task, but the system around it determines what happens when it is wrong: does the error appear in a reviewable log and stop at an approval gate, or does it go straight to the customer? Moving from AL3 to AL4 is a system design decision before it is a question of trust in the model.

Caveats Anthropic stated itself

Fairness requires noting the limits the company put on its own measurement: no independent external party has verified it, and it was produced largely using Claude itself, meaning the model being measured participated in the measurement — which can carry its errors into the judgement. Anthropic says it intends to bring in third-party evaluators later. Freezing the task basket at a July 2026 baseline also means categories of work that emerged afterwards do not enter the calculation.

So read the number for what it is: a strong directional signal from inside a frontier lab, not a settled fact about the labor market. A jump from under 1% to 26% in six months deserves attention regardless of the precision of the decimals.

Where to start

Start with one team and one week. Record what the team actually did, sort it into thirty categories, rate each on the scale, freeze the list, and re-measure in three months. You will end up with something most companies do not have today: a written baseline to compare against instead of argue about. And if you want to turn the categories that came back at zero into systems that run, that is exactly what we build at Origami.

Sources

  • Anthropic — measuring the pace of AI development inside frontier labs: anthropic.com
  • Epoch AI — the automation level scale used in the index: epoch.ai
#AI#Business Automation#Measurement#Anthropic

Frequently asked questions

What does it mean that Claude "leads" 26% of the work? Did they replace their researchers?+

No. "Leads" in the index is the fourth step on a six-step scale: the model finishes most of a task end to end from a high-level prompt, but under human supervision. Anthropic states explicitly that no measured category of work reached full autonomy with no human in the loop.

My company is mid-sized. Can I build a similar index?+

Yes, and it needs no special tooling. Pick one team, record its actual work over one week, sort it into thirty to fifty categories, weight each by the time spent on it, then rate them on a zero-to-five scale. The only condition that really matters is freezing the category list and re-measuring against that same list every three months.

What is the practical benefit of measuring AI's share of our work?+

The map reveals two useful groups: categories that have reached a good level of collaboration and deserve deeper automation, and categories still at zero involvement despite consuming many hours — usually the fastest return available to you. Without a written measure you invest on impression, and the report's own data shows human impressions diverge sharply.

Is the number trustworthy?+

Read it as a directional signal, not a final fact. Anthropic states plainly that the index has not been verified by an independent external party and was produced largely using Claude itself, and that it intends to involve third-party evaluators later. Even so, moving from under 1% in February 2026 to 26% in August is a clear trend worth noting.

Follow Origami in Google

Pin Origami as a preferred source and our articles will surface first for you in Google Search and Top Stories.

Add as a preferred source on Google

Related articles

Have a project in mind?

We build custom systems, apps and websites for your business. Tell us your idea and we will give you a straight answer on it.

One session. Twenty minutes. No commitments.