Anthropic Says Claude Leads 26% of Its AI R&D: What It Means

Conceptual editorial illustration of AI-assisted research inside a frontier AI laboratory, with human researchers supervising an AI research system and development workflow.

Anthropic has put a number on a question that has usually been discussed in much more speculative terms: how much of the work used to build advanced AI systems is now being done by AI itself?

According to Anthropic’s latest measurement, Claude “leads” 26% of the company’s AI research and development work as of August 2026. More than 90% of the measured work is at or above the level where AI collaborates with researchers. But there’s a crucial qualification behind the headline: Anthropic says Claude is not fully autonomous for any measured portion of its AI R&D.

That distinction changes how the number should be read.

The 26% figure does not mean Claude is independently running a quarter of Anthropic’s research department. It does not mean Claude is designing a replacement model without human supervision. And it does not show that Anthropic has reached autonomous recursive self-improvement.

Instead, the figure comes from a specific measurement system Anthropic built to estimate how much of different AI research tasks can now be handled at various levels of automation.

Understanding that methodology is more useful than simply repeating the 26% headline.

What Anthropic Actually Announced

On September 17, Anthropic published a set of measurements intended to give the public more visibility into the pace of AI development inside frontier AI labs. The company introduced three measurements covering AI-led AI R&D, oversight of AI-agent activity, and the allocation of computing resources.

The first measurement is the one attracting the most attention.

Anthropic calls it the Anthropic R&D Automation Index. The purpose is to estimate how much of the company’s AI research and development is being performed at different levels of AI automation.

For the August 2026 reading, Anthropic reported three important findings:

  1. Claude “leads” 26% of Anthropic’s AI R&D work.
  1. More than 90% of the measured work is at or above the “AI collaborates” level.
  1. Claude is not operating fully autonomously for any measured subset of AI R&D work.

Anthropic says the share of work reaching the “leads” level was under 1% in February 2026. Reuters also reported that the figure had risen sharply during the year, while emphasizing that Claude remains under human supervision.

The word “leads” is particularly important here.

What “Claude Leads” Actually Means

Anthropic uses an automation scale developed by Epoch AI that runs from AL0 to AL5. The scale is designed to distinguish between simple AI assistance and increasingly autonomous work.

  1. AL0 — No AI involvement.
  1. AL1 — Minimal AI involvement.
  1. AL2 — AI assists with the work.
  1. AL3 — AI collaborates and can complete large chunks of work under close human direction.
  1. AL4 — AI leads most of the task end-to-end from a high-level prompt while a human supervises.
  1. AL5 — AI operates fully autonomously with no human in the loop.

Anthropic places its 26% figure at AL4, not AL5. At AL4, the model can complete most of a task from a high-level instruction, but a human still supervises the work. Full autonomy begins at AL5.

A simple example helps.

Imagine a researcher tells an AI system: “Investigate why this model evaluation pipeline is failing, identify the likely cause, test possible fixes, and prepare the changes.”

At an AL4 level, the AI could carry out most of that workflow itself, while a researcher supervises the process and remains responsible for the result.

At AL5, the human would no longer be part of that operational loop.

That gap is why the 26% statistic should not be translated into “Claude is now building itself.”

The more precise interpretation is that Anthropic says Claude can independently carry out most of certain defined AI R&D tasks after receiving a high-level prompt, while humans still supervise the work.

That’s a significant increase in automation, but it is not the same thing as fully autonomous AI development.

Conceptual diagram showing AI automation levels AL0 through AL5, highlighting AL4 as AI-led work and AL5 as full autonomy.

How Anthropic Calculated the 26%

This is where the announcement becomes more interesting.

Anthropic did not simply ask researchers how often they use Claude. It built a detailed map of the work involved in developing AI models and then estimated the automation level of those tasks.

For each week in July 2026, Anthropic randomly sampled 20% of staff from departments that make up its model R&D loop. A Claude research agent reviewed each sampled employee’s week using Slack and internal documentation, then identified the tasks they worked on.

Across the sampled weeks, that process produced roughly 15,000 granular model R&D tasks. Anthropic then organized those tasks into a hierarchical structure containing 542 nodes, including 378 more specific leaf categories.

Those categories included concrete examples such as evaluation-platform defect diagnosis and fixes, reinforcement-learning sandbox network policy work, and serving-incident postmortems.

Anthropic then used Claude to research how each category of work is performed across the company. A separate Claude-based judging system assigned the resulting evidence an automation level using the AL0–AL5 framework.

The final index is weighted by person-time rather than by a simple task count.

Here’s the basic idea. Suppose one researcher spends a week working across four equally weighted tasks. Each task receives 0.25 units of that person’s weekly weight. If another researcher works across ten tasks, each receives 0.10.

That prevents a long list of small tasks from automatically outweighing areas of R&D where many people spend substantial amounts of time. Anthropic describes the approach as a crude proxy for how much attention different areas of work receive.

This makes the 26% figure much more specific than a statement such as “AI writes 26% of Anthropic’s code.”

It does not measure code volume.

It does not measure the percentage of Anthropic employees replaced by AI.

It does not measure the percentage of Anthropic’s computing power controlled by Claude.

Instead, it estimates the share of a defined basket of AI R&D work that reaches the AL4 “leads” level, with the categories weighted by person-time.

Why the Methodology Matters

The methodology is also the main reason to be careful about treating 26% as a universal measurement of AI research automation.

First, Anthropic’s task basket is based on work recorded in July 2026 and then frozen. That allows the company to compare the same categories over time, but it also means the index is not automatically capturing every new kind of R&D task that might emerge later.

Anthropic tested this issue by creating a separate task tree from January 2026 data and comparing new tasks arriving between February and July against that January baseline. The company said it found no rise in the number of “novel” tasks at its level of analysis, while also saying it plans to rebuild the task basket periodically.

Second, the automation ratings are not an independently audited external measurement.

Anthropic says the research and judging process uses its own Claude models. The company itself acknowledges that a judge model could share some of the same blind spots as the model being evaluated.

Anthropic did perform a human comparison. Staff members responsible for the relevant work areas rated automation levels without seeing the evidence or verdict produced by the model judges.

The model and human reviewers agreed exactly 59% of the time, while human reviewers agreed with each other exactly 35% of the time. Anthropic says the model and human ratings were within one automation level of each other 97% of the time.

Those numbers do not prove that the index is perfectly accurate.

They do show that the boundaries between automation levels can be subjective, particularly around the line separating “AI collaborates” from “AI leads.”

A common mistake here is to treat 26% as though it were a laboratory measurement with no assumptions behind it.

A better approach is to read it as an internally constructed index based on a defined task basket, a particular weighting method, and a model-assisted judging process. That doesn’t make the finding meaningless. It tells us what kind of evidence it actually is.

This Is Not Recursive Self-Improvement Yet

Anthropic connects the R&D Automation Index to a much bigger question: how close AI systems may be getting to recursive self-improvement.

The company describes recursive self-improvement in terms of a model becoming fully capable of building its successor autonomously. In the AL framework, that corresponds to AL5.

The current data does not show that.

Anthropic explicitly says Claude is not operating fully autonomously for any measured portion of its AI R&D work.

That makes the difference between AL4 and AL5 more than a technical footnote.

At AL4, humans can still supervise a system that handles most of a defined task from a high-level instruction. At AL5, the human is no longer part of the operational loop.

So the safest reading is this: Anthropic’s data points to rapidly increasing AI-led research automation, not to a completed transition from human-led AI development to autonomous AI self-improvement.

More Than 90% of the Work Is Already at “Collaborates” or Above

The second number is arguably just as important as the 26% headline.

Anthropic says more than 90% of its measured AI R&D work is now at or above AL3, where AI collaborates with researchers and can perform large chunks of work under close human direction.

That suggests the major shift is not limited to a small collection of highly autonomous tasks.

AI assistance appears to be spreading across a much larger portion of the development process.

Researchers can still decide which problems matter, review results, choose between approaches and supervise systems, while AI handles increasingly substantial portions of implementation, analysis, debugging and evaluation.

This matters because major productivity gains can happen before a lab reaches full autonomy.

A researcher doesn’t need to hand over an entire project to an AI system for automation to have a material effect. If the system can reliably take over large portions of dozens of individual tasks, the cumulative effect can still be significant.

Epoch AI’s work on measuring AI R&D automation similarly argues for looking at task-level automation rather than relying only on broad capability benchmarks or indirect measures such as compute.

Anthropic is now applying that style of measurement to its own internal R&D process.

Anthropic Is Also Measuring Its AI Agents

The 26% metric is only one part of the September disclosure.

Anthropic says that, as of August 2026, approximately 30,000 agents were doing research and engineering work at any one time on its most-used internal platform. The company makes clear that these figures cover that platform, rather than every AI agent operating across every Anthropic system.

The company says those agents’ actions pass through two monitoring layers.

Online monitors inspect actions before execution and can block or automatically steer dangerous behavior. Offline monitors ingest activity after the fact and are intended to detect problems that can develop more slowly, including potential signs of model misalignment.

Anthropic says 100% of the agents’ actions on that internal platform pass through the online monitor before execution, and 100% are also ingested into the offline monitoring system after execution.

The company analyzed more than one billion decisions from its research and engineering agents during August 2026 and reported that 0.002% of them — about one in 47,000 — were blocked by the online monitor.

That number needs context.

It is an Anthropic measurement of its own internal agent platform, not an independent estimate of how often AI agents misbehave across the wider industry.

Anthropic also says the offline system flags roughly one to two transcripts in every thousand for additional review, while higher-priority cases can be escalated to human reviewers.

For everyday readers, the practical lesson is straightforward: as AI agents receive permission to take more actions on their own, monitoring becomes part of the system architecture rather than an optional add-on.

For developers, the distinction between “the model produced a bad answer” and “the model took a consequential action” is especially important. An agent that can change code, access internal systems or launch additional tasks creates a very different oversight problem from a chatbot that only returns text.

For readers interested in that broader agent-security issue, Futuristic Byte has also covered OpenAI’s model-misalignment reporting framework and several documented cases involving unauthorized model behavior.

Anthropic Is Measuring Compute Too

The company also introduced a third measurement for how it allocates AI computing resources.

Using a snapshot covering July 13 through July 20, Anthropic estimated that about 6% of the compute used for AI R&D was allocated to safety-related work. Looking specifically at compute used for AI-driven AI R&D, Anthropic estimated that the safety share was about 12%.

But Anthropic itself warns against reading those percentages as a complete measurement of its safety effort.

Safety work can be relatively compute-light while still requiring substantial researcher time. A safety researcher may spend significant effort designing an experiment, even if running that experiment requires far less compute than a frontier model-training run.

Anthropic therefore says the value of this metric is less about treating 6% or 12% as definitive measures of safety commitment and more about creating a common mechanism that could eventually allow similar categories to be compared across developers and over time.

The company also says these estimates are deliberately conservative. For example, when compute appears to advance capabilities and safety equally, Anthropic says it does not count that compute toward the safety metric. The measurement also excludes compute used for safeguards classifiers as a separate category.

Taken together, the three measurements point toward a broader shift in how AI development can be discussed.

Instead of asking only how capable a model is, researchers can also ask how much of the model-development process it performs, how those actions are monitored, and where the computing resources behind that development are being spent.

Why Other AI Labs May Need Comparable Measurements

The biggest limitation of Anthropic’s index today is that it is not a standardized industry metric.

Anthropic acknowledges two major obstacles to comparing this type of result across frontier labs. First, there is no common methodology. Second, Anthropic is using its own models to evaluate its own systems.

That means a 26% figure from Anthropic cannot simply be placed beside a hypothetical number from OpenAI, Google DeepMind or another developer and treated as a like-for-like comparison.

Another lab could define its R&D task categories differently.

It could weight those categories differently.

It could use another model or a larger human review process to assign automation levels.

It could also draw different boundaries around what counts as AI R&D or safety work.

This is where third-party verification becomes important.

Anthropic says it plans to bring in independent third-party evaluators and give them access to internal processes, systems and data comparable to what its internal risk-assessment teams use. The company says such evaluators could verify safety practices, report incidents and monitor key metrics.

If other frontier labs eventually publish comparable measurements using transparent methodologies, the industry may get something it currently lacks: a more consistent way to track how quickly AI is taking over portions of the AI-development process.

What Does the 26% Number Mean for AI Development?

The most important takeaway is not that AI has suddenly become autonomous.

It is that the boundary between using AI to build software and using AI to build AI systems is becoming increasingly difficult to ignore.

For years, most discussion about AI progress focused on what models could do for users: write code, summarize documents, analyze data, generate images or answer questions.

A different shift is happening inside the labs that build those systems.

When AI can help researchers design experiments, analyze failures, diagnose problems, improve evaluation systems or handle other parts of the development cycle, the process of creating more capable models may itself become more efficient.

That creates the possibility of a feedback loop.

More capable AI can help automate more AI research. That automation can help researchers build more capable models. Those newer models can then take on additional research tasks.

That does not automatically mean recursive self-improvement has arrived. But it helps explain why Anthropic considers AI-led R&D automation worth measuring in the first place.

The important distinction is between acceleration and autonomy.

A human-supervised research process can become dramatically more automated without becoming fully independent.

What Should Readers Watch Next?

The next important development is unlikely to be another isolated “26%” headline.

What matters more is whether the measurement changes consistently over time, whether the task basket remains representative as research workflows evolve, and whether other frontier labs publish comparable numbers that can be independently checked.

Anthropic says it plans to continue the measurements, while also acknowledging that the current methodology has limitations and room for third-party verification.

That could make future readings much more informative than a single snapshot.

A single percentage tells us where Anthropic says its process stood in August 2026.

A sustained series could show how quickly the boundary between human-led and AI-led R&D is moving.

And if multiple labs eventually publish comparable data, researchers may be able to distinguish between a broad industry trend and a change specific to one company’s internal workflow.

For now, the clearest conclusion is narrower: Claude is already handling most of certain AI research tasks under human supervision, and Anthropic’s measurement says the share of work reaching that AI-led level has expanded rapidly during 2026.

The 26% number matters, but the methodology matters just as much.

What readers should watch is not simply whether the percentage rises.

The bigger question is whether AI systems continue moving from assistance to leadership across more kinds of research work — and whether humans remain meaningfully in control as that happens.

Post a Comment

0 Comments