Anthropic Debuts R&D Index: Claude Handled 26% of Internal AI R&D in August

Anthropic released a prototype R&D automation index showing that by August 2026, Claude was handling 26% of the company's AI research and engineering work at the AL4 level, up from under 1% in February. The framework, built on a frozen task tree of 542 nodes, also details oversight data and compute allocation for safety work.

On September 17, 2026, Anthropic released the first prototype version of its R&D automation index, showing that as of August 2026, Claude was responsible for 26% of the company's AI research and engineering work at the AL4 level, a sharp rise from under 1% in February.

The Facts

The index uses an AL0 to AL5 scale: AL3 means AI completes most of a task under close human guidance, while AL4 means AI drives the vast majority of a task from high-level instructions with human oversight. Anthropic's report shows that in August, tasks at AL3 and above accounted for more than 90% of the total, and no task reached AL5, the fully autonomous level. The task catalogue came from a July 2026 sample of work records from 20% of employees across departments, generating about 15,000 fine-grained tasks that were later organized into a task tree of 542 nodes. Each node was assessed for its automation level by a Claude agent and scored by an independent Claude judge, with weights allocated by person-hours.

On oversight data: in August, the internal platform ran about 30,000 agents simultaneously, with online monitoring covering all actions; of the more than 1 billion agent decisions analyzed that month, 0.002% were blocked. Offline monitoring flagged roughly 100,000 records per week, with about 50 high-priority flags ultimately escalated to human review. On compute allocation: during the week of July 13 to 20, about 6% of AI R&D-related compute went to safety work, and within AI-driven AI R&D compute, safety accounted for about 12%.

How It Works

Once the task tree is frozen, each weekly measurement targets the same basket of work, avoiding interference from new task types. Manual comparison between the judges and domain leads showed 59% exact agreement between model and human, compared with 35% agreement among humans, with 97% agreement within one level of difference. This bottom-up catalogue allows automation measurements to track progress on specific tasks such as "evaluating platform defect diagnosis and repair."

Persistent agent identity and a shared open messaging system support continuous monitoring across model upgrades. Online monitoring targets fast, irreversible harms, while offline monitoring targets slowly developing behaviors; together they form layered protection. For compute classification, a Claude classifier sampled nearly 10,000 training and evaluation runs, conservatively counting tokens that advance both capability and safety as AI R&D.

Industry Impact

This quantitative framework provides internal reference for understanding the steep recent curve in model capabilities. Multiple media outlets issued warnings on the day of release about a self-accelerating flywheel, focusing on the rapid rise from February to August. The limitation of a frozen task basket means new types of work are not yet captured; Anthropic plans to rebuild the basket periodically and re-release new versions.

For other frontier labs, if such internal measurement becomes routine, it could push the industry to establish comparable automation indices. The concern among safety researchers is whether a rising share of AL4-level work will accelerate a further automation loop.

Strategic Assessment

[Analysis] Based on current measurements, Anthropic is completing a large amount of internal R&D work through its own models, which in the short term may reduce reliance on external human labor, but at the same time requires a stricter monitoring system to control risk. Over the long run, if similar indices are adopted by multiple labs, industry competition may shift toward who can safely expand the automation share, rather than simply pursuing single-model performance. The conservative estimate of safety's share of compute reflects the company's cautious approach in balancing progress and protection; if this share continues to be published, it could become an important external indicator for assessing a lab's safety investment.