AppVolt AI news, read from the source

Anthropic measured how much of its own AI research its AI now leads. It says 26%.

The same report says over 90% of that work is at or above "AI collaborates", and that about one in 47,000 agent decisions gets blocked.

September 19, 2026 · 2 min read

Two researchers working at computers in a laboratory.
Photo: Unsplash

Anthropic has published a set of measurements of how far AI has moved into its own research and engineering. The headline figure, measured as of August 2026 on data collected in July 2026, is that Claude "leads" 26% of the company's AI R&D work. The share of that work sitting at or above the level the report calls "AI collaborates" is above 90%.

The method is unusually explicit for this kind of claim. Researchers sampled 20% of staff weekly across departments, had Claude research each person's tasks through Slack and internal documentation, and built a hierarchical task tree of 542 nodes. Automation levels on a scale from AL0 to AL5 were then assigned to 378 leaf categories, weighted by how much person-time each one consumes.

A second measurement covers oversight. Anthropic says it analysed over a billion decisions from its research and engineering agents during August 2026, across roughly 30,000 agents on its most-used internal platform, and that 0.002% of them, about one in 47,000, were blocked.

A third covers compute. For the week of 13 to 20 July, the report puts about 6% of the compute that went to AI R&D toward safety, and about 12% of the compute that went to AI-driven AI R&D toward safety. The authors call these "deliberately conservative estimates".

The report also states plainly that "Claude is not operating fully autonomously for any measured subset of AI R&D work".

What this does not mean. It does not mean a quarter of the work happens without people. The report says Claude is not operating fully autonomously for any measured subset, and "leads" is one level on a six-point scale the authors defined themselves. The ratings come from a judge model: it agreed exactly with human raters 59% of the time, and came within one level 97% of the time. The authors also note the task basket is frozen, so newly emerging kinds of work may be missing from it.
Primary source Anthropic — Measurements for understanding the pace of AI development inside frontier labs
Open it and check the claim yourself. That is the point.

One email a week

A digest of the week's verified stories. No daily flood, no tracking pixels, unsubscribe in one click.

We send the digest and nothing else.