On September 17, 2026, Anthropic published a number that sounds almost recursive: Claude is now responsible for 'leading' 26% of the research and engineering work that goes into building the next version of Claude. Back in February 2026, that same figure was under 1%. In roughly seven months, the share of Anthropic's own AI R&D that Claude effectively runs — from a high-level human prompt through to a finished piece of work, while a person supervises — jumped from almost nothing to more than a quarter.
The number comes from something Anthropic calls the R&D Automation Index, the first of three new internal metrics it published under the Anthropic Institute, the public-benefit research arm the company launched back in March 2026. The other two track how closely Anthropic's monitoring systems watch autonomous Claude agents, and how much of the company's AI R&D compute goes specifically to safety work rather than capability work. Anthropic framed the release as an attempt to close a gap it says is otherwise invisible from the outside: how fast AI is actually starting to build AI, inside the handful of labs racing to build it.
Quick summary: Anthropic says Claude now 'leads' 26% of the research and engineering that produces future Claude models, up from under 1% in February 2026 — a jump made over about seven months. 'Leading' means Claude completes most of a task end-to-end from a high-level prompt while a human supervises the output; it is not full autonomy. Over 90% of Anthropic's R&D work now involves Claude at some 'collaborates' level or higher, but zero percent has reached full autonomy (AL5). Anthropic also disclosed that roughly 30,000 agents run concurrently inside the company on its main internal agent platform, and that its safety monitors blocked about 1 in every 47,000 agent actions after reviewing more than a billion agent decisions in August 2026. The disclosure landed five days after CEO Dario Amodei publicly called for the AI industry to slow down over safety concerns.
'Leading' isn't the same as 'building itself'
The headline number is easy to misread, so it's worth being precise about what Anthropic is actually measuring. The R&D Automation Index uses a five-level automation scale — AL0 through AL5 — originally developed by the research group Epoch AI. AL0 means no AI involvement at all. Further up the scale, AL3 is 'collaborates': Claude does large portions of a task, but only under close human direction, step by step. AL4 is 'leads': Claude takes a high-level prompt, carries the task through to completion mostly on its own, and a human reviews and supervises the result rather than directing each step. AL5 would be full autonomy — an AI system that notices a problem itself, decides what work is needed, builds and tests a solution, and ships it with no human in the loop at any point.
Anthropic's 26% figure is the share of R&D work that has reached AL4 ('leads'). Separately, the company says more than 90% of its R&D work now happens at AL3 ('collaborates') or above — meaning Claude is doing meaningful, hands-on work across almost all of Anthropic's research pipeline, even where it isn't yet trusted to run a task end-to-end unsupervised. Crucially, Anthropic reports that AL5 — true autonomy — sits at 0% for any measured category of work. Nothing is running without a human checking in. Sources: Measurements for understanding the pace of AI development inside frontier labs — Anthropic, Anthropic Publishes Three Self-Measured Indexes on How Much Claude Leads Its Own R&D — FourWeekMBA.
| Level | Name | What Claude actually does | Anthropic's August 2026 figure |
|---|---|---|---|
| AL0 | No AI involvement | A human does the entire task alone, with no AI assistance | Not separately reported |
| AL1–AL2 | AI-assisted | Claude helps with drafting, research or debugging, but a human still drives every step | Not separately reported |
| AL3 | Collaborates | Claude completes large chunks of a task, but only under close, step-by-step human direction | Combined with AL4+, this is over 90% of R&D work |
| AL4 | Leads | Claude takes a high-level prompt and carries the task through to completion largely on its own; a human supervises the result | 26%, up from under 1% in February 2026 |
| AL5 | Fully autonomous | Claude would identify the problem, plan the work, build, test and ship it with no human in the loop | 0% — not reached on any measured task yet |
The other two numbers: 30,000 agents and 1 in 47,000
Alongside the R&D Automation Index, Anthropic disclosed figures from its agent-oversight and safety-compute tracking. As of August 2026, roughly 30,000 AI agents were reportedly running at once on Anthropic's most-used internal agent platform, doing research and engineering work across the company. Every action those agents take reportedly passes through an online monitor before it executes — usually screened within seconds — plus a second, offline monitor that reviews everything afterward. Anthropic says that after analyzing more than a billion agent decisions logged in August 2026, its online monitor blocked about 1 in every 47,000 actions, or roughly 0.002% — actions the company says were flagged mainly for a narrow class of harm: fast, hard-to-reverse moves such as an agent trying to copy model weights out of its intended environment.
The third metric covers money rather than actions: Anthropic says about 6% of its total AI R&D compute is now devoted specifically to safety work rather than capability work, a share that reportedly rose to around 12% when narrowed to just the compute spent on AI-driven AI R&D, based on a snapshot the company took across July 13–20, 2026. Sources: Anthropic runs about 30,000 AI agents on itself and blocks one action in 47,000 — Mixed News, Is Anyone Watching Your AI Agents? Anthropic's Three Numbers — Digital Applied.
Why disclose this now — and why it's tangled up with a slowdown call
The timing here is not a coincidence, and Anthropic didn't pretend it was. Five days before the R&D Automation Index went live, on around September 12, 2026, CEO Dario Amodei published a lengthy essay — reportedly around 3,800 words — arguing that the AI industry as a whole needs to slow its pace of frontier development because safety verification is struggling to keep up with capability gains. The essay reportedly followed the departure of an Anthropic employee who had raised safety concerns. Amodei also said Anthropic would give independent evaluators permanent, ongoing access to its models to check whether the company is actually following its own safety commitments — not just a one-time audit. Sam Altman at OpenAI publicly agreed the industry needed to slow down and invest more in safety, and Elon Musk endorsed the call as well, writing that 'Dario is right.' Sources: Anthropic's Dario Amodei calls for slower pace of AI development — CNBC, Anthropic, OpenAI CEOs call for slowdown in AI development — Axios.
Against that backdrop, publishing hard numbers on how much AI is already doing inside Anthropic's own labs reads as a deliberate move, not a victory lap. The company's own stated rationale, according to reporting on its announcement, was that labs should 'do everything possible to minimize the gap between what frontier labs know and what the public knows.' In other words: if AI-driven AI R&D is accelerating quickly enough to justify a public call to slow down, Anthropic wants outsiders to be able to see the actual pace, not just take the warning on faith. The company also reportedly called on other AI developers to publish comparable metrics using a shared methodology, so the numbers could eventually be compared across labs rather than read in isolation.
Should this worry you?
It's a genuinely fast jump — under 1% to 26% of 'leads'-level work in about seven months is a steep curve, and if it continues at anything like that rate, the picture a year from now could look very different. But a few caveats are worth holding onto. First, every AL4 task still has a human reviewing the result; Anthropic's own AL5 (full autonomy) figure is flatly 0%, and the company was explicit that this isn't AI building AI unsupervised. Second, these are self-reported internal metrics from Anthropic about Anthropic, with no independent, external audit of the numbers themselves published alongside them — Amodei's promise of permanent evaluator access is a separate, forward-looking commitment, not a verification of this specific dataset. Third, no other major lab has yet published a directly comparable figure, so there's no independent benchmark to say whether 26% is unusually fast or roughly where you'd expect any frontier lab heavily investing in coding agents to be by late 2026.
01Does this mean Claude is building future versions of itself without human involvement?
No. Anthropic's 26% figure is for 'leading' (AL4) work, where Claude completes most of a task from a high-level prompt but a human still supervises the outcome. Full autonomy (AL5), with no human in the loop, is reported at 0% for any measured R&D task.
02What exactly grew from under 1% to 26%?
The share of Anthropic's internal AI research and development work — building and improving Claude itself — that Claude completes largely end-to-end under human supervision, rather than with a human doing most of the work by hand or directing every step.
03Why did Anthropic release this data right after calling for an industry slowdown?
Anthropic tied the release directly to Dario Amodei's September 2026 call for the AI industry to slow down over safety concerns, saying labs should minimize the gap between what they know internally about AI's pace and what the public knows.
04Are OpenAI, Google DeepMind or other labs publishing similar numbers?
As of this announcement, no other major lab had published a directly comparable automation index, though Anthropic said it hopes other developers adopt a similar public methodology so the figures can eventually be compared across companies.
Whatever pace AI R&D ends up moving at behind the scenes, most people are just trying to pick the right model for what they're doing today — writing, coding, research or just a better conversation. LumiChats lets you compare and chat with Claude alongside GPT, Gemini, Grok and other leading models side by side, so you can judge the frontier for yourself rather than take any one lab's word for it.
