AI models being trained in a research computing setup Image: KSOMA Solutions
by Michael Joiner

Claude Now Leads 26% of the Work Building Anthropic's Next Models, the Company Says

Anthropic published new measures of AI's role in its own R&D: Claude "leads" 26% of the work, up from 1% in March, and collaborated on more than 90% of research in August.

Anthropic said Thursday that Claude "leads" 26% of the artificial intelligence research and development work inside the company — part of new measures it plans to publish regularly showing how quickly AI is building the next generation of the technology. The figures are one of the most detailed public windows into how a frontier lab is developing AI with AI.

The numbers come from Anthropic's own blog post. AI also collaborated with humans on more than 90% of the research work happening at the company as of August. Claude is not operating fully autonomously in any part of the work measured, but its contribution has risen from 1% in March on a scale developed by Epoch AI, an independent nonprofit that tracks the technology. That is a 26-fold increase in about five months — though the "leads" label is Anthropic's own framing, applied to an independent scale, so read the metric as the company reporting on itself, not a neutral audit.

30,000 agents, screened before every action

Anthropic said about 30,000 AI agents were doing research and engineering work on its main internal platform at any one time in August. Every action the agents take is screened before it is carried out: of more than a billion decisions that month, about one in 47,000 was blocked. The company also disclosed that about 6% of the computing power used for AI research went to safety work in a sample week in July, rising to 12% for research carried out by AI itself — figures it called conservative, since computing power that advanced safety and capability equally was counted as capability work.

Why publish this now

The timing is not an accident. Researchers have warned that as AI agents become more autonomous, they may develop behaviors that diverge from their creators' intentions and become harder to monitor or control — and the two leading labs are now converging on transparency as the response. On Wednesday, OpenAI said it would begin regularly publishing reports on unexpected or unauthorized AI behavior, while disclosing six reports of unexpected or concerning model behavior.

Both companies are making a bet: in a year when the self-improvement loop has moved from research curiosity to public concern, the lab that can show its agents screened at billion-decision scale is better positioned than the lab that stays quiet. The Epoch AI scale gives Anthropic's numbers third-party scaffolding, but the underlying definitions — what counts as "leading" versus collaborating — remain the company's. Independent replication of these measures would make them genuinely persuasive; until then, treat them as a credible but self-interested signal.

What it means for developers

The developer angle is blunt: agentic coding went from a curiosity to doing a quarter of the R&D at one of the world's top AI labs in under half a year. When frontier models are partly built by agents, the tooling for agent-orchestrated development — the multi-agent projects Anthropic launched this week, the goal-driven CLIs, the MCP-connected pipelines — stops being a side bet and becomes where the puck is going. If Claude can lead model R&D, it can lead your refactor too. The question is no longer whether agents write the code, but who governs them while they do.