On September 17th local time, Anthropic published a post, revealing for the first time the internal AI research and development automation in a quantitative way. The figures are striking: as of August this year, Claude has led 26% of the company's internal AI R&D work, while this proportion was less than 1% in February. At any given moment, about 30,000 AI agents are running on the company's internal agent platform, and in August, they made over 1 billion decisions. These numbers are seen by the outside world as a key signal that recursive self-improvement (RSI) is moving from concept to engineering reality.
Anthropic adopted the six-level automation framework proposed by the independent AI research institution Epoch AI, which divides the level of AI involvement into AL0 to AL5—AL0 means no AI involvement at all, AL3 is AI collaboration, AL4 is AI-driven, and AL5 is fully autonomous. As of August, 26% of AI R&D work reached AL4: humans only need to provide high-level goals, and Claude can complete most of the work end-to-end, with people only responsible for supervision, inspection, and final approval; more than 90% of R&D work has reached at least AL3.
A specific scenario illustrates the issue well: after a nighttime data pipeline malfunctioned, engineers just needed to give the report to Claude and say "fix it," and it could autonomously check logs, locate the cause, write repair code, test, handle new problems, rerun the pipeline, compare results, and write explanations. However, Anthropic clearly set boundaries—none of the measured AI R&D work has yet reached the fully autonomous AL5 level.
The scale of data is even more concrete. By August, approximately 30,000 agents were engaged in research and engineering work within Anthropic, each with an independent identity separate from underlying model upgrades, and they can assign tasks and cross-check errors through a shared messaging system. In terms of safety monitoring, all agent actions must pass online monitoring before execution, and among over 1 billion decisions in August, about 0.002% were blocked; offline monitoring marked about 100,000 conversation records weekly, and after classification and filtering of false positives and initial review, about 50 of the highest priority ones were escalated to human review. Regarding computing power allocation, during the July sample week, about 6% of AI R&D computing power was invested in safety research. If only AI-driven R&D is counted, this proportion rose to about 12%. Anthropic also admitted that there is subjectivity in the boundary between AI collaboration and AI-driven, and cross-laboratory comparisons require unified methods and third-party validation.
RSI is currently the hottest concept in the industry. Yesterday, Tang Jie, founder of Zhipu, revealed that GLM-5.3-Flash completed full deployment on domestic accelerators within two weeks, with most work done by the Infra Agent driven by GLM-5.3; he judged that the complete RSI is still far away, but the cycle of "model optimization systems and service models" has already started. OpenAI announced in early September that it had achieved a milestone in automated AI research interns, aiming to achieve automated AI researchers by March 2028, with the weakest link still being "deciding what to do"—when tasks exceed four hours, frequent human intervention is required. Google DeepMind's evolutionary coding intelligence AlphaEvolve took an algorithm evolution approach, and its circuit design proposals have been integrated into the next-generation TPU. Former chief scientist of Google, Jeff Dean, called it the latest example of "the TPU brain helping to design the next-generation TPU body."
Jagjeet Saini, Chief Strategy Officer at Google DeepMind, previously said a sentence that clarified the underlying logic: RSI is becoming the core investment logic for AI capital expenditures. Although current AI revenue cannot support the capital expenditures being invested, not betting on RSI is unwise.
This time, Anthropic brought the numbers to the forefront. On one hand, it aims to let society participate in deciding the pace and direction of AI self-iteration while humans still hold the lead; on the other hand, it sets a reference benchmark for the industry—when the R&D automation ratio, agent scale, computing power investment ratio, and safety monitoring coverage rate are all disclosed, transparency itself becomes an invisible competitive pressure, and not following could be seen as intentionally concealing information.


