
Anthropic has warned that automated artificial intelligence research could become a major safety concern within the next six to 12 months as AI systems take on a growing share of the work involved in developing new AI models.
In its August 2026 Risk Report, published on August 14, Anthropic said it still considers the current catastrophic risk from automated research and development to be low. However, the company said its confidence in that assessment has declined as existing evaluations begin to saturate and evidence shows that AI is already contributing to faster AI development.
Anthropic’s concern is not that AI systems have already achieved recursive self-improvement. The company said that has not happened and is not inevitable. Instead, it is tracking a possible progression in which AI systems increasingly perform the engineering, experimentation and analysis needed to develop more capable AI systems.
The company defines its automated R&D threshold through two possible conditions. One would be reached if AI systems could replace Anthropic’s research scientists and research engineers at competitive cost. The other would be reached if AI-driven automation produced roughly a doubling in the rate of AI capability progress compared with the expected baseline, with the acceleration substantially attributable to automated AI research and engineering.
Anthropic said it has not reached either threshold. Its current assessment indicates that AI is contributing to meaningful acceleration, but that the overall rate of AI progress remains below the approximately twofold acceleration used in its threshold.
The company nevertheless said the evidence is becoming harder to interpret because several of its existing capability evaluations are approaching saturation. Anthropic said this makes it more difficult to determine how quickly frontier systems are improving and contributes to its increased uncertainty about the current risk level.
Evidence from inside Anthropic shows how much AI-assisted development has already changed the company’s engineering work. In a separate report from the Anthropic Institute, the company said more than 80 percent of the code merged into its codebase as of May 2026 had been authored by Claude. It also reported that, during the second quarter of 2026, the typical engineer was merging eight times as much code per day as in 2024.
Anthropic cautioned that lines of code are an imperfect measure because they capture quantity rather than quality. The company said the eightfold increase therefore overstates the true productivity gain, although it still provides evidence of acceleration. A March 2026 survey of 130 Anthropic employees found that the median respondent estimated producing about four times as much output with its Mythos Preview system as without AI assistance on the projects they would otherwise have worked on.
AI systems are also taking on increasingly complex engineering tasks. Anthropic reported that Claude’s success rate on its most open-ended coding tasks reached 76 percent in May 2026, an increase of 50 percentage points over six months. In one example, Claude investigated a problem that was causing tens of thousands of model-training jobs to crash, identified the relevant debugging setting and confirmed a fix in about two hours. Anthropic said the work would normally have taken two to three days.
The company’s systems have also become more capable at executing defined research experiments. Anthropic said Claude Mythos Preview was able to increase the speed of a model-training task by about 52 times in April 2026, compared with an average improvement of about three times from Claude Opus 4 in May 2025. Anthropic said a skilled human researcher would need four to eight hours to achieve a fourfold speedup on the same type of task.
Anthropic has additionally tested whether AI agents can conduct more open-ended research. In an April 2026 project, Claude-powered agents were given an AI safety research problem and allowed to propose hypotheses, conduct experiments, share findings with other agents and iterate. Anthropic said two human researchers recovered about 23 percent of a defined performance gap over roughly a week, while the agents recovered 97 percent over 800 cumulative hours at a computing cost of about $18,000.
The company stressed significant limitations to that result. The problem was chosen by humans, humans created the scoring framework, and the findings did not transfer cleanly to production-scale models. Within those limits, however, Anthropic said the agents designed the individual experiments themselves.
Anthropic continues to identify research judgment as a major limitation. Its systems can execute a well-defined experiment and can increasingly solve underspecified engineering problems, but humans still play an important role in deciding which problems deserve attention, which results should be trusted and when a line of investigation should be abandoned.
The company also examined 129 real research-session decisions in which human researchers had taken a detour before eventually returning to a productive path. On this limited comparison, Anthropic said its strongest model in November 2025 suggested a better next step than the human choice 51 percent of the time, while Mythos Preview reached 64 percent in April 2026. Anthropic described the result as an early indication that AI systems are improving at a form of judgment relevant to research, while noting that the evaluation was not a direct head-to-head test of researchers and models.
Anthropic’s concern extends beyond AI development itself, but the company sees automated AI research as a particularly important early indicator because AI systems are already well suited to computational research and because the company can observe their effect directly within its own development process.
The company interviewed 31 academics, scientists, technology executives, government officials and other experts working across areas including robotics, energy, biotechnology, semiconductors, weapons development, neurotechnology and nanotechnology. Anthropic said it did not find evidence that automated research in those fields is currently producing the kind of rapid acceleration it is concerned about in AI.
One reason is that many other areas still depend on physical-world activities such as laboratory experiments, manufacturing, clinical trials and physical testing. AI can accelerate computational parts of those processes, but the overall pace can still be constrained by physical infrastructure, equipment and experiments.
Anthropic is also preparing for the possibility that automated AI research crosses its defined threshold. Its Responsible Scaling Policy calls for stronger safeguards when models reach specified capability levels, including stronger security, monitoring and controls around highly capable systems.
The company said it has not yet achieved its planned internal standard of having comprehensive visibility into activity across its AI development systems. Anthropic has set January 1, 2027, as the target date for reaching its “eyes on everything” goal.
Anthropic’s warning comes as the company increasingly describes AI systems as participants in the development process rather than simply tools used by researchers. Its June 2026 report, “When AI builds itself,” said the progression from code assistance to coding agents and then autonomous agents could eventually lead to systems capable of building and training future models themselves.
Anthropic said it is not claiming that such recursive self-improvement is imminent or certain. Its current evidence instead points to a narrower conclusion: AI is already helping accelerate AI development, the systems are becoming better at carrying out longer and more complex research tasks, and the tools used to measure those capabilities are becoming less able to distinguish between successive generations of frontier models.
That combination is why Anthropic now considers it plausible that automated AI research could become a major concern within 6 to 12 months, even though the company still assesses the present level of catastrophic risk from the technology as low.
Discover more from Aree Blog
Subscribe now to keep reading and get access to the full archive.


