Loss of control through self-improvement
A system that improves AI systems accelerates its own development until human oversight can no longer keep pace.
- Severity
- Existential
- Horizon
- 3–10 years
- Evidence
- Speculative
- Consensus
- Low
The hypothesis is that a system capable of improving AI systems accelerates its own development faster than human oversight can keep up with. It is the risk with the highest severity in the observatory and the weakest empirical backing, and both things have to be said together.
The inputs to the argument are real and measured. Anthropic reports that more than 80% of the code merged into its codebase was written by Claude in May 2026, that the typical engineer merged eight times more code per day than in 2024, and that in an experiment where two human researchers recovered 23% of a performance gap in a week, the agentsAI agentAn AI system that does more than answer: it takes a goal and acts on its own to reach it, step by step, using tools such as a browser, email or a terminal, without anyone approving each step.For exampleAsking an assistant to suggest flights is using a chatbot. Asking it to search, compare, buy the ticket and put it in your calendar, all by itself, is using an agent. recovered 97% [75]When AI builds itselfView source ↗. METR measures that the time horizon of tasks doubles every 89 days since 2024, with confidence intervalsConfidence intervalThe range within which the true value probably lies, given what was measured. The wider it is, the less precise the measurement. It is abbreviated CI.For exampleLike saying you will arrive “between 7 and 7:20”: you do not know the exact minute, but you do know the margin. that METR itself describes as very wide [712]Time Horizon 1.1View source ↗.
What this does not demonstrate. “Claude writes 80% of the code” is not “Claude does the research”, and the one pointing this out is the same document: the lines-of-code metric is almost certainly an overestimate of the real gain, the 52× speed-up should not be read as a speed-up of training in the real world, and large gaps persist when it comes to choosing objectives [75]When AI builds itselfView source ↗. Three months later, the same company determines that its reference model does not cross the autonomy threshold because it does not observe a sustained 2× acceleration attributable to AI in the pace of its own progress [82]System Card: Claude Fable 5.1 & Claude Mythos 5.1View source ↗. And there is direct evidence against: a consortium led from Princeton gave frontier agentsFrontier AIThe most advanced AI systems in existence at a given time, the ones pushing the limit of what the technology can do. A handful of companies with enormous resources build them.For exampleLike Formula 1 cars: there are few of them, they are extremely expensive and only a few teams build them, but what gets tested there ends up in everyone's car. the questions from two unpublished papers submitted to NeurIPS 2026, with six days and thousands of dollars of computeComputeThe computing power used to train and run an AI: thousands of specialised chips working for weeks in data centres. It is expensive and concentrated in a few companies.For exampleIf AI were a bakery, compute would be the ovens: without big ovens, it does not matter how good the recipe is., and the original authors rejected both results without ambiguity [598]Can AI agents conduct open-ended AI research? Early evidence from two case studiesView source ↗. Jack Clark, co-author of the self-improvementSelf-improvementAn AI helping to design or train the version that replaces it, which then does the same for the next one, faster each time and with fewer people involved.For exampleLike an apprentice who, once skilled, trains the next one, who learns faster and teaches even better. document, calls it a bearish signal for short timelines to recursive self-improvement [736]AI's recursive self-improvement might not come so quickly after allView source ↗.
Chain of materialisation
PreconditionObserved
AI already writes most of the code of those who build it
Anthropic reports that over 80% of the code merged into its codebase was authored by Claude in May 2026, that the typical engineer merged eight times as much code per day as in 2024, and that success on the most open-ended tasks reached 76%, up fifty points in six months.
Precedents: Anthropic publishes When AI builds itself · METR publishes Time Horizon 1.1 with an expanded suite
TriggerLab
Agents close performance gaps that humans do not
In a performance-gap recovery experiment, two human researchers recovered about 23% in a week and the agents 97%. The company itself qualifies the result: humans chose the problem and created the scoring rubric.
Precedents: Anthropic publishes When AI builds itself
Observed and demonstrated evidence ends here. What follows is projection.
CascadeProjected
The jump from executing research to directing it
Thomas Kwa's eight-parameter timelines model has a median prediction of over 99% automation of AI R&D by around 2032; the author himself says he does not weight the exact timelines much because he has not thought hard about parameter values, and the model does not separately treat research taste.
Precedents: The AI 2027 scenario is published
ImpactSpeculative
Strict loss of control, severe and permanent
In Apollo's taxonomy this is Strict LoC, maximally severe and permanent, including extinction. There is no observed loss-of-control incident outside the lab, and none of the known real cases goes beyond the Deviation category.
Related measures
See on the map →Report a mistake in this entry →
Sources
- [75] When AI builds itself · Anthropic 2026
- [82] System Card: Claude Fable 5.1 & Claude Mythos 5.1 · Anthropic 2026
- [598] Can AI agents conduct open-ended AI research? Early evidence from two case studies · Princeton University y otros 2026
- [736] AI's recursive self-improvement might not come so quickly after all · MIT Technology Review 2026
- [712] Time Horizon 1.1 · METR 2026
- [713] Task-Completion Time Horizons of Frontier AI Models · METR 2026
- [619] A simpler AI timelines model predicts 99% AI R&D automation in ~2032 · METR 2026
- [1001] The Loss of Control Playbook: Degrees, Dynamics, and Preparedness · Apollo Research 2025
- [196] Is Power-Seeking AI an Existential Risk? · Carlsmith, Joseph 2022
- [1037] What power-seeking theorems do not show · Vanderbilt University / Global Priorities Institute 2024 archived copy only
- [633] The case for AI doom isn't very convincing · Understanding AI 2025
- [867] Anthropic Alignment Lead Issues Warning About AI Killing Humans As Researcher Resigns · Forbes 2026 archived copy only