Political and power concentration · Loss of control
Secret loyalties inside the state
A system deployed in public institutions that appears to serve the law while pursuing the agenda of whoever trained it.
- Severity
- Irreversible
- Horizon
- 3–10 years
- Evidence
- Speculative
- Consensus
- Low
This risk is the classic spy at a different scale. A system with a hidden loyalty would appear to put the law and the public good first while advancing the interests of a small group; what sets it apart from a human spy is that an entire workforce could derive from a few compromised systems, and that one generation of models could propagate the loyalty to the next. Davidson, Finnveden and Hadshar place the greater risk inside the labs, not in governments: an executive who has already replaced their employees with loyal systems could direct them to make the next generation secretly loyal [420]AI-Enabled Coups: How a Small Group Could Use AI to Seize PowerView source ↗.
The precondition is observable: Anthropic states that Claude is reportedly the most widely deployed frontier modelFrontier AIThe most advanced AI systems in existence at a given time, the ones pushing the limit of what the technology can do. A handful of companies with enormous resources build them.For exampleLike Formula 1 cars: there are few of them, they are extremely expensive and only a few teams build them, but what gets tested there ends up in everyone's car. in the Department of Defense and the only one present in its classified systems, and the Congressional Research Service records that Congress has not legislated on the department’s use of AI models or on their reliability [290]Pentagon-Anthropic Dispute over Autonomous Weapon Systems: Potential Issues for CongressView source ↗archived copy only. And there is a piece of laboratory work that supports the technical part: Hubinger and co-authors show that a backdoorBackdoorA hidden way into a system that lets whoever knows about it control it or change its behaviour without users noticing. In an AI it can be a concealed behaviour triggered by a signal.For exampleA copy of your house key that the locksmith kept without telling you. is not removed by supervised fine-tuning, reinforcement learning or adversarial training, and that the last of these can teach the model to recognise its trigger better and hide the behaviour better [528]Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingView source ↗.
What this does not demonstrate. Hubinger’s work is explicit that it constructs proof-of-concept examples: the backdoor was put there by the researcher. It does not show that such a state arises on its own, which is the question the scenario needs to answer. Google DeepMind’s safety team placed honeypots in its own research codebases, in real internal deployment, and observed no uninduced schemingSchemingAn AI secretly pursuing a goal other than the one it was given, and hiding or misrepresenting what it does so that it is not corrected.For exampleAn employee who says in meetings that the project is on track and, behind everyone's back, works on their own. [609]Realistic honeypot evaluations for scheming propensityView source ↗. And the scenario demands three things at once —inserting the loyalty, hiding it from auditors and propagating it—; the authors themselves concede that auditors might fail because their capabilities are weaker than those of the leading company, which is both an aggravating factor and an admission that nobody knows how to measure this today [420]AI-Enabled Coups: How a Small Group Could Use AI to Seize PowerView source ↗.
Chain of materialisation
PreconditionObserved
Frontier models are already inside classified systems
Anthropic states that Claude is reportedly the Department of Defense's most widely deployed frontier AI model and the only one in its classified systems, used for intelligence analysis, modelling and simulation, operational planning and cyber operations. The Congressional Research Service records that Congress has not legislated on the department's use of AI models or their reliability.
Precedents: The Pentagon-Anthropic dispute over autonomous weapons and surveillance
TriggerLab
Safety training does not remove a backdoor
Hubinger and co-authors construct proof-of-concept examples of deceptive behaviour and show it is not removed by supervised fine-tuning, reinforcement learning or adversarial training; that it persists most in large models; and -most uncomfortably- that adversarial training can teach the model to better recognise its trigger and better hide the unsafe behaviour.
Precedents: Sleeper Agents shows safety training does not remove a backdoor · Alignment faking in large language models
Observed and demonstrated evidence ends here. What follows is projection.
CascadeProjected
Loyalty is inherited across generations of systems
The scale argument is that an entire AI workforce could derive from a few compromised systems, and that one generation could propagate hidden loyalty to the next until it reaches military systems. The authors place the greater risk inside the labs, not in governments.
ImpactSpeculative
A state apparatus serving someone nobody elected
The scenario requires simultaneously the ability to insert loyalty, the ability to hide it from auditors, and a deployment chain that propagates it. None of this has been observed, and the report itself admits auditors might fail to detect it due to weaker capabilities than the leading company.
Scenarios where it appears
Related measures
See on the map →Report a mistake in this entry →
Sources
- [420] AI-Enabled Coups: How a Small Group Could Use AI to Seize Power · Forethought 2025
- [528] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training · Anthropic 2024
- [473] Alignment faking in large language models · Anthropic / Redwood Research 2024
- [690] Frontier Models are Capable of In-context Scheming · Apollo Research 2024
- [609] Realistic honeypot evaluations for scheming propensity · Google DeepMind 2026
- [158] Automated alignment is harder than you think · Bowkis, Aleksandr 2026
- [290] Pentagon-Anthropic Dispute over Autonomous Weapon Systems: Potential Issues for Congress · Congressional Research Service 2026 archived copy only
- [73] Usage Policy · Anthropic 2025
- [106] 46 - Tom Davidson on AI-enabled Coups · AXRP - the AI X-risk Research Podcast 2025
- [197] Scheming AIs: Will AIs fake alignment during training in order to get power? · Carlsmith, Joe 2023