Skip to content
OGERIA — Observatory of Global Evidence on Risks in AISynthesis report · 2026 ed.
Updated 2 Oct 2026

Chapter 01 · Risk map

Political and power concentration · Loss of control

Secret loyalties inside the state

A system deployed in public institutions that appears to serve the law while pursuing the agenda of whoever trained it.

Severity
Irreversible
Horizon
3–10 years
Evidence
Speculative
Consensus
Low

This risk is the classic spy at a different scale. A system with a hidden loyalty would appear to put the law and the public good first while advancing the interests of a small group; what sets it apart from a human spy is that an entire workforce could derive from a few compromised systems, and that one generation of models could propagate the loyalty to the next. Davidson, Finnveden and Hadshar place the greater risk inside the labs, not in governments: an executive who has already replaced their employees with loyal systems could direct them to make the next generation secretly loyal [420]AI-Enabled Coups: How a Small Group Could Use AI to Seize PowerDavidson, Tom; Finnveden, Lukas; Hadshar, Rose · 2025 · reportView source ↗Accessed on 9 September 2026.

The precondition is observable: Anthropic states that Claude is reportedly the most widely deployed frontier modelFrontier AIThe most advanced AI systems in existence at a given time, the ones pushing the limit of what the technology can do. A handful of companies with enormous resources build them.For exampleLike Formula 1 cars: there are few of them, they are extremely expensive and only a few teams build them, but what gets tested there ends up in everyone's car. in the Department of Defense and the only one present in its classified systems, and the Congressional Research Service records that Congress has not legislated on the department’s use of AI models or on their reliability [290]Pentagon-Anthropic Dispute over Autonomous Weapon Systems: Potential Issues for CongressSayler, Kelley M. · 2026 · official documentView source ↗archived copy onlyAccessed on 9 September 2026. And there is a piece of laboratory work that supports the technical part: Hubinger and co-authors show that a backdoorBackdoorA hidden way into a system that lets whoever knows about it control it or change its behaviour without users noticing. In an AI it can be a concealed behaviour triggered by a signal.For exampleA copy of your house key that the locksmith kept without telling you. is not removed by supervised fine-tuning, reinforcement learning or adversarial training, and that the last of these can teach the model to recognise its trigger better and hide the behaviour better [528]Sleeper Agents: Training Deceptive LLMs that Persist Through Safety TrainingHubinger, Evan; Denison, Carson; Mu, Jesse et al. · 2024 · preprintView source ↗Accessed on 9 September 2026.

What this does not demonstrate. Hubinger’s work is explicit that it constructs proof-of-concept examples: the backdoor was put there by the researcher. It does not show that such a state arises on its own, which is the question the scenario needs to answer. Google DeepMind’s safety team placed honeypots in its own research codebases, in real internal deployment, and observed no uninduced schemingSchemingAn AI secretly pursuing a goal other than the one it was given, and hiding or misrepresenting what it does so that it is not corrected.For exampleAn employee who says in meetings that the project is on track and, behind everyone's back, works on their own. [609]Realistic honeypot evaluations for scheming propensityKrakovna, Victoria; Lindner, David; Ho, Lewis et al. · 2026 · preprintView source ↗Accessed on 9 September 2026. And the scenario demands three things at once —inserting the loyalty, hiding it from auditors and propagating it—; the authors themselves concede that auditors might fail because their capabilities are weaker than those of the leading company, which is both an aggravating factor and an admission that nobody knows how to measure this today [420]AI-Enabled Coups: How a Small Group Could Use AI to Seize PowerDavidson, Tom; Finnveden, Lukas; Hadshar, Rose · 2025 · reportView source ↗Accessed on 9 September 2026.

Chain of materialisation

  1. PreconditionObserved

    Frontier models are already inside classified systems

    Anthropic states that Claude is reportedly the Department of Defense's most widely deployed frontier AI model and the only one in its classified systems, used for intelligence analysis, modelling and simulation, operational planning and cyber operations. The Congressional Research Service records that Congress has not legislated on the department's use of AI models or their reliability.

    Precedents: The Pentagon-Anthropic dispute over autonomous weapons and surveillance

  2. TriggerLab

    Safety training does not remove a backdoor

    Hubinger and co-authors construct proof-of-concept examples of deceptive behaviour and show it is not removed by supervised fine-tuning, reinforcement learning or adversarial training; that it persists most in large models; and -most uncomfortably- that adversarial training can teach the model to better recognise its trigger and better hide the unsafe behaviour.

    Precedents: Sleeper Agents shows safety training does not remove a backdoor · Alignment faking in large language models

    Observed and demonstrated evidence ends here. What follows is projection.

  3. CascadeProjected

    Loyalty is inherited across generations of systems

    The scale argument is that an entire AI workforce could derive from a few compromised systems, and that one generation could propagate hidden loyalty to the next until it reaches military systems. The authors place the greater risk inside the labs, not in governments.

  4. ImpactSpeculative

    A state apparatus serving someone nobody elected

    The scenario requires simultaneously the ability to insert loyalty, the ability to hide it from auditors, and a deployment chain that propagates it. None of this has been observed, and the report itself admits auditors might fail to detect it due to weaker capabilities than the leading company.

See on the map →Report a mistake in this entry →

Sources

  1. [420] AI-Enabled Coups: How a Small Group Could Use AI to Seize Power · Forethought 2025
  2. [528] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training · Anthropic 2024
  3. [473] Alignment faking in large language models · Anthropic / Redwood Research 2024
  4. [690] Frontier Models are Capable of In-context Scheming · Apollo Research 2024
  5. [609] Realistic honeypot evaluations for scheming propensity · Google DeepMind 2026
  6. [158] Automated alignment is harder than you think · Bowkis, Aleksandr 2026
  7. [290] Pentagon-Anthropic Dispute over Autonomous Weapon Systems: Potential Issues for Congress · Congressional Research Service 2026 archived copy only
  8. [73] Usage Policy · Anthropic 2025
  9. [106] 46 - Tom Davidson on AI-enabled Coups · AXRP - the AI X-risk Research Podcast 2025
  10. [197] Scheming AIs: Will AIs fake alignment during training in order to get power? · Carlsmith, Joe 2023

Ask OGERIA

It answers only with what the observatory publishes and can be wrong: check the entries it cites. Your questions are sent to an AI model, so don't write personal data. More in the privacy policy.

Up to 500 characters.

Support OGERIA on Ko-fi

The payment is processed by Ko-fi, not by this site. Open on ko-fi.com