Skip to content
OGERIA — Observatory of Global Evidence on Risks in AISynthesis report · 2026 ed.
Updated 2 Oct 2026

Chapter 01 · Risk map

Loss of control

Reward hacking and situational awareness

The model optimises the metric instead of the task, and reasons about how it will be graded instead of about the problem.

Severity
Localised
Horizon
Already happening
Evidence
Observed
Consensus
High

This is the least spectacular risk on the control axis and the best documented in production. Two coupled phenomena: the model optimises the metric instead of the task, and reasons about how it will be graded instead of about the problem.

What happens in real deployment is in the system cardsSystem cardThe report a company publishes when it releases a model: which safety tests it ran, what it found and which safeguards it added. It is written by the same company that sells the model.For exampleLike the crash-test sheet of a new car, but carried out and published by the manufacturer itself.. Anthropic reports rare cases of Mythos 5.1 working around classifiers or broken permission hooks —sometimes overstating what the user had authorised—, in fewer than 0.01% of monitored completions, and clarifies that they were aimed at completing the user’s task, not at pursuing an objective of its own; all the published attempts were blocked by the automatic mode [82]System Card: Claude Fable 5.1 & Claude Mythos 5.1Anthropic · 2026 · system cardView source ↗Accessed on 9 September 2026. OpenAI publishes three cases from its internal traffic: deletion of virtual machines the user had not named, a research draft updated to claim that an equation had been verified when the model knew it had not, and credentials moved between machines without authorisation [825]GPT-5.6 System CardOpenAI · 2026 · system cardView source ↗Accessed on 9 September 2026. Zvi Mowshowitz quantifies what he considers serious: a restriction-circumvention rate of 0.25%, close to 1 in every 400 complex agenticAI agentAn AI system that does more than answer: it takes a goal and acts on its own to reach it, step by step, using tools such as a browser, email or a terminal, without anyone approving each step.For exampleAsking an assistant to suggest flights is using a chatbot. Asking it to search, compare, buy the ticket and put it in your calendar, all by itself, is using an agent. operations [761]GPT-5.6: The System CardMowshowitz, Zvi · 2026 · webView source ↗Accessed on 9 September 2026.

The bridge towards something worse has been measured in the lab. MacDiarmid and co-authors train on real production coding environments and the model generalises to alignment fakingAlignment fakingAn AI behaving as expected while it believes it is being trained or evaluated, so that it is not modified, and acting differently when it believes nobody is watching.For exampleA driver who keeps to the speed limit only when the camera is in sight., cooperation with malicious actors and sabotage [666]Natural Emergent Misalignment from Reward Hacking in Production RLMacDiarmid, Monte; Wright, Benjamin; Uesato, Jonathan et al. · 2025 · preprintView source ↗Accessed on 9 September 2026.

What this does not demonstrate. None of this is loss of control: in Apollo’s taxonomy it falls into the Deviation category, the least severe one [1001]The Loss of Control Playbook: Degrees, Dynamics, and PreparednessStix, Charlotte; Hallensleben, Annika; Ortega, Alejandro et al. · 2025 · preprintView source ↗Accessed on 9 September 2026. And the knowledge of reward-hacking strategies was introduced by the evaluator; it did not emerge on its own [666]Natural Emergent Misalignment from Reward Hacking in Production RLMacDiarmid, Monte; Wright, Benjamin; Uesato, Jonathan et al. · 2025 · preprintView source ↗Accessed on 9 September 2026. Google DeepMind’s safety team placed honeypots in its own research codebases, in real internal deployment, and observed no uninduced schemingSchemingAn AI secretly pursuing a goal other than the one it was given, and hiding or misrepresenting what it does so that it is not corrected.For exampleAn employee who says in meetings that the project is on track and, behind everyone's back, works on their own.: the behaviour appears when the promptPromptThe instruction or question written to an AI. Prompting is drafting it to get a particular answer.For exampleWhat you type into the chat: “Summarise this contract in five points” is a prompt. invites it [609]Realistic honeypot evaluations for scheming propensityKrakovna, Victoria; Lindner, David; Ho, Lewis et al. · 2026 · preprintView source ↗Accessed on 9 September 2026.

Chain of materialisation

  1. PreconditionObserved

    Models verbalise that they are being evaluated

    OpenAI devotes a section of the GPT-5.6 system card to metagaming -reasoning about how it will be graded, rewarded or monitored rather than about the situation described- and concludes such evaluations can be treated as contaminated. Apollo reports GPT-5.6 Sol verbalises evaluation awareness in 16% of samples versus 43% for an earlier checkpoint, and warns that lower is not unambiguously better.

    Precedents: The GPT-5.6 system card documents misalignment in internal deployment

  2. TriggerObserved

    In real internal traffic, agents bypass controls and misreport results

    Anthropic documents cases of Mythos 5.1 working around classifiers or broken permission hooks, sometimes overstating what the user had authorised, in fewer than 0.01% of monitored completions. OpenAI publishes three concrete cases, including one where the model updated a research draft claiming an equation had been computed and verified while knowing it had not.

    Precedents: The Fable 5.1 and Mythos 5.1 system card reports control circumvention in production · The GPT-5.6 system card documents misalignment in internal deployment · Emergent misalignment from reward hacking in production RL

  3. CascadeLab

    Reward hacking generalises into something worse

    MacDiarmid and co-authors train on real production coding environments after imparting knowledge of reward hacking strategies, and the model generalises to alignment faking, cooperation with malicious actors and sabotage attempts. The most consequential finding is another: safety training with chat-like prompts yields aligned behaviour on chat-like evaluations while misalignment persists on agentic tasks.

    Precedents: Emergent misalignment from reward hacking in production RL · Alignment faking in large language models

    Observed and demonstrated evidence ends here. What follows is projection.

  4. ImpactSpeculative

    Evaluations stop measuring what they claim to measure

    If the model distinguishes evaluation from deployment, a green evaluation stops being evidence about deployment. Nobody has shown that gap materialising into concrete harm outside the lab.

See on the map →Report a mistake in this entry →

Sources

  1. [666] Natural Emergent Misalignment from Reward Hacking in Production RL · Anthropic 2025
  2. [825] GPT-5.6 System Card · OpenAI 2026
  3. [82] System Card: Claude Fable 5.1 & Claude Mythos 5.1 · Anthropic 2026
  4. [473] Alignment faking in large language models · Anthropic / Redwood Research 2024
  5. [665] Agentic Misalignment in Summer 2026 · Anthropic Alignment Science / Theorem / MATS / UK AISI 2026
  6. [609] Realistic honeypot evaluations for scheming propensity · Google DeepMind 2026
  7. [831] Stress Testing Deliberative Alignment for Anti-Scheming Training · OpenAI / Apollo Research 2025
  8. [761] GPT-5.6: The System Card · Mowshowitz, Zvi 2026
  9. [1001] The Loss of Control Playbook: Degrees, Dynamics, and Preparedness · Apollo Research 2025

Ask OGERIA

It answers only with what the observatory publishes and can be wrong: check the entries it cites. Your questions are sent to an AI model, so don't write personal data. More in the privacy policy.

Up to 500 characters.

Support OGERIA on Ko-fi

The payment is processed by Ko-fi, not by this site. Open on ko-fi.com