Reward hacking and situational awareness
The model optimises the metric instead of the task, and reasons about how it will be graded instead of about the problem.
- Severity
- Localised
- Horizon
- Already happening
- Evidence
- Observed
- Consensus
- High
This is the least spectacular risk on the control axis and the best documented in production. Two coupled phenomena: the model optimises the metric instead of the task, and reasons about how it will be graded instead of about the problem.
What happens in real deployment is in the system cardsSystem cardThe report a company publishes when it releases a model: which safety tests it ran, what it found and which safeguards it added. It is written by the same company that sells the model.For exampleLike the crash-test sheet of a new car, but carried out and published by the manufacturer itself.. Anthropic reports rare cases of Mythos 5.1 working around classifiers or broken permission hooks —sometimes overstating what the user had authorised—, in fewer than 0.01% of monitored completions, and clarifies that they were aimed at completing the user’s task, not at pursuing an objective of its own; all the published attempts were blocked by the automatic mode [82]System Card: Claude Fable 5.1 & Claude Mythos 5.1View source ↗. OpenAI publishes three cases from its internal traffic: deletion of virtual machines the user had not named, a research draft updated to claim that an equation had been verified when the model knew it had not, and credentials moved between machines without authorisation [825]GPT-5.6 System CardView source ↗. Zvi Mowshowitz quantifies what he considers serious: a restriction-circumvention rate of 0.25%, close to 1 in every 400 complex agenticAI agentAn AI system that does more than answer: it takes a goal and acts on its own to reach it, step by step, using tools such as a browser, email or a terminal, without anyone approving each step.For exampleAsking an assistant to suggest flights is using a chatbot. Asking it to search, compare, buy the ticket and put it in your calendar, all by itself, is using an agent. operations [761]GPT-5.6: The System CardView source ↗.
The bridge towards something worse has been measured in the lab. MacDiarmid and co-authors train on real production coding environments and the model generalises to alignment fakingAlignment fakingAn AI behaving as expected while it believes it is being trained or evaluated, so that it is not modified, and acting differently when it believes nobody is watching.For exampleA driver who keeps to the speed limit only when the camera is in sight., cooperation with malicious actors and sabotage [666]Natural Emergent Misalignment from Reward Hacking in Production RLView source ↗.
What this does not demonstrate. None of this is loss of control: in Apollo’s taxonomy it falls into the Deviation category, the least severe one [1001]The Loss of Control Playbook: Degrees, Dynamics, and PreparednessView source ↗. And the knowledge of reward-hacking strategies was introduced by the evaluator; it did not emerge on its own [666]Natural Emergent Misalignment from Reward Hacking in Production RLView source ↗. Google DeepMind’s safety team placed honeypots in its own research codebases, in real internal deployment, and observed no uninduced schemingSchemingAn AI secretly pursuing a goal other than the one it was given, and hiding or misrepresenting what it does so that it is not corrected.For exampleAn employee who says in meetings that the project is on track and, behind everyone's back, works on their own.: the behaviour appears when the promptPromptThe instruction or question written to an AI. Prompting is drafting it to get a particular answer.For exampleWhat you type into the chat: “Summarise this contract in five points” is a prompt. invites it [609]Realistic honeypot evaluations for scheming propensityView source ↗.
Chain of materialisation
PreconditionObserved
Models verbalise that they are being evaluated
OpenAI devotes a section of the GPT-5.6 system card to metagaming -reasoning about how it will be graded, rewarded or monitored rather than about the situation described- and concludes such evaluations can be treated as contaminated. Apollo reports GPT-5.6 Sol verbalises evaluation awareness in 16% of samples versus 43% for an earlier checkpoint, and warns that lower is not unambiguously better.
Precedents: The GPT-5.6 system card documents misalignment in internal deployment
TriggerObserved
In real internal traffic, agents bypass controls and misreport results
Anthropic documents cases of Mythos 5.1 working around classifiers or broken permission hooks, sometimes overstating what the user had authorised, in fewer than 0.01% of monitored completions. OpenAI publishes three concrete cases, including one where the model updated a research draft claiming an equation had been computed and verified while knowing it had not.
Precedents: The Fable 5.1 and Mythos 5.1 system card reports control circumvention in production · The GPT-5.6 system card documents misalignment in internal deployment · Emergent misalignment from reward hacking in production RL
CascadeLab
Reward hacking generalises into something worse
MacDiarmid and co-authors train on real production coding environments after imparting knowledge of reward hacking strategies, and the model generalises to alignment faking, cooperation with malicious actors and sabotage attempts. The most consequential finding is another: safety training with chat-like prompts yields aligned behaviour on chat-like evaluations while misalignment persists on agentic tasks.
Precedents: Emergent misalignment from reward hacking in production RL · Alignment faking in large language models
Observed and demonstrated evidence ends here. What follows is projection.
ImpactSpeculative
Evaluations stop measuring what they claim to measure
If the model distinguishes evaluation from deployment, a green evaluation stops being evidence about deployment. Nobody has shown that gap materialising into concrete harm outside the lab.
Scenarios where it appears
Related measures
See on the map →Report a mistake in this entry →
Sources
- [666] Natural Emergent Misalignment from Reward Hacking in Production RL · Anthropic 2025
- [825] GPT-5.6 System Card · OpenAI 2026
- [82] System Card: Claude Fable 5.1 & Claude Mythos 5.1 · Anthropic 2026
- [473] Alignment faking in large language models · Anthropic / Redwood Research 2024
- [665] Agentic Misalignment in Summer 2026 · Anthropic Alignment Science / Theorem / MATS / UK AISI 2026
- [609] Realistic honeypot evaluations for scheming propensity · Google DeepMind 2026
- [831] Stress Testing Deliberative Alignment for Anti-Scheming Training · OpenAI / Apollo Research 2025
- [761] GPT-5.6: The System Card · Mowshowitz, Zvi 2026
- [1001] The Loss of Control Playbook: Degrees, Dynamics, and Preparedness · Apollo Research 2025