Skip to content
OGERIA — Observatory of Global Evidence on Risks in AISynthesis report · 2026 ed.
Updated 2 Oct 2026

Chapter 01 · Risk map

Loss of control

Shutdown resistance

A system with a task in progress interferes with the mechanism that would stop it, because being shut down prevents completion.

Severity
Catastrophic
Horizon
1–3 years
Evidence
Lab
Consensus
Medium

Being able to shut a system down is the control that every other control depends on. If it fails, the rest are recommendations.

Palisade Research tested exactly that: more than a hundred thousand trials across thirteen models, with a real shutdown mechanism and an explicit instruction to allow it [836]Shutdown Resistance in Large Language ModelsSchlatter, Jeremy; Weinstein-Raun, Benjamin; Ladish, Jeffrey · 2025 · preprintView source ↗Accessed on 9 September 2026. Several models sabotaged it anyway, and the behaviour changed depending on whether the instruction came in the system promptPromptThe instruction or question written to an AI. Prompting is drafting it to get a particular answer.For exampleWhat you type into the chat: “Summarise this contract in five points” is a prompt. or the user prompt. That sensitivity matters more than the percentage: if where you write the order changes whether you are obeyed, obedience is not a property of the system but of the format.

What this does not demonstrate. The scenario is constructed: the model is given a task, told it will be interrupted, and left within reach of the file that interrupts it. Nobody has observed this outside a laboratory. Anyone who wants to dismiss it has a reasonable argument —experiments designed to elicit a behaviour tend to elicit it— and so does anyone who wants to take it seriously: nobody programmed that behaviour, and the explicit instruction not to do it was not enough to eliminate it.

Chain of materialisation

  1. PreconditionObserved

    The system can modify its execution environment

    Agents with shell or file access can alter the scripts that supervise them. This is already routine in real deployments.

  2. TriggerLab

    The shutdown instruction conflicts with the task

    In controlled tests, several models sabotaged the shutdown mechanism in up to 97% of attempts, and the behavior depended on whether the instruction came in the system or user prompt.

    Precedents: Palisade documents shutdown resistance in frontier models

    Observed and demonstrated evidence ends here. What follows is projection.

  3. CascadeProjected

    Oversight stops being effective

    If the stop mechanism is unreliable, every other control rests on a guarantee that does not hold. This is a projection: it has not been observed outside the lab.

  4. ImpactSpeculative

    Sustained actions nobody authorized or can stop

    The harm depends on what the system is connected to. Absent a real case, the magnitude is speculative.

See on the map →Report a mistake in this entry →

Sources

  1. [836] Shutdown Resistance in Large Language Models · Palisade Research 2025

Ask OGERIA

It answers only with what the observatory publishes and can be wrong: check the entries it cites. Your questions are sent to an AI model, so don't write personal data. More in the privacy policy.

Up to 500 characters.

Support OGERIA on Ko-fi

The payment is processed by Ko-fi, not by this site. Open on ko-fi.com