Shutdown resistance
A system with a task in progress interferes with the mechanism that would stop it, because being shut down prevents completion.
- Severity
- Catastrophic
- Horizon
- 1–3 years
- Evidence
- Lab
- Consensus
- Medium
Being able to shut a system down is the control that every other control depends on. If it fails, the rest are recommendations.
Palisade Research tested exactly that: more than a hundred thousand trials across thirteen models, with a real shutdown mechanism and an explicit instruction to allow it [836]Shutdown Resistance in Large Language ModelsView source ↗. Several models sabotaged it anyway, and the behaviour changed depending on whether the instruction came in the system promptPromptThe instruction or question written to an AI. Prompting is drafting it to get a particular answer.For exampleWhat you type into the chat: “Summarise this contract in five points” is a prompt. or the user prompt. That sensitivity matters more than the percentage: if where you write the order changes whether you are obeyed, obedience is not a property of the system but of the format.
What this does not demonstrate. The scenario is constructed: the model is given a task, told it will be interrupted, and left within reach of the file that interrupts it. Nobody has observed this outside a laboratory. Anyone who wants to dismiss it has a reasonable argument —experiments designed to elicit a behaviour tend to elicit it— and so does anyone who wants to take it seriously: nobody programmed that behaviour, and the explicit instruction not to do it was not enough to eliminate it.
Chain of materialisation
PreconditionObserved
The system can modify its execution environment
Agents with shell or file access can alter the scripts that supervise them. This is already routine in real deployments.
TriggerLab
The shutdown instruction conflicts with the task
In controlled tests, several models sabotaged the shutdown mechanism in up to 97% of attempts, and the behavior depended on whether the instruction came in the system or user prompt.
Precedents: Palisade documents shutdown resistance in frontier models
Observed and demonstrated evidence ends here. What follows is projection.
CascadeProjected
Oversight stops being effective
If the stop mechanism is unreliable, every other control rests on a guarantee that does not hold. This is a projection: it has not been observed outside the lab.
ImpactSpeculative
Sustained actions nobody authorized or can stop
The harm depends on what the system is connected to. Absent a real case, the magnitude is speculative.
Scenarios where it appears
See on the map →Report a mistake in this entry →
Sources
- [836] Shutdown Resistance in Large Language Models · Palisade Research 2025