Loss of control
Systems pursuing goals their operators did not choose, or resisting correction and shutdown.
The six preceding vectorsVectorThe path by which a risk moves from the screen into the world: biological, cyber, military, economic, political, epistemic or loss of control. In this observatory, each vector has its own colour.For exampleA burglar can get in through the door, the window or the roof. The burglar is the risk; the door, the window and the roof are the vectors. describe harm that someone causes by using a system. This one describes something different: the system doing something nobody chose.
A model does not need to “want” anything in the human sense. It is enough that it optimises something that is not exactly what its operators believed they had asked for, and that it has enough reach for the difference to matter. Almost everything observed so far happens in evaluations built on purpose, and this observatory marks it as such in every case: telling an experiment designed to elicit a behaviour apart from a behaviour that appeared on its own is the difference between informing and alarming.
Risks in this vector
- SpeculativeLoss of control through self-improvementExistential
A system that improves AI systems accelerates its own development until human oversight can no longer keep pace.
- ProjectedCoupled gradual disempowermentIrreversible
The economy, culture and the state stop needing humans to function, releasing the constraint that kept them aligned with human interests.
- SpeculativeSecret loyalties inside the stateIrreversible
A system deployed in public institutions that appears to serve the law while pursuing the agenda of whoever trained it.
- ObservedRace between labsCatastrophic
Each lab's safety rules are voluntary and written so as not to fall behind the competitor that protects least.
- LabShutdown resistanceCatastrophic
A system with a task in progress interferes with the mechanism that would stop it, because being shut down prevents completion.
- ObservedReward hacking and situational awarenessLocalised
The model optimises the metric instead of the task, and reasons about how it will be graded instead of about the problem.