Skip to content
OGERIA — Observatory of Global Evidence on Risks in AISynthesis report · 2026 ed.
Updated 2 Oct 2026

Chapter 06 · Protection

Isolate agents that run unattended

No automated process gets a capability whose worst case you cannot absorb: read-only credentials, no payment access, and a kill switch that lives outside the agent's reach.

Level
Personal
Cost
Low cost
Effort
Hours
Evidence
Observed

What it does not solve

It protects you from nothing happening outside your own machines, which is where almost all of this scenario's risk lives.

This measure assumes nothing about superintelligenceSuperintelligenceA hypothetical AI that would far outperform the most capable people at almost any intellectual task, including improving itself.For exampleIf playing chess against the world champion means certain defeat, imagine someone that far ahead, but in science, strategy and negotiation all at once.. It holds equally whether you think the risk is 1% or 30%, because it answers something that already happens: automated processes with more permissions than their worst case justifies.

Concretely: read-only credentials except where writing is the point; no access to payment methods; the policy of what the agentAI agentAn AI system that does more than answer: it takes a goal and acts on its own to reach it, step by step, using tools such as a browser, email or a terminal, without anyone approving each step.For exampleAsking an assistant to suggest flights is using a chatbot. Asking it to search, compare, buy the ticket and put it in your calendar, all by itself, is using an agent. may do kept where the agent cannot edit it; and a shutdown that does not depend on the cooperation of the process it shuts down. If the agent can rewrite its own permission list, you do not have a policy, you have a suggestion.

What you gain by applying it. That a failure is an incident and not a loss. It is not a measure against AI: it is the same least-privilege discipline already applied to any automated process, extended to systems that act with more initiative than a cron job.

Works if…

  • Rapid loss of control · Partial

    Bounds what your own agent can damage. Does nothing about systems you do not control.

Risks it addresses

See in the protection matrix →Report a mistake in this entry →

Sources

  1. [836] Shutdown Resistance in Large Language Models · Palisade Research 2025

Ask OGERIA

It answers only with what the observatory publishes and can be wrong: check the entries it cites. Your questions are sent to an AI model, so don't write personal data. More in the privacy policy.

Up to 500 characters.

Support OGERIA on Ko-fi

The payment is processed by Ko-fi, not by this site. Open on ko-fi.com