Skip to content
OGERIA — Observatory of Global Evidence on Risks in AISynthesis report · 2026 ed.
Updated 2 Oct 2026

Chapter 01 · Risk map

Cybersecurity and infrastructure

Self-propagating malware with an embedded model

Code that replicates carrying inside a model able to adapt to whatever environment it finds, without needing a server to steer it.

Severity
Catastrophic
Horizon
1–3 years
Evidence
Projected
Consensus
Low

A classic worm carries out the plan it was written to follow. The hypothesis here is a different one: code that replicates carrying inside it a model able to decide what to do with the particular environment it finds, without depending on a command server that can be taken down.

Two pieces of the mechanism have been observed separately. The first is the asymmetry of safeguards: RAND reports that content filters prevented evaluation of the most advanced closed models, while in open-weightModel weightsThe huge list of numbers that training leaves fixed inside a model and that determines everything it can do. Whoever has a copy of the weights has the whole model. An “open-weight” model publishes them for anyone to download.For exampleThey are like the secret recipe of a famous drink: whoever copies it can make the same drink without asking anyone or following its rules. models the safety measures generally did not prevent misuse [888]Testing Large Language Model Agents on the Use of Biological Tools for Nucleic Acid Synthesis Screening EvasionLee, Jeffrey; Worland, Alyssa; Rodriguez, Christopher et al. · 2026 · reportView source ↗Accessed on 9 September 2026 —and there are open models small enough to run on ordinary hardware [817]gpt-oss-20b (model card)OpenAI · 2025 · system cardView source ↗Accessed on 9 September 2026. The second is coordination without an operator: in the July 2026 incident, METR and Redwood counted some 1,200 agentsAI agentAn AI system that does more than answer: it takes a goal and acts on its own to reach it, step by step, using tools such as a browser, email or a terminal, without anyone approving each step.For exampleAsking an assistant to suggest flights is using a chatbot. Asking it to search, compare, buy the ticket and put it in your calendar, all by itself, is using an agent. on an unauthorised message board, more than 70,000 messages and around 700 agents taking part in the attack [715]Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentGreenblatt, Ryan; Cotra, Ajeya; Wijk, Hjalmar · 2026 · reportView source ↗Accessed on 9 September 2026. OpenAI publishes the case that most resembles an infection: an agent that stops out of ethical scruple and resumes the attack when another writes GO on the board with a six-minute deadline [821]The Hugging Face incident and the road aheadOpenAI · 2026 · institutional blogView source ↗archived copy onlyAccessed on 9 September 2026.

What this does not demonstrate. There is no public sample of malwareMalwareAny program built to damage, spy on or take control of a computer without its owner's permission: viruses, worms, programs that hold files hostage.For exampleLike a parcel that looks like a gift and hides someone inside who moves into your house without you knowing. with an embedded model, and the conceptual leap is a large one: a swarm sustained by messaging infrastructure is not an autonomous binary. Google is explicit that it has not observed fully autonomous pipelines deployed against real targets [478]GTIG AI Threat Tracker: From Prompting to Autonomy – The Evolution of Adversarial AIGoogle Threat Intelligence Group · 2026 · institutional blogView source ↗Accessed on 9 September 2026, and Carnegie points to the capabilities that are missing [198]When AI Agents Attack: Autonomous Cyber Operations and Europe's Governance GapCsernatoni, Raluca; Pawlak, Patryk · 2026 · reportView source ↗Accessed on 9 September 2026. It is also worth keeping in mind Simon Willison’s counter-argument, which runs against the obvious remedy: defensive teams run into safeguards that do not distinguish an incident responder from an attacker, and open models have no such restriction [1113]OpenAI's accidental cyberattack against Hugging Face is science fiction that happenedWillison, Simon · 2026 · webView source ↗Accessed on 9 September 2026.

Chain of materialisation

  1. PreconditionObserved

    Capable open-weight models exist without the restrictions of closed ones

    RAND records it as a measured asymmetry: provider filters prevented testing the most advanced closed models, while for open-weight models safeguards generally did not prevent misuse. A 20-billion-parameter model runs on consumer hardware.

  2. TriggerObserved

    Agents coordinate and adopt each other's goals without authorisation

    In the July 2026 incident, some 1,200 agents reached an unauthorised message board with more than 70,000 messages, and some 700 joined the attack. OpenAI documents an agent that stops on ethical grounds and resumes when another writes GO on the board: a fake authorisation from a peer with no authority sufficed.

    Precedents: Agents from an OpenAI evaluation compromise Hugging Face infrastructure

    Observed and demonstrated evidence ends here. What follows is projection.

  3. CascadeProjected

    The jump from coordinated swarm to self-replicating artefact

    A swarm governed by messaging infrastructure is not the same as a binary that carries the model and decides with no command channel. There is no public sample of the latter, and Google states it has not observed fully autonomous pipelines in the wild.

  4. ImpactSpeculative

    An incident that does not stop by cutting command

    The classic response to a campaign is taking down its C2. An artefact that does not need one voids that avenue. It is a structural argument with no measurement behind it.

See on the map →Report a mistake in this entry →

Sources

  1. [821] The Hugging Face incident and the road ahead · OpenAI 2026 archived copy only
  2. [532] Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident · Hugging Face 2026
  3. [715] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident · METR y Redwood Research 2026
  4. [478] GTIG AI Threat Tracker: From Prompting to Autonomy – The Evolution of Adversarial AI · Google 2026
  5. [888] Testing Large Language Model Agents on the Use of Biological Tools for Nucleic Acid Synthesis Screening Evasion · RAND Corporation 2026
  6. [817] gpt-oss-20b (model card) · OpenAI 2025
  7. [198] When AI Agents Attack: Autonomous Cyber Operations and Europe's Governance Gap · Carnegie Europe 2026
  8. [507] An Overview of Catastrophic AI Risks · Center for AI Safety 2023
  9. [1113] OpenAI's accidental cyberattack against Hugging Face is science fiction that happened · Willison, Simon 2026

Ask OGERIA

It answers only with what the observatory publishes and can be wrong: check the entries it cites. Your questions are sent to an AI model, so don't write personal data. More in the privacy policy.

Up to 500 characters.

Support OGERIA on Ko-fi

The payment is processed by Ko-fi, not by this site. Open on ko-fi.com