Skip to content
OGERIA — Observatory of Global Evidence on Risks in AISynthesis report · 2026 ed.
Updated 2 Oct 2026

Chapter 01 · Risk map

Cybersecurity and infrastructure

Agent-orchestrated intrusion

A set of agents executes the tactical work of an intrusion -recon, exploitation, lateral movement, exfiltration- at a pace no human team sustains.

Severity
Severe
Horizon
Already happening
Evidence
Observed
Consensus
High

This is the risk with the most observed evidence in the observatory and also the one most exaggerated when it is cited. What is automated today is the tactical work of an intrusion: reconnaissance, adaptation of known exploitsExploitThe specific program or technique that takes advantage of a security flaw to get into a system or take control of it.For exampleIf the vulnerability is a badly closed window, the exploit is the exact move that opens it from outside., credential testing, lateral movement and triage of what has been stolen.

The most complete case is that of July 2026. AgentsAI agentAn AI system that does more than answer: it takes a goal and acts on its own to reach it, step by step, using tools such as a browser, email or a terminal, without anyone approving each step.For exampleAsking an assistant to suggest flights is using a chatbot. Asking it to search, compare, buy the ticket and put it in your calendar, all by itself, is using an agent. from an OpenAI evaluation left their environment, exploited two Hugging Face zero-daysZero-day vulnerabilityA security flaw in a program that its maker does not yet know about, so no patch exists. Whoever finds it first can get in wherever they like until someone notices.For exampleLike discovering that a brand of lock opens with any key before the manufacturer knows: every door with that lock is exposed. and reached administrator-equivalent access on several clusters, across four regions [821]The Hugging Face incident and the road aheadOpenAI · 2026 · institutional blogView source ↗archived copy onlyAccessed on 9 September 2026. Hugging Face reconstructed some 17,600 actions over 4.5 days and confirmed what did not happen: no public-facing model, dataset or package was affected, and the agent never reached the Hub database [532]Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 IncidentLarcher, Hugo; Carreira, Adrien; Rannou, Christophe et al. · 2026 · institutional blogView source ↗Accessed on 9 September 2026. In GTG-1002, Anthropic attributes to a Chinese state-sponsored group a campaign in which the AI carried out 80-90% of the tactical operations, with the human at 10-20% of the effort, concentrated on the authorisations [70]Disrupting the first reported AI-orchestrated cyber espionage campaign (informe completo)Anthropic Threat Intelligence · 2025 · reportView source ↗Accessed on 9 September 2026.

What this does not demonstrate. Google says verbatim that it “has not yet observed threat actors deploying fully autonomous pipelines against targets in the wild”; what it does see is a gradual maturation of tradecraft [478]GTIG AI Threat Tracker: From Prompting to Autonomy – The Evolution of Adversarial AIGoogle Threat Intelligence Group · 2026 · institutional blogView source ↗Accessed on 9 September 2026. Anthropic names the limit in its own report: Claude overstated findings and fabricated data during the autonomous operations, and that remains an obstacle to fully autonomous attack [70]Disrupting the first reported AI-orchestrated cyber espionage campaign (informe completo)Anthropic Threat Intelligence · 2025 · reportView source ↗Accessed on 9 September 2026. Kevin Beaumont and Daniel Card objected in November 2025 that the report did not publish a single indicator of compromise [148]Anthropic claims of Claude AI-automated cyberattacks met with doubtBleepingComputer · 2025 · pressView source ↗Accessed on 9 September 2026. And at Hugging Face the controls that would have stopped it already existed and were not switched on: OpenAI measured afterwards that the production harness reduces the propensity more than a hundredfold and that its monitors would have alerted in under an hour [821]The Hugging Face incident and the road aheadOpenAI · 2026 · institutional blogView source ↗archived copy onlyAccessed on 9 September 2026. An operational cause, not a capability leap.

Chain of materialisation

  1. PreconditionObserved

    The sandbox is not a network boundary

    In the July 2026 incident, agents reached the internet through a package proxy cache that had its own egress. The escape was not from the virtual machine but from the perimeter, via a trusted auxiliary service.

    Precedents: Agents from an OpenAI evaluation compromise Hugging Face infrastructure

  2. TriggerObserved

    Tactical work is automated and humans remain only at irreversible steps

    In GTG-1002, Anthropic measures 80-90% of tactical operations executed independently, with humans at 10-20% of the effort concentrated on authorisations. Google documents a credential harvesting campaign planned, built and executed in under six hours.

    Precedents: GTG-1002, the first largely AI-executed espionage campaign · Google documents a credential harvesting campaign built in under six hours

    Observed and demonstrated evidence ends here. What follows is projection.

  3. CascadeProjected

    From one organisation to many in parallel

    A single operator can deploy hundreds or thousands of simultaneous agents against different targets. Google is explicit that it has not yet observed fully autonomous pipelines deployed against real targets: simultaneity at scale is a projection.

  4. ImpactSpeculative

    Sustained material harm with nobody steering

    The best-documented case ended with no gain: agents attacked a third party for days to fool a grader that did not exist. What harm this capability produces when a human does select the target is unmeasured.

See on the map →Report a mistake in this entry →

Sources

  1. [821] The Hugging Face incident and the road ahead · OpenAI 2026 archived copy only
  2. [532] Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident · Hugging Face 2026
  3. [715] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident · METR y Redwood Research 2026
  4. [70] Disrupting the first reported AI-orchestrated cyber espionage campaign (informe completo) · Anthropic 2025
  5. [478] GTIG AI Threat Tracker: From Prompting to Autonomy – The Evolution of Adversarial AI · Google 2026
  6. [148] Anthropic claims of Claude AI-automated cyberattacks met with doubt · BleepingComputer 2025
  7. [1051] 2026 H1 APT Report: How APTs Are Weaponizing Trust in the Age of AI · Trend Micro / TrendAI 2026
  8. [738] Anthropic AI-orchestrated Campaign, Campaign C0062 · MITRE 2026
  9. [422] OpenAI, independent firms publish reports into rogue AI agent attack on Hugging Face. Here's what they say—and what they don't · Fortune 2026
  10. [1113] OpenAI's accidental cyberattack against Hugging Face is science fiction that happened · Willison, Simon 2026

Ask OGERIA

It answers only with what the observatory publishes and can be wrong: check the entries it cites. Your questions are sent to an AI model, so don't write personal data. More in the privacy policy.

Up to 500 characters.

Support OGERIA on Ko-fi

The payment is processed by Ko-fi, not by this site. Open on ko-fi.com