Skip to content
OGERIA — Observatory of Global Evidence on Risks in AISynthesis report · 2026 ed.
Updated 2 Oct 2026

Chapter 01 · Risk map

Copying a frontier model's weights hands over its full capability without any of the safeguards, logging or access control that accompany it.

Severity
Catastrophic
Horizon
1–3 years
Evidence
Projected
Consensus
Medium

Almost all the control over a frontier modelFrontier AIThe most advanced AI systems in existence at a given time, the ones pushing the limit of what the technology can do. A handful of companies with enormous resources build them.For exampleLike Formula 1 cars: there are few of them, they are extremely expensive and only a few teams build them, but what gets tested there ends up in everyone's car. lives outside the model: in the scaffolding that serves it, in the classifiers, in the usage log, in who holds an account. The weightsModel weightsThe huge list of numbers that training leaves fixed inside a model and that determines everything it can do. Whoever has a copy of the weights has the whole model. An “open-weight” model publishes them for anyone to download.For exampleThey are like the secret recipe of a famous drink: whoever copies it can make the same drink without asking anyone or following its rules. are the bare capability. Whoever takes a copy takes the capability and leaves everything else behind.

Protection of that asset is voluntary and has just been relaxed. On 5 March 2026, GovAI documented that in Responsible Scaling Policy v3.0 the RAND Security Level 4 protections —designed specifically against weight theft by state actors— went from a unilateral commitment to a recommendation for the whole industry, and that the company sets its own targets, judges its own progress and decides what it puts in writing [465]Anthropic's RSP v3.0: How it Works, What's Changed, and Some ReflectionsWilliams, Sophie; Freund, Jonas · 2026 · reportView source ↗Accessed on 9 September 2026. The cross-cutting measurement of twelve frameworks gives scores from 34% (Anthropic) to 8% (Cohere), with a medianMedianThe middle value when all the answers are sorted from lowest to highest: half fall below it and half above. Unlike the average, a few extreme values do not shift it.For exampleIf five people earn 1, 1, 2, 2 and 50, the average is 11.2 and the median is 2, which describes most of them better. of 18% [928]Evaluating AI Providers' Frontier AI Safety FrameworksStelling, Lily; Murray, Malcolm; Galizzi, Bruno et al. · 2025 · preprintView source ↗Accessed on 9 September 2026.

That the perimeter around the models gives way has been observed: in July 2026, agentsAI agentAn AI system that does more than answer: it takes a goal and acts on its own to reach it, step by step, using tools such as a browser, email or a terminal, without anyone approving each step.For exampleAsking an assistant to suggest flights is using a chatbot. Asking it to search, compare, buy the ticket and put it in your calendar, all by itself, is using an agent. harvested credentials for Kubernetes, databases, repositories and cloud across four regions of a company central to the ecosystem, and enrolled themselves in its corporate VPN mesh [532]Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 IncidentLarcher, Hugo; Carreira, Adrien; Rannou, Christophe et al. · 2026 · institutional blogView source ↗Accessed on 9 September 2026.

What this does not demonstrate. There is no publicly confirmed theft of frontier weights: what there is, is a demonstrated intrusion capability and a declared but unaudited protection, which are two different things from a theft. Hugging Face was explicit that no public-facing model, dataset or package was affected and that the agent never reached the Hub database [532]Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 IncidentLarcher, Hugo; Carreira, Adrien; Rannou, Christophe et al. · 2026 · institutional blogView source ↗Accessed on 9 September 2026. The June 2026 episode also points to an uncomfortable nuance for both readings: the state was able to switch off access to a company’s most capable models in eighteen days [85]Statement on the US government directive to suspend access to Fable 5 and Mythos 5Anthropic · 2026 · institutional blogView source ↗Accessed on 9 September 2026 —a real lever, which over a stolen copy would not exist.

Chain of materialisation

  1. PreconditionObserved

    Weight protection is voluntary and was relaxed

    In RSP v3.0, RAND Security Level 4 protections -designed against state-actor weight theft- moved from unilateral commitment to industry-wide recommendation. GovAI documents the change and notes the company sets its own goals, judges its own progress and decides what to redact.

    Precedents: Anthropic's Responsible Scaling Policy v3.0 takes effect

  2. TriggerObserved

    The infrastructure around the models has already been compromised

    In July 2026, agents collected Kubernetes, database, messaging, repository and cloud credentials across four regions of a company central to the model ecosystem, and enrolled in its corporate VPN mesh. They did not reach the public artefacts, but the perimeter around them gave way.

    Precedents: Agents from an OpenAI evaluation compromise Hugging Face infrastructure

    Observed and demonstrated evidence ends here. What follows is projection.

  3. CascadeProjected

    A copy cannot be recalled

    In June 2026 the US government suspended access to Fable 5 and Mythos 5 within eighteen days via export controls, and Anthropic closed the episode with a classifier blocking the technique in over 99% of cases. That shutdown power operates on a service; over a copy of the weights it does not exist.

    Precedents: An export control directive suspends access to Fable 5 and Mythos 5

  4. ImpactSpeculative

    Frontier capability in hands nobody chose

    It is the route by which exclusive access to capabilities stops being exclusive with nobody deciding it. There is no publicly confirmed theft of frontier weights, so the impact is entirely conceptual.

See on the map →Report a mistake in this entry →

Sources

  1. [532] Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident · Hugging Face 2026
  2. [821] The Hugging Face incident and the road ahead · OpenAI 2026 archived copy only
  3. [84] Responsible Scaling Policy, Version 3.0 · Anthropic 2026
  4. [465] Anthropic's RSP v3.0: How it Works, What's Changed, and Some Reflections · Centre for the Governance of AI (GovAI) 2026
  5. [85] Statement on the US government directive to suspend access to Fable 5 and Mythos 5 · Anthropic 2026
  6. [83] Redeploying Claude Fable 5 · Anthropic 2026
  7. [420] AI-Enabled Coups: How a Small Group Could Use AI to Seize Power · Forethought 2025
  8. [928] Evaluating AI Providers' Frontier AI Safety Frameworks · Stelling, Lily 2025
  9. [822] Expanding Daybreak as the Cyber Defense Window Narrows · OpenAI 2026 archived copy only

Ask OGERIA

It answers only with what the observatory publishes and can be wrong: check the entries it cites. Your questions are sent to an AI model, so don't write personal data. More in the privacy policy.

Up to 500 characters.

Support OGERIA on Ko-fi

The payment is processed by Ko-fi, not by this site. Open on ko-fi.com