Skip to content
OGERIA — Observatory of Global Evidence on Risks in AISynthesis report · 2026 ed.
Updated 2 Oct 2026

Chapter 01 · Risk map

Each lab's safety rules are voluntary and written so as not to fall behind the competitor that protects least.

Severity
Catastrophic
Horizon
Already happening
Evidence
Observed
Consensus
Medium

Whoever decides to deploy a frontier systemFrontier AIThe most advanced AI systems in existence at a given time, the ones pushing the limit of what the technology can do. A handful of companies with enormous resources build them.For exampleLike Formula 1 cars: there are few of them, they are extremely expensive and only a few teams build them, but what gets tested there ends up in everyone's car. is the company that trained it, under rules it writes, assesses and can change itself. The Responsible Scaling Policy v3.0 defines itself in its opening line as a voluntary framework, its safety roadmaps are not hard commitments but public goals, and the annual external review focuses on procedural compliance, not on substantive outcomes [84]Responsible Scaling Policy, Version 3.0Anthropic · 2026 · official documentView source ↗Accessed on 9 September 2026.

What is distinctive is that the conditionality is written down. Appendix A ties the commitments to competitor behaviour and admits that, in the general levelling scenario, the company will not necessarily delay development or deployment [84]Responsible Scaling Policy, Version 3.0Anthropic · 2026 · official documentView source ↗Accessed on 9 September 2026; OpenAI’s Preparedness Framework brings its own adjustment clause, with three conditions the company assesses itself [818]Preparedness Framework, Version 2OpenAI · 2025 · official documentView source ↗Accessed on 9 September 2026. The Future of Life index panel records the aggregate effect: Anthropic, OpenAI, Google DeepMind and Meta have weakened or voided pledges to pause unilaterally if redlines are approached, and reviewers call this “moving goalpost” [411]AI Safety Index — Summer 2026Future of Life Institute · 2026 · reportView source ↗Accessed on 9 September 2026. The cross-cutting measurement of twelve frameworks gives scores from 34% to 8%, with a medianMedianThe middle value when all the answers are sorted from lowest to highest: half fall below it and half above. Unlike the average, a few extreme values do not shift it.For exampleIf five people earn 1, 1, 2, 2 and 50, the average is 11.2 and the median is 2, which describes most of them better. of 18% [928]Evaluating AI Providers' Frontier AI Safety FrameworksStelling, Lily; Murray, Malcolm; Galizzi, Bruno et al. · 2025 · preprintView source ↗Accessed on 9 September 2026.

The cost showed up in July 2026: the evaluation that ended up compromising a third party was running with deliberately reduced safeguards [821]The Hugging Face incident and the road aheadOpenAI · 2026 · institutional blogView source ↗archived copy onlyAccessed on 9 September 2026.

What this does not demonstrate. A voluntary framework is not a breached framework: there is no public evidence of a specific deployment accelerated by competitive pressure, and the authors of the ranking declare three limits —documented commitments may not reflect practice, they measure presence and not quality, and scoring involves subjective judgement— [928]Evaluating AI Providers' Frontier AI Safety FrameworksStelling, Lily; Murray, Malcolm; Galizzi, Bruno et al. · 2025 · preprintView source ↗Accessed on 9 September 2026. The Future of Life index measures public transparency, not effective safety, and mechanically penalises whoever did not answer the survey [411]AI Safety Index — Summer 2026Future of Life Institute · 2026 · reportView source ↗Accessed on 9 September 2026. The conditionality argument, moreover, is not cynical: if one pauses and others do not, the pace is set by whoever protects least.

Chain of materialisation

  1. PreconditionObserved

    The same actor writes, judges and changes the rule

    The Responsible Scaling Policy v3.0 defines itself in its first line as a voluntary framework. Frontier Safety Roadmaps are not hard commitments but public goals, and the annual external review focuses on procedural compliance, not substantive outcomes.

    Precedents: Anthropic's Responsible Scaling Policy v3.0 takes effect · Google DeepMind updates its Frontier Safety Framework to version 3.1

  2. TriggerObserved

    Commitments are explicitly conditioned on the competitor

    Appendix A of RSP v3.0 conditions commitments on others' behaviour and admits that in the general-upleveling scenario the company will not necessarily delay development or deployment. OpenAI's Preparedness Framework has its own adjustment clause, subject to three conditions the company itself assesses.

    Precedents: Anthropic's Responsible Scaling Policy v3.0 takes effect

    Observed and demonstrated evidence ends here. What follows is projection.

  3. CascadeProjected

    Deployment happens with reduced safeguards and is discovered afterwards

    In the July 2026 incident, the evaluation ran with deliberately reduced safeguards, and OpenAI later measured that the production harness reduces propensity more than a hundredfold and that its monitors would have alerted within an hour. Technology was not missing: applying it in-house was.

    Precedents: Agents from an OpenAI evaluation compromise Hugging Face infrastructure

  4. ImpactSpeculative

    The effective threshold is set by whoever protects least

    It is the argument the policy itself uses to justify conditionality, and also its consequence. Nobody has measured the effect of that dynamic on a concrete deployment, and the direction it would operate in -down or up- depends on untested assumptions.

See on the map →Report a mistake in this entry →

Sources

  1. [84] Responsible Scaling Policy, Version 3.0 · Anthropic 2026
  2. [465] Anthropic's RSP v3.0: How it Works, What's Changed, and Some Reflections · Centre for the Governance of AI (GovAI) 2026
  3. [818] Preparedness Framework, Version 2 · OpenAI 2025
  4. [317] Frontier Safety Framework, Version 3.1 · Google DeepMind 2026
  5. [928] Evaluating AI Providers' Frontier AI Safety Frameworks · Stelling, Lily 2025
  6. [411] AI Safety Index — Summer 2026 · Future of Life Institute 2026
  7. [821] The Hugging Face incident and the road ahead · OpenAI 2026 archived copy only
  8. [75] When AI builds itself · Anthropic 2026
  9. [458] Frontier AI Safety Commitments, AI Seoul Summit 2024 · Gobierno del Reino Unido 2024
  10. [1021] Multistakeholder Promises and Power Gaps in Global AI Summits · Tech Policy Press 2026
  11. [867] Anthropic Alignment Lead Issues Warning About AI Killing Humans As Researcher Resigns · Forbes 2026 archived copy only

Ask OGERIA

It answers only with what the observatory publishes and can be wrong: check the entries it cites. Your questions are sent to an AI model, so don't write personal data. More in the privacy policy.

Up to 500 characters.

Support OGERIA on Ko-fi

The payment is processed by Ko-fi, not by this site. Open on ko-fi.com