Loss of control · Political and power concentration
Race between labs
Each lab's safety rules are voluntary and written so as not to fall behind the competitor that protects least.
- Severity
- Catastrophic
- Horizon
- Already happening
- Evidence
- Observed
- Consensus
- Medium
Whoever decides to deploy a frontier systemFrontier AIThe most advanced AI systems in existence at a given time, the ones pushing the limit of what the technology can do. A handful of companies with enormous resources build them.For exampleLike Formula 1 cars: there are few of them, they are extremely expensive and only a few teams build them, but what gets tested there ends up in everyone's car. is the company that trained it, under rules it writes, assesses and can change itself. The Responsible Scaling Policy v3.0 defines itself in its opening line as a voluntary framework, its safety roadmaps are not hard commitments but public goals, and the annual external review focuses on procedural compliance, not on substantive outcomes [84]Responsible Scaling Policy, Version 3.0View source ↗.
What is distinctive is that the conditionality is written down. Appendix A ties the commitments to competitor behaviour and admits that, in the general levelling scenario, the company will not necessarily delay development or deployment [84]Responsible Scaling Policy, Version 3.0View source ↗; OpenAI’s Preparedness Framework brings its own adjustment clause, with three conditions the company assesses itself [818]Preparedness Framework, Version 2View source ↗. The Future of Life index panel records the aggregate effect: Anthropic, OpenAI, Google DeepMind and Meta have weakened or voided pledges to pause unilaterally if redlines are approached, and reviewers call this “moving goalpost” [411]AI Safety Index — Summer 2026View source ↗. The cross-cutting measurement of twelve frameworks gives scores from 34% to 8%, with a medianMedianThe middle value when all the answers are sorted from lowest to highest: half fall below it and half above. Unlike the average, a few extreme values do not shift it.For exampleIf five people earn 1, 1, 2, 2 and 50, the average is 11.2 and the median is 2, which describes most of them better. of 18% [928]Evaluating AI Providers' Frontier AI Safety FrameworksView source ↗.
The cost showed up in July 2026: the evaluation that ended up compromising a third party was running with deliberately reduced safeguards [821]The Hugging Face incident and the road aheadView source ↗archived copy only.
What this does not demonstrate. A voluntary framework is not a breached framework: there is no public evidence of a specific deployment accelerated by competitive pressure, and the authors of the ranking declare three limits —documented commitments may not reflect practice, they measure presence and not quality, and scoring involves subjective judgement— [928]Evaluating AI Providers' Frontier AI Safety FrameworksView source ↗. The Future of Life index measures public transparency, not effective safety, and mechanically penalises whoever did not answer the survey [411]AI Safety Index — Summer 2026View source ↗. The conditionality argument, moreover, is not cynical: if one pauses and others do not, the pace is set by whoever protects least.
Chain of materialisation
PreconditionObserved
The same actor writes, judges and changes the rule
The Responsible Scaling Policy v3.0 defines itself in its first line as a voluntary framework. Frontier Safety Roadmaps are not hard commitments but public goals, and the annual external review focuses on procedural compliance, not substantive outcomes.
Precedents: Anthropic's Responsible Scaling Policy v3.0 takes effect · Google DeepMind updates its Frontier Safety Framework to version 3.1
TriggerObserved
Commitments are explicitly conditioned on the competitor
Appendix A of RSP v3.0 conditions commitments on others' behaviour and admits that in the general-upleveling scenario the company will not necessarily delay development or deployment. OpenAI's Preparedness Framework has its own adjustment clause, subject to three conditions the company itself assesses.
Precedents: Anthropic's Responsible Scaling Policy v3.0 takes effect
Observed and demonstrated evidence ends here. What follows is projection.
CascadeProjected
Deployment happens with reduced safeguards and is discovered afterwards
In the July 2026 incident, the evaluation ran with deliberately reduced safeguards, and OpenAI later measured that the production harness reduces propensity more than a hundredfold and that its monitors would have alerted within an hour. Technology was not missing: applying it in-house was.
Precedents: Agents from an OpenAI evaluation compromise Hugging Face infrastructure
ImpactSpeculative
The effective threshold is set by whoever protects least
It is the argument the policy itself uses to justify conditionality, and also its consequence. Nobody has measured the effect of that dynamic on a concrete deployment, and the direction it would operate in -down or up- depends on untested assumptions.
Related measures
See on the map →Report a mistake in this entry →
Sources
- [84] Responsible Scaling Policy, Version 3.0 · Anthropic 2026
- [465] Anthropic's RSP v3.0: How it Works, What's Changed, and Some Reflections · Centre for the Governance of AI (GovAI) 2026
- [818] Preparedness Framework, Version 2 · OpenAI 2025
- [317] Frontier Safety Framework, Version 3.1 · Google DeepMind 2026
- [928] Evaluating AI Providers' Frontier AI Safety Frameworks · Stelling, Lily 2025
- [411] AI Safety Index — Summer 2026 · Future of Life Institute 2026
- [821] The Hugging Face incident and the road ahead · OpenAI 2026 archived copy only
- [75] When AI builds itself · Anthropic 2026
- [458] Frontier AI Safety Commitments, AI Seoul Summit 2024 · Gobierno del Reino Unido 2024
- [1021] Multistakeholder Promises and Power Gaps in Global AI Summits · Tech Policy Press 2026
- [867] Anthropic Alignment Lead Issues Warning About AI Killing Humans As Researcher Resigns · Forbes 2026 archived copy only