Demanding that evaluations be published
That dangerous-capability thresholds, adversarial testing and their results be public and auditable. Twenty companies committed to publishing a frontier safety framework and twelve did: the gap between those two numbers is the indicator that matters.
- Level
- Civic
- Cost
- Low cost
- Effort
- Hours
- Evidence
- Observed
What it does not solve
A published framework measures the promise, not compliance, and independent evaluation of those frameworks finds very uneven quality. Nor does it fix that whoever writes the threshold is whoever decides if it was crossed, and it does not cover developers who never committed to anything.
It works if what you want is for someone on the outside to be able to contradict the developer. Today the best public window onto the real state of dangerous capabilities is the voluntary frameworks and their threshold activations, and that window is smaller than it looks: at the Seoul summit of May 2024, sixteen companies committed to publishing a frontier safety framework and another four joined afterwards; twelve published one [708]Common Elements of Frontier AI Safety PoliciesView source ↗. The gap between committed and published is a better indicator than the count, because it measures failure to meet an explicit commitment and admits no optimistic reading.
It does not work if you stop at the count. The independent assessment of those frameworks finds very uneven quality across providers [928]Evaluating AI Providers' Frontier AI Safety FrameworksView source ↗, and a published framework measures the promise, not compliance.
Evidence. Observed, in the precise sense that it exists and can be counted. In the European Union part of this is already a legal obligation for models with systemic risk, including adversarial testing [22]EU AI Act — Article 55: Obligations for providers of general-purpose AI models with systemic riskView source ↗.
Cost. Low for one person: demanding it, citing it, preferring those who publish.
What it does NOT solve. The underlying problem, which is that whoever writes the threshold is whoever decides whether it has been crossed. Nor does it cover developers who never committed to anything, or stop a framework from being relaxed just when it starts to get in the way.
Works if…
- Catastrophe through misuse · Works
Threshold activations and their justifications are today the best public window into the real state of biological uplift.
- Rapid loss of control · Partial
- Power grab by a small group · Partial
It attacks head-on the opacity assumption about internal capabilities and uses that scenario depends on.
- AI as normal technology · Works
Does not work if…
- The intelligence curse · Does not work
See in the protection matrix →Report a mistake in this entry →
Sources
- [708] Common Elements of Frontier AI Safety Policies · METR 2025
- [928] Evaluating AI Providers' Frontier AI Safety Frameworks · Stelling, Lily 2025
- [22] EU AI Act — Article 55: Obligations for providers of general-purpose AI models with systemic risk · Unión Europea 2026