Skip to content
OGERIA — Observatory of Global Evidence on Risks in AISynthesis report · 2026 ed.
Updated 2 Oct 2026

Chapter 06 · Protection

Demanding that evaluations be published

That dangerous-capability thresholds, adversarial testing and their results be public and auditable. Twenty companies committed to publishing a frontier safety framework and twelve did: the gap between those two numbers is the indicator that matters.

Level
Civic
Cost
Low cost
Effort
Hours
Evidence
Observed

What it does not solve

A published framework measures the promise, not compliance, and independent evaluation of those frameworks finds very uneven quality. Nor does it fix that whoever writes the threshold is whoever decides if it was crossed, and it does not cover developers who never committed to anything.

It works if what you want is for someone on the outside to be able to contradict the developer. Today the best public window onto the real state of dangerous capabilities is the voluntary frameworks and their threshold activations, and that window is smaller than it looks: at the Seoul summit of May 2024, sixteen companies committed to publishing a frontier safety framework and another four joined afterwards; twelve published one [708]Common Elements of Frontier AI Safety PoliciesMETR · 2025 · reportView source ↗Accessed on 9 September 2026. The gap between committed and published is a better indicator than the count, because it measures failure to meet an explicit commitment and admits no optimistic reading.

It does not work if you stop at the count. The independent assessment of those frameworks finds very uneven quality across providers [928]Evaluating AI Providers' Frontier AI Safety FrameworksStelling, Lily; Murray, Malcolm; Galizzi, Bruno et al. · 2025 · preprintView source ↗Accessed on 9 September 2026, and a published framework measures the promise, not compliance.

Evidence. Observed, in the precise sense that it exists and can be counted. In the European Union part of this is already a legal obligation for models with systemic risk, including adversarial testing [22]EU AI Act — Article 55: Obligations for providers of general-purpose AI models with systemic riskUnión Europea · 2026 · official documentView source ↗Accessed on 9 September 2026.

Cost. Low for one person: demanding it, citing it, preferring those who publish.

What it does NOT solve. The underlying problem, which is that whoever writes the threshold is whoever decides whether it has been crossed. Nor does it cover developers who never committed to anything, or stop a framework from being relaxed just when it starts to get in the way.

Works if…

  • Catastrophe through misuse · Works

    Threshold activations and their justifications are today the best public window into the real state of biological uplift.

  • Rapid loss of control · Partial
  • Power grab by a small group · Partial

    It attacks head-on the opacity assumption about internal capabilities and uses that scenario depends on.

  • AI as normal technology · Works

Does not work if…

  • The intelligence curse · Does not work

See in the protection matrix →Report a mistake in this entry →

Sources

  1. [708] Common Elements of Frontier AI Safety Policies · METR 2025
  2. [928] Evaluating AI Providers' Frontier AI Safety Frameworks · Stelling, Lily 2025
  3. [22] EU AI Act — Article 55: Obligations for providers of general-purpose AI models with systemic risk · Unión Europea 2026

Ask OGERIA

It answers only with what the observatory publishes and can be wrong: check the entries it cites. Your questions are sent to an AI model, so don't write personal data. More in the privacy policy.

Up to 500 characters.

Support OGERIA on Ko-fi

The payment is processed by Ko-fi, not by this site. Open on ko-fi.com