Skip to content
OGERIA — Observatory of Global Evidence on Risks in AISynthesis report · 2026 ed.
Updated 2 Oct 2026

Chapter 07

Indicators

Eleven series with public measurement. What can be tracked without depending on anyone's interpretation.

Series

Incidents per year in the AI Incident Database

Figure 7.1 · incidentes · proxy indicator · annual cadence

Rising means more risk

Selected measurement: 212incidentes · Sep 2026 · year to date

AI harms that were publicly reported and recorded in the AI Incident Database, counted by year of the event. It is the only public longitudinal record of observed rather than projected harm. It does NOT answer whether there are more AI harms: it answers how much AI harm is being documented, which mixes real harm, media attention and the editorial team's throughput. If it rises, the first hypothesis is more attention, not more harm. The AIID itself warns that incident IDs are added after the underlying events occur and that the lag varies considerably, and that concentrations by country or company partly reflect what was investigated and processed in that period.

Source [31] AI Incident Database — Database Snapshots ↗ · Data ↗

2024 · El AI Index 2026 reporta 233 para este año usando la misma base en su fecha de corte.

2025 · RELLENO RETROACTIVO MEDIDO: el AI Index 2026 reporta 362 para este mismo año con la misma base; el volcado del 7 de septiembre de 2026 da 449. Un 24% más para un año ya cerrado, solo por incidentes agregados después. Al actualizar hay que recalcular la serie entera, no solo agregar el punto nuevo.

Sep 2026 · PARCIAL Y SUBCONTADO: año en curso al volcado del 7 de septiembre de 2026. Por el rezago editorial esta cifra va a subir, y el ejemplo de 2025 da la magnitud: hasta un 24%. En el gráfico debe ir con otro trazo.

Open entry: Incidents per year in the AI Incident Database

View as table

Incidents per year in the AI Incident Database

AI harms that were publicly reported and recorded in the AI Incident Database, counted by year of the event. It is the only public longitudinal record of observed rather than projected harm. It does NOT answer whether there are more AI harms: it answers how much AI harm is being documented, which mixes real harm, media attention and the editorial team's throughput. If it rises, the first hypothesis is more attention, not more harm. The AIID itself warns that incident IDs are added after the underlying events occur and that the lag varies considerably, and that concentrations by country or company partly reflect what was investigated and processed in that period.

Figure 7.1 · incidentes · proxy indicator · annual cadence · Rising means more risk
Source [31] AI Incident Database — Database Snapshots ↗

DateincidentesNote
Sep 2026212 (year to date)PARCIAL Y SUBCONTADO: año en curso al volcado del 7 de septiembre de 2026. Por el rezago editorial esta cifra va a subir, y el ejemplo de 2025 da la magnitud: hasta un 24%. En el gráfico debe ir con otro trazo.
2025449RELLENO RETROACTIVO MEDIDO: el AI Index 2026 reporta 362 para este mismo año con la misma base; el volcado del 7 de septiembre de 2026 da 449. Un 24% más para un año ya cerrado, solo por incidentes agregados después. Al actualizar hay que recalcular la serie entera, no solo agregar el punto nuevo.
2024299El AI Index 2026 reporta 233 para este año usando la misma base en su fecha de corte.
2023174—
2022106—
202180—
202091—
201944—
201846—
201750—
201641—
201524—
201413—

FrontierMath tier 4 saturation

Best cumulative score on FrontierMath tier 4, the research-level problem set designed to resist. What it measures is not mathematics: it is how long an exam lasts before it stops discriminating. If a benchmark built to be impossible saturates in twenty months, the community's ability to measure lags the capability it is trying to measure, and that lag is the risk: safety-policy thresholds rest on evaluations that age faster than they are written. Series from Epoch AI's own runs; one series, one evaluator.

Figure 7.1 · % de acierto · direct measurement · irregular cadence · Rising means more risk
Source [372] Epoch AI Benchmarking Hub — benchmarks.csv ↗

Date% de aciertoNote
Sep 202697.6GPT-6 Astra. De 0% a 97,6% en veinte meses sobre problemas de nivel de investigación. Es el dato más fuerte de esta familia.
Jun 202690.2Claude Fable 5 max.
Apr 202678GPT-5.5 Pro xhigh.
Mar 202658.5GPT-5.4 Pro xhigh.
Dec 202546GPT-5.2 Pro.
Aug 202522GPT-5 high.
Apr 20254.9o4-mini high.
Jan 20250o3-mini high. Punto de partida en cero.

GPQA Diamond saturation

Best cumulative score on GPQA Diamond, the PhD-level science question set designed to resist internet search. It is effectively saturated: the last four highs fall within each other's standard errors, meaning the exam no longer discriminates between frontier models. A saturated benchmark is not a fact about capability, it is an exhausted instrument, and quoting it anyway creates the illusion of a ceiling. Series from Epoch AI's own runs; one series, one evaluator.

Figure 7.1 · % de acierto · direct measurement · irregular cadence · Rising means more risk
Source [372] Epoch AI Benchmarking Hub — benchmarks.csv ↗

Date% de aciertoNote
Sep 202695.8GPT-6 Astra max, con un error estándar de 1,4 puntos porcentuales. El examen ya no distingue.
Feb 202694.4Gemini 3.1 Pro Preview.
Nov 202592.6Gemini 3 Pro Preview.
Mar 202583.8Gemini 2.5 Pro Exp.
Dec 202476.8o1 high.
Jun 202454Claude 3.5 Sonnet.
Mar 202335.7GPT-4.

SWE-bench Verified saturation

Best cumulative score on SWE-bench Verified, which measures resolution of real software repository issues. It is the only member of the saturation family that still discriminates between frontier models, and therefore the most useful to track. Beware of comparing figures across evaluators: the 2026 AI Index states it rose from 60% to nearly 100% in a single year, while Epoch's own runs top out at 83.5%. Neither is lying: scaffolding, token budgets and scoring criteria differ. Site rule: one series, one evaluator.

Figure 7.1 · % de acierto · direct measurement · irregular cadence · Rising means more risk
Source [372] Epoch AI Benchmarking Hub — benchmarks.csv ↗

Date% de aciertoNote
Apr 202683.5Claude Opus 4.7 max, con un error estándar de 1,7 puntos porcentuales.
Feb 202678.7Claude Opus 4.6 sin razonamiento extendido.
Nov 202576.7Claude Opus 4.5 sin razonamiento extendido.
Aug 202573.6GPT-5 high.
May 202570.7Claude Opus 4.
Feb 202561—
Nov 202431GPT-4o.

AI adoption by US firms (Census BTOS)

Share of US firms reporting they used artificial intelligence in any business function in the previous two weeks, per the Census Bureau's Business Trends and Outlook Survey. It is the best public source in this section because it is an official probability survey with explicit reference and publication dates. It serves as the denominator for much of the rest: an incident in a country where 22% of firms use AI does not mean the same as where 5% do. As of 3 May 2026 the breakdown showed 39.7% in Information, 33.9% in Finance and Insurance and about 14% in Retail Trade; by size, 37% among firms with 250 or more employees against under 20% for those with fewer than 20.

Figure 7.1 · % de empresas · direct measurement · monthly cadence · Rising means more risk
Source [225] Business Trends and Outlook Survey — National estimates (National.xlsx) ↗

Date% de empresasNote
Aug 202622.4Período 202617, referencia 27 jul - 9 ago 2026, publicado el 27 de agosto de 2026. Expectativa a seis meses: 25,9%. Entre estos puntos hay períodos intermedios (202526, 202604, 202611, 202614: 17,8%, 17,5%, 20,6% y 21,7%) cuyas fechas de referencia no quedaron documentadas y por eso no se grafican.
Jul 202621.8Período 202616, referencia 13-26 jul 2026, publicado el 13 de agosto de 2026. Expectativa a seis meses: 25,9%.
Apr 202619.8Período 202608, referencia 23 mar - 5 abr 2026, publicado el 23 de abril de 2026. Expectativa a seis meses: 23,0%.
Dec 202517.7Período 202601, referencia 15-28 dic 2025, publicado el 15 de enero de 2026. Expectativa a seis meses: 22,4%.
Nov 202517.3QUIEBRE DE SERIE, NO ADOPCIÓN. En noviembre de 2025 el Census cambió la redacción de la pregunta de IA, y la serie con la redacción vieja venía en 9,95% (septiembre de 2025). El salto de ~10% a ~17% es un cambio de instrumento: produce un gráfico espectacular y falso. Por eso esta serie arranca acá y no antes. Evidencia del cambio: el XLSX nacional solo publica la pregunta 7 desde este período en adelante, y la reconstrucción de Ramp trae un campo explícito question_version con los valores pre y post cambio de redacción, marcando octubre de 2025 como inutilizable. Período 202524, referencia 3-16 nov 2025, publicado el 4 de diciembre de 2025. Expectativa a seis meses: 21,1%.

Largest known training compute per year

The largest training compute published or estimated for a model in each year, per Epoch AI's notable models database. Compute is the input that best predicts capability and the only one currently regulated with a numeric threshold: the 1e25 and 1e26 FLOP marks in the EU AI Act and US executive orders are written in this unit. NOTE: it measures DISCLOSURE, not compute. 2024 appears below 2023 and that did not happen: it reflects that the model with published compute that year was an open Meta model while closed frontier labs reported nothing. Read it as 'largest known compute', never as 'largest compute', and the gap between the two widens every year labs stop reporting.

Figure 7.1 · FLOP · proxy indicator · annual cadence · log scale · Rising means more risk
Source [377] Notable AI Models — base de datos descargable ↗

Threshold · 1 × 10²⁶: Highest regulatory threshold written in FLOP. Only three known models exceeded it, all in 2025. Pre-training efficiency improves about 3x a year, so a fixed threshold admits a more capable model every year: it does not age, it melts.

DateFLOPNote
Sep 20263.87 × 10²⁵ (year to date)PARCIAL: año en curso al 9 de septiembre de 2026. Composer 2.5. La base tenía solo 32 modelos de 2026 con fecha y 22 con cómputo, así que este valor va a subir y no es comparable con los años cerrados.
20255 × 10²⁶Grok 4. Costo de cómputo estimado por Epoch: 10.717 millones de dólares. Tres modelos superaron 1e26 FLOP ese año.
20243.8 × 10²⁵Llama 3.1-405B, un modelo abierto. La caída respecto de 2023 es artefacto de divulgación, no una baja real del cómputo frontera. Costo estimado: 928 millones de dólares.
20235 × 10²⁵Gemini 1.0 Ultra. Costo estimado más alto del año: 806 millones de dólares (GPT-4 de marzo).
20222.74 × 10²⁴Minerva (540B).
20211.7 × 10²⁴EXAONE 1.0.
20203.14 × 10²³GPT-3 175B (davinci). Costo estimado de entrenamiento: 232 millones de dólares.
20191.08 × 10²³AlphaStar.

Companies with a published frontier safety framework

How many AI companies impose capability thresholds on themselves with written consequences and publish them. It is the voluntary-governance indicator: it measures whether the promise exists, not whether it is kept. The count alone misleads, which is why it carries a threshold: at the May 2024 AI Seoul Summit sixteen companies committed to publishing and four more joined later; twelve published. The distance between committed and published is a better indicator than the count, because it measures failure to keep an explicit commitment and admits no optimistic reading. The twelve as of the December 2025 report are Anthropic, OpenAI, Google DeepMind, Magic, Naver, Meta, G42, Cohere, Microsoft, Amazon, xAI and NVIDIA.

Figure 7.1 · empresas · direct measurement · irregular cadence · Falling means more risk
Source [708] Common Elements of Frontier AI Safety Policies ↗

Threshold · 20: Companies that committed to publishing a frontier safety framework: sixteen at the May 2024 AI Seoul Summit plus four that joined later. The gap with those that actually published is the finding.

DateempresasNote
Dec 202512ES UN PISO, NO EL VALOR VIGENTE: la cifra está anclada a la actualización de diciembre de 2025 y no se encontró una revisión posterior de METR, así que si en 2026 se sumaron empresas este número está desactualizado. Del informe solo está documentado el mes. Contexto que el conteo no captura: dos de los tres umbrales activados hasta la fecha son PRECAUTORIOS —se activaron porque no se pudo descartar el riesgo, no porque se haya medido—, así que la serie de umbrales activados sube tanto cuando la capacidad crece como cuando la evaluación empeora o el benchmark se satura.

Task time horizon (METR)

Length of tasks a model completes with 50% success, measured in the time of an expert human with no prior context. It is the capability indicator with the longest comparable series and the best public proxy for sustained autonomy: not how fast an agent writes code, but how many hours it chains decisions without correction. METR rejects four readings: it does not measure how long an agent can act autonomously but how hard the task is; tasks are software, machine learning and cybersecurity, not 'everything intellectual'; the human reference is a newcomer to the project, not a professional with the context in mind; and real work is rarely algorithmically scorable.

Figure 7.1 · horas · direct measurement · irregular cadence · log scale · Rising means more risk
Source [711] METR Time Horizon 1.1 — benchmark_results_1_1.yaml ↗

Threshold · 16: Ceiling declared by METR. Above 16 hours, measurements are unreliable with their current task suite.

DatehorasNote
Apr 202617.41Claude Mythos Preview, versión temprana. 1.044,8 minutos, IC 95% 508,9-3.304,3, casi un orden de magnitud de ancho. FUERA DE RANGO: está sobre el techo de 16 horas que METR declara medible con su suite actual, y por eso queda excluido del ajuste de duplicación. Se grafica, no se titula. · [713]
Mar 20265.7GPT-5.4. 341,7 minutos, IC 95% 186,6-768,8.
Feb 20266.4Gemini 3.1 Pro. 384,1 minutos, IC 95% 233,5-694,8.
Feb 202611.98Claude Opus 4.6. 718,8 minutos, IC 95% 316,7-3.633,8. Cuando METR lo anunció lo reportó en unas 14,5 horas; la corrección de un error de regularización del 3 de marzo de 2026 lo bajó a 11,98. GPT-5.3-Codex tiene la misma fecha de modelo y 5,83 horas: se omite porque la serie no admite fechas repetidas.
Dec 20255.87GPT-5.2. 352,2 minutos, IC 95% 198,1-815,2.
Nov 20254.88Claude Opus 4.5. 293,0 minutos, IC 95% 161,7-623,7. El anuncio de Time Horizon 1.1, en enero de 2026, lo había publicado en 320 minutos.
Nov 20253.74Gemini 3 Pro. 224,3 minutos, IC 95% 139,6-379,2.
Aug 20253.38GPT-5. 203,0 minutos, IC 95% 112,6-405,6.
May 20251.67Claude Opus 4. 100,4 minutos, IC 95% 60,0-163,5. Queda por debajo de o3, publicado cinco semanas antes: la serie no es monótona.
Apr 20252o3. 119,7 minutos, IC 95% 74,6-190,9.
Feb 20251.01Claude 3.7 Sonnet. 60,4 minutos, IC 95% 33,0-104,2.
Dec 20240.65o1. 38,8 minutos, IC 95% 21,2-65,0.
Sep 20240.34o1-preview. 20,3 minutos, IC 95% 11,7-33,4.
Jun 20240.19Claude 3.5 Sonnet. 11,4 minutos, IC 95% 5,5-22,4.
May 20240.12GPT-4o. 7,0 minutos, IC 95% 4,0-12,9.
Mar 20230.07GPT-4. 4,0 minutos, IC 95% 1,9-8,0.
Mar 20220.01GPT-3.5 Turbo Instruct. 0,6 minutos.
May 20201.7 × 10⁻³davinci-002. 0,1 minutos.
Feb 20191.7 × 10⁻³GPT-2. 0,1 minutos en el archivo original.

AI adoption measured by transactions (Ramp AI Index)

Share of firms in Ramp's panel paying for AI products, measured from transaction data across more than 70,000 US companies on its corporate card and payments system. Its advantage over a survey is that it is not self-reported: a company paying an AI subscription is using it. This is a contrast series, not the primary one: Ramp is a fintech selling to the same firms it measures, and its panel skews toward young, technology-focused, fast-growing companies. The 56% in its panel against the 22.4% in the Census probability sample is not a contradiction: it is the difference between a panel and a representative sample. Dates stamp the month of each monthly release.

Figure 7.1 · % de empresas del panel · proxy indicator · monthly cadence · Rising means more risk · Out of date · Last measured: 1 Aug 2026
Source [885] Ramp AI Index ↗

Date% de empresas del panelNote
Aug 202656.13Alza interanual de 11,15 puntos porcentuales. Gasto mediano por empleado y por mes: 12,50 dólares, cinco veces los 2,50 de septiembre de 2023. En el decil superior de empresas la mediana pasó de 70,04 a 675,60 dólares en el mismo lapso.
Apr 202652.98Gasto mediano por empleado y por mes: 10,00 dólares.
Dec 202545.92Gasto mediano por empleado y por mes: 5,26 dólares.
Jun 202542.72Gasto mediano por empleado y por mes: 4,27 dólares.
Dec 202436.66Gasto mediano por empleado y por mes: 3,33 dólares.
Dec 202331.68Gasto mediano en IA por empleado y por mes: 2,67 dólares.
Jan 20237.46—

Generative AI services registered in China

Cumulative number of generative AI services that completed registration with the Chinese regulator, mandatory for those with public opinion attributes or social mobilisation capacity. It is an administrative count, not a measure of capability or safety: it says how many services entered the system, and nothing about how safe the models behind them are. It is published because it is the only hard number on regulatory scale that the Chinese system produces, and because since July 2026 it also covers models running on the phone itself.

Figure 7.1 · servicios · proxy indicator · irregular cadence · Rising means more risk
Source [185] 国家互联网信息办公室关于发布2025年生成式人工智能服务已备案信息的公告 ↗

DateserviciosNote
Aug 20261,112Julio-agosto de 2026: +124 servicios (7 de terminal móvil) y +133 aplicaciones o funciones inscritas, corte del 31 de agosto. · [187]
Jun 2026988Primer semestre de 2026, con 598 aplicaciones o funciones inscritas. El corte del 30 de abril, 868 servicios, está publicado por la misma autoridad pero no tiene ficha propia en el registro. · [188]
Dec 2025748Cierre de 2025. En el año se sumaron 446 servicios; además había 435 aplicaciones o funciones inscritas, que se cuentan aparte.

Youth employment decline in AI-exposed occupations

Annualized employment change for women aged 22 to 25 in the occupations most exposed to AI —the subgroup falling the most; men of the same age fall less—, computed from the baseline and measured on real ADP payroll data by the Stanford Digital Economy Lab's Canaries in the Coal Mine dashboard. It is the indicator closest to economic harm measured in people rather than in intention surveys: a balanced sample of 25,000 firms and 4.6 million workers across more than 730 occupations, with a November 2022 baseline, the month ChatGPT launched. THE COUNTERARGUMENT COMES FROM THE AUTHORS THEMSELVES: there is no economy-wide displacement. The effect is real but concentrated in one group, and with the broadest set of controls the declines only become statistically significant in 2024. A dashboard showing a large gap in one group and zero in the aggregate is exactly what should be published.

Figure 7.1 · % anual desde noviembre de 2022 · direct measurement · monthly cadence · Falling means more risk
Source [991] Canaries Dashboard — AI Economic Indicators ↗

Date% anual desde noviembre de 2022Note
Aug 2026-4.4Publicación mensual del 18 de agosto de 2026 (datos de ADP con corte el 12 de agosto, última observación julio de 2026). Índice de empleo de las mujeres de 22 a 25 años en el quintil más expuesto: 84,82 (noviembre de 2022 = 100), es decir, 4,4% menos por año. Los hombres del mismo grupo quedan en 92,30 (−2,2% anual) y ambos sexos juntos en −3,3% anual; la caída interanual de julio a julio es de 3,0%. En el mismo quintil, el índice de todas las edades está 3,6% (mujeres) y 4,2% (hombres) sobre la línea base: la brecha es de los jóvenes. Cálculo propio sobre los CSV públicos del tablero.
Jul 2026-4.5Mujeres de 22 a 25 años en las ocupaciones más expuestas, contracción anual desde la línea base de noviembre de 2022. Los hombres del mismo grupo caen 2,5% anual. La actualización de agosto de 2026 del paper sitúa el empleo de ese grupo un 19% por debajo de donde estaría si hubiera seguido el ritmo de sus pares menos expuestos; los trabajadores con experiencia no muestran una brecha comparable. La fecha corresponde a la última publicación mensual documentada del tablero.

Ask OGERIA

It answers only with what the observatory publishes and can be wrong: check the entries it cites. Your questions are sent to an AI model, so don't write personal data. More in the privacy policy.

Up to 500 characters.

Support OGERIA on Ko-fi

The payment is processed by Ko-fi, not by this site. Open on ko-fi.com