Chapter 07
Indicators
Eleven series with public measurement. What can be tracked without depending on anyone's interpretation.
Series
Incidents per year in the AI Incident Database
Figure 7.1 · incidentes · proxy indicator · annual cadence
Rising means more risk
AI harms that were publicly reported and recorded in the AI Incident Database, counted by year of the event. It is the only public longitudinal record of observed rather than projected harm. It does NOT answer whether there are more AI harms: it answers how much AI harm is being documented, which mixes real harm, media attention and the editorial team's throughput. If it rises, the first hypothesis is more attention, not more harm. The AIID itself warns that incident IDs are added after the underlying events occur and that the lag varies considerably, and that concentrations by country or company partly reflect what was investigated and processed in that period.
Source [31] AI Incident Database — Database Snapshots ↗ · Data ↗
2024 · El AI Index 2026 reporta 233 para este año usando la misma base en su fecha de corte.
2025 · RELLENO RETROACTIVO MEDIDO: el AI Index 2026 reporta 362 para este mismo año con la misma base; el volcado del 7 de septiembre de 2026 da 449. Un 24% más para un año ya cerrado, solo por incidentes agregados después. Al actualizar hay que recalcular la serie entera, no solo agregar el punto nuevo.
Sep 2026 · PARCIAL Y SUBCONTADO: año en curso al volcado del 7 de septiembre de 2026. Por el rezago editorial esta cifra va a subir, y el ejemplo de 2025 da la magnitud: hasta un 24%. En el gráfico debe ir con otro trazo.
Open entry: Incidents per year in the AI Incident Database
View as table
Incidents per year in the AI Incident Database
AI harms that were publicly reported and recorded in the AI Incident Database, counted by year of the event. It is the only public longitudinal record of observed rather than projected harm. It does NOT answer whether there are more AI harms: it answers how much AI harm is being documented, which mixes real harm, media attention and the editorial team's throughput. If it rises, the first hypothesis is more attention, not more harm. The AIID itself warns that incident IDs are added after the underlying events occur and that the lag varies considerably, and that concentrations by country or company partly reflect what was investigated and processed in that period.
Figure 7.1 · incidentes · proxy indicator · annual cadence · Rising means more risk
Source [31] AI Incident Database — Database Snapshots ↗
| Date | incidentes | Note |
|---|---|---|
| Sep 2026 | 212 (year to date) | PARCIAL Y SUBCONTADO: año en curso al volcado del 7 de septiembre de 2026. Por el rezago editorial esta cifra va a subir, y el ejemplo de 2025 da la magnitud: hasta un 24%. En el gráfico debe ir con otro trazo. |
| 2025 | 449 | RELLENO RETROACTIVO MEDIDO: el AI Index 2026 reporta 362 para este mismo año con la misma base; el volcado del 7 de septiembre de 2026 da 449. Un 24% más para un año ya cerrado, solo por incidentes agregados después. Al actualizar hay que recalcular la serie entera, no solo agregar el punto nuevo. |
| 2024 | 299 | El AI Index 2026 reporta 233 para este año usando la misma base en su fecha de corte. |
| 2023 | 174 | — |
| 2022 | 106 | — |
| 2021 | 80 | — |
| 2020 | 91 | — |
| 2019 | 44 | — |
| 2018 | 46 | — |
| 2017 | 50 | — |
| 2016 | 41 | — |
| 2015 | 24 | — |
| 2014 | 13 | — |
FrontierMath tier 4 saturation
Best cumulative score on FrontierMath tier 4, the research-level problem set designed to resist. What it measures is not mathematics: it is how long an exam lasts before it stops discriminating. If a benchmark built to be impossible saturates in twenty months, the community's ability to measure lags the capability it is trying to measure, and that lag is the risk: safety-policy thresholds rest on evaluations that age faster than they are written. Series from Epoch AI's own runs; one series, one evaluator.
Figure 7.1 · % de acierto · direct measurement · irregular cadence · Rising means more risk
Source [372] Epoch AI Benchmarking Hub — benchmarks.csv ↗
| Date | % de acierto | Note |
|---|---|---|
| Sep 2026 | 97.6 | GPT-6 Astra. De 0% a 97,6% en veinte meses sobre problemas de nivel de investigación. Es el dato más fuerte de esta familia. |
| Jun 2026 | 90.2 | Claude Fable 5 max. |
| Apr 2026 | 78 | GPT-5.5 Pro xhigh. |
| Mar 2026 | 58.5 | GPT-5.4 Pro xhigh. |
| Dec 2025 | 46 | GPT-5.2 Pro. |
| Aug 2025 | 22 | GPT-5 high. |
| Apr 2025 | 4.9 | o4-mini high. |
| Jan 2025 | 0 | o3-mini high. Punto de partida en cero. |
GPQA Diamond saturation
Best cumulative score on GPQA Diamond, the PhD-level science question set designed to resist internet search. It is effectively saturated: the last four highs fall within each other's standard errors, meaning the exam no longer discriminates between frontier models. A saturated benchmark is not a fact about capability, it is an exhausted instrument, and quoting it anyway creates the illusion of a ceiling. Series from Epoch AI's own runs; one series, one evaluator.
Figure 7.1 · % de acierto · direct measurement · irregular cadence · Rising means more risk
Source [372] Epoch AI Benchmarking Hub — benchmarks.csv ↗
| Date | % de acierto | Note |
|---|---|---|
| Sep 2026 | 95.8 | GPT-6 Astra max, con un error estándar de 1,4 puntos porcentuales. El examen ya no distingue. |
| Feb 2026 | 94.4 | Gemini 3.1 Pro Preview. |
| Nov 2025 | 92.6 | Gemini 3 Pro Preview. |
| Mar 2025 | 83.8 | Gemini 2.5 Pro Exp. |
| Dec 2024 | 76.8 | o1 high. |
| Jun 2024 | 54 | Claude 3.5 Sonnet. |
| Mar 2023 | 35.7 | GPT-4. |
SWE-bench Verified saturation
Best cumulative score on SWE-bench Verified, which measures resolution of real software repository issues. It is the only member of the saturation family that still discriminates between frontier models, and therefore the most useful to track. Beware of comparing figures across evaluators: the 2026 AI Index states it rose from 60% to nearly 100% in a single year, while Epoch's own runs top out at 83.5%. Neither is lying: scaffolding, token budgets and scoring criteria differ. Site rule: one series, one evaluator.
Figure 7.1 · % de acierto · direct measurement · irregular cadence · Rising means more risk
Source [372] Epoch AI Benchmarking Hub — benchmarks.csv ↗
| Date | % de acierto | Note |
|---|---|---|
| Apr 2026 | 83.5 | Claude Opus 4.7 max, con un error estándar de 1,7 puntos porcentuales. |
| Feb 2026 | 78.7 | Claude Opus 4.6 sin razonamiento extendido. |
| Nov 2025 | 76.7 | Claude Opus 4.5 sin razonamiento extendido. |
| Aug 2025 | 73.6 | GPT-5 high. |
| May 2025 | 70.7 | Claude Opus 4. |
| Feb 2025 | 61 | — |
| Nov 2024 | 31 | GPT-4o. |
AI adoption by US firms (Census BTOS)
Share of US firms reporting they used artificial intelligence in any business function in the previous two weeks, per the Census Bureau's Business Trends and Outlook Survey. It is the best public source in this section because it is an official probability survey with explicit reference and publication dates. It serves as the denominator for much of the rest: an incident in a country where 22% of firms use AI does not mean the same as where 5% do. As of 3 May 2026 the breakdown showed 39.7% in Information, 33.9% in Finance and Insurance and about 14% in Retail Trade; by size, 37% among firms with 250 or more employees against under 20% for those with fewer than 20.
Figure 7.1 · % de empresas · direct measurement · monthly cadence · Rising means more risk
Source [225] Business Trends and Outlook Survey — National estimates (National.xlsx) ↗
| Date | % de empresas | Note |
|---|---|---|
| Aug 2026 | 22.4 | Período 202617, referencia 27 jul - 9 ago 2026, publicado el 27 de agosto de 2026. Expectativa a seis meses: 25,9%. Entre estos puntos hay períodos intermedios (202526, 202604, 202611, 202614: 17,8%, 17,5%, 20,6% y 21,7%) cuyas fechas de referencia no quedaron documentadas y por eso no se grafican. |
| Jul 2026 | 21.8 | Período 202616, referencia 13-26 jul 2026, publicado el 13 de agosto de 2026. Expectativa a seis meses: 25,9%. |
| Apr 2026 | 19.8 | Período 202608, referencia 23 mar - 5 abr 2026, publicado el 23 de abril de 2026. Expectativa a seis meses: 23,0%. |
| Dec 2025 | 17.7 | Período 202601, referencia 15-28 dic 2025, publicado el 15 de enero de 2026. Expectativa a seis meses: 22,4%. |
| Nov 2025 | 17.3 | QUIEBRE DE SERIE, NO ADOPCIÓN. En noviembre de 2025 el Census cambió la redacción de la pregunta de IA, y la serie con la redacción vieja venía en 9,95% (septiembre de 2025). El salto de ~10% a ~17% es un cambio de instrumento: produce un gráfico espectacular y falso. Por eso esta serie arranca acá y no antes. Evidencia del cambio: el XLSX nacional solo publica la pregunta 7 desde este período en adelante, y la reconstrucción de Ramp trae un campo explícito question_version con los valores pre y post cambio de redacción, marcando octubre de 2025 como inutilizable. Período 202524, referencia 3-16 nov 2025, publicado el 4 de diciembre de 2025. Expectativa a seis meses: 21,1%. |
Largest known training compute per year
The largest training compute published or estimated for a model in each year, per Epoch AI's notable models database. Compute is the input that best predicts capability and the only one currently regulated with a numeric threshold: the 1e25 and 1e26 FLOP marks in the EU AI Act and US executive orders are written in this unit. NOTE: it measures DISCLOSURE, not compute. 2024 appears below 2023 and that did not happen: it reflects that the model with published compute that year was an open Meta model while closed frontier labs reported nothing. Read it as 'largest known compute', never as 'largest compute', and the gap between the two widens every year labs stop reporting.
Figure 7.1 · FLOP · proxy indicator · annual cadence · log scale · Rising means more risk
Source [377] Notable AI Models — base de datos descargable ↗
Threshold · 1 × 10²⁶: Highest regulatory threshold written in FLOP. Only three known models exceeded it, all in 2025. Pre-training efficiency improves about 3x a year, so a fixed threshold admits a more capable model every year: it does not age, it melts.
| Date | FLOP | Note |
|---|---|---|
| Sep 2026 | 3.87 × 10²⁵ (year to date) | PARCIAL: año en curso al 9 de septiembre de 2026. Composer 2.5. La base tenía solo 32 modelos de 2026 con fecha y 22 con cómputo, así que este valor va a subir y no es comparable con los años cerrados. |
| 2025 | 5 × 10²⁶ | Grok 4. Costo de cómputo estimado por Epoch: 10.717 millones de dólares. Tres modelos superaron 1e26 FLOP ese año. |
| 2024 | 3.8 × 10²⁵ | Llama 3.1-405B, un modelo abierto. La caída respecto de 2023 es artefacto de divulgación, no una baja real del cómputo frontera. Costo estimado: 928 millones de dólares. |
| 2023 | 5 × 10²⁵ | Gemini 1.0 Ultra. Costo estimado más alto del año: 806 millones de dólares (GPT-4 de marzo). |
| 2022 | 2.74 × 10²⁴ | Minerva (540B). |
| 2021 | 1.7 × 10²⁴ | EXAONE 1.0. |
| 2020 | 3.14 × 10²³ | GPT-3 175B (davinci). Costo estimado de entrenamiento: 232 millones de dólares. |
| 2019 | 1.08 × 10²³ | AlphaStar. |
Companies with a published frontier safety framework
How many AI companies impose capability thresholds on themselves with written consequences and publish them. It is the voluntary-governance indicator: it measures whether the promise exists, not whether it is kept. The count alone misleads, which is why it carries a threshold: at the May 2024 AI Seoul Summit sixteen companies committed to publishing and four more joined later; twelve published. The distance between committed and published is a better indicator than the count, because it measures failure to keep an explicit commitment and admits no optimistic reading. The twelve as of the December 2025 report are Anthropic, OpenAI, Google DeepMind, Magic, Naver, Meta, G42, Cohere, Microsoft, Amazon, xAI and NVIDIA.
Figure 7.1 · empresas · direct measurement · irregular cadence · Falling means more risk
Source [708] Common Elements of Frontier AI Safety Policies ↗
Threshold · 20: Companies that committed to publishing a frontier safety framework: sixteen at the May 2024 AI Seoul Summit plus four that joined later. The gap with those that actually published is the finding.
| Date | empresas | Note |
|---|---|---|
| Dec 2025 | 12 | ES UN PISO, NO EL VALOR VIGENTE: la cifra está anclada a la actualización de diciembre de 2025 y no se encontró una revisión posterior de METR, así que si en 2026 se sumaron empresas este número está desactualizado. Del informe solo está documentado el mes. Contexto que el conteo no captura: dos de los tres umbrales activados hasta la fecha son PRECAUTORIOS —se activaron porque no se pudo descartar el riesgo, no porque se haya medido—, así que la serie de umbrales activados sube tanto cuando la capacidad crece como cuando la evaluación empeora o el benchmark se satura. |
Task time horizon (METR)
Length of tasks a model completes with 50% success, measured in the time of an expert human with no prior context. It is the capability indicator with the longest comparable series and the best public proxy for sustained autonomy: not how fast an agent writes code, but how many hours it chains decisions without correction. METR rejects four readings: it does not measure how long an agent can act autonomously but how hard the task is; tasks are software, machine learning and cybersecurity, not 'everything intellectual'; the human reference is a newcomer to the project, not a professional with the context in mind; and real work is rarely algorithmically scorable.
Figure 7.1 · horas · direct measurement · irregular cadence · log scale · Rising means more risk
Source [711] METR Time Horizon 1.1 — benchmark_results_1_1.yaml ↗
Threshold · 16: Ceiling declared by METR. Above 16 hours, measurements are unreliable with their current task suite.
| Date | horas | Note |
|---|---|---|
| Apr 2026 | 17.41 | Claude Mythos Preview, versión temprana. 1.044,8 minutos, IC 95% 508,9-3.304,3, casi un orden de magnitud de ancho. FUERA DE RANGO: está sobre el techo de 16 horas que METR declara medible con su suite actual, y por eso queda excluido del ajuste de duplicación. Se grafica, no se titula. · [713] |
| Mar 2026 | 5.7 | GPT-5.4. 341,7 minutos, IC 95% 186,6-768,8. |
| Feb 2026 | 6.4 | Gemini 3.1 Pro. 384,1 minutos, IC 95% 233,5-694,8. |
| Feb 2026 | 11.98 | Claude Opus 4.6. 718,8 minutos, IC 95% 316,7-3.633,8. Cuando METR lo anunció lo reportó en unas 14,5 horas; la corrección de un error de regularización del 3 de marzo de 2026 lo bajó a 11,98. GPT-5.3-Codex tiene la misma fecha de modelo y 5,83 horas: se omite porque la serie no admite fechas repetidas. |
| Dec 2025 | 5.87 | GPT-5.2. 352,2 minutos, IC 95% 198,1-815,2. |
| Nov 2025 | 4.88 | Claude Opus 4.5. 293,0 minutos, IC 95% 161,7-623,7. El anuncio de Time Horizon 1.1, en enero de 2026, lo había publicado en 320 minutos. |
| Nov 2025 | 3.74 | Gemini 3 Pro. 224,3 minutos, IC 95% 139,6-379,2. |
| Aug 2025 | 3.38 | GPT-5. 203,0 minutos, IC 95% 112,6-405,6. |
| May 2025 | 1.67 | Claude Opus 4. 100,4 minutos, IC 95% 60,0-163,5. Queda por debajo de o3, publicado cinco semanas antes: la serie no es monótona. |
| Apr 2025 | 2 | o3. 119,7 minutos, IC 95% 74,6-190,9. |
| Feb 2025 | 1.01 | Claude 3.7 Sonnet. 60,4 minutos, IC 95% 33,0-104,2. |
| Dec 2024 | 0.65 | o1. 38,8 minutos, IC 95% 21,2-65,0. |
| Sep 2024 | 0.34 | o1-preview. 20,3 minutos, IC 95% 11,7-33,4. |
| Jun 2024 | 0.19 | Claude 3.5 Sonnet. 11,4 minutos, IC 95% 5,5-22,4. |
| May 2024 | 0.12 | GPT-4o. 7,0 minutos, IC 95% 4,0-12,9. |
| Mar 2023 | 0.07 | GPT-4. 4,0 minutos, IC 95% 1,9-8,0. |
| Mar 2022 | 0.01 | GPT-3.5 Turbo Instruct. 0,6 minutos. |
| May 2020 | 1.7 × 10⁻³ | davinci-002. 0,1 minutos. |
| Feb 2019 | 1.7 × 10⁻³ | GPT-2. 0,1 minutos en el archivo original. |
AI adoption measured by transactions (Ramp AI Index)
Share of firms in Ramp's panel paying for AI products, measured from transaction data across more than 70,000 US companies on its corporate card and payments system. Its advantage over a survey is that it is not self-reported: a company paying an AI subscription is using it. This is a contrast series, not the primary one: Ramp is a fintech selling to the same firms it measures, and its panel skews toward young, technology-focused, fast-growing companies. The 56% in its panel against the 22.4% in the Census probability sample is not a contradiction: it is the difference between a panel and a representative sample. Dates stamp the month of each monthly release.
Figure 7.1 · % de empresas del panel · proxy indicator · monthly cadence · Rising means more risk · Out of date · Last measured: 1 Aug 2026
Source [885] Ramp AI Index ↗
| Date | % de empresas del panel | Note |
|---|---|---|
| Aug 2026 | 56.13 | Alza interanual de 11,15 puntos porcentuales. Gasto mediano por empleado y por mes: 12,50 dólares, cinco veces los 2,50 de septiembre de 2023. En el decil superior de empresas la mediana pasó de 70,04 a 675,60 dólares en el mismo lapso. |
| Apr 2026 | 52.98 | Gasto mediano por empleado y por mes: 10,00 dólares. |
| Dec 2025 | 45.92 | Gasto mediano por empleado y por mes: 5,26 dólares. |
| Jun 2025 | 42.72 | Gasto mediano por empleado y por mes: 4,27 dólares. |
| Dec 2024 | 36.66 | Gasto mediano por empleado y por mes: 3,33 dólares. |
| Dec 2023 | 31.68 | Gasto mediano en IA por empleado y por mes: 2,67 dólares. |
| Jan 2023 | 7.46 | — |
Generative AI services registered in China
Cumulative number of generative AI services that completed registration with the Chinese regulator, mandatory for those with public opinion attributes or social mobilisation capacity. It is an administrative count, not a measure of capability or safety: it says how many services entered the system, and nothing about how safe the models behind them are. It is published because it is the only hard number on regulatory scale that the Chinese system produces, and because since July 2026 it also covers models running on the phone itself.
Figure 7.1 · servicios · proxy indicator · irregular cadence · Rising means more risk
Source [185] 国家互联网信息办公室关于发布2025年生成式人工智能服务已备案信息的公告 ↗
| Date | servicios | Note |
|---|---|---|
| Aug 2026 | 1,112 | Julio-agosto de 2026: +124 servicios (7 de terminal móvil) y +133 aplicaciones o funciones inscritas, corte del 31 de agosto. · [187] |
| Jun 2026 | 988 | Primer semestre de 2026, con 598 aplicaciones o funciones inscritas. El corte del 30 de abril, 868 servicios, está publicado por la misma autoridad pero no tiene ficha propia en el registro. · [188] |
| Dec 2025 | 748 | Cierre de 2025. En el año se sumaron 446 servicios; además había 435 aplicaciones o funciones inscritas, que se cuentan aparte. |
Youth employment decline in AI-exposed occupations
Annualized employment change for women aged 22 to 25 in the occupations most exposed to AI —the subgroup falling the most; men of the same age fall less—, computed from the baseline and measured on real ADP payroll data by the Stanford Digital Economy Lab's Canaries in the Coal Mine dashboard. It is the indicator closest to economic harm measured in people rather than in intention surveys: a balanced sample of 25,000 firms and 4.6 million workers across more than 730 occupations, with a November 2022 baseline, the month ChatGPT launched. THE COUNTERARGUMENT COMES FROM THE AUTHORS THEMSELVES: there is no economy-wide displacement. The effect is real but concentrated in one group, and with the broadest set of controls the declines only become statistically significant in 2024. A dashboard showing a large gap in one group and zero in the aggregate is exactly what should be published.
Figure 7.1 · % anual desde noviembre de 2022 · direct measurement · monthly cadence · Falling means more risk
Source [991] Canaries Dashboard — AI Economic Indicators ↗
| Date | % anual desde noviembre de 2022 | Note |
|---|---|---|
| Aug 2026 | -4.4 | Publicación mensual del 18 de agosto de 2026 (datos de ADP con corte el 12 de agosto, última observación julio de 2026). Índice de empleo de las mujeres de 22 a 25 años en el quintil más expuesto: 84,82 (noviembre de 2022 = 100), es decir, 4,4% menos por año. Los hombres del mismo grupo quedan en 92,30 (−2,2% anual) y ambos sexos juntos en −3,3% anual; la caída interanual de julio a julio es de 3,0%. En el mismo quintil, el índice de todas las edades está 3,6% (mujeres) y 4,2% (hombres) sobre la línea base: la brecha es de los jóvenes. Cálculo propio sobre los CSV públicos del tablero. |
| Jul 2026 | -4.5 | Mujeres de 22 a 25 años en las ocupaciones más expuestas, contracción anual desde la línea base de noviembre de 2022. Los hombres del mismo grupo caen 2,5% anual. La actualización de agosto de 2026 del paper sitúa el empleo de ese grupo un 19% por debajo de donde estaría si hubiera seguido el ritmo de sus pares menos expuestos; los trabajadores con experiencia no muestran una brecha comparable. La fecha corresponde a la última publicación mensual documentada del tablero. |