Contamination of the public record
Synthetic material enters the shared record -web, encyclopedias, literature- and stops being distinguishable from the original.
- Severity
- Severe
- Horizon
- 1–3 years
- Evidence
- Projected
- Consensus
- Medium
The harm would not be believing something false, but losing the procedure by which anything used to be checked. The only time series with a public method is produced by Graphite —an SEO company, that is, an interested party—: over 55,400 Common Crawl URLs and with three detectors averaged, mostly AI-generated articles overtook human-written ones in the fourth quarter of 2025 at 50.9%, and fell back to 49.9% in the first quarter of 2026 [472]AI Now Writes as Many Online Articles as Humans DoView source ↗. “About half of new indexable articles in English” is defensible; “half of the internet” is not.
The defence that worked is instructive. Wikipedia created the WP:G15 speedy-deletion criterion to remove without discussion pages clearly generated by a model without human review, and what served as the signature was not a statistical detector but two behavioural traces: the conversational residue pasted into the text and the invented references [1104]Wikipedia:WikiProject AI CleanupView source ↗. Automatic detection ages fast: a method with more than 99.8% AUROC on contemporary deepfakesDeepfakeA fake video, audio clip or image made with AI, showing a real person saying or doing something that never happened.For exampleA voice message from your boss asking for an urgent transfer, which your boss never recorded. loses more than 30% of its recall against techniques barely six months newer [912]Performance Decay in Deepfake Detection: The Limitations of Training on Outdated DataView source ↗.
What this does not demonstrate. The measurement inherits the error of the tool it is measured with: Graphite itself revised its figures down by some 3.3 percentage points when it moved from one detector to three [472]AI Now Writes as Many Online Articles as Humans DoView source ↗, and the curve flattened from 2024 onwards instead of shooting up. The result cited as the collapse mechanism also has a rebuttal: Shumailov and co-authors show irreversible defects when training recursively on synthetic data [967]AI models collapse when trained on recursively generated dataView source ↗, but that depends on the new data replacing the old; when it is accumulated —which is what the web does— the collapse does not occur [450]Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic DataView source ↗. And the most complete security analysis of C2PA concludes that its current specifications do not meet its own stated objectives [459]Verifying Provenance of Digital Media: Why the C2PA Specifications Fall ShortView source ↗.
Chain of materialisation
PreconditionObserved
About half of new indexable articles are already synthetic
With 55,400 Common Crawl URLs and three averaged detectors, primarily AI-generated articles surpassed human-written ones in Q4 2025 at 50.9%, returning to 49.9% in Q1 2026. The sample is English-language articles with schema markup and at least a hundred words: it is not half the internet.
TriggerObserved
The institutions that curate the record have already had to defend themselves
Wikipedia created a speedy deletion criterion, WP:G15, allowing deletion without discussion of pages clearly generated by a model without human review. What is interesting is what worked as a signature: not a statistical detector but two behavioural traces -conversational residue left in the text and fabricated references.
Observed and demonstrated evidence ends here. What follows is projection.
CascadeProjected
Detection degrades faster than it is rebuilt
A method reaching over 99.8% AUROC on contemporary deepfakes loses more than 30% recall when evaluated against generation techniques from just six months later. And a detector with 5% false positives applied to millions of items produces more false accusations than correct detections when true prevalence is low.
ImpactSpeculative
A shared record that can no longer be audited
The harm would not be believing something false but losing the procedure by which anything was checked. There is no measurement of that state, and the very estimate of how much synthetic material exists inherits the error rate of the detector used.
Related measures
See on the map →Report a mistake in this entry →
Sources
- [472] AI Now Writes as Many Online Articles as Humans Do · Graphite 2026
- [471] More Articles Are Now Created by AI Than Humans · Graphite 2025
- [1104] Wikipedia:WikiProject AI Cleanup · Wikimedia Foundation 2026
- [912] Performance Decay in Deepfake Detection: The Limitations of Training on Outdated Data · The Alan Turing Institute 2025
- [815] Beyond benchmark accuracy: Evaluating deepfake detection tools for digital forensic admissibility through a systematic review · Onyekwere, Chukwudi George-Linus 2026
- [967] AI models collapse when trained on recursively generated data · University of Oxford / University of Cambridge / Imperial College London / University of Toronto 2024
- [450] Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data · Stanford University / Harvard University / MIT 2024
- [459] Verifying Provenance of Digital Media: Why the C2PA Specifications Fall Short · University of Maryland Baltimore County 2026
- [1140] Invisible Image Watermarks Are Provably Removable Using Generative AI · UC Santa Barbara / Carnegie Mellon University 2023
- [1060] Reglamento (UE) 2024/1689 (Reglamento de Inteligencia Artificial), Artículo 50 — Obligaciones de transparencia · Unión Europea 2024
- [213] Code of Practice on Transparency of AI-generated Content · European Commission — Shaping Europe's digital future 2026