Personalised persuasion at scale
Conversations optimised to convince, held with millions of people at once, without it showing that a model is on the other side.
- Severity
- Severe
- Horizon
- Already happening
- Evidence
- Lab
- Consensus
- Low
The defensible concern is not the viral deepfakeDeepfakeA fake video, audio clip or image made with AI, showing a real person saying or doing something that never happened.For exampleA voice message from your boss asking for an urgent transfer, which your boss never recorded. but the conversation: personalised, sustained and optimised to persuade. The effect exists and has been measured. Salvi and co-authors, in Nature Human Behaviour, with a 2×2×3 design and N = 900, find that in the debate pairs where AI and humans were not equally persuasive, GPT-4 with personalisation was more persuasive 64.4% of the time [930]On the conversational persuasiveness of GPT-4View source ↗. The mechanism works in both directions: Costello, Pennycook and Rand reduced belief in conspiracy theories by around 20% with three rounds of dialogue, with the effect persisting for two months [284]Durably reducing conspiracy beliefs through dialogues with AIView source ↗verified through Crossref. The capacity to persuade is neutral with respect to content. Detection does not protect either: in the unauthorised University of Zurich experiment on a subreddit of nearly four million users, nobody ever suggested that the comments might be AI-generated [1146]Can AI Change Your View? Evidence from a Large-Scale Online Field Experiment (extended abstract)View source ↗.
What this does not demonstrate. This is the worst-supported part of the epistemic axis, and what fails is precisely the word “personalised”. Hackenburg and Margetts, in PNAS, with n = 8,587 and live microtargetingMicrotargetingSending each person, or very small groups, a different political message, tailored to what is known about their tastes, fears and data.For exampleYou and your neighbour seeing ads from the same candidate with opposite promises, each the one you wanted to hear., find that the persuasive impact of microtargeted messages was not statistically different from that of untargeted ones (4.83 against 6.20 percentage points, P = 0.226) [492]Evaluating the persuasive influence of political microtargeting with large language modelsView source ↗verified through Crossref. And the study that puts the field in order —N = 76,977, 19 models, 707 topics, 466,769 verified claims— concludes that persuasion comes from post-training and promptingPromptThe instruction or question written to an AI. Prompting is drafting it to get a particular answer.For exampleWhat you type into the chat: “Summarise this contract in five points” is a prompt., with smaller effects from personalisation and model size, and adds the finding that defines the shape of the risk: where those methods increased persuasion, they systematically lowered factual accuracy [493]The levers of political persuasion with conversational artificial intelligenceView source ↗verified through Crossref. A note on method: on 3 September 2026 an author correction to Salvi’s article was published whose content could not be verified [931]Author Correction: On the conversational persuasiveness of GPT-4View source ↗.
Chain of materialisation
PreconditionObserved
Nobody detects that there is a model on the other side
University of Zurich researchers deployed semi-automated accounts in a subreddit of almost four million users between November 2024 and March 2025, without notice. Their own summary records that throughout the intervention users never raised the possibility that AI had generated the comments.
TriggerLab
In controlled conditions, the effect exists and is measurable
Salvi and co-authors, with a 2x2x3 design and N=900, find that in debate pairs where AI and humans were not equally persuasive, GPT-4 with personalisation was more persuasive 64.4% of the time. Without personalisation it did not significantly beat humans, and humans given the same sociodemographic data did not improve either.
Observed and demonstrated evidence ends here. What follows is projection.
CascadeProjected
The jump from measured effect to population-level opinion change
Effect sizes in large, well-controlled studies are a few percentage points of attitude change, not mass conversion. That this, aggregated over millions of sustained conversations, produces a shift in public opinion is extrapolation, and the international report warns that AI-generated manipulative content is hard to detect, which makes evidence-gathering difficult in either direction.
ImpactSpeculative
An electorate persuaded by something less accurate
The shape of the risk is not scale but the trade-off: the methods that raise persuasiveness systematically lower factual accuracy. That coupling is measured; its aggregate consequence is not.
See on the map →Report a mistake in this entry →
Sources
- [930] On the conversational persuasiveness of GPT-4 · EPFL 2025
- [931] Author Correction: On the conversational persuasiveness of GPT-4 · Nature Human Behaviour 2026
- [493] The levers of political persuasion with conversational artificial intelligence · UK AI Security Institute / University of Oxford 2025 verified through Crossref
- [492] Evaluating the persuasive influence of political microtargeting with large language models · University of Oxford / Alan Turing Institute 2024 verified through Crossref
- [494] The Levers of Political Persuasion with Conversational AI · Hackenburg, Kobi 2025
- [284] Durably reducing conspiracy beliefs through dialogues with AI · MIT Sloan / American University / Cornell 2024 verified through Crossref
- [1146] Can AI Change Your View? Evidence from a Large-Scale Online Field Experiment (extended abstract) · Universität Zürich 2025
- [897] AI-Reddit study leader gets warning as ethics committee moves to 'stricter review process' · The Center for Scientific Integrity 2025
- [538] International AI Safety Report 2026 · International AI Safety Report (panel con representantes nominados por más de 30 países) 2026
- [166] Misunderstanding the harms of online misinformation · University of Michigan / Dartmouth College / Microsoft Research / Syracuse University / University of Pennsylvania 2024