Expert-enhanced pathogens
A team that already knows virology uses biological design models to produce an agent with properties that do not exist in nature.
- Severity
- Irreversible
- Horizon
- 3–10 years
- Evidence
- Projected
- Consensus
- Low
This risk does not run through the amateur but through those who already know. The question is whether biological design models allow a competent team to produce an agent with properties that nature does not offer —and which therefore has no countermeasures ready.
What has been demonstrated is the preceding step. King and co-authors published in Science, on 6 August 2026, sixteen viable bacteriophages designed with genomic language modelsLLM (large language model)A large language model: the kind of AI behind assistants such as ChatGPT, Claude or Gemini, trained on enormous amounts of text to predict which word comes next.For exampleLike your phone's autocomplete, but trained on vastly more text: that is why it can carry a whole conversation and not just the next word. on the ΦX174 template [597]Generative design of bacteriophages with genome language modelsView source ↗verified through Crossref. And Anthropic reports that Mythos 5.1 outperforms a pre-trained protein language model at predicting AAV capsid assembly using only its reasoning [82]System Card: Claude Fable 5.1 & Claude Mythos 5.1View source ↗. Genome design works, and general reasoning is starting to catch up with specialised tools.
What is missing to reach the risk is everything else, and the labs say so. Anthropic places Mythos 5.1 at CB-1 —significant help in synthesising a known weapon— and explicitly below CB-2, the threshold of replacing the scarce expert, which is the limiting factor for novel weapons [82]System Card: Claude Fable 5.1 & Claude Mythos 5.1View source ↗. On OpenAI’s Critical axis, zero of three evaluations cross the threshold [825]GPT-5.6 System CardView source ↗.
What this does not demonstrate. The phage work was built so as not to produce a human pathogen: it started from non-pathogenic systems in laboratory strains of E. coli, and the Arc Institute states that Evo cannot generate human viral sequences because of deliberate exclusions from the training data [93]How We Built the First AI-Generated GenomesView source ↗. The success rate also constrains the reading: Simon Jackson calculates that around 5% of the designs worked, and Jordi García Ojalvo points out that designed genomes have to be tested one by one [976]Expert reaction to generative design of bacteriophages with genome language modelsView source ↗. The CB-2 threshold is not a measurement but a boundary defined looking forward.
Chain of materialisation
PreconditionObserved
Genomic models already produce functional viral genomes
King and co-authors published in Science, on 6 August 2026, sixteen viable bacteriophages designed with genome language models on the ΦX174 template. Of hundreds of thousands of candidates, 285 went to synthesis and 16 worked.
Precedents: Genome language models produce sixteen viable bacteriophages
TriggerLab
General reasoning catches up with specialised tools
Anthropic reports that Mythos 5.1 exceeds the AAV capsid assembly prediction benchmark using reasoning alone, with no labelled data and no internet, beating a naive ESM-2 application. The document itself calls it an early indicator, necessary but insufficient.
Precedents: Anthropic releases Claude Fable 5 and Claude Mythos 5 · The Fable 5.1 and Mythos 5.1 system card reports control circumvention in production
Observed and demonstrated evidence ends here. What follows is projection.
CascadeProjected
The threshold that replaces scarce expertise
Anthropic defines CB-2 as functionally substituting scarce human expertise, today the primary barrier, and finds Mythos 5.1 does not cross it due to weak novel ideation, poor strategic judgement and errors requiring expertise to catch. Nobody has measured a crossing: it is a frontier defined forward.
ImpactSpeculative
An agent with no countermeasures and sustained transmission
This is the only risk in the observatory whose severity is irreversible because of the agent rather than the actor: a pathogen with sustained transmission cannot be recalled. Any magnitude estimate is entirely conceptual.
Scenarios where it appears
Related measures
See on the map →Report a mistake in this entry →
Sources
- [597] Generative design of bacteriophages with genome language models · Stanford University / Arc Institute 2026 verified through Crossref
- [596] Generative design of novel bacteriophages with genome language models · King, Samuel H. 2025
- [93] How We Built the First AI-Generated Genomes · Arc Institute 2025
- [976] Expert reaction to generative design of bacteriophages with genome language models · Science Media Centre 2026
- [82] System Card: Claude Fable 5.1 & Claude Mythos 5.1 · Anthropic 2026
- [78] Claude Fable 5 and Claude Mythos 5 · Anthropic 2026
- [825] GPT-5.6 System Card · OpenAI 2026
- [161] Contemporary AI foundation models increase biological weapons risk · Brent, Roger 2025
- [156] Rethinking the De-skilling Narrative in AI and Biological Weapons Policy · Georgetown Journal of International Affairs 2026
- [507] An Overview of Catastrophic AI Risks · Center for AI Safety 2023