AI-Enabled Pathogen Design

In 2024, the Nobel Committee awarded its Chemistry Prize to David Baker, Demis Hassabis, and John Jumper for computational protein design and structure prediction. These methods have shortened some discovery workflows, but experimental validation remains necessary. The same tools that support therapeutic protein research can theoretically be repurposed for harmful design, although most computational designs still fail or require substantial refinement in wet-lab testing.

Learning Objectives
  • Distinguish between protein structure prediction (AlphaFold) and generative design (RFdiffusion, ESM3) in biosecurity context.
  • Explain the “Screening Gap” and why traditional DNA screening may fail against AI-designed sequences.
  • Analyze the three-tier risk framework from the 2025 NASEM report.
  • Evaluate the dual-use potential of biological design tools using real case studies.
  • Apply practical governance questions to assess emerging AI-biology platforms.
  • Recognize the persistent wet-lab barriers that limit AI-enabled pathogen design today.
Scope of This Chapter

This chapter discusses biosecurity risks at a conceptual level appropriate for education and policy analysis. Consistent with responsible information practices:

  • Omitted: Actionable protocols, specific synthesis routes, exact pathogen sequences
  • Included: Risk frameworks, governance mechanisms, policy recommendations

For detailed biosafety protocols, consult your Institutional Biosafety Committee and relevant regulatory guidance.

The Shift: While LLMs (LLMs and Information Hazards) lower barriers to knowledge, Biological Design Tools (BDTs) lower barriers to expertise. We have moved from predicting structures to generating entirely new proteins from scratch.

The 2024 Nobel Prize Context: The Nobel Prize in Chemistry was awarded to David Baker, Demis Hassabis, and John Jumper for computational protein design and structure prediction - recognition that AI is fundamentally changing biological engineering.

Three Categories of Risk (NASEM 2025):

Risk Category Current AI Capability Barrier Level
Toxin design Can assist Low-Medium
Pathogen modification May partially assist Medium-High
De novo virus design Far beyond current capabilities Very High

The MegaSyn Wake-Up Call: The 2022 Urbina experiment reported about 40,000 candidate molecules predicted to be toxic in 6 hours, including structures related to known chemical-warfare agents and novel candidates. The study did not validate synthesizability or function.

The Screening Gap: Generative models can produce sequence-distant candidates predicted to retain selected properties, creating challenges for homology-based screening. Predicted retention is not proof of biological function. In one workshop study, three PPI prediction tools detected none of four validated SARS-CoV-2 binding mutants under the tested setup (Feldman & Feldman, 2025, workshop paper); the result should not be generalized to every filter, interaction, or threat.

Reality Check: Many computational protein designs fail in wet-lab validation. Biology’s complexity, including folding errors, expression failures, and unexpected interactions, remains a significant barrier. AI has lowered the cost of proposing designs, but not the difficulty of validating them.

Bottom Line: These tools offer immense benefits for drug discovery while introducing dual-use concerns requiring ongoing vigilance rather than alarm.

Introduction: From Discovery to Engineering

In October 2024, the Nobel Committee awarded its Chemistry Prize to David Baker, Demis Hassabis, and John Jumper for computational protein design and structure prediction. The announcement recognized advances in computational protein design and structure prediction, including methods that use machine learning.

In the early days of structural biology, determining the 3D structure of a protein was a PhD-level project. It involved X-ray crystallography, years of trial and error, and often ended in failure. We treated biology as a discovery science - we went out and found things.

Today, biology is becoming an engineering discipline.

The first time researchers use AlphaFold can be striking. A structure prediction can be produced quickly, but experimental structure determination, validation, and interpretation remain necessary. That change in the order and cost of some discovery steps is both the promise and the concern at the heart of this chapter.

For a physician, this is a miracle (custom therapeutics). For a biosecurity expert, it is a concern (custom toxins). The same capabilities that accelerate drug discovery can theoretically be repurposed toward harmful ends.

The Revolution in Protein Structure Prediction

For over 50 years, predicting how a string of amino acids would fold into a three-dimensional protein structure was one of biology’s grand challenges. Then came AlphaFold.

AlphaFold2, released in 2020 by Google DeepMind, achieved accuracy comparable to experimental methods in the biennial Critical Assessment of Protein Structure Prediction (CASP) competition. By 2022, DeepMind and EMBL-EBI had released predicted structures for essentially all known proteins - roughly 200 million structures - transforming what was once a bottleneck into a freely available resource.

What AlphaFold Does and Does Not Do

Understanding the limitations is essential for calibrated risk assessment:

AlphaFold can: - Predict static protein structures from amino acid sequences with high accuracy - Map mutations onto structures to reason about binding or antigenicity - Accelerate drug discovery by revealing therapeutic targets

AlphaFold cannot: - Design proteins with specific functions from scratch - Reliably predict how mutations affect pathogenicity or transmissibility - Model the complex dynamics of viral assembly or host interactions - Generate novel harmful sequences without additional tools

AlphaFold2 also struggles with intrinsically disordered regions, conformational dynamics, and multimeric assemblies; predicted models are static snapshots, not full interaction or motion maps.

As DeepMind’s biosecurity analysis for AlphaFold3 notes, structure prediction is necessary but not sufficient for most dual-use scenarios of concern.

The “Unmasking” Risk

Structure prediction’s primary biosecurity concern is not about creating new weapons - it is about understanding existing ones better.

Target Identification: To optimize a toxin, you need to know exactly how it binds to human receptors. AlphaFold provides this “lock and key” map.

The Unknown Function Problem: Public genomic databases contain millions of sequences labeled “hypothetical protein.” AlphaFold can, in principle, reveal that an innocuous-looking sequence from soil bacteria is structurally similar to a lethal toxin. This “unmasks” potential threats hidden by our previous ignorance.

The marginal risk here is real but manageable - we already live in a world of mapped toxins.

Generative Design: From Prediction to Creation

If structure prediction was the first revolution, generative design is the second. Rather than predicting what natural proteins look like, generative models create entirely new proteins that have never existed in nature.

RFdiffusion and De Novo Protein Design

RFdiffusion, developed at the Baker Lab in 2023, adapts diffusion-model methods to protein backbones. The tool starts with random noise and iteratively refines it into candidate structures, guided by user specifications.

The RFdiffusion paper reports de novo designs, including experimentally validated high-affinity binders, while also showing that computational designs require wet-lab screening. Performance varies by target and design task, so the paper’s reported rates should not be generalized to all protein-design problems (Watson et al., 2023).

A developer-authored preprint introduced RFdiffusion3, an all-atom model that conditions protein generation on ligands, nucleic acids, and other non-protein atoms. The authors reported improved in silico performance at approximately one-tenth the computational cost of the preceding approach for typical protein lengths, while experimental validation was limited to selected DNA-binding and enzyme-design tasks (Butcher et al., 2025, preprint). Training code, inference code, and model weights are publicly available, increasing both research utility and the importance of downstream screening and access controls.

ESM3: Sequence, Structure, and Function

ESM3, released by EvolutionaryScale (blog 2024; peer-reviewed publication in Science 2025), goes further by integrating sequence, structure, and function into a single generative model. Trained on 2.78 billion proteins, ESM3 can follow complex prompts combining multiple modalities.

In a demonstration, the team reported an ESM3-generated green fluorescent protein with 58% sequence identity to known fluorescent proteins. The associated estimate that this divergence would require over 500 million years of natural evolution is a company estimate, not an independently established evolutionary timescale (EvolutionaryScale, 2024).

Case Study: The MegaSyn Experiment (2022)

This is the clearest demonstration of dual-use potential from generative AI.

The Context: Collaborations Pharmaceuticals, a small drug discovery company, was invited to present at Spiez CONVERGENCE, a Swiss conference on emerging threats. Asked to discuss how AI might be misused, they decided to actually try inverting their drug discovery platform.

The Experiment: They took MegaSyn, a model trained to avoid toxicity in drug candidates, and simply flipped the reward function to maximize toxicity instead.

The Result: In less than 6 hours, on a consumer laptop, the model generated roughly 40,000 potentially toxic molecules. The model reproduced structures overlapping known CWAs (e.g., VX analogs) and proposed novel toxic analogs; synthesizability and ADMET properties were not validated.

The Lesson: The same scoring functions that guide molecules away from toxicity can be redirected toward toxicity. The demonstration shows why objective choice and downstream screening matter.

The researchers deliberately did not assess synthesizability or explore how to make the molecules. They recognized an ethical boundary and stopped. But the proof of concept was clear.

The “Screening Gap”: Unknown Unknowns

This is one important technical challenge for modern biosecurity.

How DNA Screening Works Today

When you order DNA from a synthesis company, they run your order through screening algorithms:

  • The Check: Does this sequence match Smallpox? Ricin? Anything on controlled lists?
  • The Mechanism: Homology search - matching genetic letters against known dangerous sequences

How AI Breaks This Model

Generative AI can design a protein that has the same shape (and therefore the same function) as a known toxin, but a completely different genetic sequence.

The Result: A sequence-similarity screen may not recognize a functionally similar but distant sequence. Whether an order is approved depends on the provider’s screening methods, customer review, and other controls. Most novel AI-generated sequences will not fold or function as intended, but a minority could expose a screening blind spot.

The Problem: We have moved from a world of “known unknowns” to “unknown unknowns” (sequences that function dangerously but do not match anything in our databases).

The Evidence: The October 2025 “Paraphrase Project” (Wittmann et al., Science) demonstrated this vulnerability against unpatched tools. Using EvoDiff, researchers generated thousands of toxin variants that slipped through commercial screening. Pre-patch detection was as low as 23%; after collaborative patching with IGSC members, detection improved to 97% for the highest-risk sequences. A 2026 follow-up tested fragmented synthetic homologs against four patched tools and found them effective against the design capabilities tested, while urging alternative approaches because detection declines at lower sequence similarity (Wittmann et al., 2026). Neither study produced physical proteins, so retained biological function was predicted rather than shown. See DNA Synthesis Screening for implementation details.

The Fix: Function-Based Screening

SecureDNA is a new approach. Rather than relying on sequence similarity alone, it uses:

  • Random adversarial threshold search looking for exact matches to short functional subsequences
  • Predicted variants that would preserve function
  • Screening down to 30 base pairs - far shorter than traditional approaches
  • Cryptographic protections to preserve both customer privacy and the hazard database

Baker and Church’s 2024 Science commentary proposes extending this specifically to AI-designed proteins - maintaining records of AI-generated sequences and screening them against synthesis orders.

Note: Deployment of function-based screening is uneven; most synthesis providers still rely on homology-based screening, so coverage remains incomplete.

The NASEM 2025 Framework: Calibrated Risk Assessment

The 2025 NASEM report “The Age of AI in the Life Sciences”, commissioned by the Department of Defense, provides the most comprehensive assessment to date. The committee examined three categories of risk:

1. Design of Biomolecules (Proteins and Toxins)

AI Capability: Can assist with this today.

Concern: Generative systems could produce candidate molecules or proteins outside known screening references. MegaSyn demonstrated inversion of a small-molecule model toward predicted toxicity, not validated synthesis, toxicity, or biological deployment.

Mitigation: Baker and Church’s proposal for universal screening that includes structure-based checks.

2. Modification of Existing Pathogens

AI Capability: May partially assist.

Concern: AI might help identify mutations that increase pathogenicity or transmissibility.

Limitation: Predicting how specific mutations affect complex phenotypes requires training data that largely does not exist. The biological mechanisms connecting sequence changes to pandemic potential are poorly understood even by experts.

3. De Novo Virus Design

AI Capability: Demonstrated for bacteriophages in a constrained experimental system; not established for mammalian or human-pathogenic viruses.

Concern: Creating a functional virus from scratch.

Evidence boundary: King et al. reported experimentally functional bacteriophages generated with genome language models in a supervised bacterial-virus research system (King et al., 2026). This result ends the categorical claim that functional virus design is beyond current capability for every virus, but it does not establish design of mammalian viruses, human pathogens, enhanced virulence or transmissibility, or autonomous end-to-end execution. Experimental expertise and physical validation remain substantial barriers.

The Complexity Barrier

A recurring theme in biosecurity risk assessment is biological complexity:

  • Molecules have defined structures; viruses are dynamic machines
  • Predicted molecular properties require experimental validation; transmission involves host interactions across multiple systems
  • Small molecule synthesis is routine; creating functional viral genomes requires specialized expertise
  • Drug activity is measurable in vitro; pathogen behavior can only be assessed in complex biological systems

This complexity is both a natural barrier and a reason for humility in prediction.

What We Know vs. What Remains Uncertain

Demonstrated (supported by published evidence):

  • AlphaFold predicts static protein structures with high accuracy
  • Generative design tools such as RFdiffusion can produce experimentally validated proteins for selected tasks, but reported success rates are task-specific and require wet-lab confirmation
  • Drug-discovery objectives can be inverted to generate candidates predicted to be toxic, as demonstrated computationally in MegaSyn
  • ESM3 generated a novel GFP with only 58% sequence identity to known proteins
  • Most AI-designed proteins fail in wet-lab validation
  • Safe-proxy testing, evaluation, validation, and verification (TEVV) can evaluate AI-assisted protein design risk without testing sequences of concern: NIST researchers found current systems could generate synthetic homologs with similar predicted structures, but did not reliably preserve biological activity while evading screening (Ikonomova et al., 2025)
  • In one workshop study, three inference-time PPI prediction tools failed to detect four experimentally validated SARS-CoV-2 binding mutants under the tested conditions (Feldman & Feldman, 2025, workshop paper)
  • Genome language models generated experimentally functional bacteriophages in a constrained bacterial-virus system (King et al., 2026)
  • AlphaGenome (Avsec et al., Nature, 2026) predicts how noncoding DNA variants affect gene regulation across tissue types, extending AI-bio analysis from protein structure to the regulatory genome

Theoretical (plausible but not yet demonstrated):

  • AI-designed toxins evading function-based screening at scale
  • AI predicting gain-of-function mutations for pathogens with useful accuracy
  • AI-designed sequences being successfully synthesized and weaponized
  • AI substantially accelerating pathogen modification by sophisticated actors

Beyond current capabilities (no credible pathway with existing technology):

  • De novo design of mammalian or human-pathogenic viruses with prespecified high-consequence phenotypes
  • AI systems that can execute wet-lab work autonomously
  • AI replacing the tacit knowledge required for pathogen work
  • AI predicting pandemic potential from sequence alone

The NASEM framework provided a useful 2025 baseline, but the 2026 bacteriophage result requires organism- and endpoint-specific wording. Molecular design assistance is demonstrated, pathogen modification may be partially assisted, and functional phage generation is now demonstrated in a constrained system. High-consequence mammalian-virus design and end-to-end execution remain unestablished.

A 2025 RAND Delphi study surveyed biology and AI experts on theoretical limits of AI-enabled pathogen design. Key findings: (1) limits are interdependent and context-dependent, (2) AI effectiveness depends heavily on biological data quality, and (3) no strong fundamental limit to AI capabilities was identified, though significant near-term barriers exist. Through 2027, AI is expected to remain an assistive tool rather than an autonomous driver of biological design.

A 2025 GovAI framework translates such capability findings into quantified risk estimates, suggesting that even modest capability increases could result in significant population-level harm when aggregated across potential actors.

Inference-Time Filters: A Demonstrated Weakness

The screening gap extends beyond DNA synthesis to the AI design layer itself. Protein-protein interaction (PPI) prediction tools have been proposed as inference-time biosafety filters: screen AI-generated designs for dangerous binding interactions before they reach synthesis ordering.

A 2025 study presented at the NeurIPS Biosecurity Safeguards for Generative AI workshop tested three leading PPI prediction tools (AlphaFold 3, AF3Complex, and SpatialPPIv2) against well-characterized viral-host interactions for Hepatitis B and SARS-CoV-2. The results:

  • None of the three tools detected any of four experimentally validated SARS-CoV-2 mutants with confirmed binding
  • Failures occurred even for viruses the models were trained on
  • The authors concluded the filters are “inadequate for reliably flagging even known biological threats and are even more unlikely to detect novel ones”

This finding reinforces the defense-in-depth argument: no single screening layer (DNA synthesis screening, model-level guardrails, or inference-time PPI filters) is sufficient alone. The authors argue for response-oriented infrastructure, including rapid experimental validation, adaptable biomanufacturing, and regulatory frameworks operating at AI development speeds, rather than relying on predictive filters that may provide false assurance (Feldman & Feldman, NeurIPS BioSafe GenAI Workshop, 2025, workshop paper).

Reality Check: The Wet Lab Failure Rate

Just because an AI designs a protein does not mean it will work. Published proxy studies report substantial failure during experimental validation.

The Hallucinated Binder Problem

Recent preprints testing generative design tools found that while they could generate thousands of candidate designs, the vast majority failed to bind their targets in actual wet lab experiments. Biology is governed by physics, water dynamics, temperature, and cellular context - factors that AI models still struggle to simulate.

NIST’s 2025 safe-proxy TEVV study makes this calibration more concrete. The team tested AI-assisted protein design against non-dangerous proxy proteins and found that predicted structural similarity did not reliably imply retained biological activity. Their conclusion was not that the screening gap is harmless, but that current systems are not yet reliable at rewriting a protein to both preserve activity and evade biosecurity screening (Ikonomova et al., 2025).

This result supports a narrower conclusion: predicted structure does not reliably confer biological activity, and hands-on experimental skill remains consequential. The size of that barrier depends on the task, actor, automation, and available support.

Responsible Development Practices

The AI biology community has increasingly adopted explicit biosecurity practices.

DeepMind: AlphaFold Biosecurity Assessment

Google DeepMind’s approach to AlphaFold3 included:

  • Pre-release biosecurity consultations with external experts
  • Publication of their risk analysis
  • Implementation of targeted server-side restrictions for potentially concerning use cases

EvolutionaryScale: Tiered Access

ESM3 was released with safeguards - a smaller open version (ESM3-open) with the full model available through a controlled API. This tiered approach allows broad scientific use while maintaining oversight of more capable versions.

The “If-Then” Strategy

Perhaps the most practical recommendation from the 2025 NASEM report is the proposed “if-then” strategy for ongoing assessment:

  1. Track data availability: The quality and quantity of training data determines capabilities. Monitoring what biological data becomes available provides early warning.

  2. Define capability benchmarks: Rather than vague concerns, establish specific testable thresholds. Can the model predict gain-of-function mutations? Design immune-evasive proteins?

  3. Establish trigger thresholds: When capabilities cross defined thresholds, predetermined responses activate - enhanced access controls, expanded screening, updated oversight.

  4. Regular reassessment: This should be continuous, not a one-time evaluation. As AI and biology both advance, the intersection requires ongoing attention.

Training Data Governance: The Biosecurity Data Levels Proposal

The if-then strategy’s first step, tracking data availability, points to an intervention that most governance frameworks have overlooked: controlling the training data itself, rather than the models or their outputs.

A February 2026 Science Policy Forum proposed a five-tier Biosecurity Data Level (BDL) framework for classifying pathogen datasets by their expected contribution to AI capabilities of concern (Bloomfield et al., 2026). More than 100 researchers from Johns Hopkins, Oxford, Stanford, Columbia, and NYU endorsed the approach at the 50th anniversary Asilomar Conference.

The five tiers:

Level Data Type Access Requirement
BDL-0 Vast majority of biological data No controls
BDL-1 General viral infection patterns Registered account with government-issued ID
BDL-2 Functional data on pandemic-capable virus properties (host range, environmental stability) Institutional affiliation check and bad-actor screening
BDL-3 Functional data linking genetic sequences to transmissibility, virulence, and immune evasion in human-infecting viruses Justified use case, mandatory trusted research environment, prepublication risk assessment
BDL-4 Data enabling AI design of enhanced pandemic pathogens All lower controls plus government pre-publication review

The empirical case for data-level governance is strengthening. ESM3 performed substantially worse on virus-related tasks when trained without viral protein data. Evo 2 (Brixi et al., 2026) showed similar capability degradation when eukaryote-infecting viral sequences were withheld, though fine-tuning can recover these capabilities with modest compute, reinforcing that data controls are necessary but not sufficient on their own.

Key design choices in the proposal:

  • Only new data would be restricted. Existing datasets remain untouched; restrictions apply only to data collected after governments adopt the framework.
  • No expert panel exists today to classify biological data by risk level. The framework requires building new institutional capacity.
  • International harmonization is unresolved. A framework adopted by one country but not others creates migration rather than reduction of risk.

The FY2026 NDAA’s Section 245, “Biological Data for Artificial Intelligence,” directs the Department of Defense to develop requirements for how biological data from DoD-funded research is collected and stored for AI use, including tiered cybersecurity safeguards and access controls. This is early federal recognition that data governance belongs in the biosecurity toolkit, though Sec. 245 addresses DoD-funded research specifically, not the broader open-science ecosystem the BDL framework targets.

Data-level governance sits upstream of model guardrails and synthesis screening: it reduces the capabilities a model acquires before deployment, rather than attempting to restrict outputs after training. Combined with the defense-in-depth approach (model evaluation, access controls, DNA synthesis screening), it addresses the full pipeline from data to physical capability.

Practical Governance Questions

Most readers will not be training protein models or building BDTs. Instead, you might be:

  • Sitting on an ethics or DURC committee
  • Advising a health ministry on AI investments
  • Reviewing funding proposals that mention “AI for biological design”
  • Contributing to international discussions on biosecurity norms

In those roles, you can ask sharp, concrete questions:

Checklist: Evaluating AI Biology Platforms

1. What kind of tool is this? - Structure prediction service, generative design platform, full BDT with synthesis integration, or a multi-scale simulator that models biology bottom-up from molecular data through organ-system or whole-body response? The last category is trained on biological data and can inform which molecules or biological perturbations are worth pursuing, so it belongs in this checklist even though it does not design sequences or structures directly (see Biological Design Tools vs. LLMs). - What user groups is it intended for?

2. What guardrails exist at the system level? - Is there identity verification for users? - Are high-risk features (direct synthesis ordering, batch design) limited to vetted entities? - Are logs kept, and who can audit them? - Are prompts/designs logged with anomaly detection and auditable retention to trace potential misuse?

3. How does this interact with existing biosafety? - Does the system assume synthesis screening at downstream providers? - Are users given biosafety guidance when designing constructs?

4. How is misuse potential evaluated? - Has anyone tested whether the tool can bypass current screening? - Is there a plan for repeat evaluations as the system updates?

5. What is the escalation path? - Who can suspend accounts or notify authorities? - Are there clear thresholds for concern?

These questions are practical tools for steering investments toward platforms that take biosecurity seriously.

Benefits for Biosecurity

The same tools that create dual-use concerns can strengthen defense. This symmetry is often lost in discussions focused exclusively on risks.

Accelerated Countermeasures: AI-enabled medical countermeasure development could dramatically accelerate responses to novel pathogens. During COVID-19, the unprecedented speed of vaccine development still took nearly a year. AI tools could compress timelines further - though wet-lab throughput, regulatory timelines, and manufacturing scale-up remain rate-limiting.

Improved Biosurveillance: AI can analyze genetic sequences from environmental samples and clinical cases to detect emerging threats earlier. Pattern recognition across vast datasets could identify unusual clusters that human analysis might miss.

Regulatory Genomics: AlphaGenome (Google DeepMind, 2025) processes up to one million DNA base pairs and predicts how sequence variants affect gene regulation, RNA splicing, and transcription factor binding across tissue types. Current applications are primarily defensive: understanding how noncoding mutations influence viral biology, identifying host genetic factors affecting pathogen susceptibility, and accelerating discovery of therapeutic targets for medical countermeasures. Dual-use relevance remains theoretical; engineering noncoding regulatory mutations to enhance pathogen transmissibility or virulence requires biological complexity that AlphaGenome, like current protein-design tools, has not demonstrated.

Enhanced Screening: SecureDNA uses computational methods including random adversarial threshold search and exact matching to short subsequences, enabling screening that catches not just known hazards but their predicted equivalents.

The challenge is ensuring that defensive applications keep pace with - or stay ahead of - potential offensive uses.

AI can also be applied defensively for biosecurity, accelerating detection, attribution, and countermeasure development (see AI for Biosecurity Defense).

What is AlphaGenome and how does it relate to biosecurity?

AlphaGenome (Google DeepMind, 2026) predicts how noncoding DNA variants affect gene regulation across tissue types, processing up to one million DNA base pairs. It extends AI-bio analysis beyond protein structure to the regulatory genome. Primary applications are defensive: understanding how noncoding mutations affect viral biology and accelerating countermeasure target discovery. Dual-use relevance remains theoretical.

What is the difference between AlphaFold and RFdiffusion?

AlphaFold predicts the 3D structure of existing proteins from their amino acid sequences. RFdiffusion generates entirely new protein designs from scratch based on user specifications. AlphaFold is a prediction tool; RFdiffusion is a creation tool. This distinction matters for biosecurity because generative design tools can create novel sequences that evade traditional screening.

Can AI currently design functional bioweapons?

AI-assisted generation of experimentally functional bacteriophages has been demonstrated in a constrained research setting (King et al., 2026). That result does not establish design of mammalian viruses, human pathogens, enhanced virulence or transmissibility, weaponization, or autonomous end-to-end execution. Most candidate designs still require extensive experimental validation.

What is the “screening gap” and why does it matter?

Traditional DNA screening includes sequence comparison against references of concern. Generative models can produce sequence-distant candidates predicted to retain selected properties, which can reduce sensitivity of homology-based tools. The relevant studies did not establish retained function for the physical proteins. Updated screening methods improved detection in the tested settings, while coverage remains incomplete.

What was the MegaSyn experiment and what did it prove?

The 2022 MegaSyn experiment inverted a drug discovery AI’s reward function from avoiding toxicity to maximizing it. In 6 hours on a consumer laptop, it generated 40,000 potentially toxic molecules including VX analogs and novel structures. This demonstrated that dual-use tools can be trivially repurposed for harm, though synthesizability was not validated and most designs would likely fail in practice.

Do inference-time PPI filters work for biosafety screening?

Evidence remains limited. In one workshop study, AlphaFold 3, AF3Complex, and SpatialPPIv2 failed to detect four experimentally validated SARS-CoV-2 binding mutants under the tested setup. That finding supports defense in depth, not a universal conclusion about every filter or interaction.


This chapter is part of The Biosecurity Handbook. For related content, see the previous chapters on AI as a Biosecurity Risk Amplifier and LLMs and Information Hazards.