AI as a Biosecurity Risk Amplifier
In July 2022, researchers at Collaborations Pharmaceuticals inverted their drug-discovery AI’s optimization objective from avoiding toxicity to maximizing it. In the published demonstration, the system generated about 40,000 candidate molecules predicted to be toxic, including structures related to known chemical-warfare agents and novel candidates. The study did not synthesize or otherwise validate those candidates, so it is evidence of rapid design-space exploration, not of a completed weapon pathway (Urbina et al., 2022).
Decision-makers who need the handbook-wide overview should start with the Executive Summary, which places AI uplift alongside DNA synthesis screening, cloud laboratory governance, and BWC implementation priorities.
- Differentiate between “Information Hazards” (LLMs) and “Design Hazards” (BDTs).
- Analyze the concept of “Tacit Knowledge” and why it acts as a primary barrier to AI-driven bioterrorism.
- Evaluate the findings of key empirical studies: RAND Red Team, OpenAI/Anthropic uplift assessments, and the Urbina toxic molecule generation experiment.
- Critique the “Uplift” metrics used by major AI labs to measure biosecurity risk.
- Assess DNA synthesis screening as a critical chokepoint for AI-enabled biological threats.
- Apply threat modeling frameworks to different actor categories (state, non-state, lone actors).
This chapter discusses biosecurity risks at a conceptual level appropriate for education and policy analysis. Consistent with responsible information practices:
- Omitted: Actionable protocols, specific synthesis routes, exact pathogen sequences
- Included: Risk frameworks, governance mechanisms, policy recommendations
For detailed biosafety protocols, consult your Institutional Biosafety Committee and relevant regulatory guidance.
Introduction: The AI-Biology Convergence
In July 2023, Anthropic CEO Dario Amodei warned in congressional testimony that AI could “greatly widen the range of actors with the technical capability to conduct a large-scale biological attack” within two to three years. Former UK Prime Minister Rishi Sunak similarly warned that AI could make it easier “to build chemical or biological weapons” and that “terrorist groups could use AI to spread fear and destruction on an even greater scale.”
These concerns are not merely hypothetical. President Biden’s October 2023 Executive Order 14110 on AI Safety explicitly tasked agencies with assessing AI-enabled biosecurity risks. The order notably established a lower compute threshold for models trained primarily on biological data, recognizing these systems warrant heightened oversight.
However, we must be careful not to confuse speed with possibility.
The media often portrays AI as a tool that will allow a teenager in a basement to engineer a pandemic. This narrative is dangerously distracting. The reality is more nuanced: AI is a risk amplifier. It takes existing capabilities and lowers the cost, time, and expertise required to execute them, but it does not (yet) solve the fundamental physical challenges of biology.
Anthropic’s survey of 80,508 AI users across 159 countries provides empirical context for these fears: 13% of respondents cited malicious use (cyberattacks, weapons, bioweapons) as a primary concern, ranking eighth among thirteen concern categories, behind unreliability (26.7%), job displacement (22.3%), and loss of autonomy (21.9%) (Anthropic, 2025). Governance gaps ranked fifth at 14.7%. The gap between public perception (malicious use as a moderate concern) and expert assessment (biosecurity as a critical frontier risk) underscores why governance cannot rely on public pressure alone to drive adequate safeguards.
How AI Lowers Barriers: Two Classes of Risk
Jonas Sandbrink’s influential 2023 preprint differentiates two classes of AI tools that pose biosecurity risks: large language models (LLMs) and biological design tools (BDTs). This distinction is crucial because these tool types create different risk profiles and require different mitigation strategies.
Large Language Models (LLMs)
Frontier LLMs are trained on natural-language data that include scientific literature. According to Sandbrink, LLMs can “democratize access to biological knowledge” with possible implications for misuse. The potential mechanisms include:
- Information Access: Synthesizing and explaining complex dual-use biological concepts in accessible language
- Planning Assistance: Helping structure approaches to biological experiments or agent acquisition
- Lab Assistance: Providing troubleshooting guidance for laboratory procedures
- Synthesis Evasion: Potentially helping actors circumvent DNA synthesis screening protocols
A classroom exercise reported in a preprint found that LLMs produced dual-use planning and sourcing suggestions during a biosecurity course. The exercise did not test whether participants could execute a physical workflow or overcome biosafety, access, tacit-knowledge, or synthesis-screening barriers.
However, the operational significance of this information access remains contested.
Biological Design Tools (BDTs)
While LLMs get the headlines, Biological Design Tools (BDTs) warrant greater concern. These are AI systems trained specifically on biological data, including protein structures, chemical properties, and genomic sequences, rather than text.
If LLMs are “Google on steroids,” BDTs are “calculators for biology.” They can predict how a protein folds (AlphaFold), how a molecule binds to a receptor, or how to design novel proteins with specific functions (RFdiffusion).
Notable examples include:
- AlphaFold/AlphaFold3: DeepMind’s protein structure prediction tools; Demis Hassabis and John Jumper shared the 2024 Nobel Prize in Chemistry for this work, alongside David Baker for computational protein design
- RFdiffusion: Enables de novo design of protein structures and functions
- ESM3 and Similar Models: Bridge gaps between sequence, structure, and function
- LigandMPNN: Atomic context-conditioned protein sequence design
Unlike LLMs, which primarily lower barriers for less sophisticated actors, BDTs could enable creation of agents “substantially worse than anything seen to date” by expanding the capabilities of already-sophisticated actors. The RAND Europe Global Risk Index (2025) found that AI-enabled biological tools’ “dual-use nature could lower barriers to biological weapon development or raise the ceiling of potential harm by enabling the design of novel biological agents.”
Key risk characteristics of BDTs include:
- Predictive assistance: Can reduce time spent prioritizing candidates, but performance is task- and dataset-dependent and still requires experimental validation
- Novel Design Capabilities: Enable creation of pathogens more transmissible, virulent, or capable of evading countermeasures
- Open-Source Proliferation: Unlike frontier LLMs controlled by major companies, many high-risk BDTs are released open-source. The 2025 Global Risk Index for AI-enabled Biological Tools assessed 57 advanced tools across eight functional categories, scoring each for misuse-relevant capabilities using Red/Amber/Green ratings. 13 tools were flagged as “Red” requiring action, and 61.5% of these are fully open-sourced
- Lower Compute Requirements: BDTs can be trained with fewer computational resources than frontier LLMs
For infectious-disease data, the most actionable control point is often the dataset itself. If a model is trained on pathogen sequences, host range, virulence, or immune-evasion data, the risk profile changes before any prompt or output exists. The Biosecurity Data Levels proposal in AI-Enabled Pathogen Design spells out a tiered approach to those biosecurity data-level controls and offers a detailed current framework for infectious-disease data-level controls.
This is a useful case study for understanding design-stage dual-use risk.
The Context: Researchers at Collaborations Pharmaceuticals used an AI model called “MegaSyn,” designed to avoid toxicity in drug discovery. For a biosecurity conference presentation, they simply flipped the logic - instead of penalizing toxicity, they asked the AI to maximize it.
The Result: In less than 6 hours, running on a standard consumer laptop, the AI generated about 40,000 candidate molecules predicted to be toxic. The list included structures related to known chemical-warfare agents and novel candidates, but the study did not synthesize or functionally validate them.
The Implication: The design barrier for chemical weapons has collapsed. However, the synthesis barrier remains. Knowing the structure of a novel nerve agent is dangerous, but you still need the precursors and laboratory to synthesize it.
The MegaSyn experiment illustrates rapid design-space exploration. Complementary molecular-property models can prioritize candidates before synthesis. One benchmark study reported balanced accuracies in the 80–95% range for selected acute-toxicity endpoints, but those results are task-specific and are not a direct comparison with animal-test reproducibility (Luechtefeld et al., Toxicological Sciences, 2018).
The MegaSyn case has policy implications that existing dual-use research of concern (DURC) frameworks have not absorbed. Bobier et al., 2025, in Journal of Medical Ethics, argue that current DURC policy fails to account for AI-driven pharmaceutical and chemical design and call for broadening DURC scope to cover the category of AI-accelerated work the MegaSyn experiment exemplifies. See Mapping DURC Categories to AI Capabilities for how the seven Fink Report categories now have direct AI parallels.
Comparing LLMs and BDTs: Risk Profiles
| Dimension | LLMs | BDTs |
|---|---|---|
| Primary Risk | Lower barriers for novices | Expand capabilities of sophisticated actors |
| Training Data | Natural language, internet text | Biological sequences, protein structures |
| Access Model | Mostly commercial APIs with guardrails | Largely open-source |
| Regulation Focus | Compute thresholds, safety alignment | DNA synthesis screening, structured access |
| Current Evidence | Task-dependent uplift; the largest observed uplift is on bounded digital tasks, not complete physical execution | Limited public testing; risks remain speculative |
The “Tacit Knowledge” Gap
The primary fear regarding LLMs is that they will serve as “super-mentors” for bioterrorism. To evaluate this, we must understand tacit knowledge - the unwritten, hands-on skills learned only through physical practice.
Examples of tacit knowledge in biology:
- How to pipette a liquid without creating aerosols
- How to visually identify a healthy cell culture
- How to troubleshoot contamination in real-time
- When an experiment “looks wrong” despite correct protocols
Tacit laboratory knowledge remains incompletely captured in written instructions because competence depends partly on physical practice, feedback, and situated judgment. Detailed model output does not by itself confer laboratory proficiency.
The Sociology of Tacit Knowledge
The significance of tacit knowledge barriers is not merely intuitive; it is well-established in the sociology of scientific knowledge. Harry Collins’ foundational research demonstrated that laboratory skills transfer only through direct social contact with practitioners, not through written protocols alone. His studies of laser construction showed that “only those who had significant social contact with successful laser builders could do the job,” regardless of access to written instructions (Collins, Science Studies, 1974).
This work supports a broader distinction between written instructions and contributory expertise, the ability to perform an activity competently. Detailed protocols can omit situated judgments that are learned through supervised practice.
The biosecurity implication is narrower: AI-generated instructions do not by themselves establish physical competence or successful execution. The size of the residual barrier depends on the task, actor, tools, supervision, and automation. OpenAI’s 2024 evaluation likewise concluded that information access alone was insufficient under its tested conditions.
However, emerging multimodal AI capabilities may challenge this barrier. Vision-enabled systems can observe laboratory technique via video and provide interactive feedback. Whether this translates to meaningful erosion of tacit-knowledge barriers remains an open question.
Reported case study: In December 2025, OpenAI reported a collaboration in which GPT-5 proposed protocol modifications, analyzed results, and refined experimental plans while human scientists executed the physical work. The reported efficiency improvement is a company case-study result, not an independently replicated estimate. AI-directed experimental design with human execution is not autonomous pathogen work.
Theoretical implication: This represents a step beyond information access toward operational capability. If AI systems can effectively coach physical laboratory tasks in real time, the tacit knowledge barrier may erode faster than previously anticipated. However, the demonstrated capability (AI-driven experimental design with human execution) differs from the hypothetical concern (AI enabling untrained actors to execute dangerous protocols independently). The gap between these scenarios remains significant.
Case Study: The RAND Corporation Red Team Analysis (2024)
In a landmark study, RAND Corporation researchers conducted a controlled experiment to assess whether LLMs actually helped malicious actors plan biological attacks.
The Setup: They recruited multiple “Red Teams” and gave some access to the open internet only, while others had access to the internet plus an LLM.
The Task: Plan a viable biological attack.
The Result: The study found no statistically significant difference in the quality or viability of the plans generated with AI assistance versus those without.
The Takeaway: The LLMs were helpful for brainstorming, but they did not provide the “secret sauce” needed to bypass technical hurdles. The information was already available through internet search; the AI just summarized it faster.
Read the full report: Mouton, C. A., Lucas, C., & Guest, E. “The Operational Risks of AI in Large-Scale Biological Attacks.” RAND Corporation, 2024
The “Uplift” Debate: What AI Labs Have Found
While RAND found no significant uplift for attack planning, recent technical reports from AI labs suggest capabilities are advancing:
OpenAI (2024):
The GPT-4 bio-risk evaluation observed mean uplift in accuracy of 0.88 out of 10 for experts with GPT-4 access, but differences were not statistically significant. The study concluded that GPT-4 provides “at most a mild uplift” over standard search engines.
Anthropic (2024):
The Claude 3 Model Card reported that Claude 3 models “substantially increased risk in certain parts of the bioweapons acquisition pathway” for novices but did not appear capable of uplifting experts “to a substantially concerning degree.” This led to development of their ASL safety framework.
Zhang et al. (2026):
A preprint evaluated 57 participants classified as biology novices across eight digital biology benchmark suites. A mixed-effects analysis estimated higher odds of correct responses with access to multiple frontier LLMs than with internet-only access (odds ratio 4.16, 95% CI 2.63–6.87). The heterogeneous cohorts, mixed assignment designs, potential benchmark exposure, and absence of physical laboratory outcomes limit generalization (Zhang et al., 2026, preprint). Detailed comparison with physical-world evidence appears in LLMs and Information Hazards.
DHS Assessment (2025):
A DHS-commissioned RAND study assessed risks where AI and chemical/biological weapons converge, concluding that mitigating these dangers requires coordinated effort across industry, government, and international stakeholders. The report emphasizes that developers must recognize dual-use potential in their research products, even when not designing with misuse in mind.
Critical Limitation:
Evaluation endpoints differ materially. Digital benchmarks can measure knowledge, coding, planning, or troubleshooting, but they do not establish whether a participant can execute a complex biological workflow in the physical world.
The studies cited above represent snapshots of specific model versions tested with specific evaluation methods. AI capabilities and deployment safeguards change rapidly, so the 2024 “mild uplift” results should be treated as baselines rather than permanent assurances.
OpenAI’s Preparedness Framework v2 defines a “High” biological and chemical capability threshold in terms of meaningful counterfactual assistance to novice actors, and Anthropic’s Responsible Scaling Policy likewise ties safeguards to capability thresholds. These are developer-defined governance criteria, not independent findings that a particular model has crossed a biological-weapon threshold.
Continuous monitoring and updated evaluations are essential. Citing only the 2024 results without caveat risks giving policymakers false security about a fast-moving target.
Quantifying Risk from Capability Evaluations
The uplift studies above tell us whether AI improves attack capability, but policymakers need to know how much that matters. A 2025 GovAI report by Luca Righetti provides a framework for translating capability evaluations into quantified risk estimates.
Uplift is a pathway-specific measurement, not a single property of a model. A defensible evaluation identifies the user population, baseline tools, task, assistance conditions, endpoint, and physical constraints. Better performance on knowledge questions or digital design tasks does not establish successful acquisition, synthesis, laboratory execution, dissemination, or population harm. Conversely, a null result on one end-to-end workflow does not prove that every narrower step is unaffected. Decision-makers should map each measured endpoint to the stage of the threat pathway it actually tests.
The analysis combines historical case studies, expert elicitation, and reference-class forecasting. Its estimates are conditional scenario outputs rather than measured probabilities. They should be used for sensitivity analysis, not as predictions of annual attack likelihood or expected deaths.
Six subject-matter experts and five superforecasters reviewed the methodology, finding similar median estimates. All forecasts displayed high uncertainty bands, and the authors note the continued need for better underlying evidence.
This framework matters for two reasons. First, it moves beyond qualitative terms like “mild uplift” to numbers that inform resource allocation and policy prioritization. Second, it highlights that even modest capability increases can translate to significant population-level risk when applied across many potential actors.
OpenAI’s 2024 study explicitly notes: “information access alone is insufficient to create a biological threat” and “studies of information access alone do not test for success in the physical construction of the threats.”
Physical constraints remain significant. Specialist training and access to well-resourced laboratories is critical. Estimates suggest the pool of individuals with both the technical skills and materials access to execute a sophisticated biological attack numbers in the tens of thousands globally - a significant barrier despite AI assistance.
Demonstrated (supported by published evidence):
- LLM uplift varies by endpoint: earlier planning studies reported no significant or mild uplift, while a 2026 preprint found substantial improvement across bounded digital biology tasks (Zhang et al., 2026, preprint)
- BDTs can generate toxic molecule designs rapidly (MegaSyn 2022 - 40,000 molecules in 6 hours)
- Tacit knowledge barriers remain significant for laboratory work
- Synthesis screening can detect many sequences represented in its references, with performance dependent on the tool, sequence, order, and review process
- Multimodal AI can provide real-time laboratory coaching via video
Theoretical (plausible but not yet demonstrated):
- AI-designed sequences evading function-based screening at scale
- Autonomous AI agents completing wet-lab work without human oversight (see Autonomous AI Agents)
- Digital-task uplift translating into successful end-to-end physical execution
- AI enabling lone actors to overcome tacit knowledge barriers entirely
Unknown (insufficient evidence to assess):
- Whether next-generation models will cross capability thresholds
- The true size of the population capable of exploiting AI for biological harm
- How quickly multimodal AI will erode tacit knowledge barriers
- Whether cloud laboratories will become accessible to malicious actors
This distinction matters for policy: demonstrated risks warrant immediate action, while theoretical risks require monitoring and contingency planning.
Grounding AI-Bio Risk in Real-World Actors and Constraints
An AI capability result is not, by itself, an estimate of biological risk. The August 2026 National Academy of Medicine workshop and the ongoing National Academies consensus study are examining actor capabilities, incentives, and pathway constraints, together with implications for medical-countermeasure readiness. These activities establish a current research agenda, not completed guidance or evidence that a projected threat pathway is feasible.
Threat assessment should begin with a defined actor, counterfactual baseline, pathway, and outcome. Vogel’s sociotechnical analysis argues that technical possibility must be assessed alongside scientific, organizational, economic, and infrastructural conditions (Vogel, 2013). Historical analysis of state and non-state biological programs likewise found that expertise, work organization, management, and secrecy constrained outcomes even when information and material resources were available (Ben Ouagrham-Gormley, 2012). Applying those findings to AI is an inference, not a measured estimate of current AI uplift.
| Assessment dimension | Decision question | Evidence rule |
|---|---|---|
| Actor baseline and intent | What expertise, people, infrastructure, access, incentives, and tolerance for failure exist without AI? | Describe a scenario rather than assigning a risk label to an actor category. |
| Marginal AI contribution | Which step becomes more accurate, faster, cheaper, or newly possible relative to a defined non-AI baseline? | Report the model, access conditions, task, participant population, comparison condition, and date. |
| Pathway completion | Can an actor move from model access and design through acquisition or synthesis, experimental validation, production, deployment, and population-level effect? | Success at one stage does not establish success at the next. Separate digital proxies from physical outcomes. |
| Binding constraints | Which tacit, organizational, material, validation, scale, delivery, detection, or response constraints remain? | Treat constraints as contingent and testable, not permanent barriers. |
| Outcome and uncertainty | What harm is being estimated, and how much of the pathway is directly observed? | Distinguish demonstrated capability, plausible but unvalidated pathways, and claims beyond current evidence. |
| Intervention points | Which controls reduce risk at the model, data, synthesis, laboratory, deployment, detection, or response stage? | Use overlapping safeguards and define evidence-based triggers for reassessment. |
A 2026 RAND report proposes one implementation of this logic by scoring both the potential biological modification and the capability of the actor who could use it (Williams et al., 2026). The method is not a validated prediction instrument. RAND states that its thresholds still require empirical data or expert consensus and calibration against real-world cases.
Actor Categories Are Scenario Inputs, Not Scores
| Scenario | Plausible AI contribution | Constraints and evidence posture |
|---|---|---|
| State or other well-resourced program | AI may assist analysis, candidate prioritization, or design-stage work. | Public evidence does not establish end-to-end uplift. Organizational reliability, experimental validation, production, strategic incentives, detectability, and consequences remain scenario-specific. |
| Organized non-state group | General-purpose models may assist early information and planning tasks. | Recruitment, tacit expertise, materials, facilities, validation, security, logistics, and delivery can remain binding. Historical program evidence should not be converted into a fixed present-day probability. |
| Lone actor or novice | Relative gains may be larger on bounded knowledge or digital tasks when the non-AI baseline is low. | Task-level uplift does not establish physical execution. Current studies are model-, participant-, and endpoint-specific and do not justify a general High, Medium, or Low rating. |
| Legitimate research organization | AI may increase research speed or workflow complexity in authorized settings. | Accident and misuse risk depends on biosafety, dual-use review, access controls, system interfaces, human oversight, and incident response rather than malicious intent alone. |
The policy implication is to evaluate scenarios, not stereotypes. A measured gain on an early-stage task should trigger closer assessment of the next stage and its safeguards, not an assumption that the full pathway either succeeds or fails. The red-teaming threat model provides the corresponding evaluation structure, while layered defense maps interventions across the full pathway.
DNA Synthesis Screening: The Critical Chokepoint
DNA synthesis represents a critical “digital-to-physical frontier” where AI-designed biological agents must be converted into physical materials. As a recent Science paper notes: “Synthesis of nucleic acids is a choke point in AI-assisted protein engineering pipelines.”
This makes screening of synthesis orders one of the most tractable intervention points.
The International Gene Synthesis Consortium (IGSC)
The IGSC, founded in 2009, is a voluntary coalition of synthetic DNA providers operating under a Harmonized Screening Protocol. Members together represent a majority of commercial gene synthesis capacity worldwide. The Protocol requires:
- Screening for sequences of concern (SOCs) matching regulated agents (e.g., Select Agents and Toxins)
- Customer verification before fulfilling orders
However, significant gaps remain. A study of DNA provider practices found “significant heterogeneity in security practice throughout the field, reflective of the current lack of codified oversight for DNA synthesis.”
Key limitations include:
- Voluntary Nature: No country legally requires nucleic acid synthesis screening
- Coverage Gaps: Non-IGSC providers and benchtop synthesis devices may not screen orders
- AI-Generated Sequence Diversity: Protein-design tools can generate sequence-distant candidates that may challenge homology-based screening; retained function must be established experimentally
- Order Splitting: Screening could be evaded by distributing orders across multiple providers or time periods
The 2024 OSTP Framework
In April 2024, the White House Office of Science and Technology Policy released the Framework for Nucleic Acid Synthesis Screening, implementing Section 4.4(b) of Executive Order 14110.
Key provisions include:
- Federal Funding Framework: The 2024 OSTP framework tied federal procurement conditions to compliant providers, but Executive Order 14292 directed that framework to be revised or replaced. Current legal effect should be verified against agency instructions.
- Enhanced Screening Window: By October 2026, providers should screen each 50-nucleotide window for SOCs (reduced from 200 bp)
- Expanded SOC Definition: Includes sequences “known to contribute to pathogenicity or toxicity, even when not derived from or encoding regulated biological agents”
- Manufacturer Requirements: Extends screening expectations to benchtop nucleic acid synthesis equipment manufacturers
The 2024 OSTP Framework remains the federal funding baseline while agencies revise or replace it. Executive Order 14292 directed OSTP to incorporate enforcement terms and to address non-federally funded research. ASPR’s current status page says that agencies will revise or replace the framework and will post an update when a successor is available. As of August 3, 2026, no successor is identified on that page. The order and later legislative proposals establish policy direction, not a substitute for implementing rules. For current practice, follow applicable award terms, institutional requirements, and provider screening standards while monitoring agency guidance.
Model-Level Safeguards and Evaluations
Frontier AI labs have begun implementing model-level safety strategies.
Capability Post-Training Is Not Safety Post-Training
Reinforcement learning can increase capabilities when the training objective has a reliable, automatically checkable outcome. DeepSeek-R1 provides a peer-reviewed example in mathematics and coding. Its authors also report that model-based rewards are more susceptible to reward hacking and that reliable reward signals are difficult to construct for open-ended tasks (Guo et al., 2025). This is evidence about reasoning post-training, not a demonstration that verifier-based rewards improve biological safety.
Process supervision can expose intermediate errors that an outcome-only score misses. In an ICLR study, process-supervised reward models outperformed outcome-supervised models on the MATH dataset (Lightman et al., 2024). That result is domain-specific. Biological validity often depends on incomplete measurements, tacit experimental context, and delayed physical outcomes, so a plausible reasoning trace or final digital answer is not a sufficient verifier.
For biosecurity evaluations, separate four questions:
| Layer | Evaluation question |
|---|---|
| Capability post-training | Did the intervention increase task performance, planning, tool use, or error recovery? |
| Safety-behavior post-training | Did refusals and policy adherence improve without unacceptable over-refusal or capability leakage? |
| Reward or judge validity | Does the verifier agree with independent expert or physical evidence, including rare failure cases? |
| Deployment safeguards | Do classifiers, permissions, access controls, monitoring, and human gates work around the post-trained model? |
Improved performance against a verifier is not evidence of improved safety. Re-evaluate all four layers whenever the base model, post-training data, reward, system prompt, tools, or deployment controls change.
DeepMind: AlphaFold Biosecurity Assessment
DeepMind assembled a multidisciplinary panel to assess AlphaFold 3’s biosecurity implications, concluding it “does not significantly elevate risk compared to prior structure prediction tools” while committing to explore additional safeguards. AlphaFold3 was among early adopters of experimental refusal mechanisms, though “these efforts were preliminary and highlighted the challenges of balancing functionality with security.”
Anthropic: AI Safety Levels
Anthropic has developed AI Safety Levels (ASL) tied to capability thresholds. Claude 3’s biological capabilities informed development of their ASL-3 protections.
OpenAI: Preparedness Framework
OpenAI’s Preparedness Framework v2 (April 2025) defines two thresholds for the combined biological and chemical risk category:
- High: The model can provide “meaningful counterfactual assistance (relative to unlimited access to baseline of tools available in 2021) to ‘novice’ actors (anyone with a basic relevant technical background) that enables them to create known biological or chemical threats.” Models at High capability must have safeguards sufficiently minimizing harm before deployment.
- Critical: The model introduces unprecedented new pathways to severe harm with no ready precedent under the threat model. Critical capability systems require safeguards during development, not only at deployment.
The “2021 baseline” anchors the evaluation operationally: the question is not whether a model provides any biological or chemical information, but whether it provides meaningful uplift beyond what was achievable without AI in 2021. Given the higher potential severity of biological threats, OpenAI uses biological evaluations as indicators for High and Critical thresholds across the combined bio/chemical category. The OPCW Scientific Advisory Board’s Temporary Working Group on AI (SAB/REP/1/26, March 2026) identified this as a gap leaving chemical-specific risks undertested. See Chemical Threats: A Distinct Evaluation Gap for analysis.
Government Evaluation Efforts
The UK AI Security Institute (renamed from AI Safety Institute in February 2025) and U.S. Center for AI Standards and Innovation (CAISI) (renamed from AI Safety Institute in June 2025) are developing biorisk-related tests and guidance for advanced AI models. CAISI’s pre-deployment testing program expanded in May 2026 when Google DeepMind, Microsoft, and xAI signed agreements joining OpenAI and Anthropic in providing frontier models (including versions with reduced safeguards) for government evaluation; CAISI reports having completed more than 40 such evaluations to date, with assessments conducted through the interagency TRAINS Taskforce. Key evaluation approaches include:
- Virology Capabilities Test (VCT): Multiple-choice questions measuring AI troubleshooting of complex virology protocols. In limited benchmarks, frontier models have scored comparably to or above domain experts on certain question types, though the operational significance of these scores remains debated
- Human Uplift Trials: Studies measuring whether AI access improves human performance on biosecurity-relevant tasks
- Red-Team Exercises: Experts role-playing as threat actors to assess operational feasibility
- WMDP-Bio Benchmark: The Center for AI Safety’s Weapons of Mass Destruction Proxy benchmark includes 1,273 biosecurity-specific multiple-choice questions designed to measure hazardous biological knowledge in LLMs. The benchmark also evaluates unlearning methods (such as Representation Misdirection for Unlearning) that attempt to remove dangerous knowledge while preserving general capabilities (Li, Mazeika, Hendrycks et al., 2024)
International AI Safety Report 2026
The International AI Safety Report 2026 is an evidence synthesis prepared with guidance from more than 100 experts nominated by over 30 countries and intergovernmental organizations. Its biological-risk section reports that general-purpose systems can provide information about biological and chemical weapons development, including expert-level laboratory instructions, while noting that material and operational barriers remain difficult to assess. It also records that several developers added safeguards in 2025 after pre-deployment testing could not rule out assistance to novice weapon developers. The report does not establish a universal capability threshold or an independently validated survey of all biological AI tools, so those stronger claims should not be attributed to it.
Evaluating Safety Measures: The Evo 2 Case Study
While model-level safeguards are increasingly common, their robustness under adversarial conditions remains underexplored. Evo 2 (Brixi et al., 2026), a genomic foundation model from Arc Institute and NVIDIA trained on over 9 trillion nucleotides, provides an instructive case study. The developers deliberately excluded eukaryotic viral sequences from training data to prevent the model from acquiring capabilities relevant to human pathogen design.
A 2025 Scale AI/SecureBio preprint introduced BIORISKEVAL, a framework for testing whether such data filtering actually works against determined adversaries. The framework evaluates bio-foundation models across three dimensions: sequence modeling capability, mutational effect prediction, and virulence prediction.
The findings suggest data filtering is necessary but not sufficient:
- Rapid capability recovery via fine-tuning: When researchers fine-tuned Evo 2 on related viral sequences, the model generalized to the filtered virus types within approximately 50 training steps (less than 1 H100 GPU hour). Inter-genus generalization required more compute but remained achievable.
- Latent knowledge persists despite filtering: Linear probing of Evo 2’s hidden layer representations achieved 0.46 Pearson correlation for virulence prediction, even without any fine-tuning, suggesting the model acquired predictive signals during pretraining despite data exclusion.
- Modest current risk: The authors emphasize that Evo 2’s predictive capabilities remain too modest for reliable weaponization. Its Spearman correlation of approximately 0.2 for mutational effect prediction is far below the threshold for practical misuse.
The broader evidence base also cautions against treating aggregate protein-model benchmarks as proof of viral generalization. Across 41 viral and 33 cellular deep-mutational-scanning datasets, a 2026 preprint found lower average performance for supervised protein-language-model predictors on viral datasets. A site-mean baseline matched or exceeded supervised models on many datasets, and performance declined when mutations from the same site could not appear in both training and test sets. The analysis was limited to existing single-substitution datasets and does not establish that protein language models generally fail on viral biology (Vieira et al., 2026, preprint).
A subsequent GovAI experiment extended these findings on the dimension of accessibility. In March 2026, a researcher with AI engineering experience but no prior biology background used Claude Code to independently replicate the fine-tuning approach, constructing a dataset of human-infecting viral sequences from NCBI RefSeq and running training using code from a concurrent bacteriophage fine-tuning study (King et al., 2025, preprint). Perplexity on training-set sequences improved significantly (median 3.61 to 2.44, p<2.2e-16); improvement on the held-out test set (3.60 to 3.09) did not reach statistical significance (p=0.10). The agent required no biological expertise from the user and encountered no refusals. Total cost was approximately $760 in compute, completed over one weekend (Righetti, Lukosiute, and Black, GovAI, April 2026). The authors note that other public models already outperform this fine-tuned version on relevant tasks, limiting immediate risk. The experiment illustrates how LLM coding agents can erode safeguards premised on fine-tuning difficulty, independently of biological expertise.
These results inform the broader “defense-in-depth” principle: no single safety measure should be assumed robust against adversarial manipulation. Data filtering reduces default capabilities but does not eliminate them when model weights are publicly available. For BDTs released as open-weight models, additional safeguards including DNA synthesis screening and access controls remain essential. For the governance framework proposing how to implement data-level controls at scale, see Training Data Governance: The Biosecurity Data Levels Proposal.
Public Health Implications and Preparedness
For public health practitioners, AI-biosecurity risks intersect with existing pandemic preparedness and outbreak response frameworks.
Implications for Surveillance and Detection
AI-designed pathogens could potentially evade existing detection and surveillance systems:
- Genomic Surveillance: Novel sequences may not match reference databases used for pathogen identification
- Syndromic Surveillance: Engineered agents with altered clinical presentations may evade pattern recognition
- Attribution: Distinguishing natural emergence from deliberate release becomes more challenging with AI-optimized designs
Countermeasure Development
Paradoxically, the same AI capabilities that enable threat creation could accelerate countermeasure development. AI-assisted drug discovery and vaccine platform technologies could potentially respond to threats faster than traditional approaches.
However, the asymmetry between offense and defense in biological threats - where creating harm is often easier than preventing it - remains a fundamental challenge. For connections to defensive AI applications, see AI for Biosecurity Defense.
Recommendations for Practitioners and Policymakers
Based on the current evidence and expert recommendations from NTI, CNAS, and the Federation of American Scientists, a multi-layered approach is recommended:
For Public Health Practitioners
- Integrate AI-biosecurity awareness into pandemic preparedness planning - Consider scenarios involving AI-designed or AI-optimized biological agents in tabletop exercises
- Strengthen genomic surveillance capabilities - Ensure systems can detect novel sequences that may not match existing databases
- Engage with dual-use research governance - Participate in Institutional Biosafety Committee (IBC) oversight and stay informed about emerging AI tools
- Build relationships with biosecurity experts - Establish connections with organizations like Johns Hopkins Center for Health Security, NTI | bio, and CSET before crises occur
For Policymakers
- Mandate DNA synthesis screening - Move beyond voluntary frameworks to legally require screening for all commercial synthesis providers
- Fund third-party AI evaluations - Support independent assessment of biological AI tools before release
- Develop BDT-specific governance - Recognize that regulations designed for LLMs (compute thresholds) may not adequately address BDTs
- Strengthen international coordination - Engage with the Biological Weapons Convention, Australia Group, and WHO frameworks to harmonize global biosecurity norms
- Invest in defense as well as prevention - Current U.S. biodefenses are insufficient to address large-scale biological threats; AI could help accelerate countermeasure development
- Expand LLM safety tests to include BAIM modification - Current evaluations assess whether LLMs provide dual-use biological information or help users operate BAIMs; few test whether LLMs help modify BAIMs by removing safeguards via fine-tuning, or help build new narrow-capability BAIMs. These represent distinct and growing risk channels (Righetti, Lukosiute, and Black, GovAI, April 2026)
The next chapters in Part IV examine specific aspects of the AI-biosecurity landscape: LLMs and information hazards, AI-enabled pathogen design, and defensive AI applications.
How does AI amplify biosecurity risks?
AI amplifies biosecurity risks by lowering barriers to accessing dual-use biological knowledge and accelerating design capabilities. Large Language Models democratize access to dangerous information while Biological Design Tools enable sophisticated actors to design novel threats. However, AI amplifies existing risks rather than creating fundamentally new ones.
What is the difference between LLMs and BDTs in biosecurity contexts?
LLMs (Large Language Models) lower barriers for novices by making dual-use biological knowledge more accessible through natural language synthesis. BDTs (Biological Design Tools) raise the ceiling of what sophisticated actors can achieve by enabling novel pathogen design using protein structure prediction and biological optimization. LLMs primarily affect information access, while BDTs enable capability expansion.
What did RAND and OpenAI studies find about AI biosecurity uplift?
RAND 2024 found no statistically significant difference in attack-plan viability, and OpenAI 2024 reported at most mild uplift with GPT-4. Anthropic 2024 reported novice uplift in certain acquisition steps. A 2026 preprint found substantially higher performance across bounded digital biology tasks (Zhang et al., 2026, preprint), while a separate randomized trial found no significant increase in full physical workflow completion (Hong et al., 2026, preprint).
What is DNA synthesis screening and why does it matter for AI biosecurity?
DNA synthesis screening represents a critical “digital-to-physical frontier” where AI-designed biological agents must be converted into physical materials. The 2024 OSTP Framework mandated screening for federally funded research, though implementation faces uncertainty after 2025 policy changes. This chokepoint remains one of the most tractable intervention points despite potential AI-enabled evasion strategies.
What are biosecurity data-level controls for infectious-disease AI?
Biosecurity data-level controls classify pathogen and infectious-disease datasets by their likely contribution to dangerous capabilities, then apply stronger access requirements to higher-risk tiers. The basic idea is to manage risk upstream, before model training, not just at the output layer. See the Biosecurity Data Levels proposal for a tiered framework.
Why does dataset governance matter in biosecurity?
Because filtering outputs after training is too late if the model has already learned risky signals from the data. Data governance reduces the capabilities a model acquires in the first place, and works best alongside synthesis screening and access controls.
This chapter is part of The Biosecurity Handbook. For handbook-wide priorities, see the Executive Summary. For related content, see Dual-Use Research of Concern, DNA Synthesis Screening, and Red-Teaming AI Systems.