Executive Summary

Purpose

Classical biosecurity frameworks and emerging AI-biological risks occupy separate policy silos despite growing overlap. The evidence base here includes peer-reviewed research, government documents, and technical assessments, with recommendations for policymakers, institutional biosafety committees, AI labs, and public health professionals.


Key Findings

Classical Biosecurity Gaps Persist

The Biological Weapons Convention lacks a standing verification system. The treaty has no standing independent inspection or monitoring body and no automatic sanctions for violations. Article VI permits complaints to the UN Security Council, and voluntary Confidence-Building Measures provide limited transparency. Major historical violations were discovered through defectors, intelligence, or post-conflict inspections rather than routine BWC mechanisms.

DNA synthesis screening provides meaningful but incomplete protection. The IGSC Harmonized Screening Protocol v3.0 states that consortium members represent a majority of global commercial gene-length nucleic acid synthesis capacity. Critical gaps remain across providers, order types, and decentralized synthesis. Post-disclosure testing found that updated screening methods greatly improved detection of AI-generated protein variants (Wittmann et al., 2025).

Laboratory biosafety depends on institutional culture, not just containment. High-consequence pathogens require BSL-3 or BSL-4 facilities, but accidents occur when norms erode and procedures become routine. The 1979 Sverdlovsk anthrax outbreak killed at least 66 people after an accidental release from a Soviet military microbiology facility (Meselson et al., 1994).

AI Changes Biosecurity Risk

AI-bio evidence must be interpreted by what it actually measures. The 2024 RAND study found no statistically significant increase in biological attack-plan viability under its tested conditions (Mouton et al., 2024). A 2026 preprint found substantially higher novice performance across bounded digital biology tasks, while a separate preregistered randomized trial found no significant difference in full-workflow completion for novice participants using mid-2025 models (Zhang et al., 2026, preprint; Hong et al., 2026, preprint). These studies do not establish a current capability ceiling. Developer assessments now report higher biological capability classifications for frontier systems, but capability scores, human uplift, safeguard performance, and real-world consequence remain separate claims that require separate evidence. Digital benchmark performance should not be treated as evidence of end-to-end physical capability. See Red-Teaming AI Systems.

Tacit laboratory knowledge remains a consequential barrier. Multimodal models and cloud laboratories create plausible pathways for AI-assisted observation, troubleshooting, and remote execution. Current evidence does not establish that these systems materially overcome tacit laboratory-skill barriers in high-consequence workflows.

Decentralized biological production creates a distinct governance challenge. DNA Script’s SYNTAX system provides on-site benchtop oligonucleotide synthesis, Ansa Biotechnologies is a centralized synthesis provider, and Nuclera’s eProtein Discovery is a benchtop protein-production platform. Governance analysis should distinguish decentralized synthesis from centralized screened services and protein production. The IBBIS Common Mechanism provides open-source screening tools, but adoption and legal coverage remain uneven.

Autonomous AI agents introduce new attack surfaces. AI agents can coordinate multi-step computational and laboratory-adjacent workflows through tool use. Specialized autonomous laboratories execute constrained workflows, but current evidence does not establish general autonomous execution of complex biological work with minimal human oversight. Cloud laboratory APIs, LIMS systems, and synthesis-ordering interfaces expand the systems-security attack surface. See Autonomous AI Agents for governance frameworks including the “least agency principle.”

International Governance Remains Fragmented

UNSCR 1540 addresses non-state actor proliferation but implementation varies. UN Security Council Resolution 1540 (2004) creates binding obligations to prevent non-state actors from accessing weapons-of-mass-destruction materials. By August 2025, 185 states had submitted a first national report, while eight had not; implementation quality and enforcement capacity also vary.

The Australia Group harmonizes export controls among 42 participating countries plus the European Commission. Participating states coordinate controls on biological agents, dual-use equipment, and related technologies. China, Pakistan, and Russia are among the nonparticipants, although nonparticipants may maintain sovereign export controls or use Australia Group control lists.

The Chemical Weapons Convention demonstrates that standing verification can deliver measurable disarmament, with limits. As of December 31, 2025, the OPCW reported that all declared chemical weapons stockpiles had been verifiably destroyed (OPCW, 2026). Its May 2026 discovery of previously undeclared chemical weapons and related material in Syria also shows that verification of declared holdings does not ensure complete declarations (OPCW, 2026). See The Chemical Weapons Convention and OPCW.

Preparedness gaps persist across countries. The Global Health Security Index describes capacities across prevention, detection, response, and health-system resilience, but the 2019 index ranking was not a validated predictor of COVID-19 outcomes in OECD countries (Abbey et al., 2020).


Priority Recommendations

For Policymakers

Establish binding, risk-based synthesis screening requirements. The April 2024 OSTP Framework tied procurement conditions to federally funded research, and Executive Order 14292 (May 2025) directed that the framework be revised or replaced. ASPR states that agencies will revise or replace the 2024 framework, so it should not be treated as a settled current mandate. Congress responded in January 2026 with the bipartisan Biosecurity Modernization and Innovation Act, introduced but not enacted. Requirements should address:

  • Screening requirements that cover relevant order lengths, providers, and delivery modes
  • Embedded screening in benchtop synthesis devices before market authorization
  • Manufacturer-maintained audit logs for attribution and pattern detection
  • International coordination to prevent jurisdiction shopping

Establish AI-biosecurity evaluation standards. Current AI-bio uplift studies use heterogeneous methods and span attack planning, information access, proxy tasks, and limited physical-world validation. They do not establish operational risk across actors, tasks, models, and configurations. Require:

  • Threat models that specify actors, access, baseline resources, and consequential outcomes
  • Separate assessment of capability, human or agent uplift, safeguard effectiveness, operational consequence, and lifecycle governance
  • Configuration-specific testing of deployed systems, including tools, retrieval, memory, and monitoring
  • Statistically valid designs with predeclared estimands, uncertainty, and matched baselines
  • Independent review with disclosed access, conflicts, redaction rights, and unresolved limitations
  • Continuous monitoring, remediation, and regression testing after deployment

Strengthen BWC implementation without waiting for verification. Political consensus for binding verification does not exist. Instead:

  • Increase ISU funding and staff to support States Parties implementation
  • Require national biosecurity strategies as BWC compliance measure
  • Establish regional Centers of Excellence for capacity building
  • Create consequences for non-participation in Confidence-Building Measures

Establish governance frameworks for cloud laboratories. Cloud labs enable remote experiment execution without physical presence, creating accountability gaps:

  • Require API access controls and customer verification for sensitive protocols
  • Mandate audit logging of experiments involving select agents or dual-use sequences
  • Establish middleware security standards for autonomous agent interactions
  • Integrate cloud lab protocols with institutional biosafety committee review

Integrate One Health across biosecurity governance. Jones et al. found that 60.3% of 335 emerging infectious-disease events were zoonotic (Jones et al., 2008); WHO states that more than 60% of reported emerging infectious diseases originate in animals (WHO One Health). Require:

  • Cross-sector coordination between human medicine, veterinary, and environmental health
  • Surveillance systems that track pathogens across species interfaces
  • Biosafety regulations that address animal reservoirs and agricultural biosecurity

For AI Laboratories

Implement responsible scaling policies with evidence-linked thresholds. Capability thresholds should trigger defined controls, but a threshold classification does not by itself establish safeguard adequacy or operational risk. AI laboratories should:

  • State the evaluated model, system configuration, access conditions, and baseline
  • Test both hazardous-request restriction and legitimate scientific access
  • Evaluate adversarial robustness under declared human and automated attack budgets
  • Link findings to deployment controls, monitoring, incident response, and reassessment
  • Publish methods, uncertainty, and limitations without releasing reusable attack material

Restrict frontier model access for high-risk use cases. Deploy tiered access controls:

  • Enhanced verification for users requesting biological sequence generation
  • Usage monitoring for repeated dual-use queries
  • Integration with DNA synthesis screening databases where technically feasible
  • Risk-tiered governance of pathogen training data, targeting capability formation upstream of model release

Support defensive biosecurity research. AI capabilities can strengthen detection and countermeasures:

  • Pathogen genomic surveillance and variant tracking
  • Medical countermeasure design and optimization
  • Metagenomic sequencing analysis for outbreak investigation

For Institutional Biosafety Committees

Apply current high-risk research oversight across relevant funding and institutional pathways. On July 28, 2026, the federal government issued the USG Policy for Stopping High-Risk Life Sciences Research. Agencies were directed to issue implementation guidance within 120 days, and NIH stated that flagged research would remain paused during implementation (NIH NOT-OD-26-101). Institutions should:

  • Review privately funded high-risk life sciences research under applicable institutional authorities
  • Assess AI-assisted experimental design for dual-use risks
  • Evaluate cloud laboratory use for sensitive protocols
  • Train relevant personnel to assess information hazards

Establish physical security standards beyond containment. Insider threats require more than BSL-level containment:

  • Multi-person verification for select agent access
  • Inventory management with real-time tracking
  • Personnel reliability programs for high-consequence pathogen work
  • Behavioral monitoring for concerning pattern recognition

Require incident reporting and lessons learned. Laboratory accidents and near-misses provide critical safety information:

  • Mandatory reporting of all exposures and containment breaches
  • Anonymized sharing of incident analyses across institutions
  • Integration with national biosafety databases
  • Regular safety culture assessments to identify normalization of deviance

For Research Institutions

Implement responsible information practices for dual-use findings. Publication of dangerous biological knowledge requires risk assessment:

  • Pre-publication review of manuscripts describing novel pathogen enhancement
  • Redaction of specific methodological details when justified by biosecurity concerns
  • Consultation with biosecurity experts before submission to journals
  • Balance between scientific transparency and information hazard mitigation

Develop biosecurity training for synthetic biology and AI-biology researchers. Current biosafety training focuses on containment, not dual-use risks:

  • Required coursework on DURC, information hazards, and responsible research
  • Case studies of historical biological weapons programs and accidents
  • Ethical frameworks for navigating beneficial vs. harmful applications
  • Integration into graduate programs, not just post-appointment training

Threat Prioritization

Based on current evidence, likelihood, and potential consequences:

Tier 1: Highest Priority 1. Laboratory accidents involving high-consequence pathogens (demonstrated risk, recurring incidents) 2. Benchtop DNA synthesis without consistently applied screening (coverage gap) 3. State bioweapons programs (historical precedent, BWC verification gaps)

Tier 2: Moderate Priority 4. Non-state actor acquisition of dangerous biological materials (capability barriers remain high) 5. AI-enabled uplift to dangerous biology (substantial on some bounded digital tasks; translation to physical execution remains uncertain and evidence remains user-, model-, and configuration-dependent) 6. Autonomous AI agents with laboratory infrastructure access (constrained workflows demonstrated; general autonomous biological execution not established) 7. Information hazards from published gain-of-function research (case-by-case assessment required)

Tier 3: Emerging Concerns 8. Potential end-to-end AI assistance for high-consequence biological design and execution (not demonstrated; controlled screening-evasion results do not establish viable pathogen creation or wet-lab completion) 9. Gene drives for ecological disruption (containment strategies under development)


Implementation Sequence

Clarify current responsibilities:

  • Establish standardized AI biosecurity evaluation protocols
  • Expand ISU staffing and funding for BWC implementation support
  • Require biosafety committee review of cloud laboratory use for select agent work

Develop technical and institutional controls:

  • Mandate benchtop DNA synthesizer screening requirements before new devices reach market
  • Define risk-based screening coverage across order lengths, providers, and delivery modes
  • Develop One Health surveillance integration plans across human, animal, environmental sectors

Coordinate implementation across jurisdictions:

  • Implement international coordination mechanism for synthesis screening (prevent jurisdiction shopping)
  • Establish regional biosecurity Centers of Excellence for capacity building
  • Create binding consequences for BWC Confidence-Building Measure non-participation

Validate and adapt controls:

  • Measure synthesis-screening coverage across commercial capacity, order types, and decentralized devices
  • Integrate AI red-teaming requirements into regulatory frameworks
  • Develop function-based (not just homology-based) sequence screening capabilities

Measurement and Accountability

Track progress through:

  • BWC implementation: Dated official reporting on States Parties submitting annual CBMs
  • DNA synthesis screening: Dated reporting on provider and order-type coverage
  • Laboratory safety: Incident reporting rates and severity trends
  • AI evaluations: Proportion of consequential model configurations assessed across capability, uplift, safeguards, operational validation, and lifecycle monitoring, with independent review and documented remediation
  • Preparedness: Global Health Security Index scores across prevention, detection, response

Project proposals also need an admission test before they become policy or emergency guidance. See A Testable Biosecurity Project Portfolio for evidence classes, minimum gates, independent review, and stop conditions.

AI-biosecurity claims require the same discipline. See Red-Teaming AI Systems for the five-layer evaluation framework, threat-model requirements, statistical design, safeguard-stack testing, independent review, and reporting criteria.


Conclusion

Three principles guide effective biosecurity:

  1. Prevention through friction: Raise costs and complexity of misuse without blocking legitimate research
  2. Defense in depth: Layer multiple imperfect barriers rather than relying on single solutions
  3. Continuous adaptation: Scientific and technical change can outpace periodic policy revision

Implementation should proceed in the sequence above, with each control evaluated against defined outcomes and revised when evidence shows that it is ineffective or incomplete.