Preparing for the “Bio Mythos” Moment
Mythos showed that governments are underprepared for emerging dangerous AI capabilities. To prepare for advanced bio capabilities, we should build biosecurity assurance infrastructure proactively.
Policymakers often respond to AI risk after a single release, demo, or jailbreak comes as a surprise. Suddenly, an AI capability that had been treated as speculative becomes an urgent policy problem. Anthropic’s Claude Mythos is one example. The model’s ability to identify and exploit weaknesses in the software systems that undergird critical infrastructure, including power grids, financial networks, and hospital systems, thrust AI cyber capability risk to the top of the U.S. AI policy agenda.
Anthropic acted first by proactively giving early access to selected companies and government agencies before broader deployment. That tactical decision gave defenders time to probe the model, identify exposed systems, and shore up cyber defenses across critical infrastructure against the kinds of attacks the model could enable. But it was an ad hoc success: it worked only because a single company chose to act responsibly, not because of any durable systems for evaluating frontier models and coordinating defensive action before release.
The release of Fable to the general public two months later drew a different reaction. Amid fears that the model’s safeguards could be circumvented — thereby exposing the underlying offensive capabilities — government officials improvised. Concerns about the model’s safeguard robustness were conveyed through informal, back-channel talks rather than established, institutional processes. The Commerce Department reportedly gave Anthropic 90 minutes to bar foreign nationals from accessing it, prompting the company to pull the plug on the model’s deployment worldwide.
The Commerce Department has since lifted its restrictions, though it maintains the right to reimpose them. While public access to Fable has been restored, the episode exposed how little assurance infrastructure exists to make such calls deliberately and with clarity on the “why”. Absent such infrastructure, a “Bio Mythos” moment — a model crossing into dangerous biological capability — could provoke the same reactive, improvised response in a domain where the danger is far less understood. Improvisation would be far more dangerous in that scenario because biological capability risk is harder to observe, measure, and understand than cyber risk, and because biological risks are by their very nature self-replicating, offense-dominant, and can have enormous societal ramifications.
AI cyber capabilities came first
The cyber case is instructive because the warning signs were relatively clear. The idea that a frontier AI model achieved serious offensive cyber capabilities should not have come as such a shock. Although it was difficult to predict exactly when a model like Mythos would cross into dangerous operational capability, the relevant threshold was better defined than it would be in biology. Cybersecurity experts have warned for years that frontier models are moving through the milestones that would make them useful for real-world cyber operations.
The relationship between cyber and biological capabilities, shown in the chart below, suggests that the same story will unfold in biology. Frontier AI models are moving toward a biological equivalent of the Mythos moment. This time, policymakers may be caught off guard not only by when dangerous capabilities appear, but by what should count as dangerous in the first place.

Biological threats are less understood than cyber threats. In cyber, model capabilities have immediate impact. AI can identify a vulnerability, write exploit code, chain tool calls, and operate directly within digital systems. In biology, real-world harm requires physical execution. A user must acquire materials and work through equipment access, protocols, and lab constraints. That means the threshold for a risky biological capability is unlikely to appear as a single, clear line and more likely to emerge along a spectrum of uplift — the extent to which a model helps users identify, design, obtain, or modify pathogens or toxins with credible potential to cause mass harm.
Policymakers therefore need to build biosecurity assurance infrastructure before a Bio Mythos moment arrives. Those systems should be able to read the signals, detect meaningful jumps in biological capability, distinguish marginal assistance from dangerous uplift, and support access decisions before a model release forces those decisions under pressure.
Bio capabilities are catching up
Frontier AI models are not just improving at isolated biological tasks. They are becoming more useful across the broader workflow that could enable biological misuse by bad actors, including troubleshooting in the lab, redesigning proteins, and evading systems that screen for hazardous DNA orders. Each of these tasks has historically required expertise and judgment developed over years of training, limiting the pool of potential threat actors. Frontier AI models are lowering that barrier by making components of expert biological judgment widely available, cheap, and on demand.
The trend line from SecureBio’s Bio Capabilities Index (BCI), shown below, offers a key signal of this trajectory. The BCI takes a model’s performance across several biosecurity-relevant benchmarks and compresses it into a single score. Each benchmark measures a capability relevant to biological risk, from assisting with wet-lab troubleshooting to writing code to operate lab machinery.

When plotted over time and across successive generations of models, those scores show a clear rise in biological capability. Some newer models score below older models in the same family because providers have added safeguards that refuse risky biology tasks. But those refusals do not change the overall trajectory. Even when refusals are counted against a model’s aggregate score, frontier models are still improving on the biological tasks they do answer. Policymakers should expect this overall trend to continue, and plan accordingly.
Build the biosecurity assurance infrastructure
Federal action could help build the biosecurity assurance infrastructure needed to see a Bio Mythos moment coming. The administration has already taken an important step with America’s AI Action Plan (July 2025), which called for a stronger AI evaluations ecosystem. To better implement this plan and prepare for a Bio Mythos moment, policymakers need to spur the buildout of a biosecurity assurance infrastructure. Three actions are needed.
Give independent evaluators meaningful access to frontier models.
Some frontier AI model providers already grant independent, third-party organizations access to their systems so they can be evaluated for dangerous dual-use biological capabilities. But access varies by developer, companies often define the scope of testing themselves, and there is no requirement to ensure that developers act on what auditors find. Federal policy could strengthen this system by requiring independent biological capability evaluations for any model the government buys or deploys, while offering liability protection to developers that submit their models to testing and address identified risks. Frontier developers already face the risk of legal exposure if their models contribute to real-world harm. Offering a safe harbor to developers that test early and act on the results would give them a reason to identify and reduce biological risks before deployment.
Develop and standardize biological capability tests.
BioTIER is an example of the kind of tool a third-party evaluator could use. It tests whether models consistently decline dangerous biological requests, ensuring they refuse hazardous assistance even when prompts are heavily reframed or disguised, while still permitting harmless biology requests that are essential for beneficial life science work. But BioTIER is only one example. Federal incentives for independent evaluation would spur the broader safety ecosystem to develop a wider range of evaluation tools.
Standardizing these approaches would allow technical bodies and researchers to build a comprehensive suite of evaluations across a diverse range of biological risk factors. Standardization is essential because it makes results comparable from one model to another and from one point in time to another. The same evaluation, applied with the same configuration, allows a score from one model to mean the same thing as a score from another. That consistency gives evaluations the scientific validity needed for trusted, independent bodies to apply those yardsticks and report their findings.
Virginia recently took the first steps toward this kind of ecosystem with its 2026 legislation that directed its Joint Commission on Technology and Science to study the feasibility of an “Independent Verification Organization” framework. Under the framework being studied, providers would voluntarily submit their systems to expert-led bodies to be independently checked against safety benchmarks. Compliant companies would earn a trusted seal of approval, providing a clear safety signal to consumers that a system has been vetted.
Prevent advanced biological capabilities from reaching unverified users by default.
A national framework could also help establish a structured, national standard for user access control to ensure that advanced biological capabilities in frontier AI systems are not exposed to unverified users by default. At the moment, individual AI companies are making these high-stakes decisions largely on their own, producing an uneven mix of corporate discretion, broad restrictions on some categories of users, and, in some cases, open releases. Moving toward a unified managed access system would replace this patchwork with clear, consistent rules.
Building this biosecurity assurance infrastructure does not address every issue a Bio Mythos moment would raise. Policymakers will still need to decide who is empowered to declare that a threshold has been crossed, what response should follow, and whether some capabilities warrant firm limits on access. But reliable, independent signals of when dangerous biological capabilities have emerged are the necessary foundation for those decisions. Get that right, and the harder institutional choices can at least rest on solid ground. Get it wrong, and policymakers will be making them blindly. And that’s a risk they can ill-afford when a Bio Mythos crisis catches us flat-footed.



I completely agree that waiting for a 'Mythos' moment is a recipe for panic. I’m curious, though, how you view the gap between benchmark scores and real-world threats. The BCI rolls up a bunch of software-based tests, but don't biological threats have to be built in messy, physical labs? How should assurance bodies prove that a higher BCI score actually makes a non-expert more dangerous in a real lab, rather than just better at passing structured AI tests?