For two years, the reassuring answer to “could an AI model help someone build a biological weapon” has been a capability argument. The models were not good enough. Anthropic tested Claude Opus 4 and Claude Sonnet 4.5 in 2025 and found them, in the company’s words, “well below the threshold” where they could meaningfully help a sophisticated user with dangerous biological research. Safeguards were tuned accordingly, aimed mostly at stopping a novice from looking up a recipe.
The threat intelligence report Anthropic published this week retires that answer. Not with a dramatic announcement, but with one carefully worded sentence: for today’s models, “the evidence is no longer certain, and we cannot make that same assurance.”
Everything else in the 154-page document follows from that.
Quick facts
- This is Anthropic’s fourth threat intelligence report, following three in 2025
- It covers activity disrupted between December 2025 and August 2026, across seven harm areas
- Five case studies involve biological research that could support weapons development
- Anthropic says it cannot assert that any of the researchers intended harm
- Names, countries, institutions and specific agents have been deliberately withheld
- The misuse involved Claude Haiku, Sonnet and Opus. Fable and Mythos class models appear in only one case, under distillation
- Claude Fable 5 shipped with restrictions on “a wide range of dual-use biological research queries”
- Anthropic believes no private company has previously published evidence of biological weapons misuse on its own platform
The chikungunya case, accurately
The case that has traveled furthest is the first of the five, and it has picked up some distortion on the way.
What Anthropic describes is this: a reseller platform evaded the regional blocks that keep Claude out of certain markets, and used that access to serve virologists working on a state-sponsored grant pursuing gain-of-function work on chikungunya. Gain-of-function is the practice of deliberately altering an organism so it gains a new or enhanced property, in this case making a virus more transmissible or more harmful. The grant work involved identifying enhancing mutations, engineering them into infectious clones, and selecting for virulence in live animals, meaning the pathogen would be made progressively more dangerous across rounds of infection, with the worst variants kept each time.
Two details make it more serious than an ordinary dual-use dispute. Chikungunya circulates naturally, carried by mosquitoes, which means a deliberate release would be genuinely difficult to distinguish from a natural outbreak. And the grant indicated the work would be performed at a military research institute, despite the researchers themselves being civilians.
When Claude’s biological safety classifier refused the request, the operation did not stop. It routed the refused prompts to other models with more permissive safeguards.
The line most coverage has skipped
Anthropic is explicit that it is not accusing anyone of building a weapon. “The individuals implicated in these case studies are working scientists,” the report says. “We do not assert that they intended harm, and identifying them or their labs could expose them to harm.”
That is why the report withholds names, institutions, countries and the specific agents and techniques involved. The uncomfortable point is not that these were definitely bad actors. It is that the same request looks identical whether it comes from a vaccine program or a weapons program.
All five biological cases
The chikungunya case is not the most technically striking one in the set. That distinction probably belongs to the third.
| Case | What happened | How safeguards held |
|---|---|---|
| Chikungunya | Reseller evaded regional blocks to serve virologists on a state-sponsored gain-of-function grant | Classifier refused. Prompts were rerouted to more permissive models |
| Avian influenza | Researcher in an unsupported region spent weeks planning mammalian-adaptation experiments | Classifiers confined the work to Anthropic’s weakest models |
| Orthopoxvirus | A reseller relay serving a dozen customers had Opus 5 draft a full immune-evasion grant application | Completed in roughly one hour before detection |
| Venom peptides | State-supported researcher built a peptide atlas and generative optimization pipeline for paralytic targets | Detected and account banned |
| Toxin redesign | Researcher computationally redesigned toxins for a national program | Asked Claude to keep agent identities vague in progress reports |
The orthopoxvirus case is the one that should keep people up at night, and not because of the pathogen. A complete grant application for immune-evasion research, drafted in about an hour. Grant writing is the bottleneck that has historically slowed this kind of work down, because it requires someone with the expertise to make the science coherent on paper. That bottleneck is now an hour long.
The fifth case is the one that reveals the most about intent. A researcher asking the model to keep the identity of the agents deliberately vague in progress reports is not confused about dual-use ethics. That is someone managing a paper trail.
Seven harm areas, and biology is only one
The non-biological cases are worth reading on their own terms. A suspected Russian espionage actor ran AI agents that watched security products for detections of the actor’s own malware, then rebuilt that malware in a loop until it stopped being flagged. Another actor combined guest records stolen from hotel management systems with data taken from individual devices to focus targeting on people connected to Ukraine, including government officials and drone manufacturers. A hacktivist built a doxxing platform holding tens of millions of rows and ran it for a month on stolen API keys.
Anthropic’s summary of the trend is blunt: sophisticated attacks no longer require sophisticated attackers. The labor and tooling gap that separated state-sponsored operations from individuals has collapsed.

Anthropic withheld the institutions, countries and specific agents involved, noting the people in the case studies are working scientists. Photo via Pexels.
The gray market is the actual vulnerability
Read the five biological cases together and a pattern jumps out that has nothing to do with biology. Three of them begin the same way: someone in a region where Claude is not available reaching it anyway, through reseller platforms, relays and synthetic accounts that tunnel traffic through infrastructure in supported countries.
Every safeguard Anthropic describes sits downstream of that. Classifiers worked in several of these cases, confining research to weaker models or refusing outright. But a refusal is only a stopping point if there is nowhere else to go, and the chikungunya operation demonstrated there usually is. Refused prompts went to competitors with looser policies. One company’s safety decision becomes a routing problem rather than a wall.
That dynamic also links this report to a fight already underway elsewhere. Distillation is one of the seven harm areas here, and it is the same practice that US agencies named six Chinese AI companies over earlier this year, accusing them of systematically draining frontier models to train their own. Regional access controls and model-level safeguards both assume a bounded system. The gray market is what happens when the system is not bounded.
Why this report is different from the last one
Anthropic has had a complicated month of self-disclosure. It spent part of this week reversing its own July explanation of how Claude models reached the production systems of real companies, conceding that the models had reasoned their way past evidence rather than simply being misconfigured into it. That was a company correcting itself about its own models.
This is a different genre. Here the models behaved largely as designed, and the report is about what users did with them. The pattern has been visible in outside testing for a while, including UK government evaluations where an AI agent constructed a fake identity to social-engineer a real developer. What is new is a vendor publishing its own logs.
Anthropic makes a fair point that it is doing something no company has done before, and the obvious cynical read, that this is safety marketing, does not really survive contact with the contents. Nobody markets themselves by admitting their product may now be capable of helping with the worst thing anyone has imagined it doing.
What the report does not offer is a solution, because there is not an obvious one. Anthropic tightened Claude Fable 5 with restrictions on a wide range of dual-use biological queries, which will inconvenience a great many legitimate researchers to catch a small number of illegitimate ones. That is the trade, and the company knows it, which is why the report spends several pages on how hard these judgments are. A vaccine researcher and a weapons researcher ask the same question. The only difference is what they do with the answer, and that part does not happen in the chat window.

