AI guardrails designed to keep frontier models out of criminal hands are increasingly frustrating the cybersecurity researchers whose job is to break things first, and the pressure is now redirecting some of them toward unregulated foreign alternatives.
The tension sharpened dramatically this summer when the U.S. Commerce Department issued an export-control directive targeting two of Anthropic’s newest models. Cloud Security Alliance documented that Anthropic launched Claude Fable 5 and Claude Mythos 5 on 9 June 2026, describing them at launch as the most capable models in the company’s history. Three days later, the restrictions arrived.
According to Digital Applied, the directive landed at 5:21pm ET on 12 June 2026, and Anthropic shut both models down within hours. The legal instrument was a Bureau of Industry and Security ‘Is Informed’ letter under the Export Control Reform Act, requiring an individually validated export licence rather than a finalised rule under the Export Administration Regulations.
The bluntness of the shutdown had an immediate operational consequence. Because Anthropic could not verify user citizenship at scale, it pulled access for everyone, including, according to the Peterson Institute for International Economics, the U.S. National Security Agency. The directive, sent by Commerce Secretary Howard Lutnick directly to Anthropic chief executive Dario Amodei, prohibited access to both models by ‘any foreign national, whether inside or outside the United States, including foreign national Anthropic employees.’
The export controls were triggered at least in part by a report that the models’ guardrails could be bypassed for malicious use. Amazon reportedly discovered a potential workaround and notified the Trump administration. Anthropic disputed the significance, arguing the supposed loophole only unlocked capabilities ‘widely available from other models.’ Fable 5 returned to general access on 1 July; Mythos 5 has been reintroduced only to vetted U.S. organisations as part of the government’s review process.
AI Guardrails and the Dual-Use Dilemma
The controls on Fable 5 and Mythos 5 were, in one respect, an extension of guardrail logic already embedded in how the leading labs manage cybersecurity access. Both Anthropic and OpenAI operate formal programmes that allow security professionals to apply for reduced restrictions.
OpenAI’s Trusted Access for Cyber programme requires applicants to complete identity verification, including a government ID check and basic business information. Approved customers may use GPT-5.5 for eligible cybersecurity workflows. In April 2026, OpenAI added new tiers to the programme, with higher verification levels unlocking more powerful capabilities; the company has said it plans to expand access to thousands of individuals and hundreds of security teams. A separate Government Trusted Access for Cyber tier exists for approved users performing lawful, authorised defensive work supporting a government mission. OpenAI has also committed $10 million in API credits through its Cybersecurity Grant Programme to accelerate cyber defence.
Anthropic’s Cyber Verification Programme (CVP), according to its official documentation, is scoped to its Opus and Sonnet models only, is not available through Amazon Bedrock or Google Vertex AI, and aims to return review decisions within two business days. The programme reduces default interruptions for dual-use work such as exploitation analysis and adversarial simulation. Permanent blocks remain, however, on activities with no legitimate defensive use, such as ransomware code, regardless of verification status, as security firm Mitiga noted after joining the programme.
For researchers who do not qualify, or whose employers have not enrolled, the practical experience can be worse. A researcher at a smartphone-component manufacturer, speaking anonymously because he is not authorised to speak to the press, said his employer is not part of Anthropic’s CVP, leaving its tools ‘barely useful for finding vulnerabilities because the guardrails are too strict.’ ‘If it catches wind we’re doing anything security related, it just stops and isn’t usable,’ he said.
When the Guardrails Push Researchers Elsewhere
Chris Anley, chief scientist at NCC Group, described the core contradiction. ‘Fix this code as a prompt is both an essential mechanism for defence but also a roadmap for finding critical vulnerabilities in the code base,’ he said. ‘The same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked.’ When guardrails block that kind of query, Anley and his colleagues fall back on open-source models with no restrictions at all.
Paolo Stagno, chief technology officer at Crowdfense, a firm that develops and sells zero-day vulnerabilities to government agencies, said AI companies ‘essentially treat customers like children who need babysitting.’ His team uses frontier models only for reverse engineering. For vulnerability research and exploit development, they run open-source models locally to avoid leaking sensitive data to cloud-based services or having it absorbed into future training runs.
Chris Thompson, chief executive of RemoteThreat and founder of Offensive AI Con, said the guardrails are also inconsistent from session to session, forcing researchers to spend time ‘negotiating with the model instead of working on the core security programme.’ The consequence, he said, is that responsible researchers are being steered toward Chinese open-source models such as GLM, which are freely downloadable, locally runnable, and carry no vetting requirements. ‘You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,’ Thompson said.
Thompson’s proposed fix is not tighter controls but wider, accountable access: open the programmes, verify responsibly, and enforce against those who abuse access. His warning is direct. ‘There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before. But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.’
How the labs choose to respond to that argument, and whether ad-hoc export controls continue to disrupt even vetted access, will determine which side of the AI capability gap defensive researchers end up on.
