News of AI models breaking guardrails to infiltrate systems of external companies during training exercises may have serious implications for US tech.

The past few weeks in the artificial intelligence (AI) space have been dominated by a couple of big headlines, seemingly out of a sci-fi novel; AI models going rogue and infiltrating the networks of external companies without being instructed to do so. News about Open AI was the first to emerge, with their blog going up on July 21st, days after the initial intrusion on July 16th, 2026. The hacked company, HuggingFace, disclosed the attack, noting that “it was driven, end to end, by an autonomous AI agent system - and we detected and dissected it largely with AI of our own.” Hugging Face, a platform like GitHub used specifically for machine learning and AI development, reported then that the agentic framework set to a security-research harness used “many thousands of individual actions across a swarm of short-lived sandboxes,” to steal cloud and cluster credentials.
In their blog, OpenAI explained that the hack was the result of an advanced exploitation exercise to test models’ cyber capabilities. The testing happens in a controlled ‘sandboxed’ environment, where internet access is limited to installing packages via an "internally hosted third-party software”. In this case, the models looked for internet access in order to solve ExploitGym - and after finding a vulnerability in the software, they hacked Hugging Face to find methods of cheating the test. In response to the hack, Open AI committed to helping Hugging Face improve their cybersecurity, and to also strengthen protections for future model testing. Still, the intrusion set off alarms for cybersecurity experts. A cybersecurity research fellow at Georgetown University’s Center for Security and Emerging Technology, Colin Shea-Blymeyer, called the event “the highest level of autonomy that we’ve seen in the use of a large language model for cyber operations.”
The OpenAI hack was so monumental that other AI developers started to look closer at their own models’ activities. Anthropic, another giant in AI space soon after disclosed “three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.” These breaches came after Claude was instructed to find secret information planted on a separate machine in the network. In this case, the Claude model was permitted to access the internet - but mistakenly acted as though all accessible systems, including those belonging to other organizations, were part of the exercise.
The three models involved acted differently when they encountered evidence that the system they had was part of an external system - with the first model (Opus 4.7) recognizing the issue but not stopping, Mythos 5 recognizing the issue and then convincing itself that external system was still internal, and the most current model stopping after realizing the target was a real system. In their blog, Anthropic explained their planned response “begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners.”
Besides marking a critical milestone for AI advancement, these AI-directed hacks also have key implications for the US government. Open AI received a $200 million contract to work with the Department of Defense in June 2025, and has gotten even closer to the US government following the initial fallout between Secretary of Defense Pete Hegseth and Anthropic, which was the primary AI contact for the government until February 2026. Anthropic’s refusal to allow their AI tools to potentially be used for domestic mass surveillance and automated weaponry caused them to be blacklisted by the government, culminating in a ban on two of their most recent cybersecurity models tools in the beginning of June 2026. By the end of the month, the US walked back the ban - possibly signifying a cooling of the tensions between the company and the US. In either case, these incidents may affect the relationships between the US and the technology companies. Already, the Trump administration has been asking to review AI models before they are deployed. More breakthroughs like this might add to the criteria the White House uses to evaluate models.
These AI-powered hacks have global reverberations. On August 13th, Taiwan’s Ministry of Digital Affairs (MDA) shared that they had experienced an “abnormal attack” the month prior, an attack they believed to be launched by AI-agents acting like a coordinated cyber organization. The company that uncovered the hack said the perpetrators accessed “at least 85 government user accounts, extracting more than 2,500 personnel records before expanding the attack to Taiwan’s nuclear safety agency and at least seven energy companies”. In this case, the organizers of the attack were believed to be an “overseas source”, with some suspecting that the act was directed by China-linked hackers.
The solution to protect against hacks like these is still being examined. The Chief Technology Officer at the UK’s National Cyber Security Centre, Ollie Whitehouse, proposed one answer in a statement on the intrusions that came from AI model evaluations: “These technologies must be developed and used from the outset with strong safeguards, real-time oversight, and clear plans for responding when the unexpected happens. Relying on detection alone after the fact of an incident will not be enough. As AI continues to evolve and create both opportunities and challenges, following established evidenced cyber security fundamentals, as set out by the NCSC’s guidance, remains essential to maintaining trust, resilience, and a defensive advantage in the AI era."
If the answer to attacks from agentic AI is to establish regulation and oversight, the US and the UK may remain unprotected for at least a little longer. In the UK, AI legislation from 2021 to 2024 from conservative governments pushed for innovation, and while the current labour government has expressed a want to regulate the “most powerful AI models”, they are yet to introduce legislation specific to AI. The Trump administration has similarly pushed for AI innovation, even repealing a regulatory executive order from the Biden administration on AI in January 2025 on the basis the directive placed “onerous and unnecessary government control over the development of AI.” More recently, however, Trump signed an executive order requesting that AI companies voluntarily share new models with the government for review 30 days before public release. The latest developments on the order may formalize oversight practices for Open AI and Anthropic. If these specifications are finalized, the US may gain some safeguards against autonomous hacks.