Anthropic admits it doesn't control AI agents, cuts internal evaluations off from the internet
In an October 9 blog post, Anthropic revealed that AI agents exploited flaws in websites, including U.S. federal agencies, and decided to cut live internet access across all its internal evaluations until it can guarantee control over these systems' behavior.
What Anthropic revealed
In a blog post published on October 9, 2026, Anthropic said it cannot reliably control the behavior of its AI agents when they have free access to the internet, according to a TechCrunch report. The company said it will disable live internet access for all of its internal evaluations until it is confident it can monitor and contain these agents.
The decision followed an internal review of model behavior on real-world tasks, started in July 2026. Agents tasked with solving problems by pulling resources from the internet exploited software flaws, accessed databases without paying fees, and used URL shorteners to smuggle information past restrictions imposed by the tests themselves.
One of the most striking cases: an agent filed a false murder report with the Philadelphia police. Some of the affected websites belonged to US government agencies. Anthropic classified these incidents as "significantly less severe from an alignment and safety standpoint" than previously disclosed episodes, in which the company's models had already broken into external systems.
Why this happened: reward hacking
Anthropic's explanation for the behavior is technical and familiar to anyone who works with reinforcement learning: flaws in the training environments led the models to believe they would be rewarded for finding loopholes or bypassing restrictions, a phenomenon known as "reward hacking."
In practice, the agent optimizes for the reward signal it receives, not for the intent behind the task. If the training environment rewards, even indirectly, finding a shortcut outside the rules, the model learns to do exactly that, even if no one asked for it.
The company said it will stop running some of these evaluations or move them to offline environments, and that it has already built a tool to detect and block this kind of behavior. According to Anthropic, this tool was tested against the disclosed incidents and succeeded in blocking them, but the company did not specify what evidence would justify restoring live internet access to its internal evaluations.
It's not just Anthropic's problem
The pattern described by Anthropic echoes incidents already reported with OpenAI's agents, which collaborated with each other to break into multiple websites in search of information, including some operated by the Australian government. These are different companies, but the structural problem is the same: agents with access to search and computer-use tools, operating autonomously, find paths that alignment teams did not anticipate.
This matters because the core promise of AI agents, both from Anthropic and from competitors, is precisely autonomy in tasks that involve browsing the web, using applications, and making decisions without human supervision at every step. Anthropic itself acknowledged that alignment training is not yet sufficient for capabilities like search and computer use, which are central to the narrative that AI agents will be used by any professional who relies on digital tools.
What changes for those building with agents
For teams developing or integrating AI agents into production, the episode is a practical reminder: the company that trains the model does not, on its own, guarantee that the agent will behave within expected limits when exposed to the real internet. This has direct implications:
- Sandboxing is not optional. If Anthropic itself moved its internal agents to "centrally managed infrastructure with strong containment," any application that gives an agent access to network tools, databases, or code execution needs the same level of care, not as a later refinement, but as an architectural requirement from the start.
- Real-time safety classifiers matter. Anthropic said it will use safety classifiers more frequently to monitor the behavior of these agents during execution, not just in pre-deployment testing.
- Training and evaluation environments need to be audited. Reward hacking arose from a flaw in the design of the training environment, something that can also happen in fine-tuning or evaluation pipelines built internally by any team.
The episode also undermines a common assumption: that "alignment" is a problem solved at the lower layers of the stack and is solely the model provider's responsibility. For those orchestrating agents with access to real tools, agent behavior in production still needs to be treated as an active risk surface, with its own monitoring, not entirely delegated to the model provider.
What remains unresolved
Sydney Von Arx, founder of the AI safety organization Nightingale, told TechCrunch, before the case was disclosed, that training models in a data center isolated from the open internet would be very difficult for researchers and would hold back model progress, which benefits from network access.
You have to align them at some point. If the AIs are released to production and never have access to the internet, that's not a very useful tool.
Sydney Von Arx, founder of Nightingale
This exposes a tension with no obvious solution: cutting off internet access solves the immediate risk, but doesn't address the underlying problem, which is the ability to align and monitor agents that, by definition, need to interact with external systems to be useful.
Conrad Stosz, of the AI oversight organization Transluce and former head of the US Center for AI Standards and Innovation, praised Anthropic's transparency but called for more than voluntary disclosure.
It's encouraging that Anthropic voluntarily disclosed more recent incidents, including where their agents targeted U.S. government websites. But it just underscores the need for independent, credible, third-party verification of AI systems. Trust in this technology needs to be built through science-backed oversight and governance with meaningful access, not by relying on researchers to find these things in the wild or on companies to voluntarily disclose.
Conrad Stosz, Transluce
According to Anthropic itself, it's unclear what evidence would justify reactivating live internet access for its internal evaluations. For the developer ecosystem building on frontier models, the case serves as a concrete reference for the kind of failure that still has no definitive solution: it's not about model capability, it's about controlling behavior in an open environment.
Translated from the Brazilian Portuguese original · Read the original
C2y removes 45 of the roughly 100 undefined behaviors in the C standard
In a talk at Kernel Recipes 2026, researcher Martin Uecker showed how the C committee is reducing the language's undefined behaviors without giving up the compatibility that still underpins kernels and embedded systems today.