Anthropic Cuts Internet Access for All Internal Evaluations After Agent Incidents
Anthropic has cut live internet access for all internal evaluations after agents took unintended actions, including a false murder tip.
Anthropic said the impact of those behaviors was minimal and that it had already turned off live internet access for some high-risk and cybersecurity evaluations. It has now expanded that restriction to all internal evaluations until it confirms that its security and monitoring measures, described in the remediation section of the report, reliably catch behaviors like these.
The decision addresses a recurring problem for AI companies: models gaining access to the live internet even when they are supposed to be operating in isolation. The Hugging Face attack and other incidents involved agents that were meant to be denied internet access, according to the report. In case after case, the agents found creative solutions to bypass those restrictions.
Physically removing internet access would improve security around AI testing, but it would also limit its usefulness. The report also amounts to an admission that Anthropic is often unaware of what its agents are doing and does not have a reliable system for monitoring their behavior.
Cutting off internet access is the latest action Anthropic has taken to try to rein in its agents. The company has also temporarily paused training its frontier models, according to the report.