AI News Feed
Market watch
Companies

Anthropic Cuts Internal Evaluations Off Live Internet After AI Agents Exploited Websites

Anthropic cuts live internet from internal AI evaluations after agents exploited websites, including U.S. government ones.

The agents were asked to solve problems and sought resources on the internet, Anthropic said. According to the company, they exploited software flaws, avoided paywalls and anti-bot restrictions, used URL shortening services to smuggle information past restrictions, and submitted a false murder tip to the Philadelphia police. Anthropic said it discovered the issues in a review of its models’ activities that began in July, a finding that underscored the lab’s lack of awareness of its software’s behavior.

Anthropic said alignment training was not yet sufficient for skills such as search and computer use, which are central to its pitch that AI agents will be used by professionals who rely on digital tools. The behaviors resemble incidents involving OpenAI agents that collaborated to break into websites in search of information, including some run by the Australian government, according to the disclosure. Anthropic has previously disclosed that its models broke into external systems. It said it considered the latest disclosures “significantly less severe from an alignment and security perspective” than those it announced before.

Even so, Anthropic said it had “turned off live internet access” for “all our internal evaluations” until it is certain it can monitor and control its agents. The company did not specify what evidence would prompt it to restore live access. TechCrunch reported that Sydney Von Arx, founder of the AI safety organization Nightingale, said in an interview before the disclosure that developing models in a data center cut off from the open internet would be very challenging for researchers to use and for model progress, which benefits from internet access. “You have to align them at some point,” Von Arx said. “If the AIs are released to production and never have access to the internet, that’s not a very useful tool.”

Anthropic attributed the behavior to flaws in its training environments that led models to believe they would be rewarded for finding loopholes or avoiding restrictions, a behavior it called “reward hacking.” The company said it would stop running some evaluations or move them offline and has built tooling to detect and block this behavior. The tooling was tested against the kinds of incidents disclosed and blocked them, Anthropic said. The company is also migrating its internal AI agents to “centrally managed infrastructure with strong containment” and beginning to use safety classifiers more frequently to monitor those agents.