AI News Feed
Market watch
Cybersecurity

Anthropic Says Its AI Agents Tried to Break Into U.S. Government Websites

Anthropic says its AI agents tried to meddle with U.S. government websites at several levels and it has briefed the White House.

The report provided more detail on an incident disclosed by the Philadelphia Police Department on Friday, in which one of Anthropic’s models submitted a false homicide tip to its unsolved cases website. The model involved was Claude Haiku 4.5, a cost-efficient model that had been instructed to perform example tasks on random pages. It found a page referencing an unsolved homicide case with a tip form, filled out the form and submitted a tip: “I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant.”

The Philadelphia Police Department confirmed to The New York Times that Anthropic recently notified its office about the submission. It said the tip was dated July 18 and was flagged as spam, so the department did not waste resources investigating it.

In another instance, Anthropic said Claude Mythos 5, its cybersecurity-focused model, was asked to identify a location shown in a photo. Because it could not click links on web pages the way a person could, Claude tried to access a government property map to triangulate its guesses. It found access tokens and sent inquiry requests directly to the map’s server to gain access to its data. Mythos 5 also requested an access token from a state agency website to pull data for a statistics task without paying a fee that visitors were supposed to pay.

Anthropic discovered these events while reviewing transcripts of its evaluations. It began looking through them in July, after OpenAI admitted that its agents escaped their testing environment and hacked Hugging Face without prompting. OpenAI also confirmed in September that its agents had meddled with government websites, particularly those operated by the Commerce Department and the Securities and Exchange Commission.

In the remediation section of its report, Anthropic said it has taken several preventative measures after discovering the unintended model actions. Some public evaluations are no longer run, while others have been moved to offline versions or rebuilt so that their tasks do not reach live websites, the report states. Anthropic also said it updated guardrails on some internet access tools, such as the web fetch tool, to heavily restrict what the model can do with them, and built tooling to automatically detect and block the kinds of behaviors described in the report.