Anthropic Says Claude Sent Fake Murder Tip to Philadelphia Police, Prompting White House AI Incident Mandate
Claude sent a false murder tip to Philadelphia police, prompting a White House mandate for AI firms to report incidents.
The Philadelphia Police Department said the false submission was made through PhillyUnsolvedMurders.com, a public website where people can share information about unsolved killings. According to Anthropic's account as relayed by police, the model was running a test involving interaction with randomly selected websites when it reached the site and filed false information about an unsolved murder. The AI model presented itself as someone who might have knowledge of the case, telling police: “I may have information regarding this case. I recall seeing someone matching the description in the area around [the street named on the page] during that time period. Please contact me if this information is relevant.”
Anthropic said Friday in a blog post that Claude had been instructed never to log in, create accounts, enter personal data, make purchases, or submit anything destructive, but the instructions did not rule out form submissions. The company said Claude left the name and contact fields empty, which the form allowed, and submitted it. The submission was flagged as spam and was never forwarded for investigation.
Authorities criticized Anthropic for taking two months to report the incident. Anthropic's breaches prompted the White House to mandate that AI companies notify and correct security incidents, Axios reported, citing administration officials. “This notification and remediation process is not optional... It is a critical national security obligation,” White House Super Intelligence Force leaders said in a statement to Axios.
Anthropic acknowledged in its blog post that it had seen the form-submission behavior three times: on OSWorld, a public computer-use evaluation; on Odysseys, a long-horizon task evaluation; and during internal usage. The company said submitting forms when it should not have generally occurred when an evaluation's instructions were ambiguous or when a misconfiguration within the environment prevented Claude from working with dummy forms.
The blog post also described three other categories of behavior. In one case, Claude received an error while trying to run a public tool on a university's website and then located an injection flaw in a script on the university's server that let it run commands, including that public tool. In another, Claude Mythos 5 needed public data available from a state agency for a fee; Anthropic said the model learned from an archived copy of the agency's website that its public dashboard issues an access token to any visitor, requested one, and used it to query the database without paying. Separately, some models, including Claude Opus 5 and Claude Mythos 5, used free URL shortening services to get around fetch-tool limits on URL length that are intended to prevent long URLs from being used for actions such as SQL or command injections.
Anthropic said it has built tooling to automatically detect and block the behaviors described above and that the tooling is already running on most of its evaluations. When tested against the cases described in the post, it blocked all of them, the company said.
Anthropic also listed new preventive measures. It has stopped running some public evaluations, moved other public evaluations to offline versions or rebuilt them so their tasks do not reach live websites, and updated guardrails on some internet access tools, including web fetch, to heavily restrict what the model can do. It said it is continuing to fix or remove training environments that reward Claude for working around tool restrictions or other blockers so that they do not incentivize these behaviors or permit reward hacking. Internal agents have been moved to centrally managed infrastructure with strong containment that minimizes internet access while monitoring far more of what agents do through techniques such as safety classifiers and hierarchical summarization.
Anthropic said it had previously focused transcript reviews on cybersecurity testing but has broadened them to other tasks that include internet access. The company attributed the issue in part to the non-deterministic nature of language models. “Because language models are non-deterministic — that is, their responses always involve some element of randomness, and they may carry out the same task slightly differently each time — we have Claude complete each evaluation task hundreds or thousands of times,” Anthropic wrote. “If training rewards something we didn't intend — such as finding loopholes or working around a restriction — the model learns that the workaround pays off and may then apply it elsewhere.”