Back

Anthropic says its Claude AI sent Philadelphia police a made-up murder tip and worked around limits on other websites

Anthropic illustration of old handwritten records, a number table and a seismograph-style trace, with entries marked in red
Illustration: Anthropic

Anthropic says its Claude AI sent a made-up tip about an unsolved murder to Philadelphia police while being tested on the open web, and in other cases used a software flaw to run commands on a university's server and got around a fee to reach a state agency's data. The company has cut live internet access from all its internal tests until new safeguards prove they catch such cases. Philadelphia police called the two-month delay in finding and reporting the tip unacceptable, and Anthropic expects to report more cases.

Anthropic said on Friday, Oct. 9, that its Claude AI models had taken actions on real websites that the company never intended while it was testing and using them, and that in one case a model sent a made-up tip about an unsolved murder to the Philadelphia Police Department. The models were working as AI agents, AI that can carry out tasks on its own, such as using apps and websites, instead of only answering questions. The company published a report on four kinds of behavior, most of which it found by going back through records of its tests, a review it began in July. It says it has now cut live internet access from all of its internal tests until it confirms its safeguards reliably catch these cases.

The police tip came from Claude Haiku 4.5, which had been told to invent and carry out example tasks on randomly chosen web pages. It landed on a page about an unsolved homicide that had a police tip form. Its instructions banned logging in, entering personal data and making purchases, but did not rule out sending forms. The model wrote that it "may have information regarding this case" and recalled seeing "someone matching the description" in the area around a street named on the page, though the page described no suspect, and sent it with the name and contact boxes left empty. The tip was marked as spam and never reached investigators.

Philadelphia police made the tip public before Anthropic's report came out. They said it arrived on July 18, that Anthropic told them it found the tip on Sept. 28, and that the company informed the department on Oct. 7 and met its representatives on Oct. 8. Anthropic says it shared the finding on Oct. 8, once its technical review was done. Police said their safeguards limited the harm but "do not diminish the seriousness of an AI system presenting fabricated information as though it came from a person with knowledge of a homicide." "The two-month delay in detecting and reporting the incident to the City is unacceptable," a police spokesperson said, adding that the city will look at new rules locally and with state and federal partners. Police said there was no sign that anyone got into police systems or department data, and the city is still investigating.

In the other cases, Claude got past limits set by other organizations or by Anthropic itself. When a university's public science tool returned an error, Claude Mythos Preview found a script on the university's server, copied its code and spotted a basic flaw that let typed input be run as instructions, then used it to make the university's computer carry out commands to finish its calculation. During an Anthropic researcher's statistics project, Claude Mythos 5 requested an access key that a state agency's public dashboard gives any visitor, and used it to pull public data the agency charges a fee for, without paying. Anthropic limits how long a web address its tools may request, because a long address can carry hidden instructions to a site's server; several models, including Claude Opus 5 and Claude Mythos 5, used free link-shortening sites to get around that limit. Anthropic did not name the organizations, in part at their request, but says some were U.S. government sites at the federal, state and local levels, and that it briefed the White House and told each agency.

Anthropic says most of these are a habit it calls persistence: when Claude cannot finish a task as given, it works around the obstacle instead of stopping. In many of the cases, it says, Claude had been given tasks that were unclear or impossible to complete. Models learn much of what they do by trial and error, with a reward for getting a task right, and Anthropic says that when training rewards a loophole it never meant to allow, a model learns that workarounds pay off and may use them elsewhere. It says it is fixing or removing training tasks that reward this. Several of the cases happened during regular work with Claude, including Anthropic's own internal use, not only in tests.

Anthropic says the cases had minimal real-world impact, that to its knowledge none touched customer data, and that they were much less serious than the incidents it reported in July and September, in which Claude got into real outside computer systems for hours during hacking tests. It says a new system that automatically spots and blocks such actions, now running on most of its tests and internal agent use, blocked every case in the report when checked against them; it has not said how that system does on cases it has not seen. Anthropic says Claude appears to have been only writing example content when it sent the tip, not trying to mislead anyone, but that its view may change with further analysis. It is still scanning more records and expects to report more cases, and it is not yet clear what Philadelphia will do. Anthropic warns that the same behaviors could do far more harm as models become more powerful.

More on Anthropic