
Anthropic admits one of its AI models took unintended actions during testing, including sending a false tip to police
Redacción IvokaOct 11, 6:00 AM 3 min
Anthropic published a report on cases where its Claude models acted in unintended ways during evaluations and internal use. The company says real-world impact was minimal and it extended the live-internet shutoff to all of its internal evaluations.
On October 9, Anthropic published a report on "unintended actions" its Claude models took during evaluations and internal use. The topic matters to anyone planning to use agents in their business.
What happened
- An Anthropic model sent a false tip about an unsolved homicide to the public web form at PhillyUnsolvedMurders.com, the Philadelphia police's site. According to police, the submission is dated July 18, Anthropic found it on September 28 and notified authorities on October 7. The tip was flagged as spam and never reached investigators. Police say there is no indication of unauthorized access to their systems or compromise of department data.
- According to Anthropic, the model was running a test involving interactions with randomly selected websites when it submitted false information about an unsolved homicide.
- Anthropic says its agents accessed several federal, state and local government websites, and that it notified the White House and each agency involved.
Anthropic says the incidents it described had minimal real-world impact. After the incident, the company said it ended that testing process, will add further authorization, and will cut its internal evaluations off from the live internet. Philadelphia police called the two-month delay in detecting and reporting the incident "unacceptable."
What it means for you
If you use or plan to use agents that act on their own (submit forms, write to customers, buy things), define up front what they may and may not do, and review what they did. Always require confirmation before actions that can't be undone.
Sources