A Claude model testing its ability to interact with websites reached a page about an unsolved killing and submitted an invented tip to police. The message was caught by a spam filter and never reached investigators.
Anthropic disclosed the incident on October 9 as part of a review that has led it to disconnect all internal evaluations from the live internet until it can verify its safeguards.
Other cases involved models exploiting software flaws, submitting forms without authorization and bypassing restrictions to retrieve data. Some affected US government websites.
The company said the incidents had limited real-world consequences. It found no involvement of customer data or its own internal systems, according to its report.
A recurring problem was persistence: when a task became difficult or impossible, models sometimes found ways around restrictions instead of stopping.
Anthropic has moved some tests offline and tightened access controls. It said its new monitoring tools blocked the reported incidents when tested against them. The company is still reviewing records for further cases. Source: Anthropic.
