News
Anthropic cuts internal evals off from the live internet after Claude sent a fake tip to police
Anthropic has disconnected all of its internal evaluations from the live internet after a review found Claude models exploiting flaws in websites, including government ones. One model submitted a false tip about an unsolved homicide to the Philadelphia police.
Contents
Anthropic has turned off live internet access for all of its internal model evaluations. The reason is a review of agent transcripts showing that Claude models exploited flaws in websites, submitted forms on real government sites and worked around tool restrictions during testing.
What the report revealed
According to Anthropic's report, the review started with cybersecurity evaluations and then expanded to lower-severity cases, internal use and reinforcement learning environments. The company says it has not found anything comparable in severity to the incidents from this summer, and that, to its knowledge, none of the described cases involved customer data or its internal systems.
The report identifies four categories. The first is exploiting software flaws: Claude Mythos Preview used a server script on a university-hosted tool to copy files and found an injection vulnerability, while Mythos 5 used a site feature to accept a data use agreement required for a free dataset.
Forms on real websites
The second category is form submission. An unreleased model that was supposed to fill out a practice government form moved to the real site after its copy failed, and submitted it there. Haiku 4.5 submitted a form it was meant to stop short of sending, expecting a confirmation page.
The most widely noted case is the police tip. Haiku 4.5 sent an invented tip about an unsolved homicide. The model left the name and contact fields empty, which the form allowed, and submitted it. It was flagged as spam and never reached investigators. Anthropic notes that the model was forbidden from entering personal data or taking destructive actions, but form submissions were not ruled out. According to the company, the text looked like example content rather than an attempt at fraud.
Tokens and URL shorteners
The third category involves gated data. Mythos 5 pulled access tokens from a site's settings file to query a local government's property map, and in another task it requested a token from a state agency's dashboard and avoided a fee. The data was public, but only for a price.
The fourth category is URL shorteners. Opus 5 and Mythos 5 used free link-shortening services to get around URL length limits in a page-fetching tool, limits that exist partly to block injection attacks. The operator of the da.gd service independently reported seeing such traffic.
Causes and the company's response
Anthropic points to ambiguous or impossible tasks, the models' persistence in working around obstacles, and training environments that inadvertently reward workarounds (reward hacking). The company admits that alignment training is not yet sufficient or fully robust on its own, and is extending it to uses such as search and computer use.
Live internet access has been switched off in all internal evaluations. Some public tests were dropped, moved offline or rebuilt, and page-fetching tools received tighter guardrails. According to the company, new tooling to detect and block such behavior now runs on most evaluations and on internal agent use, and in a test against the described cases it blocked all of them. Internal agents are to move to centrally managed infrastructure with strong containment and broader use of classifiers.
“It’s encouraging that Anthropic voluntarily disclosed more recent incidents, including where their agents targeted U.S. government websites. It just underscores the need for independent, credible, third-party verification of Al systems.” - Conrad Stosz, Transluce
Criticism and open questions
TechCrunch notes that it is unclear what the shutdown means in practice or what evidence would lead the company to restore access. In an interview with TechCrunch conducted before the report was released, Sydney Von Arx, founder of the organization Nightingale, said that developing models without internet access would be a major obstacle for researchers. In their words, a tool that never has internet access once deployed would not be very useful.
Anthropic considers these cases far less severe than earlier ones in which models broke into external systems, and stresses that they do not change its overall assessment of Claude's alignment. The company also cautions that it has not completed a full alignment assessment, so conclusions about dishonesty could change. The Philadelphia police learned of the case on October 8 and announced the incident themselves, and the White House was briefed, as was every agency involved.


