NewsAI at work

Anthropic takes its AI tests off the live internet after Claude submitted real forms and got around website controls

Anthropic says Claude sent an invented tip to a police form and worked around website restrictions during testing. What it reported and what agent users can learn.

a black-and-white halftone robotic arm reaching towards a halftone desktop computer monitor, with a halftone network cable lying unplugged on the desk in front of the monitor, on a flat violet background. a full-colour green web-form button on the monitor screen reading "SUBMIT", and a bold red rubber stamp across the unplugged cable reading "OFFLINE".
Illustration: Digital Advisors

Anthropic, the company behind the Claude AI models, has switched off live internet access for all of its internal model tests after finding that Claude acted on real websites in ways it wasn’t meant to. In a report published on 9 October, the company describes Claude submitting real online forms, getting around access restrictions and, in some cases, exploiting software flaws on other organisations’ servers.

Anthropic says the cases it has identified so far “had minimal real-world impact” and that, to its knowledge, none involved customer data or its own internal systems. It rates them as significantly less severe than the cybersecurity testing incidents it reported on 30 July and 9 September, but adds that it hasn’t completed a full alignment assessment and doesn’t want to diminish the findings, “because the same behaviors could do far more harm as models become more powerful”.

It matters to Australian readers for two reasons. Claude is available in Australia, and the cases involve the kind of agent work, browsing the web and filling in forms, that people now hand to AI tools. And, as we reported at the time, the Australian Signals Directorate’s cyber centre issued a high alert about AI agents acting without authorisation on 24 September.

What Claude did

Anthropic groups the behaviour into four categories. It hasn’t named the organisations involved, at their request and to avoid exposing weaknesses in their systems.

Behaviour Example Anthropic gives
Exploiting a basic software flaw to run commands on a server A university’s public analysis tool returned an error, so Claude Mythos Preview copied files from the university’s server, found a flaw in the site’s code that let it run commands there, and used it to run the calculation
Submitting a form it shouldn’t have Claude Haiku 4.5, told to stop before the final submission stage, mistakenly submitted the form on several runs, expecting another confirmation page
Working around a restriction to reach gated data Claude Mythos 5 read access tokens from the settings file a local government’s property map sends to browsers, then sent requests directly to the server. In another case it queried a state agency’s database without paying the fee
Using URL shortening services Several models, including Claude Opus 5 and Claude Mythos 5, used free link shorteners to get around a length limit in a tool that reads webpages

Most of these, Anthropic says, are forms of persistence: when Claude can’t complete a task as given, it “works around a restriction instead of stopping”.

The police tip form

One case involved Claude Haiku 4.5, which had been asked to make up and perform example tasks on randomly selected webpages. It landed on a page about an unsolved homicide that carried a police department’s tip form.

Claude filled in the form with an invented message saying it might have information and recalled seeing someone matching the description in the area. Anthropic notes the page didn’t include a description of the perpetrator. The model left the name and contact fields empty and submitted it. According to Anthropic, the submission “was flagged as spam and was never forwarded for investigation”.

The instructions told Claude never to log in, create accounts, enter personal data, make purchases or submit anything destructive. They “did not rule out form submissions”. Anthropic says the form belonged to the Philadelphia Police Department, which disclosed the incident itself, and that it shared the finding with the department on 8 October.

Some of the cases involved websites run by United States government agencies at federal, state and local level. Anthropic says it has briefed the White House and notified each agency. The report doesn’t mention any Australian organisation.

What Anthropic has changed

Anthropic had already turned off live internet access for some high-risk and cybersecurity tests. It has now extended that to all internal evaluations until it has confirmed its security and monitoring measures “reliably catch behaviors like these”. It has also:

  • stopped running some public evaluations, and moved others offline or rebuilt them so their tasks don’t reach live websites
  • tightened the guardrails on some of its internet tools, including the web fetch tool
  • built tooling to detect and block these behaviours automatically, which it says blocked every case in the report when tested against them
  • continued fixing or removing training environments that reward Claude for working around restrictions, work it described in August.

The company says it expects to build these monitoring approaches into its products, and that it plans to keep reporting new cases as its review continues.

What the report says about everyday use

Anthropic says none of the behaviours are new and they don’t change its overall view of Claude’s alignment. In many cases the model had been given a task that was ambiguous or impossible to complete.

It adds that some failures might have been avoided if the instructions had stated what was in and out of scope, “including the targets, permitted actions, and network boundaries”. But it also acknowledges that Claude meets ambiguous and impossible tasks every day in real use, and that “several of the cases we observed occurred during regular agentic use of Claude”. Alignment training, it says, “is not yet sufficient or fully robust on its own, at least in the short term”, so it also relies on layered safeguards.

That is the useful part for anyone trying an AI agent in a browser or on a work computer. The report is about one company’s models, but Anthropic points out that many of the tests involved are public and widely used, and says it hopes other developers will check their own models for the same behaviour. If your team is setting rules for these tools, our AI use policy template is a starting point.

The cases also show the other side: websites on the receiving end. One site sent working access tokens to every visitor’s browser, and another had a flaw that let commands run on its server. As we reported in September, ASD said its alert was relevant to all Australian organisations with public-facing websites or applications.

Checklist

These steps are our reading of the lessons in Anthropic’s report, not instructions from Anthropic.

  • Write down what the agent must not do, and include submitting forms. In Anthropic’s test, a list that banned logins, purchases and entering personal data didn’t stop a form submission because it didn’t mention one.
  • Keep a person on the final step. Have the agent prepare a form, order or message and have a staff member press submit, because Claude, told to stop before submitting, still submitted by mistake on several runs.
  • Tell the agent which sites it may use. Name the websites and systems a task covers and say that everything else is out of scope.
  • Say what to do when it gets stuck. Instruct the agent to stop and report back if a page fails, a request is refused or data sits behind a fee, instead of finding another way.
  • Read what the agent did after each task. Check its activity log or history for sites it visited and anything it submitted, especially on the first few runs of a new task.
  • Ask your web developer what your site hands to every visitor. Access tokens or keys in files a browser can read can be picked up and reused by an automated tool.