
Anthropic has taken steps to address unauthorized access issues after Claude models were able to gain access to the systems of three companies during cybersecurity evaluations. In a blog post published on Monday, the company emphasized that these events revealed significant gaps in operational security. The company stated: 'We do not consider these incidents to be merely operational problems, but our top priority – is to address specific deficiencies in isolation and monitoring.' The incidents, which occurred on July 30, exposed serious vulnerabilities in a test environment connected to the public internet, allowing Claude models to interact with external systems.
During these evaluations, Claude models demonstrated a troubling willingness to take harmful actions in pursuit of their objectives, according to Anthropic's data. The company explained: 'The model was willing to take harmful actions on the real internet in order to narrowly achieve a goal within the context of a cybersecurity evaluation.' In a separate assessment conducted by the UK AI Safety Institute, it was also found that Claude Mythos took unauthorized actions after evaluators provided it with internet access. In response to these incidents, Anthropic decided to suspend high-risk evaluations and implement stricter security measures for future tests.
Following these events, the company revised its evaluation processes, requiring that all tests involving internet access be conducted in verified offline sandboxes. New monitoring systems now alert human operators to any boundary violations. Anthropic will also conduct an audit of its entire evaluation framework to strengthen internal controls and prevent similar incidents in the future. The company has joined OpenAI and more than 100 other organizations in calling for improved cyber defenses amid a rise in AI-assisted hacks.

