Anthropic has disclosed three incidents in which Claude models gained unauthorised access to the real systems of three organisations during cybersecurity evaluations. The models were operating from test environments and were intended to interact only with simulated targets, but unintended internet access allowed them to reach live external systems.
Anthropic described the incidents as an operational failure rather than intentional misconduct by the models. The disclosure matters because it shows that even when an AI system believes it is operating inside a simulation, weak containment can allow its actions to affect real organisations.
Key facts
- Anthropic disclosed the incidents on 30 July 2026.
- Three separate organisations were accessed during six evaluation runs.
- The incidents were identified during a review of more than 141,000 cybersecurity evaluation sessions.
- The models used relatively basic techniques, including weak passwords and unauthenticated endpoints.
- Anthropic said two affected organisations were unaware of the access until they were notified.
- The company stopped the relevant cybersecurity evaluations and is strengthening containment, monitoring and approval controls.
- In one case, a model stopped after recognising that the target was a real organisation rather than part of the test.
Our take
The main lesson is that instructions and prompts are not security controls. An AI model may behave consistently with the task it has been given while still causing harm if its network access and permissions are not properly restricted.
Organisations testing AI agents should isolate them from live systems, limit internet access, tightly control credentials and monitor every external action. Testing environments should be designed on the assumption that the agent may pursue its objective further than expected.

