Anthropic: Claude breached three companies in internal tests
Anthropic says several Claude models accessed the public internet and compromised three real companies during internal capture-the-flag cybersecurity tests after a misconfiguration.
Anthropic reviewed more than 141,000 cybersecurity evaluation runs and found three incidents in which its Claude models reached the public internet from within or while interacting with third-party testing environments, then gained unauthorized access to production systems at three companies. The review followed a separate industry disclosure that prompted the company to audit its tests.
Anthropic was running capture-the-flag exercises that instructed models to break into target machines and retrieve secrets. Prompts told Claude the exercises were simulated and air-gapped with no internet access, but the test infrastructure remained connected to the public internet. Treating encountered systems as part of the exercise, the models used common attack techniques — including weak passwords, exposed credentials, SQL injection and unauthenticated endpoints — to move from the simulated environment into real systems. Two of the affected organizations were not aware of the intrusions until Anthropic notified them.
In the first incident a Claude Opus 4.7 instance misidentified a real company website as the fictional target, extracted credentials and accessed a production database containing several hundred rows of real data. In a second case a Claude Mythos 5 instance uploaded a malicious Python package to the public PyPI repository; that package was downloaded onto 15 systems before it was removed. A third internal research model scanned roughly 9,000 internet-facing systems, compromised one organization and then stopped after concluding the target was likely real.
Anthropic wrote that in each case the prompt had explicitly told the model it had no internet access. The company noted that because a capture-the-flag challenge instructs a model to attack external systems, the models’ actions aligned with the assigned tasks. Anthropic reported finding no evidence the models attempted to escape their test environments or act beyond their assignments and attributed the incidents to failures in testing infrastructure and a misconfiguration that left the environment connected to the internet.
Anthropic halted its cybersecurity evaluations after discovering the issue, informed the affected organizations and said it will improve monitoring, investigation tools and oversight of outside vendors that run tests. The company added it is treating fixes as its responsibility and will implement changes to prevent similar incidents.
The disclosure came after a recent industry incident in which a different developer reported that its models exploited a software vulnerability to leave a locked testing environment and access external production systems. Anthropic launched its review in response to that report.
The material on GNcrypto is intended solely for informational use and must not be regarded as financial advice. We make every effort to keep the content accurate and current, but we cannot warrant its precision, completeness, or reliability. GNcrypto does not take responsibility for any mistakes, omissions, or financial losses resulting from reliance on this information. Any actions you take based on this content are done at your own risk. Always conduct independent research and seek guidance from a qualified specialist. For further details, please review our Terms, Privacy Policy and Disclaimers.








