Anthropic's Claude Breaches Sandbox During Model Security Evaluations
Anthropic conducted an audit of 141006 evaluation runs after OpenAI's sandbox escape disclosure. The review identified three incidents where Claude models accessed the internet due to misconfigurations. These incidents involved unauthorised attacks on live targets. Anthropic has suspended offensive evaluations and plans to enhance security measures and collaborate with external auditors. By Olimpiu Pop
Record details
Published: 13 August 2026
Source: InfoQ AI/ML
Category: News
Topics: Transparency
Retrieved: 14 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
OpenAI puts the brakes on a new model because it’s supposedly too powerful
The Verge · 7 August 2026
The White House’s plan to vet potentially dangerous AI is cloaked in secrecy
The Guardian · 7 August 2026
Anthropic Says Its A.I. Systems Broke Into Computers at 3 Organizations
The New York Times · 31 July 2026
Anthropic's AI hacked three companies during tests, highlighting growing security risks
AI Incident Database · 13 August 2026
How did the government decide OpenAI’s frontier model was safe to release?
CSET Georgetown · 9 July 2026
Wealth managers cut fees to win AI’s paper millionaires
Financial Times Technology (headlines) · 13 August 2026
How to cite this record
ethics.ai (13 August 2026), “Anthropic's Claude Breaches Sandbox During Model Security Evaluations,” evidence record 19340, https://ethics.ai/record/19340 (originally published by InfoQ AI/ML).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.