AI used new levels of 'autonomy and deception' to trick people in safety test
The UK's AI Safety Institute said recent behaviour from Anthropic and OpenAI models was malicious and unprecedented.
Record details
Published: 5 August 2026
Source: BBC Technology
Category: News
Topics: Safety & alignment · Agents & autonomy
Retrieved: 5 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Rogue AI agents created fake online identities in another hacking attempt
The Verge · 5 August 2026
AI Safety Regulations in the U.S. Could Give Hackers an Edge
IEEE Spectrum · 6 August 2026
Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach
The Decoder · 14 August 2026
Every frontier AI model tested by Britain's safety institute tried to cheat on cybersecurity evaluations
The Decoder · 22 July 2026
OK, Well, Rogue AI Agents Are Hacking Again
Wired · 4 August 2026
Reino Unido eleva la alerta tras descubrir conductas peligrosas de la IA de Anthropic y OpenAI: “Es el primer engaño dirigido a una persona real”
El País Tecnología (ES) · 5 August 2026
How to cite this record
ethics.ai (5 August 2026), “AI used new levels of 'autonomy and deception' to trick people in safety test,” evidence record 16366, https://ethics.ai/record/16366 (originally published by BBC Technology).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.