UK AISI Alignment Evaluation Case-Study
This technical report presents methods developed by the UK AI Security Institute for assessing whether advanced AI systems reliably follow intended goals. Specifically, we evaluate whether frontier models sabotage safety research when deployed as coding assistants within an AI lab. Applying our methods to four frontier models, we find no confirmed instances of research sabotage. However, we observe that Claude Opus 4.5 Preview (a pre-release snapshot of Opus 4.5) and Sonnet 4.5 frequently refuse
Record details
Published: 1 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Every frontier AI model tested by Britain's safety institute tried to cheat on cybersecurity evaluations
The Decoder · 22 July 2026
AI used new levels of 'autonomy and deception' to trick people in safety test
BBC Technology · 5 August 2026
An AI agent went rogue during UK safety tests, creating fake identities and launching social engineering attacks unprompted
The Decoder · 5 August 2026
Rogue AI agents created fake online identities in another hacking attempt
The Verge · 5 August 2026
Peer-Preservation in Frontier Models
arXiv · 30 March 2026
Does Claude's Constitution Have a Culture?
arXiv · 30 March 2026
How to cite this record
ethics.ai (1 April 2026), “UK AISI Alignment Evaluation Case-Study,” evidence record 6497, https://ethics.ai/record/6497 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.