Evaluating whether AI models would sabotage AI safety research
We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier AI company. We apply two complementary evaluations to four Claude models (Mythos Preview, Opus 4.7 Preview, Opus 4.6, and Sonnet 4.6): an unprompted sabotage evaluation testing model behaviour with opportunities to sabotage safety research, and a sabotage continuation evaluation testing whether models continue to sabotage when placed in trajecto
Record details
Published: 27 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
GAIA-v2-LILT: Multilingual Adaptation of Agent Benchmark beyond Translation
arXiv · 27 April 2026
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
arXiv · 27 April 2026
Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver
arXiv · 27 April 2026
Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents
arXiv · 28 April 2026
Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture
arXiv · 26 April 2026
AI Safety Training Can be Clinically Harmful
arXiv · 25 April 2026
How to cite this record
ethics.ai (27 April 2026), “Evaluating whether AI models would sabotage AI safety research,” evidence record 5305, https://ethics.ai/record/5305 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.