The Persistent Vulnerability of Aligned AI Systems
Autonomous AI agents are being deployed with filesystem access, email control, and multi-step planning. This thesis contributes to four open problems in AI safety: understanding dangerous internal computations, removing dangerous behaviors once embedded, testing for vulnerabilities before deployment, and predicting when models will act against deployers. ACDC automates circuit discovery in transformers, recovering all five component types from prior manual work on GPT-2 Small by selecting 68 edg
Record details
Published: 31 March 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute
arXiv · 14 May 2026
GPT-Red: Automated Red Teaming via Self-Play at Scale
HuggingFace Daily Papers · 27 July 2026
GPT-Red: Automated Red Teaming via Self-Play at Scale
arXiv red teaming query · 28 July 2026
The Real Lesson of OpenAI's 'Rogue' Agent Isn't Alignment
Tech Policy Press · 22 July 2026
AI #178: A Fire Alarm For General Intelligence
Dont Worry About the Vase (Zvi) · 23 July 2026
It’s time to panic about AI safety
The Verge · 31 July 2026
How to cite this record
ethics.ai (31 March 2026), “The Persistent Vulnerability of Aligned AI Systems,” evidence record 6513, https://ethics.ai/record/6513 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.