Sustaining AI safety: Control-theoretic external impossibility, intrinsic necessity, and structural requirements
As AI systems become increasingly capable, safety strategies must be evaluated not only by how much they reduce present risk, but by whether they could sustain safety once external control can no longer reliably constrain system behavior. This paper addresses that problem by using control theory to clarify, at a structural level, whether externally enforced safety-sustaining strategies can succeed and, if not, what any alternative strategy would have to satisfy in order to be viable. It establis
Record details
Published: 13 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
arXiv · 13 May 2026
Lying Is Just a Phase: The Hidden Alignment Transition in Language Model Scaling
arXiv · 13 May 2026
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
arXiv · 13 May 2026
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
arXiv · 13 May 2026
GRACE: Gradient-aligned Reasoning Data Curation for Efficient Post-training
arXiv · 13 May 2026
Exact Linear Attention
arXiv · 13 May 2026
How to cite this record
ethics.ai (13 May 2026), “Sustaining AI safety: Control-theoretic external impossibility, intrinsic necessity, and structural requirements,” evidence record 4420, https://ethics.ai/record/4420 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.