The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems
AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems. The dominant approach places controls inside the agent's own runtime: system prompts, output filters, and guardrail libraries. Any control in the agent's address space is reachable by inputs that influence it; this generalizes to any AI system with sufficient reach into its own runtime, a class we term escapable AI systems. We identify four properties that an authorization mecha
Record details
Published: 24 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Learning Action Priors for Cross-embodiment Robot Manipulation
arXiv · 24 June 2026
IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control
arXiv · 25 June 2026
GCT-MARL: Graph-Based Contrastive Transfer for Sample-Efficient Cooperative Multi-Agent Reinforcement Learning
arXiv · 23 June 2026
ASALT: Adaptive State Alignment for Lateral Transfer in Multi-agent Reinforcement Learning
arXiv · 23 June 2026
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior
arXiv · 26 June 2026
Agent-Native Immune System: Architecture, Taxonomy, and Engineering
arXiv · 26 June 2026
How to cite this record
ethics.ai (24 June 2026), “The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems,” evidence record 603, https://ethics.ai/record/603 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.