Agent Safety Is Action Alignment
Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe, practitioners imported the chatbot-era recipe (train the model to refuse unsafe inputs) into the agentic setting, and treat the resulting capability loss as a manageable ``alignment tax.'' We argue this is a \emph{category error}. Refusal is a primitive for \emph{content safety}, where the harm is in the model's output and is therefore a learnabl
Record details
Published: 27 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Autonomy Tax: Defense Training Breaks LLM Agents
arXiv · 19 March 2026
Agent-Native Immune System: Architecture, Taxonomy, and Engineering
arXiv · 26 June 2026
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior
arXiv · 26 June 2026
IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control
arXiv · 25 June 2026
Learning Action Priors for Cross-embodiment Robot Manipulation
arXiv · 24 June 2026
The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems
arXiv · 24 June 2026
How to cite this record
ethics.ai (27 June 2026), “Agent Safety Is Action Alignment,” evidence record 502, https://ethics.ai/record/502 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.