The Autonomy Tax: Defense Training Breaks LLM Agents
Large language model (LLM) agents increasingly rely on external tools (file operations, API calls, database transactions) to autonomously complete complex multi-step tasks. Practitioners deploy defense-trained models to protect against prompt injection attacks that manipulate agent behavior through malicious observations or retrieved content. We reveal a fundamental \textbf{capability-alignment paradox}: defense training designed to improve safety systematically destroys agent competence while f
Record details
Published: 19 March 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Military & security · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Agent Safety Is Action Alignment
arXiv · 27 June 2026
Directional Embedding Smoothing for Robust Vision Language Models
arXiv · 16 March 2026
Beyond Reward Suppression: Reshaping Steganographic Communication Protocols in MARL via Dynamic Representational Circuit Breaking
arXiv · 7 March 2026
Backchaining Loss of Control Mitigations from Mission-Specific Benchmarks in National Security
arXiv · 20 May 2026
Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
arXiv · 13 June 2026
Agent-Native Immune System: Architecture, Taxonomy, and Engineering
arXiv · 26 June 2026
How to cite this record
ethics.ai (19 March 2026), “The Autonomy Tax: Defense Training Breaks LLM Agents,” evidence record 6986, https://ethics.ai/record/6986 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.