On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment
Tool-using LLM agents fail through trajectories rather than only final responses, as they may execute unsafe tool calls, follow injected instructions, comply with harmful requests, or over-refuse benign tasks despite producing a seemingly safe answer. Existing safety-alignment signals are largely response-level or off-policy, and often incur a safety-utility trade-off: improving agent safety comes at the cost of degraded task performance. Such sparse and single-objective rewards severely limit r
Record details
Published: 12 May 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Guided Streaming Stochastic Interpolant Policy
arXiv · 11 May 2026
Containment Verification: AI Safety Guarantees Independent of Alignment
arXiv · 9 May 2026
Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents
arXiv · 7 May 2026
REBAR: Reference Ethical Benchmark for Autonomy Readiness
arXiv · 18 May 2026
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
arXiv · 5 May 2026
AI Alignment via Incentives and Correction
arXiv · 2 May 2026
How to cite this record
ethics.ai (12 May 2026), “On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment,” evidence record 4481, https://ethics.ai/record/4481 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.