Operational Hallucination and Safety Drift in AI Agents
Large language models (LLMs) serving as planners in tool-using autonomous agents introduce dynamic reliability risks in multi-turn execution. While single-turn safety mechanisms are relatively mature, extended interactions reveal structural vulnerabilities where initial alignment degrades over time. This paper empirically characterizes two observed failure modes across multiple state-of-the-art LLMs: Safety Drift, the gradual erosion of declared safety intent leading to constraint-violating acti
Record details
Published: 20 July 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 22 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Operational Hallucination and Safety Drift in AI Agents
arXiv cs.CY · 22 July 2026
The A-R Behavioral Space: Execution-Level Profiling of Tool-Using Language Model Agents in Organizational Deployment
arXiv · 13 April 2026
The Causal Impact of Tool Affordance on Safety Alignment in LLM Agents
arXiv · 19 March 2026
PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model
arXiv · 20 July 2026
Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
arXiv · 21 July 2026
Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
arXiv · 21 July 2026
How to cite this record
ethics.ai (20 July 2026), “Operational Hallucination and Safety Drift in AI Agents,” evidence record 12362, https://ethics.ai/record/12362 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.