AI Safety Training Can be Clinically Harmful
Large language models are being deployed as mental health support agents at scale, yet only 16% of LLM-based chatbot interventions have undergone rigorous clinical efficacy testing, and simulations reveal psychological deterioration in over one-third of cases. We evaluate four generative models on 250 Prolonged Exposure (PE) therapy scenarios and 146 CBT cognitive restructuring exercises (plus 29 severity-escalated variants), scored by a three-judge LLM panel. All models scored near-perfectly on
Record details
Published: 25 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Healthcare · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Think Before You Act -- A Neurocognitive Governance Model for Autonomous AI Agents
arXiv · 28 April 2026
Four-Axis Decision Alignment for Long-Horizon Enterprise AI Agents
arXiv · 21 April 2026
Hidden Coalitions in Multi-Agent AI: A Spectral Diagnostic from Internal Representations
arXiv · 4 May 2026
Market-Alignment Risk in Pricing Agents: Trace Diagnostics and Trace-Prior RL under Hidden Competitor State
arXiv · 7 May 2026
A Framework for Longitudinal Health AI Agents
arXiv · 13 April 2026
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
arXiv · 12 April 2026
How to cite this record
ethics.ai (25 April 2026), “AI Safety Training Can be Clinically Harmful,” evidence record 5362, https://ethics.ai/record/5362 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.