Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
arXiv:2601.17003v2 Announce Type: replace Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual diversity of deployment. We pair four benchmark replications with an ecological audit of real-world conversations to evaluate a purpose-built mental-health AI alongside six frontier general-purpose models spanning four families (OpenAI GPT-5, GPT-5.1, GPT-5.2; DeepSeek V3; Google Gemini 3 Flash; Moonshot Kimi
Record details
Published: 5 August 2026
Source: arXiv cs.CY
Category: Research
Topics: Safety & alignment · Healthcare · Transparency
Retrieved: 5 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Performance of 5 Large Language Models in Perioperative Consultation for Pediatric Hypospadias: Cross-Sectional Comparative Study
JMIR (Journal of Medical Internet Research) · 29 July 2026
Generative AI in Academic Writing: A Comparison of DeepSeek, Qwen, ChatGPT, Gemini, Llama, Mistral, and Gemma
OpenAlex · 30 April 2026
Hallucinating with AI: Distributed Delusions and “AI Psychosis”
OpenAlex · 11 February 2026
Six Institutional Intervention Areas to Support Ethical and Effective Student Use of Generative AI in Higher Education: A Narrative Review
OpenAlex · 16 January 2026
Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations
arXiv cs.CY · 7 August 2026
When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
arXiv cs.CY · 21 July 2026
How to cite this record
ethics.ai (5 August 2026), “Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety,” evidence record 16207, https://ethics.ai/record/16207 (originally published by arXiv cs.CY).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.