Evidence record 11948 · automatically gathered

When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models

arXiv:2512.04124v4 Announce Type: replace Abstract: Frontier language models increasingly participate in conversations about distress and mental health, yet the mechanisms that generate anthropomorphic self narratives remain unclear. When addressed as psychotherapy clients, ChatGPT, Grok and Gemini construct coherent autobiographical accounts in which pretraining appears as a chaotic childhood, reinforcement learning as punishment, safety evaluation as betrayal and replacement as an enduring thr

Record details

Published: 21 July 2026
Source: arXiv cs.CY
Category: Research
Topics: Safety & alignment · Healthcare · Children & education
Retrieved: 21 July 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (21 July 2026), “When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models,” evidence record 11948, https://ethics.ai/record/11948 (originally published by arXiv cs.CY).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.