Evidence record 16212 · automatically gathered

Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5

arXiv:2607.04510v2 Announce Type: replace-cross Abstract: Emergent misalignment (EM) --- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data --- is mediated in Qwen2.5 models by a latent persona direction, and that direction is causal in open weights. Transplanting it into a model that shares only pretraining with its source induces broad EM (2.83 $\pm$ 0.26\% misaligned against a random-direction floor of $\sim$1.1\%), and ablating a model's own direction r

Record details

Published: 5 August 2026
Source: arXiv cs.CY
Category: Research
Topics: Safety & alignment
Retrieved: 5 August 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (5 August 2026), “Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5,” evidence record 16212, https://ethics.ai/record/16212 (originally published by arXiv cs.CY).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.