Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits
Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as targets and as adversaries, exposing mechanistic vulnerabilities in diffusion-based alignment. We first show that safety alignment in DLLMs remains sparse and transferable across architectures. DLLMs initialized from autoregressive predecessors inherit the same mechanistic
Record details
Published: 7 August 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Finance, VC & PE
Retrieved: 10 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe
arXiv · 7 August 2026
Divided Attention Amplifies the Importance of Expectation-Aligned Visualization Design
arXiv cs.HC · 10 August 2026
ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization
arXiv red teaming query · 4 August 2026
Effect of a robotic insole-type active assist device on horizontal ground reaction force and center-of-pressure stability during stepping in patients with medial knee osteoarthritis
Frontiers in Robotics and AI · 12 August 2026
NAE: Normalizing AutoEncoder
arXiv cs.LG · 12 August 2026
A Data-Centric Perspective on Tree Visualizations
arXiv cs.HC · 2 August 2026
How to cite this record
ethics.ai (7 August 2026), “Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits,” evidence record 17795, https://ethics.ai/record/17795 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.