What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?
Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models interpret different types of compliance demonstrations. We study this by mixing benign compliance demonstrations (non-harmful request, helpful response) with harmful compliance demonstrations (harmful request, helpful response) and testing three hypotheses about how demonstration composition drives harmful compliance. Across four models, we find that benign and harmful demonstrati
Record details
Published: 18 June 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
FinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming
arXiv · 18 June 2026
Challenges to Grassroots Organization Engagement with AI Policy
arXiv · 18 June 2026
Understanding Censorship in Large Language Models: From Mechanisms to Governance
arXiv · 16 June 2026
The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs
arXiv · 21 June 2026
Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts
arXiv · 22 June 2026
GAS-Leak-LLM: Genetic Algorithm-Based Suffix Optimization for Black-Box LLM Jailbreaking
arXiv · 14 June 2026
How to cite this record
ethics.ai (18 June 2026), “What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?,” evidence record 792, https://ethics.ai/record/792 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.