When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs
Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, code, or domain specialists into a safety-aligned base using task arithmetic, TIES, or DARE. This convenience is known to carry a safety cost, but almost all of that evidence rests on static refusal tests: fixed harmful prompts scored for compliance. We argue this is misleading. Because safety alignment is "shallow," concentrated in the first few
Record details
Published: 9 August 2026
Source: arXiv red teaming query
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 11 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast
arXiv cs.LG · 9 August 2026
Faithful or evasive? An empirical study on translation norm preferences of Chinese and American LLMs in Chinese official political and policy discourse
Frontiers in Artificial Intelligence · 10 August 2026
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks
arXiv cs.AI · 10 August 2026
MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents
arXiv · 7 August 2026
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning
arXiv · 11 August 2026
Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training
arXiv · 7 August 2026
How to cite this record
ethics.ai (9 August 2026), “When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMs,” evidence record 18313, https://ethics.ai/record/18313 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.