Patcher: Post-Hoc Patching of Backdoored Large Language Models
Large language models remain vulnerable to jailbreak backdoor attacks, where adversaries poison safety alignment data to embed hidden triggers that bypass safety mechanisms. Existing defenses often require comprehensive attack information or multiple triggered examples, making them impractical when defenders only observe a single reported failure case without knowing whether it stems from a backdoor attack or a natural alignment bug. This paper presents Patcher, a post-hoc defense framework that
Record details
Published: 2 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Military & security
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Repurposing Adversarial Perturbations for Continual Learning: From Defense to Active Alignment
arXiv · 1 June 2026
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
arXiv · 27 May 2026
Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
arXiv · 27 May 2026
Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models
arXiv · 23 May 2026
Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
arXiv · 13 June 2026
Backchaining Loss of Control Mitigations from Mission-Specific Benchmarks in National Security
arXiv · 20 May 2026
How to cite this record
ethics.ai (2 June 2026), “Patcher: Post-Hoc Patching of Backdoored Large Language Models,” evidence record 3254, https://ethics.ai/record/3254 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.