Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security
Large Language Models (LLMs) are increasingly vulnerable to adversarial prompts that exploit semantic ambiguities to bypass safety mechanisms, resulting in harmful or inappropriate outputs. Such attacks, including jailbreaking and prompt injection, pose significant risks to the integrity and availability of LLMs in security-critical applications. This paper proposes the Adversarial Prompt Disentanglement (APD) framework, a novel defense mechanism that proactively identifies and neutralizes malic
Record details
Published: 27 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Military & security
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
SPARD: Defending Harmful Fine-Tuning Attack via Safety Projection with Relevance-Diversity Data Selection
arXiv · 27 May 2026
Jailbreak to Protect: Buffering and Reinforcing via Temporary Jailbreaking for Safe Fine-Tuning in Large Language Models
arXiv · 23 May 2026
Repurposing Adversarial Perturbations for Continual Learning: From Defense to Active Alignment
arXiv · 1 June 2026
Patcher: Post-Hoc Patching of Backdoored Large Language Models
arXiv · 2 June 2026
Backchaining Loss of Control Mitigations from Mission-Specific Benchmarks in National Security
arXiv · 20 May 2026
Fusion-fission forecasts when AI will shift to undesirable behavior
arXiv · 14 May 2026
How to cite this record
ethics.ai (27 May 2026), “Disentangling Adversarial Prompts: A Semantic-Graph Defense for Robust LLM Security,” evidence record 3634, https://ethics.ai/record/3634 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.