Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks
We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR). The operative change is in the defense pipeline, not in the image. Across five frontier VLMs, two encoded-attack families, and three black-box defenses, a caption-mediated defense (ECSO) that leaves ASR essentially unchanged on text-only encoded input drops i
Record details
Published: 2 August 2026
Source: arXiv red teaming query
Category: Research
Topics: Safety & alignment · Military & security
Retrieved: 4 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models
arXiv · 30 July 2026
Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning
arXiv · 29 July 2026
Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks
arXiv red teaming query · 29 July 2026
FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition
arXiv · 7 August 2026
The Anatomy of a Prompt Injection: A Component Model for Structured Analysis
arXiv red teaming query · 7 August 2026
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
arXiv red teaming query · 27 July 2026
How to cite this record
ethics.ai (2 August 2026), “Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks,” evidence record 16170, https://ethics.ai/record/16170 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.