Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap. The natural fix is a guard-agnostic recover-and-decode amplifier that transcribes image content and restates encoded text into its plain payload before the guard, so any off
Record details
Published: 29 July 2026
Source: arXiv red teaming query
Category: Research
Topics: Safety & alignment · Military & security
Retrieved: 30 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning
arXiv · 29 July 2026
A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models
arXiv · 30 July 2026
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
arXiv red teaming query · 27 July 2026
Continuous surrogates versus threshold Boolean networks for modeling Arabidopsis ISR gene regulation
arXiv cs.LG · 25 July 2026
Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks
arXiv red teaming query · 2 August 2026
DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection
arXiv red teaming query · 22 July 2026
How to cite this record
ethics.ai (29 July 2026), “Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks,” evidence record 14849, https://ethics.ai/record/14849 (originally published by arXiv red teaming query).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.