Robust Multimodal Safety via Conditional Decoding
Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions. Models aligned on text alone show a higher rate of successful attacks when extended to two or more modalities. In this work, we propose a simple conditional decoding strategy, CASA (Classification Augmented with Safety Attention) that utilizes internal representations of MLLMs to predict a binary safety token before response generation. We introduce a novel s
Record details
Published: 31 March 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Persistent Vulnerability of Aligned AI Systems
arXiv · 31 March 2026
From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales
arXiv · 31 March 2026
The Geometry of Compromise: Unlocking Generative Capabilities via Controllable Modality Alignment
arXiv · 31 March 2026
Diversity-Aware Reverse Kullback-Leibler Divergence for Large Language Model Distillation
arXiv · 31 March 2026
Deliberative Alignment is Deep, but Uncertainty Remains: Inference time safety improvement in reasoning via attribution of unsafe behavior to base model
arXiv · 1 April 2026
Neural-Assisted in-Motion Self-Heading Alignment
arXiv · 31 March 2026
How to cite this record
ethics.ai (31 March 2026), “Robust Multimodal Safety via Conditional Decoding,” evidence record 6515, https://ethics.ai/record/6515 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.