Harnessing Textual Refusal Directions for Multimodal Safety
To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. Both strategies are less feasible in Multimodal LLMs (MLLMs) as they require unsafe multimodal data, harder to collect than their unimodal counterpart. In this work, we relax this constraint and investigate whether textual refusal directions, extracted directly from the LLM backbone, generalize across modalities (i.e., image, video). Preliminary f
Record details
Published: 30 June 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Internal Pluralism and the Limits of Pairwise Comparisons
arXiv · 2 July 2026
Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs
HuggingFace Daily Papers · 3 July 2026
Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning
arXiv · 26 June 2026
AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing
arXiv · 25 June 2026
What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics
arXiv · 23 June 2026
ASALT: Adaptive State Alignment for Lateral Transfer in Multi-agent Reinforcement Learning
arXiv · 23 June 2026
How to cite this record
ethics.ai (30 June 2026), “Harnessing Textual Refusal Directions for Multimodal Safety,” evidence record 385, https://ethics.ai/record/385 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.