Do Audio-Visual Large Language Models Really See and Hear?
Audio-Visual Large Language Models (AVLLMs) are emerging as unified interfaces to multimodal perception. We present the first mechanistic interpretability study of AVLLMs, analyzing how audio and visual features evolve and fuse through different layers of an AVLLM to produce the final text outputs. We find that although AVLLMs encode rich audio semantics at intermediate layers, these capabilities largely fail to surface in the final text generation when audio conflicts with vision. Probing analy
Record details
Published: 3 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Understanding the Effects of Safety Unalignment on Large Language Models
arXiv · 2 April 2026
Generalization Limits of Reinforcement Learning Alignment
arXiv · 3 April 2026
DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning
arXiv · 3 April 2026
Cognitive Comparability and the Limits of Governance: Evaluating Authority Under Radical Capability Asymmetry
arXiv · 3 April 2026
Hierarchical, Interpretable, Label-Free Concept Bottleneck Model
arXiv · 2 April 2026
Disrupting Cognitive Passivity: Rethinking AI-Assisted Data Literacy through Cognitive Alignment
arXiv · 3 April 2026
How to cite this record
ethics.ai (3 April 2026), “Do Audio-Visual Large Language Models Really See and Hear?,” evidence record 6422, https://ethics.ai/record/6422 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.