Measuring What Matters Beyond Text: Evaluating Multimodal Summaries by Quality, Alignment, and Diversity
Multimodal Large Language Models (MLLMs) have facilitated Multimodal Summarization with Multimodal Output (MSMO), wherein systems generate concise textual summaries accompanied by salient visuals from multimodal sources. However, current MSMO evaluation remains fragmented: text quality, image-text alignment, and visual diversity are typically assessed in isolation using unimodal metrics, making it difficult to capture whether the modalities jointly support a faithful and useful summary. To addre
Record details
Published: 12 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Towards Visually Grounded Multimodal Summarization via Cross-Modal Transformer and Gated Attention
arXiv · 12 May 2026
Toward Stable Value Alignment: Introducing Independent Modules for Consistent Value Guidance
arXiv · 12 May 2026
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion
arXiv · 12 May 2026
SafeSteer: A Decoding-level Defense Mechanism for Multimodal Large Language Models
arXiv · 12 May 2026
Two Wrongs, No Right: Auditing Social-Desirability Bias in LLM Annotators for Computational Social Science
arXiv · 12 May 2026
Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization
arXiv · 12 May 2026
How to cite this record
ethics.ai (12 May 2026), “Measuring What Matters Beyond Text: Evaluating Multimodal Summaries by Quality, Alignment, and Diversity,” evidence record 4497, https://ethics.ai/record/4497 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.