Towards Visually Grounded Multimodal Summarization via Cross-Modal Transformer and Gated Attention
Multimodal summarization requires models to jointly understand textual and visual inputs to generate concise, semantically coherent summaries. Existing methods often inject shallow visual features into deep language models, leading to representational mismatches and weak cross-modal grounding. We propose a unified framework that jointly performs text summarization and representative image selection. Our system, SPeCTrA-Sum (Sampler Perceiver with Cross-modal Transformer and gated Attention for S
Record details
Published: 12 May 2026
Source: arXiv
Category: Research
Topics: unclassified
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
How to cite this record
ethics.ai (12 May 2026), “Towards Visually Grounded Multimodal Summarization via Cross-Modal Transformer and Gated Attention,” evidence record 4488, https://ethics.ai/record/4488 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.