VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
Long video question answering requires locating sparse, time-scattered visual evidence within highly redundant content. Although current MLLMs perform well on short videos, long videos introduce long-horizon search and verification, which often necessitates multi-turn, agentic interaction. We show that existing LVU agents can exhibit "evidence misalignment": they produce correct answers that are not supported by the retrieved or inspected evidence. To characterize this failure, we introduce two
Record details
Published: 12 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignment
arXiv · 12 May 2026
No More, No Less: Task Alignment in Terminal Agents
arXiv · 12 May 2026
Leveraging RAG for Training-Free Alignment of LLMs
arXiv · 11 May 2026
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
arXiv · 13 May 2026
Position: Assistive Agents Need Accessibility Alignment
arXiv · 13 May 2026
Unweighted ranking for value-based decision making with uncertainty
arXiv · 13 May 2026
How to cite this record
ethics.ai (12 May 2026), “VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority,” evidence record 4475, https://ethics.ai/record/4475 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.