Evidence-RL: Towards Evidence-intensive Visual Reasoning
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response,
Record details
Published: 7 August 2026
Source: HuggingFace Daily Papers
Category: Research
Topics: Transparency
Retrieved: 11 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings
arXiv cs.AI · 7 August 2026
Layered agency as a basis for allocating accountability for AI
AI & Society · 8 August 2026
Assessing AI-generated music detection in real-world broadcast monitoring
arXiv · 7 August 2026
An End-to-End Agent Auditing Engine
arXiv cs.AI · 7 August 2026
Beyond the Black Box: Interpretable Models of Human Randomisation Failures
arXiv cs.AI · 7 August 2026
Human-Centered Explainable AI for TinyML Edge Devices: A Pareto-Based Selection Framework with LLM-Guided Design
arXiv cs.HC · 7 August 2026
How to cite this record
ethics.ai (7 August 2026), “Evidence-RL: Towards Evidence-intensive Visual Reasoning,” evidence record 17977, https://ethics.ai/record/17977 (originally published by HuggingFace Daily Papers).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.