Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models often rely on implicit inference without sufficient visual evidence, leading to a disconnect between perception and reasoning. Meanwhile, existing outcome-oriented benchmarks evaluate only final predictions and fail to diagnose failures in the underlying reasoning process
Record details
Published: 11 August 2026
Source: arXiv
Category: Research
Topics: Healthcare
Retrieved: 12 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
CARE: Confidence-Aware Reasoning for Reliable Medical VQA
arXiv cs.AI · 11 August 2026
SCOUT: Symmetric Consensus Outlier Detection for Failure Localization in LLM Pre-Training
arXiv cs.LG · 11 August 2026
Finite State Machine–Guided Retrieval-Augmented Generation Improves Expert-Rated Acceptability of a Peripherally Inserted Central Catheter Self-Management Chatbot: Single-Center Content Validation Study
JMIR (Journal of Medical Internet Research) · 11 August 2026
MIRA: Medical Image Reflection for Agentic Diagnosis
arXiv cs.AI · 11 August 2026
Trump wants MMR vaccine split up: the science behind why it’s a bad idea
Nature Machine Intelligence · 11 August 2026
Internet Attachment–Based Compassion Therapy for Adults With Chronic Medical Conditions: Randomized Controlled Trial
JMIR (Journal of Medical Internet Research) · 11 August 2026
How to cite this record
ethics.ai (11 August 2026), “Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes,” evidence record 18409, https://ethics.ai/record/18409 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.