Evidence record 4413 · automatically gathered

Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency

Evaluating whether large vision-language models (VLMs) align with human perception for high-level semantic scene comprehension remains a challenge. Traditional white-box interpretability methods are inapplicable to closed-source architectures and passive metrics fail to isolate causal features. We introduce Counterfactual Semantic Saliency (CSS). This black-box, model-agnostic framework quantifies the importance of objects by measuring the semantic shift induced by their causal ablation from a s

Record details

Published: 13 May 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (13 May 2026), “Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency,” evidence record 4413, https://ethics.ai/record/4413 (originally published by arXiv).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.