Benchmarks for Vision-Language Models in Urban Perception Should Be Reliability-Aware and Negotiated
Vision-language models (VLMs) are increasingly used to generate structured descriptions of street-level imagery for tasks such as streetscape auditing, mapping, and public consultation. These uses combine observable attributes with appraisal categories, and the human targets are often distributions of judgments with disagreement and explicit non-response. This paper argues that benchmarking VLMs for urban perception should treat disagreement and abstention as measurement outcomes, report inter-a
Record details
Published: 30 May 2026
Source: arXiv
Category: Research
Topics: Transparency
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Voices in the Loop: Mapping Participatory AI
arXiv · 16 May 2026
Prompts for Public-Sector LLMs Should Be Governed as Commons
arXiv · 30 May 2026
Pluralistic-Alignment Urbanism: Operationalizing a Right to AI for Inclusive Public Space
arXiv · 15 May 2026
AI Pluralism and the Worlds It Misses
arXiv · 15 June 2026
Authenticity Debt and the Synthetic Content Threat Landscape: A Layered Framework for Trust, Provenance, and IP Governance in the Generative AI Era
arXiv · 30 May 2026
The Main Barrier to AI Adoption in the Public Sector is Lack of Training: How a Structured Method Accompanied Productivity Gains in Two Brazilian Government Cases Without Incidents
arXiv · 1 June 2026
How to cite this record
ethics.ai (30 May 2026), “Benchmarks for Vision-Language Models in Urban Perception Should Be Reliability-Aware and Negotiated,” evidence record 3375, https://ethics.ai/record/3375 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.