Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity
MLLM-as-a-Judge is conventionally validated by agreement with human annotations, but this metric is undefined when the human pool is culturally heterogeneous. We introduce VOIR DIRE, a multimodal benchmark of 626 culturally paired image--prompt artifacts spanning U.S. and mainland Chinese contexts across food, fashion, and architecture, with annotator pools that are within-pool reliable (a = 0.86/0.74) but cross-pool divergent on evaluation (Q1 r = -0.12). Across six MLLMs, the bias decomposes i
Record details
Published: 12 June 2026
Source: arXiv
Category: Research
Topics: Bias & fairness
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
It's the humans, not the data: Geopolitical bias in LLMs originates in post-training, amplified by the language of the prompt
arXiv · 22 May 2026
CN-Buzz2Portfolio: A Chinese-Market Dataset and Benchmark for LLM-Based Macro and Sector Asset Allocation from Daily Trending Financial News
arXiv · 18 March 2026
Total cholesterol, high-density lipoprotein, and glucose (CHG) index and diabetic retinopathy in middle-aged and elderly Chinese adults with diabetes: a cross-sectional study
OpenAlex · 22 January 2026
Alibaba reportedly seeks sale of gaming unit Lingxi Games, valuation starts at $1.03 billion
TechNode (CN) · 24 June 2026
Inside the ‘Culture of Fear’ at One American-Chinese University
Inside Higher Ed Tech · 20 July 2026
Fintech firm Ant International raises US$1.2b to fuel global growth
SCMP Tech (HK/CN) · 21 July 2026
How to cite this record
ethics.ai (12 June 2026), “Jury Duty: Calibration and Orientation Failures in MLLM-as-a-Judge Under Cultural Ambiguity,” evidence record 1050, https://ethics.ai/record/1050 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.