Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation
Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples. We introduce a dataset-centric meta-evaluation framework that audits benchmark datasets at the sample level along five latent dimensions: 1. Cognitive and Knowledge Demands, 2. Language and Content Quality, 3. Task Properties, 4. Context, and 5. Ethics, Safety, and Fairness. Applying this framework,
Record details
Published: 30 July 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Transparency
Retrieved: 3 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
KAISEN: Reproducible Subgroup Fairness Auditing for Clinical Risk Models
arXiv cs.LG · 30 July 2026
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
arXiv · 31 July 2026
When AI Does the Work, What Is Learning For? Post-Instrumental Learning and the Risk of Capacity Dissolution
arXiv cs.CY · 31 July 2026
When AI Does the Work, What Is Learning For? Post-Instrumental Learning and the Risk of Capacity Dissolution
arXiv · 30 July 2026
Fairness Auditing: Lower Bounds on Company Manipulation
arXiv fairness query · 1 August 2026
Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening
arXiv cs.CY · 29 July 2026
How to cite this record
ethics.ai (30 July 2026), “Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation,” evidence record 15692, https://ethics.ai/record/15692 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.