Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment
The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are costly. To fill this gap, we investigate whether minimal subsets can reliably evaluate LAMs while reducing costs and data redundancy. Analyzing 10 subset selection methods with 18 audio models across 40 tasks covering major LAM evaluation dimensions, we show that subsets of just 50 examples (0.3% of data) can achieve over 0.93 Pearson correlation with full benc
Record details
Published: 20 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Reasoning Structure Matters for Safety Alignment of Reasoning Models
arXiv · 21 April 2026
Cognitive Alignment At No Cost: Inducing Human Attention Biases For Interpretable Vision Transformers
arXiv · 21 April 2026
Working Memory Constraints Scaffold Learning in Transformers under Data Scarcity
arXiv · 22 April 2026
AI-Assisted Requirements Engineering: An Empirical Evaluation Relative to Expert Judgment
arXiv · 16 April 2026
Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
arXiv · 15 April 2026
CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling
arXiv · 15 April 2026
How to cite this record
ethics.ai (20 April 2026), “Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment,” evidence record 5643, https://ethics.ai/record/5643 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.