Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-confidence set of on-policy responses can provide a reliable learning signal. Our method, DMAPO (Data-centric Multi-evaluator Agreement for Preference Optimization), generates candidate responses from the target policy, evaluates helpfulness, factuality, and conciseness with rubric-specialized evaluators, applies a process-critic correction, and retains only
Record details
Published: 27 July 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 29 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
arXiv · 27 July 2026
Share No More Than the Request Requires: Federated Disclosure for Perspective-Aware AI
arXiv cs.CY · 28 July 2026
Regulating for AI Legitimacy
arXiv cs.CY · 28 July 2026
Inverse RL Helps Align AI by Imitating Humans
arXiv cs.LG · 27 July 2026
Regulating for AI Legitimacy
arXiv cs.AI · 27 July 2026
Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
arXiv · 28 July 2026
How to cite this record
ethics.ai (27 July 2026), “Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization,” evidence record 14165, https://ethics.ai/record/14165 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.