S-SPPO: Semantic-Calibrated Self-Play Preference Optimization
Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO). However, the standard Bradley-Terry instantiation of DPO is limited in modeling common departures from transitivity in human preferences. To address this, recent work has introduced Self-Play Preference Optimization (SPPO), which iteratively refines the policy by training on self-generated win-lose pairs. Our investigation, however, reveals a critical instability in SPPO: th
Record details
Published: 1 June 2026
Source: arXiv
Category: Research
Topics: Regulation · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Federating Governance: How Community Rules Scale with Mastodon Instances
arXiv · 3 June 2026
EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization in Open-Ended QA
arXiv · 27 May 2026
A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation
arXiv · 9 June 2026
AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks
arXiv · 10 June 2026
Demographic Patterns in Cybersecurity Culture: Insights from a Global Organisation Supporting Safety-Critical and Critical Infrastructure Sectors
arXiv · 12 June 2026
Chunking German Legal Code
arXiv · 19 May 2026
How to cite this record
ethics.ai (1 June 2026), “S-SPPO: Semantic-Calibrated Self-Play Preference Optimization,” evidence record 3315, https://ethics.ai/record/3315 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.