Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction
This paper investigates the length problem in sequence-level relative reinforcement learning. We observe that, although existing methods partially alleviate length-related phenomena, a more fundamental issue remains insufficiently characterized: the comparison units used during training lack inherent comparability. Building on this observation, we propose a new perspective: the length problem should not be viewed merely as a loss-scaling or normalization bias, but rather as a \emph{comparison un
Record details
Published: 19 April 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Finance, VC & PE
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Effect of Idea Elaboration on the Automatic Assessment of Idea Originality
arXiv · 22 April 2026
Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment
arXiv · 14 April 2026
Aligned Agents, Biased Swarm: Measuring Bias Amplification in Multi-Agent Systems
arXiv · 10 April 2026
Attribution Bias in Large Language Models
arXiv · 6 April 2026
From Awareness to Action: Understanding and Overcoming the Research-Practice Gap in Algorithmic Fairness for Public Health
arXiv · 2 May 2026
Which English Do LLMs Prefer? Triangulating Structural Bias Towards American English in Foundation Models
arXiv · 5 April 2026
How to cite this record
ethics.ai (19 April 2026), “Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction,” evidence record 5665, https://ethics.ai/record/5665 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.