LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition
The evolution of Large Language Model (LLM) reasoning is bottlenecked by the scarcity of high-quality process data. While self-alignment via endogenous rewards offers a solution, mining valid supervision faces three challenges: (1) Label Noise via Mimetic Bias, where rewards prioritize statistical likelihood over logical truth, creating a "correctness illusion" that masks compounding errors; (2) Coarse-Grained Supervision, where sparse global outcomes (e.g., in GRPO) fail to provide granular gui
Record details
Published: 19 May 2026
Source: arXiv
Category: Research
Topics: Bias & fairness · Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective
arXiv · 13 May 2026
Silent Failures in Federated Personalization of Foundation Models
arXiv · 31 May 2026
BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization
arXiv · 3 June 2026
OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models
arXiv · 25 April 2026
Prompt Programming for Cultural Bias and Alignment of Large Language Models
arXiv · 17 March 2026
LLM Constitutional Multi-Agent Governance
arXiv · 13 March 2026
How to cite this record
ethics.ai (19 May 2026), “LC-ERD: Mining Latent Logic for Self-Evolving Reasoning via Consistency-Regulated Reward Decomposition,” evidence record 4045, https://ethics.ai/record/4045 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.