ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy
Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards still make token-level credit assignment difficult. Existing scalable RL methods typically assign trajectory-level rewards uniformly across tokens, while recent entropy-aware approaches either rely on coarse detached heuristics or directly optimize true entropy, which can introduce non-local gradient components misaligned with sampled-token policy updates. We p
Record details
Published: 3 July 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
arXiv · 3 July 2026
Explainable Reinforcement Learning for Adaptive Traffic Signal Control
arXiv · 4 July 2026
From Battlefield to Boardroom: Strategic Red Teaming as an Epistemic Governance Instrument in the Age of AI
arXiv · 2 July 2026
Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity
arXiv · 1 July 2026
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
arXiv · 6 July 2026
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment
arXiv · 6 July 2026
How to cite this record
ethics.ai (3 July 2026), “ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy,” evidence record 285, https://ethics.ai/record/285 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.