Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment
Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of "rollout then update", which inevitably results in off-policy training data. To resolve this, Importance sampling (IS) is proposed, while the token-level ratios compound over long sequences, causing severe variance exploded. A natural idea is "transferring" these off-policy token into on-policy token, so that the importance scores for correction are unnecessary. Following this idea, we prop
Record details
Published: 6 July 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
arXiv · 6 July 2026
Explainable Reinforcement Learning for Adaptive Traffic Signal Control
arXiv · 4 July 2026
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
arXiv · 3 July 2026
ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy
arXiv · 3 July 2026
From Battlefield to Boardroom: Strategic Red Teaming as an Epistemic Governance Instrument in the Age of AI
arXiv · 2 July 2026
Geopolitical alignment: Endorsement effects in large language models
arXiv · 10 July 2026
How to cite this record
ethics.ai (6 July 2026), “Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment,” evidence record 207, https://ethics.ai/record/207 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.