Predictive Divergence Masks for LLM RL
Reinforcement learning for large language models (LLMs) typically relies on trust-region masks to stabilize off-policy updates. The dominant PPO-style approach uses the sampled-token importance ratio for two criteria: a proximity criterion, which asks whether the policy has moved too far from the behavior policy, and a direction criterion, which asks whether the update pushes it farther away. Recent work DPPO improves the proximity criterion by replacing PPO's ratio-based test with a probability
Record details
Published: 11 July 2026
Source: HuggingFace Daily Papers
Category: Research
Topics: Regulation
Retrieved: 25 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Revising research practices for singing data collection
AI & Society · 12 July 2026
Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification
arXiv · 12 July 2026
When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control
arXiv · 11 July 2026
WasteAssistant: Regulation-Guided Visual Question Answering Framework for Intelligent Waste Segregation and Sustainable Managemen
arXiv · 12 July 2026
Sequential compliance decisions of firms on cross-border data flows: An institutionally anchored decision support system
arXiv · 12 July 2026
WattCouncil: Context-Aware Household Energy Scenario Generation With Governed LLMs
arXiv · 12 July 2026
How to cite this record
ethics.ai (11 July 2026), “Predictive Divergence Masks for LLM RL,” evidence record 13020, https://ethics.ai/record/13020 (originally published by HuggingFace Daily Papers).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.