Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models
Consider an auto-regressive model that produces outputs x (e.g., answers to questions, molecules) each of which can be summarized by an attribute vector y (e.g., helpfulness vs. harmlessness, or bio-availability vs. lipophilicity). An arbitrary reward function r(y) encodes tradeoffs between these properties. Typically, tilting the model's sampling distribution to increase this reward is done at training time via reinforcement learning. However, if the reward function changes, re-alignment requir
Record details
Published: 16 April 2026
Source: arXiv
Category: Research
Topics: Regulation · Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Decentralized autonomous organization and blockchain-based incentivization framework for community-based facilities management
arXiv · 16 April 2026
Layered Mutability: Continuity and Governance in Persistent Self-Modifying Agents
arXiv · 16 April 2026
The Granularity Gap: A Multi-Dimensional Longitudinal Audit of Sycophancy in Gemini Models
arXiv · 19 April 2026
Demystifying the unreasonable effectiveness of online alignment methods
arXiv · 19 April 2026
AISafetyBenchExplorer: A Metric-Aware Catalogue of AI Safety Benchmarks Reveals Fragmented Measurement and Weak Benchmark Governance
arXiv · 14 April 2026
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
arXiv · 20 April 2026
How to cite this record
ethics.ai (16 April 2026), “Reward Weighted Classifier-Free Guidance as Policy Improvement in Autoregressive Models,” evidence record 5753, https://ethics.ai/record/5753 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.