Generalization Limits of Reinforcement Learning Alignment
The safety of large language models (LLMs) relies on alignment techniques such as reinforcement learning from human feedback (RLHF). However, recent theoretical analyses suggest that reinforcement learning-based training does not acquire new capabilities but merely redistributes the utilization probabilities of existing ones. In this study, we propose ``compound jailbreaks'' targeting OpenAI gpt-oss-20b, which exploit the generalization failures of alignment. This approach combines multiple atta
Record details
Published: 3 April 2026
Source: arXiv
Category: Research
Topics: Safety & alignment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Persistent Vulnerability of Aligned AI Systems
arXiv · 31 March 2026
Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment
arXiv · 14 April 2026
Working Memory Constraints Scaffold Learning in Transformers under Data Scarcity
arXiv · 22 April 2026
Self-Regulated Personal Contracts as a Harm Reduction Approach to Generative AI in Undergraduate Programming Education
arXiv · 11 March 2026
Fusion-fission forecasts when AI will shift to undesirable behavior
arXiv · 14 May 2026
Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute
arXiv · 14 May 2026
How to cite this record
ethics.ai (3 April 2026), “Generalization Limits of Reinforcement Learning Alignment,” evidence record 6418, https://ethics.ai/record/6418 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.