Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback
Reinforcement Learning from Human Feedback (RLHF) is the standard approach for aligning large language models with human preferences, but its quality is limited by static, task-agnostic reward models. This mismatch leads to sparse learning signals and suboptimal alignment. We introduce MeRLa (Meta-Learned Reward Shaping), a principled framework that meta-learns a task-aware shaping function $Φ(x,y;φ)$ across auxiliary tasks before RLHF training. The learned shaping produces a composite reward th
Record details
Published: 28 July 2026
Source: arXiv cs.LG
Category: Research
Topics: Safety & alignment
Retrieved: 30 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Constructing Executable Analytical Knowledge Representations for Meta-Analysis Synthesis Using an Agentic Harness
arXiv · 3 August 2026
SMETA-ZSL:Semantic Meta-Alignment for Zero-Shot Threat Classification
arXiv · 10 July 2026
A systematic review of toxicity in large language models: definitions, datasets, detectors, detoxification methods and challenges
Artificial Intelligence Review · 2 July 2026
Meta-Transfer Learning for mmWave Beam Alignment
arXiv · 1 July 2026
AMR-SD: Asymmetric Meta-Reflective Self-Distillation for Token-Level Credit Assignment
arXiv · 18 May 2026
Meta-Aligner: Bidirectional Preference-Policy Optimization for Multi-Objective LLMs Alignment
arXiv · 27 April 2026
How to cite this record
ethics.ai (28 July 2026), “Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback,” evidence record 14858, https://ethics.ai/record/14858 (originally published by arXiv cs.LG).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.