BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards
Critic-free reinforcement learning with verifiable rewards (RLVR), exemplified by Group Relative Policy Optimization (GRPO), avoids training a value function (critic) and reduces memory and compute overhead relative to critic-based PPO pipelines for aligning large language models. However, GRPO-style advantage estimation depends on prompt-local (within-prompt-group) reward statistics and can be unstable. In particular, when all rollouts in a prompt group receive identical rewards, the within-gro
Record details
Published: 27 June 2026
Source: arXiv
Category: Research
Topics: Regulation
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance
arXiv · 27 June 2026
The registrar's function in a hybrid society. AI value chain,smart data and the concept of property
arXiv · 27 June 2026
Defeat Devices in AI Systems
arXiv · 27 June 2026
Evidence-Based Text-Conditioned 3D CT Synthesis for Ovarian Cancer
arXiv · 27 June 2026
Fine-Tuning General-Purpose Large Language Models for Agricultural Applications:A Reproducible Framework and Evaluation Protocol Based on Qwen3-8B
arXiv · 27 June 2026
Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations
arXiv · 28 June 2026
How to cite this record
ethics.ai (27 June 2026), “BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards,” evidence record 505, https://ethics.ai/record/505 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.