Stage-Transition Dense Reward Modeling for Reinforcement Learning
Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense shaping signals is costly and brittle to changes in environments and object configurations. This work proposes Stage-Transition Dense Reward (STDR), a visual reward-learning framework that converts unstructured expert videos into logically grounded dense rewards for training RL agents from scratch. STDR leverages semantic understanding to infer a task's stag
Record details
Published: 30 June 2026
Source: arXiv
Category: Research
Topics: Agents & autonomy · Environment
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems
arXiv · 30 June 2026
Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling
arXiv · 30 June 2026
A new paradigm for marine ecological monitoring through swarm intelligence, digital twins, and Human–Swarm interaction
Frontiers in Robotics and AI · 30 June 2026
Projection surface detection and pose selection for autonomously displaying multimedia on walls using mobile robots
Frontiers in Robotics and AI · 29 June 2026
Recent advances in AI-based mobile robots for human companionship: survey
Artificial Intelligence Review · 29 June 2026
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior
arXiv · 26 June 2026
How to cite this record
ethics.ai (30 June 2026), “Stage-Transition Dense Reward Modeling for Reinforcement Learning,” evidence record 410, https://ethics.ai/record/410 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.