Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by the slow, manual process of dense reward shaping. To bypass this limitation, we leverage Sample-based Model Predictive Control (SMPC) entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets. Because this data solves the fundamental exploration problem, we can train an off-policy RL agent
Record details
Published: 12 August 2026
Source: arXiv
Category: Research
Topics: Regulation · Agents & autonomy
Retrieved: 13 August 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
arXiv cs.AI · 12 August 2026
Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents
arXiv · 12 August 2026
Most biomedical publications show signs of LLM-assisted writing
arXiv cs.CY · 12 August 2026
Governing Agentic AI in FinTech
arXiv cs.CY · 13 August 2026
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
arXiv · 13 August 2026
AI Guardrail Survival under Single-Cycle Agentic Self-Summarization
arXiv · 11 August 2026
How to cite this record
ethics.ai (12 August 2026), “Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL,” evidence record 18778, https://ethics.ai/record/18778 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.