HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models
World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to calibration, perception, and contact-dynamics errors in real-world precision tasks, often failing in the final few millimeters of alignment or insertion. We propose HALO-WA, a hybrid-attention latent-guided online reinforcement learning (RL) framework for WA models, which leverages latent features and action priors from the WA generation process through a light
Record details
Published: 5 July 2026
Source: arXiv
Category: Research
Topics: Safety & alignment · Agents & autonomy
Retrieved: 14 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
Evaluating calibrated refusal and safe usefulness in dual-use biology settings
arXiv red teaming query · 6 July 2026
Governed Individuation: Cryptographically Decoupling an Agent's Learning from Its Authority
arXiv · 6 July 2026
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
arXiv · 6 July 2026
Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social Dilemmas
arXiv · 6 July 2026
Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents
arXiv red teaming query · 4 July 2026
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
arXiv · 3 July 2026
How to cite this record
ethics.ai (5 July 2026), “HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models,” evidence record 236, https://ethics.ai/record/236 (originally published by arXiv).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.