MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization
To operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interaction setting. We ask whether this execution-level variation can be learned as a reusable behavioral factor shared across tasks. We present MoMo, a two-stage imitation-learning framework consisting of a spatiotemporal action tokenizer and a behavior-cloning transformer that takes task and a continuous motion-mode conditi
Record details
Published: 30 July 2026
Source: Apple Machine Learning Research
Category: Field notes
Topics: Agents & autonomy
Retrieved: 31 July 2026
Related evidence
These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.
How avatarin built a 24/7 retail agent with GPT-Realtime
OpenAI · 30 July 2026
AI and the jobs apocalypse that wasn’t
R Street Institute · 30 July 2026
Measuring the Tendency of AI Agents to Go Rogue
Bruce Schneier — Schneier on Security · 29 July 2026
A new benchmark for evaluating patient-facing health AI agents
Amazon Science · 29 July 2026
What’s new in Microsoft Security: July 2026
Microsoft Responsible AI · 30 July 2026
Poolside’s Laguna S 2.1: the model factory delivers
Air Street Capital (State of AI) · 29 July 2026
How to cite this record
ethics.ai (30 July 2026), “MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization,” evidence record 15182, https://ethics.ai/record/15182 (originally published by Apple Machine Learning Research).
Use and limitations
This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.