Evidence record 15182 · automatically gathered

MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization

To operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interaction setting. We ask whether this execution-level variation can be learned as a reusable behavioral factor shared across tasks. We present MoMo, a two-stage imitation-learning framework consisting of a spatiotemporal action tokenizer and a behavior-cloning transformer that takes task and a continuous motion-mode conditi

Record details

Published: 30 July 2026
Source: Apple Machine Learning Research
Category: Field notes
Topics: Agents & autonomy
Retrieved: 31 July 2026

source-onlyevidence status

These records share source-supplied organisations, an exact publisher byline, automatic topics or regions. The reason is shown on every link; related does not mean supporting, agreeing with or verifying this record.

How to cite this record

ethics.ai (30 July 2026), “MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization,” evidence record 15182, https://ethics.ai/record/15182 (originally published by Apple Machine Learning Research).

JSON

Use and limitations

This page is a stable index and citation surface for a source record. ethics.ai did not author the underlying report and has not independently verified every claim. Automatic topics may be imperfect. For consequential use, quote and cite the original publisher.