What if non-expert operators could adapt an industrial robot’s skills as easily as touching, talking to, or clicking on it? Our framework MOMO (Motion Modulation) unifies three complementary interaction modalities for seamless robot skill learning and adaptation.
MOMO integrates five components around a central motion modulation module:
- Kinesthetic touch — physically guiding the robot’s arm — for precise spatial corrections, using energy-tank-based human intention detection that automatically inserts via-points (intermediate points the motion must pass through) into the underlying motion model.
- Natural language for high-level semantic modifications through a tool-based LLM architecture that selects and parameterizes pre-validated functions — never generating executable code.
- A graphical web interface (Vue.js/Three.js) for visualizing geometric relations, inspecting parameters, and editing via-points via drag-and-drop on a real-time digital twin.
- Probabilistic Virtual Fixtures for guided demonstration recording.
- Ergodic control — a control method that spreads the robot’s motion evenly over a surface — for finishing tasks like polishing.
A key result: the tool-based LLM architecture generalizes beyond skills based on Kernelized Movement Primitives (KMPs) to ergodic control, enabling the same chat interface to drive voice-commanded surface finishing. Users freely switch between modalities — voice for obstacle avoidance, kinesthetic for fine corrections, graphical for verification.
Validated on a 7-DoF torque-controlled DLR robot in two industrial use cases — bearing ring measurement and surface finishing — and showcased live at the Automatica 2025 trade fair and at DLR, where visitors operated the robot themselves. The system runs a local LLM backend (Qwen2.5-VL-72B-Instruct) for data privacy and low latency.
Building MOMO meant merging four independent PhD research threads (13 authors from DLR and the Technical University of Munich), each with its own assumptions, into one modular framework — and then integrating it into the existing Human Factory Interface of the DLR SARA robot. I led that integration effort. The paper will be presented at ICRA 2027 in Seoul.
Code
MOMO’s components are open source. The verbal modality builds on the IROSA code, the kinesthetic via-point modulation on the LOCI interactive incremental learning code, and the Probabilistic Virtual Fixtures on PGLfD (Probabilistic and Geometric Learning from Demonstration). Links to all three repositories and the runnable Code Ocean capsules are listed below.