Research · JSAI 2026 · 2026-06

JSAI 2026: Multimodal Imitation Learning Using Acoustic Spatial Information

音空間情報を用いたマニピュレータのためのマルチモーダル模倣学習 — sound source localization and separation over microphone arrays, fused with vision for manipulation policies.

Overview 音空間情報を用いたマニピュレータのためのマルチモーダル模倣学習 (Multimodal Imitation Learning for Robotic Manipulators Using Acoustic Spatial Information) Imitation learning has made complex manipulation learnable mostly from visual observations, but non visual cues — contact sounds, and the ambient sound produced while an object is being manipulated — remain underused. This work proposes a multimodal imitation learning method for robotic manipulators that exploits acoustic spatial information . Approach Apply sound source localization and sound source separation to the signals captured by several microphone arrays. Extract an acoustic representation that reflects the spatial distribution of the sound sources. Fuse it with visual observations so that the manipulator learns from sound and vision together. Experiments show the method improves both learning efficiency and success rate on manipulation tasks where the acoustic environment is a necessary cue. Related This is the domestic conference presentation of the line of work developed in S2A2 and started in RSJ 2025. Co authors: Ryosuke Kojima (Kyoto University, RIKEN), Benjamin Yen (RIKEN, Institute of Science Tokyo).

#Imitation Learning #Robot Audition #Multimodal #Manipulation #JSAI

Kaneyoshi Hiratsuka