Research · 2026-07

S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information

A multimodal imitation learning framework that fuses acoustic spatial maps and spotformed spectrograms with vision, plus a new suite of acoustic-aware manipulation tasks.

Abstract Acoustic information provides rich cues about object location, material properties, and changes caused by contact or motion. This work introduces a new set of acoustic aware manipulation tasks for imitation learning, in which a robot must use auditory cues to determine what to manipulate and where to put it. These tasks require sound source localization and identification for active exploration in robotic manipulation. We propose Spatial Spectral Audio Action (S2A2) , a multimodal imitation learning framework that integrates visual features with acoustic spatial information (an acoustic spatial map computed with MUSIC over multiple microphone arrays) and acoustic signal information (a spectrogram extracted by Spotforming). S2A2 is implemented on top of four policies — ACT, Diffusion Policy, VQ BeT, and . Simulation experiments show that the full framework is the most effective for tasks that require both position and timbre, and real robot experiments confirm that both the tasks and the framework transfer to real world manipulation. Tasks Localization — two visually identical objects; only one emits sound. Spatial information alone determines the target. Identification — one object emitting one of two sounds; the timbre determines the destination box. Localization & Identification (L&I) — both cues are required at once. Exploratory — neither object sounds at rest; the robot must shake each one to find the one that rattles. Results On the L&I task the full S2A2 model is the only configuration that performs well across policies (e.g. 89.7% with Diffusion Policy vs.…

#Imitation Learning #Robot Audition #Multimodal #Manipulation #Sound Source Localization #Diffusion Policy

Kaneyoshi Hiratsuka