Project · SusHi Tech Tokyo · 2026-04

SusHi Tech Tokyo 2026: Cloth-Folding VLA on a Self-Built Bimanual Robot

Exhibited a handkerchief-folding demo at SusHi Tech Tokyo 2026 with KUPAC, running a fine-tuned X-VLA policy on ILOHA, our self-built bimanual robot.

Overview With KUPAC, the student led Physical AI community at Kyoto University, we exhibited a cloth folding demonstration at SusHi Tech Tokyo 2026, running a Vision Language Action policy on ILOHA — the open source bimanual manipulator we build ourselves. Summary Movie What we did Collected roughly 300 teleoperated episodes of handkerchief folding. Compared candidate VLA models and settled on X VLA , fine tuned on a dataset in which high and low quality demonstrations carry an explicit quality label. On site, the lighting, background, and camera angles differed enough from our lab that the policy degraded. We collected 50 extra episodes at the venue and re fine tuned for 3.5 hours, taking the success rate from roughly 60% to over 90%. The Robot in Action Lesson learned The bottleneck was not the model. Cable reliability and USB bandwidth were what actually limited the demo — a reminder that in Physical AI the hardware plumbing decides how much of your policy the audience ever gets to see. Write up 【SusHi Tech Tokyo 2026】自作双腕ロボット×VLAで布畳みロボを展示した話 (Qiita)

#VLA #Imitation Learning #Robotics #ILOHA #Exhibition #KUPAC

Kaneyoshi Hiratsuka