Project · AIRoA Competition · 2025-09

AIRoA VLA Competition (Diffusion/Flow Policies)

Implemented and evaluated advanced methods like ReinFlow and DSRL to stably train VLA policies based on Diffusion and Flow Matching.

Overview In the AIRoA VLA Competition, I focused on training VLA models that use Diffusion or Flow Matching as their policy, which often suffer from policy collapse. Key Achievements Implemented and evaluated advanced methods like ReinFlow, which adapts DPPO for Flow Matching models by framing the denoising process as an MDP. Worked on reproducing DSRL, a method that applies an RL framework to generative models by treating noise inputs as actions. Gained a deep theoretical and practical understanding of how to stably train these cutting edge models.

#VLA #Diffusion Models #Flow Matching #Reinforcement Learning #ReinFlow #DSRL

Kaneyoshi Hiratsuka