Project · Matsuo Lab · Internship · 2024-09

Video Generation World Models

Developed and optimized Transformer-Diffusion models for autonomous driving simulators at Matsuo Institute.

Internship Project at Matsuo Institute During my internship, I helped develop video generation world models (combining a Transformer Based autoregressive model and a Diffusion model) for autonomous driving. Key Contributions Solved "Color Shift": Identified and fixed unnatural colors in camera generation by implementing 'v prediction,' which stabilizes training and prevents color drift. Prevented VQ Collapse: Implemented the "rotation trick" as an alternative to STE, which redefines gradient flow during quantization to ensure more balanced codebook use and richer representations. Awarded "Best Poster" at an internal research presentation for this work. Currently involved in a new project to develop a VLA for autonomous driving based on the Simlingo architecture. Award 最優秀賞 (Best Poster Award) — 松尾・岩澤研究室主催 ポスターDay, The University of Tokyo, 25 September 2024, for 「自動運転のための世界モデル構築」 (Building World Models for Autonomous Driving).

#Diffusion Models #Transformers #World Models #Autonomous Driving #Generative AI

Kaneyoshi Hiratsuka