Google
Incoming Student Researcher
Will work with Dr. Yale Song on video editing improvements for Veo 3.1.
Google
Incoming Student Researcher
Tongyi MAI
Research Intern
Biography
Hi, I'm Xiangpeng. I am currently a Ph.D. student at the University of Technology Sydney (UTS). I am also a research intern with the Tongyi MAI Z-Image Team, advised by Dr. Peng Gao and Prof. Steven Hoi. I am an incoming Student Researcher at Google, working with Dr. Yale Song on video editing improvements for Veo 3.1.
My research interests involve Generative AI, Video Generation, and Multi-modal Learning. Specifically, I focus on video world models, video generation, and multi-modal foundation models.
Looking ahead, I am deeply motivated to build unified video models capable of jointly understanding dynamic visual environments and generating coherent future content within a single framework. I believe this direction is a crucial step toward world models, where systems can reason about and interact with the physical world through continuous video understanding and prediction.
I am seeking full-time opportunities starting in Fall 2026 in image/video generation, world models, embodied AI, and unified models. Please feel free to reach out!
News
Selected Publications



Industry Experience
Google
Will work with Dr. Yale Song on video editing improvements for Veo 3.1.
Working with Dr. Peng Gao and Prof. Steven C. H. Hoi on generative foundation models in the Z-Image Team, including Z-Multi-Shot Video.
Advised by Dr. Yifan Sun on multimodal learning and video-text retrieval. This internship led to DGL, published at AAAI 2024.
Invited Talks
Academic Service
Peer reviewer
Visitor Map