I am a master's student in Information and Communication Engineering at the
University of Science and Technology of China,
advised by Prof. Zhiwei Xiong.
Before joining USTC, I received a bachelor's degree in Data Science and Big Data Technology from
Harbin Institute of Technology, Shenzhen.
My research interests include controllable video generation, video world models, and embodied AI, with earlier work in medical AI.
I study whether video generation models can preserve coherent task-relevant event states as a camera moves, looks away,
and later returns. I also investigate how visual representations in vision-language-action models shape decisions and actions,
and how medical AI systems can remain efficient, reliable, and clinically reviewable.
My recent work includes WRBench,
a benchmark for testing whether generated events remain coherent across camera changes. I also co-authored
VLA-Trace, which traces how
vision-language-action models turn representations into behavior, and
Pelican-Unify 1.0, which jointly
generates future video and actions.
Related work has appeared at ICLR, EMNLP, and MICCAI, and in Medical Image Analysis.
📢 I am looking for a full-time video generation research internship (6+ months, available immediately; Beijing, Shanghai, or Hangzhou) and welcome research collaborations. CV: English / 中文. Please feel free to email me.