👨‍🎓 About Me

I am a master's student in Information and Communication Engineering at the University of Science and Technology of China, advised by Prof. Zhiwei Xiong. Before joining USTC, I received a bachelor's degree in Data Science and Big Data Technology from Harbin Institute of Technology, Shenzhen.

My research interests include controllable video generation, video world models, and embodied AI, with earlier work in medical AI. I study whether video generation models can preserve coherent task-relevant event states as a camera moves, looks away, and later returns. I also investigate how visual representations in vision-language-action models shape decisions and actions, and how medical AI systems can remain efficient, reliable, and clinically reviewable.

My recent work includes WRBench, a benchmark for testing whether generated events remain coherent across camera changes. I also co-authored VLA-Trace, which traces how vision-language-action models turn representations into behavior, and Pelican-Unify 1.0, which jointly generates future video and actions.

Related work has appeared at ICLR, EMNLP, and MICCAI, and in Medical Image Analysis.

📢 I am looking for a full-time video generation research internship (6+ months, available immediately; Beijing, Shanghai, or Hangzhou) and welcome research collaborations. CV: English / 中文. Please feel free to email me.

  • Video World Models I design evaluations that test whether generated events remain coherent as the camera moves and returns.
  • Embodied AI I study how visual representations in vision-language-action models drive reasoning, action, and future prediction.
  • Medical AI I have worked on efficient 3D medical imaging and pathology, with an emphasis on reliable model analysis.

🔥 News

💼 Experience

2025.11 - 2026.09

Beijing Innovation Center of Humanoid Robotics (X-Humanoid)

Algorithm Intern, Embodied World Model Team, Beijing. Led camera-controllable video generation, action-conditioned world modeling, and few-step action prediction; proposed WRBench as first author; contributed to VLA-Trace (EMNLP 2026) and Pelican-Unify 1.0.

📘 Selected Publications

Full list: Google Scholar Google Scholar citations: 161 h-index 6 · i10-index 3

🌍 Video World Models & Generative Video Evaluation

arXiv 2026 WRBench source figure showing viewpoint changes and state-consistency evaluation
First Author 2026.06

Current World Models Lack a Persistent State Core

J. Lu, D. Zhu, H. Shi, L. Cai, G. Tang, Y. Chen, J. Cao, D. Tang, Y. Zhang, Y. Dai, X. Ju.

  • Question: When the camera changes only what is observed, does one task-relevant event state continue to evolve coherently?
  • Approach: WRBench separates camera execution, visual integrity, visible evolution, re-observation support, and conditional endpoint consistency across four native camera-control interfaces.
  • Finding: Across 9,600 videos from 23 models, better view synthesis and re-observation support did not reliably yield correct event endpoints.

🤖 Embodied AI & VLA Analysis

1 / 2
EMNLP 2026 VLA-Trace source figure showing representation and behavior tracing framework
Co-author CCF B 2026.08

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

  • Question: How do internal VLA representations translate into decisions and rollout behavior?
  • Approach: VLA-Trace combines representation analysis, causal attention interventions, and behavioral probes across π0.5 and OpenVLA.
  • Finding: The study identifies distinct adaptation and action-routing patterns, while revealing limits in fine-grained semantic following.
arXiv 2026 Pelican-Unify source figure showing embodied intelligence model components
Co-author 2026.05

Pelican-Unify 1.0: A Unified Embodied Intelligence Model

  • Question: Can one embodied model connect visual understanding, reasoning, future prediction, and action generation?
  • Approach: Pelican-Unify uses a shared vision-language model, then jointly generates future video and action.
  • Training: Language, video, and action objectives are optimized together instead of as isolated specialist capabilities.

🏥 Medical AI & Efficient 3D Perception

1 / 3
ICLR 2026 VeloxSeg overview source figure
First Author CCF A 2026.01

Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation

J. Lu, L. Cai, Y. Chen, G. Tang, S. Jiang, H. Shi, Z. Xiong.

  • Task: Perform efficient multimodal 3D medical segmentation for volumetric PET/CT perception.
  • Method: Combine JL-guided convolution, paired window attention, and spatially decoupled knowledge transfer in VeloxSeg.
  • Result: Delivers a 1.66M-parameter segmentation model that keeps the architecture lightweight for 3D medical-image analysis.
Tech Report 2025 DINOv3 PET/CT feature source figure
Co-first Author 2025.10

Does DINOv3 Set a New Medical Vision Standard?

C. Liu*, Y. Chen*, H. Shi*, J. Lu*, B. Jian*, J. Pan*, L. Cai*, et al.

  • Task: Test whether DINOv3-style self-supervised visual features transfer reliably to medical vision tasks.
  • Method: Benchmark 2D and 3D classification, segmentation, and registration across multiple medical modalities.
  • Result: Identifies where transfer is useful, where task limits remain, and which modality gaps matter for medical foundation-model use.
MICCAI 2024 H2ASeg source figure showing PET/CT hierarchical adaptive interaction architecture
First Author CCF B 2024.03

H2ASeg: Hierarchical Adaptive Interaction and Weighting Network for Tumor Segmentation in PET/CT Images

J. Lu, J. Chen, L. Cai, S. Jiang, Y. Zhang.

  • Task: Segment tumor regions in paired PET/CT images under heterogeneous lesion appearance and modality imbalance.
  • Method: Use hierarchical adaptive interaction and weighting to fuse PET metabolic cues with CT anatomical structure.
  • Result: Improves robust lesion-region segmentation and establishes the medical-segmentation line that later led to VeloxSeg.

Additional Medical AI Publications

🧰 Open-source Work

I build practical tools for agent and research workflows.

Three projects per page.

1 / 2

🎓 Education

2025.09 - 2028.06

University of Science and Technology of China

M.S. in Information and Communication Engineering. Advisor: Prof. Zhiwei Xiong.

📝 Academic Service

Reviewer for ICLR 2027, NeurIPS 2026, and IEEE TNNLS.

👣 Site Visitors

Approximate visitor map and page-view counter.

Approximate Location

Approximate location: loading...

Location is estimated from public IP geolocation and may be imprecise; this page does not display your raw IP address.

Visitor map and counter powered by Flag Counter