👨‍🎓 About Me

I am a master's student in Information and Communication Engineering at the University of Science and Technology of China, advised by Prof. Zhiwei Xiong. Before joining USTC, I received a bachelor's degree in Data Science and Big Data Technology from Harbin Institute of Technology, Shenzhen.

My research interests include video world models, embodied AI, and medical AI. I study whether video generation models can preserve coherent task-relevant event states as a camera moves, looks away, and later returns. I also investigate how visual representations in vision-language-action models shape decisions and actions, and how medical AI systems can remain efficient, reliable, and clinically reviewable.

My recent work includes WRBench, a benchmark for testing whether generated events remain coherent across camera changes. I also co-authored VLA-Trace, which traces how vision-language-action models turn representations into behavior, and Pelican-Unify 1.0, which jointly generates future video and actions.

Related work has appeared at ICLR and MICCAI, and in Medical Image Analysis.

📢 I am actively seeking research collaborations and research internship opportunities. Please feel free to email me.

  • Video World Models I design evaluations that test whether generated events remain coherent as the camera moves and returns.
  • Embodied AI I study how visual representations in vision-language-action models drive reasoning, action, and future prediction.
  • Medical AI I have worked on efficient 3D medical imaging and pathology, with an emphasis on reliable model analysis.

🔥 News

📘 Selected Publications

Full list: Google Scholar Google Scholar citations: 120 h-index 5 · i10-index 3

🌍 Video World Models & Generative Video Evaluation

arXiv 2026 WRBench source figure showing viewpoint changes and state-consistency evaluation
First Author 2026.06

Current World Models Lack a Persistent State Core

J. Lu, D. Zhu, H. Shi, L. Cai, G. Tang, Y. Chen, J. Cao, D. Tang, Y. Zhang, Y. Dai, X. Ju.

  • Question: When the camera changes only what is observed, does one task-relevant event state continue to evolve coherently?
  • Approach: WRBench separates camera execution, visual integrity, visible evolution, re-observation support, and conditional endpoint consistency across four native camera-control interfaces.
  • Finding: Across 9,600 videos from 23 models, better view synthesis and re-observation support did not reliably yield correct event endpoints.

🤖 Embodied AI & VLA Analysis

1 / 2
arXiv 2026 VLA-Trace source figure showing representation and behavior tracing framework
Co-author 2026.05

VLA-Trace: Diagnosing Vision-Language-Action Models through Representation and Behavior Tracing

  • Question: How do internal VLA representations translate into decisions and rollout behavior?
  • Approach: VLA-Trace combines representation analysis, causal attention interventions, and behavioral probes across π0.5 and OpenVLA.
  • Finding: The study identifies distinct adaptation and action-routing patterns, while revealing limits in fine-grained semantic following.
arXiv 2026 Pelican-Unify source figure showing embodied intelligence model components
Co-author 2026.05

Pelican-Unify 1.0: A Unified Embodied Intelligence Model

  • Question: Can one embodied model connect visual understanding, reasoning, future prediction, and action generation?
  • Approach: Pelican-Unify uses a shared vision-language model, then jointly generates future video and action.
  • Training: Language, video, and action objectives are optimized together instead of as isolated specialist capabilities.

🏥 Medical AI & Efficient 3D Perception

1 / 3
ICLR 2026 VeloxSeg overview source figure
First Author CCF A 2026.01

Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation

J. Lu, L. Cai, Y. Chen, G. Tang, S. Jiang, H. Shi, Z. Xiong.

  • Task: Perform efficient multimodal 3D medical segmentation for volumetric PET/CT perception.
  • Method: Combine JL-guided convolution, paired window attention, and spatially decoupled knowledge transfer in VeloxSeg.
  • Result: Delivers a 1.66M-parameter segmentation model that keeps the architecture lightweight for 3D medical-image analysis.
Tech Report 2025 DINOv3 PET/CT feature source figure
Co-first Author 2025.10

Does DINOv3 Set a New Medical Vision Standard?

C. Liu*, Y. Chen*, H. Shi*, J. Lu*, B. Jian*, J. Pan*, L. Cai*, et al.

  • Task: Test whether DINOv3-style self-supervised visual features transfer reliably to medical vision tasks.
  • Method: Benchmark 2D and 3D classification, segmentation, and registration across multiple medical modalities.
  • Result: Identifies where transfer is useful, where task limits remain, and which modality gaps matter for medical foundation-model use.
MICCAI 2024 H2ASeg source figure showing PET/CT hierarchical adaptive interaction architecture
First Author CCF B 2024.03

H2ASeg: Hierarchical Adaptive Interaction and Weighting Network for Tumor Segmentation in PET/CT Images

J. Lu, J. Chen, L. Cai, S. Jiang, Y. Zhang.

  • Task: Segment tumor regions in paired PET/CT images under heterogeneous lesion appearance and modality imbalance.
  • Method: Use hierarchical adaptive interaction and weighting to fuse PET metabolic cues with CT anatomical structure.
  • Result: Improves robust lesion-region segmentation and establishes the medical-segmentation line that later led to VeloxSeg.

Additional Medical AI Publications

🧰 Open-source Work

I build practical tools for agent and research workflows.

Three projects per page.

1 / 1

🎓 Education

2025.09 - 2028.06

University of Science and Technology of China

M.S. in Information and Communication Engineering. Advisor: Prof. Zhiwei Xiong.

👣 Site Visitors

Approximate visitor map and page-view counter.

Approximate Location

Approximate location: loading...

Location is estimated from public IP geolocation and may be imprecise; this page does not display your raw IP address.

Visitor map and counter powered by Flag Counter