Hello, I'm Che Liu

I am a final-year PhD student at Imperial College London, supervised by Prof. Rossella Arcucci and Prof. Wenjia Bai. I was a visiting student at Technical University of Munich (TUM), advised by Prof. Dr. Daniel Rückert.

I have also conducted research at X-Humanoid on Embodied Understanding and Generation, StepFun on Omni-modal learning, DAMO Academy on Visual Pre-training, and AstraZeneca Cambridge on VLM.

I am actively seeking a full-time research position based in the US, UK, or China, with a focus on multimodal learning, unified visual understanding and generation, and their applications.

Email: che.liu21 [at] imperial.ac.uk

Che Liu portrait

Selected Publications

† - Equal contribution, ‡ - Supervision. More on Google Scholar.

Pelican-VL 1.0: A Foundation Brain Model for Embodied Intelligence

C. Liu, X-Humanoid

Technical Report (X-humanoid), 2025

Step-Audio 2

StepFun Audio Team

Technical Report (StepFun), 2025

CogDoc: Towards Unified thinking in Documents

Q. Xu, H. Wang, C. Liu, F. Lin, W. Chen

Technical Report, 2025

Does DINOv3 Set a New Medical Vision Standard?

C. Liu, et al.

Technical Report, 2025

NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI

Cosmin I. Bercea, C. Liu, et al.

NeurIPS Dataset and Benchmark Track 2025 (Oral)

SuPreME: A Supervised Pre-training Framework for Multimodal ECG Representation Learning

M. Cai, J. Jiang, W. Huang, C. Liu‡, R. Arcucci

EMNLP Findings 2025

An Electrocardiogram Foundation Model Built on over 10 Million Recordings with External Evaluation across Multiple Domains

J. Li, A. Aguirre, J. Moura, C. Liu, L. Zhong, C. Sun, G. Clifford, B. Westover, S. Hong

NEJM AI 2024

IMITATE: Clinical Prior Guided Hierarchical Vision-Language Pre-training

C. Liu, et al.

IEEE Transactions on Medical Imaging 2024

Frozen Language Model Helps ECG Zero-Shot Learning

J. Li†, C. Liu†, et al.

MIDL 2023 (Oral)

Experiences

X-humanoid, Remote
Visiting Researcher (Embodied VLM) — Apr 2025 to Present
Leading VLM post-training (SFT + self-evolving RL) on large-scale video.
X-humanoid logo
StepFun, Remote
Visiting Researcher (OmniLLM) — May 2025 to Present
Audio-visual-language SFT/RL and omni-capability benchmarks.
StepFun logo
DAMO Academy, Beijing
Research Intern — Nov 2024 to Apr 2025
Unified multimodal vision pretraining across 2D/3D/video.
DAMO Academy logo
AstraZeneca, Cambridge, UK
Research Intern (Vision-Language Models) — Jul 2024 to Sep 2024
Synthetic-data VLP; showed synthetic pretraining can beat real-data baselines.
AstraZeneca logo

Education

Imperial College London, UK
Ph.D. in Multimodal Learning — Feb 2022 to 2026 (expected)
Supervisors: Rossella Arcucci, Wenjia Bai.
Imperial College London logo
Technical University of Munich, Germany
Visiting Ph.D. — Apr 2024 to Jun 2024
Supervisor: Daniel Rückert.
Technical University of Munich logo
Swansea University, UK
M.Sc. (Distinction) in Computational Mechanics — Sep 2019 to Sep 2021.
Supervisor: Dunhui Xiao.
Swansea University logo
Shanghai University of Engineering and Science, China
B.Sc. in Automotive Engineering — Sep 2016 to Jul 2018.
Shanghai University of Engineering and Science logo