Primary supervisor
Jianfei CaiDeep learning has achieved ground-breaking performance in many vision tasks in the recent years. The objective of this project is to apply state-of-the-art visual-language foundation models such as Qwen for medical image analysis and report generation.
This is a "research project" best for students who are independent and willing to take up challenges with high expectation in the grade when fulfilled the somewhat challenging requirements. Under-performing is likely to fail to meet the passing requirements. It is also a good practice for students who wish to pursue further study at a postgraduate/PhD level.
Aim/outline
The objective of this project is to apply state-of-the-art VLM-based foundation models such as Qwen for understanding and analyzing medical images. It is for the purpose of assisting doctors in analyzing medical images, detecting lesions, and predicting future situations, etc.
URLs/references
Y. Wu, T. Song, Z. Wu, J. Ye, Z. Ge, W. Bai, Z. Chen, and J. Cai, “Virtual full-stack scanning of brain MRI via imputing any quantised code”, CVPR 2026.
P. Chen, J. Ye, et al., “GMAI-MMBench: A comprehensive multimodal evaluation benchmark towards general medical AI”, NeurIPS 2024.
Required knowledge
The student must have knowledge of deep learning (e.g., taking online Stanford deep learning, computer vision-related courses) and be skilled in Python programming and vibe coding.