Skip to main content

Primary supervisor

Jianfei Cai

Deep learning has achieved groundbreaking performance in many 2D vision tasks in recent years. With more and more 3D data available, such as that captured by LiDAR, the next research trend is to conduct advanced 3D perception and generation tasks. The objective of this project is to study the state-of-the-art 3D foundation models and apply them for tasks that require spatial intelligence.

This is a "research project" best suited for students who are independent and willing to take on challenges with high expectations for the grade when fulfilled the somewhat challenging requirements. Under-performing is likely to fail to meet the passing requirements. It is also a good practice for students who wish to pursue further study at a postgraduate/PhD level.

Aim/outline

The objective of this project is to study state-of-the-art 3D foundation models such as VGGT. They can be applied to 3D reconstruction and 3D understanding tasks, such as robotic applications.

URLs/references

- MVSplat: https://github.com/donydchen/mvsplat

- VGGT: https://vgg-t.github.io/

 

Required knowledge

The student must have knowledge of deep learning (e.g., taking online Stanford deep learning, computer vision-related courses) and be skillful in Python programming and vibe coding.