Primary supervisor
Jesmin NaharThis project uses machine learning and predictive analytics to support the early detection and risk prediction of breast cancer using publicly available or clinical healthcare datasets. Students will clean and analyse clinical and diagnostic health data, apply classification and risk modeling algorithms such as Logistic Regression, Random Forest, Support Vector Machines, or deep learning, and identify key clinical risk factors and diagnostic markers. The project aims to show how data-driven methods can help healthcare practitioners enable timely clinical intervention, improve screening accuracy, and optimize patient outcomes.
Title: AI-Driven Early Detection and Risk Prediction of Breast Cancer Using Machine Learning and Statistical Modelling for Timely Intervention.
Description: This project focuses on analysing clinical registries, diagnostic tabular parameters, and radiological data to identify distinct risk profiles and early malignancy patterns associated with breast cancer. Students will work with public or clinical health datasets (such as structured diagnostic features, electronic health records, or imaging data) to clean and prepare the data, then apply machine learning models to discover critical risk indicators and predict malignancy risk. Depending on the project scope, Honours students will focus on feature selection, classical machine learning classification, and decision-curve analysis using structured tabular data, while Master's students will implement multimodal deep learning frameworks (e.g., multi-view mammography combined with clinical markers) for multi-year incidence and risk prediction. The results will help clinicians and researchers develop data-driven screening tools and personalized intervention strategies based on objective patient risk factors.
Aim/outline
Expected Outcomes:
-
Well-defined breast cancer risk profiles with transparent clinical and diagnostic indicators.
-
An AI-driven prediction pipeline and evaluation report offering data-driven insights for early clinical decision support, risk stratification, and timely intervention.
Skills Required: Proficiency in Python or R; basic knowledge of machine learning and statistical modeling algorithms (e.g., Logistic Regression, Random Forest, Support Vector Machines); familiarity with data preprocessing and clinical evaluation metrics (e.g., ROC-AUC, Decision Curve Analysis).
Benefits: This project helps healthcare practitioners improve early detection accuracy and screening efficiency by identifying critical diagnostic risk factors, leading to earlier clinical intervention and improved patient survival outcomes.
Required knowledge
General Project Pre-requisite Skills and/or Knowledge:
-
Basic understanding of machine learning concepts, particularly classification algorithms (e.g., Logistic Regression, Support Vector Machines, Random Forest) and predictive modeling.
-
Familiarity with Python or R for data handling, statistical analysis, and model development.
-
Basic knowledge of data preprocessing (cleaning, feature selection, and preparing tabular or clinical datasets).
-
Interest in clinical evaluation metrics and data visualisation (e.g., ROC-AUC curves, Decision Curve Analysis) to evaluate screening utility and model performance.