Primary supervisor
Adamu Muhammad BuhariResearch area
Human Centred AIWorld models are becoming an important direction in artificial intelligence and robotics. Instead of responding only to what is currently observed, an intelligent agent can use an internal model of the world to represent its surroundings, anticipate what may happen next, and reason about the consequences of possible actions. This ability is particularly important for robots operating around people, where behaviour is dynamic, uncertain, and strongly influenced by context.
Most existing world models concentrate on physical environments, object dynamics, navigation, or robotic manipulation. Human-centred environments introduce a different challenge. A robot interacting with people needs to understand more than objects and geometry. It may need to recognise where a person is looking, what they are doing, how they are responding to the robot, whether they intend to interact, and how these behaviours are likely to change over time.
This PhD project will explore human-centred multimodal world models for socially intelligent embodied agents. The central idea is to build predictive models that jointly represent people, the surrounding environment, and the actions of the embodied agent.
Information from video, facial behaviour, body pose, gaze, speech, language, environmental context, and robot state may be combined to form a continuously evolving representation of an interaction. Rather than simply classifying the current human state, the model should be able to anticipate possible future states and use these predictions to support decision-making.
For example, an embodied agent approaching a person could reason about whether that person has noticed the robot, whether they are likely to engage, how their behaviour may change if the robot approaches, and which action would be appropriate next. This moves beyond conventional perception pipelines towards AI systems that can observe, remember, anticipate, and act.
The project will examine questions around multimodal representation learning, temporal modelling, predictive learning, uncertainty, memory, action-conditioned prediction, and human-aware planning. Depending on the direction taken, generative models, multimodal foundation models, self-supervised learning, or latent world models may form part of the methodological framework.
Experiments will initially use established datasets and simulation environments to enable controlled comparison with existing approaches. Selected methods can then be evaluated through real-world human-robot interaction using humanoid and mobile robotic platforms.
Aim/outline
The aim of this project is to develop multimodal world models that allow embodied agents to understand human behaviour, anticipate how interactions may evolve, and use these predictions to make better decisions in dynamic environments.
The project may explore the following areas:
- Multimodal Representation Learning
- World Models and Predictive Learning
- Human Behaviour Understanding
- Temporal and Long-Horizon Prediction
- Human Action and Intention Anticipation
- Vision-Language and Multimodal Foundation Models
- Uncertainty-Aware Prediction
- Memory and Temporal Context
- Action-Conditioned Future Prediction
- Human-Aware Planning
- Socially Intelligent Human-Robot Interaction
The exact research questions will be refined with the PhD candidate following the initial literature review and exploratory experiments.
Research deliverables
- Critical review of multimodal world models and human-centred embodied AI
- New methods for modelling human behaviour and interaction dynamics
- Multimodal temporal and predictive learning algorithms
- Experimental evaluation against strong contemporary baselines
- A multimodal HRI dataset and/or benchmark, where appropriate
- Evaluation of prediction, robustness and generalisation across different environments
- Validation through simulation and/or physical robotic platforms
- Reproducible implementations and experimental results
- Research papers arising from the main contributions
- PhD thesis
Publication expectations
This is a research-oriented PhD project aimed at developing methodological contributions at the intersection of computer vision, multimodal learning, predictive world models, embodied AI, and human-robot interaction.
The candidate will work towards models that go beyond recognising the current state of an interaction and instead learn representations that support prediction, anticipation, and decision-making. Strong experimental validation will be important, including comparisons with contemporary approaches, ablation studies, cross-dataset evaluation, and testing under previously unseen conditions.
Where appropriate, the project may also develop a multimodal human-robot interaction dataset or benchmark to support reproducibility and comparison with future work.
The work is expected to lead to publications in established computer vision, machine learning, artificial intelligence, and robotics venues. Depending on the contribution and maturity of the work, relevant conferences include CVPR, NeurIPS, etc., while extended studies may be suitable for journals such as IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), Artificial Intelligence, IEEE Transactions on Robotics, and Nature Machine Intelligence.
Required knowledge
Applicants should have:
- A good background in machine learning, deep learning, artificial intelligence, or computer vision
- Strong Python programming skills
- Experience with at least one deep-learning framework, preferably PyTorch
- An interest in multimodal learning, generative AI, human behaviour analysis, robotics, or embodied intelligence
- Good analytical, mathematical, and problem-solving skills
Experience with Transformers, vision-language models, generative models, self-supervised learning, reinforcement learning, OpenCV, ROS/ROS2, or Linux would be useful but is not essential.
Knowledge of facial behaviour analysis, human pose estimation, gaze estimation, speech/audio processing, temporal modelling, or human-robot interaction would also be beneficial.
Candidates with a strong machine-learning or computer-vision background who would like to move into multimodal embodied AI are encouraged to apply.