Skip to main content

Trustworthy and Resource-Adaptive Vision-Language-Action Models for Embodied AI

Primary supervisor

Adamu Muhammad Buhari

Research area

Machine Learning

Vision-Language-Action (VLA) models are changing the way robots perceive their surroundings, interpret instructions, reason about tasks, and interact with the physical world. By bringing vision, language, and action into a common framework, these models offer a promising route towards robots that can operate more naturally in complex and unfamiliar environments.

Despite this progress, several important problems remain. Modern VLA models can be computationally expensive, their confidence does not always reflect whether a decision is correct, and performance can deteriorate when a robot encounters conditions that differ from its training data. A robot operating in the real world therefore needs more than an accurate model: it needs to recognise when information is unreliable and decide how much computation is necessary before taking an action.

This PhD project will address these challenges by developing trustworthy and resource-adaptive VLA models. A central question is whether an embodied agent can estimate the reliability of its own perception and reasoning, then use that information to adjust how it processes a task.

For example, a straightforward scene may be handled by a lightweight model, while an ambiguous instruction, unfamiliar object, conflicting sensory information, or uncertain prediction could trigger additional reasoning or a more capable model. Possible directions include uncertainty estimation, confidence calibration, model routing, token selection, early exiting, adaptive visual processing, and efficient use of multiple foundation models.

The work will combine algorithm development with experimental validation. Established datasets and simulation environments can be used for systematic benchmarking, followed by deployment on humanoid or mobile robots using NVIDIA Jetson edge computing platforms. This provides an opportunity to study not only accuracy, but also reliability, latency, memory use, computational cost, and behaviour under realistic operating conditions.

The broader goal is to move towards embodied AI systems that can make better decisions about what they know, when they are uncertain, and how much computation they need before acting.

Aim/outline

The aim of this project is to develop VLA models that can operate reliably and efficiently under uncertainty and limited computational resources.

The project may explore the following areas:

  • Vision-Language-Action models
  • Multimodal foundation models
  • Uncertainty estimation and calibration
  • Failure and out-of-distribution detection
  • Adaptive and selective computation
  • Model routing and dynamic inference
  • Generalisation to unseen environments
  • Embodied reasoning and decision-making
  • Edge AI for real-time robotic systems

The exact direction will be refined with the PhD candidate based on their background, interests, and findings from the initial literature review.

Research deliverables

  • Critical review of the state of the art in VLA and embodied AI
  • New methods for reliable and computationally adaptive VLA inference
  • Experimental evaluation against strong contemporary baselines
  • Benchmarking under uncertainty, distribution shifts, and resource constraints
  • Validation using simulation and/or physical robotic platforms
  • Reproducible implementations and experimental results
  • Research papers arising from the major contributions
  • PhD thesis

Publication expectations

This is a research-oriented PhD project, with an emphasis on developing methods that contribute beyond the integration of existing robotic and AI components.

The candidate will be encouraged to identify clear research gaps, develop technically sound solutions, and evaluate them against strong baselines using recognised benchmarks. Where appropriate, experiments will also examine robustness, calibration, generalisation, computational efficiency, and real-world performance.

The work is expected to lead to publications in established machine learning, computer vision, artificial intelligence, and robotics venues. Depending on the contribution and maturity of the work, relevant conferences include NeurIPS, CVPR, etc, while extended studies may be suitable for journals such as IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), Artificial Intelligence, IEEE Transactions on Robotics, and Nature Machine Intelligence.

Required knowledge

Applicants should have:

  • A good background in machine learning, deep learning, artificial intelligence, or computer vision
  • Strong Python programming skills
  • Experience with at least one deep-learning framework, preferably PyTorch
  • An interest in foundation models, multimodal AI, robotics, or embodied intelligence
  • Good analytical and problem-solving skills

Experience with Transformers, vision-language models, large language models, reinforcement learning, ROS/ROS2, CUDA, NVIDIA Jetson, or Linux would be useful but is not essential. Candidates with a strong machine-learning background who are interested in moving into embodied AI are also encouraged to apply.

Project funding

Other

Learn more about minimum entry requirements.