Primary supervisor
Julian Garcia GallegoAI safety often asks whether one system follows instructions or reaches the correct answer. Many important failures arise only when agents interact. Individually capable agents may exploit one another, fail to cooperate with unfamiliar partners, manipulate signals of trust, or pursue immediate rewards at the expense of accurate beliefs.
This project studies strategic AI safety: how incentives and social interaction shape the behaviour of AI agents. The student will choose one focused question, such as whether cooperation survives a change of partner, whether reputation identifies reliable agents, or whether language models respond to risk and temptation in the same way as people.
The project may use game-theoretic models, multi-agent reinforcement learning, or experiments with language models. An Honours project would normally select one game family and one safety property, then test when that property survives changes in partners, incentives or information.
URLs/references
References
- Perera, I., de Nijs, F. and García, J. “Learning to cooperate against ensembles of diverse opponents.” Neural Computing and Applications (2025). https://doi.org/10.1007/s00521-024-10511-9
- de Arruda, H. F., Gracia-Lázaro, C., Aleta, A. and Moreno, Y. “Collective cooperation without individual fidelity in LLM agents.” arXiv (2026). https://arxiv.org/abs/2606.30454
- García, J. and Traulsen, A. “Evolution of coordinated punishment to enforce cooperation from an unbiased strategy space.” Journal of the Royal Society Interface 16 (2019): 20190127. https://doi.org/10.1098/rsif.2019.0127
Required knowledge
- An interest in AI safety, strategic interaction and the behaviour of learning agents.
- Good Python programming skills and a solid mathematical background.
- Prior experience with reinforcement learning or game theory is useful but not required.
- Having taken FIT3139 is useful but not required.