I am an Associate Professor in Artificial Intelligence in the School of Computing and Information Systems, Faculty of Engineering and Information Technology, at The University of Melbourne (Australia), where I lead the Artificial Intelligence group. I serve as Course Director of the Master of Artificial Intelligence (Online), and as a member of the ICAPS Executive Council, where I am the Mentoring and Diversity Chair. My research and teaching interests span AI planning, search, learning, reasoning with large language models, intention recognition, and autonomous systems. I’m a member of the AI and Autonomous Agents Lab and the Digital Agriculture, Food and Wine lab.
My research focuses on how to introduce different approaches to the problem of inference in sequential decision problems, as well as applications to autonomous systems in agriculture.
I completed my PhD at the Artificial Intelligence and Machine Learning Group, Universitat Pompeu Fabra, under the supervision of Prof. Hector Geffner. I was a research fellow for 3 years under the supervision of Prof. Peter Stuckey and Prof. Adrian Pearce, working on solving Mining Scheduling problems through automated planning, constraint programming and operations research techniques. Since then, I have built a research program around width-based search and novelty, whose planners have been awarded in several International Planning Competitions, most recently winning the Satisficing and Agile tracks of the IPC 2026 Numeric Tracks. This program has grown to intersect with other areas, from goal and intention recognition with humans, to the integration of reasoning and learning, where planning meets machine learning and large language models.
Graduate Certificate in University Teaching, 2020
The University of Melbourne
PhD in Artificial Intelligence, 2012
Universitat Pompeu Fabra
MEng in Artificial Intelligence, 2007
Universitat Pompeu Fabra
BSc in Computer Science, 2004
Universitat Pompeu Fabra
[10/26] New NeurIPS 2026 paper on SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning
[10/26] New EMNLP 2026 paper on Mind the Perspective: Let’s Reason Recursively for Theory of Mind
[10/26] New EMNLP 2026 Findings paper on From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
[07/26] Our Panino planner won the Satisficing and Agile tracks (Overall) at the IPC 2026 Numeric Tracks ( planner abstract)
[06/26] New DIS 2026 Best Paper Honourable Mention on Design and Evaluation of AR-Based Real-Time Feedback System for Kinesthetic Robot Teaching
[06/26] New ICAPS paper on Planning as Goal Recognition: Deriving Heuristics from Intention Models
[06/26] New ICAPS paper on Computing Planning Width: How Hard Is It, Really?
[06/26] New HSDIP workshop paper on A Notion of Width for Numeric Planning Problems
[06/26] New PR-BGI workshop paper on Divergence Identities for Probabilistic Goal Recognition
[06/26] New CAiSE paper on Mining Role-based Behavioral Patterns From Event Data for Effective Process Simulation
Planning as a Service (PaaS) is an extendable API to deploy planners online in local or cloud servers
Farm.bot is an open-source robotic platform to explore problems on AI and Automation (Planning, Vision, Learning) for small scale …
Width Based Planning searches for solutions through a general measure of state novelty. Performs well over black-box simulators and …
Planimation is a framework to visualise sequential solutions of planning problems specified in PDDL
Award-winning classical and numeric planners in several International Planning Competitions 2008 - 2026
Invariants, Traps, Un-reachability Certificates, and Dead-end Detection
Software to support AI courses in Mel & RMIT Unis (Melbourne, AUS)
Classical Planners playing Atari 2600 games as well as Deep Reinforcement Learning
classical planners computing infinite loopy plans, and FOND planners synthesizing controllers expressed as policies.
Lightweight Automated Planning ToolKiT (LAPKT) to build, use or extend basic to advanced Automated Planners
Width-based algorithms search for solutions through a general definition of state novelty. These algorithms have been shown to result in state-of-the-art performance in classical planning, and have been successfully applied to model-based and model-free settings where the dynamics of the problem are given through simulation engines. Width-based algorithms performance is understood theoretically through the notion of planning width, providing polynomial guarantees on their runtime and memory consumption. To facilitate synergies across research communities, this paper summarizes the area of width-based planning, and surveys current and future research directions.
Recent reasoning-oriented LLMs have demonstrated strong performance on challenging tasks such as mathematics and science examinations. However, core cognitive faculties of human intelligence, such as abstract reasoning and generalization, remain underexplored. To address this, we evaluate recent reasoning-oriented LLMs on the Abstraction and Reasoning Corpus (ARC) benchmark, which explicitly demands both faculties. We formulate ARC as a program synthesis task and propose nine candidate solvers. Experimental results show that repeated-sampling planning-aided code generation (RSPC) achieves the highest test accuracy and demonstrates consistent generalization across most LLMs. To further improve performance, we introduce an ARC solver, Knowledge Augmentation for Abstract Reasoning (KAAR), which encodes core knowledge priors within an ontology that classifies priors into three hierarchical levels based on their dependencies. KAAR progressively expands LLM reasoning capacity by gradually augmenting priors at each level, and invokes RSPC to generate candidate solutions after each augmentation stage. This stage-wise reasoning reduces interference from irrelevant priors and improves LLM performance. Empirical results show that KAAR maintains strong generalization and consistently outperforms non-augmented RSPC across all evaluated LLMs, achieving around 5% absolute gains and up to 64.52% relative improvement. Despite these achievements, ARC remains a challenging benchmark for reasoning-oriented LLMs, highlighting future avenues of progress in LLMs. Our code is available at https://github.com/you68681/kaar.
Theory of Mind (ToM) reasoning requires inferring agents’ beliefs from partial and asymmetric observations, which remains an open challenge for LLMs. Existing prompting-based approaches improve ToM reasoning through observable-event filtering or temporal belief chains, without explicitly modeling nested beliefs. We introduce RecToM, an inference-time framework for ToM reasoning that models nested beliefs via recursive perspective construction. RecToM constructs each character perspective from the preceding character perspective along the character chain specified by the question, reducing higher-order belief questions to actual-world questions within the final constructed perspective. We further provide a KD45 analysis showing that RecToM’s perspective construction induces a well-formed belief modality beyond simple event filtering. Experiments on ToM benchmarks, including Hi-ToM, Big-ToM, and FanToM, across multiple LLM backbones show that RecToM consistently outperforms recent advanced approaches, achieving state-of-the-art performance. Notably, RecToM reaches 100% accuracy on Hi-ToM with GPT-5.4 and Qwen3.5, a benchmark requiring higher-order ToM reasoning. Our code is available at https://github.com/you68681/rectom.
Spatial reasoning remains a challenge for Multimodal Large Language Models (MLLMs), as it requires reliable multi-hop inference over both intermediate states and state transitions. Current studies often leave intermediate states unverified and treat state transitions as implicit processes, which limits reliability in multi-hop spatial reasoning. To address this, we propose State-aware Visualization-of-Thought (SVoT), a reinforcement learning framework that generates interleaved, verifiable intermediate states and visualizations. SVoT integrates transition reasoning chains into the generation processes, enabling the model to verify action preconditions and effects through interleaved textual and visual reasoning. We train SVoT via Group Relative Policy Optimization (GRPO), instantiating verification through reward design and evaluating the efficacy of different fine-grained rewards. As existing benchmarks reduce state transitions to single-variable updates, substantially simplifying the problems, we establish five domains by extending classical environments and introducing two novel domains, Pacman and Gather, that require multi-object interactions and numerical reasoning. These domains support systematic evaluation of multi-hop spatial reasoning with quantitative verification of generated intermediate states and transition reasoning. SVoT with transition-aware supervision achieves state-of-the-art performance across the introduced domains, yielding up to a 65% absolute accuracy gain on out-of-distribution test sets.
Winner of the Satisficing and Agile tracks (Overall) at IPC 2026 Numeric Tracks.
Best Paper Honourable Mention at DIS 2026.
Benjamin Grayland [2026 - current] co-supervised with Prof. Sebastian Sardina and Dr. Steven Korevaar, Topic: Structured Learning approaches for multi-agent problems
Giacomo Rosa [2024 - current] co-supervised with Prof. Sebastian Sardina and Dr. Jean Honorio, Topic: Exploration methods for Planning
Jiajia Song [2024 - current] co-supervised with Prof. Sebastian Sardina and Dr. William Umboh, Topic: What Makes AI Planning Hard? From Complexity Analysis to Algorithm Design
David Adams [2024 - current] co-supervised with Dr. Renata Borovica-Gajic, Topic: Exploration Methods for Databases
Qingtan Shen [2023 - current] co-supervised with A/Prof. Artem Polyvyanyy and Dr. Timotheus Kampik, Topic: Multi-agent system discovery
Ciao Lei [2022 - current]. co-supervised with Dr. Kris Ehinger and A/Prof Sigfredo Fuentes, Topic: Generalized vision planning problems and their applications in Agriculture
Zhiaho Pei [2022 - current]. co-supervized with Dr. Angela Rojas, Dr. Fjalar De Haan and Dr. Enayat A. Moallemi, Topic: Robust decision making for complex systems
Muhammad Bilal [2023 - 2026], co-supervised with Dr. Wafa Johal and Prof. Denny Oetomo. Thesis:
Understanding and Improving Novice Kinesthetic Teaching for Robot Learning from Demonstration First Employment: Post-Doc @ The University of Texas at Austin [2026 - current]
Sukai Huang [2022 - 2025]. co-supervized with Prof. Trevor Cohn, Thesis:
Integrating Natural Language in Sequential Decision Problems First Employment: Post-Doc @ Monash University [2025- current]
Lingfei Wang [2021 - 2025], co-supervised with Dr.Maria Rodriguez. Thesis:
Job Scheduling in High Performance Computing Clusters with Deep Reinforcement Learning First Employment: TBA
Guang Hu [2021 - 2025], co-supervised with Dr.Tim Miller. Thesis:
“Seen Is Believing”: Modeling and Solving Epistemic Planning Problems using Justified Perspectives First Employment: Education Specialist @ The University of Melbourne [2026 - Current]
Zihang Su [2020 - 2024], co-supervised with Dr. Artem Polyvyanyy and Prof. Sebastian Sardina. Thesis:
Evidence-Based Goal Recognition Using Process Mining Techniques First Employment: Post-Doc @ Tsinghua University [2024 - current]
Chenyuan Zhang [2020 - 2024], co-supervised with A/Prof. Charles Kemp (Psychology). Thesis:
Planning and Goal Recognition in Humans and Machines First Employment: Post-Doc @ Monash University [2024 - current] Best Student Paper Award AAMAS (2024)
Anubhav Singh [2019 - 2024], co-supervized with Dr. Miquel Ramirez and Prof. Peter Stuckey. Thesis:
Lazy Constraint Generation and Tractable Approximations for Large-scale Planning Problems First Employment: Post-Doc @ Universtiy of Toronto [2024 - current]
Stefan O'Toole [2018 - 2022], co-supervized with Dr. Miquel Ramirez and Prof. Adrian Pearce. Thesis:
The Intersection of Planning and Learning through Cost-to-go Approximations, Imitation and Symbolic Regression First Employment: Meta [2022 - current]
Toby Davies, [2013-2017], co-supervized with Prof. Adrian Pearce, Prof. Peter Stuckey and Prof. Harald Sondergaard. Thesis:
Learning from Conflict in Multi-Agent, Classical, and Temporal Planning. First Employment: Google [2017 - current]. Best Paper Award ICAPS (2015), Best PhD Thesis, Melbourne School of Engineering 2018
Giacomo Rosa [2023-2024]. Thesis:
Count-Based Novelty Exploration and
ECAI24 paper
Zhiaho Pei [2021]. co-supervized with Dr. Angela Rojas, Dr. Fjalar De Haan and Dr. Enayat A. Moallemi, Thesis:
Robust decision making for complex systems
Marco Marasco [2021]. co-supervized with Dr. Angela Rojas, Dr. Fjalar De Haan and Dr. Enayat A. Moallemi, Thesis:
Adaptive Policy making for systems of electricity provision
Jiayuan Chang [2021]. co-supervized with A/Prof Sigfredo Fuentes, Thesis:
FarmBot.io Automated Planning: simulation and integration
Yajing Ma [2021]. co-supervized with A/Prof Sigfredo Fuentes, Thesis:
Electronic Nose for pest detection
Dmitry Grebenyuk [2018-2020], co-supervised with Dr. Miquel Ramirez, and Dr. Kris Ehinger. Thesis:
Agnostic Features for generalized policies computed with Deep Reinforcement Learning (DRL). First Employment: Start-up working on Image Processing using DRL
Guang Hu [2018-2020], co-supervised with Dr.Tim Miller. Thesis:
What you get is what you see: Decomposing Epistemic Planning using Functional STRIPS. PhD Candidate [2021 - current]
Ciao Lei [2019-2020]. Thesis:
Regression and Width in Classical Planning and
ICAPS21 paper
Artificial Intelligence group Lead, School of Computing and Information Systems, The University of Melbourne, (2025 - ongoing)
Course Director,
Master of Artificial Intelligence (Online), The University of Melbourne, (2025 - ongoing)
ICAPS Executive Council – Member, and Mentoring and Diversity Chair, (2025 - 2031)
ICAPS (2025)ICAPS (2019)
Optimisation and Planning ICAPS 2025 Summer School – Organizer, (2025)
AgentsVic Autumn Symposium on Reasoning and Learning for Autonomous Agents – Organizer, (2024)
International Conference on Automated Planning and Scheduling – Publicity co-chair, ICAPS (2010)
First Unsolvability International Planning Competition – Co-Organizer, UIPC-1 (2016)
Heuristics and Search for Domain-independent Planning – Co-Organizer, ICAPS workshop HSDIP (2015,2016,2017,2018)
Demonstration track – Co-Chair AAAI (2023)
Student Abstract track – Co-Chair, AAAI (2018,2019)
Journal Presentation track – Co-Chair ICAPS (2018)
AAAI (2027)AAAI (2020,2021,2022,2023)IJCAI (2021,2023)(Jan 2026)International Joint Conferences on Artificial Intelligence IJCAI (2011,2013,2015,2017,2018,2020,2022)
Association for the Advancement of Artificial Intelligence, AAAI (2013,2015,2016,2017,2018,2019)
European Conference on Artificial Intelligence, ECAI (2014,2016)
International Conference on Automated Planning and Scheduling, ICAPS (2015,2016,2017,2018,2020)
Symposium on Combinatorial Search SOCS (2020,2021,2022,2023)
Journal of Artificial Intelligence Research, JAIR
Reviewer Artificial Intelligence, Elsevier AIJ
Reviewer Communications of the ACM, CACM
(2024,2025)Master of Artificial Intelligence - Online (Course Director), at The University of Melbourne,
2025 - ongoing
Pacman Capture the flag Inter-University Contest, run for Unimelb AI coure and
Hall of Fame contest, 2016 - current
AI Planning for Autonomy (Lecturer), at M.Sc. AI specialization, The University of Melbourne,
2016 - current
Data Structures and Algorithms (Lecturer), at The University of Melbourne,
2016 - current
Software Agents (Lecturer), at M.Sc. Software, The University of Melbourne, 2013, 2014, 2015
Autonomous Systems, at M.Sc. Intelligent Interactive Systems, University Pompeu Fabra, 2012
Advanced course on AI: workshop on RoboSoccer simulator, at Polytechnic School, University Pompeu Fabra, 2009, 2010, 2011
Artificial Intelligence course, at Polytechnic School, University Pompeu Fabra, 2010, 2011
Introduction to Data Structures and Algorithms course, at Polytechnic School, University Pompeu Fabra, 2008
Programming course, at Polytechnic School, University Pompeu Fabra, 2008, 2009, 2010, 2011