VLDB 2026 Research / reviewers in the wild / expert
Alexandre Laterre
dblp:223/4200
· DBLP profile ↗
7ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 62% Planning, search and constraint satisfaction · 30% Motion planning and robot control · 8% | |
| Theoretical computer science
3 papers |
Mathematical optimization · 83% Algorithmic game theory and mechanism design · 17% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › reinforcement learning environment
benchmark environments |
0.8 | 1 | 2024 | Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX · ICLR 2024 |
Machine learning › Reinforcement learning
model-based reinforcement learning |
0.8 | 1 | 2024 | SPO: Sequential Monte Carlo Policy Optimisation · NeurIPS 2024 |
Machine learning › Reinforcement learning
reinforcement learning environment |
0.8 | 1 | 2024 | Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX · ICLR 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › search control
learning to branch |
0.7 | 1 | 2023 | Reinforcement Learning for Branch-and-Bound Optimisation Using Retrospective Trajectories · AAAI 2023 |
Machine learning › Reinforcement learning
policy learning |
0.7 | 1 | 2023 | Combinatorial Optimization with Policy Adaptation using Latent Space Search · NeurIPS 2023 |
Mathematical optimization › integer programming
branch-and-bound |
0.7 | 1 | 2023 | Reinforcement Learning for Branch-and-Bound Optimisation Using Retrospective Trajectories · AAAI 2023 |
Mathematical optimization
combinatorial optimization |
0.7 | 1 | 2023 | Combinatorial Optimization with Policy Adaptation using Latent Space Search · NeurIPS 2023 |
Mathematical optimization › discrete optimization
mixed integer linear programming |
0.7 | 1 | 2023 | Reinforcement Learning for Branch-and-Bound Optimisation Using Retrospective Trajectories · AAAI 2023 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.4 | 1 | 2020 | A game-theoretic analysis of networked system control for common-pool resource management using multi-agent reinforcement learning · NeurIPS 2020 |
Robotics › Motion planning and robot control
networked system control |
0.4 | 1 | 2020 | A game-theoretic analysis of networked system control for common-pool resource management using multi-agent reinforcement learning · NeurIPS 2020 |
Algorithmic game theory and mechanism design › computational game theory
empirical game-theoretic analysis |
0.4 | 1 | 2020 | A game-theoretic analysis of networked system control for common-pool resource management using multi-agent reinforcement learning · NeurIPS 2020 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search |
0.4 | 1 | 2019 | Learning Compositional Neural Programs with Recursive Tree Search and Planning · NeurIPS 2019 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
search-based planning |
0.4 | 1 | 2019 | Learning Compositional Neural Programs with Recursive Tree Search and Planning · NeurIPS 2019 |
Program synthesis and code generation
neural program synthesis |
0.4 | 1 | 2019 | Learning Compositional Neural Programs with Recursive Tree Search and Planning · NeurIPS 2019 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
tree search |
0.2 | 1 | 2024 | SPO: Sequential Monte Carlo Policy Optimisation · NeurIPS 2024 |
Mathematical optimization › combinatorial optimization › network optimization
routing and scheduling |
0.2 | 1 | 2023 | Combinatorial Optimization with Policy Adaptation using Latent Space Search · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 3.4retrospective trajectory construction · 1.3latent space search · 1.3imitation learning · 1.3multi-agent reinforcement learning · 0.9empirical game theory · 0.9sequential monte carlo · 0.8expectation-maximisation · 0.8actor-critic · 0.8JAX · 0.8neural programmer-interpreter · 0.4alphazero · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bootstrap Your Own Teacher: Online Policy Distillation for Multi-Game Reinforcement LearningabstractTraining generalist agents capable of performing well across diverse environments is a significant goal of reinforcement learning (RL). Current state-of-the-art methods for multi-game RL rely on offline datasets, and often discard the policy used to gather trajectories despite its potential to provide a rich learning signal. In this paper, we revisit policy distillation (PD) for multi-game RL and introduce a new framework called Bootstrap Your Own Teacher (BYOT) that extends policy distillation to the online-RL setting. BYOT alternates between two phases: (i) game-specific finetuning and (ii) distilling bootstrapped teachers back into a shared multi-game policy. By directly regulating the multi-game learning dynamics in policy space, BYOT balances training without explicit gradient adjustments or reward normalization, whilst being highly parameter efficient. Our framework is empirically validated for both online and offline multi-game learning on the Atari-40 benchmark. BYOT outperforms all prior online Atari-40 multigame agents, achieving an IQM human-normalized-score (HNS) of 152.7 %. Moreover, by adopting state-of-the-art PPO teacher agents—contrasting the widely-used datasets from weaker DQN agents—and policy distillation, we more than triple the leading IQM-HNS in offline settings to 369.5 %, whilst using significantly fewer parameters. Overall, our results emphasise the power of distillation in multi-game settings. Donal Byrne, Marko Tot, Paul Duckworth, Clément Bonnet, Alexandre Laterre, Thomas D. Barrett |
CoG | 5 |
| 2024 | Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAXabstractOpen-source reinforcement learning (RL) environments have played a crucial role in driving progress in the development of AI algorithms.
In modern RL research, there is a need for simulated environments that are performant, scalable, and modular to enable their utilization in a wider range of potential real-world applications.
Therefore, we present Jumanji, a suite of diverse RL environments specifically designed to be fast, flexible, and scalable.
Jumanji provides a suite of environments focusing on combinatorial problems frequently encountered in industry, as well as challenging general decision-making tasks.
By leveraging the efficiency of JAX and hardware accelerators like GPUs and TPUs, Jumanji enables rapid iteration of research ideas and large-scale experimentation, ultimately empowering more capable agents.
Unlike existing RL environment suites, Jumanji is highly customizable, allowing users to tailor the initial state distribution and problem complexity to their needs.
Furthermore, we provide actor-critic baselines for each environment, accompanied by preliminary findings on scaling and generalization scenarios.
Jumanji aims to set a new standard for speed, adaptability, and scalability of RL environments. Clément Bonnet, Daniel Luo, Donal Byrne, Shikha Surana, Sasha Abramowitz, Paul Duckworth, Vincent Coyette, Laurence Illing Midgley, Elshadai Tegegn, Tristan Kalloniatis, Omayma Mahjoub, Matthew Macfarlane, Andries P. Smit, Nathan Grinsztajn, Raphaël Boige, Cemlyn N. Waters, Mohamed A. Mimouni, Ulrich A. Mbou Sob, Ruan de Kock, Siddarth Singh, Daniel Furelos-Blanco, Victor Le, Arnu Pretorius, Alexandre Laterre |
ICLR | 24 |
| 2024 | SPO: Sequential Monte Carlo Policy OptimisationabstractLeveraging planning during learning and decision-making is central to the long-term development of intelligent agents. Recent works have successfully combined tree-based search methods and self-play learning mechanisms to this end. However, these methods typically face scaling challenges due to the sequential nature of their search. While practical engineering solutions can partly overcome this, they often result in a negative impact on performance. In this paper, we introduce SPO: Sequential Monte Carlo Policy Optimisation, a model-based reinforcement learning algorithm grounded within the Expectation Maximisation (EM) framework. We show that SPO provides robust policy improvement and efficient scaling properties. The sample-based search makes it directly applicable to both discrete and continuous action spaces without modifications. We demonstrate statistically significant improvements in performance relative to model-free and model-based baselines across both continuous and discrete environments. Furthermore, the parallel nature of SPO’s search enables effective utilisation of hardware accelerators, yielding favourable scaling laws. Matthew Macfarlane, Edan Toledo, Donal Byrne, Paul Duckworth, Alexandre Laterre |
NeurIPS | 5 |
| 2023 | Reinforcement Learning for Branch-and-Bound Optimisation Using Retrospective TrajectoriesabstractCombinatorial optimisation problems framed as mixed integer linear programmes (MILPs) are ubiquitous across a range of real-world applications. The canonical branch-and-bound algorithm seeks to exactly solve MILPs by constructing a search tree of increasingly constrained sub-problems. In practice, its solving time performance is dependent on heuristics, such as the choice of the next variable to constrain ('branching'). Recently, machine learning (ML) has emerged as a promising paradigm for branching. However, prior works have struggled to apply reinforcement learning (RL), citing sparse rewards, difficult exploration, and partial observability as significant challenges. Instead, leading ML methodologies resort to approximating high quality handcrafted heuristics with imitation learning (IL), which precludes the discovery of novel policies and requires expensive data labelling. In this work, we propose retro branching; a simple yet effective approach to RL for branching. By retrospectively deconstructing the search tree into multiple paths each contained within a sub-tree, we enable the agent to learn from shorter trajectories with more predictable next states. In experiments on four combinatorial tasks, our approach enables learning-to-branch without any expert guidance or pre-training. We outperform the current state-of-the-art RL branching algorithm by 3-5x and come within 20% of the best IL method's performance on MILPs with 500 constraints and 1000 variables, with ablations verifying that our retrospectively constructed trajectories are essential to achieving these results. Christopher Parsonson, Alexandre Laterre, Thomas D. Barrett |
AAAI | 2 |
| 2023 | Combinatorial Optimization with Policy Adaptation using Latent Space SearchabstractCombinatorial Optimization underpins many real-world applications and yet, designing performant algorithms to solve these complex, typically NP-hard, problems remains a significant research challenge. Reinforcement Learning (RL) provides a versatile framework for designing heuristics across a broad spectrum of problem domains. However, despite notable progress, RL has not yet supplanted industrial solvers as the go-to solution. Current approaches emphasize pre-training heuristics that construct solutions, but often rely on search procedures with limited variance, such as stochastically sampling numerous solutions from a single policy, or employing computationally expensive fine-tuning of the policy on individual problem instances. Building on the intuition that performant search at inference time should be anticipated during pre-training, we propose COMPASS, a novel RL approach that parameterizes a distribution of diverse and specialized policies conditioned on a continuous latent space. We evaluate COMPASS across three canonical problems - Travelling Salesman, Capacitated Vehicle Routing, and Job-Shop Scheduling - and demonstrate that our search strategy (i) outperforms state-of-the-art approaches in 9 out of 11 standard benchmarking tasks and (ii) generalizes better, surpassing all other approaches on a set of 18 procedurally transformed instance distributions. Félix Chalumeau, Shikha Surana, Clément Bonnet, Nathan Grinsztajn, Arnu Pretorius, Alexandre Laterre, Tom Barrett |
NeurIPS | 6 |
| 2020 | A game-theoretic analysis of networked system control for common-pool resource management using multi-agent reinforcement learningabstractMulti-agent reinforcement learning has recently shown great promise as an approach to networked system control. Arguably, one of the most difficult and important tasks for which large scale networked system control is applicable is common-pool resource management. Crucial common-pool resources include arable land, fresh water, wetlands, wildlife, fish stock, forests and the atmosphere, of which proper management is related to some of society's greatest challenges such as food security, inequality and climate change. Here we take inspiration from a recent research program investigating the game-theoretic incentives of humans in social dilemma situations such as the well-known \textit{tragedy of the commons}. However, instead of focusing on biologically evolved human-like agents, our concern is rather to better understand the learning and operating behaviour of engineered networked systems comprising general-purpose reinforcement learning agents, subject only to nonbiological constraints such as memory, computation and communication bandwidth. Harnessing tools from empirical game-theoretic analysis, we analyse the differences in resulting solution concepts that stem from employing different information structures in the design of networked multi-agent systems. These information structures pertain to the type of information shared between agents as well as the employed communication protocol and network topology. Our analysis contributes new insights into the consequences associated with certain design choices and provides an additional dimension of comparison between systems beyond efficiency, robustness, scalability and mean control performance. Arnu Pretorius, Scott Alexander Cameron, Elan Van Biljon, Tom Makkink, Shahil Mawjee, Jeremy du Plessis, Jonathan Shock, Alexandre Laterre, Karim Beguir |
NeurIPS | 8 |
| 2019 | Learning Compositional Neural Programs with Recursive Tree Search and PlanningabstractWe propose a novel reinforcement learning algorithm, AlphaNPI, that incorpo- rates the strengths of Neural Programmer-Interpreters (NPI) and AlphaZero. NPI contributes structural biases in the form of modularity, hierarchy and recursion, which are helpful to reduce sample complexity, improve generalization and in- crease interpretability. AlphaZero contributes powerful neural network guided search algorithms, which we augment with recursion. AlphaNPI only assumes a hierarchical program specification with sparse rewards: 1 when the program execution satisfies the specification, and 0 otherwise. This specification enables us to overcome the need for strong supervision in the form of execution traces and consequently train NPI models effectively with reinforcement learning. The experiments show that AlphaNPI can sort as well as previous strongly supervised NPI variants. The AlphaNPI agent is also trained on a Tower of Hanoi puzzle with two disks and is shown to generalize to puzzles with an arbitrary number of disks. The experiments also show that when deploying our neural network policies, it is advantageous to do planning with guided Monte Carlo tree search. Thomas Pierrot, Guillaume Ligner, Scott E. Reed, Olivier Sigaud, Nicolas Perrin-Gilbert, Alexandre Laterre, David Kas, Karim Beguir, Nando de Freitas |
NeurIPS | 6 |