Alexandre Laterre

dblp:223/4200 · DBLP profile ↗
← Back
7ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 62% Planning, search and constraint satisfaction · 30% Motion planning and robot control · 8%
Theoretical computer science
3 papers
Mathematical optimization · 83% Algorithmic game theory and mechanism design · 17%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 100%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › reinforcement learning environment
benchmark environments
0.812024
Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX · ICLR 2024
Machine learning › Reinforcement learning
model-based reinforcement learning
0.812024
SPO: Sequential Monte Carlo Policy Optimisation · NeurIPS 2024
Machine learning › Reinforcement learning
reinforcement learning environment
0.812024
Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX · ICLR 2024
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › search control
learning to branch
0.712023
Reinforcement Learning for Branch-and-Bound Optimisation Using Retrospective Trajectories · AAAI 2023
Machine learning › Reinforcement learning
policy learning
0.712023
Combinatorial Optimization with Policy Adaptation using Latent Space Search · NeurIPS 2023
Mathematical optimization › integer programming
branch-and-bound
0.712023
Reinforcement Learning for Branch-and-Bound Optimisation Using Retrospective Trajectories · AAAI 2023
Mathematical optimization
combinatorial optimization
0.712023
Combinatorial Optimization with Policy Adaptation using Latent Space Search · NeurIPS 2023
Mathematical optimization › discrete optimization
mixed integer linear programming
0.712023
Reinforcement Learning for Branch-and-Bound Optimisation Using Retrospective Trajectories · AAAI 2023
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.412020
A game-theoretic analysis of networked system control for common-pool resource management using multi-agent reinforcement learning · NeurIPS 2020
Robotics › Motion planning and robot control
networked system control
0.412020
A game-theoretic analysis of networked system control for common-pool resource management using multi-agent reinforcement learning · NeurIPS 2020
Algorithmic game theory and mechanism design › computational game theory
empirical game-theoretic analysis
0.412020
A game-theoretic analysis of networked system control for common-pool resource management using multi-agent reinforcement learning · NeurIPS 2020
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › game tree search
monte carlo tree search
0.412019
Learning Compositional Neural Programs with Recursive Tree Search and Planning · NeurIPS 2019
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
search-based planning
0.412019
Learning Compositional Neural Programs with Recursive Tree Search and Planning · NeurIPS 2019
Program synthesis and code generation
neural program synthesis
0.412019
Learning Compositional Neural Programs with Recursive Tree Search and Planning · NeurIPS 2019
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
tree search
0.212024
SPO: Sequential Monte Carlo Policy Optimisation · NeurIPS 2024
Mathematical optimization › combinatorial optimization › network optimization
routing and scheduling
0.212023
Combinatorial Optimization with Policy Adaptation using Latent Space Search · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 3.4retrospective trajectory construction · 1.3latent space search · 1.3imitation learning · 1.3multi-agent reinforcement learning · 0.9empirical game theory · 0.9sequential monte carlo · 0.8expectation-maximisation · 0.8actor-critic · 0.8JAX · 0.8neural programmer-interpreter · 0.4alphazero · 0.4
YearPublicationVenuePosition
2025 Bootstrap Your Own Teacher: Online Policy Distillation for Multi-Game Reinforcement Learning
abstract
Training generalist agents capable of performing well across diverse environments is a significant goal of reinforcement learning (RL). Current state-of-the-art methods for multi-game RL rely on offline datasets, and often discard the policy used to gather trajectories despite its potential to provide a rich learning signal. In this paper, we revisit policy distillation (PD) for multi-game RL and introduce a new framework called Bootstrap Your Own Teacher (BYOT) that extends policy distillation to the online-RL setting. BYOT alternates between two phases: (i) game-specific finetuning and (ii) distilling bootstrapped teachers back into a shared multi-game policy. By directly regulating the multi-game learning dynamics in policy space, BYOT balances training without explicit gradient adjustments or reward normalization, whilst being highly parameter efficient. Our framework is empirically validated for both online and offline multi-game learning on the Atari-40 benchmark. BYOT outperforms all prior online Atari-40 multigame agents, achieving an IQM human-normalized-score (HNS) of 152.7 %. Moreover, by adopting state-of-the-art PPO teacher agents—contrasting the widely-used datasets from weaker DQN agents—and policy distillation, we more than triple the leading IQM-HNS in offline settings to 369.5 %, whilst using significantly fewer parameters. Overall, our results emphasise the power of distillation in multi-game settings.
Donal Byrne, Marko Tot, Paul Duckworth, Clément Bonnet, Alexandre Laterre, Thomas D. Barrett
CoG5
2024 Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAX
abstract
Open-source reinforcement learning (RL) environments have played a crucial role in driving progress in the development of AI algorithms. In modern RL research, there is a need for simulated environments that are performant, scalable, and modular to enable their utilization in a wider range of potential real-world applications. Therefore, we present Jumanji, a suite of diverse RL environments specifically designed to be fast, flexible, and scalable. Jumanji provides a suite of environments focusing on combinatorial problems frequently encountered in industry, as well as challenging general decision-making tasks. By leveraging the efficiency of JAX and hardware accelerators like GPUs and TPUs, Jumanji enables rapid iteration of research ideas and large-scale experimentation, ultimately empowering more capable agents. Unlike existing RL environment suites, Jumanji is highly customizable, allowing users to tailor the initial state distribution and problem complexity to their needs. Furthermore, we provide actor-critic baselines for each environment, accompanied by preliminary findings on scaling and generalization scenarios. Jumanji aims to set a new standard for speed, adaptability, and scalability of RL environments.
Clément Bonnet, Daniel Luo, Donal Byrne, Shikha Surana, Sasha Abramowitz, Paul Duckworth, Vincent Coyette, Laurence Illing Midgley, Elshadai Tegegn, Tristan Kalloniatis, Omayma Mahjoub, Matthew Macfarlane, Andries P. Smit, Nathan Grinsztajn, Raphaël Boige, Cemlyn N. Waters, Mohamed A. Mimouni, Ulrich A. Mbou Sob, Ruan de Kock, Siddarth Singh, Daniel Furelos-Blanco, Victor Le, Arnu Pretorius, Alexandre Laterre
ICLR24
2024 SPO: Sequential Monte Carlo Policy Optimisation
abstract
Leveraging planning during learning and decision-making is central to the long-term development of intelligent agents. Recent works have successfully combined tree-based search methods and self-play learning mechanisms to this end. However, these methods typically face scaling challenges due to the sequential nature of their search. While practical engineering solutions can partly overcome this, they often result in a negative impact on performance. In this paper, we introduce SPO: Sequential Monte Carlo Policy Optimisation, a model-based reinforcement learning algorithm grounded within the Expectation Maximisation (EM) framework. We show that SPO provides robust policy improvement and efficient scaling properties. The sample-based search makes it directly applicable to both discrete and continuous action spaces without modifications. We demonstrate statistically significant improvements in performance relative to model-free and model-based baselines across both continuous and discrete environments. Furthermore, the parallel nature of SPO’s search enables effective utilisation of hardware accelerators, yielding favourable scaling laws.
Matthew Macfarlane, Edan Toledo, Donal Byrne, Paul Duckworth, Alexandre Laterre
NeurIPS5
2023 Reinforcement Learning for Branch-and-Bound Optimisation Using Retrospective Trajectories
abstract
Combinatorial optimisation problems framed as mixed integer linear programmes (MILPs) are ubiquitous across a range of real-world applications. The canonical branch-and-bound algorithm seeks to exactly solve MILPs by constructing a search tree of increasingly constrained sub-problems. In practice, its solving time performance is dependent on heuristics, such as the choice of the next variable to constrain ('branching'). Recently, machine learning (ML) has emerged as a promising paradigm for branching. However, prior works have struggled to apply reinforcement learning (RL), citing sparse rewards, difficult exploration, and partial observability as significant challenges. Instead, leading ML methodologies resort to approximating high quality handcrafted heuristics with imitation learning (IL), which precludes the discovery of novel policies and requires expensive data labelling. In this work, we propose retro branching; a simple yet effective approach to RL for branching. By retrospectively deconstructing the search tree into multiple paths each contained within a sub-tree, we enable the agent to learn from shorter trajectories with more predictable next states. In experiments on four combinatorial tasks, our approach enables learning-to-branch without any expert guidance or pre-training. We outperform the current state-of-the-art RL branching algorithm by 3-5x and come within 20% of the best IL method's performance on MILPs with 500 constraints and 1000 variables, with ablations verifying that our retrospectively constructed trajectories are essential to achieving these results.
Christopher Parsonson, Alexandre Laterre, Thomas D. Barrett
AAAI2
2023 Combinatorial Optimization with Policy Adaptation using Latent Space Search
abstract
Combinatorial Optimization underpins many real-world applications and yet, designing performant algorithms to solve these complex, typically NP-hard, problems remains a significant research challenge. Reinforcement Learning (RL) provides a versatile framework for designing heuristics across a broad spectrum of problem domains. However, despite notable progress, RL has not yet supplanted industrial solvers as the go-to solution. Current approaches emphasize pre-training heuristics that construct solutions, but often rely on search procedures with limited variance, such as stochastically sampling numerous solutions from a single policy, or employing computationally expensive fine-tuning of the policy on individual problem instances. Building on the intuition that performant search at inference time should be anticipated during pre-training, we propose COMPASS, a novel RL approach that parameterizes a distribution of diverse and specialized policies conditioned on a continuous latent space. We evaluate COMPASS across three canonical problems - Travelling Salesman, Capacitated Vehicle Routing, and Job-Shop Scheduling - and demonstrate that our search strategy (i) outperforms state-of-the-art approaches in 9 out of 11 standard benchmarking tasks and (ii) generalizes better, surpassing all other approaches on a set of 18 procedurally transformed instance distributions.
Félix Chalumeau, Shikha Surana, Clément Bonnet, Nathan Grinsztajn, Arnu Pretorius, Alexandre Laterre, Tom Barrett
NeurIPS6
2020 A game-theoretic analysis of networked system control for common-pool resource management using multi-agent reinforcement learning
abstract
Multi-agent reinforcement learning has recently shown great promise as an approach to networked system control. Arguably, one of the most difficult and important tasks for which large scale networked system control is applicable is common-pool resource management. Crucial common-pool resources include arable land, fresh water, wetlands, wildlife, fish stock, forests and the atmosphere, of which proper management is related to some of society's greatest challenges such as food security, inequality and climate change. Here we take inspiration from a recent research program investigating the game-theoretic incentives of humans in social dilemma situations such as the well-known \textit{tragedy of the commons}. However, instead of focusing on biologically evolved human-like agents, our concern is rather to better understand the learning and operating behaviour of engineered networked systems comprising general-purpose reinforcement learning agents, subject only to nonbiological constraints such as memory, computation and communication bandwidth. Harnessing tools from empirical game-theoretic analysis, we analyse the differences in resulting solution concepts that stem from employing different information structures in the design of networked multi-agent systems. These information structures pertain to the type of information shared between agents as well as the employed communication protocol and network topology. Our analysis contributes new insights into the consequences associated with certain design choices and provides an additional dimension of comparison between systems beyond efficiency, robustness, scalability and mean control performance.
Arnu Pretorius, Scott Alexander Cameron, Elan Van Biljon, Tom Makkink, Shahil Mawjee, Jeremy du Plessis, Jonathan Shock, Alexandre Laterre, Karim Beguir
NeurIPS8
2019 Learning Compositional Neural Programs with Recursive Tree Search and Planning
abstract
We propose a novel reinforcement learning algorithm, AlphaNPI, that incorpo- rates the strengths of Neural Programmer-Interpreters (NPI) and AlphaZero. NPI contributes structural biases in the form of modularity, hierarchy and recursion, which are helpful to reduce sample complexity, improve generalization and in- crease interpretability. AlphaZero contributes powerful neural network guided search algorithms, which we augment with recursion. AlphaNPI only assumes a hierarchical program specification with sparse rewards: 1 when the program execution satisfies the specification, and 0 otherwise. This specification enables us to overcome the need for strong supervision in the form of execution traces and consequently train NPI models effectively with reinforcement learning. The experiments show that AlphaNPI can sort as well as previous strongly supervised NPI variants. The AlphaNPI agent is also trained on a Tower of Hanoi puzzle with two disks and is shown to generalize to puzzles with an arbitrary number of disks. The experiments also show that when deploying our neural network policies, it is advantageous to do planning with guided Monte Carlo tree search.
Thomas Pierrot, Guillaume Ligner, Scott E. Reed, Olivier Sigaud, Nicolas Perrin-Gilbert, Alexandre Laterre, David Kas, Karim Beguir, Nando de Freitas
NeurIPS6