VLDB 2026 Research / reviewers in the wild / expert
Thomas D. Barrett
dblp:248/8263
· DBLP profile ↗
8ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0001-6241-3028ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Reinforcement learning · 27% Multi-agent systems · 25% Language models and text generation · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% | |
| Theoretical computer science
2 papers |
Mathematical optimization · 94% Graph algorithms and graph theory · 6% |
Topics — the 18 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › protein sequence analysis › protein sequence representation
protein language model |
1.1 | 2 | 2025 | Metalic: Meta-Learning In-Context with Protein Language Models · ICLR 2025 ManyFold: an efficient and flexible library for training and validating protein folding models · Bioinform. 2023 |
Bioinformatics and computational biology › protein function prediction › protein variant effect prediction
protein fitness prediction |
0.9 | 1 | 2025 | Metalic: Meta-Learning In-Context with Protein Language Models · ICLR 2025 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.8 | 1 | 2024 | Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs · ICML 2024 |
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination |
0.8 | 1 | 2024 | Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs · ICML 2024 |
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
multi-agent debate |
0.8 | 1 | 2024 | Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs · ICML 2024 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › search control
learning to branch |
0.7 | 1 | 2023 | Reinforcement Learning for Branch-and-Bound Optimisation Using Retrospective Trajectories · AAAI 2023 |
Bioinformatics and computational biology
protein structure prediction |
0.7 | 1 | 2023 | ManyFold: an efficient and flexible library for training and validating protein folding models · Bioinform. 2023 |
Mathematical optimization › integer programming
branch-and-bound |
0.7 | 1 | 2023 | Reinforcement Learning for Branch-and-Bound Optimisation Using Retrospective Trajectories · AAAI 2023 |
Mathematical optimization › discrete optimization
mixed integer linear programming |
0.7 | 1 | 2023 | Reinforcement Learning for Branch-and-Bound Optimisation Using Retrospective Trajectories · AAAI 2023 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › multi-agent communication
communication protocols |
0.6 | 1 | 2022 | Universally Expressive Communication in Multi-Agent Reinforcement Learning · NeurIPS 2022 |
Machine learning › Graph learning
graph neural network |
0.6 | 1 | 2022 | Universally Expressive Communication in Multi-Agent Reinforcement Learning · NeurIPS 2022 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.6 | 1 | 2022 | Universally Expressive Communication in Multi-Agent Reinforcement Learning · NeurIPS 2022 |
Machine learning › Representation and self-supervised learning › representation learning
disentangled representation learning |
0.4 | 1 | 2020 | Learning Disentangled Representations and Group Structure of Dynamical Environments · NeurIPS 2020 |
Robotics › Motion planning and robot control
dynamic modeling |
0.4 | 1 | 2020 | Learning Disentangled Representations and Group Structure of Dynamical Environments · NeurIPS 2020 |
Machine learning › Reinforcement learning
reinforcement learning for combinatorial optimization |
0.4 | 1 | 2020 | Exploratory Combinatorial Optimization with Reinforcement Learning · AAAI 2020 |
Mathematical optimization
combinatorial optimization |
0.4 | 1 | 2020 | Exploratory Combinatorial Optimization with Reinforcement Learning · AAAI 2020 |
Mathematical optimization › combinatorial optimization
graph combinatorial optimization |
0.4 | 1 | 2020 | Exploratory Combinatorial Optimization with Reinforcement Learning · AAAI 2020 |
Graph algorithms and graph theory › graph cut
max-cut |
0.1 | 1 | 2020 | Exploratory Combinatorial Optimization with Reinforcement Learning · AAAI 2020 |
Methods — techniques the papers use, named apart from their topics
retrospective trajectory construction · 1.3reinforcement learning · 1.3imitation learning · 1.3protein language model · 0.9meta-learning · 0.9in-context learning · 0.9fine-tuning · 0.9exploratory search · 0.9deep q-network · 0.9self-consistency · 0.8prompting · 0.8ensembling · 0.8multiple sequence alignment · 0.7deep learning · 0.7JAX · 0.7theoretical analysis · 0.6graph neural network · 0.6random search · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bootstrap Your Own Teacher: Online Policy Distillation for Multi-Game Reinforcement LearningabstractTraining generalist agents capable of performing well across diverse environments is a significant goal of reinforcement learning (RL). Current state-of-the-art methods for multi-game RL rely on offline datasets, and often discard the policy used to gather trajectories despite its potential to provide a rich learning signal. In this paper, we revisit policy distillation (PD) for multi-game RL and introduce a new framework called Bootstrap Your Own Teacher (BYOT) that extends policy distillation to the online-RL setting. BYOT alternates between two phases: (i) game-specific finetuning and (ii) distilling bootstrapped teachers back into a shared multi-game policy. By directly regulating the multi-game learning dynamics in policy space, BYOT balances training without explicit gradient adjustments or reward normalization, whilst being highly parameter efficient. Our framework is empirically validated for both online and offline multi-game learning on the Atari-40 benchmark. BYOT outperforms all prior online Atari-40 multigame agents, achieving an IQM human-normalized-score (HNS) of 152.7 %. Moreover, by adopting state-of-the-art PPO teacher agents—contrasting the widely-used datasets from weaker DQN agents—and policy distillation, we more than triple the leading IQM-HNS in offline settings to 369.5 %, whilst using significantly fewer parameters. Overall, our results emphasise the power of distillation in multi-game settings. Donal Byrne, Marko Tot, Paul Duckworth, Clément Bonnet, Alexandre Laterre, Thomas D. Barrett |
CoG | 6 |
| 2025 | Metalic: Meta-Learning In-Context with Protein Language ModelsabstractPredicting the biophysical and functional properties of proteins is essential for in silico protein design. Machine learning has emerged as a promising technique for such prediction tasks. However, the relative scarcity of in vitro annotations means that these models often have little, or no, specific data on the desired fitness prediction task. As a result of limited data, protein language models (PLMs) are typically trained on general protein sequence modeling tasks, and then fine-tuned, or applied zero-shot, to protein fitness prediction. When no task data is available, the models make strong assumptions about the correlation between the protein sequence likelihood and fitness scores. In contrast, we propose meta-learning over a distribution of standard fitness prediction tasks, and demonstrate positive transfer to unseen fitness prediction tasks. Our method, called Metalic (Meta-Learning In-Context), uses in-context learning and fine-tuning, when data is available, to adapt to new tasks. Crucially, fine-tuning enables considerable generalization, even though it is not accounted for during meta-training. Our fine-tuned models achieve strong results with 18 times fewer parameters than state-of-the-art models. Moreover, our method sets a new state-of-the-art in low-data settings on ProteinGym, an established fitness-prediction benchmark. Due to data scarcity, we believe meta-learning will play a pivotal role in advancing protein engineering. Jacob Beck, Shikha Surana, Manus McAuliffe, Oliver Bent, Thomas D. Barrett, Juan Jose Garau Luis, Paul Duckworth |
ICLR | 5 |
| 2024 | Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMsabstractRecent advancements in large language models (LLMs) underscore their potential for responding to inquiries in various domains. However, ensuring that generative agents provide accurate and reliable answers remains an ongoing challenge. In this context, multi-agent debate (MAD) has emerged as a promising strategy for enhancing the truthfulness of LLMs. We benchmark a range of debating and prompting strategies to explore the trade-offs between cost, time, and accuracy. Importantly, we find that multi-agent debating systems, in their current form, do not reliably outperform other proposed prompting strategies, such as self-consistency and ensembling using multiple reasoning paths. However, when performing hyperparameter tuning, several MAD systems, such as Multi-Persona, perform better. This suggests that MAD protocols might not be inherently worse than other approaches, but that they are more sensitive to different hyperparameter settings and difficult to optimize. We build on these results to offer insights into improving debating strategies, such as adjusting agent agreement levels, which can significantly enhance performance and even surpass all other non-debate protocols we evaluated. We provide an open-source repository to the community with several state-of-the-art protocols together with evaluation scripts to benchmark across popular research datasets. Andries P. Smit, Nathan Grinsztajn, Paul Duckworth, Thomas D. Barrett, Arnu Pretorius |
ICML | 4 |
| 2023 | Reinforcement Learning for Branch-and-Bound Optimisation Using Retrospective TrajectoriesabstractCombinatorial optimisation problems framed as mixed integer linear programmes (MILPs) are ubiquitous across a range of real-world applications. The canonical branch-and-bound algorithm seeks to exactly solve MILPs by constructing a search tree of increasingly constrained sub-problems. In practice, its solving time performance is dependent on heuristics, such as the choice of the next variable to constrain ('branching'). Recently, machine learning (ML) has emerged as a promising paradigm for branching. However, prior works have struggled to apply reinforcement learning (RL), citing sparse rewards, difficult exploration, and partial observability as significant challenges. Instead, leading ML methodologies resort to approximating high quality handcrafted heuristics with imitation learning (IL), which precludes the discovery of novel policies and requires expensive data labelling. In this work, we propose retro branching; a simple yet effective approach to RL for branching. By retrospectively deconstructing the search tree into multiple paths each contained within a sub-tree, we enable the agent to learn from shorter trajectories with more predictable next states. In experiments on four combinatorial tasks, our approach enables learning-to-branch without any expert guidance or pre-training. We outperform the current state-of-the-art RL branching algorithm by 3-5x and come within 20% of the best IL method's performance on MILPs with 500 constraints and 1000 variables, with ablations verifying that our retrospectively constructed trajectories are essential to achieving these results. Christopher Parsonson, Alexandre Laterre, Thomas D. Barrett |
AAAI | 3 |
| 2023 | ManyFold: an efficient and flexible library for training and validating protein folding modelsabstractSUMMARY: ManyFold is a flexible library for protein structure prediction with deep learning that (i) supports models that use both multiple sequence alignments (MSAs) and protein language model (pLM) embedding as inputs, (ii) allows inference of existing models (AlphaFold and OpenFold), (iii) is fully trainable, allowing for both fine-tuning and the training of new models from scratch and (iv) is written in Jax to support efficient batched operation in distributed settings. A proof-of-concept pLM-based model, pLMFold, is trained from scratch to obtain reasonable results with reduced computational overheads in comparison to AlphaFold. AVAILABILITY AND IMPLEMENTATION: The source code for ManyFold, the validation dataset and a small sample of training data are available at https://github.com/instadeepai/manyfold. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Amelia Villegas-Morcillo, Louis Robinson, Arthur Flajolet, Thomas D. Barrett |
Bioinform. | 4 |
| 2022 | Universally Expressive Communication in Multi-Agent Reinforcement LearningabstractAllowing agents to share information through communication is crucial for solving complex tasks in multi-agent reinforcement learning. In this work, we consider the question of whether a given communication protocol can express an arbitrary policy. By observing that many existing protocols can be viewed as instances of graph neural networks (GNNs), we demonstrate the equivalence of joint action selection to node labelling. With standard GNN approaches provably limited in their expressive capacity, we draw from existing GNN literature and consider augmenting agent observations with: (1) unique agent IDs and (2) random noise. We provide a theoretical analysis as to how these approaches yield universally expressive communication, and also prove them capable of targeting arbitrary sets of actions for identical agents. Empirically, these augmentations are found to improve performance on tasks where expressive communication is required, whilst, in general, the optimal communication protocol is found to be task-dependent. Matthew Morris, Thomas D. Barrett, Arnu Pretorius |
NeurIPS | 2 |
| 2020 | Exploratory Combinatorial Optimization with Reinforcement LearningabstractMany real-world problems can be reduced to combinatorial optimization on a graph, where the subset or ordering of vertices that maximize some objective function must be found. With such tasks often NP-hard and analytically intractable, reinforcement learning (RL) has shown promise as a framework with which efficient heuristic methods to tackle these problems can be learned. Previous works construct the solution subset incrementally, adding one element at a time, however, the irreversible nature of this approach prevents the agent from revising its earlier decisions, which may be necessary given the complexity of the optimization task. We instead propose that the agent should seek to continuously improve the solution by learning to explore at test time. Our approach of exploratory combinatorial optimization (ECO-DQN) is, in principle, applicable to any combinatorial problem that can be defined on a graph. Experimentally, we show our method to produce state-of-the-art RL performance on the Maximum Cut problem. Moreover, because ECO-DQN can start from any arbitrary configuration, it can be combined with other search methods to further improve performance, which we demonstrate using a simple random search. Thomas D. Barrett, William R. Clements, Jakob N. Foerster, A. I. Lvovsky 0001 |
AAAI | 1 |
| 2020 | Learning Disentangled Representations and Group Structure of Dynamical EnvironmentsabstractLearning disentangled representations is a key step towards effectively discovering and modelling the underlying structure of environments. In the natural sciences, physics has found great success by describing the universe in terms of symmetry preserving transformations. Inspired by this formalism, we propose a framework, built upon the theory of group representation, for learning representations of a dynamical environment structured around the transformations that generate its evolution. Experimentally, we learn the structure of explicitly symmetric environments without supervision from observational data generated by sequential interactions. We further introduce an intuitive disentanglement regularisation to ensure the interpretability of the learnt representations. We show that our method enables accurate long-horizon predictions, and demonstrate a correlation between the quality of predictions and disentanglement in the latent space. Robin Quessard, Thomas D. Barrett, William R. Clements |
NeurIPS | 2 |