EDBT 2026 Demo / reviewers in the wild / expert
Matthew Aitchison
dblp:235/5682
· DBLP profile ↗
10ranked-venue papers
3as first author
8since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 66% Learning theory · 10% Efficient and distributed learning · 8% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 100% |
Topics — the 15 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › dynamic programming
bellman operator |
0.8 | 1 | 2024 | Distributional Bellman Operators over Mean Embeddings · ICML 2024 |
Machine learning › Efficient and distributed learning
compression |
0.8 | 1 | 2024 | Language Modeling Is Compression · ICLR 2024 |
Machine learning › Reinforcement learning › value-based reinforcement learning
distributional reinforcement learning |
0.8 | 1 | 2024 | Distributional Bellman Operators over Mean Embeddings · ICML 2024 |
Machine learning › Kernel, tree and ensemble methods › kernel embedding
mean embedding |
0.8 | 1 | 2024 | Distributional Bellman Operators over Mean Embeddings · ICML 2024 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.8 | 1 | 2024 | Learning Universal Predictors · ICML 2024 |
Machine learning › Learning theory › online learning › sequence prediction
universal prediction |
0.8 | 1 | 2024 | Learning Universal Predictors · ICML 2024 |
Machine learning › Reinforcement learning
value function approximation |
0.8 | 1 | 2024 | Distributional Bellman Operators over Mean Embeddings · ICML 2024 |
Machine learning › Reinforcement learning
benchmark design |
0.7 | 1 | 2023 | Atari-5: Distilling the Arcade Learning Environment down to Five Games · ICML 2023 |
Machine learning › Reinforcement learning
planning and learning |
0.7 | 1 | 2023 | Self-Predictive Universal AI · NeurIPS 2023 |
Performance modeling and evaluation
benchmarking |
0.7 | 1 | 2023 | Atari-5: Distilling the Arcade Learning Environment down to Five Games · ICML 2023 |
Machine learning › Reinforcement learning
actor-critic methods |
0.6 | 1 | 2022 | DNA: Proximal Policy Optimization with a Dual Network Architecture · NeurIPS 2022 |
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization |
0.6 | 1 | 2022 | DNA: Proximal Policy Optimization with a Dual Network Architecture · NeurIPS 2022 |
Machine learning › Reinforcement learning
value function estimation |
0.6 | 1 | 2022 | DNA: Proximal Policy Optimization with a Dual Network Architecture · NeurIPS 2022 |
Coding theory › source coding
lossless compression |
0.2 | 1 | 2024 | Language Modeling Is Compression · ICLR 2024 |
Machine learning › Reinforcement learning
temporal difference learning |
0.2 | 1 | 2023 | Self-Predictive Universal AI · NeurIPS 2023 |
Methods — techniques the papers use, named apart from their topics
subset selection · 1.3correlation analysis · 1.3universal turing machine · 0.8temporal difference learning · 0.8sketch bellman operator · 0.8meta-learning · 0.8universal mixture-policy · 0.7self-prediction · 0.7q-value estimation · 0.7dual-network architecture · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Language Modeling Is CompressionabstractIt has long been established that predictive models can be transformed into lossless compressors and vice versa. Incidentally, in recent years, the machine learning community has focused on training increasingly large and powerful self-supervised (language) models. Since these large language models exhibit impressive predictive capabilities, they are well-positioned to be strong compressors. In this work, we advocate for viewing the prediction problem through the lens of compression and evaluate the compression capabilities of large (foundation) models. We show that large language models are powerful general-purpose predictors and that the compression viewpoint provides novel insights into scaling laws, tokenization, and in-context learning. For example, Chinchilla 70B, while trained primarily on text, compresses ImageNet patches to 43.4% and LibriSpeech samples to 16.4% of their raw size, beating domain-specific compressors like PNG (58.5%) or FLAC (30.3%), respectively. Finally, we show that the prediction-compression equivalence allows us to use any compressor (like gzip) to build a conditional generative model. Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, Joel Veness |
ICLR | 9 |
| 2024 | Learning Universal PredictorsabstractMeta-learning has emerged as a powerful approach to train neural networks to learn new tasks quickly from limited data by pre-training them on a broad set of tasks. But, what are the limits of meta-learning? In this work, we explore the potential of amortizing the most powerful universal predictor, namely Solomonoff Induction (SI), into neural networks via leveraging (memory-based) meta-learning to its limits. We use Universal Turing Machines (UTMs) to generate training data used to expose networks to a broad range of patterns. We provide theoretical analysis of the UTM data generation processes and meta-training protocols. We conduct comprehensive experiments with neural architectures (e.g. LSTMs, Transformers) and algorithmic data generators of varying complexity and universality. Our results suggest that UTM data is a valuable resource for meta-learning, and that it can be used to train neural networks capable of learning universal prediction strategies. Jordi Grau-Moya, Tim Genewein, Marcus Hutter, Laurent Orseau, Grégoire Delétang, Elliot Catt, Anian Ruoss, Li Kevin Wenliang, Christopher Mattern, Matthew Aitchison, Joel Veness |
ICML | 10 |
| 2024 | Distributional Bellman Operators over Mean EmbeddingsabstractWe propose a novel algorithmic framework for distributional reinforcement learning, based on learning finite-dimensional mean embeddings of return distributions. The framework reveals a wide variety of new algorithms for dynamic programming and temporal-difference algorithms that rely on the sketch Bellman operator, which updates mean embeddings with simple linear-algebraic computations. We provide asymptotic convergence theory, and examine the empirical performance of the algorithms on a suite of tabular tasks. Further, we show that this approach can be straightforwardly combined with deep reinforcement learning. Li Kevin Wenliang, Grégoire Delétang, Matthew Aitchison, Marcus Hutter, Anian Ruoss, Arthur Gretton, Mark Rowland 0001 |
ICML | 3 |
| 2023 | HiveMind: Learning to Play the Cooperative Chess Variant Bughouse with DNNs and MCTSabstractIn 2017, the AlphaZero algorithm achieved superhuman performance in chess, outperforming the then-reigning chess engine Stockfish 8. AlphaZero has since been applied to other variants of chess, such as Crazyhouse, with similarly impressive results. However, limited work has been done on the chess variant Bughouse, which has both cooperative and real-time aspects, as well as a far higher game tree complexity than chess. In this paper, we present HiveMind, a neural network Bughouse engine that focuses on cooperation and decision-making. We trained HiveMind via supervised learning on human Bughouse games. We then used the AlphaZero algorithm by incorporating domain knowledge, using clock times to determine the optimal turn sequence, to perform a tree search over both boards. This two-board search incorporated all aspects of Bughouse, including being time aware and capable of cross-board coordination without heuristics. Finally, we evaluated the strength of HiveMind by playing matches with different time settings against Fairy-Stockfish, the current state-of-the-art alpha-beta Bughouse engine. HiveMind convincingly defeated Fairy-Stockfish achieving a win rate of over 95% with a search time of 2 seconds, showing significantly better scaling. Benjamin Woo, Penny Kyburz, Matthew Aitchison |
CoG | 3 |
| 2023 | Atari-5: Distilling the Arcade Learning Environment down to Five GamesabstractThe Arcade Learning Environment (ALE) has become an essential benchmark for assessing the performance of reinforcement learning algorithms. However, the computational cost of generating results on the entire 57-game dataset limits ALE’s use and makes the reproducibility of many results infeasible. We propose a novel solution to this problem in the form of a principled methodology for selecting small but representative subsets of environments within a benchmark suite. We applied our method to identify a subset of five ALE games, we call Atari-5, which produces 57-game median score estimates within 10% of their true values. Extending the subset to 10-games recovers 80% of the variance for log-scores for all games within the 57-game set. We show this level of compression is possible due to a high degree of correlation between many of the games in ALE. Matthew Aitchison, Penny Kyburz, Marcus Hutter |
ICML | 1 |
| 2023 | Self-Predictive Universal AIabstractReinforcement Learning (RL) algorithms typically utilize learning and/or planning techniques to derive effective policies. The integration of both approaches has proven to be highly successful in addressing complex sequential decision-making challenges, as evidenced by algorithms such as AlphaZero and MuZero, which consolidate the planning process into a parametric search-policy. AIXI, the most potent theoretical universal agent, leverages planning through comprehensive search as its primary means to find an optimal policy. Here we define an alternative universal agent, which we call Self-AIXI, that on the contrary to AIXI, maximally exploits learning to obtain good policies. It does so by self-predicting its own stream of action data, which is generated, similarly to other TD(0) agents, by taking an action maximization step over the current on-policy (universal mixture-policy) Q-value estimates. We prove that Self-AIXI converges to AIXI, and inherits a series of properties like maximal Legg-Hutter intelligence and the self-optimizing property. Elliot Catt, Jordi Grau-Moya, Marcus Hutter, Matthew Aitchison, Tim Genewein, Grégoire Delétang, Joel Veness |
NeurIPS | 4 |
| 2022 | DNA: Proximal Policy Optimization with a Dual Network ArchitectureabstractThis paper explores the problem of simultaneously learning a value function and policy in deep actor-critic reinforcement learning models. We find that the common practice of learning these functions jointly is sub-optimal due to an order-of-magnitude difference in noise levels between the two tasks. Instead, we show that learning these tasks independently, but with a constrained distillation phase, significantly improves performance. Furthermore, we find that policy gradient noise levels decrease when using a lower \textit{variance} return estimate. Whereas, value learning noise level decreases with a lower \textit{bias} estimate. Together these insights inform an extension to Proximal Policy Optimization we call \textit{Dual Network Architecture} (DNA), which significantly outperforms its predecessor. DNA also exceeds the performance of the popular Rainbow DQN algorithm on four of the five environments tested, even under more difficult stochastic control settings. Matthew Aitchison, Penny Kyburz |
NeurIPS | 1 |
| 2022 | An Agile New Research Framework for Hybrid Human-AI Teaming: Trust, Transparency, and TransferabilityabstractWe propose a new research framework by which the nascent discipline of human-AI teaming can be explored within experimental environments in preparation for transferal to real-world contexts. We examine the existing literature and unanswered research questions through the lens of an Agile approach to construct our proposed framework. Our framework aims to provide a structure for understanding the macro features of this research landscape, supporting holistic research into the acceptability of human-AI teaming to human team members and the affordances of AI team members. The framework has the potential to enhance decision-making and performance of hybrid human-AI teams. Further, our framework proposes the application of Agile methodology for research management and knowledge discovery. We propose a transferability pathway for hybrid teaming to be initially tested in a safe environment, such as a real-time strategy video game, with elements of lessons learned that can be transferred to real-world situations. Sabrina B. Caldwell, Penny Kyburz, Nicholas O'Donnell, Matthew James Knight, Matthew Aitchison, Tom Gedeon, Daniel Johnson 0001, Margot Brereton, Marcus Gallagher, David Conroy |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2020 | Do Game Bots Dream of Electric Rewards?: The universality of intrinsic motivationabstractThe purpose of this paper is to draw together theories, ideas, and observations related to rewards, motivation, and play to develop and question our understanding and practice of designing reward-based systems and technology. Our exploration includes reinforcement, rewards, motivational theory, flow, play, games, gamification, and machine learning. We examine the design and psychology of reward-based systems in society and technology, using gamification and machine learning as case studies. We propose that the problems that exist with reward-based systems in our society are also present and pertinent when designing technology. We suggest that motivation, exploration, and play are not just fundamental to human learning and behaviour, but that they could transcend nature into machine learning. Finally, we question the value and potential harm of the reward-based systems that permeate every aspect of our lives and assert the importance of ethics in the design of all systems and technology. Penny Kyburz, Matthew Aitchison |
FDG | 2 |
| 2019 | Optimal Use of Experience in First Person Shooter EnvironmentsabstractAlthough reinforcement learning has made great strides recently, a continuing limitation is that it requires an extremely high number of interactions with the environment. In this paper, we explore the effectiveness of reusing experience from the experience replay buffer in the Deep Q-Learning algorithm. We test the effectiveness of applying learning update steps multiple times per environmental step in the VizDoom environment and show first, this requires a change in the learning rate, and second that it does not improve the performance of the agent. Furthermore, we show that updating less frequently is effective up to a ratio of 4:1, after which performance degrades significantly. These results quantitatively confirm the widespread practice of performing learning updates every 4th environmental step. Matthew Aitchison |
CoG | 1 |