Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Matthew Aitchison

dblp:235/5682 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Reinforcement learning · 66% Learning theory · 10% Efficient and distributed learning · 8%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › dynamic programming
bellman operator
0.812024
Distributional Bellman Operators over Mean Embeddings · ICML 2024
Machine learning › Efficient and distributed learning
compression
0.812024
Language Modeling Is Compression · ICLR 2024
Machine learning › Reinforcement learning › value-based reinforcement learning
distributional reinforcement learning
0.812024
Distributional Bellman Operators over Mean Embeddings · ICML 2024
Machine learning › Kernel, tree and ensemble methods › kernel embedding
mean embedding
0.812024
Distributional Bellman Operators over Mean Embeddings · ICML 2024
Machine learning › Transfer learning and domain adaptation
meta-learning
0.812024
Learning Universal Predictors · ICML 2024
Machine learning › Learning theory › online learning › sequence prediction
universal prediction
0.812024
Learning Universal Predictors · ICML 2024
Machine learning › Reinforcement learning
value function approximation
0.812024
Distributional Bellman Operators over Mean Embeddings · ICML 2024
Machine learning › Reinforcement learning
benchmark design
0.712023
Atari-5: Distilling the Arcade Learning Environment down to Five Games · ICML 2023
Machine learning › Reinforcement learning
planning and learning
0.712023
Self-Predictive Universal AI · NeurIPS 2023
Performance modeling and evaluation
benchmarking
0.712023
Atari-5: Distilling the Arcade Learning Environment down to Five Games · ICML 2023
Machine learning › Reinforcement learning
actor-critic methods
0.612022
DNA: Proximal Policy Optimization with a Dual Network Architecture · NeurIPS 2022
Machine learning › Reinforcement learning › policy optimization
proximal policy optimization
0.612022
DNA: Proximal Policy Optimization with a Dual Network Architecture · NeurIPS 2022
Machine learning › Reinforcement learning
value function estimation
0.612022
DNA: Proximal Policy Optimization with a Dual Network Architecture · NeurIPS 2022
Coding theory › source coding
lossless compression
0.212024
Language Modeling Is Compression · ICLR 2024
Machine learning › Reinforcement learning
temporal difference learning
0.212023
Self-Predictive Universal AI · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

subset selection · 1.3correlation analysis · 1.3universal turing machine · 0.8temporal difference learning · 0.8sketch bellman operator · 0.8meta-learning · 0.8universal mixture-policy · 0.7self-prediction · 0.7q-value estimation · 0.7dual-network architecture · 0.6
YearPublicationVenuePosition
2024 Language Modeling Is Compression
abstract
It has long been established that predictive models can be transformed into lossless compressors and vice versa. Incidentally, in recent years, the machine learning community has focused on training increasingly large and powerful self-supervised (language) models. Since these large language models exhibit impressive predictive capabilities, they are well-positioned to be strong compressors. In this work, we advocate for viewing the prediction problem through the lens of compression and evaluate the compression capabilities of large (foundation) models. We show that large language models are powerful general-purpose predictors and that the compression viewpoint provides novel insights into scaling laws, tokenization, and in-context learning. For example, Chinchilla 70B, while trained primarily on text, compresses ImageNet patches to 43.4% and LibriSpeech samples to 16.4% of their raw size, beating domain-specific compressors like PNG (58.5%) or FLAC (30.3%), respectively. Finally, we show that the prediction-compression equivalence allows us to use any compressor (like gzip) to build a conditional generative model.
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, Joel Veness
ICLR9
2024 Learning Universal Predictors
abstract
Meta-learning has emerged as a powerful approach to train neural networks to learn new tasks quickly from limited data by pre-training them on a broad set of tasks. But, what are the limits of meta-learning? In this work, we explore the potential of amortizing the most powerful universal predictor, namely Solomonoff Induction (SI), into neural networks via leveraging (memory-based) meta-learning to its limits. We use Universal Turing Machines (UTMs) to generate training data used to expose networks to a broad range of patterns. We provide theoretical analysis of the UTM data generation processes and meta-training protocols. We conduct comprehensive experiments with neural architectures (e.g. LSTMs, Transformers) and algorithmic data generators of varying complexity and universality. Our results suggest that UTM data is a valuable resource for meta-learning, and that it can be used to train neural networks capable of learning universal prediction strategies.
Jordi Grau-Moya, Tim Genewein, Marcus Hutter, Laurent Orseau, Grégoire Delétang, Elliot Catt, Anian Ruoss, Li Kevin Wenliang, Christopher Mattern, Matthew Aitchison, Joel Veness
ICML10
2024 Distributional Bellman Operators over Mean Embeddings
abstract
We propose a novel algorithmic framework for distributional reinforcement learning, based on learning finite-dimensional mean embeddings of return distributions. The framework reveals a wide variety of new algorithms for dynamic programming and temporal-difference algorithms that rely on the sketch Bellman operator, which updates mean embeddings with simple linear-algebraic computations. We provide asymptotic convergence theory, and examine the empirical performance of the algorithms on a suite of tabular tasks. Further, we show that this approach can be straightforwardly combined with deep reinforcement learning.
Li Kevin Wenliang, Grégoire Delétang, Matthew Aitchison, Marcus Hutter, Anian Ruoss, Arthur Gretton, Mark Rowland 0001
ICML3
2023 HiveMind: Learning to Play the Cooperative Chess Variant Bughouse with DNNs and MCTS
abstract
In 2017, the AlphaZero algorithm achieved superhuman performance in chess, outperforming the then-reigning chess engine Stockfish 8. AlphaZero has since been applied to other variants of chess, such as Crazyhouse, with similarly impressive results. However, limited work has been done on the chess variant Bughouse, which has both cooperative and real-time aspects, as well as a far higher game tree complexity than chess. In this paper, we present HiveMind, a neural network Bughouse engine that focuses on cooperation and decision-making. We trained HiveMind via supervised learning on human Bughouse games. We then used the AlphaZero algorithm by incorporating domain knowledge, using clock times to determine the optimal turn sequence, to perform a tree search over both boards. This two-board search incorporated all aspects of Bughouse, including being time aware and capable of cross-board coordination without heuristics. Finally, we evaluated the strength of HiveMind by playing matches with different time settings against Fairy-Stockfish, the current state-of-the-art alpha-beta Bughouse engine. HiveMind convincingly defeated Fairy-Stockfish achieving a win rate of over 95% with a search time of 2 seconds, showing significantly better scaling.
Benjamin Woo, Penny Kyburz, Matthew Aitchison
CoG3
2023 Atari-5: Distilling the Arcade Learning Environment down to Five Games
abstract
The Arcade Learning Environment (ALE) has become an essential benchmark for assessing the performance of reinforcement learning algorithms. However, the computational cost of generating results on the entire 57-game dataset limits ALE’s use and makes the reproducibility of many results infeasible. We propose a novel solution to this problem in the form of a principled methodology for selecting small but representative subsets of environments within a benchmark suite. We applied our method to identify a subset of five ALE games, we call Atari-5, which produces 57-game median score estimates within 10% of their true values. Extending the subset to 10-games recovers 80% of the variance for log-scores for all games within the 57-game set. We show this level of compression is possible due to a high degree of correlation between many of the games in ALE.
Matthew Aitchison, Penny Kyburz, Marcus Hutter
ICML1
2023 Self-Predictive Universal AI
abstract
Reinforcement Learning (RL) algorithms typically utilize learning and/or planning techniques to derive effective policies. The integration of both approaches has proven to be highly successful in addressing complex sequential decision-making challenges, as evidenced by algorithms such as AlphaZero and MuZero, which consolidate the planning process into a parametric search-policy. AIXI, the most potent theoretical universal agent, leverages planning through comprehensive search as its primary means to find an optimal policy. Here we define an alternative universal agent, which we call Self-AIXI, that on the contrary to AIXI, maximally exploits learning to obtain good policies. It does so by self-predicting its own stream of action data, which is generated, similarly to other TD(0) agents, by taking an action maximization step over the current on-policy (universal mixture-policy) Q-value estimates. We prove that Self-AIXI converges to AIXI, and inherits a series of properties like maximal Legg-Hutter intelligence and the self-optimizing property.
Elliot Catt, Jordi Grau-Moya, Marcus Hutter, Matthew Aitchison, Tim Genewein, Grégoire Delétang, Joel Veness
NeurIPS4
2022 DNA: Proximal Policy Optimization with a Dual Network Architecture
abstract
This paper explores the problem of simultaneously learning a value function and policy in deep actor-critic reinforcement learning models. We find that the common practice of learning these functions jointly is sub-optimal due to an order-of-magnitude difference in noise levels between the two tasks. Instead, we show that learning these tasks independently, but with a constrained distillation phase, significantly improves performance. Furthermore, we find that policy gradient noise levels decrease when using a lower \textit{variance} return estimate. Whereas, value learning noise level decreases with a lower \textit{bias} estimate. Together these insights inform an extension to Proximal Policy Optimization we call \textit{Dual Network Architecture} (DNA), which significantly outperforms its predecessor. DNA also exceeds the performance of the popular Rainbow DQN algorithm on four of the five environments tested, even under more difficult stochastic control settings.
Matthew Aitchison, Penny Kyburz
NeurIPS1
2022 An Agile New Research Framework for Hybrid Human-AI Teaming: Trust, Transparency, and Transferability
abstract
We propose a new research framework by which the nascent discipline of human-AI teaming can be explored within experimental environments in preparation for transferal to real-world contexts. We examine the existing literature and unanswered research questions through the lens of an Agile approach to construct our proposed framework. Our framework aims to provide a structure for understanding the macro features of this research landscape, supporting holistic research into the acceptability of human-AI teaming to human team members and the affordances of AI team members. The framework has the potential to enhance decision-making and performance of hybrid human-AI teams. Further, our framework proposes the application of Agile methodology for research management and knowledge discovery. We propose a transferability pathway for hybrid teaming to be initially tested in a safe environment, such as a real-time strategy video game, with elements of lessons learned that can be transferred to real-world situations.
Sabrina B. Caldwell, Penny Kyburz, Nicholas O'Donnell, Matthew James Knight, Matthew Aitchison, Tom Gedeon, Daniel Johnson 0001, Margot Brereton, Marcus Gallagher, David Conroy
ACM Trans. Interact. Intell. Syst.5
2020 Do Game Bots Dream of Electric Rewards?: The universality of intrinsic motivation
abstract
The purpose of this paper is to draw together theories, ideas, and observations related to rewards, motivation, and play to develop and question our understanding and practice of designing reward-based systems and technology. Our exploration includes reinforcement, rewards, motivational theory, flow, play, games, gamification, and machine learning. We examine the design and psychology of reward-based systems in society and technology, using gamification and machine learning as case studies. We propose that the problems that exist with reward-based systems in our society are also present and pertinent when designing technology. We suggest that motivation, exploration, and play are not just fundamental to human learning and behaviour, but that they could transcend nature into machine learning. Finally, we question the value and potential harm of the reward-based systems that permeate every aspect of our lives and assert the importance of ethics in the design of all systems and technology.
Penny Kyburz, Matthew Aitchison
FDG2
2019 Optimal Use of Experience in First Person Shooter Environments
abstract
Although reinforcement learning has made great strides recently, a continuing limitation is that it requires an extremely high number of interactions with the environment. In this paper, we explore the effectiveness of reusing experience from the experience replay buffer in the Deep Q-Learning algorithm. We test the effectiveness of applying learning update steps multiple times per environmental step in the VizDoom environment and show first, this requires a change in the learning rate, and second that it does not improve the performance of the agent. Furthermore, we show that updating less frequently is effective up to a ratio of 4:1, after which performance degrades significantly. These results quantitatively confirm the widespread practice of performing learning updates every 4th environmental step.
Matthew Aitchison
CoG1