Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Christopher Mattern

dblp:19/10437 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 5 first-authorArtificial intelligence and machine learning · 3 · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Learning theory · 34% Deep learning architectures and training · 23% Efficient and distributed learning · 18%
Theoretical computer science
1 paper
Coding theory · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
compression
0.812024
Language Modeling Is Compression · ICLR 2024
Machine learning › Transfer learning and domain adaptation
meta-learning
0.812024
Learning Universal Predictors · ICML 2024
Machine learning › Learning theory › online learning › sequence prediction
universal prediction
0.812024
Learning Universal Predictors · ICML 2024
Machine learning › Deep learning architectures and training › neural network training
backpropagation-free training
0.512021
Gated Linear Networks · AAAI 2021
Machine learning › Deep learning architectures and training › feedforward neural network › piecewise linear network
gated linear networks
0.512021
Gated Linear Networks · AAAI 2021
Machine learning › Learning theory
online learning
0.512021
Gated Linear Networks · AAAI 2021
Coding theory › source coding
lossless compression
0.212024
Language Modeling Is Compression · ICLR 2024
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.112021
Gated Linear Networks · AAAI 2021
Machine learning › Trustworthy machine learning
robustness
0.112021
Gated Linear Networks · AAAI 2021

Methods — techniques the papers use, named apart from their topics

universal turing machine · 0.8meta-learning · 0.8online convex optimization · 0.5data-dependent gating · 0.5
YearPublicationVenuePosition
2024 Language Modeling Is Compression
abstract
It has long been established that predictive models can be transformed into lossless compressors and vice versa. Incidentally, in recent years, the machine learning community has focused on training increasingly large and powerful self-supervised (language) models. Since these large language models exhibit impressive predictive capabilities, they are well-positioned to be strong compressors. In this work, we advocate for viewing the prediction problem through the lens of compression and evaluate the compression capabilities of large (foundation) models. We show that large language models are powerful general-purpose predictors and that the compression viewpoint provides novel insights into scaling laws, tokenization, and in-context learning. For example, Chinchilla 70B, while trained primarily on text, compresses ImageNet patches to 43.4% and LibriSpeech samples to 16.4% of their raw size, beating domain-specific compressors like PNG (58.5%) or FLAC (30.3%), respectively. Finally, we show that the prediction-compression equivalence allows us to use any compressor (like gzip) to build a conditional generative model.
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, Joel Veness
ICLR6
2024 Learning Universal Predictors
abstract
Meta-learning has emerged as a powerful approach to train neural networks to learn new tasks quickly from limited data by pre-training them on a broad set of tasks. But, what are the limits of meta-learning? In this work, we explore the potential of amortizing the most powerful universal predictor, namely Solomonoff Induction (SI), into neural networks via leveraging (memory-based) meta-learning to its limits. We use Universal Turing Machines (UTMs) to generate training data used to expose networks to a broad range of patterns. We provide theoretical analysis of the UTM data generation processes and meta-training protocols. We conduct comprehensive experiments with neural architectures (e.g. LSTMs, Transformers) and algorithmic data generators of varying complexity and universality. Our results suggest that UTM data is a valuable resource for meta-learning, and that it can be used to train neural networks capable of learning universal prediction strategies.
Jordi Grau-Moya, Tim Genewein, Marcus Hutter, Laurent Orseau, Grégoire Delétang, Elliot Catt, Anian Ruoss, Li Kevin Wenliang, Christopher Mattern, Matthew Aitchison, Joel Veness
ICML9
2021 Gated Linear Networks
abstract
This paper presents a new family of backpropagation-free neural architectures, Gated Linear Networks (GLNs). What distinguishes GLNs from contemporary neural networks is the distributed and local nature of their credit assignment mechanism; each neuron directly predicts the target, forgoing the ability to learn feature representations in favor of rapid online learning. Individual neurons are able to model nonlinear functions via the use of data-dependent gating in conjunction with online convex optimization. We show that this architecture gives rise to universal learning capabilities in the limit, with effective model capacity increasing as a function of network size in a manner comparable with deep ReLU networks. Furthermore, we demonstrate that the GLN learning mechanism possesses extraordinary resilience to catastrophic forgetting, performing almost on par to an MLP with dropout and Elastic Weight Consolidation on standard benchmarks.
Joel Veness, Tor Lattimore, David Budden, Avishkar Bhoopchand, Christopher Mattern, Agnieszka Grabska-Barwinska, Eren Sezener, Peter Toth, Simon Schmitt, Marcus Hutter
AAAI5
2018 Generalized Probability Smoothing
abstract
In this work we consider a generalized version of Probability Smoothing, the core elementary model for sequential prediction in the state of the art PAQ family of data compression algorithms. Our main contribution is a code length analysis that considers the redundancy of Probability Smoothing with respect to a Piecewise Stationary Source. The analysis holds for a finite alphabet and expresses redundancy in terms of the total variation in probability mass of the stationary distributions of a Piecewise Stationary Source. By choosing parameters appropriately Probability Smoothing has redundancy O(S · √T log T) for sequences of length T with respect to a Piecewise Stationary Source with S segments.
Christopher Mattern
DCC1
2015 On Probability Estimation via Relative Frequencies and Discount
abstract
Probability estimation is an elementary building block of every statistical data compression algorithm. In practice probability estimation is often based on relative letter frequencies which get scaled down, when their sum is too large. Such algorithms are attractive in terms of memory requirements, running time and practical performance. However, there still is a lack of theoretical understanding. In this work we formulate a typical probability estimation algorithm based on relative frequencies and frequency discount, Algorithm RFD. Our main contribution is its theoretical analysis. We show that Algorithm RFD performs almost as good as any piecewise stationary model with either bounded or unbounded letter probabilities. This theoretically confirms the recency effect of periodic frequency discount, which has often been observed empirically.
Christopher Mattern
DCC1
2015 On Probability Estimation by Exponential Smoothing
abstract
Probability estimation is an elementary building block of every statistical data compression algorithm. In practice probability estimation should be adaptive, recent observations should receive a higher weight than older observations. We present a probability estimation method based on exponential smoothing that satisfies this requirement. Our main contribution is a theoretical analysis for various smoothing rate sequences: We show that the redundancy w.r.t. A piecewise stationary model with s segments is O (s n0.5) for any bit sequence of length n, an improvement over previous approaches with a similar complexity.
Christopher Mattern
DCC1
2013 Linear and Geometric Mixtures - Analysis
abstract
Linear and geometric mixtures are two methods to combine arbitrary models in data compression. Geometric mixtures generalize the empirically well-performing PAQ7 mixture. Both mixture schemes rely on weight vectors, which heavily determine their performance. Typically weight vectors are identified via Online Gradient Descent. In this work we show that one can obtain strong code length bounds for such a weight estimation scheme. These bounds hold for arbitrary input sequences. For this purpose we introduce the class of nice mixtures and analyze how Online Gradient Descent with a fixed step size combined with a nice mixture performs. These results translate to linear and geometric mixtures, which are nice, as we show. The results hold for PAQ7 mixtures as well, thus we provide the first theoretical analysis of PAQ7.
Christopher Mattern
DCC1
2012 Mixing Strategies in Data Compression
abstract
We propose geometric weighting as a novel method to combine multiple models in data compression. Our results reveal the rationale behind PAQ-weighting and generalize it to a non-binary alphabet. Based on a similar technique we present a new, generic linear mixture technique. All novel mixture techniques rely on given weight vectors. We consider the problem of finding optimal weights and show that the weight optimization leads to a strictly convex (and thus, good-natured) optimization problem. Finally, an experimental evaluation compares the two presented mixture techniques for a binary alphabet. The results indicate that geometric weighting is superior to linear weighting.
Christopher Mattern
DCC1