EDBT 2026 Demo / reviewers in the wild / expert
Christopher Mattern
dblp:19/10437
· DBLP profile ↗
8ranked-venue papers
5as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 5 first-authorArtificial intelligence and machine learning · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Learning theory · 34% Deep learning architectures and training · 23% Efficient and distributed learning · 18% | |
| Theoretical computer science
1 paper |
Coding theory · 100% |
Topics — the 9 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
compression |
0.8 | 1 | 2024 | Language Modeling Is Compression · ICLR 2024 |
Machine learning › Transfer learning and domain adaptation
meta-learning |
0.8 | 1 | 2024 | Learning Universal Predictors · ICML 2024 |
Machine learning › Learning theory › online learning › sequence prediction
universal prediction |
0.8 | 1 | 2024 | Learning Universal Predictors · ICML 2024 |
Machine learning › Deep learning architectures and training › neural network training
backpropagation-free training |
0.5 | 1 | 2021 | Gated Linear Networks · AAAI 2021 |
Machine learning › Deep learning architectures and training › feedforward neural network › piecewise linear network
gated linear networks |
0.5 | 1 | 2021 | Gated Linear Networks · AAAI 2021 |
Machine learning › Learning theory
online learning |
0.5 | 1 | 2021 | Gated Linear Networks · AAAI 2021 |
Coding theory › source coding
lossless compression |
0.2 | 1 | 2024 | Language Modeling Is Compression · ICLR 2024 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.1 | 1 | 2021 | Gated Linear Networks · AAAI 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.1 | 1 | 2021 | Gated Linear Networks · AAAI 2021 |
Methods — techniques the papers use, named apart from their topics
universal turing machine · 0.8meta-learning · 0.8online convex optimization · 0.5data-dependent gating · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Language Modeling Is CompressionabstractIt has long been established that predictive models can be transformed into lossless compressors and vice versa. Incidentally, in recent years, the machine learning community has focused on training increasingly large and powerful self-supervised (language) models. Since these large language models exhibit impressive predictive capabilities, they are well-positioned to be strong compressors. In this work, we advocate for viewing the prediction problem through the lens of compression and evaluate the compression capabilities of large (foundation) models. We show that large language models are powerful general-purpose predictors and that the compression viewpoint provides novel insights into scaling laws, tokenization, and in-context learning. For example, Chinchilla 70B, while trained primarily on text, compresses ImageNet patches to 43.4% and LibriSpeech samples to 16.4% of their raw size, beating domain-specific compressors like PNG (58.5%) or FLAC (30.3%), respectively. Finally, we show that the prediction-compression equivalence allows us to use any compressor (like gzip) to build a conditional generative model. Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, Joel Veness |
ICLR | 6 |
| 2024 | Learning Universal PredictorsabstractMeta-learning has emerged as a powerful approach to train neural networks to learn new tasks quickly from limited data by pre-training them on a broad set of tasks. But, what are the limits of meta-learning? In this work, we explore the potential of amortizing the most powerful universal predictor, namely Solomonoff Induction (SI), into neural networks via leveraging (memory-based) meta-learning to its limits. We use Universal Turing Machines (UTMs) to generate training data used to expose networks to a broad range of patterns. We provide theoretical analysis of the UTM data generation processes and meta-training protocols. We conduct comprehensive experiments with neural architectures (e.g. LSTMs, Transformers) and algorithmic data generators of varying complexity and universality. Our results suggest that UTM data is a valuable resource for meta-learning, and that it can be used to train neural networks capable of learning universal prediction strategies. Jordi Grau-Moya, Tim Genewein, Marcus Hutter, Laurent Orseau, Grégoire Delétang, Elliot Catt, Anian Ruoss, Li Kevin Wenliang, Christopher Mattern, Matthew Aitchison, Joel Veness |
ICML | 9 |
| 2021 | Gated Linear NetworksabstractThis paper presents a new family of backpropagation-free neural architectures, Gated Linear Networks (GLNs). What distinguishes GLNs from contemporary neural networks is the distributed and local nature of their credit assignment mechanism; each neuron directly predicts the target, forgoing the ability to learn feature representations in favor of rapid online learning. Individual neurons are able to model nonlinear functions via the use of data-dependent gating in conjunction with online convex optimization. We show that this architecture gives rise to universal learning capabilities in the limit, with effective model capacity increasing as a function of network size in a manner comparable with deep ReLU networks. Furthermore, we demonstrate that the GLN learning mechanism possesses extraordinary resilience to catastrophic forgetting, performing almost on par to an MLP with dropout and Elastic Weight Consolidation on standard benchmarks. Joel Veness, Tor Lattimore, David Budden, Avishkar Bhoopchand, Christopher Mattern, Agnieszka Grabska-Barwinska, Eren Sezener, Peter Toth, Simon Schmitt, Marcus Hutter |
AAAI | 5 |
| 2018 | Generalized Probability SmoothingabstractIn this work we consider a generalized version of Probability Smoothing, the core elementary model for sequential prediction in the state of the art PAQ family of data compression algorithms. Our main contribution is a code length analysis that considers the redundancy of Probability Smoothing with respect to a Piecewise Stationary Source. The analysis holds for a finite alphabet and expresses redundancy in terms of the total variation in probability mass of the stationary distributions of a Piecewise Stationary Source. By choosing parameters appropriately Probability Smoothing has redundancy O(S · √T log T) for sequences of length T with respect to a Piecewise Stationary Source with S segments. Christopher Mattern |
DCC | 1 |
| 2015 | On Probability Estimation via Relative Frequencies and DiscountabstractProbability estimation is an elementary building block of every statistical data compression algorithm. In practice probability estimation is often based on relative letter frequencies which get scaled down, when their sum is too large. Such algorithms are attractive in terms of memory requirements, running time and practical performance. However, there still is a lack of theoretical understanding. In this work we formulate a typical probability estimation algorithm based on relative frequencies and frequency discount, Algorithm RFD. Our main contribution is its theoretical analysis. We show that Algorithm RFD performs almost as good as any piecewise stationary model with either bounded or unbounded letter probabilities. This theoretically confirms the recency effect of periodic frequency discount, which has often been observed empirically. Christopher Mattern |
DCC | 1 |
| 2015 | On Probability Estimation by Exponential SmoothingabstractProbability estimation is an elementary building block of every statistical data compression algorithm. In practice probability estimation should be adaptive, recent observations should receive a higher weight than older observations. We present a probability estimation method based on exponential smoothing that satisfies this requirement. Our main contribution is a theoretical analysis for various smoothing rate sequences: We show that the redundancy w.r.t. A piecewise stationary model with s segments is O (s n0.5) for any bit sequence of length n, an improvement over previous approaches with a similar complexity. Christopher Mattern |
DCC | 1 |
| 2013 | Linear and Geometric Mixtures - AnalysisabstractLinear and geometric mixtures are two methods to combine arbitrary models in data compression. Geometric mixtures generalize the empirically well-performing PAQ7 mixture. Both mixture schemes rely on weight vectors, which heavily determine their performance. Typically weight vectors are identified via Online Gradient Descent. In this work we show that one can obtain strong code length bounds for such a weight estimation scheme. These bounds hold for arbitrary input sequences. For this purpose we introduce the class of nice mixtures and analyze how Online Gradient Descent with a fixed step size combined with a nice mixture performs. These results translate to linear and geometric mixtures, which are nice, as we show. The results hold for PAQ7 mixtures as well, thus we provide the first theoretical analysis of PAQ7. Christopher Mattern |
DCC | 1 |
| 2012 | Mixing Strategies in Data CompressionabstractWe propose geometric weighting as a novel method to combine multiple models in data compression. Our results reveal the rationale behind PAQ-weighting and generalize it to a non-binary alphabet. Based on a similar technique we present a new, generic linear mixture technique. All novel mixture techniques rely on given weight vectors. We consider the problem of finding optimal weights and show that the weight optimization leads to a strictly convex (and thus, good-natured) optimization problem. Finally, an experimental evaluation compares the two presented mixture techniques for a binary alphabet. The results indicate that geometric weighting is superior to linear weighting. Christopher Mattern |
DCC | 1 |