VLDB 2026 Research / reviewers in the wild / expert
Andreas Kirsch 0002
dblp:56/2914-2
· DBLP profile ↗
6ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0001-8244-7700ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Efficient and distributed learning · 43% Trustworthy machine learning · 22% Probabilistic and Bayesian machine learning · 17% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
data-efficient learning |
1.3 | 2 | 2024 | CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training · NeurIPS 2024 Prioritized Training on Points that are Learnable, Worth Learning, and not yet Learnt · ICML 2022 |
Machine learning › Efficient and distributed learning
data selection |
1.3 | 2 | 2024 | CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training · NeurIPS 2024 Prioritized Training on Points that are Learnable, Worth Learning, and not yet Learnt · ICML 2022 |
Machine learning › Probabilistic and Bayesian machine learning › experimental design › bayesian experimental design
bayesian active learning |
0.9 | 2 | 2021 | Causal-BALD: Deep Bayesian Active Learning of Outcomes to Infer Treatment-Effects from Observational Data · NeurIPS 2021 BatchBALD: Efficient and Diverse Batch Acquisition for Deep Bayesian Active Learning · NeurIPS 2019 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.8 | 2 | 2023 | Deep Deterministic Uncertainty: A New Simple Baseline · CVPR 2023 Causal-BALD: Deep Bayesian Active Learning of Outcomes to Infer Treatment-Effects from Observational Data · NeurIPS 2021 |
Machine learning › Probabilistic and Bayesian machine learning
data filtering |
0.8 | 1 | 2024 | CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training · NeurIPS 2024 |
Natural language and speech › Language models and text generation › large language model training
language model pretraining |
0.8 | 1 | 2024 | CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training · NeurIPS 2024 |
Machine learning › Trustworthy machine learning › uncertainty estimation
epistemic uncertainty |
0.7 | 1 | 2023 | Deep Deterministic Uncertainty: A New Simple Baseline · CVPR 2023 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.7 | 1 | 2023 | Deep Deterministic Uncertainty: A New Simple Baseline · CVPR 2023 |
Machine learning › Efficient and distributed learning › efficient training
training acceleration |
0.6 | 1 | 2022 | Prioritized Training on Points that are Learnable, Worth Learning, and not yet Learnt · ICML 2022 |
Machine learning › Efficient and distributed learning
active learning |
0.5 | 1 | 2021 | Causal-BALD: Deep Bayesian Active Learning of Outcomes to Infer Treatment-Effects from Observational Data · NeurIPS 2021 |
Computational social science and digital humanities
causal inference |
0.5 | 1 | 2021 | Causal-BALD: Deep Bayesian Active Learning of Outcomes to Infer Treatment-Effects from Observational Data · NeurIPS 2021 |
Machine learning › Optimization for machine learning › black-box optimization
batch acquisition |
0.4 | 1 | 2019 | BatchBALD: Efficient and Diverse Batch Acquisition for Deep Bayesian Active Learning · NeurIPS 2019 |
Machine learning › Efficient and distributed learning › active learning › deep active learning
deep bayesian active learning |
0.4 | 1 | 2019 | BatchBALD: Efficient and Diverse Batch Acquisition for Deep Bayesian Active Learning · NeurIPS 2019 |
Machine learning › Transfer learning and domain adaptation
domain adaptation |
0.2 | 1 | 2024 | CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › normalization
spectral normalization |
0.2 | 1 | 2023 | Deep Deterministic Uncertainty: A New Simple Baseline · CVPR 2023 |
Machine learning › Optimization for machine learning › model-based optimization › bayesian optimization
information-theoretic acquisition function |
0.1 | 1 | 2021 | Causal-BALD: Deep Bayesian Active Learning of Outcomes to Infer Treatment-Effects from Observational Data · NeurIPS 2021 |
Methods — techniques the papers use, named apart from their topics
information theory · 1.0bayesian acquisition function · 1.0empirical bayes · 0.8auxiliary model loss · 0.8spectral normalization · 0.7residual connections · 0.7gaussian discriminant analysis · 0.7importance sampling · 0.6holdout loss · 0.6curriculum learning · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-trainingabstractSelecting high-quality data for pre-training is crucial in shaping the downstream task performance of language models. A major challenge lies in identifying this optimal subset, a problem generally considered intractable, thus necessitating scalable and effective heuristics. In this work, we propose a data selection method, CoLoR-Filter (Conditional Loss Reduction Filtering), which leverages an empirical Bayes-inspired approach to derive a simple and computationally efficient selection criterion based on the relative loss values of two auxiliary models.
In addition to the modeling rationale, we evaluate CoLoR-Filter empirically on two language modeling tasks: (1) selecting data from C4 for domain adaptation to evaluation on Books and (2) selecting data from C4 for a suite of downstream multiple-choice question answering tasks. We demonstrate favorable scaling both as we subselect more aggressively and using small auxiliary models to select data for large target models. As one headline result, CoLoR-Filter data selected using a pair of 150m parameter auxiliary models can train a 1.2b parameter target model to match a 1.2b parameter model trained on 25b randomly selected tokens with 25x less data for Books and 11x less data for the downstream tasks.
Code: https://github.com/davidbrandfonbrener/color-filter-olmo
Filtered data: https://huggingface.co/datasets/davidbrandfonbrener/color-filtered-c4 David Brandfonbrener, Hanlin Zhang 0002, Andreas Kirsch 0002, Jonathan Schwarz, Sham M. Kakade |
NeurIPS | 3 |
| 2023 | Prediction-Oriented Bayesian Active Learning
Freddie Bickford Smith, Andreas Kirsch 0002, Sebastian Farquhar, Yarin Gal, Adam Foster 0001, Tom Rainforth |
AISTATS | 2 |
| 2023 | Deep Deterministic Uncertainty: A New Simple BaselineabstractReliable uncertainty from deterministic single-forward pass models is sought after because conventional methods of uncertainty quantification are computationally expensive. We take two complex single-forward-pass uncertainty approaches, DUQ and SNGP, and examine whether they mainly rely on a well-regularized feature space. Crucially, without using their more complex methods for estimating uncertainty, we find that a single softmax neural net with such a regularized feature-space, achieved via residual connections and spectral normalization, outperforms DUQ and SNGP's epistemic uncertainty predictions using simple Gaussian Discriminant Analysis post-training as a separate feature-space density estimator-without fine-tuning on OoD data, feature ensembling, or input pre-procressing. Our conceptually simple Deep Deterministic Uncertainty (DDU) baseline can also be used to disentangle aleatoric and epistemic uncertainty and performs as well as Deep Ensembles, the state-of-the art for uncertainty prediction, on several OoD bench-marks (CIFAR-10/100 vs SVHN/Tiny-ImageNet, ImageNet vs ImageNet-O), active learning settings across different model architectures, as well as in large scale vision tasks like semantic segmentation, while being computationally cheaper. Jishnu Mukhoti, Andreas Kirsch 0002, Joost van Amersfoort, Philip Torr 0001, Yarin Gal |
CVPR | 2 |
| 2022 | Prioritized Training on Points that are Learnable, Worth Learning, and not yet LearntabstractTraining on web-scale data can take months. But much computation and time is wasted on redundant and noisy points that are already learnt or not learnable. To accelerate training, we introduce Reducible Holdout Loss Selection (RHO-LOSS), a simple but principled technique which selects approximately those points for training that most reduce the model’s generalization loss. As a result, RHO-LOSS mitigates the weaknesses of existing data selection methods: techniques from the optimization literature typically select "hard" (e.g. high loss) points, but such points are often noisy (not learnable) or less task-relevant. Conversely, curriculum learning prioritizes "easy" points, but such points need not be trained on once learned. In contrast, RHO-LOSS selects points that are learnable, worth learning, and not yet learnt. RHO-LOSS trains in far fewer steps than prior art, improves accuracy, and speeds up training on a wide range of datasets, hyperparameters, and architectures (MLPs, CNNs, and BERT). On the large web-scraped image dataset Clothing-1M, RHO-LOSS trains in 18x fewer steps and reaches 2% higher final accuracy than uniform data shuffling. Sören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma, Andreas Kirsch 0002, Winnie Xu, Benedikt Höltgen, Aidan N. Gomez, Adrien Morisot, Sebastian Farquhar, Yarin Gal |
ICML | 5 |
| 2021 | Causal-BALD: Deep Bayesian Active Learning of Outcomes to Infer Treatment-Effects from Observational DataabstractEstimating personalized treatment effects from high-dimensional observational data is essential in situations where experimental designs are infeasible, unethical, or expensive. Existing approaches rely on fitting deep models on outcomes observed for treated and control populations. However, when measuring individual outcomes is costly, as is the case of a tumor biopsy, a sample-efficient strategy for acquiring each result is required. Deep Bayesian active learning provides a framework for efficient data acquisition by selecting points with high uncertainty. However, existing methods bias training data acquisition towards regions of non-overlapping support between the treated and control populations. These are not sample-efficient because the treatment effect is not identifiable in such regions. We introduce causal, Bayesian acquisition functions grounded in information theory that bias data acquisition towards regions with overlapping support to maximize sample efficiency for learning personalized treatment effects. We demonstrate the performance of the proposed acquisition strategies on synthetic and semi-synthetic datasets IHDP and CMNIST and their extensions, which aim to simulate common dataset biases and pathologies. Andrew Jesson, Panagiotis Tigas, Joost van Amersfoort, Andreas Kirsch 0002, Uri Shalit, Yarin Gal |
NeurIPS | 4 |
| 2019 | BatchBALD: Efficient and Diverse Batch Acquisition for Deep Bayesian Active LearningabstractWe develop BatchBALD, a tractable approximation to the mutual information between a batch of points and model parameters, which we use as an acquisition function to select multiple informative points jointly for the task of deep Bayesian active learning. BatchBALD is a greedy linear-time $1 - \nicefrac{1}{e}$-approximate algorithm amenable to dynamic programming and efficient caching. We compare BatchBALD to the commonly used approach for batch data acquisition and find that the current approach acquires similar and redundant points, sometimes performing worse than randomly acquiring data. We finish by showing that, using BatchBALD to consider dependencies within an acquisition batch, we achieve new state of the art performance on standard benchmarks, providing substantial data efficiency improvements in batch acquisition. Andreas Kirsch 0002, Joost van Amersfoort, Yarin Gal |
NeurIPS | 1 |