Jörg Bornschein

dblp:13/8510 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
7since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
13 papers
Probabilistic and Bayesian machine learning · 19% Learning theory · 18% Deep learning architectures and training · 18%

Topics — the 28 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
model selection
1.122023
Evaluating Representations with Readout Model Switching · ICLR 2023
Small Data, Big Decisions: Model Selection in the Small-Data Regime · ICML 2020
Machine learning › Generative modeling
diffusion model
0.812024
Denoising Autoregressive Representation Learning · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian filtering
kalman filtering
0.812024
Kalman Filter for Online Classification of Non-Stationary Data · ICLR 2024
Machine learning › Probabilistic and Bayesian machine learning
non-stationary data
0.812024
Kalman Filter for Online Classification of Non-Stationary Data · ICLR 2024
Machine learning › Learning paradigms › continual learning
online continual learning
0.812024
Kalman Filter for Online Classification of Non-Stationary Data · ICLR 2024
Machine learning › Learning theory
online learning
0.812024
Kalman Filter for Online Classification of Non-Stationary Data · ICLR 2024
Machine learning › Deep learning architectures and training
state space model
0.812024
Kalman Filter for Online Classification of Non-Stationary Data · ICLR 2024
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.812024
Imitating Language via Scalable Inverse Reinforcement Learning · NeurIPS 2024
Machine learning › Deep learning architectures and training › transformer
transformer decoder
0.812024
Denoising Autoregressive Representation Learning · ICML 2024
Machine learning › Representation and self-supervised learning › representation learning
visual representation learning
0.812024
Denoising Autoregressive Representation Learning · ICML 2024
Machine learning › Probabilistic and Bayesian machine learning › causal inference
causal discovery
0.712023
Learning to Induce Causal Structure · ICLR 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning
causal reasoning
0.712023
Learning to Induce Causal Structure · ICLR 2023
Machine learning › Learning paradigms
continual learning
0.712023
Nevis'22: A Stream of 100 Tasks Sampled from 30 Years of Computer Vision Research · J. Mach. Learn. Res. 2023
Machine learning › Transfer learning and domain adaptation
meta-learning
0.712023
Nevis'22: A Stream of 100 Tasks Sampled from 30 Years of Computer Vision Research · J. Mach. Learn. Res. 2023
Machine learning › Learning theory › model selection
minimum description length
0.712023
Sequential Learning of Neural Networks for Prequential MDL · ICLR 2023
Machine learning › Representation and self-supervised learning › representation analysis
representation evaluation
0.712023
Evaluating Representations with Readout Model Switching · ICLR 2023
Machine learning › Deep learning architectures and training
sequence modeling
0.712023
Sequential Learning of Neural Networks for Prequential MDL · ICLR 2023
Machine learning › Generative modeling
variational autoencoder
0.522017
Variational Memory Addressing in Generative Models · NIPS 2017
Bidirectional Helmholtz Machines · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.532016
Bidirectional Helmholtz Machines · ICML 2016
Why MCA? Nonlinear sparse coding with spike-and-slab prior for neurally plausible image encoding · NIPS 2012
Select and Sample - A Model of Efficient Neural Inference and Learning · NIPS 2011
Machine learning › Learning theory
generalization
0.412020
Small Data, Big Decisions: Model Selection in the Small-Data Regime · ICML 2020
Machine learning › Deep learning architectures and training
small-data regime
0.412020
Small Data, Big Decisions: Model Selection in the Small-Data Regime · ICML 2020
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.432012
Why MCA? Nonlinear sparse coding with spike-and-slab prior for neurally plausible image encoding · NIPS 2012
Select and Sample - A Model of Efficient Neural Inference and Learning · NIPS 2011
The Maximal Causes of Natural Scenes are Edge Filters · NIPS 2010
Machine learning › Deep learning architectures and training › deep generative model
memory-augmented generative model
0.312017
Variational Memory Addressing in Generative Models · NIPS 2017
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › deep latent variable model
helmholtz machine
0.212016
Bidirectional Helmholtz Machines · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › sparse bayesian learning
spike-and-slab prior
0.112012
Why MCA? Nonlinear sparse coding with spike-and-slab prior for neurally plausible image encoding · NIPS 2012
Machine learning › Representation and self-supervised learning › blind source separation
independent component analysis
0.112010
The Maximal Causes of Natural Scenes are Edge Filters · NIPS 2010
Machine learning › Generative modeling › generative model › probabilistic generative model
nonlinear generative models
0.112010
The Maximal Causes of Natural Scenes are Edge Filters · NIPS 2010
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.112017
Variational Memory Addressing in Generative Models · NIPS 2017

Methods — techniques the papers use, named apart from their topics

temporal difference regularization · 0.8stochastic gradient descent · 0.8mean squared error · 0.8maximum likelihood estimation · 0.8kalman filter · 0.8inverse soft-q learning · 0.8diffusion objective · 0.8bayesian inference · 0.8autoregressive modeling · 0.8causal induction · 0.7
YearPublicationVenuePosition
2024 Kalman Filter for Online Classification of Non-Stationary Data
abstract
In Online Continual Learning (OCL) a learning system receives a stream of data and sequentially performs prediction and training steps. Key challenges in OCL include automatic adaptation to the specific non-stationary structure of the data and maintaining appropriate predictive uncertainty. To address these challenges we introduce a probabilistic Bayesian online learning approach that utilizes a (possibly pretrained) neural representation and a state space model over the linear predictor weights. Non-stationarity in the linear predictor weights is modelled using a “parameter drift” transition density, parametrized by a coefficient that quantifies forgetting. Inference in the model is implemented with efficient Kalman filter recursions which track the posterior distribution over the linear weights, while online SGD updates over the transition dynamics coefficient allow for adaptation to the non-stationarity observed in the data. While the framework is developed assuming a linear Gaussian model, we extend it to deal with classification problems and for fine-tuning the deep learning representation. In a set of experiments in multi-class classification using data sets such as CIFAR-100 and CLOC we demonstrate the model's predictive ability and its flexibility in capturing non-stationarity.
Michalis K. Titsias, Alexandre Galashov, Amal Rannen Triki, Razvan Pascanu, Yee Whye Teh, Jörg Bornschein
ICLR6
2024 Denoising Autoregressive Representation Learning
abstract
In this paper, we explore a new generative approach for learning visual representations. Our method, DARL, employs a decoder-only Transformer to predict image patches autoregressively. We find that training with Mean Squared Error (MSE) alone leads to strong representations. To enhance the image generation ability, we replace the MSE loss with the diffusion objective by using a denoising patch decoder. We show that the learned representation can be improved by using tailored noise schedules and longer training in larger models. Notably, the optimal schedule differs significantly from the typical ones used in standard image diffusion models. Overall, despite its simple architecture, DARL delivers performance remarkably close to state-of-the-art masked prediction models under the fine-tuning protocol. This marks an important step towards a unified model capable of both visual perception and generation, effectively combining the strengths of autoregressive and denoising diffusion models.
Yazhe Li, Jörg Bornschein
ICML2
2024 Imitating Language via Scalable Inverse Reinforcement Learning
abstract
The majority of language model training builds on imitation learning. It covers pretraining, supervised fine-tuning, and affects the starting conditions for reinforcement learning from human feedback (RLHF). The simplicity and scalability of maximum likelihood estimation (MLE) for next token prediction led to its role as predominant paradigm. However, the broader field of imitation learning can more effectively utilize the sequential structure underlying autoregressive generation. We focus on investigating the inverse reinforcement learning (IRL) perspective to imitation, extracting rewards and directly optimizing sequences instead of individual token likelihoods and evaluate its benefits for fine-tuning large language models. We provide a new angle, reformulating inverse soft-Q-learning as a temporal difference regularized extension of MLE. This creates a principled connection between MLE and IRL and allows trading off added complexity with increased performance and diversity of generations in the supervised fine-tuning (SFT) setting. We find clear advantages for IRL-based imitation, in particular for retaining diversity while maximizing task performance, rendering IRL a strong alternative on fixed SFT datasets even without online data generation. Our analysis of IRL-extracted reward functions further indicates benefits for more robust reward functions via tighter integration of supervised and preference-based LLM post-training.
Markus Wulfmeier, Michael Bloesch, Nino Vieillard, Arun Ahuja, Jörg Bornschein, Sandy H. Huang, Artem Sokolov 0001, Matt Barnes 0001, Guillaume Desjardins, Alex Bewley, Sarah Bechtle, Jost Tobias Springenberg, Nikola Momchev, Olivier Bachem, Matthieu Geist, Martin A. Riedmiller
NeurIPS5
2023 Sequential Learning of Neural Networks for Prequential MDL
Jörg Bornschein, Yazhe Li, Marcus Hutter
ICLR1
2023 Learning to Induce Causal Structure
Nan Rosemary Ke, Silvia Chiappa, Jane X. Wang, Jörg Bornschein, Anirudh Goyal, Mélanie Rey, Theophane Weber, Matt M. Botvinick, Michael C. Mozer, Danilo Jimenez Rezende
ICLR4
2023 Evaluating Representations with Readout Model Switching
Yazhe Li, Jörg Bornschein, Marcus Hutter
ICLR2
2023 Nevis'22: A Stream of 100 Tasks Sampled from 30 Years of Computer Vision Research
abstract
A shared goal of several machine learning communities like continual learning, meta-learning and transfer learning, is to design algorithms and models that efficiently and robustly adapt to unseen tasks. An even more ambitious goal is to build models that never stop adapting, and that become increasingly more efficient through time by suitably transferring the accrued knowledge. Beyond the study of the actual learning algorithm and model architecture, there are several hurdles towards our quest to build such models, such as the choice of learning protocol, metric of success and data needed to validate research hypotheses. In this work, we introduce the Never-Ending VIsual-classification Stream (NEVIS'22), a benchmark consisting of a stream of over 100 visual classification tasks, sorted chronologically and extracted from papers sampled uniformly from computer vision proceedings spanning the last three decades. The resulting stream reflects what the research community thought was meaningful at any point in time, and it serves as an ideal test bed to assess how well models can adapt to new tasks, and do so better and more efficiently as time goes by. Despite being limited to classification, the resulting stream has a rich diversity of tasks from OCR, to texture analysis, scene recognition, and so forth. The diversity is also reflected in the wide range of dataset sizes, spanning over four orders of magnitude. Overall, NEVIS'22 poses an unprecedented challenge for current sequential learning approaches due to the scale and diversity of tasks, yet with a low entry barrier as it is limited to a single modality and well understood supervised learning problems. Moreover, we provide a reference implementation including strong baselines and an evaluation protocol to compare methods in terms of their trade-off between accuracy and compute. We hope that NEVIS'22 can be useful to researchers working on continual learning, meta-learning, AutoML and more generally sequential learning, and help these communities join forces towards more robust models that efficiently adapt to a never ending stream of data.
Jörg Bornschein, Alexandre Galashov, Ross Hemsley, Amal Rannen Triki, Yutian Chen 0001, Arslan Chaudhry, Xu Owen He, Arthur Douillard, Massimo Caccia, Qixuan Feng, Sylvestre-Alvise Rebuffi, Kitty Stacpoole, Diego de Las Casas, Will Hawkins, Angeliki Lazaridou, Yee Whye Teh, Andrei A. Rusu, Razvan Pascanu, Marc'Aurelio Ranzato
J. Mach. Learn. Res.1
2020 Small Data, Big Decisions: Model Selection in the Small-Data Regime
abstract
Highly overparametrized neural networks can display curiously strong generalization performance – a phenomenon that has recently garnered a wealth of theoretical and empirical research in order to better understand it. In contrast to most previous work, which typically considers the performance as a function of the model size, in this paper we empirically study the generalization performance as the size of the training set varies over multiple orders of magnitude. These systematic experiments lead to some interesting and potentially very useful observations; perhaps most notably that training on smaller subsets of the data can lead to more reliable model selection decisions whilst simultaneously enjoying smaller computational overheads. Our experiments furthermore allow us to estimate Minimum Description Lengths for common datasets given modern neural network architectures, thereby paving the way for principled model selection taking into account Occams-razor.
Jörg Bornschein, Francesco Visin, Simon Osindero
ICML1
2017 Variational Memory Addressing in Generative Models
abstract
Aiming to augment generative models with external memory, we interpret the output of a memory module with stochastic addressing as a conditional mixture distribution, where a read operation corresponds to sampling a discrete memory address and retrieving the corresponding content from memory. This perspective allows us to apply variational inference to memory addressing, which enables effective training of the memory module by using the target information to guide memory lookups. Stochastic addressing is particularly well-suited for generative models as it naturally encourages multimodality which is a prominent aspect of most high-dimensional datasets. Treating the chosen address as a latent variable also allows us to quantify the amount of information gained with a memory lookup and measure the contribution of the memory module to the generative process. To illustrate the advantages of this approach we incorporate it into a variational autoencoder and apply the resulting model to the task of generative few-shot learning. The intuition behind this architecture is that the memory module can pick a relevant template from memory and the continuous part of the model can concentrate on modeling remaining variations. We demonstrate empirically that our model is able to identify and access the relevant memory contents even with hundreds of unseen Omniglot characters in memory.
Jörg Bornschein, Andriy Mnih, Daniel Zoran, Danilo Jimenez Rezende
NIPS1
2016 Bidirectional Helmholtz Machines
abstract
Efficient unsupervised training and inference in deep generative models remains a challenging problem. One basic approach, called Helmholtz machine or Variational Autoencoder, involves training a top-down directed generative model together with a bottom-up auxiliary model used for approximate inference. Recent results indicate that better generative models can be obtained with better approximate inference procedures. Instead of improving the inference procedure, we here propose a new model, the bidirectional Helmholtz machine, which guarantees that the top-down and bottom-up distributions can efficiently invert each other. We achieve this by interpreting both the top-down and the bottom-up directed models as approximate inference distributions and by defining the model distribution to be the geometric mean of these two. We present a lower-bound for the likelihood of this model and we show that optimizing this bound regularizes the model so that the Bhattacharyya distance between the bottom-up and top-down approximate distributions is minimized. This approach results in state of the art generative models which prefer significantly deeper architectures while it allows for orders of magnitude more efficient likelihood estimation.
Jörg Bornschein, Samira Shabanian, Asja Fischer, Yoshua Bengio
ICML1
2013 Are V1 Simple Cells Optimized for Visual Occlusions? A Comparative Study
abstract
Simple cells in primary visual cortex were famously found to respond to low-level image components such as edges. Sparse coding and independent component analysis (ICA) emerged as the standard computational models for simple cell coding because they linked their receptive fields to the statistics of visual stimuli. However, a salient feature of image statistics, occlusions of image components, is not considered by these models. Here we ask if occlusions have an effect on the predicted shapes of simple cell receptive fields. We use a comparative approach to answer this question and investigate two models for simple cells: a standard linear model and an occlusive model. For both models we simultaneously estimate optimal receptive fields, sparsity and stimulus noise. The two models are identical except for their component superposition assumption. We find the image encoding and receptive fields predicted by the models to differ significantly. While both models predict many Gabor-like fields, the occlusive model predicts a much sparser encoding and high percentages of 'globular' receptive fields. This relatively new center-surround type of simple cell response is observed since reverse correlation is used in experimental studies. While high percentages of 'globular' fields can be obtained using specific choices of sparsity and overcompleteness in linear sparse coding, no or only low proportions are reported in the vast majority of studies on linear models (including all ICA models). Likewise, for the here investigated linear model and optimal sparsity, only low proportions of 'globular' fields are observed. In comparison, the occlusive model robustly infers high proportions and can match the experimentally observed high proportions of 'globular' fields well. Our computational study, therefore, suggests that 'globular' fields may be evidence for an optimal encoding of visual occlusions in primary visual cortex.
Jörg Bornschein, Marc Henniges, Jörg Lücke
PLoS Comput. Biol.1
2012 Why MCA? Nonlinear sparse coding with spike-and-slab prior for neurally plausible image encoding
abstract
Modelling natural images with sparse coding (SC) has faced two main challenges: flexibly representing varying pixel intensities and realistically representing low- level image components. This paper proposes a novel multiple-cause generative model of low-level image statistics that generalizes the standard SC model in two crucial points: (1) it uses a spike-and-slab prior distribution for a more realistic representation of component absence/intensity, and (2) the model uses the highly nonlinear combination rule of maximal causes analysis (MCA) instead of a lin- ear combination. The major challenge is parameter optimization because a model with either (1) or (2) results in strongly multimodal posteriors. We show for the first time that a model combining both improvements can be trained efficiently while retaining the rich structure of the posteriors. We design an exact piece- wise Gibbs sampling method and combine this with a variational method based on preselection of latent dimensions. This combined training scheme tackles both analytical and computational intractability and enables application of the model to a large number of observed and hidden dimensions. Applying the model to image patches we study the optimal encoding of images by simple cells in V1 and compare the model’s predictions with in vivo neural recordings. In contrast to standard SC, we find that the optimal prior favors asymmetric and bimodal ac- tivity of simple cells. Testing our model for consistency we find that the average posterior is approximately equal to the prior. Furthermore, we find that the model predicts a high percentage of globular receptive fields alongside Gabor-like fields. Similarly high percentages are observed in vivo. Our results thus argue in favor of improvements of the standard sparse coding model for simple cells by using flexible priors and nonlinear combinations.
Jacquelyn Shelton, Philip Sterne, Jörg Bornschein, Abdul-Saboor Sheikh, Jörg Lücke
NIPS3
2011 A Gabor Wavelet Pyramid-Based Object Detection Algorithm
Yasuomi D. Sato, Jenia Jitsev, Jörg Bornschein, Daniela Pamplona, Christian Keck, Christoph von der Malsburg
ISNN (2)3
2011 Select and Sample - A Model of Efficient Neural Inference and Learning
abstract
An increasing number of experimental studies indicate that perception encodes a posterior probability distribution over possible causes of sensory stimuli, which is used to act close to optimally in the environment. One outstanding difficulty with this hypothesis is that the exact posterior will in general be too complex to be represented directly, and thus neurons will have to represent an approximation of this distribution. Two influential proposals of efficient posterior representation by neural populations are: 1) neural activity represents samples of the underlying distribution, or 2) they represent a parametric representation of a variational approximation of the posterior. We show that these approaches can be combined for an inference scheme that retains the advantages of both: it is able to represent multiple modes and arbitrary correlations, a feature of sampling methods, and it reduces the represented space to regions of high probability mass, a strength of variational approximations. Neurally, the combined method can be interpreted as a feed-forward preselection of the relevant state space, followed by a neural dynamics implementation of Markov Chain Monte Carlo (MCMC) to approximate the posterior over the relevant states. We demonstrate the effectiveness and efficiency of this approach on a sparse coding model. In numerical experiments on artificial data and image patches, we compare the performance of the algorithms to that of exact EM, variational state space selection alone, MCMC alone, and the combined select and sample approach. The select and sample approach integrates the advantages of the sampling and variational approximations, and forms a robust, neurally plausible, and very efficient model of processing and learning in cortical networks. For sparse coding we show applications easily exceeding a thousand observed and a thousand hidden dimensions.
Jacquelyn Shelton, Jörg Bornschein, Abdul-Saboor Sheikh, Pietro Berkes, Jörg Lücke
NIPS2
2010 The Maximal Causes of Natural Scenes are Edge Filters
abstract
We study the application of a strongly non-linear generative model to image patches. As in standard approaches such as Sparse Coding or Independent Component Analysis, the model assumes a sparse prior with independent hidden variables. However, in the place where standard approaches use the sum to combine basis functions we use the maximum. To derive tractable approximations for parameter estimation we apply a novel approach based on variational Expectation Maximization. The derived learning algorithm can be applied to large-scale problems with hundreds of observed and hidden variables. Furthermore, we can infer all model parameters including observation noise and the degree of sparseness. In applications to image patches we find that Gabor-like basis functions are obtained. Gabor-like functions are thus not a feature exclusive to approaches assuming linear superposition. Quantitatively, the inferred basis functions show a large diversity of shapes with many strongly elongated and many circular symmetric functions. The distribution of basis function shapes reflects properties of simple cell receptive fields that are not reproduced by standard linear approaches. In the study of natural image statistics, the implications of using different superposition assumptions have so far not been investigated systematically because models with strong non-linearities have been found analytically and computationally challenging. The presented algorithm represents the first large-scale application of such an approach.
Gervasio Puertas, Jörg Bornschein, Jörg Lücke
NIPS2