VLDB 2026 Research / reviewers in the wild / expert
Lucas Theis
dblp:28/8772
· DBLP profile ↗
23ranked-venue papers
8as first author
7since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 7 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Generative modeling · 44% Probabilistic and Bayesian machine learning · 15% Information extraction and text analysis · 10% | |
| Computer graphics and multimedia
7 papers |
Image and video coding · 68% Image and video processing · 17% Visual content generation and editing · 14% | |
| Theoretical computer science
2 papers |
Coding theory · 78% Information theory · 22% |
Topics — the 30 heaviest of 45, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning theory › computational complexity
algorithmic information theory |
0.8 | 1 | 2024 | Position: What makes an image realistic? · ICML 2024 |
Machine learning › Generative modeling › generative model evaluation
realism evaluation |
0.8 | 1 | 2024 | Position: What makes an image realistic? · ICML 2024 |
Image and video coding
image quality assessment |
0.8 | 1 | 2024 | The Unreasonable Effectiveness of Linear Prediction as a Perceptual Metric · ICLR 2024 |
Image and video coding › neural compression
neural image and video compression |
0.8 | 1 | 2024 | C3: High-Performance and Low-Complexity Neural Compression from a Single Image or Video · CVPR 2024 |
Image and video coding › image quality assessment
perceptual similarity |
0.8 | 1 | 2024 | The Unreasonable Effectiveness of Linear Prediction as a Perceptual Metric · ICLR 2024 |
Machine learning › Generative modeling
generative adversarial network |
0.7 | 2 | 2019 | HoloGAN: Unsupervised Learning of 3D Representations From Natural Images · ICCV 2019 Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network · CVPR 2017 |
Coding theory › channel coding
channel simulation |
0.6 | 1 | 2022 | Algorithms for the Communication of Samples · ICML 2022 |
Coding theory › error-correcting codes
hybrid coding |
0.6 | 1 | 2022 | Algorithms for the Communication of Samples · ICML 2022 |
Image and video coding
image compression |
0.4 | 1 | 2020 | Universally Quantized Neural Compression · NeurIPS 2020 |
Image and video coding
neural compression |
0.4 | 1 | 2020 | Universally Quantized Neural Compression · NeurIPS 2020 |
Coding theory › source coding
quantization |
0.4 | 1 | 2020 | Universally Quantized Neural Compression · NeurIPS 2020 |
Coding theory › source coding › quantization
universal quantization |
0.4 | 1 | 2020 | Universally Quantized Neural Compression · NeurIPS 2020 |
Machine learning › Generative modeling › generative adversarial network
3d-aware image synthesis |
0.4 | 1 | 2019 | HoloGAN: Unsupervised Learning of 3D Representations From Natural Images · ICCV 2019 |
Natural language and speech › Information extraction and text analysis › topic model
latent dirichlet allocation |
0.4 | 1 | 2019 | Discriminative Topic Modeling with Logistic LDA · NeurIPS 2019 |
Natural language and speech › Information extraction and text analysis
topic model |
0.4 | 1 | 2019 | Discriminative Topic Modeling with Logistic LDA · NeurIPS 2019 |
Computer vision › 3D vision › geometric deep learning
unsupervised 3d representation learning |
0.4 | 1 | 2019 | HoloGAN: Unsupervised Learning of 3D Representations From Natural Images · ICCV 2019 |
Machine learning › Deep learning architectures and training
autoencoder |
0.3 | 1 | 2017 | Lossy Image Compression with Compressive Autoencoders · ICLR (Poster) 2017 |
Machine learning › Generative modeling
diffusion model |
0.3 | 1 | 2017 | Amortised MAP Inference for Image Super-resolution · ICLR 2017 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
MAP inference |
0.3 | 1 | 2017 | Amortised MAP Inference for Image Super-resolution · ICLR 2017 |
Machine learning › Generative modeling
style transfer |
0.3 | 1 | 2017 | Fast Face-Swap Using Convolutional Neural Networks · ICCV 2017 |
Machine learning › Generative modeling › image reconstruction
super-resolution |
0.3 | 1 | 2017 | Amortised MAP Inference for Image Super-resolution · ICLR 2017 |
Machine learning › Generative modeling › generative adversarial network › image-to-image translation
super-resolution GAN |
0.3 | 1 | 2017 | Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network · CVPR 2017 |
Visual content generation and editing › face editing
face swapping |
0.3 | 1 | 2017 | Fast Face-Swap Using Convolutional Neural Networks · ICCV 2017 |
Image and video processing › super-resolution
image super-resolution |
0.3 | 1 | 2017 | Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network · CVPR 2017 |
Image and video coding › image compression
lossy image compression |
0.3 | 1 | 2017 | Lossy Image Compression with Compressive Autoencoders · ICLR (Poster) 2017 |
Image and video processing
perceptual loss |
0.3 | 1 | 2017 | Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network · CVPR 2017 |
Image and video processing › super-resolution › image super-resolution
perceptual super-resolution |
0.3 | 1 | 2017 | Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network · CVPR 2017 |
Machine learning › Efficient and distributed learning
model compression |
0.2 | 1 | 2024 | C3: High-Performance and Low-Complexity Neural Compression from a Single Image or Video · CVPR 2024 |
Machine learning › Generative modeling
autoregressive model |
0.2 | 1 | 2015 | Generative Image Modeling Using Spatial LSTMs · NIPS 2015 |
Machine learning › Generative modeling
image generation |
0.2 | 1 | 2015 | Generative Image Modeling Using Spatial LSTMs · NIPS 2015 |
Methods — techniques the papers use, named apart from their topics
rate-distortion optimization · 1.5overfitting · 1.5universal quantization · 1.3uniform noise channel · 1.3soft quantizer · 1.3weighted least squares · 0.8linear autoregressive model · 0.8algorithmic information theory · 0.8poisson functional representation · 0.6importance sampling · 0.6dithered quantization · 0.6rigid body transformation · 0.4logistic LDA · 0.4generative adversarial network · 0.4deep neural network · 0.4residual network · 0.3content loss · 0.3adversarial loss · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | C3: High-Performance and Low-Complexity Neural Compression from a Single Image or VideoabstractMost neural compression models are trained on large datasets of images or videos in order to generalize to unseen data. Such generalization typically requires large and expressive architectures with a high decoding complexity. Here we introduce C3, a neural compression method with strong rate-distortion (RD) performance that instead overfits a small model to each image or video separately. The resulting decoding complexity of C3 can be an order of magnitude lower than neural baselines with similar RD performance. C3 builds on Cool-chic [43] and makes several simple and effective improvements for images. We further develop new methodology to apply C3 to videos. On the CLIC2020 image benchmark, we match the RD performance of VTM, the reference implementation of the H.266 codec, with less than 3k MACs/pixel for decoding. On the UVG video benchmark, we match the RD performance of the Video Compression Transformer [60], a well-established neural video codec, with less than 5k MACs/pixel for decoding. Hyunjik Kim, Matthias Bauer 0001, Lucas Theis, Jonathan Schwarz, Emilien Dupont |
CVPR | 3 |
| 2024 | The Unreasonable Effectiveness of Linear Prediction as a Perceptual MetricabstractWe show how perceptual embeddings of the visual system can be constructed at inference-time with no training data or deep neural network features. Our perceptual embeddings are solutions to a weighted least squares (WLS) problem, defined at the pixel-level, and solved at inference-time, that can capture global and local image characteristics. The distance in embedding space is used to define a perceptual similary metric which we call \emph{LASI: Linear Autoregressive Similarity Index}. Experiments on full-reference image quality assessment datasets show LASI performs competitively with learned deep feature based methods like LPIPS \citep{zhang2018unreasonable} and PIM \citep{bhardwaj2020unsupervised}, at a similar computational cost to hand-crafted methods such as MS-SSIM \citep{wang2003multiscale}. We found that increasing the dimensionality of the embedding space consistently reduces the WLS loss while increasing performance on perceptual tasks, at the cost of increasing the computational complexity. LASI is fully differentiable, scales cubically with the number of embedding dimensions, and can be parallelized at the pixel-level. A Maximum Differentiation (MAD) competition \citep{wang2008maximum} between LASI and LPIPS shows that both methods are capable of finding failure points for the other, suggesting these metrics can be combined. Daniel Severo 0001, Lucas Theis, Jona Ballé |
ICLR | 2 |
| 2024 | Position: What makes an image realistic?abstractThe last decade has seen tremendous progress in our ability to *generate* realistic-looking data, be it images, text, audio, or video. Here, we discuss the closely related problem of *quantifying* realism, that is, designing functions that can reliably tell realistic data from unrealistic data. This problem turns out to be significantly harder to solve and remains poorly understood, despite its prevalence in machine learning and recent breakthroughs in generative AI. Drawing on insights from algorithmic information theory, we discuss why this problem is challenging, why a good generative model alone is insufficient to solve it, and what a good solution would look like. In particular, we introduce the notion of a *universal critic*, which unlike adversarial critics does not require adversarial training. While universal critics are not immediately practical, they can serve both as a North Star for guiding practical implementations and as a tool for analyzing existing attempts to capture realism. Lucas Theis |
ICML | 1 |
| 2024 | Gaussian Channel Simulation with Rotated Dithered QuantizationabstractChannel simulation involves generating a sample$Y$from the conditional distribution$P_{Y\vert X}$, where$X$is a remote realization sampled from$P_{X}$. This paper introduces a novel approach to approximate Gaussian channel simulation using dithered quantization. Our method concurrently simulates$n$channels, reducing the upper bound on the excess information by half compared to one-dimensional methods. When used with higher-dimensional lattices, our approach achieves up to six times reduction on the upper bound. Furthermore, we demonstrate that the KL divergence between the distributions of the simulated and Gaussian channels decreases with the number of dimensions at a rate of$O(n^{-1})$. Szymon Kobus, Lucas Theis, Deniz Gündüz |
ISIT | 2 |
| 2023 | Adaptive Greedy Rejection SamplingabstractWe consider channel simulation protocols between two communicating parties, Alice and Bob. First, Alice receives a target distribution Q, unknown to Bob. Then, she employs a shared coding distribution P to send the minimum amount of information to Bob so that he can simulate a single sample X ~ Q. For discrete distributions, Harsha et al. [1] developed a well-known channel simulation protocol – greedy rejection sampling (GRS) – with a bound of ${D_{{\text{KL}}}}\left[ {Q\left\| P \right.} \right] + 2\ln \left( {{D_{{\text{KL}}}}\left[ {Q\left\| P \right.} \right] + 1} \right) + \mathcal{O}\left( 1 \right)$ on the expected code-length of the protocol. In this paper, we extend the definition of GRS to general probability spaces and allow it to adapt its proposal distribution after each step. We call this new procedure Adaptive GRS (AGRS) and prove its correctness. Furthermore, we prove the surprising result that the expected runtime of GRS is exactly exp(D∞[Q║P]), where D∞[Q║P] denotes the Rényi ∞-divergence. We then apply AGRS to Gaussian channel simulation problems. We show that the expected runtime of GRS is infinite when averaged over target distributions and propose a solution that trades off a slight increase in the coding cost for a finite runtime. Finally, we describe a specific instance of AGRS for 1D Gaussian channels inspired by hybrid coding [2]. We conjecture and demonstrate empirically that the runtime of AGRS is $\mathcal{O}\left( {{D_{KL}}\left[ {Q\left\| P \right.} \right]} \right)$ in this case. Gergely Flamich, Lucas Theis |
ISIT | 2 |
| 2022 | Optimal Compression of Locally Differentially Private MechanismsabstractCompressing the output of $\epsilon$-locally differentially private (LDP) randomizers naively leads to suboptimal utility. In this work, we demonstrate the benefits of using schemes that jointly compress and privatize the data using shared randomness. In particular, we investigate a family of schemes based on Minimal Random Coding (Havasi et al., 2019) and prove that they offer optimal privacy-accuracy-communication tradeoffs. Our theoretical and empirical findings show that our approach can compress PrivUnit (Bhowmick et al., 2018) and Subset Selection (Ye et al., 2018), the best known LDP algorithms for mean and frequency estimation, to the order of $\epsilon$ bits of communication while preserving their privacy and accuracy guarantees. Abhin Shah, Wei-Ning Chen, Jona Ballé, Peter Kairouz, Lucas Theis |
AISTATS | 5 |
| 2022 | Algorithms for the Communication of SamplesabstractThe efficient communication of noisy data has applications in several areas of machine learning, such as neural compression or differential privacy, and is also known as reverse channel coding or the channel simulation problem. Here we propose two new coding schemes with practical advantages over existing approaches. First, we introduce ordered random coding (ORC) which uses a simple trick to reduce the coding cost of previous approaches. This scheme further illuminates a connection between schemes based on importance sampling and the so-called Poisson functional representation. Second, we describe a hybrid coding scheme which uses dithered quantization to more efficiently communicate samples from distributions with bounded support. Lucas Theis, Noureldin Y. Ahmed |
ICML | 1 |
| 2020 | Universally Quantized Neural CompressionabstractA popular approach to learning encoders for lossy compression is to use additive uniform noise during training as a differentiable approximation to test-time quantization. We demonstrate that a uniform noise channel can also be implemented at test time using universal quantization (Ziv, 1985). This allows us to eliminate the mismatch between training and test phases while maintaining a completely differentiable loss function. Implementing the uniform noise channel is a special case of the more general problem of communicating a sample, which we prove is computationally hard if we do not make assumptions about its distribution. However, the uniform special case is efficient as well as easy to implement and thus of great interest from a practical point of view. Finally, we show that quantization can be obtained as a limiting case of a soft quantizer applied to the uniform noise channel, bridging compression with and without quantization. Eirikur Agustsson, Lucas Theis |
NeurIPS | 2 |
| 2019 | HoloGAN: Unsupervised Learning of 3D Representations From Natural ImagesabstractWe propose a novel generative adversarial network (GAN) for the task of unsupervised learning of 3D representations from natural images. Most generative models rely on 2D kernels to generate images and make few assumptions about the 3D world. These models therefore tend to create blurry images or artefacts in tasks that require a strong 3D understanding, such as novel-view synthesis. HoloGAN instead learns a 3D representation of the world, and to render this representation in a realistic manner. Unlike other GANs, HoloGAN provides explicit control over the pose of generated objects through rigid-body transformations of the learnt 3D features. Our experiments show that using explicit 3D features enables HoloGAN to disentangle 3D pose and identity, which is further decomposed into shape and appearance, while still being able to generate images with similar or higher visual quality than other generative models. HoloGAN can be trained end-to-end from unlabelled 2D images only. Particularly, we do not require pose labels, 3D shapes, or multiple views of the same objects. This shows that HoloGAN is the first generative model that learns 3D representations from natural images in an entirely unsupervised manner. Thu Nguyen-Phuoc, Lucas Theis, Christian Richardt |
ICCV | 3 |
| 2019 | Discriminative Topic Modeling with Logistic LDAabstractDespite many years of research into latent Dirichlet allocation (LDA), applying LDA to collections of non-categorical items is still challenging for practitioners. Yet many problems with much richer data share a similar structure and could benefit from the vast literature on LDA. We propose logistic LDA, a novel discriminative variant of latent Dirichlet allocation which is easy to apply to arbitrary inputs. In particular, our model can easily be applied to groups of images, arbitrary text embeddings, or integrate deep neural networks. Although it is a discriminative model, we show that logistic LDA can learn from unlabeled data in an unsupervised manner by exploiting the group structure present in the data. In contrast to other recent topic models designed to handle arbitrary inputs, our model does not sacrifice the interpretability and principled motivation of LDA. Iryna Korshunova, Hanchen Xiong, Mateusz Fedoryszak, Lucas Theis |
NeurIPS | 4 |
| 2019 | Addressing delayed feedback for continuous training with neural networks in CTR predictionabstractOne of the challenges in display advertising is that the distribution of features and click through rate (CTR) can exhibit large shifts over time due to seasonality, changes to ad campaigns and other factors. The predominant strategy to keep up with these shifts is to train predictive models continuously, on fresh data, in order to prevent them from becoming stale. However, in many ad systems positive labels are only observed after a possibly long and random delay. These delayed labels pose a challenge to data freshness in continuous training: fresh data may not have complete label information at the time they are ingested by the training algorithm. Naive strategies which consider any data point a negative example until a positive label becomes available tend to underestimate CTR, resulting in inferior user experience and suboptimal performance for advertisers. The focus of this paper is to identify the best combination of loss functions and models that enable large-scale learning from a continuous stream of data in the presence of delayed labels. In this work, we compare 5 different loss functions, 3 of them applied to this problem for the first time. We benchmark their performance in offline settings on both public and proprietary datasets in conjunction with shallow and deep model architectures. We also discuss the engineering cost associated with implementing each loss function in a production environment. Finally, we carried out online experiments with the top performing methods, in order to validate their performance in a continuous training scheme. While training on 668 million in-house data points offline, our proposed methods outperform previous state-of-the-art by 3% relative cross entropy (RCE). During online experiments, we observed 55% gain in revenue per thousand requests (RPMq) against naive log loss. Sofia Ira Ktena, Alykhan Tejani, Lucas Theis, Pranay Kumar Myana, Deepak Dilipkumar, Ferenc Huszar, Steven Yoo, Wenzhe Shi |
RecSys | 3 |
| 2018 | Adaptive Paired-Comparison Method for Subjective Video Quality Assessment on Mobile DevicesabstractTo effectively evaluate subjective visual quality in weakly-controlled environments, we propose an Adaptive Paired Comparison method based on particle filtering. As our approach requires each sample to be rated only once, the test time compared to regular paired comparison can be reduced. The method works with non-experts and improves reliability compared to MOS and DS-MOS methods. Katherine Storrs, Sebastiaan Van Leuven, Steve Kojder, Lucas Theis, Ferenc Huszar |
PCS | 4 |
| 2018 | Community-based benchmarking improves spike rate inference from two-photon calcium imaging dataabstractIn recent years, two-photon calcium imaging has become a standard tool to probe the function of neural circuits and to study computations in neuronal populations. However, the acquired signal is only an indirect measurement of neural activity due to the comparatively slow dynamics of fluorescent calcium indicators. Different algorithms for estimating spike rates from noisy calcium measurements have been proposed in the past, but it is an open question how far performance can be improved. Here, we report the results of the spikefinder challenge, launched to catalyze the development of new spike rate inference algorithms through crowd-sourcing. We present ten of the submitted algorithms which show improved performance compared to previously evaluated methods. Interestingly, the top-performing algorithms are based on a wide range of principles from deep neural networks to generative models, yet provide highly correlated estimates of the neural activity. The competition shows that benchmark challenges can drive algorithmic developments in neuroscience. Philipp Berens, Jeremy Freeman, Thomas Deneux, Nicolay Chenkov, Thomas McColgan, Artur Speiser, Jakob H. Macke, Srinivas C. Turaga, Patrick J. Mineault, Peter Rupprecht, Stephan Gerhard, Rainer W. Friedrich, Johannes Friedrich, Liam Paninski, Marius Pachitariu, Kenneth D. Harris, Ben Bolte, Timothy A. Machado, Dario Ringach, Jasmine Stone, Luke E. Rogerson, Nicolas J. Sofroniew, Jacob Reimer, Emmanouil Froudarakis, Thomas Euler, Miroslav Román Rosón, Lucas Theis, Andreas S. Tolias, Matthias Bethge |
PLoS Comput. Biol. | 27 |
| 2017 | Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial NetworkabstractDespite the breakthroughs in accuracy and speed of single image super-resolution using faster and deeper convolutional neural networks, one central problem remains largely unsolved: how do we recover the finer texture details when we super-resolve at large upscaling factors? The behavior of optimization-based super-resolution methods is principally driven by the choice of the objective function. Recent work has largely focused on minimizing the mean squared reconstruction error. The resulting estimates have high peak signal-to-noise ratios, but they are often lacking high-frequency details and are perceptually unsatisfying in the sense that they fail to match the fidelity expected at the higher resolution. In this paper, we present SRGAN, a generative adversarial network (GAN) for image super-resolution (SR). To our knowledge, it is the first framework capable of inferring photo-realistic natural images for 4x upscaling factors. To achieve this, we propose a perceptual loss function which consists of an adversarial loss and a content loss. The adversarial loss pushes our solution to the natural image manifold using a discriminator network that is trained to differentiate between the super-resolved images and original photo-realistic images. In addition, we use a content loss motivated by perceptual similarity instead of similarity in pixel space. Our deep residual network is able to recover photo-realistic textures from heavily downsampled images on public benchmarks. An extensive mean-opinion-score (MOS) test shows hugely significant gains in perceptual quality using SRGAN. The MOS scores obtained with SRGAN are closer to those of the original high-resolution images than to those obtained with any state-of-the-art method. Christian Ledig, Lucas Theis, Ferenc Huszar, Jose Caballero, Andrew Cunningham, Alejandro Acosta, Andrew P. Aitken, Alykhan Tejani, Johannes Totz, Wenzhe Shi |
CVPR | 2 |
| 2017 | Fast Face-Swap Using Convolutional Neural NetworksabstractWe consider the problem of face swapping in images, where an input identity is transformed into a target identity while preserving pose, facial expression and lighting. To perform this mapping, we use convolutional neural networks trained to capture the appearance of the target identity from an unstructured collection of his/her photographs. This approach is enabled by framing the face swapping problem in terms of style transfer, where the goal is to render an image in the style of another one. Building on recent advances in this area, we devise a new loss function that enables the network to produce highly photorealistic results. By combining neural networks with simple pre- and post-processing steps, we aim at making face swap work in real-time with no input from the user. Iryna Korshunova, Wenzhe Shi, Joni Dambre, Lucas Theis |
ICCV | 4 |
| 2017 | Amortised MAP Inference for Image Super-resolution
Casper Kaae Sønderby, Jose Caballero, Lucas Theis, Wenzhe Shi, Ferenc Huszar |
ICLR | 3 |
| 2017 | Lossy Image Compression with Compressive Autoencoders
Lucas Theis, Wenzhe Shi, Andrew Cunningham, Ferenc Huszar |
ICLR (Poster) | 1 |
| 2015 | Data modeling with the elliptical gamma distributionabstractWe study mixture modeling using the elliptical gamma (EG) distribution, a non-Gaussian distribution that allows heavy and light tail and peak behaviors. We first consider maximum likelihood parameter estimation, a task that turns out to be very challenging: we must handle positive definiteness constraints, and more crucially, we must handle possibly nonconcave log-likelihoods, which makes maximization hard. We overcome these difficulties by developing algorithms based on fixed-point theory; our methods respect the psd constraint, while also efficiently solving the (possibly) nonconcave maximization to global optimality. Subsequently, we focus on mixture modeling using EG distributions: we present a closed-form expression of the KL-divergence between two EG distributions, which we then combine with our ML estimation methods to obtain an efficient split-and-merge expectation maximization algorithm. We illustrate the use of our model and algorithms on a dataset of natural image patches. Suvrit Sra, Reshad Hosseini, Lucas Theis, Matthias Bethge |
AISTATS | 3 |
| 2015 | A trust-region method for stochastic variational inference with applications to streaming dataabstractStochastic variational inference allows for fast posterior inference in complex Bayesian models. However, the algorithm is prone to local optima which can make the quality of the posterior approximation sensitive to the choice of hyperparameters and initialization. We address this problem by replacing the natural gradient step of stochastic varitional inference with a trust-region update. We show that this leads to generally better results and reduced sensitivity to hyperparameters. We also describe a new strategy for variational inference on streaming data and show that here our trust-region method is crucial for getting good performance. Lucas Theis, Matthew Hoffman 0001 |
ICML | 1 |
| 2015 | Generative Image Modeling Using Spatial LSTMsabstractModeling the distribution of natural images is challenging, partly because of strong statistical dependencies which can extend over hundreds of pixels. Recurrent neural networks have been successful in capturing long-range dependencies in a number of problems but only recently have found their way into generative image models. We here introduce a recurrent image model based on multi-dimensional long short-term memory units which are particularly suited for image modeling due to their spatial structure. Our model scales to images of arbitrary size and its likelihood is computationally tractable. We find that it outperforms the state of the art in quantitative comparisons on several image datasets and produces promising results when used for texture synthesis and inpainting. Lucas Theis, Matthias Bethge |
NIPS | 1 |
| 2013 | Beyond GLMs: A Generative Mixture Modeling Approach to Neural System IdentificationabstractGeneralized linear models (GLMs) represent a popular choice for the probabilistic characterization of neural spike responses. While GLMs are attractive for their computational tractability, they also impose strong assumptions and thus only allow for a limited range of stimulus-response relationships to be discovered. Alternative approaches exist that make only very weak assumptions but scale poorly to high-dimensional stimulus spaces. Here we seek an approach which can gracefully interpolate between the two extremes. We extend two frequently used special cases of the GLM-a linear and a quadratic model-by assuming that the spike-triggered and non-spike-triggered distributions can be adequately represented using Gaussian mixtures. Because we derive the model from a generative perspective, its components are easy to interpret as they correspond to, for example, the spike-triggered distribution and the interspike interval distribution. The model is able to capture complex dependencies on high-dimensional stimuli with far fewer parameters than other approaches such as histogram-based methods. The added flexibility comes at the cost of a non-concave log-likelihood. We show that in practice this does not have to be an issue and the mixture-based model is able to outperform generalized linear and quadratic models. Lucas Theis, André Maia Chagas, Daniel Arnstein, Cornelius Schwarz, Matthias Bethge |
PLoS Comput. Biol. | 1 |
| 2012 | Training sparse natural image models with a fast Gibbs sampler of an extended state spaceabstractWe present a new learning strategy based on an efficient blocked Gibbs sampler for sparse overcomplete linear models. Particular emphasis is placed on statistical image modeling, where overcomplete models have played an important role in discovering sparse representations. Our Gibbs sampler is faster than general purpose sampling schemes while also requiring no tuning as it is free of parameters. Using the Gibbs sampler and a persistent variant of expectation maximization, we are able to extract highly sparse distributions over latent sources from data. When applied to natural images, our algorithm learns source distributions which resemble spike-and-slab distributions. We evaluate the likelihood and quantitatively compare the performance of the overcomplete linear model to its complete counterpart as well as a product of experts model, which represents another overcomplete generalization of the complete linear model. In contrast to previous claims, we find that overcomplete representations lead to significant improvements, but that the overcomplete linear model still underperforms other models. Lucas Theis, Jascha Sohl-Dickstein, Matthias Bethge |
NIPS | 1 |
| 2011 | In All Likelihood, Deep Belief Is Not Enough
Lucas Theis, Sebastian Gerwinn, Fabian H. Sinz, Matthias Bethge |
J. Mach. Learn. Res. | 1 |