VLDB 2026 Research / reviewers in the wild / expert
Jona Ballé
dblp:84/4973 · also Johannes Ballé
· DBLP profile ↗
30ranked-venue papers
10as first author
12since 2021 · last 2025
0000-0003-0769-8985ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Good, Cheap, and Fast: Overfitted Image Compression with Wasserstein DistortionabstractInspired by the success of generative image models, recent work on learned image compression increasingly focuses on better probabilistic models of the natural image distribution, leading to excellent image quality. This, however, comes at the expense of a computational complexity that is several orders of magnitude higher than today’s commercial codecs, and thus prohibitive for most practical applications. With this paper, we demonstrate that by focusing on modeling visual perception rather than the data distribution, we can achieve a very good trade-off between visual quality and bit rate similar to “generative” compression models such as HiFiC, while requiring less than 1% of the multiply–accumulate operations (MACs) for decompression. We do this by optimizing C3, an overfitted image codec, for Wasserstein Distortion (WD), and evaluating the image reconstructions with a human rater study, showing that WD clearly outperforms LPIPS as an optimization objective. The study also reveals that WD outperforms other perceptual metrics such as LPIPS, DISTS, and MS-SSIM as a predictor of human ratings, remarkably achieving over 94% Pearson correlation with Elo scores. Jona Ballé, Luca Versari, Emilien Dupont, Hyunjik Kim |
CVPR | 1 |
| 2025 | Discretized Approximate Ancestral SamplingabstractThe Fourier Basis Density Model (FBM) [1] was recently introduced as a flexible probability model for band-limited distributions, i.e. ones which are smooth in the sense of having a characteristic function with limited support around the origin. Its density and cumulative distribution functions can be efficiently evaluated and trained with stochastic optimization methods, which makes the model suitable for deep learning applications. However, the model lacked support for sampling. Here, we introduce a method inspired by discretization-interpolation methods common in Digital Signal Processing, which directly take advantage of the band-limited property. We review mathematical properties of the FBM, and prove quality bounds of the sampled distribution in terms of the total variation (TV) and Wasserstein-1 divergences from the model. These bounds can be used to inform the choice of hyperparameters to reach any desired sample quality. We discuss these results in comparison to a variety of other sampling techniques, highlighting tradeoffs between computational complexity and sampling quality. Alfredo De la Fuente, Jona Ballé |
ISIT | 3 |
| 2024 | The Unreasonable Effectiveness of Linear Prediction as a Perceptual MetricabstractWe show how perceptual embeddings of the visual system can be constructed at inference-time with no training data or deep neural network features. Our perceptual embeddings are solutions to a weighted least squares (WLS) problem, defined at the pixel-level, and solved at inference-time, that can capture global and local image characteristics. The distance in embedding space is used to define a perceptual similary metric which we call \emph{LASI: Linear Autoregressive Similarity Index}. Experiments on full-reference image quality assessment datasets show LASI performs competitively with learned deep feature based methods like LPIPS \citep{zhang2018unreasonable} and PIM \citep{bhardwaj2020unsupervised}, at a similar computational cost to hand-crafted methods such as MS-SSIM \citep{wang2003multiscale}. We found that increasing the dimensionality of the embedding space consistently reduces the WLS loss while increasing performance on perceptual tasks, at the cost of increasing the computational complexity. LASI is fully differentiable, scales cubically with the number of embedding dimensions, and can be parallelized at the pixel-level. A Maximum Differentiation (MAD) competition \citep{wang2008maximum} between LASI and LPIPS shows that both methods are capable of finding failure points for the other, suggesting these metrics can be combined. Daniel Severo 0001, Lucas Theis, Jona Ballé |
ICLR | 3 |
| 2024 | Fourier Basis Density ModelabstractWe introduce a lightweight, flexible and end-to-end trainable probability density model parameterized by a constrained Fourier basis. We assess its performance at approximating a range of multimodal 1D densities, which are generally difficult to fit. In comparison to the deep factorized model introduced in [1], our model achieves a lower cross entropy at a similar computational budget. In addition, we also evaluate our method on a toy compression task, demonstrating its utility in learned compression. Alfredo De la Fuente, Jona Ballé |
PCS | 3 |
| 2023 | Learned Wyner-Ziv Compressors Recover BinningabstractWe consider lossy compression of an information source when the decoder has lossless access to a correlated one. This setup, also known as the Wyner-Ziv problem, is a special case of distributed source coding. To this day, real-world applications of this problem have neither been fully developed nor heavily investigated. We propose a data-driven method based on machine learning that leverages the universal function approximation capability of artificial neural networks. We find that our neural network-based compression scheme re-discovers some principles of the optimum theoretical solution of the Wyner-Ziv setup, such as binning in the source space as well as linear decoder behavior within each quantization index, for the quadratic-Gaussian case. These behaviors emerge although no structure exploiting knowledge of the source distributions was imposed. Binning is a widely used tool in information theoretic proofs and methods, and to our knowledge, this is the first time it has been explicitly observed to emerge from data-driven learning. Ezgi Özyilkan, Jona Ballé, Elza Erkip |
ISIT | 2 |
| 2022 | Optimal Compression of Locally Differentially Private MechanismsabstractCompressing the output of $\epsilon$-locally differentially private (LDP) randomizers naively leads to suboptimal utility. In this work, we demonstrate the benefits of using schemes that jointly compress and privatize the data using shared randomness. In particular, we investigate a family of schemes based on Minimal Random Coding (Havasi et al., 2019) and prove that they offer optimal privacy-accuracy-communication tradeoffs. Our theoretical and empirical findings show that our approach can compress PrivUnit (Bhowmick et al., 2018) and Subset Selection (Ye et al., 2018), the best known LDP algorithms for mean and frequency estimation, to the order of $\epsilon$ bits of communication while preserving their privacy and accuracy guarantees. Abhin Shah, Wei-Ning Chen, Jona Ballé, Peter Kairouz, Lucas Theis |
AISTATS | 3 |
| 2022 | Hyperspectral remote sensing data compression with neural networksabstractHyperspectral images are typically highly correlated along their spectrum, and this similarity is usually found to cluster in intervals of consecutive bands. We identified 5 such intervals in AVIRIS uncalibrated data (i.e., as captured on-board). These 5 intervals maximised the average spectral correlation along the 224 band spectrum. The resulting in-tervals were composed of bands 1–40, 41–96, 97–155, 156–165, and 166–224, as seen in the figure to the right. Sebastià Mijares i Verdú, Jona Ballé, Valero Laparra, Joan Bartrina-Rapesta, Miguel Hernández-Cabronero, Joan Serra-Sagristà |
DCC | 2 |
| 2022 | Neural Video Compression Using GANs for Detail Synthesis and Propagation
Fabian Mentzer, Eirikur Agustsson, Jona Ballé, David Minnen, Nicholas Johnston, George Toderici |
ECCV (26) | 3 |
| 2022 | On the relation between statistical learning and perceptual distances
Alexander Hepburn, Valero Laparra, Raúl Santos-Rodríguez, Jona Ballé, Jesús Malo |
ICLR | 4 |
| 2022 | Do Neural Networks Compress Manifolds Optimally?abstractArtifical Neural-Network-based (ANN-based) lossy compressors have recently obtained striking results on several sources. Their success may be ascribed to an ability to identify the structure of low-dimensional manifolds in high-dimensional ambient spaces. Indeed, prior work has shown that ANN-based compressors can achieve the optimal entropy-distortion curve for some such sources. In contrast, we determine the optimal entropy-distortion tradeoffs for two low-dimensional manifolds with circular structure and show that state-of-the-art ANN-based compressors fail to optimally compress them. Sourbh Bhadane, Aaron B. Wagner, Jona Ballé |
ITW | 3 |
| 2021 | Neural Networks Optimally Compress the SawbridgeabstractNeural-network-based compressors have proven to be remarkably effective at compressing sources, such as images, that are nominally high-dimensional but presumed to be concentrated on a low-dimensional manifold. We consider a continuous-time random process that models an extreme version of such a source, wherein the realizations fall along a one-dimensional “curve” in function space that has infinite-dimensional linear span. We precisely characterize the optimal entropy-distortion tradeoff for this source and show numerically that it is achieved by neural-network-based compressors trained via stochastic gradient descent. In contrast, we show both analytically and experimentally that compressors based on the classical Karhunen-Loeve transform are highly suboptimal at high rates. Aaron B. Wagner, Jona Ballé |
DCC | 2 |
| 2021 | 3D Scene Compression through Entropy Penalized Neural Representation FunctionsabstractSome forms of novel visual media enable the viewer to explore a 3D scene from essentially arbitrary viewpoints, by interpolating between a discrete set of original views. Compared to 2D imagery, these types of applications require much larger amounts of storage space, which we seek to reduce. Existing approaches for compressing 3D scenes are often based on a separation of compression and rendering: each of the original views is compressed using traditional 2D image formats; the receiver decompresses the views and then performs the rendering. We unify these steps by directly compressing an implicit representation of the scene, a function that maps spatial coordinates to a radiance vector field, which can then be queried to render arbitrary viewpoints. The function is implemented as a neural network and jointly trained for reconstruction as well as compressibility, in an end-to-end manner, with the use of an entropy penalty on the parameters. Our method significantly outperforms a state-of-the-art conventional approach for scene compression, achieving simultaneously higher quality reconstructions and lower bitrates. Furthermore, we show that the performance at lower bitrates can be improved by jointly representing multiple scenes using a soft form of parameter sharing. Thomas Bird, Jona Ballé, Philip A. Chou |
PCS | 2 |
| 2020 | Scale-Space Flow for End-to-End Optimized Video CompressionabstractDespite considerable progress on end-to-end optimized deep networks for image compression, video coding remains a challenging task. Recently proposed methods for learned video compression use optical flow and bilinear warping for motion compensation and show competitive rate-distortion performance relative to hand-engineered codecs like H.264 and HEVC. However, these learning-based methods rely on complex architectures and training schemes including the use of pre-trained optical flow networks, sequential training of sub-networks, adaptive rate control, and buffering intermediate reconstructions to disk during training. In this paper, we show that a generalized warping operator that better handles common failure cases, e.g. disocclusions and fast motion, can provide competitive compression results with a greatly simplified model and training procedure. Specifically, we propose scale-space flow, an intuitive generalization of optical flow that adds a scale parameter to allow the network to better model uncertainty. Our experiments show that a low-latency video compression model (no B-frames) using scale-space flow for motion compensation can outperform analogous state-of-the art learned video compression models while being trained using a much simpler procedure and without any pre-trained optical flow networks. Eirikur Agustsson, David Minnen, Nicholas Johnston, Jona Ballé, Sung Jin Hwang, George Toderici |
CVPR | 4 |
| 2020 | End-to-End Learning of Compressible FeaturesabstractPre-trained convolutional neural networks (CNNs) are powerful off-the-shelf feature generators and have been shown to perform very well on a variety of tasks. Unfortunately, the generated features are high dimensional and expensive to store: potentially hundreds of thousands of floats per example when processing videos. Traditional entropy based lossless compression methods are of little help as they do not yield desired level of compression, while general purpose lossy compression methods based on energy compaction (e.g. PCA followed by quantization and entropy coding) are sub-optimal, as they are not tuned to task specific objective. We propose a learned method that jointly optimizes for compressibility along with the task objective for learning the features. The plug-in nature of our method makes it straight-forward to integrate with any target objective and trade-off against compressibility. We present results on multiple benchmarks and demonstrate that our method produces features that are an order of magnitude more compressible, while having a regularization effect that leads to a consistent improvement in accuracy. Sami Abu-El-Haija, Nicholas Johnston, Jona Ballé, Abhinav Shrivastava, George Toderici |
ICIP | 4 |
| 2020 | Scalable Model Compression by Entropy Penalized Reparameterization
Deniz Oktay, Jona Ballé, Abhinav Shrivastava |
ICLR | 2 |
| 2020 | An Unsupervised Information-Theoretic Perceptual Quality MetricabstractTractable models of human perception have proved to be challenging to build. Hand-designed models such as MS-SSIM remain popular predictors of human image quality judgements due to their simplicity and speed. Recent modern deep learning approaches can perform better, but they rely on supervised data which can be costly to gather: large sets of class labels such as ImageNet, image quality ratings, or both. We combine recent advances in information-theoretic objective functions with a computational architecture informed by the physiology of the human visual system and unsupervised training on pairs of video frames, yielding our Perceptual Information Metric (PIM). We show that PIM is competitive with supervised metrics on the recent and challenging BAPPS image quality assessment dataset and outperforms them in predicting the ranking of image compression methods in CLIC 2020. We also perform qualitative experiments using the ImageNet-C dataset, and establish that PIM is robust with respect to architectural details. Sangnie Bhardwaj, Ian Fischer, Jona Ballé, Troy T. Chinen |
NeurIPS | 3 |
| 2019 | Integer Networks for Data Compression with Latent-Variable Models
Jona Ballé, Nicholas Johnston, David Minnen |
ICLR (Poster) | 1 |
| 2018 | Towards A Semantic Perceptual Image MetricabstractWe present a full reference, perceptual image metric based on VGG-16, an artificial neural network trained on object classification. We fit the metric to a new database based on 140k unique images annotated with ground truth by human raters who received minimal instruction. The resulting metric shows competitive performance on TID 2013, a database widely used to assess image quality assessments methods. More interestingly, it shows strong responses to objects potentially carrying semantic relevance such as faces and text, which we demonstrate using a visualization technique and ablation experiments. In effect, the metric appears to model a higher influence of semantic context on judgments, which we observe particularly in untrained raters. As the vast majority of users of image processing systems are unfamiliar with Image Quality Assessment (IQA) tasks, these findings may have significant impact on real-world applications of perceptual metrics. Troy T. Chinen, Jona Ballé, Chunhui Gu, Sung Jin Hwang, Sergey Ioffe, Nicholas Johnston, Thomas K. Leung, David Minnen, Sean M. O'Malley, Charles Rosenberg 0001, George Toderici |
ICIP | 2 |
| 2018 | Variational image compression with a scale hyperprior
Jona Ballé, David Minnen, Sung Jin Hwang, Nicholas Johnston |
ICLR (Poster) | 1 |
| 2018 | Joint Autoregressive and Hierarchical Priors for Learned Image CompressionabstractRecent models for learned image compression are based on autoencoders that learn approximately invertible mappings from pixels to a quantized latent representation. The transforms are combined with an entropy model, which is a prior on the latent representation that can be used with standard arithmetic coding algorithms to generate a compressed bitstream. Recently, hierarchical entropy models were introduced as a way to exploit more structure in the latents than previous fully factorized priors, improving compression performance while maintaining end-to-end optimization. Inspired by the success of autoregressive priors in probabilistic generative models, we examine autoregressive, hierarchical, and combined priors as alternatives, weighing their costs and benefits in the context of image compression. While it is well known that autoregressive models can incur a significant computational penalty, we find that in terms of compression performance, autoregressive and hierarchical priors are complementary and can be combined to exploit the probabilistic structure in the latents better than all previous learned models. The combined model yields state-of-the-art rate-distortion performance and generates smaller files than existing methods: 15.8% rate reductions over the baseline hierarchical model and 59.8%, 35%, and 8.4% savings over JPEG, JPEG2000, and BPG, respectively. To the best of our knowledge, our model is the first learning-based method to outperform the top standard image codec (BPG) on both the PSNR and MS-SSIM distortion metrics. David Minnen, Jona Ballé, George Toderici |
NeurIPS | 2 |
| 2018 | Efficient Nonlinear Transforms for Lossy Image CompressionabstractWe assess the performance of two techniques in the context of nonlinear transform coding with artificial neural networks, Sadam and GDN. Both techniques have been success- fully used in state-of-the-art image compression methods, but their performance has not been individually assessed to this point. Together, the techniques stabilize the training procedure of nonlinear image transforms and increase their capacity to approximate the (unknown) rate-distortion optimal transform functions. Besides comparing their performance to established alternatives, we detail the implementation of both methods and provide open-source code along with the paper. Jona Ballé |
PCS | 1 |
| 2017 | End-to-end Optimized Image Compression
Jona Ballé, Valero Laparra, Eero P. Simoncelli |
ICLR | 1 |
| 2017 | Eigen-Distortions of Hierarchical RepresentationsabstractWe develop a method for comparing hierarchical image representations in terms of their ability to explain perceptual sensitivity in humans. Specifically, we utilize Fisher information to establish a model-derived prediction of sensitivity to local perturbations of an image. For a given image, we compute the eigenvectors of the Fisher information matrix with largest and smallest eigenvalues, corresponding to the model-predicted most- and least-noticeable image distortions, respectively. For human subjects, we then measure the amount of each distortion that can be reliably detected when added to the image. We use this method to test the ability of a variety of representations to mimic human perceptual sensitivity. We find that the early layers of VGG16, a deep neural network optimized for object recognition, provide a better match to human perception than later layers, and a better match than a 4-stage convolutional neural network (CNN) trained on a database of human ratings of distorted image quality. On the other hand, we find that simple models of early visual processing, incorporating one or more stages of local gain control, trained on the same database of distortion ratings, provide substantially better predictions of human sensitivity than either the CNN, or any combination of layers of VGG16. Alexander Berardino, Valero Laparra, Jona Ballé, Eero P. Simoncelli |
NIPS | 3 |
| 2016 | End-to-end optimization of nonlinear transform codes for perceptual qualityabstractWe introduce a general framework for end-to-end optimization of the rate-distortion performance of nonlinear transform codes assuming scalar quantization. The framework can be used to optimize any differentiable pair of analysis and synthesis transforms in combination with any differentiable perceptual metric. As an example, we consider a code built from a linear transform followed by a form of multi-dimensional local gain control. Distortion is measured with a state-of-the-art perceptual metric. When optimized over a large database of images, this representation offers substantial improvements in bitrate and perceptual appearance over fixed (DCT) codes, and over linear transform codes optimized for mean squared error. Jona Ballé, Valero Laparra, Eero P. Simoncelli |
PCS | 1 |
| 2014 | Learning sparse filter bank transforms with convolutional ICAabstractIndependent Component Analysis (ICA) is a generalization of Principal Component Analysis that optimizes a linear transformation to whiten and sparsify a family of source signals. The computational costs of ICA grow rapidly with dimensionality, and application to high-dimensional data is generally achieved by restricting to small windows, violating the translation-invariant nature of many real-world signals, and producing blocking artifacts in applications. Here, we reformulate the ICA problem for transformations computed through convolution with a bank of filters, and develop a generalization of the fastICA algorithm for optimizing the filters over a set of example signals. This results in a substantial reduction of computational complexity and memory requirements. When applied to a database of photographic images, the method yields bandpass oriented filters, whose responses are sparser than those of orthogonal wavelets or block DCT, and slightly more heavy-tailed than those of block ICA, despite fewer model parameters. Jona Ballé, Eero P. Simoncelli |
ICIP | 1 |
| 2012 | Subjective evaluation of texture similarity metrics for compression applicationsabstractThe paper summarizes the results of an experimental subjective evaluation of texture similarity metrics with 25 test subjects. The compared metrics comprise the frequency-weighted log-spectral and Itakura distances, as well as a set of metrics based on an overcomplete Gabor-like filterbank - the STSIM as well as two new ones. The test set consists of 30 synthetic GMRF texture pairs. It turns out that all metrics perform well, but the weighted log-spectral distance tends to outperform the filterbank-based metrics. Jona Ballé |
PCS | 1 |
| 2012 | Median trilateral loop filter for depth map video codingabstractEmerging extensions to conventional stereo video technologies like 3D Video require to add depth information to 2D video data. This supplementary data needs to be coded efficiently and transmitted to the receiver where arbitrary viewpoints are generated by using this additional information. The depth maps are characterized by piecewise smooth regions, which are bounded by sharp edges describing depth discontinuities along object boundaries. Preserving these characteristics and especially depth discontinuities is a crucial requirement for depth map coding. When coding depth maps by means of a conventional hybrid video coder, ringing artifacts are introduced along the sharp edges and result in quality degradation when using the reconstructed depth maps for view synthesis. To reduce these ringing artifacts and also to better align object boundaries in video and depth data, a new in-loop filter is proposed, which reconstructs the described characteristics of depth maps. Fabian Jäger, Jona Ballé |
PCS | 2 |
| 2011 | Improved entropy coding for component-based image codingabstractIn this paper, we improve on our previous work regarding component-based image coding, a hybrid transform-based/perceptual image coding scheme based on a decomposition of the image into structure and texture characterized by a Gaussian Markov random field. The 2D Itakura distance allows us to evaluate the performance of our texture model in terms of rate vs. distortion. A minimal quantization step size for near-lossless coding of model parameters is determined. Furthermore, we show that texture contrast can be efficiently coded using transform-based techniques. Christian Feldmann, Jona Ballé |
ICIP | 2 |
| 2009 | Component-based image coding using non-local means filtering and an autoregressive texture modelabstractWhile noise is usually regarded as a problem of the image formation process, we observe that it is also frequently part of natural texture. In this paper, we present a concept for improved compression of noisy texture in natural images. Since noise is problematic to decorrelation-based compression methods, we propose to perform image decomposition by denoising, followed by separate compression of the components. The denoised component is encoded using conventional methods, while the texture component is compressed by encoding parameters of a texture model. It turns out that at similar bit rates, our method can improve visual quality. Jona Ballé, Bastian Jurczyk, Aleksandar Stojanovic 0002 |
ICIP | 1 |
| 2007 | Extended Texture Prediction for H.264/AVC Intra CodingabstractEfficient intra prediction is an important aspect of video coding with high compression efficiency. H.264/AVC applies directional prediction from neighboring pixels on an adjustable block size for local decorrelation. In this paper, we present an extended prediction scheme in the context of H.264/AVC that comprises two additional prediction methods exploiting self-similar properties of the encoded texture. A new macroblock type is implemented, allowing for flexible selection of the available prediction methods for sub-partitions of the macroblock. Depending on the content of the encoded video sequence, substantial gains in rate-distortion performance are achieved. The results may indicate directions towards an enhanced intra coding scheme with improved rate-distortion performance. Jona Ballé, Mathias Wien |
ICIP (6) | 1 |