Eero P. Simoncelli

dblp:30/5604 · DBLP profile ↗
← Back
99ranked-venue papers
13as first author
19since 2021 · last 2025
0000-0002-1206-527XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 59 · 4 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 10 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 Discriminating image representations with principal distortions
abstract
Image representations (artificial or biological) are often compared in terms of their global geometric structure; however, representations with similar global structure can have strikingly different local geometries. Here, we propose a framework for comparing a set of image representations in terms of their local geometries. We quantify the local geometry of a representation using the Fisher information matrix, a standard statistical tool for characterizing the sensitivity to local stimulus distortions, and use this as a substrate for a metric on the local geometry in the vicinity of a base image. This metric may then be used to optimally differentiate a set of models, by finding a pair of "principal distortions" that maximize the variance of the models under this metric. As an example, we use this framework to compare a set of simple models of the early visual system, identifying a novel set of image distortions that allow immediate comparison of the models by visual inspection. In a second example, we apply our method to a set of deep neural network models and reveal differences in the local geometry that arise due to architecture and training types. These examples demonstrate how our framework can be used to probe for informative differences in local sensitivities between complex models, and suggest how it could be used to compare model representations with human perception.
Jenelle Feather, David Lipshutz, Sarah E. Harvey, Alex H. Williams, Eero P. Simoncelli
ICLR5
2025 Learning normalized image densities via dual score matching
abstract
Learning probability models from data is at the heart of many machine learning endeavors, but is notoriously difficult due to the curse of dimensionality. We introduce a new framework for learning \emph{normalized} energy (log probability) models that is inspired by diffusion generative models, which rely on networks optimized to estimate the score. We modify a score network architecture to compute an energy while preserving its inductive biases. The gradient of this energy network with respect to its input image is the score of the learned density, which can be optimized using a denoising objective. Importantly, the gradient with respect to the noise level provides an additional score that can be optimized with a novel secondary objective, ensuring consistent and normalized energies across noise levels. We train an energy network with this \emph{dual} score matching objective on the ImageNet64 dataset, and obtain a cross-entropy (negative log likelihood) value comparable to the state of the art. We further validate our approach by showing that our energy model \emph{strongly generalizes}: log probabilities estimated with two networks trained on non-overlapping data subsets are nearly identical. Finally, we demonstrate that both image probability and dimensionality of local neighborhoods vary substantially depending on image content, in contrast with conventional assumptions such as concentration of measure or support on a low-dimensional manifold.
Florentin Guth, Zahra Kadkhodaie, Eero P. Simoncelli
NeurIPS3
2025 Perceptual learning improves discrimination but does not reduce distortions in appearance
abstract
Human perceptual sensitivity often improves with training, a phenomenon known as "perceptual learning." Another important perceptual dimension is appearance, the subjective sense of stimulus magnitude. Are training-induced improvements in sensitivity accompanied by more accurate appearance? Here, we examined this question by measuring both discrimination (sensitivity) and estimation (appearance) responses to near-horizontal motion directions, which are known to be repulsed away from horizontal. Participants performed discrimination and estimation tasks before and after training in either the discrimination or the estimation task or none (control group). Human observers who trained in either discrimination or estimation exhibited improvements in discrimination accuracy, but estimation repulsion did not decrease; instead, it either persisted or increased. Hence, distortions in perception can be exacerbated after perceptual learning. We developed a computational observer model in which perceptual learning arises from increases in the precision of underlying neural representations, which explains this counterintuitive finding. For each observer, the fitted model accounted for discrimination performance, the distribution of estimates, and their changes with training. Our empirical findings and modeling suggest that learning enhances distinctions between categories, a potentially important aspect of real-world perception and perceptual learning.
Sarit Szpiro, Charlie S. Burlingham, Eero P. Simoncelli, Marisa Carrasco
PLoS Comput. Biol.3
2024 Generalization in diffusion models arises from geometry-adaptive harmonic representations
abstract
Deep neural networks (DNNs) trained for image denoising are able to generate high-quality samples with score-based reverse diffusion algorithms. These impressive capabilities seem to imply an escape from the curse of dimensionality, but recent reports of memorization of the training set raise the question of whether these networks are learning the "true" continuous density of the data. Here, we show that two DNNs trained on non-overlapping subsets of a dataset learn nearly the same score function, and thus the same density, when the number of training images is large enough. In this regime of strong generalization, diffusion-generated images are distinct from the training set, and are of high visual quality, suggesting that the inductive biases of the DNNs are well-aligned with the data density. We analyze the learned denoising functions and show that the inductive biases give rise to a shrinkage operation in a basis adapted to the underlying image. Examination of these bases reveals oscillating harmonic structures along contours and in homogeneous regions. We demonstrate that trained denoisers are inductively biased towards these geometry-adaptive harmonic bases since they arise not only when the network is trained on photographic images, but also when it is trained on image classes supported on low-dimensional manifolds for which the harmonic basis is suboptimal. Finally, we show that when trained on regular image classes for which the optimal basis is known to be geometry-adaptive and harmonic, the denoising performance of the networks is near-optimal.
Zahra Kadkhodaie, Florentin Guth, Eero P. Simoncelli, Stéphane Mallat
ICLR3
2024 Shaping the distribution of neural responses with interneurons in a recurrent circuit model
abstract
Efficient coding theory posits that sensory circuits transform natural signals into neural representations that maximize information transmission subject to resource constraints. Local interneurons are thought to play an important role in these transformations, shaping patterns of circuit activity to facilitate and direct information flow. However, the relationship between these coordinated, nonlinear, circuit-level transformations and the properties of interneurons (e.g., connectivity, activation functions) remains unknown. Here, we propose a normative computational model that establishes such a relationship. Our model is derived from an optimal transport objective that conceptualizes the circuit's input-response function as transforming the inputs to achieve a target response distribution. The circuit, which is comprised of primary neurons that are recurrently connected to a set of local interneurons, continuously optimizes this objective by dynamically adjusting both the synaptic connections between neurons as well as the interneuron activation functions. In an application motivated by redundancy reduction theory, we demonstrate that when the inputs are natural image statistics and the target distribution is a spherical Gaussian, the circuit learns a nonlinear transformation that significantly reduces statistical dependencies in neural responses. Overall, our results provide a framework in which the distribution of circuit responses is systematically and nonlinearly controlled by adjustment of interneuron connectivity and activation functions.
David Lipshutz, Eero P. Simoncelli
NeurIPS2
2024 Learning predictable and robust neural representations by straightening image sequences
abstract
Prediction is a fundamental capability of all living organisms, and has been proposed as an objective for learning sensory representations. Recent work demonstrates that in primate visual systems, prediction is facilitated by neural representations that follow straighter temporal trajectories than their initial photoreceptor encoding, which allows for prediction by linear extrapolation. Inspired by these experimental findings, we develop a self-supervised learning (SSL) objective that explicitly quantifies and promotes straightening. We demonstrate the power of this objective in training deep feedforward neural networks on smoothly-rendered synthetic image sequences that mimic commonly-occurring properties of natural videos. The learned model contains neural embeddings that are predictive, but also factorize the geometric, photometric, and semantic attributes of objects. The representations also prove more robust to noise and adversarial attacks compared to previous SSL methods that optimize for invariance to random augmentations. Moreover, these beneficial properties can be transferred to other training procedures by using the straightening objective as a regularizer, suggesting a broader utility for straightening as a principle for robust unsupervised learning.
Xueyan Niu 0002, Cristina Savin, Eero P. Simoncelli
NeurIPS3
2024 Contrastive-Equivariant Self-Supervised Learning Improves Alignment with Primate Visual Area IT
abstract
Models trained with self-supervised learning objectives have recently matched or surpassed models trained with traditional supervised object recognition in their ability to predict neural responses of object-selective neurons in the primate visual system. A self-supervised learning objective is arguably a more biologically plausible organizing principle, as the optimization does not require a large number of labeled examples. However, typical self-supervised objectives may result in network representations that are overly invariant to changes in the input. Here, we show that a representation with structured variability to the input transformations is better aligned with known features of visual perception and neural computation. We introduce a novel framework for converting standard invariant SSL losses into "contrastive-equivariant" versions that encourage preserving aspects of the input transformation without supervised access to the transformation parameters. We further demonstrate that our proposed method systematically increases models' ability to predict responses in macaque inferior temporal cortex. Our results demonstrate the promise of incorporating known features of neural computation into task-optimization for building better models of visual cortex.
Thomas E. Yerxa, Jenelle Feather, Eero P. Simoncelli, SueYeon Chung
NeurIPS3
2023 Learning multi-scale local conditional probability models of images
Zahra Kadkhodaie, Florentin Guth, Stéphane Mallat, Eero P. Simoncelli
ICLR4
2023 Adaptive Whitening in Neural Populations with Gain-modulating Interneurons
abstract
Statistical whitening transformations play a fundamental role in many computational systems, and may also play an important role in biological sensory systems. Existing neural circuit models of adaptive whitening operate by modifying synaptic interactions; however, such modifications would seem both too slow and insufficiently reversible. Motivated by the extensive neuroscience literature on gain modulation, we propose an alternative model that adaptively whitens its responses by modulating the gains of individual neurons. Starting from a novel whitening objective, we derive an online algorithm that whitens its outputs by adjusting the marginal variances of an overcomplete set of projections. We map the algorithm onto a recurrent neural network with fixed synaptic weights and gain-modulating interneurons. We demonstrate numerically that sign-constraining the gains improves robustness of the network to ill-conditioned inputs, and a generalization of the circuit achieves a form of local whitening in convolutional populations, such as those found throughout the visual or auditory systems.
Lyndon R. Duong, David Lipshutz, David J. Heeger, Dmitri B. Chklovskii, Eero P. Simoncelli
ICML5
2023 Adaptive whitening with fast gain modulation and slow synaptic plasticity
abstract
Neurons in early sensory areas rapidly adapt to changing sensory statistics, both by normalizing the variance of their individual responses and by reducing correlations between their responses. Together, these transformations may be viewed as an adaptive form of statistical whitening. Existing mechanistic models of adaptive whitening exclusively use either synaptic plasticity or gain modulation as the biological substrate for adaptation; however, on their own, each of these models has significant limitations. In this work, we unify these approaches in a normative multi-timescale mechanistic model that adaptively whitens its responses with complementary computational roles for synaptic plasticity and gain modulation. Gains are modified on a fast timescale to adapt to the current statistical context, whereas synapses are modified on a slow timescale to match structural properties of the input statistics that are invariant across contexts. Our model is derived from a novel multi-timescale whitening objective that factorizes the inverse whitening matrix into basis vectors, which correspond to synaptic weights, and a diagonal matrix, which corresponds to neuronal gains. We test our model on synthetic and natural datasets and find that the synapses learn optimal configurations over long timescales that enable adaptive whitening on short timescales using gain modulation.
Lyndon R. Duong, Eero P. Simoncelli, Dmitri B. Chklovskii, David Lipshutz
NeurIPS2
2023 A polar prediction model for learning to represent visual transformations
abstract
All organisms make temporal predictions, and their evolutionary fitness level depends on the accuracy of these predictions. In the context of visual perception, the motions of both the observer and objects in the scene structure the dynamics of sensory signals, allowing for partial prediction of future signals based on past ones. Here, we propose a self-supervised representation-learning framework that extracts and exploits the regularities of natural videos to compute accurate predictions. We motivate the polar architecture by appealing to the Fourier shift theorem and its group-theoretic generalization, and we optimize its parameters on next-frame prediction. Through controlled experiments, we demonstrate that this approach can discover the representation of simple transformation groups acting in data. When trained on natural video datasets, our framework achieves better prediction performance than traditional motion compensation and rivals conventional deep networks, while maintaining interpretability and speed. Furthermore, the polar computations can be restructured into components resembling normalized simple and direction-selective complex cell models of primate V1 neurons. Thus, polar prediction offers a principled framework for understanding how the visual system represents sensory inputs in a form that simplifies temporal prediction.
Pierre-Étienne H. Fiquet, Eero P. Simoncelli
NeurIPS2
2023 Learning Efficient Coding of Natural Images with Maximum Manifold Capacity Representations
abstract
The efficient coding hypothesis proposes that the response properties of sensory systems are adapted to the statistics of their inputs such that they capture maximal information about the environment, subject to biological constraints. While elegant, information theoretic properties are notoriously difficult to measure in practical settings or to employ as objective functions in optimization. This difficulty has necessitated that computational models designed to test the hypothesis employ several different information metrics ranging from approximations and lower bounds to proxy measures like reconstruction error. Recent theoretical advances have characterized a novel and ecologically relevant efficiency metric, the ``manifold capacity,” which is the number of object categories that may be represented in a linearly separable fashion. However, calculating manifold capacity is a computationally intensive iterative procedure that until now has precluded its use as an objective. Here we outline the simplifying assumptions that allow manifold capacity to be optimized directly, yielding Maximum Manifold Capacity Representations (MMCR). The resulting method is closely related to and inspired by advances in the field of self supervised learning (SSL), and we demonstrate that MMCRs are competitive with state of the art results on standard SSL benchmarks. Empirical analyses reveal differences between MMCRs and representations learned by other SSL frameworks, and suggest a mechanism by which manifold compression gives rise to class separability. Finally we evaluate a set of SSL methods on a suite of neural predicitivity benchmarks, and find MMCRs are higly competitive as models of the ventral stream.
Thomas E. Yerxa, Yilun Kuang, Eero P. Simoncelli, SueYeon Chung
NeurIPS3
2022 Maximum a posteriori natural scene reconstruction from retinal ganglion cells with deep denoiser priors
abstract
Visual information arriving at the retina is transmitted to the brain by signals in the optic nerve, and the brain must rely solely on these signals to make inferences about the visual world. Previous work has probed the content of these signals by directly reconstructing images from retinal activity using linear regression or nonlinear regression with neural networks. Maximum a posteriori (MAP) reconstruction using retinal encoding models and separately-trained natural image priors offers a more general and principled approach. We develop a novel method for approximate MAP reconstruction that combines a generalized linear model for retinal responses to light, including their dependence on spike history and spikes of neighboring cells, with the image prior implicitly embedded in a deep convolutional neural network trained for image denoising. We use this method to reconstruct natural images from ex vivo simultaneously-recorded spikes of hundreds of retinal ganglion cells uniformly sampling a region of the retina. The method produces reconstructions that match or exceed the state-of-the-art in perceptual similarity and exhibit additional fine detail, while using substantially fewer model parameters than previous approaches. The use of more rudimentary encoding models (a linear-nonlinear-Poisson cascade) or image priors (a 1/f spectral model) significantly reduces reconstruction performance, indicating the essential role of both components in achieving high-quality reconstructed images from the retinal signal.
Eric Wu, Nora Brackbill, Alexander Sher, Alan M. Litke, Eero P. Simoncelli, E. J. Chichilnisky
NeurIPS5
2022 Image Quality Assessment: Unifying Structure and Texture Similarity
abstract
Objective measures of image quality generally operate by comparing pixels of a "degraded" image to those of the original. Relative to human observers, these measures are overly sensitive to resampling of texture regions (e.g., replacing one patch of grass with another). Here, we develop the first full-reference image quality model with explicit tolerance to texture resampling. Using a convolutional neural network, we construct an injective and differentiable function that transforms images to multi-scale overcomplete representations. We demonstrate empirically that the spatial averages of the feature maps in this representation capture texture appearance, in that they provide a set of sufficient statistical constraints to synthesize a wide variety of texture patterns. We then describe an image quality method that combines correlations of these spatial averages ("texture similarity") with correlations of the feature maps ("structure similarity"). The parameters of the proposed measure are jointly optimized to match human ratings of image quality, while minimizing the reported distances between subimages cropped from the same texture images. Experiments show that the optimized method explains human perceptual scores, both on conventional image quality databases, as well as on texture databases. The measure also offers competitive performance on related tasks such as texture classification and retrieval. Finally, we show that our method is relatively insensitive to geometric transformations (e.g., translation and dilation), without use of any specialized training or data augmentation. Code is available at https://github.com/dingkeyan93/DISTS.
Keyan Ding, Kede Ma, Shiqi Wang 0001, Eero P. Simoncelli
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Unsupervised Deep Video Denoising
abstract
Deep convolutional neural networks (CNNs) for video denoising are typically trained with supervision, assuming the availability of clean videos. However, in many applications, such as microscopy, noiseless videos are not available. To address this, we propose an Unsupervised Deep Video Denoiser (UDVD1), a CNN architecture designed to be trained exclusively with noisy data. The performance of UDVD is comparable to the supervised state-of-the-art, even when trained only on a single short noisy video. We demonstrate the promise of our approach in real-world imaging applications by denoising raw video, fluorescence-microscopy and electron-microscopy data. In contrast to many current approaches to video denoising, UDVD does not require explicit motion compensation. This is advantageous because motion compensation is computationally expensive, and can be unreliable when the input data are noisy. A gradient-based analysis reveals that UDVD automatically adapts to local motion in the input noisy videos. Thus, the network learns to perform implicit motion compensation, even though it is only trained for denoising.
Dev Yashpal Sheth, Sreyas Mohan, Joshua L. Vincent, Ramon Manzorro, Peter A. Crozier, Mitesh M. Khapra, Eero P. Simoncelli, Carlos Fernandez-Granda
ICCV7
2021 Impression learning: Online representation learning with synaptic plasticity
abstract
Understanding how the brain constructs statistical models of the sensory world remains a longstanding challenge for computational neuroscience. Here, we derive an unsupervised local synaptic plasticity rule that trains neural circuits to infer latent structure from sensory stimuli via a novel loss function for approximate online Bayesian inference. The learning algorithm is driven by a local error signal computed between two factors that jointly contribute to neural activity: stimulus drive and internal predictions --- the network's 'impression' of the stimulus. Physiologically, we associate these two components with the basal and apical dendrites of pyramidal neurons, respectively. We show that learning can be implemented online, is capable of capturing temporal dependencies in continuous input streams, and generalizes to hierarchical architectures. Furthermore, we demonstrate both analytically and empirically that the algorithm is more data-efficient than a three-factor plasticity alternative, enabling it to learn statistics of high-dimensional, naturalistic inputs. Overall, the model provides a bridge from mechanistic accounts of synaptic plasticity to algorithmic descriptions of unsupervised probabilistic learning and inference.
Colin Bredenberg, Benjamin Lyo, Eero P. Simoncelli, Cristina Savin
NeurIPS3
2021 Stochastic Solutions for Linear Inverse Problems using the Prior Implicit in a Denoiser
abstract
Deep neural networks have provided state-of-the-art solutions for problems such as image denoising, which implicitly rely on a prior probability model of natural images. Two recent lines of work – Denoising Score Matching and Plug-and-Play – propose methodologies for drawing samples from this implicit prior and using it to solve inverse problems, respectively. Here, we develop a parsimonious and robust generalization of these ideas. We rely on a classic statistical result that shows the least-squares solution for removing additive Gaussian noise can be written directly in terms of the gradient of the log of the noisy signal density. We use this to derive a stochastic coarse-to-fine gradient ascent procedure for drawing high-probability samples from the implicit prior embedded within a CNN trained to perform blind denoising. A generalization of this algorithm to constrained sampling provides a method for using the implicit prior to solve any deterministic linear inverse problem, with no additional training, thus extending the power of supervised learning for denoising to a much broader set of problems. The algorithm relies on minimal assumptions and exhibits robust convergence over a wide range of parameter choices. To demonstrate the generality of our method, we use it to obtain state-of-the-art levels of unsupervised performance for deblurring, super-resolution, and compressive sensing.
Zahra Kadkhodaie, Eero P. Simoncelli
NeurIPS2
2021 Adaptive Denoising via GainTuning
abstract
Deep convolutional neural networks (CNNs) for image denoising are typically trained on large datasets. These models achieve the current state of the art, but they do not generalize well to data that deviate from the training distribution. Recent work has shown that it is possible to train denoisers on a single noisy image. These models adapt to the features of the test image, but their performance is limited by the small amount of information used to train them. Here we propose "GainTuning'', a methodology by which CNN models pre-trained on large datasets can be adaptively and selectively adjusted for individual test images. To avoid overfitting, GainTuning optimizes a single multiplicative scaling parameter (the “Gain”) of each channel in the convolutional layers of the CNN. We show that GainTuning improves state-of-the-art CNNs on standard image-denoising benchmarks, boosting their denoising performance on nearly every image in a held-out test set. These adaptive improvements are even more substantial for test images differing systematically from the training data, either in noise level or image type. We illustrate the potential of adaptive GainTuning in a scientific application to transmission-electron-microscope images, using a CNN that is pre-trained on synthetic data. In contrast to the existing methodology, GainTuning is able to faithfully reconstruct the structure of catalytic nanoparticles from these data at extremely low signal-to-noise ratios.
Sreyas Mohan, Joshua L. Vincent, Ramon Manzorro, Peter A. Crozier, Carlos Fernandez-Granda, Eero P. Simoncelli
NeurIPS6
2021 Comparison of Full-Reference Image Quality Models for Optimization of Image Processing Systems
Keyan Ding, Kede Ma, Shiqi Wang 0001, Eero P. Simoncelli
Int. J. Comput. Vis.4
2020 Robust And Interpretable Blind Image Denoising Via Bias-Free Convolutional Neural Networks
Sreyas Mohan, Zahra Kadkhodaie, Eero P. Simoncelli, Carlos Fernandez-Granda
ICLR3
2020 Learning efficient task-dependent representations with synaptic plasticity
abstract
Neural populations encode the sensory world imperfectly: their capacity is limited by the number of neurons, availability of metabolic and other biophysical resources, and intrinsic noise. The brain is presumably shaped by these limitations, improving efficiency by discarding some aspects of incoming sensory streams, while preferentially preserving commonly occurring, behaviorally-relevant information. Here we construct a stochastic recurrent neural circuit model that can learn efficient, task-specific sensory codes using a novel form of reward-modulated Hebbian synaptic plasticity. We illustrate the flexibility of the model by training an initially unstructured neural network to solve two different tasks: stimulus estimation, and stimulus discrimination. The network achieves high performance in both tasks by appropriately allocating resources and using its recurrent circuitry to best compensate for different levels of noise. We also show how the interaction between stimulus priors and task structure dictates the emergent network representations.
Colin Bredenberg, Eero P. Simoncelli, Cristina Savin
NeurIPS2
2019 Blind Image Quality Assessment by Learning from Multiple Annotators
abstract
Models for image quality assessment (IQA) are generally optimized and tested by comparing to human ratings, which are expensive to obtain. Here, we develop a blind IQA (BIQA) model, and a method of training it without human ratings. We first generate a large number of corrupted image pairs, and use a set of existing IQA models to identify which image of each pair has higher quality. We then train a convolutional neural network to estimate perceived image quality along with the uncertainty, optimizing for consistency with the binary labels. The reliability of each IQA annotator is also estimated during training. Experiments demonstrate that our model outperforms state-of-the-art BIQA models in terms of correlation with human ratings in existing databases, as well in group maximum differentiation (gMAD) competition.
Kede Ma, Xuelin Liu, Yuming Fang 0001, Eero P. Simoncelli
ICIP4
2019 Flexible information routing in neural populations through stochastic comodulation
abstract
Humans and animals are capable of flexibly switching between a multitude of tasks, each requiring rapid, sensory-informed decision making. Incoming stimuli are processed by a hierarchy of neural circuits consisting of millions of neurons with diverse feature selectivity. At any given moment, only a small subset of these carry task-relevant information. In principle, downstream processing stages could identify the relevant neurons through supervised learning, but this would require many example trials. Such extensive learning periods are inconsistent with the observed flexibility of humans or animals, who can adjust to changes in task parameters or structure almost immediately. Here, we propose a novel solution based on functionally-targeted stochastic modulation. It has been observed that trial-to-trial neural activity is modulated by a shared, low-dimensional, stochastic signal that introduces task-irrelevant noise. Counter-intuitively this noise is preferentially targeted towards task-informative neurons, corrupting the encoded signal. However, we hypothesize that this modulation offers a solution to the identification problem, labeling task-informative neurons so as to facilitate decoding. We simulate an encoding population of spiking neurons whose rates are modulated by a shared stochastic signal, and show that a linear decoder with readout weights approximating neuron-specific modulation strength can achieve near-optimal accuracy. Such a decoder allows fast and flexible task-dependent information routing without relying on hardwired knowledge of the task-informative neurons (as in maximum likelihood) or unrealistically many supervised training trials (as in regression).
Caroline Haimerl, Cristina Savin, Eero P. Simoncelli
NeurIPS3
2017 End-to-end Optimized Image Compression
Jona Ballé, Valero Laparra, Eero P. Simoncelli
ICLR3
2017 Eigen-Distortions of Hierarchical Representations
abstract
We develop a method for comparing hierarchical image representations in terms of their ability to explain perceptual sensitivity in humans. Specifically, we utilize Fisher information to establish a model-derived prediction of sensitivity to local perturbations of an image. For a given image, we compute the eigenvectors of the Fisher information matrix with largest and smallest eigenvalues, corresponding to the model-predicted most- and least-noticeable image distortions, respectively. For human subjects, we then measure the amount of each distortion that can be reliably detected when added to the image. We use this method to test the ability of a variety of representations to mimic human perceptual sensitivity. We find that the early layers of VGG16, a deep neural network optimized for object recognition, provide a better match to human perception than later layers, and a better match than a 4-stage convolutional neural network (CNN) trained on a database of human ratings of distorted image quality. On the other hand, we find that simple models of early visual processing, incorporating one or more stages of local gain control, trained on the same database of distortion ratings, provide substantially better predictions of human sensitivity than either the CNN, or any combination of layers of VGG16.
Alexander Berardino, Valero Laparra, Jona Ballé, Eero P. Simoncelli
NIPS4
2016 End-to-end optimization of nonlinear transform codes for perceptual quality
abstract
We introduce a general framework for end-to-end optimization of the rate-distortion performance of nonlinear transform codes assuming scalar quantization. The framework can be used to optimize any differentiable pair of analysis and synthesis transforms in combination with any differentiable perceptual metric. As an example, we consider a code built from a linear transform followed by a form of multi-dimensional local gain control. Distortion is measured with a state-of-the-art perceptual metric. When optimized over a large database of images, this representation offers substantial improvements in bitrate and perceptual appearance over fixed (DCT) codes, and over linear transform codes optimized for mean squared error.
Jona Ballé, Valero Laparra, Eero P. Simoncelli
PCS3
2016 Neural Quadratic Discriminant Analysis: Nonlinear Decoding with V1-Like Computation
abstract
Linear-nonlinear (LN) models and their extensions have proven successful in describing transformations from stimuli to spiking responses of neurons in early stages of sensory hierarchies. Neural responses at later stages are highly nonlinear and have generally been better characterized in terms of their decoding performance on prespecified tasks. Here we develop a biologically plausible decoding model for classification tasks, that we refer to as neural quadratic discriminant analysis (nQDA). Specifically, we reformulate an optimal quadratic classifier as an LN-LN computation, analogous to "subunit" encoding models that have been used to describe responses in retina and primary visual cortex. We propose a physiological mechanism by which the parameters of the nQDA classifier could be optimized, using a supervised variant of a Hebbian learning rule. As an example of its applicability, we show that nQDA provides a better account than many comparable alternatives for the transformation between neural representations in two high-level brain areas recorded as monkeys performed a visual delayed-match-to-sample task.
Marino Pagan, Eero P. Simoncelli, Nicole C. Rust
Neural Comput.2
2014 Learning sparse filter bank transforms with convolutional ICA
abstract
Independent Component Analysis (ICA) is a generalization of Principal Component Analysis that optimizes a linear transformation to whiten and sparsify a family of source signals. The computational costs of ICA grow rapidly with dimensionality, and application to high-dimensional data is generally achieved by restricting to small windows, violating the translation-invariant nature of many real-world signals, and producing blocking artifacts in applications. Here, we reformulate the ICA problem for transformations computed through convolution with a bank of filters, and develop a generalization of the fastICA algorithm for optimizing the filters over a set of example signals. This results in a substantial reduction of computational complexity and memory requirements. When applied to a database of photographic images, the method yields bandpass oriented filters, whose responses are sparser than those of orthogonal wavelets or block DCT, and slightly more heavy-tailed than those of block ICA, despite fewer model parameters.
Jona Ballé, Eero P. Simoncelli
ICIP2
2014 Efficient Sensory Encoding and Bayesian Inference with Heterogeneous Neural Populations
abstract
The efficient coding hypothesis posits that sensory systems maximize information transmitted to the brain about the environment. We develop a precise and testable form of this hypothesis in the context of encoding a sensory variable with a population of noisy neurons, each characterized by a tuning curve. We parameterize the population with two continuous functions that control the density and amplitude of the tuning curves, assuming that the tuning widths vary inversely with the cell density. This parameterization allows us to solve, in closed form, for the information-maximizing allocation of tuning curves as a function of the prior probability distribution of sensory variables. For the optimal population, the cell density is proportional to the prior, such that more cells with narrower tuning are allocated to encode higher-probability stimuli and that each cell transmits an equal portion of the stimulus probability mass. We also compute the stimulus discrimination capabilities of a perceptual system that relies on this neural representation and find that the best achievable discrimination thresholds are inversely proportional to the sensory prior. We examine how the prior information that is implicitly encoded in the tuning curves of the optimal population may be used for perceptual inference and derive a novel decoder, the Bayesian population vector, that closely approximates a Bayesian least-squares estimator that has explicit access to the prior. Finally, we generalize these results to sigmoidal tuning curves, correlated neural variability, and a broader class of objective functions. These results provide a principled embedding of sensory prior information in neural populations and yield predictions that are readily testable with environmental, physiological, and perceptual data.
Deep Ganguli, Eero P. Simoncelli
Neural Comput.2
2012 Hierarchical spike coding of sound
abstract
We develop a probabilistic generative model for representing acoustic event structure at multiple scales via a two-stage hierarchy. The first stage consists of a spiking representation which encodes a sound with a sparse set of kernels at different frequencies positioned precisely in time. The coarse time and frequency statistical structure of the first-stage spikes is encoded by a second stage spiking representation, while fine-scale statistical regularities are encoded by recurrent interactions within the first-stage. When fitted to speech data, the model encodes acoustic features such as harmonic stacks, sweeps, and frequency modulations, that can be composed to represent complex acoustic events. The model is also able to synthesize sounds from the higher-level representation and provides significant improvement over wavelet thresholding techniques on a denoising task.
Yan Karklin, Chaitanya Ekanadham, Eero P. Simoncelli
NIPS3
2012 Efficient and direct estimation of a neural subunit model for sensory coding
abstract
Many visual and auditory neurons have response properties that are well explained by pooling the rectified responses of a set of self-similar linear filters. These filters cannot be found using spike-triggered averaging (STA), which estimates only a single filter. Other methods, like spike-triggered covariance (STC), define a multi-dimensional response subspace, but require substantial amounts of data and do not produce unique estimates of the linear filters. Rather, they provide a linear basis for the subspace in which the filters reside. Here, we define a 'subunit' model as an LN-LN cascade, in which the first linear stage is restricted to a set of shifted ("convolutional") copies of a common filter, and the first nonlinear stage consists of rectifying nonlinearities that are identical for all filter outputs; we refer to these initial LN elements as the 'subunits' of the receptive field. The second linear stage then computes a weighted sum of the responses of the rectified subunits. We present a method for directly fitting this model to spike data. The method performs well for both simulated and real data (from primate V1), and the resulting model outperforms STA and STC in terms of both cross-validated accuracy and efficiency.
Brett Vintch, Andrew D. Zaharia, J. Anthony Movshon, Eero P. Simoncelli
NIPS4
2011 Sparse decomposition of transformation-invariant signals with continuous basis pursuit
abstract
Consider the decomposition of a signal into features that undergo transformations drawn from a continuous family. Current methods discretely sample the transformations and apply sparse recovery methods to the resulting finite dictionary. These methods do not exploit the underlying continuous structure, thereby limiting the ability to produce sparse solutions. Instead, we employ interpolation functions which linearly approximate the manifold of scaled and transformed features. Coefficients are interpreted as interpolation weights, and we formulate a convex optimization problem for obtaining them, enforcing both reconstruction accuracy and sparsity. We compare our method, which we call continuous basis pursuit (CBP) with the standard basis pursuit approach on a sparse deconvolution task. CBP yields substantially sparser solutions without sacrificing accuracy, and does so with a smaller dictionary. We conclude that for signals generated by transformation-invariant processes, a representation that explicitly accommodates the transformation(s) can yield sparser and more interpretable decompositions.
Chaitanya Ekanadham, Daniel Tranchina, Eero P. Simoncelli
ICASSP3
2011 A blind sparse deconvolution method for neural spike identification
abstract
We consider the problem of estimating neural spikes from extracellular voltage recordings. Most current methods are based on clustering, which requires substantial human supervision and produces systematic errors by failing to properly handle temporally overlapping spikes. We formulate the problem as one of statistical inference, in which the recorded voltage is a noisy sum of the spike trains of each neuron convolved with its associated spike waveform. Joint maximum-a-posteriori (MAP) estimation of the waveforms and spikes is then a blind deconvolution problem in which the coefficients are sparse. We develop a block-coordinate descent method for approximating the MAP solution. We validate our method on data simulated according to the generative model, as well as on real data for which ground truth is available via simultaneous intracellular recordings. In both cases, our method substantially reduces the number of missed spikes and false positives when compared to a standard clustering algorithm, primarily by recovering temporally overlapping spikes. The method offers a fully automated alternative to clustering methods that is less susceptible to systematic errors.
Chaitanya Ekanadham, Daniel Tranchina, Eero P. Simoncelli
NIPS3
2011 Efficient coding of natural images with a population of noisy Linear-Nonlinear neurons
abstract
Efficient coding provides a powerful principle for explaining early sensory coding. Most attempts to test this principle have been limited to linear, noiseless models, and when applied to natural images, have yielded oriented filters consistent with responses in primary visual cortex. Here we show that an efficient coding model that incorporates biologically realistic ingredients input and output noise, nonlinear response functions, and a metabolic cost on the firing rate predicts receptive fields and response nonlinearities similar to those observed in the retina. Specifically, we develop numerical methods for simultaneously learning the linear filters and response nonlinearities of a population of model neurons, so as to maximize information transmission subject to metabolic costs. When applied to an ensemble of natural images, the method yields filters that are center-surround and nonlinearities that are rectifying. The filters are organized into two populations, with On- and Off-centers, which independently tile the visual space. As observed in the primate retina, the Off-center neurons are more numerous and have filters with smaller spatial extent. In the absence of noise, our method reduces to a generalized version of independent components analysis, with an adapted nonlinear "contrast" function; in this case, the optimal filters are localized and oriented.
Yan Karklin, Eero P. Simoncelli
NIPS2
2011 Least Squares Estimation Without Priors or Supervision
abstract
Selection of an optimal estimator typically relies on either supervised training samples (pairs of measurements and their associated true values) or a prior probability model for the true values. Here, we consider the problem of obtaining a least squares estimator given a measurement process with known statistics (i.e., a likelihood function) and a set of unsupervised measurements, each arising from a corresponding true value drawn randomly from an unknown distribution. We develop a general expression for a nonparametric empirical Bayes least squares (NEBLS) estimator, which expresses the optimal least squares estimator in terms of the measurement density, with no explicit reference to the unknown (prior) density. We study the conditions under which such estimators exist and derive specific forms for a variety of different measurement processes. We further show that each of these NEBLS estimators may be used to express the mean squared estimation error as an expectation over the measurement density alone, thus generalizing Stein's unbiased risk estimator (SURE), which provides such an expression for the additive gaussian noise case. This error expression may then be optimized over noisy measurement samples, in the absence of supervised training data, yielding a generalized SURE-optimized parametric least squares (SURE2PLS) estimator. In the special case of a linear parameterization (i.e., a sum of nonlinear kernel functions), the objective function is quadratic, and we derive an incremental form for learning this estimator from data. We also show that combining the NEBLS form with its corresponding generalized SURE expression produces a generalization of the score-matching procedure for parametric density estimation. Finally, we have implemented several examples of such estimators, and we show that their performance is comparable to their optimal Bayesian or supervised regression counterparts for moderate to large amounts of data.
Martin Raphan, Eero P. Simoncelli
Neural Comput.2
2010 Implicit encoding of prior probabilities in optimal neural populations
abstract
Optimal coding provides a guiding principle for understanding the representation of sensory variables in neural populations. Here we consider the influence of a prior probability distribution over sensory variables on the optimal allocation of cells and spikes in a neural population. We model the spikes of each cell as samples from an independent Poisson process with rate governed by an associated tuning curve. For this response model, we approximate the Fisher information in terms of the density and amplitude of the tuning curves, under the assumption that tuning width varies inversely with cell density. We consider a family of objective functions based on the expected value, over the sensory prior, of a functional of the Fisher information. This family includes lower bounds on mutual information and perceptual discriminability as special cases. In all cases, we find a closed form expression for the optimum, in which the density and gain of the cells in the population are power law functions of the stimulus prior. This also implies a power law relationship between the prior and perceptual discriminability. We show preliminary evidence that the theory successfully predicts the relationship between empirically measured stimulus priors, physiologically measured neural response properties (cell density, tuning widths, and firing rates), and psychophysically measured discrimination thresholds.
Deep Ganguli, Eero P. Simoncelli
NIPS2
2009 Quantifying color image distortions based on adaptive spatio-chromatic signal decompositions
abstract
We describe a framework for quantifying color image distortion based on an adaptive signal decomposition. Specifically, local blocks of the image error are decomposed using a set of spatio-chromatic basis functions that are adapted to the spatial and color structure of the original image. The adaptive functions are chosen to isolate specific distortions such as luminance, hue, and saturation changes. These adaptive basis functions are used to augment a generic orthonormal basis, and the overall distortion is computed from the weighted sum of the coefficients of the resulting overcomplete decomposition, with smaller weights chosen for the adaptive terms. A set of preliminary experiments show that the proposed distortion measure is consistent with human perception of color images subjected to a variety of different common distortions. The framework may be easily extended to include any form of continuous spatio-chromatic distortion.
Umesh Rajashekar, Zhou Wang 0001, Eero P. Simoncelli
ICIP3
2009 Hierarchical Modeling of Local Image Features through $L_p$-Nested Symmetric Distributions
abstract
We introduce a new family of distributions, called $L_p${\em -nested symmetric distributions}, whose densities access the data exclusively through a hierarchical cascade of $L_p$-norms. This class generalizes the family of spherically and $L_p$-spherically symmetric distributions which have recently been successfully used for natural image modeling. Similar to those distributions it allows for a nonlinear mechanism to reduce the dependencies between its variables. With suitable choices of the parameters and norms, this family also includes the Independent Subspace Analysis (ISA) model, which has been proposed as a means of deriving filters that mimic complex cells found in mammalian primary visual cortex. $L_p$-nested distributions are easy to estimate and allow us to explore the variety of models between ISA and the $L_p$-spherically symmetric models. Our main findings are that, without a preprocessing step of contrast gain control, the independent subspaces of ISA are in fact more dependent than the individual filter coefficients within a subspace and, with contrast gain control, where ISA finds more than one subspace, the filter responses were almost independent anyway.
Fabian H. Sinz, Eero P. Simoncelli, Matthias Bethge
NIPS2
2009 Nonlinear Extraction of Independent Components of Natural Images Using Radial Gaussianization
abstract
We consider the problem of efficiently encoding a signal by transforming it to a new representation whose components are statistically independent. A widely studied linear solution, known as independent component analysis (ICA), exists for the case when the signal is generated as a linear transformation of independent nongaussian sources. Here, we examine a complementary case, in which the source is nongaussian and elliptically symmetric. In this case, no invertible linear transform suffices to decompose the signal into independent components, but we show that a simple nonlinear transformation, which we call radial gaussianization (RG), is able to remove all dependencies. We then examine this methodology in the context of natural image statistics. We first show that distributions of spatially proximal bandpass filter responses are better described as elliptical than as linearly transformed independent sources. Consistent with this, we demonstrate that the reduction in dependency achieved by applying RG to either nearby pairs or blocks of bandpass filter responses is significantly greater than that achieved by ICA. Finally, we show that the RG transformation may be closely approximated by divisive normalization, which has been used to model the nonlinear response properties of visual neurons.
Siwei Lyu, Eero P. Simoncelli
Neural Comput.2
2009 Is the Homunculus "Aware" of Sensory Adaptation?
abstract
Neural activity and perception are both affected by sensory history. The work presented here explores the relationship between the physiological effects of adaptation and their perceptual consequences. Perception is modeled as arising from an encoder-decoder cascade, in which the encoder is defined by the probabilistic response of a population of neurons, and the decoder transforms this population activity into a perceptual estimate. Adaptation is assumed to produce changes in the encoder, and we examine the conditions under which the decoder behavior is consistent with observed perceptual effects in terms of both bias and discriminability. We show that for all decoders, discriminability is bounded from below by the inverse Fisher information. Estimation bias, on the other hand, can arise for a variety of different reasons and can range from zero to substantial. We specifically examine biases that arise when the decoder is fixed, "unaware" of the changes in the encoding population (as opposed to "aware" of the adaptation and changing accordingly). We simulate the effects of adaptation on two well-studied sensory attributes, motion direction and contrast, assuming a gain change description of encoder adaptation. Although we cannot uniquely constrain the source of decoder bias, we find for both motion and contrast that an "unaware" decoder that maximizes the likelihood of the percept given by the preadaptation encoder leads to predictions that are consistent with behavioral data. This model implies that adaptation-induced biases arise as a result of temporary suboptimality of the decoder.
Peggy Seriès, Alan A. Stocker, Eero P. Simoncelli
Neural Comput.3
2009 Modeling Multiscale Subbands of Photographic Images with Fields of Gaussian Scale Mixtures
abstract
The local statistical properties of photographic images, when represented in a multi-scale basis, have been described using Gaussian scale mixtures. Here, we use this local description as a substrate for constructing a global field of Gaussian scale mixtures (FoGSMs). Specifically, we model multi-scale subbands as a product of an exponentiated homogeneous Gaussian Markov random field (hGMRF) and a second independent hGMRF. We show that parameter estimation for this model is feasible, and that samples drawn from a FoGSM model have marginal and joint statistics similar to subband coefficients of photographic images. We develop an algorithm for removing additive Gaussian white noise based on the FoGSM model, and demonstrate denoising performance comparable with state-of-the-art methods.
Siwei Lyu, Eero P. Simoncelli
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Nonlinear image representation using divisive normalization
abstract
In this paper, we describe a nonlinear image representation based on divisive normalization that is designed to match the statistical properties of photographic images, as well as the perceptual sensitivity of biological visual systems. We decompose an image using a multi-scale oriented representation, and use Student's t as a model of the dependencies within local clusters of coefficients. We then show that normalization of each coefficient by the square root of a linear combination of the amplitudes of the coefficients in the cluster reduces statistical dependencies. We further show that the resulting divisive normalization transform is invertible and provide an efficient iterative inversion algorithm. Finally, we probe the statistical and perceptual advantages of this image representation by examining its robustness to added noise, and using it to enhance image contrast.
Siwei Lyu, Eero P. Simoncelli
CVPR2
2008 Contextually adaptive signal representation using conditional principal component analysis
abstract
The conventional method of generating a basis that is optimally adapted (in MSE) for representation of an ensemble of signals is Principal Component Analysis (PCA). A more ambitious modern goal is the construction of bases that are adapted to individual signal instances. Here we develop a new framework for instance-adaptive signal representation by exploiting the fact that many real-world signals exhibit local self-similarity. Specifically, we decompose the signal into multiscale subbands, and then represent local blocks of each subband using basis functions that are linearly derived from the surrounding context. The linear mappings that generate these basis functions are learned sequentially, with each one optimized to account for as much variance as possible in the local blocks. We apply this methodology to learning a coarse-to-fine representation of images within a multi-scale basis, demonstrating that the adaptive basis can account for significantly more variance than a PCA basis of the same dimensionality.
Rosa M. Figueras i Ventura, Umesh Rajashekar, Zhou Wang 0001, Eero P. Simoncelli
ICASSP4
2008 Image denoising using mixtures of Gaussian scale mixtures
abstract
The local statistical properties of photographic images, when represented in a multi-scale basis, have been described using Gaussian scale mixtures (GSMs). In that model, each spatial neighborhood of coefficients is described as a Gaussian random vector modulated by a random hidden positive scaling variable. Here, we introduce a more powerful model in which neighborhoods of each subband are described as a finite mixture of GSMs. We develop methods to learn the mixing densities and covariance matrices associated with each of the GSM components from a single image, and show that this process naturally segments the image into regions of similar content. The model parameters can also be learned in the presence of additive Gaussian noise, and the resulting fitted model may be used as a prior for Bayesian noise removal. Simulations demonstrate this model substantially outperforms the original GSM model.
Jose A. Guerrero-Colon, Eero P. Simoncelli, Javier Portilla
ICIP2
2008 Reducing statistical dependencies in natural signals using radial Gaussianization
abstract
We consider the problem of efficiently encoding a signal by transforming it to a new representation whose components are statistically independent. A widely studied linear solution, independent components analysis (ICA), exists for the case when the signal is generated as a linear transformation of independent non- Gaussian sources. Here, we examine a complementary case, in which the source is non-Gaussian but elliptically symmetric. In this case, no linear transform suffices to properly decompose the signal into independent components, but we show that a simple nonlinear transformation, which we call radial Gaussianization (RG), is able to remove all dependencies. We then demonstrate this methodology in the context of natural signal statistics. We first show that the joint distributions of bandpass filter responses, for both sound and images, are better described as elliptical than linearly transformed independent sources. Consistent with this, we demonstrate that the reduction in dependency achieved by applying RG to either pairs or blocks of bandpass filter responses is significantly greater than that achieved by PCA or ICA.
Siwei Lyu, Eero P. Simoncelli
NIPS2
2008 Image Modeling and Denoising With Orientation-Adapted Gaussian Scale Mixtures
abstract
We develop a statistical model to describe the spatially varying behavior of local neighborhoods of coefficients in a multiscale image representation. Neighborhoods are modeled as samples of a multivariate Gaussian density that are modulated and rotated according to the values of two hidden random variables, thus allowing the model to adapt to the local amplitude and orientation of the signal. A third hidden variable selects between this oriented process and a nonoriented scale mixture of Gaussians process, thus providing adaptability to the local orientedness of the signal. Based on this model, we develop an optimal Bayesian least squares estimator for denoising images and show through simulations that the resulting method exhibits significant improvement over previously published results obtained with Gaussian scale mixtures.
David K. Hammond, Eero P. Simoncelli
IEEE Trans. Image Process.2
2008 Optimal Denoising in Redundant Representations
abstract
Image denoising methods are often designed to minimize mean-squared error (MSE) within the subbands of a multiscale decomposition. However, most high-quality denoising results have been obtained with overcomplete representations, for which minimization of MSE in the subband domain does not guarantee optimal MSE performance in the image domain. We prove that, despite this suboptimality, the expected image-domain MSE resulting from applying estimators to subbands that are made redundant through spatial replication of basis functions (e.g., cycle spinning) is always less than or equal to that resulting from applying the same estimators to the original nonredundant representation. In addition, we show that it is possible to further exploit overcompleteness by jointly optimizing the subband estimators for image-domain MSE. We develop an extended version of Stein's unbiased risk estimate (SURE) that allows us to perform this optimization adaptively, for each observed noisy image. We demonstrate this methodology using a new class of estimator formed from linear combinations of localized "bump" functions that are applied either pointwise or on local neighborhoods of subband coefficients. We show through simulations that the performance of these estimators applied to overcomplete subbands and optimized for image-domain MSE is substantially better than that obtained when they are optimized within each subband. This performance is, in turn, substantially better than that obtained when they are optimized for use on a nonredundant representation.
Martin Raphan, Eero P. Simoncelli
IEEE Trans. Image Process.2
2007 A Machine Learning Framework for Adaptive Combination of Signal Denoising Methods
abstract
We present a general framework for combination of two distinct local denoising methods. Interpolation between the two methods is controlled by a spatially varying decision function. Assuming the availability of clean training data, we formulate a learning problem for determining the decision function. As an example application we use Weighted Kernel Ridge Regression to solve this learning problem for a pair of wavelet-based image denoising algorithms, yielding a "hybrid" denoising algorithm whose performance surpasses that of either initial method.
David K. Hammond, Eero P. Simoncelli
ICIP (6)2
2007 Optimal Denoising in Redundant Bases
abstract
Image denoising methods are often based on estimators chosen to minimize mean squared error (MSE) within the sub-bands of a multi-scale decomposition. But this does not guarantee optimal MSE performance in the image domain, unless the decomposition is orthonormal. We prove that despite this suboptimality, the expected image-domain MSE resulting from a representation that is made redundant through spatial replication of basis functions (e.g., cycle-spinning) is less than or equal to that resulting from the original non-redundant representation. We also develop an extension of Stein's unbiased risk estimator (SURE) that allows minimization of the image-domain MSE for estimators that operate on subbands of a redundant decomposition. We implement an example, jointly optimizing the parameters of scalar estimators applied to each subband of an overcomplete representation, and demonstrate substantial MSE improvement over the sub-optimal application of SURE within individual subbands.
Martin Raphan, Eero P. Simoncelli
ICIP (3)2
2007 Statistically Driven Sparse Image Approximation
abstract
Finding the sparsest approximation of an image as a sum of basis functions drawn from a redundant dictionary is an NP-hard problem. In the case of a dictionary whose elements form an overcomplete basis, a recently developed method, based on alternating thresholding and projection operations, provides an appealing approximate solution. When applied to images, this method produces sparser results and requires less computation than current alternative methods. Motivated by recent developments in statistical image modeling, we develop an enhancement of this method based on a locally adaptive threshold operation, and demonstrate that the enhanced algorithm is capable of finding sparser approximations with a decrease in computational complexity.
Rosa M. Figueras i Ventura, Eero P. Simoncelli
ICIP (1)2
2007 A Bayesian Model of Conditioned Perception
abstract
We propose an extended probabilistic model for human perception. We argue that in many circumstances, human observers simultaneously evaluate sensory evidence under different hypotheses regarding the underlying physical process that might have generated the sensory information. Within this context, inference can be optimal if the observer weighs each hypothesis according to the correct belief in that hypothesis. But if the observer commits to a particular hypothesis, the belief in that hypothesis is converted into subjective certainty, and subsequent perceptual behavior is suboptimal, conditioned only on the chosen hypothesis. We demonstrate that this framework can explain psychophysical data of a recently reported decision-estimation experiment. The model well accounts for the data, predicting the same estimation bias as a consequence of the preceding decision step. The power of the framework is that it has no free parameters except the degree of the observer's uncertainty about its internal sensory representation. All other parameters are defined by the particular experiment which allows us to make quantitative predictions of human perception to two modifications of the original experiment.
Alan A. Stocker, Eero P. Simoncelli
NIPS2
2006 Image Denoising with an Orientation-Adaptive Gaussian Scale Mixture Model
abstract
We develop a statistical model for images that explicitly captures variations in local orientation and contrast. Patches of wavelet coefficients are described as samples of a fixed Gaussian process that are rotated and scaled according to a set of hidden variables representing the local image contrast and orientation. An optimal Bayesian least squares estimator is developed by conditioning upon and integrating over the hidden orientation and scale variables. The resulting denoising procedure gives results that are visually superior to those obtained with a Gaussian scale mixture model that does not explicitly incorporate local image orientation.
David K. Hammond, Eero P. Simoncelli
ICIP2
2006 Statistical Modeling of Images with Fields of Gaussian Scale Mixtures
abstract
The local statistical properties of photographic images, when represented in a multi-scale basis, have been described using Gaussian scale mixtures (GSMs). Here, we use this local description to construct a global field of Gaussian scale mixtures (FoGSM). Specifically, we model subbands of wavelet coefficients as a product of an exponentiated homogeneous Gaussian Markov random field (hGMRF) and a second independent hGMRF. We show that parameter estimation for FoGSM is feasible, and that samples drawn from an estimated FoGSM model have marginal and joint statistics similar to wavelet coefficients of photographic images. We develop an algorithm for image denoising based on the FoGSM model, and demonstrate substantial improvements over current state-ofthe-art denoising method based on the local GSM model. Many successful methods in image processing and computer vision rely on statistical models for images, and it is thus of continuing interest to develop improved models, both in terms of their ability to precisely capture image structures, and in terms of their tractability when used in applications. Constructing such a model is difficult, primarily because of the intrinsic high dimensionality of the space of images. Two simplifying assumptions are usually made to reduce model complexity. The first is Markovianity: the density of a pixel conditioned on a small neighborhood, is assumed to be independent from the rest of the image. The second assumption is homogeneity: the local density is assumed to be independent of its absolute position within the image. The set of models satisfying both of these assumptions constitute the class of homogeneous Markov random fields (hMRFs). Over the past two decades, studies of photographic images represented with multi-scale multiorientation image decompositions (loosely referred to as "wavelets") have revealed striking nonGaussian regularities and inter and intra-subband dependencies. For instance, wavelet coefficients generally have highly kurtotic marginal distributions [1, 2], and their amplitudes exhibit strong correlations with the amplitudes of nearby coefficients [3, 4]. One model that can capture the nonGaussian marginal behaviors is a product of non-Gaussian scalar variables [5]. A number of authors have developed non-Gaussian MRF models based on this sort of local description [6, 7, 8], among which the recently developed fields of experts model [7] has demonstrated impressive performance in denoising (albeit at an extremely high computational cost in learning model parameters). An alternative model that can capture non-Gaussian local structure is a scale mixture model [9, 10, 11]. An important special case is Gaussian scale mixtures (GSM), which consists of a Gaussian random vector whose amplitude is modulated by a hidden scaling variable. The GSM model provides a particularly good description of local image statistics, and the Gaussian substructure of the model leads to efficient algorithms for parameter estimation and inference. Local GSM-based methods represent the current state-of-the-art in image denoising [12]. The power of GSM models should be substantially improved when extended to describe more than a small neighborhood of wavelet coefficients. To this end, several authors have embedded local Gaussian mixtures into tree-structured MRF models [e.g., 13, 14]. In order to maintain tractability, these models are arranged such that coefficients are grouped in non-overlapping clusters, allowing a graphical probability model with no loops. Despite their global consistency, the artificially imposed cluster boundaries lead to substantial artifacts in applications such as denoising. In this paper, we use a local GSM as a basis for a globally consistent and spatially homogeneous field of Gaussian scale mixtures (FoGSM). Specifically, the FoGSM is formulated as the product of two mutually independent MRFs: a positive multiplier field obtained by exponentiating a homogeneous Gaussian MRF (hGMRF), and a second hGMRF. We develop a parameter estimation procedure, and show that the model is able to capture important statistical regularities in the marginal and joint wavelet statistics of a photographic image. We apply the FoGSM to image denoising, demonstrating substantial improvement over the previous state-of-the-art results obtained with a local GSM model. 1 Gaussian scale mixtures A GSM random vector x is formed as the product of a zero-mean Gaussian random vector u and an d d independent random variable z, as x = zu, where = denotes equality in distribution. The density of x is determined by the covariance of the Gaussian vector, , and the density of the multiplier, p z (z), through the integral - T -1 p z z xx 1 exp (1) p(x) = Nx (0, z) pz (z)dz z (z)d z. 2z z|| A key property of GSMs is that when z determines the scale of the conditional variance of x given z, wich is a Gaussian variable with zero mean and covariance z. In addition, the normalized variable h x z is a zero mean Gaussian with covariance matrix . The GSM model has been used to describe the marginal and joint densities of local clusters of wavelet coefficients, both within and across subbands [9], where the embedded Gaussian structure affords simple and efficient computation. This local GSM model has been be used for denoising, by independently estimating each coefficient conditioned on its surrounding cluster [12]. This method achieves state-of-the-art performances, despite the fact that treating overlapping clusters as independent does not give rise to a globally consistent statistical model that satisfies all the local constraints. 2 Fields of Gaussian scale mixtures In this section, we develop fields of Gaussian scale mixtures (FoGSM) as a framework for modeling wavelet coefficients of photographic images. Analogous to the local GSM model, we use a latent multiplier field to modulate a homogeneous Gaussian MRF (hGMRF). Formally, we define a FoGSM x as the product of two mutually independent MRFs, d x = u z, (2) where u is a zero-mean hGMRF, and z is a field of positive multipliers that control the local coefficient variances. The operator denotes element-wise multiplication, and the square root operation is applied to each component. Note that x has a one-dimensional GSM marginal distributions, while its components have dependencies captured by the MRF structures of u and z. Analogous to the local GSM, when conditioned on z, x is an inhomogeneous GMRF | | - - -1 x = 1 T D -1 1 Qu | Qu | p(x|z) x (x z)T Qu (x i exp z Qu D z i exp zi 2 zi 2 , z) (3) where Qu is the inverse covariance matrix of u (also known as the precision matrix), and D() denotes the operator that form a diagonal matrix from an input vector. Note also that the elementwise division of the two fields, x z, yields a hGMRF with precision matrix Q u . To complete the FoGSM model, we need to specify the structure of the multiplier field z. For tractability, we use another hGMRF as a substrate, and map it into positive values by exponentiation,
Siwei Lyu, Eero P. Simoncelli
NIPS2
2006 Learning to be Bayesian without Supervision
Martin Raphan, Eero P. Simoncelli
NIPS2
2006 Nonlinear image representation for efficient perceptual coding
abstract
Image compression systems commonly operate by transforming the input signal into a new representation whose elements are independently quantized. The success of such a system depends on two properties of the representation. First, the coding rate is minimized only if the elements of the representation are statistically independent. Second, the perceived coding distortion is minimized only if the errors in a reconstructed image arising from quantization of the different elements of the representation are perceptually independent. We argue that linear transforms cannot achieve either of these goals and propose, instead, an adaptive nonlinear image representation in which each coefficient of a linear transform is divided by a weighted sum of coefficient amplitudes in a generalized neighborhood. We then show that the divisive operation greatly reduces both the statistical and the perceptual redundancy amongst representation elements. We develop an efficient method of inverting this transformation, and we demonstrate through simulations that the dual reduction in dependency can greatly improve the visual quality of compressed images.
Jesús Malo, Irene Epifanio, Rafael Fonolla Navarro, Eero P. Simoncelli
IEEE Trans. Image Process.4
2006 Quality-aware images
abstract
We propose the concept of quality-aware image, in which certain extracted features of the original (high-quality) image are embedded into the image data as invisible hidden messages. When a distorted version of such an image is received, users can decode the hidden messages and use them to provide an objective measure of the quality of the distorted image. To demonstrate the idea, we build a practical quality-aware image encoding, decoding and quality analysis system, which employs: 1) a novel reduced-reference image quality assessment algorithm based on a statistical model of natural images and 2) a previously developed quantization watermarking-based data hiding technique in the wavelet transform domain.
Zhou Wang 0001, Guixing Wu, Hamid R. Sheikh, Eero P. Simoncelli, En-Hui Yang, Alan C. Bovik
IEEE Trans. Image Process.4
2005 Translation Insensitive Image Similarity in Complex Wavelet Domain
abstract
We propose a complex wavelet domain image similarity measure, which is simultaneously insensitive to luminance change, contrast change and spatial translation. The key idea is to make use of the fact that these image distortions lead to consistent magnitude and/or phase changes of local wavelet coefficients. Since small scaling and rotation of images can be locally approximated by translation, the proposed measure also shows robustness to spatial scaling and rotation when these geometric distortions are small relative to the size of the wavelet filters. Compared with previous methods, the proposed measure is computationally efficient, and can evaluate the similarity of two images without a precise registration process at the front end.
Zhou Wang 0001, Eero P. Simoncelli
ICASSP (2)2
2005 Locally adaptive multiscale contrast optimization
abstract
We describe a method for automatically and adaptively boosting the visibility of local features in an image. A log intensity image is first decomposed into a set of subbands at multiple scales and orientations. Operating successively from coarse frequency bands to fine, the coefficients of each subband are adjusted so as to move their locally averaged amplitudes toward a target value using a gamma operation. Target values are chosen to fall linearly over scale, consistent with a scale-invariant spectral model. To avoid enlarging the range of image intensity values, in those locations where the local mean is near the minimal or maximal values of the image and the local contrast is being boosted significantly, the local mean is moved toward the global mean. Finally, a spatial mask is applied in the pixel domain to ensure that the enhancements are applied only in the vicinity of image features. The resulting image appears to be both sharper and of higher contrast.
Nicolas Bonnier, Eero P. Simoncelli
ICIP (1)2
2005 An adaptive linear system framework for image distortion analysis
abstract
We describe a framework for decomposing the distortion between two images into a linear combination of components. Unlike conventional linear bases such as those in Fourier or wavelet decompositions, a subset of the components in our representation are not fixed, but are adaptively computed from the input images. We show that this framework is a generalization of a number of existing image comparison approaches. As an example of a specific implementation, we select the components based on the structural similarity principle, separating the overall image distortions into non-structural distortions (those that do not change the structures of the objects in the scene) and the remaining structural distortions. We demonstrate that the resulting measure is effective in predicting image distortions as perceived by human observers.
Zhou Wang 0001, Eero P. Simoncelli
ICIP (3)2
2005 Sensory Adaptation within a Bayesian Framework for Perception
abstract
We extend a previously developed Bayesian framework for perception to account for sensory adaptation. We first note that the perceptual ef- fects of adaptation seems inconsistent with an adjustment of the inter- nally represented prior distribution. Instead, we postulate that adaptation increases the signal-to-noise ratio of the measurements by adapting the operational range of the measurement stage to the input range. We show that this changes the likelihood function in such a way that the Bayesian estimator model can account for reported perceptual behavior. In particu- lar, we compare the model’s predictions to human motion discrimination data and demonstrate that the model accounts for the commonly observed perceptual adaptation effects of repulsion and enhanced discriminability.
Alan A. Stocker, Eero P. Simoncelli
NIPS2
2005 Comparing integrate-and-fire models estimated using intracellular and extracellular data
Liam Paninski, Jonathan W. Pillow, Eero P. Simoncelli
Neurocomputing3
2004 Constraining a Bayesian Model of Human Visual Speed Perception
abstract
It has been demonstrated that basic aspects of human visual motion per- ception are qualitatively consistent with a Bayesian estimation frame- work, where the prior probability distribution on velocity favors slow speeds. Here, we present a refined probabilistic model that can account for the typical trial-to-trial variabilities observed in psychophysical speed perception experiments. We also show that data from such experiments can be used to constrain both the likelihood and prior functions of the model. Specifically, we measured matching speeds and thresholds in a two-alternative forced choice speed discrimination task. Parametric fits to the data reveal that the likelihood function is well approximated by a LogNormal distribution with a characteristic contrast-dependent vari- ance, and that the prior distribution on velocity exhibits significantly heavier tails than a Gaussian, and approximately follows a power-law function. Humans do not perceive visual motion veridically. Various psychophysical experiments have shown that the perceived speed of visual stimuli is affected by stimulus contrast, with low contrast stimuli being perceived to move slower than high contrast ones [1, 2]. Computational models have been suggested that can qualitatively explain these perceptual effects. Commonly, they assume the perception of visual motion to be optimal either within a deterministic framework with a regularization constraint that biases the solution toward zero motion [3, 4], or within a probabilistic framework of Bayesian estimation with a prior that favors slow velocities [5, 6]. The solutions resulting from these two frameworks are similar (and in some cases identi- cal), but the probabilistic framework provides a more principled formulation of the problem in terms of meaningful probabilistic components. Specifically, Bayesian approaches rely on a likelihood function that expresses the relationship between the noisy measurements and the quantity to be estimated, and a prior distribution that expresses the probability of encountering any particular value of that quantity. A probabilistic model can also provide a richer description, by defining a full probability density over the set of possible "percepts", rather than just a single value. Numerous analyses of psychophysical experiments have made use of such distributions within the framework of signal detection theory in order to model perceptual behavior [7]. Previous work has shown that an ideal Bayesian observer model based on Gaussian forms high contrast low contrast y y posterior likelihood y densit y densit posterior likelihood obabilit prior prior obabilit pr pr v^ v^ a visual speed b visual speed Figure 1: Bayesian model of visual speed perception. a) For a high contrast stimulus, the likelihood has a narrow width (a high signal-to-noise ratio) and the prior induces only a small shift of the mean ^ v of the posterior. b) For a low contrast stimuli, the measurement is noisy, leading to a wider likelihood. The shift is much larger and the perceived speed lower than under condition (a). for both likelihood and prior is sufficient to capture the basic qualitative features of global translational motion perception [5, 6]. But the behavior of the resulting model deviates systematically from human perceptual data, most importantly with regard to trial-to-trial variability and the precise form of interaction between contrast and perceived speed. A recent article achieved better fits for the model under the assumption that human contrast perception saturates [8]. In order to advance the theory of Bayesian perception and provide significant constraints on models of neural implementation, it seems essential to constrain quantitatively both the likelihood function and the prior probability distribution. In previous work, the proposed likelihood functions were derived from the brightness constancy con- straint [5, 6] or other generative principles [9]. Also, previous approaches defined the prior distribution based on general assumptions and computational convenience, typically choos- ing a Gaussian with zero mean, although a Laplacian prior has also been suggested [4]. In this paper, we develop a more general form of Bayesian model for speed perception that can account for trial-to-trial variability. We use psychophysical speed discrimination data in order to constrain both the likelihood and the prior function. 1 Probabilistic Model of Visual Speed Perception 1.1 Ideal Bayesian Observer Assume that an observer wants to obtain an estimate for a variable v based on a measure- ment m that she/he performs. A Bayesian observer "knows" that the measurement device is not ideal and therefore, the measurement m is affected by noise. Hence, this observer combines the information gained by the measurement m with a priori knowledge about v. Doing so (and assuming that the prior knowledge is valid), the observer will on average perform better in estimating v than just trusting the measurements m. According to Bayes' rule 1 p(v|m) = p(m|v)p(v) (1) the probability of perceiving v given m (posterior) is the product of the likelihood of v for a particular measurements m and the a priori knowledge about the estimated variable v (prior). is a normalization constant independent of v that ensures that the posterior is a proper probability distribution. 1 Pcum=0.875 )1^ P + cum=0.5 > v2^ P(v 0 v2 a b vmatch vthres Figure 2: 2AFC speed discrimination experiment. a) Two patches of drifting gratings were displayed simultaneously (motion without movement). The subject was asked to fixate the center cross and decide after the presentation which of the two gratings was moving faster. b) A typical psychometric curve obtained under such paradigm. The dots represent the empirical probability that the subject perceived stimulus2 moving faster than stimulus1. The speed of stimulus1 was fixed while v2 is varied. The point of subjective equality, vmatch, is the value of v2 for which Pcum = 0.5. The threshold velocity vthresh is the velocity for which Pcum = 0.875. It is important to note that the measurement m is an internal variable of the observer and is not necessarily represented in the same space as v. The likelihood embodies both the mapping from v to m and the noise in this mapping. So far, we assume that there is a monotonic function f (v) : v vm that maps v into the same space as m (m-space). Doing so allows us to analytically treat m and vm in the same space. We will later propose a suitable form of the mapping function f (v). An ideal Bayesian observer selects the estimate that minimizes the expected loss, given the posterior and a loss function. We assume a least-squares loss function. Then, the optimal estimate ^ v is the mean of the posterior in Equation (1). It is easy to see why this model of a Bayesian observer is consistent with the fact that perceived speed decreases with con- trast. The width of the likelihood varies inversely with the accuracy of the measurements performed by the observer, which presumably decreases with decreasing contrast due to a decreasing signal-to-noise ratio. As illustrated in Figure 1, the shift in perceived speed towards slow velocities grows with the width of the likelihood, and thus a Bayesian model can qualitatively explain the psychophysical results [1]. 1.2 Two Alternative Forced Choice Experiment We would like to examine perceived speeds under a wide range of conditions in order to constrain a Bayesian model. Unfortunately, perceived speed is an internal variable, and it is not obvious how to design an experiment that would allow subjects to express it directly 1. Perceived speed can only be accessed indirectly by asking the subject to compare the speed of two stimuli. For a given trial, an ideal Bayesian observer in such a two-alternative forced choice (2AFC) experimental paradigm simply decides on the basis of the two trial estimates ^v1 (stimulus1) and ^v2 (stimulus2) which stimulus moves faster. Each estimate ^v is based on a particular measurement m. For a given stimulus with speed v, an ideal Bayesian observer will produce a distribution of estimates p(^ v|v) because m is noisy. Over trials, the observers behavior can be described by classical signal detection theory based on the distributions of the estimates, hence e.g. the probability of perceiving stimulus2 moving 1Although see [10] for an example of determining and even changing the prior of a Bayesian model for a sensorimotor task, where the estimates are more directly accessible. faster than stimulus1 is given as the cumulative probability ^v2 Pcum(^ v2 > ^v1) = p(^ v2|v2) p(^ v1|v1) d^v1 d^v2 (2) 0 0 Pcum describes the full psychometric curve. Figure 2b illustrates the measured psychomet- ric curve and its fit from such an experimental situation. 2 Experimental Methods We measured matching speeds (Pcum = 0.5) and thresholds (Pcum = 0.875) in a 2AFC speed discrimination task. Subjects were presented simultaneously with two circular patches of horizontally drifting sine-wave gratings for the duration of one second (Fig- ure 2a). Patches were 3deg in diameter, and were displayed at 6deg eccentricity to either side of a fixation cross. The stimuli had an identical spatial frequency of 1.5 cycle/deg. One stimulus was considered to be the reference stimulus having one of two different contrast values (c1=[0.075 0.5]) and one of five different speed values (u1=[1 2 4 8 12] deg/sec) while the second stimulus (test) had one of five different contrast values (c2=[0.05 0.1 0.2 0.4 0.8]) and a varying speed that was determined by an interleaved staircase procedure. For each condition there were 96 trials. Conditions were randomly interleaved, including a random choice of stimulus identity (test vs. reference) and motion direction (right vs. left). Subjects were asked to fixate during stimulus presentation and select the faster mov- ing stimulus. The threshold experiment differed only in that auditory feedback was given to indicate the correctness of their decision. This did not change the outcome of the ex- periment but increased significantly the quality of the data and thus reduced the number of trials needed.
Alan A. Stocker, Eero P. Simoncelli
NIPS2
2004 Machine Learning Applied to Perception: Decision Images for Gender Classification
abstract
We study gender discrimination of human faces using a combination of psychophysical classification and discrimination experiments together with methods from machine learning. We reduce the dimensionality of a set of face images using principal component analysis, and then train a set of linear classifiers on this reduced representation (linear support vec- tor machines (SVMs), relevance vector machines (RVMs), Fisher linear discriminant (FLD), and prototype (prot) classifiers) using human clas- sification data. Because we combine a linear preprocessor with linear classifiers, the entire system acts as a linear classifier, allowing us to visu- alise the decision-image corresponding to the normal vector of the separ- ating hyperplanes (SH) of each classifier. We predict that the female-to- maleness transition along the normal vector for classifiers closely mim- icking human classification (SVM and RVM [1]) should be faster than the transition along any other direction. A psychophysical discrimina- tion experiment using the decision images as stimuli is consistent with this prediction.
Felix A. Wichmann, Arnulf B. A. Graf, Eero P. Simoncelli, Heinrich H. Bülthoff, Bernhard Schölkopf
NIPS3
2004 Spike-triggered characterization of excitatory and suppressive stimulus dimensions in monkey V1
Nicole C. Rust, Odelia Schwartz, J. Anthony Movshon, Eero P. Simoncelli
Neurocomputing4
2004 Maximum Likelihood Estimation of a Stochastic Integrate-and-Fire Neural Encoding Model
abstract
We examine a cascade encoding model for neural response in which a linear filtering stage is followed by a noisy, leaky, integrate-and-fire spike generation mechanism. This model provides a biophysically more realistic alternative to models based on Poisson (memoryless) spike generation, and can effectively reproduce a variety of spiking behaviors seen in vivo. We describe the maximum likelihood estimator for the model parameters, given only extracellular spike train responses (not intracellular voltage data). Specifically, we prove that the log-likelihood function is concave and thus has an essentially unique global maximum that can be found using gradient ascent techniques. We develop an efficient algorithm for computing the maximum likelihood solution, demonstrate the effectiveness of the resulting estimator with numerical simulations, and discuss a method of testing the model's validity using time-rescaling and density evolution techniques.
Liam Paninski, Jonathan W. Pillow, Eero P. Simoncelli
Neural Comput.3
2004 Differentiation of discrete multidimensional signals
abstract
We describe the design of finite-size linear-phase separable kernels for differentiation of discrete multidimensional signals. The problem is formulated as an optimization of the rotation-invariance of the gradient operator, which results in a simultaneous constraint on a set of one-dimensional low-pass prefilter and differentiator filters up to the desired order. We also develop extensions of this formulation to both higher dimensions and higher order directional derivatives. We develop a numerical procedure for optimizing the constraint, and demonstrate its use in constructing a set of example filters. The resulting filters are significantly more accurate than those commonly used in the image and multidimensional signal processing literature.
Hany Farid, Eero P. Simoncelli
IEEE Trans. Image Process.2
2004 Image quality assessment: from error visibility to structural similarity
abstract
Objective methods for assessing perceptual image quality traditionally attempted to quantify the visibility of errors (differences) between a distorted image and a reference image using a variety of known properties of the human visual system. Under the assumption that human visual perception is highly adapted for extracting structural information from a scene, we introduce an alternative complementary framework for quality assessment based on the degradation of structural information. As a specific example of this concept, we develop a Structural Similarity Index and demonstrate its promise through a set of intuitive examples, as well as comparison to both subjective ratings and state-of-the-art objective methods on a database of images compressed with JPEG and JPEG2000.
Zhou Wang 0001, Alan C. Bovik, Hamid R. Sheikh, Eero P. Simoncelli
IEEE Trans. Image Process.4
2003 Image restoration using Gaussian scale mixtures in the wavelet domain
abstract
A statistical model for images decomposed in an overcomplete wavelet pyramid is described. Each neighborhood of pyramid coefficients is modeled as the product of a Gaussian vector of known covariance, and an independent hidden positive scalar random variable. We propose an efficient Bayesian estimator for the pyramid coefficients of an image degraded by linear distortion (e.g., blur) and additive Gaussian noise. We demonstrate the quality of our results in simulations over a wide range of blur and noise levels.
Javier Portilla, Eero P. Simoncelli
ICIP (2)2
2003 Maximum Likelihood Estimation of a Stochastic Integrate-and-Fire Neural Model
abstract
Recent work has examined the estimation of models of stimulus-driven neural activity in which some linear filtering process is followed by a nonlinear, probabilistic spiking stage. We analyze the estimation of one such model for which this nonlinear step is implemented by a noisy, leaky, integrate-and-fire mechanism with a spike-dependent after- current. This model is a biophysically plausible alternative to models with Poisson (memory-less) spiking, and has been shown to effectively reproduce various spiking statistics of neurons in vivo. However, the problem of estimating the model from extracellular spike train data has not been examined in depth. We formulate the problem in terms of max- imum likelihood estimation, and show that the computational problem of maximizing the likelihood is tractable. Our main contribution is an algorithm and a proof that this algorithm is guaranteed to find the global optimum with reasonable speed. We demonstrate the effectiveness of our estimator with numerical simulations. A central issue in computational neuroscience is the characterization of the functional re- lationship between sensory stimuli and neural spike trains. A common model for this re- lationship consists of linear filtering of the stimulus, followed by a nonlinear, probabilistic spike generation process. The linear filter is typically interpreted as the neuron’s “receptive field,” while the spiking mechanism accounts for simple nonlinearities like rectification and response saturation. Given a set of stimuli and (extracellularly) recorded spike times, the characterization problem consists of estimating both the linear filter and the parameters governing the spiking mechanism. One widely used model of this type is the Linear-Nonlinear-Poisson (LNP) cascade model, in which spikes are generated according to an inhomogeneous Poisson process, with rate determined by an instantaneous (“memoryless”) nonlinear function of the filtered input. This model has a number of desirable features, including conceptual simplicity and com- putational tractability. Additionally, reverse correlation analysis provides a simple unbi- ased estimator for the linear filter [5], and the properties of estimators (for both the linear filter and static nonlinearity) have been thoroughly analyzed, even for the case of highly non-symmetric or “naturalistic” stimuli [12]. One important drawback of the LNP model, JWP and LP contributed equally to this work. We thank E.J. Chichilnisky for helpful discussions. l
Jonathan W. Pillow, Liam Paninski, Eero P. Simoncelli
NIPS3
2003 Local Phase Coherence and the Perception of Blur
Zhou Wang 0001, Eero P. Simoncelli
NIPS2
2003 Biases in white noise analysis due to non-Poisson spike generation
Jonathan W. Pillow, Eero P. Simoncelli
Neurocomputing2
2003 Image denoising using scale mixtures of Gaussians in the wavelet domain
abstract
We describe a method for removing noise from digital images, based on a statistical model of the coefficients of an overcomplete multiscale oriented basis. Neighborhoods of coefficients at adjacent positions and scales are modeled as the product of two independent random variables: a Gaussian vector and a hidden positive scalar multiplier. The latter modulates the local variance of the coefficients in the neighborhood, and is thus able to account for the empirically observed correlation between the coefficient amplitudes. Under this model, the Bayesian least squares estimate of each coefficient reduces to a weighted average of the local linear estimates over all possible values of the hidden multiplier variable. We demonstrate through simulations with images contaminated by additive white Gaussian noise that the performance of this method substantially surpasses that of previously published methods, both visually and in terms of mean squared error.
Javier Portilla, Vasily Strela, Martin J. Wainwright, Eero P. Simoncelli
IEEE Trans. Image Process.4
2001 Adaptive Wiener denoising using a Gaussian scale mixture model in the wavelet domain
abstract
We describe a statistical model for images decomposed in an overcomplete wavelet pyramid. Each coefficient of the pyramid is modeled as the product of two independent random variables: an element of a Gaussian random field, and a hidden multiplier with a marginal log-normal prior. The latter modulates the local variance of the coefficients. We assume subband coefficients are contaminated with additive Gaussian noise of known covariance, and compute a MAP estimate of each multiplier variable based on observation of a local neighborhood of coefficients. Conditioned on this multiplier, we then estimate the subband coefficients with a local Wiener estimator. Unlike previous approaches, we (a) empirically motivate our choice for the prior on the multiplier; (b) use the full covariance of signal and noise in the estimation; (c) include adjacent scales in the conditioning neighborhood. To our knowledge, the results are the best in the literature, both visually and in terms of squared error.
Vasily Strela, Javier Portilla, Martin J. Wainwright, Eero P. Simoncelli
ICIP (2)4
2001 Characterizing Neural Gain Control using Spike-triggered Covariance
abstract
Spike-triggered averaging techniques are effective for linear characterization of neural responses. But neurons exhibit important nonlinear behaviors, such as gain control, that are not captured by such analyses. We describe a spike-triggered covariance method for retrieving suppressive components of the gain control signal in a neuron. We demonstrate the method in simulation and on retinal ganglion cell data. Analysis of physiological data reveals significant suppressive axes and explains neural nonlinearities. This method should be applicable to other sensory areas and modalities.
Odelia Schwartz, E. J. Chichilnisky, Eero P. Simoncelli
NIPS3
2001 Modeling temporal response characteristics of V1 neurons with a dynamic normalization model
Samuel Mikaelian, Eero P. Simoncelli
Neurocomputing2
2000 Color Channels Decorrelation by ICA Transformation in the Wavelet Domain for Color Texture Analysis and Synthesis
abstract
In this paper, we developed a color model to cancel the dependency between color channels, which enables us to separate spectral processing from, spatial processing. We introduced Independent Component Analysis (ICA) transformation in the wavelet domain to decorrelate the subband color joint statistics. The decorrelated joint color conditional histograms display scaling of variance. Gaussian Scale Mixture (GSM) was used to model the subband color statistics and a normalization scheme was adapted to cancel the pair-wise color subband statistical dependency. This color model was combined with the Portilla/Simoncelli texture model to construct the color texture model. Based on this model, features were extracted and the corresponding color texture synthesis scheme was developed.
Yufeng Liang, Eero P. Simoncelli, Zhibin Lei
CVPR2
2000 Image Denoising via Adjustment of Wavelet Coefficient Magnitude Correlation
abstract
We describe a novel method of removing additive white noise of known variance from photographic images. The method is based on a characterization of the statistical properties of natural images represented in a complex wavelet decomposition. Specifically, we decompose the noisy image into wavelet subbands, estimate the autocorrelation of both the noise-free raw coefficients and their magnitudes within each subband, impose these statistics by projecting onto the space of images having the desired autocorrelations, and reconstruct an image from the modified wavelet coefficients. This process is applied repeatedly, and can be accelerated to produce optimal results in only a few iterations. Denoising results compare favorably to three reference methods, both perceptually and in terms of mean squared error.
Javier Portilla, Eero P. Simoncelli
ICIP2
2000 Random Cascades of Gaussian Scale Mixtures and Their Use in Modeling Natural Images With Application to Denoising
abstract
Multiresolution representations play an important role in image processing and computer vision, as well as in modeling stochastic processes. We have developed a semi-parametric class of non-Gaussian multiscale statistical processes defined by random cascades on wavelet trees. This model class is rich enough to accurately capture the remarkably regular non-Gaussian features of natural images, but sufficiently structured to permit estimation of the underlying state variables. We showed that our models accurately fit both the marginal and joint histograms of wavelet coefficients from natural images. We developed a Newton-like method for exact MAP state estimation that exploits fast algorithms for tree estimation, and hence is very efficient. Applications of this algorithm to denoising of both 1D signals and natural images were presented. The GSM-tree model class is related to a number of previous approaches to image coding and denoising.
Martin J. Wainwright, Eero P. Simoncelli, Alan S. Willsky
ICIP2
2000 Natural Sound Statistics and Divisive Normalization in the Auditory System
abstract
We explore the statistical properties of natural sound stimuli pre(cid:173) processed with a bank of linear filters. The responses of such filters exhibit a striking form of statistical dependency, in which the response variance of each filter grows with the response amplitude of filters tuned for nearby frequencies. These dependencies may be substantially re(cid:173) duced using an operation known as divisive normalization, in which the response of each filter is divided by a weighted sum of the recti(cid:173) fied responses of other filters. The weights may be chosen to maximize the independence of the normalized responses for an ensemble of natu(cid:173) ral sounds. We demonstrate that the resulting model accounts for non(cid:173) linearities in the response characteristics of the auditory nerve, by com(cid:173) paring model simulations to electrophysiological recordings. In previous work (NIPS, 1998) we demonstrated that an analogous model derived from the statistics of natural images accounts for non-linear properties of neurons in primary visual cortex. Thus, divisive normalization appears to be a generic mechanism for eliminating a type of statistical dependency that is prevalent in natural signals of different modalities. Signals in the real world are highly structured. For example, natural sounds typically con(cid:173) tain both harmonic and rythmic structure. It is reasonable to assume that biological auditory systems are designed to represent these structures in an efficient manner [e.g., 1,2]. Specif(cid:173) ically, Barlow hypothesized that a role of early sensory processing is to remove redundancy in the sensory input, resulting in a set of neural responses that are statistically independent. Experimentally, one can test this hypothesis by examining the statistical properties of neural responses under natural stimulation conditions [e.g., 3,4], or the statistical dependency of pairs (or groups) of neural responses. Due to their technical difficulty, such multi-cellular experiments are only recently becoming possible, and the earliest reports in vision appear consistent with the hypothesis [e.g., 5]. An alternative approach, which we follow here, is to develop a neural model from the statistics of natural signals and show that response properties of this model are similar to those of biological sensory neurons. A number of researchers have derived linear filter models using statistical criterion. For vi(cid:173) sual images, this results in linear filters localized in frequency, orientation and phase [6, 7]. Similar work in audition has yielded filters localized in frequency and phase [8]. Although these linear models provide an important starting point for neural modeling, sensory neu(cid:173) rons are highly nonlinear. In addition, the statistical properties of natural signals are too complex to expect a linear transformation to result in an independent set of components. Recent results indicate that nonlinear gain control plays an important role in neural pro(cid:173) cessing. Ruderman and Bialek [9] have shown that division by a local estimate of standard deviation can increase the entropy of responses of center-surround filters to natural images. Such a model is consistent with the properties of neurons in the retina and lateral genicu(cid:173) late nucleus. Heeger and colleagues have shown that the nonlinear behaviors of neurons in primary visual cortex may be described using a form of gain control known as divisive normalization [10], in which the response of a linear kernel is rectified and divided by the sum of other rectified kernel responses and a constant. We have recently shown that the responses of oriented linear filters exhibit nonlinear statistical dependencies that may be substantially reduced using a variant of this model, in which the normalization signal is computed from a weighted sum of other rectified kernel responses [11, 12]. The resulting model, with weighting parameters determined from image statistics, accounts qualitatively for physiological nonlinearities observed in primary visual cortex. In this paper, we demonstrate that the responses of bandpass linear filters to natural sounds exhibit striking statistical dependencies, analogous to those found in visual images. A divisive normalization procedure can substantially remove these dependencies. We show that this model, with parameters optimized for a collection of natural sounds, can account for nonlinear behaviors of neurons at the level of the auditory nerve. Specifically, we show that: 1) the shape offrequency tuning curves varies with sound pressure level, even though the underlying linear filters are fixed; and 2) superposition of a non-optimal tone suppresses the response of a linear filter in a divisive fashion, and the amount of suppression depends on the distance between the frequency of the tone and the preferred frequency of the filter. 1 Empirical observations of natural sound statistics The basic statistical properties of natural sounds, as observed through a linear filter, have been previously documented by Attias [13]. In particular, he showed that, as with visual images, the spectral energy falls roughly according to a power law, and that the histograms of filter responses are more kurtotic than a Gaussian (i.e., they have a sharp peak at zero, and very long tails). Here we examine the joint statistical properties of a pair of linear filters tuned for nearby temporal frequencies. We choose a fixed set of filters that have been widely used in mod(cid:173) eling the peripheral auditory system [14]. Figure 1 shows joint histograms of the instanta(cid:173) neous responses of a particular pair of linear filters to five different types of natural sound, and white noise. First note that the responses are approximately decorrelated: the expected value of the y-axis value is roughly zero for all values of the x-axis variable. The responses are not, however, statistically independent: the width of the distribution of responses of one filter increases with the response amplitude of the other filter. If the two responses were statistically independent, then the response of the first filter should not provide any information about the distribution of responses of the other filter. We have found that this type of variance dependency (sometimes accompanied by linear correlation) occurs in a wide range of natural sounds, ranging from animal sounds to music. We emphasize that this dependency is a property of natural sounds, and is not due purely to our choice of lin(cid:173) ear filters. For example, no such dependency is observed when the input consists of white noise (see Fig. 1). The strength of this dependency varies for different pairs of linear filters . In addition, we see this type of dependency between instantaneous responses of a single filter at two
Odelia Schwartz, Eero P. Simoncelli
NIPS2
2000 A Parametric Texture Model Based on Joint Statistics of Complex Wavelet Coefficients
Javier Portilla, Eero P. Simoncelli
Int. J. Comput. Vis.2
1999 Scale Mixtures of Gaussians and the Statistics of Natural Images
Martin J. Wainwright, Eero P. Simoncelli
NIPS2
1999 Image compression via joint statistical characterization in the wavelet domain
abstract
We develop a probability model for natural images, based on empirical observation of their statistics in the wavelet transform domain. Pairs of wavelet coefficients, corresponding to basis functions at adjacent spatial locations, orientations, and scales, are found to be non-Gaussian in both their marginal and joint statistical properties. Specifically, their marginals are heavy-tailed, and although they are typically decorrelated, their magnitudes are highly correlated. We propose a Markov model that explains these dependencies using a linear predictor for magnitude coupled with both multiplicative and additive uncertainties, and show that it accounts for the statistics of a wide variety of images including photographic images, graphical images, and medical images. In order to directly demonstrate the power of the model, we construct an image coder called EPWIC (embedded predictive wavelet image coder), in which subband coefficients are encoded one bitplane at a time using a nonadaptive arithmetic encoder that utilizes conditional probabilities calculated from the model. Bitplanes are ordered using a greedy algorithm that considers the MSE reduction per encoded bit. The decoder uses the statistical model to predict coefficient values based on the bits it has received. Despite the simplicity of the model, the rate-distortion performance of the coder is roughly comparable to the best image coders in the literature.
Robert W. Buccigrossi, Eero P. Simoncelli
IEEE Trans. Image Process.2
1998 Texture Characterization via Joint Statistics of Wavelet Coefficient Magnitudes
abstract
We present a parametric statistical characterization of texture images in the context of an overcomplete complex wavelet frame. The characterization consists of the local autocorrelation of the coefficients in each subband, the local autocorrelation of the coefficent magnitudes, and the cross-correlation of coefficient magnitudes at all orientations and adjacent spatial scales. We develop an efficient algorithm for sampling from an implicit probability density conforming to these statistics, and demonstrate its effectiveness in synthesizing artificial and natural texture images.
Eero P. Simoncelli, Javier Portilla
ICIP (1)1
1998 Modeling Surround Suppression in V1 Neurons with a Statistically Derived Normalization Model
Eero P. Simoncelli, Odelia Schwartz
NIPS1
1997 Optimally Rotation-Equivariant Directional Derivative Kernels
Hany Farid, Eero P. Simoncelli
CAIP2
1997 Discrete-Time Rigidity-Constrained Optical Flow
Jeffrey Mendelsohn, Eero P. Simoncelli, Ruzena Bajcsy
CAIP2
1997 Progressive wavelet image coding based on a conditional probability model
abstract
We present a wavelet image coder based on an explicit model of the conditional statistical relationships between coefficients in different subbands. In particular, we construct a parameterized model for the conditional probability of a coefficient given coefficients at a coarser scale. Subband coefficients are encoded one bitplane at a time using a non-adaptive arithmetic encoder. The overall ordering of bitplanes is determined by the ratio of their encoded variance to compressed size. We show rate-distortion comparisons of the coder to first and second-order theoretical entropy bounds and the EZW coder. The coder is inherently embedded, and should prove useful in applications requiring progressive transmission.
Robert W. Buccigrossi, Eero P. Simoncelli
ICASSP2
1997 Embedded Wavelet Image Compression Based On A Joint Probability Model
abstract
We present an embedded image coder based on a statistical characterization of natural images in the wavelet transform domain. We describe the joint distribution between pairs of coefficients at adjacent spatial locations, orientations, and scales. Although the raw coefficients are nearly, uncorrelated, their magnitudes are highly correlated. A linear magnitude predictor coupled with both multiplicative and additive uncertainties, provides a reasonable description of the conditional probability densities. We use this model to construct an image coder called EPWIC (embedded predictive wavelet image coder), in which subband coefficients are encoded one bit-plane at a time using a non-adaptive arithmetic encoder. Bit-planes are ordered using a greedy algorithm that considers the MSE reduction per encoded bit. We demonstrate the quality of the statistical characterization by comparing rate-distortion curves of the coder to several standard coders.
Eero P. Simoncelli, Robert W. Buccigrossi
ICIP (1)1
1996 Direct Differential Range Estimation Using Optical Masks
Eero P. Simoncelli, Hany Farid
ECCV (2)1
1996 A filter design technique for steerable pyramid image transforms
abstract
We describe a novel recursive filter design technique for multi-scale "pyramid" transforms. The recursion in the design technique follows that of the pyramid construction, and allows us to solve a reduced design problem at each step. We demonstrate the use of this technique by designing filters of various orientation bandwidths for use in a "steerable pyramid" image transform.
Anestis Karasaridis, Eero P. Simoncelli
ICASSP2
1996 A rotation invariant pattern signature
abstract
We propose a "signature" for rotation-invariant representation of local image structure. The signature is a complex-valued vector constructed analytically from the projections of the image onto a set of oriented basis kernels. The components of the signature form an over-complete set of algebraic invariants, but are chosen to avoid instabilities associated with previously developed algebraic invariants. We demonstrate the use of this signature for representing and classifying junctions in grayscale imagery.
Eero P. Simoncelli
ICIP (3)1
1996 Noise removal via Bayesian wavelet coring
abstract
The classical solution to the noise removal problem is the Wiener filter, which utilizes the second-order statistics of the Fourier decomposition. Subband decompositions of natural images have significantly non-Gaussian higher-order point statistics; these statistics capture image properties that elude Fourier-based techniques. We develop a Bayesian estimator that is a natural extension of the Wiener solution, and that exploits these higher-order statistics. The resulting nonlinear estimator performs a "coring" operation. We provide a simple model for the subband statistics, and use it to develop a semi-blind noise removal algorithm based on a steerable wavelet pyramid.
Eero P. Simoncelli, Edward H. Adelson
ICIP (1)1
1996 Steerable wedge filters for local orientation analysis
abstract
Steerable filters have been used to analyze local orientation patterns in imagery. Such filters are typically based on directional derivatives, whose symmetry produces orientation responses that are periodic with period pi, independent of image structure. We present a more general set of steerable filters that alleviate this problem.
Eero P. Simoncelli, Hany Farid
IEEE Trans. Image Process.1
1995 Steerable Wedge Filters
abstract
Steerable filters, as developed by Freeman and Adelson (1991), are a class of rotation-invariant linear operators that may be used to analyze local orientation patterns in imagery. The most common examples of such operators are directional derivatives of Gaussians and their 2D Hilbert transforms. The inherent symmetry of these filters produces an orientation response that is periodic with period /spl pi/, even when the underlying image structure does not have such symmetry. This problem may be alleviated by reconsidering the full class of steerable filters. We develop a family of even- and odd-symmetric steerable filters that have a spatially asymmetric "wedge-like" shape and are optimally localized in their orientation response. Unlike the original steerable filters, these filters are not based on directional derivatives and the Hilbert transform relationship is imposed on their angular components. We demonstrate the ability of these filters to properly represent oriented structures.>
Eero P. Simoncelli, Hany Farid
ICCV1
1995 The steerable pyramid: a flexible architecture for multi-scale derivative computation
abstract
We describe an architecture for efficient and accurate linear decomposition of an image into scale and orientation subbands. The basis functions of this decomposition are directional derivative operators of any desired order. We describe the construction and implementation of the transform.
Eero P. Simoncelli, William T. Freeman
ICIP (3)1
1994 Design of Multi-Dimensional Derivative Filters
abstract
Many multi-dimensional signal processing problems require the computation of signal gradients or directional derivatives. Traditional derivative estimates based on adjacent or central differences are often inappropriate for multi-dimensional problems. As replacements for these traditional operators, the author designs a set of matched pairs of derivative filters and lowpass prefilters. The author demonstrates the superiority of these filters over simple difference operators.>
Eero P. Simoncelli
ICIP (1)1
1992 Shiftable multiscale transforms
abstract
One of the major drawbacks of orthogonal wavelet transforms is their lack of translation invariance: the content of wavelet subbands is unstable under translations of the input signal. Wavelet transforms are also unstable with respect to dilations of the input signal and, in two dimensions, rotations of the input signal. The authors formalize these problems by defining a type of translation invariance called shiftability. In the spatial domain, shiftability corresponds to a lack of aliasing; thus, the conditions under which the property holds are specified by the sampling theorem. Shiftability may also be applied in the context of other domains, particularly orientation and scale. Jointly shiftable transforms that are simultaneously shiftable in more than one domain are explored. Two examples of jointly shiftable transforms are designed and implemented: a 1-D transform that is jointly shiftable in position and scale, and a 2-D transform that is jointly shiftable in position and orientation. The usefulness of these image representations for scale-space analysis, stereo disparity measurement, and image enhancement is demonstrated.>
Eero P. Simoncelli, William T. Freeman, Edward H. Adelson, David J. Heeger
IEEE Trans. Inf. Theory1
1991 Probability distributions of optical flow
abstract
Gradient methods are widely used in the computation of optical flow. The authors discuss extensions of these methods which compute probability distributions of optical flow. The use of distributions allows representation of the uncertainties inherent in the optical flow computation, facilitating the combination with information from other sources. Distributed optical flow for a synthetic image sequence is computed, and it is demonstrated that the probabilistic model accounts for the errors in the flow estimates. The distributed optical flow for a real image sequence is computed.>
Eero P. Simoncelli, Edward H. Adelson, David J. Heeger
CVPR1
1990 Non-separable extensions of quadrature mirror filters to multiple dimensions
abstract
Generalized non-separable extensions of quadrature mirror filter (QMF) banks to two and three dimensions, in which the orientation specificity of the high-pass filters is greatly improved, are described. In particular, extensions to two dimensions with hexagonal symmetry, and 3-D spatiotemporal extensions with rhombic-dodecahedral symmetry, are discussed. Although these filters are conceived and designed on nonstandard sampling lattices, they can be applied to rectangularly sampled images. As in one dimension, these transformations can be hierarchically cascaded to form a multiscale pyramid representation. A set of example filters is designed and applied to the problems of image compression, progressive transmission, orientation analysis, and motion analysis.
Eero P. Simoncelli, Edward H. Adelson
Proc. IEEE1