VLDB 2026 Research / reviewers in the wild / expert
Bruno A. Olshausen
dblp:30/3869
· DBLP profile ↗
39ranked-venue papers
4as first author
13since 2021 · last 2025
0000-0002-4738-3067ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sparse coding generates efficient representations for autoassociative memories
Yazhou Zhao, Zeyu Yun, Bruno A. Olshausen, Christopher J. Kymn |
CogSci | 3 |
| 2024 | Binding in hippocampal-entorhinal circuits enables compositionality in cognitive mapsabstractWe propose a normative model for spatial representation in the hippocampal formation that combines optimality principles, such as maximizing coding range and spatial information per neuron, with an algebraic framework for computing in distributed representation. Spatial position is encoded in a residue number system, with individual residues represented by high-dimensional, complex-valued vectors. These are composed into a single vector representing position by a similarity-preserving, conjunctive vector-binding operation. Self-consistency between the vectors representing position and the individual residues is enforced by a modular attractor network whose modules correspond to the grid cell modules in entorhinal cortex. The vector binding operation can also be used to bind different contexts to spatial representations, yielding a model for entorhinal cortex and hippocampus. We provide model analysis of scaling, similarity preservation and convergence behavior as well as experiments demonstrating noise robustness, sub-integer resolution in representing position, and path integration. The model formalizes the computations in the cognitive map and makes testable experimental predictions. Christopher J. Kymn, Sonia Mazelet, Anthony Thomas, Denis Kleyko, Edward Paxon Frady, Fritz Sommer, Bruno A. Olshausen |
NeurIPS | 7 |
| 2024 | Learning Internal Representations of 3D Transformations From 2D Projected InputsabstractWe describe a computational model for inferring 3D structure from the motion of projected 2D points in an image, with the aim of understanding how biological vision systems learn and internally represent 3D transformations from the statistics of their input. The model uses manifold transport operators to describe the action of 3D points in a scene as they undergo transformation. We show that the model can learn the generator of the Lie group for these transformations from purely 2D input, providing a proof-of-concept demonstration for how biological systems could adapt their internal representations based on sensory input. Focusing on a rotational model, we evaluate the ability of the model to infer depth from moving 2D projected points and to learn rotational transformations from 2D training stimuli. Finally, we compare the model performance to psychophysical performance on structure-from-motion tasks. Marissa Connor, Bruno A. Olshausen, Christopher J. Rozell |
Neural Comput. | 2 |
| 2024 | Computing With Residue Numbers in High-Dimensional Representationabstract, a computing framework that unifies residue number systems with an algebra defined over random, high-dimensional vectors. We show how residue numbers can be represented as high-dimensional vectors in a manner that allows algebraic operations to be performed with component-wise, parallelizable operations on the vector elements. The resulting framework, when combined with an efficient method for factorizing high-dimensional vectors, can represent and operate on numerical values over a large dynamic range using vastly fewer resources than previous methods, and it exhibits impressive robustness to noise. We demonstrate the potential for this framework to solve computationally difficult problems in visual perception and combinatorial optimization, showing improvement over baseline methods. More broadly, the framework provides a possible account for the computational operations of grid cells in the brain, and it suggests new machine learning architectures for representing and manipulating numerical data. Christopher J. Kymn, Denis Kleyko, Edward Paxon Frady, Connor Bybee, Pentti Kanerva, Friedrich T. Sommer, Bruno A. Olshausen |
Neural Comput. | 7 |
| 2023 | Minimalistic Unsupervised Representation Learning with the Sparse Manifold Transform
Yubei Chen, Zeyu Yun, Yi Ma 0001, Bruno A. Olshausen, Yann LeCun |
ICLR | 4 |
| 2023 | Bispectral Neural Networks
Sophia Sanborn, Christian Shewmake, Bruno A. Olshausen, Christopher Hillar |
ICLR | 3 |
| 2023 | Emergence of Sparse Representations from NoiseabstractA hallmark of biological neural networks, which distinguishes them from their artificial counterparts, is the high degree of sparsity in their activations. This discrepancy raises three questions our work helps to answer: (i) Why are biological networks so sparse? (ii) What are the benefits of this sparsity? (iii) How can these benefits be utilized by deep learning models? Our answers to all of these questions center around training networks to handle random noise. Surprisingly, we discover that noisy training introduces three implicit loss terms that result in sparsely firing neurons specializing to high variance features of the dataset. When trained to reconstruct noisy-CIFAR10, neurons learn biological receptive fields. More broadly, noisy training presents a new approach to potentially increase model interpretability with additional benefits to robustness and computational efficiency. Trenton Bricken, Rylan Schaeffer, Bruno A. Olshausen, Gabriel Kreiman |
ICML | 3 |
| 2023 | Efficient Decoding of Compositional Structure in Holistic RepresentationsabstractWe investigate the task of retrieving information from compositional distributed representations formed by hyperdimensional computing/vector symbolic architectures and present novel techniques that achieve new information rate bounds. First, we provide an overview of the decoding techniques that can be used to approach the retrieval task. The techniques are categorized into four groups. We then evaluate the considered techniques in several settings that involve, for example, inclusion of external noise and storage elements with reduced precision. In particular, we find that the decoding techniques from the sparse coding and compressed sensing literature (rarely used for hyperdimensional computing/vector symbolic architectures) are also well suited for decoding information from the compositional distributed representations. Combining these decoding techniques with interference cancellation ideas from communications improves previously reported bounds (Hersche et al., 2021) of the information rate of the distributed representations from 1.20 to 1.40 bits per dimension for smaller codebooks and from 0.60 to 1.26 bits per dimension for larger codebooks. Denis Kleyko, Connor Bybee, Ping-Chen Huang, Christopher J. Kymn, Bruno A. Olshausen, Edward Paxon Frady, Friedrich T. Sommer |
Neural Comput. | 5 |
| 2022 | An Interactive Annotation Tool for Perceptual Video CompressionabstractHuman perception is at the core of lossy video compression and yet, it is challenging to collect data that is sufficiently dense to drive compression. In perceptual quality assessment, human feedback is typically collected as a single scalar quality score indicating preference of one distorted video over another. In reality, some videos may be better in some parts but not in others. We propose an approach for collecting finer-grained user feedback through an interactive tool that allows direct optimization of perceptual quality given a fixed bitrate. To this end, we built a novel web-tool which allows users to paint spatio-temporal importance maps over videos. The tool allows for interactive successive refinement: we iteratively re-encode the original video according to the painted importance maps, while maintaining the same bitrate, thus allowing the user to visually see the trade-off of assigning higher importance to one spatio-temporal part of the video at the cost of others. We use this tool to collect data in-the-wild (10 videos, 17 users) and utilize the obtained importance maps in the context of x264 coding to demonstrate that the tool can indeed be used to generate videos which, at the same bitrate, look perceptually better through a subjective study (n = 26) - and are 1.9 times more likely to be preferred by viewers. We plan on collecting a large-scale dataset using the tool for automated perceptual compression in the future. The code for the tool and dataset can be found at https://github.com/jenyap/video-annotation-tool.git. Evgenya Pergament, Pulkit Tandon, Kedar Tatwawadi, Oren Rippel, Lubomir D. Bourdev, Bruno A. Olshausen, Tsachy Weissman, Sachin Katti, Alexander G. Anderson |
QoMEX | 6 |
| 2022 | Learning and Inference in Sparse Coding Models With Langevin DynamicsabstractWe describe a stochastic, dynamical system capable of inference and learning in a probabilistic latent variable model. The most challenging problem in such models-sampling the posterior distribution over latent variables-is proposed to be solved by harnessing natural sources of stochasticity inherent in electronic and neural systems. We demonstrate this idea for a sparse coding model by deriving a continuous-time equation for inferring its latent variables via Langevin dynamics. The model parameters are learned by simultaneously evolving according to another continuous-time equation, thus bypassing the need for digital accumulators or a global clock. Moreover, we show that Langevin dynamics lead to an efficient procedure for sampling from the posterior distribution in the L0 sparse regime, where latent variables are encouraged to be set to zero as opposed to having a small L1 norm. This allows the model to properly incorporate the notion of sparsity rather than having to resort to a relaxed version of sparsity to make optimization tractable. Simulations of the proposed dynamical system on both synthetic and natural image data sets demonstrate that the model is capable of probabilistically correct inference, enabling learning of the dictionary as well as parameters of the prior. Michael Y.-S. Fang, Mayur Mudigonda, Ryan Zarcone, Amir Khosrowshahi, Bruno A. Olshausen |
Neural Comput. | 5 |
| 2022 | Vector Symbolic Architectures as a Computing Framework for Emerging Hardwareabstract(also known as Hyperdimensional Computing). This framework is well suited for implementation in stochastic, emerging hardware and it naturally expresses the types of cognitive operations required for Artificial Intelligence (AI). We demonstrate in this article that the field-like algebraic structure of Vector Symbolic Architectures offers simple but powerful operations on high-dimensional vectors that can support all data structures and manipulations relevant to modern computing. In addition, we illustrate the distinguishing feature of Vector Symbolic Architectures, "computing in superposition," which sets it apart from conventional computing. It also opens the door to efficient solutions to the difficult combinatorial search problems inherent in AI applications. We sketch ways of demonstrating that Vector Symbolic Architectures are computationally universal. We see them acting as a framework for computing with distributed representations that can play a role of an abstraction layer for emerging computing hardware. This article serves as a reference for computer architects by illustrating the philosophy behind Vector Symbolic Architectures, techniques of distributed computing with them, and their relevance to emerging computing hardware, such as neuromorphic computing. Denis Kleyko, Mike Davies 0002, Edward Paxon Frady, Pentti Kanerva, Spencer J. Kent, Bruno A. Olshausen, Evgeny Osipov, Jan M. Rabaey, Dmitri A. Rachkovskij, Abbas Rahimi, Friedrich T. Sommer |
Proc. IEEE | 6 |
| 2021 | Tent: Fully Test-Time Adaptation by Entropy Minimization
Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen, Trevor Darrell |
ICLR | 4 |
| 2021 | Generalized Learning Vector Quantization for Classification in Randomized Neural Networks and Hyperdimensional ComputingabstractMachine learning algorithms deployed on edge devices must meet certain resource constraints and efficiency requirements. Random Vector Functional Link (RVFL) networks are favored for such applications due to their simple design and training efficiency. We propose a modified RVFL network that avoids computationally expensive matrix operations during training, thus expanding the network's range of potential applications. Our modification replaces the least-squares classifier with the Generalized Learning Vector Quantization (GLVQ) classifier, which only employs simple vector and distance calculations. The GLVQ classifier can also be considered an improvement upon certain classification algorithms popularly used in the area of Hyperdimensional Computing. The proposed approach achieved state-of-the-art accuracy on a collection of datasets from the UCI Machine Learning Repository - higher than previously proposed RVFL networks. We further demonstrate that our approach still achieves high accuracy while severely limited in training iterations (using on average only 21% of the least-squares classifier computational costs). Cameron Diao, Denis Kleyko, Jan M. Rabaey, Bruno A. Olshausen |
IJCNN | 4 |
| 2020 | Resonator Networks, 1: An Efficient Solution for Factoring High-Dimensional, Distributed Representations of Data StructuresabstractThe ability to encode and manipulate data structures with distributed neural representations could qualitatively enhance the capabilities of traditional neural networks by supporting rule-based symbolic reasoning, a central property of cognition. Here we show how this may be accomplished within the framework of Vector Symbolic Architectures (VSAs) (Plate, 1991; Gayler, 1998; Kanerva, 1996), whereby data structures are encoded by combining high-dimensional vectors with operations that together form an algebra on the space of distributed representations. In particular, we propose an efficient solution to a hard combinatorial search problem that arises when decoding elements of a VSA data structure: the factorization of products of multiple codevectors. Our proposed algorithm, called a resonator network, is a new type of recurrent neural network that interleaves VSA multiplication operations and pattern completion. We show in two examples-parsing of a tree-like data structure and parsing of a visual scene-how the factorization problem arises and how the resonator network can solve it. More broadly, resonator networks open the possibility of applying VSAs to myriad artificial intelligence problems in real-world domains. The companion article in this issue (Kent, Frady, Sommer, & Olshausen, 2020) presents a rigorous analysis and evaluation of the performance of resonator networks, showing it outperforms alternative approaches. Edward Paxon Frady, Spencer J. Kent, Bruno A. Olshausen, Friedrich T. Sommer |
Neural Comput. | 3 |
| 2020 | Resonator Networks, 2: Factorization Performance and Capacity Compared to Optimization-Based MethodsabstractWe develop theoretical foundations of resonator networks, a new type of recurrent neural network introduced in Frady, Kent, Olshausen, and Sommer (2020), a companion article in this issue, to solve a high-dimensional vector factorization problem arising in Vector Symbolic Architectures. Given a composite vector formed by the Hadamard product between a discrete set of high-dimensional vectors, a resonator network can efficiently decompose the composite into these factors. We compare the performance of resonator networks against optimization-based methods, including Alternating Least Squares and several gradient-based algorithms, showing that resonator networks are superior in several important ways. This advantage is achieved by leveraging a combination of nonlinear dynamics and searching in superposition, by which estimates of the correct solution are formed from a weighted superposition of all possible solutions. While the alternative methods also search in superposition, the dynamics of resonator networks allow them to strike a more effective balance between exploring the solution space and exploiting local information to drive the network toward probable solutions. Resonator networks are not guaranteed to converge, but within a particular regime they almost always do. In exchange for relaxing the guarantee of global convergence, resonator networks are dramatically more effective at finding factorizations than all alternative approaches considered. Spencer J. Kent, Edward Paxon Frady, Friedrich T. Sommer, Bruno A. Olshausen |
Neural Comput. | 4 |
| 2019 | Superposition of many models into oneabstractWe present a method for storing multiple models within a single set of parameters. Models can coexist in superposition and still be retrieved individually. In experiments with neural networks, we show that a surprisingly large number of models can be effectively stored within a single parameter instance. Furthermore, each of these models can undergo thousands of training steps without significantly interfering with other models within the superposition. This approach may be viewed as the online complement of compression: rather than reducing the size of a network after training, we make use of the unrealized capacity of a network during training. Brian Cheung, Alexander Terekhov, Yubei Chen, Pulkit Agrawal 0001, Bruno A. Olshausen |
NeurIPS | 5 |
| 2018 | Joint Source-Channel Coding with Neural Networks for Analog Data Compression and StorageabstractWe provide an encoding and decoding strategy for efficient storage of analog data onto an array of Phase-Change Memory (PCM) devices. The PCM array is treated as an analog channel, with the stochastic relationship between write voltage and read resistance for each device determining its theoretical capacity. The encoder and decoder are implemented as neural networks with parameters that are trained end-to-end to minimize distortion for a fixed number of devices. To minimize distortion, the encoder and decoder must adapt jointly to the statistics of images and the statistics of the channel. Similar to Balle et al. (2017), we find that incorporating divisive normalization in the encoder, paired with de-normalization in the decoder, improves model performance. We show that the autoencoder achieves a rate-distortion performance above that achieved by a separate JPEG source coding and binary channel coding scheme. These results demonstrate the feasibility of exploiting the full analog dynamic range of PCM or other emerging memory devices for efficient storage of analog image data. Ryan Zarcone, Dylan M. Paiton, Alex Anderson, Jesse H. Engel, H.-S. Philip Wong, Bruno A. Olshausen |
DCC | 6 |
| 2018 | The Sparse Manifold TransformabstractWe present a signal representation framework called the sparse manifold transform that combines key ideas from sparse coding, manifold learning, and slow feature analysis. It turns non-linear transformations in the primary sensory signal space into linear interpolations in a representational embedding space while maintaining approximate invertibility. The sparse manifold transform is an unsupervised and generative framework that explicitly and simultaneously models the sparse discreteness and low-dimensional manifold structure found in natural scenes. When stacked, it also models hierarchical composition. We provide a theoretical description of the transform and demonstrate properties of the learned representation on both synthetic data and natural videos. Yubei Chen, Dylan M. Paiton, Bruno A. Olshausen |
NeurIPS | 3 |
| 2017 | Emergence of foveal image sampling from learning to attend in visual scenes
Brian Cheung, Eric Weiss, Bruno A. Olshausen |
ICLR (Poster) | 3 |
| 2014 | Modeling Higher-Order Correlations within Cortical MicrocolumnsabstractWe statistically characterize the population spiking activity obtained from simultaneous recordings of neurons across all layers of a cortical microcolumn. Three types of models are compared: an Ising model which captures pairwise correlations between units, a Restricted Boltzmann Machine (RBM) which allows for modeling of higher-order correlations, and a semi-Restricted Boltzmann Machine which is a combination of Ising and RBM models. Model parameters were estimated in a fast and efficient manner using minimum probability flow, and log likelihoods were compared using annealed importance sampling. The higher-order models reveal localized activity patterns which reflect the laminar organization of neurons within a cortical column. The higher-order models also outperformed the Ising model in log-likelihood: On populations of 20 cells, the RBM had 10% higher log-likelihood (relative to an independent model) than a pairwise model, increasing to 45% gain in a larger network with 100 spatiotemporal elements, consisting of 10 neurons over 10 time steps. We further removed the need to model stimulus-induced correlations by incorporating a peri-stimulus time histogram term, in which case the higher order models continued to perform best. These results demonstrate the importance of higher-order interactions to describe the structure of correlated activity in cortical networks. Boltzmann Machines with hidden units provide a succinct and effective way to capture these dependencies without increasing the difficulty of model estimation and evaluation. Urs Köster, Jascha Sohl-Dickstein, Charles M. Gray, Bruno A. Olshausen |
PLoS Comput. Biol. | 4 |
| 2012 | Learning Intermediate-Level Representations of Form and Motion from Natural MoviesabstractWe present a model of intermediate-level visual representation that is based on learning invariances from movies of the natural environment. The model is composed of two stages of processing: an early feature representation layer and a second layer in which invariances are explicitly represented. Invariances are learned as the result of factoring apart the temporally stable and dynamic components embedded in the early feature representation. The structure contained in these components is made explicit in the activities of second-layer units that capture invariances in both form and motion. When trained on natural movies, the first layer produces a factorization, or separation, of image content into a temporally persistent part representing local edge structure and a dynamic part representing local motion structure, consistent with known response properties in early visual cortex (area V1). This factorization linearizes statistical dependencies among the first-layer units, making them learnable by the second layer. The second-layer units are split into two populations according to the factorization in the first layer. The form-selective units receive their input from the temporally persistent part (local edge structure) and after training result in a diverse set of higher-order shape features consisting of extended contours, multiscale edges, textures, and texture boundaries. The motion-selective units receive their input from the dynamic part (local motion structure) and after training result in a representation of image translation over different spatial scales and directions, in addition to more complex deformations. These representations provide a rich description of dynamic natural images and testable hypotheses regarding intermediate-level representation in visual cortex. Charles F. Cadieu, Bruno A. Olshausen |
Neural Comput. | 2 |
| 2011 | Lie Group Transformation Models for Predictive Video CodingabstractWe propose a new method for modeling the temporal correlation in videos, based on local transforms realized by Lie group operators. A large class of transforms can be theoretically described by these operators, however, we propose to learn from natural movies a subset of transforms that are statistically relevant for video representation. The proposed transformation modeling is further exploited to remove inter-view redundancy, i.e., as the prediction step of video encoding. Since the Lie group transformation coefficients are continuous, a quantization step is necessary for each transform. Therefore, we derive theoretical bounds on the distortion due to coefficient quantization. The experimental results demonstrate that the new prediction method with learned transforms leads to better rate-distortion performance at higher bit-rates, and competitive performance at lower bit-rates, compared to the standard prediction based on block-based motion estimation. Ching Ming Wang, Jascha Sohl-Dickstein, Ivana Tosic, Bruno A. Olshausen |
DCC | 4 |
| 2011 | Building a better probabilistic model of images by factorizationabstractWe describe a directed bilinear model that learns higher-order groupings among features of natural images. The model represents images in terms of two sets of latent variables: one set of variables represents which feature groups are active, while the other specifies the relative activity within groups. Such a factorized representation is beneficial because it is stable in response to small variations in the placement of features while still preserving information about relative spatial relationships. When trained on MNIST digits, the resulting representation provides state of the art performance in classification using a simple classifier. When trained on natural images, the model learns to group features according to proximity in position, orientation, and scale. The model achieves high log-likelihood (-94 nats), surpassing the current state of the art for natural images achievable with an mcRBM model. Benjamin J. Culpepper, Jascha Sohl-Dickstein, Bruno A. Olshausen |
ICCV | 3 |
| 2010 | Group Sparse Coding with a Laplacian Scale Mixture PriorabstractWe propose a class of sparse coding models that utilizes a Laplacian Scale Mixture (LSM) prior to model dependencies among coefficients. Each coefficient is modeled as a Laplacian distribution with a variable scale parameter, with a Gamma distribution prior over the scale parameter. We show that, due to the conjugacy of the Gamma prior, it is possible to derive efficient inference procedures for both the coefficients and the scale parameter. When the scale parameters of a group of coefficients are combined into a single variable, it is possible to describe the dependencies that occur due to common amplitude fluctuations among coefficients, which have been shown to constitute a large fraction of the redundancy in natural images. We show that, as a consequence of this group sparse coding, the resulting inference of the coefficients follows a divisive normalization rule, and that this may be efficiently implemented a network architecture similar to that which has been proposed to occur in primary visual cortex. We also demonstrate improvements in image coding and compressive sensing recovery using the LSM model. Pierre Garrigues, Bruno A. Olshausen |
NIPS | 2 |
| 2009 | Learning transport operators for image manifoldsabstractWe describe a method for learning a group of continuous transformation operators to traverse smooth nonlinear manifolds. The method is applied to model how natural images change over time and scale. The group of continuous transform operators is represented by a basis that is adapted to the statistics of the data so that the infinitesimal generator for a measurement orbit can be produced by a linear combination of a few basis elements. We illustrate how the method can be used to efficiently code time-varying images by describing changes across time and scale in terms of the learned operators. Benjamin J. Culpepper, Bruno A. Olshausen |
NIPS | 2 |
| 2008 | Learning Transformational Invariants from Natural MoviesabstractWe describe a hierarchical, probabilistic model that learns to extract complex motion from movies of the natural environment. The model consists of two hidden layers: the first layer produces a sparse representation of the image that is expressed in terms of local amplitude and phase variables. The second layer learns the higher-order structure among the time-varying phase variables. After training on natural movies, the top layer units discover the structure of phase-shifts within the first layer. We show that the top layer units encode transformational invariants: they are selective for the speed and direction of a moving pattern, but are invariant to its spatial structure (orientation/spatial-frequency). The diversity of units in both the intermediate and top layers of the model provides a set of testable predictions for representations that might be found in V1 and MT. In addition, the model demonstrates how feedback from higher levels can influence representations at lower levels as a by-product of inference in a graphical model. Charles F. Cadieu, Bruno A. Olshausen |
NIPS | 2 |
| 2008 | Sparse Coding via Thresholding and Local Competition in Neural CircuitsabstractWhile evidence indicates that neural systems may be employing sparse approximations to represent sensed stimuli, the mechanisms underlying this ability are not understood. We describe a locally competitive algorithm (LCA) that solves a collection of sparse coding principles minimizing a weighted combination of mean-squared error and a coefficient cost function. LCAs are designed to be implemented in a dynamical system composed of many neuron-like elements operating in parallel. These algorithms use thresholding functions to induce local (usually one-way) inhibitory competitions between nodes to produce sparse representations. LCAs produce coefficients with sparsity levels comparable to the most popular centralized sparse coding algorithms while being readily suited for neural implementation. Additionally, LCA coefficients for video sequences demonstrate inertial properties that are both qualitatively and quantitatively more regular (i.e., smoother and more predictable) than the coefficients produced by greedy algorithms. Christopher J. Rozell, Don H. Johnson, Richard G. Baraniuk, Bruno A. Olshausen |
Neural Comput. | 4 |
| 2007 | Locally Competitive Algorithms for Sparse ApproximationabstractPractical sparse approximation algorithms (particularly greedy algorithms) suffer two significant drawbacks: they are difficult to implement in hardware, and they are inefficient for time-varying stimuli (e.g., video) because they produce erratic temporal coefficient sequences. We present a class of locally competitive algorithms (LCAs) that correspond to a collection of sparse approximation principles minimizing a weighted combination of reconstruction MSE and a coefficient cost function. These systems use thresholding functions to induce local nonlinear competitions in a dynamical system. Simple analog hardware can implement the required nonlinearities and competitions. We show that our LCAs are stable under normal operating conditions and can produce sparsity levels comparable to existing methods. Additionally, these LCAs can produce coefficients for video sequences that are more regular (i.e., smoother and more predictable) than the coefficients produced by greedy algorithms. Christopher J. Rozell, Don H. Johnson, Richard G. Baraniuk, Bruno A. Olshausen |
ICIP (4) | 4 |
| 2007 | Learning Horizontal Connections in a Sparse Coding Model of Natural ImagesabstractIt has been shown that adapting a dictionary of basis functions to the statistics of natural images so as to maximize sparsity in the coefficients results in a set of dictionary elements whose spatial properties resemble those of V1 (primary visual cortex) receptive fields. However, the resulting sparse coefficients still exhibit pronounced statistical dependencies, thus violating the independence assumption of the sparse coding model. Here, we propose a model that attempts to capture the dependencies among the basis function coefficients by including a pairwise coupling term in the prior over the coefficient activity states. When adapted to the statistics of natural images, the coupling terms learn a combination of facilitatory and inhibitory interactions among neighboring basis functions. These learned interactions may offer an explanation for the function of horizontal connections in V1, and we discuss the implications of our findings for physiological experiments. Pierre Garrigues, Bruno A. Olshausen |
NIPS | 2 |
| 2005 | Neuroanatomical image alignment
Issac J. Trotts, Bruno A. Olshausen, Edward Jones |
IEEE Visualization | 2 |
| 2005 | How Close Are We to Understanding V1?abstractA wide variety of papers have reviewed what is known about the function of primary visual cortex. In this review, rather than stating what is known, we attempt to estimate how much is still unknown about V1 function. In particular, we identify five problems with the current view of V1 that stem largely from experimental and theoretical biases, in addition to the contributions of nonlinearities in the cortex that are not well understood. Our purpose is to open the door to new theories, a number of which we describe, along with some proposals for testing them. Bruno A. Olshausen, David J. Field |
Neural Comput. | 1 |
| 2003 | Learning sparse, overcomplete representations of time-varying natural imagesabstractI show how to adapt an overcomplete dictionary of space-time functions so as to represent time-varying natural images with maximum sparsity. The basis functions are considered as part of a probabilistic model of image sequences, with a sparse prior imposed over the coefficients. Learning is accomplished by maximizing the log-likelihood of the model, using natural movies as training data. The basis functions that emerge are space-time inseparable functions that resemble the motion-selective receptive fields of simple-cells in mammalian visual cortex. When the coefficients are computed via matching-pursuit in space and time, one obtains a punctuate, spike-like representation of continuous time-varying images. It is suggested that such a coding scheme may be at work in the visual cortex. Bruno A. Olshausen |
ICIP (1) | 1 |
| 2003 | Image denoising using learned overcomplete representationsabstractWe describe a method for learning sparse multiscale image representations using a sparse prior distribution over the basis function coefficients. The prior consists of a mixture of a Gaussian and a Dirac delta function, and thus encourages coefficients to have exact zero values. Coefficients for an image are computed by sampling from the resulting posterior distribution with a Gibbs sampler. Denoising using the learned image model is demonstrated for some standard test images, with results that compare favorably with other denoising methods. Phil Sallee, Bruno A. Olshausen |
ICIP (3) | 2 |
| 2002 | Learning Sparse Multiscale Image RepresentationsabstractWe describe a method for learning sparse multiscale image repre- sentations using a sparse prior distribution over the basis function coe(cid:14)cients. The prior consists of a mixture of a Gaussian and a Dirac delta function, and thus encourages coe(cid:14)cients to have exact zero values. Coe(cid:14)cients for an image are computed by sampling from the resulting posterior distribution with a Gibbs sampler. The learned basis is similar to the Steerable Pyramid basis, and yields slightly higher SNR for the same number of active coe(cid:14)cients. De- noising using the learned image model is demonstrated for some standard test images, with results that compare favorably with other denoising methods. Phil Sallee, Bruno A. Olshausen |
NIPS | 2 |
| 2000 | Learning Sparse Image Codes using a Wavelet Pyramid ArchitectureabstractWe show how a wavelet basis may be adapted to best represent natural images in terms of sparse coefficients. The wavelet basis, which may be either complete or overcomplete, is specified by a small number of spatial functions which are repeated across space and combined in a recursive fashion so as to be self-similar across scale. These functions are adapted to minimize the estimated code length under a model that assumes images are composed of a linear superposition of sparse, independent components. When adapted to natural images, the wavelet bases take on different orientations and they evenly tile the orientation domain, in stark contrast to the standard, non-oriented wavelet bases used in image compression. When the basis set is allowed to be overcomplete, it also yields higher coding efficiency than standard wavelet bases. Bruno A. Olshausen, Phil Sallee, Michael S. Lewicki |
NIPS | 1 |
| 1999 | Learning Sparse Codes with a Mixture-of-Gaussians Prior
Bruno A. Olshausen, K. Jarrod Millman |
NIPS | 1 |
| 1997 | Inferring Sparse, Overcomplete Image Codes Using an Efficient Coding Framework
Michael S. Lewicki, Bruno A. Olshausen |
NIPS | 2 |
| 1996 | A nonlinear Hebbian network that learns to detect disparity in random- dot stereogramsabstractAn intrinsic limitation of linear, Hebbian networks is that they are capable of learning only from the linear pairwise correlations within an input stream. To explore what higher forms of structure could be learned with a nonlinear Hebbian network, we constructed a model network containing a simple form of nonlinearity and we applied it to the problem of learning to detect the disparities present in random-dot stereograms. The network consists of three layers, with nonlinear sigmoidal activation functions in the second-layer units. The nonlinearities allow the second layer to transform the pixel-based representation in the input layer into a new representation based on coupled pairs of left-right inputs. The third layer of the network then clusters patterns occurring on the second-layer outputs according to their disparity via a standard competitive learning rule. Analysis of the network dynamics shows that the second-layer units' nonlinearities interact with the Hebbian learning rule to expand the region over which pairs of left-right inputs are stable. The learning rule is neurobiologically inspired and plausible, and the model may shed light on how the nervous system learns to use coincidence detection in general. Bruno A. Olshausen |
Neural Comput. | 2 |
| 1993 | Neurobiology, Psychophysics, and Computational Models of Visual Attention
Ernst Niebur, Bruno A. Olshausen |
NIPS | 2 |