VLDB 2026 Research / reviewers in the wild / expert
Haim Sompolinsky
dblp:33/5545
· DBLP profile ↗
31ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-0322-0629ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
13 papers |
Learning theory · 40% Deep learning architectures and training · 27% Probabilistic and Bayesian machine learning · 18% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Bioinformatics and computational biology · 100% |
Topics — the 30 heaviest of 34, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
bayesian inference |
1.3 | 2 | 2024 | Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of Transformers · NeurIPS 2024 Globally Gated Deep Linear Networks · NeurIPS 2022 |
Machine learning › Learning theory › statistical learning theory
statistical physics of learning |
1.3 | 2 | 2024 | Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of Transformers · NeurIPS 2024 A theory of weight distribution-constrained learning · NeurIPS 2022 |
Machine learning › Learning theory
neural network theory |
0.9 | 1 | 2025 | When narrower is better: the narrow width limit of Bayesian parallel branching neural networks · ICLR 2025 |
Machine learning › Deep learning architectures and training › attention mechanism
self-attention |
0.8 | 1 | 2024 | Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of Transformers · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of Transformers · NeurIPS 2024 |
Machine learning › Deep learning architectures and training › feedforward neural network › piecewise linear network
gated linear networks |
0.6 | 1 | 2022 | Globally Gated Deep Linear Networks · NeurIPS 2022 |
Machine learning › Learning theory › generalization
generalization theory |
0.6 | 1 | 2022 | Globally Gated Deep Linear Networks · NeurIPS 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge engineering › knowledge integration
prior knowledge integration |
0.6 | 1 | 2022 | A theory of weight distribution-constrained learning · NeurIPS 2022 |
Bioinformatics and computational biology
computational neuroscience |
0.4 | 3 | 2022 | A theory of weight distribution-constrained learning · NeurIPS 2022 Inferring Stimulus Selectivity from the Spatial Structure of Neural Network Dynamics · NIPS 2010 Short-term memory in neuronal networks through dynamical compressed sensing · NIPS 2010 |
Machine learning › Graph learning
graph neural network |
0.3 | 1 | 2025 | When narrower is better: the narrow width limit of Bayesian parallel branching neural networks · ICLR 2025 |
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction |
0.1 | 1 | 2010 | Inferring Stimulus Selectivity from the Spatial Structure of Neural Network Dynamics · NIPS 2010 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
subspace learning |
0.1 | 1 | 2010 | Inferring Stimulus Selectivity from the Spatial Structure of Neural Network Dynamics · NIPS 2010 |
Bioinformatics and computational biology › computational neuroscience
neural population dynamics |
0.1 | 1 | 2010 | Inferring Stimulus Selectivity from the Spatial Structure of Neural Network Dynamics · NIPS 2010 |
Information theory › signal processing
compressed sensing |
0.1 | 1 | 2010 | Short-term memory in neuronal networks through dynamical compressed sensing · NIPS 2010 |
Bioinformatics and computational biology › computational neuroscience
neural coding |
0.0 | 1 | 2001 | Correlation Codes in Neuronal Populations · NIPS 2001 |
Bioinformatics and computational biology › computational neuroscience › neural coding
neural population coding |
0.0 | 1 | 2001 | Correlation Codes in Neuronal Populations · NIPS 2001 |
Machine learning › Representation and self-supervised learning
mutual information maximization |
0.0 | 1 | 2000 | An Information Maximization Approach to Overcomplete and Recurrent Representations · NIPS 2000 |
Machine learning › Representation and self-supervised learning
overcomplete representation |
0.0 | 1 | 2000 | An Information Maximization Approach to Overcomplete and Recurrent Representations · NIPS 2000 |
Computer vision › 3D vision
higher-order statistics |
0.0 | 1 | 1999 | Algorithms for Independent Components Analysis and Higher Order Statistics · NIPS 1999 |
Machine learning › Representation and self-supervised learning › blind source separation
independent component analysis |
0.0 | 1 | 1999 | Algorithms for Independent Components Analysis and Higher Order Statistics · NIPS 1999 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.0 | 1 | 1998 | Learning a Continuous Hidden Variable Model for Binary Data · NIPS 1998 |
Machine learning › Representation and self-supervised learning › computational neuroscience › neural coding
population codes |
0.0 | 1 | 1998 | The Effect of Correlations on the Fisher Information of Population Codes · NIPS 1998 |
Information theory › information measures
fisher information |
0.0 | 1 | 1998 | The Effect of Correlations on the Fisher Information of Population Codes · NIPS 1998 |
Machine learning › Learning theory
generalization error |
0.0 | 2 | 1994 | Query by Committee · COLT 1992 On-line Learning of Dichotomies · NIPS 1994 |
Machine learning › Learning theory
online learning |
0.0 | 1 | 1994 | On-line Learning of Dichotomies · NIPS 1994 |
Machine learning › Learning theory › online learning
perceptron |
0.0 | 1 | 1994 | On-line Learning of Dichotomies · NIPS 1994 |
Machine learning › Learning theory
fisher information |
0.0 | 1 | 2001 | Correlation Codes in Neuronal Populations · NIPS 2001 |
Machine learning › Efficient and distributed learning
active learning |
0.0 | 1 | 1992 | Query by Committee · COLT 1992 |
Machine learning › Learning theory › information-theoretic learning
information gain |
0.0 | 1 | 1992 | Query by Committee · COLT 1992 |
Machine learning › Efficient and distributed learning › active learning › disagreement-based active learning
query by committee |
0.0 | 1 | 1992 | Query by Committee · COLT 1992 |
Methods — techniques the papers use, named apart from their topics
kernel theory · 1.3wasserstein distance · 1.1stochastic gradient descent · 1.1optimal transport · 1.1information geometry · 1.1statistical mechanics · 1.0bayesian parallel branching neural network · 0.9NNGP · 0.9gradient descent · 0.6bayesian learning · 0.6statistical physics of disordered systems · 0.2asymptotic analysis · 0.2subspace angle analysis · 0.1principal component analysis · 0.1fisher information · 0.0bilinear readout model · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | When narrower is better: the narrow width limit of Bayesian parallel branching neural networksabstractThe infinite width limit of random neural networks is known to result in Neural Networks as Gaussian Process (NNGP) (Lee et al. (2018)), characterized by task-independent kernels. It is widely accepted that larger network widths contribute to improved generalization (Park et al. (2019)). However, this work challenges this notion by investigating the narrow width limit of the Bayesian Parallel Branching
Neural Network (BPB-NN), an architecture that resembles neural networks with residual blocks. We demonstrate that when the width of a BPB-NN is significantly smaller compared to the number of training examples, each branch exhibits more robust learning due to a symmetry breaking of branches in kernel renormalization. Surprisingly, the performance of a BPB-NN in the narrow width limit is generally superior to or comparable to that achieved in the wide width limit in bias-limited scenarios. Furthermore, the readout norms of each branch in the narrow width limit are mostly independent of the architectural hyperparameters but generally reflective of the nature of the data. We demonstrate such phenomenon primarily in the branching graph neural networks, where each branch represents a different order of convolutions of the graph; we also extend the results to other more general architectures such as the residual-MLP and demonstrate that the narrow width effect is a general feature of the branching networks. Our results characterize a newly defined narrow-width regime for parallel branching networks in general. Haim Sompolinsky |
ICLR | 2 |
| 2024 | Dissecting the Interplay of Attention Paths in a Statistical Mechanics Theory of TransformersabstractDespite the remarkable empirical performance of Transformers, their theoretical understanding remains elusive. Here, we consider a deep multi-head self-attention network, that is closely related to Transformers yet analytically tractable. We develop a statistical mechanics theory of Bayesian learning in this model, deriving exact equations for the network's predictor statistics under the finite-width thermodynamic limit, i.e., $N,P\rightarrow\infty$, $P/N=\mathcal{O}(1)$, where $N$ is the network width and $P$ is the number of training examples. Our theory shows that the predictor statistics are expressed as a sum of independent kernels, each one pairing different "attention paths", defined as information pathways through different attention heads across layers. The kernels are weighted according to a "task-relevant kernel combination" mechanism that aligns the total kernel with the task labels. As a consequence, this interplay between attention paths enhances generalization performance. Experiments confirm our findings on both synthetic and real-world sequence classification tasks. Finally, our theory explicitly relates the kernel combination mechanism to properties of the learned weights, allowing for a qualitative transfer of its insights to models trained via gradient descent. As an illustration, we demonstrate an efficient size reduction of the network, by pruning those attention heads that are deemed less relevant by our theory. Lorenzo Tiberi, Francesca Mignacco, Kazuki Irie, Haim Sompolinsky |
NeurIPS | 4 |
| 2023 | Optimal Quadratic Binding for Relational Reasoning in Vector Symbolic Neural ArchitecturesabstractBinding operation is fundamental to many cognitive processes, such as cognitive map formation, relational reasoning, and language comprehension. In these processes, two different modalities, such as location and objects, events and their contextual cues, and words and their roles, need to be bound together, but little is known about the underlying neural mechanisms. Previous work has introduced a binding model based on quadratic functions of bound pairs, followed by vector summation of multiple pairs. Based on this framework, we address the following questions: Which classes of quadratic matrices are optimal for decoding relational structures? And what is the resultant accuracy? We introduce a new class of binding matrices based on a matrix representation of octonion algebra, an eight-dimensional extension of complex numbers. We show that these matrices enable a more accurate unbinding than previously known methods when a small number of pairs are present. Moreover, numerical optimization of a binding operator converges to this octonion binding. We also show that when there are a large number of bound pairs, however, a random quadratic binding performs, as well as the octonion and previously proposed binding methods. This study thus provides new insight into potential neural mechanisms of binding operations in the brain. Naoki Hiratani, Haim Sompolinsky |
Neural Comput. | 2 |
| 2022 | Globally Gated Deep Linear NetworksabstractRecently proposed Gated Linear Networks (GLNs) present a tractable nonlinear network architecture, and exhibit interesting capabilities such as learning with local error signals and reduced forgetting in sequential learning. In this work, we introduce a novel gating architecture, named Globally Gated Deep Linear Networks (GGDLNs) where gating units are shared among all processing units in each layer, thereby decoupling the architectures of the nonlinear but unlearned gating and the learned linear processing motifs. We derive exact equations for the generalization properties of Bayesian Learning in these networks in the finite-width thermodynamic limit, defined by $N, P\rightarrow\infty$ while $P/N=O(1)$ where $N$ and $P$ are the hidden layers' width and size of training data sets respectfully. We find that the statistics of the network predictor can be expressed in terms of kernels that undergo shape renormalization through a data-dependent order-parameter matrix compared to the infinite-width Gaussian Process (GP) kernels. Our theory accurately captures the behavior of finite width GGDLNs trained with gradient descent (GD) dynamics. We show that kernel shape renormalization gives rise to rich generalization properties w.r.t. network width, depth, and $L_2$ regularization amplitude. Interestingly, networks with a large number of gating units behave similarly to standard ReLU architectures. Although gating units in the model do not participate in supervised learning, we show the utility of unsupervised learning of the gating parameters. Additionally, our theory allows the evaluation of the network capacity for learning multiple tasks by incorporating task-relevant information into the gating units. In summary, our work is the first exact theoretical solution of learning in a family of nonlinear networks with finite width. The rich and diverse behavior of the GGDLNs suggests that they are helpful analytically tractable models of learning single and multiple tasks, in finite-width nonlinear deep networks. Qianyi Li, Haim Sompolinsky |
NeurIPS | 2 |
| 2022 | A theory of weight distribution-constrained learningabstractA central question in computational neuroscience is how structure determines function in neural networks. Recent large-scale connectomic studies have started to provide a wealth of structural information such as the distribution of excitatory/inhibitory cell and synapse types as well as the distribution of synaptic weights in the brains of different species. The emerging high-quality large structural datasets raise the question of what general functional principles can be gleaned from them. Motivated by this question, we developed a statistical mechanical theory of learning in neural networks that incorporates structural information as constraints. We derived an analytical solution for the memory capacity of the perceptron, a basic feedforward model of supervised learning, with constraint on the distribution of its weights. Interestingly, the theory predicts that the reduction in capacity due to the constrained weight-distribution is related to the Wasserstein distance between the cumulative distribution function of the constrained weights and that of the standard normal distribution. To test the theoretical predictions, we use optimal transport theory and information geometry to develop an SGD-based algorithm to find weights that simultaneously learn the input-output task and satisfy the distribution constraint. We show that training in our algorithm can be interpreted as geodesic flows in the Wasserstein space of probability distributions. Given a parameterized family of weight distributions, our theory predicts the shape of the distribution with optimal parameters. We apply our theory to map out the experimental parameter landscape for the estimated distribution of synaptic weights in mammalian cortex and show that our theory’s prediction for optimal distribution is close to the experimentally measured value. We further developed a statistical mechanical theory for teacher-student perceptron rule learning and ask for the best way for the student to incorporate prior knowledge of the rule (i.e., the teacher). Our theory shows that it is beneficial for the learner to adopt different prior weight distributions during learning, and shows that distribution-constrained learning outperforms unconstrained and sign-constrained learning. Our theory and algorithm provide novel strategies for incorporating prior knowledge about weights into learning, and reveal a powerful connection between structure and function in neural networks. Weishun Zhong, Ben Sorscher, Haim Sompolinsky |
NeurIPS | 4 |
| 2022 | The spectrum of covariance matrices of randomly connected recurrent neuronal networks with linear dynamicsabstractA key question in theoretical neuroscience is the relation between the connectivity structure and the collective dynamics of a network of neurons. Here we study the connectivity-dynamics relation as reflected in the distribution of eigenvalues of the covariance matrix of the dynamic fluctuations of the neuronal activities, which is closely related to the network dynamics' Principal Component Analysis (PCA) and the associated effective dimensionality. We consider the spontaneous fluctuations around a steady state in a randomly connected recurrent network of stochastic neurons. An exact analytical expression for the covariance eigenvalue distribution in the large-network limit can be obtained using results from random matrices. The distribution has a finitely supported smooth bulk spectrum and exhibits an approximate power-law tail for coupling matrices near the critical edge. We generalize the results to include second-order connectivity motifs and discuss extensions to excitatory-inhibitory networks. The theoretical results are compared with those from finite-size networks and the effects of temporal and spatial sampling are studied. Preliminary application to whole-brain imaging data is presented. Using simple connectivity models, our work provides theoretical predictions for the covariance spectrum, a fundamental property of recurrent neuronal dynamics, that can be compared with experimental data. Yu Hu 0008, Haim Sompolinsky |
PLoS Comput. Biol. | 2 |
| 2020 | High-dimensional dynamics of generalization error in neural networksabstractWe perform an analysis of the average generalization dynamics of large neural networks trained using gradient descent. We study the practically-relevant "high-dimensional" regime where the number of free parameters in the network is on the order of or even larger than the number of examples in the dataset. Using random matrix theory and exact solutions in linear models, we derive the generalization error and training error dynamics of learning and analyze how they depend on the dimensionality of data and signal to noise ratio of the learning problem. We find that the dynamics of gradient descent learning naturally protect against overtraining and overfitting in large networks. Overtraining is worst at intermediate network sizes, when the effective number of free parameters equals the number of samples, and thus can be reduced by making a network smaller or larger. Additionally, in the high-dimensional regime, low generalization error requires starting with small initial weights. We then turn to non-linear neural networks, and show that making networks very large does not harm their generalization performance. On the contrary, it can in fact reduce overtraining, even without early stopping or regularization of any sort. We identify two novel phenomena underlying this behavior in overcomplete models: first, there is a frozen subspace of the weights in which no learning occurs under gradient descent; and second, the statistical properties of the high-dimensional regime yield better-conditioned input correlations which protect against overtraining. We demonstrate that standard application of theories such as Rademacher complexity are inaccurate in predicting the generalization performance of deep neural networks, and derive an alternative bound which incorporates the frozen subspace and conditioning effects and qualitatively matches the behavior observed in simulation. Madhu Advani, Andrew M. Saxe, Haim Sompolinsky |
Neural Networks | 3 |
| 2019 | Functional diversity among sensory neurons from efficient coding principlesabstractIn many sensory systems the neural signal is coded by the coordinated response of heterogeneous populations of neurons. What computational benefit does this diversity confer on information processing? We derive an efficient coding framework assuming that neurons have evolved to communicate signals optimally given natural stimulus statistics and metabolic constraints. Incorporating nonlinearities and realistic noise, we study optimal population coding of the same sensory variable using two measures: maximizing the mutual information between stimuli and responses, and minimizing the error incurred by the optimal linear decoder of responses. Our theory is applied to a commonly observed splitting of sensory neurons into ON and OFF that signal stimulus increases or decreases, and to populations of monotonically increasing responses of the same type, ON. Depending on the optimality measure, we make different predictions about how to optimally split a population into ON and OFF, and how to allocate the firing thresholds of individual neurons given realistic stimulus distributions and noise, which accord with certain biases observed experimentally. Julijana Gjorgjieva, Markus Meister, Haim Sompolinsky |
PLoS Comput. Biol. | 3 |
| 2018 | Learning Data Manifolds with a Cutting Plane MethodabstractWe consider the problem of classifying data manifolds where each manifold represents invariances that are parameterized by continuous degrees of freedom. Conventional data augmentation methods rely on sampling large numbers of training examples from these manifolds. Instead, we propose an iterative algorithm, [Formula: see text], based on a cutting plane approach that efficiently solves a quadratic semi-infinite programming problem to find the maximum margin solution. We provide a proof of convergence as well as a polynomial bound on the number of iterations required for a desired tolerance in the objective function. The efficiency and performance of [Formula: see text] are demonstrated in high-dimensional simulations and on image manifolds generated from the ImageNet data set. Our results indicate that [Formula: see text] is able to rapidly learn good classifiers and shows superior generalization performance compared with conventional maximum margin methods using data augmentation methods. SueYeon Chung, Uri Cohen, Haim Sompolinsky, Daniel D. Lee |
Neural Comput. | 3 |
| 2018 | Coherent chaos in a recurrent neural network with structured connectivityabstractWe present a simple model for coherent, spatially correlated chaos in a recurrent neural network. Networks of randomly connected neurons exhibit chaotic fluctuations and have been studied as a model for capturing the temporal variability of cortical activity. The dynamics generated by such networks, however, are spatially uncorrelated and do not generate coherent fluctuations, which are commonly observed across spatial scales of the neocortex. In our model we introduce a structured component of connectivity, in addition to random connections, which effectively embeds a feedforward structure via unidirectional coupling between a pair of orthogonal modes. Local fluctuations driven by the random connectivity are summed by an output mode and drive coherent activity along an input mode. The orthogonality between input and output mode preserves chaotic fluctuations by preventing feedback loops. In the regime of weak structured connectivity we apply a perturbative approach to solve the dynamic mean-field equations, showing that in this regime coherent fluctuations are driven passively by the chaos of local residual fluctuations. When we introduce a row balance constraint on the random connectivity, stronger structured connectivity puts the network in a distinct dynamical regime of self-tuned coherent chaos. In this regime the coherent component of the dynamics self-adjusts intermittently to yield periods of slow, highly coherent chaos. The dynamics display longer time-scales and switching-like activity. We show how in this regime the dynamics depend qualitatively on the particular realization of the connectivity matrix: a complex leading eigenvalue can yield coherent oscillatory chaos while a real leading eigenvalue can yield chaos with broken symmetry. The level of coherence grows with increasing strength of structured connectivity until the dynamics are almost entirely constrained to a single spatial mode. We examine the effects of network-size scaling and show that these results are not finite-size effects. Finally, we show that in the regime of weak structured connectivity, coherent chaos emerges also for a generalized structured connectivity with multiple input-output modes. Itamar Daniel Landau, Haim Sompolinsky |
PLoS Comput. Biol. | 2 |
| 2016 | Optimal Architectures in a Solvable Model of Deep NetworksabstractDeep neural networks have received a considerable attention due to the success of their training for real world machine learning applications. They are also of great interest to the understanding of sensory processing in cortical sensory hierarchies. The purpose of this work is to advance our theoretical understanding of the computational benefits of these architectures. Using a simple model of clustered noisy inputs and a simple learning rule, we provide analytically derived recursion relations describing the propagation of the signals along the deep network. By analysis of these equations, and defining performance measures, we show that these model networks have optimal depths. We further explore the dependence of the optimal architecture on the system parameters. Jonathan Kadmon, Haim Sompolinsky |
NIPS | 2 |
| 2012 | How the Brain Generates MovementabstractIn this study, we assume that the brain uses a general-purpose pattern generator to transform static commands into basic movement segments. We hypothesize that this pattern generator includes an oscillator whose complete cycle generates a single movement segment. In order to demonstrate this hypothesis, we construct an oscillator-based model of movement generation. The model includes an oscillator that generates harmonic outputs whose frequency and amplitudes can be modulated by external inputs. The harmonic outputs drive a number of integrators, each activating a single muscle. The model generates muscle activation patterns composed of rectilinear and harmonic terms. We show that rectilinear and fundamental harmonic terms account for known properties of natural movements, such as the invariant bell-shaped hand velocity profile during reaching. We implement these dynamics by a neural network model and characterize the tuning properties of the neural integrator cells, the neural oscillator cells, and the inputs to the system. Finally, we propose a method to test our hypothesis that a neural oscillator is a central component in the generation of voluntary movement. Uri Rokni, Haim Sompolinsky |
Neural Comput. | 2 |
| 2010 | Short-term memory in neuronal networks through dynamical compressed sensingabstractRecent proposals suggest that large, generic neuronal networks could store memory traces of past input sequences in their instantaneous state. Such a proposal raises important theoretical questions about the duration of these memory traces and their dependence on network size, connectivity and signal statistics. Prior work, in the case of gaussian input sequences and linear neuronal networks, shows that the duration of memory traces in a network cannot exceed the number of neurons (in units of the neuronal time constant), and that no network can out-perform an equivalent feedforward network. However a more ethologically relevant scenario is that of sparse input sequences. In this scenario, we show how linear neural networks can essentially perform compressed sensing (CS) of past inputs, thereby attaining a memory capacity that {\it exceeds} the number of neurons. This enhanced capacity is achieved by a class of ``orthogonal recurrent networks and not by feedforward networks or generic recurrent networks. We exploit techniques from the statistical physics of disordered systems to analytically compute the decay of memory traces in such networks as a function of network size, signal sparsity and integration time. Alternately, viewed purely from the perspective of CS, this work introduces a new ensemble of measurement matrices derived from dynamical systems, and provides a theoretical analysis of their asymptotic performance." Surya Ganguli, Haim Sompolinsky |
NIPS | 2 |
| 2010 | Inferring Stimulus Selectivity from the Spatial Structure of Neural Network DynamicsabstractHow are the spatial patterns of spontaneous and evoked population responses related? We study the impact of connectivity on the spatial pattern of fluctuations in the input-generated response of a neural network, by comparing the distribution of evoked and intrinsically generated activity across the different units. We develop a complementary approach to principal component analysis in which separate high-variance directions are typically derived for each input condition. We analyze subspace angles to compute the difference between the shapes of trajectories corresponding to different network states, and the orientation of the low-dimensional subspaces that driven trajectories occupy within the full space of neuronal activity. In addition to revealing how the spatiotemporal structure of spontaneous activity affects input-evoked responses, these methods can be used to infer input selectivity induced by network dynamics from experimentally accessible measures of spontaneous activity (e.g. from voltage- or calcium-sensitive optical imaging experiments). We conclude that the absence of a detailed spatial map of afferent inputs and cortical connectivity does not limit our ability to design spatially extended stimuli that evoke strong responses. Kanaka Rajan, L. F. Abbott, Haim Sompolinsky |
NIPS | 3 |
| 2009 | Stimulus-Dependent Correlations in Threshold-Crossing Spiking NeuronsabstractWe consider a threshold-crossing spiking process as a simple model for the activity within a population of neurons. Assuming that these neurons are driven by a common fluctuating input with gaussian statistics, we evaluate the cross-correlation of spike trains in pairs of model neurons with different thresholds. This correlation function tends to be asymmetric in time, indicating a preference for the neuron with the lower threshold to fire before the one with the higher threshold, even if their inputs are identical. The relationship between these results and spike statistics in other models of neural activity is explored. In particular, we compare our model with an integrate-and-fire model in which the membrane voltage resets following each spike. The qualitative properties of spike cross-correlations, emerging from the threshold-crossing model, are similar to those of bursting events in the integrate-and-fire model. This is particularly true for generalized integrate-and-fire models in which spikes tend to occur in bursts, as observed, for example, in retinal ganglion cells driven by a rapidly fluctuating visual stimulus. The threshold-crossing model thus provides a simple, analytically tractable description of event onsets in these neurons. Yoram Burak, Sam Lewallen, Haim Sompolinsky |
Neural Comput. | 3 |
| 2006 | Implications of Neuronal Diversity on Population CodingabstractIn many cortical and subcortical areas, neurons are known to modulate their average firing rate in response to certain external stimulus features. It is widely believed that information about the stimulus features is coded by a weighted average of the neural responses. Recent theoretical studies have shown that the information capacity of such a coding scheme is very limited in the presence of the experimentally observed pairwise correlations. However, central to the analysis of these studies was the assumption of a homogeneous population of neurons. Experimental findings show a considerable measure of heterogeneity in the response properties of different neurons. In this study, we investigate the effect of neuronal heterogeneity on the information capacity of a correlated population of neurons. We show that information capacity of a heterogeneous network is not limited by the correlated noise, but scales linearly with the number of cells in the population. This information cannot be extracted by the population vector readout, whose accuracy is greatly suppressed by the correlated noise. On the other hand, we show that an optimal linear readout that takes into account the neuronal heterogeneity can extract most of this information. We study analytically the nature of the dependence of the optimal linear readout weights on the neuronal diversity. We show that simple online learning can generate readout weights with the appropriate dependence on the neuronal diversity, thereby yielding efficient readout. Maoz Shamir, Haim Sompolinsky |
Neural Comput. | 2 |
| 2004 | Nonlinear Population CodesabstractTheoretical and experimental studies of distributed neuronal representations of sensory and behavioral variables usually assume that the tuning of the mean firing rates is the main source of information. However, recent theoretical studies have investigated the effect of cross-correlations in the trial-to-trial fluctuations of the neuronal responses on the accuracy of the representation. Assuming that only the first-order statistics of the neuronal responses are tuned to the stimulus, these studies have shown that in the presence of correlations, similar to those observed experimentally in cortical ensembles of neurons, the amount of information in the population is limited, yielding nonzero error levels even in the limit of infinitely large populations of neurons. In this letter, we study correlated neuronal populations whose higher-order statistics, and in particular response variances, are also modulated by the stimulus. Weask two questions: Does the correlated noise limit the accuracy of the neuronal representation of the stimulus? and, How can a biological mechanism extract most of the information embedded in the higher-order statistics of the neuronal responses? Specifically, we address these questions in the context of a population of neurons coding an angular variable. We show that the information embedded in the variances grows linearly with the population size despite the presence of strong correlated noise. This information cannot be extracted by linear readout schemes, including the linear population vector. Instead, we propose a bilinear readout scheme that involves spatial decorrelation, quadratic nonlinearity, and population vector summation. We show that this nonlinear population vector scheme yields accurate estimates of stimulus parameters, with an efficiency that grows linearly with the population size. This code can be implemented using biologically plausible neurons. Maoz Shamir, Haim Sompolinsky |
Neural Comput. | 2 |
| 2003 | Rate Models for Conductance-Based Cortical Neuronal NetworksabstractPopulation rate models provide powerful tools for investigating the principles that underlie the cooperative function of large neuronal systems. However, biophysical interpretations of these models have been ambiguous. Hence, their applicability to real neuronal systems and their experimental validation have been severely limited. In this work, we show that conductance-based models of large cortical neuronal networks can be described by simplified rate models, provided that the network state does not possess a high degree of synchrony. We first derive a precise mapping between the parameters of the rate equations and those of the conductance-based network models for time-independent inputs. This mapping is based on the assumption that the effect of increasing the cell's input conductance on its f-I curve is mainly subtractive. This assumption is confirmed by a single compartment Hodgkin-Huxley type model with a transient potassium A-current. This approach is applied to the study of a network model of a hypercolumn in primary visual cortex. We also explore extensions of the rate model to the dynamic domain by studying the firing-rate response of our conductance-based neuron to time-dependent noisy inputs. We show that the dynamics of this response can be approximated by a time-dependent second-order differential equation. This phenomenological single-cell rate model is used to calculate the response of a conductance-based network to time-dependent inputs. Oren Shriki, David Hansel, Haim Sompolinsky |
Neural Comput. | 3 |
| 2001 | Correlation Codes in Neuronal PopulationsabstractPopulation codes often rely on the tuning of the mean responses to the stimulus parameters. However, this information can be greatly sup- pressed by long range correlations. Here we study the efficiency of cod- ing information in the second order statistics of the population responses. We show that the Fisher Information of this system grows linearly with the size of the system. We propose a bilinear readout model for extract- ing information from correlation codes, and evaluate its performance in discrimination and estimation tasks. It is shown that the main source of information in this system is the stimulus dependence of the variances of the single neuron responses. Maoz Shamir, Haim Sompolinsky |
NIPS | 2 |
| 2000 | An Information Maximization Approach to Overcomplete and Recurrent RepresentationsabstractThe principle of maximizing mutual information is applied to learning overcomplete and recurrent representations. The underlying model con(cid:173) sists of a network of input units driving a larger number of output units with recurrent interactions. In the limit of zero noise, the network is de(cid:173) terministic and the mutual information can be related to the entropy of the output units. Maximizing this entropy with respect to both the feed(cid:173) forward connections as well as the recurrent interactions results in simple learning rules for both sets of parameters. The conventional independent components (ICA) learning algorithm can be recovered as a special case where there is an equal number of output units and no recurrent con(cid:173) nections. The application of these new learning rules is illustrated on a simple two-dimensional input example. Oren Shriki, Haim Sompolinsky, Daniel D. Lee |
NIPS | 2 |
| 1999 | Algorithms for Independent Components Analysis and Higher Order Statistics
Daniel D. Lee, Uri Rokni, Haim Sompolinsky |
NIPS | 3 |
| 1998 | Learning a Continuous Hidden Variable Model for Binary Data
Daniel D. Lee, Haim Sompolinsky |
NIPS | 2 |
| 1998 | The Effect of Correlations on the Fisher Information of Population Codes
Hyoungsoo Yoon, Haim Sompolinsky |
NIPS | 2 |
| 1998 | Chaotic Balanced State in a Model Of Cortical CircuitsabstractThe nature and origin of the temporal irregularity in the electrical activity of cortical neurons in vivo are not well understood. We consider the hypothesis that this irregularity is due to a balance of excitatory and inhibitory currents into the cortical cells. We study a network model with excitatory and inhibitory populations of simple binary units. The internal feedback is mediated by relatively large synaptic strengths, so that the magnitude of the total excitatory and inhibitory feedback is much larger than the neuronal threshold. The connectivity is random and sparse. The mean number of connections per unit is large, though small compared to the total number of cells in the network. The network also receives a large, temporally regular input from external sources. We present an analytical solution of the mean-field theory of this model, which is exact in the limit of large network size. This theory reveals a new cooperative stationary state of large networks, which we term a balanced state. In this state, a balance between the excitatory and inhibitory inputs emerges dynamically for a wide range of parameters, resulting in a net input whose temporal fluctuations are of the same order as its mean. The internal synaptic inputs act as a strong negative feedback, which linearizes the population responses to the external drive despite the strong nonlinearity of the individual cells. This feedback also greatly stabilizes the system's state and enables it to track a time-dependent input on time scales much shorter than the time constant of a single cell. The spatiotemporal statistics of the balanced state are calculated. It is shown that the autocorrelations decay on a short time scale, yielding an approximate Poissonian temporal statistics. The activity levels of single cells are broadly distributed, and their distribution exhibits a skewed shape with a long power-law tail. The chaotic nature of the balanced state is revealed by showing that the evolution of the microscopic state of the network is extremely sensitive to small deviations in its initial conditions. The balanced state generated by the sparse, strong connections is an asynchronous chaotic state. It is accompanied by weak spatial cross-correlations, the strength of which vanishes in the limit of large network size. This is in contrast to the synchronized chaotic states exhibited by more conventional network models with high connectivity of weak synapses. Carl van Vreeswijk, Haim Sompolinsky |
Neural Comput. | 2 |
| 1996 | Neural network models of perceptual learning of angle discriminationabstractWe study neural network models of discriminating between stimuli with two similar angles, using the two-alternative forced choice (2AFC) paradigm. Two network architectures are investigated: a two-layer perceptron network and a gating network. In the two-layer network all hidden units contribute to the decision at all angles, while in the other architecture the gating units select, for each stimulus, the appropriate hidden units that will dominate the decision. We find that both architectures can perform the task reasonably well for all angles. Perceptual learning has been modeled by training the networks to perform the task, using unsupervised Hebb learning algorithms with pairs of stimuli at fixed angles theta and delta theta. Perceptual transfer is studied by measuring the performance of the network on stimuli with theta' not equal to theta. The two-layer perceptron shows a partial transfer for angles that are within a distance a from theta, where a is the angular width of the input tuning curves. The change in performance due to learning is positive for angles close to theta, but for magnitude of theta-theta' approximately a it is negative, i.e., its performance after training is worse than before. In contrast, negative transfer can be avoided in the gating network by limiting the effects of learning to hidden units that are optimized for angles that are close to the trained angle. Germán Mato, Haim Sompolinsky |
Neural Comput. | 2 |
| 1994 | On-line Learning of DichotomiesabstractThe performance of on-line algorithms for learning dichotomies is studied. In on-line learn(cid:173) ing, the number of examples P is equivalent to the learning time, since each example is presented only once. The learning curve, or generalization error as a function of P, depends on the schedule at which the learning rate is lowered. For a target that is a perceptron rule, the learning curve of the perceptron algorithm can decrease as fast as p- 1 , if the sched(cid:173) ule is optimized. If the target is not realizable by a perceptron, the perceptron algorithm does not generally converge to the solution with lowest generalization error. For the case of unrealizability due to a simple output noise, we propose a new on-line algorithm for a perceptron yielding a learning curve that can approach the optimal generalization error as fast as p-l/2. We then generalize the perceptron algorithm to any class of thresholded smooth functions learning a target from that class. For "well-behaved" input distributions, if this algorithm converges to the optimal solution, its learning curve can decrease as fast as p-l. N. Barkai, H. Sebastian Seung, Haim Sompolinsky |
NIPS | 3 |
| 1994 | Segmentation by a Network of Oscillators with Stored MemoriesabstractWe propose a model of coupled oscillators with noise that performs segmentation of stimuli using a set of stored images, each consisting of objects and a background. The oscillators' amplitudes encode the spatial and featural distribution of the external stimulus. The coherence of their phases signifies their belonging to the same object. In the learning stage, the couplings between phases are modified in a Hebb-like manner. By mean-field analysis and simulations, we show that an external stimulus whose local features resemble those of one or several of the stored objects generates a selective phase coherence that represents the stored pattern of segmentation. Haim Sompolinsky, Michail Tsodyks |
Neural Comput. | 1 |
| 1993 | Correlation Functions in a Large Stochastic Network
Iris Ginzburg, Haim Sompolinsky |
NIPS | 2 |
| 1993 | Stimulus-Dependent Synchronization of Neuronal AssembliesabstractWe study theoretically how an interaction between assemblies of neuronal oscillators can be modulated by the pattern of external stimuli. It is shown that spatial variations in the stimuli can control the magnitude and phase of the synchronization between the output of neurons with different receptive fields. This modulation emerges from cooperative dynamics in the network, without the need for specialized, activity-dependent synapses. Our results further suggest that the modulation of neuronal interactions by extended features of a stimulus may give rise to complex spatiotemporal fluctuations in the phases of neuronal oscillations. E. R. Grannan, D. Kleinfeld, Haim Sompolinsky |
Neural Comput. | 3 |
| 1992 | Query by CommitteeabstractWe propose an algorithm called query by commitee, in which a committee of students is trained on the same data set. The next query is chosen according to the principle of maximal disagreement. The algorithm is studied for two toy models: the high-low game and perceptron learning of another perceptron. As the number of queries goes to infinity, the committee algorithm yields asymptotically finite information gain. This leads to generalization error that decreases exponentially with the number of examples. This in marked contrast to learning from randomly chosen inputs, for which the information gain approaches zero and the generalization error decreases with a relatively slow inverse power law. We suggest that asymptotically finite information gain may be an important characteristic of good query algorithms. H. Sebastian Seung, Manfred Opper, Haim Sompolinsky |
COLT | 3 |
| 1992 | Processing of Sensory Information by a Network of Oscillators with MemoryabstractWe propose a model of coupled oscillators with noise that performs binding and segmentation of objects using a set of stored images each consisting of figures and a background. The amplitudes of the oscillators encode the spatial and featural distribution of the external stimulus. In the learning stage the couplings between the phases are modified in a Hebb-like manner. By meanfield analysis and simulations we show that an external stimulus whose local features resemble those of one or several of the stored figures causes a selective phase coherence that retrieves the stored pattern of segmentation. Haim Sompolinsky, Michail Tsodyks |
Int. J. Neural Syst. | 1 |