Friedrich T. Sommer

dblp:41/406 · DBLP profile ↗
← Back
32ranked-venue papers
7as first author
6since 2021 · last 2024
0000-0002-6738-9263ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 28 · 7 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Representation and self-supervised learning · 94% Learning theory · 6%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Emerging computing paradigms · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.732019
Learning Overcomplete, Low Coherence Dictionaries with Linear Inference · J. Mach. Learn. Res. 2019
When Can Dictionary Learning Uniquely Recover Sparse Data From Subsamples? · IEEE Trans. Inf. Theory 2015
Deciphering subsampled data: adaptive compressive sampling as a principle of brain communication · NIPS 2010
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › sparse coding
dictionary learning
0.622019
Learning Overcomplete, Low Coherence Dictionaries with Linear Inference · J. Mach. Learn. Res. 2019
When Can Dictionary Learning Uniquely Recover Sparse Data From Subsamples? · IEEE Trans. Inf. Theory 2015
Emerging computing paradigms
neuromorphic computing
0.612022
Vector Symbolic Architectures as a Computing Framework for Emerging Hardware · Proc. IEEE 2022
Emerging computing paradigms › bio-inspired computing
vector symbolic architecture
0.612022
Vector Symbolic Architectures as a Computing Framework for Emerging Hardware · Proc. IEEE 2022
Machine learning › Representation and self-supervised learning › blind source separation
independent component analysis
0.412019
Learning Overcomplete, Low Coherence Dictionaries with Linear Inference · J. Mach. Learn. Res. 2019
Emerging computing paradigms › neuromorphic computing
hyperdimensional computing
0.212022
Vector Symbolic Architectures as a Computing Framework for Emerging Hardware · Proc. IEEE 2022
Machine learning › Learning theory
compressed sensing
0.112010
Deciphering subsampled data: adaptive compressive sampling as a principle of brain communication · NIPS 2010
Bioinformatics and computational biology › computational neuroscience
neural coding
0.112010
Deciphering subsampled data: adaptive compressive sampling as a principle of brain communication · NIPS 2010
Machine learning › Representation and self-supervised learning
associative memory
0.011997
Bidirectional Retrieval from Associative Memory · NIPS 1997
Information retrieval
retrieval models
0.011997
Bidirectional Retrieval from Associative Memory · NIPS 1997

Methods — techniques the papers use, named apart from their topics

combinatorial matrix theory · 0.4overcomplete representation · 0.4linear inference · 0.4coherence control · 0.4sparse coding · 0.2compressive sampling · 0.2bidirectional retrieval · 0.0
YearPublicationVenuePosition
2024 Computing With Residue Numbers in High-Dimensional Representation
abstract
, a computing framework that unifies residue number systems with an algebra defined over random, high-dimensional vectors. We show how residue numbers can be represented as high-dimensional vectors in a manner that allows algebraic operations to be performed with component-wise, parallelizable operations on the vector elements. The resulting framework, when combined with an efficient method for factorizing high-dimensional vectors, can represent and operate on numerical values over a large dynamic range using vastly fewer resources than previous methods, and it exhibits impressive robustness to noise. We demonstrate the potential for this framework to solve computationally difficult problems in visual perception and combinatorial optimization, showing improvement over baseline methods. More broadly, the framework provides a possible account for the computational operations of grid cells in the brain, and it suggests new machine learning architectures for representing and manipulating numerical data.
Christopher J. Kymn, Denis Kleyko, Edward Paxon Frady, Connor Bybee, Pentti Kanerva, Friedrich T. Sommer, Bruno A. Olshausen
Neural Comput.6
2024 Perceptron Theory Can Predict the Accuracy of Neural Networks
abstract
Multilayer neural networks set the current state of the art for many technical classification problems. But, these networks are still, essentially, black boxes in terms of analyzing them and predicting their performance. Here, we develop a statistical theory for the one-layer perceptron and show that it can predict performances of a surprisingly large variety of neural networks with different architectures. A general theory of classification with perceptrons is developed by generalizing an existing theory for analyzing reservoir computing models and connectionist models for symbolic reasoning known as vector symbolic architectures. Our statistical theory offers three formulas leveraging the signal statistics with increasing detail. The formulas are analytically intractable, but can be evaluated numerically. The description level that captures maximum details requires stochastic sampling methods. Depending on the network model, the simpler formulas already yield high prediction accuracy. The quality of the theory predictions is assessed in three experimental settings, a memorization task for echo state networks (ESNs) from reservoir computing literature, a collection of classification datasets for shallow randomly connected networks, and the ImageNet dataset for deep convolutional neural networks. We find that the second description level of the perceptron theory can predict the performance of types of ESNs, which could not be described previously. Furthermore, the theory can predict deep multilayer neural networks by being applied to their output layer. While other methods for prediction of neural networks performance commonly require to train an estimator model, the proposed theory requires only the first two moments of the distribution of the postsynaptic sums in the output neurons. Moreover, the perceptron theory compares favorably to other methods that do not rely on training an estimator model.
Denis Kleyko, Antonello Rosato, Edward Paxon Frady, Massimo Panella, Friedrich T. Sommer
IEEE Trans. Neural Networks Learn. Syst.5
2023 Efficient Decoding of Compositional Structure in Holistic Representations
abstract
We investigate the task of retrieving information from compositional distributed representations formed by hyperdimensional computing/vector symbolic architectures and present novel techniques that achieve new information rate bounds. First, we provide an overview of the decoding techniques that can be used to approach the retrieval task. The techniques are categorized into four groups. We then evaluate the considered techniques in several settings that involve, for example, inclusion of external noise and storage elements with reduced precision. In particular, we find that the decoding techniques from the sparse coding and compressed sensing literature (rarely used for hyperdimensional computing/vector symbolic architectures) are also well suited for decoding information from the compositional distributed representations. Combining these decoding techniques with interference cancellation ideas from communications improves previously reported bounds (Hersche et al., 2021) of the information rate of the distributed representations from 1.20 to 1.40 bits per dimension for smaller codebooks and from 0.60 to 1.26 bits per dimension for larger codebooks.
Denis Kleyko, Connor Bybee, Ping-Chen Huang, Christopher J. Kymn, Bruno A. Olshausen, Edward Paxon Frady, Friedrich T. Sommer
Neural Comput.7
2023 Variable Binding for Sparse Distributed Representations: Theory and Applications
abstract
Variable binding is a cornerstone of symbolic reasoning and cognition. But how binding can be implemented in connectionist models has puzzled neuroscientists, cognitive psychologists, and neural network researchers for many decades. One type of connectionist model that naturally includes a binding operation is vector symbolic architectures (VSAs). In contrast to other proposals for variable binding, the binding operation in VSAs is dimensionality-preserving, which enables representing complex hierarchical data structures, such as trees, while avoiding a combinatoric expansion of dimensionality. Classical VSAs encode symbols by dense randomized vectors, in which information is distributed throughout the entire neuron population. By contrast, in the brain, features are encoded more locally, by the activity of single neurons or small groups of neurons, often forming sparse vectors of neural activation. Following Laiho et al. (2015), we explore symbolic reasoning with a special case of sparse distributed representations. Using techniques from compressed sensing, we first show that variable binding in classical VSAs is mathematically equivalent to tensor product binding between sparse feature vectors, another well-known binding operation which increases dimensionality. This theoretical result motivates us to study two dimensionality-preserving binding methods that include a reduction of the tensor matrix into a single sparse vector. One binding method for general sparse vectors uses random projections, the other, block-local circular convolution, is defined for sparse vectors with block structure, sparse block-codes. Our experiments reveal that block-local circular convolution binding has ideal properties, whereas random projection based binding also works, but is lossy. We demonstrate in example applications that a VSA with block-local circular convolution and sparse block-codes reaches similar performance as classical VSAs. Finally, we discuss our results in the context of neuroscience and neural networks.
Edward Paxon Frady, Denis Kleyko, Friedrich T. Sommer
IEEE Trans. Neural Networks Learn. Syst.3
2022 Vector Symbolic Architectures as a Computing Framework for Emerging Hardware
abstract
(also known as Hyperdimensional Computing). This framework is well suited for implementation in stochastic, emerging hardware and it naturally expresses the types of cognitive operations required for Artificial Intelligence (AI). We demonstrate in this article that the field-like algebraic structure of Vector Symbolic Architectures offers simple but powerful operations on high-dimensional vectors that can support all data structures and manipulations relevant to modern computing. In addition, we illustrate the distinguishing feature of Vector Symbolic Architectures, "computing in superposition," which sets it apart from conventional computing. It also opens the door to efficient solutions to the difficult combinatorial search problems inherent in AI applications. We sketch ways of demonstrating that Vector Symbolic Architectures are computationally universal. We see them acting as a framework for computing with distributed representations that can play a role of an abstraction layer for emerging computing hardware. This article serves as a reference for computer architects by illustrating the philosophy behind Vector Symbolic Architectures, techniques of distributed computing with them, and their relevance to emerging computing hardware, such as neuromorphic computing.
Denis Kleyko, Mike Davies 0002, Edward Paxon Frady, Pentti Kanerva, Spencer J. Kent, Bruno A. Olshausen, Evgeny Osipov, Jan M. Rabaey, Dmitri A. Rachkovskij, Abbas Rahimi, Friedrich T. Sommer
Proc. IEEE11
2022 Cellular Automata Can Reduce Memory Requirements of Collective-State Computing
abstract
Various non-classical approaches of distributed information processing, such as neural networks, computation with Ising models, reservoir computing, vector symbolic architectures, and others, employ the principle of collective-state computing. In this type of computing, the variables relevant in a computation are superimposed into a single high-dimensional state vector, the collective-state. The variable encoding uses a fixed set of random patterns, which has to be stored and kept available during the computation. Here we show that an elementary cellular automaton with rule 90 (CA90) enables space-time tradeoff for collective-state computing models that use random dense binary representations, i.e., memory requirements can be traded off with computation running CA90. We investigate the randomization behavior of CA90, in particular, the relation between the length of the randomization period and the size of the grid, and how CA90 preserves similarity in the presence of the initialization noise. Based on these analyses we discuss how to optimize a collective-state computing model, in which CA90 expands representations on the fly from short seed patterns - rather than storing the full set of random patterns. The CA90 expansion is applied and tested in concrete scenarios using reservoir computing and vector symbolic architectures. Our experimental results show that collective-state computing with CA90 expansion performs similarly compared to traditional collective-state models, in which random patterns are generated initially by a pseudo-random number generator and then stored in a large memory.
Denis Kleyko, Edward Paxon Frady, Friedrich T. Sommer
IEEE Trans. Neural Networks Learn. Syst.3
2020 Resonator Networks, 1: An Efficient Solution for Factoring High-Dimensional, Distributed Representations of Data Structures
abstract
The ability to encode and manipulate data structures with distributed neural representations could qualitatively enhance the capabilities of traditional neural networks by supporting rule-based symbolic reasoning, a central property of cognition. Here we show how this may be accomplished within the framework of Vector Symbolic Architectures (VSAs) (Plate, 1991; Gayler, 1998; Kanerva, 1996), whereby data structures are encoded by combining high-dimensional vectors with operations that together form an algebra on the space of distributed representations. In particular, we propose an efficient solution to a hard combinatorial search problem that arises when decoding elements of a VSA data structure: the factorization of products of multiple codevectors. Our proposed algorithm, called a resonator network, is a new type of recurrent neural network that interleaves VSA multiplication operations and pattern completion. We show in two examples-parsing of a tree-like data structure and parsing of a visual scene-how the factorization problem arises and how the resonator network can solve it. More broadly, resonator networks open the possibility of applying VSAs to myriad artificial intelligence problems in real-world domains. The companion article in this issue (Kent, Frady, Sommer, & Olshausen, 2020) presents a rigorous analysis and evaluation of the performance of resonator networks, showing it outperforms alternative approaches.
Edward Paxon Frady, Spencer J. Kent, Bruno A. Olshausen, Friedrich T. Sommer
Neural Comput.4
2020 Resonator Networks, 2: Factorization Performance and Capacity Compared to Optimization-Based Methods
abstract
We develop theoretical foundations of resonator networks, a new type of recurrent neural network introduced in Frady, Kent, Olshausen, and Sommer (2020), a companion article in this issue, to solve a high-dimensional vector factorization problem arising in Vector Symbolic Architectures. Given a composite vector formed by the Hadamard product between a discrete set of high-dimensional vectors, a resonator network can efficiently decompose the composite into these factors. We compare the performance of resonator networks against optimization-based methods, including Alternating Least Squares and several gradient-based algorithms, showing that resonator networks are superior in several important ways. This advantage is achieved by leveraging a combination of nonlinear dynamics and searching in superposition, by which estimates of the correct solution are formed from a weighted superposition of all possible solutions. While the alternative methods also search in superposition, the dynamics of resonator networks allow them to strike a more effective balance between exploring the solution space and exploiting local information to drive the network toward probable solutions. Resonator networks are not guaranteed to converge, but within a particular regime they almost always do. In exchange for relaxing the guarantee of global convergence, resonator networks are dramatically more effective at finding factorizations than all alternative approaches considered.
Spencer J. Kent, Edward Paxon Frady, Friedrich T. Sommer, Bruno A. Olshausen
Neural Comput.3
2019 Learning Overcomplete, Low Coherence Dictionaries with Linear Inference
abstract
Finding overcomplete latent representations of data has applications in data analysis, signal processing, machine learning, theoretical neuroscience and many other fields. In an overcomplete representation, the number of latent features exceeds the data dimensionality, which is useful when the data is undersampled by the measurements (compressed sensing or information bottlenecks in neural systems) or composed from multiple complete sets of linear features, each spanning the data space. Independent Components Analysis (ICA) is a linear technique for learning sparse latent representations, which typically has a lower computational cost than sparse coding, a linear generative model which requires an iterative, nonlinear inference step. While well suited for finding complete representations, we show that overcompleteness poses a challenge to existing ICA algorithms. Specifically, the coherence control used in existing ICA and other dictionary learning algorithms, necessary to prevent the formation of duplicate dictionary features, is ill-suited in the overcomplete case. We show that in the overcomplete case, several existing ICA algorithms have undesirable global minima that maximize coherence. We provide a theoretical explanation of these failures and, based on the theory, propose improved coherence control costs for overcomplete ICA algorithms. Further, by comparing ICA algorithms to the computationally more expensive sparse coding on synthetic data, we show that the limited applicability of overcomplete, linear inference can be extended with the proposed cost functions. Finally, when trained on natural images, we show that the coherence control biases the exploration of the data manifold, sometimes yielding suboptimal, coherent solutions. All told, this study contributes new insights into and methods for coherence control for linear ICA, some of which are applicable to many other nonlinear models.
Jesse A. Livezey, Alejandro F. Bujan, Friedrich T. Sommer
J. Mach. Learn. Res.3
2019 Information integration in large brain networks
abstract
An outstanding problem in neuroscience is to understand how information is integrated across the many modules of the brain. While classic information-theoretic measures have transformed our understanding of feedforward information processing in the brain's sensory periphery, comparable measures for information flow in the massively recurrent networks of the rest of the brain have been lacking. To address this, recent work in information theory has produced a sound measure of network-wide "integrated information", which can be estimated from time-series data. But, a computational hurdle has stymied attempts to measure large-scale information integration in real brains. Specifically, the measurement of integrated information involves a combinatorial search for the informational "weakest link" of a network, a process whose computation time explodes super-exponentially with network size. Here, we show that spectral clustering, applied on the correlation matrix of time-series data, provides an approximate but robust solution to the search for the informational weakest link of large networks. This reduces the computation time for integrated information in large systems from longer than the lifespan of the universe to just minutes. We evaluate this solution in brain-like systems of coupled oscillators as well as in high-density electrocortigraphy data from two macaque monkeys, and show that the informational "weakest link" of the monkey cortex splits posterior sensory areas from anterior association areas. Finally, we use our solution to provide evidence in support of the long-standing hypothesis that information integration is maximized by networks with a high global efficiency, and that modular network structures promote the segregation of information.
Daniel Toker, Friedrich T. Sommer
PLoS Comput. Biol.2
2018 A Theory of Sequence Indexing and Working Memory in Recurrent Neural Networks
abstract
To accommodate structured approaches of neural computation, we propose a class of recurrent neural networks for indexing and storing sequences of symbols or analog data vectors. These networks with randomized input weights and orthogonal recurrent weights implement coding principles previously described in vector symbolic architectures (VSA) and leverage properties of reservoir computing. In general, the storage in reservoir computing is lossy, and crosstalk noise limits the retrieval accuracy and information capacity. A novel theory to optimize memory performance in such networks is presented and compared with simulation experiments. The theory describes linear readout of analog data and readout with winner-take-all error correction of symbolic data as proposed in VSA models. We find that diverse VSA models from the literature have universal performance properties, which are superior to what previous analyses predicted. Further, we propose novel VSA models with the statistically optimal Wiener filter in the readout that exhibit much higher information capacity, in particular for storing analog data. The theory we present also applies to memory buffers, networks with gradual forgetting, which can operate on infinite data streams without memory overflow. Interestingly, we find that different forgetting mechanisms, such as attenuating recurrent weights or neural nonlinearities, produce very similar behavior if the forgetting time constants are matched. Such models exhibit extensive capacity when their forgetting time constant is optimized for given noise conditions and network size. These results enable the design of new types of VSA models for the online processing of data streams.
Edward Paxon Frady, Denis Kleyko, Friedrich T. Sommer
Neural Comput.3
2015 When Can Dictionary Learning Uniquely Recover Sparse Data From Subsamples?
abstract
Sparse coding or sparse dictionary learning has been widely used to recover underlying structure in many kinds of natural data. Here, we provide conditions guaranteeing when this recovery is universal; that is, when sparse codes and dictionaries are unique (up to natural symmetries). Our main tool is a useful lemma in combinatorial matrix theory that allows us to derive bounds on the sample sizes guaranteeing such uniqueness under various assumptions for how training data are generated. Whenever the conditions to one of our theorems are met, any sparsity-constrained learning algorithm that succeeds in reconstructing the data recovers the original sparse codes and dictionary. We also discuss potential applications to neuroscience and data analysis.
Christopher Hillar, Friedrich T. Sommer
IEEE Trans. Inf. Theory2
2010 Adaptive compressed sensing - A new class of self-organizing coding models for neuroscience
abstract
Sparse coding networks, which utilize unsupervised learning to maximize coding efficiency, have successfully reproduced response properties found in primary visual cortex [1]. However, conventional sparse coding models require that the coding circuit can fully sample the sensory data in a one-to-one fashion, a requirement not supported by experimental data from the thalamo-cortical projection. To relieve these strict wiring requirements, we propose a sparse coding network constructed by introducing synaptic learning in the framework of compressed sensing. We demonstrate a new model that evolves biologically realistic, spatially smooth receptive fields despite the fact that the feedforward connectivity subsamples the input and thus the learning must rely on an impoverished and distorted account of the original visual data. Further, we demonstrate that the model could form a general scheme of cortical communication: it can form meaningful representations in a secondary sensory area, which receives input from the primary sensory area through a “compressing” cortico-cortical projection. Finally, we prove that our model belongs to a new class of sparse coding algorithms in which recurrent connections are essential in forming the spatial receptive fields.
William K. Coulter, Christopher Hillar, Guy Isley, Friedrich T. Sommer
ICASSP4
2010 Deciphering subsampled data: adaptive compressive sampling as a principle of brain communication
abstract
A new algorithm is proposed for a) unsupervised learning of sparse representations from subsampled measurements and b) estimating the parameters required for linearly reconstructing signals from the sparse codes. We verify that the new algorithm performs efficient data compression on par with the recent method of compressive sampling. Further, we demonstrate that the algorithm performs robustly when stacked in several stages or when applied in undercomplete or overcomplete situations. The new algorithm can explain how neural populations in the brain that receive subsampled input through fiber bottlenecks are able to form coherent response properties.
Guy Isley, Christopher Hillar, Friedrich T. Sommer
NIPS3
2010 Memory Capacities for Synaptic and Structural Plasticity
abstract
Neural associative networks with plastic synapses have been proposed as computational models of brain functions and also for applications such as pattern recognition and information retrieval. To guide biological models and optimize technical applications, several definitions of memory capacity have been used to measure the efficiency of associative memory. Here we explain why the currently used performance measures bias the comparison between models and cannot serve as a theoretical benchmark. We introduce fair measures for information-theoretic capacity in associative memory that also provide a theoretical benchmark. In neural networks, two types of manipulating synapses can be discerned: synaptic plasticity, the change in strength of existing synapses, and structural plasticity, the creation and pruning of synapses. One of the new types of memory capacity we introduce permits quantifying how structural plasticity can increase the network efficiency by compressing the network structure, for example, by pruning unused synapses. Specifically, we analyze operating regimes in the Willshaw model in which structural plasticity can compress the network structure and push performance to the theoretical benchmark. The amount C of information stored in each synapse can scale with the logarithm of the network size rather than being constant, as in classical Willshaw and Hopfield nets (< or = ln 2 approximately 0.7). Further, the review contains novel technical material: a capacity analysis of the Willshaw model that rigorously controls for the level of retrieval quality, an analysis for memories with a nonconstant number of active units (where C < or = 1/e ln 2 approximately 0.53), and the analysis of the computational complexity of associative memories with and without network compression.
Andreas Knoblauch, Günther Palm, Friedrich T. Sommer
Neural Comput.3
2009 Learning Bimodal Structure in Audio-Visual Data
abstract
A novel model is presented to learn bimodally informative structures from audio-visual signals. The signal is represented as a sparse sum of audio-visual kernels. Each kernel is a bimodal function consisting of synchronous snippets of an audio waveform and a spatio-temporal visual basis function. To represent an audio-visual signal, the kernels can be positioned independently and arbitrarily in space and time. The proposed algorithm uses unsupervised learning to form dictionaries of bimodal kernels from audio-visual material. The basis functions that emerge during learning capture salient audio-visual data structures. In addition, it is demonstrated that the learned dictionary can be used to locate sources of sound in the movie frame. Specifically, in sequences containing two speakers, the algorithm can robustly localize a speaker even in the presence of severe acoustic and visual distracters.
Gianluca Monaci, Pierre Vandergheynst, Friedrich T. Sommer
IEEE Trans. Neural Networks3
2006 Storing and restoring visual input with collaborative rank coding and associative memory
Martin Rehn, Friedrich T. Sommer
Neurocomputing2
2005 Computing with inter-spike interval codes in networks of integrate and fire neurons
Dileep George, Friedrich T. Sommer
Neurocomputing2
2005 Synfire chains with conductance-based neurons: internal timing and coordination with timed input
Friedrich T. Sommer, Thomas Wennekers
Neurocomputing1
2004 Spike-timing-dependent synaptic plasticity can form "zero lag links" for cortical oscillations
Andreas Knoblauch, Friedrich T. Sommer
Neurocomputing2
2003 Synaptic plasticity, conduction delays, and inter-areal phase relations of spike activity in a model of reciprocally connected areas
Andreas Knoblauch, Friedrich T. Sommer
Neurocomputing2
2003 The impact of thalamo-cortical projections on activity spread in cortex
Volker Schmitt, Rolf Kötter, Friedrich T. Sommer
Neurocomputing3
2002 Is voltage-dependent synaptic transmission in NMDA receptors a robust mechanism for working memory?
Andreas Knoblauch, Thomas Wennekers, Friedrich T. Sommer
Neurocomputing3
2001 Associative memory in a pair of cortical cell groups with reciprocal projections
Friedrich T. Sommer, Thomas Wennekers
Neurocomputing1
2001 Coexistence of short and long term memory in a model network of realistic neurons
Urs Vollmer, Friedrich T. Sommer
Neurocomputing2
2001 Associative memory in networks of spiking neurons
Friedrich T. Sommer, Thomas Wennekers
Neural Networks1
2000 On cell assemblies in a cortical column
Friedrich T. Sommer
Neurocomputing1
1999 Gamma-oscillations support optimal retrieval in associative memories of two-compartment neurons
Thomas Wennekers, Friedrich T. Sommer
Neurocomputing2
1999 Improved bidirectional retrieval of sparse patterns stored by Hebbian learning
Friedrich T. Sommer, Günther Palm
Neural Networks1
1998 Bayesian retrieval in associative memories with storage errors
abstract
It is well known that for finite-sized networks, onestep retrieval in the autoassociative Willshaw net is a suboptimal way to extract the information stored in the synapses. Iterative retrieval strategies are much better, but have hitherto only had heuristic justification. We show how they emerge naturally from considerations of probabilistic inference under conditions of noisy and partial input and a corrupted weight matrix. We start from the conditional probability distribution over possible patterns for retrieval. This contains all possible information that is available to an observer of the network and the initial input. Since this distribution is over exponentially many patterns, we use it to develop two approximate, but tractable, iterative retrieval methods. One performs maximum likelihood inference to find the single most likely pattern, using the (negative log of the) conditional probability as a Lyapunov function for retrieval. In physics terms, if storage errors are present, then the modified iterative update equations contain an additional antiferromagnetic interaction term and site dependent threshold values. The second method makes a mean field assumption to optimize a tractable estimate of the full conditional probability distribution. This leads to iterative mean field equations which can be interpreted in terms of a network of neurons with sigmoidal responses but with the same interactions and thresholds as in the maximum likelihood update equations. In the absence of storage errors, both models become very similiar to the Willshaw model, where standard retrieval is iterated using a particular form of linear threshold strategy.
Friedrich T. Sommer, Peter Dayan
IEEE Trans. Neural Networks1
1997 Bidirectional Retrieval from Associative Memory
Friedrich T. Sommer, Günther Palm
NIPS1
1996 Iterative retrieval of sparsely coded associative memory patterns
Friedhelm Schwenker, Friedrich T. Sommer, Günther Palm
Neural Networks2