Sara A. Solla

dblp:67/1208 · DBLP profile ↗
← Back
20ranked-venue papers
1as first author
1since 2021 · last 2021
0000-0001-7696-447XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
Transfer learning and domain adaptation · 79% Deep learning architectures and training · 8% Learning theory · 5%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Bioinformatics and computational biology · 55% Medical and health informatics · 45%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Transfer learning and domain adaptation › domain adaptation › distribution adaptation
adversarial domain adaptation
0.412019
Adversarial Domain Adaptation for Stable Brain-Machine Interfaces · ICLR (Poster) 2019
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.412019
Adversarial Domain Adaptation for Stable Brain-Machine Interfaces · ICLR (Poster) 2019
Medical and health informatics
brain-computer interface
0.412019
Adversarial Domain Adaptation for Stable Brain-Machine Interfaces · ICLR (Poster) 2019
Bioinformatics and computational biology › neuroscience › neuroinformatics › neural data analysis
neural signal processing
0.412019
Adversarial Domain Adaptation for Stable Brain-Machine Interfaces · ICLR (Poster) 2019
Bioinformatics and computational biology
computational neuroscience
0.122003
Dopamine Modulation in a Basal Ganglio-cortical Network Implements Saliency-based Gating of Working Memory · NIPS 2003
Dopamine Induced Bistability Enhances Signal Processing in Spiny Neurons · NIPS 2002
Machine learning › Deep learning architectures and training › feedforward neural network
multilayer neural network
0.031996
Learning with Noise and Regularizers in Multilayer Neural Networks · NIPS 1996
Dynamics of On-Line Gradient Descent Learning for Multilayer Neural Networks · NIPS 1995
A statistical approach to learning and generalization in layered neural networks · Proc. IEEE 1990
Machine learning › Learning theory
learning curves
0.021993
Learning Curves: Asymptotic Values and Rate of Convergence · NIPS 1993
A statistical approach to learning and generalization in layered neural networks · Proc. IEEE 1990
Machine learning › Deep learning architectures and training › regularization
noise injection
0.011996
Learning with Noise and Regularizers in Multilayer Neural Networks · NIPS 1996
Machine learning › Deep learning architectures and training
regularization
0.011996
Learning with Noise and Regularizers in Multilayer Neural Networks · NIPS 1996
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
0.011995
Dynamics of On-Line Gradient Descent Learning for Multilayer Neural Networks · NIPS 1995
Machine learning › Optimization for machine learning
online gradient descent
0.011995
Dynamics of On-Line Gradient Descent Learning for Multilayer Neural Networks · NIPS 1995
Knowledge, reasoning and agents › Knowledge representation and reasoning
cognitive modeling
0.012003
Dopamine Modulation in a Basal Ganglio-cortical Network Implements Saliency-based Gating of Working Memory · NIPS 2003
Computer vision › Image recognition and object detection
character recognition
0.011991
Structural Risk Minimization for Character Recognition · NIPS 1991
Machine learning › Learning theory › model selection
structural risk minimization
0.011991
Structural Risk Minimization for Character Recognition · NIPS 1991
Machine learning › Deep learning architectures and training › architecture learning
architecture selection
0.011990
A statistical approach to learning and generalization in layered neural networks · Proc. IEEE 1990
Machine learning › Learning theory › generalization
generalization theory
0.011990
A statistical approach to learning and generalization in layered neural networks · Proc. IEEE 1990
Machine learning › Deep learning architectures and training
loss landscape
0.011990
Second Order Properties of Error Surfaces · NIPS 1990
Machine learning › Learning theory
statistical learning theory
0.011990
A statistical approach to learning and generalization in layered neural networks · Proc. IEEE 1990
Network optimization and economics
admission control
0.011990
Neural Network Implementation of Admission Control · NIPS 1990
Machine learning › Efficient and distributed learning
model compression
0.011989
Optimal Brain Damage · NIPS 1989
Machine learning › Efficient and distributed learning › model compression
pruning
0.011989
Optimal Brain Damage · NIPS 1989
Machine learning › Optimization for machine learning
second-order optimization
0.011989
Optimal Brain Damage · NIPS 1989

Methods — techniques the papers use, named apart from their topics

adversarial training · 0.8computational modeling · 0.2network simulation · 0.1ion channel modeling · 0.1regularization · 0.0noise injection · 0.0gradient descent · 0.0asymptotic analysis · 0.0structural risk minimization · 0.0hessian analysis · 0.0
YearPublicationVenuePosition
2021 Estimating the dimensionality of the manifold underlying multi-electrode neural recordings
abstract
It is generally accepted that the number of neurons in a given brain area far exceeds the number of neurons needed to carry any specific function controlled by that area. For example, motor areas of the human brain contain tens of millions of neurons that control the activation of tens or at most hundreds of muscles. This massive redundancy implies the covariation of many neurons, which constrains the population activity to a low-dimensional manifold within the space of all possible patterns of neural activity. To gain a conceptual understanding of the complexity of the neural activity within a manifold, it is useful to estimate its dimensionality, which quantifies the number of degrees of freedom required to describe the observed population activity without significant information loss. While there are many algorithms for dimensionality estimation, we do not know which are well suited for analyzing neural activity. The objective of this study was to evaluate the efficacy of several representative algorithms for estimating the dimensionality of linearly and nonlinearly embedded data. We generated synthetic neural recordings with known intrinsic dimensionality and used them to test the algorithms' accuracy and robustness. We emulated some of the important challenges associated with experimental data by adding noise, altering the nature of the embedding of the low-dimensional manifold within the high-dimensional recordings, varying the dimensionality of the manifold, and limiting the amount of available data. We demonstrated that linear algorithms overestimate the dimensionality of nonlinear, noise-free data. In cases of high noise, most algorithms overestimated the dimensionality. We thus developed a denoising algorithm based on deep learning, the "Joint Autoencoder", which significantly improved subsequent dimensionality estimation. Critically, we found that all algorithms failed when the intrinsic dimensionality was high (above 20) or when the amount of data used for estimation was low. Based on the challenges we observed, we formulated a pipeline for estimating the dimensionality of experimental neural data.
Ege Altan, Sara A. Solla, Lee E. Miller, Eric J. Perreault
PLoS Comput. Biol.2
2019 Adversarial Domain Adaptation for Stable Brain-Machine Interfaces
Ali Farshchian, Juan Alvaro Gallego, Joseph Paul Cohen, Yoshua Bengio, Lee E. Miller, Sara A. Solla
ICLR (Poster)6
2019 The dynamics of motor learning through the formation of internal models
abstract
A medical student learning to perform a laparoscopic procedure or a recently paralyzed user of a powered wheelchair must learn to operate machinery via interfaces that translate their actions into commands for an external device. Since the user's actions are selected from a number of alternatives that would result in the same effect in the control space of the external device, learning to use such interfaces involves dealing with redundancy. Subjects need to learn an externally chosen many-to-one map that transforms their actions into device commands. Mathematically, we describe this type of learning as a deterministic dynamical process, whose state is the evolving forward and inverse internal models of the interface. The forward model predicts the outcomes of actions, while the inverse model generates actions designed to attain desired outcomes. Both the mathematical analysis of the proposed model of learning dynamics and the learning performance observed in a group of subjects demonstrate a first-order exponential convergence of the learning process toward a particular state that depends only on the initial state of the inverse and forward models and on the sequence of targets supplied to the users. Noise is not only present but necessary for the convergence of learning through the minimization of the difference between actual and predicted outcomes.
Camilla Pierella, Maura Casadio, Ferdinando A. Mussa-Ivaldi, Sara A. Solla
PLoS Comput. Biol.4
2006 Identification of Multiple-Input Systems with Highly Coupled Inputs: Application to EMG Prediction from Multiple Intracortical Electrodes
abstract
A robust identification algorithm has been developed for linear, time-invariant, multiple-input single-output systems, with an emphasis on how this algorithm can be used to estimate the dynamic relationship between a set of neural recordings and related physiological signals. The identification algorithm provides a decomposition of the system output such that each component is uniquely attributable to a specific input signal, and then reduces the complexity of the estimation problem by discarding those input signals that are deemed to be insignificant. Numerical difficulties due to limited input bandwidth and correlations among the inputs are addressed using a robust estimation technique based on singular value decomposition. The algorithm has been evaluated on both simulated and experimental data. The latter involved estimating the relationship between up to 40 simultaneously recorded motor cortical signals and peripheral electromyograms (EMGs) from four upper limb muscles in a freely moving primate. The algorithm performed well in both cases: it provided reliable estimates of the system output and significantly reduced the number of inputs needed for output prediction. For example, although physiological recordings from up to 40 different neuronal signals were available, the input selection algorithm reduced this to 10 neuronal signals that made significant contributions to the recorded EMGs.
David T. Westwick, Eric A. Pohlmeyer, Sara A. Solla, Lee E. Miller, Eric J. Perreault
Neural Comput.3
2003 Dopamine Modulation in a Basal Ganglio-cortical Network Implements Saliency-based Gating of Working Memory
Aaron J. Gruber, Peter Dayan, Boris Gutkin, Sara A. Solla
NIPS4
2002 Dopamine Induced Bistability Enhances Signal Processing in Spiny Neurons
abstract
Single unit activity in the striatum of awake monkeys shows a marked dependence on the expected reward that a behavior will elicit. We present a computational model of spiny neurons, the principal neurons of the striatum, to assess the hypothesis that di(cid:173) rect neuromodulatory effects of dopamine through the activation of D 1 receptors mediate the reward dependency of spiny neuron activity. Dopamine release results in the amplification of key ion currents, leading to the emergence of bistability, which not only modulates the peak firing rate but also introduces a temporal and state dependence of the model's response, thus improving the de(cid:173) tectability of temporally correlated inputs.
Aaron J. Gruber, Sara A. Solla, James C. Houk
NIPS2
1997 Bayesian online learning in the perceptron
Ole Winther, Sara A. Solla
ESANN2
1997 Universal Distribution of Saliencies for Pruning in Layered Neural Networks
abstract
A better understanding of pruning methods based on a ranking of weights according to their saliency in a trained network requires further information on the statistical properties of such saliencies. We focus on two-layer networks with either a linear or nonlinear output unit, and obtain analytic expressions for the distribution of saliencies and their logarithms. Our results reveal unexpected universal properties of the log-saliency distribution and suggest a novel algorithm for saliency-based weight ranking that avoids the numerical cost of second derivative evaluations.
Jan Gorodkin, Lars Kai Hansen, Benny Lautrup, Sara A. Solla
Int. J. Neural Syst.4
1996 Learning with Noise and Regularizers in Multilayer Neural Networks
David Saad, Sara A. Solla
NIPS2
1995 Dynamics of On-Line Gradient Descent Learning for Multilayer Neural Networks
David Saad, Sara A. Solla
NIPS2
1993 Learning Curves: Asymptotic Values and Rate of Convergence
Corinna Cortes, Lawrence D. Jackel, Sara A. Solla, Vladimir Vapnik, John S. Denker
NIPS3
1992 Capacity control in linear classifiers for pattern recognition
abstract
Achieving good performance in statistical pattern recognition requires matching the capacity of the classifier to the amount of training data. If the classifier has too many adjustable parameters (large capacity), it is likely to learn the training data without difficulty, but will probably not generalize properly to patterns that do not belong to the training set. Conversely, if the capacity of the classifier is not large enough, it might not be able to learn the task at all. In between, there is an optimal classifier capacity which ensures the best expected generalization for a given amount of training data. The method of structural risk minimization (SRM) refers to tuning the capacity of the classifier to the available amount of training data. This paper illustrates the method of SRM with several examples of algorithms. Experiments confirm theoretical predictions of performance improvement in application to handwritten digit recognition.>
Isabelle Guyon, Vladimir Vapnik, Bernhard E. Boser, Léon Bottou, Sara A. Solla
ICPR (2)5
1991 Structural Risk Minimization for Character Recognition
Isabelle Guyon, Vladimir Vapnik, Bernhard E. Boser, Léon Bottou, Sara A. Solla
NIPS5
1990 Hardware requirements for neural-net optical character recognition
abstract
Hardware architectures for character recognition are discussed, and choices for possible circuits are outlined. An advanced (and working) reconfigurable neural-net chip that mixes analog and digital processing is described. It is found that different approaches to image recognition often lead to neural-net architectures that have limited connectivity and repeated use of the same set of weights. This architecture is ideal for time-multiplexing (a combined parallel-series processing) on hardware systems that would be too small to evaluate the entire network in parallel. To make this process efficient, a chip needs to have shift registers to format the input data and additional registers to store intermediate results. Within this framework, it is possible to design chips that have broad utility, large connection capacity, and high speed. This was demonstrated by a new chip with 32000 reconfigurable connections
Lawrence D. Jackel, Bernhard E. Boser, John S. Denker, Hans Peter Graf, Yann LeCun, Isabelle Guyon, Donnie Henderson, Richard E. Howard, Wayne E. Hubbard, Sara A. Solla
IJCNN10
1990 Second Order Properties of Error Surfaces
Yann LeCun, Ido Kanter, Sara A. Solla
NIPS3
1990 Neural Network Implementation of Admission Control
Rodolfo A. Milito, Isabelle Guyon, Sara A. Solla
NIPS3
1990 Exhaustive Learning
abstract
Exhaustive exploration of an ensemble of networks is used to model learning and generalization in layered neural networks. A simple Boolean learning problem involving networks with binary weights is numerically solved to obtain the entropy Sm and the average generalization ability Gm as a function of the size m of the training set. Learning curves Gm vs m are shown to depend solely on the distribution of generalization abilities over the ensemble of networks. Such distribution is determined prior to learning, and provides a novel theoretical tool for the prediction of network performance on a specific task.
Daniel B. Schwartz, Vijay K. Samalam, Sara A. Solla, John S. Denker
Neural Comput.3
1990 A statistical approach to learning and generalization in layered neural networks
abstract
A general statistical description of the problem of learning from examples is presented. Learning in layered networks is posed as a search in the network parameter space for a network that minimizes an additive error function of a statistically independent examples. By imposing the equivalence of the minimum error and the maximum likelihood criteria for training the network, the Gibbs distribution on the ensemble of networks with a fixed architecture is derived. The probability of correct prediction of a novel example can be expressed using the ensemble, serving as a measure to the network's generalization ability. The entropy of the prediction distribution is shown to be a consistent measure of the network's performance. The proposed formalism is applied to the problems of selecting an optimal architecture and the prediction of learning curves.>
Esther Levin, Naftali Tishby, Sara A. Solla
Proc. IEEE3
1989 Optimal Brain Damage
Yann LeCun, John S. Denker, Sara A. Solla
NIPS3
1988 Learning contiguity with layered neural networks
Sara A. Solla
Neural Networks1