EDBT 2026 Demo / reviewers in the wild / expert
Todd K. Leen
dblp:82/4320
· DBLP profile ↗
32ranked-venue papers
10as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 10 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
15 papers |
Efficient and distributed learning · 43% Learning theory · 33% Probabilistic and Bayesian machine learning · 9% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 100% |
Topics — the 30 heaviest of 47, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
active learning |
0.6 | 2 | 2018 | A Probabilistic Active Learning Algorithm Based on Fisher Information Ratio · IEEE Trans. Pattern Anal. Mach. Intell. 2018 Asymptotic Analysis of Objectives Based on Fisher Information in Active Learning · J. Mach. Learn. Res. 2017 |
Machine learning › Efficient and distributed learning › active learning
query selection |
0.6 | 2 | 2018 | A Probabilistic Active Learning Algorithm Based on Fisher Information Ratio · IEEE Trans. Pattern Anal. Mach. Intell. 2018 Asymptotic Analysis of Objectives Based on Fisher Information in Active Learning · J. Mach. Learn. Res. 2017 |
Machine learning › Learning theory
fisher information |
0.3 | 1 | 2018 | A Probabilistic Active Learning Algorithm Based on Fisher Information Ratio · IEEE Trans. Pattern Anal. Mach. Intell. 2018 |
Machine learning › Learning theory
information-theoretic learning |
0.3 | 1 | 2018 | A Probabilistic Active Learning Algorithm Based on Fisher Information Ratio · IEEE Trans. Pattern Anal. Mach. Intell. 2018 |
Machine learning › Learning theory › statistical learning theory
asymptotic analysis |
0.3 | 1 | 2017 | Asymptotic Analysis of Objectives Based on Fisher Information in Active Learning · J. Mach. Learn. Res. 2017 |
Machine learning › Kernel, tree and ensemble methods › kernel function
fisher kernel |
0.1 | 1 | 2008 | Hierarchical Fisher Kernels for Longitudinal Data · NIPS 2008 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.1 | 1 | 2008 | A reproducing kernel Hilbert space framework for pairwise time series distances · ICML 2008 |
Machine learning › Probabilistic and Bayesian machine learning › hierarchical modeling
hierarchical bayesian model |
0.1 | 1 | 2008 | Hierarchical Fisher Kernels for Longitudinal Data · NIPS 2008 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel learning |
0.1 | 1 | 2008 | A reproducing kernel Hilbert space framework for pairwise time series distances · ICML 2008 |
Machine learning › Kernel, tree and ensemble methods › kernel methods
reproducing kernel hilbert space |
0.1 | 1 | 2008 | A reproducing kernel Hilbert space framework for pairwise time series distances · ICML 2008 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture model |
0.1 | 2 | 2004 | Semi-supervised Learning with Penalized Probabilistic Clustering · NIPS 2004 Classifying with Gaussian Mixtures and Clusters · NIPS 1994 |
Data mining
clustering |
0.0 | 1 | 2004 | Semi-supervised Learning with Penalized Probabilistic Clustering · NIPS 2004 |
Data mining › clustering
semi-supervised clustering |
0.0 | 1 | 2004 | Semi-supervised Learning with Penalized Probabilistic Clustering · NIPS 2004 |
Data mining
anomaly detection |
0.0 | 1 | 2003 | Parameterized Novelty Detectors for Environmental Sensor Monitoring · NIPS 2003 |
Machine learning › Optimization for machine learning
stochastic optimization |
0.0 | 2 | 1997 | Two Approaches to Optimal Annealing · NIPS 1997 Optimal Stochastic Search and Adaptive Momentum · NIPS 1993 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
mixture model |
0.0 | 1 | 2000 | From Mixtures of Mixtures to Adaptive Transform Coding · NIPS 2000 |
Image and video coding
transform coding |
0.0 | 1 | 2000 | From Mixtures of Mixtures to Adaptive Transform Coding · NIPS 2000 |
Machine learning › Optimization for machine learning
stochastic search |
0.0 | 2 | 1996 | Using Curvature Information for Fast Stochastic Search · NIPS 1996 Optimal Stochastic Search and Adaptive Momentum · NIPS 1993 |
Machine learning › Time series and sequential data › time series analysis
time series classification |
0.0 | 1 | 2008 | A reproducing kernel Hilbert space framework for pairwise time series distances · ICML 2008 |
Image and video processing › image fusion
multisensor image fusion |
0.0 | 1 | 1998 | Probabilistic Image Sensor Fusion · NIPS 1998 |
Machine learning › Optimization for machine learning › stochastic search
annealing |
0.0 | 1 | 1997 | Two Approaches to Optimal Annealing · NIPS 1997 |
Machine learning › Optimization for machine learning › black-box optimization › zeroth-order optimization
simulated annealing |
0.0 | 1 | 1997 | Two Approaches to Optimal Annealing · NIPS 1997 |
Machine learning › Learning theory
stochastic learning |
0.0 | 2 | 1992 | Weight Space Probability Densities in Stochastic Learning: II. Transients and Basin Hopping Times · NIPS 1992 Weight Space Probability Densities in Stochastic Learning: I. Dynamics and Equilibria · NIPS 1992 |
Internet of things and sensor networks › wireless sensor network › network diagnosis
sensor fault detection |
0.0 | 1 | 2003 | Parameterized Novelty Detectors for Environmental Sensor Monitoring · NIPS 2003 |
Machine learning › Learning theory
classification |
0.0 | 1 | 1994 | Classifying with Gaussian Mixtures and Clusters · NIPS 1994 |
Machine learning › Probabilistic and Bayesian machine learning
clustering |
0.0 | 1 | 1994 | Classifying with Gaussian Mixtures and Clusters · NIPS 1994 |
Machine learning › Deep learning architectures and training
data augmentation |
0.0 | 1 | 1994 | From Data Distributions to Regularization in Invariant Learning · NIPS 1994 |
Machine learning › Trustworthy machine learning › out-of-distribution generalization
invariant learning |
0.0 | 1 | 1994 | From Data Distributions to Regularization in Invariant Learning · NIPS 1994 |
Machine learning › Generative modeling
winner-takes-all |
0.0 | 1 | 1994 | Classifying with Gaussian Mixtures and Clusters · NIPS 1994 |
Machine learning › Optimization for machine learning › adaptive optimization
adaptive momentum methods |
0.0 | 1 | 1993 | Optimal Stochastic Search and Adaptive Momentum · NIPS 1993 |
Methods — techniques the papers use, named apart from their topics
fisher information ratio · 0.3fisher information · 0.3asymptotic analysis · 0.3mixed-effect model · 0.2fisher information matrix · 0.2parameterized bio-fouling model · 0.1novelty detection · 0.1classifier · 0.1penalized probabilistic clustering · 0.1EM algorithm · 0.1non-parametric model · 0.1generalized lloyd algorithm · 0.1bit allocation · 0.1mixture of constrained gaussian mixtures · 0.0probabilistic modeling · 0.0manifold learning · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | A Probabilistic Active Learning Algorithm Based on Fisher Information RatioabstractThe task of labeling samples is demanding and expensive. Active learning aims to generate the smallest possible training data set that results in a classifier with high performance in the test phase. It usually consists of two steps of selecting a set of queries and requesting their labels. Among the suggested objectives to score the query sets, information theoretic measures have become very popular. Yet among them, those based on Fisher information (FI) have the advantage of considering the diversity among the queries and tractable computations. In this work, we provide a practical algorithm based on Fisher information ratio to obtain query distribution for a general framework where, in contrast to the previous FI-based querying methods, we make no assumptions over the test distribution. The empirical results on synthetic and real-world data sets indicate that this algorithm gives competitive results. Jamshid Sourati, Murat Akçakaya, Deniz Erdogmus, Todd K. Leen, Jennifer G. Dy |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Asymptotic Analysis of Objectives Based on Fisher Information in Active LearningabstractObtaining labels can be costly and time-consuming. Active learning allows a learning algorithm to intelligently query samples to be labeled for a more efficient learning. Fisher information ratio (FIR) has been used as an objective for selecting queries. However, little is known about the theory behind the use of FIR for active learning. There is a gap between the underlying theory and the motivation of its usage in practice. In this paper, we attempt to fill this gap and provide a rigorous framework for analyzing existing FIR-based active learning methods. In particular, we show that FIR can be asymptotically viewed as an upper bound of the expected variance of the log-likelihood ratio. Additionally, our analysis suggests a unifying framework that not only enables us to make theoretical comparisons among the existing querying methods based on FIR, but also allows us to give insight into the development of new active learning approaches based on this objective. Jamshid Sourati, Murat Akçakaya, Todd K. Leen, Deniz Erdogmus, Jennifer G. Dy |
J. Mach. Learn. Res. | 3 |
| 2012 | Stochastic Perturbation Methods for Spike-Timing-Dependent PlasticityabstractOnline machine learning rules and many biological spike-timing-dependent plasticity (STDP) learning rules generate jump process Markov chains for the synaptic weights. We give a perturbation expansion for the dynamics that, unlike the usual approximation by a Fokker-Planck equation (FPE), is well justified. Our approach extends the related system size expansion by giving an expansion for the probability density as well as its moments. We apply the approach to two observed STDP learning rules and show that in regimes where the FPE breaks down, the new perturbation expansion agrees well with Monte Carlo simulations. The methods are also applicable to the dynamics of stochastic neural activity. Like previous ensemble analyses of STDP, we focus on equilibrium solutions, although the methods can in principle be applied to transients as well. Todd K. Leen, Robert Friel |
Neural Comput. | 1 |
| 2012 | Approximating distributions in stochastic learning
Todd K. Leen, Robert Friel, David Nielsen |
Neural Networks | 1 |
| 2011 | Perturbation theory for stochastic learning dynamicsabstractOn-line machine learning and biological spike-timing-dependent plasticity (STDP) rules both generate Markov chains for the synaptic weights. We give a perturbation expansion (in powers of the learning rate) for the dynamics that, unlike the usual approximation by a Fokker-Planck equation (FPE), is rigorous. Our approach extends the related system size expansion by giving an expansion for the probability density as well as its moments. Applied to two observed STDP learning rules, our approach provides better agreement with Monte-Carlo simulations than either the FPE or a simple linearized theory. The approach is also applicable to stochastic neural dynamics. Todd K. Leen, Robert Friel |
IJCNN | 1 |
| 2011 | Kernels for Longitudinal Data with Variable Sequence Length and Sampling IntervalsabstractWe develop several kernel methods for classification of longitudinal data and apply them to detect cognitive decline in the elderly. We first develop mixed-effects models, a type of hierarchical empirical Bayes generative models, for the time series. After demonstrating their utility in likelihood ratio classifiers (and the improvement over standard regression models for such classifiers), we develop novel Fisher kernels based on mixture of mixed-effects models and use them in support vector machine classifiers. The hierarchical generative model allows us to handle variations in sequence length and sampling interval gracefully. We also give nonparametric kernels not based on generative models, but rather on the reproducing kernel Hilbert space. We apply the methods to detecting cognitive decline from longitudinal clinical data on motor and neuropsychological tests. The likelihood ratio classifiers based on the neuropsychological tests perform better than than classifiers based on the motor behavior. Discriminant classifiers performed better than likelihood ratio classifiers for the motor behavior tests. Zhengdong Lu, Todd K. Leen, Jeffrey A. Kaye |
Neural Comput. | 2 |
| 2008 | Detecting mild cognitive loss with continuous monitoring of medication adherenceabstractThis paper describes an approach for detecting early cognitive loss using medication adherence behavior. We investigate the discriminative power of a comprehensive set of recurrent medication timing features extracted from time-of-day and inter-dose timing statistics. We adopt information theoretic measures for feature ranking for initial dimensionality reduction and conduct exhaustive leave-one-out cross validation for final feature selection and regularization. The selected feature set is subjected to a support vector machine for classification. The results demonstrate that patterns of adherence based on the data from relatively unobtrusive behavior monitoring can make reliable inference for mild cognitive loss individuals. Yonghong Huang, Deniz Erdogmus, Zhengdong Lu, Todd K. Leen |
ICASSP | 4 |
| 2008 | A reproducing kernel Hilbert space framework for pairwise time series distancesabstractA good distance measure for time series needs to properly incorporate the temporal structure, and should be applicable to sequences with unequal lengths. In this paper, we propose a distance measure as a principled solution to the two requirements. Unlike the conventional feature vector representation, our approach represents each time series with a summarizing smooth curve in a reproducing kernel Hilbert space (RKHS), and therefore translate the distance between time series into distances between curves. Moreover we propose to learn the kernel of this RKHS from a population of time series with discrete observations using Gaussian process-based non-parametric mixed-effect models. Experiments on two vastly different real-world problems show that the proposed distance measure leads to improved classification accuracy over the conventional distance measures. Zhengdong Lu, Todd K. Leen, Yonghong Huang, Deniz Erdogmus |
ICML | 2 |
| 2008 | Hierarchical Fisher Kernels for Longitudinal DataabstractWe develop new techniques for time series classification based on hierarchical Bayesian generative models (called mixed-effect models) and the Fisher kernel derived from them. A key advantage of the new formulation is that one can compute the Fisher information matrix despite varying sequence lengths and sampling times. We therefore can avoid the ad hoc replacement of Fisher information matrix with the identity matrix commonly used in literature, which destroys the geometrical grounding of the kernel construction. In contrast, our construction retains the proper geometric structure resulting in a kernel that is properly invariant under change of coordinates in the model parameter space. Experiments on detecting cognitive decline show that classifiers based on the proposed kernel out-perform those based on generative models and other feature extraction routines. Zhengdong Lu, Todd K. Leen, Jeffrey A. Kaye |
NIPS | 2 |
| 2007 | Penalized Probabilistic ClusteringabstractWhile clustering is usually an unsupervised operation, there are circumstances in which we believe (with varying degrees of certainty) that items A and B should be assigned to the same cluster, while items A and C should not. We would like such pairwise relations to influence cluster assignments of out-of-sample data in a manner consistent with the prior knowledge expressed in the training set. Our starting point is probabilistic clustering based on gaussian mixture models (GMM) of the data distribution. We express clustering preferences in a prior distribution over assignments of data points to clusters. This prior penalizes cluster assignments according to the degree with which they violate the preferences. The model parameters are fit with the expectation-maximization (EM) algorithm. Our model provides a flexible framework that encompasses several other semisupervised clustering models as its special cases. Experiments on artificial and real-world problems show that our model can consistently improve clustering results when pairwise relations are incorporated. The experiments also demonstrate the superiority of our model to other semisupervised clustering methods on handling noisy pairwise relations. Zhengdong Lu, Todd K. Leen |
Neural Comput. | 2 |
| 2007 | Fast neural network surrogates for very high dimensional physics-based models in computational oceanography
Rudolph van der Merwe, Todd K. Leen, Zhengdong Lu, Sergey Frolov, António M. Baptista |
Neural Networks | 2 |
| 2007 | Erratum to "Fast neural network surrogates for very high dimensional physics-based models in computational oceanography" [Neural Netw. 20(4) (2007) 462-478]
Rudolph van der Merwe, Todd K. Leen, Zhengdong Lu, Sergey Frolov, António M. Baptista |
Neural Networks | 2 |
| 2004 | Semi-supervised Learning with Penalized Probabilistic ClusteringabstractWhile clustering is usually an unsupervised operation, there are circum- stances in which we believe (with varying degrees of certainty) that items A and B should be assigned to the same cluster, while items A and C should not. We would like such pairwise relations to influence cluster assignments of out-of-sample data in a manner consistent with the prior knowledge expressed in the training set. Our starting point is proba- bilistic clustering based on Gaussian mixture models (GMM) of the data distribution. We express clustering preferences in the prior distribution over assignments of data points to clusters. This prior penalizes cluster assignments according to the degree with which they violate the prefer- ences. We fit the model parameters with EM. Experiments on a variety of data sets show that PPC can consistently improve clustering results. Zhengdong Lu, Todd K. Leen |
NIPS | 2 |
| 2003 | Parameterized Novelty Detectors for Environmental Sensor MonitoringabstractAs part of an environmental observation and forecasting system, sensors deployed in the Columbia RIver Estuary (CORIE) gather information on physical dynamics and changes in estuary habi- tat. Of these, salinity sensors are particularly susceptible to bio- fouling, which gradually degrades sensor response and corrupts crit- ical data. Automatic fault detectors have the capability to identify bio-fouling early and minimize data loss. Complicating the devel- opment of discriminatory classi(cid:12)ers is the scarcity of bio-fouling onset examples and the variability of the bio-fouling signature. To solve these problems, we take a novelty detection approach that incorporates a parameterized bio-fouling model. These detectors identify the occurrence of bio-fouling, and its onset time as reliably as human experts. Real-time detectors installed during the sum- mer of 2001 produced no false alarms, yet detected all episodes of sensor degradation before the (cid:12)eld sta(cid:11) scheduled these sensors for cleaning. From this initial deployment through February 2003, our bio-fouling detectors have essentially doubled the amount of useful data coming from the CORIE sensors. Cynthia Archer, Todd K. Leen, António M. Baptista |
NIPS | 2 |
| 2001 | The Coding-Optimal TransformabstractWe propose a new transform coding algorithm that integrates all optimization steps into a coherent and consistent framework. Each iteration of the algorithm is designed to minimize coding distortion as a function of both the transform and quantizer designs. Our algorithm is a constrained version of the Linde-Buzo-Gray (LBG) algorithm for vector quantizer design. The reproduction vectors are constrained to lie at the vertices of a rectangular grid. A significant result of our approach is a new transform basis specifically designed to minimize mean-squared quantization distortion for both fixed-rate and entropy-constrained coding. For Gaussian distributed data, this transform reduces to the Karhunen-Loeve transform (KLT). However, in general the coding-optimal transform (COT) differs from the KLT enough to provide up to 1 dB improvement in compressed signal-to-noise ratio (SNR) on images. We describe a practical algorithm that finds the COT for a given signal. In addition, we present image compression results demonstrating the SNR improvement achieved with our algorithm relative to KLT based transform coding. Cynthia Archer, Todd K. Leen |
Data Compression Conference | 2 |
| 2000 | From Mixtures of Mixtures to Adaptive Transform CodingabstractWe establish a principled framework for adaptive transform cod(cid:173) ing. Transform coders are often constructed by concatenating an ad hoc choice of transform with suboptimal bit allocation and quan(cid:173) tizer design. Instead, we start from a probabilistic latent variable model in the form of a mixture of constrained Gaussian mixtures. From this model we derive a transform coding algorithm, which is a constrained version of the generalized Lloyd algorithm for vector quantizer design. A byproduct of our derivation is the introduc(cid:173) tion of a new transform basis, which unlike other transforms (PCA, DCT, etc.) is explicitly optimized for coding. Image compression experiments show adaptive transform coders designed with our al(cid:173) gorithm improve compressed image signal-to-noise ratio up to 3 dB compared to global transform coding and 0.5 to 2 dB compared to other adaptive transform coders. Cynthia Archer, Todd K. Leen |
NIPS | 2 |
| 1999 | Optimal dimension reduction and transform coding with mixture principal componentsabstractThe paper addresses the problem of resource allocation in local linear models for nonlinear principal component analysis (PCA). In the local PCA model, the data space is partitioned into regions and PCA is performed in each region. Our primary result indicates that the advantage of these models over conventional PCA has been significantly underestimated in previous work. We apply local PCA models to the problems of image dimension reduction and transform coding. Our results show that by allocating representation or coding resources to the different image regions, instead of using a fixed arbitrary dimension everywhere, substantial increases in dimension reduced or compressed image qualify can be achieved. Cynthia Archer, Todd K. Leen |
IJCNN | 2 |
| 1999 | A Fast Histogram-Based Postprocessor That Improves Posterior Probability EstimatesabstractAlthough the outputs of neural network classifiers are often considered to be estimates of posterior class probabilities, the literature that assesses the calibration accuracy of these estimates illustrates that practical networks often fall far short of being ideal estimators. The theorems used to justify treating network outputs as good posterior estimates are based on several assumptions: that the network is sufficiently complex to model the posterior distribution accurately, that there are sufficient training data to specify the network, and that the optimization routine is capable of finding the global minimum of the cost function. Any or all of these assumptions may be violated in practice. This article does three things. First, we apply a simple, previously used histogram technique to assess graphically the accuracy of posterior estimates with respect to individual classes. Second, we introduce a simple and fast remapping procedure that transforms network outputs to provide better estimates of posteriors. Third, we use the remapping in a real-world telephone speech recognition system. The remapping results in a 10% reduction of both word-level error rates (from 4.53% to 4.06%) and sentence-level error rates (from 16.38% to 14.69%) on one corpus, and a 29% reduction at sentence-level error (from 6.3% to 4.5%) on another. The remapping required negligible additional overhead (in terms of both parameters and calculations). McNemar's test shows that these levels of improvement are statistically significant. Todd K. Leen, Etienne Barnard |
Neural Comput. | 2 |
| 1998 | Probabilistic Image Sensor Fusion
Ravi K. Sharma, Todd K. Leen, Misha Pavel |
NIPS | 2 |
| 1997 | Two Approaches to Optimal Annealing
Todd K. Leen, Bernhard Schottky, David Saad |
NIPS | 1 |
| 1997 | Dimension Reduction by Local Principal Component AnalysisabstractReducing or eliminating statistical redundancy between the components of high-dimensional vector data enables a lower-dimensional representation without significant loss of information. Recognizing the limitations of principal component analysis (PCA), researchers in the statistics and neural network communities have developed nonlinear extensions of PCA. This article develops a local linear approach to dimension reduction that provides accurate representations and is fast to compute. We exercise the algorithms on speech and image data, and compare performance with PCA and with neural network implementations of nonlinear PCA. We find that both nonlinear techniques can provide more accurate representations than PCA and show that the local linear techniques outperform neural network implementations. Nanda Kambhatla, Todd K. Leen |
Neural Comput. | 2 |
| 1996 | Using Curvature Information for Fast Stochastic Search
Genevieve B. Orr, Todd K. Leen |
NIPS | 2 |
| 1995 | From Data Distributions to Regularization in Invariant LearningabstractIdeally pattern recognition machines provide constant output when the inputs are transformed under a group G of desired invariances. These invariances can be achieved by enhancing the training data to include examples of inputs transformed by elements of G, while leaving the corresponding targets unchanged. Alternatively the cost function for training can include a regularization term that penalizes changes in the output when the input is transformed under the group. This paper relates the two approaches, showing precisely the sense in which the regularized cost function approximates the result of adding transformed examples to the training data. We introduce the notion of a probability distribution over the group transformations, and use this to rewrite the cost function for the enhanced training data. Under certain conditions, the new cost function is equivalent to the sum of the original cost function plus a regularizer. For unbiased models, the regularizer reduces to the intuitively obvious choice—a term that penalizes changes in the output when the inputs are transformed under the group. For infinitesimal transformations, the coefficient of the regularization term reduces to the variance of the distortions introduced into the training data. This correspondence provides a simple bridge between the two approaches. Todd K. Leen |
Neural Comput. | 1 |
| 1994 | Classifying with Gaussian Mixtures and ClustersabstractIn this paper, we derive classifiers which are winner-take-all (WTA) approximations to a Bayes classifier with Gaussian mixtures for class conditional densities. The derived classifiers include clustering based algorithms like LVQ and k-Means. We propose a constrained rank Gaussian mixtures model and derive a WTA algorithm for it. Our experiments with two speech classification tasks indicate that the constrained rank model and the WTA approximations improve the performance over the unconstrained models. Nanda Kambhatla, Todd K. Leen |
NIPS | 2 |
| 1994 | From Data Distributions to Regularization in Invariant LearningabstractIdeally pattern recognition machines provide constant output when the inputs are transformed under a group 9 of desired invariances. These invariances can be achieved by enhancing the training data to include examples of inputs transformed by elements of g, while leaving the corresponding targets unchanged. Alternatively the cost function for training can include a regularization term that penalizes changes in the output when the input is transformed un(cid:173) der the group. This paper relates the two approaches, showing precisely the sense in which the regularized cost function approximates the result of adding transformed (or distorted) examples to the training data. The cost function for the enhanced training set is equivalent to the sum of the original cost function plus a regularizer. For unbiased models, the regularizer reduces to the intuitively obvious choice - a term that penalizes changes in the output when the inputs are transformed under the group. For infinitesimal transformations, the coefficient of the regularization term reduces to the variance of the distortions introduced into the training data. This correspon(cid:173) dence provides a simple bridge between the two approaches. Todd K. Leen |
NIPS | 1 |
| 1993 | Fast Non-Linear Dimension Reduction
Nanda Kambhatla, Todd K. Leen |
NIPS | 2 |
| 1993 | Optimal Stochastic Search and Adaptive Momentum
Todd K. Leen, Genevieve B. Orr |
NIPS | 1 |
| 1993 | Fast Pruning Using Principal Components
Asriel U. Levin, Todd K. Leen, John E. Moody |
NIPS | 2 |
| 1992 | Weight Space Probability Densities in Stochastic Learning: I. Dynamics and Equilibria
Todd K. Leen, John E. Moody |
NIPS | 1 |
| 1992 | Weight Space Probability Densities in Stochastic Learning: II. Transients and Basin Hopping Times
Genevieve B. Orr, Todd K. Leen |
NIPS | 2 |
| 1990 | Hebbian feature discovery improves classifier efficiencyabstractTwo neural network implementations of principal component analysis (PCA) are used to reduce the dimension of speech signals. The compressed signals are then used to train a feedforward classification network for vowel recognition. A comparison is made of classification performance, network size, and training time for networks trained with both compressed and uncompressed data. Results show that a significant reduction in training time, fivefold in the present case, can be achieved without a sacrifice in classifier accuracy. This reduction includes the time required to train the compression network. Thus, dimension reduction, as performed by unsupervised neural networks, is a viable tool for enhancing the efficiency of neural classifiers Todd K. Leen, Mike Rudnick, Dan W. Hammerstrom |
IJCNN | 1 |
| 1990 | Dynamics of Learning in Recurrent Feature-Discovery Networks
Todd K. Leen |
NIPS | 1 |