EDBT 2026 Demo / reviewers in the wild / expert
José C. Príncipe
dblp:p/JoseCPrincipe · also José Carlos Príncipe
· DBLP profile ↗
428ranked-venue papers
23as first author
40since 2021 · last 2026
0000-0002-3449-3531ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 261 · 12 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 142 · 6 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 5 first-author · 1 since 2021Systems, architecture and hardware · 12Human-computer interaction and ubiquitous computing · 6 · 2 since 2021Computer networks · 5Databases, data management, data science and information retrieval · 4Software engineering, systems software and programming languages · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training sparse convolutional deep predictive coding networks with attention
Chi Ding, José C. Príncipe |
Neural Networks | 3 |
| 2026 | A closed-form solution for kernel adaptive filtering
Benjamin Colburn, Luis Gonzalo Sánchez Giraldo, Kan Li 0002, José C. Príncipe |
Signal Process. | 4 |
| 2026 | Online filtering in kernel adaptive principal subspace
Kan Li 0002, José C. Príncipe |
Signal Process. | 2 |
| 2025 | Trimformer: A Novel Sequence Compression Mechanism with Local AttentionabstractTraining on extremely long sequences poses significant challenges for attention mechanisms. In this paper, we introduce a novel trim attention mechanism that capitalizes on the inherent sparsity within attention processes. This mechanism effectively compresses the sequence length, thereby reducing the overall computational complexity without altering the number of trainable parameters. Our experimental results demonstrate that this approach not only decreases computational demands but also outperforms the Vision Transformer in image classification tasks. The trim attention mechanism can seamlessly replace any standard attention layer. Ran Dou, Liyang Ru, José C. Príncipe |
ICASSP | 3 |
| 2025 | Schoenberg Kernel Loss for Spiking Neural Network TrainingabstractIn this paper, we propose the Schoenberg kernel loss, a loss function that measures the divergence between spike trains for spiking neural network training. Our method addresses the challenges of non-differentiable spiking neurons and leverages both spatial and temporal dynamics in information processing. We evaluated our approach on the neuromorphic benchmark datasets, N-MNIST, and DVS Gesture, demonstrating our method outperforms previously published loss functions on our network architecture. In addition, we extend our investigation to real-time classification simulations in real-world scenarios. This work contributes to the advancement of direct SNN training methods, offering improved performance, particularly on low spike density data, and potential applications in neuromorphic classification. Liyang Ru, Kan Li 0002, José C. Príncipe |
ICASSP | 3 |
| 2025 | Fast DPCNs for Feature Extraction without LabelsabstractDeep predictive coding networks (DPCNs) effectively model and capture video features through a bi-directional inference without labels. They are based on an overcomplete description of video scenes, and one of the bottlenecks has been the lack of effective sparsification techniques to find discriminative and robust dictionaries. This paper proposes a DPCN with a fast inference of internal dictionaries and variables that achieve high sparsity and feature clustering accuracy. The proposed unsupervised learning procedure uses majorization-minimization (MM) to smooth sparsity constraints in optimization and admits explainability and convergence. Experiments in the image and video data sets CIFAR-10, Super Mario Bros, and Coil-100 validate that the approach outperforms previous versions of DPCNs on learning rate, sparsity ratio, and feature clustering accuracy. This advance opens the door for general applications in object recognition in video without labels. Wenqian Xue, Chi Ding, José C. Príncipe |
ICASSP | 3 |
| 2025 | A Simple and Effective Method for Uncertainty Quantification and OOD DetectionabstractBayesian neural networks and deep ensemble methods have been proposed for uncertainty quantification; however, they are computationally intensive and require large storage. By utilizing a single deterministic model, we can solve the above issue. We propose an effective method based on feature space density to quantify uncertainty for distributional shifts and out-of-distribution (OOD) detection. Specifically, we leverage the information potential field derived from kernel density estimation to approximate the feature space density of the training set. By comparing this density with the feature space representation of test samples, we can effectively determine whether a distributional shift has occurred. Experiments were conducted on a 2D synthetic dataset (Two Moons and Three Spirals) as well as an OOD detection task (CIFAR-10 vs. SVHN). The results demonstrate that our method outperforms baseline models. Yaxin Ma 0001, Benjamin Colburn, José C. Príncipe |
IJCNN | 3 |
| 2025 | The Conditional Cauchy-Schwarz Divergence With Applications to Time-Series Data and Sequential Decision MakingabstractThe Cauchy-Schwarz (CS) divergence was developed by Príncipe et al. in 2000. In this paper, we extend the classic CS divergence to quantify the closeness between two conditional distributions and show that the developed conditional CS divergence can be elegantly estimated by a kernel density estimator from given samples. We illustrate the advantages (e.g., rigorous faithfulness guarantee, lower computational complexity, higher statistical power, and much more flexibility in a wide range of applications) of our conditional CS divergence over previous proposals, such as the conditional Kullback-Leibler divergence and the conditional maximum mean discrepancy. We also demonstrate the compelling performance of conditional CS divergence in two machine learning tasks related to time series data and sequential inference, namely time series clustering and uncertainty-guided exploration for sequential decision making. Shujian Yu, Sigurd Løkse, Robert Jenssen, José C. Príncipe |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Improved PCA Reconstruction-Based Unsupervised Anomaly Detection in Uncontrolled Structural Health Monitoring With CorrentropyabstractGuided wave-based structural health monitoring is extensively utilized in various industrial applications to ensure the integrity of components within industrial systems. Among these monitoring techniques, principal component analysis (PCA) reconstruction methods are widely used for anomaly detection due to their computational efficiency and interoperability. However, existing PCA reconstruction methods are semisupervised anomaly detection approaches that require training on historical normal data and fail to detect anomalous signals within the training set. To address this limitation, this work proposes a correntropy PCA (C-PCA), enabling fully unsupervised anomaly detection on raw training data without requiring label information, when the dataset contains a high proportion of abnormal signals. This method allows anomaly detection on real-time measurements without the need for precleaned historical normal data or can also be used to generate clean data for existing semisupervised anomaly detection methods. In correntropy PCA, principal components are extracted from the correntropy matrix rather than the correlation matrix. The correntropy, representing the statistical dependence between samples of guided waves, is estimated utilizing a Gaussian kernel with a specified kernel width. Through the optimization of the kernel width, the correntropy PCA reconstruction method demonstrates superior anomaly detection performance compared with the standard PCA reconstruction method, especially in scenarios where training data are contaminated by a significant proportion of abnormal signals. Guidelines for the optimization of the kernel width are provided. The effectiveness of the correntropy PCA reconstruction-based anomaly detection method is validated using data collected from ten regions over an 80-day period, encompassing guided waves induced by damage occurring over durations ranging from 2 to 20 days. Zhenhan Lin, Zhihui Tian, José C. Príncipe, Joel B. Harley |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | Guest Editorial: Special Issue on Information Theoretic Methods for the Generalization, Robustness, and Interpretability of Machine Learning
Badong Chen, Shujian Yu, Robert Jenssen, José C. Príncipe, Klaus-Robert Müller |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Learning Orthonormal Features in Self-Supervised Learning using Functional Maximal CorrelationabstractThis paper applies statistical dependence measures to interpret self-supervised learning (SSL). Conventional applications of measures like mutual information commonly use separate procedures for feature extraction and dependence estimation, where the relationship between optimal features and the strength of dependence is unclear. This causes limitations in tasks requiring multivariate feature representations, particularly in SSL. The recently introduced multivariate measure, functional maximal correlation, is a unified framework based on orthonormal decomposition of density ratios, wherein the spectrum and the bases become the measure and the features, respectively. This paper proposes that features in SSL can also be interpreted as basis functions of the density ratio. We introduce the Hierarchical Functional Maximal Correlation Algorithm (HFMCA), a theoretically justified approach that ensures faster convergence, enhanced stability, and prevents feature collapse by learning orthonormal bases as multivariate features. Yuheng Bu, José C. Príncipe |
ICIP | 3 |
| 2024 | Cauchy-Schwarz Divergence Information Bottleneck for RegressionabstractThe information bottleneck (IB) approach is popular to improve the generalization, robustness and explainability of deep neural networks. Essentially, it aims to find a minimum sufficient representation $\mathbf{t}$ by striking a trade-off between a compression term $I(\mathbf{x};\mathbf{t})$ and a prediction term $I(y;\mathbf{t})$, where $I(\cdot;\cdot)$ refers to the mutual information (MI). MI is for the IB for the most part expressed in terms of the Kullback-Leibler (KL) divergence, which in the regression case corresponds to prediction based on mean squared error (MSE) loss with Gaussian assumption and compression approximated by variational inference.
In this paper, we study the IB principle for the regression problem and develop a new way to parameterize the IB with deep neural networks by exploiting favorable properties of the Cauchy-Schwarz (CS) divergence. By doing so, we move away from MSE-based regression and ease estimation by avoiding variational approximations or distributional assumptions. We investigate the improved generalization ability of our proposed CS-IB and demonstrate strong adversarial robustness guarantees. We demonstrate its superior performance on six real-world regression tasks over other popular deep IB approaches. We additionally observe that the solutions discovered by CS-IB always achieve the best trade-off between prediction accuracy and compression ratio in the information plane. The code is available at \url{https://github.com/SJYuCNEL/Cauchy-Schwarz-Information-Bottleneck}. Shujian Yu, Sigurd Løkse, Robert Jenssen, José C. Príncipe |
ICLR | 5 |
| 2024 | End-to-end Image Classification in Linear Hybrid Cellular AutomataabstractCellular automata have ideal properties for scalable and efficient computing, but a lack of a training method limits their real-world applications. First, we propose to partition the cellular lattice into three regions: input, output, and processing. Second, we propose a novel synthesis method to train a linear hybrid cellular automaton. Third, we show image classification on the MNIST dataset using only logic operations. By mapping local states over the globally linear lattice, the proposed model achieved above 90% test accuracy in binary image classification. Our method does not require any pre or post-processors to perform computation over the lattice. Hence, the lattice maintains its massive parallelism and locality of computation, ideal for ultra-low power processing in machine learning. Naoki Sawahashi, José C. Príncipe |
IJCNN | 2 |
| 2024 | Finding Local Dependent Regions in PDFs using RKHS Uncertainty Moments and Optimal TransportabstractReliable measurement of dependence between random variables is essential in many applications of statistics and machine learning. Current approaches for dependence estimation employ the full probability density function (PDF) and are unable to quantify local dependence, which is required to improve precision, robustness and/or interpretability. We propose a two-step approach for local dependence quantification between random variables: 1) First decompose the PDF of the variables involved in orthogonal regional moments; 2) Compute an optimal transport map to measure the similarity, in the space of the data, between the corresponding sets of local low density moments, which correspond to uncertainty. Statistical dependence is then determined by the degree of one-to-one correspondence between the respective uncertainty moments decomposition. The proposed approach is robust towards outliers and monotone transformations of data, while the multiple moments of uncertainty provide high resolution and interpretability of the type of dependence being quantified. We support these claims through preliminary results using simulated data. Rishabh Singh, Yaxin Ma 0001, José C. Príncipe |
IJCNN | 3 |
| 2024 | Learning Cortico-Muscular Dependence through Orthonormal Decomposition of Density RatiosabstractThe cortico-spinal neural pathway is fundamental for motor control and movement execution, and in humans it is typically studied using concurrent electroencephalography (EEG) and electromyography (EMG) recordings. However, current approaches for capturing high-level and contextual connectivity between these recordings have important limitations. Here, we present a novel application of statistical dependence estimators based on orthonormal decomposition of density ratios to model the relationship between cortical and muscle oscillations. Our method extends from traditional scalar-valued measures by learning eigenvalues, eigenfunctions, and projection spaces of density ratios from realizations of the signal, addressing the interpretability, scalability, and local temporal dependence of cortico-muscular connectivity. We experimentally demonstrate that eigenfunctions learned from cortico-muscular connectivity can accurately classify movements and subjects. Moreover, they reveal channel and temporal dependencies that confirm the activation of specific EEG channels during movement. Shihan Ma, Alex Clarke 0001, Blanka Zicher, Arnault H. Caillet, Dario Farina, José C. Príncipe |
NeurIPS | 8 |
| 2024 | IA-LSTM: Interaction-Aware LSTM for Pedestrian Trajectory PredictionabstractPredicting the trajectory of pedestrians in crowd scenarios is indispensable in self-driving or autonomous mobile robot field because estimating the future locations of pedestrians around is beneficial for policy decision to avoid collision. It is a challenging issue because humans have different walking motions, and the interactions between humans and objects in the current environment, especially between humans themselves, are complex. Previous researchers focused on how to model human-human interactions but neglected the relative importance of interactions. To address this issue, a novel mechanism based on correntropy is introduced. The proposed mechanism not only can measure the relative importance of human-human interactions but also can build personal space for each pedestrian. An interaction module, including this data-driven mechanism, is further proposed. In the proposed module, the data-driven mechanism can effectively extract the feature representations of dynamic human-human interactions in the scene and calculate the corresponding weights to represent the importance of different interactions. To share such social messages among pedestrians, an interaction-aware architecture based on long short-term memory network for trajectory prediction is designed. Experiments are conducted on two public datasets. Experimental results demonstrate that our model can achieve better performance than several latest methods with good performance. Jing Yang 0014, Yuehai Chen, Shaoyi Du, Badong Chen, José C. Príncipe |
IEEE Trans. Cybern. | 5 |
| 2023 | Causal Recurrent Variational Autoencoder for Medical Time Series GenerationabstractWe propose causal recurrent variational autoencoder (CR-VAE), a novel generative model that is able to learn a Granger causal graph from a multivariate time series x and incorporates the underlying causal mechanism into its data generation process. Distinct to the classical recurrent VAEs, our CR-VAE uses a multi-head decoder, in which the p-th head is responsible for generating the p-th dimension of x (i.e., x^p). By imposing a sparsity-inducing penalty on the weights (of the decoder) and encouraging specific sets of weights to be zero, our CR-VAE learns a sparse adjacency matrix that encodes causal relations between all pairs of variables. Thanks to this causal matrix, our decoder strictly obeys the underlying principles of Granger causality, thereby making the data generating process transparent. We develop a two-stage approach to train the overall objective. Empirically, we evaluate the behavior of our model in synthetic data and two real-world human brain datasets involving, respectively, the electroencephalography (EEG) signals and the functional magnetic resonance imaging (fMRI) data. Our model consistently outperforms state-of-the-art time series generative models both qualitatively and quantitatively. Moreover, it also discovers a faithful causal graph with similar or improved accuracy over existing Granger causality-based causal inference methods. Code of CR-VAE is publicly available at https://github.com/hongmingli1995/CR-VAE. Shujian Yu, José C. Príncipe |
AAAI | 3 |
| 2023 | Universal Recurrent Event Memories for Streaming DataabstractIn this paper, we propose a new event memory architecture (MemNet) for recurrent neural networks, which is universal for different types of time series data such as scalar, multivariate or symbolic. Unlike other external neural memory architectures, it stores key-value pairs, which separate the information for addressing and for content to improve the representation, as in the digital archetype. Moreover, the key-value pairs also avoid the compromise between memory depth and resolution that applies to memories constructed by the model state. One of the MemNet key characteristics is that it requires only linear adaptive mapping functions while implementing a nonlinear operation on the input data. MemNet architecture can be applied without modifications to scalar time series, logic operators on strings, and also to natural language processing, providing state-of-the-art results in all application domains such as the chaotic time series, the symbolic operation tasks, and the question-answering tasks (bAbI). Finally, controlled by five linear layers, MemNet requires a much smaller number of training parameters than other external memory networks as well as the transformer network. The space complexity of MemNet equals a single self-attention layer. It greatly improves the efficiency of the attention mechanism and opens the door for IoT applications. Ran Dou, José C. Príncipe |
IJCNN | 2 |
| 2023 | Dynamic Analysis and an Eigen Initializer for Recurrent Neural NetworksabstractIn recurrent neural networks, learning long-term dependency is the main difficulty due to the vanishing and exploding gradient problem. Many researchers are dedicated to solving this issue and they proposed many algorithms. Although these algorithms have achieved great success, understanding how the information decays remains an open problem. In this paper, we study the dynamics of the hidden state in recurrent neural networks. We propose a new perspective to analyze the hidden state space based on an eigen decomposition of the weight matrix. We start the analysis by linear state space model and explain the function of preserving information in activation functions. We provide an explanation for long-term dependency based on the eigen analysis. We also point out the different behavior of eigenvalues for regression tasks and classification tasks. From the observations on well-trained recurrent neural networks, we proposed a new initialization method for recurrent neural networks, which improves consistently the performance. It can be applied to vanilla-RNN, LSTM, and GRU. We test is on many datasets, such as Tomita Grammars, pixel-by-pixel MNIST dataset, and machine translation dataset (Multi30k). It outperforms Xavier initializer and kaiming initializer as well as other RNN-only initializers like IRNN and sp-RNN in several tasks. Ran Dou, José C. Príncipe |
IJCNN | 2 |
| 2023 | Labels, Information, and Computation: Efficient Learning Using Sufficient LabelsabstractIn supervised learning, obtaining a large set of fully-labeled training data is expensive. We show that we do not always need full label information on every single training example to train a competent classifier. Specifically, inspired by the principle of sufficiency in statistics, we present a statistic (a summary) of the fully-labeled training set that captures almost all the relevant information for classification but at the same time is easier to obtain directly. We call this statistic "sufficiently-labeled data" and prove its sufficiency and efficiency for finding the optimal hidden representations, on which competent classifier heads can be trained using as few as a single randomly-chosen fully-labeled example per class. Sufficiently-labeled data can be obtained from annotators directly without collecting the fully-labeled data first. And we prove that it is easier to directly obtain sufficiently-labeled data than obtaining fully-labeled data. Furthermore, sufficiently-labeled data is naturally more secure since it stores relative, instead of absolute, information. Extensive experimental results are provided to support our theory. Shiyu Duan, Spencer Chang, José C. Príncipe |
J. Mach. Learn. Res. | 3 |
| 2023 | Multiscale principle of relevant information for hyperspectral image classification
Yantao Wei, Shujian Yu, Luis Gonzalo Sánchez Giraldo, José C. Príncipe |
Mach. Learn. | 4 |
| 2023 | Faster Convergence in Deep-Predictive-Coding Networks to Learn Deeper RepresentationsabstractDeep-predictive-coding networks (DPCNs) are hierarchical, generative models. They rely on feed-forward and feedback connections to modulate latent feature representations of stimuli in a dynamic and context-sensitive manner. A crucial element of DPCNs is a forward-backward inference procedure to uncover sparse, invariant features. However, this inference is a major computational bottleneck. It severely limits the network depth due to learning stagnation. Here, we prove why this bottleneck occurs. We then propose a new forward-inference strategy based on accelerated proximal gradients. This strategy has faster theoretical convergence guarantees than the one used for DPCNs. It overcomes learning stagnation. We also demonstrate that it permits constructing deep and wide predictive-coding networks. Such convolutional networks implement receptive fields that capture well the entire classes of objects on which the networks are trained. This improves the feature representations compared with our lab's previous nonconvolutional and convolutional DPCNs. It yields unsupervised object recognition that surpass convolutional autoencoders and is on par with convolutional networks trained in a supervised manner. Isaac J. Sledge, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | Deep Deterministic Independent Component Analysis for Hyperspectral UnmixingabstractWe develop a new neural network based independent component analysis (ICA) method by directly minimizing the dependence amongst all extracted components. Using the matrix-based Rényi’s α-order entropy functional, our network can be directly optimized by stochastic gradient descent (SGD), without any variational approximation or adversarial training. As a solid application, we evaluate our ICA in the problem of hyperspectral unmixing (HU) and refute a statement that "ICA does not play a role in unmixing hyperspectral data", which was initially suggested by [1]. Code and additional remarks of our DDICA is available at https://github.com/hongmingli1995/DDICA. Shujian Yu, José C. Príncipe |
ICASSP | 3 |
| 2022 | Explaining Deep and ResNet Architecture Choices with Information FlowabstractRecently, Information Theoretic Learning (ITL) has helped explain the learning dynamics for deep learning models such as multilayer perceptrons (MLP), convolutional neural networks (CNN), and stacked autoencoders (SAE). It is understood that for MLPs and CNNs, where the desired signal and input are independent of each other, the set of consecutive layers in the primary and adjoint networks represent individual Markov Chains (MC) that effect two data processing inequalities (DPI) in their respective directions. For the SAE, the desired signal is the input, so the DPI is only confirmed until the bottleneck layer. In this paper, we propose using the adjoint network to compute conditional mutual information with the backpropagated errors to demonstrate the DPI until the SAE's last layer. Also, we present an ITL-based analysis of the residual network (ResNet) architecture and propose an explanation for why the identity mapping is the optimal shortcut connection. Spencer Chang, José C. Príncipe |
IJCNN | 2 |
| 2022 | The Extended Kernel Adaptive Autoregressive-Moving-Average AlgorithmabstractIn this paper, we proposed the Extended KAARMA algorithm, which substitutes gradient descent (SGD) by the Extended Kalman (EKF) update equations. Comparing with the stochastic gradient descent method, the EKF method provides a higher rate of convergence. By creating more centers in the memory, it can explore the error space in an efficient manner. Besides, the Extended Kalman method can adjust the learning rate automatically. The decreasing learning rate solves the gradient explosion problem, where a gradient clipping technique is needed in the SGD method instead. Ran Dou, José C. Príncipe |
IJCNN | 2 |
| 2022 | Kernel Nonlinear Dynamic System Identification Based on Expectation-Maximization MethodabstractThis paper develops a kernel nonlinear system identification approach to dual estimation problems. Given the observation model and the corresponding observation sequence, the unknown state transition function is approximated in the reproducing kernel Hilbert space (RKHS), while the corresponding hidden states are also estimated with the unscented Kalman smoother (UKS). To implement the dual estimation, the optimization solution of the proposed kernel expectation-maximization (EM) cost function is approximated based on the unscented transform (UT). Finally, the kernel nonlinear system identification method is applied to the noisy time-series estimation, where the data are generated by the IKEDA chaotic dynamical system with unknown system parameters. The simulation results show that the noisy data's signal-noise ratio (SNR) is improved significantly and is close to the SNR obtained with the accurate system model. Furthermore, the proposed non-parametric approach is compared with the parametric approach and outperforms obviously the parametric approach, even if the parameterized system model is assumed to be known. Pingping Zhu, José C. Príncipe |
IJCNN | 2 |
| 2022 | Principle of relevant information for graph sparsificationabstractGraph sparsification aims to reduce the number of edges of a graph while maintaining its structural properties. In this paper, we propose the first general and effective information-theoretic formulation of graph sparsification, by taking inspiration from the Principle of Relevant Information (PRI). To this end, we extend the PRI from a standard scalar random variable setting to structured data (i.e., graphs). Our Graph-PRI objective is achieved by operating on the graph Laplacian, made possible by expressing the graph Laplacian of a subgraph in terms of a sparse edge selection vector w. We provide both theoretical and empirical justifications on the validity of our Graph-PRI approach. We also analyze its analytical solutions in a few special cases. We finally present three representative real-world applications, namely graph sparsification, graph regularized multi-task learning, and medical imaging-derived brain network classification, to demonstrate the effectiveness, the versatility and the enhanced interpretability of our approach over prevalent sparsification techniques. Code of Graph-PRI is available at https://github.com/SJYuCNEL/PRI-Graphs. Shujian Yu, Francesco Alesiani, Wenzhe Yin, Robert Jenssen, José C. Príncipe |
UAI | 5 |
| 2022 | A self-learning cognitive architecture exploiting causality from rewards
Ran Dou, Andreas Keil, José C. Príncipe |
Neural Networks | 4 |
| 2022 | Modularizing Deep Learning via Pairwise Learning With KernelsabstractBy redefining the conventional notions of layers, we present an alternative view on finitely wide, fully trainable deep neural networks as stacked linear models in feature spaces, leading to a kernel machine interpretation. Based on this construction, we then propose a provably optimal modular learning framework for classification that does not require between-module backpropagation. This modular approach brings new insights into the label requirement of deep learning (DL). It leverages only implicit pairwise labels (weak supervision) when learning the hidden modules. When training the output module, on the other hand, it requires full supervision but achieves high label efficiency, needing as few as ten randomly selected labeled examples (one from each class) to achieve 94.88% accuracy on CIFAR-10 using a ResNet-18 backbone. Moreover, modular training enables fully modularized DL workflows, which then simplify the design and implementation of pipelines and improve the maintainability and reusability of models. To showcase the advantages of such a modularized workflow, we describe a simple yet reliable method for estimating reusability of pretrained modules as well as task transferability in a transfer learning setting. At practically no computation overhead, it precisely described the task space structure of 15 binary classification tasks from CIFAR-10. Shiyu Duan, Shujian Yu, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Measuring Dependence with Matrix-based Entropy FunctionalabstractMeasuring the dependence of data plays a central role in statistics and machine learning. In this work, we summarize and generalize the main idea of existing information-theoretic dependence measures into a higher-level perspective by the Shearer's inequality. Based on our generalization, we then propose two measures, namely the matrix-based normalized total correlation and the matrix-based normalized dual total correlation, to quantify the dependence of multiple variables in arbitrary dimensional space, without explicit estimation of the underlying data distributions. We show that our measures are differentiable and statistically more powerful than prevalent ones. We also show the impact of our measures in four different machine learning problems, namely the gene regulatory network inference, the robust machine learning under covariate shift and non-Gaussian noises, the subspace outlier detection, and the understanding of the learning dynamics of convolutional neural networks, to demonstrate their utilities, advantages, as well as implications to those problems. Shujian Yu, Francesco Alesiani, Robert Jenssen, José C. Príncipe |
AAAI | 5 |
| 2021 | Training a Bank of Wiener Models with a Novel Quadratic Mutual Information Cost FunctionabstractThis paper presents a novel training methodology to adapt parameters of a bank of Wiener models (BWMs), i.e., a bank of linear filters followed by a static memoryless nonlinearity, using full pdf information of the projected outputs and the desired signal. BWMs also share the same architecture with the first layer of a time-delay neural networks (TDNN) with a single hidden layer, which is often trained with backpropagation. To optimize BWMs, we develop a novel cost function called the empirical embedding of quadratic mutual information (E-QMI) that is metric-driven and efficient in characterizing the statistical dependency. We demonstrate experimentally that by applying this cost function to the proposed model, our method is comparable with state-of-the-art neural network architectures for regressions tasks without using backpropagation of the error. José C. Príncipe |
ICASSP | 2 |
| 2021 | Deep Deterministic Information Bottleneck with Matrix-Based Entropy FunctionalabstractWe introduce the matrix-based Rényi’s α-order entropy functional to parameterize Tishby et al. information bottleneck (IB) principle [1] with a neural network. We term our methodology Deep Deterministic Information Bottleneck (DIB), as it avoids variational inference and distribution assumption. We show that deep neural networks trained with DIB outperform the variational objective counterpart and those that are trained with other forms of regularization, in terms of generalization performance and robustness to adversarial attack. Code available at https://github.com/yuxi120407/DIB. Shujian Yu, José C. Príncipe |
ICASSP | 3 |
| 2021 | Information-Theoretic Methods in Deep Neural Networks: Recent Advances and Emerging OpportunitiesabstractWe present a review on the recent advances and emerging opportunities around the theme of analyzing deep neural networks (DNNs) with information-theoretic methods. We first discuss popular information-theoretic quantities and their estimators. We then introduce recent developments on information-theoretic learning principles (e.g., loss functions, regularizers and objectives) and their parameterization with DNNs. We finally briefly review current usages of information-theoretic concepts in a few modern machine learning problems and list a few emerging opportunities. Shujian Yu, Luis Gonzalo Sánchez Giraldo, José C. Príncipe |
IJCAI | 3 |
| 2021 | Speeding Up Reinforcement Learning by Exploiting Causality in Reward SequencesabstractThis paper presents a methodology to exploit causation in deep reinforcement learning (DRL). We take advantage of a cognitive architecture that automatically decomposes the game world in proto-objects and records their position in 2D space across time. Therefore, while playing the game, proto-objects' locations define internal time sequences that can be compared with the reward sequence to select the proto-object that caused the reward. We propose a novel non-parametric information theoretic learning Granger causality (ITL-GC) estimator of directed information using Reny's entropy that is accurate in high dimensions. We integrate this module in a state-of-the-art DRL architecture (A3C) and show substantial improvement in the speed of convergence compared with conventional training. José C. Príncipe |
IJCNN | 2 |
| 2021 | Associations between MSE and SSIM as cost functions in linear decomposition with application to bit allocation for sparse coding
Jianji Wang 0001, Nanning Zheng 0001, Badong Chen, José C. Príncipe, Fei-Yue Wang 0001 |
Neurocomputing | 5 |
| 2021 | Toward a Kernel-Based Uncertainty Decomposition Framework for Data and ModelsabstractThis letter introduces a new framework for quantifying predictive uncertainty for both data and models that relies on projecting the data into a gaussian reproducing kernel Hilbert space (RKHS) and transforming the data probability density function (PDF) in a way that quantifies the flow of its gradient as a topological potential field (quantified at all points in the sample space). This enables the decomposition of the PDF gradient flow by formulating it as a moment decomposition problem using operators from quantum physics, specifically Schrödinger's formulation. We experimentally show that the higher-order moments systematically cluster the different tail regions of the PDF, thereby providing unprecedented discriminative resolution of data regions having high epistemic uncertainty. In essence, this approach decomposes local realizations of the data PDF in terms of uncertainty moments. We apply this framework as a surrogate tool for predictive uncertainty quantification of point-prediction neural network models, overcoming various limitations of conventional Bayesian-based uncertainty quantification methods. Experimental comparisons with some established methods illustrate performance advantages that our framework exhibits. Rishabh Singh, José C. Príncipe |
Neural Comput. | 2 |
| 2021 | Unsupervised foveal vision neural architecture with top-down attention
Ryan Burt, Nina Thigpen, Andreas Keil, José C. Príncipe |
Neural Networks | 4 |
| 2021 | Understanding Convolutional Neural Networks With Information Theory: An Initial ExplorationabstractA novel functional estimator for Rényi's α -entropy and its multivariate extension was recently proposed in terms of the normalized eigenspectrum of a Hermitian matrix of the projected data in a reproducing kernel Hilbert space (RKHS). However, the utility and possible applications of these new estimators are rather new and mostly unknown to practitioners. In this brief, we first show that this estimator enables straightforward measurement of information flow in realistic convolutional neural networks (CNNs) without any approximation. Then, we introduce the partial information decomposition (PID) framework and develop three quantities to analyze the synergy and redundancy in convolutional layer representations. Our results validate two fundamental data processing inequalities and reveal more inner properties concerning CNN training. Shujian Yu, Kristoffer Wickstrøm, Robert Jenssen, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Minimum Error Entropy Kalman FilterabstractTo date, most linear and nonlinear Kalman filters (KFs) have been developed under the Gaussian assumption and the well-known minimum mean square error (MMSE) criterion. In order to improve the robustness with respect to impulsive (or heavy-tailed) non-Gaussian noises, the maximum correntropy criterion (MCC) has recently been used to replace the MMSE criterion in developing several robust Kalman-type filters. To deal with more complicated non-Gaussian noises such as noises from multimodal distributions, in this article, we develop a new Kalman-type filter, called minimum error entropy KF (MEE-KF), by using the minimum error entropy (MEE) criterion instead of the MMSE or MCC. Similar to the MCC-based KFs, the proposed filter is also an online algorithm with the recursive process, in which the propagation equations are used to give prior estimates of the state and covariance matrix, and a fixed-point algorithm is used to update the posterior estimates. In addition, the MEE extended KF (MEE-EKF) is also developed for performance improvement in the nonlinear situations. The high accuracy and strong robustness of MEE-KF and MEE-EKF are confirmed by experimental results. Badong Chen, Lujuan Dang, Yuantao Gu, Nanning Zheng 0001, José C. Príncipe |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2021 | Effects of Outliers on the Maximum Correntropy Estimation: A Robustness AnalysisabstractRecently, maximum correntropy criterion (MCC) has been widely and successfully used in robust signal processing and machine learning, in which the correntropy is maximized instead of minimizing the popular mean square error (MSE) to improve the robustness with respect to outliers or impulsive noises. A lot of efforts have been devoted to derive different adaptive algorithms under MCC, but to date, little insight has been gained as to how the MCC solution will be influenced by outliers. In this paper, we investigate this problem and our focus is mainly on the parameter estimation of a simple linear errors-in-variables (EIVs) model with scalar variables. Under some conditions, we derive an upper bound on the absolute value of the estimation error and show that the MCC solution can get very close to the true value of the unknown parameter even with arbitrarily large outliers in both the input and output variables. Illustrative examples are provided to verify and clarify the theory. Badong Chen, Lei Xing 0003, Haiquan Zhao 0001, Shaoyi Du, José C. Príncipe |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2020 | Composite Dynamic Texture Synthesis Using Hierarchical Linear Dynamical SystemabstractWe demonstrate that a systematic inclusion of prior structural constraints on the states of a linear dynamical system significantly improves its ability to model complex multidimensional sequences. This constrained LDS, typically termed as the hierarchical linear dynamical system (HLDS), is a Kalman filter based topology that extracts relevant self-segmenting information from the input signal in an unsupervised manner by hierarchically constraining its information representing state subspaces thereby slowing down the signal dynamics. We highlight some of its practical advantages over the existing methods in real-world video applications. As a concrete application, we show that the HLDS, despite being a linear model trained in an unsupervised setting, is able to capture the dynamics of complex texture sequences consisting of multiple co-occurring textures. We compare its performance with a similarly trained LDS model in the reconstruction and synthesis of such signals. Rishabh Singh, Shujian Yu, José C. Príncipe |
ICASSP | 3 |
| 2020 | A Cognitive Architecture for Object Recognition in VideoabstractSummary form only given, as follows. The complete presentation was not made available for publication as part of the conference proceedings. This talk describes our efforts to abstract from the animal visual system the computational principles to explain images in video. We develop a hierarchical, distributed architecture of dynamical systems that self-organizes to explain the input imagery using an empirical Bayes criterion with sparseness constraints and dual state estimation. The interpretation of the images is mediated through causes that flow top down and change the priors for the bottom up processing. We will present preliminary results in several data sets. José C. Príncipe |
ICMLA | 1 |
| 2020 | Measuring the Discrepancy between Conditional Distributions: Methods, Properties and ApplicationsabstractWe propose a simple yet powerful test statistic to quantify the discrepancy between two conditional distributions. The new statistic avoids the explicit estimation of the underlying distributions in high-dimensional space and it operates on the cone of symmetric positive semidefinite (SPS) matrix using the Bregman matrix divergence. Moreover, it inherits the merits of the correntropy function to explicitly incorporate high-order statistics in the data. We present the properties of our new statistic and illustrate its connections to prior art. We finally show the applications of our new statistic on three different machine learning problems, namely the multi-task learning over graphs, the concept drift detection, and the information-theoretic feature selection, to demonstrate its utility and advantage. Code of our statistic is available at https://bit.ly/BregmanCorrentropy. Shujian Yu, Ammar Shaker, Francesco Alesiani, José C. Príncipe |
IJCAI | 4 |
| 2020 | Cognitive Architecture for Video GamesabstractThere has been an increasing interest in Frame-oriented reinforcement learning (FORL) in recent year. However, most of the works in the literature show little inspiration from human's perception action reward cycle (PARC) and causation.Inspired by human's vision system and learning strategy, we propose a novel architecture for FORL that understands the content of raw frames. The architecture achieves four objectives: 1. Extracting information from the environment by exploiting only unsupervised learning and reinforcement learning. 2. Understanding the content of a raw frame. 3. Exploiting a Folvea vision strategy which is analogous to human's vision system. 4. Establishing self-awareness and collecting new training data subset automatically to learn new objects without forgetting previous ones..The architecture is developed in the Super Mario Brothers video game.. At first, Mario is the only object recognized by the architecture. After automatic data subset collection and memory update, the architecture can recognize both Goomba and Mario and classify them using incremental training.We exemplify performance of each piece of the architecture with snippets obtained from the video game. José C. Príncipe |
IJCNN | 3 |
| 2020 | Regularized Training of Convolutional Autoencoders using the Rényi-Stratonovich Value of InformationabstractWe propose an information-theoretic cost function or the regularized training of convolutional auto encoders that imposes an organization on the bottleneck-layer-projected samples so as to facilitate discrimination. This function is based on a continuous-space, Rényi-mutual-information version of Stratonovich's value of information. It quantifies the maximum benefit that can be obtained for a given bottleneck-layer representation compression amount. The compression amount is controlled by a single hyperparameter that trades off between the autoencoder reconstruction quality and the hidden-layer representation uncertainty. Isaac J. Sledge, José C. Príncipe |
IJCNN | 2 |
| 2020 | Time Series Analysis using a Kernel based Multi-Modal Uncertainty Decomposition FrameworkabstractThis paper proposes a kernel based information theoretic framework with quantum physical underpinnings for data characterization that is relevant to online time series applications such as unsupervised change point detection and whole sequence clustering. In this framework, we utilize the Gaussian kernel mean embedding metric for universal characterization of data PDF. We then utilize concepts of quantum physics to impart a local dynamical structure to characterized data PDF, resulting in a new energy based formulation. This facilitates a multi-modal physics based uncertainty representation of the signal PDF at each sample using Hermite polynomial projections. We demonstrate in this paper using synthesized datasets that such uncertainty features provide a better ability for online detection of statistical change points in time series data when compared to existing non-parametric and unsupervised methods. We also demonstrate a better ability of the framework in clustering time series sequences when compared to discrete wavelet transform features on a subset of VidTIMIT speaker recognition corpus. Rishabh Singh, José C. Príncipe |
UAI | 2 |
| 2020 | On Kernel Method-Based Connectionist Models and Supervised Deep Learning Without BackpropagationabstractWe propose a novel family of connectionist models based on kernel machines and consider the problem of learning layer by layer a compositional hypothesis class (i.e., a feedforward, multilayer architecture) in a supervised setting. In terms of the models, we present a principled method to “kernelize” (partly or completely) any neural network (NN). With this method, we obtain a counterpart of any given NN that is powered by kernel machines instead of neurons. In terms of learning, when learning a feedforward deep architecture in a supervised setting, one needs to train all the components simultaneously using backpropagation (BP) since there are no explicit targets for the hidden layers (Rumelhart, Hinton, & Williams, 1986 ). We consider without loss of generality the two-layer case and present a general framework that explicitly characterizes a target for the hidden layer that is optimal for minimizing the objective function of the network. This characterization then makes possible a purely greedy training scheme that learns one layer at a time, starting from the input layer. We provide instantiations of the abstract framework under certain architectures and objective functions. Based on these instantiations, we present a layer-wise training algorithm for an [Formula: see text]-layer feedforward network for classification, where [Formula: see text] can be arbitrary. This algorithm can be given an intuitive geometric interpretation that makes the learning dynamics transparent. Empirical results are provided to complement our theory. We show that the kernelized networks, trained layer-wise, compare favorably with classical kernel machines as well as other connectionist models trained by BP. We also visualize the inner workings of the greedy kernelized models to validate our claim on the transparency of the layer-wise algorithm. Shiyu Duan, Shujian Yu, Yunmei Chen, José C. Príncipe |
Neural Comput. | 4 |
| 2020 | Multivariate Extension of Matrix-Based Rényi's $\alpha$α-Order Entropy FunctionalabstractThe matrix-based Rényi's α-order entropy functional was recently introduced using the normalized eigenspectrum of a Hermitian matrix of the projected data in a reproducing kernel Hilbert space (RKHS). However, the current theory in the matrix-based Rényi's α-order entropy functional only defines the entropy of a single variable or mutual information between two random variables. In information theory and machine learning communities, one is also frequently interested in multivariate information quantities, such as the multivariate joint entropy and different interactive quantities among multiple variables. In this paper, we first define the matrix-based Rényi's α-order joint entropy among multiple variables. We then show how this definition can ease the estimation of various information quantities that measure the interactions among multiple variables, such as interactive information and total correlation. We finally present an application to feature selection to show how our definition provides a simple yet powerful way to estimate a widely-acknowledged intractable quantity from data. A real example on hyperspectral image (HSI) band selection is also provided. Shujian Yu, Luis Gonzalo Sánchez Giraldo, Robert Jenssen, José C. Príncipe |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | A Taxonomy for Neural Memory NetworksabstractAn increasing number of neural memory networks have been developed, leading to the need for a systematic approach to analyze and compare their underlying memory structures. Thus, in this paper, we first create a framework for memory organization and then compare four popular dynamic models: vanilla recurrent neural network, long short-term memory, neural stack, and neural RAM. This analysis helps to open the dynamic neural networks' black box from the memory usage prospective. Accordingly, a taxonomy for these networks and their variants is proposed and proved using a unifying architecture. With the taxonomy, both network architectures and learning tasks are classified into four classes, and a one-to-one mapping is built between them to help practitioners select the appropriate architecture. To exemplify each task type, four synthetic tasks with different memory requirements are selected. Moreover, we use some signal processing applications and two natural language processing applications to evaluate the methodology in a realistic setting. José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Probability Density Rank-Based Quantization for Convex Universal Learning MachinesabstractThe distributions of input data are very important for learning machines, such as the convex universal learning machines (CULMs). The CULMs are a family of universal learning machines with convex optimization. However, the computational complexity is a crucial problem in CULMs, because the dimension of the nonlinear mapping layer (the hidden layer) of the CULMs is usually rather large in complex system modeling. In this article, we propose an efficient quantization method called Probability density Rank-based Quantization (PRQ) to decrease the computational complexity of CULMs. The PRQ ranks the data according to the estimated probability densities and then selects a subset whose elements are equally spaced in the ranked data sequence. We apply the PRQ to kernel ridge regression (KRR) and random Fourier feature recursive least squares (RFF-RLS), which are two typical algorithms of CULMs. The proposed method not only keeps the similarity of data distribution between the code book and data set but also reduces the computational cost by using the kd-tree. Meanwhile, for a given data set, the method yields deterministic quantization results, and it can also exclude the outliers and avoid too many borders in the code book. This brings great convenience to practical applications of the CULMs. The proposed PRQ is evaluated on several real-world benchmark data sets. Experimental results show satisfactory performance of PRQ compared with some state-of-the-art methods. Zhengda Qin, Badong Chen, Yuantao Gu, Nanning Zheng 0001, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | An Exact Reformulation of Feature-Vector-Based Radial-Basis-Function Networks for Graph-Based ObservationsabstractRadial basis function (RBF) networks are traditionally defined for sets of vector-based observations. In this brief, we reformulate such networks so that they can be applied to adjacency-matrix representations of weighted, directed graphs that represent the relationships between object pairs. We restate the sum-of-squares objective function so that it is purely dependent on entries from the adjacency matrix. From this objective function, we derive a gradient descent update for the network weights. We also derive a gradient update that simulates the repositioning of the radial basis prototypes and changes in the radial basis prototype parameters. An important property of our radial basis function networks is that they are guaranteed to yield the same responses as conventional radial basis networks trained on a corresponding vector realization of the relationships encoded by the adjacency matrix. Such a vector realization only needs to provably exist for this property to hold, which occurs whenever the relationships correspond to distances from some arbitrary metric applied to a latent set of vectors. We, therefore, completely avoid needing to actually construct vectorial realizations via multidimensional scaling, which ensures that the underlying relationships are totally preserved. Isaac J. Sledge, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | A Differential-geometric Approach for Globally Solving a Non-convex, Discontinuous Depth Estimation Problem for Plenoptic Camera ImagesabstractIn this paper, we address the problem of estimating a scene's three-dimensional geometry from plenoptic camera images. Existing approaches for this problem have emphasized the development of sharpness and contrast measures for distinguishing between in-/out-of-focus image regions. The ways in which these measures are aggregated can yield erroneous, localized distance fluctuations, though. To deal with such fluctuations, post-processing smoothing techniques can be applied. However, they may remove fine-scale, non-erroneous depth structures and edges. Here, we propose a non-convex, discontinuous cost-function that simultaneously combines and regularizes sharpness and contrast so that valid depth transitions are better preserved. We implicitly convert this function into one that is continuous and (quasi-)convex by optimizing it on the non-positively-curved Riemannian manifold of depth maps with a learned metric. Isaac J. Sledge, José C. Príncipe |
ICASSP | 2 |
| 2019 | An Information-theoretic Approach for Automatically Determining the Number of State Groups When Aggregating Markov ChainsabstractA fundamental problem when aggregating Markov chains is the specification of the number of state groups. Too few state groups may fail to sufficiently capture the pertinent dynamics of the original, high-order Markov chain. Too many state groups may lead to a non-parsimonious, reduced-order Markov chain whose complexity rivals that of the original. In this paper, we show that an augmented value-of-information-based approach to aggregating Markov chains facilitates the determination of the number of state groups. The optimal state-group count coincides with the case where the complexity of the reduced-order chain is balanced against the mutual dependence between the original- and reduced-order chain dynamics. Isaac J. Sledge, José C. Príncipe |
ICASSP | 2 |
| 2019 | Using a Recurrent Kernel Learning Machine for Small-Sample Image ClassificationabstractMany machine learning algorithms, like Convolutional Neural Networks (CNNs), have excelled in image processing tasks; however, they have many practical limitations. For one, these systems require large datasets that accurately represent the sample distribution in order to optimize performance. Secondly, they have difficulty transferring previously learned knowledge when evaluating data from slightly different sample distributions. To overcome these drawbacks, we propose a recurrent kernel-based approach for image processing using the Kernel Adaptive Autoregressive Moving Average algorithm (KAARMA). KAARMA minimizes the amount of training data required by using the Reproducing Kernel Hilbert Space to build inference into the system. The recurrent nature of KAARMA additionally allows the system to better learn the spatial correlations in the images through one-shot or near one-shot learning. We demonstrate KAARMA's superiority for small-sample image classification using the JAFFE Face Dataset and the UCI hand written digit dataset. Mihael Cudic, José C. Príncipe |
IJCNN | 2 |
| 2019 | Fast segmentation for large and sparsely labeled coral imagesabstractMarine organism datasets often present sparse annotated labels and with many objects in cluttered background. Therefore, there are two challenges to do image segmentation on these sparsely labeled datasets: one is to obtain denser labeled training data and the other is to improve the speed of testing on large images. In this paper, we propose a label augmentation method to generate more labels for training based on the superpixel algorithm, and we also create coarse-to-fine approach to detect the coral areas quickly in the large images. Our experiments run on coral image dataset collected in Pulley Ridge1, proving that this label augmentation and coarse-to-fine approach allows us to speed up the process of quantifying the percent of corals in large images while preserving accuracy. Stephanie Farrington, John Reed, Bing Ouyang, José C. Príncipe |
IJCNN | 6 |
| 2019 | Understanding autoencoders with information theoretic concepts
Shujian Yu, José C. Príncipe |
Neural Networks | 2 |
| 2019 | Maximum Correntropy Criterion With Variable CenterabstractCorrentropy is a local similarity measure defined in kernel space, and the maximum correntropy criterion (mcc) has been successfully applied in many areas of signal processing and machine learning in recent years. The kernel function in correntropy is usually restricted to the Gaussian function with the center located at zero. However, the zero-mean Gaussian function may not be a good choice for many practical applications. In this letter, we propose an extended version of correntropy, whose center can be located at any position. Accordingly, we propose a new optimization criterion called maximum correntropy criterion with variable center (MCC-VC). We also propose an efficient approach to optimize the kernel width and center location in the MCC-VC. Simulation results of regression with linear-in-parameter (LIP) models confirm the desirable performance of the new method. Badong Chen, Yingsong Li 0001, José C. Príncipe |
IEEE Signal Process. Lett. | 4 |
| 2019 | Analysis of Agent Expertise in Ms. Pac-Man Using Value-of-Information-Based PoliciesabstractConventional reinforcement-learning methods for Markov decision processes rely on weakly guided, stochastic searches to drive the learning process. It can therefore be difficult to predict what agent behaviors might emerge. In this paper, we consider an information-theoretic cost function for performing constrained stochastic searches that promote the formation of risk-averse to risk-favoring behaviors. This cost function is the value of information, which provides the optimal tradeoff between the expected return of a policy and the policy's complexity; policy complexity is measured by number of bits and controlled by a single hyperparameter on the cost function. As the policy complexity is reduced, the agents will increasingly eschew risky actions. This reduces the potential for high accrued rewards. As the policy complexity increases, the agents will take actions, regardless of the risk, that can raise the long-term rewards. The obtainable reward depends on a single, tunable hyperparameter that regulates the degree of policy complexity. We evaluate the performance of value-of-information-based policies on a stochastic version of Ms. Pac-Man. A major component of this paper is the demonstration that ranges of policy complexity values yield different game-play styles and explaining why this occurs. We also show that our reinforcement-learning search mechanism is more efficient than the others we utilize. This result implies that the value-of-information theory is appropriate for framing the exploitation-exploration tradeoff in reinforcement learning. Isaac J. Sledge, José C. Príncipe |
IEEE Trans. Games | 2 |
| 2019 | Quantized Minimum Error Entropy CriterionabstractComparing with traditional learning criteria, such as mean square error, the minimum error entropy (MEE) criterion is superior in nonlinear and non-Gaussian signal processing and machine learning. The argument of the logarithm in Renyi's entropy estimator, called information potential (IP), is a popular MEE cost in information theoretic learning. The computational complexity of IP is, however, quadratic in terms of sample number due to double summation. This creates the computational bottlenecks, especially for large-scale data sets. To address this problem, in this paper, we propose an efficient quantization approach to reduce the computational burden of IP, which decreases the complexity from O(N2) to O(MN) with M ≪ N. The new learning criterion is called the quantized MEE (QMEE). Some basic properties of QMEE are presented. Illustrative examples with linear-in-parameter models are provided to verify the excellent performance of QMEE. Badong Chen, Lei Xing 0003, Nanning Zheng 0001, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Nearest-Instance-Centroid-Estimation Linear Discriminant Analysis (Nice Lda)abstractWe propose a novel cascaded classification technique called the Nearest Instance Centroid Estimation (NICE) LDA algorithm. Our algorithm (inspired from NICE KLMS) performs a cascade combination of two weak classifiers - threshold based class-wise clustering and linear discriminant classification to achieve state-of-the-art results on various high dimensional UCI datasets. We show how our method is more robust towards skewed data and computationally more efficient than previous methods of combining clustering with classification techniques. We also develop an efficient aggregation method based on instance based learning that implements this cascade combination of classifiers in a much simpler manner computationally. We demonstrate that our method of data clustering and LDA implementation, while introducing only one free parameter, leads to results that are similar and often better than those achieved by the state-of-the-art kernel RBF SVMs. Rishabh Singh, Kan Li 0002, José C. Príncipe |
ICASSP | 3 |
| 2018 | Partitioning Relational Matrices of Similarities or Dissimilarities Using the Value of InformationabstractIn this paper, we provide an approach to clustering relational matrices whose entries correspond to either similarities or dissimilarities between objects. Our approach is based on the value of information, a parameterized, information-theoretic criterion that measures the change in costs associated with changes in information. Optimizing the value of information yields a deterministic annealing style of clustering with many benefits. For instance, investigators avoid needing to a priori specify the number of clusters, as the partitions naturally undergo phase changes, during the annealing process, whereby the number of clusters changes in a data-driven fashion. The global-best partition can also often be identified. Isaac J. Sledge, José C. Príncipe |
ICASSP | 2 |
| 2018 | Request-and-Reverify: Hierarchical Hypothesis Testing for Concept Drift Detection with Expensive LabelsabstractOne important assumption underlying common classification models is the stationarity of the data. However, in real-world streaming applications, the data concept indicated by the joint distribution of feature and label is not stationary but drifting over time. Concept drift detection aims to detect such drifts and adapt the model so as to mitigate any deterioration in the model's predictive performance. Unfortunately, most existing concept drift detection methods rely on a strong and over-optimistic condition that the true labels are available immediately for all already classified instances. In this paper, a novel Hierarchical Hypothesis Testing framework with Request-and-Reverify strategy is developed to detect concept drifts by requesting labels only when necessary. Two methods, namely Hierarchical Hypothesis Testing with Classification Uncertainty (HHT-CU) and Hierarchical Hypothesis Testing with Attribute-wise "Goodness-of-fit" (HHT-AG), are proposed respectively under the novel framework. In experiments with benchmark datasets, our methods demonstrate overwhelming advantages over state-of-the-art unsupervised drift detectors. More importantly, our methods even outperform DDM (the widely used supervised drift detector) when we use significantly fewer labels. Shujian Yu, José C. Príncipe |
IJCAI | 3 |
| 2018 | Top-down Gamma Saliency - Learning to Search for Objects in Complex ScenesabstractSaliency measures are often used to predict fixation location in images. However, a pure bottom up saliency is not useful for visual search in a complex scene with many objects since it is only driven by the input image. Alternatively, neural networks can localize objects within scenes, but rely on a brute force classification of heuristic bounding boxes. We propose a top-down attention mechanism that combines the traditional saliency measures with the learned ability of neural networks to distinguish between objects. To do this, we will use a set of feature maps produced by the convolutional layers of a trained classification network as the inputs to our saliency measure instead of a traditional RBG or LAB image. On top of these feature maps, we can learn a set of weights to bias the saliency towards specific objects. We test this top-down approach against the traditional bottom-up approach in a synthetic environment where it proves to be more adept at finding specific objects quickly in crowded scenes. Ryan Burt, José C. Príncipe |
IJCNN | 2 |
| 2018 | Surprise-Novelty Information Processing for Gaussian Online Active Learning (SNIP-GOAL)abstractIn this paper, we propose a novel, combined surprise-novelty approach to online active learning for kernel adaptive filters. Surprise and novelty criteria have always been used individually in designing sparse kernel machines. While closely related, they are in fact two distinct concepts. In this paper, we highlight the key differences between surprise and novelty, define quantitative measures, and propose a unifying framework that leverages the complementary properties of the two concepts combined. We test the information theoretic approach of designing sparse kernel adaptive filters using the surprise-novelty information processing for Gaussian online active learning (SNIP-GOAL) on the task of nonlinear chaotic time series prediction. The proposed method outperforms existing algorithms using the measures individually. Results show that combining surprise and novelty can be advantageous in terms of efficiency and performance. Leveraging both measures allows the system to be not only sparse but also generalizes better. Kan Li 0002, José C. Príncipe |
IJCNN | 2 |
| 2018 | Comparison of Static Neural Network with External Memory and RNNs for Deterministic Context Free Language LearningabstractIn this paper, a learning model for prediction is introduced by coupling a static neural network with an external stack memory, creating a new type of recurrent system. We analyze the differences between this external memory recurrent network and recurrent neural network, which possesses internal memory. Internal memory remembers the last state while external memory remembers past useful contents. For a specific automaton, the internal memory is needed if the last state is a variable of the state transition function and the external memory is needed if the past content is a variable of state transition function. Our arguments are verified by comparing the prediction accuracy of different models with internal memory, external memory and the combination of them on counting and reversing tasks. The results shows that: network with an external stack works best for counting tasks since the variables of state transition function is composed of the current input and one past input, while network with the combination of internal and external memory works best for the reversing task since the variables of state transition function is composed of the current input, current state and the past inputs. José C. Príncipe |
IJCNN | 2 |
| 2018 | Augmented Space Linear ModelabstractThe linear model uses the space defined by the input to project the target or desired signal and find the optimal set of model parameters. When the problem is nonlinear, the adaption requires nonlinear models for good performance, but it becomes slower and more cumbersome. In this paper, we propose a linear model called Augmented Space Linear Model (ASLM), which uses the full joint space of input and desired signal as the projection space and approaches the performance of nonlinear models. This new algorithm takes advantage of the linear solution, and corrects the estimate for the current testing phase input with the error assigned to the input space neighborhood in the training phase. This algorithm can solve the nonlinear problem with the computational efficiency of linear methods, which can be regarded as a trade off between accuracy and computational complexity. Making full use of the training data, the proposed augmented space model may provide a new way to improve many modeling tasks. Zhengda Qin, Badong Chen, Nanning Zheng 0001, José C. Príncipe |
IJCNN | 4 |
| 2018 | Correntropy Based Hierarchical Linear Dynamical System For Speech RecognitionabstractHierarchical Linear Dynamical System (HLDS) is a recently introduced Kalman filter based generative state model that extracts relevant self-segmenting information from input time series signal by hierarchically constraining the information representing subspaces of its states thus slowing down the dynamics of the input signal. Despite the simplicity of its nested architecture and its dependance on linear Kalman update rules, the HLDS has been shown to have state-of-the-art performance in the classification of musical notes. However, it was observed that the application scope of this state based model was only limited to linearly separable signals since the representations of non-linear and non-stationary signals (such as speech) in the state space of the HLDS was highly intermingled and hence, non-discriminative. This paper proposes a kernel based extension of the HLDS that shows promising results in the sparse and discriminative representation of speech phonemes in its information representing state spaces. Specifically, we use correntropy as an additional non-linear constraint on top of the linear constraints already provided by the nested architecture of the states. We show that by using correntropy as the cost function in the Kalman update equations, we are able to adaptively restrict the different phonemes of a speech signal into localized areas of the top state space. Our training results, along with their authentication through top-down inference of the states, provide valid credibility to the use of HLDS as a promising speech recognition model. Rishabh Singh, José C. Príncipe |
IJCNN | 2 |
| 2018 | Kernelized Q-Learning for Large-Scale, Potentially Continuous, Markov Decision ProcessesabstractWe introduce a novel means of generalizing experi- ences agent experiences for large-scale Markov decision processes. Our approach is based on a kernel local linear regression function approximation, which we combine with Q-learning. Through this kernelized regression process, value function estimates from visited portions of the state-action space can be generalized to those areas that have not yet been visited in a non-linear, non-parametric fash- ion. This can be done when the state-action space is either discrete or continuous. We assess the performance of our approach on the game Su- per Mario Land 2 for the Nintendo GameBoy system. We show thatbetter performance is obtained with our kernelized Q-learning approach compared to linear function approximators for this com- plicated environment. Better performance is also witnessed with our approach compared to other non-linear approximators. Isaac J. Sledge, José C. Príncipe |
IJCNN | 2 |
| 2018 | Complex correntropy function: Properties, and application to a channel equalization problem
João P. F. Guimarães, Aluisio I. Rêgo Fontes, Joilson B. A. Rego, Allan de Medeiros Martins, José C. Príncipe |
Expert Syst. Appl. | 5 |
| 2018 | A flexible testing environment for visual question answering with performance evaluation
Mihael Cudic, Ryan Burt, Eder Santana, José C. Príncipe |
Neurocomputing | 4 |
| 2018 | Self-Organised direction aware data partitioning algorithm
Xiaowei Gu 0001, Plamen Angelov 0001, Dmitry Kangin, José C. Príncipe |
Inf. Sci. | 4 |
| 2018 | A method for autonomous data partitioning
Xiaowei Gu 0001, Plamen Angelov 0001, José C. Príncipe |
Inf. Sci. | 3 |
| 2018 | A Generalized Methodology for Data AnalysisabstractBased on a critical analysis of data analytics and its foundations, we propose a functional approach to estimate data ensemble properties, which is based entirely on the empirical observations of discrete data samples and the relative proximity of these points in the data space and hence named empirical data analysis (EDA). The ensemble functions include the nonparametric square centrality (a measure of closeness used in graph theory) and typicality (an empirically derived quantity which resembles probability). A distinctive feature of the proposed new functional approach to data analysis is that it does not assume randomness or determinism of the empirically observed data, nor independence. The typicality is derived from the discrete data directly in contrast to the traditional approach, where a continuous probability density function is assumed a priori. The typicality is expressed in a closed analytical form that can be calculated recursively and, thus, is computationally very efficient. The proposed nonparametric estimators of the ensemble properties of the data can also be interpreted as a discrete form of the information potential (known from the information theoretic learning theory as well as the Parzen windows). Therefore, EDA is very suitable for the current move to a data-rich environment, where the understanding of the underlying phenomena behind the available vast amounts of data is often not clear. We also present an extension of EDA for inference. The areas of applications of the new methodology of the EDA are wide because it concerns the very foundation of data analysis. Preliminary tests show its good performance in comparison to traditional techniques. Plamen Angelov 0001, Xiaowei Gu 0001, José C. Príncipe |
IEEE Trans. Cybern. | 3 |
| 2018 | Autonomous Learning Multimodel Systems From Data StreamsabstractIn this paper, an approach to autonomous learning of a multimodel system from streaming data, named ALMMo, is proposed. The proposed approach is generic and can easily be applied also to probabilistic or other types of local models forming multimodel systems. It is fully data driven and its structure is decided by the nonparametric data clouds extracted from the empirically observed data without making any prior assumptions concerning data distribution and other data properties. All metaparameters of the proposed system are obtained directly from the data and can be updated recursively, which improves memory and calculation efficiencies of the proposed algorithm. The structural evolution mechanism and online data cloud quality monitoring mechanism of the ALMMo system largely enhance the ability of handling shifts and/or drifts in the streaming data pattern. Numerical examples of the use of ALMMo system for streaming data analytics, classification, and prediction are presented as a proof of the proposed concept. Plamen Angelov 0001, Xiaowei Gu 0001, José C. Príncipe |
IEEE Trans. Fuzzy Syst. | 3 |
| 2018 | Insights Into the Robustness of Minimum Error Entropy EstimationabstractThe minimum error entropy (MEE) is an important and highly effective optimization criterion in information theoretic learning (ITL). For regression problems, MEE aims at minimizing the entropy of the prediction error such that the estimated model preserves the information of the data generating system as much as possible. In many real world applications, the MEE estimator can outperform significantly the well-known minimum mean square error (MMSE) estimator and show strong robustness to noises especially when data are contaminated by non-Gaussian (multimodal, heavy tailed, discrete valued, and so on) noises. In this brief, we present some theoretical results on the robustness of MEE. For a one-parameter linear errors-in-variables (EIV) model and under some conditions, we derive a region that contains the MEE solution, which suggests that the MEE estimate can be very close to the true value of the unknown parameter even in presence of arbitrarily large outliers in both input and output variables. Theoretical prediction is verified by an illustrative example. Badong Chen, Lei Xing 0003, Bin Xu 0003, Haiquan Zhao 0001, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | Exploiting Spatio-Temporal Structure With Recurrent Winner-Take-All NetworksabstractWe propose a convolutional recurrent neural network (ConvRNNs), with winner-take-all (WTA) dropout for high-dimensional unsupervised feature learning in multidimensional time series. We apply the proposed method for object recognition using temporal context in videos and obtain better results than comparable methods in the literature, including the deep predictive coding networks (DPCNs) previously proposed by Chalasani and Principe. Our contributions can be summarized as a scalable reinterpretation of the DPCNs trained end-to-end with backpropagation through time, an extension of the previously proposed WTA autoencoders to sequences in time, and a new technique for initializing and regularizing ConvRNNs. Eder Santana, Matthew Emigh, Pablo Zegers, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2018 | Guided Policy Exploration for Markov Decision Processes Using an Uncertainty-Based Value-of-Information CriterionabstractReinforcement learning in environments with many action-state pairs is challenging. The issue is the number of episodes needed to thoroughly search the policy space. Most conventional heuristics address this search problem in a stochastic manner. This can leave large portions of the policy space unvisited during the early training stages. In this paper, we propose an uncertainty-based, information-theoretic approach for performing guided stochastic searches that more effectively cover the policy space. Our approach is based on the value of information, a criterion that provides the optimal tradeoff between expected costs and the granularity of the search process. The value of information yields a stochastic routine for choosing actions during learning that can explore the policy space in a coarse to fine manner. We augment this criterion with a state-transition uncertainty factor, which guides the search process into previously unexplored regions of the policy space. We evaluate the uncertainty-based value-of-information policies on the games Centipede and Crossy Road. Our results indicate that our approach yields better performing policies in fewer episodes than stochastic-based exploration strategies. We show that the training rate for our approach can be further improved by using the policy cross entropy to guide our criterion's hyperparameter selection. Isaac J. Sledge, Matthew Emigh, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Robust C-Loss Kernel ClassifiersabstractThe correntropy-induced loss (C-loss) function has the nice property of being robust to outliers. In this paper, we study the C-loss kernel classifier with the Tikhonov regularization term, which is used to avoid overfitting. After using the half-quadratic optimization algorithm, which converges much faster than the gradient optimization algorithm, we find out that the resulting C-loss kernel classifier is equivalent to an iterative weighted least square support vector machine (LS-SVM). This relationship helps explain the robustness of iterative weighted LS-SVM from the correntropy and density estimation perspectives. On the large-scale data sets which have low-rank Gram matrices, we suggest to use incomplete Cholesky decomposition to speed up the training process. Moreover, we use the representer theorem to improve the sparseness of the resulting C-loss kernel classifier. Experimental results confirm that our methods are more robust to outliers than the existing common classifiers. Guibiao Xu, Bao-Gang Hu, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Special Issue on Deep Reinforcement Learning and Adaptive Dynamic ProgrammingabstractThe sixteen papers in this special section focus on deep reinforcement learning and adaptive dynamic programming (deep RL/ADP). Deep RL is able to output control signal directly based on input images, which incorporates both the advantages of the perception of deep learning (DL) and the decision making of RL or adaptive dynamic programming (ADP). This mechanism makes the artificial intelligence much closer to human thinking modes. Deep RL/ADP has achieved remarkable success in terms of theory and applications since it was proposed. Successful applications cover video games, Go, robotics, smart driving, healthcare, and so on. However, it is still an open problem to perform the theoretical analysis on deep RL/ADP, e.g., the convergence, stability, and optimality analyses. The learning efficiency needs to be improved by proposing new algorithms or combined with other methods. More practical demonstrations are encouraged to be presented. Therefore, the aim of this special issue is to call for the most advanced research and state-of-the-art works in the field of deep RL/ADP. Dongbin Zhao, Derong Liu 0001, Frank L. Lewis, José C. Príncipe, Stefano Squartini |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Group-Wise Point-Set Registration Based on Rényi's Second Order EntropyabstractIn this paper, we describe a set of robust algorithms for group-wise registration using both rigid and non-rigid transformations of multiple unlabelled point-sets with no bias toward a given set. These methods mitigate the need to establish a correspondence among the point-sets by representing them as probability density functions where the registration is treated as a multiple distribution alignment. Holder's and Jensen's inequalities provide a notion of similarity/distance among point-sets and Rényi's second order entropy yields a closed-form solution to the cost function and update equations. We also show that the methods can be improved by normalizing the entropy with a scale factor. These provide simple, fast and accurate algorithms to compute the spatial transformation function needed to register multiple point-sets. The algorithms are compared against two well-known methods for group-wise point-set registration. The results show an improvement in both accuracy and computational complexity. Luis Gonzalo Sánchez Giraldo, Erion Hasanbelliu, Murali Rao, José C. Príncipe |
CVPR | 4 |
| 2017 | Automatic insect recognition using optical flight dynamics modeled by kernel adaptive ARMA networkabstractAutomatic insect recognition (AIR), using noninvasive methods in situ, has far-reaching implications in entomology, agriculture, and disease control and prevention. An emerging technology in computational entomology uses flight information captured by laser sensors. Current methods treat these optical signals as static patterns, rather than time series. We propose a novel approach to AIR by evaluating each insect passage as a nonstationary process involving a sequence of pseudo-acoustic frames and modeling the short-term flight dynamics using the kernel adaptive autoregressive-moving average (KAARMA) algorithm. Since flight behavior is both nonlinear and nonstationary in nature, dynamic modeling provides a general framework that fully exploits the transitional and contextual information. Results show KAARMA classifier outperforms the state-of-the-art AIR methods, using support vector machine (SVM), deep-learning autoencoder, and batch learning, in identifying Zika vector mosquito Aedes aegypti among five species of flying insects, while using significantly more efficient data representation. Kan Li 0002, José C. Príncipe |
ICASSP | 2 |
| 2017 | A novel methodology to quantify dense EEG in cognitive tasksabstractCognition emerges from complex interaction amongst widespread brain areas. In this paper, we use a novel methodology for temporal networks quantification for EEG. We model the spatiotemporal structure of dependencies across different electrodes with respect to a single electrode as a local probability density function. This enables immediately the use of information theoretic quantities (information divergences) to quantify brain connectivity in simple two-dimensional graphs. We show that for a visual-motor-driven task, we are able to cluster subjects that performed the task with higher attention-coefficient, in an unsupervised-fashion. We test this methodology with two measures of functional connectivity: correlation coefficient and a measure of association. Catia S. Silva, José C. Príncipe, Andreas Keil |
ICASSP | 2 |
| 2017 | Balancing exploration and exploitation in reinforcement learning using a value of information criterionabstractIn this paper, we consider an information-theoretic approach for addressing the exploration-exploitation dilemma in reinforcement learning. We employ the value of information, a criterion that provides the optimal trade-off between the expected returns and a policy's degrees of freedom. As the degrees of freedom are reduced, an agent will exploit more than explore. As the policy degrees of freedom increase, an agent will explore more than exploit. We provide an efficient computational procedure for constructing policies using the value of information. The performance is demonstrated on a standard reinforcement learning benchmark problem. Isaac J. Sledge, José C. Príncipe |
ICASSP | 2 |
| 2017 | Autoencoders trained with relevant information: Blending Shannon and Wiener's perspectivesabstractIt is almost seventy years after the publication of Claude Shannon's “A Mathematical Theory of Communication” [1] and Norbert Wiener's “Extrapolation, Interpolation and Smoothing of Stationary Time Series” [2]. The pioneering works of Shannon and Wiener lay the foundation of communication, data storage, control, and other information technologies. This paper briefly reviews Shannon and Wiener's perspectives on the problem of message transmission over noisy channel and also experimentally evaluates the feasibility of integrating these two perspectives to train autoencoders close to the information limit. To this end, the principle of relevant information (PRI) is used and validated to optimally encode input imagery in the presence of noise. Shujian Yu, Matthew Emigh, Eder Santana, José C. Príncipe |
ICASSP | 4 |
| 2017 | Fast feedforward non-parametric deep learning network with automatic feature extractionabstractIn this paper, a new type of feedforward non-parametric deep learning network with automatic feature extraction is proposed. The proposed network is based on human-understandable local aggregations extracted directly from the images. There is no need for any feature selection and parameter tuning. The proposed network involves nonlinear transformation, segmentation operations to select the most distinctive features from the training images and builds RBF neurons based on them to perform classification with no weights to train. The design of the proposed network is very efficient (computation and time wise) and produces highly accurate classification results. Moreover, the training process is parallelizable, and the time consumption can be further reduced with more processors involved. Numerical examples demonstrate the high performance and very short training process of the proposed network for different applications. Plamen Angelov 0001, Xiaowei Gu 0001, José C. Príncipe |
IJCNN | 3 |
| 2017 | Fusing attention with visual question answeringabstractVisual Question Answering is a complex problem that fuses natural language and image processing to answer a question based on information from the image. The basic architecture for accomplishing this is using a CNN to extract features from the image and an RNN for the language processing, then combine the two in an MLP to produce an answer. These architectures perform well at identifying content, but fail at higher level reasoning such as spatial awareness and combining objects. To help remedy this, we propose using attention to divide the image into separate objects, then using the extracted features along with the location and size information to learn the MLP. Ryan Burt, Mihael Cudic, José C. Príncipe |
IJCNN | 3 |
| 2017 | Flight dynamics modeling and recognition using finite state machine for automatic insect recognitionabstractIn this paper, we propose a novel state-based approach to insect flight dynamics modeling and automatic insect recognition (AIR). An emerging technology uses noninvasive laser sensors to capture flight information of flying insects in situ. However current computational entomology methods treat these optical signals as static patterns, rather than time series. We propose to evaluate each insect passage as a nonstationary process involving a sequence of pseudo-acoustic frames and model the short-term flight dynamics using a kernel adaptive autoregressive-moving-average (KAARMA) network. Since flight behavior is both nonlinear and nonstationary in nature, dynamic modeling provides a general framework that fully exploits the transitional and contextual information. Using the kernel adaptive ARMA algorithm and spatial clustering, each flight passage is defined to be an ordered sequence of states in a discretized spatiotemporal space. The state transition trajectories are used to construct a finite state machine (FSM) recognizer. The computational efficiency of the automaton classifier eliminates the need to perform numeric computations and enables realtime performance with far-reaching implications in entomology, agriculture, and disease control and prevention. Results show multiclass KAARMA classifier outperforms the state-of-the-art AIR methods, using support vector machine (SVM), deep learning autoencoder, and batch learning, in identifying Zika vector mosquito Aedes aegypti among five species of flying insects, while using significantly more efficient data representation. The extracted FSM from trained KAARMA network retained competitive performance compared to the current computationally expensive implementations. Kan Li 0002, José C. Príncipe |
IJCNN | 2 |
| 2017 | Adaptive filtering based on extended kernel recursive maximum correntropyabstractIn this paper, an adaptive filtering algorithm, termed the extended kernel recursive maximum correntropy (EX-KRMC) algorithm is proposed as a novel approach of traditional recursion based adaptive filtering algorithms. Maximum correntropy criterion is employed to better the robustness to non-Gaussian noise and kernel methods are used to enable the capacity for nonlinear systems. It is verified by simulation experiments that EX-KRMC outperforms existing adaptive filtering algorithms when dealing with non-Gaussian noise for nonlinear time-variant systems. Shengyang Luan, Tianshuang Qiu, José C. Príncipe |
IJCNN | 3 |
| 2017 | Cyclostationary correntropy: Definition and applications
Aluisio I. Rêgo Fontes, Joilson B. A. Rego, Allan de Medeiros Martins, Luiz Felipe Q. Silveira, José C. Príncipe |
Expert Syst. Appl. | 5 |
| 2017 | Corrigendum to "Cyclostationary Correntropy: Definition and applications" [Expert Systems with Applications 69 (2017) 110-117]
Aluisio I. Rêgo Fontes, Joilson B. A. Rego, Allan de Medeiros Martins, Luiz Felipe Q. Silveira, José C. Príncipe |
Expert Syst. Appl. | 5 |
| 2017 | Marine animal classification using UMSLI in HBOI optical test facility
José C. Príncipe, Bing Ouyang, Fraser R. Dalgleish, Anni K. Vuorenkoski, Brian Ramos, Gabriel Alsenas |
Multim. Tools Appl. | 2 |
| 2017 | Robust support vector machines based on the rescaled hinge loss function
Guibiao Xu, Bao-Gang Hu, José C. Príncipe |
Pattern Recognit. | 4 |
| 2017 | Complex Correntropy: Probabilistic Interpretation and Application to Complex-Valued DataabstractRecent studies have demonstrated that correntropy is an efficient tool for analyzing higher order statistical moments in non-Gaussian noise environments. Although correntropy has been used with complex data, no theoretical study was pursued to elucidate its properties, nor how to best use it for optimization. By using a probabilistic interpretation, this work presents a novel similarity measure between two complex random variables, which is defined as complex correntropy. A new recursive solution for the maximum complex correntropy criterion is introduced based on a fixed-point solution. This technique is applied to a system identification, and the results demonstrate prominent advantages when compared against three other algorithms: the complex least mean square, complex recursive least squares, and least absolute deviation. By the aforementioned probabilistic interpretation, correntropy can now be applied to solve several problems involving complex data in a more straightforward way. João P. F. Guimarães, Aluisio I. Rêgo Fontes, Joilson B. A. Rego, Allan de Medeiros Martins, José C. Príncipe |
IEEE Signal Process. Lett. | 5 |
| 2017 | Quantized Attention-Gated Kernel Reinforcement Learning for Brain-Machine Interface DecodingabstractReinforcement learning (RL)-based decoders in brain-machine interfaces (BMIs) interpret dynamic neural activity without patients' real limb movements. In conventional RL, the goal state is selected by the user or defined by the physics of the problem, and the decoder finds an optimal policy essentially by assigning credit over time, which is normally very time-consuming. However, BMI tasks require finding a good policy in very few trials, which impose a limit on the complexity of the tasks that can be learned before the animal quits. Therefore, this paper explores the possibility of letting the agent infer potential goals through actions over space with multiple objects, using the instantaneous reward to assign credit spatially. A previous method, attention-gated RL employs a multilayer perceptron trained with backpropagation, but it is prone to local minima entrapment. We propose a quantized attention-gated kernel RL (QAGKRL) to avoid the local minima adaptation in spatial credit assignment and sparsify the network topology. The experimental results show that the QAGKRL achieves higher successful rates and more stable performance, indicating its powerful decoding ability for more sophisticated BMI tasks as required in clinical applications. Yiwen Wang 0002, Hongbao Li, Yuxi Liao, Qiaosheng Zhang 0001, Shaomin Zhang, Xiaoxiang Zheng, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2016 | Predicting visual attention using gamma kernelsabstractSaliency measures are a popular way to predict visual attention. However, saliency is normally tested on sets of single resolution images that are unlike what the human vision system sees. We propose a new saliency measure based on convolving images with 2D gamma kernels which function as a comparison between a center and a surrounding neighborhood. The two parameters in the gamma kernel provide an ideal way to change the size of both the center and the surrounding neighborhood, which makes finding saliency at different scales simple and fast. We test the new saliency measure on both the CAT2000 database and the Toronto database and compare the results with other simple saliency methods. In addition, we test the methods on a foveated version of the Toronto database to test whether these methods perform well in a fixation system similar to the human vision system. Gamma saliency is shown to both perform better and compute faster than the competing methods in both the standard databases and the foveated version. Ryan Burt, Eder Santana, José C. Príncipe, Nina Thigpen, Andreas Keil |
ICASSP | 3 |
| 2016 | Information point set registration for shape recognitionabstractThis paper proposes a way of enhancing shape recognition through point set registration. Firstly, a modified version of shape context (SC) is developed, which is invariant to rigid transformation and flipping. With the point correspondence obtained by the modified SC, an affine transformation based on the maximum correntropy criterion (MCC) is performed on the query shape. This point set registration could be further refined by non-rigid morphing with the minimization of Cauchy-Schwarz divergence (DCS). Not only does this information theoretical learning (ITL) approach renders excellent registration result, but a new shape similarity measure can also be derived from the registration. José C. Príncipe, Bing Ouyang |
ICASSP | 2 |
| 2016 | Transient model of EEG using Gini Index-based matching pursuitabstractWe introduce a novel, transient model for the electroencephalogram (EEG) as the noisy addition of linear filters responding to trains of delta functions. We set the synthesis part as a parameter-tuning problem and obtain synthetic EEG-like data that visually resembles brain activity in the time and frequency domains. For the analysis counterpart, we use sparse approximation to decompose the signal in relevant events via Matching Pursuit. We improve this algorithm by incorporating the Gini Index as a stopping criteria; in this way, we promote sparse sources while, at the same time, eliminating one of the free parameters of Matching Pursuit. Results are presented using synthetic EEG and BCI competition data. Statistics of the model parameters are more informative and posses finer temporal resolution than classical methods such as Power Spectral Density (PSD) estimation. Carlos A. Loza, José C. Príncipe |
ICASSP | 2 |
| 2016 | Multiple adaptive kernel size KLMS for Beijing PM2.5 predictionabstractThe kernel least mean square (KLMS) algorithm is an efficient non-linear adaptive filter that operates in the reproducing kernel Hilbert space (RKHS). In realistic applications of system identification or time series prediction, there are usually multiple inputs that demand multiple kernels or kernel parameters. This paper proposes to use a tensor product kernel for KLMS that accommodates multiple inputs. Furthermore, instead of arbitrarily setting kernel parameters, appropriate kernel sizes can be chosen by a gradient descent based adaptive algorithm that minimizes the square of instant error, which helps KLMS to better capture the underlying system mechanism. Effectiveness of the proposed algorithm is shown by experiments conducted for both simulated dataset and an important real-world problem - Beijing PM2.5 prediction. Shujian Yu, Guibiao Xu, Badong Chen, José C. Príncipe |
IJCNN | 5 |
| 2016 | Generalized Correntropy Matching Pursuit: A novel, robust algorithm for sparse decompositionabstractWe introduce a novel variation on the well-known Matching Pursuit (MP) algorithm. In particular, the sparse approximation problem is solved in a greedy scheme using estimated higher-order statistics as similarity measures instead of the somehow limited second-order statistics that perform optimally only under Gaussian assumptions. This is conveyed via the generalized correntropy (GC) function instead of the cross-correlation approach usually utilized in stochastic random processes applications. Additionally, extra flexibility is achieved by the GC parameters that control the behavior of the induced metric. The result is the robust Generalized Correntropy Matching Pursuit (GCMP) algorithm. Furthermore, we present results on two different frameworks dealing with detection and sparse approximation and highlight the robustness of this method in the presence of high-tailed impulsive noise. Carlos A. Loza, José C. Príncipe |
IJCNN | 2 |
| 2016 | Information Theoretic-Learning auto-encoderabstractWe propose Information Theoretic-Learning (ITL) divergence measures for variational regularization of neural networks. We also explore ITL-regularized autoencoders as an alternative to variational autoencoding bayes, adversarial autoencoders and generative adversarial networks for randomly generating sample data without explicitly defining a paritition function. This paper also formalizes, generative moment matching networks under the ITL framework. Eder Santana, Matthew Emigh, José C. Príncipe |
IJCNN | 3 |
| 2016 | Density-dependent quantized kernel least mean squareabstractKernel least mean square is a simple and effective adaptive algorithm, but dragged by its unlimited growing network size. Many schemes have been proposed to reduce the network size, but few takes the distribution of the input data into account. Input data distribution is generally important in view of both model sparsification and generalization performance promotion. In this paper, we introduce an online density-dependent vector quantization scheme, which adopts a shrinkage threshold to adapt its output to the input data distribution. This scheme is then incorporated into the quantized kernel least mean square (QKLMS) to develop a density-dependent QKLMS (DQKLMS). Experiments on static function estimation and short-term chaotic time series prediction are presented to demonstrate the desirable performance of DQKLMS. Bao Xi, Lei Sun 0006, Badong Chen, Jianji Wang 0001, Nanning Zheng 0001, José C. Príncipe |
IJCNN | 6 |
| 2016 | Robust bounded logistic regression in the class imbalance problemabstractIn this paper, we propose to deal with the problems of logistic regression with outliers and class imbalance, which are common in a wide range of practical applications. The robust bounded logistic regression with different error costs is developed to reduce the combined influence of outliers and class imbalance. First, inspired by the Correntropy induced loss function, we develop the bounded logistic loss function which is a monotonic, bounded and nonconvex loss and thus robust to outliers. With the bounded logistic loss, we construct a new robust logistic regression. Second, under the principle of cost-sensitive learning, we assign different error costs for different classes in order to reduce the sensitiveness of the new robust logistic regression to class imbalance. Using the half-quadratic optimization method, it is easy to optimize the proposed logistic regression model. Experimental results demonstrate that our proposed method improves the performance of logistic regression on the datasets with outliers and class imbalance. Guibiao Xu, Bao-Gang Hu, José C. Príncipe |
IJCNN | 3 |
| 2016 | A reconfigurable parallel FPGA accelerator for the adapt-then-combine diffusion LMS algorithmabstractThe combination of diffusion strategies and least-mean-square (LMS) algorithm provides many advantages for adaptive-filter to solve distributed optimization, estimation and inference problems. However, suffering from high computation complexity, software implementation of diffusion LMS algorithm is unsuitable for real-time and portable applications. In order to extend its availability, we design a reconfigurable parallel FPG accelerator by exploring multiple dimensions of parallelism, including: parallel execution of agents state updating, data combining, data training and multi-stages pipeline to speedup the execution time. The accelerator for networks with various number of agents and different input dimensions is implemented. Results demonstrate that, it can achieve a speedup of three orders of magnitude at 100Mhz compared with C implementation for a 32-nodes network with 16-dimensional input-data. Qihang Yu, Badong Chen, José C. Príncipe, Nanning Zheng 0001, Pengju Ren |
ISCAS | 4 |
| 2016 | Empirical data analysis: A new tool for data analyticsabstractIn this paper, a novel empirical data analysis approach (abbreviated as EDA) is introduced which is entirely data-driven and free from restricting assumptions and pre-defined problem- or user-specific parameters and thresholds. It is well known that the traditional probability theory is restricted by strong prior assumptions which are often impractical and do not hold in real problems. Machine learning methods, on the other hand, are closer to the real problems but they usually rely on problem- or user-specific parameters or thresholds making it rather art than science. In this paper we introduce a theoretically sound yet practically unrestricted and widely applicable approach that is based on the density in the data space. Since the data may have exactly the same value multiple times we distinguish between the data points and unique locations in the data space. The number of data points, k is larger or equal to the number of unique locations, l and at least one data point occupies each unique location. The number of different data points that have exactly the same location in the data space (equal value), f can be seen as frequency. Through the combination of the spatial density and the frequency of occurrence of discrete data points, a new concept called multimodal typicality, τMMis proposed in this paper. It offers a closed analytical form that represents ensemble properties derived entirely from the empirical observations of data. Moreover, it is very close (yet different) from the histograms, from the probability density function (pdf) as well as from fuzzy set membership functions. Remarkably, there is no need to perform complicated pre-processing like clustering to get the multimodal representation. Moreover, the closed form for the case of Euclidean, Mahalanobis type of distance as well as some other forms (e.g. cosine-based dissimilarity) can be expressed recursively making it applicable to data streams and online algorithms. Inference/estimation of the typicality of data points that were not present in the data so far can be made. This new concept allows to rethink the very foundations of statistical and machine learning as well as to develop a series of anomaly detection, clustering, classification, prediction, control and other algorithms. Plamen Angelov 0001, Xiaowei Gu 0001, Dmitry Kangin, José C. Príncipe |
SMC | 4 |
| 2016 | Correntropy induced joint power and admission control algorithm for dense small cell networkabstractThe authors consider the joint admission and power control problem in a dense small cell network, which contains multiple interference links. The goal is to mainly maximise the number of the admitted links, and at the same time minimise the transmit power. The authors formulate the admission control and power control problem as a joint optimisation problem, which is however non‐deterministic polynomial hard (NP‐hard). Such NP‐hard problem can be relaxed to a p ‐norm problem (0 < p < 1) by using the correntropy induced metric. The correntropy is a novel non‐linear similarity measure, which has been successfully used in the robust and spares signal processing, especially when the data contain large outliers. Thus, in this work the authors propose a new correntropy induced joint power and admission control algorithm. To achieve a faster convergence speed, the authors also propose an adaptive kernel size method, in which the kernel size is determined by the error so that the convergence speed is the fastest during the iterations. Simulation results show that the proposed approach can achieve much better results than the existing works. Zhirong Luan, Hua Qu, Jihong Zhao 0001, Badong Chen, José C. Príncipe |
IET Commun. | 5 |
| 2016 | Kernel least mean square with adaptive kernel size
Badong Chen, Junli Liang, Nanning Zheng 0001, José C. Príncipe |
Neurocomputing | 4 |
| 2016 | Efficient and robust deep learning with Correntropy-induced loss function
Liangjun Chen, Hua Qu, Jihong Zhao 0001, Badong Chen, José C. Príncipe |
Neural Comput. Appl. | 5 |
| 2016 | Neurally Encoding Time for Olfactory NavigationabstractAccurately encoding time is one of the fundamental challenges faced by the nervous system in mediating behavior. We recently reported that some animals have a specialized population of rhythmically active neurons in their olfactory organs with the potential to peripherally encode temporal information about odor encounters. If these neurons do indeed encode the timing of odor arrivals, it should be possible to demonstrate that this capacity has some functional significance. Here we show how this sensory input can profoundly influence an animal's ability to locate the source of odor cues in realistic turbulent environments-a common task faced by species that rely on olfactory cues for navigation. Using detailed data from a turbulent plume created in the laboratory, we reconstruct the spatiotemporal behavior of a real odor field. We use recurrence theory to show that information about position relative to the source of the odor plume is embedded in the timing between odor pulses. Then, using a parameterized computational model, we show how an animal can use populations of rhythmically active neurons to capture and encode this temporal information in real time, and use it to efficiently navigate to an odor source. Our results demonstrate that the capacity to accurately encode temporal information about sensory cues may be crucial for efficient olfactory navigation. More generally, our results suggest a mechanism for extracting and encoding temporal information from the sensory environment that could have broad utility for neural information processing. In Jun Park, Andrew M. Hein, Yuriy V. Bobkov, Matthew A. Reidenbach, Barry W. Ache, José C. Príncipe |
PLoS Comput. Biol. | 6 |
| 2016 | Fault detection via recurrence time statistics and one-class classification
David Martínez-Rego, Oscar Fontenla-Romero, Amparo Alonso-Betanzos, José C. Príncipe |
Pattern Recognit. Lett. | 4 |
| 2016 | Reinforcement Learning in Video Games Using Nearest Neighbor Interpolation and Metric LearningabstractReinforcement learning (RL) has had mixed success when applied to games. Large state spaces and the curse of dimensionality have limited the ability for RL techniques to learn to play complex games in a reasonable length of time. We discuss a modification of Q-learning to use nearest neighbor states to exploit previous experience in the early stages of learning. A weighting on the state features is learned using metric learning techniques, such that neighboring states represent similar game situations. Our method is tested on the arcade game Frogger, and it is shown that some of the effects of the curse of dimensionality can be mitigated. Matthew Emigh, Evan Kriminger, Austin J. Brockmeier, José C. Príncipe, Panos M. Pardalos |
IEEE Trans. Comput. Intell. AI Games | 4 |
| 2016 | Kernel Learning for Dynamic Texture SynthesisabstractDynamic textures (DTs) that represent moving scenes such as flames, smoke, and waves, exhibit fixed dynamics within a period of time and have been successfully modeled using linear dynamic systems (LDS). In this paper, we show that the widely used LDS model can be approximated using a principal component regression (PCR) model with the main advantage of simplicity. Furthermore, to capture the nonlinearity of training frames, we extend traditional PCR to its kernelized version and introduce kernel principal component regression (KPCR) to model and synthesize DTs. To ensure algorithm stability, we remove the standard state model and directly apply the quantized kernel least mean squares algorithm from signal processing domain to approximate the performance achieved with KPCR. We term this improvement kernel adaptive dynamic texture synthesis (KADTS), which also has the benefits of computational and memory efficiency. These advantages make KADTS ideally suited for real-world applications, since the majority of electronic devices, including cell phones and laptops, suffer from limited memory and real-time constraints. We demonstrate, via both theoretical and experimental analyses, the connections between DT synthesis using KPCR and KADTS with a regularization network theory. We also show the superiority of our proposed algorithms for DT synthesis compared with other dynamic system-based benchmarks. MATLAB code is available from our project homepage http://bmal.hust.edu.cn/project/dts.html. Xinge You, Weigang Guo, Shujian Yu, Kan Li 0002, José C. Príncipe, Dacheng Tao |
IEEE Trans. Image Process. | 5 |
| 2016 | The Kernel Adaptive Autoregressive-Moving-Average AlgorithmabstractIn this paper, we present a novel kernel adaptive recurrent filtering algorithm based on the autoregressive-moving-average (ARMA) model, which is trained with recurrent stochastic gradient descent in the reproducing kernel Hilbert spaces. This kernelized recurrent system, the kernel adaptive ARMA (KAARMA) algorithm, brings together the theories of adaptive signal processing and recurrent neural networks (RNNs), extending the current theory of kernel adaptive filtering (KAF) using the representer theorem to include feedback. Compared with classical feedforward KAF methods, the KAARMA algorithm provides general nonlinear solutions for complex dynamical systems in a state-space representation, with a deferred teacher signal, by propagating forward the hidden states. We demonstrate its capabilities to provide exact solutions with compact structures by solving a set of benchmark nondeterministic polynomial-complete problems involving grammatical inference. Simulation results show that the KAARMA algorithm outperforms equivalent input-space recurrent architectures using first- and second-order RNNs, demonstrating its potential as an effective learning solution for the identification and synthesis of deterministic finite automata. Kan Li 0002, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | Explicit versus implicit source estimation for blind multiple input single output system identificationabstractSparsely-activated time series are found in many physical systems. In these cases, the signals can be approximated by convolution of sparse sources with a set of shift-invariant filters. When there is access to only one sensor, such that there is a single observation signal, identifying the source signals appears to be an ill-posed problem, but for very sparse sources it is still possible to learn the system. We discuss analysis techniques for sparsely activated signals, which retrieve sparse sources given the filters, and identify conditions when algorithms based on independent component analysis (ICA) and sparse coding can blindly estimate filters from a single noisy time-series. Many qualitative results have been made for learning shift-invariant bases on natural signals, but for a thorough understanding of the effect of sparsity, we quantitatively analyze results on synthetic examples, comparing how ICA and shift-invariant sparse coding approaches perform for multiple-source blind system identification. Austin J. Brockmeier, José C. Príncipe |
ICASSP | 2 |
| 2015 | Sparsity aware minimum error entropy algorithmsabstractSparse estimation has received a lot of attention due to its broad applicability. In sparse channel estimat ion, the parameter vector with sparsity characteristic can be well estimated from noisy measurements through sparse adaptive filters. In previous studies, most works use the mean square error (MSE) based cost to develop sparse filters, which is rat ional under the assumption of Gaussian distributions. However, Gaussian assumption does not always hold in real-world environments. To address this issue, we incorporate in this work l1-norm and reweighted l1-norm into the minimum error entropy (MEE) criterion to develop new sparse adaptive filters, which may perform much better than the MSE based methods especially in non-Gaussian situations, since the error entropy can capture higher-order statistics of the errors . Furthermore, a new approximator of l0-norm based on the Correntropy Induced Metric (CIM) is also used as a sparsity penalty term (SPT). Simulation results show the excellent performance of the proposed algorithms. Wentao Ma 0007, Hua Qu, Jihong Zhao 0001, Badong Chen, José C. Príncipe |
ICASSP | 5 |
| 2015 | Learning joint features for color and depth images with Convolutional Neural Networks for object classificationabstractIn this paper we investigate the advantages of learning representations of color plus depth images (Red-Blue-Green-Depth, RGB-D) over color only images (RGB) for computer vision. Specifically, we investigate the advantages on the task of object recognition. For this purpose, we applied the state-of-art deep convolutional neural networks (CNN) for classification of images on the RGB-D dataset published by (Bo et al., 2011). We show that this approach provides better results than those that use separate features for color and depth. Also, we probe the resulting CNN to gain intuition about how filters for depth and color channels iterate to generate useful features. Eder Santana, Karl P. Dockendorf, José C. Príncipe |
ICASSP | 3 |
| 2015 | Cognitive Workload Discrimination in Flight Simulation Task Using a Generalized Measure of Association
Zhongxiang Dai, José C. Príncipe, Anastasios Bezerianos, Nitish V. Thakor |
ICONIP (3) | 2 |
| 2015 | Group feature selection in image classification with multiple kernel learningabstractClassification of large amount of images calls for diverse types of features, but employing all possible feature types will create unnecessary computation burden, and may result in reduced classification accuracy. Selecting feature vectors individually is not a feasible solution in this scenario due to the high amount of feature vectors needed for reasonable performance. Instead, this paper proposes a measure that effectively evaluates the relative significance of a feature group, employing the minimum redundancy maximum relevance (mRMR) feature selection. Multiple kernel learning (MKL) is used for combining different feature types in classification, which implicitly also serves an alternative way for weighing the feature groups' importance. Results show the proposed group feature selection better reflects a feature type's importance, and improve upon MKL performance. This study also finds that the convolutional neural network (CNN) features have the best discriminative power among all features, but it is still possible to improve classification accuracy with other well-designed features. José C. Príncipe, Bing Ouyang |
IJCNN | 2 |
| 2015 | Exponential C-Loss for data fittingabstractAs a robust measure of similarity, C-Loss can be successfully used for data fitting such as regression and classification, especially when data contain large outliers. In this paper, we propose a modified C-Loss function, called exponential C-Loss (EC-Loss), which is defined as an exponential function of the C-Loss. The EC-Loss inherits the robustness and smoothness of the C-Loss but may have a better performance surface that favors the usage of a gradient-based learning algorithm, particularly at a region far from the optimal solution. In order to avoid the flatness of the performance surface near the optimal solution and obtain a fast convergence speed during the overall adaptation process, we also propose a novel switching strategy between C-Loss and EC-Loss. A simple simulation example is presented to demonstrate the performance surface and desirable performance of the new method. Badong Chen, Ren Wang 0010, Nanning Zheng 0001, José C. Príncipe |
IJCNN | 4 |
| 2015 | On initial convergence behavior of the kernel least mean square algorithmabstractThe mean square convergence of the kernel least mean square (KLMS) algorithm has been studied in a recent paper [B. Chen, S. Zhao, P. Zhu, J. C. Principe, Mean square convergence analysis of the kernel least mean square algorithm, Signal Processing, vol. 92, pp. 2624-2632, 2012]. In this paper, we continue this study and focus mainly on the initial convergence behavior. Two measures of the convergence performance are considered, namely the weight error power (WEP) and excess mean square error (EMSE). The analytical expressions of the initial decreases of the WEP and EMSE are derived, and several interesting facts about the initial convergence are presented. An illustration example is given to support our observation. Badong Chen, Ren Wang 0010, Nanning Zheng 0001, José C. Príncipe |
IJCNN | 4 |
| 2015 | Linear discriminant analysis with an information divergence criterionabstractLinear discriminant analysis seeks to find a one-dimensional projection of a dataset to alleviate the problems associated with classifying high-dimensional data. The earliest methods, based on second-order statistics often fail on multimodal datasets. Information-theoretic criteria do not suffer in such cases, and allow for projections to spaces higher than one dimension and with multiple classes. These approaches are based on maximizing mutual information between the projected data and the labels. However, mutual information is computationally demanding and vulnerable to datasets with class imbalance. In this paper we propose an information-theoretic criterion for learning discriminants based on the Euclidean distance divergence between classes. This objective more directly seeks projections which separate classes and performs well in the midst of class imbalance. We demonstrate the effectiveness on real datasets, and provide extensions to the multi-class and multi-dimension cases. Matthew Emigh, Evan Kriminger, José C. Príncipe |
IJCNN | 3 |
| 2015 | Directed generalized measure of association: A data driven approach towards causal inferenceabstractIn this paper we propose a new statistical concept called directed generalized measure of association (dGMA) to quantify the amount of association transferred between subsystems in a system evolving in time. This paper is an improvement of the previously established method called generalized measure of association (GMA) by taking the conditional dependence between time-delay representation of subsystems into account. Directed-GMA is a rank-based pair-wise measure which can deal with dynamic data sets. This is done by calculating the rank permutation in a conditional scheme under the framework of conditional causality. In this paper we present a bivariate case and assume that the cause-effect relationship is one directional, e.g. there is no feedback. The preliminary results on the synthetic data sets reveal that the proposed method is able to extract nonlinear causal relationship, which cannot be extracted by traditional approaches. To further assess the performance we compared the results with another rank-based method called symbolic transfer entropy (STE). Our approach can be a promising tool to infer causal relationship in complex systems, e. g. human brain, to reveal their underlying effective connectivity. Mehrnaz Khodam Hazrati, Andreas Keil, José C. Príncipe |
IJCNN | 3 |
| 2015 | Effective insect recognition using a stacked autoencoder with maximum correntropy criterionabstractThroughout the history, insects had been intimately connected to humanity, in both positive and negative ways. Insects play an important part in crop pollination, on the other hand, some of them spread diseases that kill millions of people every year. Effective control of harmful insects while having little impact to beneficial insects and environment is extremely important. Recently, an intelligent trap that uses laser sensors was presented to control the population of target insects. The device could record and analyze sensor signals when an insect passes through the trap and make quick decisions whether to catch it or not. The effectiveness of the trap relies on the correct choice of classification algorithm to perform the insect detection. In this paper, we propose to use a deep neural network with maximum correntropy criterion (MCC) for reliable classification of insects in real-time. Experimental results show that, deep networks are effective for learning stable features from brief insect passage signals. By replacing the mean square error cost with MCC, the robustness of autoencoders against noise is improved significantly and robust features could be learned. The method is tested on five species of insects and a total of 5325 passages. High classification accuracy of 92.1% is achieved. Compared with previously applied methods, better classification performance is obtained using only 10% of the computation time. Therefore, our method is efficient and reliable for online insect detection. Goktug T. Cinar, Vinícius M. A. de Souza, Gustavo Batista, Yueming Wang 0001, José C. Príncipe |
IJCNN | 6 |
| 2015 | Parallel flow in Deep Predictive Coding NetworksabstractThis paper proposes a cognitive architecture for sensory processing of multimodal data. The cognitive architectures, referred to as Deep Predictive Coding Networks (DPCN) were first used to model video streams. Here we use DPCNs with two input sources, for example: video and speech recordings. We train DPCNs as generative models of both sensors. Since we constrain the network to have a single hidden code for both inputs, we name the proposed architecture as Multimodal DPCN (MDPCN). Experimental results show that the “parallel” flow between the two sensory modes increases the interclass separability achieved by unsupervised clustering. We validate the proposed method with a multimodal classification task using part of the VIDTIMIT dataset. Eder Santana, Goktug T. Cinar, José C. Príncipe |
IJCNN | 3 |
| 2015 | Mixed generative and supervised learning modes in Deep Predictive Coding NetworksabstractIn this paper we propose a modification of the Cognitive Architectures for Sensory Processing proposed by Chalasani and Principe. Here we keep the bottom-up data representation through generative models as before, but propose a top-down flow based on backpropagation of gradients for recognition. By treating the bottom-up procedure involved in the inference step as a recursive neural network, we show that supervised learning can be used in conjunction with other layers commonly used for Deep Learning. Also, this allows us to learn models that incorporate at the same time data classification and statistical modeling of the input. We show that this combination provides classification results that are robust to input noise. Eder Santana, José C. Príncipe |
IJCNN | 2 |
| 2015 | A variable step-size adaptive algorithm under maximum correntropy criterionabstractCorrentropy, a novel localized similarity measure defined in kernel space, has been successfully used as a cost function in adaptive system training. The adaptive algorithms under the maximum correntropy criterion (MCC) have been shown to be robust to impulsive non-Gaussian noises. However, they may converge slowly especially at a region far from the optimal solution. In this paper, we propose a new MCC algorithm with a variable step-size (VSS) called the VSS-MCC algorithm, which may achieve a much faster convergence speed while maintaining similar steady-state performance. In the new algorithm, the step-size is updated based on an approximation for the curvature of performance surface. Simulation results demonstrate the superior performance of VSS-MCC compared with the original MCC algorithm. Ren Wang 0010, Badong Chen, Nanning Zheng 0001, José C. Príncipe |
IJCNN | 4 |
| 2015 | A switch kernel width method of correntropy for channel estimationabstractCorrentropy has been successfully applied in non-Gaussian signal processing, but the superior performance achieved is depends on appropriate selection of the kernel width. How to select a proper kernel width is a crucial problem in correntropy applications. In this paper, we propose an adaptive algorithm to update the kernel width, which is set at a maximum between the absolute value of instantaneous error divided by square root of 2 and a predetermined kernel width. The new algorithm involves no extra free parameters and keeps the simplicity and robustness of the original maximum correntropy criterion (MCC) algorithm. Simulation results confirm that the proposed algorithm can achieve excellent performance in channel estimation under impulsive noises. Jihong Zhao 0001, Hua Qu, Badong Chen, José C. Príncipe |
IJCNN | 5 |
| 2015 | Performance evaluation of the correntropy coefficient in automatic modulation classification
Aluisio I. Rêgo Fontes, Allan de Medeiros Martins, Luiz Felipe Q. Silveira, José C. Príncipe |
Expert Syst. Appl. | 4 |
| 2015 | Self-organizing maps with information theoretic learning
Rakesh Chalasani, José C. Príncipe |
Neurocomputing | 2 |
| 2015 | Special Issue on Advances in Self-organizing Maps
Pablo A. Estévez, José C. Príncipe |
Neurocomputing | 2 |
| 2015 | An algorithm based on non-squared sum of the errors
Cristiane Cristina Sousa da Silva, Allan Kardec Barros, Ewaldo E. C. Santana, Marcos A. F. de Araújo, Marcus V. de S. Lopes, João Viana da Fonseca Neto, José C. Príncipe |
Signal Process. | 7 |
| 2015 | Convergence of a Fixed-Point Algorithm under Maximum Correntropy CriterionabstractThe maximum correntropy criterion (MCC) has received increasing attention in signal processing and machine learning due to its robustness against outliers (or impulsive noises). Some gradient based adaptive filtering algorithms under MCC have been developed and available for practical use. The fixed-point algorithms under MCC are, however, seldom studied. In particular, too little attention has been paid to the convergence issue of the fixed-point MCC algorithms. In this letter, we will study this problem and give a sufficient condition to guarantee the convergence of a fixed-point MCC algorithm. Badong Chen, Jianji Wang 0001, Haiquan Zhao 0001, Nanning Zheng 0001, José C. Príncipe |
IEEE Signal Process. Lett. | 5 |
| 2015 | Downscaling Satellite-Based Soil Moisture in Heterogeneous Regions Using High-Resolution Remote Sensing Products and Information Theory: A Synthetic StudyabstractIn this study, a novel methodology based upon the information-theoretic measures of entropy and mutual information was implemented to downscale soil moisture (SM) observations from 10 km to 1 km. It included a transformation function that related auxiliary remotely sensed (RS) products at high resolution to in situ SM observations to obtain first estimates of SM at 1 km and merging this estimate with SM at coarse resolutions through Principle of Relevant Information (PRI). The PRI-based estimates were evaluated using synthetic observations in NC Florida for heterogeneous agricultural land covers (LC), with two growing seasons of sweet corn and one of cotton, annually. The cumulative density function showed an overall error in SM of <; 0.03 cubic meter/cubic meter in the region, with a confidence interval of 95% during the simulation period. The PRI estimates at 1 km were also compared with those from the method based upon Universal Triangle (UT). The spatially averaged root mean square error (RMSE) aggregated over the vegetative LC were 0.01 cubic meter/cubic meter and 0.15 cubic meter/cubic meter using the PRI and UT methods, respectively. The RMSE for downscaled estimates using the UT method increased to 0.28 cubic meter/cubic meter when Laplacian errors are used, while the corresponding RMSE for the PRI remains the same for both Laplacian or Gaussian errors. The Kullback-Liebler divergence (KLD) for estimates using PRI is about 50% lower than those using the method based upon UT indicating that the probability density function (PDF) of the PRI estimate is closer to PDF of the true SM, than the UT method. Subit Chakrabarti, Tara Bongiovanni, Jasmeet Judge, Karthik Nagarajan, José C. Príncipe |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2015 | Measures of Entropy From Data Using Infinitely Divisible KernelsabstractInformation theory provides principled ways to analyze different inference and learning problems, such as hypothesis testing, clustering, dimensionality reduction, classification, and so forth. However, the use of information theoretic quantities as test statistics, that is, as quantities obtained from empirical data, poses a challenging estimation problem that often leads to strong simplifications, such as Gaussian models, or the use of plug in density estimators that are restricted to certain representation of the data. In this paper, a framework to nonparametrically obtain measures of entropy directly from data using operators in reproducing kernel Hilbert spaces defined by infinitely divisible kernels is presented. The entropy functionals, which bear resemblance with quantum entropies, are defined on positive definite matrices and satisfy similar axioms to those of Renyi's definition of entropy. Convergence of the proposed estimators follows from concentration results on the difference between the ordered spectrum of the Gram matrices and the integral operators associated to the population quantities. In this way, capitalizing on both the axiomatic definition of entropy and on the representation power of positive definite kernels, the proposed measure of entropy avoids the estimation of the probability distribution underlying the data. Moreover, estimators of kernel-based conditional entropy and mutual information are also defined. Numerical experiments on independence tests compare favorably with state-of-the-art. Luis Gonzalo Sánchez Giraldo, Murali Rao, José C. Príncipe |
IEEE Trans. Inf. Theory | 3 |
| 2015 | Context Dependent Encoding Using Convolutional Dynamic NetworksabstractPerception of sensory signals is strongly influenced by their context, both in space and time. In this paper, we propose a novel hierarchical model, called convolutional dynamic networks, that effectively utilizes this contextual information, while inferring the representations of the visual inputs. We build this model based on a predictive coding framework and use the idea of empirical priors to incorporate recurrent and top-down connections. These connections endow the model with contextual information coming from temporal as well as abstract knowledge from higher layers. To perform inference efficiently in this hierarchical model, we rely on a novel scheme based on a smoothing proximal gradient method. When trained on unlabeled video sequences, the model learns a hierarchy of stable attractors, representing low-level to high-level parts of the objects. We demonstrate that the model effectively utilizes contextual information to produce robust and stable representations for object recognition in video sequences, even in case of highly corrupted inputs. Rakesh Chalasani, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2014 | Functional relevant multichannel kernel adaptive filter for human activity analysisabstractA multichannel kernel adaptive filtering framework is presented that highlights relevant channels for the task of analyzing Motion Capture (MoCap) data. Functional relevance analysis is performed over input multichannel data by computing the pair-wise channel similarities to describe the main behavior of the considered applications. Particularly, the well-known Kernel Least Mean Square filter is enhanced using a correntropy-based similarity criterion between channel pairs. Besides, two sparseness criteria are studied to extract a sample subset that constructs a learning model displaying a good trade-off between filter complexity and accuracy. The proposed approach allows devising complex relationship among multi-channel time-series, revealing dependencies among the channels and the process time-structure. The method is tested in a well-known MoCap data set. Results show that our framework is an adequate alternative for finding functional relevance amongst multi-channel time-series. Andrés Marino Álvarez-Meza, Germán Castellanos-Domínguez, José C. Príncipe |
ICASSP | 3 |
| 2014 | Projentropy: Using entropy to optimize spatial projectionsabstractMethods for hypothesis testing on zero-mean vector-valued signals often rely on a Gaussian assumption, where the second-order statistics of the observed sample are sufficient statistics of the conditional distribution. This yields fast and simple tests, but by using information-theoretic statistics one can relax the Gaussian assumption. We propose using Rényi's quadratic entropy as an alternative to the covariance and show how a linear projection can be optimized to maximize the difference between the conditional entropies. In addition, if the observed sample is actually a window of a multivariate time-series, then the temporal structure can be exploited using the generalized auto-correlation function, correntropy, of the projected sample. This both reduces the computational complexity and increases the performance. These tests can be applied for decoding the brain state from electroencephalogram (EEG) recordings. Preliminary results are demonstrated on a brain-computer interface competition dataset. On unfiltered signals, the projections optimized with the entropy-based statistic perform better than those of common spatial pattern (CSP) algorithm in terms of classification performance. Austin J. Brockmeier, Eder Santana, Luis Gonzalo Sánchez Giraldo, José C. Príncipe |
ICASSP | 4 |
| 2014 | Dynamic sparse coding with smoothing proximal gradient methodabstractIn this work we focus on the problem of estimating time-varying sparse signals from a sequence of under-sampled observations. We formulate this problem as estimating hidden states in a dynamic model and exploit the underlying temporal structure to find a more accurate solution, particularly when the information in the observations is at scarce. We propose an optimization procedure based on smoothing proximal gradient method to estimate these hidden states. We show that the proposed model is efficient and more robust to the noise in the system. Rakesh Chalasani, José C. Príncipe |
ICASSP | 2 |
| 2014 | Feature selection based on survival Cauchy-Schwartz mutual informationabstractFeature selection techniques play a crucial role in machine learning tasks such as regression and classification. Many filter methods of feature selection are based on the mutual information (e.g. MIFS, MIFS-U, NMIFS, and mRMR methods). In this work, a new mutual information is defined based on the cross survival information potential (CSIP) and Cauchy-Schwartz divergence (CSD), called the survival Cauchy-Schwartz mutual information (SCS-MI). We apply this new mutual information to select an informative subset of features for a SVM classifier. Experimental results illustrate the desirable performance of the new method. Badong Chen, Hua Qu, Jihong Zhao 0001, Nanning Zheng 0001, José C. Príncipe |
ICASSP | 6 |
| 2014 | Sparse kernel recursive least squares using L1 regularization and a fixed-point sub-iterationabstractA new kernel adaptive filtering (KAF) algorithm, namely the sparse kernel recursive least squares (SKRLS), is derived by adding a ℓ1-norm penalty on the center coefficients to the least squares (LS) cost (i.e. the sum of the squared errors). In each iteration, the center coefficients are updated by a fixed-point sub-iteration. Compared with the original KRLS algorithm, the proposed algorithm can produce a much sparser network, in which many coefficients are negligibly small. A much more compact structure can thus be achieved by pruning these negligible centers. Simulation results show that the SKRLS performs very well, yielding a very sparse network while preserving a desirable performance. Badong Chen, Nanning Zheng 0001, José C. Príncipe |
ICASSP | 3 |
| 2014 | Clustering of time series using a hierarchical linear dynamical systemabstractThe auditory cortex in the brain does effortlessly a better job of extracting information from the acoustic world than our current generation of signal processing algorithms. Abstracting the principles of the auditory cortex, the proposed architecture is based on Kalman filters with hierarchically coupled state models that stabilize the input dynamics and provide a representation space. This approach extracts information from the input and self-organizes it in the higher layers leading to an algorithm capable of clustering time series in an unsupervised manner. An important characteristic of the methodology is that it is adaptive and self-organizing, i.e. previous exposure to the acoustic input is the only requirement for learning and recognition, so there is no need of selecting the number of clusters. Goktug T. Cinar, José C. Príncipe |
ICASSP | 2 |
| 2014 | Online Nonlinear Granger Causality Detection by Quantized Kernel Least Mean Square
Badong Chen, Zejian Yuan, Nanning Zheng 0001, Andreas Keil, José C. Príncipe |
ICONIP (2) | 6 |
| 2014 | Correntropy kernel temporal differences for reinforcement learning brain machine interfacesabstractThis paper introduces a novel temporal difference algorithm to estimate a value function in reinforcement learning. This is a kernel adaptive system using a robust cost function called correntropy. We call this system correntropy kernel temporal differences (CKTD). This algorithm is integrated with Q-learning to find a proper policy (Q-learning via correntropy kernel temporal differences). The proposed method was tested with a synthetic problem, and its robustness under a changing policy was quantified. The same algorithm was applied to the decoding of a monkey's neural states in a reinforcement learning brain machine interface (RLBMI) in a center-out reaching task. The results showed the potential advantage of the proposed algorithm in the RLBMI framework. Jihye Bae, Luis Gonzalo Sánchez Giraldo, José C. Príncipe, Joseph T. Francis |
IJCNN | 3 |
| 2014 | Pitch estimation using non-negative matrix factorizationabstractThe problem of pitch detection consists of estimating the dominant frequency present in a certain time window. This paper demonstrates and analyzes the use of a non-negative matrix factorization technique with a frequency basis formed with a correntropy kernel. This offers the advantage that the frequency basis is adaptable, allowing the matrix factorization to fit the data precisely, as well as including a dictionary specifically to account for noise. Using non-negative matrix factorization also allows an increase in dimensionality, which increases the frequency resolution of the algorithm. The method is tested on a database of trumpet notes and compared to other current methods, improving on their performance for noisy signals. Ryan Burt, Goktug T. Cinar, José C. Príncipe |
IJCNN | 3 |
| 2014 | Trimmed affine projection algorithmsabstractThe least trimmed squares (LTS) estimator is a robust estimator as it can avoid undue influence from outliers. The exact solution of the LTS estimation is however hard to And and if the number of data is large then the method is unfeasible. In this work, we apply the LTS criterion to adaptive Altering and develop the trimmed affine projection algorithm (TAPA) and kernel trimmed affine projection algorithm (KTAPA). The proposed adaptive algorithms are very robust to outliers and have low computational complexity. Simulation results conflrm their excellent and robust performance. Badong Chen, Hua Qu, Nanning Zheng 0001, José C. Príncipe |
IJCNN | 6 |
| 2014 | Hierarchical Linear Dynamical Systems: A new model for clustering of time seriesabstractThe auditory cortex in the brain does effortlessly a better job of extracting information from the acoustic world than our current generation of signal processing algorithms. The proposed architecture, Hierarchical Linear Dynamical System (HLDS), is based on Kalman filters with hierarchically coupled state models that stabilize the input dynamics and provide a representation space. This approach extracts information from the input and self-organizes it in the higher layers leading to an algorithm capable of clustering time series in an unsupervised manner. In this paper we further investigate the properties of HLDS, demonstrate its performance on music rather than isolated notes and propose the time domain implementation to overcome one of its current bottlenecks. Goktug T. Cinar, Carlos A. Loza, José C. Príncipe |
IJCNN | 3 |
| 2014 | Quantized mixture kernel least mean squareabstractUse of multiple kernels in the conventional kernel algorithms is gaining much popularity as it addresses the kernel selection problem as well as improves the performance. Kernel least mean square (KLMS) has been extended to multiple kernels recently using different approaches, one of which is mixture kernel least mean square (MxKLMS). Although this method addresses the kernel selection problem, and improves the performance, it suffers from a problem of linearly growing dictionary like in KLMS. In this paper, we present the quantized MxKLMS (QMxKLMS) algorithm to achieve sub-linear growth in dictionary. This method quantizes the input space based on the conventional criteria using Euclidean distance in input space as well as a new criteria using Euclidean distance in RKHS induced by the sum kernel. The empirical results suggest that QMxKLMS using the latter metric is suitable in a non-stationary environment with abruptly changing modes as they are able to utilize the information regarding the relative importance of kernels. Moreover, the QMxKLMS using both metrics are compared with the QKLMS and the existing multi-kernel methods MKLMS and MKNLMS-CS, showing an improved performance over these methods. Rosha Pokharel, Sohan Seth, José C. Príncipe |
IJCNN | 3 |
| 2014 | An asymmetric stagewise least square loss function for imbalanced classificationabstractIn this paper, we present an asymmetric stagewise least square (ASLS) loss function for imbalanced classification. While keeping all the advantages of the stagewise least square (SLS) loss function, such as, better robustness, computational efficiency and sparseness, the ASLS loss extends the SLS loss by adding another two parameters, namely, ramp coefficient and margin coefficient. Therefore, asymmetric ramps and margins can be formed which makes the ASLS loss be more flexible and appropriate for processing class imbalance problems. A reduced kernel classifier of the ASLS loss is also developed which only uses a small part of the dataset to generate an efficient nonlinear classifier. Experimental results confirm the effectiveness of the ASLS loss in imbalanced classification. Guibiao Xu, Bao-Gang Hu, José C. Príncipe |
IJCNN | 3 |
| 2014 | Neural Decoding with Kernel-Based Metric LearningabstractIn studies of the nervous system, the choice of metric for the neural responses is a pivotal assumption. For instance, a well-suited distance metric enables us to gauge the similarity of neural responses to various stimuli and assess the variability of responses to a repeated stimulus-exploratory steps in understanding how the stimuli are encoded neurally. Here we introduce an approach where the metric is tuned for a particular neural decoding task. Neural spike train metrics have been used to quantify the information content carried by the timing of action potentials. While a number of metrics for individual neurons exist, a method to optimally combine single-neuron metrics into multineuron, or population-based, metrics is lacking. We pose the problem of optimizing multineuron metrics and other metrics using centered alignment, a kernel-based dependence measure. The approach is demonstrated on invasively recorded neural data consisting of both spike trains and local field potentials. The experimental paradigm consists of decoding the location of tactile stimulation on the forepaws of anesthetized rats. We show that the optimized metrics highlight the distinguishing dimensions of the neural response, significantly increase the decoding accuracy, and improve nonlinear dimensionality reduction methods for exploratory neural analysis. Austin J. Brockmeier, John S. Choi, Evan Kriminger, Joseph T. Francis, José C. Príncipe |
Neural Comput. | 5 |
| 2014 | Information Theoretic Shape MatchingabstractIn this paper, we describe two related algorithms that provide both rigid and non-rigid point set registration with different computational complexity and accuracy. The first algorithm utilizes a nonlinear similarity measure known as correntropy. The measure combines second and high order moments in its decision statistic showing improvements especially in the presence of impulsive noise. The algorithm assumes that the correspondence between the point sets is known, which is determined with the surprise metric. The second algorithm mitigates the need to establish a correspondence by representing the point sets as probability density functions (PDF). The registration problem is then treated as a distribution alignment. The method utilizes the Cauchy-Schwarz divergence to measure the similarity/distance between the point sets and recover the spatial transformation function needed to register them. Both algorithms utilize information theoretic descriptors; however, correntropy works at the realizations level, whereas Cauchy-Schwarz divergence works at the PDF level. This allows correntropy to be less computationally expensive, and for correct correspondence, more accurate. The two algorithms are robust against noise and outliers and perform well under varying levels of distortion. They outperform several well-known and state-of-the-art methods for point set registration. Erion Hasanbelliu, Luis Gonzalo Sánchez Giraldo, José C. Príncipe |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2014 | On the Importance of Synergisms Between Neuroscience and Engineering [Further Thoughts]
José C. Príncipe |
Proc. IEEE | 1 |
| 2014 | Cognitive Architectures for Sensory ProcessingabstractThis paper describes our efforts to design a cognitive architecture for object recognition in video. Unlike most efforts in computer vision, our work proposes a Bayesian approach to object recognition in video, using a hierarchical, distributed architecture of dynamic processing elements that learns in a self-organizing way to cluster objects in the video input. A biologically inspired innovation is to implement a top-down pathway across layers in the form of causes, creating effectively a bidirectional processing architecture with feedback. To simplify discrimination, overcomplete representations are utilized. Both inference and parameter learning are performed using empirical priors, while imposing appropriate sparseness constraints. Preliminary results show that the cognitive architecture has features that resemble the functional organization of the early visual cortex. One example showing the use of top-down connections is given to disambiguate a synthetic video from correlated noise. José C. Príncipe, Rakesh Chalasani |
Proc. IEEE | 1 |
| 2014 | The C-loss function for pattern classification
Rosha Pokharel, José C. Príncipe |
Pattern Recognit. | 3 |
| 2014 | Steady-State Mean-Square Error Analysis for Adaptive Filtering under the Maximum Correntropy CriterionabstractThe steady-state excess mean square error (EMSE) of the adaptive filtering under the maximum correntropy criterion (MCC) has been studied. For Gaussian noise case, we establish a fixed-point equation to solve the exact value of the steady-state EMSE, while for non-Gaussian noise case, we derive an approximate analytical expression for the steady-state EMSE, based on a Taylor expansion approach. Simulation results agree with the theoretical calculations quite well. Badong Chen, Lei Xing 0003, Junli Liang, Nanning Zheng 0001, José C. Príncipe |
IEEE Signal Process. Lett. | 5 |
| 2013 | A scalable RC architecture for mean-shift clusteringabstractThe mean-shift algorithm provides a unique non-parametric and unsupervised clustering solution to image segmentation and has a proven record of very good performance for a wide variety of input images. It is essential to image processing because it provides the initial and vital steps to numerous object recognition and tracking applications. However, image segmentation using mean-shift clustering is widely recognized as one of the most compute-intensive tasks in image processing, and suffers from poor scalability with respect to the image size (N pixels) and number of iterations (k): O(kN2). Our novel approach focuses on creating a scalable hardware architecture fine-tuned to the computational requirements of the mean-shift clustering algorithm. By efficiently parallelizing and mapping the algorithm to reconfigurable hardware, we can effectively cluster hundreds of pixels independently. Each pixel can benefit from its own dedicated pipeline and can move independently of all other pixels towards its respective cluster. By using our mean-shift FPGA architecture, we achieve a speedup of three orders of magnitude with respect to our software baseline. Stefan Craciun, Gongyu Wang, Alan D. George, Herman Lam, José C. Príncipe |
ASAP | 5 |
| 2013 | A greedy algorithm for model selection of tensor decompositionsabstractVarious tensor decompositions use different arrangements of factors to explain multi-way data. Components from different decompositions can vary in the number of parameters. Allowing a model to contain components from different decompositions results in a combinatoric number of possible models. Model selection balances approximation error and the number of parameters, but due to the number of possible models, post-hoc model selection is infeasible. Instead, we incrementally build a model. This approach is analogous to sparse coding with a union of dictionaries. The proposed greedy approach can estimate a model consisting of a combination of tensor decompositions. Austin J. Brockmeier, José C. Príncipe, Anh Huy Phan 0001, Andrzej Cichocki |
ICASSP | 2 |
| 2013 | Kernel recurrent system trained by real-time recurrent learning algorithmabstractThis paper presents a kernelized version of recurrent systems (KRS) and develops a kernel real-time recurrent learning (KRTRL) algorithm to train KRS. To avoid instabilities during training, the teacher forcing technique is adopted to modify the KRTRL learning. The proposed algorithms compared with the KLMS in Lorenz time series prediction. The prediction performances of the proposed algorithm outperform the KLMS significantly. Pingping Zhu, José C. Príncipe |
ICASSP | 2 |
| 2013 | A fast proximal method for convolutional sparse codingabstractSparse coding, an unsupervised feature learning technique, is often used as a basic building block to construct deep networks. Convolutional sparse coding is proposed in the literature to overcome the scalability issues of sparse coding techniques to large images. In this paper we propose an efficient algorithm, based on the fast iterative shrinkage thresholding algorithm (FISTA), for learning sparse convolutional features. Through numerical experiments, we show that the proposed convolutional extension of FISTA can not only lead to faster convergence compared to existing methods but can also easily generalize to other cost functions. Rakesh Chalasani, José C. Príncipe, Naveen Ramakrishnan |
IJCNN | 2 |
| 2013 | Survival kernel with application to kernel adaptive filteringabstractIn this paper, we define a new Mercer kernel, namely survival kernel, which is closely related to our recently proposed survival information potential (SIP). The new kernel function is parameter free, simple in calculation, and strictly positive-definite (SPD) over ℝm+, hence it has potential utility in machine learning especially in online kernel learning. In this work we apply the survival kernel to kernel adaptive filtering, in particular the kernel least mean square (KLMS) algorithm. Simulation results show that KLMS with survival kernel may achieve satisfactory performance with little computational time and without the choice of free parameters. Badong Chen, Nanning Zheng 0001, José C. Príncipe |
IJCNN | 3 |
| 2013 | Diffusion least-mean squares over adaptive networks with dynamic topologiesabstractThe purpose of this paper is to study the performance of diffusion-based distributed adaptive algorithms when relaxing the static assumption on network topology. Adaptive networks with topologies that change across time are useful to model a wide class of real-time sensor networks. This includes topologies where the number of nodes is variable and links are dynamic. We propose two schemes to accommodate dynamic topologies. The first proceeds following a probabilistic mechanism to add or remove nodes at each point in time. The second allows edges to be re-assigned at each iteration. The suggested changes allow retaining fundamental characteristics of the sensor graph, like maximum and minimum node degrees, as their significance impacts the network's throughput and its resistance to failures of neighbors. Simulations were carried for different node collaboration settings and averaged over Monte Carlo runs. Results show that no significant deterioration in performance is observed despite changes in the network size and connectivity, with a gained affinity to accommodate real sensor network systems. Bilal Fadlallah, José C. Príncipe |
IJCNN | 2 |
| 2013 | A parameter-free kernel design based on cumulative distribution function for correntropyabstractThis paper proposes a parameter-free kernel that is translation invariant and positive definite. The new kernel is based on the data cumulative distribution function (CDF) that provides all the statistical information about the observed samples. Without an explicit kernel size parameter, this novel kernel is used to define the autocorrentropy function, which is a generalized similarity measure, and spectral density estimator. Numerical examples show that the proposed method provides comparable performance to the existing Gaussian kernel with optimized kernel size. Pingping Zhu, José C. Príncipe |
IJCNN | 3 |
| 2013 | Kernel adaptive filtering with confidence intervalsabstractSince its introduction, kernel adaptive filtering (KAF) has attracted considerable attention in recent years. Its main advantages include universal nonlinear approximation using kernel methods, linearity with convex learning in the Reproducing Kernel Hilbert Space (RKHS), and online adaptation with moderate complexity. Among its applications, the kernel least mean square (KLMS) algorithm deserves particular attention due to its simplicity and effectiveness for learning complex systems. A major drawback of current implementations of KAF is the lack of a simple determination of the certainty of each estimate. In this paper, we present a novel kernel adaptive filtering architecture with confidence intervals. By introducing an auxiliary filter, the variance of each estimate can be computed using stochastic gradient descent in O(N). Results show that the proposed algorithm produces comparable estimates of the mean and the variance functions, using only a fraction of the computation associated with the Gaussian process (GP) prediction, and is more versatile in the cases of time-varying noise variance or heteroskedasticity. Kan Li 0002, Badong Chen, José C. Príncipe |
IJCNN | 3 |
| 2013 | Mixture kernel least mean squareabstractInstead of using single kernel, different approaches of using multiple kernels have been proposed recently in kernel learning literature, one of which is multiple kernel learning (MKL). In this paper, we propose an alternative to MKL in order to select the appropriate kernel given a pool of predefined kernels, for a family of online kernel filters called kernel adaptive filters (KAF). The need for an alternative is that, in a sequential learning method where the hypothesis is updated at every incoming sample, MKL would provide a new kernel, and thus a new hypothesis in the new reproducing kernel Hilbert space (RKHS) associated with the kernel. This does not fit well in the KAF framework, as learning a hypothesis in a fixed RKHS is the core of the KAF algorithms. Hence, we introduce an adaptive learning method to address the kernel selection problem for the KAF, based on competitive mixture of models. We propose mixture kernel least mean square (MxKLMS) adaptive filtering algorithm, where the kernel least mean square (KLMS) filters learned with different kernels, act in parallel at each input instance and are competitively combined such that the filter with the best kernel is an expert for each input regime. The competition among these experts is created by using a performance based gating, that chooses the appropriate expert locally. Therefore, the individual filter parameters as well as the weights for combination of these filters are learned simultaneously in an online fashion. The results obtained suggest that the model not only selects the best kernel, but also significantly improves the prediction accuracy. Rosha Pokharel, Sohan Seth, José C. Príncipe |
IJCNN | 3 |
| 2013 | Analysis on extended kernel recursive least squares algorithmabstractIn this paper, the extended kernel recursive least squares (Ex-KRLS) algorithm is reviewed and analyzed. We point out that the Theorem 1 in [10] is not always correct for general cases. Furthermore, the Ex-KRLS algorithm for tracking model is just a random walk KRLS algorithm. Finally, this algorithm is explained as a special Kalman filter in the reproducing kernel Hilbert space. Pingping Zhu, José C. Príncipe |
IJCNN | 2 |
| 2013 | Kernel minimum error entropy algorithm
Badong Chen, Zejian Yuan, Nanning Zheng 0001, José C. Príncipe |
Neurocomputing | 4 |
| 2013 | Fixed budget quantized kernel least-mean-square algorithm
Songlin Zhao, Badong Chen, Pingping Zhu, José C. Príncipe |
Signal Process. | 4 |
| 2013 | Invexity of the Minimum Error Entropy CriterionabstractIn this letter, optimization properties of Minimization of Error Entropy (MEE) and Minimization of Error Entropy with Fiducial points (MEEF) are presented. It is proved that by varying the kernel parameter of the MEE and/or MEEF objective function, the resulting problem, in general leads to an invex problem. Furthermore, for certain values of the kernel parameter it is shown that the problems may transform to convex or pseudo-convex problems. Mujahid N. Syed, Panos M. Pardalos, José C. Príncipe |
IEEE Signal Process. Lett. | 3 |
| 2013 | Guest Editorial: Brain/neuronal - Computer game interfaces and interactionabstractWhile, to date, there has been successful research into brain-computer game interaction (BCGI), the algorithms and techniques developed are limited in scope and may not utilize all available data in the appropriate contexts, e.g., optimizing for genre-specific games. This special issue was, therefore, solicited to gain insights into new biosignal processing algorithms, tested in gaming applications, which exploit BCI and neural signals to enhance gameplay experience and playermotivation, be the players ablebodied or physically impaired. A snapshot of the current trends in BCIcontrolled computer games is presented across 11 manuscripts. Each is briefly summarized in this editorial introduction. Damien Coyle, José C. Príncipe, Fabien Lotte, Anton Nijholt |
IEEE Trans. Comput. Intell. AI Games | 2 |
| 2013 | Quantized Kernel Recursive Least Squares AlgorithmabstractIn a recent paper, we developed a novel quantized kernel least mean square algorithm, in which the input space is quantized (partitioned into smaller regions) and the network size is upper bounded by the quantization codebook size (number of the regions). In this paper, we propose the quantized kernel least squares regression, and derive the optimal solution. By incorporating a simple online vector quantization method, we derive a recursive algorithm to update the solution, namely the quantized kernel recursive least squares algorithm. The good performance of the new algorithm is demonstrated by Monte Carlo simulations. Badong Chen, Songlin Zhao, Pingping Zhu, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2012 | One-class classifier based on extreme value statistics
David Martínez-Rego, Evan Kriminger, José C. Príncipe, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
ESANN | 3 |
| 2012 | Analyzing dependence structure of the human brain in response to visual stimuliabstractCommunication between cortices mediated by deep brain structures such as the amygdala and fusiform gyrus has been suggested to explain the enhanced perception of stimuli bearing emotional content or having facial features. In this paper, we analyze the dependence structure of the relevant brain regions to assess their connectivity in response to a facial stimulus, and to discriminate it from a mock stimulus. The proposed approach treats the brain as a graphical network where vertices correspond to locations of electroencephalogram (EEG) recordings, and weights of the edges correspond to dependence values. We employ a novel measure of dependence, called generalized measure of association (GMA), due to its underlying simplicity, and compare its performance against Pearson's correlation. The performance is assessed in terms of the discriminability between the face and mock stimuli. We observe that GMA successfully exhibits higher dependence in regions that might reflect the activity of the amygdaloid complex and the right fusiform gyrus when the stimulus is face. Furthermore, the distributions of the dependence values show that GMA also achieves a better separation between face and mock, compared to correlation. Bilal Fadlallah, Sohan Seth, Andreas Keil, José C. Príncipe |
ICASSP | 4 |
| 2012 | Data-driven tree-structured Bayesian network for image segmentationabstractThis paper presents Data-Driven Tree-structured Bayesian network (DDT), a novel probabilistic graphical model for hierarchical unsupervised image segmentation. The DDT captures long and short-ranged correlations between neighboring regions in each image using a tree-structured prior. Unlike other previous work, DDT first segments an input image into superpixels and learn a tree-structured prior based on the topology of superpixels in different scales. Such a tree structure is referred to as data-driven tree structure. Each superpixel is represented by a variable node taking a discrete value of class/label of the segmentation. The probabilistic relationships among the nodes are represented by edges in the network. The unsupervised image segmentation, hence, can be viewed as an inference problem of the nodes in the tree structure of DDT, which can be carried out efficiently. We evaluate quantitatively our results with respect to the ground-truth segmentation, demonstrating that our proposed framework performs competitively with the state of the art in unsupervised image segmentation and contour detection. Kittipat Kampa, José C. Príncipe, Duangmanee Putthividhya, Anand Rangarajan 0001 |
ICASSP | 2 |
| 2012 | Measure of Statistical Dependence
José C. Príncipe |
ICPRAM (1) | 1 |
| 2012 | Sequential causal estimation and learning from time-varying imagesabstractDynamic models are used in modeling the perceptual systems with hierarchies. But most of the models assume Gaussian statistics on the underlying causes. In this paper we try to develop a basic building block for these hierarchical models where the causes are assumed to be non-Gaussian. We describe a sequential dual estimation framework for inferring the hidden states and unknown causes/inputs while learning the parameters of the model. It is observed that the algorithm is able to extract bases from the time varying image sequence that resembles receptive fields of the simple cells in V1. In addition, the dynamical model gives us the ability to deconvolve spatial and temporal changes in the image sequence. Rakesh Chalasani, Goktug T. Cinar, José C. Príncipe |
IJCNN | 3 |
| 2012 | Online efficient learning with quantized KLMS and L1 regularizationabstractIn a recent work, we have proposed the quantized kernel least mean square (QKLMS) algorithm, which is quite effective in online learning sequentially a nonlinear mapping with a slowly growing radial basis function (RBF) structure. In this paper, in order to further reduce the network size, we propose a sparse QKLMS algorithm, which is derived by adding a sparsity inducing l 1 norm penalty of the coefficients to the squared error cost. Simulation examples show that the new algorithm works efficiently, and results in a much sparser network while preserving a desirable performance. Badong Chen, Songlin Zhao, Sohan Seth, José C. Príncipe |
IJCNN | 4 |
| 2012 | Hidden state estimation using the Correntropy Filter with fixed point update and adaptive kernel sizeabstractIn this paper we review the Correntropy Filter for hidden state estimation and we introduce the fixed point update rule for the Correntropy Filter instead of using gradient ascent for faster convergence. We further propose an adaptive kernel bandwidth selection algorithm. It is shown that the new filter outperforms the Kalman Filter and has no free parameters. The algorithm's capabilities are demonstrated on a simulated experiment and a vehicle tracking problem. Goktug T. Cinar, José C. Príncipe |
IJCNN | 2 |
| 2012 | Online learning using a Bayesian surprise metricabstractDictionary.com defines learning as the process of acquiring knowledge. In psychology, learning is defined as the modification of behavior through training. In our work, we combine these definitions to define learning as the modification of a system model to incorporate the knowledge acquired by new observations. During learning, the system creates and modifies a model to improve its performance. As new samples are introduced, the system updates its model based on the new information provided by the samples. However, this update may not necessarily improve the model. We propose a Bayesian surprise metric to differentiate good data (beneficial) from outliers (detrimental), and thus help to selectively adapt the model parameters. The surprise metric is calculated based on the difference between the prior and the posterior distributions of the model when a new sample is introduced. The metric is useful not only to identify outlier data, but also to differentiate between the data carrying useful information for improving the model and those carrying no new information (redundant). Allowing only the relevant data to update the model would speed up the learning process and prevent the system from overfitting. The method is demonstrated in all three learning procedures: supervised, semi-supervised and unsupervised. The results show the benefit of surprise in both clustering and outlier detection. Erion Hasanbelliu, Kittipat Kampa, José C. Príncipe, James Tory Cobb |
IJCNN | 3 |
| 2012 | Nearest Neighbor Distributions for imbalanced classificationabstractThe class imbalance problem is pervasive in machine learning. To accurately classify the minority class, current methods rely on sampling schemes to close the gap between classes, or on the application of error costs to create algorithms which favor the minority class. Since the sampling schemes and costs must be specified, these methods are highly dependent on the class distributions present in the training set. This makes them difficult to apply in settings where the level of imbalance changes, such as in online streaming data. Often they cannot handle multi-class problems. We present a novel single-class algorithm called Class Conditional Nearest Neighbor Distribution (CCNND), which mitigates the effects of class imbalance through local geometric structure in the data. Our algorithm can be applied seamlessly to problems with any level of imbalance or number of classes, and new examples are simply added to the training set. We show that it performs as well as or better than top sampling and cost-weighting methods on four imbalanced datasets from the UCI Machine Learning Repository, and then apply it to streaming data from the oil and gas industry alongside a modified nearest neighbor algorithm. Our algorithm's competitive performance relative to the state-of-the-art, coupled with its extremely simple implementation and automatic adjustment for minority classes, demonstrates that it is worth further study. Evan Kriminger, José C. Príncipe, Choudur Lakshminarayan |
IJCNN | 2 |
| 2012 | Kernel classifier with Correntropy lossabstractClassification can be seen as a mapping problem where some function of x n predicts the expectation of a class variable y n . This paper uses kernel methods for the prediction of class variable, together with a recently proposed cost function for classification, called Correntropy-loss (C-loss) function. C-Loss is a non-convex loss function based on a similarity measure called correntropy and is known to closely approximate the ideal 0-1 loss function for classification. This paper shows via experimental results that, by replacing the cost function - Mean Square Error (MSE) in a conventional kernel based functional mapping, by a non-convex loss function C-Loss, a non-overfitting, and hence, a better classifier can be obtained. Since gradient descent can still be used with the C-loss and the kernel mapper, the classifier can be easily trained without performance penalty, compared to the SVM, which makes the approach very practical. Rosha Pokharel, José C. Príncipe |
IJCNN | 2 |
| 2012 | An adaptive kernel width update for correntropyabstractCorrentropy, as an adaptive criterion of Information Theoretic Learning (ITL), has been successfully used in signal processing and machine learning. How to appropriately select the kernel width of correntropy is a crucial problem in correntropy applications. Existing kernel width selection methods are not suitable enough for this problem. In this paper, we develop an adaptive method for kernel width selection in correntropy. Based on the Middleton's non-Gaussian models, this method utilizes the kurtosis as a ratio to adjust the standard deviation of the prediction error to obtain the kernel width online. The superior performance of the new method has been demonstrated by simulation examples in the noisy frequency doubling and echo cancelation problems. Songlin Zhao, Badong Chen, José C. Príncipe |
IJCNN | 3 |
| 2012 | Locally linear embedding based on correntropy measure for visualization and classification
Genaro Daza-Santacoloma, Germán Castellanos-Domínguez, José C. Príncipe |
Neurocomputing | 3 |
| 2012 | Strictly Positive-Definite Spike Train Kernels for Point-Process DivergencesabstractExploratory tools that are sensitive to arbitrary statistical variations in spike train observations open up the possibility of novel neuroscientific discoveries. Developing such tools, however, is difficult due to the lack of Euclidean structure of the spike train space, and an experimenter usually prefers simpler tools that capture only limited statistical features of the spike train, such as mean spike count or mean firing rate. We explore strictly positive-definite kernels on the space of spike trains to offer both a structural representation of this space and a platform for developing statistical measures that explore features beyond count or rate. We apply these kernels to construct measures of divergence between two point processes and use them for hypothesis testing, that is, to observe if two sets of spike trains originate from the same underlying probability law. Although there exist positive-definite spike train kernels in the literature, we establish that these kernels are not strictly definite and thus do not induce measures of divergence. We discuss the properties of both of these existing nonstrict kernels and the novel strict kernels in terms of their computational complexity, choice of free parameters, and performance on both synthetic and real data through kernel principal component analysis and hypothesis testing. Il Park 0002, Sohan Seth, Murali Rao, José C. Príncipe |
Neural Comput. | 4 |
| 2012 | Conditional AssociationabstractEstimating conditional dependence between two random variables given the knowledge of a third random variable is essential in neuroscientific applications to understand the causal architecture of a distributed network. However, existing methods of assessing conditional dependence, such as the conditional mutual information, are computationally expensive, involve free parameters, and are difficult to understand in the context of realizations. In this letter, we discuss a novel approach to this problem and develop a computationally simple and parameter-free estimator. The difference between the proposed approach and the existing ones is that the former expresses conditional dependence in terms of a finite set of realizations, whereas the latter use random variables, which are not available in practice. We call this approach conditional association, since it is based on a generalization of the concept of association to arbitrary metric spaces. We also discuss a novel and computationally efficient approach of generating surrogate data for evaluating the significance of the acquired association value. Sohan Seth, José C. Príncipe |
Neural Comput. | 2 |
| 2012 | A novel extended kernel recursive least squares algorithm
Pingping Zhu, Badong Chen, José C. Príncipe |
Neural Networks | 3 |
| 2012 | Mean square convergence analysis for kernel least mean square algorithm
Badong Chen, Songlin Zhao, Pingping Zhu, José C. Príncipe |
Signal Process. | 4 |
| 2012 | Extraction of signals with higher order temporal structure using Correntropy
Eder Santana, José C. Príncipe, Ewaldo E. C. Santana, Allan Kardec Barros |
Signal Process. | 2 |
| 2012 | Maximum Correntropy Estimation Is a Smoothed MAP EstimationabstractAs a new measure of similarity, the correntropy can be used as an objective function for many applications. In this letter, we study Bayesian estimation under maximum correntropy (MC) criterion. We show that the MC estimation is, in essence, a smoothed maximum a posteriori (MAP) estimation, including the MAP and the minimum mean square error (MMSE) estimation as the extreme cases. We also prove that under a certain condition, when the kernel size in correntropy is larger than some value, the MC estimation will have a unique optimal solution lying in a strictly concave region of the smoothed posterior distribution. Badong Chen, José C. Príncipe |
IEEE Signal Process. Lett. | 2 |
| 2012 | Quantized Kernel Least Mean Square AlgorithmabstractIn this paper, we propose a quantization approach, as an alternative of sparsification, to curb the growth of the radial basis function structure in kernel adaptive filtering. The basic idea behind this method is to quantize and hence compress the input (or feature) space. Different from sparsification, the new approach uses the "redundant" data to update the coefficient of the closest center. In particular, a quantized kernel least mean square (QKLMS) algorithm is developed, which is based on a simple online vector quantization method. The analytical study of the mean square convergence has been carried out. The energy conservation relation for QKLMS is established, and on this basis we arrive at a sufficient condition for mean square convergence, and a lower and upper bound on the theoretical value of the steady-state excess mean square error. Static function estimation and short-term chaotic time-series prediction examples are presented to demonstrate the excellent performance. Badong Chen, Songlin Zhao, Pingping Zhu, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2012 | Assessing Granger Non-Causality Using Nonparametric Measure of Conditional IndependenceabstractIn recent years, Granger causality has become a popular method in a variety of research areas including engineering, neuroscience, and economics. However, despite its simplicity and wide applicability, the linear Granger causality is an insufficient tool for analyzing exotic stochastic processes such as processes involving non-linear dynamics or processes involving causality in higher order statistics. In order to analyze such processes more reliably, a different approach toward Granger causality has become increasingly popular. This new approach employs conditional independence as a tool to discover Granger non-causality without any assumption on the underlying stochastic process. This paper discusses the concept of discovering Granger non-causality using measures of conditional independence, and proposes a novel measure of conditional independence. In brief, the proposed approach estimates the conditional distribution function through a kernel based least square regression approach. This paper also explores the strengths and weaknesses of the proposed method compared to other available methods, and provides a detailed comparison of these methods using a variety of synthetic data sets. Sohan Seth, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2011 | Statistical dependence measure for feature selection in microarray datasets
Verónica Bolón-Canedo, Sohan Seth, Noelia Sánchez-Maroño, Amparo Alonso-Betanzos, José C. Príncipe |
ESANN | 5 |
| 2011 | Information theory related learning
Thomas Villmann, José C. Príncipe, Andrzej Cichocki |
ESANN | 2 |
| 2011 | An efficient rank-deficient computation of the Principle of Relevant InformationabstractOne of the main difficulties in computing information theoretic learning (ITL) estimators is the computational complexity that grows quadratically with data. Considerable amount of work has been done on computation of low rank approximations of Gram matrices without accessing all their elements. In this paper we discuss how these techniques can be applied to reduce computational complexity of Principle of Relevant Information (PRI). This particular objective function involves estimators of Renyi's second order entropy and cross-entropy and their gradients, therefore posing a technical challenge for implementation in a realistic scenario. Moreover, we introduce a simple modification to the Nyström method motivated by the idea that our estimator must perform accurately only for certain vectors not for all possible cases. We show some results on how this rank deficient decompositions allow the application of the PRI on moderately large datasets. Luis Gonzalo Sánchez Giraldo, José C. Príncipe |
ICASSP | 2 |
| 2011 | Modified embedding for multi-regime detection in nonstationary streaming dataabstractMany practical data streams are typically composed of several states known as regimes. In this paper, we invoke phase space reconstruction methods from non-linear time series and dynamical systems for regime detection. But the data collected from sensors is normally noisy, does not have constant amplitude and is sometimes plagued by shifts in the mean. All these aspects make modeling even more difficult. We propose a representation of the time series in the phase space with a modified embedding, which is invariant to translation and scale. The features we use for regime detection are based on comparing trajectory segments in the modified embedding space with cross-correntropy, which is a generalized correlation function. We apply our algorithm to non-linear oscillations, and compare its performance with the standard time delay embedding. Evan Kriminger, José C. Príncipe, Choudur Lakshminarayan |
ICASSP | 2 |
| 2011 | Estimation of symmetric chi-square divergence for point processesabstractThis paper addresses the estimation of symmetric χ2-divergence between two point processes. We propose a novel approach by, first, mapping the space of spike trains in an appropriate functional space, and then, estimating the divergence in this functional space using a least square regression approach. We compare the proposed approach with other available methods on simulated data, and discuss its pros and cons. Il Park 0002, Sohan Seth, Murali Rao, José C. Príncipe |
ICASSP | 4 |
| 2011 | A metric approach toward point process divergenceabstractEstimating divergence between two point processes, i.e. probability laws on the space of spike trains, is an essential tool in many computational neuroscience applications, such as change detection and neural coding. However, the problem of estimating divergence, although well studied in the Euclidean space, has seldom been addressed in a more general setting. Since the space of spike trains can be viewed as a metric space, we address the problem of estimating Jensen-Shannon divergence in a metric space using a nearest neighbor based approach. We empirically demonstrate the validity of the proposed estimator, and compare it against other available methods in the context of two-sample problem. Sohan Seth, Austin J. Brockmeier, José C. Príncipe |
ICASSP | 3 |
| 2011 | From compressive to adaptive sampling of neural and ECG recordingsabstractThe miniaturization required for interfacing with the brain demands new methods of transforming neuron responses (spikes) into digital representations. The sparse nature of neural recordings is evident when represented in a shift invariant basis. Although a compressive sensing (CS) framework may seem suitable in reducing the data rates, we show that the time varying sparsity in the signals makes it difficult to apply. Furthermore, we present an adaptive sampling scheme which takes advantage of the local characteristics of the neural spike trains and electrocardiograms (ECG). In contrast to the global constraints imposed in CS our solution is sensitive to the local time structure of the input. The simplicity in the design of the integrate-and-fire (IF) make it a viable solution in current brain machine interfaces (BMI) and ambulatory cardiac monitoring. Alexander Singh-Alvarado, José C. Príncipe |
ICASSP | 2 |
| 2011 | Sparse analog associative memory via L1-regularization and thresholdingabstractThe CA3 region of the hippocampus acts as an auto-associative memory and is responsible for the consolidation of episodic memory. Two important characteristics of such a network is the sparsity of the stored patterns and the nonsaturating firing rate dynamics. To construct such a network, here we use a maximum a posteriori based cost function, regularized with L1-norm, to change the internal state of the neurons. Then a linear thresholding function is used to obtain the desired output firing rate. We show how such a model leads to a more biologically reasonable dynamic model which can produce a sparse output and recalls with good accuracy when the network is presented with a corrupted input. Rakesh Chalasani, José C. Príncipe |
IJCNN | 2 |
| 2011 | Adaptive background estimation using an information theoretic cost for hidden state estimationabstractHidden state estimation in linear systems is a popular and broad research topic which became a mainstream research area after Rudolf Kalman's seminal paper. The Kalman Filter (KF) gives the optimal solution to the estimation problem in a setting where all the processes are Gaussian random processes. However because of the suboptimal behavior of the KF in non-Gaussian settings, there is a need for a new filter that can extract higher order information from the signals. In this paper we propose using an information theoretic cost function utilizing the similarity measure Correntropy as a performance index. This results in a different perspective on hidden state estimation. We present the superior performance of the new filter on both synthetic data and on adaptive background estimation problem and discuss future research directions. Goktug T. Cinar, José C. Príncipe |
IJCNN | 2 |
| 2011 | Closed-form cauchy-schwarz PDF divergence for mixture of GaussiansabstractThis paper presents an efficient approach to calculate the difference between two probability density functions (pdfs), each of which is a mixture of Gaussians (MoG). Unlike Kullback-Leibler divergence (DKL), the authors propose that the Cauchy-Schwarz (CS) pdf divergence measure (DCS) can give an analytic, closed-form expression for MoG. This property of the DCSmakes fast and efficient calculations possible, which is tremendously desired in real-world applications where the dimensionality of the data/features is very high. We show that DCSfollows similar trends to DKL, but can be computed much faster, especially when the dimensionality is high. Moreover, the proposed method is shown to significantly outperform DKLin classifying real-world 2D and 3D objects, and static hand posture recognition based on distances alone. Kittipat Kampa, Erion Hasanbelliu, José C. Príncipe |
IJCNN | 3 |
| 2011 | Evaluating dependence in spike train metric spacesabstractAssessing dependence between two sets of spike trains or between a set of input stimuli and the corresponding generated spike trains is crucial in many neuroscientific applications, such as in analyzing functional connectivity among neural assemblies, and in neural coding. Dependence between two random variables is traditionally assessed in terms of mutual information. However, although well explored in the context of real or vector valued random variables, estimating mutual information still remains a challenging issue when the random variables exist in more exotic spaces such as the space of spike trains. In the statistical literature, on the other hand, the concept of dependence between two random variables has been presented in many other ways, e.g. using copula, or using measures of association such as Spearman's ρ, and Kendall's τ. Although these methods are usually applied on the real line, their simplicity, both in terms of understanding and estimating, make them worth investigating in the context of spike train dependence. In this paper, we generalize the concept of association to any abstract metric spaces. This new approach is an attractive alternative to mutual information, since it can be easily estimated from realizations without binning or clustering. It also provides an intuitive understanding of what dependence implies in the context of realizations. We show that this new methodology effectively captures dependence between sets of stimuli and spike trains. Moreover, the estimator has desirable small sample characteristic, and it often outperforms an existing similar metric based approach. Sohan Seth, Austin J. Brockmeier, John S. Choi, Mulugeta Semework, Joseph T. Francis, José C. Príncipe |
IJCNN | 6 |
| 2011 | Kernel adaptive filtering with maximum correntropy criterionabstractKernel adaptive filters have drawn increasing attention due to their advantages such as universal nonlinear approximation with universal kernels, linearity and convexity in Reproducing Kernel Hilbert Space (RKHS). Among them, the kernel least mean square (KLMS) algorithm deserves particular attention because of its simplicity and sequential learning approach. Similar to most conventional adaptive filtering algorithms, the KLMS adopts the mean square error (MSE) as the adaptation cost. However, the mere second-order statistics is often not suitable for nonlinear and non-Gaussian situations. Therefore, various non-MSE criteria, which involve higher-order statistics, have received an increasing interest. Recently, the correntropy, as an alternative of MSE, has been successfully used in nonlinear and non-Gaussian signal processing and machine learning domains. This fact motivates us in this paper to develop a new kernel adaptive algorithm, called the kernel maximum correntropy (KMC), which combines the advantages of the KLMS and maximum correntropy criterion (MCC). We also study its convergence and self-regularization properties by using the energy conservation relation. The superior performance of the new algorithm has been demonstrated by simulation experiments in the noisy frequency doubling problem. Songlin Zhao, Badong Chen, José C. Príncipe |
IJCNN | 3 |
| 2011 | A nonparametric information theoretic approach for change detection in time seriesabstractThis paper presents an online nonparametric methodology based on the Kernel Least Mean Square (KLMS) algorithm and the surprise criterion, which is based on an information theoretic framework. Surprise quantifies the amount of information a datum contains given a known system state, and can be estimated online using Gaussian Process Theory. Based on this concept, we use the KLMS algorithm together with surprise criterion to detect regime change in nonstationary time series. We test the methodology on a synthesized chaotic time series to illustrate this criterion. The results show that surprise criterion is better than the conventional segmentation based on the error criterion. Songlin Zhao, José C. Príncipe |
IJCNN | 2 |
| 2011 | Extended Kalman filter using a kernel recursive least squares observerabstractIn this paper, a novel methodology is proposed to solve the state estimation problem combining the extended Kalman filter (EKF) with a kernel recursive least squares (KRLS) algorithm (EKF-KRLS). The EKF algorithm estimates hidden states in the input space, while the KRLS algorithm estimates the measurement model. The algorithm works well without knowing the linear or nonlinear measurement model. We apply this algorithm to vehicle tracking, and compare the performances with traditional Kalman filter, EKF and KRLS algorithms. Results demonstrate that the performance of the EKF-KRLS algorithm outperforms these existing algorithms. Especially when nonlinear measurement functions are applied, the advantage of the EKF-KRLS algorithm is very obvious. Pingping Zhu, Badong Chen, José C. Príncipe |
IJCNN | 3 |
| 2011 | Integrate and fire circuit as an ADC replacementabstractWe evaluate the performance of an asynchronous analog-to-digital converter known as the Integrate and Fire (I&F) circuit. The reconstruction algorithm for the sampler is presented along with the analytical results for the signal-to-noise ratio (SNR) of the reconstructed signal. We present the power vs. SNR tradeoff, based on the measurements from the chip fabricated in AMI 0.6 µm technology. The figure-of-merit (FOM) of the I&F sampler is 0.6 pJ/conversion, which is comparable to the reported FOM for other synchronous and asynchronous ADCs, even though the I&F is implemented in a much older technology as compared to the other ADCs. Manu Rastogi, Alexander Singh-Alvarado, John G. Harris, José C. Príncipe |
ISCAS | 4 |
| 2011 | The integrate-and-fire sampler: A special type of asynchronous Σ - Δ modulatorabstractWe present the integrate-and-fire (IF) model as a time encoding scheme and show that it corresponds to a special case of the asynchronous Σ-Δ modulator (ASDM). We depart from typical rate based encoding used by conventional ASDM towards time based encoding. We show that for certain families of signals, such as neural recordings, the IF sampler accurately encodes the regions of interest at sub-Nyquist data rates. Nevertheless, the encoder requires nonlinear recovery algorithms, instead of the typical low pass filter. Alexander Singh-Alvarado, Manu Rastogi, John G. Harris, José C. Príncipe |
ISCAS | 4 |
| 2011 | Autocorrelation features for synthetic aperture sonar image seabed segmentationabstractHigh-resolution synthetic aperture sonar (SAS) systems yield richly detailed images of seabed environments. Algorithms that automatically segment and label seabed textures such as coral, sea grass, sand ripple, and mud, require suitable features that discriminate between the texture classes. Here we present a robust, parameterized SAS image texture model based on the autocorrelation function (ACF) of the intensity image. This ACF texture model has been shown to accurately model first- and second-order statistical features of various seabed environments. An unsupervised multi-class k-means segmentation algorithm that uses the features derived from the ACF model is employed to label rock and ripple textures from a set of textured SAS images. The results of the segmentation are compared against the performance of the segmentation approach using biorthogonal wavelets and Haralick features. In the described experiments, the ACF model features are shown to produce better segmentations than the features based on wavelet coefficients and Haralick features for classifiers of low complexity. James Tory Cobb, José C. Príncipe |
SMC | 2 |
| 2011 | Δ-Entropy: Definition, properties and applications in system identification with quantized data
Badong Chen, Yu Zhu 0001, Jinchun Hu, José C. Príncipe |
Inf. Sci. | 4 |
| 2011 | A test of independence based on a generalized correlation function
Murali Rao, Sohan Seth, Jianwu Xu, Yunmei Chen, Hemant D. Tagare, José C. Príncipe |
Signal Process. | 6 |
| 2011 | Information theoretic learning with adaptive kernels
José C. Príncipe |
Signal Process. | 2 |
| 2011 | Period Estimation in Astronomical Time Series Using Slotted CorrentropyabstractIn this letter, we propose a method for period estimation in light curves from periodic variable stars using correntropy. Light curves are astronomical time series of stellar brightness over time, and are characterized as being noisy and unevenly sampled. We propose to use slotted time lags in order to estimate correntropy directly from irregularly sampled time series. A new information theoretic metric is proposed for discriminating among the peaks of the correntropy spectral density. The slotted correntropy method outperformed slotted correlation, string length, VarTools (Lomb-Scargle periodogram and Analysis of Variance), and SigSpec applications on a set of light curves drawn from the MACHO survey. Pablo Huijse, Pablo A. Estévez, Pablo Zegers, José C. Príncipe, Pavlos Protopapas |
IEEE Signal Process. Lett. | 4 |
| 2011 | An Augmented Echo State Network for Nonlinear Adaptive Filtering of Complex Noncircular SignalsabstractA novel complex echo state network (ESN), utilizing full second-order statistical information in the complex domain, is introduced. This is achieved through the use of the so-called augmented complex statistics, thus making complex ESNs suitable for processing the generality of complex-valued signals, both second-order circular (proper) and noncircular (improper). Next, in order to deal with nonstationary processes with large nonlinear dynamics, a nonlinear readout layer is introduced and is further equipped with an adaptive amplitude of the nonlinearity. This combination of augmented complex statistics and enhanced adaptivity within ESNs also facilitates the processing of bivariate signals with strong component correlations. Simulations in the prediction setting on both circular and noncircular synthetic benchmark processes and real-world noncircular and nonstationary wind signals support the analysis. Yili Xia, Beth Jelfs, Marc M. Van Hulle, José C. Príncipe, Danilo P. Mandic |
IEEE Trans. Neural Networks | 4 |
| 2010 | Information Theoretic fuzzy modeling for regressionabstractThis paper presents a novel, Information Theoretic Learning (ITL) method to model a fuzzy system for regression tasks that minimizes the Renyi's entropy of the error signal. An architecture based on a generalization of the well-known Adaptive-Network-Based Fuzzy Inference System (ANFIS) was used to perform such a modeling. The resulting method was tested on the prediction of future values for the Mackey-Glass chaotic time series. The results show that, when using the ITL cost function, the method returns better models in comparison with a Mean Squared Error (MSE)-guided cost function. Diego Álvarez-Estévez, José C. Príncipe, Vicente Moret-Bonillo |
FUZZ-IEEE | 2 |
| 2010 | Linear Projection Method Based on Information Theoretic Learning
Pablo A. Vera, Pablo A. Estévez, José C. Príncipe |
ICANN (3) | 3 |
| 2010 | Dynamic factor graphs: A novel framework for multiple features data fusionabstractThe Dynamic Tree (DT) Bayesian Network is a powerful analytical tool for image segmentation and object segmentation tasks. Its hierarchical nature makes it possible to analyze and incorporate information from different scales, which is desirable in many applications. Having a flexible structure enables model selection, concurrent with parameter inference. In this paper, we propose a novel framework, dynamic factor graphs (DFG), where data segmentation and fusion tasks are combined in the same framework. Factor graphs (FGs) enable us to have a broader range of modeling applications than Bayesian networks (BNs) since FGs include both directed acyclic and undirected graphs in the same setting. The example in this paper will focus on segmentation and fusion of 2D image features with a linear Gaussian model assumption. Kittipat Kampa, José C. Príncipe, K. Clint Slatton |
ICASSP | 2 |
| 2010 | Quantification of inter-trial non-stationarity in spike trains from periodically stimulated neural culturesabstractIn neuroscience, non-stationarity detection of spike trains is useful for ensuring stability of experimental condition, and detecting plasticity. A novel method for estimating point process divergence and its application for non-stationarity detection in spike trains is proposed. The method for measuring divergence is based on decomposition of finite point process and Hilbertian metrics. The method is demonstrated by detecting short-term and long-term plasticity in neural culture probed with periodic stimulations. Il Park 0002, José C. Príncipe |
ICASSP | 2 |
| 2010 | A conditional distribution function based approach to design nonparametric tests of independence and conditional independenceabstractMeasures of independence and conditional independence are two important statistical concepts that have found profound applications in engineering such as in feature selection and causality detection, respectively. Therefore, designing efficient ways, typically nonparametric, to estimate these measures has been an active research area in the last decade. In this paper, we propose a novel framework to test (conditional) independence, using the concept of conditional distribution function. Although, estimating conditional distribution function is a difficult task on its own, we show that the proposed measures can be estimated efficiently and actually can be expressed as the Frobenius norm of a matrix. We compare the proposed methods with other state-of-the-art techniques and show that they yield very promising results. Sohan Seth, José C. Príncipe |
ICASSP | 2 |
| 2010 | Kernel width adaptation in information theoretic cost functionsabstractThis paper presents an algorithm for online adaptation of the kernel width parameter in information theoretic cost functions used for adaptive system training. Training algorithms which optimize information theoretic quantities like entropy involve choosing a kernel size for their sample estimators. The kernel size essentially dictates the nature of the performance surface of the cost function over which the system parameters adapt. This, in turn, governs factors like speed of adaptation and presence of local minima. We show results of using the Minimum Error Entropy (MEE) criterion with the proposed adaptive kernel algorithm for training a time delay neural network. Our simulations show that having an adaptive kernel width results in faster convergence of parameters as compared to having fixed values. José C. Príncipe |
ICASSP | 2 |
| 2010 | A closed form recursive solution for Maximum Correntropy trainingabstractThis paper presents a closed form recursive solution for training adaptive filters using the Maximum Correntropy Criterion (MCC). Correntropy has been recently proposed as a robust similarity measure between two random variables or signals, when the pdfs involved are heavy tailed and non-Gaussian. Maximizing the cross-correntropy between the output of an adaptive filter and the desired response leads to the Maximum Correntropy Criterion for adaptive systems training. We show that a closed form, recursive solution of the filter weights using this criterion yields a simple weighted least squares like formulation. Our simulations show that training the filter weights using this recursive solution is much faster than gradient based training, and more accurate than the RLS algorithm in cases where the error pdf is non-Gaussian and heavy tailed. José C. Príncipe |
ICASSP | 2 |
| 2010 | Fixed-budget kernel recursive least-squaresabstractWe present a kernel-based recursive least-squares (KRLS) algorithm on a fixed memory budget, capable of recursively learning a nonlinear mapping and tracking changes over time. In order to deal with the growing support inherent to online kernel methods, the proposed method uses a combined strategy of growing and pruning the support. In contrast to a previous sliding-window based technique, the presented algorithm does not prune the oldest data point in every time instant but it instead aims to prune the least significant data point. We also introduce a label update procedure to equip the algorithm with tracking capability. Simulations show that the proposed method obtains better performance than state-of-the-art kernel adaptive filtering techniques given similar memory requirements. Steven Van Vaerenbergh, Ignacio Santamaría, Weifeng Liu 0016, José C. Príncipe |
ICASSP | 4 |
| 2010 | Variable Selection: A Statistical Dependence PerspectiveabstractMeasures of statistical dependence such as the correlation coefficient and mutual information have been widely used in variable selection. The use of correlation has been inspired by the concept of regression whereas the use of mutual information has been largely motivated by information theory. In a statistical sense, however, the concept of dependence is much broader, and extends beyond correlation and mutual information. In this paper, we explore the fundamental notion of statistical dependence in the context of variable selection. In particular, we discuss the properties of dependence as proposed by Rényi, and evaluate their significance in the variable selection context. We, also, explore a measure of dependence that satisfies most of these desired properties, and discuss its applicability as a substitute for correlation coefficient and mutual information. Finally, we compare these measures of dependence to select important variables for regression with real world data. Sohan Seth, José C. Príncipe |
ICMLA | 2 |
| 2010 | A Recursive Online Kernel PCA AlgorithmabstractIn this paper, we describe a new method for performing kernel principal component analysis which is online and also has a fast convergence rate. The method follows the Rayleigh quotient to obtain a fixed point update rule to extract the leading eigenvalue and eigenvector. Online deflation is used to estimate the remaining components. These operations are performed in reproducing kernel Hilbert space (RKHS) with linear order memory and computation complexity. The derivation of the method and several applications are presented. Erion Hasanbelliu, Luis Gonzalo Sánchez Giraldo, José C. Príncipe |
ICPR | 3 |
| 2010 | A Test of Granger Non-causality Based on Nonparametric Conditional IndependenceabstractIn this paper we describe a test of Granger non-causality from the perspective of a new measure of nonparametric conditional independence. We apply the proposed test on two synthetic nonlinear problems where linear Granger causality fails and show that the proposed method is able to derive the true causal connectivity effectively. Sohan Seth, José C. Príncipe |
ICPR | 2 |
| 2010 | Information theoretic learning applied to wind power modelingabstractThis paper reports new results in adopting information theoretic learning concepts in the training of neural networks to perform wind power forecasts. The forecast “goodness” is discussed under two paradigms: one is only concerned in measuring the deviation between the forecasted and realized values, the other is related with the value of the forecast in the electricity market for different agents. The results and conclusions are supported by a real case example. Ricardo J. Bessa, Vladimiro Miranda, José C. Príncipe, Audun Botterud |
IJCNN | 3 |
| 2010 | Self organizing maps with the correntropy induced metricabstractThe similarity measure popularly used in Kohonen's self organizing maps and several of its other variants is the mean square error (MSE). It is shown that this leads to, in information theoretic sense, a suboptimal solution of distributing the centers of the map. Here we show that using a similarity measure called the correntropy induced metric (CIM) can lead to a solution with better magnification of the input density. It provides an insight into how the type of the kernel effects the mapping and also under what condition is using SOM with CIM (SOM-CIM) can perform better than SOM with MSE. We also show that the use of this in clustering and data visualization can provide better results. Rakesh Chalasani, José C. Príncipe |
IJCNN | 2 |
| 2010 | Period detection in light curves from astronomical objects using correntropyabstractIn this paper we propose a new method for determining the period in astronomical time series using correntropy, an information theoretical concept recently developed in the computational intelligence field. The time series correspond to the stellar brightness over time, so-called light curves, and are characterized as being noisy and unevenly sampled. The advantages of using correntropy instead of correlation are to escape from the constraints of linearity and Gaussianity and are clearly demonstrated. The performance of the proposed method is compared with other algorithms published in the literature on a set of light curves drawn from the MACHO survey. The results show that the correntropy-based method obtains the correct periods more frequently than the Lomb-Scargle periodogram and the Period04 program. Pablo A. Estévez, Pablo Huijse, Pablo Zegers, José C. Príncipe, Pavlos Protopapas |
IJCNN | 4 |
| 2010 | Binary classification based on SVDD projection and nearest neighborsabstractThe SVDD (support vector data description) is one of the most well-known one-class support vector learning methods, in which one tries the strategy of utilizing balls defined on the feature space in order to distinguish a set of normal data from all other possible abnormal objects. The usual strategy of the SVDD depends on the process of finding the region for the normal-class training data with somewhat neglecting the distribution of the abnormal data, thus it may not work well if applied to the binary classification problems in which two classes have similar number of data. In this paper, we consider the problem of performing binary classification based on the SVDD techniques, and in order to overcome the possible drawback of the usual SVDD strategy focusing on the normal-class data only, we propose a new SVDD algorithm which is based on the use of two different SVDD balls for the positive and negative classes along with the SVDD projection and nearest neighbor rule. To investigate how the proposed method works, we compared the performance of the proposed method with SVC (support vector classifier) and conventional SVDD using several real datasets. Daesung Kang, Jooyoung Park 0001, José C. Príncipe |
IJCNN | 3 |
| 2010 | A loss function for classification based on a robust similarity metricabstractWe present a margin-based loss function for classification, inspired by the recently proposed similarity measure called correntropy. We show that correntropy induces a nonconvex loss function that is a closer approximation to the misclassification loss (ideal 0-1 loss). We show that the discriminant function obtained by optimizing the proposed loss function using a neural network is insensitive to outliers and has better generalization performance as compared to using the squared loss function which is common in neural network classifiers. The proposed method of training classifiers is a practical way of obtaining better results on real world classification problems, that uses a simple gradient based online training procedure for minimizing the empirical risk. José C. Príncipe |
IJCNN | 2 |
| 2010 | A novel family of non-parametric cumulative based divergences for point processesabstractHypothesis testing on point processes has several applications such as model fitting, plasticity detection, and non-stationarity detection. Standard tools for hypothesis testing include tests on mean firing rate and time varying rate function. However, these statistics do not fully describe a point process and thus the tests can be misleading. In this paper, we introduce a family of non-parametric divergence measures for hypothesis testing. We extend the traditional Kolmogorov--Smirnov and Cramer--von-Mises tests for point process via stratification. The proposed divergence measures compare the underlying probability structure and, thus, is zero if and only if the point processes are the same. This leads to a more robust test of hypothesis. We prove consistency and show that these measures can be efficiently estimated from data. We demonstrate an application of using the proposed divergence as a cost function to find optimally matched spike trains. Sohan Seth, Il Park 0002, Austin J. Brockmeier, Mulugeta Semework, John S. Choi, Joseph T. Francis, José C. Príncipe |
NIPS | 7 |
| 2010 | A comparison of binless spike train measures
António R. C. Paiva, Il Park 0002, José C. Príncipe |
Neural Comput. Appl. | 3 |
| 2010 | Modulation Selection from a Battery Power Efficiency PerspectiveabstractIn this paper, we compare the battery power efficiencies of various pulse-based modulations widely adopted for their low complexity. Taking into account circuit modules and battery imperfectness, we establish simple closed-form analytical formulas which can be used to conveniently determine the relative preference between arbitrary pulse-based modulation pairs in terms of their actual average battery energy consumption. Dongliang Duan, Fengzhong Qu, Liuqing Yang 0001, Ananthram Swami, José C. Príncipe |
IEEE Trans. Commun. | 5 |
| 2009 | Hierarchical clustering of neural data using Linked-Mixtures of Hidden Markov Models for Brain Machine InterfacesabstractIn this paper, we build upon previous brain machine interface (BMI) signal processing models that require a-priori knowledge about the patient's arm kinematics. Specifically, we propose an unsupervised hierarchical clustering model that attempts to discover both the interdependencies between neural channels and the self-organized clusters represented in the spatial-temporal neural data. Given that BMIs must work with disabled patients who lack arm kinematic information, the clustering work describe within this paper is very relevant for future BMIs. Shalom Darmanjian, José C. Príncipe |
ICASSP | 2 |
| 2009 | A new nonparametric measure of conditional independenceabstractIn this paper we propose a new measure of conditional independence that is loosely based on measuring the L2distance between the conditional joint and the product of the conditional marginal density functions. However, we propose to smooth the arguments prior to measuring the distance and use kernel density estimation to derive the estimator. We show that under suitable conditions the proposed smoothing does not affect the conditional independence but using proper smoothing function helps in choosing the bandwidth parameter robustly. We discuss the computational issues and propose an approximation to evaluate the estimator efficiently. We apply the proposed measure in different experiments to show its validity. Sohan Seth, Il Park 0002, José C. Príncipe |
ICASSP | 3 |
| 2009 | On speeding up computation in information theoretic learningabstractWith the recent progress in kernel based learning methods, computation with Gram matrices has received immense attention. However, the complexity of computing the entire Gram matrix is quadratic in terms of number of samples. Therefore, a considerable amount of work has been focused on extracting relevant information from the Gram matrix without accessing all the elements. Most of these methods exploits the positive definiteness and rapidly decaying eigenstructure of the Gram matrix. Although information theoretic learning (ITL) is conceptually different from kernel based learning, several ITL estimators can be written in terms of Gram matrices. However, the difference between ITL and kernel based methods is that a few ITL estimators include a special type of matrix which is neither positive definite nor symmetric. In this paper we discuss how the techniques applied in kernel based learning can be applied to reduce computational complexity of the ITL estimators involving both Gram matrices and these other matrices. Sohan Seth, José C. Príncipe |
IJCNN | 2 |
| 2009 | Using Correntropy as a cost function in linear adaptive filtersabstractCorrentropy has been recently defined as a localised similarity measure between two random variables, exploiting higher order moments of the data. This paper presents the use of correntropy as a cost function for minimizing the error between the desired signal and the output of an adaptive filter, in order to train the filter weights.We have shown that this cost function has the computational simplicity of the popular LMS algorithm, along with the robustness that is obtained by using higher order moments for error minimization. We apply this technique for system identification and noise cancellation configurations. The results demonstrate the advantages of the proposed cost function as compared to LMS algorithm, and the recently proposed minimum error entropy (MEE) cost function. José C. Príncipe |
IJCNN | 2 |
| 2009 | Selecting neural subsets for kinematics decoding by information theoretical analysis in motor Brain Machine InterfacesabstractPrevious decoding algorithms for Brain Machine Interfaces (BMIs) reconstruct the kinematics from recorded activities of hundreds of neurons, which are not all related to the movement task. Decoding from all neurons not only brings problem towards model generalization but also a significant computation burden. Knowledge of neural receptive fields helps ascertain the neuron importance associate with the movements. We propose to apply information theoretical analysis based on an instantaneous tuning model to extract the candidate neuron subsets, which also reduces the computation complexity for the decoding process. The cortical distribution of extracted neuron subsets is analyzed and the statistical decoding performances using neuron subset selection are compared to the one by the full neuron ensemble. Yiwen Wang 0002, Justin C. Sanchez, José C. Príncipe |
IJCNN | 3 |
| 2009 | A Biphasic Integrate-and-fire SystemabstractA neuronal recording system for brain-machine interfaces (BMI) based on asynchronous biphasic pulse coding is described. It demonstrates the first step in the development of a completely implanted wireless solution with a fully integrated circuit architecture. A recording experiment comparing, in parallel, a commercial recording system (Tucker-Davis technology (TDT)) and the UF's custom solution (FWIRE) is set up to compare performance. The novel aspect of the UF system is that the analog signal is represented by an asynchronous pulse train, which provides a low-power, low-bandwidth, noise-resistant means for coding and transmission. Taking advantage of neural firing features, the pulse-based approach uses only 3 K pulses/second to record a 25 kHz bandwidth signal from a hardware neural simulator. Sheng-Feng Yen, Jie Xu 0035, Manu Rastogi, John G. Harris, José C. Príncipe, Justin C. Sanchez |
ISCAS | 5 |
| 2009 | Correntropy Based Matched Filtering for Classification in Sidescan Sonar ImageryabstractThis paper presents an automated way of classifying mines in sidescan sonar imagery. A nonlinear extension to the matched filter is introduced using a new metric called correntropy. This method features high order moments in the decision statistic showing improvements in classification especially in the presence of noise. Templates have been designed using prior knowledge about the objects in the dataset. During classification, these templates are linearly transformed to accommodate for the shape variability in the observation. The template resulting in the largest correntropy cost function is chosen as the object category. The method is tested on real sonar images producing promising results considering the low number of images required to design the templates. Erion Hasanbelliu, José C. Príncipe, K. Clint Slatton |
SMC | 2 |
| 2009 | Prerequesites for Symbiotic Brain-Machine InterfacesabstractRecent advancements in the neuroscience and engineering of brain-machine interfaces are providing a blueprint for how new co-adaptive designs based on reinforcement learning change the nature of a user's ability to accomplish tasks that were not possible using static methodologies. By designing adaptive controls and artificial intelligence into the neural interface, computers can become active assistants in goal-directed behavior and further enhance human performance. This paper presents a set of minimal prerequisites that enable a cooperative symbiosis and dialogue between biological and artificial systems. Justin C. Sanchez, José C. Príncipe |
SMC | 2 |
| 2009 | Modulation selection from a battery power efficiency perspective: a case study of PPM and OOKabstractSensor nodes in wireless sensor networks (WSNs) are often expected to operate on batteries for a long period of time. Battery power efficiency (BPE) is therefore a critical factor dictating the lifetime of WSNs. In this paper, we aim to select the appropriate modulation scheme from a battery power efficiency perspective. Pulse position modulation (PPM) and on-off keying (OOK), as low-complexity pulse-based modulation schemes, are used for a case study of our methodology. The analysis is based on a general model that integrates typical WSN transmission and reception modules with a realistic nonlinear battery model. We first present the quantitative comparison results under general system design criteria. Then, we illustrate the comparisons with theoretical and numerical results under the bit error rate (BER) system design criterion. Dongliang Duan, Fengzhong Qu, Liuqing Yang 0001, Ananthram Swami, José C. Príncipe |
WCNC | 5 |
| 2009 | Cooperative diversity of spectrum sensing in cognitive radio networksabstractSpectrum sensing is a critical issue in cognitive radio networks. Cooperation among the secondary users is utilized to improve the performance of spectrum sensing. In this paper, we quantify the gain of cooperation in spectrum sensing by introducing the concept of diversity order. With different system performance metrics, we introduce different diversity quantities. We analyze the single-user sensing and the multi-user sensing with soft and hard information fusion strategies using the diversity quantities as our figure of merit. In particular, we discuss the selection of threshold in each sensing scheme with respect to the diversity performance and obtain the quantitative relationship among the diversity performance, threshold, and the number of cooperative users. We also observe and quantify the tradeoff between false alarm and missed detection performance in all spectrum sensing schemes. Dongliang Duan, Liuqing Yang 0001, José C. Príncipe |
WCNC | 3 |
| 2009 | A Reproducing Kernel Hilbert Space Framework for Spike Train Signal ProcessingabstractThis letter presents a general framework based on reproducing kernel Hilbert spaces (RKHS) to mathematically describe and manipulate spike trains. The main idea is the definition of inner products to allow spike train signal processing from basic principles while incorporating their statistical description as point processes. Moreover, because many inner products can be formulated, a particular definition can be crafted to best fit an application. These ideas are illustrated by the definition of a number of spike train inner products. To further elicit the advantages of the RKHS framework, a family of these inner products, the cross-intensity (CI) kernels, is analyzed in detail. This inner product family encapsulates the statistical description from the conditional intensity functions of spike trains. The problem of their estimation is also addressed. The simplest of the spike train kernels in this family provide an interesting perspective to others' work, as will be demonstrated in terms of spike train distance measures. Finally, as an application example, the RKHS framework is used to derive a clustering algorithm for spike trains from simple principles. António R. C. Paiva, Il Park 0002, José C. Príncipe |
Neural Comput. | 3 |
| 2009 | Sequential Monte Carlo Point-Process Estimation of Kinematics from Neural Spiking Activity for Brain-Machine InterfacesabstractMany decoding algorithms for brain machine interfaces' (BMIs) estimate hand movement from binned spike rates, which do not fully exploit the resolution contained in spike timing and may exclude rich neural dynamics from the modeling. More recently, an adaptive filtering method based on a Bayesian approach to reconstruct the neural state from the observed spike times has been proposed. However, it assumes and propagates a gaussian distributed state posterior density, which in general is too restrictive. We have also proposed a sequential Monte Carlo estimation methodology to reconstruct the kinematic states directly from the multichannel spike trains. This letter presents a systematic testing of this algorithm in a simulated neural spike train decoding experiment and then in BMI data. Compared to a point-process adaptive filtering algorithm with a linear observation model and a gaussian approximation (the counterpart for point processes of the Kalman filter), our sequential Monte Carlo estimation methodology exploits a detailed encoding model (tuning function) derived for each neuron from training data. However, this added complexity is translated into higher performance with real data. To deal with the intrinsic spike randomness in online modeling, several synthetic spike trains are generated from the intensity function estimated from the neurons and utilized as extra model inputs in an attempt to decrease the variance in the kinematic predictions. The performance of the sequential Monte Carlo estimation methodology augmented with this synthetic spike input provides improved reconstruction, which raises interesting questions and helps explain the overall modeling requirements better. Yiwen Wang 0002, António R. C. Paiva, José C. Príncipe, Justin C. Sanchez |
Neural Comput. | 3 |
| 2009 | Mapping broadband electrocorticographic recordings to two-dimensional hand trajectories in humans: Motor control features
Aysegul Gunduz, Justin C. Sanchez, Paul R. Carney, José C. Príncipe |
Neural Networks | 4 |
| 2009 | Exploiting co-adaptation for the design of symbiotic neuroprosthetic assistants
Justin C. Sanchez, Babak Mahmoudi, Jack DiGiovanna, José C. Príncipe |
Neural Networks | 4 |
| 2009 | Ascertaining neuron importance by information theoretical analysis in motor Brain-Machine Interfaces
Yiwen Wang 0002, José C. Príncipe, Justin C. Sanchez |
Neural Networks | 2 |
| 2009 | The correntropy MACE filter
Kyu-Hwa Jeong, Weifeng Liu 0016, Seungju Han 0001, Erion Hasanbelliu, José C. Príncipe |
Pattern Recognit. | 5 |
| 2009 | Mean shift: An information theoretic perspective
Sudhir Rao, Allan de Medeiros Martins, José C. Príncipe |
Pattern Recognit. Lett. | 3 |
| 2009 | Correntropy as a novel measure for nonlinearity tests
Aysegul Gunduz, José C. Príncipe |
Signal Process. | 2 |
| 2009 | Kernel least mean square algorithm with constrained growth
Puskal P. Pokharel, Weifeng Liu 0016, José C. Príncipe |
Signal Process. | 3 |
| 2009 | A low complexity robust detector in impulsive noise
Puskal P. Pokharel, Weifeng Liu 0016, José C. Príncipe |
Signal Process. | 3 |
| 2009 | An Information Theoretic Approach of Designing Sparse Kernel Adaptive FiltersabstractThis paper discusses an information theoretic approach of designing sparse kernel adaptive filters. To determine useful data to be learned and remove redundant ones, a subjective information measure called surprise is introduced. Surprise captures the amount of information a datum contains which is transferable to a learning system. Based on this concept, we propose a systematic sparsification scheme, which can drastically reduce the time and space complexity without harming the performance of kernel adaptive filters. Nonlinear regression, short term chaotic time-series prediction, and long term time-series forecasting examples are presented. Weifeng Liu 0016, Il Park 0002, José C. Príncipe |
IEEE Trans. Neural Networks | 3 |
| 2008 | Reproducing kernel Hilbert spaces for spike train analysisabstractThis paper introduces a generalized cross-correlation (GCC) measure for spike train analysis derived from reproducing kernel Hilbert spaces (RKHS) theory. An estimator for GCC is derived that does not depend on binning or a specific kernel and it operates directly and efficiently on spike times. For instantaneous analysis as required for real-time use, an instantaneous estimator is proposed and proved to yield the GCC on average. We finalize with two experiments illustrating the usefulness of the techniques derived. António R. C. Paiva, Il Park 0002, José C. Príncipe |
ICASSP | 3 |
| 2008 | Correntropy based Granger causalityabstractWe propose a novel nonlinear extension to Granger causality. It is derived from a nonlinear mapping of a stochastic process using the recently introduced generalized correlation measure called correntropy. The method is demonstrated by detecting the direction of coupling in a chaotic system where the original Granger causality failed. Il Park 0002, José C. Príncipe |
ICASSP | 2 |
| 2008 | Compressed signal reconstruction using the correntropy induced metricabstractRecovering a sparse signal from insufficient number of measurements has become a popular area of research under the name of compressed sensing or compressive sampling. The reconstruction algorithm of compressed sensing tries to find the sparsest vector (minimum lo-norm) satisfying a series of linear constraints. However, lo-norm minimization, being a NP hard problem is replaced by li-norm minimization with the cost of higher number of measurements in the sampling process. In this paper we propose to minimize an approximation of lo-norm to reduce the required number of measurements. We use the recently introduced correntropy induced metric (CIM) as an approximation of lo-norm, which is also a novel application of CIM. We show that by reducing the kernel size appropriately we can approximate the lo-norm, theoretically, with arbitary accuracy. Sohan Seth, José C. Príncipe |
ICASSP | 2 |
| 2008 | The wellposedness analysis of the kernel adalineabstractIn this paper, we investigate the wellposedness of the kernel adaline. The kernel adaline finds the linear coefficients in a radial basis function network using deterministic gradient descent. We will show that the gradient descent provides an inherent regularization as long as the training is properly early-stopped. Along with other popular regularization techniques, this result is investigated in a unifying regularization-function concept. This understanding provides an alternative and possibly simpler way to obtain regularized solutions comparing with the cross-validation approach in regularization networks. Weifeng Liu 0016, José C. Príncipe |
IJCNN | 2 |
| 2008 | Pulse-based signal compression for implanted neural recording systemsabstractToday's implanted neural systems are bound by tight constraints on power and communication bandwidth. Most conventional ADC-based approaches fall into two categories. Either they transmit all of the information at the Nyquist rate but are ultimately limited to only a handful of channels due to communication bandwidth constraints. Or they perform spike detection on the front-end which allows a scale up to 100 or more channels but prevents the use of spike sorting on the backend. Spike sorting is an important step that provides a labeling to multiple neurons on each channel and further improves the accuracy of spike detection. In this paper we describe the pulse-based approach used in the FWIRE (Florida wireless implantable recording electrodes) project. A hardware spiking neuron on each channel is configured either to transmit pulses for full reconstruction on the back-end, or to transmit dramatically fewer pulses but still allow for spike sorting on the back-end. Spike sorting results show that the pulse-based spike sorting accuracy is competitive with conventional methods used in daily practice. John G. Harris, José C. Príncipe, Justin C. Sanchez, Christy She |
ISCAS | 2 |
| 2008 | Real time signal reconstruction from spikes on a digital signal processorabstractThere has been significant interest in encoding information into asynchronous timing events or spikes using models of biological neurons. It has been shown that these neurons can act as lossless encoders and software algorithms can reconstruct the encoded information. We present here, a real time implementation of a reconstruction algorithm on a TI TMS320C6713 DSK board. This implementation reconstructs the signal from a spike train which was encoded by a fabricated neuron chip. A signal to error ratio of 32 dB with an average firing rate of 8 KHz is obtained. We discuss the practical implementation, challenges and possible future directions. John G. Harris, Jie Xu 0035, Manu Rastogi, Alexander Singh-Alvarado, José C. Príncipe, Kalyana Vuppamandla |
ISCAS | 6 |
| 2008 | Enhancing the correntropy MACE filter with random projections
Kyu-Hwa Jeong, José C. Príncipe |
Neurocomputing | 2 |
| 2008 | A new nonlinear similarity measure for multichannel signals
Jianwu Xu, Hovagim Bakardjian, Andrzej Cichocki, José C. Príncipe |
Neural Networks | 4 |
| 2008 | Recursive complex BSS via generalized eigendecomposition and application in image rejection for BPSK
Puskal P. Pokharel, Umut Ozertem, Deniz Erdogmus, José C. Príncipe |
Signal Process. | 4 |
| 2008 | A Pitch Detector Based on a Generalized Correlation FunctionabstractThis paper proposes a novel pitch determination algorithm (PDA) based on the newly introduced concept of a generalized correlation function called correntropy. Correntropy is a positive definite kernel function which implicitly transforms the original signal into a high-dimensional reproducing kernel Hilbert space (RKHS) in a nonlinear way, and calculates very efficiently the generalized correlation in that RKHS. By incorporating the kernel function, correntropy is able to utilize higher order statistics to enhance the resolution of pitch estimation. The proposed PDA computes the summary of correntropy functions from the outputs of an equivalent rectangular bandwidth (ERB) filter bank. We present simulations on pitch determination for a single vowel, double vowels, and a benchmark database test. Simulations show that correntropy exhibits much better resolution than conventional autocorrelation in pitch determination and outperforms other PDAs in the benchmark database test. Jianwu Xu, José C. Príncipe |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Robust switching blind equalizer for wireless cognitive receiversabstractKurtosis minimization has been applied for the existing blind equalization schemes but the corresponding optimization procedures are very sensitive to the channel conditions and the initial conditions. In this paper, we introduce a new cognitive receiver front-end, which includes a novel switching blind equalizer and an automatic modulation classifier. We design a switching criterion based on the kurtosis/normalized moment ratio threshold to select the better signal between the raw data and the equalized sequence. Simulations demonstrate that our proposed robust switching blind equalization scheme can significantly outperform the existing blind equalizer and would not degrade the subsequent modulation classification accuracy. Hsiao-Chun Wu, Yiyan Wu 0001, José C. Príncipe, Xianbin Wang 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2007 | A New Information Theoretic Measure of PDF SymmetryabstractIn this paper, a new quantity called symmetric information potential (SIP) is proposed to measure the reflection symmetry and to estimate the location parameter of probability density functions. SIP is defined as an inner product in the probability density function space and has a close relation to information theoretic learning. A simple nonparametric estimator directly from the data exists. Experiments demonstrate that this concept can be very useful dealing with impulsive data distributions, in particular, α-stable distributions. Weifeng Liu 0016, Puskal P. Pokharel, José C. Príncipe |
ICASSP (2) | 3 |
| 2007 | The Fast Correntropy Mace FilterabstractIn this paper, we implement the newly introduced correntropy MACE filter using the fast Gauss transform (FGT). The correntropy MACE filter is a nonlinear extension to the MACE filter using the correntropy function in a feature space nonlinearly related to the input. The correntropy MACE outperforms the traditional linear MACE in both generalization and rejection abilities. However, in practice, the drawback of the correntropy MACE filter is its computation complexity. This paper present a fast version of the correntropy MACE by using the FGT idea and validates the approximation with results in synthetic aperture radar (SAR) image recognition. Kyu-Hwa Jeong, Seungju Han 0001, José C. Príncipe |
ICASSP (2) | 3 |
| 2007 | Kernel LMSabstractIn this paper a nonlinear adaptive algorithm based on a kernel space least mean squares (LMS) approach is presented. With most of the neural network based methods for time series modeling it is difficult to implement a sample-by-sample adaptation method. This puts a serious limitation on the applicability of adaptive nonlinear filters in many optimal signal processing and communication applications where data arrives sequentially. This paper shows that the kernel LMS algorithm provides a computational simple and an effective algorithm to train nonlinear systems for system modeling without the need for regularization, without convergence to local minima and without the need for a separate book of data as a training set. Puskal P. Pokharel, Weifeng Liu 0016, José C. Príncipe |
ICASSP (3) | 3 |
| 2007 | Recursive Complex Blind Source Separation via Eigendecomposition of Cumulant MatricesabstractUnder the assumptions of non-Gaussian, non-stationary, or non-white independent sources, linear blind source separation can be formulated as a generalized eigenvalue decomposition problem. Here we provide an elegant method of doing this online, instead of waiting for a sufficiently large batch of data. This is done through a recursive generalized eigendecomposition algorithm that tracks the optimal solution, which is obtained using all the data observed. The algorithms proposed in this paper follow the well-known recursive least squares (RLS) algorithm in nature. Puskal P. Pokharel, Umut Ozertem, Deniz Erdogmus, José C. Príncipe |
ICASSP (2) | 4 |
| 2007 | Hierarchal Decomposition of Neural Data using Boosted Mixtures of Hidden Markov Chains and its application to a BMIabstractIn this paper, we propose a simple algorithm that takes multidimensional neural input data and decomposes the joint likelihood into marginals using boosted mixtures of hidden Markov chains (BM-HMM). The algorithm applies techniques from boosting to create hierarchal dependencies between these marginal subspaces. Finally, borrowing ideas from mixture of experts, the local information is weighted and incorporated into an ensemble decision. Our results show that this algorithm is very simple to train and computationally efficient, while also providing the ability to reduce the input dimensionality for brain machine interfaces (BMIs). Shalom Darmanjian, António R. C. Paiva, José C. Príncipe, Justin C. Sanchez |
IJCNN | 3 |
| 2007 | An Associative Memory Readout in ESN for Neural Action Potential DetectionabstractThis paper describes how echo state networks (ESN) can be used in conjunction with minimum average correlation energy (MACE) filters in order to create a system that can identify spikes in neural recordings. Various experiments using real-world data were used to compare the performance of the ESN-MACE against threshold and matched filter detectors to ascertain the capabilities of such a system in detecting neural action potentials. The experiments demonstrate that the ESN-MACE can correctly detect spikes with lower false alarm rates than established detection techniques since it captures the inherent variability and the covariance information in spike shapes by training. Nicolas J. Dedual, Mustafa C. Ozturk, Justin C. Sanchez, José C. Príncipe |
IJCNN | 4 |
| 2007 | A Novel Switching Scheme Between Adaptive Information AlgorithmsabstractSwitching approaches can improve the performance of adaptive schemes, however a data driven criterion to accomplish the task is unclear. In this paper, we propose a new optimization criterion for switching which is estimated directly from data. We apply the method to the recently introduced MEE and MEE-SAS algorithms. Using this novel switching scheme, we develop a single algorithm which effectively combines the strengths of MEE and MEE-SAS without sacrificing the simplicity and stability properties of MEE. We explain these results analytically, and through simulations. Seungju Han 0001, Sudhir Rao, Deniz Erdogmus, José C. Príncipe |
IJCNN | 4 |
| 2007 | Spectral Clustering of Synchronous Spike TrainsabstractIn this paper a clustering algorithm that learns the groups of synchronized spike trains directly from data is proposed. Clustering of spike trains based on the presence of synchronous neural activity is of high relevance in neurophys-iological studies. In this context such activity is thought to be associated with functional structures in the brain. In addition, clustering has the potential to analyze large volumes of data. The algorithm couples a distance between two spike trains recently proposed in the literature with spectral clustering. Finally, the algorithm is illustrated in sets of computer generated spike trains and analyzed for the dependence on its parameters and accuracy with respect to features of interest. António R. C. Paiva, Sudhir Rao, Il Park 0002, José C. Príncipe |
IJCNN | 4 |
| 2007 | A Closed Form Solution for Multiple-Input Spike Based Adaptive FiltersabstractNeurons are point process systems, in the sense that the inputs and output which are spike trains can be treated as point processes. System identification of a point process system has been studied mostly with single input. However, multiple input is required in many applications such as liquid state machines or neural prosthetics. We propose a simple multiple-input spike based adaptive filter which is based on an integrate-and-fire neuron model. The optimal closed solution is derived, and the performance is analyzed with respect to noise in various parameters and measurement. Il Park 0002, António R. C. Paiva, José C. Príncipe, John G. Harris |
IJCNN | 3 |
| 2007 | Information Theoretic Vector Quantization with Fixed Point UpdatesabstractIn this paper, we revisit information theoretic vector quantization (ITVQ) algorithm introduced in (T. Lehn-Schioler et al., 2005) and make it practical. We derive a fixed point update rule to minimize the Cauchy-Schwartz(CS) pdf divergence between the set of codewords and the actual data. In doing so, we overcome two severe deficiencies of the previous gradient based method namely, the number of parameters to be optimized and slow convergence rate, thus making this algorithm more efficient and useful as a compression algorithm. Sudhir Rao, Seungju Han 0001, José C. Príncipe |
IJCNN | 3 |
| 2007 | A Novel Weighted LBG Algorithm for Neural Spike CompressionabstractIn this paper, we present a weighted Linde-Buzo-Gray algorithm (WLBG) as a powerful and efficient technique for compressing neural spike data. We compare this technique with the recently proposed self-organizing map with dynamic learning (SOM-DL) and the traditional SOM. A significant achievement of WLBG over SOM-DL is a 15 dB increase in the SNR of the spike data apart from having a compression ratio of 150 : 1. Being simple and extremely fast, this algorithm allows real-time implementation on DSP chips opening new opportunities in BMI applications. Sudhir Rao, António R. C. Paiva, José C. Príncipe |
IJCNN | 3 |
| 2007 | Water Inflow Forecasting using the Echo State Network: a Brazilian Case StudyabstractA type of recurrent neural network has been proposed by H. Jaeger. This model, called Echo State Network (ESN), possesses a highly interconnected and recurrent topology of nonlinear processing elements, which constitutes a "reservoir of rich dynamics" and contains information about the history of input or/and output patterns. The interesting property of ESN is that only the memoryless readout is trained, whereas the recurrent topology has fixed connection weights. This reduces the complexity of recurrent neural network training to simple linear regression while preserving a recurrent topology. In this paper, the ESN is used to forecast hydropower plant reservoir water inflow, which is a fundamental information to the hydrothermal power system operation planning. A database of average monthly water inflows of Furnas plant, one of the Brazilian hydropower plants, was used as source of training and test data. The performance of the ESN is compared with SONARX network, RBF network and ANFIS model. The results show that the Echo State Network provides pretty good results for one-step ahead water inflow forecasting, providing a valuable information for the system operator. Rodrigo Sacchi, Mustafa C. Ozturk, José C. Príncipe, Adriano A. F. Carneiro, Ivan Nunes da Silva |
IJCNN | 3 |
| 2007 | A Monte Carlo Sequential Estimation of Point Process Optimum Filtering for Brain Machine InterfacesabstractThe previous decoding algorithms for brain machine interfaces are normally utilized to estimate animal's movement from binned spike rates, which loses spike timing resolution and may exclude rich neural dynamics due to single spikes. Based on recently proposed Monte Carlo sequential estimation algorithm on point process, we present a decoding framework to reconstruct the kinematic states directly from the multi-channel spike trains. Starting with analysis on the differences between the simulation and real BMI data, neural tuning properties are modeled to encode the movement information of the experimental primate as the pre-knowledge for Monte-Carlo sequential estimation for BMI. The preliminary kinematics reconstruction shows better results when compared with Kalman filter. Yiwen Wang 0002, António R. C. Paiva, José C. Príncipe, Justin C. Sanchez |
IJCNN | 3 |
| 2007 | A New Nonlinear Similarity Measure for Multichannel Biological SignalsabstractWe propose a novel similarity measure, called the correntropy coefficient, sensitive to higher order moments of the signal statistics based on a similarity function called crosscorrentopy. Crossorrentropy nonlinearly maps the original time series into a high-dimensional reproducing kernel Hilbert space (RKHS). The correntropy coefficient computes the cosine of the angle between the transformed vectors. Preliminary experiments with simulated data and multichannel electroencephalogram (EEG) signals during behavior studies elucidate the performance of the new measure versus the well established correlation coefficient. Jianwu Xu, Hovagim Bakardjian, Andrzej Cichocki, José C. Príncipe |
IJCNN | 4 |
| 2007 | Florida Wireless Implantable Recording Electrodes (FWIRE) for Brain Machine InterfacesabstractThis paper reviews on-going efforts towards the development of the Florida wireless implantable recording electrodes (FWIRE). The FWIRE microsystem platform is a fully implantable flexible substrate microelectrode array that employs state-of-the-art integrate-and-fire (IF) signal representation and wireless interface circuitry for recording neural activity from behaving rodents. The modular nature of the implantable neural recording electrode allows future enhancements to be seamlessly added to improve functionality including but not limited to rechargeability through inductive coupling, custom microelectrode arrays, higher capacity batteries, and more advanced integrated circuit technologies. This paper concentrates on custom integrated circuits such as neural interfacing amplifiers, baseband signal processing, wireless data and power interfaces, and battery management system. Rizwan Bashirullah, John G. Harris, Justin C. Sanchez, Toshikazu Nishida, José C. Príncipe |
ISCAS | 5 |
| 2007 | Quasi-sliding mode control strategy based on multiple-linear models
Jeongho Cho, José C. Príncipe, Deniz Erdogmus, Mark A. Motter |
Neurocomputing | 2 |
| 2007 | Analysis and Design of Echo State NetworksabstractThe design of echo state network (ESN) parameters relies on the selection of the maximum eigenvalue of the linearized system around zero (spectral radius). However, this procedure does not quantify in a systematic manner the performance of the ESN in terms of approximation error. This article presents a functional space approximation framework to better understand the operation of ESNs and proposes an information-theoretic metric, the average entropy of echo states, to assess the richness of the ESN dynamics. Furthermore, it provides an interpretation of the ESN dynamics rooted in system theory as families of coupled linearized systems whose poles move according to the input signal dynamics. With this interpretation, a design methodology for functional approximation is put forward where ESNs are designed with uniform pole distributions covering the frequency spectrum to abide by the richness metric, irrespective of the spectral radius. A single bias parameter at the ESN input, adapted with the modeling error, configures the ESN spectral radius to the input-output joint space. Function approximation examples compare the proposed design methodology versus the conventional design. Mustafa C. Ozturk, José C. Príncipe |
Neural Comput. | 3 |
| 2007 | Self-organizing maps with dynamic learning for signal reconstruction
Jeongho Cho, António R. C. Paiva, Sung-Phil Kim, Justin C. Sanchez, José C. Príncipe |
Neural Networks | 5 |
| 2007 | Special issue on echo state networks and liquid state machines
Herbert Jaeger, Wolfgang Maass 0001, José C. Príncipe |
Neural Networks | 3 |
| 2007 | An associative memory readout for ESNs with applications to dynamical pattern recognition
Mustafa C. Ozturk, José C. Príncipe |
Neural Networks | 2 |
| 2007 | Information cut for clustering using a gradient descent approach
Robert Jenssen, Deniz Erdogmus, Kenneth E. Hild II, José C. Príncipe, Torbjørn Eltoft |
Pattern Recognit. | 4 |
| 2007 | Independent Component Analysis and Blind Source Separation
Allan Kardec Barros, José C. Príncipe, Deniz Erdogmus |
Signal Process. | 2 |
| 2007 | A minimum-error entropy criterion with self-adjusting step-size (MEE-SAS)
Seungju Han 0001, Sudhir Rao, Deniz Erdogmus, Kyu-Hwa Jeong, José C. Príncipe |
Signal Process. | 5 |
| 2007 | A unifying criterion for instantaneous blind source separation based on correntropy
Ruijiang Li, Weifeng Liu 0016, José C. Príncipe |
Signal Process. | 3 |
| 2006 | Robust Blind Simo Channel Estimation Using AdatronabstractIn this paper we apply the structural risk minimization (SRM) principle to derive a blind single-input multiple-output (SIMO) channel estimation algorithm, which is robust to channel order overestimation. Specifically, the blind estimation is formulated as a support vector regression (SVR) problem in which the channel coefficients are the Lagrange multipliers of the dual problem. In this paper, we show that the SRM principle pushes to zero the small leading and trailing terms of the channel impulse response even when its order is highly overestimated. The main drawback of this approach is the high computational cost of the resulting quadratic programming (QP) problem. To alleviate this, in this paper we propose to use a simple and fast algorithm called the Adatron to solve the QP problem. Simulation results are provided to demonstrate the performance of our channel estimator. Dongho Han, José C. Príncipe, Liuqing Yang 0001, Ignacio Santamaría, Javier Vía |
ICASSP (4) | 2 |
| 2006 | A Normalized Minimum Error Entropy Stochastic AlgorithmabstractWe propose in this paper the normalized Minimum Error Entropy (NMEE). Following the same rational that lead to the normalized LMS, the weight update adjustment for Minimum Error Entropy (MEE) is constrained by the principle of minimum disturbance. Unexpectedly, we obtained an algorithm that not only is insensitive to the power of the input, but is also faster than the MEE for the same misadjustment, and also that is less sensitive to the kernel size. We explain these results analytically, and through system identification simulations. Seungju Han 0001, Sudhir Rao, Kyu-Hwa Jeong, José C. Príncipe |
ICASSP (5) | 4 |
| 2006 | Kernel Based Synthetic Discriminant Function for Object RecognitionabstractIn this paper a non-linear extension to the synthetic discriminant function (SDF) is proposed. The SDF is a well known 2-D correlation filter for object recognition. The proposed non-linear version of the SDF is derived from kernel-based learning. The kernel SDF is implemented in a nonlinear high dimensional space by using the kernel trick and it can improve the performance of the linear SDF by incorporating the image's class higher order moments. We show that this kernelized composite correlation filter has an intrinsic connection with the recently proposed correntropy function. We apply this kernel SDF to face recognition and simulations show that the kernel SDF significantly outperforms the traditional SDF as well as is robust in noisy data environments. Kyu-Hwa Jeong, Puskal P. Pokharel, Jianwu Xu, Seungju Han 0001, José C. Príncipe |
ICASSP (5) | 5 |
| 2006 | A Closed Form Solution for a Nonlinear Wiener FilterabstractIn this paper a nonlinear extension to the Wiener filter is presented. A direct approach has been devised of replacing the autocorrelation function with a novel function called correntropy, derived from ideas on kernel-based learning theory and information theoretic learning. The linear Wiener filter, widely used because of its simplicity and optimality for linear systems and Gaussian distribution, is no longer effective when dealing with nonlinear time series data. The proposed method incorporates higher order moments in the general form of autocorrelation and improves upon the linear filter. Moreover, the computation cost is still lower than some kernel based methods and has a closed form solution to the problem unlike neural network based methods. Puskal P. Pokharel, Jianwu Xu, Deniz Erdogmus, José C. Príncipe |
ICASSP (3) | 4 |
| 2006 | Spike Sorting Using non Parametric Clustering VIA Cauchy Schwartz PDF DivergenceabstractWe propose a new method of clustering neural spike waveforms for spike sorting. After detecting the spikes using a threshold detector, we use principal component analysis (PCA) to get the first few PCA components of the data. Clustering on these PCA components is achieved by maximizing the Cauchy Schwartz PDF divergence measure which uses the Parzen window method to non parametrically estimate the pdf of the clusters. Comparison with other clustering techniques in spike sorting like k-means and Gaussian mixture elucidates the superiority of our method in terms of classification results and computational complexity. Sudhir Rao, Justin C. Sanchez, Seungju Han 0001, José C. Príncipe |
ICASSP (5) | 4 |
| 2006 | An Explicit Construction Of A Reproducing Gaussian Kernel Hilbert SpaceabstractIn this paper, we propose a method to explicitly construct a reproducing kernel Hilbert space (RKHS) associated with a Gaussian kernel by means of polynomial spaces. In contrast to the conventional Mercer's theorem approach that implicitly defines kernels by an eigendecomposition, the functionals in this reproducing kernel Hilbert space are explicitly constructed and are not necessary orthonormal. We also point out an intriguing connection between this reproducing kernel Hilbert space and a generalized Fock space. We give an experimental result on approximation of the constructed kernel to a Gaussian kernel. Jianwu Xu, Puskal P. Pokharel, Kyu-Hwa Jeong, José C. Príncipe |
ICASSP (5) | 4 |
| 2006 | Arm Motion Reconstruction via Feature Clustering in Joint Angle SpaceabstractWe hypothesize that a set of movements can be used to reconstruct biomechanically realistic movements. Using parameters from a reaching and grasping task we create a representative three-dimensional motion. From this motion we extract features from the joint angle space. We believe that the physiological importance of these features makes them worth investigating as possible movements. Machine learning techniques are employed to cluster similar features. The clusters are then used to recursively reconstruct the motion trajectory. Even with only twenty clusters, the average trajectory reconstruction error in Cartesian space is less than 1% of the dynamic range of motion. Our ability to create and analyze realistic motions may be crucial to both future BMI experiments where a desired signal is not available and our understanding of motor control. Jack DiGiovanna, Justin C. Sanchez, B. J. Fregly, José C. Príncipe |
IJCNN | 4 |
| 2006 | Correntropy as a Novel Measure for Nonlinearity TestsabstractStatistical tests have become an essential step in nonlinear system modeling due to the complexities involved in their analysis. Correntropy is a kernel-based similarity measure which includes the information of both distribution and time structure of a stochastic process. The correntropy function's capability of preserving nonlinear characteristics and high order moments makes it a suitable candidate as a statistic for determining whether a nonlinear structure exists within the system that created the observed time series. Experiments based on surrogate data methods have confirmed that correntropy can be employed as a discriminant measure for detecting nonlinear characteristics in time series. Aysegul Gunduz, Anant Hegde, José C. Príncipe |
IJCNN | 3 |
| 2006 | Information Theoretic Angle-Based Spectral Clustering: A Theoretical Analysis and an AlgorithmabstractRecent work has revealed a close connection between certain information theoretic divergence measures and properties of Mercer kernel feature spaces. Specifically, it has been proposed that an information theoretic measure may be used as a cost function for clustering in a kernel space, approximated by the spectral properties of the Laplacian matrix. In this paper we extend this result to other kernel matrices. We develop an algorithm for the actual clustering which is based on comparing angles between data points, and demonstrate that the proposed method performs equally good as a state-of-the art spectral clustering method. We point out some drawbacks of spectral clustering related to outliers, and suggest measures to be taken. Robert Jenssen, Deniz Erdogmus, José C. Príncipe |
IJCNN | 3 |
| 2006 | Correntropy: A Localized Similarity MeasureabstractThe measure of similarity normally utilized in statistical signal processing is based on second order moments. In this paper, we reveal the probabilistic meaning of correntropy as a new localized similarity measure based on information theoretic learning (ITL) and kernel methods. As such it has vastly different properties when compared with mean square error (MSE) that can be very useful in nonlinear, non-Gaussian signal processing. Two examples are presented to illustrate the technique. Weifeng Liu 0016, Puskal P. Pokharel, José C. Príncipe |
IJCNN | 3 |
| 2006 | Modeling of Synchronized Burst in Dissociated Cortical Tissue: An Exploration of Parameter SpaceabstractIn vitro neuronal culture is a potentially powerful tool for investigation of information processing of neurons. However, in a mature dissociated neuronal culture, synchronized bursting is often observed. Bursting strongly drives the dynamics of the network, thus making it difficult to investigate the computational properties. A quantitative modeling of the dynamics would provide a means for understanding the processes that produce bursts and a means to manipulate the dynamics to control them. A simulated network composed of leaky integrate and Are neurons and dynamic synapses is used to model the phenomenon. We investigate the variability of burst dynamics according to parameters such as connectivity, recovery time constant, and noise. Il Park 0002, Thomas B. DeMarse, José C. Príncipe |
IJCNN | 4 |
| 2006 | A Monte Carlo Sequential Estimation for Point Process Optimum FilteringabstractAdaptive filtering is normally utilized to estimate system states or outputs from continuous valued observations, and it is of limited use when the observations are discrete events. Recently a Bayesian approach to reconstruct the state from the discrete point observations has been proposed. However, it assumes the posterior density of the state given the observations is Gaussian distributed, which is in general restrictive. We propose a Monte Carlo sequential estimation methodology to estimate directly the posterior density. Sample observations are generated at each time to recursively evaluate the posterior density more accurately. The state estimation is obtained easily by collapse, i.e. by smoothing the posterior density with Gaussian kernels to estimate its mean. The algorithm is tested in a simulated neural spike train decoding experiment and reconstructs better the velocity when compared with point process adaptive filtering algorithm with the Gaussian assumption. Yiwen Wang 0002, António R. C. Paiva, José C. Príncipe |
IJCNN | 3 |
| 2006 | Nonlinear Component Analysis Based on CorrentropyabstractIn this paper, we propose a new nonlinear prin- cipal component analysis based on a generalized correlation function which we call correntropy. The data is nonlinearly transformed to a feature space, and the principal directions are found by eigen-decomposition of the correntropy matrix, which has the same dimension as the standard covariance matrix for the original input data. The correntropy matrix characterizes the nonlinear correlations between the data. With the correntropy function, one can efficiently compute the principal components in the feature space by projecting the transformed data onto those principal directions. We give the derivation of the new method and present simulation results. Jianwu Xu, Puskal P. Pokharel, António R. C. Paiva, José C. Príncipe |
IJCNN | 4 |
| 2006 | Asynchronous biphasic pulse signal coding and its CMOS realizationabstractAn asynchronous biphasic pulse train signal representation mechanism is proposed for low bandwidth, low power sensor applications. We prove that this novel pulse coding method provides lossless encoding theoretically. We also present a CMOS circuit realization using the AMI 0.6 mum process. Cadence simulation results show that the reconstructed signal has 8 effective bits of resolution with less than 100 muW of total power consumption John G. Harris, José C. Príncipe |
ISCAS | 5 |
| 2006 | A low power battery management system for rechargeable wireless implantable electronicsabstractAn integrated battery management system for low power wireless implantable electronics is presented herein. The system is designed for miniature implantable Li-ion rechargeable batteries with limited cell capacity and employs a new control loop that relaxes comparator resolution requirements, provides simultaneous operation of constant-current and constant voltage loops, and eliminates the external current sense resistor from the charging path. The accuracy of the end-of-charge detection is primarily determined by the voltage drop across matched resistors and current-sources and the offset voltage of the sense comparator. Preliminary simulations in 0.5mum bulk CMOS technology indicate that plusmn1% (or plusmn15muA) end-of-charge accuracy can be obtained under worst-case conditions for a comparator offset voltage of plusmn5mV. The proposed battery control loop measures roughly 170mum by 960mum and dissipates approximately 150muWatts Pengfei Li 0001, Rizwan Bashirullah, José C. Príncipe |
ISCAS | 3 |
| 2006 | Computation in a reduced KII network based on synchronizationabstractWhen implementing dynamical systems for information processing, a well-defined set of computational primitives should be constructed to represent the state space. Based on the specific collective behavior of Freeman's computational model of the olfactory cortex, we study the computation of reduced KII networks by using the synchronization of output channels. The design of coupling coefficients in the network for obtaining a desired output state response is presented. We demonstrate the computational power of reduced KII networks by two applications that address logic computation and associative memory, respectively. © 2006 Wiley Periodicals, Inc. Int J Int Syst 21: 919–935, 2006. José C. Príncipe |
Int. J. Intell. Syst. | 2 |
| 2006 | Using non-linear even functions for error minimization in adaptive filters
Allan Kardec Barros, José C. Príncipe, Yoshinori Takeuchi, Noboru Ohnishi |
Neurocomputing | 2 |
| 2006 | Feature Extraction Using Information-Theoretic LearningabstractA classification system typically consists of both a feature extractor (preprocessor) and a classifier. These two components can be trained either independently or simultaneously. The former option has an implementation advantage since the extractor need only be trained once for use with any classifier, whereas the latter has an advantage since it can be used to minimize classification error directly. Certain criteria, such as Minimum Classification Error, are better suited for simultaneous training, whereas other criteria, such as Mutual Information, are amenable for training the feature extractor either independently or simultaneously. Herein, an information-theoretic criterion is introduced and is evaluated for training the extractor independently of the classifier. The proposed method uses nonparametric estimation of Renyi's entropy to train the extractor by maximizing an approximation of the mutual information between the class labels and the output of the feature extractor. The evaluations show that the proposed method, even though it uses independent training, performs at least as well as three feature extraction methods that train the extractor and classifier simultaneously. Kenneth E. Hild II, Deniz Erdogmus, Kari Torkkola, José C. Príncipe |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2006 | An analysis of entropy estimators for blind source separation
Kenneth E. Hild II, Deniz Erdogmus, José C. Príncipe |
Signal Process. | 3 |
| 2006 | Modeling and inverse controller design for an unmanned aerial vehicle based on the self-organizing mapabstractThe next generation of aircraft will have dynamics that vary considerably over the operating regime. A single controller will have difficulty to meet the design specifications. In this paper, a self-organizing map (SOM)-based local linear modeling scheme of an unmanned aerial vehicle (UAV) is developed to design a set of inverse controllers. The SOM selects the operating regime depending only on the embedded output space information and avoids normalization of the input data. Each local linear model is associated with a linear controller, which is easy to design. Switching of the controllers is done synchronously with the active local linear model that tracks the different operating conditions. The proposed multiple modeling and control strategy has been successfully tested in a simulator that models the LoFLYTE UAV. Jeongho Cho, José C. Príncipe, Deniz Erdogmus, Mark A. Motter |
IEEE Trans. Neural Networks | 2 |
| 2005 | Supervised training of adaptive systems with partially labeled dataabstractSupervised adaptive system training is traditionally performed with available pairs of input-output data and the system weights are fixed following this training procedure. Recently, in the context of machine learning, where the desired outputs are discrete-valued, the idea of exploiting unlabeled samples for improving classification performance has been proposed. We introduce an information theoretic framework based on density divergence minimization to obtain extended training algorithms. Our goal is to provide a theoretical framework upon which we can build efficient algorithms to this end. Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe |
ICASSP (5) | 3 |
| 2005 | The Laplacian spectral classifierabstractWe develop a novel classifier in a kernel feature space defined by the eigenspectrum of the Laplacian data matrix. The classification cost function is derived from a distance measure between probability densities. The Laplacian data matrix is obtained based on a training set, while test data is mapped to the kernel space using the Nystrom routine. In that space, the test data is classified based on the angle between the test point and the training data class means. We illustrate the performance of the new classifier on synthetic and real data. Robert Jenssen, Deniz Erdogmus, José C. Príncipe, Torbjørn Eltoft |
ICASSP (5) | 3 |
| 2005 | Learning mappings in brain machine interfaces with echo state networksabstractBrain machine interfaces (BMI) utilize linear or non-linear models to map the neural activity to the associated behavior which is typically the 2D or 3D hand position of a primate. Linear models are plagued by the massive disparity of the input and output dimensions thereby leading to poor generalization. A solution would be to use non-linear models like the recurrent multi-layer perceptron (RMLP) that provide parsimonious mapping functions with better generalization. However, this results in a drastic increase in the training complexity, which can be critical for practical use of a BMI. This paper bridges the gap between superior performance per trained weight and model learning complexity. Towards this end, we propose to use echo state networks (ESN) to transform the neuronal firing activity into a higher dimensional space and then derive an optimal sparse linear mapping in the transformed space to match the hand position. The sparse mapping is obtained using a weight constrained cost function whose optimal solution is determined using a stochastic gradient algorithm. Yadunandana N. Rao, Sung-Phil Kim, Justin C. Sanchez, Deniz Erdogmus, José C. Príncipe, Jose M. Carmena, Mikhail Lebedev 0001, Miguel A. L. Nicolelis |
ICASSP (5) | 5 |
| 2005 | An information-theoretic perspective to kernel independent components analysisabstractIn this paper, we investigate the intriguing relationship between information-theoretic learning (ITL), based on weighted Parzen window density estimator, and kernel-based learning algorithms. We prove the equivalence between kernel independent component analysis (kernel ICA) and the Cauchy-Schwartz (C-S) independence measure. This link gives a theoretical motivation for the selection of the Mercer kernel, based on density estimation. Demonstrating this equivalence requires introducing a weighted kernel density estimator, a modification of Parzen windowing. We also discuss the role of the weights in the weighted Parzen windowing and kernel ICA. Jianwu Xu, Deniz Erdogmus, Robert Jenssen, José C. Príncipe |
ICASSP (5) | 4 |
| 2005 | Separating Spatial and Temporal Activation Patterns in fMRI Using Competitive Subspace ProjectionabstractThe challenge for functional magnetic resonance imaging (fMRI) is to determine when and where the response due to an external stimulus occurs in the images. The temporal clustering analysis (TCA) method has been used to study brain activity after eating and drinking, in both the time and spatial domains. We propose a new method, competitive subspace projection (CSP), to represent data optimally compared to other subspace projection methods. This method is used to detect spatial and temporal activation patterns in fMRI associated with such behavior. The CSP and TCA methods are compared using both synthetic and real fMRI data. The results on both data sets show consistent conclusions can be drawn from these two methods while CSP is observed to have a better noise rejection capability than TCA. Guojun He, Deniz Erdogmus, Sung-Phil Kim, José C. Príncipe |
ICASSP (2) | 5 |
| 2005 | Towards the modeling of dissociated cortical tissue in the liquid state machine frameworkabstractUnderstanding biological information processing systems is key in overcoming the limitations of traditional engineered systems. The advent of the liquid state machine (LSM) as a computational model for cortical processing, as well as the ability to experiment with dissociated cortical tissue (DCT) cultures in a controlled environment provide new opportunities in advancing our understanding of such systems. In this paper we examine the possibilities for modeling the behavior of the DCT cultures in the LSM framework. We show that the LSM framework has the capability to model the spontaneous activity and burstiness of the DCT cultures. Finally, multifractal measures are used to characterize the long range dependencies of the recorded and simulated data. Though the detrended fluctuation analysis (DFA) used to estimate the long range dependencies in the data does not show wholly consistent similarities, the endogenously active neurons that drive the biological and model networks are found to have similar fractal structure. Dilip Goswami, Klaus Schuch, Thomas B. DeMarse, José C. Príncipe |
IJCNN | 5 |
| 2005 | Sparse channel estimation with regularization method using convolution inequality for entropyabstractIn this paper, we show that the sparse channel estimation problem can be formulated as a regularization problem between mean squared error (MSE) and the L/sub 1/-norm constraint of the channel impulse response. A simple adaptive method to solve regularization problem using the convolution inequality for entropy is proposed. Performance of this proposed regularization method is compared to the Wiener filter, the matching pursuit (IMP) algorithm and the information criterion based method. The results show that the estimate of the sparse channel using the MSE criterion with the L/sub 1/-norm constraint outperforms the Wiener filter and the conventional sparse solution methods in terms of MSE of the estimates and the generalization performance. Dongho Han, Sung-Phil Kim, José C. Príncipe |
IJCNN | 3 |
| 2005 | An information theoretic approach to adaptive system training using unlabeled dataabstractTraditionally, supervised learning is performed with pairwise input-output labelled data. After the training procedure, the adaptive system weights are fixed and the system is tested with unlabelled data. Recently, exploiting the unlabeled data to improve classification performance has been proposed in the machine learning community. In this paper, we present an information theoretic approach based on density divergence minimization to obtain an extended training algorithm using unlabeled data during testing. The simulations for classification problems suggest that our method can improve the performance of adaptive system in the application phase. Kyu-Hwa Jeong, Jianwu Xu, José C. Príncipe |
IJCNN | 3 |
| 2005 | Computing with transiently stable statesabstractStability is an essential constraint in the design of linear dynamical systems. Similar stability restrictions on nonlinear dynamical systems, such as echo state network, have been enforced in order to use them for reliable computation. In this paper we introduce a novel computational mode for nonlinear systems with sigmoidal nonlinearity, which does not require global stability. In this mode, although the autonomous system is unstable, the input signal forces the system dynamics to become "transiently stable". We demonstrate with a function approximation experiment that the transiently stable system can still do useful computation. We explain the principles of computation with the stability of local dynamics obtained from linearization of the system at the operating point. Mustafa C. Ozturk, José C. Príncipe |
IJCNN | 2 |
| 2005 | Comparison of TDNN training algorithms in brain machine interfacesabstractLinear or non-linear models are used in brain machine interfaces (BIMIs) to map the neural activity to the associated behavior, typically the primate's hand position. Linear models assume a linear relationship between neural activity and hand position that may not be the case. A solution would be time-delay neural network (TDNN) that provides effectively a nonlinear combination of linear models. However, this model results in a drastic increase of free parameters and slow convergence when trained by an error backpropagation learning rule. We propose to train the TDNN by scaled conjugate gradient, which avoids time-consuming linear search, coupled with weight decay to reduce the free parameters number and produce generally faster convergence. Yiwen Wang 0002, Sung-Phil Kim, José C. Príncipe |
IJCNN | 3 |
| 2005 | Direct adaptive control: an echo state network and genetic algorithm approachabstractThis paper presents a direct adaptive approach to design controllers for nonlinear dynamical systems, where system identification of the unknown dynamical system is not required. The solution is powered by both echo state network (ESN) and genetic algorithm (GA). ESN enables a simple modeling of the controller, with which only a linear readout needs to be trained. GA is used to optimize ESN's linear readout directly so that system identification is not required. Simulation results reveal that the algorithm is capable of achieving very good control performance with computational efficiency. Jing Lan, José C. Príncipe |
IJCNN | 3 |
| 2005 | Fast error whitening algorithms for system identification and control with noisy data
Yadunandana N. Rao, Deniz Erdogmus, Geetha Y. Rao, José C. Príncipe |
Neurocomputing | 4 |
| 2005 | Vector quantization using information theoretic concepts
Tue Lehn-Schiøler, Anant Hegde, Deniz Erdogmus, José C. Príncipe |
Nat. Comput. | 4 |
| 2005 | A new classifier based on information theoretic learning with unlabeled data
Kyu-Hwa Jeong, Jianwu Xu, Deniz Erdogmus, José C. Príncipe |
Neural Networks | 4 |
| 2005 | A mutual information extension to the matched filter
Deniz Erdogmus, Rati Agrawal, José C. Príncipe |
Signal Process. | 3 |
| 2005 | Information Theoretic Signal Processing
Deniz Erdogmus, José C. Príncipe |
Signal Process. | 2 |
| 2005 | Quantifying spatio-temporal dependencies in epileptic ECOG
Anant Hegde, Deniz Erdogmus, Deng-Shan Shiau, José C. Príncipe, J. Chris Sackellares |
Signal Process. | 4 |
| 2005 | Linear-least-squares initialization of multilayer perceptrons through backpropagation of the desired responseabstractTraining multilayer neural networks is typically carried out using descent techniques such as the gradient-based backpropagation (BP) of error or the quasi-Newton approaches including the Levenberg-Marquardt algorithm. This is basically due to the fact that there are no analytical methods to find the optimal weights, so iterative local or global optimization techniques are necessary. The success of iterative optimization procedures is strictly dependent on the initial conditions, therefore, in this paper, we devise a principled novel method of backpropagating the desired response through the layers of a multilayer perceptron (MLP), which enables us to accurately initialize these neural networks in the minimum mean-square-error sense, using the analytic linear least squares solution. The generated solution can be used as an initial condition to standard iterative optimization algorithms. However, simulations demonstrate that in most cases, the performance achieved through the proposed initialization scheme leaves little room for further improvement in the mean-square-error (MSE) over the training set. In addition, the performance of the network optimized with the proposed approach also generalizes well to testing data. A rigorous derivation of the initialization algorithm is presented and its high performance is verified with a number of benchmark training problems including chaotic time-series prediction, classification, and nonlinear system identification with MLPs. Deniz Erdogmus, Oscar Fontenla-Romero, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo |
IEEE Trans. Neural Networks | 3 |
| 2004 | Mixture of competitive linear models for phased-array magnetic resonance imagingabstractPhased-array magnetic resonance imaging is an important contemporary research field in terms of the expected clinical gains in medical imaging technology. Recent research focused on heuristic coil image recombination methods as well as statistical signal processing approaches. In this paper, we investigate the performance of an adaptive signal processing approach, namely mixture of competitively trained models. The proposed method has the ability to train on a set of images and generalize its performance to previously unseen images. Performance evaluations on real data validate the effectiveness of this method. Deniz Erdogmus, Erik G. Larsson, José C. Príncipe, Jeffrey R. Fitzsimmons |
ICASSP (5) | 4 |
| 2004 | Parzen particle filtersabstractUsing a Parzen density estimator, any distribution can be approximated arbitrarily close by a sum of kernels. In particle filtering, this fact is utilized to estimate a probability density function with Dirac delta kernels; when the distribution is discretized it becomes possible to solve an otherwise intractable integral. In this work, we propose to extend the idea and use any kernel to approximate the distribution. The extra work involved in propagating small kernels through the nonlinear function can be made up for by decreasing the number of kernels needed, especially for high dimensional problems. A further advantage of using kernels with nonzero width is that the density estimate becomes continuous. Tue Lehn-Schiøler, Deniz Erdogmus, José C. Príncipe |
ICASSP (5) | 3 |
| 2004 | Accurate linear parameter estimation in colored noiseabstractEstimation of the parameters of an unknown system is an important problem in signal processing. The classical mean squared error (MSE) criterion and its variants have been widely used to solve this problem. However, it is well known that the MSE criterion produces biased parameter estimates when the signals of interest (especially the input) are corrupted with additive noise having arbitrary or no coloring (white). Alternative approaches require additional system constraints and explicit estimation of the noise covariances. Recently, we proposed a new criterion called the error whitening criterion (EWC) along with associated algorithms that solved the problem when the additive disturbances are white. However, the performance of EWC is not satisfactory when the disturbances are correlated (colored). In this paper, we propose a method based on the principles of the EWC that can consistently estimate the parameters of an unknown arbitrary linear system in colored input noise without estimating the noise covariances. We then present a novel stochastic gradient algorithm that estimates the optimal parameters in an on-line fashion. We briefly discuss the convergence of this algorithm and present extensive simulation results to show the superiority of this criterion over MSE. Yadunandana N. Rao, Deniz Erdogmus, José C. Príncipe |
ICASSP (2) | 3 |
| 2004 | Minimizing Fisher information of the error in supervised adaptive filter trainingabstractIn this paper, we propose minimizing the Fisher information of the error in supervised training of linear and nonlinear adaptive filters. Fisher information considers the local structure of the error probability distribution and therefore it is a criterion that deserves to be investigated as an alternative to more common statistics such as minimum mean-square-error or minimum-error-entropy. A gradient-based training algorithm, based on a nonparametric estimator of Fisher information is presented and the performances of the three mentioned optimization criteria are compared using Monte Carlo simulations. Jianwu Xu, Deniz Erdogmus, José C. Príncipe |
ICASSP (5) | 3 |
| 2004 | Nonlinear independent component analysis by homomorphic transformation of the mixturesabstractIndependent component analysis is often approached from an information theoretic perspective employing specific sample estimates for the mutual information between the separated outputs. These approximations involve the nonparametric estimation of signal entropies. The common approach involves the estimation of these quantities and adaptation based on these criteria. In contrast, in this paper, we propose a Gaussianization-based approach, where the separation is performed in two stages: Gaussianization of the mixtures using a homomorphic nonlinearity and separation of the independent components using principal component analysis (both stages possibly adaptive). Due to the rotation uncertainty in nonlinear ICA, the original sources cannot be recovered solely by the independence assumption. The proposed ICA methodology is applicable to instantaneous linear and nonlinear mixtures. The idea also generalizes easily to complex-valued nonlinear ICA. Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe |
IJCNN | 3 |
| 2004 | Real-time PCA (principal component analysis) implementation on DSPabstractPCA (principal component analysis) is a wellknown statistical technique used in many signal processing applications. An on-line temporal PCA learning algorithm is implemented on a floating-point DSP for real-time applications. This algorithm is coded in assembly language to optimize. The experimental results showed that the implemented on-line temporal PCA algorithm not only can accurately estimate the principal components from the input but also can track the principal components from the time varying input. And this algorithm can be applied in space easily by using spacial signals as its inputs instead of using the past inputs as in temporal PCA. Dongho Han, Yadunandana N. Rao, José C. Príncipe, Karl S. Gugel |
IJCNN | 3 |
| 2004 | Vector-quantization by density matching in the minimum Kullback-Leibler divergence senseabstractRepresentation of a large set of high-dimensional data is a fundamental problem in many applications such as communications and biomedical systems. The problem has been tackled by encoding the data with a compact set of code-vectors called processing elements. In this study, we propose a vector quantization technique that encodes the information in the data using concepts derived from information theoretic learning. The algorithm minimizes a cost function based on the Kullback-Liebler divergence to match the distribution of the processing elements with the distribution of the data. The performance of this algorithm is demonstrated on synthetic data as well as on an edge-image of a face. Comparisons are provided with some of the existing algorithms such as LBG and SOM. Anant Hegde, Deniz Erdogmus, Tue Lehn-Schiøler, Yadunandana N. Rao, José C. Príncipe |
IJCNN | 5 |
| 2004 | Information theoretic spectral clusteringabstractWe discuss a new information-theoretic framework for spectral clustering that is founded on the recently introduced information cut. A novel spectral clustering algorithm is proposed, where the clustering solution is given as a linearly weighted combination of certain top eigenvectors of the data affinity matrix. The information cut provides us with a theoretically well-defined graph-spectral cost function, and also establishes a close link between spectral clustering, and non-parametric density estimation. As a result, a natural criterion for creating the data affinity matrix is provided. We present preliminary clustering results to illustrate some of the properties of our algorithm, and we also make comparative remarks. Robert Jenssen, Torbjørn Eltoft, José C. Príncipe |
IJCNN | 3 |
| 2004 | Modified Freeman model: a stability analysis and application to pattern recognitionabstractThe biologically realistic Freeman model of the olfactory cortex has been used to solve some engineering problems. However, due to the nature of the nonlinear function in the model, only numerical computer simulations can help explore the behavior of the system for different sets of control parameters. We modify the nonlinear function with a piecewise linear model and show that this simplified model exhibits the same qualitative behavior as the original one. Moreover, for this modified model, we employ the analytical tools of nonlinear dynamics to understand the system response for different parameter values. Finally, similar to the original system, we show that the modified system can be used as an auto-associative memory. Mustafa C. Ozturk, José C. Príncipe |
IJCNN | 3 |
| 2004 | WelcomeabstractPresents the welcome message from the conference proceedings. José C. Príncipe |
IJCNN | 1 |
| 2004 | The Laplacian PDF Distance: A Cost Function for Clustering in a Kernel Feature SpaceabstractA new distance measure between probability density functions (pdfs) is introduced, which we refer to as the Laplacian pdf dis- tance. The Laplacian pdf distance exhibits a remarkable connec- tion to Mercer kernel based learning theory via the Parzen window technique for density estimation. In a kernel feature space defined by the eigenspectrum of the Laplacian data matrix, this pdf dis- tance is shown to measure the cosine of the angle between cluster mean vectors. The Laplacian data matrix, and hence its eigenspec- trum, can be obtained automatically based on the data at hand, by optimal Parzen window selection. We show that the Laplacian pdf distance has an interesting interpretation as a risk function connected to the probability of error. 1 Introduction In recent years, spectral clustering methods, i.e. data partitioning based on the eigenspectrum of kernel matrices, have received a lot of attention [1, 2]. Some unresolved questions associated with these methods are for example that it is not always clear which cost function that is being optimized and that is not clear how to construct a proper kernel matrix. In this paper, we introduce a well-defined cost function for spectral clustering. This cost function is derived from a new information theoretic distance measure between cluster pdfs, named the Laplacian pdf distance. The information theoretic/spectral duality is established via the Parzen window methodology for density estimation. The resulting spectral clustering cost function measures the cosine of the angle between cluster mean vectors in a Mercer kernel feature space, where the feature space is determined by the eigenspectrum of the Laplacian matrix. A principled approach to spectral clustering would be to optimize this cost function in the feature space by assigning cluster memberships. Because of space limitations, we leave it to a future paper to present an actual clustering algorithm optimizing this cost function, and focus in this paper on the theoretical properties of the new measure. Corresponding author. Phone: (+47) 776 46493. Email: [email protected] An important by-product of the theory presented is that a method for learning the Mercer kernel matrix via optimal Parzen windowing is provided. This means that the Laplacian matrix, its eigenspectrum and hence the feature space mapping can be determined automatically. We illustrate this property by an example. We also show that the Laplacian pdf distance has an interesting relationship to the probability of error. In section 2, we briefly review kernel feature space theory. In section 3, we utilize the Parzen window technique for function approximation, in order to introduce the new Laplacian pdf distance and discuss some properties in sections 4 and 5. Section 6 concludes the paper. 2 Kernel Feature Spaces Mercer kernel-based learning algorithms [3] make use of the following idea: via a nonlinear mapping : Rd F, x (x) (1) the data x1, . . . , xN Rd is mapped into a potentially much higher dimensional feature space F. For a given learning problem one now considers the same algorithm in F instead of in Rd, that is, one works with (x1),...,(xN) F. Consider a symmetric kernel function k(x, y). If k : C C R is a continuous kernel of a positive integral operator in a Hilbert space L2(C) on a compact set C Rd, i.e. L2(C) : k(x,y)(x)(y)dxdy 0, (2) C then there exists a space F and a mapping : Rd F, such that by Mercer's theorem [4] NF k(x, y) = (x), (y) = ii(x)i(y), (3) i=1 where , denotes an inner product, the i's are the orthonormal eigenfunctions of the kernel and NF [3]. In this case (x) = [ 11(x), 22(x), . . . ]T , (4) can potentially be realized. In some cases, it may be desirable to realize this mapping. This issue has been addressed in [5]. Define the (N N) Gram matrix, K, also called the affinity, or kernel matrix, with elements Kij = k(xi, xj), i, j = 1, . . . , N . This matrix can be diagonalized as ET KE = , where the columns of E contains the eigenvectors of K and is a diagonal matrix containing the non-negative eigenvalues ~ 1, . . . , ~ N , ~ 1 ~N. In [5], it was shown that the eigenfunctions and eigenvalues of (4) can ~ be approximated as j j (xi) Neji, j , where e N ji denotes the ith element of the jth eigenvector. Hence, the mapping (4), can be approximated as (xi) [ ~1e1i,..., ~NeNi]T. (5) Thus, the mapping is based on the eigenspectrum of K. The feature space data set may be represented in matrix form as NN = [(x1), . . . , (xN )]. Hence, = 1 2 ET . It may be desirable to truncate the mapping (5) to C-dimensions. Thus, T only the C first rows of are kept, yielding ^ . It is well-known that ^ K = ^ ^ is the best rank-C approximation to K wrt. the Frobenius norm [6]. The most widely used Mercer kernel is the radial-basis-function (RBF) k(x, y) = exp -||x - y||2 . (6) 22 3 Function Approximation using Parzen Windowing Parzen windowing is a kernel-based density estimation method, where the resulting density estimate is continuous and differentiable provided that the selected kernel is continuous and differentiable [7]. Given a set of iid samples {x1,...,xN} drawn from the true density f (x), the Parzen window estimate for this distribution is [7] N ^ 1 f (x) = W N 2 (x, xi), (7) i=1 where W2 is the Parzen window, or kernel, and 2 controls the width of the kernel. The Parzen window must integrate to one, and is typically chosen to be a pdf itself with mean xi, such as the Gaussian kernel 1 W2 (x, xi) = exp , (8) d -||x - xi||2 (22) 2 22 which we will assume in the rest of this paper. In the conclusion, we briefly discuss the use of other kernels. Consider a function h(x) = v(x)f (x), for some function v(x). We propose to estimate h(x) by the following generalized Parzen estimator N ^ 1 h(x) = v(xi)W N 2 (x, xi). (9) i=1 This estimator is asymptotically unbiased, which can be shown as follows 1 N Ef v(xi)W N 2 (x, xi) = v(z)f (z)W2 (x, z)dz = [v(x)f (x)] W2(x), i=1 (10) where Ef () denotes expectation with respect to the density f(x). In the limit as N and (N) 0, we have lim [v(x)f (x)] W2(x) = v(x)f(x). (11) N (N )0 Of course, if v(x) = 1 x, then (9) is nothing but the traditional Parzen estimator of h(x) = f (x). The estimator (9) is also asymptotically consistent provided that the kernel width (N ) is annealed at a sufficiently slow rate. The proof will be presented in another paper. Many approaches have been proposed in order to optimally determine the size of the Parzen window, given a finite sample data set. A simple selection rule was proposed by Silverman [8], using the mean integrated square error (MISE) between the estimated and the actual pdf as the optimality metric: 1 d+4 opt = X 4N -1(2d + 1)-1 , (12) where d is the dimensionality of the data and 2 = d-1 , where are the X i Xii Xii diagonal elements of the sample covariance matrix. More advanced approximations to the MISE solution also exist. 4 The Laplacian PDF Distance Cost functions for clustering are often based on distance measures between pdfs. The goal is to assign memberships to the data patterns with respect to a set of clusters, such that the cost function is optimized. Assume that a data set consists of two clusters. Associate the probability density function p(x) with one of the clusters, and the density q(x) with the other cluster. Let f (x) be the overall probability density function of the data set. Now define the f -1 weighted inner product between p(x) and q(x) as p, q f p(x)q(x)f-1(x)dx. In such an inner product space, the Cauchy-Schwarz inequality holds, that is, p, q 2 q, q . Based on this discussion, an information theoretic distance f p, p f f measure between the two pdfs can be expressed as p, q D f L = - log 0. (13) p, p q, q f f We refer to this measure as the Laplacian pdf distance, for reasons that we discuss next. It can be seen that the distance DL is zero if and only if the two densities are equal. It is non-negative, and increases as the overlap between the two pdfs decreases. However, it does not obey the triangle inequality, and is thus not a distance measure in the strict mathematical sense. We will now show that the Laplacian pdf distance is also a cost function for clus- tering in a kernel feature space, using the generalized Parzen estimators discussed in the previous section. Since the logarithm is a monotonic function, we will derive the expression for the argument of the log in (13). This quantity will for simplicity be denoted by the letter "L" in equations. Assume that we have available the iid data points {xi}, i = 1,...,N1, drawn from p(x), which is the density of cluster C1, and the iid {xj}, j = 1, . . ., N2, drawn from q(x), the density of C2. Let h(x) = f - 12 (x)p(x) and g(x) = f - 12 (x)q(x). Hence, we may write h(x)g(x)dx L = . (14) h2(x)dx g2(x)dx We estimate h(x) and g(x) by the generalized Parzen kernel estimators, as follows N1 N2 ^ 1 1 h(x) = f - 12 (xi)W f - 12 (xj )W N 2 (x, xi ), ^ g(x) = 2 (x, xj ). (15) 1 N2 i=1 j=1 The approach taken, is to substitute these estimators into (14), to obtain N N 1 1 1 2 h(x)g(x)dx f - 12 (xi)W f - 12 (xj )W N 2 (x, xi ) 2 (x, xj ) 1 N2 i=1 j=1 N 1 1 ,N2 = f - 12 (xi)f - 12 (xj ) W N 2 (x, xi )W2 (x, xj )dx 1N2 i,j=1 N 1 1 ,N2 = f - 12 (xi)f - 12 (xj )W N 22 (xi, xj ), (16) 1N2 i,j=1 where in the last step, the convolution theorem for Gaussians has been employed. Similarly, we have N 1 1 ,N1 h2(x)dx f - 12 (xi)f - 12 (xi )W N 2 22 (xi, xi ), (17) 1 i,i =1 N 1 2 ,N2 g2(x)dx f - 12 (xj)f - 12 (xj )W N 2 22 (xj , xj ). (18) 2 j,j =1 Now we define the matrix Kf , such that Kf = K ij f (xi, xj ) = f - 1 2 (xi)f - 12 (xj )K(xi, xj ), (19) where K(xi, xj ) = W22 (xi, xj) for i, j = 1, . . . , N and N = N1 + N2. As a consequence, (14) can be re-written as follows N1,N2 Kf (xi, xj) L = i,j=1 (20) N1,N1 K K i,i =1 f (xi, xi ) N2,N2 j,j =1 f (xj , xj ) The key point of this paper, is to note that the matrix K = Kij = K(xi, xj), i, j = 1, . . . , N , is the data affinity matrix, and that K(xi, xj) is a Gaussian RBF kernel function. Hence, it is also a kernel function that satisfies Mercer's theorem. Since K(xi, xj) satisfies Mercer's theorem, the following by definition holds [4]. For any set of examples {x1,...,xN} and any set of real numbers 1,...,N N N ijK(xi, xj) 0, (21) i=1 j=1 in analogy to (3). Moreover, this means that N N N N ijf - 12 (xi)f - 12 (xj )K(xi, xj) = ijKf (xi, xj) 0, (22) i=1 j=1 i=1 j=1 hence Kf (xi, xj ) is also a Mercer kernel. Now, it is readily observed that the Laplacian pdf distance can be analyzed in terms of inner products in a Mercer kernel-based Hilbert feature space, since Kf (xi, xj) = f (xi), f (xj) . Consequently, (20) can be written as follows N1,N2 f (xi), f (xj) L = i,j=1 N1,N1 i,i =1 f (xi), f (xi ) N2,N2 j,j =1 f (xj ), f (xj ) 1 N1 N2 N i=1 f (xi ), 1 N j=1 f (xj ) = 1 2 1 N1 N1 N2 N2 N f (xi), 1 f (xi ) 1 f (xj ), 1 f (xj ) 1 i=1 N1 i =1 N2 j=1 N2 j =1 m1 , m2 = f f = cos (m , m ), (23) ||m 1f 2f 1f ||||m2f || where m Ni i = 1 f N f (xl), i = 1, 2, that is, the sample mean of the ith cluster i l=1 in feature space. This is a very interesting result. We started out with a distance measure between densities in the input space. By utilizing the Parzen window method, this distance measure turned out to have an equivalent expression as a measure of the distance between two clusters of data points in a Mercer kernel feature space. In the feature space, the distance that is measured is the cosine of the angle between the cluster mean vectors. The actual mapping of a data point to the kernel feature space is given by the eigendecomposition of Kf , via (5). Let us examine this mapping in more detail. 1 Note that f 2 (xi) can be estimated from the data by the traditional Parzen pdf estimator as follows N 1 1 f 2 (xi) = W (xi, xl) = di. (24) N 2 f l=1 Define the matrix D = diag(d1, . . . , dN ). Then Kf can be expressed as Kf = D- 12 KD- 12 . (25) Quite interestingly, for 2 = 22, this is in fact the Laplacian data matrix. 1 f The above discussion explicitly connects the Parzen kernel and the Mercer kernel. Moreover, automatic procedures exist in the density estimation literature to opti- mally determine the Parzen kernel given a data set. Thus, the Mercer kernel is also determined by the same procedure. Therefore, the mapping by the Laplacian matrix to the kernel feature space can also be determined automatically. We regard this as a significant result in the kernel based learning theory. As an example, consider Fig. 1 (a) which shows a data set consisting of a ring with a dense cluster in the middle. The MISE kernel size is opt = 0.16, and the Parzen pdf estimate is shown in Fig. 1 (b). The data mapping given by the corresponding Laplacian matrix is shown in Fig. 1 (c) (truncated to two dimensions for visualization purposes). It can be seen that the data is distributed along two lines radially from the origin, indicating that clustering based on the angular measure we have derived makes sense. The above analysis can easily be extended to any number of pdfs/clusters. In the C-cluster case, we define the Laplacian pdf distance as C-1 pi, pj L = f . (26) i=1 j=i C pi, pi p f j , pj f In the kernel feature space, (26), corresponds to all cluster mean vectors being pairwise as orthogonal to each other as possible, for all possible unique pairs. 4.1 Connection to the Ng et al. [2] algorithm Recently, Ng et al. [2] proposed to map the input data to a feature space determined by the eigenvectors corresponding to the C largest eigenvalues of the Laplacian ma- trix. In that space, the data was normalized to unit norm and clustered by the C-means algorithm. We have shown that the Laplacian pdf distance provides a 1It is a bit imprecise to refer to Kf as the Laplacian matrix, as readers familiar with spectral graph theory may recognize, since the definition of the Laplacian matrix is L = I - Kf . However, replacing Kf by L does not change the eigenvectors, it only changes the eigenvalues from i to 1 - i. 0 0 (a) Data set (b) Parzen pdf estimate (c) Feature space data Figure 1: The kernel size is automatically determined (MISE), yielding the Parzen estimate (b) with the corresponding feature space mapping (c). clustering cost function, measuring the cosine of the angle between cluster means, in a related kernel feature space, which in our case can be determined automati- cally. A more principled approach to clustering than that taken by Ng et al. is to optimize (23) in the feature space, instead of using C-means. However, because of the normalization of the data in the feature space, C-means can be interpreted as clustering the data based on an angular measure. This may explain some of the success of the Ng et al. algorithm; it achieves more or less the same goal as cluster- ing based on the Laplacian distance would be expected to do. We will investigate this claim in our future work. Note that we in our framework may choose to use only the C largest eigenvalues/eigenvectors in the mapping, as discussed in section 2. Since we incorporate the eigenvalues in the mapping, in contrast to Ng et al., the actual mapping will in general be different in the two cases. 5 The Laplacian PDF distance as a risk function We now give an analysis of the Laplacian pdf distance that may further motivate its use as a clustering cost function. Consider again the two cluster case. The overall data distribution can be expressed as f (x) = P1p(x) + P2q(x), were Pi, i = 1, 2, are the priors. Assume that the two clusters are well separated, such that for xi C1, f (xi) P1p(xi), while for xi C2, f(xi) P2q(xi). Let us examine the numerator of (14) in this case. It can be approximated as p(x)q(x) dx f (x) p(x)q(x) p(x)q(x) 1 1 dx + dx q(x)dx + p(x)dx. (27) C f (x) f (x) P1 P2 1 C2 C1 C2 By performing a similar calculation for the denominator of (14), it can be shown to be approximately equal to 1 . Hence, the Laplacian pdf distance can be written P1P1 as a risk function, given by 1 1 L P1P2 q(x)dx + p(x)dx . (28) P1 C P2 1 C2 Note that if P1 = P2 = 1 , then L = 2P 2 e, where Pe is the probability of error when assigning data points to the two clusters, that is Pe = P1 q(x)dx + P2 p(x)dx. (29) C1 C2 Thus, in this case, minimizing L is equivalent to minimizing Pe. However, in the case that P1 = P2, (28) has an even more interesting interpretation. In that situation, it can be seen that the two integrals in the expressions (28) and (29) are weighted exactly oppositely. For example, if P1 is close to one, L p(x)dx, while P C e 2 q(x)dx. Thus, the Laplacian pdf distance emphasizes to cluster the most un- C1 likely data points correctly. In many real world applications, this property may be crucial. For example, in medical applications, the most important points to classify correctly are often the least probable, such as detecting some rare disease in a group of patients. 6 Conclusions We have introduced a new pdf distance measure that we refer to as the Laplacian pdf distance, and we have shown that it is in fact a clustering cost function in a kernel feature space determined by the eigenspectrum of the Laplacian data matrix. In our exposition, the Mercer kernel and the Parzen kernel is equivalent, making it possible to determine the Mercer kernel based on automatic selection procedures for the Parzen kernel. Hence, the Laplacian data matrix and its eigenspectrum can be determined automatically too. We have shown that the new pdf distance has an interesting property as a risk function. The results we have derived can only be obtained analytically using Gaussian ker- nels. The same results may be obtained using other Mercer kernels, but it requires an additional approximation wrt. the expectation operator. This discussion is left for future work. Acknowledgments. This work was partially supported by NSF grant ECS- 0300340. Robert Jenssen, Deniz Erdogmus, José C. Príncipe, Torbjørn Eltoft |
NIPS | 3 |
| 2004 | Minimax Mutual Information Approach for Independent Component AnalysisabstractMinimum output mutual information is regarded as a natural criterion for independent component analysis (ICA) and is used as the performance measure in many ICA algorithms. Two common approaches in information-theoretic ICA algorithms are minimum mutual information and maximum output entropy approaches. In the former approach, we substitute some form of probability density function (pdf) estimate into the mutual information expression, and in the latter we incorporate the source pdf assumption in the algorithm through the use of nonlinearities matched to the corresponding cumulative density functions (cdf). Alternative solutions to ICA use higher-order cumulant-based optimization criteria, which are related to either one of these approaches through truncated series approximations for densities. In this article, we propose a new ICA algorithm motivated by the maximum entropy principle (for estimating signal distributions). The optimality criterion is the minimum output mutual information, where the estimated pdfs are from the exponential family and are approximate solutions to a constrained entropy maximization problem. This approach yields an upper bound for the actual mutual information of the output signals - hence, the name minimax mutual information ICA algorithm. In addition, we demonstrate that for a specific selection of the constraint functions in the maximum entropy density estimation procedure, the algorithm relates strongly to ICA methods using higher-order cumulants. Deniz Erdogmus, Kenneth E. Hild II, Yadunandana N. Rao, José C. Príncipe |
Neural Comput. | 4 |
| 2004 | Asymptotic SNR-performance of some image combination techniques for phased-array MRI
Deniz Erdogmus, Erik G. Larsson, José C. Príncipe, Jeffrey R. Fitzsimmons |
Signal Process. | 4 |
| 2004 | Measuring the signal-to-noise ratio in magnetic resonance imaging: a caveat
Deniz Erdogmus, Erik G. Larsson, José C. Príncipe, Jeffrey R. Fitzsimmons |
Signal Process. | 4 |
| 2004 | Advanced search algorithms for information-theoretic learning with kernel-based estimatorsabstractRecent publications have proposed various information-theoretic learning (ITL) criteria based on Renyi's quadratic entropy with nonparametric kernel-based density estimation as alternative performance metrics for both supervised and unsupervised adaptive system training. These metrics, based on entropy and mutual information, take into account higher order statistics unlike the mean-square error (MSE) criterion. The drawback of these information-based metrics is the increased computational complexity, which underscores the importance of efficient training algorithms. In this paper, we examine familiar advanced-parameter search algorithms and propose modifications to allow training of systems with these ITL criteria. The well known algorithms tailored here for ITL include various improved gradient-descent methods, conjugate gradient approaches, and the Levenberg-Marquardt (LM) algorithm. Sample problems and metrics are presented to illustrate the computational efficiency attained by employing the proposed algorithms. Rodney A. Morejon, José C. Príncipe |
IEEE Trans. Neural Networks | 2 |
| 2004 | Guest Editorial Special Issue on Information Theoretic Learning
José C. Príncipe, Erkki Oja, Lei Xu 0001, Andrzej Cichocki, Deniz Erdogmus |
IEEE Trans. Neural Networks | 1 |
| 2004 | Feature selection in MLPs and SVMs based on maximum output informationabstractThis paper presents feature selection algorithms for multilayer perceptrons (MLPs) and multiclass support vector machines (SVMs), using mutual information between class labels and classifier outputs, as an objective function. This objective function involves inexpensive computation of information measures only on discrete variables; provides immunity to prior class probabilities; and brackets the probability of error of the classifier. The maximum output information (MOI) algorithms employ this function for feature subset selection by greedy elimination and directed search. The output of the MOI algorithms is a feature subset of user-defined size and an associated trained classifier (MLP/SVM). These algorithms compare favorably with a number of other methods in terms of performance on various artificial and real-world data sets. Vikas Sindhwani, Subrata Rakshit, Dipti Deodhare, Deniz Erdogmus, José C. Príncipe, Partha Niyogi |
IEEE Trans. Neural Networks | 5 |
| 2004 | Dynamical analysis of neural oscillators in an olfactory cortex modelabstractThis paper presents a theoretical approach to understand the basic dynamics of a hierarchical and realistic computational model of the olfactory system proposed by W. J. Freeman. While the system's parameter space could be scanned to obtain the desired dynamical behavior, our approach exploits the hierarchical organization and focuses on understanding the simplest building block of this highly connected network. Based on bifurcation analysis, we obtain analytical solutions of how to control the qualitative behavior of a reduced KII set taking into consideration both the internal coupling coefficients and the external stimulus. This also provides useful insights for investigating higher level structures that are composed of the same basic structure. Experimental results are presented to verify our theoretical analysis. José C. Príncipe |
IEEE Trans. Neural Networks | 2 |
| 2003 | Recursive Least Squares for an Entropy Regularized MSE Cost Function
Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
ESANN | 3 |
| 2003 | Accelerating the convergence speed of neural networks learning methods using least squares
Oscar Fontenla-Romero, Deniz Erdogmus, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo |
ESANN | 3 |
| 2003 | Linear Least-Squares Based Methods for Neural Networks Learning
Oscar Fontenla-Romero, Deniz Erdogmus, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo |
ICANN | 3 |
| 2003 | On the convergence of SIPEX: a simultaneous principal components extraction algorithmabstractWe have previously proposed SIPEX as a fast-converging and accurate principal components algorithm (Erdogmus, D. et al., Proc. ICASSP'02, vol.1, p.1069-72, 2002; Proc. EUSIPCO'02, vol.2, p.335-8, 2002). Its superiority in terms of data efficiency and solution accuracy was demonstrated through Monte Carlo simulations. We focus on the convergence properties of the original gradient-based algorithm as well as two modified versions of SIPEX based on approximations to the Hessian matrix of the cost function. We provide practical bounds on the step sizes of these algorithms and compare their convergence properties. Deniz Erdogmus, Yadunandana N. Rao, M. Can Ozturk, Luis Vielva, José C. Príncipe |
ICASSP (2) | 5 |
| 2003 | A hybrid subspace projection method for system identificationabstractPrincipal components analysis (PCA), being the most optimal linear mapper in a least-squares (LS) sense, has been predominantly used in subspace-based signal processing methods. In system identification problems, optimal subspace projections must span the joint space of the input and output of the unknown system. In this scenario, subspaces determined by the principal components of the input or the desired signal alone do not embed key information, which lies in the joint space. We first propose a hybrid subspace projection method that finds optimal projections in the joint space. The concepts behind this method are firmly rooted in statistical theory. We then derive adaptive learning algorithms to estimate the subspace projections. Finally, we show the superiority of the new framework in solving system identification problems in noisy environments. Sung-Phil Kim, Yadunandana N. Rao, Deniz Erdogmus, José C. Príncipe |
ICASSP (6) | 4 |
| 2003 | Echo cancellation by global optimization of Kautz filters using an information theoretic criterionabstractIn practical settings, the echo cancellation problem generally requires the adaptation of an IIR filter using some optimality criterion. This brings two problems: direct adaptation of numerator and denominator polynomial coefficients of IIR filters might result in unstable systems and/or the optimization might result in a suboptimal local minimum of the criterion. These two issues are addressed in this paper. To resolve the first problem, orthogonal Kautz filters are utilized for their stability is easily controlled through the pole locations. The second problem is addressed by employing an information theoretic optimality criterion, which has a parameter that is annealed to ensure global optimization. Ching-An Lai, Deniz Erdogmus, José C. Príncipe |
ICASSP (6) | 3 |
| 2003 | Matched pdf-based blind equalizationabstractIn this paper, a new blind equalization algorithm for multilevel modulations is proposed. It is based on maximizing the correlation between the probability density function (pdf) of the signal at the output of the equalizer and the desired pdf. The algorithm employs the Parzen window method to estimate the pdf of the squared modulus of the equalizer output. A stochastic gradient-based algorithm is used to maximize the correlation between this pdf and the pdf of the corresponding modulation. The proposed algorithm shows an excellent performance when compared with conventional adaptive blind algorithms, such as CMA, in quadrature amplitude modulation (QAM) schemes. Marcelino Lázaro, Ignacio Santamaría, Carlos Pantaleón, Deniz Erdogmus, José C. Príncipe |
ICASSP (4) | 5 |
| 2003 | Image combination for high-field phased-array MRIabstractWe consider signal processing methods for phased-array magnetic resonance imaging (MRI). A theoretical description of a phased-array MRI data model is presented and three image reconstruction algorithms are proposed to estimate the effective image pixels. An analysis is provided to show how the new algorithms compare to conventional reconstruction methods. Deniz Erdogmus, Erik G. Larsson, José C. Príncipe, Jeffrey R. Fitzsimmons |
ICASSP (5) | 4 |
| 2003 | A local linear modeling paradigm with a modified counterpropagation networkabstractThe counter-propagation neural network (CPN) was selected to investigate the modeling problems because it integrates both supervised and unsupervised learning. The network is taught to have clusters that are described by codebook vectors in the training phase. The basic CPN algorithm is modified incorporating the local linear models (LLMs) to provide functional mappings and identify potentially nonlinear plants. The objective is a reduction of the approximation error in a CPN. In this framework the quantization error in the output space serves as a basis for the LLMs in the output space. The performance of the proposed algorithms is tested on the nonlinear dynamic system and shows the influence of this modification on the system identification quality. Jeongho Cho, José C. Príncipe, Mark A. Motter |
IJCNN | 2 |
| 2003 | Supervised synaptic weight adaptation for a spiking neuronabstractA novel algorithm named Spike-LMS is described that adapts the synaptic weights of an artificial spiking neuron to produce a desired response. The derivation of Spike-LMS follows from the derivation of the least-mean squares (LMS) algorithm used in adaptive filter theory. Spike-LMS works directly in the domain of spike trains, and therefore makes no assumptions about any particular neural encoding method. This algorithm is able to identify the synaptic weights of a spiking neuron given the pre-synaptic and post-synaptic spike trains. Bryan A. Davis, Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe |
IJCNN | 4 |
| 2003 | Accurate initialization of neural network weights by backpropagation of the desired responseabstractProper initialization of neural networks is critical for a successful training of its weights. Many methods have been proposed to achieve this, including heuristic least squares approaches. In this paper, inspired by these previous attempts to train (or initialize) neural networks, we formulate a mathematically sound algorithm based on backpropagating the desired output through the layers of a multilayer perceptron. The approach is accurate up to local first order approximations of the nonlinearities. It is shown to provide successful weight initialization for many data sets by Monte Carlo experiments. Deniz Erdogmus, Oscar Fontenla-Romero, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo, Robert Jenssen |
IJCNN | 3 |
| 2003 | Clustering using Renyi's entropyabstractWe propose a new clustering algorithm using Renyi's entropy as our similarity metric. The main idea is to assign a data pattern to the cluster, which among all possible clusters, increases its within-cluster entropy the least, upon inclusion of the pattern. We refer to this procedure as differential entropy clustering. Not knowing the true number of clusters in advance, initially a number of small clusters are "seeded" randomly in the data set, labeling a small subset of the data. Thereafter all remaining patterns are labeled by differential entropy clustering. Subsequently, we identify the "worst cluster" by a quantity we name as the between-cluster entropy. Its members are re-clustered, again by differential entropy clustering, reducing the overall number of clusters by one. This procedure is repeated until only two clusters remain. At each step we store the current labels, thus producing a hierarchy of clusters. The between-cluster entropy also enables us to select our final set of clusters in other cluster hierarchy. We demonstrate the clustering algorithm when applied both to artificially created data sets and a real data set. Robert Jenssen, Kenneth E. Hild II, Deniz Erdogmus, José C. Príncipe, Torbjørn Eltoft |
IJCNN | 4 |
| 2003 | A two-processing element adaptable linear oscillating recurrent system with single-weight plasticityabstractA simple recursive linear two-processing element adaptable oscillator is developed and demonstrated. Simple harmonic motion and its mathematical description is the foundation of this study. State equations for an undamped spring-mass system are converted to a discrete time system, followed by eigenvalue analysis to convert the four-degree-of-freedom system to a recursive network with plasticity in one variable. The oscillator is initialized to a frequency in the neighborhood of the desired frequency, and tuned by using resilient backpropagation modified for backpropagation through time. It is shown that the network can track frequencies in a spectrum of operation as defined by the sampling frequency, given the network is initialized to frequencies in the neighborhood of those of interest. Michael R. Johnson, José C. Príncipe |
IJCNN | 2 |
| 2003 | Modeling the relation from motor cortical neuronal firing to hand movements using competitive linear filters and a MLPabstractRecent research has demonstrated that linear model are able to estimate hand positions using populations of action potentials collected in the pre-motor and motor cortical areas of a primate's brain. One of the applications of this result is to restore movement in patients suffering from paralysis. To implement this technology in real-time, reliable and accurate signal processing models that produce sufficient small error in the estimated hand positions are required. In this paper, we propose the hybrid model approach that combines competitive linear filters with a neural network. The mapping performance of our approach is compared with a single Wiener filter during reaching movements. Our approach demonstrates more accurate estimations. Sung-Phil Kim, Justin C. Sanchez, Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe, Miguel A. L. Nicolelis |
IJCNN | 5 |
| 2003 | Identification of dynamical systems using GMM with VQ initializationabstractWe are using Gaussian mixture models (GMM) as a tool to construct local mappings of nonlinear multi-input multi-output (MIMO) systems. In this work, we combine the advantages of GMM with the Kalman filter. To improve the accuracy of the local linear mappings in a potentially large dimensional state space, we propose to initialize the GMM parameters with vector quantization (VQ) or its more parsimonious counterpart growing self-organizing maps (G-SOM). The performance of the proposed modeling algorithm on simulated data obtained from a realistic aircraft model show improvements in both converge speed and accuracy. Jing Lan, José C. Príncipe, Mark A. Motter |
IJCNN | 2 |
| 2003 | Simulation of the Freeman model of the olfactory cortex: a quantitative performance analysis for the DSP approachabstractIt has been shown that a framework composed of digital signal processing (DSP) elements can be used to simulate and study Freeman's model of the biologically realistic olfactory cortex. In this paper, based on impulsive invariant transformation, a DSP environment has been developed corresponding to the original continuous-time dynamical system. The performance of the DSP system is quantitatively evaluated by comparing with the original system, which is implemented using the traditional Runge-Kutta integration technique. The discrete-time architecture is shown to have high performance in approximating the dynamical behavior of the original system. Mustafa C. Ozturk, José C. Príncipe, Bryan A. Davis, Deniz Erdogmus |
IJCNN | 2 |
| 2003 | Error whitening criterion for linear filter estimationabstractMean square error (MSE) has been the most widely used tool to solve the linear filter estimation or system identification problem. However, MSE gives biased results when the input signals are noisy. This paper presents a novel error whitening criterion (EWC) to tackle the problem of linear system identification in the presence of additive white disturbances. We would motivate the theory behind the new criterion and derive an online stochastic gradient algorithm based on EWC. Convergence proof of the stochastic gradient algorithm is derived making mild assumptions. Simulation results show the effectiveness of this criterion. We compare its performance with MSE as well as the powerful total least squares method. Yadunandana N. Rao, Deniz Erdogmus, José C. Príncipe |
IJCNN | 3 |
| 2003 | Bimodal brain-machine interface for motor control of robotic prostheticabstractWe are working on mapping multi-channel neural spike data, recorded from multiple cortical areas of an owl monkey, to corresponding 3D monkey arm positions. In earlier work on this mapping task, we observed that continuous function approximators (such as artificial neural networks) have difficulty in jointly estimating 3D arm positions for two distinct cases-namely, when the monkey's arm is stationary and when it is moving. Therefore, we propose a multiple-model approach that first classifies neural spike data into two classes, corresponding to two states of the monkey's arm: (1) stationary and (2) moving. Then, the output of this classifier is used as a gating mechanism for subsequent continuous models, with one model per class. In this paper, we first motivate and discuss our approach. Next, we present encouraging results for the classifier stage, based on hidden Markov models (HMMs), and also for the entire bimodal mapping system. Finally, we conclude with a discussion of the results and suggest future avenues of research. Shalom Darmanjian, Sung-Phil Kim, Michael C. Nechyba, Scott Morrison, José C. Príncipe, Johan Wessberg, Miguel A. L. Nicolelis |
IROS | 5 |
| 2003 | Divide-and-conquer approach for brain machine interfaces: nonlinear mixture of competitive linear models
Sung-Phil Kim, Justin C. Sanchez, Deniz Erdogmus, Yadunandana N. Rao, Johan Wessberg, José C. Príncipe, Miguel A. L. Nicolelis |
Neural Networks | 6 |
| 2003 | Stochastic error whitening algorithm for linear filter estimation with noisy data
Yadunandana N. Rao, Deniz Erdogmus, Geetha Y. Rao, José C. Príncipe |
Neural Networks | 4 |
| 2003 | Maximum margin equalizers trained with the Adatron algorithm
Ignacio Santamaría, Rafael González Ayestarán, Carlos Pantaleón, José C. Príncipe |
Signal Process. | 4 |
| 2003 | Online entropy manipulation: stochastic information gradientabstractEntropy has found significant applications in numerous signal processing problems including independent components analysis and blind deconvolution. In general, entropy estimators require O(N/sup 2/) operations, N being the number of samples. For practical online entropy manipulation, it is desirable to determine a stochastic gradient for entropy, which has O(N) complexity. In this paper, we propose a stochastic Shannon's entropy estimator. We determine the corresponding stochastic gradient and investigate its performance. The proposed stochastic gradient for Shannon's entropy can be used in online adaptation problems where the optimization of an entropy-based cost function is necessary. Deniz Erdogmus, Kenneth E. Hild II, José C. Príncipe |
IEEE Signal Process. Lett. | 3 |
| 2002 | Potential Energy and Particle Interaction Approach for Learning in Adaptive Systems
Deniz Erdogmus, José C. Príncipe, Luis Vielva, David Luengo |
ICANN | 2 |
| 2002 | Local Modeling Using Self-Organizing Maps and Single Layer Neural Networks
Oscar Fontenla-Romero, Amparo Alonso-Betanzos, Enrique F. Castillo, José C. Príncipe, Bertha Guijarro-Berdiñas |
ICANN | 4 |
| 2002 | Simultaneous extraction of Principal Components using givens rotations and output variancesabstractPrincipal Components Analysis (PCA) is an invaluable statistical tool in signal processing. In many cases, an on-line algorithm to adapt the PCA network to determine the principal projections in the input space is desired. Algorithms proposed until now use the traditional deflation or the inflation procedure to determine the intermediate components sequentially, after the convergence of the principal or minor component is achieved. In this paper, we propose a constrained linear network and a robust cost function to determine any number of principal components simultaneously. The topology exploits the fact that the eigenvector matrix sought is orthonormal. A gradient-based algorithm named SIPEX-G is also presented. Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe, Kenneth E. Hild II |
ICASSP | 3 |
| 2002 | Blind source separation of time-varying instantaneous mixtures using an on-line algorithmabstractA blind source separation algorithm is presented that performs online separation of an unknown, time-varying, instantaneous mixture of independent sources. This procedure utilizes the information-theoretic MeRMald-SIG criterion and an on-line PCA algorithm, referred to as SIPEX-G, that has been recently submitted for publication. Results indicate that the combination of the two on-line criteria is able to track a rapidly changing mixing matrix. A performance comparison of separation criteria is also made using the same, aforementioned on-line PCA algorithm. Results show the superior performance of the proposed method. Kenneth E. Hild II, Deniz Erdogmus, José C. Príncipe |
ICASSP | 3 |
| 2002 | Robust on-line Principal Component Analysis based on a fixed-point approachabstractPrincipal Component Analysis (PCA) is a widely used statistical tool in many signal-processing applications. In this paper we will present a new on-line algorithm for computing the principal components. The new algorithm belongs to a class of fixed-point methods. We mathematically investigate the convergence properties of the method and also verify the robustness of the algorithm with simulations. Yadunandana N. Rao, José C. Príncipe |
ICASSP | 2 |
| 2002 | Fast algorithm for adaptive blind equalization using order-α Renyi's entropyabstractIn this paper a novel blind equalization algorithm based on stochastic gradient descent minimization of order-α Renyi's entropy and designed for constant modulus signals is introduced. The algorithm applies a new nonparametric estimator for Renyi's entropy, which has been recently proposed and allows to compute any order of entropy. In comparison with conventional adaptive blind techniques, such us CMA, the proposed algorithm shows a remarkable increase in convergence speed with only a moderate increase in computational cost. Ignacio Santamaría, Carlos Pantaleón, Luis Vielva, José C. Príncipe |
ICASSP | 4 |
| 2002 | Underdetermined blind source separation in a time-varying environmentabstractThe problem of estimating n source signals from m measurements that are an unknown mixture of the sources is known as blind source separation. In the underdetermined —less measurements than sources— linear case, the solution process can be conveniently divided in three stages: represent the signals in a sparse domain, find the mixing matrix, and estimate the sources. In this paper we adhere to that approach and parametrize the performance of these stages as a function of the sparsity of the signals. To find the mixing matrix and track its variations in the dynamic case a nonparametric maximum-likelihood approach based on Parzen windowing is presented. To invert the underdetermined linear problem we present an estimator that chooses the “best” demixing matrix in a sample by sample basis by using some previous knowledge of the statistics of the sources. The results are validated by Montecarlo simulations. Luis Vielva, Deniz Erdogmus, Carlos Pantaleón, Ignacio Santamaría, J. A. Pereda, José C. Príncipe |
ICASSP | 6 |
| 2002 | Blind source separation using Renyi's -marginal entropies
Deniz Erdogmus, Kenneth E. Hild II, José C. Príncipe |
Neurocomputing | 3 |
| 2002 | Beyond second-order statistics for learning: A pairwise interaction model for entropy estimation
Deniz Erdogmus, José C. Príncipe, Kenneth E. Hild II |
Nat. Comput. | 2 |
| 2002 | Principles and networks for self-organization in space-time
José C. Príncipe, Neil R. Euliano, Shayan Garani |
Neural Networks | 1 |
| 2002 | Information Theoretic ClusteringabstractClustering is an important topic in pattern recognition. Since only the structure of the data dictates the grouping (unsupervised learning), information theory is an obvious criteria to establish the clustering rule. The paper describes a novel valley seeking clustering algorithm using an information theoretic measure to estimate the cost of partitioning the data set. The information theoretic criteria developed here evolved from a Renyi entropy estimator (A. Renyi, 1960) that was proposed recently and has been successfully applied to other machine learning applications (J.C. Principe et al., 2000). An improved version of the k-change algorithm is used in optimization because of the stepwise nature of the cost function and existence of local minima. Even when applied to nonlinearly separable data, the new algorithm performs well, and was able to find nonlinear boundaries between clusters. The algorithm is also applied to the segmentation of magnetic resonance imaging data (MRI) with very promising results. Erhan Gokcay, José C. Príncipe |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | Generalized information potential criterion for adaptive system trainingabstractWe have previously proposed the quadratic Renyi's error entropy as an alternative cost function for supervised adaptive system training. An entropy criterion instructs the minimization of the average information content of the error signal rather than merely trying to minimize its energy. In this paper, we propose a generalization of the error entropy criterion that enables the use of any order of Renyi's entropy and any suitable kernel function in density estimation. It is shown that the proposed entropy estimator preserves the global minimum of actual entropy. The equivalence between global optimization by convolution smoothing and the convolution by the kernel in Parzen windowing is also discussed. Simulation results are presented for time-series prediction and classification where experimental demonstration of all the theoretical concepts is presented. Deniz Erdogmus, José C. Príncipe |
IEEE Trans. Neural Networks | 2 |
| 2002 | Guest editorial special issue on intelligent multimedia processing
Ling Guan, Tülay Adali, Shigeru Katagiri, Jan Larsen, José C. Príncipe |
IEEE Trans. Neural Networks | 5 |
| 2001 | Optimization in companion search spaces: the case of cross-entropy and the Levenberg-Marquardt algorithmabstractWe present a new learning algorithm for the supervised training of multilayer perceptions for classification that is significantly faster than any previously known method. Like existing methods, the algorithm assumes a multilayer perceptron with a normalized exponential (softmax) output trained under a cross-entropy criterion. However, this output-criteria pairing turns out to have poor properties for existing optimization methods (backpropagation and its second order extensions) because second-order expansion of the network weights about the optimal solution is not a good approximation. The proposed algorithm overcomes this limitation by defining a new search space for which a second-order expansion is valid and such that the optimal solution in the new space coincides with the original criterion. This allows the application of the Levenberg-Marquardt search procedure to the cross-entropy criterion, which was previously thought applicable only to a mean square error criteria. Craig L. Fancourt, José C. Príncipe |
ICASSP | 2 |
| 2001 | Design and implementation of a biologically realistic olfactory cortex in analog VLSIabstractThis paper reviews the problem of translating signals into symbols preserving maximally the information contained in the signal time structure. In this context, we motivate the use of nonconvergent dynamics for the signal to symbol translator. We then describe a biologically realistic model of the olfactory system proposed by W. Freeman (1975) that has locally stable dynamics but is globally chaotic. We show how we can discretize Freeman's model using digital signal processing techniques, providing an alternative to the more conventional Runge-Kutta integration. This analysis leads to a direct mixed-signal (analog amplitude/discrete time) implementation of the dynamical building block that simplifies the implementation of the interconnect. We present results of simulations and measurements obtained from a fabricated analog VLSI chip. José C. Príncipe, Vítor Grade Tavares, John G. Harris, Walter J. Freeman |
Proc. IEEE | 1 |
| 2001 | Blind source separation using Renyi's mutual informationabstractA blind source separation algorithm is proposed that is based on minimizing Renyi's mutual information by means of nonparametric probability density function (PDF) estimation. The two-stage process consists of spatial whitening and a series of Givens rotations and produces a cost function consisting only of marginal entropies. This formulation avoids the problems of PDF inaccuracy due to truncation of series expansion and the estimation of joint PDFs in high-dimensional spaces given the typical paucity of data. Simulations illustrate the superior efficiency, in terms of data length, of the proposed method compared to fast independent component analysis (FastICA), Comon's (1994) minimum mutual information, and Bell and Sejnowski's (1995) Infomax. Kenneth E. Hild II, Deniz Erdogmus, José C. Príncipe |
IEEE Signal Process. Lett. | 3 |
| 2000 | Multiresolution using principal component analysisabstractThis paper proposes principal component analysis (PCA) to find adaptive bases for multiresolution. An input image is decomposed into components (compressed images) which are uncorrelated and have maximum l/sub 2/ energy. With only minor modification, a single layer linear network using the generalized Hebbian algorithm (GHA) is used for multiresolution PCA. The decomposition has been successfully applied to face classification. Good results with biological signals have also been reported. Vic Brennan, José C. Príncipe |
ICASSP | 2 |
| 2000 | Dynamic subgrouping in RTRL provides a faster O(N2) algorithmabstractStatic grouping of processing elements (PEs) has been proposed to reduce the computational complexity of real time recurrent learning (RTRL) from O(n/sup 4/) to O(n/sup 2/), but performance suffers. This paper proposes a dynamic subgrouping of PEs estimated from a local approximation of the /spl pi/ matrix based on temporal Hebbian of sensitivities during training. The method is O(n/sup 2/) and leads to better performance. Neil R. Euliano, José C. Príncipe |
ICASSP | 2 |
| 2000 | On the relationship between the Karhunen-Loeve transform and the prolate spheroidal wave functionsabstractWe find a close relationship between the discrete Karhunen-Loeve transform (KLT) and the discrete prolate spheroidal wave functions (DPSWF). We show that the DPSWF form a natural basis for an expansion of the eigenfunctions of the KLT in the frequency domain, and then determine more general conditions that any set of functions must obey to be a valid basis. We also present approximate solutions for small, medium, and large filter orders. The medium order solution suggests that the principal eigenfunction is, to a high degree of approximation, the principal DPSWF modulated so that its center frequency coincides with the peak of maximum energy in the signal spectrum. We then use this result to propose a new basis. Craig L. Fancourt, José C. Príncipe |
ICASSP | 2 |
| 2000 | A new clustering evaluation function using Renyi's information potentialabstractClustering is an important unsupervised learning paradigm, but so far the traditional methodologies are mostly based on the minimization of the variance between the data and the cluster means. Here we propose a new evaluation function based on a previously developed information theoretic measure defined from Renyi's (1960) entropy. We show how to apply Renyi's entropy to clustering and analyze the resulting staircase nature of the performance function that can be expected during learning. We suggest simulated annealing as a possible optimization criterion. Erhan Gokcay, José C. Príncipe |
ICASSP | 2 |
| 2000 | On the Use of Neural Networks in the Generalized Likelihood Ratio Test for Detecting Abrupt Changes in SignalsabstractWith the advent of efficient algorithms and fast computers for training neural networks, it is now feasible to employ neural network predictors in the generalized likelihood ratio test for the purpose of detecting abrupt nonstationary changes in the dynamics of a time series. We examine some of the special issues involved and present some simulation results validating the new hybrid algorithm. Craig L. Fancourt, José C. Príncipe |
IJCNN (2) | 2 |
| 2000 | A silicon olfactory bulb oscillatorabstractThis paper presents a low power MOS-VLSI implementation of an oscillator proposed by Freeman (1975) to model a very important component of the olfactory cortex. The model dynamics has time constants in the order of 1/220 (s). To accomplish the long time constants a new filtering technique recently proposed was utilized. All the blocks involving information processing were designed to operate below threshold (weak inversion). Vítor Grade Tavares, José C. Príncipe, John G. Harris |
ISCAS | 2 |
| 2000 | Innovating adaptive and neural systems instruction with interactive electronic booksabstractThis paper describes an integrated strategy to innovate teaching in the undergraduate classroom by appropriately utilizing information technologies. The innovation is an interactive learning environment built around an interactive electronic book (i-book). The i-book is a tight integration of a hypertext document with a simulator. The hypertext specifies the presentation order of the theoretical topics, and the simulator is an essential piece of learning. The lectures exploit the i-book, transforming the classroom into an interactive teaching laboratory. An electronic white board is used in every lecture to teach from the text and run the simulations. Our goal is to teach adaptive systems to undergraduate electrical engineering students. Undergraduates do not have the mathematical background to comprehend the equation-based approach utilized in graduate-level adaptive systems courses. We have used the i-book for undergraduate instruction with very positive results. We are now ready to extend the format for distance learning. This new teaching methodology transcends adaptive systems and can be applied to other engineering and science courses which have access to simulators. José C. Príncipe, Neil R. Euliano, Curt Lefebvre |
Proc. IEEE | 1 |
| 1999 | Nonlinear dynamic modeling of the voiced excitation for improved speech synthesisabstractThis paper describes the implementation of a waveform-based global dynamic model with the goal of capturing vocal folds variability. The residue extracted from speech by inverse filtering is pre-processed to remove phoneme dependence and is used as the input time series to the dynamic model. After training, the dynamic model is seeded with a point from the trajectory of the time series, and iterated to produce the synthetic excitation waveform. The output of the dynamic model is compared with the input time series. These comparisons confirmed that the dynamic model had captured the variability in the residue. The output of the dynamic models is used to synthesize speech using a pitch-synchronous speech synthesizer, and the output is observed to be close to natural speech. Karthik Narasimhan, José C. Príncipe, Donald G. Childers |
ICASSP | 2 |
| 1999 | Generalized anti-Hebbian learning for source separationabstractThe information-theoretic framework for source separation is highly suitable. However the choice of the nonlinearity or the estimation of the multidimensional joint probability density function are nontrivial. We propose here a generalized Gaussian model to construct a generalized blind source separation network based on the minimum entropy principle. This new separation network can suppress the interference to a significant amount compared to the traditional LMS-echo-canceler. The simulation is given to show the disparity of the performance as a varies. Finally how to choose the appropriate a in our generalized anti-Hebbian rule is discussed. Hsiao-Chun Wu, José C. Príncipe |
ICASSP | 2 |
| 1999 | Training MLPs layer-by-layer with the information potentialabstractIn the area of information processing one fundamental issue is how to measure the relationship between two variables based only on their samples. In a previous paper, the idea of information potential which was formulated from the so called quadratic mutual information was introduced, and successfully applied to problems such as blind source separation and pose estimation of SAR (synthetic aperture radar) Images. This paper shows how information potential can be used to train a MLP (multilayer perceptron) layer-by-layer, which provides evidence that the hidden layer of a MLP serves as an "information filter" which tries to best represent the desired output in that layer in the statistical sense of mutual information. Dongxin Xu, José C. Príncipe |
ICASSP | 2 |
| 1999 | Blind separation of convolutive mixturesabstractReverberant signals recorded by multiple microphones can be described as sums of sources convolved with different parameters. Blind source separation of this unknown linear system can be transformed to a set of instantaneous mixtures for every frequency band. In each frequency band, we may use the simultaneous diagonalization algorithms to separate the sources. In addition to our previous simultaneous diagonalization to minimize the Frobenius norm, we now propose another set of efficient simultaneous diagonalization algorithms based on Hadamard's inequality to make the source separation feasible in the frequency domain. José C. Príncipe, Hsiao-Chun Wu |
IJCNN | 1 |
| 1999 | An introduction to information theoretic learningabstractLearning from examples has been traditionally done with correlation or with the mean square error (MSE) criterion, in spite of the fact that learning is intrinsically related with the extraction of information from examples. The problem is that Shannon (1948) introduced the idea of information entropy which has a sound theoretical foundation but is not easy to implement in a learning from examples scenario. In this paper Renyi's entropy definition (1976) is used and integrated with a nonparametric estimator of the probability density function (Parzen window). The experimental results on blind source separation confirm the theory. Although the work is preliminary, the "information potential" method is rather general and will have many applications. José C. Príncipe, Dongxin Xu |
IJCNN | 1 |
| 1999 | Loss function for blind source separation-minimum entropy criterion and its generalized anti-Hebbian rulesabstractIn adaptive signal processing, the least-mean squares (LMS) algorithm has long been used in signal enhancement and noise cancellation but it cannot overcome the difficulty caused by the signal leakage into the reference input. Hence we have to explore more general statistical properties about the observed signals. This view corresponds to a statistical modeling of the signals using statistical measures such as a loss function, which is different from the mutual information. This paper proposes a new loss function based on generalized Gaussian distribution family, and derives new simple adaptive learning rules. Our separator based on the new generalized "anti-Hebbian rules" is also justified by the simulation on both artificial and real data with good performance. Hsiao-Chun Wu, José C. Príncipe, John G. Harris, Jui-Kuo Juan |
IJCNN | 2 |
| 1999 | Training MLPs layer-by-layer with the information potentialabstractIn the area of information processing one fundamental issue is how to measure the statistical relationship between two variables based only on their samples. The authors previously (1998) presented the idea of information potential which was formulated from the quadratic mutual information, and successfully applied it to problems such as blind source separation and pose estimation of SAR images. This paper shows how information potential can be used to train a MLP (multilayer perceptron) layer-by-layer, which provides evidence that the hidden layer of a MLP serves as an "information filter" which tries to best represent the desired output in that layer in the statistical sense of mutual information. Dongxin Xu, José C. Príncipe |
IJCNN | 2 |
| 1999 | Improving ATR performance by incorporating virtual negative examplesabstractOne common problem in learning from examples is the insufficient size of the training set. Many researchers have proposed methods to counteract this shortcoming, such as the noisy interpolation theory, hints, new distance measure (tangent distance), virtual examples, etc. This paper presents the idea of creating virtual negative examples as severe distortions of the known class patterns. Two classifiers are studied, a perceptron and a support vector machine trained to recognize objects in synthetic aperture radar (SAR) images. They utilize the training set (positive examples) to create the discriminant function of each class in the conventional way. On the other hand, the virtual negative examples will help determine the regions where the discriminant function should yield a low value. The experimental results show that incorporating the negative examples improves greatly (up to 50 percent improvement) the confuser rejection rates. José C. Príncipe |
IJCNN | 2 |
| 1999 | Effects of source-tract interaction in perception of nasalityabstractIn this paper we study the effect of source changes, caused by vocal tract load, in perception of nasality. For that we have developed an articulatory speech synthesizer, including a comprehensive nasal tract model and an interactive glottal source model. Our main objective was to investigate to what extent is necessary, in systems aimed to produce high quality synthetic sounds, to include the effect of source-tract interaction in the glottal source model when synthesizing nasal vowels. In our studies we used Portuguese nasal vowels. Portuguese uses nasalization of vowels in its phonological inventory. Changes in glottal wave, caused by the additional load of the nasal tract are more significant in vowels like [i] with low F1 and high F2. Effects are more dramatic in time rather frequency domain. Perception tests favor the idea that listener aren’t able to detect the perceptual effect of source-tract interaction changes caused by the additional coupling of the nasal tract. More tests are needed to support, or reject, this. António J. S. Teixeira, Francisco A. C. Vaz, José C. Príncipe |
EUROSPEECH | 3 |
| 1999 | Modeling the precedence effect for speech using the gamma filter
Odelia Schwartz, John G. Harris, José C. Príncipe |
Neural Networks | 3 |
| 1999 | Super-resolution of images based on local correlationsabstractAn adaptive two-step paradigm for the superresolution of optical images is developed in this paper. The procedure locally projects image samples onto a family of kernels that are learned from image data. First, an unsupervised feature extraction is performed on local neighborhood information from a training image. These features are then used to cluster the neighborhoods into disjoint sets for which an optimal mapping relating homologous neighborhoods across scales can be learned in a supervised manner. A super-resolved image is obtained through the convolution of a low-resolution test image with the established family of kernels. Results demonstrate the effectiveness of the approach. Frank M. Candocia, José C. Príncipe |
IEEE Trans. Neural Networks | 2 |
| 1999 | Training neural networks with additive noise in the desired signalabstractA new global optimization strategy for training adaptive systems such as neural networks and adaptive filters [finite or infinite impulse response (FIR or IIR)] is proposed in this paper. Instead of adding random noise to the weights as proposed in the past, additive random noise is injected directly into the desired signal. Experimental results show that this procedure also speeds up greatly the backpropagation algorithm. The method is very easy to implement in practice, preserving the backpropagation algorithm and requiring a single random generator with a monotonically decreasing step size per output channel. Hence, this is an ideal strategy to speed up supervised learning, and avoid local minima entrapment when the noise variance is appropriately scheduled. José C. Príncipe |
IEEE Trans. Neural Networks | 2 |
| 1998 | An interactive learning environment for adaptive systems instructionabstractWe have developed a new computer based learning environment to teach adaptive systems in the electrical engineering (EE) undergraduate curriculum. The learning environment is based upon an electronic book containing a hypertext document linked with a software simulator which runs interactive examples on-line, and is fully controlled by the student. This paper presents the concepts behind the project, its implementation, and discusses the teaching methodology and an evaluation of the first course offering. José C. Príncipe, Neil R. Euliano, Curt Lefebvre |
ICASSP | 1 |
| 1998 | Exploring the time-frequency microstructure of speech for blind source separationabstractThis paper explores the different frequency contents in short time segments (temporal microstructure) of speech to identify the mixing matrix in blind source separation. We propose a new method based on the eigenspread in different frequency bands to identify the segments which contain only one of the mixtures. It is much simpler to accurately estimate the mixing matrices from these segments. This short-time subband analysis trains very fast and estimates reliably the column vectors of the linear mixture. Simulation results show that our proposed method outperforms the existing model-based and competitive learning approaches in the identification of the mixing matrix for both sensor-sufficient (as many sensors as sources) and sensor-deficient (less sensors than sources) cases. Hsiao-Chun Wu, José C. Príncipe, Dongxin Xu |
ICASSP | 2 |