EDBT 2026 Demo / reviewers in the wild / expert
Deniz Erdogmus
dblp:57/3284
· DBLP profile ↗
161ranked-venue papers
22as first author
25since 2021 · last 2026
0000-0002-1114-3539ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 76 · 11 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 72 · 11 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 since 2021Systems, architecture and hardware · 5 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Computer networks · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SSplain: Sparse and Smooth Explainer for Retinopathy of Prematurity ClassificationabstractNeural networks are frequently used in medical diagnosis. However, due to their black-box nature, model explainers are used to help clinicians understand better and trust model outputs. This paper introduces an explainer method for classifying Retinopathy of Prematurity (ROP) from fundus images. Previous methods fail to generate explanations that preserve input image structures such as smoothness and sparsity. We introduce Sparse and Smooth Explainer (SSplain), a method that generates pixel-wise explanations while preserving image structures by enforcing smoothness and sparsity. This results in realistic explanations to enhance the understanding of the given black-box model. To achieve this goal, we define an optimization problem with combinatorial constraints and solve it using the Alternating Direction Method of Multipliers (ADMM). Experimental results show that SSplain outperforms commonly used explainers in terms of both post-hoc accuracy and smoothness analyses. Additionally, SSplain identifies features that are consistent with domain-understandable features that clinicians consider as discriminative factors for ROP. We also show SSplain’s generalization by applying it to additional publicly available datasets. Code is available at https://github.com/neu-spiral/SSplain. Elifnur Sunger, Tales Imbiriba, J. Peter Campbell, Deniz Erdogmus, Stratis Ioannidis, Jennifer G. Dy |
WACV | 4 |
| 2025 | Learning Physics Informed Neural ODEs with Partial MeasurementsabstractLearning dynamics governing physical and spatiotemporal processes is a challenging problem, especially in scenarios where states are partially measured. In this work, we tackle the problem of learning dynamics governing these systems when parts of the system's states are not measured, specifically when the dynamics generating the non-measured states are unknown. Inspired by state estimation theory and Physics Informed Neural ODEs, we present a sequential optimization framework in which dynamics governing unmeasured processes can be learned. We demonstrate the performance of the proposed approach leveraging numerical simulations and a real dataset extracted from an electro-mechanical positioning system. We show how the underlying equations fit into our formalism and demonstrate the improved performance of the proposed method when compared with baselines. Paul Ghanem, Ahmet Demirkaya, Tales Imbiriba, Alireza Ramezani, Zachary Danziger, Deniz Erdogmus |
AAAI | 6 |
| 2025 | MarkovType: A Markov Decision Process Strategy for Non-Invasive Brain-Computer Interfaces Typing SystemsabstractBrain-Computer Interfaces (BCIs) help people with severe speech and motor disabilities communicate and interact with their environment using neural activity. This work focuses on the Rapid Serial Visual Presentation (RSVP) paradigm of BCIs using noninvasive electroencephalography (EEG). The RSVP typing task is a recursive task with multiple sequences, where users see only a subset of symbols in each sequence. Extensive research has been conducted to improve classification in the RSVP typing task, achieving fast classification. However, these methods struggle to achieve high accuracy and do not consider the typing mechanism in the learning procedure. They apply binary target and non-target classification without including recursive training. To improve performance in the classification of symbols while controlling the classification speed, we incorporate the typing setup into training by proposing a Partially Observable Markov Decision Process (POMDP) approach. To the best of our knowledge, this is the first work to formulate the RSVP typing task as a POMDP for recursive classification. Experiments show that the proposed approach, MarkovType, results in a more accurate typing system compared to competitors. Additionally, our experiments demonstrate that while there is a trade-off between accuracy and speed, MarkovType achieves the optimal balance between these factors compared to other methods. Elifnur Sunger, Yunus Bicer, Deniz Erdogmus, Tales Imbiriba |
AAAI | 3 |
| 2024 | Learning semilinear neural operators: A unified recursive framework for prediction and data assimilationabstractRecent advances in the theory of Neural Operators (NOs) have enabled fast and accurate computation of the solutions to complex systems described by partial differential equations (PDEs). Despite their great success, current NO-based solutions face important challenges when dealing with spatio-temporal PDEs over long time scales. Specifically, the current theory of NOs does not present a systematic framework to perform data assimilation and efficiently correct the evolution of PDE solutions over time based on sparsely sampled noisy measurements. In this paper, we propose a learning-based state-space approach to compute the solution operators to infinite-dimensional semilinear PDEs. Exploiting the structure of semilinear PDEs and the theory of nonlinear observers in function spaces, we develop a flexible recursive method that allows for both prediction and data assimilation by combining prediction and correction operations. The proposed framework is capable of producing fast and accurate predictions over long time horizons, dealing with irregularly sampled noisy measurements to correct the solution, and benefits from the decoupling between the spatial and temporal dynamics of this class of PDEs. We show through experiments on the Kuramoto-Sivashinsky, Navier-Stokes and Korteweg-de Vries equations that the proposed model is robust to noise and can leverage arbitrary amounts of measurements to correct its prediction over a long time horizon with little computational overhead. Ricardo Augusto Borsoi, Deniz Erdogmus, Tales Imbiriba |
ICLR | 3 |
| 2023 | A Deep Disentangled Approach for Interpretable Hyperspectral UnmixingabstractDeep learning-based frameworks have been recently applied to hyperspectral umixing due to their flexibility and powerful representation capabilities. However, such techniques either use black-box models which are not physically interpretable, or fail to address the non-idealities of the unmixing problem. In this paper, we propose a physically interpretable deep learning method for hyperspectral unmixing accounting for nonlinearity and the variability of the endmembers. The proposed method is based on a probabilistic variational deep learning framework which employs semi-supervised disentanglement learning to properly separate the abundances and endmembers. A self-supervised strategy is used to generate labeled training data, and the model is learned end-to-end using stochastic backpropagation. Experimental results on both synthetic and real datasets illustrate the performance of the proposed method compared to state-of-the-art algorithms. Ricardo Augusto Borsoi, Tales Imbiriba, Deniz Erdogmus |
ICASSP | 3 |
| 2023 | Inv-Senet: Invariant Self Expression Network for Clustering Under Biased DataabstractSubspace clustering algorithms are used for understanding the cluster structure that explains the patterns prevalent in the dataset well. These methods are extensively used for data-exploration tasks in various areas of Natural Sciences. However, most of these methods fail to handle confounding attributes in the dataset. For datasets where a data sample represent multiple attributes, naively applying any clustering approach can result in undesired output. To this end, we propose a novel framework for jointly removing confounding attributes while learning to cluster data points in individual subspaces. Assuming we have label information about these confounding attributes, we regularize the clustering method by adversarially learning to minimize the mutual information between the data representation and the confounding attribute labels. Our experimental result on synthetic and real-world datasets demonstrate the effectiveness of our approach. Aria Masoomi, Tales Imbiriba, Erik G. Learned-Miller, Deniz Erdogmus |
ICASSP | 6 |
| 2023 | Recursive Estimation of User Intent From Noninvasive Electroencephalography Using Discriminative ModelsabstractWe study the problem of inferring user intent from noninvasive electroencephalography (EEG) to restore communication for people with severe speech and physical impairments (SSPI). The focus of this work is improving the estimation of posterior symbol probabilities in a typing task. At each iteration of the typing procedure, a subset of symbols is chosen for the next query based on the current probability estimate. Evidence about the user’s response is collected from event-related potentials (ERP) in order to update symbol probabilities, until one symbol exceeds a predefined confidence threshold. We provide a graphical model describing this task, and derive a recursive Bayesian update rule based on a discriminative probability over label vectors for each query, which we approximate using a neural network classifier. We evaluate the proposed method in a simulated typing task and show that it outperforms previous approaches based on generative modeling. Niklas Smedemark-Margulies, Basak Celik, Tales Imbiriba, Aziz Kocanaogullari, Deniz Erdogmus |
ICASSP | 5 |
| 2023 | Gaussian Process-Based Prediction of Human Trajectories to Promote Seamless Human-Robot HandoversabstractHumans can perform seamless object handovers with little to no effort. These handovers are characterized by an early movement onset that anticipates the handover location and a smooth velocity profile with minimal trajectory corrections. Replicating these characteristics in an object handover task between humans and robots presents a significant modeling challenge. In this paper we implement a Gaussian Process prediction model to serve as a robotic surrogate of human inference, and investigate how this model affects the kinematics of a human giver handing an object to the robot. Additionally, we analyze how the resulting robot kinematics compare to those of a human, and gauge human comfort through subjective reporting. Human giver kinematics during human-robot handover compared closely to human-human giver kinematics with respect to movement speed, movement timing, movement smoothness, and handover distance. Notable differences were observed in reach time and receiver peak transport velocity. When asked how well four attributes of their human-robot handovers (receiver competence, handover comfort, handover naturalness, handover safety) compared to those attributes in human-human handovers, subjects gave mean scores ranging from 4.43 (naturalness) to 5.13 (safety) on a 7 point Likert scale. Kyle Lockwood, Garrit Strenge, Yunus Bicer, Tales Imbiriba, Mariusz P. Furmanek, Taskin Padir, Deniz Erdogmus, Eugene Tunik, Mathew Yarossi |
RO-MAN | 7 |
| 2023 | Circular-symmetric correlation layer
Bahar Azari, Deniz Erdogmus |
Mach. Learn. | 2 |
| 2022 | Equivariant Deep Dynamical Model for Motion PredictionabstractLearning representations through deep generative modeling is a powerful approach for dynamical modeling to discover the most simplified and compressed underlying description of the data, to then use it for other tasks such as prediction. Most learning tasks have intrinsic symmetries, i.e., the input transformations leave the output unchanged, or the output undergoes a similar transformation. The learning process is, however, usually uninformed of these symmetries. Therefore, the learned representations for individually transformed inputs may not be meaningfully related. In this paper, we propose an SO(3) equivariant deep dynamical model (EqDDM) for motion prediction that learns a structured representation of the input space in the sense that the embedding varies with symmetry transformations. EqDDM is equipped with equivariant networks to parameterize the state-space emission and transition models. We demonstrate the superior predictive performance of the proposed model on various motion data. Bahar Azari, Deniz Erdogmus |
AISTATS | 2 |
| 2022 | NN-key: A Neural Network-Based Secret Key for Demapping OFDM SymbolsabstractGenerating custom modulation patterns as well as dynamically varying the mapping of the constellation points to their corresponding bit representations are some existing methods for mitigating eavesdropping attacks. In such cases, the custom symbol to bit mapping needs to be conveyed to the receiver through a secure and reliable channel. Instead of sending the representations of the modified symbols in regular information fields, we propose a machine learning-based approach, in which the modified symbols are encoded in the parameters of a light-weight neural network (NN). This NN is trained at the transmitter-side, sent as a secret key to the receiver, where it serves as a demapping block to recover the received symbols correctly. In addition, this paper explores the role of data augmentation during the training stage to increase the robustness of the NN with respect to the noise in the channel, as well as architecture compression to reduce transmission overhead. We validate the robustness of the proposed NN-based custom-modulation demapping approach by comparing it with demapping of a standard scheme (e.g., 16QAM), which reveals no appreciable loss in performance. We further quantitatively analyze the impact of channel and noise impairments on the demapping performance. Nasim Soltani, Yanyu Li, Deniz Erdogmus, Yanzhi Wang 0001, Kaushik R. Chowdhury |
CCNC | 3 |
| 2022 | Hybrid Neural Network Augmented Physics-based Models for Nonlinear Filtering
Tales Imbiriba, Ahmet Demirkaya, Jindrich Duník, Ondrej Straka, Deniz Erdogmus, Pau Closas |
FUSION | 5 |
| 2022 | VAST: Visual and Spectral Terrain Classification in Unstructured Multi-Class EnvironmentsabstractTerrain classification is a challenging task for robots operating in unstructured environments. Existing classification methods make simplifying assumptions, such as a reduced number of classes, clearly segmentable roads, or good lighting conditions, and focus primarily on one sensor type. These assumptions do not translate well to off-road vehicles, which operate in varying terrain conditions. To provide mobile robots with the capability to identify the terrain being traversed and avoid undesirable surface types, we propose a multimodal sensor suite capable of classifying different terrains. We capture high resolution macro images of surface texture, spectral reflectance curves, and localization data from a 9 degrees of freedom (DOF) inertial measurement unit (IMU) on 11 different terrains at different times of day. Using this dataset, we train individual neural networks on each of the modalities, and then combine their outputs in a fusion network. The fused network achieved an accuracy of 99.98% percent on the test set, exceeding the results of the best individual network component by 0.98%. We conclude that a combination of visual, spectral, and IMU data provides meaningful improvement over state of the art in terrain classification approaches. The data created for this research is available at https://github.com/RIVeR-Lab/vast_data. Nathaniel Hanson, Michael Shaham, Deniz Erdogmus, Taskin Padir |
IROS | 3 |
| 2022 | Automated deep learning-based wide-band receiver
Bahar Azari, Hai Cheng, Nasim Soltani, Haoqing Li 0001, Yanyu Li, Mauro Belgiovine, Tales Imbiriba, Salvatore D'Oro, Tommaso Melodia, Yanzhi Wang 0001, Pau Closas, Kaushik R. Chowdhury, Deniz Erdogmus |
Comput. Networks | 13 |
| 2022 | Model-Based Deep Autoencoder Networks for Nonlinear Hyperspectral UnmixingabstractAutoencoder (AEC) networks have recently emerged as a promising approach to perform unsupervised hyperspectral unmixing (HU) by associating the latent representations with the abundances, the decoder with the mixing model, and the encoder with its inverse. AECs are especially appealing for nonlinear HU since they lead to unsupervised and model-free algorithms. However, existing approaches fail to explore the fact that the encoder should invert the mixing process, which might reduce their robustness. In this letter, we propose a model-based AEC for nonlinear HU by considering the mixing model a nonlinear fluctuation over a linear mixture. Different from previous works, we show that this restriction naturally imposes a particular structure to both the encoder and decoder networks. This introduces prior information in the AEC without reducing the flexibility of the mixing model. Simulations with synthetic and real data indicate that the proposed strategy improves nonlinear HU. Haoqing Li 0001, Ricardo Augusto Borsoi, Tales Imbiriba, Pau Closas, José Carlos M. Bermudez, Deniz Erdogmus |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | Stopping Criterion Design for Recursive Bayesian Classification: Analysis and Decision GeometryabstractSystems that are based on recursive Bayesian updates for classification limit the cost of evidence collection through certain stopping/termination criteria and accordingly enforce decision making. Conventionally, two termination criteria based on pre-defined thresholds over (i) the maximum of the state posterior distribution; and (ii) the state posterior uncertainty are commonly used. In this paper, we propose a geometric interpretation over the state posterior progression and accordingly we provide a point-by-point analysis over the disadvantages of using such conventional termination criteria. For example, through the proposed geometric interpretation we show that confidence thresholds defined over maximum of the state posteriors suffer from stiffness that results in unnecessary evidence collection whereas uncertainty based thresholding methods are fragile to number of categories and terminate prematurely if some state candidates are already discovered to be unfavorable. Moreover, both types of termination methods neglect the evolution of posterior updates. We then propose a new stopping/termination criterion with a geometrical insight to overcome the limitations of these conventional methods and provide a comparison in terms of decision accuracy and speed. We validate our claims using simulations and using real experimental data obtained through a brain computer interfaced typing system. Aziz Kocanaogullari, Murat Akçakaya, Deniz Erdogmus |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Sample complexity of rank regression using pairwise comparisons
Berkan Kadioglu, Jennifer G. Dy, Deniz Erdogmus, Stratis Ioannidis |
Pattern Recognit. | 4 |
| 2022 | Active recursive Bayesian inference using Rényi information measures
Yeganeh M. Marghi, Aziz Kocanaogullari, Murat Akçakaya, Deniz Erdogmus |
Pattern Recognit. Lett. | 4 |
| 2022 | Spectral Ranking RegressionabstractWe study the problem of ranking regression, in which a dataset of rankings is used to learn Plackett–Luce scores as functions of sample features. We propose a novel spectral algorithm to accelerate learning in ranking regression. Our main technical contribution is to show that the Plackett–Luce negative log-likelihood augmented with a proximal penalty has stationary points that satisfy the balance equations of a Markov Chain. This allows us to tackle the ranking regression problem via an efficient spectral algorithm by using the Alternating Directions Method of Multipliers (ADMM). ADMM separates the learning of scores and model parameters, and in turn, enables us to devise fast spectral algorithms for ranking regression via both shallow and deep neural network (DNN) models. For shallow models, our algorithms are up to 579 times faster than the Newton’s method. For DNN models, we extend the standard ADMM via a Kullback–Leibler proximal penalty and show that this is still amenable to fast inference via a spectral approach. Compared to a state-of-the-art siamese network, our resulting algorithms are up to 175 times faster and attain better predictions by up to 26% Top-1 Accuracy and 6% Kendall-Tau correlation over five real-life ranking datasets. Ilkay Yildiz, Jennifer G. Dy, Deniz Erdogmus, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Stratis Ioannidis |
ACM Trans. Knowl. Discov. Data | 3 |
| 2021 | Deep Spectral RankingabstractLearning from ranking observations arises in many domains, and siamese deep neural networks have shown excellent inference performance in this setting. However, SGD does not scale well, as an epoch grows exponentially with the ranking observation size. We show that a spectral algorithm can be combined with deep learning methods to significantly accelerate training. We combine a spectral estimate of Plackett-Luce ranking scores with a deep model via the Alternating Directions Method of Multipliers with a Kullback-Leibler proximal penalty. Compared to a state-of-the-art siamese network, our algorithms are up to 175 times faster and attain better predictions by up to 26% Top-1 Accuracy and 6% Kendall-Tau correlation over five real-life ranking datasets. Ilkay Yildiz, Jennifer G. Dy, Deniz Erdogmus, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Stratis Ioannidis |
AISTATS | 3 |
| 2021 | NetCut: Real-Time DNN Inference Using Layer RemovalabstractDeep Learning plays a significant role in assisting humans in many aspects of their lives. As these networks tend to get deeper over time, they extract more features to increase accuracy at the cost of additional inference latency. This accuracy-performance trade-off makes it more challenging for Embedded Systems, as resource-constrained processors with strict deadlines, to deploy them efficiently. This can lead to selection of networks that can prematurely meet a specified deadline with excess slack time that could have potentially contributed to increased accuracy. In this work, we propose: (i) the concept of layer removal as a means of constructing TRimmed Networks (TRNs) that are based on removing problem-specific features of a pretrained network used in transfer learning, and (ii) NetCut, a methodology based on an empirical or an analytical latency estimator, which only proposes and retrains TRNs that can meet the application's deadline, hence reducing the exploration time significantly. We demonstrate that TRNs can expand the Pareto frontier that trades off latency and accuracy to provide networks that can meet arbitrary deadlines with potential accuracy improvement over off-the-shelf networks. Our experimental results show that such utilization of TRNs, while transferring to a simpler dataset, in combination with NetCut, can lead to the proposal of networks that can achieve relative accuracy improvement of up to 10.43% among existing off-the-shelf neural architectures while meeting a specific deadline, and 27x speedup in exploration time. Mehrshad Zandigohar, Deniz Erdogmus, Gunar Schirner |
DATE | 2 |
| 2021 | End-to-end grasping policies for human-in-the-loop robots via deep reinforcement learning*abstractState-of-the-art human-in-the-loop robot grasping is hugely suffered by Electromyography (EMG) inference robustness issues. As a workaround, researchers have been looking into integrating EMG with other signals, often in an ad hoc manner. In this paper, we are presenting a method for end-to-end training of a policy for human-in-the-loop robot grasping on real reaching trajectories. For this purpose we use Reinforcement Learning (RL) and Imitation Learning (IL) in DEXTRON (DEXTerity enviRONment), a stochastic simulation environment with real human trajectories that are augmented and selected using a Monte Carlo (MC) simulation method. We also offer a success model which once trained on the expert policy data and the RL policy roll-out transitions, can provide transparency to how the deep policy works and when it is probably going to fail. Mohammadreza Sharif, Deniz Erdogmus, Christopher Amato, Taskin Padir |
ICRA | 2 |
| 2021 | Stochastic mutual information gradient estimation for dimensionality reduction networks
Ozan Özdenizci, Deniz Erdogmus |
Inf. Sci. | 2 |
| 2021 | Geometric Analysis of Uncertainty Sampling for Dense Neural Network LayerabstractFor model adaptation of fully connected neural network layers, we provide an information geometric and sample behavioral active learning uncertainty sampling objective analysis. We identify conditions under which several uncertainty-based methods have the same performance and show that such conditions are more likely to appear in the early stages of learning. We define riskier samples for adaptation, and demonstrate that, as the set of labeled samples increases, margin-based sampling outperforms other uncertainty sampling methods by preferentially selecting these risky samples. We support our derivations and illustrations with experiments using Meta-Dataset, a benchmark for few-shot learning. We compare uncertainty-based active learning objectives using features produced by SimpleCNAPS (a state-of-the-art few-shot classifier) as input for a fully-connected adaptation layer. Our results indicate that margin-based uncertainty sampling achieves similar performance as other uncertainty based sampling methods with fewer labelled samples as discussed in the novel geometric analysis. Aziz Kocanaogullari, Niklas Smedemark-Margulies, Murat Akçakaya, Deniz Erdogmus |
IEEE Signal Process. Lett. | 4 |
| 2021 | Universal Physiological Representation Learning With Soft-Disentangled Rateless AutoencodersabstractHuman computer interaction (HCI) involves a multidisciplinary fusion of technologies, through which the control of external devices could be achieved by monitoring physiological status of users. However, physiological biosignals often vary across users and recording sessions due to unstable physical/mental conditions and task-irrelevant activities. To deal with this challenge, we propose a method of adversarial feature encoding with the concept of a Rateless Autoencoder (RAE), in order to exploit disentangled, nuisance-robust, and universal representations. We achieve a good trade-off between user-specific and task-relevant features by making use of the stochastic disentanglement of the latent representations by adopting additional adversarial networks. The proposed model is applicable to a wider range of unknown users and tasks as well as different classifiers. Results on cross-subject transfer evaluations show the advantages of the proposed framework, with up to an 11.6% improvement in the average subject-transfer classification accuracy. Mo Han, Ozan Özdenizci, Toshiaki Koike-Akino, Ye Wang 0001, Deniz Erdogmus |
IEEE J. Biomed. Health Informatics | 5 |
| 2020 | Fast and Accurate Ranking RegressionabstractWe consider a ranking regression problem in which we use a dataset of ranked choices to learn Plackett-Luce scores as functions of sample features. We solve the maximum likelihood estimation problem by using the Alternating Directions Method of Multipliers (ADMM), effectively separating the learning of scores and model parameters. This separation allows us to express scores as the stationary distribution of a continuous-time Markov Chain. Using this equivalence, we propose two spectral algorithms for ranking regression that learn model parameters up to 579 times faster than the Newton’s method. Ilkay Yildiz, Jennifer G. Dy, Deniz Erdogmus, Jayashree Kalpathy-Cramer, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Stratis Ioannidis |
AISTATS | 3 |
| 2020 | Kinematic Optimization of an Underactuated Anthropomorphic Prosthetic HandabstractThe human hand serves as an inspiration for robotic grippers. However, the dimensions of the human hand evolved under a different set of constraints and requirements than that of robots today. This paper discusses a method of kinematically optimizing the design of an anthropomorphic robotic hand. We focus on maximizing the workspace intersection of the thumb and the other fingers as well as maximizing the size of the largest graspable object. We perform this optimization and use the resulting dimensions to construct a flexible, underactuated 3D printed prototype. We verify the results of the optimization through experimentation, demonstrating that the optimized hand is capable of grasping objects ranging from less than 1 mm to 12.8 cm in diameter with a high degree of reliability. The hand is lightweight and inexpensive, weighing 333 g and costing less than 175 USD, and strong enough to lift over 1.1 lb (500 g). We demonstrate that the optimized hand outperforms an open-source 3D printed anthropomorphic hand on multiple tasks. Finally, we demonstrate the performance of our hand by employing a classification-based user intent decision system which predicts the grasp type using real-time electromyographic (EMG) activity patterns. Ann Marie Votta, Sezen Yagmur Günay, Brian Zylich, Erik H. Skorina, Raagini Rameshwar, Deniz Erdogmus, Cagdas D. Onal |
IROS | 6 |
| 2020 | Disentangled Adversarial Autoencoder for Subject-Invariant Physiological Feature ExtractionabstractRecent developments in biosignal processing have enabled users to exploit their physiological status for manipulating devices in a reliable and safe manner. One major challenge of physiological sensing lies in the variability of biosignals across different users and tasks. To address this issue, we propose an adversarial feature extractor for transfer learning to exploit disentangled universal representations. We consider the trade-off between task-relevant features and user-discriminative information by introducing additional adversary and nuisance networks in order to manipulate the latent representations such that the learned feature extractor is applicable to unknown users and various tasks. Results on cross-subject transfer evaluations exhibit the benefits of the proposed framework, with up to 8.8% improvement in average accuracy of classification, and demonstrate adaptability to a broader range of subjects. Mo Han, Ozan Özdenizci, Ye Wang 0001, Toshiaki Koike-Akino, Deniz Erdogmus |
IEEE Signal Process. Lett. | 5 |
| 2019 | Variational Inference from Ranked Samples with FeaturesabstractIn many supervised learning settings, elicited labels comprise pairwise comparisons or rankings of samples. We propose a Bayesian inference model for ranking datasets, allowing us to take a probabilistic approach to ranking inference. Our probabilistic assumptions are motivated by, and consistent with, the so-called Plackett-Luce model. We propose a variational inference method to extract a closed-form Gaussian posterior distribution. We show experimentally that the resulting posterior yields more reliable ranking predictions compared to predictions via point estimates. Jennifer G. Dy, Deniz Erdogmus, Jayashree Kalpathy-Cramer, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Stratis Ioannidis |
ACML | 3 |
| 2019 | A History-based Stopping Criterion in Recursive Bayesian State EstimationabstractIn dynamic state-space models, the state can be estimated through recursive computation of the posterior distribution of the state given all measurements. In scenarios where active sensing/querying is possible, a hard decision is made when the state posterior achieves a pre-set confidence threshold. This mandate to meet a hard threshold may sometimes unnecessarily require more queries. In application domains where sensing/querying cost is of concern, some potential accuracy may be sacrificed for greater gains in sensing cost. In this paper, we (a) propose a criterion based on a linear combination of state posterior and its changes, (b) show that for discrete-valued state estimation scenarios the proposed objective is more likely to sort correct and incorrect estimates appropriately compared to just looking at the posterior, and finally (c) demonstrate that the method can lead to significant human intent estimation speed increase without significant loss of accuracy in a brain-computer interface application. Yeganeh M. Marghi, Aziz Kocanaogullari, Murat Akçakaya, Deniz Erdogmus |
ICASSP | 4 |
| 2019 | Structured Adversarial Attack: Towards General Implementation and Better Interpretability
Kaidi Xu, Sijia Liu 0001, Pu Zhao 0001, Huan Zhang 0001, Quanfu Fan, Deniz Erdogmus, Yanzhi Wang 0001, Xue Lin 0001 |
ICLR (Poster) | 7 |
| 2019 | A Severity Score for Retinopathy of PrematurityabstractRetinopathy of Prematurity (ROP) is a leading cause for childhood blindness worldwide. An automated ROP detection system could significantly improve the chance of a child receiving proper diagnosis and treatment. We propose a means of producing a continuous severity score in an automated fashion, regressed from both (a) diagnostic class labels as well as (b) comparison outcomes. Our generative model combines the two sources, and successfully addresses inherent variability in diagnostic outcomes. In particular, our method exhibits an excellent predictive performance of both diagnostic and comparison outcomes over a broad array of metrics, including AUC, precision, and recall. Jayashree Kalpathy-Cramer, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Jennifer G. Dy, Deniz Erdogmus, Stratis Ioannidis |
KDD | 8 |
| 2019 | Accelerated Experimental Design for Pairwise ComparisonsabstractPairwise comparison labels are more informative and less variable than class labels, but generating them poses a challenge: their number grows quadratically in the dataset size. We study a natural experimental design objective, namely, D-optimality, that can be used to identify which K pairwise comparisons to generate. This objective is known to perform well in practice, and is submodular, making the selection approximable via the greedy algorithm. A naïve greedy implementation has O(N2 d2 K) complexity, where N is the dataset size, d is the feature space dimension, and K is the number of generated comparisons. We show that, by exploiting the inherent geometry of the dataset–namely, that it consists of pairwise comparisons–the greedy algorithm's complexity can be reduced to O(N2 (K + d) + N(dK + d2) + d2 K). We apply the same acceleration also to the so-called lazy greedy algorithm. When combined, the above improvements lead to an execution time of less than 1 hour for a dataset with 108 comparisons; the naïve greedy algorithm on the same dataset would require more than 10 days to terminate. Jennifer G. Dy, Deniz Erdogmus, Jayashree Kalpathy-Cramer, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Stratis Ioannidis |
SDM | 3 |
| 2019 | Classification and comparison via neural networks
Ilkay Yildiz, Jennifer G. Dy, Deniz Erdogmus, James M. Brown 0001, Jayashree Kalpathy-Cramer, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Stratis Ioannidis |
Neural Networks | 4 |
| 2019 | Target tracking via recursive Bayesian state estimation in cognitive radar networks
Yijian Xiang, Murat Akçakaya, Satyabrata Sen, Deniz Erdogmus, Arye Nehorai |
Signal Process. | 4 |
| 2019 | Adversarial Deep Learning in EEG BiometricsabstractDeep learning methods for person identification based on electroencephalographic (EEG) brain activity encounters the problem of exploiting the temporally correlated structures or recording session specific variability within EEG. Furthermore, recent methods have mostly trained and evaluated based on single session EEG data. We address this problem from an invariant representation learning perspective. We propose an adversarial inference approach to extend such deep learning models to learn session-invariant person-discriminative representations that can provide robustness in terms of longitudinal usability. Using adversarial learning within a deep convolutional network, we empirically assess and show improvements with our approach based on longitudinally collected EEG data for person identification from half-second EEG epochs. Ozan Özdenizci, Ye Wang 0001, Toshiaki Koike-Akino, Deniz Erdogmus |
IEEE Signal Process. Lett. | 4 |
| 2019 | Real-Time Deep Pose Estimation With Geodesic Loss for Image-to-Template Rigid RegistrationabstractWith an aim to increase the capture range and accelerate the performance of state-of-the-art inter-subject and subject-to-template 3-D rigid registration, we propose deep learning-based methods that are trained to find the 3-D position of arbitrarily-oriented subjects or anatomy in a canonical space based on slices or volumes of medical images. For this, we propose regression convolutional neural networks (CNNs) that learn to predict the angle-axis representation of 3-D rotations and translations using image features. We use and compare mean square error and geodesic loss to train regression CNNs for 3-D pose estimation used in two different scenarios: slice-to-volume registration and volume-to-volume registration. As an exemplary application, we applied the proposed methods to register arbitrarily oriented reconstructed images of fetuses scanned in-utero at a wide gestational age range to a standard atlas space. Our results show that in such registration applications that are amendable to learning, the proposed deep learning methods with geodesic loss minimization achieved 3-D pose estimation with a wide capture range in real-time (<100ms). We also tested the generalization capability of the trained CNNs on an expanded age range and on images of newborn subjects with similar and different MR image contrasts. We trained our models on T2-weighted fetal brain MRI scans and used them to predict the 3-D pose of newborn brains based on T1-weighted MRI scans. We showed that the trained models generalized well for the new domain when we performed image contrast transfer through a conditional generative adversarial network. This indicates that the domain of application of the trained deep regression CNNs can be further expanded to image modalities and contrasts other than those used in training. A combination of our proposed methods with accelerated optimization-based registration algorithms can dramatically enhance the performance of automatic imaging devices and image processing methods of the future. Seyed Sadegh Mohseni Salehi, Shadab Khan, Deniz Erdogmus, Ali Gholipour |
IEEE Trans. Medical Imaging | 3 |
| 2018 | Experimental Design under the Bradley-Terry ModelabstractLabels generated by human experts via comparisons exhibit smaller variance compared to traditional sample labels. Collecting comparison labels is challenging over large datasets, as the number of comparisons grows quadratically with the dataset size. We study the following experimental design problem: given a budget of expert comparisons, and a set of existing sample labels, we determine the comparison labels to collect that lead to the highest classification improvement. We study several experimental design objectives motivated by the Bradley-Terry model. The resulting optimization problems amount to maximizing submodular functions. We experimentally evaluate the performance of these methods over synthetic and real-life datasets. Jayashree Kalpathy-Cramer, Susan Ostmo, J. Peter Campbell, Michael F. Chiang, Deniz Erdogmus, Jennifer G. Dy, Stratis Ioannidis |
IJCAI | 7 |
| 2018 | A Probabilistic Active Learning Algorithm Based on Fisher Information RatioabstractThe task of labeling samples is demanding and expensive. Active learning aims to generate the smallest possible training data set that results in a classifier with high performance in the test phase. It usually consists of two steps of selecting a set of queries and requesting their labels. Among the suggested objectives to score the query sets, information theoretic measures have become very popular. Yet among them, those based on Fisher information (FI) have the advantage of considering the diversity among the queries and tractable computations. In this work, we provide a practical algorithm based on Fisher information ratio to obtain query distribution for a general framework where, in contrast to the previous FI-based querying methods, we make no assumptions over the test distribution. The empirical results on synthetic and real-world data sets indicate that this algorithm gives competitive results. Jamshid Sourati, Murat Akçakaya, Deniz Erdogmus, Todd K. Leen, Jennifer G. Dy |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | On Analysis of Active Querying for Recursive State EstimationabstractIn stochastic linear/non-linear active dynamic systems, states are estimated with the evidence through recursive measurements in response to queries of the system about the state to be estimated. Therefore, query selection is essential for such systems to improve state estimation accuracy and time. Query selection is conventionally achieved by minimization of the evidence variance or optimization of various information theoretic objectives. It was shown that optimization of mutual information-based objectives and variance-based objectives arrive at the same solution. However, existing approaches optimize approximations to the intended objectives rather than solving the exact optimization problems. To overcome these shortcomings, we propose an active querying procedure using mutual information maximization in recursive state estimation. First we show that mutual information generalizes variance based query selection methods and show the equivalence between objectives if the evidence likelihoods have unimodal distributions. We then solve the exact optimization problem for query selection and propose a query (measurement) selection algorithm. We specifically formulate the mutual information maximization for query selection as a combinatorial optimization problem and show that the objective is sub-modular, therefore can be solved efficiently with guaranteed convergence bounds through a greedy approach. Additionally, we analyze the performance of the query selection algorithm by testing it through a brain computer interface typing system. Aziz Kocanaogullari, Deniz Erdogmus, Murat Akçakaya |
IEEE Signal Process. Lett. | 2 |
| 2018 | Optimal Query Selection Using Multi-Armed BanditsabstractQuery selection for latent variable estimation is conventionally performed by opting for observations with low noise or optimizing information theoretic objectives related to reducing the level of estimated uncertainty based on the current best estimate. In these approaches, typically the system makes a decision by leveraging the current available information about the state. However, trusting the current best estimate results in poor query selection when truth is far from the current estimate, and this negatively impacts the speed and accuracy of the latent variable estimation procedure. We introduce a novel sequential adaptive action value function for query selection using the multi-armed bandit (MAB) framework which allows us to find a tractable solution. For this adaptive-sequential query selection method, we analytically show: (i) performance improvement in the query selection for a dynamical system, (ii) the conditions where the model outperforms competitors. We also present favorable empirical assessments of the performance for this method, compared to alternative methods, both using Monte Carlo simulations and human-in-the-loop experiments with a brain computer interface (BCI) typing system where the language model provides the prior information. Aziz Kocanaogullari, Yeganeh M. Marghi, Murat Akçakaya, Deniz Erdogmus |
IEEE Signal Process. Lett. | 4 |
| 2017 | Asymptotic Analysis of Objectives Based on Fisher Information in Active LearningabstractObtaining labels can be costly and time-consuming. Active learning allows a learning algorithm to intelligently query samples to be labeled for a more efficient learning. Fisher information ratio (FIR) has been used as an objective for selecting queries. However, little is known about the theory behind the use of FIR for active learning. There is a gap between the underlying theory and the motivation of its usage in practice. In this paper, we attempt to fill this gap and provide a rigorous framework for analyzing existing FIR-based active learning methods. In particular, we show that FIR can be asymptotically viewed as an upper bound of the expected variance of the log-likelihood ratio. Additionally, our analysis suggests a unifying framework that not only enables us to make theoretical comparisons among the existing querying methods based on FIR, but also allows us to give insight into the development of new active learning approaches based on this objective. Jamshid Sourati, Murat Akçakaya, Todd K. Leen, Deniz Erdogmus, Jennifer G. Dy |
J. Mach. Learn. Res. | 4 |
| 2017 | Subject-specific abnormal region detection in traumatic brain injury using sparse model selection on high dimensional diffusion data
Matineh Shaker, Deniz Erdogmus, Jennifer G. Dy, Sylvain Bouix |
Medical Image Anal. | 2 |
| 2017 | Spatio-temporal EEG models for brain interfaces
Paula Gonzalez-Navarro, Mohammad Moghadamfalahi, Murat Akçakaya, Deniz Erdogmus |
Signal Process. | 4 |
| 2017 | Auto-Context Convolutional Neural Network (Auto-Net) for Brain Extraction in Magnetic Resonance ImagingabstractBrain extraction or whole brain segmentation is an important first step in many of the neuroimage analysis pipelines. The accuracy and the robustness of brain extraction, therefore, are crucial for the accuracy of the entire brain analysis process. The state-of-the-art brain extraction techniques rely heavily on the accuracy of alignment or registration between brain atlases and query brain anatomy, and/or make assumptions about the image geometry, and therefore have limited success when these assumptions do not hold or image registration fails. With the aim of designing an accurate, learning-based, geometry-independent, and registration-free brain extraction tool, in this paper, we present a technique based on an auto-context convolutional neural network (CNN), in which intrinsic local and global image features are learned through 2-D patches of different window sizes. We consider two different architectures: 1) a voxelwise approach based on three parallel 2-D convolutional pathways for three different directions (axial, coronal, and sagittal) that implicitly learn 3-D image information without the need for computationally expensive 3-D convolutions and 2) a fully convolutional network based on the U-net architecture. Posterior probability maps generated by the networks are used iteratively as context information along with the original image patches to learn the local shape and connectedness of the brain to extract it from non-brain tissue. The brain extraction results we have obtained from our CNNs are superior to the recently reported results in the literature on two publicly available benchmark data sets, namely, LPBA40 and OASIS, in which we obtained the Dice overlap coefficients of 97.73% and 97.62%, respectively. Significant improvement was achieved via our auto-context algorithm. Furthermore, we evaluated the performance of our algorithm in the challenging problem of extracting arbitrarily oriented fetal brains in reconstructed fetal brain magnetic resonance imaging (MRI) data sets. In this application, our voxelwise auto-context CNN performed much better than the other methods (Dice coefficient: 95.97%), where the other methods performed poorly due to the non-standard orientation and geometry of the fetal brain in MRI. Through training, our method can provide accurate brain extraction in challenging applications. This, in turn, may reduce the problems associated with image registration in segmentation tasks. Seyed Sadegh Mohseni Salehi, Deniz Erdogmus, Ali Gholipour |
IEEE Trans. Medical Imaging | 2 |
| 2015 | Minor Surfaces are Boundaries of Mode-Based ClustersabstractWe show that mode-based cluster boundaries exhibit themselves as minor surfaces of the data probability density function. Based on this result, we provide a connectivity measure depending on minor surface search between sample pairs. Accordingly, we build a connectivity graph among data samples. The use of graph construction is particularly demonstrated for clustering, but applications in other machine learning areas are possible. On Gaussian mixture and kernel density estimate type probability density models, we illustrate the theoretical results with examples and demonstrate that cluster boundaries between sample pairs can be detected using a line integral. We also demonstrate an example where the data distribution has a continuous line segment as its set of local maxima (not strict), for which mean-shift like gradient flow and other mode-seeking algorithms fail to identify a single cluster, while the proposed approach successfully determines this fact. Esra Ataer Cansizoglu, Murat Akçakaya, Deniz Erdogmus |
IEEE Signal Process. Lett. | 3 |
| 2015 | A Bayesian Framework for Intent Detection and Stimulation Selection in SSVEP BCIsabstractCurrently, many Brain Computer Interfaces (BCI) classifiers output point estimates of user intent which make it difficult to incorporate context prior information or assign a principled confidence measurement to a decision. We propose a Bayesian framework to extend current Steady State Visually Evoked Potential (SSVEP) classifiers to a maximum a posteriori (MAP) classifiers by using a Kernel Density Estimate (KDE) to learn the distribution of features conditioned on stimulation class. To demonstrate our framework we extend Canonical Correlation Analysis (CCA) and Power Spectral Density (PSD) style methods. Traditionally, in either example, the class is estimated as the class associated with the maximum feature. Our framework increases performance by relaxing the assumption that a stimulation class's sample often maximizes its class-associated feature. Further, by leveraging the KDE, we present a method which estimates the performance of a classifier under different stimulation frequency sets. Using this, we optimize the selection of stimulation frequencies from those present in a training set. Matt Higger, Murat Akçakaya, Hooman Nezamfar, Gerald LaMountain, Umut Orhan, Deniz Erdogmus |
IEEE Signal Process. Lett. | 6 |
| 2014 | Moving towards a real-time system for automatically recognizing stereotypical motor movements in individuals on the autism spectrum using wireless accelerometryabstractThis paper extends previous work automatically detecting stereotypical motor movements (SMM) in individuals on the autism spectrum. Using three-axis accelerometer data obtained through wearable wireless sensors, we compare recognition results for two different classifiers -- Support Vector Machine and Decision Tree -- in combination with different feature sets based on time-frequency characteristics of accelerometer data. We use data collected from six individuals on the autism spectrum who participated in two different studies conducted three years apart in classroom settings, and observe an average accuracy across all participants over time ranging from 81.2% (TPR: 0.91; FPR: 0.21) to 99.1% (TPR: 0.99; FPR: 0.01) for all combinations of classifiers and feature sets. We also provide analyses of kinematic parameters associated with observed movements in an attempt to explain classifier-feature specific performance. Based on our results, we conclude that real-time, person-dependent, adaptive algorithms are needed in order to accurately and consistently measure SMM automatically in individuals on the autism spectrum over time in real-word settings. Matthew S. Goodwin, Marzieh Haghighi, Qu Tang, Murat Akçakaya, Deniz Erdogmus, Stephen S. Intille |
UbiComp | 5 |
| 2014 | Manifold learning by preserving distance orders
Esra Ataer Cansizoglu, Murat Akçakaya, Umut Orhan, Deniz Erdogmus |
Pattern Recognit. Lett. | 4 |
| 2014 | Accelerated Learning-Based Interactive Image Segmentation Using Pairwise ConstraintsabstractAlgorithms for fully automatic segmentation of images are often not sufficiently generic with suitable accuracy, and fully manual segmentation is not practical in many settings. There is a need for semiautomatic algorithms, which are capable of interacting with the user and taking into account the collected feedback. Typically, such methods have simply incorporated user feedback directly. Here, we employ active learning of optimal queries to guide user interaction. Our work in this paper is based on constrained spectral clustering that iteratively incorporates user feedback by propagating it through the calculated affinities. The original framework does not scale well to large data sets, and hence is not straightforward to apply to interactive image segmentation. In order to address this issue, we adopt advanced numerical methods for eigen-decomposition implemented over a subsampling scheme. Our key innovation, however, is an active learning strategy that chooses pairwise queries to present to the user in order to increase the rate of learning from the feedback. Performance evaluation is carried out on the Berkeley segmentation and Graz-02 image data sets, confirming that convergence to high accuracy levels is realizable in relatively few iterations. Jamshid Sourati, Deniz Erdogmus, Jennifer G. Dy, Dana H. Brooks |
IEEE Trans. Image Process. | 2 |
| 2013 | Improved inference and autotyping in EEG-based BCI typing systemsabstractThe RSVP Keyboard™ is a brain-computer interface (BCI)-based typing system for people with severe physical disabilities, specifically those with locked-in syndrome (LIS). It uses signals from an electroencephalogram (EEG) combined with information from an n-gram language model to select letters to be typed. One characteristic of the system as currently configured is that it does not keep track of past EEG observations, i.e., observations of user intent made while the user was in a different part of a typed message. We present a principled approach for taking all past observations into account, and show that this method results in a 20% increase in simulated typing speed under a variety of conditions on realistic stimuli. We also show that this method allows for a principled and improved estimate of the probability of the backspace symbol, by which mis-typed symbols are corrected. Finally, we demonstrate the utility of automatically typing likely letters in certain contexts, a technique that achieves increased typing speed under our new method, though not under the baseline approach. Andrew Fowler, Brian Roark, Umut Orhan, Deniz Erdogmus, Melanie Fried-Oken |
ASSETS | 4 |
| 2013 | Contour-based shape representation using principal curves
Esra Ataer Cansizoglu, Erhan Bas, Jayashree Kalpathy-Cramer, Gregory C. Sharp, Deniz Erdogmus |
Pattern Recognit. | 5 |
| 2013 | A Robust Fusion Algorithm for Sensor FailureabstractAccurate multimodal and multisensor detection of a target phenomenon requires knowledge of probabilistic sensor characteristics to determine an appropriate fusion rule which optimizes an objective of interest, traditionally the expected Bayesian risk. However, a particular sensor characteristic can change online, introducing unaccounted additional risk to the fusion rule that was based on assumed sensor specifications. To mitigate such changes, we propose a sensor-failure-robust fusion rule assuming that only first order characteristics of a probabilistic sensor failure model are known. Under this failure model, we compute the expected Bayesian risk and minimize this risk to develop the proposed fusion method. Matt Higger, Murat Akçakaya, Deniz Erdogmus |
IEEE Signal Process. Lett. | 3 |
| 2012 | A mode-based clustering algorithm without mode seekingabstractMode-based clustering approaches such as mean-shift and its variants are extremely successful. They are also computationally expensive due to their iterative hill-climbing strategy when determining cluster labels for samples. We identify the fact that mode-based cluster boundaries exhibit themselves as minor surfaces of the data distribution. Based on this observation, we develop a mode-based clustering methodology that does not involve iterative hill climbing for each sample. The method, instead, is based on searching for the presence of a minor surface on a path that connects pairs of samples. The pairwise data connections, when evaluated efficiently, may lead to a simple graph connectivity matrix based on which clusters can be identified using connected components. This search efficiency is achieved by an agglomerative clustering approach in the particular proposition presented in this paper. Illustrative experiments are carried out on synthetic datasets using Gaussian mixture models and kernel density estimates. Esra Ataer Cansizoglu, Deniz Erdogmus |
ICASSP | 2 |
| 2012 | RSVP keyboard: An EEG based typing interfaceabstractHumans need communication. The desire to communicate remains one of the primary issues for people with locked-in syndrome (LIS). While many assistive and augmentative communication systems that use various physiological signals are available commercially, the need is not satisfactorily met. Brain interfaces, in particular, those that utilize event related potentials (ERP) in electroencephalography (EEG) to detect the intent of a person noninvasively, are emerging as a promising communication interface to meet this need where existing options are insufficient. Existing brain interfaces for typing use many repetitions of the visual stimuli in order to increase accuracy at the cost of speed. However, speed is also crucial and is an integral portion of peer-to-peer communication; a message that is not delivered timely often looses its importance. Consequently, we utilize rapid serial visual presentation (RSVP) in conjunction with language models in order to assist letter selection during the brain-typing process with the final goal of developing a system that achieves high accuracy and speed simultaneously. This paper presents initial results from the RSVP Keyboard system that is under development. These initial results on healthy and locked-in subjects show that single-trial or few-trial accurate letter selection may be possible with the RSVP Keyboard paradigm. Umut Orhan, Kenneth E. Hild II, Deniz Erdogmus, Brian Roark, Barry Oken, Melanie Fried-Oken |
ICASSP | 3 |
| 2012 | Unsupervised wrinkle detection in reflectance confocal microscopy images of the human skinabstractReflectance confocal microscopy (RCM) is a non-invasive and in-vivo imaging modality, which can take images from different depths of the human skin. A challenging problem is to detect a clinically important subsurface section of the skin, the Dermis/Epidermis junction, in RCM images. This is a tough problem because of the huge variation of texture and intensity features across both intersubject and intrasubject tissues. On the other hand, there's almost no wrinkle-free part of the skin. This well-known phenomenon can be used as a histological clue for guessing the probability of being Dermis or Epidermis in the neighboring regions. In this paper, we develop a two-step wrinkle detector for RCM images. By analyzing the results on different RCM images, we conclude it has high sensitivity and specificity, but a relatively lower Jaccard index. Jamshid Sourati, Dana H. Brooks, Jennifer G. Dy, Esra Ataer Cansizoglu, Deniz Erdogmus, Milind Rajadhyaksha |
ICASSP | 5 |
| 2012 | Local tracing of curvilinear structures in volumetric color images: Application to the Brainbow analysis
Erhan Bas, Deniz Erdogmus, R. W. Draft, Jeff Lichtman |
J. Vis. Commun. Image Represent. | 2 |
| 2011 | Polytope kernel density estimates on Delaunay graphsabstractWe present a polytope-kernel density estimation (PKDE) methodology that allows us to perform exact mean-shift up dates along the edges of the Delaunay graph of the data. We discuss explicit and implicit constructions of such a PKDE, where in the implicit construction one can exploit a smoother kernel such as the standard isotropic Gaussian. The resulting density estimate allows us to perform mean-shift clustering in a computationally efficient manner (similar to mediod shift), but in a manner that is exact and consistent with the underlying density assumption. The procedure also yields a hierarchical connectivity structure, a tree, that spans the dataset. We demonstrate how this tree, combined with density-weighted geodesic distance calculations between modal samples can be used to select number of clusters as well as a distance preserving dimension reduction technique. Erhan Bas, Deniz Erdogmus |
ICASSP | 2 |
| 2011 | Sampling on locally defined principal manifoldsabstractWe start with a locally defined principal curve definition for a given probability density function (pdf) and define a pairwise manifold score based on local derivatives of the pdf. Proposed manifold score can be used to check if data pairs lie on the same manifold. We use this score to (i) cluster nonlinear manifolds having irregular shapes, and (ii) (down)sample a selected principal curve with sufficient accuracy sparsely. Our goal is to provide a heuristic-free formulation for principal graph generation and curve parametrization in order to form a basis for a principled principal manifold unwrapping method. Erhan Bas, Deniz Erdogmus |
ICASSP | 2 |
| 2011 | On visually evoked potentials in eeg induced by multiple pseudorandom binary sequences for brain computer interface designabstractVisually evoked potentials have attracted great attention in the last two decades for the purpose of brain computer interface design. Visually evoked P300 response is a major signal of interest that has been widely studied. Steady state visual evoked potentials that occur in response to periodically flickering visual stimuli have been primarily investigated as an alternative. There also exists some work on the use of an m-sequence and its shifted versions to induce responses that are primarily in the visual cortex but are not periodic. In this paper, we study the use of multiple m-sequences for intent discrimination in the brain interface, as opposed to a single m-sequence whose shifted versions are to be discriminated from each other. Specifically we used four different m-sequences of length 31. Our main goal is to study if the bit presentation rate of the m-sequences have an impact on classification accuracy and speed. In this initial study, where we compared two basic classifier schemes using EEG data acquired with 15Hz and 30Hz bit presentation rates, our results are mixed; while on one subject, we got promising results indicating bit presentation rate could be increased without decrease in classification accuracy; thus leading to a faster decision-rate in the brain interface, on our second subject, this conclusion is not supported. Further detailed experimental studies as well as signal processing methodology design, especially for information fusion across EEG channels, will be conducted to investigate this question further. Hooman Nezamfar, Umut Orhan, Deniz Erdogmus, Kenneth E. Hild II, Shalini Purwar, Barry Oken, Melanie Fried-Oken |
ICASSP | 3 |
| 2011 | A framework for rapid visual image search using single-trial brain evoked responses
Yonghong Huang, Deniz Erdogmus, Misha Pavel, Santosh Mathan, Kenneth E. Hild II |
Neurocomputing | 2 |
| 2011 | Locally Defined Principal Curves and Surfaces
Umut Ozertem, Deniz Erdogmus |
J. Mach. Learn. Res. | 2 |
| 2011 | Connectivity of projected high dimensional data charts on one-dimensional curves
Erhan Bas, Deniz Erdogmus |
Signal Process. | 2 |
| 2010 | Principal curve tracing
Erhan Bas, Deniz Erdogmus |
ESANN | 2 |
| 2010 | Identifying informative features for ERP speller systems based on RSVP paradigm
Tian Lan 0007, Deniz Erdogmus, Lois M. Black, Jan P. H. van Santen |
ESANN | 2 |
| 2010 | A Novel Application of Principal Surfaces to Segmentation in 4D-CT for Radiation Treatment PlanningabstractRadiation therapy is one of the most effective options used in the treatment of about half of all people with cancer. A critical goal in radiation therapy is to deliver optimal radiation doses to the observed tumor while sparing the surrounding healthy tissues. Radiation oncologists typically manually delineate normal and diseased structures on three-dimensional computed tomography~(3D-CT) scans. Manual delineation is a labor intensive, tedious and time-consuming task. In recent years, concerns about respiration induced motion have led to the popularity of four-dimensional computed tomography~(4D-CT) for the tracking of tumors and deformation of organs. However, as manually contouring in all phases would be prohibitively expensive, the development of fast, robust, and automatic segmentation tools has been an active area of research in 4D radiotherapy. In this paper, we describe a novel application of principal surfaces for the propagation of contours in 4D-CT studies. Regions of interest~(ROIs) are manually delineated slice-by-slice in the reference 3D-CT scans. Edges are detected on all of the slices of the target 3D-CT phase. A kernel density estimation~(KDE) based on the detected edges is then calculated. The principal surface algorithm is applied to find the ridges of the edge KDE to provide the object contours. Manually drawn contours from the reference phase are used as an initialization. Contours of ROIs are propagated recursively in all consecutive phases to complete a respiration cycle. Results are provided for a phantom data set of simulated tumor motion as well as on a de-identified data set of the lung of a patient. Evaluation of the efficacy of automatic segmentation in organs and tumors are based on the comparison between manually drawn contours and automatically delineated contours. The Dice coefficients are approximately 0.97 for the lung tumor on the phantom data sets and 0.95 for the patient data sets. The centroid distances between manually delineated lung volume and automatically segmented lung volume in each CT direction are <; 1 mm for both phantom data sets and patient data sets. Sheng You, Esra Ataer Cansizoglu, Deniz Erdogmus, James Tanyi, Jayashree Kalpathy-Cramer |
ICMLA | 3 |
| 2010 | Numerical optimization of a sum-of-rank-1 decomposition for n-dimensional order-p symmetric tensors
Olexiy O. Kyrgyzov, Deniz Erdogmus |
Neurocomputing | 2 |
| 2009 | Target detection using incremental learning on single-trial evoked responseabstractThe human neural responses associated with cognitive events, referred as event related potentials (ERPs), can provide reliable inference for target image detection. Incremental learning has been widely investigated to deal with large datasets. To solve the problem of data growing over time in cross session studies, we apply an incremental learning support vector machines (SVM) method on single-trial ERP detection for identifying targets in satellite images. We implement the incremental learning SVM by keeping only the support vectors, instead of all the data, from the previous sessions and incorporating them with the data of the current session. Thus the incremental learning dramatically reduces the computational load. The results demonstrate that the incremental learning ERP detection system performs as well as the naive method, which uses only the current training session, and the batch mode, which uses all training data. Furthermore, it is more computationally efficient, which allows it to better cope with a continuous stream of EEG data. Yonghong Huang, Deniz Erdogmus, Misha Pavel, Kenneth E. Hild II, Santosh Mathan |
ICASSP | 2 |
| 2009 | A new approach for the reassignment of time-frequency representationsabstractThe reassignment method is a widespread approach for obtaining high resolution time-frequency representations. Nevertheless, its performance is not always optimal and can deteriorate for low signal-to-noise ratio (SNR) values. In order to overcome these obstacles, a novel method for obtaining high resolution time-frequency representations is proposed in this paper. The new method implements proposed nonparametric snakes in order to obtain accurate locations of the signal ridges in the time-frequency domain. The results of numerical analysis show that the proposed method is capable of achieving significantly higher concentration of signals in the time-frequency domain in comparison to the spectrogram and the traditional reassignment method. Furthermore, the new scheme also maintains good performance for low SNR values, while the performance of the other two considered methods significantly diminishes. It is clear from the results that the proposed method might be of significance in applications where accurate estimation of the signal components is required for low SNR values. Ervin Sejdic, Umut Ozertem, Igor Djurovic, Deniz Erdogmus |
ICASSP | 4 |
| 2009 | A Novel LMS Algorithm Applied to Adaptive Noise CancellationabstractIn this letter, we propose a novel least-mean-square (LMS) algorithm for filtering speech sounds in the adaptive noise cancellation (ANC) problem. It is based on the minimization of the squared Euclidean norm of the difference weight vector under a stability constraint defined over theaposterioriestimation error. To this purpose, the Lagrangian methodology has been used in order to propose a nonlinear adaptation rule defined in terms of the product of differential inputs and errors which means a generalization of the normalized (N)LMS algorithm. The proposed method yields better tracking ability in this context as shown in the experiments which are carried out on the AURORA 2 and 3 speech databases. They provide an extensive performance evaluation along with an exhaustive comparison to standard LMS algorithms with almost the same computational load, including the NLMS and other recently reported LMS algorithms such as the modified (M)-NLMS, the error nonlinearity (EN)-LMS, or the normalized data nonlinearity (NDN)-LMS adaptation. Juan Manuel Górriz, Javier Ramírez 0001, Sergio Cruces, Carlos García Puntonet, Elmar Wolfgang Lang, Deniz Erdogmus |
IEEE Signal Process. Lett. | 6 |
| 2009 | Second-Order Volterra System Identification With Noisy Input-Output MeasurementsabstractSystem identification with noisy input-output measurements has been dominantly addressed through the optimization of the mean-squared-error criterion (MSE), especially in adaptive filtering. MSE is known to provide models that approximate the conditional expectation of the target output given the input; however, when the input signal is also contaminated by noise - a frequent occurrence - MSE yields biased estimates of the model parameters with the severity of the bias dependent on the noise power. This drawback has been addressed in various ways, including errors-in-variables techniques. Recently, error whitening criterion (EWC) and associated adaptation algorithms were proposed to address this issue in linear system identification. We extend the applicability of the main concept behind EWC to the unbiased identification of order-2 Volterra series models of nonlinear dynamical systems. The extension does not apply to higher order Volterra models. The main contribution of this letter is a statistical criterion that can be utilized to identify analytically the true parameters of an order-2 Volterra model from noisy input-output data. We also support the theoretical results with simulations; however online learning algorithms that can be derived for the proposed criterion will not be addressed. Umut Ozertem, Deniz Erdogmus |
IEEE Signal Process. Lett. | 2 |
| 2009 | RKHS Bayes Discriminant: A Subspace Constrained Nonlinear Feature Projection for Signal DetectionabstractGiven the knowledge of class probability densities, a priori probabilities, and relative risk levels, Bayes classifier provides the optimal minimum-risk decision rule. Specifically, focusing on the two-class (detection) scenario, under certain symmetry assumptions, matched filters provide optimal results for the detection problem. Noticing that the Bayes classifier is in fact a nonlinear projection of the feature vector to a single-dimensional statistic, in this paper, we develop a smooth nonlinear projection filter constrained to the estimated span of class conditional distributions as does the Bayes classifier. The nonlinear projection filter is designed in a reproducing kernel Hilbert space leading to an analytical solution both for the filter and the optimal threshold. The proposed approach is tested on typical detection problems, such as neural spike detection or automatic target detection in synthetic aperture radar (SAR) imagery. Results are compared with linear and kernel discriminant analysis, as well as classification algorithms such as support vector machine, AdaBoost and LogitBoost. Umut Ozertem, Deniz Erdogmus |
IEEE Trans. Neural Networks | 2 |
| 2008 | Detecting mild cognitive loss with continuous monitoring of medication adherenceabstractThis paper describes an approach for detecting early cognitive loss using medication adherence behavior. We investigate the discriminative power of a comprehensive set of recurrent medication timing features extracted from time-of-day and inter-dose timing statistics. We adopt information theoretic measures for feature ranking for initial dimensionality reduction and conduct exhaustive leave-one-out cross validation for final feature selection and regularization. The selected feature set is subjected to a support vector machine for classification. The results demonstrate that patterns of adherence based on the data from relatively unobtrusive behavior monitoring can make reliable inference for mild cognitive loss individuals. Yonghong Huang, Deniz Erdogmus, Zhengdong Lu, Todd K. Leen |
ICASSP | 2 |
| 2008 | Large-scale image database triage via EEG evoked responsesabstractThis paper describes an approach for target image search using human brain signals generated by perceptual processes in the brain. The human brain generates event related potentials (ERPs) in response to critical events, such as interesting/novel visual stimuli in the form of a target image. In this paper, we describe experiments involving six professional image analysts and summarize the ERP detection performance as they search for targets within a large image database. We develop a disjoint windowing scheme for data preprocessing to discard irrelevant and redundant information from the raw data to get clean training data. We apply support vector machines to detect ERPs and conduct 10-fold cross validation for parameter regularization. The results demonstrate that the ERP pattern recognition can provide reliable inference for image triage. Yonghong Huang, Deniz Erdogmus, Santosh Mathan, Misha Pavel |
ICASSP | 2 |
| 2008 | Local conditions for critical and principal manifoldsabstractPrincipal manifolds are essential underlying structures that manifest canonical solutions for significant problems such as data de- noising and dimensionality reduction. The traditional definition of self-consistent manifolds rely on a least-squares construction error approach that utilizes semi-global expectations across hyperplanes orthogonal to the solution. This definition creates various practical difficulties for algorithmic solutions to identify such manifolds, besides the theoretical shortcoming that self-intersecting or nonsmooth manifolds are not acceptable in this framework. We present local conditions for critical and principal manifolds by introducing the concept of subspace local maxima. The conditions generalize the two conditions that characterize stationary points of a function to stationary surfaces. The proposed framework yields a unique set of principal points which can be partitioned into principal curves and manifolds of any intrinsic dimensionality. A subspace-constrained fixed-point algorithm is proposed to determine the principal graph. Umut Ozertem, Deniz Erdogmus |
ICASSP | 2 |
| 2008 | Signal denoising using principal curves: Application to timewarpingabstractOne of the most important problems with current time warping algorithms in the literature is sensitivity to noise. To improve the noise robustness of the current algorithms, we propose a denoising step based on the likelihood maximization of the pairwise signals. This approach is independent of the selection of the particular time warping algorithm, and can be coupled with any algorithm in the literature. Improvement in noise robustness not only brings increased robustness to current time warping applications, but also may trigger new application areas where the signals that need to be compared are buried in noise. Umut Ozertem, Deniz Erdogmus |
ICASSP | 2 |
| 2008 | Density geodesics for similarity clusteringabstractWe address the problem of similarity metric selection in pairwise affinity clustering. Traditional techniques employ standard algebraic context-independent sample-distance measures, such as the Euclidean distance. More recent context-dependent metric modifications employ the bottleneck principle to develop path-bottleneck or path- average distances and define similarities based on geodesies determined according to these metrics. This paper develops a principled context-adaptive similarity metric for pairs of feature vectors utilizing the probability density of all data. Specifically, based on the postulate that Euclidean distance is the canonical metric for data drawn from a unit-hypercube uniform density, a density-geodesic distance measure stemming from Riemannian geometry of curved surfaces is derived. Comparisons with alternative metrics demonstrate the superior properties such as robustness. Umut Ozertem, Deniz Erdogmus, Miguel Á. Carreira-Perpiñán |
ICASSP | 2 |
| 2008 | A reproducing kernel Hilbert space framework for pairwise time series distancesabstractA good distance measure for time series needs to properly incorporate the temporal structure, and should be applicable to sequences with unequal lengths. In this paper, we propose a distance measure as a principled solution to the two requirements. Unlike the conventional feature vector representation, our approach represents each time series with a summarizing smooth curve in a reproducing kernel Hilbert space (RKHS), and therefore translate the distance between time series into distances between curves. Moreover we propose to learn the kernel of this RKHS from a population of time series with discrete observations using Gaussian process-based non-parametric mixed-effect models. Experiments on two vastly different real-world problems show that the proposed distance measure leads to improved classification accuracy over the conventional distance measures. Zhengdong Lu, Todd K. Leen, Yonghong Huang, Deniz Erdogmus |
ICML | 4 |
| 2008 | Geometric structure of sum-of-rank-1 decompositions for n-dimensional order-p symmetric tensorsabstractThe canonical sum-of-rank-one decomposition of tensors is a fundamental linear algebraic problem encountered in signal processing, machine learning, and other scientific fields. Current algorithms that emerge from CANDECOMP or PARAFAC formalisms rely on the basic definition of tensor decomposition that describes rank as the minimum number of vectors that are needed to reconstruct the tensor using outer product linear combinations, which is an extension of the same property of matrix rank. In this paper, we reinterpret the orthogonality condition of symmetric matrix eigenvectors as a geometric constraint on the coordinate frame formed by the eigenvectors and relaxing the orthogonality, we develop a set of structured-bases that can be utilized to decompose any symmetric tensor into its sum-of-rank-one (canonical) decomposition. The eigenvectors of order-p tensors are observed to form a frame where the angle between various pairs of eigenvectors are integer multiples of pi/p. Validation of the proposed geometric structure and demonstration of decomposition accuracies obtained using these frames (at the level of a computer's numerical-epsiv) are provided. Olexiy O. Kyrgyzov, Deniz Erdogmus |
ISCAS | 2 |
| 2008 | Advances in blind signal processing
Deniz Erdogmus, Danilo P. Mandic, Toshihisa Tanaka 0001 |
Neurocomputing | 1 |
| 2008 | Mean shift spectral clustering
Umut Ozertem, Deniz Erdogmus, Robert Jenssen |
Pattern Recognit. | 2 |
| 2008 | Recursive complex BSS via generalized eigendecomposition and application in image rejection for BPSK
Puskal P. Pokharel, Umut Ozertem, Deniz Erdogmus, José C. Príncipe |
Signal Process. | 3 |
| 2008 | Continuously Differentiable Sample-Spacing Entropy EstimationabstractThe insufficiency of using only second-order statistics and premise of exploiting higher order statistics of the data has been well understood, and more advanced objectives including higher order statistics, especially those stemming from information theory, such as error entropy minimization, are now being studied and applied in many contexts of machine learning and signal processing. In the adaptive system training context, the main drawback of utilizing output error entropy as compared to correlation-estimation-based second-order statistics is the computational load of the entropy estimation, which is usually obtained via a plug-in kernel estimator. Sample-spacing estimates offer computationally inexpensive entropy estimators; however, resulting estimates are not differentiable, hence, not suitable for gradient-based adaptation. In this brief paper, we propose a nonparametric entropy estimator that captures the desirable properties of both approaches. The resulting estimator yields continuously differentiable estimates with a computational complexity at the order of those of the sample-spacing techniques. The proposed estimator is compared with the kernel density estimation (KDE)-based entropy estimator in the supervised neural network training framework with computation time and performance comparisons. Umut Ozertem, Ismail Uysal, Deniz Erdogmus |
IEEE Trans. Neural Networks | 3 |
| 2007 | Self-Consistent Locally Defined Principal SurfacesabstractPrincipal curves and surfaces play an important role in dimensionality reduction applications of machine learning and signal processing. Vaguely defined, principal curves are smooth curves that pass through the middle of the data distribution. This intuitive definition is ill posed and to this day researchers have struggled with its practical implications. Two main causes of these difficulties are: (i) the desire to build a self-consistent definition using global statistics (for instance conditional expectations), and (ii) not decoupling the definition of the principal curve from the data samples. In this paper, we introduce the concept of principal sets, which are the union of all principal surfaces with a particular dimensionality. The proposed definition of principal surfaces provides rigorous conditions for a point to satisfy that can be evaluated using only the gradient and Hessian of the probability density at the point of interest. Since the definition is decoupled from the data samples, any density estimator could be employed to obtain a probability distribution expression and identify the principal surfaces of the data under this particular model. Deniz Erdogmus, Umut Ozertem |
ICASSP (2) | 1 |
| 2007 | Information Regularized Maximum Likelihood for Binary Motion SensorsabstractWe propose a pairwise mutual information based regularization technique for maximum likelihood sensor fusion in dense distributed sensor networks. The principle is demonstrated in target localization and tracking using a dense binary motion sensor network under a centralized data fusion framework. Simulations demonstrate that the information regularization enables the maximum likelihood localization procedure to provide significantly more accurate target position estimates compared to its unregularized counterpart, which is the current benchmark. The extensions of the information regularization principle to various sensor and data fusion problems such as outlier detection, and sensor failure identification are discussed. Umut Ozertem, Deniz Erdogmus |
ICASSP (2) | 2 |
| 2007 | Recursive Complex Blind Source Separation via Eigendecomposition of Cumulant MatricesabstractUnder the assumptions of non-Gaussian, non-stationary, or non-white independent sources, linear blind source separation can be formulated as a generalized eigenvalue decomposition problem. Here we provide an elegant method of doing this online, instead of waiting for a sufficiently large batch of data. This is done through a recursive generalized eigendecomposition algorithm that tracks the optimal solution, which is obtained using all the data observed. The algorithms proposed in this paper follow the well-known recursive least squares (RLS) algorithm in nature. Puskal P. Pokharel, Umut Ozertem, Deniz Erdogmus, José C. Príncipe |
ICASSP (2) | 3 |
| 2007 | Nonlinear Coordinate Unfolding Via Principal Curve Projections with Application to Nonlinear BSS
Deniz Erdogmus, Umut Ozertem |
ICONIP (2) | 1 |
| 2007 | A Novel Switching Scheme Between Adaptive Information AlgorithmsabstractSwitching approaches can improve the performance of adaptive schemes, however a data driven criterion to accomplish the task is unclear. In this paper, we propose a new optimization criterion for switching which is estimated directly from data. We apply the method to the recently introduced MEE and MEE-SAS algorithms. Using this novel switching scheme, we develop a single algorithm which effectively combines the strengths of MEE and MEE-SAS without sacrificing the simplicity and stability properties of MEE. We explain these results analytically, and through simulations. Seungju Han 0001, Sudhir Rao, Deniz Erdogmus, José C. Príncipe |
IJCNN | 3 |
| 2007 | A Nonparametric Approach for Active ContoursabstractActive contours are commonly used in many image segmentation applications. There are different active contour definitions, but all active contour definitions in the literature use parametric forms to determine the shape priors or adjust the weighting of internal and external forces acting on the active contour. However, the evaluation or estimation of the optimal values of these parameters is impossible in a general sense, and the algorithms are run with different parameters until a satisfactory result is obtained. To get rid of this exhaustive parameter search, we approach the same problem in a nonparametric way to translate the problem of seeking good values of these unknown parameters into seeking for a good density estimate. We tested the proposed method and compared with earlier approaches and obtained better results. Umut Ozertem, Deniz Erdogmus |
IJCNN | 2 |
| 2007 | Automatic Brain Image Segmentation for Evaluation of Experimental Ischemic Stroke Using Gradient vector flow and kernel annealingabstractIschemic stroke is the most prevalent catastrophic disease of the brain. Various animal models have been used to study the disease. The majority of the models are based on induction of focal ischemic cerebral necrosis, followed by exhaustive morphometric analysis of the tissues. Despite recent advances in machine learning and image processing, neurological damage evaluations are still based on tedious manual or semi-automatic segmentation of brain images. We demonstrate a method that uses active contours combined with a kernel annealing approach to automatically segment the brain organs of interest, as well as a simple feature that highlights the contrast between normal and infarct brain tissue for automated analysis. The automated segmentation and analysis solution will be useful for increasing the productivity of experimentation and removing investigator bias from the data analysis. Umut Ozertem, Andras Gruber, Deniz Erdogmus |
IJCNN | 3 |
| 2007 | Quasi-sliding mode control strategy based on multiple-linear models
Jeongho Cho, José C. Príncipe, Deniz Erdogmus, Mark A. Motter |
Neurocomputing | 3 |
| 2007 | Information cut for clustering using a gradient descent approach
Robert Jenssen, Deniz Erdogmus, Kenneth E. Hild II, José C. Príncipe, Torbjørn Eltoft |
Pattern Recognit. | 2 |
| 2007 | Independent Component Analysis and Blind Source Separation
Allan Kardec Barros, José C. Príncipe, Deniz Erdogmus |
Signal Process. | 3 |
| 2007 | A minimum-error entropy criterion with self-adjusting step-size (MEE-SAS)
Seungju Han 0001, Sudhir Rao, Deniz Erdogmus, Kyu-Hwa Jeong, José C. Príncipe |
Signal Process. | 3 |
| 2007 | Nonparametric SnakesabstractActive contours, or so-called snakes, require some parameters to determine the form of the external force or to adjust the tradeoff between the internal forces and the external forces acting on the active contour. However, the optimal values of these parameters cannot be easily identified in a general sense. The usual way to find these required parameters is to run the algorithm several times for a different set of parameters, until a satisfactory performance is obtained. Our nonparametric formulation translates the problem of seeking these unknown parameters into the problem of seeking a good edge probability density estimate. Density estimation is a well-researched field, and our nonparametric formulation allows using well-known concepts of density estimation to get rid of the exhaustive parameter search. Indeed, with the use of kernel density estimation these parameters can be defined locally, whereas, in the original snake approach, all the shape parameters are defined globally. We tested the proposed method on synthetic and real images and obtained comparatively better results. Umut Ozertem, Deniz Erdogmus |
IEEE Trans. Image Process. | 2 |
| 2007 | Adaptive motion estimation schemes using maximum mutual information criterionabstractAbstract We consider the motion estimation problem in video coding. In our previous work, we proposed a new motion estimation method where motion estimation is formulated as an optimization problem and an adaptive system under the minimum error entropy (MEE) criterion is used for motion estimation. In this paper, we develop an adaptive system under the criterion of maximum mutual information to address the motion estimation problem. Our proposed motion estimation algorithms have very low encoding complexity and hence are ideally suited for wireless video sensor networks where limited bandwidth, restricted computational capability, and limited battery power supply impose stringent constraints on the video encoding system. Copyright © 2007 John Wiley & Sons, Ltd. Dapeng Oliver Wu, Deniz Erdogmus, Yuguang Fang, Zhihai He |
Wirel. Commun. Mob. Comput. | 3 |
| 2006 | Kernel Density Estimation, Affinity-Based Clustering, And Typical CutsabstractThe typical cut is a clustering method that is based on the probability pnmthat points xnand xmare in the same cluster over all possible partitions (under the Boltzmann distribution for the mincut cost function). We present two contributions regarding this algorithm. (1) We show that, given a kernel density estimate of the data, minimising the overlap between cluster densities is equivalent to the mincut criterion. This gives a principled way to determine what affinities and scales to use in the typical-cut algorithm, and more generally in clustering and dimensionality reduction algorithms based on pair-wise affinities. (2) We introduce an iterated version of the typical-cut algorithm, where the estimated pnmare used to refine the affinities. We show this procedure is equivalent to finding stationary points of a certain objective function over clusterings; and that at the stationary points the value of pnmis 1 if n and m are in the same cluster and a small value otherwise. Thus, the iterated typical-cut algorithm sharpens the pnmmatrix and makes the cluster structure more obvious. Deniz Erdogmus, Miguel Á. Carreira-Perpiñán, Umut Ozertem |
ICASSP (5) | 1 |
| 2006 | Mean Shift Spectral Clustering for Perceptual Image SegmentationabstractSegmentation is a fundamental problem in image processing having a wide range of applications. Image segmentation algorithms in the literature range from a cost criterion based optimization techniques to various heuristic methods. In this paper, we propose utilizing mean shift spectral clustering for perceptually better image segmentation results Umut Ozertem, Deniz Erdogmus, Tian Lan 0007 |
ICASSP (2) | 2 |
| 2006 | A Closed Form Solution for a Nonlinear Wiener FilterabstractIn this paper a nonlinear extension to the Wiener filter is presented. A direct approach has been devised of replacing the autocorrelation function with a novel function called correntropy, derived from ideas on kernel-based learning theory and information theoretic learning. The linear Wiener filter, widely used because of its simplicity and optimality for linear systems and Gaussian distribution, is no longer effective when dealing with nonlinear time series data. The proposed method incorporates higher order moments in the general form of autocorrelation and improves upon the linear filter. Moreover, the computation cost is still lower than some kernel based methods and has a closed form solution to the problem unlike neural network based methods. Puskal P. Pokharel, Jianwu Xu, Deniz Erdogmus, José C. Príncipe |
ICASSP (3) | 3 |
| 2006 | Comparison of Linear and Nonlinear Approaches on Single Trial ERP Detection in Rapid Serial Visual Presentation TasksabstractIn this paper, we describe a system for detecting encephalography (EEG) signatures of visual recognition events evoked in a single trial during rapid serial visual presentation (RSVP). In order to investigate the viability of nonlinear approaches in EEG detection and assess the performance comparison, we applied three classifiers (linear logistic regression model, Laplacian classifier, and spectral maximum mutual information projection) in the detection tasks. The EEG was recorded using 32 electrodes during the rapid image presentation (50ms/100ms per image). Subjects were instructed to push a button when they recognize a target image. The results suggest that while the detection of single trial EEG-based recognition is possible, taking advantage of the nonlinear techniques requires data representation that would overcome the non-stationarity of the EEG signals. Yonghong Huang, Deniz Erdogmus, Santosh Mathan, Misha Pavel |
IJCNN | 2 |
| 2006 | Information Theoretic Angle-Based Spectral Clustering: A Theoretical Analysis and an AlgorithmabstractRecent work has revealed a close connection between certain information theoretic divergence measures and properties of Mercer kernel feature spaces. Specifically, it has been proposed that an information theoretic measure may be used as a cost function for clustering in a kernel space, approximated by the spectral properties of the Laplacian matrix. In this paper we extend this result to other kernel matrices. We develop an algorithm for the actual clustering which is based on comparing angles between data points, and demonstrate that the proposed method performs equally good as a state-of-the art spectral clustering method. We point out some drawbacks of spectral clustering related to outliers, and suggest measures to be taken. Robert Jenssen, Deniz Erdogmus, José C. Príncipe |
IJCNN | 2 |
| 2006 | Estimating Mutual Information Using Gaussian Mixture Model for Feature Ranking and SelectionabstractFeature selection is a critical step for pattern recognition and many other applications. Typically, feature selection strategies can be categorized into wrapper and filter approaches. Filter approach has attracted much attention because of its flexibility and computational efficiency. Previously, we have developed an ICA-MI framework for feature selection, in which the Mutual Information (MI) between features and class labels was used as the criterion. However, since this method depends on the linearity assumption, it is not applicable for an arbitrary distribution. In this paper, exploiting the fact that Gaussian Mixture Model (GMM) is generally a suitable tool for estimating probability densities, we propose GMM-MI method for feature ranking and selection. We will discuss the details of GMM-MI algorithm and demonstrate the experimental results. We will also compare the GMM-MI method with the ICA-MI method in terms of performance and computational efficiency. Tian Lan 0007, Deniz Erdogmus, Umut Ozertem, Yonghong Huang |
IJCNN | 2 |
| 2006 | Automatic Frequency Bands Segmentation Using Statistical Similarity for Power Spectrum Density Based Brain Computer InterfacesabstractPower spectrum density (PSD) of electroencephalogram (EEG) signals is a widely used feature for brain computer interfaces (BCI). Usually, PSD features are integrated over different frequency bands, such as delta, theta, alpha, beta, gamma, which are based on well-established interpretations of EEG signals in prior experimental and clinical contexts. However, these predefined frequency bands do not necessarily relate to the optimal features for various BCI applications. In this paper, we propose an alternative feature dimensionality reduction method, which automatically determines the optimal number and the range of frequency bands. We applied the proposed method on EEG classification in the context of Augmented Cognition (AugCog) using BCI. The experimental results show that the proposed method can extract more robust features than features manually extracted from predefined frequency bands. Tian Lan 0007, Deniz Erdogmus, Misha Pavel, Santosh Mathan |
IJCNN | 2 |
| 2006 | Clustering with Normalized Information Potential Constrained Maximum Entropy Boltzmann DistributionabstractFrom a probabilistic perspective, the question of clustering is 'what is the probability that two data samples belong to the same cluster?'' Accepting the natural preclustering of samples into corresponding modes of the data probability distribution and answering the question posed above for these modes can reduce the problem complexity. Under the maximum entropy principle, a Boltzmann distribution model can be employed to evaluate the required mode-connectivity probabilities. An algorithm is developed using kernel density estimation in this framework. Its performance is demonstrated on benchmark datasets. Umut Ozertem, Deniz Erdogmus |
IJCNN | 2 |
| 2006 | Kernel Maximum Entropy Data Transformation and an Enhanced Spectral Clustering AlgorithmabstractWe propose a new kernel-based data transformation technique. It is founded on the principle of maximum entropy (MaxEnt) preservation, hence named kernel MaxEnt. The key measure is Renyi's entropy estimated via Parzen windowing. We show that kernel MaxEnt is based on eigenvectors, and is in that sense similar to kernel PCA, but may produce strikingly different transformed data sets. An enhanced spectral clustering algorithm is proposed, by replacing kernel PCA by kernel MaxEnt as an intermediate step. This has a major impact on performance. Robert Jenssen, Torbjørn Eltoft, Mark A. Girolami, Deniz Erdogmus |
NIPS | 4 |
| 2006 | Feature Extraction Using Information-Theoretic LearningabstractA classification system typically consists of both a feature extractor (preprocessor) and a classifier. These two components can be trained either independently or simultaneously. The former option has an implementation advantage since the extractor need only be trained once for use with any classifier, whereas the latter has an advantage since it can be used to minimize classification error directly. Certain criteria, such as Minimum Classification Error, are better suited for simultaneous training, whereas other criteria, such as Mutual Information, are amenable for training the feature extractor either independently or simultaneously. Herein, an information-theoretic criterion is introduced and is evaluated for training the extractor independently of the classifier. The proposed method uses nonparametric estimation of Renyi's entropy to train the extractor by maximizing an approximation of the mutual information between the class labels and the output of the feature extractor. The evaluations show that the proposed method, even though it uses independent training, performs at least as well as three feature extraction methods that train the extractor and classifier simultaneously. Kenneth E. Hild II, Deniz Erdogmus, Kari Torkkola, José C. Príncipe |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Spectral feature projections that maximize Shannon mutual information with class labels
Umut Ozertem, Deniz Erdogmus, Robert Jenssen |
Pattern Recognit. | 2 |
| 2006 | An analysis of entropy estimators for blind source separation
Kenneth E. Hild II, Deniz Erdogmus, José C. Príncipe |
Signal Process. | 2 |
| 2006 | Modeling and inverse controller design for an unmanned aerial vehicle based on the self-organizing mapabstractThe next generation of aircraft will have dynamics that vary considerably over the operating regime. A single controller will have difficulty to meet the design specifications. In this paper, a self-organizing map (SOM)-based local linear modeling scheme of an unmanned aerial vehicle (UAV) is developed to design a set of inverse controllers. The SOM selects the operating regime depending only on the embedded output space information and avoids normalization of the input data. Each local linear model is associated with a linear controller, which is easy to design. Switching of the controllers is done synchronously with the active local linear model that tracks the different operating conditions. The proposed multiple modeling and control strategy has been successfully tested in a simulator that models the LoFLYTE UAV. Jeongho Cho, José C. Príncipe, Deniz Erdogmus, Mark A. Motter |
IEEE Trans. Neural Networks | 3 |
| 2005 | Supervised training of adaptive systems with partially labeled dataabstractSupervised adaptive system training is traditionally performed with available pairs of input-output data and the system weights are fixed following this training procedure. Recently, in the context of machine learning, where the desired outputs are discrete-valued, the idea of exploiting unlabeled samples for improving classification performance has been proposed. We introduce an information theoretic framework based on density divergence minimization to obtain extended training algorithms. Our goal is to provide a theoretical framework upon which we can build efficient algorithms to this end. Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe |
ICASSP (5) | 1 |
| 2005 | MRI Image Reconstruction via Homomorphic Signal ProcessingabstractWe consider the use of homomorphic signal processing for image reconstruction in phased-array magnetic resonance imaging (MRI). Based on the prior information provided from the spectral analysis of the estimated true pixels and the coil sensitivities, homomorphic signal processing is used to partly filter out the coil sensitivities based on the criterion of image contrast in the reconstructed image in specified pixel locations. The image quality, quantized as the entropy of the pixel value distribution, demonstrates a better image contrast compared with the widely used sum-of-squares (SoS) method. Jeffrey R. Fitzsimmons, Deniz Erdogmus |
ICASSP (2) | 3 |
| 2005 | The Laplacian spectral classifierabstractWe develop a novel classifier in a kernel feature space defined by the eigenspectrum of the Laplacian data matrix. The classification cost function is derived from a distance measure between probability densities. The Laplacian data matrix is obtained based on a training set, while test data is mapped to the kernel space using the Nystrom routine. In that space, the test data is classified based on the angle between the test point and the training data class means. We illustrate the performance of the new classifier on synthetic and real data. Robert Jenssen, Deniz Erdogmus, José C. Príncipe, Torbjørn Eltoft |
ICASSP (5) | 2 |
| 2005 | Learning mappings in brain machine interfaces with echo state networksabstractBrain machine interfaces (BMI) utilize linear or non-linear models to map the neural activity to the associated behavior which is typically the 2D or 3D hand position of a primate. Linear models are plagued by the massive disparity of the input and output dimensions thereby leading to poor generalization. A solution would be to use non-linear models like the recurrent multi-layer perceptron (RMLP) that provide parsimonious mapping functions with better generalization. However, this results in a drastic increase in the training complexity, which can be critical for practical use of a BMI. This paper bridges the gap between superior performance per trained weight and model learning complexity. Towards this end, we propose to use echo state networks (ESN) to transform the neuronal firing activity into a higher dimensional space and then derive an optimal sparse linear mapping in the transformed space to match the hand position. The sparse mapping is obtained using a weight constrained cost function whose optimal solution is determined using a stochastic gradient algorithm. Yadunandana N. Rao, Sung-Phil Kim, Justin C. Sanchez, Deniz Erdogmus, José C. Príncipe, Jose M. Carmena, Mikhail Lebedev 0001, Miguel A. L. Nicolelis |
ICASSP (5) | 4 |
| 2005 | An information-theoretic perspective to kernel independent components analysisabstractIn this paper, we investigate the intriguing relationship between information-theoretic learning (ITL), based on weighted Parzen window density estimator, and kernel-based learning algorithms. We prove the equivalence between kernel independent component analysis (kernel ICA) and the Cauchy-Schwartz (C-S) independence measure. This link gives a theoretical motivation for the selection of the Mercer kernel, based on density estimation. Demonstrating this equivalence requires introducing a weighted kernel density estimator, a modification of Parzen windowing. We also discuss the role of the weights in the weighted Parzen windowing and kernel ICA. Jianwu Xu, Deniz Erdogmus, Robert Jenssen, José C. Príncipe |
ICASSP (5) | 2 |
| 2005 | Separating Spatial and Temporal Activation Patterns in fMRI Using Competitive Subspace ProjectionabstractThe challenge for functional magnetic resonance imaging (fMRI) is to determine when and where the response due to an external stimulus occurs in the images. The temporal clustering analysis (TCA) method has been used to study brain activity after eating and drinking, in both the time and spatial domains. We propose a new method, competitive subspace projection (CSP), to represent data optimally compared to other subspace projection methods. This method is used to detect spatial and temporal activation patterns in fMRI associated with such behavior. The CSP and TCA methods are compared using both synthetic and real fMRI data. The results on both data sets show consistent conclusions can be drawn from these two methods while CSP is observed to have a better noise rejection capability than TCA. Guojun He, Deniz Erdogmus, Sung-Phil Kim, José C. Príncipe |
ICASSP (2) | 3 |
| 2005 | Feature selection by independent component analysis and mutual information maximization in EEG signal classificationabstractFeature selection and dimensionality reduction are important steps in pattern recognition. In this paper, we propose a scheme for feature selection using linear independent component analysis and mutual information maximization method. The method is theoretically motivated by the fact that the classification error rate is related to the mutual information between the feature vectors and the class labels. The feasibility of the principle is illustrated on a synthetic dataset and its performance is demonstrated using EEG signal classification. Experimental results show that this method works well for feature selection. Tian Lan 0007, Deniz Erdogmus, André Adami, Michael Pavel |
IJCNN | 2 |
| 2005 | Maximally discriminative spectral feature projections using mutual informationabstractDetermining the optimal subspace projections, which maintains the best representation of the original data, is an important problem in machine learning and pattern recognition. In this paper, we propose a nonparametric nonlinear subspace projection technique that employs kernel density estimation based information theoretic methods and kernel machines, in order to maintain class separability maximally under the Shannon mutual information criterion. Umut Ozertem, Deniz Erdogmus |
IJCNN | 2 |
| 2005 | Fast error whitening algorithms for system identification and control with noisy data
Yadunandana N. Rao, Deniz Erdogmus, Geetha Y. Rao, José C. Príncipe |
Neurocomputing | 2 |
| 2005 | Vector quantization using information theoretic concepts
Tue Lehn-Schiøler, Anant Hegde, Deniz Erdogmus, José C. Príncipe |
Nat. Comput. | 3 |
| 2005 | A new classifier based on information theoretic learning with unlabeled data
Kyu-Hwa Jeong, Jianwu Xu, Deniz Erdogmus, José C. Príncipe |
Neural Networks | 3 |
| 2005 | A mutual information extension to the matched filter
Deniz Erdogmus, Rati Agrawal, José C. Príncipe |
Signal Process. | 1 |
| 2005 | Information Theoretic Signal Processing
Deniz Erdogmus, José C. Príncipe |
Signal Process. | 1 |
| 2005 | Quantifying spatio-temporal dependencies in epileptic ECOG
Anant Hegde, Deniz Erdogmus, Deng-Shan Shiau, José C. Príncipe, J. Chris Sackellares |
Signal Process. | 2 |
| 2005 | Linear-least-squares initialization of multilayer perceptrons through backpropagation of the desired responseabstractTraining multilayer neural networks is typically carried out using descent techniques such as the gradient-based backpropagation (BP) of error or the quasi-Newton approaches including the Levenberg-Marquardt algorithm. This is basically due to the fact that there are no analytical methods to find the optimal weights, so iterative local or global optimization techniques are necessary. The success of iterative optimization procedures is strictly dependent on the initial conditions, therefore, in this paper, we devise a principled novel method of backpropagating the desired response through the layers of a multilayer perceptron (MLP), which enables us to accurately initialize these neural networks in the minimum mean-square-error sense, using the analytic linear least squares solution. The generated solution can be used as an initial condition to standard iterative optimization algorithms. However, simulations demonstrate that in most cases, the performance achieved through the proposed initialization scheme leaves little room for further improvement in the mean-square-error (MSE) over the training set. In addition, the performance of the network optimized with the proposed approach also generalizes well to testing data. A rigorous derivation of the initialization algorithm is presented and its high performance is verified with a number of benchmark training problems including chaotic time-series prediction, classification, and nonlinear system identification with MLPs. Deniz Erdogmus, Oscar Fontenla-Romero, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo |
IEEE Trans. Neural Networks | 1 |
| 2004 | Mixture of competitive linear models for phased-array magnetic resonance imagingabstractPhased-array magnetic resonance imaging is an important contemporary research field in terms of the expected clinical gains in medical imaging technology. Recent research focused on heuristic coil image recombination methods as well as statistical signal processing approaches. In this paper, we investigate the performance of an adaptive signal processing approach, namely mixture of competitively trained models. The proposed method has the ability to train on a set of images and generalize its performance to previously unseen images. Performance evaluations on real data validate the effectiveness of this method. Deniz Erdogmus, Erik G. Larsson, José C. Príncipe, Jeffrey R. Fitzsimmons |
ICASSP (5) | 1 |
| 2004 | Parzen particle filtersabstractUsing a Parzen density estimator, any distribution can be approximated arbitrarily close by a sum of kernels. In particle filtering, this fact is utilized to estimate a probability density function with Dirac delta kernels; when the distribution is discretized it becomes possible to solve an otherwise intractable integral. In this work, we propose to extend the idea and use any kernel to approximate the distribution. The extra work involved in propagating small kernels through the nonlinear function can be made up for by decreasing the number of kernels needed, especially for high dimensional problems. A further advantage of using kernels with nonzero width is that the density estimate becomes continuous. Tue Lehn-Schiøler, Deniz Erdogmus, José C. Príncipe |
ICASSP (5) | 2 |
| 2004 | Accurate linear parameter estimation in colored noiseabstractEstimation of the parameters of an unknown system is an important problem in signal processing. The classical mean squared error (MSE) criterion and its variants have been widely used to solve this problem. However, it is well known that the MSE criterion produces biased parameter estimates when the signals of interest (especially the input) are corrupted with additive noise having arbitrary or no coloring (white). Alternative approaches require additional system constraints and explicit estimation of the noise covariances. Recently, we proposed a new criterion called the error whitening criterion (EWC) along with associated algorithms that solved the problem when the additive disturbances are white. However, the performance of EWC is not satisfactory when the disturbances are correlated (colored). In this paper, we propose a method based on the principles of the EWC that can consistently estimate the parameters of an unknown arbitrary linear system in colored input noise without estimating the noise covariances. We then present a novel stochastic gradient algorithm that estimates the optimal parameters in an on-line fashion. We briefly discuss the convergence of this algorithm and present extensive simulation results to show the superiority of this criterion over MSE. Yadunandana N. Rao, Deniz Erdogmus, José C. Príncipe |
ICASSP (2) | 2 |
| 2004 | Minimizing Fisher information of the error in supervised adaptive filter trainingabstractIn this paper, we propose minimizing the Fisher information of the error in supervised training of linear and nonlinear adaptive filters. Fisher information considers the local structure of the error probability distribution and therefore it is a criterion that deserves to be investigated as an alternative to more common statistics such as minimum mean-square-error or minimum-error-entropy. A gradient-based training algorithm, based on a nonparametric estimator of Fisher information is presented and the performances of the three mentioned optimization criteria are compared using Monte Carlo simulations. Jianwu Xu, Deniz Erdogmus, José C. Príncipe |
ICASSP (5) | 2 |
| 2004 | Nonlinear independent component analysis by homomorphic transformation of the mixturesabstractIndependent component analysis is often approached from an information theoretic perspective employing specific sample estimates for the mutual information between the separated outputs. These approximations involve the nonparametric estimation of signal entropies. The common approach involves the estimation of these quantities and adaptation based on these criteria. In contrast, in this paper, we propose a Gaussianization-based approach, where the separation is performed in two stages: Gaussianization of the mixtures using a homomorphic nonlinearity and separation of the independent components using principal component analysis (both stages possibly adaptive). Due to the rotation uncertainty in nonlinear ICA, the original sources cannot be recovered solely by the independence assumption. The proposed ICA methodology is applicable to instantaneous linear and nonlinear mixtures. The idea also generalizes easily to complex-valued nonlinear ICA. Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe |
IJCNN | 1 |
| 2004 | Vector-quantization by density matching in the minimum Kullback-Leibler divergence senseabstractRepresentation of a large set of high-dimensional data is a fundamental problem in many applications such as communications and biomedical systems. The problem has been tackled by encoding the data with a compact set of code-vectors called processing elements. In this study, we propose a vector quantization technique that encodes the information in the data using concepts derived from information theoretic learning. The algorithm minimizes a cost function based on the Kullback-Liebler divergence to match the distribution of the processing elements with the distribution of the data. The performance of this algorithm is demonstrated on synthetic data as well as on an edge-image of a face. Comparisons are provided with some of the existing algorithms such as LBG and SOM. Anant Hegde, Deniz Erdogmus, Tue Lehn-Schiøler, Yadunandana N. Rao, José C. Príncipe |
IJCNN | 2 |
| 2004 | The Laplacian PDF Distance: A Cost Function for Clustering in a Kernel Feature SpaceabstractA new distance measure between probability density functions (pdfs) is introduced, which we refer to as the Laplacian pdf dis- tance. The Laplacian pdf distance exhibits a remarkable connec- tion to Mercer kernel based learning theory via the Parzen window technique for density estimation. In a kernel feature space defined by the eigenspectrum of the Laplacian data matrix, this pdf dis- tance is shown to measure the cosine of the angle between cluster mean vectors. The Laplacian data matrix, and hence its eigenspec- trum, can be obtained automatically based on the data at hand, by optimal Parzen window selection. We show that the Laplacian pdf distance has an interesting interpretation as a risk function connected to the probability of error. 1 Introduction In recent years, spectral clustering methods, i.e. data partitioning based on the eigenspectrum of kernel matrices, have received a lot of attention [1, 2]. Some unresolved questions associated with these methods are for example that it is not always clear which cost function that is being optimized and that is not clear how to construct a proper kernel matrix. In this paper, we introduce a well-defined cost function for spectral clustering. This cost function is derived from a new information theoretic distance measure between cluster pdfs, named the Laplacian pdf distance. The information theoretic/spectral duality is established via the Parzen window methodology for density estimation. The resulting spectral clustering cost function measures the cosine of the angle between cluster mean vectors in a Mercer kernel feature space, where the feature space is determined by the eigenspectrum of the Laplacian matrix. A principled approach to spectral clustering would be to optimize this cost function in the feature space by assigning cluster memberships. Because of space limitations, we leave it to a future paper to present an actual clustering algorithm optimizing this cost function, and focus in this paper on the theoretical properties of the new measure. Corresponding author. Phone: (+47) 776 46493. Email: [email protected] An important by-product of the theory presented is that a method for learning the Mercer kernel matrix via optimal Parzen windowing is provided. This means that the Laplacian matrix, its eigenspectrum and hence the feature space mapping can be determined automatically. We illustrate this property by an example. We also show that the Laplacian pdf distance has an interesting relationship to the probability of error. In section 2, we briefly review kernel feature space theory. In section 3, we utilize the Parzen window technique for function approximation, in order to introduce the new Laplacian pdf distance and discuss some properties in sections 4 and 5. Section 6 concludes the paper. 2 Kernel Feature Spaces Mercer kernel-based learning algorithms [3] make use of the following idea: via a nonlinear mapping : Rd F, x (x) (1) the data x1, . . . , xN Rd is mapped into a potentially much higher dimensional feature space F. For a given learning problem one now considers the same algorithm in F instead of in Rd, that is, one works with (x1),...,(xN) F. Consider a symmetric kernel function k(x, y). If k : C C R is a continuous kernel of a positive integral operator in a Hilbert space L2(C) on a compact set C Rd, i.e. L2(C) : k(x,y)(x)(y)dxdy 0, (2) C then there exists a space F and a mapping : Rd F, such that by Mercer's theorem [4] NF k(x, y) = (x), (y) = ii(x)i(y), (3) i=1 where , denotes an inner product, the i's are the orthonormal eigenfunctions of the kernel and NF [3]. In this case (x) = [ 11(x), 22(x), . . . ]T , (4) can potentially be realized. In some cases, it may be desirable to realize this mapping. This issue has been addressed in [5]. Define the (N N) Gram matrix, K, also called the affinity, or kernel matrix, with elements Kij = k(xi, xj), i, j = 1, . . . , N . This matrix can be diagonalized as ET KE = , where the columns of E contains the eigenvectors of K and is a diagonal matrix containing the non-negative eigenvalues ~ 1, . . . , ~ N , ~ 1 ~N. In [5], it was shown that the eigenfunctions and eigenvalues of (4) can ~ be approximated as j j (xi) Neji, j , where e N ji denotes the ith element of the jth eigenvector. Hence, the mapping (4), can be approximated as (xi) [ ~1e1i,..., ~NeNi]T. (5) Thus, the mapping is based on the eigenspectrum of K. The feature space data set may be represented in matrix form as NN = [(x1), . . . , (xN )]. Hence, = 1 2 ET . It may be desirable to truncate the mapping (5) to C-dimensions. Thus, T only the C first rows of are kept, yielding ^ . It is well-known that ^ K = ^ ^ is the best rank-C approximation to K wrt. the Frobenius norm [6]. The most widely used Mercer kernel is the radial-basis-function (RBF) k(x, y) = exp -||x - y||2 . (6) 22 3 Function Approximation using Parzen Windowing Parzen windowing is a kernel-based density estimation method, where the resulting density estimate is continuous and differentiable provided that the selected kernel is continuous and differentiable [7]. Given a set of iid samples {x1,...,xN} drawn from the true density f (x), the Parzen window estimate for this distribution is [7] N ^ 1 f (x) = W N 2 (x, xi), (7) i=1 where W2 is the Parzen window, or kernel, and 2 controls the width of the kernel. The Parzen window must integrate to one, and is typically chosen to be a pdf itself with mean xi, such as the Gaussian kernel 1 W2 (x, xi) = exp , (8) d -||x - xi||2 (22) 2 22 which we will assume in the rest of this paper. In the conclusion, we briefly discuss the use of other kernels. Consider a function h(x) = v(x)f (x), for some function v(x). We propose to estimate h(x) by the following generalized Parzen estimator N ^ 1 h(x) = v(xi)W N 2 (x, xi). (9) i=1 This estimator is asymptotically unbiased, which can be shown as follows 1 N Ef v(xi)W N 2 (x, xi) = v(z)f (z)W2 (x, z)dz = [v(x)f (x)] W2(x), i=1 (10) where Ef () denotes expectation with respect to the density f(x). In the limit as N and (N) 0, we have lim [v(x)f (x)] W2(x) = v(x)f(x). (11) N (N )0 Of course, if v(x) = 1 x, then (9) is nothing but the traditional Parzen estimator of h(x) = f (x). The estimator (9) is also asymptotically consistent provided that the kernel width (N ) is annealed at a sufficiently slow rate. The proof will be presented in another paper. Many approaches have been proposed in order to optimally determine the size of the Parzen window, given a finite sample data set. A simple selection rule was proposed by Silverman [8], using the mean integrated square error (MISE) between the estimated and the actual pdf as the optimality metric: 1 d+4 opt = X 4N -1(2d + 1)-1 , (12) where d is the dimensionality of the data and 2 = d-1 , where are the X i Xii Xii diagonal elements of the sample covariance matrix. More advanced approximations to the MISE solution also exist. 4 The Laplacian PDF Distance Cost functions for clustering are often based on distance measures between pdfs. The goal is to assign memberships to the data patterns with respect to a set of clusters, such that the cost function is optimized. Assume that a data set consists of two clusters. Associate the probability density function p(x) with one of the clusters, and the density q(x) with the other cluster. Let f (x) be the overall probability density function of the data set. Now define the f -1 weighted inner product between p(x) and q(x) as p, q f p(x)q(x)f-1(x)dx. In such an inner product space, the Cauchy-Schwarz inequality holds, that is, p, q 2 q, q . Based on this discussion, an information theoretic distance f p, p f f measure between the two pdfs can be expressed as p, q D f L = - log 0. (13) p, p q, q f f We refer to this measure as the Laplacian pdf distance, for reasons that we discuss next. It can be seen that the distance DL is zero if and only if the two densities are equal. It is non-negative, and increases as the overlap between the two pdfs decreases. However, it does not obey the triangle inequality, and is thus not a distance measure in the strict mathematical sense. We will now show that the Laplacian pdf distance is also a cost function for clus- tering in a kernel feature space, using the generalized Parzen estimators discussed in the previous section. Since the logarithm is a monotonic function, we will derive the expression for the argument of the log in (13). This quantity will for simplicity be denoted by the letter "L" in equations. Assume that we have available the iid data points {xi}, i = 1,...,N1, drawn from p(x), which is the density of cluster C1, and the iid {xj}, j = 1, . . ., N2, drawn from q(x), the density of C2. Let h(x) = f - 12 (x)p(x) and g(x) = f - 12 (x)q(x). Hence, we may write h(x)g(x)dx L = . (14) h2(x)dx g2(x)dx We estimate h(x) and g(x) by the generalized Parzen kernel estimators, as follows N1 N2 ^ 1 1 h(x) = f - 12 (xi)W f - 12 (xj )W N 2 (x, xi ), ^ g(x) = 2 (x, xj ). (15) 1 N2 i=1 j=1 The approach taken, is to substitute these estimators into (14), to obtain N N 1 1 1 2 h(x)g(x)dx f - 12 (xi)W f - 12 (xj )W N 2 (x, xi ) 2 (x, xj ) 1 N2 i=1 j=1 N 1 1 ,N2 = f - 12 (xi)f - 12 (xj ) W N 2 (x, xi )W2 (x, xj )dx 1N2 i,j=1 N 1 1 ,N2 = f - 12 (xi)f - 12 (xj )W N 22 (xi, xj ), (16) 1N2 i,j=1 where in the last step, the convolution theorem for Gaussians has been employed. Similarly, we have N 1 1 ,N1 h2(x)dx f - 12 (xi)f - 12 (xi )W N 2 22 (xi, xi ), (17) 1 i,i =1 N 1 2 ,N2 g2(x)dx f - 12 (xj)f - 12 (xj )W N 2 22 (xj , xj ). (18) 2 j,j =1 Now we define the matrix Kf , such that Kf = K ij f (xi, xj ) = f - 1 2 (xi)f - 12 (xj )K(xi, xj ), (19) where K(xi, xj ) = W22 (xi, xj) for i, j = 1, . . . , N and N = N1 + N2. As a consequence, (14) can be re-written as follows N1,N2 Kf (xi, xj) L = i,j=1 (20) N1,N1 K K i,i =1 f (xi, xi ) N2,N2 j,j =1 f (xj , xj ) The key point of this paper, is to note that the matrix K = Kij = K(xi, xj), i, j = 1, . . . , N , is the data affinity matrix, and that K(xi, xj) is a Gaussian RBF kernel function. Hence, it is also a kernel function that satisfies Mercer's theorem. Since K(xi, xj) satisfies Mercer's theorem, the following by definition holds [4]. For any set of examples {x1,...,xN} and any set of real numbers 1,...,N N N ijK(xi, xj) 0, (21) i=1 j=1 in analogy to (3). Moreover, this means that N N N N ijf - 12 (xi)f - 12 (xj )K(xi, xj) = ijKf (xi, xj) 0, (22) i=1 j=1 i=1 j=1 hence Kf (xi, xj ) is also a Mercer kernel. Now, it is readily observed that the Laplacian pdf distance can be analyzed in terms of inner products in a Mercer kernel-based Hilbert feature space, since Kf (xi, xj) = f (xi), f (xj) . Consequently, (20) can be written as follows N1,N2 f (xi), f (xj) L = i,j=1 N1,N1 i,i =1 f (xi), f (xi ) N2,N2 j,j =1 f (xj ), f (xj ) 1 N1 N2 N i=1 f (xi ), 1 N j=1 f (xj ) = 1 2 1 N1 N1 N2 N2 N f (xi), 1 f (xi ) 1 f (xj ), 1 f (xj ) 1 i=1 N1 i =1 N2 j=1 N2 j =1 m1 , m2 = f f = cos (m , m ), (23) ||m 1f 2f 1f ||||m2f || where m Ni i = 1 f N f (xl), i = 1, 2, that is, the sample mean of the ith cluster i l=1 in feature space. This is a very interesting result. We started out with a distance measure between densities in the input space. By utilizing the Parzen window method, this distance measure turned out to have an equivalent expression as a measure of the distance between two clusters of data points in a Mercer kernel feature space. In the feature space, the distance that is measured is the cosine of the angle between the cluster mean vectors. The actual mapping of a data point to the kernel feature space is given by the eigendecomposition of Kf , via (5). Let us examine this mapping in more detail. 1 Note that f 2 (xi) can be estimated from the data by the traditional Parzen pdf estimator as follows N 1 1 f 2 (xi) = W (xi, xl) = di. (24) N 2 f l=1 Define the matrix D = diag(d1, . . . , dN ). Then Kf can be expressed as Kf = D- 12 KD- 12 . (25) Quite interestingly, for 2 = 22, this is in fact the Laplacian data matrix. 1 f The above discussion explicitly connects the Parzen kernel and the Mercer kernel. Moreover, automatic procedures exist in the density estimation literature to opti- mally determine the Parzen kernel given a data set. Thus, the Mercer kernel is also determined by the same procedure. Therefore, the mapping by the Laplacian matrix to the kernel feature space can also be determined automatically. We regard this as a significant result in the kernel based learning theory. As an example, consider Fig. 1 (a) which shows a data set consisting of a ring with a dense cluster in the middle. The MISE kernel size is opt = 0.16, and the Parzen pdf estimate is shown in Fig. 1 (b). The data mapping given by the corresponding Laplacian matrix is shown in Fig. 1 (c) (truncated to two dimensions for visualization purposes). It can be seen that the data is distributed along two lines radially from the origin, indicating that clustering based on the angular measure we have derived makes sense. The above analysis can easily be extended to any number of pdfs/clusters. In the C-cluster case, we define the Laplacian pdf distance as C-1 pi, pj L = f . (26) i=1 j=i C pi, pi p f j , pj f In the kernel feature space, (26), corresponds to all cluster mean vectors being pairwise as orthogonal to each other as possible, for all possible unique pairs. 4.1 Connection to the Ng et al. [2] algorithm Recently, Ng et al. [2] proposed to map the input data to a feature space determined by the eigenvectors corresponding to the C largest eigenvalues of the Laplacian ma- trix. In that space, the data was normalized to unit norm and clustered by the C-means algorithm. We have shown that the Laplacian pdf distance provides a 1It is a bit imprecise to refer to Kf as the Laplacian matrix, as readers familiar with spectral graph theory may recognize, since the definition of the Laplacian matrix is L = I - Kf . However, replacing Kf by L does not change the eigenvectors, it only changes the eigenvalues from i to 1 - i. 0 0 (a) Data set (b) Parzen pdf estimate (c) Feature space data Figure 1: The kernel size is automatically determined (MISE), yielding the Parzen estimate (b) with the corresponding feature space mapping (c). clustering cost function, measuring the cosine of the angle between cluster means, in a related kernel feature space, which in our case can be determined automati- cally. A more principled approach to clustering than that taken by Ng et al. is to optimize (23) in the feature space, instead of using C-means. However, because of the normalization of the data in the feature space, C-means can be interpreted as clustering the data based on an angular measure. This may explain some of the success of the Ng et al. algorithm; it achieves more or less the same goal as cluster- ing based on the Laplacian distance would be expected to do. We will investigate this claim in our future work. Note that we in our framework may choose to use only the C largest eigenvalues/eigenvectors in the mapping, as discussed in section 2. Since we incorporate the eigenvalues in the mapping, in contrast to Ng et al., the actual mapping will in general be different in the two cases. 5 The Laplacian PDF distance as a risk function We now give an analysis of the Laplacian pdf distance that may further motivate its use as a clustering cost function. Consider again the two cluster case. The overall data distribution can be expressed as f (x) = P1p(x) + P2q(x), were Pi, i = 1, 2, are the priors. Assume that the two clusters are well separated, such that for xi C1, f (xi) P1p(xi), while for xi C2, f(xi) P2q(xi). Let us examine the numerator of (14) in this case. It can be approximated as p(x)q(x) dx f (x) p(x)q(x) p(x)q(x) 1 1 dx + dx q(x)dx + p(x)dx. (27) C f (x) f (x) P1 P2 1 C2 C1 C2 By performing a similar calculation for the denominator of (14), it can be shown to be approximately equal to 1 . Hence, the Laplacian pdf distance can be written P1P1 as a risk function, given by 1 1 L P1P2 q(x)dx + p(x)dx . (28) P1 C P2 1 C2 Note that if P1 = P2 = 1 , then L = 2P 2 e, where Pe is the probability of error when assigning data points to the two clusters, that is Pe = P1 q(x)dx + P2 p(x)dx. (29) C1 C2 Thus, in this case, minimizing L is equivalent to minimizing Pe. However, in the case that P1 = P2, (28) has an even more interesting interpretation. In that situation, it can be seen that the two integrals in the expressions (28) and (29) are weighted exactly oppositely. For example, if P1 is close to one, L p(x)dx, while P C e 2 q(x)dx. Thus, the Laplacian pdf distance emphasizes to cluster the most un- C1 likely data points correctly. In many real world applications, this property may be crucial. For example, in medical applications, the most important points to classify correctly are often the least probable, such as detecting some rare disease in a group of patients. 6 Conclusions We have introduced a new pdf distance measure that we refer to as the Laplacian pdf distance, and we have shown that it is in fact a clustering cost function in a kernel feature space determined by the eigenspectrum of the Laplacian data matrix. In our exposition, the Mercer kernel and the Parzen kernel is equivalent, making it possible to determine the Mercer kernel based on automatic selection procedures for the Parzen kernel. Hence, the Laplacian data matrix and its eigenspectrum can be determined automatically too. We have shown that the new pdf distance has an interesting property as a risk function. The results we have derived can only be obtained analytically using Gaussian ker- nels. The same results may be obtained using other Mercer kernels, but it requires an additional approximation wrt. the expectation operator. This discussion is left for future work. Acknowledgments. This work was partially supported by NSF grant ECS- 0300340. Robert Jenssen, Deniz Erdogmus, José C. Príncipe, Torbjørn Eltoft |
NIPS | 2 |
| 2004 | Minimax Mutual Information Approach for Independent Component AnalysisabstractMinimum output mutual information is regarded as a natural criterion for independent component analysis (ICA) and is used as the performance measure in many ICA algorithms. Two common approaches in information-theoretic ICA algorithms are minimum mutual information and maximum output entropy approaches. In the former approach, we substitute some form of probability density function (pdf) estimate into the mutual information expression, and in the latter we incorporate the source pdf assumption in the algorithm through the use of nonlinearities matched to the corresponding cumulative density functions (cdf). Alternative solutions to ICA use higher-order cumulant-based optimization criteria, which are related to either one of these approaches through truncated series approximations for densities. In this article, we propose a new ICA algorithm motivated by the maximum entropy principle (for estimating signal distributions). The optimality criterion is the minimum output mutual information, where the estimated pdfs are from the exponential family and are approximate solutions to a constrained entropy maximization problem. This approach yields an upper bound for the actual mutual information of the output signals - hence, the name minimax mutual information ICA algorithm. In addition, we demonstrate that for a specific selection of the constraint functions in the maximum entropy density estimation procedure, the algorithm relates strongly to ICA methods using higher-order cumulants. Deniz Erdogmus, Kenneth E. Hild II, Yadunandana N. Rao, José C. Príncipe |
Neural Comput. | 1 |
| 2004 | Asymptotic SNR-performance of some image combination techniques for phased-array MRI
Deniz Erdogmus, Erik G. Larsson, José C. Príncipe, Jeffrey R. Fitzsimmons |
Signal Process. | 1 |
| 2004 | Measuring the signal-to-noise ratio in magnetic resonance imaging: a caveat
Deniz Erdogmus, Erik G. Larsson, José C. Príncipe, Jeffrey R. Fitzsimmons |
Signal Process. | 1 |
| 2004 | Guest Editorial Special Issue on Information Theoretic Learning
José C. Príncipe, Erkki Oja, Lei Xu 0001, Andrzej Cichocki, Deniz Erdogmus |
IEEE Trans. Neural Networks | 5 |
| 2004 | Feature selection in MLPs and SVMs based on maximum output informationabstractThis paper presents feature selection algorithms for multilayer perceptrons (MLPs) and multiclass support vector machines (SVMs), using mutual information between class labels and classifier outputs, as an objective function. This objective function involves inexpensive computation of information measures only on discrete variables; provides immunity to prior class probabilities; and brackets the probability of error of the classifier. The maximum output information (MOI) algorithms employ this function for feature subset selection by greedy elimination and directed search. The output of the MOI algorithms is a feature subset of user-defined size and an associated trained classifier (MLP/SVM). These algorithms compare favorably with a number of other methods in terms of performance on various artificial and real-world data sets. Vikas Sindhwani, Subrata Rakshit, Dipti Deodhare, Deniz Erdogmus, José C. Príncipe, Partha Niyogi |
IEEE Trans. Neural Networks | 4 |
| 2003 | Recursive Least Squares for an Entropy Regularized MSE Cost Function
Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe, Oscar Fontenla-Romero, Amparo Alonso-Betanzos |
ESANN | 1 |
| 2003 | Accelerating the convergence speed of neural networks learning methods using least squares
Oscar Fontenla-Romero, Deniz Erdogmus, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo |
ESANN | 2 |
| 2003 | Linear Least-Squares Based Methods for Neural Networks Learning
Oscar Fontenla-Romero, Deniz Erdogmus, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo |
ICANN | 2 |
| 2003 | On the convergence of SIPEX: a simultaneous principal components extraction algorithmabstractWe have previously proposed SIPEX as a fast-converging and accurate principal components algorithm (Erdogmus, D. et al., Proc. ICASSP'02, vol.1, p.1069-72, 2002; Proc. EUSIPCO'02, vol.2, p.335-8, 2002). Its superiority in terms of data efficiency and solution accuracy was demonstrated through Monte Carlo simulations. We focus on the convergence properties of the original gradient-based algorithm as well as two modified versions of SIPEX based on approximations to the Hessian matrix of the cost function. We provide practical bounds on the step sizes of these algorithms and compare their convergence properties. Deniz Erdogmus, Yadunandana N. Rao, M. Can Ozturk, Luis Vielva, José C. Príncipe |
ICASSP (2) | 1 |
| 2003 | A hybrid subspace projection method for system identificationabstractPrincipal components analysis (PCA), being the most optimal linear mapper in a least-squares (LS) sense, has been predominantly used in subspace-based signal processing methods. In system identification problems, optimal subspace projections must span the joint space of the input and output of the unknown system. In this scenario, subspaces determined by the principal components of the input or the desired signal alone do not embed key information, which lies in the joint space. We first propose a hybrid subspace projection method that finds optimal projections in the joint space. The concepts behind this method are firmly rooted in statistical theory. We then derive adaptive learning algorithms to estimate the subspace projections. Finally, we show the superiority of the new framework in solving system identification problems in noisy environments. Sung-Phil Kim, Yadunandana N. Rao, Deniz Erdogmus, José C. Príncipe |
ICASSP (6) | 3 |
| 2003 | Echo cancellation by global optimization of Kautz filters using an information theoretic criterionabstractIn practical settings, the echo cancellation problem generally requires the adaptation of an IIR filter using some optimality criterion. This brings two problems: direct adaptation of numerator and denominator polynomial coefficients of IIR filters might result in unstable systems and/or the optimization might result in a suboptimal local minimum of the criterion. These two issues are addressed in this paper. To resolve the first problem, orthogonal Kautz filters are utilized for their stability is easily controlled through the pole locations. The second problem is addressed by employing an information theoretic optimality criterion, which has a parameter that is annealed to ensure global optimization. Ching-An Lai, Deniz Erdogmus, José C. Príncipe |
ICASSP (6) | 2 |
| 2003 | Matched pdf-based blind equalizationabstractIn this paper, a new blind equalization algorithm for multilevel modulations is proposed. It is based on maximizing the correlation between the probability density function (pdf) of the signal at the output of the equalizer and the desired pdf. The algorithm employs the Parzen window method to estimate the pdf of the squared modulus of the equalizer output. A stochastic gradient-based algorithm is used to maximize the correlation between this pdf and the pdf of the corresponding modulation. The proposed algorithm shows an excellent performance when compared with conventional adaptive blind algorithms, such as CMA, in quadrature amplitude modulation (QAM) schemes. Marcelino Lázaro, Ignacio Santamaría, Carlos Pantaleón, Deniz Erdogmus, José C. Príncipe |
ICASSP (4) | 4 |
| 2003 | Image combination for high-field phased-array MRIabstractWe consider signal processing methods for phased-array magnetic resonance imaging (MRI). A theoretical description of a phased-array MRI data model is presented and three image reconstruction algorithms are proposed to estimate the effective image pixels. An analysis is provided to show how the new algorithms compare to conventional reconstruction methods. Deniz Erdogmus, Erik G. Larsson, José C. Príncipe, Jeffrey R. Fitzsimmons |
ICASSP (5) | 2 |
| 2003 | Supervised synaptic weight adaptation for a spiking neuronabstractA novel algorithm named Spike-LMS is described that adapts the synaptic weights of an artificial spiking neuron to produce a desired response. The derivation of Spike-LMS follows from the derivation of the least-mean squares (LMS) algorithm used in adaptive filter theory. Spike-LMS works directly in the domain of spike trains, and therefore makes no assumptions about any particular neural encoding method. This algorithm is able to identify the synaptic weights of a spiking neuron given the pre-synaptic and post-synaptic spike trains. Bryan A. Davis, Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe |
IJCNN | 2 |
| 2003 | Accurate initialization of neural network weights by backpropagation of the desired responseabstractProper initialization of neural networks is critical for a successful training of its weights. Many methods have been proposed to achieve this, including heuristic least squares approaches. In this paper, inspired by these previous attempts to train (or initialize) neural networks, we formulate a mathematically sound algorithm based on backpropagating the desired output through the layers of a multilayer perceptron. The approach is accurate up to local first order approximations of the nonlinearities. It is shown to provide successful weight initialization for many data sets by Monte Carlo experiments. Deniz Erdogmus, Oscar Fontenla-Romero, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo, Robert Jenssen |
IJCNN | 1 |
| 2003 | Clustering using Renyi's entropyabstractWe propose a new clustering algorithm using Renyi's entropy as our similarity metric. The main idea is to assign a data pattern to the cluster, which among all possible clusters, increases its within-cluster entropy the least, upon inclusion of the pattern. We refer to this procedure as differential entropy clustering. Not knowing the true number of clusters in advance, initially a number of small clusters are "seeded" randomly in the data set, labeling a small subset of the data. Thereafter all remaining patterns are labeled by differential entropy clustering. Subsequently, we identify the "worst cluster" by a quantity we name as the between-cluster entropy. Its members are re-clustered, again by differential entropy clustering, reducing the overall number of clusters by one. This procedure is repeated until only two clusters remain. At each step we store the current labels, thus producing a hierarchy of clusters. The between-cluster entropy also enables us to select our final set of clusters in other cluster hierarchy. We demonstrate the clustering algorithm when applied both to artificially created data sets and a real data set. Robert Jenssen, Kenneth E. Hild II, Deniz Erdogmus, José C. Príncipe, Torbjørn Eltoft |
IJCNN | 3 |
| 2003 | Modeling the relation from motor cortical neuronal firing to hand movements using competitive linear filters and a MLPabstractRecent research has demonstrated that linear model are able to estimate hand positions using populations of action potentials collected in the pre-motor and motor cortical areas of a primate's brain. One of the applications of this result is to restore movement in patients suffering from paralysis. To implement this technology in real-time, reliable and accurate signal processing models that produce sufficient small error in the estimated hand positions are required. In this paper, we propose the hybrid model approach that combines competitive linear filters with a neural network. The mapping performance of our approach is compared with a single Wiener filter during reaching movements. Our approach demonstrates more accurate estimations. Sung-Phil Kim, Justin C. Sanchez, Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe, Miguel A. L. Nicolelis |
IJCNN | 3 |
| 2003 | Simulation of the Freeman model of the olfactory cortex: a quantitative performance analysis for the DSP approachabstractIt has been shown that a framework composed of digital signal processing (DSP) elements can be used to simulate and study Freeman's model of the biologically realistic olfactory cortex. In this paper, based on impulsive invariant transformation, a DSP environment has been developed corresponding to the original continuous-time dynamical system. The performance of the DSP system is quantitatively evaluated by comparing with the original system, which is implemented using the traditional Runge-Kutta integration technique. The discrete-time architecture is shown to have high performance in approximating the dynamical behavior of the original system. Mustafa C. Ozturk, José C. Príncipe, Bryan A. Davis, Deniz Erdogmus |
IJCNN | 4 |
| 2003 | Error whitening criterion for linear filter estimationabstractMean square error (MSE) has been the most widely used tool to solve the linear filter estimation or system identification problem. However, MSE gives biased results when the input signals are noisy. This paper presents a novel error whitening criterion (EWC) to tackle the problem of linear system identification in the presence of additive white disturbances. We would motivate the theory behind the new criterion and derive an online stochastic gradient algorithm based on EWC. Convergence proof of the stochastic gradient algorithm is derived making mild assumptions. Simulation results show the effectiveness of this criterion. We compare its performance with MSE as well as the powerful total least squares method. Yadunandana N. Rao, Deniz Erdogmus, José C. Príncipe |
IJCNN | 2 |
| 2003 | Divide-and-conquer approach for brain machine interfaces: nonlinear mixture of competitive linear models
Sung-Phil Kim, Justin C. Sanchez, Deniz Erdogmus, Yadunandana N. Rao, Johan Wessberg, José C. Príncipe, Miguel A. L. Nicolelis |
Neural Networks | 3 |
| 2003 | Stochastic error whitening algorithm for linear filter estimation with noisy data
Yadunandana N. Rao, Deniz Erdogmus, Geetha Y. Rao, José C. Príncipe |
Neural Networks | 2 |
| 2003 | Online entropy manipulation: stochastic information gradientabstractEntropy has found significant applications in numerous signal processing problems including independent components analysis and blind deconvolution. In general, entropy estimators require O(N/sup 2/) operations, N being the number of samples. For practical online entropy manipulation, it is desirable to determine a stochastic gradient for entropy, which has O(N) complexity. In this paper, we propose a stochastic Shannon's entropy estimator. We determine the corresponding stochastic gradient and investigate its performance. The proposed stochastic gradient for Shannon's entropy can be used in online adaptation problems where the optimization of an entropy-based cost function is necessary. Deniz Erdogmus, Kenneth E. Hild II, José C. Príncipe |
IEEE Signal Process. Lett. | 1 |
| 2002 | Potential Energy and Particle Interaction Approach for Learning in Adaptive Systems
Deniz Erdogmus, José C. Príncipe, Luis Vielva, David Luengo |
ICANN | 1 |
| 2002 | Simultaneous extraction of Principal Components using givens rotations and output variancesabstractPrincipal Components Analysis (PCA) is an invaluable statistical tool in signal processing. In many cases, an on-line algorithm to adapt the PCA network to determine the principal projections in the input space is desired. Algorithms proposed until now use the traditional deflation or the inflation procedure to determine the intermediate components sequentially, after the convergence of the principal or minor component is achieved. In this paper, we propose a constrained linear network and a robust cost function to determine any number of principal components simultaneously. The topology exploits the fact that the eigenvector matrix sought is orthonormal. A gradient-based algorithm named SIPEX-G is also presented. Deniz Erdogmus, Yadunandana N. Rao, José C. Príncipe, Kenneth E. Hild II |
ICASSP | 1 |
| 2002 | Blind source separation of time-varying instantaneous mixtures using an on-line algorithmabstractA blind source separation algorithm is presented that performs online separation of an unknown, time-varying, instantaneous mixture of independent sources. This procedure utilizes the information-theoretic MeRMald-SIG criterion and an on-line PCA algorithm, referred to as SIPEX-G, that has been recently submitted for publication. Results indicate that the combination of the two on-line criteria is able to track a rapidly changing mixing matrix. A performance comparison of separation criteria is also made using the same, aforementioned on-line PCA algorithm. Results show the superior performance of the proposed method. Kenneth E. Hild II, Deniz Erdogmus, José C. Príncipe |
ICASSP | 2 |
| 2002 | Underdetermined blind source separation in a time-varying environmentabstractThe problem of estimating n source signals from m measurements that are an unknown mixture of the sources is known as blind source separation. In the underdetermined —less measurements than sources— linear case, the solution process can be conveniently divided in three stages: represent the signals in a sparse domain, find the mixing matrix, and estimate the sources. In this paper we adhere to that approach and parametrize the performance of these stages as a function of the sparsity of the signals. To find the mixing matrix and track its variations in the dynamic case a nonparametric maximum-likelihood approach based on Parzen windowing is presented. To invert the underdetermined linear problem we present an estimator that chooses the “best” demixing matrix in a sample by sample basis by using some previous knowledge of the statistics of the sources. The results are validated by Montecarlo simulations. Luis Vielva, Deniz Erdogmus, Carlos Pantaleón, Ignacio Santamaría, J. A. Pereda, José C. Príncipe |
ICASSP | 2 |
| 2002 | Blind source separation using Renyi's -marginal entropies
Deniz Erdogmus, Kenneth E. Hild II, José C. Príncipe |
Neurocomputing | 1 |
| 2002 | Beyond second-order statistics for learning: A pairwise interaction model for entropy estimation
Deniz Erdogmus, José C. Príncipe, Kenneth E. Hild II |
Nat. Comput. | 1 |
| 2002 | Generalized information potential criterion for adaptive system trainingabstractWe have previously proposed the quadratic Renyi's error entropy as an alternative cost function for supervised adaptive system training. An entropy criterion instructs the minimization of the average information content of the error signal rather than merely trying to minimize its energy. In this paper, we propose a generalization of the error entropy criterion that enables the use of any order of Renyi's entropy and any suitable kernel function in density estimation. It is shown that the proposed entropy estimator preserves the global minimum of actual entropy. The equivalence between global optimization by convolution smoothing and the convolution by the kernel in Parzen windowing is also discussed. Simulation results are presented for time-series prediction and classification where experimental demonstration of all the theoretical concepts is presented. Deniz Erdogmus, José C. Príncipe |
IEEE Trans. Neural Networks | 1 |
| 2001 | Blind source separation using Renyi's mutual informationabstractA blind source separation algorithm is proposed that is based on minimizing Renyi's mutual information by means of nonparametric probability density function (PDF) estimation. The two-stage process consists of spatial whitening and a series of Givens rotations and produces a cost function consisting only of marginal entropies. This formulation avoids the problems of PDF inaccuracy due to truncation of series expansion and the estimation of joint PDFs in high-dimensional spaces given the typical paucity of data. Simulations illustrate the superior efficiency, in terms of data length, of the proposed method compared to fast independent component analysis (FastICA), Comon's (1994) minimum mutual information, and Bell and Sejnowski's (1995) Infomax. Kenneth E. Hild II, Deniz Erdogmus, José C. Príncipe |
IEEE Signal Process. Lett. | 2 |