VLDB 2026 Research / reviewers in the wild / expert
Robert Jenssen
dblp:45/5813
· DBLP profile ↗
86ranked-venue papers
12as first author
41since 2021 · last 2026
0000-0002-7496-8474ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 59 · 11 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SuperCM: Improving semi-supervised learning and domain adaptation through differentiable clusteringabstractSemi-Supervised Learning (SSL) and Unsupervised Domain Adaptation (UDA) enhance the model performance by exploiting information from labeled and unlabeled data. The clustering assumption has proven advantageous for learning with limited supervision and states that data points belonging to the same cluster in a high-dimensional space should be assigned to the same category. Recent works have utilized different training mechanisms to implicitly enforce this assumption for the SSL and UDA. In this work, we take a different approach by explicitly involving a differentiable clustering module which is extended to leverage the supervised data to compute its centroids. We demonstrate the effectiveness of our straightforward end-to-end training strategy for SSL and UDA over extensive experiments and highlight its benefits, especially in low supervision regimes, both as a standalone model and as a regularizer for existing approaches. Durgesh Kumar Singh, Ahcène Boubekki, Robert Jenssen, Michael Kampffmeyer |
Pattern Recognit. | 3 |
| 2025 | REPEAT: Improving Uncertainty Estimation in Representation Learning ExplainabilityabstractIncorporating uncertainty is crucial to provide trustworthy explanations of deep learning models. Recent works have demonstrated how uncertainty modeling can be particularly important in the unsupervised field of representation learning explainable artificial intelligence (R-XAI). Current R-XAI methods provide uncertainty by measuring variability in the importance score. However, they fail to provide meaningful estimates of whether a pixel is certainly important or not. In this work, we propose a new R-XAI method called REPEAT that addresses the key question of whether or not a pixel is certainly important. REPEAT leverages the stochasticity of current R-XAI methods to produce multiple estimates of importance, thus considering each pixel in an image as a Bernoulli random variable that is either important or unimportant. From these Bernoulli random variables we can directly estimate the importance of a pixel and its associated certainty, thus enabling users to determine certainty in pixel importance. Our extensive evaluation shows that REPEAT gives certainty estimates that are more intuitive, better at detecting out-of-distribution data, and more concise. Kristoffer Wickstrøm, Thea Brüsch, Michael Kampffmeyer, Robert Jenssen |
AAAI | 4 |
| 2025 | Aggregation of Dependent Expert Distributions in Multimodal Variational AutoencodersabstractMultimodal learning with variational autoencoders (VAEs) requires estimating joint distributions to evaluate the evidence lower bound (ELBO). Current methods, the product and mixture of experts, aggregate single-modality distributions assuming independence for simplicity, which is an overoptimistic assumption. This research introduces a novel methodology for aggregating single-modality distributions by exploiting the principle of *consensus of dependent experts* (CoDE), which circumvents the aforementioned assumption. Utilizing the CoDE method, we propose a novel ELBO that approximates the joint likelihood of the multimodal data by learning the contribution of each subset of modalities. The resulting CoDE-VAE model demonstrates better performance in terms of balancing the trade-off between generative coherence and generative quality, as well as generating more precise log-likelihood estimations. CoDE-VAE further minimizes the generative quality gap as the number of modalities increases. In certain cases, it reaches a generative quality similar to that of unimodal VAEs, which is a desirable property that is lacking in most current methods. Finally, the classification accuracy achieved by CoDE-VAE is comparable to that of state-of-the-art multimodal VAE models. Rogelio Andrade Mancisidor, Robert Jenssen, Shujian Yu, Michael Kampffmeyer |
ICML | 2 |
| 2025 | Reconsidering Explicit Longitudinal Mammography Alignment for Enhanced Breast Cancer Risk Prediction
Solveig Thrun, Stine Hansen, Zijun Sun, Nele Blum, Suaiba Amina Salahuddin, Kristoffer Wickstrøm, Elisabeth Wetzer, Robert Jenssen, Maik Stille, Michael Kampffmeyer |
MICCAI (2) | 8 |
| 2025 | Generalized Cauchy-Schwarz divergence: Efficient estimation and applications in deep learning
Mingfei Lu, Shujian Yu, Robert Jenssen, Badong Chen |
Neurocomputing | 3 |
| 2025 | The Conditional Cauchy-Schwarz Divergence With Applications to Time-Series Data and Sequential Decision MakingabstractThe Cauchy-Schwarz (CS) divergence was developed by Príncipe et al. in 2000. In this paper, we extend the classic CS divergence to quantify the closeness between two conditional distributions and show that the developed conditional CS divergence can be elegantly estimated by a kernel density estimator from given samples. We illustrate the advantages (e.g., rigorous faithfulness guarantee, lower computational complexity, higher statistical power, and much more flexibility in a wide range of applications) of our conditional CS divergence over previous proposals, such as the conditional Kullback-Leibler divergence and the conditional maximum mean discrepancy. We also demonstrate the compelling performance of conditional CS divergence in two machine learning tasks related to time series data and sequential inference, namely time series clustering and uncertainty-guided exploration for sequential decision making. Shujian Yu, Sigurd Løkse, Robert Jenssen, José C. Príncipe |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Guest Editorial: Special Issue on Information Theoretic Methods for the Generalization, Robustness, and Interpretability of Machine Learning
Badong Chen, Shujian Yu, Robert Jenssen, José C. Príncipe, Klaus-Robert Müller |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | BrainIB: Interpretable Brain Network-Based Psychiatric Diagnosis With Graph Information BottleneckabstractDeveloping new diagnostic models based on the underlying biological mechanisms rather than subjective symptoms for psychiatric disorders is an emerging consensus. Recently, machine learning (ML)-based classifiers using functional connectivity (FC) for psychiatric disorders and healthy controls (HCs) are developed to identify brain markers. However, existing ML-based diagnostic models are prone to overfitting (due to insufficient training samples) and perform poorly in new test environments. Furthermore, it is difficult to obtain explainable and reliable brain biomarkers elucidating the underlying diagnostic decisions. These issues hinder their possible clinical applications. In this work, we propose BrainIB, a new graph neural network (GNN) framework to analyze functional magnetic resonance images (fMRI), by leveraging the famed information bottleneck (IB) principle. BrainIB is able to identify the most informative edges in the brain (i.e., subgraph) and generalizes well to unseen data. We evaluate the performance of BrainIB against three baselines and seven state-of-the-art (SOTA) brain network classification methods on three psychiatric datasets and observe that our BrainIB always achieves the highest diagnosis accuracy. It also discovers the subgraph biomarkers that are consistent with clinical and neuroimaging findings. The source code and implementation details of BrainIB are freely available at the GitHub repository (https://github.com/SJYuCNEL/brain-and-Information-Bottleneck). Kaizhong Zheng, Shujian Yu, Baojuan Li, Robert Jenssen, Badong Chen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | DIB-X: Formulating Explainability Principles for a Self-Explainable Model Through Information Theoretic LearningabstractThe recent development of self-explainable deep learning approaches has focused on integrating well-defined explainability principles into learning process, with the goal of achieving these principles through optimization. In this work, we propose DIB-X, a self-explainable deep learning approach for image data, which adheres to the principles of minimal, sufficient, and interactive explanations. The minimality and sufficiency principles are rooted from the trade-off relationship within the information bottleneck framework. Distinctly, DIB-X directly quantifies the minimality principle using the recently proposed matrix-based Rényi’s α-order entropy functional, circumventing the need for variational approximation and distributional assumption. The interactivity principle is realized by incorporating existing domain knowledge as prior explanations, fostering explanations that align with established domain understanding. Empirical results on MNIST and two marine environment monitoring datasets with different modalities reveal that our approach primarily provides improved explainability with the added advantage of enhanced classification performance. Changkyu Choi, Shujian Yu, Michael Kampffmeyer, Arnt-Børre Salberg, Nils Olav Handegard, Robert Jenssen |
ICASSP | 6 |
| 2024 | MAP IT to Visualize RepresentationsabstractMAP IT visualizes representations by taking a fundamentally different approach to dimensionality reduction. MAP IT aligns distributions over discrete marginal probabilities in the input space versus the target space, thus capturing information in local regions, as opposed to current methods which align based on individual probabilities between pairs of data points (states) only. The MAP IT theory reveals that alignment based on a projective divergence avoids normalization of weights (to obtain true probabilities) entirely, and further reveals a dual viewpoint via continuous densities and kernel smoothing. MAP IT is shown to produce visualizations which capture class structure better than the current state of the art while being inherently scalable. Robert Jenssen |
ICLR | 1 |
| 2024 | Cauchy-Schwarz Divergence Information Bottleneck for RegressionabstractThe information bottleneck (IB) approach is popular to improve the generalization, robustness and explainability of deep neural networks. Essentially, it aims to find a minimum sufficient representation $\mathbf{t}$ by striking a trade-off between a compression term $I(\mathbf{x};\mathbf{t})$ and a prediction term $I(y;\mathbf{t})$, where $I(\cdot;\cdot)$ refers to the mutual information (MI). MI is for the IB for the most part expressed in terms of the Kullback-Leibler (KL) divergence, which in the regression case corresponds to prediction based on mean squared error (MSE) loss with Gaussian assumption and compression approximated by variational inference.
In this paper, we study the IB principle for the regression problem and develop a new way to parameterize the IB with deep neural networks by exploiting favorable properties of the Cauchy-Schwarz (CS) divergence. By doing so, we move away from MSE-based regression and ease estimation by avoiding variational approximations or distributional assumptions. We investigate the improved generalization ability of our proposed CS-IB and demonstrate strong adversarial robustness guarantees. We demonstrate its superior performance on six real-world regression tasks over other popular deep IB approaches. We additionally observe that the solutions discovered by CS-IB always achieve the best trade-off between prediction accuracy and compression ratio in the information plane. The code is available at \url{https://github.com/SJYuCNEL/Cauchy-Schwarz-Information-Bottleneck}. Shujian Yu, Sigurd Løkse, Robert Jenssen, José C. Príncipe |
ICLR | 4 |
| 2024 | Finding NEM-U: Explaining unsupervised representation learning through neural network generated explanation masksabstractUnsupervised representation learning has become an important ingredient of today's deep learning systems. However, only a few methods exist that explain a learned vector embedding in the sense of providing information about which parts of an input are the most important for its representation. These methods generate the explanation for a given input after the model has been evaluated and tend to produce either inaccurate explanations or are slow, which limits their practical use. To address these limitations, we introduce the Neural Explanation Masks (NEM) framework, which turns a fixed representation model into a self-explaining model by augmenting it with a masking network. This network provides occlusion-based explanations in parallel to computing the representations during inference. We present an instance of this framework, the NEM-U (NEM using U-net structure) architecture, which leverages similarities between segmentation and occlusion-based masks. Our experiments show that NEM-U generates explanations faster and with lower complexity compared to the current state-of-the-art while maintaining high accuracy as measured by locality. Bjørn Leth Møller, Christian Igel, Kristoffer Wickstrøm, Jon Sporring, Robert Jenssen, Bulat Ibragimov |
ICML | 5 |
| 2024 | LSNetv2: Improving weakly supervised power line detection with bipartite matchingabstractThis paper addresses the crucial task of power line detection and localization in electrical infrastructure inspection using Unmanned Aerial Vehicles (UAVs) from weak supervision, polyline annotations. We first identify several limitations in the state-of-the-art approach LSNet. In particular, the inability of LSNet to detect line-crossings and lines in close proximity. To overcome these limitations, we propose LSNetv2, which enhances LSNet with multi-line segment detection capability facilitated via a bipartite matching loss. Additionally, we update LSNet’s regression loss in order to stabilize training by reducing the interdependence between predicted coordinates. Finally, LSNetv2 makes use of an increased receptive field to extract global information, improving overall detection performance. Through extensive evaluations on various power line detection datasets, LSNetv2 demonstrates superior performance and robustness. On the public datasets PLDU, PLDM and TTPLA, it achieved Fβ scores of 0.857, 0.875, and 0.671, respectively, while using only modified weak polyline annotation, establishing itself as an effective and efficient solution for power line detection in UAV-based electrical infrastructure inspections. Duy Khoi Tran, Van Nhan Nguyen, Davide Roverso, Robert Jenssen, Michael Kampffmeyer |
Expert Syst. Appl. | 4 |
| 2024 | Interrogating Sea Ice Predictability With GradientsabstractPredicting sea ice concentration is an important task in climate analysis. The recently proposed deep learning system IceNet is the state of the art sea ice prediction model. IceNet takes high-dimensional climate simulations and observational data as input features and forecasts sea ice concentration for the next six months over a spatial grid over the northern hemisphere. The model has proven to be particularly good at predicting extreme sea ice events compared to previous dynamical models, but lacks interpretability. In the original IceNet paper, a permute-and-predict approach was taken for assessing feature importance. However, this approach is not capable of revealing whether a feature contributes positively or negatively to the final prediction, nor can it reveal the importance of features over the spatial grid of predictions. In this paper, we take steps to instead interrogate the effect of the IceNet input feature with a gradient-based analysis, taking advantage of developments within the deep learning literature to open the so-called black box. Our analysis focuses on the unusually large sea ice extent event in September 2013 and indicates that IceNet places a strong emphasis on previous observations of sea ice concentration, linear trends, and seasonal components when making predictions. In our analysis, we identify which input features that are most influential for the prediction, and also at which spatial location these measurements are particularly influential. Harald L. Joakimsen, Iver Martinsen, Luigi Tommaso Luppino, Andrew McDonald 0003, Scott Hosking, Robert Jenssen |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2024 | Discriminative multimodal learning via conditional priors in generative modelsabstractDeep generative models with latent variables have been used lately to learn joint representations and generative processes from multi-modal data, which depict an object from different viewpoints. These two learning mechanisms can, however, conflict with each other and representations can fail to embed information on the data modalities. This research studies the realistic scenario in which all modalities and class labels are available for model training, e.g. images or handwriting, but where some modalities and labels required for downstream tasks are missing, e.g. text or annotations. We show, in this scenario, that the variational lower bound limits mutual information between joint representations and missing modalities. We, to counteract these problems, introduce a novel conditional multi-modal discriminative model that uses an informative prior distribution and optimizes a likelihood-free objective function that maximizes mutual information between joint representations and missing modalities. Extensive experimentation demonstrates the benefits of our proposed model, empirical results show that our model achieves state-of-the-art results in representative problems such as downstream classification, acoustic inversion, and image and annotation generation. Rogelio Andrade Mancisidor, Michael Kampffmeyer, Kjersti Aas, Robert Jenssen |
Neural Networks | 4 |
| 2024 | Leveraging tensor kernels to reduce objective function mismatch in deep clusteringabstractObjective Function Mismatch (OFM) occurs when the optimization of one objective has a negative impact on the optimization of another objective. In this work we study OFM in deep clustering, and find that the popular autoencoder-based approach to deep clustering can lead to both reduced clustering performance, and a significant amount of OFM between the reconstruction and clustering objectives. To reduce the mismatch, while maintaining the structure-preserving property of an auxiliary objective, we propose a set of new auxiliary objectives for deep clustering, referred to as the Unsupervised Companion Objectives (UCOs). The UCOs rely on a kernel function to formulate a clustering objective on intermediate representations in the network. Generally, intermediate representations can include other dimensions, for instance spatial or temporal, in addition to the feature dimension. We therefore argue that the naïve approach of vectorizing and applying a vector kernel is suboptimal for such representations, as it ignores the information contained in the other dimensions. To address this drawback, we equip the UCOs with structure-exploiting tensor kernels, designed for tensors of arbitrary rank. The UCOs can thus be adapted to a broad class of network architectures. We also propose a novel, regression-based measure of OFM, allowing us to accurately quantify the amount of OFM observed during training. Our experiments show that the OFM between the UCOs and the main clustering objective is lower, compared to a similar autoencoder-based model. Further, we illustrate that the UCOs improve the clustering performance of the model, in contrast to the autoencoder-based approach. The code for our experiments is available at https://github.com/danieltrosten/tk-uco. Daniel J. Trosten, Sigurd Løkse, Robert Jenssen, Michael Kampffmeyer |
Pattern Recognit. | 3 |
| 2024 | Code-Aligned Autoencoders for Unsupervised Change Detection in Multimodal Remote Sensing ImagesabstractImage translation with convolutional autoencoders has recently been used as an approach to multimodal change detection (CD) in bitemporal satellite images. A main challenge is the alignment of the code spaces by reducing the contribution of change pixels to the learning of the translation function. Many existing approaches train the networks by exploiting supervised information of the change areas, which, however, is not always available. We propose to extract relational pixel information captured by domain-specific affinity matrices at the input and use this to enforce alignment of the code spaces and reduce the impact of change pixels on the learning objective. A change prior is derived in an unsupervised fashion from pixel pair affinities that are comparable across domains. To achieve code space alignment, we enforce pixels with similar affinity relations in the input domains to be correlated also in code space. We demonstrate the utility of this procedure in combination with cycle consistency. The proposed approach is compared with the state-of-the-art machine learning and deep learning algorithms. Experiments conducted on four real and representative datasets show the effectiveness of our methodology. Luigi Tommaso Luppino, Mads A. Hansen, Michael Kampffmeyer, Filippo Maria Bianchi, Gabriele Moser, Robert Jenssen, Stian Normann Anfinsen |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Hubs and Hyperspheres: Reducing Hubness and Improving Transductive Few-Shot Learning with Hyperspherical EmbeddingsabstractDistance-based classification is frequently used in transductive few-shot learning (FSL). However, due to the high-dimensionality of image representations, FSL classifiers are prone to suffer from the hubness problem, where a few points (hubs) occur frequently in multiple nearest neighbour lists of other points. Hubness negatively impacts distance-based classification when hubs from one class appear often among the nearest neighbors of points from another class, degrading the classifier's performance. To address the hubness problem in FSL, we first prove that hubness can be eliminated by distributing representations uniformly on the hypersphere. We then propose two new approaches to embed representations on the hypersphere, which we prove optimize a tradeoff between uniformity and local similarity preservation - reducing hubness while retaining class structure. Our experiments show that the proposed methods reduce hubness, and significantly improves transductive FSL accuracy for a wide range of classifiers11Code available at https://github.com/uitml/noHub.. Daniel J. Trosten, Rwiddhi Chakraborty, Sigurd Løkse, Kristoffer Wickstrøm, Robert Jenssen, Michael Kampffmeyer |
CVPR | 5 |
| 2023 | On the Effects of Self-supervision and Contrastive Alignment in Deep Multi-view ClusteringabstractSelf-supervised learning is a central component in recent approaches to deep multi-view clustering (MVC). However, we find large variations in the development of self-supervision-based methods for deep MVC, potentially slowing the progress of the field. To address this, we present Deep-MVC, a unified framework for deep MVC that includes many recent methods as instances. We leverage our framework to make key observations about the effect of self-supervision, and in particular, drawbacks of aligning representations with contrastive learning. Further, we prove that contrastive alignment can negatively influence cluster separability, and that this effect becomes worse when the number of views increases. Motivated by our findings, we develop several new DeepMVC instances with new forms of self-supervision. We conduct extensive experiments and find that (i) in line with our theoretical findings, contrastive alignments decreases performance on datasets with many views; (ii) all methods benefit from some form of self-supervision; and (iii) our new instances outperform previous methods on several datasets. Based on our results, we suggest several promising directions for future research. To enhance the openness of the field, we provide an open-source implementation of DeepMVC, including recent models and our new instances. Our implementation includes a consistent evaluation protocol, facilitating fair and accurate evaluation of methods and components11Code: https://github.com/DanielTrosten/DeepMVC. Daniel J. Trosten, Sigurd Løkse, Robert Jenssen, Michael Kampffmeyer |
CVPR | 3 |
| 2023 | Supercm: Revisiting Clustering for Semi-Supervised LearningabstractThe development of semi-supervised learning (SSL) has in recent years largely focused on the development of new consistency regularization or entropy minimization approaches, often resulting in models with complex training strategies to obtain the desired results. In this work, we instead propose a novel approach that explicitly incorporates the underlying clustering assumption in SSL through extending a recently proposed differentiable clustering module. Leveraging annotated data to guide the cluster centroids results in a simple end-to-end trainable deep SSL approach. We demonstrate that the proposed model improves the performance over the supervised-only baseline and show that our framework can be used in conjunction with other SSL methods to further boost their performance. Durgesh Singh 0003, Ahcène Boubekki, Robert Jenssen, Michael Kampffmeyer |
ICASSP | 3 |
| 2023 | RELAX: Representation Learning ExplainabilityabstractAbstract Despite the significant improvements that self-supervised representation learning has led to when learning from unlabeled data, no methods have been developed that explain what influences the learned representation. We address this need through our proposed approach, RELAX, which is the first approach for attribution-based explanations of representations. Our approach can also model the uncertainty in its explanations, which is essential to produce trustworthy explanations. RELAX explains representations by measuring similarities in the representation space between an input and masked out versions of itself, providing intuitive explanations that significantly outperform the gradient-based baselines. We provide theoretical interpretations of RELAX and conduct a novel analysis of feature extractors trained using supervised and unsupervised learning, providing insights into different learning strategies. Moreover, we conduct a user study to assess how well the proposed approach aligns with human intuition and show that the proposed method outperforms the baselines in both the quantitative and human evaluation studies. Finally, we illustrate the usability of RELAX in several use cases and highlight that incorporating uncertainty can be essential for providing faithful explanations, taking a crucial step towards explaining representations. Kristoffer Wickstrøm, Daniel J. Trosten, Sigurd Løkse, Ahcène Boubekki, Karl Øyvind Mikalsen, Michael Kampffmeyer, Robert Jenssen |
Int. J. Comput. Vis. | 7 |
| 2023 | ADNet++: A few-shot learning framework for multi-class medical image volume segmentation with uncertainty-guided feature refinementabstractA major barrier to applying deep segmentation models in the medical domain is their typical data-hungry nature, requiring experts to collect and label large amounts of data for training. As a reaction, prototypical few-shot segmentation (FSS) models have recently gained traction as data-efficient alternatives. Nevertheless, despite the recent progress of these models, they still have some essential shortcomings that must be addressed. In this work, we focus on three of these shortcomings: (i) the lack of uncertainty estimation, (ii) the lack of a guiding mechanism to help locate edges and encourage spatial consistency in the segmentation maps, and (iii) the models' inability to do one-step multi-class segmentation. Without modifying or requiring a specific backbone architecture, we propose a modified prototype extraction module that facilitates the computation of uncertainty maps in prototypical FSS models, and show that the resulting maps are useful indicators of the model uncertainty. To improve the segmentation around boundaries and to encourage spatial consistency, we propose a novel feature refinement module that leverages structural information in the input space to help guide the segmentation in the feature space. Furthermore, we demonstrate how uncertainty maps can be used to automatically guide this feature refinement. Finally, to avoid ambiguous voxel predictions that occur when images are segmented class-by-class, we propose a procedure to perform one-step multi-class FSS. The efficiency of our proposed methodology is evaluated on two representative datasets for abdominal organ segmentation (CHAOS dataset and BTCV dataset) and one dataset for cardiac segmentation (MS-CMRSeg dataset). The results show that our proposed methodology significantly (one-sided Wilcoxon signed rank test, p<0.05) improves the baseline, increasing the overall dice score with +5.2, +5.1, and +2.8 percentage points for the CHAOS dataset, the BTCV dataset, and the MS-CMRSeg dataset, respectively. Stine Hansen, Srishti Gautam, Suaiba Amina Salahuddin, Michael Kampffmeyer, Robert Jenssen |
Medical Image Anal. | 5 |
| 2023 | This looks More Like that: Enhancing Self-Explaining Models by Prototypical Relevance PropagationabstractCurrent machine learning models have shown high efficiency in solving a wide variety of real-world problems. However, their black box character poses a major challenge for the comprehensibility and traceability of the underlying decision-making strategies. As a remedy, numerous post-hoc and self-explanation methods have been developed to interpret the models’ behavior. Those methods, in addition, enable the identification of artifacts that, inherent in the training data, can be erroneously learned by the model as class-relevant features. In this work, we provide a detailed case study of a representative for the state-of-the-art self-explaining network, ProtoPNet, in the presence of a spectrum of artifacts. Accordingly, we identify the main drawbacks of ProtoPNet, especially its coarse and spatially imprecise explanations. We address these limitations by introducing Prototypical Relevance Propagation (PRP), a novel method for generating more precise model-aware explanations. Furthermore, in order to obtain a clean, artifact-free dataset, we propose to use multi-view clustering strategies for segregating the artifact images using the PRP explanations, thereby suppressing the potential artifact learning in the models. Srishti Gautam, Marina M.-C. Höhne, Stine Hansen, Robert Jenssen, Michael Kampffmeyer |
Pattern Recognit. | 4 |
| 2023 | Selective Imputation for Multivariate Time Series Datasets With Missing ValuesabstractMultivariate time series often contain missing values for reasons such as failures in data collection mechanisms. Since these missing values can complicate the analysis of time series data, imputation techniques are typically used to deal with this issue. However, the quality of the imputation directly affects the performance of downstream tasks. In this paper, we propose a selective imputation method that identifies a subset of timesteps with missing values to impute in a multivariate time series dataset. This selection, which will result in shorter and simpler time series, is based on both reducing the uncertainty of the imputations and representing the original time series as good as possible. In particular, the method uses multi-objective optimization techniques to select the optimal set of points, and in this selection process, we leverage the beneficial properties of the Multi-task Gaussian Process (MGP). The method is applied to different datasets to analyze the quality of the imputations and the performance obtained in downstream tasks, such as classification or anomaly detection. The results show that much shorter and simpler time series are able to maintain or even improve both the quality of the imputations and the performance of the downstream tasks. Ane Blázquez-García, Kristoffer Wickstrøm, Shujian Yu, Karl Øyvind Mikalsen, Ahcène Boubekki, Angel Conde, Usue Mori, Robert Jenssen, José Antonio Lozano 0001 |
IEEE Trans. Knowl. Data Eng. | 8 |
| 2022 | ProtoVAE: A Trustworthy Self-Explainable Prototypical Variational ModelabstractThe need for interpretable models has fostered the development of self-explainable classifiers. Prior approaches are either based on multi-stage optimization schemes, impacting the predictive performance of the model, or produce explanations that are not transparent, trustworthy or do not capture the diversity of the data. To address these shortcomings, we propose ProtoVAE, a variational autoencoder-based framework that learns class-specific prototypes in an end-to-end manner and enforces trustworthiness and diversity by regularizing the representation space and introducing an orthonormality constraint. Finally, the model is designed to be transparent by directly incorporating the prototypes into the decision process. Extensive comparisons with previous self-explainable approaches demonstrate the superiority of ProtoVAE, highlighting its ability to generate trustworthy and diverse explanations, while not degrading predictive performance. Srishti Gautam, Ahcène Boubekki, Stine Hansen, Suaiba Amina Salahuddin, Robert Jenssen, Marina M.-C. Höhne, Michael Kampffmeyer |
NeurIPS | 5 |
| 2022 | Principle of relevant information for graph sparsificationabstractGraph sparsification aims to reduce the number of edges of a graph while maintaining its structural properties. In this paper, we propose the first general and effective information-theoretic formulation of graph sparsification, by taking inspiration from the Principle of Relevant Information (PRI). To this end, we extend the PRI from a standard scalar random variable setting to structured data (i.e., graphs). Our Graph-PRI objective is achieved by operating on the graph Laplacian, made possible by expressing the graph Laplacian of a subgraph in terms of a sparse edge selection vector w. We provide both theoretical and empirical justifications on the validity of our Graph-PRI approach. We also analyze its analytical solutions in a few special cases. We finally present three representative real-world applications, namely graph sparsification, graph regularized multi-task learning, and medical imaging-derived brain network classification, to demonstrate the effectiveness, the versatility and the enhanced interpretability of our approach over prevalent sparsification techniques. Code of Graph-PRI is available at https://github.com/SJYuCNEL/PRI-Graphs. Shujian Yu, Francesco Alesiani, Wenzhe Yin, Robert Jenssen, José C. Príncipe |
UAI | 4 |
| 2022 | Generating customer's credit behavior with deep generative modelsabstractBanks collect data x1 in loan applications to decide whether to grant credit and accepted applications generate new data x2 throughout the loan period. Hence, banks have two measurement-modalities, which provide a complete picture about customers. If we can generate x2 conditioned on x1 keeping the relationship between these two modalities, credit and behavior scoring may be enabled simultaneously (at the time x1 is obtained) to support cross-selling, launching of new products or marketing campaigns. Therefore, we develop a novel conditional bi-modal discriminative (CBMD) model for credit scoring, which is able to generate x2 based on x1 and can classify the outcome of loans in an unified framework. The idea behind CBMD is to learn joint (among modalities) latent representations that are useful to generate x2 using the available data x1 during the application process. The classifier model introduced in CBMD encourages the generative process to generate x2 accurately. Further, CBMD optimizes a novel objective function introduced in this research, which maximizes mutual information between the latent representation z and the modality x2 to improve the generative process in the model. We benchmark the generative process of our proposed model and CBMD outperforms other multi-learning models. Similarly, the classification performance of CBMD is tested under different scenarios and it achieves higher or on a par model performance compared to the state-of-the-art in multi-modal learning models. Rogelio Andrade Mancisidor, Michael Kampffmeyer, Kjersti Aas, Robert Jenssen |
Knowl. Based Syst. | 4 |
| 2022 | Anomaly detection-inspired few-shot medical image segmentation through self-supervision with supervoxelsabstractRecent work has shown that label-efficient few-shot learning through self-supervision can achieve promising medical image segmentation results. However, few-shot segmentation models typically rely on prototype representations of the semantic classes, resulting in a loss of local information that can degrade performance. This is particularly problematic for the typically large and highly heterogeneous background class in medical image segmentation problems. Previous works have attempted to address this issue by learning additional prototypes for each class, but since the prototypes are based on a limited number of slices, we argue that this ad-hoc solution is insufficient to capture the background properties. Motivated by this, and the observation that the foreground class (e.g., one organ) is relatively homogeneous, we propose a novel anomaly detection-inspired approach to few-shot medical image segmentation in which we refrain from modeling the background explicitly. Instead, we rely solely on a single foreground prototype to compute anomaly scores for all query pixels. The segmentation is then performed by thresholding these anomaly scores using a learned threshold. Assisted by a novel self-supervision task that exploits the 3D structure of medical images through supervoxels, our proposed anomaly detection-inspired few-shot medical image segmentation model outperforms previous state-of-the-art approaches on two representative MRI datasets for the tasks of abdominal organ segmentation and cardiac segmentation. Stine Hansen, Srishti Gautam, Robert Jenssen, Michael Kampffmeyer |
Medical Image Anal. | 3 |
| 2022 | Mixing up contrastive learning: Self-supervised representation learning for time seriesabstractThe lack of labeled data is a key challenge for learning useful representation from time series data. However, an unsupervised representation framework that is capable of producing high quality representations could be of great value. It is key to enabling transfer learning, which is especially beneficial for medical applications, where there is an abundance of data but labeling is costly and time consuming. We propose an unsupervised contrastive learning framework that is motivated from the perspective of label smoothing. The proposed approach uses a novel contrastive loss that naturally exploits a data augmentation scheme in which new samples are generated by mixing two data samples with a mixing component. The task in the proposed framework is to predict the mixing component, which is utilized as soft targets in the loss function. Experiments demonstrate the framework’s superior performance compared to other representation learning approaches on both univariate and multivariate time series and illustrate its benefits for transfer learning for clinical time series. Kristoffer Wickstrøm, Michael Kampffmeyer, Karl Øyvind Mikalsen, Robert Jenssen |
Pattern Recognit. Lett. | 4 |
| 2022 | Deep Image Translation With an Affinity-Based Change Prior for Unsupervised Multimodal Change DetectionabstractImage translation with convolutional neural networks has recently been used as an approach to multimodal change detection. Existing approaches train the networks by exploiting supervised information of the change areas, which, however, is not always available. A main challenge in the unsupervised problem setting is to avoid that change pixels affect the learning of the translation function. We propose two new network architectures trained with loss functions weighted by priors that reduce the impact of change pixels on the learning objective. The change prior is derived in an unsupervised fashion from relational pixel information captured by domain-specific affinity matrices. Specifically, we use the vertex degrees associated with an absolute affinity difference matrix and demonstrate their utility in combination with cycle consistency and adversarial training. The proposed neural networks are compared with the state-of-the-art algorithms. Experiments conducted on three real data sets show the effectiveness of our methodology. Luigi Tommaso Luppino, Michael Kampffmeyer, Filippo Maria Bianchi, Gabriele Moser, Sebastiano B. Serpico, Robert Jenssen, Stian Normann Anfinsen |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Clinically Relevant Features for Predicting the Severity of Surgical Site InfectionsabstractSurgical site infections are hospital-acquired infections resulting in severe risk for patients and significantly increased costs for healthcare providers. In this work, we show how to leverage irregularly sampled preoperative blood tests to predict, on the day of surgery, a future surgical site infection and its severity. Our dataset is extracted from the electronic health records of patients who underwent gastrointestinal surgery and developed either deep, shallow or no infection. We represent the patients using the concentrations of fourteen common blood components collected over the four weeks preceding the surgery partitioned into six time windows. A gradient boosting based classifier trained on our new set of features reports an AUROC of 0.991 for predicting a postoperative infection and and AUROC of 0.937 for classifying the severity of the infection. Further analyses support the clinical relevance of our approach as the most important features describe the nutritional status and the liver function over the two weeks prior to surgery. Ahcène Boubekki, Jonas Nordhaug Myhre, Luigi Tommaso Luppino, Karl Øyvind Mikalsen, Arthur Revhaug, Robert Jenssen |
IEEE J. Biomed. Health Informatics | 6 |
| 2021 | Measuring Dependence with Matrix-based Entropy FunctionalabstractMeasuring the dependence of data plays a central role in statistics and machine learning. In this work, we summarize and generalize the main idea of existing information-theoretic dependence measures into a higher-level perspective by the Shearer's inequality. Based on our generalization, we then propose two measures, namely the matrix-based normalized total correlation and the matrix-based normalized dual total correlation, to quantify the dependence of multiple variables in arbitrary dimensional space, without explicit estimation of the underlying data distributions. We show that our measures are differentiable and statistically more powerful than prevalent ones. We also show the impact of our measures in four different machine learning problems, namely the gene regulatory network inference, the robust machine learning under covariate shift and non-Gaussian noises, the subspace outlier detection, and the understanding of the learning dynamics of convolutional neural networks, to demonstrate their utilities, advantages, as well as implications to those problems. Shujian Yu, Francesco Alesiani, Robert Jenssen, José C. Príncipe |
AAAI | 4 |
| 2021 | Reconsidering Representation Alignment for Multi-View ClusteringabstractAligning distributions of view representations is a core component of today’s state of the art models for deep multi-view clustering. However, we identify several drawbacks with naïvely aligning representation distributions. We demonstrate that these drawbacks both lead to less separable clusters in the representation space, and inhibit the model’s ability to prioritize views. Based on these observations, we develop a simple baseline model for deep multi-view clustering. Our baseline model avoids representation alignment altogether, while performing similar to, or better than, the current state of the art. We also expand our baseline model by adding a contrastive learning component. This introduces a selective alignment procedure that preserves the model’s ability to prioritize views. Our experiments show that the contrastive learning component enhances the baseline model, improving on the current state of the art by a large margin on several datasets1. Daniel J. Trosten, Sigurd Løkse, Robert Jenssen, Michael Kampffmeyer |
CVPR | 3 |
| 2021 | Unsupervised supervoxel-based lung tumor segmentation across patient scans in hybrid PET/MRIabstractTumor segmentation is a crucial but difficult task in treatment planning and follow-up of cancerous patients. The challenge of automating the tumor segmentation has recently received a lot of attention, but the potential of utilizing hybrid positron emission tomography (PET)/magnetic resonance imaging (MRI), a novel and promising imaging modality in oncology, is still under-explored. Recent approaches have either relied on manual user input and/or performed the segmentation patient-by-patient, whereas a fully unsupervised segmentation framework that exploits the available information from all patients is still lacking. We present an unsupervised across-patients supervoxel-based clustering framework for lung tumor segmentation in hybrid PET/MRI. The method consists of two steps: First, each patient is represented by a set of PET/MRI supervoxel-features. Then the data points from all patients are transformed and clustered on a population level into tumor and non-tumor supervoxels. The proposed framework is tested on the scans of 18 non-small cell lung cancer patients with a total of 19 tumors and evaluated with respect to manual delineations provided by clinicians. Experiments study the performance of several commonly used clustering algorithms within the framework and provide analysis of (i) the effect of tumor size, (ii) the segmentation errors, (iii) the benefit of across-patient clustering, and (iv) the noise robustness. The proposed framework detected 15 out of 19 tumors in an unsupervised manner. Moreover, performance increased considerably by segmenting across patients, with the mean dice score increasing from 0.169±0.295 (patient-by-patient) to 0.470±0.308 (across-patients). Results demonstrate that both spectral clustering and Manhattan hierarchical clustering have the potential to segment tumors in PET/MRI with a low number of missed tumors and a low number of false-positives, but that spectral clustering seems to be more robust to noise. Stine Hansen, Samuel Kuttner, Michael Kampffmeyer, Tom-Vegard Markussen, Rune Sundset, Silje Kjærnes Øen, Live Eikenes, Robert Jenssen |
Expert Syst. Appl. | 8 |
| 2021 | Learning latent representations of bank customers with the Variational AutoencoderabstractLearning data representations that reflect the customers’ creditworthiness can improve marketing campaigns, customer relationship management , data and process management or the credit risk assessment in retail banks. In this research, we show that it is possible to steer data representations in the latent space of the Variational Autoencoder (VAE) using a semi-supervised learning framework and a specific grouping of the input data called Weight of Evidence (WoE). Our proposed method learns a latent representation of the data showing a well-defied clustering structure . The clustering structure captures the customers’ creditworthiness, which is unknown a priori and cannot be identified in the input space. The main advantages of our proposed method are that it captures the natural clustering of the data, suggests the number of clusters, captures the spatial coherence of customers’ creditworthiness, generates data representations of unseen customers and assign them to one of the existing clusters. Our empirical results, based on real data sets reflecting different market and economic conditions, show that none of the well-known data representation models in the benchmark analysis are able to obtain well-defined clustering structures like our proposed method. Further, we show how banks can use our proposed methodology to improve marketing campaigns and credit risk assessment. Rogelio Andrade Mancisidor, Michael Kampffmeyer, Kjersti Aas, Robert Jenssen |
Expert Syst. Appl. | 4 |
| 2021 | Joint optimization of an autoencoder for clustering and embeddingabstractAbstract Deep embedded clustering has become a dominating approach to unsupervised categorization of objects with deep neural networks. The optimization of the most popular methods alternates between the training of a deep autoencoder and a k -means clustering of the autoencoder’s embedding. The diachronic setting, however, prevents the former to benefit from valuable information acquired by the latter. In this paper, we present an alternative where the autoencoder and the clustering are learned simultaneously. This is achieved by providing novel theoretical insight, where we show that the objective function of a certain class of Gaussian mixture models (GMM’s) can naturally be rephrased as the loss function of a one-hidden layer autoencoder thus inheriting the built-in clustering capabilities of the GMM. That simple neural network, referred to as the clustering module, can be integrated into a deep autoencoder resulting in a deep clustering model able to jointly learn a clustering and an embedding. Experiments confirm the equivalence between the clustering module and Gaussian mixture models. Further evaluations affirm the empirical relevance of our deep architecture as it outperforms related baselines on several data sets. Ahcène Boubekki, Michael Kampffmeyer, Ulf Brefeld, Robert Jenssen |
Mach. Learn. | 4 |
| 2021 | LS-Net: fast single-shot line-segment detectorabstractAbstract In unmanned aerial vehicle (UAV) flights, power lines are considered as one of the most threatening hazards and one of the most difficult obstacles to avoid. In recent years, many vision-based techniques have been proposed to detect power lines to facilitate self-driving UAVs and automatic obstacle avoidance. However, most of the proposed methods are typically based on a common three-step approach: (i) edge detection, (ii) the Hough transform, and (iii) spurious line elimination based on power line constrains. These approaches not only are slow and inaccurate but also require a huge amount of effort in post-processing to distinguish between power lines and spurious lines. In this paper, we introduce LS-Net, a fast single-shot line-segment detector, and apply it to power line detection. The LS-Net is by design fully convolutional, and it consists of three modules: (i) a fully convolutional feature extractor, (ii) a classifier, and (iii) a line segment regressor. Due to the unavailability of large datasets with annotations of power lines, we render synthetic images of power lines using the physically based rendering approach and propose a series of effective data augmentation techniques to generate more training data. With a customized version of the VGG-16 network as the backbone, the proposed approach outperforms existing state-of-the-art approaches. In addition, the LS-Net can detect power lines in near real time. This suggests that our proposed approach has a promising role in automatic obstacle avoidance and as a valuable component of self-driving UAVs, especially for automatic autonomous power line inspection. Van Nhan Nguyen, Robert Jenssen, Davide Roverso |
Mach. Vis. Appl. | 2 |
| 2021 | Time series cluster kernels to exploit informative missingness and incomplete label informationabstractThe time series cluster kernel (TCK) provides a powerful tool for analysing multivariate time series subject to missing data. TCK is designed using an ensemble learning approach in which Bayesian mixture models form the base models. Because of the Bayesian approach, TCK can naturally deal with missing values without resorting to imputation and the ensemble strategy ensures robustness to hyperparameters, making it particularly well suited for unsupervised learning. However, TCK assumes missing at random and that the underlying missingness mechanism is ignorable, i.e. uninformative, an assumption that does not hold in many real-world applications, such as e.g. medicine. To overcome this limitation, we present a kernel capable of exploiting the potentially rich information in the missing values and patterns, as well as the information from the observed data. In our approach, we create a representation of the missing pattern, which is incorporated into mixed mode mixture models in such a way that the information provided by the missing patterns is effectively exploited. Moreover, we also propose a semi-supervised kernel, capable of taking advantage of incomplete label information to learn more accurate similarities. Experiments on benchmark data, as well as a real-world case study of patients described by longitudinal electronic health record data who potentially suffer from hospital-acquired infections, demonstrate the effectiveness of the proposed methods. Karl Øyvind Mikalsen, Cristina Soguero-Ruíz, Filippo Maria Bianchi, Arthur Revhaug, Robert Jenssen |
Pattern Recognit. | 5 |
| 2021 | Uncertainty-Aware Deep Ensembles for Reliable and Explainable Predictions of Clinical Time SeriesabstractDeep learning-based support systems have demonstrated encouraging results in numerous clinical applications involving the processing of time series data. While such systems often are very accurate, they have no inherent mechanism for explaining what influenced the predictions, which is critical for clinical tasks. However, existing explainability techniques lack an important component for trustworthy and reliable decision support, namely a notion of uncertainty. In this paper, we address this lack of uncertainty by proposing a deep ensemble approach where a collection of DNNs are trained independently. A measure of uncertainty in the relevance scores is computed by taking the standard deviation across the relevance scores produced by each model in the ensemble, which in turn is used to make the explanations more reliable. The class activation mapping method is used to assign a relevance score for each time step in the time series. Results demonstrate that the proposed ensemble is more accurate in locating relevant time steps and is more consistent across random initializations, thus making the model more trustworthy. The proposed methodology paves the way for constructing trustworthy and dependable support systems for processing clinical time series for healthcare related tasks. Kristoffer Wickstrøm, Karl Øyvind Mikalsen, Michael Kampffmeyer, Arthur Revhaug, Robert Jenssen |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | Reservoir Computing Approaches for Representation and Classification of Multivariate Time SeriesabstractClassification of multivariate time series (MTS) has been tackled with a large variety of methodologies and applied to a wide range of scenarios. Reservoir computing (RC) provides efficient tools to generate a vectorial, fixed-size representation of the MTS that can be further processed by standard classifiers. Despite their unrivaled training speed, MTS classifiers based on a standard RC architecture fail to achieve the same accuracy of fully trainable neural networks. In this article, we introduce the reservoir model space, an unsupervised approach based on RC to learn vectorial representations of MTS. Each MTS is encoded within the parameters of a linear model trained to predict a low-dimensional embedding of the reservoir dynamics. Compared with other RC methods, our model space yields better representations and attains comparable computational performance due to an intermediate dimensionality reduction procedure. As a second contribution, we propose a modular RC framework for MTS classification, with an associated open-source Python library. The framework provides different modules to seamlessly implement advanced RC architectures. The architectures are compared with other MTS classifiers, including deep learning models and time series kernels. Results obtained on the benchmark and real-world MTS data sets show that RC classifiers are dramatically faster and, when implemented using our proposed representation, also achieve superior classification accuracy. Filippo Maria Bianchi, Simone Scardapane, Sigurd Løkse, Robert Jenssen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Understanding Convolutional Neural Networks With Information Theory: An Initial ExplorationabstractA novel functional estimator for Rényi's α -entropy and its multivariate extension was recently proposed in terms of the normalized eigenspectrum of a Hermitian matrix of the projected data in a reproducing kernel Hilbert space (RKHS). However, the utility and possible applications of these new estimators are rather new and mostly unknown to practitioners. In this brief, we first show that this estimator enables straightforward measurement of information flow in realistic convolutional neural networks (CNNs) without any approximation. Then, we introduce the partial information decomposition (PID) framework and develop three quantities to analyze the synergy and redundancy in convolutional layer representations. Our results validate two fundamental data processing inequalities and reveal more inner properties concerning CNN training. Shujian Yu, Kristoffer Wickstrøm, Robert Jenssen, José C. Príncipe |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | SEN: A Novel Feature Normalization Dissimilarity Measure for Prototypical Few-Shot Learning Networks
Van Nhan Nguyen, Sigurd Løkse, Kristoffer Wickstrøm, Michael Kampffmeyer, Davide Roverso, Robert Jenssen |
ECCV (23) | 6 |
| 2020 | Self-Constructing Graph Convolutional Networks for Semantic LabelingabstractGraph Neural Networks (GNNs) have received increasing attention in many fields. However, due to the lack of prior graphs, their use for semantic labeling has been limited. Here, we propose a novel architecture called the Self-Constructing Graph (SCG), which makes use of learnable latent variables to generate embeddings and to self-construct the underlying graphs directly from the input features without relying on manually built prior knowledge graphs. SCG can automatically obtain optimized non-local context graphs from complex-shaped objects in aerial imagery. We optimize SCG via an adaptive diagonal enhancement method and a variational lower bound that consists of a customized graph reconstruction term and a Kullback-Leibler divergence regularization term. We demonstrate the effectiveness and flexibility of the proposed SCG on the publicly available ISPRS Vaihingen dataset and our model SCG-Net achieves competitive results in terms of F1-score with much fewer parameters and at a lower computational cost compared to related pure-CNN based work. Qinghui Liu, Michael Kampffmeyer, Robert Jenssen, Arnt-Børre Salberg |
IGARSS | 3 |
| 2020 | Deep generative models for reject inference in credit scoring
Rogelio Andrade Mancisidor, Michael Kampffmeyer, Kjersti Aas, Robert Jenssen |
Knowl. Based Syst. | 4 |
| 2020 | Uncertainty and interpretability in convolutional neural networks for semantic segmentation of colorectal polypsabstractColorectal polyps are known to be potential precursors to colorectal cancer, which is one of the leading causes of cancer-related deaths on a global scale. Early detection and prevention of colorectal cancer is primarily enabled through manual screenings, where the intestines of a patient is visually examined. Such a procedure can be challenging and exhausting for the person performing the screening. This has resulted in numerous studies on designing automatic systems aimed at supporting physicians during the examination. Recently, such automatic systems have seen a significant improvement as a result of an increasing amount of publicly available colorectal imagery and advances in deep learning research for object image recognition. Specifically, decision support systems based on Convolutional Neural Networks (CNNs) have demonstrated state-of-the-art performance on both detection and segmentation of colorectal polyps. However, CNN-based models need to not only be precise in order to be helpful in a medical context. In addition, interpretability and uncertainty in predictions must be well understood. In this paper, we develop and evaluate recent advances in uncertainty estimation and model interpretability in the context of semantic segmentation of polyps from colonoscopy images. Furthermore, we propose a novel method for estimating the uncertainty associated with important features in the input and demonstrate how interpretability and uncertainty can be modeled in DSSs for semantic segmentation of colorectal polyps. Results indicate that deep models are utilizing the shape and edge information of polyps to make their prediction. Moreover, inaccurate predictions show a higher degree of uncertainty compared to precise predictions. Kristoffer Wickstrøm, Michael Kampffmeyer, Robert Jenssen |
Medical Image Anal. | 3 |
| 2020 | Multivariate Extension of Matrix-Based Rényi's $\alpha$α-Order Entropy FunctionalabstractThe matrix-based Rényi's α-order entropy functional was recently introduced using the normalized eigenspectrum of a Hermitian matrix of the projected data in a reproducing kernel Hilbert space (RKHS). However, the current theory in the matrix-based Rényi's α-order entropy functional only defines the entropy of a single variable or mutual information between two random variables. In information theory and machine learning communities, one is also frequently interested in multivariate information quantities, such as the multivariate joint entropy and different interactive quantities among multiple variables. In this paper, we first define the matrix-based Rényi's α-order joint entropy among multiple variables. We then show how this definition can ease the estimation of various information quantities that measure the interactions among multiple variables, such as interactive information and total correlation. We finally present an application to feature selection to show how our definition provides a simple yet powerful way to estimate a widely-acknowledged intractable quantity from data. A real example on hyperspectral image (HSI) band selection is also provided. Shujian Yu, Luis Gonzalo Sánchez Giraldo, Robert Jenssen, José C. Príncipe |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Dense Dilated Convolutions' Merging Network for Land Cover ClassificationabstractLand cover classification of remote sensing images is a challenging task due to limited amounts of annotated data, highly imbalanced classes, frequent incorrect pixel-level annotations, and an inherent complexity in the semantic segmentation task. In this article, we propose a novel architecture called the dense dilated convolutions' merging network (DDCM-Net) to address this task. The proposed DDCM-Net consists of dense dilated image convolutions merged with varying dilation rates. This effectively utilizes rich combinations of dilated convolutions that enlarge the network's receptive fields with fewer parameters and features compared with the state-of-the-art approaches in the remote sensing domain. Importantly, DDCM-Net obtains fused local- and global-context information, in effect incorporating surrounding discriminative capability for multiscale and complex-shaped objects with similar color and textures in very high-resolution aerial imagery. We demonstrate the effectiveness, robustness, and flexibility of the proposed DDCM-Net on the publicly available ISPRS Potsdam and Vaihingen data sets, as well as the DeepGlobe land cover data set. Our single model, trained on three-band Potsdam and Vaihingen data sets, achieves better accuracy in terms of both mean intersection over union (mIoU) and F1-score compared with other published models trained with more than three-band data. We further validate our model on the DeepGlobe data set, achieving state-of-the-art result 56.2% mIoU with much fewer parameters and at a lower computational cost compared with related recent work. Qinghui Liu, Michael Kampffmeyer, Robert Jenssen, Arnt-Børre Salberg |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2019 | Recurrent Deep Divergence-based Clustering for Simultaneous Feature Learning and Clustering of Variable Length Time SeriesabstractThe task of clustering unlabeled time series and sequences entails a particular set of challenges, namely to adequately model temporal relations and variable sequence lengths. If these challenges are not properly handled, the resulting clusters might be of suboptimal quality. As a key solution, we present a joint clustering and feature learning framework for time series based on deep learning. For a given set of time series, we train a recurrent network to represent, or embed, each time series in a vector space such that a divergence-based clustering loss function can discover the underlying cluster structure in an end-to-end manner. Unlike previous approaches, our model inherently handles multivariate time series of variable lengths and does not require specification of a distance-measure in the input space. On a diverse set of benchmark datasets we illustrate that our proposed Recurrent Deep Divergence-based Clustering approach outperforms, or performs comparable to, previous approaches. Daniel J. Trosten, Andreas Storvik Strauman, Michael Kampffmeyer, Robert Jenssen |
ICASSP | 4 |
| 2019 | Road Mapping in Lidar Images Using a Joint-Task Dense Dilated Convolutions Merging NetworkabstractIt is important, but challenging, for the forest industry to accurately map roads which are used for timber transport by trucks. In this work, we propose a Dense Dilated Convolutions Merging Network (DDCM-Net) to detect these roads in lidar images. The DDCM-Net can effectively recognize multi-scale and complex shaped roads with similar texture and colors, and also is shown to have superior performance over existing methods. To further improve its ability to accurately infer categories of roads, we propose the use of a joint-task learning strategy that utilizes two auxiliary output branches, i.e, multi-class classification and binary segmentation, joined with the main output of full-class segmentation. This pushes the network towards learning more robust representations that are expected to boost the ultimate performance of the main task. In addition, we introduce an iterative-random-weighting method to automatically weigh the joint losses for auxiliary tasks. This can avoid the difficult and expensive process of tuning the weights of each task's loss by hand. The experiments demonstrate that our proposed jointtask DDCM-Net can achieve better performance with fewer parameters and higher computational efficiency than previous state-of-the-art approaches. Qinghui Liu, Michael Kampffmeyer, Robert Jenssen, Arnt-Børre Salberg |
IGARSS | 3 |
| 2019 | Deep divergence-based approach to clusteringabstractA promising direction in deep learning research consists in learning representations and simultaneously discovering cluster structure in unlabeled data by optimizing a discriminative loss function. As opposed to supervised deep learning, this line of research is in its infancy, and how to design and optimize suitable loss functions to train deep neural networks for clustering is still an open question. Our contribution to this emerging field is a new deep clustering network that leverages the discriminative power of information-theoretic divergence measures, which have been shown to be effective in traditional clustering. We propose a novel loss function that incorporates geometric regularization constraints, thus avoiding degenerate structures of the resulting clustering partition. Experiments on synthetic benchmarks and real datasets show that the proposed network achieves competitive performance with respect to other state-of-the-art methods, scales well to large datasets, and does not require pre-training steps. Michael Kampffmeyer, Sigurd Løkse, Filippo Maria Bianchi, Lorenzo Livi, Arnt-Børre Salberg, Robert Jenssen |
Neural Networks | 6 |
| 2019 | Learning representations of multivariate time series with missing data
Filippo Maria Bianchi, Lorenzo Livi, Karl Øyvind Mikalsen, Michael Kampffmeyer, Robert Jenssen |
Pattern Recognit. | 5 |
| 2019 | Noisy multi-label semi-supervised dimensionality reductionabstractNoisy labeled data represent a rich source of information that often are easily accessible and cheap to obtain, but label noise might also have many negative consequences if not accounted for. How to fully utilize noisy labels has been studied extensively within the framework of standard supervised machine learning over a period of several decades. However, very little research has been conducted on solving the challenge posed by noisy labels in non-standard settings. This includes situations where only a fraction of the samples are labeled (semi-supervised) and each high-dimensional sample is associated with multiple labels. In this work, we present a novel semi-supervised and multi-label dimensionality reduction method that effectively utilizes information from both noisy multi-labels and unlabeled data. With the proposed Noisy multi-label semi-supervised dimensionality reduction (NMLSDR) method, the noisy multi-labels are denoised and unlabeled data are labeled simultaneously via a specially designed label propagation algorithm. NMLSDR then learns a projection matrix for reducing the dimensionality by maximizing the dependence between the enlarged and denoised multi-label space and the features in the projected space. Extensive experiments on synthetic data, benchmark datasets, as well as a real-world case study, demonstrate the effectiveness of the proposed algorithm and show that it outperforms state-of-the-art multi-label feature extraction algorithms. Karl Øyvind Mikalsen, Cristina Soguero-Ruíz, Filippo Maria Bianchi, Robert Jenssen |
Pattern Recognit. | 4 |
| 2018 | Using multi-anchors to identify patients suffering from multimorbidities
Karl Øyvind Mikalsen, Cristina Soguero-Ruíz, I. Mora-Jiménez, Isabel Caballero-López-Fando, Robert Jenssen |
BIBM | 5 |
| 2018 | Learning compressed representations of blood samples time series with missing data
Filippo Maria Bianchi, Karl Øyvind Mikalsen, Robert Jenssen |
ESANN | 3 |
| 2018 | Bidirectional deep-readout echo state networks
Filippo Maria Bianchi, Simone Scardapane, Sigurd Løkse, Robert Jenssen |
ESANN | 4 |
| 2018 | Ranking Using Transition Probabilities Learned from Multi-Attribute DataabstractIn this paper, as a novel approach, we learn Markov chain transition probabilities for ranking of multi -attribute data from the inherent structures in the data itself. The procedure is inspired by consensus clustering and exploits a suitable form of the PageRank algorithm. This is very much in the spirit of the original PageRank utilizing the hyperlink structure to learn such probabilities. As opposed to existing approaches for ranking multi -attribute data, our method is not dependent on tuning of critical user-specified parameters. Experiments show the benefits of the proposed method. Sigurd Løkse, Robert Jenssen |
ICASSP | 2 |
| 2018 | A Comparison of Deep Learning Architectures for Semantic Mapping of Very High Resolution ImagesabstractSemantic mapping of land cover is a key, but challenging, problem in remote sensing. Recent advances in deep learning, especially deep convolutional neural networks (CNNs), have shown outstanding performance in this task. In order to develop refined deep learning pipeline for meeting the rising need for accurate semantic mapping in remote sensing images, this paper study and compare a number of advanced deep learning segmentation architectures, which have obtained state-of-the-art results on computer vision contests like the Pascal VOC. To further analyze and compare the effectiveness of some elaborate layers and underlying structures introduced by these architectures, we evaluate them by re-implementing, train and test them on ISPRS Potsdam dataset. Our results show that a promising performance with overall Fl_score above 87% and mIoU of 79% can be obtained by only using the RGB images, without any post-processing such as conditional random field (CRF) smoothing. At last, we propose several possible approaches to further enhance the deep learning architectures to better deal with high-resolution aerial images. We therefore consider this work to be helpful for the remote sensing research community. Qinghui Liu, Arnt-Børre Salberg, Robert Jenssen |
IGARSS | 3 |
| 2018 | Time series cluster kernel for learning similarities between multivariate time series with missing data
Karl Øyvind Mikalsen, Filippo Maria Bianchi, Cristina Soguero-Ruíz, Robert Jenssen |
Pattern Recognit. | 4 |
| 2018 | Robust clustering using a kNN mode seeking ensemble
Jonas Nordhaug Myhre, Karl Øyvind Mikalsen, Sigurd Løkse, Robert Jenssen |
Pattern Recognit. | 4 |
| 2017 | Density ridge manifold traversalabstractThe density ridge framework for estimating principal curves and surfaces has in a number of recent works been shown to capture manifold structure in data in an intuitive and effective manner. However, to date there exists no efficient way to traverse these manifolds as defined by density ridges. This is unfortunate, as manifold traversal is an important problem for example for shape estimation in medical imaging, or in general for being able to characterize and understand state transitions or local variability over the data manifold. In this paper, we remedy this situation by introducing a novel manifold traversal algorithm based on geodesics within the density ridge approach. The traversal is executed in a subspace capturing the intrinsic dimensionality of the data using dimensionality reduction techniques such as principal component analysis or kernel entropy component analysis. A mapping back to the ambient space is obtained by training a neural network. We compare against maximum mean discrepancy traversal, a recent approach, and obtain promising results. Jonas Nordhaug Myhre, Michael Kampffmeyer, Robert Jenssen |
ICASSP | 3 |
| 2017 | Urban land cover classification with missing data using deep convolutional neural networksabstractFusing different sensors with different data modalities is a common technique to improve land cover classification performance in remote sensing. However, all modalities are rarely available for all test data, and this missing data problem poses severe challenges for multi-modal learning. Inspired by recent successes in deep learning, we propose as a remedy a convolutional neural network architecture for urban remote sensing image segmentation trained on data modalities which are not all available at test time. We train our architecture with a cost function particularly suited for imbalanced classes, as this is a frequent problem in remote sensing. We demonstrate the method using a benchmark dataset containing RGB and DSM images. Assuming that the DSM images are missing during testing, our method outperforms both a CNN trained on RGB images as well as an ensemble of two CNNs trained on the RGB images, by exploiting the training time information of the missing modality. Michael Kampffmeyer, Arnt-Børre Salberg, Robert Jenssen |
IGARSS | 3 |
| 2017 | Temporal overdrive recurrent neural networkabstractIn this work we present a novel recurrent neural network architecture designed to model systems characterized by multiple characteristic timescales in their dynamics. The proposed network is composed by several recurrent groups of neurons that are trained to separately adapt to each timescale, in order to improve the system identification process. We test our framework on time series prediction tasks and we show some promising, preliminary results achieved on synthetic data. To evaluate the capabilities of our network, we compare the performance with several state-of-the-art recurrent architectures. Filippo Maria Bianchi, Michael Kampffmeyer, Enrico Maiorino, Robert Jenssen |
IJCNN | 4 |
| 2017 | Critical echo state network dynamics by means of Fisher information maximizationabstractThe computational capability of an Echo State Network (ESN), expressed in terms of low prediction error and high short-term memory capacity, is maximized on the so-called “edge of criticality”. In this paper we present a novel, unsupervised approach to identify this edge and, accordingly, we determine hyperparameters configuration that maximize network performance. The proposed method is application-independent and stems from recent theoretical results consolidating the link between Fisher information and critical phase transitions. We show how to identify optimal ESN hyperparameters by relying only on the Fisher information matrix (FIM) estimated from the activations of hidden neurons. In order to take into account the particular input signal driving the network dynamics, we adopt a recently proposed non-parametric FIM estimator. Experimental results on a set of standard benchmarks are provided and discussed, demonstrating the validity of the proposed method. Filippo Maria Bianchi, Lorenzo Livi, Robert Jenssen, Cesare Alippi |
IJCNN | 3 |
| 2017 | Optimized Kernel Entropy ComponentsabstractThis brief addresses two main issues of the standard kernel entropy component analysis (KECA) algorithm: the optimization of the kernel decomposition and the optimization of the Gaussian kernel parameter. KECA roughly reduces to a sorting of the importance of kernel eigenvectors by entropy instead of variance, as in the kernel principal components analysis. In this brief, we propose an extension of the KECA method, named optimized KECA (OKECA), that directly extracts the optimal features retaining most of the data entropy by means of compacting the information in very few features (often in just one or two). The proposed method produces features which have higher expressive power. In particular, it is based on the independent component analysis framework, and introduces an extra rotation to the eigen decomposition, which is optimized via gradient-ascent search. This maximum entropy preservation suggests that OKECA features are more efficient than KECA features for density estimation. In addition, a critical issue in both the methods is the selection of the kernel parameter, since it critically affects the resulting performance. Here, we analyze the most common kernel length-scale selection criteria. The results of both the methods are illustrated in different synthetic and real problems. Results show that OKECA returns projections with more expressive power than KECA, the most successful rule for estimating the kernel parameter is based on maximum likelihood, and OKECA is more robust to the selection of the length-scale parameter in kernel density estimation. Emma Izquierdo-Verdiguier, Valero Laparra, Robert Jenssen, Luis Gómez-Chova, Gustau Camps-Valls |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2016 | Predicting colorectal surgical complications using heterogeneous clinical data and kernel methods
Cristina Soguero-Ruíz, Kristian Hindberg, I. Mora-Jiménez, José Luis Rojo-Álvarez, Stein Olav Skrøvseth, Fred Godtliebsen, Kim Mortensen, Arthur Revhaug, Rolv-Ole Lindsetmo, Knut Magne Augestad, Robert Jenssen |
J. Biomed. Informatics | 11 |
| 2016 | Support Vector Feature Selection for Early Detection of Anastomosis Leakage From Bag-of-Words in Electronic Health RecordsabstractThe free text in electronic health records (EHRs) conveys a huge amount of clinical information about health state and patient history. Despite a rapidly growing literature on the use of machine learning techniques for extracting this information, little effort has been invested toward feature selection and the features' corresponding medical interpretation. In this study, we focus on the task of early detection of anastomosis leakage (AL), a severe complication after elective surgery for colorectal cancer (CRC) surgery, using free text extracted from EHRs. We use a bag-of-words model to investigate the potential for feature selection strategies. The purpose is earlier detection of AL and prediction of AL with data generated in the EHR before the actual complication occur. Due to the high dimensionality of the data, we derive feature selection strategies using the robust support vector machine linear maximum margin classifier, by investigating: 1) a simple statistical criterion (leave-one-out-based test); 2) an intensive-computation statistical criterion (Bootstrap resampling); and 3) an advanced statistical criterion (kernel entropy). Results reveal a discriminatory power for early detection of complications after CRC (sensitivity 100%; specificity 72%). These results can be used to develop prediction models, based on EHR data, that can support surgeons and patients in the preoperative decision making phase. Cristina Soguero-Ruíz, Kristian Hindberg, José Luis Rojo-Álvarez, Stein Olav Skrøvseth, Fred Godtliebsen, Kim Mortensen, Arthur Revhaug, Rolv-Ole Lindsetmo, Knut Magne Augestad, Robert Jenssen |
IEEE J. Biomed. Health Informatics | 10 |
| 2015 | Data-driven Temporal Prediction of Surgical Site Infection
Cristina Soguero-Ruíz, Fei Wang 0001, Robert Jenssen, Knut Magne Augestad, José Luis Rojo-Álvarez, I. Mora-Jiménez, Rolv-Ole Lindsetmo, Stein Olav Skrøvseth |
AMIA | 3 |
| 2015 | Sensitivity analysis of Gaussian processes for oceanic chlorophyll predictionabstractGaussian Process Regression (GPR) for machine learning has lately been successfully introduced for chlorophyll content mapping from remotely sensed data. The method provides a fast, stable and accurate prediction of biophysical parameters. However, since GPR is a non-linear kernel regression method, the relevance of the features are not accessible. In this paper, we introduce a probabilistic approach for feature sensitivity analysis (SA) of the GPR in order to reveal the relative importance of the features (bands) being used in the regression process. We evaluated the SA on GPR ocean chlorophyll content prediction. The method revealed the importance of the spectral bands, thus allowing the discrimination between Case-1 water and Case-2 water conditions. Katalin Blix, Gustau Camps-Valls, Robert Jenssen |
IGARSS | 3 |
| 2015 | Spectral clustering with the probabilistic cluster kernel
Emma Izquierdo-Verdiguier, Robert Jenssen, Luis Gómez-Chova, Gustau Camps-Valls |
Neurocomputing | 2 |
| 2014 | Information theoretic clustering using a k-nearest neighbors approach
Vidar Vikjord, Robert Jenssen |
Pattern Recognit. | 2 |
| 2013 | Mean Vector Component Analysis for Visualization and Clustering of Nonnegative DataabstractMean vector component analysis (MVCA) is introduced as a new method for visualization and clustering of nonnegative data. The method is based on dimensionality reduction by preserving the squared length, and implicitly also the direction, of the mean vector of the original data. The optimal mean vector preserving basis is obtained from the spectral decomposition of the inner-product matrix, and it is shown to capture clustering structure. MVCA corresponds to certain uncentered principal component analysis (PCA) axes. Unlike traditional PCA, these axes are in general not corresponding to the top eigenvalues. MVCA is shown to produce different visualizations and sometimes considerably improved clustering results for nonnegative data, compared with PCA. Robert Jenssen |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2011 | Kernel entropy component analysis in remote sensing data clusteringabstractThis paper proposes the kernel entropy component analysis (KECA) for clustering remote sensing data. The method generates nonlinear features that reveal structure related to the Renyi entropy of the input space data set. Unlike other kernel feature extraction methods, the top eigenvalues and eigenvectors of the kernel matrix are not necessarily chosen. Data are interestingly mapped with a distinct angular structure, which is exploited to derive a new angle-based spectral clustering algorithm based on the mapped data. An out-of-sample extension of the method is also presented to deal with test data. We focus on cloud screening from MERIS images. Several images are considered to account for the high variability of the problem. Good results show the suitability of the proposal. Luis Gómez-Chova, Robert Jenssen, Gustau Camps-Valls |
IGARSS | 2 |
| 2010 | Kernel Entropy Component AnalysisabstractWe introduce kernel entropy component analysis (kernel ECA) as a new method for data transformation and dimensionality reduction. Kernel ECA reveals structure relating to the Renyi entropy of the input space data set, estimated via a kernel matrix using Parzen windowing. This is achieved by projections onto a subset of entropy preserving kernel principal component analysis (kernel PCA) axes. This subset does not need, in general, to correspond to the top eigenvalues of the kernel matrix, in contrast to the dimensionality reduction using kernel PCA. We show that kernel ECA may produce strikingly different transformed data sets compared to kernel PCA, with a distinct angle-based structure. A new spectral clustering algorithm utilizing this structure is developed with positive results. Furthermore, kernel ECA is shown to be an useful alternative for pattern denoising. Robert Jenssen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | A new information theoretic analysis of sum-of-squared-error kernel clustering
Robert Jenssen, Torbjørn Eltoft |
Neurocomputing | 1 |
| 2008 | Mean shift spectral clustering
Umut Ozertem, Deniz Erdogmus, Robert Jenssen |
Pattern Recognit. | 3 |
| 2007 | Information cut for clustering using a gradient descent approach
Robert Jenssen, Deniz Erdogmus, Kenneth E. Hild II, José C. Príncipe, Torbjørn Eltoft |
Pattern Recognit. | 1 |
| 2006 | Information Theoretic Angle-Based Spectral Clustering: A Theoretical Analysis and an AlgorithmabstractRecent work has revealed a close connection between certain information theoretic divergence measures and properties of Mercer kernel feature spaces. Specifically, it has been proposed that an information theoretic measure may be used as a cost function for clustering in a kernel space, approximated by the spectral properties of the Laplacian matrix. In this paper we extend this result to other kernel matrices. We develop an algorithm for the actual clustering which is based on comparing angles between data points, and demonstrate that the proposed method performs equally good as a state-of-the art spectral clustering method. We point out some drawbacks of spectral clustering related to outliers, and suggest measures to be taken. Robert Jenssen, Deniz Erdogmus, José C. Príncipe |
IJCNN | 1 |
| 2006 | Kernel Maximum Entropy Data Transformation and an Enhanced Spectral Clustering AlgorithmabstractWe propose a new kernel-based data transformation technique. It is founded on the principle of maximum entropy (MaxEnt) preservation, hence named kernel MaxEnt. The key measure is Renyi's entropy estimated via Parzen windowing. We show that kernel MaxEnt is based on eigenvectors, and is in that sense similar to kernel PCA, but may produce strikingly different transformed data sets. An enhanced spectral clustering algorithm is proposed, by replacing kernel PCA by kernel MaxEnt as an intermediate step. This has a major impact on performance. Robert Jenssen, Torbjørn Eltoft, Mark A. Girolami, Deniz Erdogmus |
NIPS | 1 |
| 2006 | Spectral feature projections that maximize Shannon mutual information with class labels
Umut Ozertem, Deniz Erdogmus, Robert Jenssen |
Pattern Recognit. | 3 |
| 2005 | The Laplacian spectral classifierabstractWe develop a novel classifier in a kernel feature space defined by the eigenspectrum of the Laplacian data matrix. The classification cost function is derived from a distance measure between probability densities. The Laplacian data matrix is obtained based on a training set, while test data is mapped to the kernel space using the Nystrom routine. In that space, the test data is classified based on the angle between the test point and the training data class means. We illustrate the performance of the new classifier on synthetic and real data. Robert Jenssen, Deniz Erdogmus, José C. Príncipe, Torbjørn Eltoft |
ICASSP (5) | 1 |
| 2005 | An information-theoretic perspective to kernel independent components analysisabstractIn this paper, we investigate the intriguing relationship between information-theoretic learning (ITL), based on weighted Parzen window density estimator, and kernel-based learning algorithms. We prove the equivalence between kernel independent component analysis (kernel ICA) and the Cauchy-Schwartz (C-S) independence measure. This link gives a theoretical motivation for the selection of the Mercer kernel, based on density estimation. Demonstrating this equivalence requires introducing a weighted kernel density estimator, a modification of Parzen windowing. We also discuss the role of the weights in the weighted Parzen windowing and kernel ICA. Jianwu Xu, Deniz Erdogmus, Robert Jenssen, José C. Príncipe |
ICASSP (5) | 3 |
| 2004 | Information theoretic spectral clusteringabstractWe discuss a new information-theoretic framework for spectral clustering that is founded on the recently introduced information cut. A novel spectral clustering algorithm is proposed, where the clustering solution is given as a linearly weighted combination of certain top eigenvectors of the data affinity matrix. The information cut provides us with a theoretically well-defined graph-spectral cost function, and also establishes a close link between spectral clustering, and non-parametric density estimation. As a result, a natural criterion for creating the data affinity matrix is provided. We present preliminary clustering results to illustrate some of the properties of our algorithm, and we also make comparative remarks. Robert Jenssen, Torbjørn Eltoft, José C. Príncipe |
IJCNN | 1 |
| 2004 | The Laplacian PDF Distance: A Cost Function for Clustering in a Kernel Feature SpaceabstractA new distance measure between probability density functions (pdfs) is introduced, which we refer to as the Laplacian pdf dis- tance. The Laplacian pdf distance exhibits a remarkable connec- tion to Mercer kernel based learning theory via the Parzen window technique for density estimation. In a kernel feature space defined by the eigenspectrum of the Laplacian data matrix, this pdf dis- tance is shown to measure the cosine of the angle between cluster mean vectors. The Laplacian data matrix, and hence its eigenspec- trum, can be obtained automatically based on the data at hand, by optimal Parzen window selection. We show that the Laplacian pdf distance has an interesting interpretation as a risk function connected to the probability of error. 1 Introduction In recent years, spectral clustering methods, i.e. data partitioning based on the eigenspectrum of kernel matrices, have received a lot of attention [1, 2]. Some unresolved questions associated with these methods are for example that it is not always clear which cost function that is being optimized and that is not clear how to construct a proper kernel matrix. In this paper, we introduce a well-defined cost function for spectral clustering. This cost function is derived from a new information theoretic distance measure between cluster pdfs, named the Laplacian pdf distance. The information theoretic/spectral duality is established via the Parzen window methodology for density estimation. The resulting spectral clustering cost function measures the cosine of the angle between cluster mean vectors in a Mercer kernel feature space, where the feature space is determined by the eigenspectrum of the Laplacian matrix. A principled approach to spectral clustering would be to optimize this cost function in the feature space by assigning cluster memberships. Because of space limitations, we leave it to a future paper to present an actual clustering algorithm optimizing this cost function, and focus in this paper on the theoretical properties of the new measure. Corresponding author. Phone: (+47) 776 46493. Email: [email protected] An important by-product of the theory presented is that a method for learning the Mercer kernel matrix via optimal Parzen windowing is provided. This means that the Laplacian matrix, its eigenspectrum and hence the feature space mapping can be determined automatically. We illustrate this property by an example. We also show that the Laplacian pdf distance has an interesting relationship to the probability of error. In section 2, we briefly review kernel feature space theory. In section 3, we utilize the Parzen window technique for function approximation, in order to introduce the new Laplacian pdf distance and discuss some properties in sections 4 and 5. Section 6 concludes the paper. 2 Kernel Feature Spaces Mercer kernel-based learning algorithms [3] make use of the following idea: via a nonlinear mapping : Rd F, x (x) (1) the data x1, . . . , xN Rd is mapped into a potentially much higher dimensional feature space F. For a given learning problem one now considers the same algorithm in F instead of in Rd, that is, one works with (x1),...,(xN) F. Consider a symmetric kernel function k(x, y). If k : C C R is a continuous kernel of a positive integral operator in a Hilbert space L2(C) on a compact set C Rd, i.e. L2(C) : k(x,y)(x)(y)dxdy 0, (2) C then there exists a space F and a mapping : Rd F, such that by Mercer's theorem [4] NF k(x, y) = (x), (y) = ii(x)i(y), (3) i=1 where , denotes an inner product, the i's are the orthonormal eigenfunctions of the kernel and NF [3]. In this case (x) = [ 11(x), 22(x), . . . ]T , (4) can potentially be realized. In some cases, it may be desirable to realize this mapping. This issue has been addressed in [5]. Define the (N N) Gram matrix, K, also called the affinity, or kernel matrix, with elements Kij = k(xi, xj), i, j = 1, . . . , N . This matrix can be diagonalized as ET KE = , where the columns of E contains the eigenvectors of K and is a diagonal matrix containing the non-negative eigenvalues ~ 1, . . . , ~ N , ~ 1 ~N. In [5], it was shown that the eigenfunctions and eigenvalues of (4) can ~ be approximated as j j (xi) Neji, j , where e N ji denotes the ith element of the jth eigenvector. Hence, the mapping (4), can be approximated as (xi) [ ~1e1i,..., ~NeNi]T. (5) Thus, the mapping is based on the eigenspectrum of K. The feature space data set may be represented in matrix form as NN = [(x1), . . . , (xN )]. Hence, = 1 2 ET . It may be desirable to truncate the mapping (5) to C-dimensions. Thus, T only the C first rows of are kept, yielding ^ . It is well-known that ^ K = ^ ^ is the best rank-C approximation to K wrt. the Frobenius norm [6]. The most widely used Mercer kernel is the radial-basis-function (RBF) k(x, y) = exp -||x - y||2 . (6) 22 3 Function Approximation using Parzen Windowing Parzen windowing is a kernel-based density estimation method, where the resulting density estimate is continuous and differentiable provided that the selected kernel is continuous and differentiable [7]. Given a set of iid samples {x1,...,xN} drawn from the true density f (x), the Parzen window estimate for this distribution is [7] N ^ 1 f (x) = W N 2 (x, xi), (7) i=1 where W2 is the Parzen window, or kernel, and 2 controls the width of the kernel. The Parzen window must integrate to one, and is typically chosen to be a pdf itself with mean xi, such as the Gaussian kernel 1 W2 (x, xi) = exp , (8) d -||x - xi||2 (22) 2 22 which we will assume in the rest of this paper. In the conclusion, we briefly discuss the use of other kernels. Consider a function h(x) = v(x)f (x), for some function v(x). We propose to estimate h(x) by the following generalized Parzen estimator N ^ 1 h(x) = v(xi)W N 2 (x, xi). (9) i=1 This estimator is asymptotically unbiased, which can be shown as follows 1 N Ef v(xi)W N 2 (x, xi) = v(z)f (z)W2 (x, z)dz = [v(x)f (x)] W2(x), i=1 (10) where Ef () denotes expectation with respect to the density f(x). In the limit as N and (N) 0, we have lim [v(x)f (x)] W2(x) = v(x)f(x). (11) N (N )0 Of course, if v(x) = 1 x, then (9) is nothing but the traditional Parzen estimator of h(x) = f (x). The estimator (9) is also asymptotically consistent provided that the kernel width (N ) is annealed at a sufficiently slow rate. The proof will be presented in another paper. Many approaches have been proposed in order to optimally determine the size of the Parzen window, given a finite sample data set. A simple selection rule was proposed by Silverman [8], using the mean integrated square error (MISE) between the estimated and the actual pdf as the optimality metric: 1 d+4 opt = X 4N -1(2d + 1)-1 , (12) where d is the dimensionality of the data and 2 = d-1 , where are the X i Xii Xii diagonal elements of the sample covariance matrix. More advanced approximations to the MISE solution also exist. 4 The Laplacian PDF Distance Cost functions for clustering are often based on distance measures between pdfs. The goal is to assign memberships to the data patterns with respect to a set of clusters, such that the cost function is optimized. Assume that a data set consists of two clusters. Associate the probability density function p(x) with one of the clusters, and the density q(x) with the other cluster. Let f (x) be the overall probability density function of the data set. Now define the f -1 weighted inner product between p(x) and q(x) as p, q f p(x)q(x)f-1(x)dx. In such an inner product space, the Cauchy-Schwarz inequality holds, that is, p, q 2 q, q . Based on this discussion, an information theoretic distance f p, p f f measure between the two pdfs can be expressed as p, q D f L = - log 0. (13) p, p q, q f f We refer to this measure as the Laplacian pdf distance, for reasons that we discuss next. It can be seen that the distance DL is zero if and only if the two densities are equal. It is non-negative, and increases as the overlap between the two pdfs decreases. However, it does not obey the triangle inequality, and is thus not a distance measure in the strict mathematical sense. We will now show that the Laplacian pdf distance is also a cost function for clus- tering in a kernel feature space, using the generalized Parzen estimators discussed in the previous section. Since the logarithm is a monotonic function, we will derive the expression for the argument of the log in (13). This quantity will for simplicity be denoted by the letter "L" in equations. Assume that we have available the iid data points {xi}, i = 1,...,N1, drawn from p(x), which is the density of cluster C1, and the iid {xj}, j = 1, . . ., N2, drawn from q(x), the density of C2. Let h(x) = f - 12 (x)p(x) and g(x) = f - 12 (x)q(x). Hence, we may write h(x)g(x)dx L = . (14) h2(x)dx g2(x)dx We estimate h(x) and g(x) by the generalized Parzen kernel estimators, as follows N1 N2 ^ 1 1 h(x) = f - 12 (xi)W f - 12 (xj )W N 2 (x, xi ), ^ g(x) = 2 (x, xj ). (15) 1 N2 i=1 j=1 The approach taken, is to substitute these estimators into (14), to obtain N N 1 1 1 2 h(x)g(x)dx f - 12 (xi)W f - 12 (xj )W N 2 (x, xi ) 2 (x, xj ) 1 N2 i=1 j=1 N 1 1 ,N2 = f - 12 (xi)f - 12 (xj ) W N 2 (x, xi )W2 (x, xj )dx 1N2 i,j=1 N 1 1 ,N2 = f - 12 (xi)f - 12 (xj )W N 22 (xi, xj ), (16) 1N2 i,j=1 where in the last step, the convolution theorem for Gaussians has been employed. Similarly, we have N 1 1 ,N1 h2(x)dx f - 12 (xi)f - 12 (xi )W N 2 22 (xi, xi ), (17) 1 i,i =1 N 1 2 ,N2 g2(x)dx f - 12 (xj)f - 12 (xj )W N 2 22 (xj , xj ). (18) 2 j,j =1 Now we define the matrix Kf , such that Kf = K ij f (xi, xj ) = f - 1 2 (xi)f - 12 (xj )K(xi, xj ), (19) where K(xi, xj ) = W22 (xi, xj) for i, j = 1, . . . , N and N = N1 + N2. As a consequence, (14) can be re-written as follows N1,N2 Kf (xi, xj) L = i,j=1 (20) N1,N1 K K i,i =1 f (xi, xi ) N2,N2 j,j =1 f (xj , xj ) The key point of this paper, is to note that the matrix K = Kij = K(xi, xj), i, j = 1, . . . , N , is the data affinity matrix, and that K(xi, xj) is a Gaussian RBF kernel function. Hence, it is also a kernel function that satisfies Mercer's theorem. Since K(xi, xj) satisfies Mercer's theorem, the following by definition holds [4]. For any set of examples {x1,...,xN} and any set of real numbers 1,...,N N N ijK(xi, xj) 0, (21) i=1 j=1 in analogy to (3). Moreover, this means that N N N N ijf - 12 (xi)f - 12 (xj )K(xi, xj) = ijKf (xi, xj) 0, (22) i=1 j=1 i=1 j=1 hence Kf (xi, xj ) is also a Mercer kernel. Now, it is readily observed that the Laplacian pdf distance can be analyzed in terms of inner products in a Mercer kernel-based Hilbert feature space, since Kf (xi, xj) = f (xi), f (xj) . Consequently, (20) can be written as follows N1,N2 f (xi), f (xj) L = i,j=1 N1,N1 i,i =1 f (xi), f (xi ) N2,N2 j,j =1 f (xj ), f (xj ) 1 N1 N2 N i=1 f (xi ), 1 N j=1 f (xj ) = 1 2 1 N1 N1 N2 N2 N f (xi), 1 f (xi ) 1 f (xj ), 1 f (xj ) 1 i=1 N1 i =1 N2 j=1 N2 j =1 m1 , m2 = f f = cos (m , m ), (23) ||m 1f 2f 1f ||||m2f || where m Ni i = 1 f N f (xl), i = 1, 2, that is, the sample mean of the ith cluster i l=1 in feature space. This is a very interesting result. We started out with a distance measure between densities in the input space. By utilizing the Parzen window method, this distance measure turned out to have an equivalent expression as a measure of the distance between two clusters of data points in a Mercer kernel feature space. In the feature space, the distance that is measured is the cosine of the angle between the cluster mean vectors. The actual mapping of a data point to the kernel feature space is given by the eigendecomposition of Kf , via (5). Let us examine this mapping in more detail. 1 Note that f 2 (xi) can be estimated from the data by the traditional Parzen pdf estimator as follows N 1 1 f 2 (xi) = W (xi, xl) = di. (24) N 2 f l=1 Define the matrix D = diag(d1, . . . , dN ). Then Kf can be expressed as Kf = D- 12 KD- 12 . (25) Quite interestingly, for 2 = 22, this is in fact the Laplacian data matrix. 1 f The above discussion explicitly connects the Parzen kernel and the Mercer kernel. Moreover, automatic procedures exist in the density estimation literature to opti- mally determine the Parzen kernel given a data set. Thus, the Mercer kernel is also determined by the same procedure. Therefore, the mapping by the Laplacian matrix to the kernel feature space can also be determined automatically. We regard this as a significant result in the kernel based learning theory. As an example, consider Fig. 1 (a) which shows a data set consisting of a ring with a dense cluster in the middle. The MISE kernel size is opt = 0.16, and the Parzen pdf estimate is shown in Fig. 1 (b). The data mapping given by the corresponding Laplacian matrix is shown in Fig. 1 (c) (truncated to two dimensions for visualization purposes). It can be seen that the data is distributed along two lines radially from the origin, indicating that clustering based on the angular measure we have derived makes sense. The above analysis can easily be extended to any number of pdfs/clusters. In the C-cluster case, we define the Laplacian pdf distance as C-1 pi, pj L = f . (26) i=1 j=i C pi, pi p f j , pj f In the kernel feature space, (26), corresponds to all cluster mean vectors being pairwise as orthogonal to each other as possible, for all possible unique pairs. 4.1 Connection to the Ng et al. [2] algorithm Recently, Ng et al. [2] proposed to map the input data to a feature space determined by the eigenvectors corresponding to the C largest eigenvalues of the Laplacian ma- trix. In that space, the data was normalized to unit norm and clustered by the C-means algorithm. We have shown that the Laplacian pdf distance provides a 1It is a bit imprecise to refer to Kf as the Laplacian matrix, as readers familiar with spectral graph theory may recognize, since the definition of the Laplacian matrix is L = I - Kf . However, replacing Kf by L does not change the eigenvectors, it only changes the eigenvalues from i to 1 - i. 0 0 (a) Data set (b) Parzen pdf estimate (c) Feature space data Figure 1: The kernel size is automatically determined (MISE), yielding the Parzen estimate (b) with the corresponding feature space mapping (c). clustering cost function, measuring the cosine of the angle between cluster means, in a related kernel feature space, which in our case can be determined automati- cally. A more principled approach to clustering than that taken by Ng et al. is to optimize (23) in the feature space, instead of using C-means. However, because of the normalization of the data in the feature space, C-means can be interpreted as clustering the data based on an angular measure. This may explain some of the success of the Ng et al. algorithm; it achieves more or less the same goal as cluster- ing based on the Laplacian distance would be expected to do. We will investigate this claim in our future work. Note that we in our framework may choose to use only the C largest eigenvalues/eigenvectors in the mapping, as discussed in section 2. Since we incorporate the eigenvalues in the mapping, in contrast to Ng et al., the actual mapping will in general be different in the two cases. 5 The Laplacian PDF distance as a risk function We now give an analysis of the Laplacian pdf distance that may further motivate its use as a clustering cost function. Consider again the two cluster case. The overall data distribution can be expressed as f (x) = P1p(x) + P2q(x), were Pi, i = 1, 2, are the priors. Assume that the two clusters are well separated, such that for xi C1, f (xi) P1p(xi), while for xi C2, f(xi) P2q(xi). Let us examine the numerator of (14) in this case. It can be approximated as p(x)q(x) dx f (x) p(x)q(x) p(x)q(x) 1 1 dx + dx q(x)dx + p(x)dx. (27) C f (x) f (x) P1 P2 1 C2 C1 C2 By performing a similar calculation for the denominator of (14), it can be shown to be approximately equal to 1 . Hence, the Laplacian pdf distance can be written P1P1 as a risk function, given by 1 1 L P1P2 q(x)dx + p(x)dx . (28) P1 C P2 1 C2 Note that if P1 = P2 = 1 , then L = 2P 2 e, where Pe is the probability of error when assigning data points to the two clusters, that is Pe = P1 q(x)dx + P2 p(x)dx. (29) C1 C2 Thus, in this case, minimizing L is equivalent to minimizing Pe. However, in the case that P1 = P2, (28) has an even more interesting interpretation. In that situation, it can be seen that the two integrals in the expressions (28) and (29) are weighted exactly oppositely. For example, if P1 is close to one, L p(x)dx, while P C e 2 q(x)dx. Thus, the Laplacian pdf distance emphasizes to cluster the most un- C1 likely data points correctly. In many real world applications, this property may be crucial. For example, in medical applications, the most important points to classify correctly are often the least probable, such as detecting some rare disease in a group of patients. 6 Conclusions We have introduced a new pdf distance measure that we refer to as the Laplacian pdf distance, and we have shown that it is in fact a clustering cost function in a kernel feature space determined by the eigenspectrum of the Laplacian data matrix. In our exposition, the Mercer kernel and the Parzen kernel is equivalent, making it possible to determine the Mercer kernel based on automatic selection procedures for the Parzen kernel. Hence, the Laplacian data matrix and its eigenspectrum can be determined automatically too. We have shown that the new pdf distance has an interesting property as a risk function. The results we have derived can only be obtained analytically using Gaussian ker- nels. The same results may be obtained using other Mercer kernels, but it requires an additional approximation wrt. the expectation operator. This discussion is left for future work. Acknowledgments. This work was partially supported by NSF grant ECS- 0300340. Robert Jenssen, Deniz Erdogmus, José C. Príncipe, Torbjørn Eltoft |
NIPS | 1 |
| 2003 | Accurate initialization of neural network weights by backpropagation of the desired responseabstractProper initialization of neural networks is critical for a successful training of its weights. Many methods have been proposed to achieve this, including heuristic least squares approaches. In this paper, inspired by these previous attempts to train (or initialize) neural networks, we formulate a mathematically sound algorithm based on backpropagating the desired output through the layers of a multilayer perceptron. The approach is accurate up to local first order approximations of the nonlinearities. It is shown to provide successful weight initialization for many data sets by Monte Carlo experiments. Deniz Erdogmus, Oscar Fontenla-Romero, José C. Príncipe, Amparo Alonso-Betanzos, Enrique F. Castillo, Robert Jenssen |
IJCNN | 6 |
| 2003 | Clustering using Renyi's entropyabstractWe propose a new clustering algorithm using Renyi's entropy as our similarity metric. The main idea is to assign a data pattern to the cluster, which among all possible clusters, increases its within-cluster entropy the least, upon inclusion of the pattern. We refer to this procedure as differential entropy clustering. Not knowing the true number of clusters in advance, initially a number of small clusters are "seeded" randomly in the data set, labeling a small subset of the data. Thereafter all remaining patterns are labeled by differential entropy clustering. Subsequently, we identify the "worst cluster" by a quantity we name as the between-cluster entropy. Its members are re-clustered, again by differential entropy clustering, reducing the overall number of clusters by one. This procedure is repeated until only two clusters remain. At each step we store the current labels, thus producing a hierarchy of clusters. The between-cluster entropy also enables us to select our final set of clusters in other cluster hierarchy. We demonstrate the clustering algorithm when applied both to artificially created data sets and a real data set. Robert Jenssen, Kenneth E. Hild II, Deniz Erdogmus, José C. Príncipe, Torbjørn Eltoft |
IJCNN | 1 |
| 2003 | Independent component analysis for texture segmentation
Robert Jenssen, Torbjørn Eltoft |
Pattern Recognit. | 1 |