VLDB 2026 Research / reviewers in the wild / expert
Klaus-Robert Müller
dblp:m/KRMuller
· DBLP profile ↗
199ranked-venue papers
3as first author
40since 2021 · last 2026
0000-0002-3861-7685ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 159 · 3 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 9 since 2021Human-computer interaction and ubiquitous computing · 5Databases, data management, data science and information retrieval · 3 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond attention heatmaps: How to get better explanations for multiple instance learning models in histopathologyabstractMultiple instance learning (MIL) has enabled substantial progress in computational histopathology, where a large amount of patches from gigapixel whole slide images are aggregated into slide-level predictions. Heatmaps are widely used to validate MIL models and to discover tissue biomarkers. Yet, the validity of these heatmaps has barely been investigated. In this work, we introduce a general framework for evaluating the quality of MIL heatmaps without requiring additional labels. We conduct a large-scale benchmark experiment to assess six explanation methods across histopathology task types (classification, regression, survival), MIL model architectures (Attention-, Transformer-, Mamba-based), and patch encoder backbones (UNI2, Virchow2). Our results show that explanation quality mostly depends on MIL model architecture and task type, with perturbation ("Single"), layer-wise relevance propagation (LRP), and integrated gradients (IG) consistently outperforming attention-based and gradient-based saliency heatmaps, which often fail to reflect model decision mechanisms. We further demonstrate the advanced capabilities of the best-performing explanation methods: (i) We provide a proof-of-concept that MIL heatmaps of a bulk gene expression prediction model can be correlated with spatial transcriptomics for biological validation, and (ii) showcase the discovery of distinct model strategies for predicting human papillomavirus (HPV) infection from head and neck cancer slides. Our work highlights the importance of validating MIL heatmaps and establishes that improved explainability can enable more reliable model validation and yield biological insights, making a case for a broader adoption of explainable AI in digital pathology. Our code is provided in a public GitHub repository: https://github.com/bifold-pathomics/xMIL. Mina Jamshidi Idaji, Julius Hense, Tom Neuhäuser, Augustin Krause, Yanqing Luo, Oliver Eberle, Thomas Schnake, Laure Ciernik, Farnoush Rezaei Jafari, Reza Vahidimajd, Jonas Dippel, Christoph Walz, Frederick Klauschen, Andreas Mock, Klaus-Robert Müller |
Medical Image Anal. | 15 |
| 2026 | Towards desiderata-driven design of visual counterfactual explainersabstractVisual counterfactual explainers (VCEs) are a straightforward and promising approach to enhancing the transparency of image classifiers. VCEs complement other types of explanations, such as feature attribution, by revealing the specific data transformations to which a machine learning model responds most strongly. In this paper, we argue that existing VCEs tend to focus too narrowly on optimizing sample quality or change minimality; they do not consider the more holistic desiderata for an explanation, such as fidelity, understandability, and sufficiency. To address this shortcoming, we explore new mechanisms for counterfactual generation and investigate how they can help fulfill these desiderata. We combine these mechanisms into a novel ‘smooth counterfactual explorer’ (SCE) algorithm and demonstrate its effectiveness through systematic evaluations on synthetic and real data. Sidney Bender, Jan Herrmann, Klaus-Robert Müller, Grégoire Montavon |
Pattern Recognit. | 3 |
| 2026 | Fast and accurate explanations of distance-based classifiers by uncovering latent explanatory structuresabstractDistance-based classifiers, such as k-nearest neighbors and support vector machines, continue to be a workhorse of machine learning, widely used in science and industry. In practice, to derive insights from these models, it is also important to ensure that their predictions are explainable. While the field of Explainable AI has supplied methods that are in principle applicable to any model, it has also emphasized the usefulness of latent structures (e.g. the sequence of layers in a neural network) to produce explanations. In this paper, we contribute by uncovering a hidden neural network structure in distance-based classifiers (consisting of linear detection units combined with nonlinear pooling layers) upon which Explainable AI techniques such as layer-wise relevance propagation (LRP) become applicable. Through quantitative evaluations, we demonstrate the advantage of our novel explanation approach over several baselines. We also show the overall usefulness of explaining distance-based models through two practical use cases. Florian Bley, Jacob R. Kauffmann, Simon León Krug, Klaus-Robert Müller, Grégoire Montavon |
Pattern Recognit. | 4 |
| 2026 | Network Measure-Enriched GNNs: A New Framework for Power Grid Stability PredictionabstractFacing climate change, the transformation to renewable energy poses stability challenges for power grids due to their reduced inertia and increased decentralization. Traditional dynamic stability assessments, crucial for safe grid operation with higher renewable shares, are computationally expensive and unsuitable for large-scale grids in the real world. Although multiple proofs in the network science have shown that network measures, which quantify the structural characteristics of networked dynamical systems, have the potential to facilitate basin stability prediction, no studies to date have demonstrated their ability to efficiently generalize to real-world grids. With recent breakthroughs in Graph Neural Networks (GNNs), we are surprised to find that there is still a lack of a common foundation about: Whether network measures can enhance GNNs' capability to predict dynamic stability and how they might help GNNs generalize to realistic grid topologies. In this paper, we conduct, for the first time, a comprehensive analysis of 48 network measures in GNN-based stability assessments, introducing two strategies for their integration into the GNN framework. We uncover that prioritizing measures with consistent distributions across different grids as the input or regarding measures as auxiliary supervised information improves the model's generalization ability to realistic grid topologies, even when models trained on only 20-node synthetic datasets are used. Our empirical results demonstrate a significant enhancement in model generalizability, increasing the$R^{2}$perforsmance from 66% to 83%. When evaluating the probabilistic stability indices on the realistic Texan grid model, GNNs reduce the time needed from 28,950 hours (Monte Carlo sampling) to just 0.06 seconds. Junyou Zhu, Christian Nauck, Michael Lindner, Langzhou He, Philip S. Yu, Klaus-Robert Müller, Jürgen Kurths, Frank Hellmann |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2026 | Enhancing Brain Source Reconstruction by Initializing 3-D Neural Networks With Physical Inverse SolutionsabstractReconstructing brain sources is a fundamental challenge in neuroscience, crucial for understanding brain function and dysfunction. Electroencephalography (EEG) signals have a high temporal resolution. However, identifying the correct spatial location of brain sources from these signals remains difficult due to the ill-posed structure of the problem. Traditional methods predominantly rely on manually crafted priors, missing the flexibility of data-driven learning, while recent deep learning approaches focus on end-to-end learning, typically using the physical information of the forward model only for generating training data. We propose the novel hybrid method 3D-PIUNet for EEG source localization that effectively integrates the strengths of traditional and deep learning techniques. 3D-PIUNet starts from an initial physics-informed estimate by using the pseudo inverse to map from measurements to source space. Secondly, by viewing the brain as a 3D volume, we use a 3D convolutional U-Net to capture spatial dependencies and refine the solution according to the learned data prior. Training the model relies on simulated pseudo-realistic brain source data, covering different source distributions. Trained on this data, our model significantly improves spatial accuracy, demonstrating superior performance over both traditional and end-to-end data-driven methods. Additionally, we validate our findings with real EEG data from a visual task, where 3D-PIUNet successfully identifies the visual cortex and reconstructs the expected temporal behavior, thereby showcasing its practical applicability. Marco Morik, Ali Hashemi 0002, Klaus-Robert Müller, Stefan Haufe, Shinichi Nakajima |
IEEE Trans. Medical Imaging | 3 |
| 2025 | MeDi: Metadata-Guided Diffusion Models for Mitigating Biases in Tumor ClassificationabstractDeep learning models have made significant advances in histological prediction tasks in recent years. However, for adaptation in clinical practice, their lack of robustness to varying conditions such as staining, scanner, hospital, and demographics is still a limiting factor: if trained on overrepresented subpopulations, models regularly struggle with less frequent patterns, leading to shortcut learning and biased predictions. Large-scale foundation models have not fully eliminated this issue. Therefore, we propose a novel approach explicitly modeling such metadata into a Me tadata-guided generative Di ffusion model framework (MeDi). MeDi allows for a targeted augmentation of underrepresented subpopulations with synthetic data, which balances limited training data and mitigates biases in downstream models. We experimentally show that MeDi generates high-quality histopathology images for unseen subpopulations in TCGA, boosts the overall fidelity of the generated images, and enables improvements in performance for downstream classifiers on datasets with subpopulation shifts. Our work is a proof-of-concept towards better mitigating data biases with generative models. David Jacob Drexlin, Jonas Dippel, Julius Hense, Niklas Prenißl, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller |
MICCAI (14) | 7 |
| 2025 | Manipulating Feature Visualizations with Gradient SlingshotsabstractFeature Visualization (FV) is a widely used technique for interpreting concepts learned by Deep Neural Networks (DNNs), which synthesizes input patterns that maximally activate a given feature. Despite its popularity, the trustworthiness of FV explanations has received limited attention. We introduce Gradient Slingshots, a novel method that enables FV manipulation without modifying model architecture or significantly degrading performance. By shaping new trajectories in off-distribution regions of a feature's activation landscape, we coerce the optimization process to converge to a predefined visualization. We evaluate our approach on several DNN architectures, demonstrating its ability to replace faithful FVs with arbitrary targets. These results expose a critical vulnerability: auditors relying solely on FV may accept entirely fabricated explanations. To mitigate this risk, we propose a straightforward defense and quantitatively demonstrate its effectiveness. Dilyara Bareeva, Marina M.-C. Höhne, Alexander Warnecke, Lukas Pirch, Klaus-Robert Müller, Konrad Rieck, Sebastian Lapuschkin, Kirill Bykov |
NeurIPS | 5 |
| 2025 | Sampling 3D Molecular Conformers with Diffusion TransformersabstractDiffusion Transformers (DiTs) have demonstrated strong performance in generative modeling, particularly in image synthesis, making them a compelling choice for molecular conformer generation. However, applying DiTs to molecules introduces novel challenges, such as integrating discrete molecular graph information with continuous 3D geometry, handling Euclidean symmetries, and designing conditioning mechanisms that generalize across molecules of varying sizes and structures. We propose DiTMC, a framework that adapts DiTs to address these challenges through a modular architecture that separates the processing of 3D coordinates from conditioning on atomic connectivity. To this end, we introduce two complementary graph-based conditioning strategies that integrate seamlessly with the DiT architecture. These are combined with different attention mechanisms, including both standard non-equivariant and SO(3)-equivariant formulations, enabling flexible control over the trade-off between between accuracy and computational efficiency.
Experiments on standard conformer generation benchmarks (GEOM-QM9, -DRUGS, -XL) demonstrate that DiTMC achieves state-of-the-art precision and physical validity. Our results highlight how architectural choices and symmetry priors affect sample quality and efficiency, suggesting promising directions for large-scale generative modeling of molecular structures. Code is available at [https://github.com/ML4MolSim/dit_mc](https://github.com/ML4MolSim/dit_mc). J. Thorben Frank, Winfried Ripken, Gregor Lied, Klaus-Robert Müller, Oliver T. Unke, Stefan Chmiela |
NeurIPS | 4 |
| 2025 | Smoothed Differentiation Efficiently Mitigates Shattered Gradients in ExplanationsabstractExplaining complex machine learning models is a fundamental challenge when developing safe and trustworthy deep learning applications. To date, a broad selection of explainable AI (XAI) algorithms exist. One popular choice is SmoothGrad, which has been conceived to alleviate the well-known shattered gradient problem by smoothing gradients through convolution. SmoothGrad proposes to solve this high-dimensional convolution integral by sampling -- typically approximating the convolution with limited precision. Higher numbers of samples would amount to higher precision in approximating the convolution but also to higher computing demand, therefore in practice only few samples are used in SmoothGrad. In this work we propose a well founded novel method _SmoothDiff_ to resolve this tradeoff yielding a _speedup of over two orders of magnitude_. Specifically, _SmoothDiff_ leverages automatic differentiation to decompose the expected values of Jacobians across a network architecture, directly targeting only the non-linearities responsible for shattered gradients and making it easy to implement. We demonstrate SmoothDiff's excellent speed and performance in a number of experiments and benchmarks. Thus, SmoothDiff greatly enhances the usability (quality and speed) of SmoothGrad -- a popular workhorse of XAI. Adrian Hill, Neal McKee, Johannes Maeß, Stefan Blücher, Klaus-Robert Müller |
NeurIPS | 5 |
| 2025 | Parameterization of intraoperative human microelectrode recordings: Linking action potential morphology to brain anatomyabstractDeep brain stimulation (DBS) is a targeted manipulation of brain circuitry to treat neurological and neuropsychiatric conditions. Optimal DBS lead placement is essential for treatment efficacy. Current targeting practice is based on preoperative and intraoperative brain imaging, intraoperative electrophysiology, and stimulation mapping. Electrophysiological mapping using extracellular microelectrode recordings aids in identifying functional subdomains, anatomical boundaries, and disease-correlated physiology. The shape of single-unit action potentials may differ due to different biophysical properties between cell-types and brain regions. Here, we describe a technique to parameterize the structure and duration of sorted spike units using a novel algorithmic approach based on canonical response parameterization, and illustrate how it may be used on DBS microelectrode recordings. Isolated spike shapes are parameterized then compared using a spike similarity metric and grouped by hierarchical clustering. When spike morphology is associated with anatomy, we find regional clustering in the human globus pallidus. This method is widely applicable for spike removal and single-unit characterization and could be integrated into intraoperative array-based technologies to enhance targeting and clinical outcomes in DBS lead placement. Matthew R. Baker, Bryan T. Klassen, Michael A. Jensen, Gabriela Ojeda Valencia, Hossein Heydari, Nuri Firat Ince, Klaus-Robert Müller, Kai J. Miller |
PLoS Comput. Biol. | 7 |
| 2025 | Guest Editorial: Special Issue on Information Theoretic Methods for the Generalization, Robustness, and Interpretability of Machine Learning
Badong Chen, Shujian Yu, Robert Jenssen, José C. Príncipe, Klaus-Robert Müller |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Set Learning for Accurate and Calibrated ModelsabstractModel overconfidence and poor calibration are common in machine learning and difficult to account for when applying standard empirical risk minimization. In this work, we propose a novel method to alleviate these problems that we call odd-$k$-out learning (OKO), which minimizes the cross-entropy error for sets rather than for single examples. This naturally allows the model to capture correlations across data examples and achieves both better accuracy and calibration, especially in limited training data and class-imbalanced regimes. Perhaps surprisingly, OKO often yields better calibration even when training with hard labels and dropping any additional calibration parameter tuning, such as temperature scaling. We demonstrate this in extensive experimental analyses and provide a mathematical theory to interpret our findings. We emphasize that OKO is a general framework that can be easily adapted to many settings and a trained model can be applied to single examples at inference time, without significant run-time overhead or architecture changes. Lukas Muttenthaler, Robert A. Vandermeulen, Qiuyi Zhang 0001, Thomas Unterthiner, Klaus-Robert Müller |
ICLR | 5 |
| 2024 | xMIL: Insightful Explanations for Multiple Instance Learning in HistopathologyabstractMultiple instance learning (MIL) is an effective and widely used approach for weakly supervised machine learning. In histopathology, MIL models have achieved remarkable success in tasks like tumor detection, biomarker prediction, and outcome prognostication. However, MIL explanation methods are still lagging behind, as they are limited to small bag sizes or disregard instance interactions. We revisit MIL through the lens of explainable AI (XAI) and introduce xMIL, a refined framework with more general assumptions. We demonstrate how to obtain improved MIL explanations using layer-wise relevance propagation (LRP) and conduct extensive evaluation experiments on three toy settings and four real-world histopathology datasets. Our approach consistently outperforms previous explanation attempts with particularly improved faithfulness scores on challenging biomarker prediction tasks. Finally, we showcase how xMIL explanations enable pathologists to extract insights from MIL models, representing a significant advance for knowledge discovery and model debugging in digital histopathology. Julius Hense, Mina Jamshidi Idaji, Oliver Eberle, Thomas Schnake, Jonas Dippel, Laure Ciernik, Oliver Buchstab, Andreas Mock, Frederick Klauschen, Klaus-Robert Müller |
NeurIPS | 10 |
| 2024 | MambaLRP: Explaining Selective State Space Sequence ModelsabstractRecent sequence modeling approaches using selective state space sequence models, referred to as Mamba models, have seen a surge of interest. These models allow efficient processing of long sequences in linear time and are rapidly being adopted in a wide range of applications such as language modeling, demonstrating promising performance. To foster their reliable use in real-world scenarios, it is crucial to augment their transparency. Our work bridges this critical gap by bringing explainability, particularly Layer-wise Relevance Propagation (LRP), to the Mamba architecture. Guided by the axiom of relevance conservation, we identify specific components in the Mamba architecture, which cause unfaithful explanations. To remedy this issue, we propose MambaLRP, a novel algorithm within the LRP framework, which ensures a more stable and reliable relevance propagation through these components. Our proposed method is theoretically sound and excels in achieving state-of-the-art explanation performance across a diverse range of models and datasets. Moreover, MambaLRP facilitates a deeper inspection of Mamba architectures, uncovering various biases and evaluating their significance. It also enables the analysis of previous speculations regarding the long-range capabilities of Mamba models. Farnoush Rezaei Jafari, Grégoire Montavon, Klaus-Robert Müller, Oliver Eberle |
NeurIPS | 3 |
| 2024 | Disentangled Explanations of Neural Network Predictions by Finding Relevant SubspacesabstractExplainable AI aims to overcome the black-box nature of complex ML models like neural networks by generating explanations for their predictions. Explanations often take the form of a heatmap identifying input features (e.g. pixels) that are relevant to the model's decision. These explanations, however, entangle the potentially multiple factors that enter into the overall complex decision strategy. We propose to disentangle explanations by extracting at some intermediate layer of a neural network, subspaces that capture the multiple and distinct activation patterns (e.g. visual concepts) that are relevant to the prediction. To automatically extract these subspaces, we propose two new analyses, extending principles found in PCA or ICA to explanations. These novel analyses, which we call principal relevant component analysis (PRCA) and disentangled relevant subspace analysis (DRSA), maximize relevance instead of e.g. variance or kurtosis. This allows for a much stronger focus of the analysis on what the ML model actually uses for predicting, ignoring activations or concepts to which the model is invariant. Our approach is general enough to work alongside common attribution techniques such as Shapley Value, Integrated Gradients, or LRP. Our proposed methods show to be practically useful and compare favorably to the state of the art as demonstrated on benchmarks and three use cases. Pattarawat Chormai, Jan Herrmann, Klaus-Robert Müller, Grégoire Montavon |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Diffeomorphic Counterfactuals With Generative ModelsabstractCounterfactuals can explain classification decisions of neural networks in a human interpretable way. We propose a simple but effective method to generate such counterfactuals. More specifically, we perform a suitable diffeomorphic coordinate transformation and then perform gradient ascent in these coordinates to find counterfactuals which are classified with great confidence as a specified target class. We propose two methods to leverage generative models to construct such suitable coordinate systems that are either exactly or approximately diffeomorphic. We analyze the generation process theoretically using Riemannian differential geometry and validate the quality of the generated counterfactuals using various qualitative and quantitative measures. Ann-Kathrin Dombrowski, Jan E. Gerken, Klaus-Robert Müller, Pan Kessel |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Joint Learning of Full-Structure Noise in Hierarchical Bayesian Regression ModelsabstractWe consider the reconstruction of brain activity from electroencephalography (EEG). This inverse problem can be formulated as a linear regression with independent Gaussian scale mixture priors for both the source and noise components. Crucial factors influencing the accuracy of the source estimation are not only the noise level but also its correlation structure, but existing approaches have not addressed the estimation of noise covariance matrices with full structure. To address this shortcoming, we develop hierarchical Bayesian (type-II maximum likelihood) models for observations with latent variables for source and noise, which are estimated jointly from data. As an extension to classical sparse Bayesian learning (SBL), where across-sensor observations are assumed to be independent and identically distributed, we consider Gaussian noise with full covariance structure. Using the majorization-maximization framework and Riemannian geometry, we derive an efficient algorithm for updating the noise covariance along the manifold of positive definite matrices. We demonstrate that our algorithm has guaranteed and fast convergence and validate it in simulations and with real MEG data. Our results demonstrate that the novel framework significantly improves upon state-of-the-art techniques in the real-world scenario where the noise is indeed non-diagonal and full-structured. Our method has applications in many domains beyond biomagnetic inverse problems. Ali Hashemi 0002, Yijing Gao, Sanjay Ghosh, Klaus-Robert Müller, Srikantan S. Nagarajan, Stefan Haufe |
IEEE Trans. Medical Imaging | 5 |
| 2024 | From Clustering to Cluster Explanations via Neural NetworksabstractA recent trend in machine learning has been to enrich learned models with the ability to explain their own predictions. The emerging field of explainable AI (XAI) has so far mainly focused on supervised learning, in particular, deep neural network classifiers. In many practical problems, however, the label information is not given and the goal is instead to discover the underlying structure of the data, for example, its clusters. While powerful methods exist for extracting the cluster structure in data, they typically do not answer the question why a certain data point has been assigned to a given cluster. We propose a new framework that can, for the first time, explain cluster assignments in terms of input features in an efficient and reliable manner. It is based on the novel insight that clustering models can be rewritten as neural networks-or "neuralized." Cluster predictions of the obtained networks can then be quickly and accurately attributed to the input features. Several showcases demonstrate the ability of our method to assess the quality of learned clusters and to extract novel insights from the analyzed data and representations. Jacob R. Kauffmann, Malte Esders, Lukas Ruff, Grégoire Montavon, Wojciech Samek, Klaus-Robert Müller |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2023 | Shortcomings of Top-Down Randomization-Based Sanity Checks for Evaluations of Deep Neural Network ExplanationsabstractWhile the evaluation of explanations is an important step towards trustworthy models, it needs to be done carefully, and the employed metrics need to be well-understood. Specifically model randomization testing can be overinterpreted if regarded as a primary criterion for selecting or discarding explanation methods. To address shortcomings of this test, we start by observing an experimental gap in the ranking of explanation methods between randomization-based sanity checks [1] and model output faithfulness measures (e.g. [20]). We identify limitations of model-randomization-based sanity checks for the purpose of evaluating explanations. Firstly, we show that uninformative attribution maps created with zero pixel-wise covariance easily achieve high scores in this type of checks. Secondly, we show that top-down model randomization preserves scales of forward pass activations with high probability. That is, channels with large activations have a high probility to contribute strongly to the output, even after randomization of the network on top of them. Hence, explanations after randomization can only be expected to differ to a certain extent. This explains the observed experimental gap. In summary, these results demonstrate the inadequacy of model-randomization-based sanity checks as a criterion to rank attribution methods. Alexander Binder, Leander Weber, Sebastian Lapuschkin, Grégoire Montavon, Klaus-Robert Müller, Wojciech Samek |
CVPR | 5 |
| 2023 | Relevant Walk Search for Explaining Graph Neural NetworksabstractGraph Neural Networks (GNNs) have become important machine learning tools for graph analysis, and its explainability is crucial for safety, fairness, and robustness. Layer-wise relevance propagation for GNNs (GNN-LRP) evaluates the relevance of walks to reveal important information flows in the network, and provides higher-order explanations, which have been shown to be superior to the lower-order, i.e., node-/edge-level, explanations. However, identifying relevant walks by GNN-LRP requires exponential computational complexity with respect to the network depth, which we will remedy in this paper. Specifically, we propose polynomial-time algorithms for finding top-$K$ relevant walks, which drastically reduces the computation and thus increases the applicability of GNN-LRP to large-scale problems. Our proposed algorithms are based on the max-product algorithm—a common tool for finding the maximum likelihood configurations in probabilistic graphical models—and can find the most relevant walks exactly at the neuron level and approximately at the node level. Our experiments demonstrate the performance of our algorithms at scale and their utility across application domains, i.e., on epidemiology, molecular, and natural language benchmarks. We provide our codes under github.com/xiong-ping/rel_walk_gnnlrp. Ping Xiong 0002, Thomas Schnake, Michael Gastegger, Grégoire Montavon, Klaus-Robert Müller, Shinichi Nakajima |
ICML | 5 |
| 2023 | Physics-Informed Bayesian Optimization of Variational Quantum CircuitsabstractIn this paper, we propose a novel and powerful method to harness Bayesian optimization for variational quantum eigensolvers (VQEs) - a hybrid quantum-classical protocol used to approximate the ground state of a quantum Hamiltonian. Specifically, we derive a *VQE-kernel* which incorporates important prior information about quantum circuits: the kernel feature map of the VQE-kernel exactly matches the known functional form of the VQE's objective function and thereby significantly reduces the posterior uncertainty.
Moreover, we propose a novel acquisition function for Bayesian optimization called \emph{Expected Maximum Improvement over Confident Regions} (EMICoRe) which can actively exploit the inductive bias of the VQE-kernel by treating regions with low predictive uncertainty as indirectly "observed". As a result, observations at as few as three points in the search domain are sufficient to determine the complete objective function along an entire one-dimensional subspace of the optimization landscape.
Our numerical experiments demonstrate that our approach improves over state-of-the-art baselines. Kim Nicoli, Christopher J. Anders, Lena Funcke, Tobias Hartung, Karl Jansen, Stefan Kühn, Klaus-Robert Müller, Paolo Stornati, Pan Kessel, Shinichi Nakajima |
NeurIPS | 7 |
| 2023 | Learning domain invariant representations by joint Wasserstein distance minimizationabstractDomain shifts in the training data are common in practical applications of machine learning; they occur for instance when the data is coming from different sources. Ideally, a ML model should work well independently of these shifts, for example, by learning a domain-invariant representation. However, common ML losses do not give strong guarantees on how consistently the ML model performs for different domains, in particular, whether the model performs well on a domain at the expense of its performance on another domain. In this paper, we build new theoretical foundations for this problem, by contributing a set of mathematical relations between classical losses for supervised ML and the Wasserstein distance in joint space (i.e. representation and output space). We show that classification or regression losses, when combined with a GAN-type discriminator between domains, form an upper-bound to the true Wasserstein distance between domains. This implies a more invariant representation and also more stable prediction performance across domains. Theoretical results are corroborated empirically on several image datasets. Our proposed approach systematically produces the highest minimum classification accuracy across domains, and the most invariant representation. Léo Andéol, Yusei Kawakami, Yuichiro Wada, Takafumi Kanamori, Klaus-Robert Müller, Grégoire Montavon |
Neural Networks | 5 |
| 2023 | Canonical Response Parameterization: Quantifying the structure of responses to single-pulse intracranial electrical brain stimulationabstractSingle-pulse electrical stimulation in the nervous system, often called cortico-cortical evoked potential (CCEP) measurement, is an important technique to understand how brain regions interact with one another. Voltages are measured from implanted electrodes in one brain area while stimulating another with brief current impulses separated by several seconds. Historically, researchers have tried to understand the significance of evoked voltage polyphasic deflections by visual inspection, but no general-purpose tool has emerged to understand their shapes or describe them mathematically. We describe and illustrate a new technique to parameterize brain stimulation data, where voltage response traces are projected into one another using a semi-normalized dot product. The length of timepoints from stimulation included in the dot product is varied to obtain a temporal profile of structural significance, and the peak of the profile uniquely identifies the duration of the response. Using linear kernel PCA, a canonical response shape is obtained over this duration, and then single-trial traces are parameterized as a projection of this canonical shape with a residual term. Such parameterization allows for dissimilar trace shapes from different brain areas to be directly compared by quantifying cross-projection magnitudes, response duration, canonical shape projection amplitudes, signal-to-noise ratios, explained variance, and statistical significance. Artifactual trials are automatically identified by outliers in sub-distributions of cross-projection magnitude, and rejected. This technique, which we call "Canonical Response Parameterization" (CRP) dramatically simplifies the study of CCEP shapes, and may also be applied in a wide range of other settings involving event-triggered data. Kai J. Miller, Klaus-Robert Müller, Gabriela Ojeda Valencia, Harvey Huang, Nicholas M. Gregg, Gregory A. Worrell, Dora Hermes |
PLoS Comput. Biol. | 2 |
| 2023 | Langevin Cooling for Unsupervised Domain TranslationabstractDomain translation is the task of finding correspondence between two domains. Several deep neural network (DNN) models, e.g., CycleGAN and cross-lingual language models, have shown remarkable successes on this task under the unsupervised setting-the mappings between the domains are learned from two independent sets of training data in both domains (without paired samples). However, those methods typically do not perform well on a significant proportion of test samples. In this article, we hypothesize that many of such unsuccessful samples lie at the fringe-relatively low-density areas-of data distribution, where the DNN was not trained very well, and propose to perform the Langevin dynamics to bring such fringe samples toward high-density areas. We demonstrate qualitatively and quantitatively that our strategy, called Langevin cooling (L-Cool), enhances state-of-the-art methods in image translation and language translation tasks. Vignesh Srinivasan, Klaus-Robert Müller, Wojciech Samek, Shinichi Nakajima |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2022 | XAI for Transformers: Better Explanations through Conservative PropagationabstractTransformers have become an important workhorse of machine learning, with numerous applications. This necessitates the development of reliable methods for increasing their transparency. Multiple interpretability methods, often based on gradient information, have been proposed. We show that the gradient in a Transformer reflects the function only locally, and thus fails to reliably identify the contribution of input features to the prediction. We identify Attention Heads and LayerNorm as main reasons for such unreliable explanations and propose a more stable way for propagation through these layers. Our proposal, which can be seen as a proper extension of the well-established LRP method to Transformers, is shown both theoretically and empirically to overcome the deficiency of a simple gradient-based approach, and achieves state-of-the-art explanation performance on a broad range of Transformer models and datasets. Ameen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon, Klaus-Robert Müller, Lior Wolf |
ICML | 5 |
| 2022 | Efficient Computation of Higher-Order Subgraph Attribution via Message PassingabstractExplaining graph neural networks (GNNs) has become more and more important recently. Higher-order interpretation schemes, such as GNN-LRP (layer-wise relevance propagation for GNN), emerged as powerful tools for unraveling how different features interact thereby contributing to explaining GNNs. GNN-LRP gives a relevance attribution of walks between nodes at each layer, and the subgraph attribution is expressed as a sum over exponentially many such walks. In this work, we demonstrate that such exponential complexity can be avoided. In particular, we propose novel algorithms that enable to attribute subgraphs with GNN-LRP in linear-time (w.r.t. the network depth). Our algorithms are derived via message passing techniques that make use of the distributive property, thereby directly computing quantities for higher-order explanations. We further adapt our efficient algorithms to compute a generalization of subgraph attributions that also takes into account the neighboring graph features. Experimental results show the significant acceleration of the proposed algorithms and demonstrate the high usefulness and scalability of our novel generalized subgraph attribution method. Ping Xiong 0002, Thomas Schnake, Grégoire Montavon, Klaus-Robert Müller, Shinichi Nakajima |
ICML | 4 |
| 2022 | So3krates: Equivariant attention for interactions on arbitrary length-scales in molecular systemsabstractThe application of machine learning methods in quantum chemistry has enabled the study of numerous chemical phenomena, which are computationally intractable with traditional ab-initio methods. However, some quantum mechanical properties of molecules and materials depend on non-local electronic effects, which are often neglected due to the difficulty of modeling them efficiently. This work proposes a modified attention mechanism adapted to the underlying physics, which allows to recover the relevant non-local effects. Namely, we introduce spherical harmonic coordinates (SPHCs) to reflect higher-order geometric information for each atom in a molecule, enabling a non-local formulation of attention in the SPHC space. Our proposed model So3krates - a self-attention based message passing neural network - uncouples geometric information from atomic features, making them independently amenable to attention mechanisms. Thereby we construct spherical filters, which extend the concept of continuous filters in Euclidean space to SPHC space and serve as foundation for a spherical self-attention mechanism. We show that in contrast to other published methods, So3krates is able to describe non-local quantum mechanical effects over arbitrary length scales. Further, we find evidence that the inclusion of higher-order geometric correlations increases data efficiency and improves generalization. So3krates matches or exceeds state-of-the-art performance on popular benchmarks, notably, requiring a significantly lower number of parameters (0.25 - 0.4x) while at the same time giving a substantial speedup (6 - 14x for training and 2 - 11x for inference) compared to other models. J. Thorben Frank, Oliver T. Unke, Klaus-Robert Müller |
NeurIPS | 3 |
| 2022 | Scrutinizing XAI using linear ground-truth data with suppressor variablesabstractMachine learning (ML) is increasingly often used to inform high-stakes decisions. As complex ML models (e.g., deep neural networks) are often considered black boxes, a wealth of procedures has been developed to shed light on their inner workings and the ways in which their predictions come about, defining the field of 'explainable AI' (XAI). Saliency methods rank input features according to some measure of 'importance'. Such methods are difficult to validate since a formal definition of feature importance is, thus far, lacking. It has been demonstrated that some saliency methods can highlight features that have no statistical association with the prediction target (suppressor variables). To avoid misinterpretations due to such behavior, we propose the actual presence of such an association as a necessary condition and objective preliminary definition for feature importance. We carefully crafted a ground-truth dataset in which all statistical dependencies are well-defined and linear, serving as a benchmark to study the problem of suppressor variables. We evaluate common explanation methods including LRP, DTD, PatternNet, PatternAttribution, LIME, Anchors, SHAP, and permutation-based methods with respect to our objective definition. We show that most of these methods are unable to distinguish important features from suppressors in this setting. Supplementary Information: The online version contains supplementary material available at 10.1007/s10994-022-06167-y. Rick Wilming, Céline Budding, Klaus-Robert Müller, Stefan Haufe |
Mach. Learn. | 3 |
| 2022 | Building and Interpreting Deep Similarity ModelsabstractMany learning algorithms such as kernel machines, nearest neighbors, clustering, or anomaly detection, are based on distances or similarities. Before similarities are used for training an actual machine learning model, we would like to verify that they are bound to meaningful patterns in the data. In this paper, we propose to make similarities interpretable by augmenting them with an explanation. We develop BiLRP, a scalable and theoretically founded method to systematically decompose the output of an already trained deep similarity model on pairs of input features. Our method can be expressed as a composition of LRP explanations, which were shown in previous works to scale to highly nonlinear models. Through an extensive set of experiments, we demonstrate that BiLRP robustly explains complex similarity models, e.g., built on VGG-16 deep neural network features. Additionally, we apply our method to an open problem in digital humanities: detailed assessment of similarity between historical documents, such as astronomical tables. Here again, BiLRP provides insight and brings verifiability into a highly engineered and problem-specific similarity model. Oliver Eberle, Jochen Büttner, Florian Kräutli, Klaus-Robert Müller, Matteo Valleriani, Grégoire Montavon |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Higher-Order Explanations of Graph Neural Networks via Relevant WalksabstractGraph Neural Networks (GNNs) are a popular approach for predicting graph structured data. As GNNs tightly entangle the input graph into the neural network structure, common explainable AI approaches are not applicable. To a large extent, GNNs have remained black-boxes for the user so far. In this paper, we show that GNNs can in fact be naturally explained using higher-order expansions, i.e., by identifying groups of edges that jointly contribute to the prediction. Practically, we find that such explanations can be extracted using a nested attribution scheme, where existing techniques such as layer-wise relevance propagation (LRP) can be applied at each step. The output is a collection of walks into the input graph that are relevant for the prediction. Our novel explanation method, which we denote by GNN-LRP, is applicable to a broad range of graph neural networks and lets us extract practically relevant insights on sentiment analysis of text data, structure-property relationships in quantum chemistry, and image classification. Thomas Schnake, Oliver Eberle, Jonas Lederer, Shinichi Nakajima, Kristof Schütt, Klaus-Robert Müller, Grégoire Montavon |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Towards robust explanations for deep neural networksabstractExplanation methods shed light on the decision process of black-box classifiers such as deep neural networks. But their usefulness can be compromised because they are susceptible to manipulations. With this work, we aim to enhance the resilience of explanations. We develop a unified theoretical framework for deriving bounds on the maximal manipulability of a model. Based on these theoretical insights, we present three different techniques to boost robustness against manipulation: training with weight decay, smoothing activation functions, and minimizing the Hessian of the network. Our experimental results confirm the effectiveness of these approaches. Ann-Kathrin Dombrowski, Christopher J. Anders, Klaus-Robert Müller, Pan Kessel |
Pattern Recognit. | 3 |
| 2021 | Explainable Deep One-Class Classification
Philipp Liznerski, Lukas Ruff, Robert A. Vandermeulen, Billy Joe Franks, Marius Kloft, Klaus-Robert Müller |
ICLR | 6 |
| 2021 | Efficient hierarchical Bayesian inference for spatio-temporal regression models in neuroimagingabstractSeveral problems in neuroimaging and beyond require inference on the parameters of multi-task sparse hierarchical regression models. Examples include M/EEG inverse problems, neural encoding models for task-based fMRI analyses, and climate science. In these domains, both the model parameters to be inferred and the measurement noise may exhibit a complex spatio-temporal structure. Existing work either neglects the temporal structure or leads to computationally demanding inference schemes. Overcoming these limitations, we devise a novel flexible hierarchical Bayesian framework within which the spatio-temporal dynamics of model parameters and noise are modeled to have Kronecker product covariance structure. Inference in our framework is based on majorization-minimization optimization and has guaranteed convergence properties. Our highly efficient algorithms exploit the intrinsic Riemannian geometry of temporal autocovariance matrices. For stationary dynamics described by Toeplitz matrices, the theory of circulant embeddings is employed. We prove convex bounding properties and derive update rules of the resulting algorithms. On both synthetic and real neural data from M/EEG, we demonstrate that our methods lead to improved performance. Ali Hashemi 0002, Yijing Gao, Sanjay Ghosh, Klaus-Robert Müller, Srikantan S. Nagarajan, Stefan Haufe |
NeurIPS | 5 |
| 2021 | SE(3)-equivariant prediction of molecular wavefunctions and electronic densitiesabstractMachine learning has enabled the prediction of quantum chemical properties with high accuracy and efficiency, allowing to bypass computationally costly ab initio calculations. Instead of training on a fixed set of properties, more recent approaches attempt to learn the electronic wavefunction (or density) as a central quantity of atomistic systems, from which all other observables can be derived. This is complicated by the fact that wavefunctions transform non-trivially under molecular rotations, which makes them a challenging prediction target. To solve this issue, we introduce general SE(3)-equivariant operations and building blocks for constructing deep learning architectures for geometric point cloud data and apply them to reconstruct wavefunctions of atomistic systems with unprecedented accuracy. Our model achieves speedups of over three orders of magnitude compared to ab initio methods and reduces prediction errors by up to two orders of magnitude compared to the previous state-of-the-art. This accuracy makes it possible to derive properties such as energies and forces directly from the wavefunction in an end-to-end manner. We demonstrate the potential of our approach in a transfer learning application, where a model trained on low accuracy reference wavefunctions implicitly learns to correct for electronic many-body interactions from observables computed at a higher level of theory. Such machine-learned wavefunction surrogates pave the way towards novel semi-empirical methods, offering resolution at an electronic level while drastically decreasing computational cost. Additionally, the predicted wavefunctions can serve as initial guess in conventional ab initio methods, decreasing the number of iterations required to arrive at a converged solution, thus leading to significant speedups without any loss of accuracy or robustness. While we focus on physics applications in this contribution, the proposed equivariant framework for deep learning on point clouds is promising also beyond, say, in computer vision or graphics. Oliver T. Unke, Mihail Bogojeski, Michael Gastegger, Mario Geiger, Tess E. Smidt, Klaus-Robert Müller |
NeurIPS | 6 |
| 2021 | Robustifying models against adversarial attacks by Langevin dynamics
Vignesh Srinivasan, Csaba Rohrer, Arturo Marbán, Klaus-Robert Müller, Wojciech Samek, Shinichi Nakajima |
Neural Networks | 4 |
| 2021 | A Unifying Review of Deep and Shallow Anomaly DetectionabstractDeep learning approaches to anomaly detection (AD) have recently improved the state of the art in detection performance on complex data sets, such as large collections of images or text. These results have sparked a renewed interest in the AD problem and led to the introduction of a great variety of new methods. With the emergence of numerous such methods, including approaches based on generative models, one-class classification, and reconstruction, there is a growing need to bring methods of this field into a systematic and unified perspective. In this review, we aim to identify the common underlying principles and the assumptions that are often made implicitly by various methods. In particular, we draw connections between classic “shallow” and novel deep approaches and show how this relation might cross-fertilize or extend both directions. We further provide an empirical assessment of major existing methods that are enriched by the use of recent explainability techniques and present specific worked-through examples together with practical advice. Finally, we outline critical open challenges and identify specific paths for future research in AD. Lukas Ruff, Jacob R. Kauffmann, Robert A. Vandermeulen, Grégoire Montavon, Wojciech Samek, Marius Kloft, Thomas G. Dietterich, Klaus-Robert Müller |
Proc. IEEE | 8 |
| 2021 | Explaining Deep Neural Networks and Beyond: A Review of Methods and ApplicationsabstractWith the broader and highly successful usage of machine learning (ML) in industry and the sciences, there has been a growing demand for explainable artificial intelligence (XAI). Interpretability and explanation methods for gaining a better understanding of the problem-solving abilities and strategies of nonlinear ML, in particular, deep neural networks, are, therefore, receiving increased attention. In this work, we aim to: 1) provide a timely overview of this active emerging field, with a focus on “post hoc” explanations, and explain its theoretical foundations; 2) put interpretability algorithms to a test both from a theory and comparative evaluation perspective using extensive simulations; 3) outline best practice aspects, i.e., how to best include interpretation methods into the standard usage of ML; and 4) demonstrate successful usage of XAI in a representative selection of application scenarios. Finally, we discuss challenges and possible future directions of this exciting foundational field of ML. Wojciech Samek, Grégoire Montavon, Sebastian Lapuschkin, Christopher J. Anders, Klaus-Robert Müller |
Proc. IEEE | 5 |
| 2021 | Basis profile curve identification to understand electrical stimulation effects in human brain networksabstractBrain networks can be explored by delivering brief pulses of electrical current in one area while measuring voltage responses in other areas. We propose a convergent paradigm to study brain dynamics, focusing on a single brain site to observe the average effect of stimulating each of many other brain sites. Viewed in this manner, visually-apparent motifs in the temporal response shape emerge from adjacent stimulation sites. This work constructs and illustrates a data-driven approach to determine characteristic spatiotemporal structure in these response shapes, summarized by a set of unique "basis profile curves" (BPCs). Each BPC may be mapped back to underlying anatomy in a natural way, quantifying projection strength from each stimulation site using simple metrics. Our technique is demonstrated for an array of implanted brain surface electrodes in a human patient. This framework enables straightforward interpretation of single-pulse brain stimulation data, and can be applied generically to explore the diverse milieu of interactions that comprise the connectome. Kai J. Miller, Klaus-Robert Müller, Dora Hermes |
PLoS Comput. Biol. | 2 |
| 2021 | Pruning by explaining: A novel criterion for deep neural network pruningabstractThe success of convolutional neural networks (CNNs) in various applications is accompanied by a significant increase in computation and parameter storage costs. Recent efforts to reduce these overheads involve pruning and compressing the weights of various layers while at the same time aiming to not sacrifice performance. In this paper, we propose a novel criterion for CNN pruning inspired by neural network interpretability: The most relevant units, i.e. weights or filters, are automatically found using their relevance scores obtained from concepts of explainable AI (XAI). By exploring this idea, we connect the lines of interpretability and model compression research. We show that our proposed method can efficiently prune CNN models in transfer-learning setups in which networks pre-trained on large corpora are adapted to specialized tasks. The method is evaluated on a broad range of computer vision datasets. Notably, our novel criterion is not only competitive or better compared to state-of-the-art pruning criteria when successive retraining is performed, but clearly outperforms these previous criteria in the resource-constrained application scenario in which the data of the task to be transferred to is very scarce and one chooses to refrain from fine-tuning. Our method is able to compress the model iteratively while maintaining or even improving accuracy. At the same time, it has a computational cost in the order of gradient computation and is comparatively simple to apply without the need for tuning hyperparameters for pruning. Seul-Ki Yeom, Philipp Seegerer, Sebastian Lapuschkin, Alexander Binder, Simon Wiedemann, Klaus-Robert Müller, Wojciech Samek |
Pattern Recognit. | 6 |
| 2021 | Clustered Federated Learning: Model-Agnostic Distributed Multitask Optimization Under Privacy ConstraintsabstractFederated learning (FL) is currently the most widely adopted framework for collaborative training of (deep) machine learning models under privacy constraints. Albeit its popularity, it has been observed that FL yields suboptimal results if the local clients' data distributions diverge. To address this issue, we present clustered FL (CFL), a novel federated multitask learning (FMTL) framework, which exploits geometric properties of the FL loss surface to group the client population into clusters with jointly trainable data distributions. In contrast to existing FMTL approaches, CFL does not require any modifications to the FL communication protocol to be made, is applicable to general nonconvex objectives (in particular, deep neural networks), does not require the number of clusters to be known a priori, and comes with strong mathematical guarantees on the clustering quality. CFL is flexible enough to handle client populations that vary over time and can be implemented in a privacy-preserving way. As clustering is only performed after FL has converged to a stationary point, CFL can be viewed as a postprocessing method that will always achieve greater or equal performance than conventional FL by allowing clients to arrive at more specialized models. We verify our theoretical analysis in experiments with deep convolutional and recurrent neural networks on commonly used FL data sets. Felix Sattler, Klaus-Robert Müller, Wojciech Samek |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Benign Examples: Imperceptible Changes Can Enhance Image Translation Performance
Vignesh Srinivasan, Klaus-Robert Müller, Wojciech Samek, Shinichi Nakajima |
AAAI | 2 |
| 2020 | On the Byzantine Robustness of Clustered Federated LearningabstractFederated Learning (FL) is currently the most widely adopted framework for collaborative training of (deep) machine learning models under privacy constraints. Albeit it's popularity, it has been observed that Federated Learning yields suboptimal results if the local clients' data distributions diverge. The recently proposed Clustered Federated Learning Framework addresses this issue, by separating the client population into different groups based on the pairwise cosine similarities between their parameter updates. In this work we investigate the application of CFL to byzantine settings, where a subset of clients behaves unpredictably or tries to disturb the joint training effort in an directed or undirected way. We perform experiments with deep neural networks on common Federated Learning datasets which demonstrate that CFL (without modifications) is able to reliably detect byzantine clients and remove them from training. Felix Sattler, Klaus-Robert Müller, Thomas Wiegand 0001, Wojciech Samek |
ICASSP | 2 |
| 2020 | EEG-Based Assessment of Perceived Quality in Complex Natural ImagesabstractPsychophysiological methods gained a lot of interest in recent years as a potential remedy for the inherent flaws of overt psychophysical quality assessment methods. Among the psychophysiological monitoring methods, electroencephalography showed to be a promising choice. Specifically, the steady state visually evoked potential (SSVEP) was shown to provide a reliable neural correlate of perceived visual quality for degraded texture image patches. This paper evaluates the feasibility of the SSVEP-based quality assessment approach for more realistic and practically relevant images. To this end, we collected overt, psychophysical responses and neural, psychophysiological responses of 14 participants to 6 HD images each compressed at 4 distortion levels. The psychophysical part followed the Degradation Category Rating procedure. In the subsequent psychophysiological part, the subjects were presented with distorted and reference images alternating at a fixed rate of $f_{stim}\,=5$ Hz to elicit the SSVEP. We show that the amplitude of the $1 ^{st}$ harmonic of the SSVEP correlates significantly with the psychophysical responses $(\vert \rho \vert = 0.85, p \lt 0.05)$ in a single channel analysis at the Oz electrode. Tamer Ajaj, Klaus-Robert Müller, Gabriel Curio, Thomas Wiegand 0001, Sebastian Bosse |
ICIP | 2 |
| 2020 | Deep Semi-Supervised Anomaly Detection
Lukas Ruff, Robert A. Vandermeulen, Nico Görnitz, Alexander Binder, Emmanuel Müller, Klaus-Robert Müller, Marius Kloft |
ICLR | 6 |
| 2020 | Fairwashing explanations with off-manifold detergentabstractExplanation methods promise to make black-box classifiers more transparent. As a result, it is hoped that they can act as proof for a sensible, fair and trustworthy decision-making process of the algorithm and thereby increase its acceptance by the end-users. In this paper, we show both theoretically and experimentally that these hopes are presently unfounded. Specifically, we show that, for any classifier $g$, one can always construct another classifier $\tilde{g}$ which has the same behavior on the data (same train, validation, and test error) but has arbitrarily manipulated explanation maps. We derive this statement theoretically using differential geometry and demonstrate it experimentally for various explanation methods, architectures, and datasets. Motivated by our theoretical insights, we then propose a modification of existing explanation methods which makes them significantly more robust. Christopher J. Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller, Pan Kessel |
ICML | 4 |
| 2020 | EEG-Based Assessment of Perceived Realness in Stylized Face ImagesabstractIn this paper, we investigate the perception of realness in rendered face images experimentally using electroencephalography. To this end, we presented ten subjects with 36 character images based on six different faces (varying in gender and emotional expression) rendered at six different levels of realness ranging from abstract, cartoon-like renderings to real photographs. In the first psychophysical part of our study, we asked participants to rate perceived realness, appeal, familiarity, reassurance, and attractiveness for the presented characters. In the second part, we recorded the electroencephalogram when presenting the character images at a stimulation frequency of fstim= 5 Hz. We show that the amplitudes of the odd harmonics of the elicited steady-state visual evoked potential correlate with the psychophysical responses (|ρ| = 0.83, p < 0.05). Milena T. Bagdasarian, Anna Hilsmann, Peter Eisert, Gabriel Curio, Klaus-Robert Müller, Thomas Wiegand 0001, Sebastian Bosse |
QoMEX | 5 |
| 2020 | Towards explaining anomalies: A deep Taylor decomposition of one-class modelsabstractDetecting anomalies in the data is a common machine learning task, with numerous applications in the sciences and industry. In practice, it is not always sufficient to reach high detection accuracy, one would also like to be able to understand why a given data point has been predicted to be anomalous. We propose a principled approach for one-class SVMs (OC-SVM), that draws on the novel insight that these models can be rewritten as distance/pooling neural networks. This ‘neuralization’ step lets us apply deep Taylor decomposition (DTD), a methodology that leverages the model structure in order to quickly and reliably explain decisions in terms of input features. The proposed method (called ‘OC-DTD’) is applicable to a number of common distance-based kernel functions, and it outperforms baselines such as sensitivity analysis, distance to nearest neighbor, or edge detection. Jacob R. Kauffmann, Klaus-Robert Müller, Grégoire Montavon |
Pattern Recognit. | 2 |
| 2020 | Optimizing for Measure of Performance in Max-Margin ParsingabstractMany learning tasks in the field of natural language processing including sequence tagging, sequence segmentation, and syntactic parsing have been successfully approached by means of structured prediction methods. An appealing property of the corresponding training algorithms is their ability to integrate the loss function of interest into the optimization process improving the final results according to the chosen measure of performance. Here, we focus on the task of constituency parsing and show how to optimize the model for the F1-score in the max-margin framework of a structural support vector machine (SVM). For reasons of computational efficiency, it is a common approach to binarize the corresponding grammar before training. Unfortunately, this introduces a bias during the training procedure as the corresponding loss function is evaluated on the binary representation, while the resulting performance is measured on the original unbinarized trees. Here, we address this problem by extending the inference procedure presented by Bauer et al. Specifically, we propose an algorithmic modification that allows evaluating the loss on the unbinarized trees. The new approach properly models the loss function of interest resulting in better prediction accuracy and still benefits from the computational efficiency due to binarized representation. The presented idea can be easily transferred to other structured loss functions. Alexander Bauer 0001, Shinichi Nakajima, Nico Görnitz, Klaus-Robert Müller |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Robust and Communication-Efficient Federated Learning From Non-i.i.d. DataabstractFederated learning allows multiple parties to jointly train a deep learning model on their combined data, without any of the participants having to reveal their local data to a centralized server. This form of privacy-preserving collaborative learning, however, comes at the cost of a significant communication overhead during training. To address this problem, several compression methods have been proposed in the distributed training literature that can reduce the amount of required communication by up to three orders of magnitude. These existing methods, however, are only of limited utility in the federated learning setting, as they either only compress the upstream communication from the clients to the server (leaving the downstream communication uncompressed) or only perform well under idealized conditions, such as i.i.d. distribution of the client data, which typically cannot be found in federated learning. In this article, we propose sparse ternary compression (STC), a new compression framework that is specifically designed to meet the requirements of the federated learning environment. STC extends the existing compression technique of top- k gradient sparsification with a novel mechanism to enable downstream compression as well as ternarization and optimal Golomb encoding of the weight updates. Our experiments on four different learning tasks demonstrate that STC distinctively outperforms federated averaging in common federated learning scenarios. These results advocate for a paradigm shift in federated optimization toward high-frequency low-bitwidth communication, in particular in the bandwidth-constrained learning environments. Felix Sattler, Simon Wiedemann, Klaus-Robert Müller, Wojciech Samek |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Compact and Computationally Efficient Representation of Deep Neural NetworksabstractAt the core of any inference procedure, deep neural networks are dot product operations, which are the component that requires the highest computational resources. For instance, deep neural networks, such as VGG-16, require up to 15-G operations in order to perform the dot products present in a single forward pass, which results in significant energy consumption and thus limits their use in resource-limited environments, e.g., on embedded devices or smartphones. One common approach to reduce the complexity of the inference is to prune and quantize the weight matrices of the neural network. Usually, this results in matrices whose entropy values are low, as measured relative to the empirical probability mass distribution of its elements. In order to efficiently exploit such matrices, one usually relies on, inter alia, sparse matrix representations. However, most of these common matrix storage formats make strong statistical assumptions about the distribution of the elements; therefore, cannot efficiently represent the entire set of matrices that exhibit low-entropy statistics (thus, the entire set of compressed neural network weight matrices). In this paper, we address this issue and present new efficient representations for matrices with low-entropy statistics. Alike sparse matrix data structures, these formats exploit the statistical properties of the data in order to reduce the size and execution complexity. Moreover, we show that the proposed data structures can not only be regarded as a generalization of sparse formats but are also more energy and time efficient under practically relevant assumptions. Finally, we test the storage requirements and execution performance of the proposed formats on compressed neural networks and compare them to dense and sparse representations. We experimentally show that we are able to attain up to ×42 compression ratios, ×5 speed ups, and ×90 energy savings when we lossless convert the state-of-the-art networks, such as AlexNet, VGG-16, ResNet152, and DenseNet, into the new data structures and benchmark their respective dot product. Simon Wiedemann, Klaus-Robert Müller, Wojciech Samek |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Partial Optimality of Dual Decomposition for MAP Inference in Pairwise MRFsabstractMarkov random fields (MRFs) are a powerful tool for modelling statistical dependencies for a set of random variables using a graphical representation. An important computational problem related to MRFs, called maximum a posteriori (MAP) inference, is finding a joint variable assignment with the maximal probability. It is well known that the two popular optimisation techniques for this task, linear programming (LP) relaxation and dual decomposition (DD), have a strong connection both providing an optimal solution to the MAP problem when a corresponding LP relaxation is tight. However, less is known about their relationship in the opposite and more realistic case. In this paper, we explain how the fully integral assignments obtained via DD partially agree with the optimal fractional assignments via LP relaxation when the latter is not tight. In particular, for binary pairwise MRFs the corresponding result suggests that both methods share the partial optimality property of their solutions. Alexander Bauer 0001, Shinichi Nakajima, Nico Görnitz, Klaus-Robert Müller |
AISTATS | 4 |
| 2019 | Sparse Binary Compression: Towards Distributed Deep Learning with minimal CommunicationabstractCurrently, progressively larger deep neural networks are trained on ever growing data corpora. In result, distributed training schemes are becoming increasingly relevant. A major issue in distributed training is the limited communication bandwidth between contributing nodes or prohibitive communication cost in general. To mitigate this problem we propose Sparse Binary Compression (SBC), a compression framework that allows for a drastic reduction of communication cost for distributed training. SBC combines existing techniques of communication delay and gradient sparsification with a novel binarization method and optimal weight update encoding to push compression gains to new limits. By doing so, our method also allows us to smoothly trade-off gradient sparsity and temporal sparsity to adapt to the requirements of the learning task. Our experiments show, that SBC can reduce the upstream communication on a variety of convolutional and recurrent neural network architectures by more than four orders of magnitude without significantly harming the convergence speed in terms of forward-backward passes. For instance, we can train ResNet50 on ImageNet in the same number of iterations to the baseline accuracy, using ×3531 less bits or train it to a 1% lower accuracy using ×37208 less bits. In the latter case, the total upstream communication required is cut from 125 terabytes to 3.35 gigabytes for every participating client. Felix Sattler, Simon Wiedemann, Klaus-Robert Müller, Wojciech Samek |
IJCNN | 3 |
| 2019 | Entropy-Constrained Training of Deep Neural NetworksabstractMotivated by the Minimum Description Length (MDL) principle, we first derive an expression for the entropy of a neural network which measures its complexity explicitly in terms of its bit-size. Then, we formalize the problem of neural network compression as an entropy-constrained optimization objective. This objective generalizes many of the currently proposed compression techniques in the literature, in that pruning or reducing the cardinality of the weight elements can be seen as special cases of entropy reduction methods. Furthermore, we derive a continuous relaxation of the objective, which allows us to minimize it using gradient-based optimization techniques. Finally, we show that we can reach compression results, which are competitive with those obtained using state-of-the-art techniques, on different network architectures and data sets, e.g. achieving×71 compression gains on a VGG-like architecture. Simon Wiedemann, Arturo Marbán, Klaus-Robert Müller, Wojciech Samek |
IJCNN | 3 |
| 2019 | Explanations can be manipulated and geometry is to blameabstractExplanation methods aim to make neural networks more trustworthy and interpretable. In this paper, we demonstrate a property of explanation methods which is disconcerting for both of these purposes. Namely, we show that explanations can be manipulated arbitrarily by applying visually hardly perceptible perturbations to the input that keep the network's output approximately constant. We establish theoretically that this phenomenon can be related to certain geometrical properties of neural networks. This allows us to derive an upper bound on the susceptibility of explanations to manipulations. Based on this result, we propose effective mechanisms to enhance the robustness of explanations. Ann-Kathrin Dombrowski, Maximilian Alber, Christopher J. Anders, Marcel Ackermann 0001, Klaus-Robert Müller, Pan Kessel |
NeurIPS | 5 |
| 2019 | Classification of structured validation data using stateless and stateful featuresabstractTo reliably identify problems impacting the service quality and system dependability of mobile communication networks, the monitored data needs to be validated. This paper proposes and evaluates analysis methods, features and learning methods for the automatic validation of such data, with a special focus on failure data of mobile communication data. This data can be analyzed for discriminating failures caused by problems in the infrastructure (valid failures) from those caused by other circumstances like device imperfections (invalid failures), with the purpose of filtering the invalid failures, which effectively increases both dependability and value of the underlying data. To represent the complex structural and temporal properties of the mobile communication data, two complementary feature representations are proposed and compared, followed by a discussion of classification methods which are suitable for these feature spaces and for an interpretation of their results to support manual auditing. Their classification performances on these feature spaces are evaluated and compared to competitive approaches. In the evaluation a classification performances of up to 97% AUC–ROC is achieved. This renders our approach a good alternative to using manual matching rules, which require costly expert-knowledge and are much more time-consuming to define and maintain — while also highlighting the relevance of combining feature spaces of different problem perspectives. Additionally it is shown that using non-proprietary data analysis can enable feature representations nearly as expressive as those created by using proprietary analysis methods, which allows a broader application of the proposed methods, due to the lower processing requirements. Guido Schwenk, Ralf Pabst, Klaus-Robert Müller |
Comput. Commun. | 3 |
| 2019 | iNNvestigate Neural Networks!abstractIn recent years, deep neural networks have revolutionized many application domains of machine learning and are key components of many critical decision or predictive processes. Therefore, it is crucial that domain specialists can understand and analyze actions and predictions, even of the most complex neural network architectures. Despite these arguments neural networks are often treated as black boxes. In the attempt to alleviate this shortcoming many analysis methods were proposed, yet the lack of reference implementations often makes a systematic comparison between the methods a major effort. The presented library innvestigate addresses this by providing a common interface and out-of-the-box implementation for many analysis methods, including the reference implementation for PatternNet and PatternAttribution as well as for LRP-methods. To demonstrate the versatility of innvestigate, we provide an analysis of image classifications for variety of state-of-the-art neural network architectures. Maximilian Alber, Sebastian Lapuschkin, Philipp Seegerer, Miriam Hägele, Kristof Schütt, Grégoire Montavon, Wojciech Samek, Klaus-Robert Müller, Sven Dähne, Pieter-Jan Kindermans |
J. Mach. Learn. Res. | 8 |
| 2019 | N-ary decomposition for multi-class classification
Joey Tianyi Zhou, Ivor W. Tsang, Shen-Shyang Ho, Klaus-Robert Müller |
Mach. Learn. | 4 |
| 2018 | How are the Centered Kernel Principal Components Relevant to Regression Task? -An Exact AnalysisabstractWe present an exact analytic expression of the contributions of the kernel principal components to the relevant information in a nonlinear regression problem. A related study has been presented by Braun, Buhmann, and Müller in 2008, where an upper bound of the contributions was given for a general supervised learning problem but with “uncentered” kernel PCAs. Our analysis clarifies that the relevant information of a kernel regression under explicit centering operation is contained in a finite number of leading kernel principal components, as in the “uncentered” kernel-Pca case, if the kernel matches the underlying nonlinear function so that the eigenvalues of the centered kernel matrix decay quickly. We compare the regression performances of the least-square-based methods with the centered and uncentered kernel PCAs by simulations. Masahiro Yukawa, Klaus-Robert Müller, Yuto Ogino |
ICASSP | 2 |
| 2018 | Learning how to explain neural networks: PatternNet and PatternAttribution
Pieter-Jan Kindermans, Kristof Schütt, Maximilian Alber, Klaus-Robert Müller, Dumitru Erhan, Been Kim, Sven Dähne |
ICLR (Poster) | 4 |
| 2018 | Curly: An AI-based Curling Robot Successfully Competing in the Olympic Discipline of CurlingabstractMost artificial intelligence (AI) based learning systems act in virtual or laboratory environments. Here we demonstrate an AI-based curling robot system named `Curly' that competes on a real-world curling ice sheet. Curly encompasses (1) an AI-based curling strategy and simulation engine under consideration of the high `icy' uncertainty, (2) the thrower robot enabled by autonomous driving with traction control, and (3) the skip robot that allows to recognize the curling field and stone configuration based on vision technology. The Curly performed well both: in classical game situations and when interacting with human opponents, namely, the top-ranked Korean amateur high school curling team. Dong-Ok Won, Byung-Do Kim, Ho-Jung Kim, Tae-San Eom, Klaus-Robert Müller, Seong-Whan Lee |
IJCAI | 5 |
| 2018 | Assessing Perceived Image Quality Using Steady-State Visual Evoked Potentials and Spatio-Spectral DecompositionabstractSteady-state visual evoked potentials (SSVEPs) are neural responses, measurable using electroencephalography (EEG), that are directly linked to sensory processing of visual stimuli. In this paper, SSVEP is used to assess the perceived quality of texture images. The EEG-based assessment method is compared with conventional methods, and recorded EEG data are correlated to obtained mean opinion scores (MOSs). A dimensionality reduction technique for EEG data called spatio-spectral decomposition (SSD) is adapted for the SSVEP framework and used to extract physiologically meaningful and plausible neural components from the EEG recordings. It is shown that the use of SSD not only increases the correlation between neural features and MOS to r = -0.93, but also solves the problem of channel selection in an EEG-based image-quality assessment. Sebastian Bosse, Laura Acqualagna, Wojciech Samek, Anne Porbadnigk, Gabriel Curio, Benjamin Blankertz, Klaus-Robert Müller, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2018 | Deep Neural Networks for No-Reference and Full-Reference Image Quality AssessmentabstractWe present a deep neural network-based approach to image quality assessment (IQA). The network is trained end-to-end and comprises ten convolutional layers and five pooling layers for feature extraction, and two fully connected layers for regression, which makes it significantly deeper than related IQA models. Unique features of the proposed architecture are that: 1) with slight adaptations it can be used in a no-reference (NR) as well as in a full-reference (FR) IQA setting and 2) it allows for joint learning of local quality and local weights, i.e., relative importance of local quality to the global quality estimate, in an unified framework. Our approach is purely data-driven and does not rely on hand-crafted features or other types of prior domain knowledge about the human visual system or image statistics. We evaluate the proposed approach on the LIVE, CISQ, and TID2013 databases as well as the LIVE In the wild image quality challenge database and show superior performance to state-of-the-art NR and FR IQA methods. Finally, cross-database evaluation shows a high ability to generalize between different databases, indicating a high robustness of the learned features. Sebastian Bosse, Dominique Maniry, Klaus-Robert Müller, Thomas Wiegand 0001, Wojciech Samek |
IEEE Trans. Image Process. | 3 |
| 2018 | Support Vector Data Descriptions and k-Means Clustering: One Class?abstractWe present ClusterSVDD, a methodology that unifies support vector data descriptions (SVDDs) and $k$ -means clustering into a single formulation. This allows both methods to benefit from one another, i.e., by adding flexibility using multiple spheres for SVDDs and increasing anomaly resistance and flexibility through kernels to $k$ -means. In particular, our approach leads to a new interpretation of $k$ -means as a regularized mode seeking algorithm. The unifying formulation further allows for deriving new algorithms by transferring knowledge from one-class learning settings to clustering settings and vice versa. As a showcase, we derive a clustering method for structured data based on a one-class learning scenario. Additionally, our formulation can be solved via a particularly simple optimization scheme. We evaluate our approach empirically to highlight some of the proposed benefits on artificially generated data, as well as on real-world problems, and provide a Python software package comprising various implementations of primal and dual SVDD as well as our proposed ClusterSVDD. Nico Görnitz, Luiz Alberto Lima, Klaus-Robert Müller, Marius Kloft, Shinichi Nakajima |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | Transductive Regression for Data With Latent Dependence StructureabstractAnalyzing data with latent spatial and/or temporal structure is a challenge for machine learning. In this paper, we propose a novel nonlinear model for studying data with latent dependence structure. It successfully combines the concepts of Markov random fields, transductive learning, and regression, making heavy use of the notion of joint feature maps. Our transductive conditional random field regression model is able to infer the latent states by combining limited labeled data of high precision with unlabeled data containing measurement uncertainty. In this manner, we can propagate accurate information and greatly reduce uncertainty. We demonstrate the usefulness of our novel framework on generated time series data with the known temporal structure and successfully validate it on synthetic as well as real-world offshore data with the spatial structure from the oil industry to predict rock porosities from acoustic impedance data. Nico Görnitz, Luiz Alberto Lima, Luiz Eduardo Varella, Klaus-Robert Müller, Shinichi Nakajima |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Interpretable human action recognition in compressed domainabstractCompressed domain human action recognition algorithms are extremely efficient, because they only require a partial decoding of the video bit stream. However, the question what exactly makes these algorithms decide for a particular action is still a mystery. In this paper, we present a general method, Layer-wise Relevance Propagation (LRP), to understand and interpret action recognition algorithms and apply it to a state-of-the-art compressed domain method based on Fisher vector encoding and SVM classification. By using LRP, the classifiers decisions are propagated back every step in the action recognition pipeline until the input is reached. This methodology allows to identify where and when the important (from the classifier's perspective) action happens in the video. To our knowledge, this is the first work to interpret a compressed domain action recognition algorithm. We evaluate our method on the HMDB51 dataset and show that in many cases a few significant frames contribute most towards the prediction of the video to a particular class. Vignesh Srinivasan, Sebastian Lapuschkin, Cornelius Hellge, Klaus-Robert Müller, Wojciech Samek |
ICASSP | 4 |
| 2017 | Minimizing Trust Leaks for Robust Sybil DetectionabstractSybil detection is a crucial task to protect online social networks (OSNs) against intruders who try to manipulate automatic services provided by OSNs to their customers. In this paper, we first discuss the robustness of graph-based Sybil detectors SybilRank and Integro and refine theoretically their security guarantees towards more realistic assumptions. After that, we formally introduce adversarial settings for the graph-based Sybil detection problem and derive a corresponding optimal attacking strategy by exploitation of trust leaks. Based on our analysis, we propose transductive Sybil ranking (TSR), a robust extension to SybilRank and Integro that directly minimizes trust leaks. Our empirical evaluation shows significant advantages of TSR over state-of-the-art competitors on a variety of attacking scenarios on artificially generated data and real-world datasets. János Höner, Shinichi Nakajima, Alexander Bauer 0001, Klaus-Robert Müller, Nico Görnitz |
ICML | 4 |
| 2017 | An Empirical Study on The Properties of Random Bases for Kernel MethodsabstractKernel machines as well as neural networks possess universal function approximation properties. Nevertheless in practice their ways of choosing the appropriate function class differ. Specifically neural networks learn a representation by adapting their basis functions to the data and the task at hand, while kernel methods typically use a basis that is not adapted during training. In this work, we contrast random features of approximated kernel machines with learned features of neural networks. Our analysis reveals how these random and adaptive basis functions affect the quality of learning. Furthermore, we present basis adaptation schemes that allow for a more compact representation, while retaining the generalization properties of kernel machines. Maximilian Alber, Pieter-Jan Kindermans, Kristof Schütt, Klaus-Robert Müller, Fei Sha |
NIPS | 4 |
| 2017 | SchNet: A continuous-filter convolutional neural network for modeling quantum interactionsabstractDeep learning has the potential to revolutionize quantum chemistry as it is ideally suited to learn representations for structured data and speed up the exploration of chemical space. While convolutional neural networks have proven to be the first choice for images, audio and video data, the atoms in molecules are not restricted to a grid. Instead, their precise locations contain essential physical information, that would get lost if discretized. Thus, we propose to use continuous-filter convolutional layers to be able to model local correlations without requiring the data to lie on a grid. We apply those layers in SchNet: a novel deep learning architecture modeling quantum interactions in molecules. We obtain a joint model for the total energy and interatomic forces that follows fundamental quantum-chemical principles. Our architecture achieves state-of-the-art performance for benchmarks of equilibrium molecules and molecular dynamics trajectories. Finally, we introduce a more challenging benchmark with chemical and structural variations that suggests the path for further work. Kristof Schütt, Pieter-Jan Kindermans, Huziel Enoc Sauceda Felix, Stefan Chmiela, Alexandre Tkatchenko, Klaus-Robert Müller |
NIPS | 6 |
| 2017 | An Easy-to-hard Learning Paradigm for Multiple Classes and Multiple LabelsabstractMany applications, such as human action recognition and object detection, can be formulated as a multiclass classification problem. One-vs-rest (OVR) is one of the most widely used approaches for multiclass classification due to its simplicity and excellent performance. However, many confusing classes in such applications will degrade its results. For example, hand clap and boxing are two confusing actions. Hand clap is easily misclassified as boxing, and vice versa. Therefore, precisely classifying confusing classes remains a challenging task. To obtain better performance for multiclass classifications that have confusing classes, we first develop a classifier chain model for multiclass classification (CCMC) to transfer class information between classifiers. Then, based on an analysis of our proposed model, we propose an easy- to-hard learning paradigm for multiclass classification to automatically identify easy and hard classes and then use the predictions from simpler classes to help solve harder classes. Similar to CCMC, the classifier chain (CC) model is also proposed by Read et al. (2009) to capture the label dependency for multi-label classification. However, CC does not consider the order of difficulty of the labels and achieves degenerated performance when there are many confusing labels. Therefore, it is non- trivial to learn the appropriate label order for CC. Motivated by our analysis for CCMC, we also propose the easy-to-hard learning paradigm for multi-label classification to automatically identify easy and hard labels, and then use the predictions from simpler labels to help solve harder labels. We also demonstrate that our proposed strategy can be successfully applied to a wide range of applications, such as ordinal classification and relationship prediction. Extensive empirical studies validate our analysis and the effectiveness of our proposed easy-to-hard learning strategies. Weiwei Liu 0003, Ivor W. Tsang, Klaus-Robert Müller |
J. Mach. Learn. Res. | 3 |
| 2017 | Explaining nonlinear classification decisions with deep Taylor decompositionabstractNonlinear methods such as Deep Neural Networks (DNNs) are the gold standard for various challenging machine learning problems such as image recognition. Although these methods perform impressively well, they have a significant disadvantage, the lack of transparency, limiting the interpretability of the solution and thus the scope of application in practice. Especially DNNs act as black boxes due to their multilayer nonlinear structure. In this paper we introduce a novel methodology for interpreting generic multilayer neural networks by decomposing the network classification decision into contributions of its input elements. Although our focus is on image classification, the method is applicable to a broad set of input data, learning tasks and network architectures. Our method called deep Taylor decomposition efficiently utilizes the structure of the network by backpropagating the explanations from the output to the input layer. We evaluate the proposed method empirically on the MNIST and ILSVRC data sets. Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, Klaus-Robert Müller |
Pattern Recognit. | 5 |
| 2017 | Accurate Maximum-Margin Training for Parsing With Context-Free GrammarsabstractThe task of natural language parsing can naturally be embedded in the maximum-margin framework for structured output prediction using an appropriate joint feature map and a suitable structured loss function. While there are efficient learning algorithms based on the cutting-plane method for optimizing the resulting quadratic objective with potentially exponential number of linear constraints, their efficiency crucially depends on the inference algorithms used to infer the most violated constraint in a current iteration. In this paper, we derive an extension of the well-known Cocke-Kasami-Younger (CKY) algorithm used for parsing with probabilistic context-free grammars for the case of loss-augmented inference enabling an effective training in the cutting-plane approach. The resulting algorithm is guaranteed to find an optimal solution in polynomial time exceeding the running time of the CKY algorithm by a term, which only depends on the number of possible loss values. In order to demonstrate the feasibility of the presented algorithm, we perform a set of experiments for parsing English sentences. Alexander Bauer 0001, Mikio L. Braun, Klaus-Robert Müller |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Efficient Exact Inference With Loss Augmented Objective in Structured LearningabstractStructural support vector machine (SVM) is an elegant approach for building complex and accurate models with structured outputs. However, its applicability relies on the availability of efficient inference algorithms--the state-of-the-art training algorithms repeatedly perform inference to compute a subgradient or to find the most violating configuration. In this paper, we propose an exact inference algorithm for maximizing nondecomposable objectives due to special type of a high-order potential having a decomposable internal structure. As an important application, our method covers the loss augmented inference, which enables the slack and margin scaling formulations of structural SVM with a variety of dissimilarity measures, e.g., Hamming loss, precision and recall, Fβ-loss, intersection over union, and many other functions that can be efficiently computed from the contingency table. We demonstrate the advantages of our approach in natural language parsing and sequence segmentation applications. Alexander Bauer 0001, Shinichi Nakajima, Klaus-Robert Müller |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Evaluating the Visualization of What a Deep Neural Network Has LearnedabstractDeep neural networks (DNNs) have demonstrated impressive performance in complex machine learning tasks such as image classification or speech recognition. However, due to their multilayer nonlinear structure, they are not transparent, i.e., it is hard to grasp what makes them arrive at a particular classification or recognition decision, given a new unseen data sample. Recently, several approaches have been proposed enabling one to understand and interpret the reasoning embodied in a DNN for a single test image. These methods quantify the “importance” of individual pixels with respect to the classification decision and allow a visualization in terms of a heatmap in pixel/input space. While the usefulness of heatmaps can be judged subjectively by a human, an objective quality measure is missing. In this paper, we present a general methodology based on region perturbation for evaluating ordered collections of pixels such as heatmaps. We compare heatmaps computed by three different methods on the SUN397, ILSVRC2012, and MIT Places data sets. Our main result is that the recently proposed layer-wise relevance propagation algorithm qualitatively and quantitatively provides a better explanation of what made a DNN arrive at a particular classification decision than the sensitivity-based approach or the deconvolution method. We provide theoretical arguments to explain this result and discuss its practical implications. Finally, we investigate the use of heatmaps for unsupervised assessment of the neural network performance. Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, Klaus-Robert Müller |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2016 | Analyzing Classifiers: Fisher Vectors and Deep Neural NetworksabstractFisher vector (FV) classifiers and Deep Neural Networks (DNNs) are popular and successful algorithms for solving image classification problems. However, both are generally considered 'black box' predictors as the non-linear transformations involved have so far prevented transparent and interpretable reasoning. Recently, a principled technique, Layer-wise Relevance Propagation (LRP), has been developed in order to better comprehend the inherent structured reasoning of complex nonlinear classification models such as Bag of Feature models or DNNs. In this paper we (1) extend the LRP framework also for Fisher vector classifiers and then use it as analysis tool to (2) quantify the importance of context for classification, (3) qualitatively compare DNNs against FV classifiers in terms of important image regions and (4) detect potential flaws and biases in data. All experiments are performed on the PASCAL VOC 2007 and ILSVRC 2012 data sets. Sebastian Lapuschkin, Alexander Binder, Grégoire Montavon, Klaus-Robert Müller, Wojciech Samek |
CVPR | 4 |
| 2016 | Layer-Wise Relevance Propagation for Neural Networks with Local Renormalization Layers
Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, Klaus-Robert Müller, Wojciech Samek |
ICANN (2) | 4 |
| 2016 | Controlling explanatory heatmap resolution and semantics via decomposition depthabstractWe present an application of the Layer-wise Relevance Propagation (LRP) algorithm to state of the art deep convolutional neural networks and Fisher Vector classifiers to compare the image perception and prediction strategies of both classifiers with the use of visualized heatmaps. Layer-wise Relevance Propagation (LRP) is a method to compute scores for individual components of an input image, denoting their contribution to the prediction of the classifier for one particular test point. We demonstrate the impact of different choices of decomposition cut-off points during the LRP-process, controlling the resolution and semantics of the heatmap on test images from the PASCAL VOC 2007 test data set. Sebastian Lapuschkin, Alexander Binder, Klaus-Robert Müller, Wojciech Samek |
ICIP | 3 |
| 2016 | Wasserstein Training of Restricted Boltzmann MachinesabstractBoltzmann machines are able to learn highly complex, multimodal, structured and multiscale real-world data distributions. Parameters of the model are usually learned by minimizing the Kullback-Leibler (KL) divergence from training samples to the learned model. We propose in this work a novel approach for Boltzmann machine training which assumes that a meaningful metric between observations is known. This metric between observations can then be used to define the Wasserstein distance between the distribution induced by the Boltzmann machine on the one hand, and that given by the training sample on the other hand. We derive a gradient of that distance with respect to the model parameters. Minimization of this new objective leads to generative models with different statistical properties. We demonstrate their practical potential on data completion and denoising, for which the metric between observations plays a crucial role. Grégoire Montavon, Klaus-Robert Müller, Marco Cuturi |
NIPS | 2 |
| 2016 | Neural network-based full-reference image quality assessmentabstractS.334-338 Sebastian Bosse, Dominique Maniry, Klaus-Robert Müller, Thomas Wiegand 0001, Wojciech Samek |
PCS | 3 |
| 2016 | Block adaptive selection of multiple core transforms for video codingabstractTransform coding tools in video coding have traditionally relied on the Discrete Cosine Transform Type II (DCT-II) to map residual signals to a new domain where quantization and entropy coding tools achieve a better coding efficiency than in the spatial domain. However, the DCT-II is not sufficient to model all different types of residual signals efficiently, especially in the intra-predicted blocks case. For this reason, the DST-VII was introduced in H.265/High Efficiency Video Coding (HEVC) in order to improve the compression performance of 4 × 4 intra-predicted blocks. In this paper we propose a multiple core transform approach, in which each transform is separable and generated by combining two one-dimensional transforms for the vertical and horizontal directions. The pair of 1-D transforms is selected from a set of three different types of Discrete Trigonometric Transforms and the Identity Transformation. Test results show that the proposed algorithm achieves bit rate reductions of 3% on average with respect to HEVC for intra-predicted residuals. Santiago De-Luxán-Hernández, Detlev Marpe, Heiko Schwarz, Klaus-Robert Müller, Mathias Wien, Jens-Rainer Ohm, Thomas Wiegand 0001 |
PCS | 4 |
| 2016 | Brain-Computer Interfacing for multimedia quality assessmentabstractThe assessment of perceived multimedia quality is a central research field in information and media technology. Conventionally, psychophysical techniques are used for determining the quality of multimedia signals. Recently, Brain-Computer Interfacing (BCI)-based methods have been proposed for the assessment of perceived multimedia signal quality. In this paper we give an overview over the shortcomings of conventional approaches, present the state-of-the art of BCI-based methods and discuss open questions and challenges relevant to the BCI community. Sebastian Bosse, Klaus-Robert Müller, Thomas Wiegand 0001, Wojciech Samek |
SMC | 2 |
| 2016 | Alternative CSP approaches for multimodal distributed BCI dataabstractBrain-Computer Interfaces (BCIs) are trained to distinguish between two (or more) mental states, e.g., left and right hand motor imagery, from the recorded brain signals. Common Spatial Patterns (CSP) is a popular method to optimally separate data from two motor imagery tasks under the assumption of an unimodal class distribution. In out of lab environments where users are distracted by additional noise sources this assumption may not hold. This paper systematically investigates BCI performance under such distractions and proposes two novel CSP variants, ensemble CSP and 2-step CSP, which can cope with multimodal class distributions. The proposed algorithms are evaluated using simulations and BCI data of 16 healthy participants performing motor imagery under 6 different types of distraction. Both methods are shown to significantly enhance the performance compared to the standard procedure. Stephanie Brandl, Klaus-Robert Müller, Wojciech Samek |
SMC | 2 |
| 2016 | The LRP Toolbox for Artificial Neural NetworksabstractThe Layer-wise Relevance Propagation (LRP) algorithm explains a classifier's prediction specific to a given data point by attributing relevance scores to important components of the input by using the topology of the learned model itself. With the LRP Toolbox we provide platform-agnostic implementations for explaining the predictions of pre-trained state of the art Caffe networks and stand-alone implementations for fully connected Neural Network models. The implementations for Matlab and python shall serve as a playing field to familiarize oneself with the LRP algorithm and are implemented with readability and transparency in mind. Models and data can be imported and exported using raw text formats, Matlab's .mat files and the .npy format for numpy or plain text. Sebastian Lapuschkin, Alexander Binder, Grégoire Montavon, Klaus-Robert Müller, Wojciech Samek |
J. Mach. Learn. Res. | 4 |
| 2016 | Why Does a Hilbertian Metric Work Efficiently in Online Learning With Kernels?abstractThe autocorrelation matrix of the kernelized input vector is well approximated by the squared Gram matrix (scaled down by the dictionary size). This holds true under the condition that the input covariance matrix in the feature space is approximated by its sample estimate based on the dictionary elements, leading to a couple of fundamental insights into online learning with kernels. First, the eigenvalue spread of the autocorrelation matrix relevant to the hyperplane projection along affine subspace algorithm is approximately a square root of that for the kernel normalized least mean square algorithm. This clarifies the mechanism behind fast convergence due to the use of a Hilbertian metric. Second, for efficient function estimation, the dictionary needs to be constructed in general by taking into account the distribution of the input vector, so as to satisfy the condition. The theoretical results are justified by computer experiments. Masahiro Yukawa, Klaus-Robert Müller |
IEEE Signal Process. Lett. | 2 |
| 2015 | Opening the Black Box: Revealing Interpretable Sequence Motifs in Kernel-Based Learning Algorithms
Marina M.-C. Vidovic, Nico Görnitz, Klaus-Robert Müller, Gunnar Rätsch, Marius Kloft |
ECML/PKDD (2) | 3 |
| 2015 | Multivariate Machine Learning Methods for Fusing Multimodal Functional Neuroimaging DataabstractMultimodal data are ubiquitous in engineering, communications, robotics, computer vision, or more generally speaking in industry and the sciences. All disciplines have developed their respective sets of analytic tools to fuse the information that is available in all measured modalities. In this paper, we provide a review of classical as well as recent machine learning methods (specifically factor models) for fusing information from functional neuroimaging techniques such as: LFP, EEG, MEG, fNIRS, and fMRI. Early and late fusion scenarios are distinguished, and appropriate factor models for the respective scenarios are presented along with example applications from selected multimodal neuroimaging studies. Further emphasis is given to the interpretability of the resulting model parameters, in particular by highlighting how factor models relate to physical models needed for source localization. The methods we discuss allow for the extraction of information from neural data, which ultimately contributes to 1) better neuroscientific understanding; 2) enhance diagnostic performance; and 3) discover neural signals of interest that correlate maximally with a given cognitive paradigm. While we clearly study the multimodal functional neuroimaging challenge, the discussed machine learning techniques have a wide applicability, i.e., in general data fusion, and may thus be informative to the general interested reader. Sven Dähne, Felix Bießmann, Wojciech Samek, Stefan Haufe, Dominique Goltz, Christopher Gundlach, Arno Villringer, Siamac Fazli, Klaus-Robert Müller |
Proc. IEEE | 9 |
| 2015 | Learning From More Than One Data Source: Data Fusion Techniques for Sensorimotor Rhythm-Based Brain-Computer InterfacesabstractBrain-computer interfaces (BCIs) are successfully used in scientific, therapeutic and other applications. Remaining challenges are among others a low signal-to-noise ratio of neural signals, lack of robustness for decoders in the presence of inter-trial and inter-subject variability, time constraints on the calibration phase and the use of BCIs outside a controlled lab environment. Recent advances in BCI research addressed these issues by novel combinations of complementary analysis as well as recording techniques, so called hybrid BCIs. In this paper, we review a number of data fusion techniques for BCI along with hybrid methods for BCI that have recently emerged. Our focus will be on sensorimotor rhythm-based BCIs. We will give an overview of the three main lines of research in this area, integration of complementary features of neural activation, integration of multiple previous sessions and of multiple subjects, and show how these techniques can be used to enhance modern BCI systems. Siamac Fazli, Sven Dähne, Wojciech Samek, Felix Bießmann, Klaus-Robert Müller |
Proc. IEEE | 5 |
| 2015 | Towards Noninvasive Hybrid Brain-Computer Interfaces: Framework, Practice, Clinical Application, and BeyondabstractIn their early days, brain-computer interfaces (BCIs) were only considered as control channel for end users with severe motor impairments such as people in the locked-in state. But, thanks to the multidisciplinary progress achieved over the last decade, the range of BCI applications has been substantially enlarged. Indeed, today BCI technology cannot only translate brain signals directly into control signals, but also can combine such kind of artificial output with a natural muscle-based output. Thus, the integration of multiple biological signals for real-time interaction holds the promise to enhance a much larger population than originally thought end users with preserved residual functions who could benefit from new generations of assistive technologies. A BCI system that combines a BCI with other physiological or technical signals is known as hybrid BCI (hBCI). In this work, we review the work of a large scale integrated project funded by the European commission which was dedicated to develop practical hybrid BCIs and introduce them in various fields of applications. This article presents an hBCI framework, which was used in studies with nonimpaired as well as end users with motor impairments. Gernot R. Müller-Putz, Robert Leeb, Michael Tangermann, Johannes Höhne, Andrea Kübler, Febo Cincotti, Donatella Mattia, Rüdiger Rupp, Klaus-Robert Müller, José del R. Millán |
Proc. IEEE | 9 |
| 2015 | The Plurality of Human Brain-Computer Interfacing [Scanning the Issue]abstractThe articles in this special issue focus on brain-computer interfacing. The papers are dedicated to this growing and diversifying research enterprise, and features important review articles as well as some important current examples of research in this area. The field of brain-computer interface (BCI) research began to develop about 25 years ago and transformed from initially isolated demonstrations by a few groups into a large scientific enterprise that is currently producing hundreds of peer-reviewed articles and several dedicated conferences and workshops each year. This level of productivity is reflective of the large and continually growing enthusiasm by the scientific community, funding agencies, and the public. Gernot R. Müller-Putz, José del R. Millán, Gerwin Schalk, Klaus-Robert Müller |
Proc. IEEE | 4 |
| 2014 | Learning and Evaluation in Presence of Non-i.i.d. Label NoiseabstractIn many real-world applications, the simplified assumption of independent and identically distributed noise breaks down, and labels can have structured, systematic noise. For example, in brain-computer interface applications, training data is often the result of lengthy experimental sessions, where the attention levels of participants can change over the course of the experiment. In such application cases, structured label noise will cause problems because most machine learning methods assume independent and identically distributed label noise. In this paper, we present a novel methodology for learning and evaluation in presence of systematic label noise. The core of which is a novel extension of support vector data description / one-class SVM that can incorporate latent variables. Controlled simulations on synthetic data and a real-world EEG experiment with 20 subjects from the domain of brain-computer-interfacing show that our method achieves accuracies that go beyond the state of the art. Nico Görnitz, Anne Porbadnigk, Alexander Binder, Claudia Sannelli, Mikio L. Braun, Klaus-Robert Müller, Marius Kloft |
AISTATS | 6 |
| 2014 | Neurally informed assessment of perceived natural texture image qualityabstractConventionally, the quality of images and related codecs are assessed using subjective tests, such as Degradation Category Rating. These quality assessments consider the behavioral level only. Recently, it has been proposed to complement this approach by investigating how quality is processed in the brain of a user (using electroencephalography, EEG), potentially leading to results that are less biased by subjective factors. In this paper, a novel method is presented for assessing how image quality is processed on a neural level, using Steady-State Visual Evoked Potentials (SSVEPs) as EEG features. We tested our approach in an EEG study with 16 participants who were presented with distorted images of natural textures. Subsequently, we compared our approach analogously to the standardized Degradation Category Rating quality assessment. Remarkably, our novel method yields a correlation of |r| = 0.93 to MOS on the recorded dataset. Sebastian Bosse, Laura Acqualagna, Anne Porbadnigk, Benjamin Blankertz, Gabriel Curio, Klaus-Robert Müller, Thomas Wiegand 0001 |
ICIP | 6 |
| 2014 | Covariance shrinkage for autocorrelated data
Daniel Bartz, Klaus-Robert Müller |
NIPS | 2 |
| 2014 | Robust Common Spatial Filters with a Maxmin ApproachabstractElectroencephalographic signals are known to be nonstationary and easily affected by artifacts; therefore, their analysis requires methods that can deal with noise. In this work, we present a way to robustify the popular common spatial patterns (CSP) algorithm under a maxmin approach. In contrast to standard CSP that maximizes the variance ratio between two conditions based on a single estimate of the class covariance matrices, we propose to robustly compute spatial filters by maximizing the minimum variance ratio within a prefixed set of covariance matrices called the tolerance set. We show that this kind of maxmin optimization makes CSP robust to outliers and reduces its tendency to overfit. We also present a data-driven approach to construct a tolerance set that captures the variability of the covariance matrices over time and shows its ability to reduce the nonstationarity of the extracted features and significantly improve classification accuracy. We test the spatial filters derived with this approach and compare them to standard CSP and a state-of-the-art method on a real-world brain-computer interface (BCI) data set in which we expect substantial fluctuations caused by environmental differences. Finally we investigate the advantages and limitations of the maxmin approach with simulations. Motoaki Kawanabe, Wojciech Samek, Klaus-Robert Müller, Carmen Vidaurre |
Neural Comput. | 3 |
| 2014 | Efficient Algorithms for Exact Inference in Sequence Labeling SVMsabstractThe task of structured output prediction deals with learning general functional dependencies between arbitrary input and output spaces. In this context, two loss-sensitive formulations for maximum-margin training have been proposed in the literature, which are referred to as margin and slack rescaling, respectively. The latter is believed to be more accurate and easier to handle. Nevertheless, it is not popular due to the lack of known efficient inference algorithms; therefore, margin rescaling--which requires a similar type of inference as normal structured prediction--is the most often used approach. Focusing on the task of label sequence learning, we here define a general framework that can handle a large class of inference problems based on Hamming-like loss functions and the concept of decomposability for the underlying joint feature map. In particular, we present an efficient generic algorithm that can handle both rescaling approaches and is guaranteed to find an optimal solution in polynomial time. Alexander Bauer 0001, Nico Görnitz, Franziska Biegler, Klaus-Robert Müller, Marius Kloft |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2013 | Generalizing Analytic Shrinkage for Arbitrary Covariance StructuresabstractAnalytic shrinkage is a statistical technique that offers a fast alternative to cross-validation for the regularization of covariance matrices and has appealing consistency properties. We show that the proof of consistency implies bounds on the growth rates of eigenvalues and their dispersion, which are often violated in data. We prove consistency under assumptions which do not restrict the covariance structure and therefore better match real world data. In addition, we propose an extension of analytic shrinkage --orthogonal complement shrinkage-- which adapts to the covariance structure. Finally we demonstrate the superior performance of our novel approach on data from the domains of finance, spoken letter and optical character recognition, and neuroscience. Daniel Bartz, Klaus-Robert Müller |
NIPS | 2 |
| 2013 | Robust Spatial Filtering with Beta DivergenceabstractThe efficiency of Brain-Computer Interfaces (BCI) largely depends upon a reliable extraction of informative features from the high-dimensional EEG signal. A crucial step in this protocol is the computation of spatial filters. The Common Spatial Patterns (CSP) algorithm computes filters that maximize the difference in band power between two conditions, thus it is tailored to extract the relevant information in motor imagery experiments. However, CSP is highly sensitive to artifacts in the EEG data, i.e. few outliers may alter the estimate drastically and decrease classification performance. Inspired by concepts from the field of information geometry we propose a novel approach for robustifying CSP. More precisely, we formulate CSP as a divergence maximization problem and utilize the property of a particular type of divergence, namely beta divergence, for robustifying the estimation of spatial filters in the presence of artifacts in the data. We demonstrate the usefulness of our method on toy data and on EEG recordings from 80 subjects. Wojciech Samek, Duncan A. J. Blythe, Klaus-Robert Müller, Motoaki Kawanabe |
NIPS | 3 |
| 2013 | Enhanced representation and multi-task learning for image annotation
Alexander Binder, Wojciech Samek, Klaus-Robert Müller, Motoaki Kawanabe |
Comput. Vis. Image Underst. | 3 |
| 2013 | Integration of Multivariate Data Streams With Bandpower SignalsabstractThe urge to further our understanding of multimodal neural data has recently become an important topic due to the ever increasing availability of simultaneously recorded data from different neural imaging modalities. In case where EEG is one of the modalities, it is of interest to relate a nonlinear function of the raw EEG time-domain signal, say, EEG band power, to another modality such as the hemodynamic response, as measured with NIRS or fMRI. In this work we tackle exactly this problem defining a novel algorithm that we denote multimodal source power correlation analysis (mSPoC). The validity and high performance of the mSPoC framework is demonstrated for simulated and real-world multimodal data. Sven Dähne, Felix Bießmann, Frank C. Meinecke, Jan Mehnert, Siamac Fazli, Klaus-Robert Müller |
IEEE Trans. Multim. | 6 |
| 2012 | Learning Invariant Representations of Molecules for Atomization Energy PredictionabstractThe accurate prediction of molecular energetics in chemical compound space is a crucial ingredient for rational compound design. The inherently graph-like, non-vectorial nature of molecular data gives rise to a unique and difficult machine learning problem. In this paper, we adopt a learning-from-scratch approach where quantum-mechanical molecular energies are predicted directly from the raw molecular geometry. The study suggests a benefit from setting flexible priors and enforcing invariance stochastically rather than structurally. Our results improve the state-of-the-art by a factor of almost three, bringing statistical methods one step closer to the holy grail of ''chemical accuracy''. Grégoire Montavon, Katja Hansen, Siamac Fazli, Matthias Rupp, Franziska Biegler, Andreas Ziehe, Alexandre Tkatchenko, O. Anatole von Lilienfeld, Klaus-Robert Müller |
NIPS | 9 |
| 2012 | On Taxonomies for Multi-class Image CategorizationabstractWe study the problem of classifying images into a given, pre-determined taxonomy. This task can be elegantly translated into the structured learning framework. However, despite its power, structured learning has known limits in scalability due to its high memory requirements and slow training process. We propose an efficient approximation of the structured learning approach by an ensemble of local support vector machines (SVMs) that can be trained efficiently with standard techniques. A first theoretical discussion and experiments on toy-data allow to shed light onto why taxonomy-based classification can outperform taxonomy-free approaches and why an appropriately combined ensemble of local SVMs might be of high practical use. Further empirical results on subsets of Caltech256 and VOC2006 data indeed show that our local SVM formulation can effectively exploit the taxonomy structure and thus outperforms standard multi-class classification algorithms while it achieves on par results with taxonomy-based structured algorithms at a significantly decreased computing time. Alexander Binder, Klaus-Robert Müller, Motoaki Kawanabe |
Int. J. Comput. Vis. | 2 |
| 2012 | Algebraic Geometric Comparison of Probability Distributions
Franz J. Király, Paul von Bünau, Frank C. Meinecke, Duncan A. J. Blythe, Klaus-Robert Müller |
J. Mach. Learn. Res. | 5 |
| 2012 | Toward a Direct Measure of Video Quality Perception Using EEGabstractAn approach to the direct measurement of perception of video quality change using electroencephalography (EEG) is presented. Subjects viewed 8-s video clips while their brain activity was registered using EEG. The video signal was either uncompressed at full length or changed from uncompressed to a lower quality level at a random time point. The distortions were introduced by a hybrid video codec. Subjects had to indicate whether they had perceived a quality change. In response to a quality change, a positive voltage change in EEG (the so-called P3 component) was observed at latency of about 400-600 ms for all subjects. The voltage change positively correlated with the magnitude of the video quality change, substantiating the P3 component as a graded neural index of the perception of video quality change within the presented paradigm. By applying machine learning techniques, we could classify on a single-trial basis whether a subject perceived a quality change. Interestingly, some video clips wherein changes were missed (i.e., not reported) by the subject were classified as quality changes, suggesting that the brain detected a change, although the subject did not press a button. In conclusion, abrupt changes of video quality give rise to specific components in the EEG that can be detected on a single-trial basis. Potentially, a neurotechnological approach to video assessment could lead to a more objective quantification of quality change detection, overcoming the limitations of subjective approaches (such as subjective bias and the requirement of an overt response). Furthermore, it allows for real-time applications wherein the brain response to a video clip is monitored while it is being viewed. Simon Scholler, Sebastian Bosse, Matthias Treder, Benjamin Blankertz, Gabriel Curio, Klaus-Robert Müller, Thomas Wiegand 0001 |
IEEE Trans. Image Process. | 6 |
| 2012 | Feature Extraction for Change-Point Detection Using Stationary Subspace AnalysisabstractDetecting changes in high-dimensional time series is difficult because it involves the comparison of probability densities that need to be estimated from finite samples. In this paper, we present the first feature extraction method tailored to change-point detection, which is based on an extended version of stationary subspace analysis. We reduce the dimensionality of the data to the most nonstationary directions, which are most informative for detecting state changes in the time series. In extensive simulations on synthetic data, we show that the accuracy of three change-point detection algorithms is significantly increased by a prior feature extraction step. These findings are confirmed in an application to industrial fault monitoring. Duncan A. J. Blythe, Paul von Bünau, Frank C. Meinecke, Klaus-Robert Müller |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2011 | ℓ1-Penalized Linear Mixed-Effects Models for BCI
Siamac Fazli, Márton Danóczy, Jürg Schelldorfer, Klaus-Robert Müller |
ICANN (1) | 4 |
| 2011 | Kernel Analysis of Deep Networks
Grégoire Montavon, Mikio L. Braun, Klaus-Robert Müller |
J. Mach. Learn. Res. | 3 |
| 2011 | The Stationary Subspace Analysis Toolbox
Jan Saputra Müller, Paul von Bünau, Frank C. Meinecke, Franz J. Király, Klaus-Robert Müller |
J. Mach. Learn. Res. | 5 |
| 2011 | Machine-Learning-Based Coadaptive Calibration for Brain-Computer InterfacesabstractBrain-computer interfaces (BCIs) allow users to control a computer application by brain activity as acquired (e.g., by EEG). In our classic machine learning approach to BCIs, the participants undertake a calibration measurement without feedback to acquire data to train the BCI system. After the training, the user can control a BCI and improve the operation through some type of feedback. However, not all BCI users are able to perform sufficiently well during feedback operation. In fact, a nonnegligible portion of participants (estimated 15%-30%) cannot control the system (a BCI illiteracy problem, generic to all motor-imagery-based BCIs). We hypothesize that one main difficulty for a BCI user is the transition from offline calibration to online feedback. In this work, we investigate adaptive machine learning methods to eliminate offline calibration and analyze the performance of 11 volunteers in a BCI based on the modulation of sensorimotor rhythms. We present an adaptation scheme that individually guides the user. It starts with a subject-independent classifier that evolves to a subject-optimized state-of-the-art classifier within one session while the user interacts continuously. These initial runs use supervised techniques for robust coadaptive learning of user and machine. Subsequent runs use unsupervised adaptation to track the features' drift during the session and provide an unbiased measure of BCI performance. Using this approach, without any offline calibration, six users, including one novice, obtained good performance after 3 to 6 minutes of adaptation. More important, this novel guided learning also allows participants with BCI illiteracy to gain significant control with the BCI in less than 60 minutes. In addition, one volunteer without sensorimotor idle rhythm peak at the beginning of the BCI experiment developed it during the course of the session and used voluntary modulation of its amplitude to control the feedback application. Carmen Vidaurre, Claudia Sannelli, Klaus-Robert Müller, Benjamin Blankertz |
Neural Comput. | 3 |
| 2010 | Layer-wise analysis of deep networks with Gaussian kernelsabstractDeep networks can potentially express a learning problem more efficiently than local learning machines. While deep networks outperform local learning machines on some problems, it is still unclear how their nice representation emerges from their complex structure. We present an analysis based on Gaussian kernels that measures how the representation of the learning problem evolves layer after layer as the deep network builds higher-level abstract representations of the input. We use this analysis to show empirically that deep networks build progressively better representations of the learning problem and that the best representations are obtained when the deep network discriminates only in the last layers. Grégoire Montavon, Mikio L. Braun, Klaus-Robert Müller |
NIPS | 3 |
| 2010 | How to Explain Individual Classification Decisions
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, Klaus-Robert Müller |
J. Mach. Learn. Res. | 6 |
| 2010 | Approximate Tree Kernels
Konrad Rieck, Tammo Krueger, Ulf Brefeld, Klaus-Robert Müller |
J. Mach. Learn. Res. | 4 |
| 2010 | Temporal kernel CCA and its application in multimodal neuronal data analysisabstractData recorded from multiple sources sometimes exhibit non-instantaneous couplings. For simple data sets, cross-correlograms may reveal the coupling dynamics. But when dealing with high-dimensional multivariate data there is no such measure as the cross-correlogram. We propose a simple algorithm based on Kernel Canonical Correlation Analysis (kCCA) that computes a multivariate temporal filter which links one data modality to another one. The filters can be used to compute a multivariate extension of the cross-correlogram, the canonical correlogram, between data sources that have different dimensionalities and temporal resolutions. The canonical correlogram reflects the coupling dynamics between the two sources. The temporal filter reveals which features in the data give rise to these couplings and when they do so. We present results from simulations and neuroscientific experiments showing that tkCCA yields easily interpretable temporal filters and correlograms. In the experiments, we simultaneously performed electrode recordings and functional magnetic resonance imaging (fMRI) in primary visual cortex of the non-human primate. While electrode recordings reflect brain activity directly, fMRI provides only an indirect view of neural activity via the Blood Oxygen Level Dependent (BOLD) response. Thus it is crucial for our understanding and the interpretation of fMRI signals in general to relate them to direct measures of neural activity acquired with electrodes. The results computed by tkCCA confirm recent models of the hemodynamic response to neural activity and allow for a more detailed analysis of neurovascular coupling dynamics. Felix Bießmann, Frank C. Meinecke, Arthur Gretton, Alexander Rauch, Gregor Rainer, Nikos K. Logothetis, Klaus-Robert Müller |
Mach. Learn. | 7 |
| 2009 | Subject independent EEG-based BCI decodingabstractIn the quest to make Brain Computer Interfacing (BCI) more usable, dry electrodes have emerged that get rid of the initial 30 minutes required for placing an electrode cap. Another time consuming step is the required individualized adaptation to the BCI user, which involves another 30 minutes calibration for assessing a subjects brain signature. In this paper we aim to also remove this calibration proceedure from BCI setup time by means of machine learning. In particular, we harvest a large database of EEG BCI motor imagination recordings (83 subjects) for constructing a library of subject-specific spatio-temporal filters and derive a subject independent BCI classifier. Our offline results indicate that BCI-na{i}ve users could start real-time BCI use with no prior calibration at only a very moderate performance loss." Siamac Fazli, Cristian Grozea, Márton Danóczy, Benjamin Blankertz, Florin Popescu, Klaus-Robert Müller |
NIPS | 6 |
| 2009 | Efficient and Accurate Lp-Norm Multiple Kernel LearningabstractLearning linear combinations of multiple kernels is an appealing strategy when the right choice of features is unknown. Previous approaches to multiple kernel learning (MKL) promote sparse kernel combinations and hence support interpretability. Unfortunately, L1-norm MKL is hardly observed to outperform trivial baselines in practical applications. To allow for robust kernel mixtures, we generalize MKL to arbitrary Lp-norms. We devise new insights on the connection between several existing MKL formulations and develop two efficient interleaved optimization strategies for arbitrary p>1. Empirically, we demonstrate that the interleaved optimization strategies are much faster compared to the traditionally used wrapper approaches. Finally, we apply Lp-norm MKL to real-world problems from computational biology, showing that non-sparse MKL achieves accuracies that go beyond the state-of-the-art. Marius Kloft, Ulf Brefeld, Sören Sonnenburg, Pavel Laskov, Klaus-Robert Müller, Alexander Zien |
NIPS | 5 |
| 2009 | Designing for uncertain, asymmetric control: Interaction design for brain-computer interfaces
John Williamson 0001, Roderick Murray-Smith, Benjamin Blankertz, Matthias Krauledat, Klaus-Robert Müller |
Int. J. Hum. Comput. Stud. | 5 |
| 2009 | Subject-independent mental state classification in single trials
Siamac Fazli, Florin Popescu, Márton Danóczy, Benjamin Blankertz, Klaus-Robert Müller, Cristian Grozea |
Neural Networks | 5 |
| 2009 | Recent advances in brain-machine interfaces
Tadashi Isa, Eberhard E. Fetz, Klaus-Robert Müller |
Neural Networks | 3 |
| 2009 | Improving BCI performance by task-related trial pruning
Claudia Sannelli, Mikio L. Braun, Klaus-Robert Müller |
Neural Networks | 3 |
| 2009 | A Generalized Framework for Quantifying the Dynamics of EEG Event-Related DesynchronizationabstractBrains were built by evolution to react swiftly to environmental challenges. Thus, sensory stimuli must be processed ad hoc, i.e., independent--to a large extent--from the momentary brain state incidentally prevailing during stimulus occurrence. Accordingly, computational neuroscience strives to model the robust processing of stimuli in the presence of dynamical cortical states. A pivotal feature of ongoing brain activity is the regional predominance of EEG eigenrhythms, such as the occipital alpha or the pericentral mu rhythm, both peaking spectrally at 10 Hz. Here, we establish a novel generalized concept to measure event-related desynchronization (ERD), which allows one to model neural oscillatory dynamics also in the presence of dynamical cortical states. Specifically, we demonstrate that a somatosensory stimulus causes a stereotypic sequence of first an ERD and then an ensuing amplitude overshoot (event-related synchronization), which at a dynamical cortical state becomes evident only if the natural relaxation dynamics of unperturbed EEG rhythms is utilized as reference dynamics. Moreover, this computational approach also encompasses the more general notion of a "conditional ERD," through which candidate explanatory variables can be scrutinized with regard to their possible impact on a particular oscillatory dynamics under study. Thus, the generalized ERD represents a powerful novel analysis tool for extending our understanding of inter-trial variability of evoked responses and therefore the robust processing of environmental stimuli. Steven Lemm, Klaus-Robert Müller, Gabriel Curio |
PLoS Comput. Biol. | 2 |
| 2008 | Stopping conditions for exact computation of leave-one-out error in support vector machinesabstractWe propose a new stopping condition for a Support Vector Machine (SVM) solver which precisely reflects the objective of the Leave-One-Out error computation. The stopping condition guarantees that the output on an intermediate SVM solution is identical to the output of the optimal SVM solution with one data point excluded from the training set. A simple augmentation of a general SVM training algorithm allows one to use a stopping criterion equivalent to the proposed sufficient condition. A comprehensive experimental evaluation of our method shows consistent speedup of the exact LOO computation by our method, up to the factor of 13 for the linear kernel. The new algorithm can be seen as an example of constructive guidance of an optimization algorithm towards achieving the best attainable expected risk at optimal computational cost. Vojtech Franc, Pavel Laskov, Klaus-Robert Müller |
ICML | 3 |
| 2008 | Estimating vector fields using sparse basis field expansionsabstractWe introduce a novel framework for estimating vector fields using sparse basis field expansions (S-FLEX). The notion of basis fields, which are an extension of scalar basis functions, arises naturally in our framework from a rotational invariance requirement. We consider a regression setting as well as inverse problems. All variants discussed lead to second-order cone programming formulations. While our framework is generally applicable to any type of vector field, we focus in this paper on applying it to solving the EEG/MEG inverse problem. It is shown that significantly more precise and neurophysiologically more plausible location and shape estimates of cerebral current sources from EEG/MEG measurements become possible with our method when comparing to the state-of-the-art. Stefan Haufe, Vadim V. Nikulin, Andreas Ziehe, Klaus-Robert Müller, Guido Nolte |
NIPS | 4 |
| 2008 | Playing Pinball with non-invasive BCIabstractCompared to invasive Brain-Computer Interfaces (BCI), non-invasive BCI systems based on Electroencephalogram (EEG) signals have not been applied successfully for complex control tasks. In the present study, however, we demonstrate this is possible and report on the interaction of a human subject with a complex real device: a pinball machine. First results in this single subject study clearly show that fast and well-timed control well beyond chance level is possible, even though the environment is extremely rich and requires complex predictive behavior. Using machine learning methods for mental state decoding, BCI-based pinball control is possible within the first session without the necessity to employ lengthy subject training. While the current study is still of anecdotal nature, it clearly shows that very compelling control with excellent timing and dynamics is possible for a non-invasive BCI. Michael Tangermann, Matthias Krauledat, Konrad Grzeska, Max Sagebaum, Benjamin Blankertz, Carmen Vidaurre, Klaus-Robert Müller |
NIPS | 7 |
| 2008 | On Relevant Dimensions in Kernel Feature Spaces
Mikio L. Braun, Joachim M. Buhmann, Klaus-Robert Müller |
J. Mach. Learn. Res. | 3 |
| 2007 | Asymptotic Bayesian generalization error when training and test distributions are differentabstractIn supervised learning, we commonly assume that training and test data are sampled from the same distribution. However, this assumption can be violated in practice and then standard machine learning techniques perform poorly. This paper focuses on revealing and improving the performance of Bayesian estimation when the training and test distributions are different. We formally analyze the asymptotic Bayesian generalization error and establish its upper bound under a very general setting. Our important finding is that lower order terms---which can be ignored in the absence of the distribution change---play an important role under the distribution change. We also propose a novel variant of stochastic complexity which can be used for choosing an appropriate model and hyper-parameters under a particular distribution change. Keisuke Yamazaki, Motoaki Kawanabe, Sumio Watanabe, Masashi Sugiyama, Klaus-Robert Müller |
ICML | 5 |
| 2007 | Invariant Common Spatial Patterns: Alleviating Nonstationarities in Brain-Computer InterfacingabstractBrain-Computer Interfaces can suffer from a large variance of the subject condi- tions within and across sessions. For example vigilance fluctuations in the indi- vidual, variable task involvement, workload etc. alter the characteristics of EEG signals and thus challenge a stable BCI operation. In the present work we aim to define features based on a variant of the common spatial patterns (CSP) algorithm that are constructed invariant with respect to such nonstationarities. We enforce invariance properties by adding terms to the denominator of a Rayleigh coefficient representation of CSP such as disturbance covariance matrices from fluctuations in visual processing. In this manner physiological prior knowledge can be used to shape the classification engine for BCI. As a proof of concept we present a BCI classifier that is robust to changes in the level of parietal a -activity. In other words, the EEG decoding still works when there are lapses in vigilance. Benjamin Blankertz, Motoaki Kawanabe, Ryota Tomioka, Friederike U. Hohlefeld, Vadim V. Nikulin, Klaus-Robert Müller |
NIPS | 6 |
| 2007 | Heterogeneous Component AnalysisabstractIn bioinformatics it is often desirable to combine data from various measurement sources and thus structured feature vectors are to be analyzed that possess different intrinsic blocking characteristics (e.g., different patterns of missing values, obser- vation noise levels, effective intrinsic dimensionalities). We propose a new ma- chine learning tool, heterogeneous component analysis (HCA), for feature extrac- tion in order to better understand the factors that underlie such complex structured heterogeneous data. HCA is a linear block-wise sparse Bayesian PCA based not only on a probabilistic model with block-wise residual variance terms but also on a Bayesian treatment of a block-wise sparse factor-loading matrix. We study vari- ous algorithms that implement our HCA concept extracting sparse heterogeneous structure by obtaining common components for the blocks and specific compo- nents within each block. Simulations on toy and bioinformatics data underline the usefulness of the proposed structured matrix factorization concept. Shigeyuki Oba, Motoaki Kawanabe, Klaus-Robert Müller, Shin Ishii |
NIPS | 3 |
| 2007 | Berlin Brain-Computer Interface - The HCI communication channel for discovery
Roman Krepki, Gabriel Curio, Benjamin Blankertz, Klaus-Robert Müller |
Int. J. Hum. Comput. Stud. | 4 |
| 2007 | The Need for Open Source Software in Machine Learning
Sören Sonnenburg, Mikio L. Braun, Cheng Soon Ong, Samy Bengio, Léon Bottou, Geoff Holmes 0001, Yann LeCun, Klaus-Robert Müller, Fernando Pereira 0003, Carl E. Rasmussen, Gunnar Rätsch, Bernhard Schölkopf, Alexander J. Smola, Pascal Vincent, Jason Weston, Robert C. Williamson |
J. Mach. Learn. Res. | 8 |
| 2007 | Covariate Shift Adaptation by Importance Weighted Cross Validation
Masashi Sugiyama, Matthias Krauledat, Klaus-Robert Müller |
J. Mach. Learn. Res. | 3 |
| 2007 | Optimal dyadic decision trees
Gilles Blanchard, Christin Schäfer, Yves Rozenholc, Klaus-Robert Müller |
Mach. Learn. | 4 |
| 2007 | The Berlin Brain-Computer Interface (BBCI) - towards a new communication channel for online control in gaming applications
Roman Krepki, Benjamin Blankertz, Gabriel Curio, Klaus-Robert Müller |
Multim. Tools Appl. | 4 |
| 2007 | Improving the Caenorhabditis elegans Genome Annotation Using Machine LearningabstractFor modern biology, precise genome annotations are of prime importance, as they allow the accurate definition of genic regions. We employ state-of-the-art machine learning methods to assay and improve the accuracy of the genome annotation of the nematode Caenorhabditis elegans. The proposed machine learning system is trained to recognize exons and introns on the unspliced mRNA, utilizing recent advances in support vector machines and label sequence learning. In 87% (coding and untranslated regions) and 95% (coding regions only) of all genes tested in several out-of-sample evaluations, our method correctly identified all exons and introns. Notably, only 37% and 50%, respectively, of the presently unconfirmed genes in the C. elegans genome annotation agree with our predictions, thus we hypothesize that a sizable fraction of those genes are not correctly annotated. A retrospective evaluation of the Wormbase WS120 annotation [] of C. elegans reveals that splice form predictions on unconfirmed genes in WS120 are inaccurate in about 18% of the considered cases, while our predictions deviate from the truth only in 10%-13%. We experimentally analyzed 20 controversial genes on which our system and the annotation disagree, confirming the superiority of our predictions. While our method correctly predicted 75% of those cases, the standard annotation was never completely correct. The accuracy of our system is further corroborated by a comparison with two other recently proposed systems that can be used for splice form prediction: SNAP and ExonHunter. We conclude that the genome annotation of C. elegans and other organisms can be greatly enhanced using modern machine learning technology. Gunnar Rätsch, Sören Sonnenburg, Jagan Srinivasan, Hanh Witte, Klaus-Robert Müller, Ralf J. Sommer, Bernhard Schölkopf |
PLoS Comput. Biol. | 5 |
| 2006 | A Model Selection Method Based on Bound of Learning Coefficient
Keisuke Yamazaki, Kenji Nagata, Sumio Watanabe, Klaus-Robert Müller |
ICANN (2) | 4 |
| 2006 | Obtaining the Best Linear Unbiased Estimator of Noisy Signals by Non-Gaussian Component AnalysisabstractObtaining the best linear unbiased estimator (BLUE) of noisy signals is a traditional but powerful approach to noise reduction. Explicitly computing BLUE usually requires the prior knowledge of the subspace to which the true signal belongs and the noise covariance matrix. However, such prior knowledge is often unavailable in reality, which prevents us from applying BLUE to real-world problems. In this paper, we therefore give a method for obtaining BLUE without such prior knowledge. Our additional assumption is that the true signal follows a non-Gaussian distribution while the noise is Gaussian Masashi Sugiyama, Motoaki Kawanabe, Gilles Blanchard, Vladimir G. Spokoiny, Klaus-Robert Müller |
ICASSP (3) | 5 |
| 2006 | Denoising and Dimension Reduction in Feature SpaceabstractWe show that the relevant information about a classification problem in feature space is contained up to negligible error in a finite number of leading kernel PCA components if the kernel matches the underlying learning problem. Thus, kernels not only transform data sets such that good generalization can be achieved even by linear discriminant functions, but this transformation is also performed in a manner which makes economic use of feature space dimensions. In the best case, kernels provide efficient implicit representations of the data to perform classification. Practically, we propose an algorithm which enables us to recover the subspace and dimensionality relevant for good classification. Our algorithm can therefore be applied (1) to analyze the interplay of data set and kernel in a geometric fashion, (2) to help in model selection, and to (3) de-noise in feature space in order to yield better classification results. Mikio L. Braun, Joachim M. Buhmann, Klaus-Robert Müller |
NIPS | 3 |
| 2006 | Reducing Calibration Time For Brain-Computer Interfaces: A Clustering ApproachabstractUp to now even subjects that are experts in the use of machine learning based BCI systems still have to undergo a calibration session of about 20-30 min. From this data their (movement) intentions are so far infered. We now propose a new paradigm that allows to completely omit such calibration and instead transfer knowledge from prior sessions. To achieve this goal we first define normalized CSP features and distances in-between. Second, we derive prototypical features across sessions: (a) by clustering or (b) by feature concatenation methods. Finally, we construct a classifier based on these individualized prototypes and show that, indeed, classifiers can be successfully transferred to a new session for a number of subjects. Matthias Krauledat, Michael Tangermann, Benjamin Blankertz, Klaus-Robert Müller |
NIPS | 4 |
| 2006 | Inducing Metric Violations in Human Similarity JudgementsabstractAttempting to model human categorization and similarity judgements is both a very interesting but also an exceedingly difficult challenge. Some of the difficulty arises because of conflicting evidence whether human categorization and similarity judgements should or should not be modelled as to operate on a mental representation that is essentially metric. Intuitively, this has a strong appeal as it would allow (dis)similarity to be represented geometrically as distance in some internal space. Here we show how a single stimulus, carefully constructed in a psychophysical experiment, introduces l2 violations in what used to be an internal similarity space that could be adequately modelled as Euclidean. We term this one influential data point a conflictual judgement. We present an algorithm of how to analyse such data and how to identify the crucial point. Thus there may not be a strict dichotomy between either a metric or a non-metric internal space but rather degrees to which potentially large subsets of stimuli are represented metrically with a small subset causing a global violation of metricity. Julian Laub, Jakob H. Macke, Klaus-Robert Müller, Felix A. Wichmann |
NIPS | 3 |
| 2006 | Logistic Regression for Single Trial EEG ClassificationabstractWe propose a novel framework for the classification of single trial ElectroEncephaloGraphy (EEG), based on regularized logistic regression. Framed in this robust statistical framework no prior feature extraction or outlier removal is required. We present two variations of parameterizing the regression function: (a) with a full rank symmetric matrix coefficient and (b) as a difference of two rank=1 matrices. In the first case, the problem is convex and the logistic regression is optimal under a generative model. The latter case is shown to be related to the Common Spatial Pattern (CSP) algorithm, which is a popular technique in Brain Computer Interfacing. The regression coefficients can also be topographically mapped onto the scalp similarly to CSP pro jections, which allows neuro-physiological interpretation. Simulations on 162 BCI datasets demonstrate that classification accuracy and robustness compares favorably against conventional CSP based classifiers. Ryota Tomioka, Kazuyuki Aihara, Klaus-Robert Müller |
NIPS | 3 |
| 2006 | From outliers to prototypes: Ordering data
Stefan Harmeling, Guido Dornhege, David M. J. Tax, Frank C. Meinecke, Klaus-Robert Müller |
Neurocomputing | 5 |
| 2006 | In Search of Non-Gaussian Components of a High-Dimensional DistributionabstractFinding non-Gaussian components of high-dimensional data is an important preprocessing step for efficient information processing. This article proposes a new linear method to identify the "non-Gaussian subspace" within a very general semi-parametric framework. Our proposed method, called NGCA (non-Gaussian component analysis), is based on a linear operator which, to any arbitrary nonlinear (smooth) function, associates a vector belonging to the low dimensional non-Gaussian target subspace, up to an estimation error. By applying this operator to a family of different nonlinear functions, one obtains a family of different vectors lying in a vicinity of the target space. As a final step, the target space itself is estimated by applying PCA to this family of vectors. We show that this procedure is consistent in the sense that the estimaton error tends to zero at a parametric rate, uniformly over the family, Numerical examples demonstrate the usefulness of our method. Gilles Blanchard, Motoaki Kawanabe, Masashi Sugiyama, Vladimir G. Spokoiny, Klaus-Robert Müller |
J. Mach. Learn. Res. | 5 |
| 2006 | Incremental Support Vector Learning: Analysis, Implementation and ApplicationsabstractIncremental Support Vector Machines (SVM) are instrumental in practical applications of online learning. This work focuses on the design and analysis of efficient incremental SVM learning, with the aim of providing a fast, numerically stable and robust implementation. A detailed analysis of convergence and of algorithmic complexity of incremental SVM learning is carried out. Based on this analysis, a new design of storage and numerical operations is proposed, which speeds up the training of an incremental SVM by a factor of 5 to 20. The performance of the new algorithm is demonstrated in two scenarios: learning with limited resources and active learning. Various applications of the algorithm, such as in drug discovery, online monitoring of industrial devices and and surveillance of network traffic, can be foreseen. Pavel Laskov, Christian Gehl, Stefan Krüger, Klaus-Robert Müller |
J. Mach. Learn. Res. | 4 |
| 2006 | On the information and representation of non-Euclidean pairwise data
Julian Laub, Volker Roth 0001, Joachim M. Buhmann, Klaus-Robert Müller |
Pattern Recognit. | 4 |
| 2005 | Model Selection Under Covariate Shift
Masashi Sugiyama, Klaus-Robert Müller |
ICANN (2) | 2 |
| 2005 | Non-Gaussian Component Analysis: a Semi-parametric Framework for Linear Dimension ReductionabstractWe propose a new linear method for dimension reduction to identify nonGaussian components in high dimensional data. Our method, NGCA (non-Gaussian component analysis), uses a very general semi-parametric framework. In contrast to existing projection methods we define what is uninteresting (Gaussian): by projecting out uninterestingness, we can estimate the relevant non-Gaussian subspace. We show that the estimation error of finding the non-Gaussian components tends to zero at a parametric rate. Once NGCA components are identified and extracted, various tasks can be applied in the data analysis process, like data visualization, clustering, denoising or classification. A numerical study demonstrates the usefulness of our method. Gilles Blanchard, Masashi Sugiyama, Motoaki Kawanabe, Vladimir G. Spokoiny, Klaus-Robert Müller |
NIPS | 5 |
| 2005 | Optimizing spatio-temporal filters for improving Brain-Computer InterfacingabstractBrain-Computer Interface (BCI) systems create a novel communication channel from the brain to an output device by bypassing conventional motor output pathways of nerves and muscles. Therefore they could provide a new communication and control option for paralyzed patients. Modern BCI technology is essentially based on techniques for the clas- sification of single-trial brain signals. Here we present a novel technique that allows the simultaneous optimization of a spatial and a spectral filter enhancing discriminability of multi-channel EEG single-trials. The eval- uation of 60 experiments involving 22 different subjects demonstrates the superiority of the proposed algorithm. Apart from the enhanced clas- sification, the spatial and/or the spectral filter that are determined by the algorithm can also be used for further analysis of the data, e.g., for source localization of the respective brain rhythms. Guido Dornhege, Benjamin Blankertz, Matthias Krauledat, Florian Losch, Gabriel Curio, Klaus-Robert Müller |
NIPS | 6 |
| 2005 | Analyzing Coupled Brain Sources: Distinguishing True from Spurious InteractionabstractWhen trying to understand the brain, it is of fundamental importance to analyse (e.g. from EEG/MEG measurements) what parts of the cortex interact with each other in order to infer more accurate models of brain activity. Common techniques like Blind Source Separation (BSS) can estimate brain sources and single out artifacts by using the underlying assumption of source signal independence. However, physiologically interesting brain sources typically interact, so BSS will--by construction-- fail to characterize them properly. Noting that there are truly interacting sources and signals that only seemingly interact due to effects of volume conduction, this work aims to contribute by distinguishing these effects. For this a new BSS technique is proposed that uses anti-symmetrized cross-correlation matrices and subsequent diagonalization. The resulting decomposition consists of the truly interacting brain sources and suppresses any spurious interaction stemming from volume conduction. Our new concept of interacting source analysis (ISA) is successfully demonstrated on MEG data. Guido Nolte, Andreas Ziehe, Frank C. Meinecke, Klaus-Robert Müller |
NIPS | 4 |
| 2005 | Estimating Functions for Blind Separation When Sources Have Variance DependenciesabstractA blind separation problem where the sources are not independent, but have variance dependencies is discussed. For this scenario Hyvärinen and Hurri (2004) proposed an algorithm which requires no assumption on distributions of sources and no parametric model of dependencies between components. In this paper, we extend the semiparametric approach of Amari and Cardoso (1997) to variance dependencies and study estimating functions for blind separation of such dependent sources. In particular, we show that many ICA algorithms are applicable to the variance-dependent model as well under mild conditions, although they should in principle not. Our results indicate that separation can be done based only on normalized sources which are adjusted to have stationary variances and is not affected by the dependent activity levels. We also study the asymptotic distribution of the quasi maximum likelihood method and the stability of the natural gradient learning in detail. Simulation results of artificial and realistic examples match well with our theoretical findings. Motoaki Kawanabe, Klaus-Robert Müller |
J. Mach. Learn. Res. | 2 |
| 2004 | Regularizing generalization error estimators: a novel approach to robust model selection
Masashi Sugiyama, Motoaki Kawanabe, Klaus-Robert Müller |
ESANN | 3 |
| 2004 | Feature Discovery in Non-Metric Pairwise Data
Julian Laub, Klaus-Robert Müller |
J. Mach. Learn. Res. | 2 |
| 2004 | A Fast Algorithm for Joint Diagonalization with Non-orthogonal Transformations and its Application to Blind Source Separation
Andreas Ziehe, Pavel Laskov, Guido Nolte, Klaus-Robert Müller |
J. Mach. Learn. Res. | 4 |
| 2004 | Trading Variance Reduction with Unbiasedness: The Regularized Subspace Information Criterion for Robust Model Selection in Kernel RegressionabstractA well-known result by Stein (1956) shows that in particular situations, biased estimators can yield better parameter estimates than their generally preferred unbiased counterparts. This letter follows the same spirit, as we will stabilize the unbiased generalization error estimates by regularization and finally obtain more robust model selection criteria for learning. We trade a small bias against a larger variance reduction, which has the beneficial effect of being more precise on a single training set. We focus on the subspace information criterion (SIC), which is an unbiased estimator of the expected generalization error measured by the reproducing kernel Hilbert space norm. SIC can be applied to the kernel regression, and it was shown in earlier experiments that a small regularization of SIC has a stabilization effect. However, it remained open how to appropriately determine the degree of regularization in SIC. In this article, we derive an unbiased estimator of the expected squared error, between SIC and the expected generalization error and propose determining the degree of regularization of SIC such that the estimator of the expected squared error is minimized. Computer simulations with artificial and real data sets illustrate that the proposed method works effectively for improving the precision of SIC, especially in the high-noise-level cases. We furthermore compare the proposed method to the original SIC, the cross-validation, and an empirical Bayesian method in ridge parameter selection, with good results. Masashi Sugiyama, Motoaki Kawanabe, Klaus-Robert Müller |
Neural Comput. | 3 |
| 2004 | Asymptotic Properties of the Fisher KernelabstractThis letter analyzes the Fisher kernel from a statistical point of view. The Fisher kernel is a particularly interesting method for constructing a model of the posterior probability that makes intelligent use of unlabeled data (i.e., of the underlying data density). It is important to analyze and ultimately understand the statistical properties of the Fisher kernel. To this end, we first establish sufficient conditions that the constructed posterior model is realizable (i.e., it contains the true distribution). Realizability immediately leads to consistency results. Subsequently, we focus on an asymptotic analysis of the generalization error, which elucidates the learning curves of the Fisher kernel and how unlabeled data contribute to learning. We also point out that the squared or log loss is theoretically preferable-because both yield consistent estimators-to other losses such as the exponential loss, when a linear classifier is used together with the Fisher kernel. Therefore, this letter underlines that the Fisher kernel should be viewed not as a heuristics but as a powerful statistical tool with well-controlled statistical properties. Koji Tsuda, Shotaro Akaho, Motoaki Kawanabe, Klaus-Robert Müller |
Neural Comput. | 4 |
| 2004 | Injecting noise for analysing the stability of ICA components
Stefan Harmeling, Frank C. Meinecke, Klaus-Robert Müller |
Signal Process. | 3 |
| 2003 | Feature Extraction for One-Class Classification
David M. J. Tax, Klaus-Robert Müller |
ICANN | 2 |
| 2003 | Increase Information Transfer Rates in BCI by CSP Extension to Multi-classabstractBrain-Computer Interfaces (BCI) are an interesting emerging technology that is driven by the motivation to develop an effective communication in- terface translating human intentions into a control signal for devices like computers or neuroprostheses. If this can be done bypassing the usual hu- man output pathways like peripheral nerves and muscles it can ultimately become a valuable tool for paralyzed patients. Most activity in BCI re- search is devoted to finding suitable features and algorithms to increase information transfer rates (ITRs). The present paper studies the implica- tions of using more classes, e.g., left vs. right hand vs. foot, for operating a BCI. We contribute by (1) a theoretical study showing under some mild assumptions that it is practically not useful to employ more than three or four classes, (2) two extensions of the common spatial pattern (CSP) algorithm, one interestingly based on simultaneous diagonalization, and (3) controlled EEG experiments that underline our theoretical findings and show excellent improved ITRs. Guido Dornhege, Benjamin Blankertz, Gabriel Curio, Klaus-Robert Müller |
NIPS | 4 |
| 2003 | Blind Separation of Post-nonlinear Mixtures using Linearizing Transformations and Temporal Decorrelation
Andreas Ziehe, Motoaki Kawanabe, Stefan Harmeling, Klaus-Robert Müller |
J. Mach. Learn. Res. | 4 |
| 2003 | Kernel-Based Nonlinear Blind Source SeparationabstractWe propose kTDSEP, a kernel-based algorithm for nonlinear blind source separation (BSS). It combines complementary research fields: kernel feature spaces and BSS using temporal information. This yields an efficient algorithm for nonlinear BSS with invertible nonlinearity. Key assumptions are that the kernel feature space is chosen rich enough to approximate the nonlinearity and that signals of interest contain temporal information. Both assumptions are fulfilled for a wide set of real-world applications. The algorithm works as follows: First, the data are (implicitly) mapped to a high (possibly infinite)—dimensional kernel feature space. In practice, however, the data form a smaller submanifold in feature space—even smaller than the number of training data points—a fact that has already been used by, for example, reduced set techniques for support vector machines. We propose to adapt to this effective dimension as a preprocessing step and to construct an orthonormal basis of this submanifold. The latter dimension-reduction step is essential for making the subsequent application of BSS methods computationally and numerically tractable. In the reduced space, we use a BSS algorithm that is based on second-order temporal decorrelation. Finally, we propose a selection procedure to obtain the original sources from the extracted nonlinear components automatically. Experiments demonstrate the excellent performance and efficiency of our kTDSEP algorithm for several problems of nonlinear BSS and for more than two sources. Stefan Harmeling, Andreas Ziehe, Motoaki Kawanabe, Klaus-Robert Müller |
Neural Comput. | 4 |
| 2003 | Constructing Descriptive and Discriminative Nonlinear Features: Rayleigh Coefficients in Kernel Feature SpacesabstractWe incorporate prior knowledge to construct nonlinear algorithms for invariant feature extraction and discrimination. Employing a unified framework in terms of a nonlinearized variant of the Rayleigh coefficient, we propose nonlinear generalizations of Fisher's discriminant and oriented PCA using support vector kernel functions. Extensive simulations show the utility of our approach. Sebastian Mika, Gunnar Rätsch, Jason Weston, Bernhard Schölkopf, Alexander J. Smola, Klaus-Robert Müller |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2002 | New Methods for Splice Site Recognition
Sören Sonnenburg, Gunnar Rätsch, Arun K. Jagota, Klaus-Robert Müller |
ICANN | 4 |
| 2002 | Selecting Ridge Parameters in Infinite Dimensional Hypothesis Spaces
Masashi Sugiyama, Klaus-Robert Müller |
ICANN | 2 |
| 2002 | Combining Features for BCIabstractRecently, interest is growing to develop an effective communication in- terface connecting the human brain to a computer, the ’Brain-Computer Interface’ (BCI). One motivation of BCI research is to provide a new communication channel substituting normal motor output in patients with severe neuromuscular disabilities. In the last decade, various neuro- physiological cortical processes, such as slow potential shifts, movement related potentials (MRPs) or event-related desynchronization (ERD) of spontaneous EEG rhythms, were shown to be suitable for BCI, and, con- sequently, different independent approaches of extracting BCI-relevant EEG-features for single-trial analysis are under investigation. Here, we present and systematically compare several concepts for combining such EEG-features to improve the single-trial classification. Feature combi- nations are evaluated on movement imagination experiments with 3 sub- jects where EEG-features are based on either MRPs or ERD, or both. Those combination methods that incorporate the assumption that the sin- gle EEG-features are physiologically mutually independent outperform the plain method of ’adding’ evidence where the single-feature vectors are simply concatenated. These results strengthen the hypothesis that MRP and ERD reflect at least partially independent aspects of cortical processes and open a new perspective to boost BCI effectiveness. Guido Dornhege, Benjamin Blankertz, Gabriel Curio, Klaus-Robert Müller |
NIPS | 4 |
| 2002 | Going Metric: Denoising Pairwise DataabstractPairwise data in empirical sciences typically violate metricity, ei(cid:173) ther due to noise or due to fallible estimates, and therefore are hard to analyze by conventional machine learning technology. In this paper we therefore study ways to work around this problem. First, we present an alternative embedding to multi-dimensional scaling (MDS) that allows us to apply a variety of classical ma(cid:173) chine learning and signal processing algorithms. The class of pair(cid:173) wise grouping algorithms which share the shift-invariance property is statistically invariant under this embedding procedure, leading to identical assignments of objects to clusters. Based on this new vectorial representation, denoising methods are applied in a sec(cid:173) ond step. Both steps provide a theoretically well controlled setup to translate from pairwise data to the respective denoised met(cid:173) ric representation. We demonstrate the practical usefulness of our theoretical reasoning by discovering structure in protein sequence data bases, visibly improving performance upon existing automatic methods. Volker Roth 0001, Julian Laub, Joachim M. Buhmann, Klaus-Robert Müller |
NIPS | 4 |
| 2002 | Clustering with the Fisher ScoreabstractRecently the Fisher score (or the Fisher kernel) is increasingly used as a feature extractor for classification problems. The Fisher score is a vector of parameter derivatives of loglikelihood of a probabilistic model. This paper gives a theoretical analysis about how class information is pre- served in the space of the Fisher score, which turns out that the Fisher score consists of a few important dimensions with class information and many nuisance dimensions. When we perform clustering with the Fisher score, K-Means type methods are obviously inappropriate because they make use of all dimensions. So we will develop a novel but simple clus- tering algorithm specialized for the Fisher score, which can exploit im- portant dimensions. This algorithm is successfully tested in experiments with artificial data and real data (amino acid sequences). as follows: Let us assume that a probabilistic model parameter estimate Koji Tsuda, Motoaki Kawanabe, Klaus-Robert Müller |
NIPS | 3 |
| 2002 | The Subspace Information Criterion for Infinite Dimensional Hypothesis Spaces
Masashi Sugiyama, Klaus-Robert Müller |
J. Mach. Learn. Res. | 2 |
| 2002 | A New Discriminative Kernel from Probabilistic ModelsabstractRecently, Jaakkola and Haussler (1999) proposed a method for constructing kernel functions from probabilistic models. Their so-called Fisher kernel has been combined with discriminative classifiers such as support vector machines and applied successfully in, for example, DNA and protein analysis. Whereas the Fisher kernel is calculated from the marginal log-likelihood, we propose the TOP kernel derived; from tangent vectors of posterior log-odds. Furthermore, we develop a theoretical framework on feature extractors from probabilistic models and use it for analyzing the TOP kernel. In experiments, our new discriminative TOP kernel compares favorably to the Fisher kernel. Koji Tsuda, Motoaki Kawanabe, Gunnar Rätsch, Sören Sonnenburg, Klaus-Robert Müller |
Neural Comput. | 5 |
| 2002 | On-line learning in changing environments with applications in supervised and unsupervised learning
Noboru Murata, Motoaki Kawanabe, Andreas Ziehe, Klaus-Robert Müller, Shun-ichi Amari |
Neural Networks | 4 |
| 2002 | Constructing Boosting Algorithms from SVMs: An Application to One-Class ClassificationabstractWe show via an equivalence of mathematical programs that a support vector (SV) algorithm can be translated into an equivalent boosting-like algorithm and vice versa. We exemplify this translation procedure for a new algorithm: one-class leveraging, starting from the one-class support vector machine (1-SVM). This is a first step toward unsupervised learning in a boosting framework. Building on so-called barrier methods known from the theory of constrained optimization, it returns a function, written as a convex combination of base hypotheses, that characterizes whether a given test point is likely to have been generated from the distribution underlying the training data. Simulations on one-class classification problems demonstrate the usefulness of our approach. Gunnar Rätsch, Sebastian Mika, Bernhard Schölkopf, Klaus-Robert Müller |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2002 | Subspace information criterion for nonquadratic regularizers-Model selection for sparse regressorsabstractNonquadratic regularizers, in particular the l(1) norm regularizer can yield sparse solutions that generalize well. In this work we propose the generalized subspace information criterion (GSIC) that allows to predict the generalization error for this useful family of regularizers. We show that under some technical assumptions GSIC is an asymptotically unbiased estimator of the generalization error. GSIC is demonstrated to have a good performance in experiments with the l(1) norm regularizer as we compare with the network information criterion (NIC) and cross- validation in relatively large sample cases. However in the small sample case, GSIC tends to fail to capture the optimal model due to its large variance. Therefore, also a biased version of GSIC is introduced,which achieves reliable model selection in the relevant and challenging scenario of high-dimensional data and few samples. Koji Tsuda, Masashi Sugiyama, Klaus-Robert Müller |
IEEE Trans. Neural Networks | 3 |
| 2001 | Learning to Predict the Leave-One-Out Error of Kernel Based Classifiers
Koji Tsuda, Gunnar Rätsch, Sebastian Mika, Klaus-Robert Müller |
ICANN | 4 |
| 2001 | Classifying Single Trial EEG: Towards Brain Computer InterfacingabstractDriven by the progress in the field of single-trial analysis of EEG, there is a growing interest in brain computer interfaces (BCIs), i.e., systems that enable human subjects to control a computer only by means of their brain signals. In a pseudo-online simulation our BCI detects upcoming finger movements in a natural keyboard typing condition and predicts their lat- erality. This can be done on average 100–230 ms before the respective key is actually pressed, i.e., long before the onset of EMG. Our approach is appealing for its short response time and high classification accuracy (>96%) in a binary decision where no human training is involved. We compare discriminative classifiers like Support Vector Machines (SVMs) and different variants of Fisher Discriminant that possess favorable reg- ularization properties for dealing with high noise cases (inter-trial vari- ablity). Benjamin Blankertz, Gabriel Curio, Klaus-Robert Müller |
NIPS | 3 |
| 2001 | Kernel Feature Spaces and Nonlinear Blind Souce SeparationabstractIn kernel based learning the data is mapped to a kernel feature space of a dimension that corresponds to the number of training data points. In practice, however, the data forms a smaller submanifold in feature space, a fact that has been used e.g. by reduced set techniques for SVMs. We propose a new mathematical construction that permits to adapt to the in- trinsic dimension and to find an orthonormal basis of this submanifold. In doing so, computations get much simpler and more important our theoretical framework allows to derive elegant kernelized blind source separation (BSS) algorithms for arbitrary invertible nonlinear mixings. Experiments demonstrate the good performance and high computational efficiency of our kTDSEP algorithm for the problem of nonlinear BSS. Stefan Harmeling, Andreas Ziehe, Motoaki Kawanabe, Klaus-Robert Müller |
NIPS | 4 |
| 2001 | Estimating the Reliability of ICA ProjectionsabstractWhen applying unsupervised learning techniques like ICA or tem(cid:173) poral decorrelation, a key question is whether the discovered pro(cid:173) jections are reliable. In other words: can we give error bars or can we assess the quality of our separation? We use resampling meth(cid:173) ods to tackle these questions and show experimentally that our proposed variance estimations are strongly correlated to the sepa(cid:173) ration error. We demonstrate that this reliability estimation can be used to choose the appropriate ICA-model, to enhance signifi(cid:173) cantly the separation performance, and, most important, to mark the components that have a actual physical meaning. Application to 49-channel-data from an magneto encephalography (MEG) ex(cid:173) periment underlines the usefulness of our approach. Frank C. Meinecke, Andreas Ziehe, Motoaki Kawanabe, Klaus-Robert Müller |
NIPS | 4 |
| 2001 | A New Discriminative Kernel From Probabilistic ModelsabstractRecently, Jaakkola and Haussler proposed a method for construct(cid:173) ing kernel functions from probabilistic models. Their so called "Fisher kernel" has been combined with discriminative classifiers such as SVM and applied successfully in e.g. DNA and protein analysis. Whereas the Fisher kernel (FK) is calculated from the marginal log-likelihood, we propose the TOP kernel derived from Tangent vectors Of Posterior log-odds. Furthermore we develop a theoretical framework on feature extractors from probabilistic models and use it for analyzing FK and TOP. In experiments our new discriminative TOP kernel compares favorably to the Fisher kernel. Koji Tsuda, Motoaki Kawanabe, Gunnar Rätsch, Sören Sonnenburg, Klaus-Robert Müller |
NIPS | 5 |
| 2001 | Soft Margins for AdaBoost
Gunnar Rätsch, Takashi Onoda, Klaus-Robert Müller |
Mach. Learn. | 3 |
| 2001 | An introduction to kernel-based learning algorithmsabstractThis paper provides an introduction to support vector machines, kernel Fisher discriminant analysis, and kernel principal component analysis, as examples for successful kernel-based learning methods. We first give a short background about Vapnik-Chervonenkis theory and kernel feature spaces and then proceed to kernel based learning in supervised and unsupervised scenarios including practical and algorithmic considerations. We illustrate the usefulness of kernel algorithms by discussing applications such as optical character recognition and DNA analysis. Klaus-Robert Müller, Sebastian Mika, Gunnar Rätsch, Koji Tsuda, Bernhard Schölkopf |
IEEE Trans. Neural Networks | 1 |
| 2000 | Barrier Boosting
Gunnar Rätsch, Manfred K. Warmuth, Sebastian Mika, Takashi Onoda, Steven Lemm, Klaus-Robert Müller |
COLT | 6 |
| 2000 | A Mathematical Programming Approach to the Kernel Fisher AlgorithmabstractWe investigate a new kernel-based classifier: the Kernel Fisher Discrim(cid:173) inant (KFD). A mathematical programming formulation based on the ob(cid:173) servation that KFD maximizes the average margin permits an interesting modification of the original KFD algorithm yielding the sparse KFD. We find that both, KFD and the proposed sparse KFD, can be understood in an unifying probabilistic context. Furthermore, we show connections to Support Vector Machines and Relevance Vector Machines. From this understanding, we are able to outline an interesting kernel-regression technique based upon the KFD algorithm. Simulations support the use(cid:173) fulness of our approach. Sebastian Mika, Gunnar Rätsch, Klaus-Robert Müller |
NIPS | 3 |
| 2000 | Robust Ensemble Learning for Data Mining
Gunnar Rätsch, Bernhard Schölkopf, Alexander J. Smola, Sebastian Mika, Takashi Onoda, Klaus-Robert Müller |
PAKDD | 6 |
| 2000 | Engineering support vector machine kernels that recognize translation initiation sitesabstractMOTIVATION: In order to extract protein sequences from nucleotide sequences, it is an important step to recognize points at which regions start that code for proteins. These points are called translation initiation sites (TIS). RESULTS: The task of finding TIS can be modeled as a classification problem. We demonstrate the applicability of support vector machines for this task, and show how to incorporate prior biological knowledge by engineering an appropriate kernel function. With the described techniques the recognition performance can be improved by 26% over leading existing approaches. We provide evidence that existing related methods (e.g. ESTScan) could profit from advanced TIS recognition. Alexander Zien, Gunnar Rätsch, Sebastian Mika, Bernhard Schölkopf, Thomas Lengauer, Klaus-Robert Müller |
Bioinform. | 6 |
| 1999 | Hidden Markov gating for prediction of change points in switching dynamical systems
Stefan Liehr, Klaus Pawelzik, Jens Kohlmorgen, Steven Lemm, Klaus-Robert Müller |
ESANN | 5 |
| 1999 | Invariant Feature Extraction and Classification in Kernel Spaces
Sebastian Mika, Gunnar Rätsch, Jason Weston, Bernhard Schölkopf, Alexander J. Smola, Klaus-Robert Müller |
NIPS | 6 |
| 1999 | Unmixing Hyperspectral Data
Lucas C. Parra, Clay Spence, Paul Sajda, Andreas Ziehe, Klaus-Robert Müller |
NIPS | 5 |
| 1999 | v-Arc: Ensemble Learning in the Presence of Outliers
Gunnar Rätsch, Bernhard Schölkopf, Alexander J. Smola, Klaus-Robert Müller, Takashi Onoda, Sebastian Mika |
NIPS | 4 |
| 1999 | Input space versus feature space in kernel-based methodsabstractThis paper collects some ideas targeted at advancing our understanding of the feature spaces associated with support vector (SV) kernel functions. We first discuss the geometry of feature space. In particular, we review what is known about the shape of the image of input space under the feature space map, and how this influences the capacity of SV methods. Following this, we describe how the metric governing the intrinsic geometry of the mapped surface can be computed in terms of the kernel, using the example of the class of inhomogeneous polynomial kernels, which are often used in SV pattern recognition. We then discuss the connection between feature space and input space by dealing with the question of how one can, given some vector in feature space, find a preimage (exact or approximate) in input space. We describe algorithms to tackle this issue, and show their utility in two applications of kernel methods. First, we use it to reduce the computational complexity of SV decision functions; second, we combine it with the Kernel PCA algorithm, thereby constructing a nonlinear statistical denoising technique which is shown to perform well on real-world data. Bernhard Schölkopf, Sebastian Mika, Christopher J. C. Burges, Phil Knirsch, Klaus-Robert Müller, Gunnar Rätsch, Alexander J. Smola |
IEEE Trans. Neural Networks | 5 |
| 1998 | An Improvement of AdaBoost to Avoid Overfitting
Gunnar Rätsch, Takashi Onoda, Klaus-Robert Müller |
ICONIP | 3 |
| 1998 | Kernel PCA and De-Noising in Feature Spaces
Sebastian Mika, Bernhard Schölkopf, Alexander J. Smola, Klaus-Robert Müller, Matthias Scholz, Gunnar Rätsch |
NIPS | 4 |
| 1998 | Regularizing AdaBoost
Gunnar Rätsch, Takashi Onoda, Klaus-Robert Müller |
NIPS | 3 |
| 1998 | Nonlinear Component Analysis as a Kernel Eigenvalue ProblemabstractA new method for performing a nonlinear form of principal component analysis is proposed. By the use of integral operator kernel functions, one can efficiently compute principal components in high-dimensional feature spaces, related to input space by some nonlinear map—for instance, the space of all possible five-pixel products in 16 × 16 images. We give the derivation of the method and present experimental results on polynomial feature extraction for pattern recognition. Bernhard Schölkopf, Alexander J. Smola, Klaus-Robert Müller |
Neural Comput. | 3 |
| 1998 | The connection between regularization operators and support vector kernels
Alexander J. Smola, Bernhard Schölkopf, Klaus-Robert Müller |
Neural Networks | 3 |
| 1998 | Data Set A is a Pattern Matching Problem
Jens Kohlmorgen, Klaus-Robert Müller |
Neural Process. Lett. | 2 |
| 1997 | Analysis of Wake/Sleep EEG with Competing Experts
Jens Kohlmorgen, Klaus-Robert Müller, Jörn Rittweger, Klaus Pawelzik |
ICANN | 2 |
| 1997 | Predicting Time Series with Support Vector Machines
Klaus-Robert Müller, Alexander J. Smola, Gunnar Rätsch, Bernhard Schölkopf, Jens Kohlmorgen, Vladimir Vapnik |
ICANN | 1 |
| 1997 | Kernel Principal Component Analysis
Bernhard Schölkopf, Alexander J. Smola, Klaus-Robert Müller |
ICANN | 3 |
| 1997 | Analysis of Drifting Dynamics with Neural Network Hidden Markov Models
Jens Kohlmorgen, Klaus-Robert Müller, Klaus Pawelzik |
NIPS | 2 |
| 1997 | Asymptotic statistical theory of overtraining and cross-validationabstractA statistical theory for overtraining is proposed. The analysis treats general realizable stochastic neural networks, trained with Kullback-Leibler divergence in the asymptotic case of a large number of training examples. It is shown that the asymptotic gain in the generalization error is small if we perform early stopping, even if we have access to the optimal stopping time. Based on the cross-validation stopping we consider the ratio the examples should be divided into training and cross-validation sets in order to obtain the optimum performance. Although cross-validated early stopping is useless in the asymptotic region, it surely decreases the generalization error in the nonasymptotic region. Our large scale simulations done on a CM5 are in good agreement with our analytical findings. Shun-ichi Amari, Noboru Murata, Klaus-Robert Müller, Michael Finke, Howard Hua Yang |
IEEE Trans. Neural Networks | 3 |
| 1996 | Analysis of Drifting Dynamics with Competing Predictors
Jens Kohlmorgen, Klaus-Robert Müller, Klaus Pawelzik |
ICANN | 2 |
| 1996 | Prediction of Mixtures
Klaus Pawelzik, Klaus-Robert Müller, Jens Kohlmorgen |
ICANN | 2 |
| 1996 | Adaptive On-line Learning in Changing Environments
Noboru Murata, Klaus-Robert Müller, Andreas Ziehe, Shun-ichi Amari |
NIPS | 2 |
| 1996 | A Numerical Study on Learning Curves in Stochastic Multilayer Feedforward NetworksabstractThe universal asymptotic scaling laws proposed by Amari et al. are studied in large scale simulations using a CM5. Small stochastic multilayer feedforward networks trained with backpropagation are investigated. In the range of a large number of training patterns t, the asymptotic generalization error scales as 1/t as predicted. For a medium range t a faster 1/t2 scaling is observed. This effect is explained by using higher order corrections of the likelihood expansion. It is shown for small t that the scaling law changes drastically, when the network undergoes a transition from strong overfitting to effective learning. Klaus-Robert Müller, Michael Finke, Noboru Murata, Klaus Schulten, Shun-ichi Amari |
Neural Comput. | 1 |
| 1996 | Annealed Competition of Experts for a Segmentation and Classification of Switching DynamicsabstractWe present a method for the unsupervised segmentation of data streams originating from different unknown sources that alternate in time. We use an architecture consisting of competing neural networks. Memory is included to resolve ambiguities of input-output relations. To obtain maximal specialization, the competition is adiabatically increased during training. Our method achieves almost perfect identification and segmentation in the case of switching chaotic dynamics where input manifolds overlap and input-output relations are ambiguous. Only a small dataset is needed for the training procedure. Applications to time series from complex systems demonstrate the potential relevance of our approach for time series analysis and short-term prediction. Klaus Pawelzik, Jens Kohlmorgen, Klaus-Robert Müller |
Neural Comput. | 3 |
| 1995 | Statistical Theory of Overtraining - Is Cross-Validation Asymptotically Effective?
Shun-ichi Amari, Noboru Murata, Klaus-Robert Müller, Michael Finke, Howard Hua Yang |
NIPS | 3 |