VLDB 2026 Research / reviewers in the wild / expert
Wojciech Samek
dblp:79/9736 · also Wojciech Wojcikiewicz
· DBLP profile ↗
95ranked-venue papers
9as first author
42since 2021 · last 2026
0000-0002-6283-3265ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 3 first-author · 27 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Weights to Activations: Is Steering the Next Frontier of Adaptation?abstractSimon Ostermann, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich, Sebastian Lapuschkin, Wojciech Samek, Vera Schmitt. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Simon Ostermann 0002, Daniil Gurgurov, Tanja Baeumel, Michael A. Hedderich, Sebastian Lapuschkin, Wojciech Samek, Vera Schmitt |
ACL (1) | 6 |
| 2026 | Leveraging Sparsity for Privacy in Collaborative InferenceabstractCollaborative inference (CI) is hampered by high communication costs and privacy risks, with existing defenses often forcing a trade-off between efficiency and formal privacy guarantees. In this work, we present a framework that leverages activation sparsity as a dual-purpose mechanism to address both challenges simultaneously. Our approach uses a lightweight Sparse Autoencoder (SAE) to learn a sparse representation, which is then protected by a novel two-channel noise mechanism grounded in information theory. This design provides a tunable privacy budget while remaining computationally inexpensive. Evaluations on CIFAR-10, Tiny-ImageNet, and FaceScrub show that our method achieves a state-of-the-art privacy-utility trade-off, sustaining high accuracy at sparsity levels of up to 97%, while offering superior resilience against strong model inversion attacks. Our results underline that sparsity can be transformed from an effective compression tool into a powerful and theoretically-grounded privacy defense, paving the way for more practical and trustworthy CI systems. We provide code at https://github.com/an7123/privacy_ci. Maximilian Andreas Hoefler, Karsten Müller 0001, Wojciech Samek |
WACV | 3 |
| 2025 | Human-Centric AI: From Explainability and Trustworthiness to Actionable EthicsabstractTo address the potential risks of AI while supporting innovation and ensuring responsible adoption, there is an urgent need for clear governance frameworks grounded in human-centric values. It is imperative that AI systems operate in ways that are transparent, trustworthy, and ethically sound. Developing truly human-centric AI goes beyond technical innovation. It requires interdisciplinary collaboration and diverse perspectives. This workshop will explore key challenges and emerging solutions in the development of human-centric AI, with a focus on explainability, trustworthiness, fairness, and privacy. We welcome both theoretical contributions and practical case studies that demonstrate how human-centered principles are realized in real-world AI systems. The official workshop webpage is available at https://xai.kaist.ac.kr/Workshop/hcai2025/, which provides comprehensive information about the program. Jaesik Choi, Bohyung Han, Myoung-Wan Koo, Kyungman Bae, Chang Dong Yoo, Simon S. Woo, Wojciech Samek |
CIKM | 7 |
| 2025 | Model Science: Getting Serious About Verification, Explanation and Control of AI SystemsabstractThe recent proliferation of foundation models calls for a paradigm shift from Data Science to Model Science. Unlike data-centric approaches, Model Science places the trained model at the core of analysis, aiming to interact, verify, explain, and control its behavior across diverse operational contexts. This paper introduces a conceptual framework for a new discipline called Model Science, along with the proposal for its four key pillars: Verification, which requires strict, context-aware evaluation protocols; Explanation, which is understood as various approaches to explore of internal model operations; Control, which integrates alignment techniques to steer model behavior; and Interface, which develops interactive and visual explanation tools to improve human calibration and decision-making. The proposed framework aims to guide the development of credible, safe, and human-aligned AI systems. Przemyslaw Biecek, Wojciech Samek |
ECAI | 2 |
| 2025 | Efficient Federated Learning of Mixed-Token Transformers for Cellular Feature PredictionabstractIn telecommunications, Autonomous Networks (ANs) automatically adjust configurations based on specific requirements (e.g., bandwidth) and available resources. These networks rely on continuous monitoring and intelligent mechanisms for self-optimization, self-repair, and self-protection, nowa-days enhanced by Neural Networks (NNs) to enable predictive modeling and pattern recognition.In this work, we propose a neural architecture called Mixed-Token Transformer (MTT) that can process and predict both numerical and textual tokens representing various mobile network features (e.g., ping, SNR or band frequency). We train the models using a Federated Learning (FL) approach, enabling multiple AN cells — each equipped with an MTT — to collaboratively learn from cellular data while preserving data privacy.Since FL involves frequent transmission of neural network updates, it requires an efficient, standardized compression strategy for reliable communication. To address this, we investigate NNCodec, our implementation of the ISO/IEC Neural Network Coding (NNC) standard. Our experimental results on the Berlin V2X dataset demonstrate that NNCodec achieves transparent compression (i.e., negligible performance loss) while reducing communication overhead to below 1%, showing the effectiveness of combining NNC with FL in collaboratively learned autonomous mobile networks. Our proposed MTT architecture speeds up model inference by ≈ 5×. Code is available at: github.com/Jostarndt/NNCodec-FL-MixedTokenTransformer. Daniel Becking, Jost Arndt, Ingo Friese, Karsten Müller 0001, Jackie Ma, Thomas Buchholz, Mandy Galkow-Schneider, Wojciech Samek, Detlev Marpe |
GLOBECOM | 8 |
| 2025 | FedXDS: Leveraging Model Attribution Methods to Counteract Data Heterogeneity in Federated Learningabstract4572 Maximilian Andreas Hoefler, Karsten Müller 0001, Wojciech Samek |
ICCV | 3 |
| 2025 | Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional DivergenceabstractWith a growing interest in understanding neural network prediction strategies, Concept Activation Vectors (CAVs) have emerged as a popular tool for modeling human-understandable concepts in the latent space.
Commonly, CAVs are computed by leveraging linear classifiers optimizing the *separability* of latent representations of samples with and without a given concept. However, in this paper we show that such a separability-oriented computation leads to solutions, which may diverge from the actual goal of precisely modeling the concept direction.
This discrepancy can be attributed to the significant influence of distractor directions, i.e., signals unrelated to the concept, which are picked up by filters (i.e., weights) of linear models to optimize class-separability.
To address this, we introduce *pattern-based CAVs*, solely focussing on concept signals, thereby providing more accurate concept directions.
We evaluate various CAV methods in terms of their alignment with the true concept direction and their impact on CAV applications, including concept sensitivity testing and model correction for shortcut behavior caused by data artifacts.
We demonstrate the benefits of pattern-based CAVs using the Pediatric Bone Age, ISIC2019, and FunnyBirds datasets with VGG, ResNet, ReXNet, EfficientNet, and Vision Transformer as model architectures. Frederik Pahde, Maximilian Dreyer, Moritz Weckbecker, Leander Weber, Christopher J. Anders, Thomas Wiegand 0001, Wojciech Samek, Sebastian Lapuschkin |
ICLR | 7 |
| 2025 | Inspecting AI Like Engineers: From Explanation to Validation with SemanticLens
Wojciech Samek |
ICSOFT | 1 |
| 2025 | The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval AugmentationabstractLarge language models are able to exploit in-context learning to access external knowledge beyond their training data through retrieval-augmentation. While promising, its inner workings remain unclear. In this work, we shed light on the mechanism of in-context retrieval augmentation for question answering by viewing a prompt as a composition of informational components. We propose an attribution-based method to identify specialized attention heads, revealing in-context heads that comprehend instructions and retrieve relevant contextual information, and parametric heads that store entities' relational knowledge. To better understand their roles, we extract function vectors and modify their attention weights to show how they can influence the answer generation process. Finally, we leverage the gained insights to trace the sources of knowledge used during inference, paving the way towards more safe and transparent language models. Patrick Kahardipraja, Reduan Achtibat, Thomas Wiegand 0001, Wojciech Samek, Sebastian Lapuschkin |
NeurIPS | 4 |
| 2025 | Fractional Diffusion Bridge ModelsabstractWe present *Fractional Diffusion Bridge Models* (FDBM), a novel generative diffusion bridge framework driven by an approximation of the rich and non-Markovian fractional Brownian motion (fBM). Real stochastic processes exhibit a degree of memory effects (correlations in time), long-range dependencies, roughness and anomalous diffusion phenomena that are not captured in standard diffusion or bridge modeling due to the use of Brownian motion (BM).
As a remedy, leveraging a recent Markovian approximation of fBM (MA-fBM), we construct FDBM that enable tractable inference while preserving the non-Markovian nature of fBM. We prove the existence of a coupling-preserving generative diffusion bridge and leverage it for future state prediction from paired training data. We then extend our formulation to the Schrödinger bridge problem and derive a principled loss function to learn the unpaired data translation. We evaluate FDBM on both tasks: predicting future protein conformations from aligned data, and unpaired image translation. In both settings, FDBM achieves superior performance compared to the Brownian baselines, yielding lower root mean squared deviation (RMSD) of C$_\alpha$ atomic positions in protein structure prediction and lower Fréchet Inception Distance (FID) in unpaired image translation. Gabriel Nobis, Maximilian Springenberg, Arina Belova, Rembert Daems, Christoph Knochenhauer, Manfred Opper, Tolga Birdal, Wojciech Samek |
NeurIPS | 8 |
| 2025 | Beyond Scalars: Concept-Based Alignment Analysis in Vision TransformersabstractMeasuring the alignment between representations lets us understand similarities between the feature spaces of different models, such as Vision Transformers trained under diverse paradigms. However, traditional measures for representational alignment yield only scalar values that obscure how these spaces agree in terms of learned features. To address this, we combine alignment analysis with concept discovery, allowing a fine-grained breakdown of alignment into individual concepts. This approach reveals both universal concepts across models and each representation’s internal concept structure. We introduce a new definition of concepts as non-linear manifolds, hypothesizing they better capture the geometry of the feature space. A sanity check demonstrates the advantage of this manifold-based definition over linear baselines for concept-based alignment. Finally, our alignment analysis of four different ViTs shows that increased supervision tends to reduce semantic organization in learned representations. Johanna Vielhaben, Dilyara Bareeva, Jim Berend, Wojciech Samek, Nils Strodthoff |
NeurIPS | 4 |
| 2025 | Ensuring medical AI safety: interpretability-driven detection and mitigation of spurious model behavior and associated dataabstractDeep neural networks are increasingly employed in high-stakes medical applications, despite their tendency for shortcut learning in the presence of spurious correlations, which can have potentially fatal consequences in practice. Whereas a multitude of works address either the detection or mitigation of such shortcut behavior in isolation, the Reveal2Revise approach provides a comprehensive bias mitigation framework combining these steps. However, effectively addressing these biases often requires substantial labeling efforts from domain experts. In this work, we review the steps of the Reveal2Revise framework and enhance it with semi-automated interpretability-based bias annotation capabilities. This includes methods for the sample- and feature-level bias annotation, providing valuable information for bias mitigation methods to unlearn the undesired shortcut behavior. We show the applicability of the framework using four medical datasets across two modalities, featuring controlled and real-world spurious correlations caused by data artifacts. We successfully identify and mitigate these biases in VGG16, ResNet50, and contemporary Vision Transformer models, ultimately increasing their robustness and applicability for real-world medical tasks. Our code is available at https://github.com/frederikpahde/medical-ai-safety. Frederik Pahde, Thomas Wiegand 0001, Sebastian Lapuschkin, Wojciech Samek |
Mach. Learn. | 4 |
| 2025 | Explaining predictive uncertainty by exposing second-order effectsabstractExplainable AI has brought transparency to complex ML black boxes, enabling us, in particular, to identify which features these models use to make predictions. So far, the question of how to explain predictive uncertainty, i.e., why a model ‘doubts’, has been scarcely studied. Our investigation reveals that predictive uncertainty is dominated by second-order effects , involving single features or product interactions between them. We contribute a new method for explaining predictive uncertainty based on these second-order effects. Computationally, our method reduces to a simple covariance computation over a collection of first-order explanations. Our method is generally applicable, allowing for turning common attribution techniques (LRP, Gradient × Input, etc.) into powerful second-order uncertainty explainers, which we call CovLRP, CovGI, etc. The accuracy of the explanations our method produces is demonstrated through systematic quantitative evaluations, and the overall usefulness of our method is demonstrated through two practical showcases. • Explaining the predictive uncertainty of an ML model is practically relevant. • Second-order effects are found to be dominant in uncertainty estimates. • An efficient method for attributing uncertainty to input features is proposed. • Evaluations and showcases demonstrate the accuracy and usefulness of the method. Florian Bley, Sebastian Lapuschkin, Wojciech Samek, Grégoire Montavon |
Pattern Recognit. | 3 |
| 2025 | A Privacy Preserving System for Movie Recommendations Using Federated LearningabstractRecommender systems have become ubiquitous in the past years. They solve the tyranny of choice problem faced by many users, and are utilized by many online businesses to drive engagement and sales. Besides other criticisms, like creating filter bubbles within social networks, recommender systems are often reproved for collecting considerable amounts of personal data. However, to personalize recommendations, personal information is fundamentally required. A recent distributed learning scheme called federated learning has made it possible to learn from personal user data without its central collection. Consequently, we present a recommender system for movie recommendations, which provides privacy and thus trustworthiness on multiple levels: First and foremost, it is trained using federated learning and thus, by its very nature, privacy-preserving, while still enabling users to benefit from global insights. Furthermore, a novel federated learning scheme, called FedQ, is employed, which not only addresses the problem of non-i.i.d.-ness and small local datasets, but also prevents input data reconstruction attacks by aggregating client updates early. Finally, to reduce the communication overhead, compression is applied, which significantly compresses the exchanged neural network parametrizations to a fraction of their original size. We conjecture that this may also improve data privacy through its lossy quantization stage. David Neumann, Andreas Lutz, Karsten Müller 0001, Wojciech Samek |
Trans. Recomm. Syst. | 4 |
| 2024 | From Hope to Safety: Unlearning Biases of Deep Models via Gradient Penalization in Latent SpaceabstractDeep Neural Networks are prone to learning spurious correlations embedded in the training data, leading to potentially biased predictions. This poses risks when deploying these models for high-stake decision-making, such as in medical applications. Current methods for post-hoc model correction either require input-level annotations which are only possible for spatially localized biases, or augment the latent feature space, thereby hoping to enforce the right reasons. We present a novel method for model correction on the concept level that explicitly reduces model sensitivity towards biases via gradient penalization. When modeling biases via Concept Activation Vectors, we highlight the importance of choosing robust directions, as traditional regression-based approaches such as Support Vector Machines tend to result in diverging directions. We effectively mitigate biases in controlled and real-world settings on the ISIC, Bone Age, ImageNet and CelebA datasets using VGG, ResNet and EfficientNet architectures. Code and Appendix are available on https://github.com/frederikpahde/rrclarc. Maximilian Dreyer, Frederik Pahde, Christopher J. Anders, Wojciech Samek, Sebastian Lapuschkin |
AAAI | 4 |
| 2024 | Boosting Federated Learning with Diffusion Models for Non-IID and Imbalanced DataabstractFederated learning (FL) has emerged as an effective paradigm in machine learning, enabling participants to extract value from diverse data sources while preserving privacy. However, FL faces significant challenges, including communication inefficiency, computational overhead, and data heterogeneity. The latter, in particular, can substantially impair FL performance by increasing convergence time and degrading generalization capabilities. Recent FL literature indicates that poor global model performance in heterogeneous conditions can be attributed to divergence between the classification layers of clients. To address this challenge, we propose a novel approach leveraging recent advancements in foundation models, specifically diffusion models, to generate synthetic data that bridges the gap between heterogeneous FL and centralized learning. Specifically, our method involves fine-tuning a diffusion model on the client side, using a small subset of images from the local dataset, to generate synthetic images. These images along with labels are shared with the server, where the final classification layer of a converged federated model is retrained to counteract classifier divergence. Our experiments demonstrate significant improvements, including a +30% increase in performance accuracy and a 9-fold acceleration in convergence for non-IID FL scenarios. Furthermore, we evaluate our method on real-world datasets prone to image artifacts and data imbalance, showcasing the effectiveness of our approach in both industrial and medical applications. Maximilian Andreas Hoefler, Tatsiana Mazouka, Wojciech Samek |
IEEE Big Data | 4 |
| 2024 | AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for TransformersabstractLarge Language Models are prone to biased predictions and hallucinations, underlining the paramount importance of understanding their model-internal reasoning process. However, achieving faithful attributions for the entirety of a black-box transformer model and maintaining computational efficiency is an unsolved challenge. By extending the Layer-wise Relevance Propagation attribution method to handle attention layers, we address these challenges effectively. While partial solutions exist, our method is the first to faithfully and holistically attribute not only input but also latent representations of transformer models with the computational efficiency similar to a single backward pass. Through extensive evaluations against existing methods on LLaMa 2, Mixtral 8x7b, Flan-T5 and vision transformer architectures, we demonstrate that our proposed approach surpasses alternative methods in terms of faithfulness and enables the understanding of latent representations, opening up the door for concept-based explanations. We provide an LRP library at https://github.com/rachtibat/LRP-eXplains-Transformers. Reduan Achtibat, Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Thomas Wiegand 0001, Sebastian Lapuschkin, Wojciech Samek |
ICML | 7 |
| 2024 | Position: Explain to Question not to JustifyabstractExplainable Artificial Intelligence (XAI) is a young but very promising field of research. Unfortunately, the progress in this field is currently slowed down by divergent and incompatible goals. We separate various threads tangled within the area of XAI into two complementary cultures of human/value-oriented explanations (BLUE XAI) and model/validation-oriented explanations (RED XAI). This position paper argues that the area of RED XAI is currently under-explored, i.e., more methods for explainability are desperately needed to question models (e.g., extract knowledge from well-performing models as well as spotting and fixing bugs in faulty models), and the area of RED XAI hides great opportunities and potential for important research necessary to ensure the safety of AI systems. We conclude this paper by presenting promising challenges in this area. Przemyslaw Biecek, Wojciech Samek |
ICML | 2 |
| 2024 | Generative Fractional Diffusion ModelsabstractWe introduce the first continuous-time score-based generative model that leverages fractional diffusion processes for its underlying dynamics. Although diffusion models have excelled at capturing data distributions, they still suffer from various limitations such as slow convergence, mode-collapse on imbalanced data, and lack of diversity. These issues are partially linked to the use of light-tailed Brownian motion (BM) with independent increments. In this paper, we replace BM with an approximation of its non-Markovian counterpart, fractional Brownian motion (fBM), characterized by correlated increments and Hurst index $H \in (0,1)$, where $H=0.5$ recovers the classical BM. To ensure tractable inference and learning, we employ a recently popularized Markov approximation of fBM (MA-fBM) and derive its reverse-time model, resulting in *generative fractional diffusion models* (GFDM). We characterize the forward dynamics using a continuous reparameterization trick and propose *augmented score matching* to efficiently learn the score function, which is partly known in closed form, at minimal added cost. The ability to drive our diffusion model via MA-fBM offers flexibility and control. $H \leq 0.5$ enters the regime of *rough paths* whereas $H>0.5$ regularizes diffusion paths and invokes long-term memory. The Markov approximation allows added control by varying the number of Markov processes linearly combined to approximate fBM. Our evaluations on real image datasets demonstrate that GFDM achieves greater pixel-wise diversity and enhanced image quality, as indicated by a lower FID, offering a promising alternative to traditional diffusion models Gabriel Nobis, Maximilian Springenberg, Marco Aversa, Michael Detzel, Rembert Daems, Roderick Murray-Smith, Shinichi Nakajima, Sebastian Lapuschkin, Stefano Ermon, Tolga Birdal, Manfred Opper, Christoph Knochenhauer, Luis Oala, Wojciech Samek |
NeurIPS | 14 |
| 2024 | Explainable AI for time series via Virtual Inspection LayersabstractThe field of eXplainable Artificial Intelligence (XAI) has witnessed significant advancements in recent years. However, the majority of progress has been concentrated in the domains of computer vision and natural language processing. For time series data, where the input itself is often not interpretable, dedicated XAI research is scarce. In this work, we put forward a virtual inspection layer for transforming the time series to an interpretable representation and allows to propagate relevance attributions to this representation via local XAI methods. In this way, we extend the applicability of XAI methods to domains (e.g. speech) where the input is only interpretable after a transformation. In this work, we focus on the Fourier transformation which, is prominently applied in the preprocessing of time series, with Layer-wise Relevance Propagation (LRP) and refer to our method as DFT-LRP. We demonstrate the usefulness of DFT-LRP in various time series classification settings like audio and electronic health records. We showcase how DFT-LRP reveals differences in the classification strategies of models trained in different domains (e.g., time vs. frequency domain) or helps to discover how models act on spurious correlations in the data. Johanna Vielhaben, Sebastian Lapuschkin, Grégoire Montavon, Wojciech Samek |
Pattern Recognit. | 4 |
| 2024 | Neural Network Coding of Difference Updates for Efficient Distributed Learning CommunicationabstractDistributed learning requires a frequent communication of neural network update data. For this, we present a set of new compression tools, jointly called differential neural network coding (dNNC). dNNC is specifically tailored to efficiently code incremental neural network updates and includes tools for federated BatchNorm folding (FedBNF), structured and unstructured sparsification, tensor row skipping, quantization optimization and temporal adaptation for improved context-adaptive binary arithmetic coding (CABAC). Furthermore, dNNC provides a new parameter update tree (PUT) mechanism, which allows to identify updates for different neural network parameter sub-sets and their relationship in synchronous and asynchronous neural network communication scenarios. Most of these tools have been included into the standardization process of the NNC standard (ISO/IEC 15938-17) edition 2. We benchmark dNNC in multiple federated and split learning scenarios using a variety of NN models and data including vision transformers and large-scale ImageNet experiments: It achieves compression efficiencies of 60% in comparison to the NNC standard edition 1 for transparent coding cases, i.e., without degrading the inference or training performance. This corresponds to a reduction in the size of the NN updates to less than 1% of their original size. Moreover, dNNC reduces the overall energy consumption required for communication in federated learning systems by up to 94%. Daniel Becking, Karsten Müller 0001, Paul Haase, Heiner Kirchhoffer, Gerhard Tech, Wojciech Samek, Heiko Schwarz, Detlev Marpe, Thomas Wiegand 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | From Clustering to Cluster Explanations via Neural NetworksabstractA recent trend in machine learning has been to enrich learned models with the ability to explain their own predictions. The emerging field of explainable AI (XAI) has so far mainly focused on supervised learning, in particular, deep neural network classifiers. In many practical problems, however, the label information is not given and the goal is instead to discover the underlying structure of the data, for example, its clusters. While powerful methods exist for extracting the cluster structure in data, they typically do not answer the question why a certain data point has been assigned to a given cluster. We propose a new framework that can, for the first time, explain cluster assignments in terms of input features in an efficient and reliable manner. It is based on the novel insight that clustering models can be rewritten as neural networks-or "neuralized." Cluster predictions of the obtained networks can then be quickly and accurately attributed to the input features. Several showcases demonstrate the ability of our method to assess the quality of learned clusters and to extract novel insights from the analyzed data and representations. Jacob R. Kauffmann, Malte Esders, Lukas Ruff, Grégoire Montavon, Wojciech Samek, Klaus-Robert Müller |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Decentralized and Incentivized Federated Learning: A Blockchain-Enabled Framework Utilising Compressed Soft-Labels and Peer ConsistencyabstractFederated Learning (FL) has emerged as a powerful paradigm in Artificial Intelligence, facilitating the parallel training of Artificial Neural Networks on edge devices while safeguarding data privacy. Nonetheless, to encourage widespread adoption, Federated Learning Frameworks (FLFs) must tackle (i) the power imbalance between a central authority and its participants, and (ii) the challenge of equitably measuring and incentivizing contributions. Existing approaches to decentralize and incentivize FL processes are hindered by (i) computational overhead and (ii) uncertainty in contribution assessment [1]), limiting FL's scalability beyond use cases where trust between participants and the server is established. This work introduces a cutting-edge, blockchain-enabled federated learning framework that incorporates Federated Knowledge Distillation (FD) with compressed 1-bit soft-labels, aggregated through a smart contract. Furthermore, we present the Peer Truth Serum for Federated Distillation (PTSFD), which cultivates an incentive-compatible ecosystem by rewarding honest participation based on an implicit yet effective comparison of worker contributions. The primary innovation stems from its lightweight architecture that simultaneously promotes decentralization and incentivization, addressing critical challenges in contemporary FL approaches. Leon Witt, Usama Zafar, KuoYeh Shen, Felix Sattler, Dan Li 0001, Wojciech Samek |
IEEE Trans. Serv. Comput. | 7 |
| 2023 | XAI-based Comparison of Audio Event Classifiers with different Input RepresentationsabstractDeep neural networks are a promising tool for Audio Event Classification. In contrast to other data like natural images, there are many sensible and non-obvious representations for audio data, which could serve as input to these models. Due to their black-box nature, the effect of different input representations has so far mostly been investigated by measuring classification performance. In this work, we leverage eXplainable AI (XAI), to understand the underlying classification strategies of models trained on different input representations. Specifically, we compare two model architectures with regard to relevant input features used for Audio Event Detection: one directly processes the signal as the raw waveform, and the other takes its time-frequency spectrogram representation as input. We show how relevance heatmaps obtained via Layer-wise Relevance Propagation uncover representation-dependent decision strategies. With these insights, we can make a well-informed decision about the best model and input representation in terms of robustness and representativity. Further, we can test whether the model’s classification strategies align with human requirements. Annika Frommholz, Fabian Seipel, Sebastian Lapuschkin, Wojciech Samek, Johanna Vielhaben |
CBMI | 4 |
| 2023 | Shortcomings of Top-Down Randomization-Based Sanity Checks for Evaluations of Deep Neural Network ExplanationsabstractWhile the evaluation of explanations is an important step towards trustworthy models, it needs to be done carefully, and the employed metrics need to be well-understood. Specifically model randomization testing can be overinterpreted if regarded as a primary criterion for selecting or discarding explanation methods. To address shortcomings of this test, we start by observing an experimental gap in the ranking of explanation methods between randomization-based sanity checks [1] and model output faithfulness measures (e.g. [20]). We identify limitations of model-randomization-based sanity checks for the purpose of evaluating explanations. Firstly, we show that uninformative attribution maps created with zero pixel-wise covariance easily achieve high scores in this type of checks. Secondly, we show that top-down model randomization preserves scales of forward pass activations with high probability. That is, channels with large activations have a high probility to contribute strongly to the output, even after randomization of the network on top of them. Hence, explanations after randomization can only be expected to differ to a certain extent. This explains the observed experimental gap. In summary, these results demonstrate the inadequacy of model-randomization-based sanity checks as a criterion to rank attribution methods. Alexander Binder, Leander Weber, Sebastian Lapuschkin, Grégoire Montavon, Klaus-Robert Müller, Wojciech Samek |
CVPR | 6 |
| 2023 | Reveal to Revise: An Explainable AI Life Cycle for Iterative Bias Correction of Deep Models
Frederik Pahde, Maximilian Dreyer, Wojciech Samek, Sebastian Lapuschkin |
MICCAI (2) | 3 |
| 2023 | DiffInfinite: Large Mask-Image Synthesis via Parallel Random Patch Diffusion in HistopathologyabstractWe present DiffInfinite, a hierarchical diffusion model that generates arbitrarily large histological images while preserving long-range correlation structural information. Our approach first generates synthetic segmentation masks, subsequently used as conditions for the high-fidelity generative diffusion process. The proposed sampling method can be scaled up to any desired image size while only requiring small patches for fast training. Moreover, it can be parallelized more efficiently than previous large-content generation methods while avoiding tiling artifacts. The training leverages classifier-free guidance to augment a small, sparsely annotated dataset with unlabelled data. Our method alleviates unique challenges in histopathological imaging practice: large-scale information, costly manual annotation, and protective data handling. The biological plausibility of DiffInfinite data is evaluated in a survey by ten experienced pathologists as well as a downstream classification and segmentation task. Samples from the model score strongly on anti-copying metrics which is relevant for the protection of patient data. Marco Aversa, Gabriel Nobis, Miriam Hägele, Kai Standvoss, Mihaela Chirica, Roderick Murray-Smith, Ahmed Alaa 0001, Lukas Ruff, Daniela Ivanova, Wojciech Samek, Frederick Klauschen, Bruno Sanguinetti, Luis Oala |
NeurIPS | 10 |
| 2023 | Decentral and Incentivized Federated Learning Frameworks: A Systematic Literature ReviewabstractThe advent of federated learning (FL) has sparked a new paradigm of parallel and confidential decentralized machine learning (ML) with the potential of utilizing the computational power of a vast number of Internet of Things (IoT), mobile, and edge devices without data leaving the respective device, thus ensuring privacy by design. Yet, simple FL frameworks (FLFs) naively assume an honest central server and altruistic client participation. In order to scale this new paradigm beyond small groups of already entrusted entities toward mass adoption, FLFs must be: 1) truly decentralized and 2) incentivized to participants. This systematic literature review is the first to analyze FLFs that holistically apply both, the blockchain technology to decentralize the process and reward mechanisms to incentivize participation. 422 publications were retrieved by querying 12 major scientific databases. After a systematic filtering process, 40 articles remained for an in-depth examination following our five research questions. To ensure the correctness of our findings, we verified the examination results with the respective authors. Although having the potential to direct the future of distributed and secure artificial intelligence, none of the analyzed FLFs is production ready. The approaches vary heavily in terms of use cases, system design, solved issues, and thoroughness. We provide a systematic approach to classify and quantify differences between FLFs, expose limitations of current works and derive future directions for research in this novel domain. Leon Witt, Mathis Heyer, Kentaroh Toyoda, Wojciech Samek, Dan Li 0001 |
IEEE Internet Things J. | 4 |
| 2023 | Quantus: An Explainable AI Toolkit for Responsible Evaluation of Neural Network Explanations and BeyondabstractThe evaluation of explanation methods is a research topic that has not yet been explored deeply, however, since explainability is supposed to strengthen trust in artificial intelligence, it is necessary to systematically review and compare explanation methods in order to confirm their correctness. Until now, no tool with focus on XAI evaluation exists that exhaustively and speedily allows researchers to evaluate the performance of explanations of neural network predictions. To increase transparency and reproducibility in the field, we therefore built Quantus—a comprehensive, evaluation toolkit in Python that includes a growing, well-organised collection of evaluation metrics and tutorials for evaluating explainable methods. The toolkit has been thoroughly tested and is available under an open-source license on PyPi (or on https://github.com/understandable-machine-intelligence-lab/Quantus/). Anna Hedström, Leander Weber, Daniel Krakowczyk, Dilyara Bareeva, Franz Motzkus, Wojciech Samek, Sebastian Lapuschkin, Marina M.-C. Höhne |
J. Mach. Learn. Res. | 6 |
| 2023 | FedAUX: Leveraging Unlabeled Auxiliary Data in Federated LearningabstractFederated distillation (FD) is a popular novel algorithmic paradigm for Federated learning (FL), which achieves training performance competitive to prior parameter averaging-based methods, while additionally allowing the clients to train different model architectures, by distilling the client predictions on an unlabeled auxiliary set of data into a student model. In this work, we propose FedAUX, an extension to FD, which, under the same set of assumptions, drastically improves the performance by deriving maximum utility from the unlabeled auxiliary data. FedAUX modifies the FD training procedure in two ways: First, unsupervised pre-training on the auxiliary data is performed to find a suitable model initialization for the distributed training. Second, (ε, δ) -differentially private certainty scoring is used to weight the ensemble predictions on the auxiliary data according to the certainty of each client model. Experiments on large-scale convolutional neural networks (CNNs) and transformer models demonstrate that our proposed method achieves remarkable performance improvements over state-of-the-art FL methods, without adding appreciable computation, communication, or privacy cost. For instance, when training ResNet8 on non-independent identically distributed (i.i.d.) subsets of CIFAR10, FedAUX raises the maximum achieved validation accuracy from 30.4% to 78.1%, further closing the gap to centralized training performance. Code is available at https://github.com/fedl-repo/fedaux. Felix Sattler, Tim Korjakow, Roman Rischke, Wojciech Samek |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Langevin Cooling for Unsupervised Domain TranslationabstractDomain translation is the task of finding correspondence between two domains. Several deep neural network (DNN) models, e.g., CycleGAN and cross-lingual language models, have shown remarkable successes on this task under the unsupervised setting-the mappings between the domains are learned from two independent sets of training data in both domains (without paired samples). However, those methods typically do not perform well on a significant proportion of test samples. In this article, we hypothesize that many of such unsuccessful samples lie at the fringe-relatively low-density areas-of data distribution, where the DNN was not trained very well, and propose to perform the Langevin dynamics to bring such fringe samples toward high-density areas. We demonstrate qualitatively and quantitatively that our strategy, called Langevin cooling (L-Cool), enhances state-of-the-art methods in image translation and language translation tasks. Vignesh Srinivasan, Klaus-Robert Müller, Wojciech Samek, Shinichi Nakajima |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Explain to Not Forget: Defending Against Catastrophic Forgetting with XAI
Sami Ede, Serop Baghdadlian, Leander Weber, Dario Zanca, Wojciech Samek, Sebastian Lapuschkin |
CD-MAKE | 6 |
| 2022 | History Dependent Significance Coding for Incremental Neural Network CompressionabstractThis paper presents an improved probability estimation scheme for the entropy coder of Incremental Neural Network Coding (INNC), which is currently under standardization in ISO/IEC MPEG. More specifically, the paper first analyzes the compression performance of INNC and how the bitstream size relates to the neural network (NN) layers. For the layers requiring the most bits, it analyzes the coded NN weight updates and their temporal dependencies. Major finding is that the probability of a significant (i.e., non-zero) update for a weight can depend considerably on whether the weight has been updated before. Based on this finding, the paper proposes a new probability estimation scheme: Depending on whether a significant update has been received before (i.e., based on the weight’s history), the entropy coder models the probability for a current significant update differently. This scheme achieves a bitstream size reduction of about 2% and 1% in a transfer and a federated learning scenario, respectively, without any accuracy loss or significant complexity increase. Therefore, MPEG adopted our history dependent significance probability (HDSP) scheme to its emerging standard for INNC. Gerhard Tech, Paul Haase, Daniel Becking, Heiner Kirchhoffer, Karsten Müller 0001, Jonathan Pfaff, Heiko Schwarz, Wojciech Samek, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 8 |
| 2022 | Explaining Machine Learning Models for Clinical Gait AnalysisabstractMachine Learning (ML) is increasingly used to support decision-making in the healthcare sector. While ML approaches provide promising results with regard to their classification performance, most share a central limitation, their black-box character. This article investigates the usefulness of Explainable Artificial Intelligence (XAI) methods to increase transparency in automated clinical gait classification based on time series. For this purpose, predictions of state-of-the-art classification methods are explained with a XAI method called Layer-wise Relevance Propagation (LRP). Our main contribution is an approach that explains class-specific characteristics learned by ML models that are trained for gait classification. We investigate several gait classification tasks and employ different classification methods, i.e., Convolutional Neural Network, Support Vector Machine, and Multi-layer Perceptron. We propose to evaluate the obtained explanations with two complementary approaches: a statistical analysis of the underlying data using Statistical Parametric Mapping and a qualitative evaluation by two clinical experts. A gait dataset comprising ground reaction force measurements from 132 patients with different lower-body gait disorders and 62 healthy controls is utilized. Our experiments show that explanations obtained by LRP exhibit promising statistical properties concerning inter-class discriminativity and are also in line with clinically relevant biomechanical gait characteristics. Djordje Slijepcevic, Fabian Horst, Sebastian Lapuschkin, Brian Horsak, Anna-Maria Raberger, Andreas Kranzl, Wojciech Samek, Christian Breiteneder, Wolfgang Immanuel Schöllhorn, Matthias Zeppelzauer |
ACM Trans. Comput. Heal. | 7 |
| 2022 | Overview of the Neural Network Compression and Representation (NNR) StandardabstractNeural Network Coding and Representation (NNR) is the first international standard for efficient compression of neural networks (NNs). The standard is designed as a toolbox of compression methods, which can be used to create coding pipelines. It can be either used as an independent coding framework (with its own bitstream format) or together with external neural network formats and frameworks. For providing the highest degree of flexibility, the network compression methods operate per parameter tensor in order to always ensure proper decoding, even if no structure information is provided. The NNR standard contains compression-efficient quantization and deep context-adaptive binary arithmetic coding (DeepCABAC) as core encoding and decoding technologies, as well as neural network parameter pre-processing methods like sparsification, pruning, low-rank decomposition, unification, local scaling and batch norm folding. NNR achieves a compression efficiency of more than 97% for transparent coding cases, i.e. without degrading classification quality, such as top-1 or top-5 accuracies. This paper provides an overview of the technical features and characteristics of NNR. Heiner Kirchhoffer, Paul Haase, Wojciech Samek, Karsten Müller 0001, Hamed Rezazadegan Tavakoli, Francesco Cricri, Emre Aksu, Miska M. Hannuksela, Wei Jiang 0001, Wei Wang 0311, Shan Liu 0001, Swayambhoo Jain, Shahab Hamidi-Rad, Fabien Racapé, Werner Bailer |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Encoder Optimizations For The NNR Standard On Neural Network CompressionabstractThe novel Neural Network Compression and Representation Standard (NNR), recently issued by ISO/IEC MPEG, achieves very high coding gains, compressing neural networks to 5% in size without accuracy loss. The underlying NNR encoder technology includes parameter quantization, followed by efficient arithmetic coding, namely DeepCABAC. In addition, NNR also allows very flexible adaptations, such as signaling specific local scaling values, setting quantization parameters per tensor rather than per network and supporting specific parameter fusion operations. This paper presents our new approach for optimally deriving these parameters, namely the derivation of parameters for local scaling adaptation (LSA), inference-optimized quantization (IOQ), and batch-norm folding (BNF). By allowing inference and fine tuning within the encoding process, quantization errors are reduced and the NNR coding efficiency is further improved to a compressed bitstream size of only 3% in comparison to the original model size. Paul Haase, Daniel Becking, Heiner Kirchhoffer, Karsten Müller 0001, Heiko Schwarz, Wojciech Samek, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 6 |
| 2021 | Robustifying models against adversarial attacks by Langevin dynamics
Vignesh Srinivasan, Csaba Rohrer, Arturo Marbán, Klaus-Robert Müller, Wojciech Samek, Shinichi Nakajima |
Neural Networks | 5 |
| 2021 | A Unifying Review of Deep and Shallow Anomaly DetectionabstractDeep learning approaches to anomaly detection (AD) have recently improved the state of the art in detection performance on complex data sets, such as large collections of images or text. These results have sparked a renewed interest in the AD problem and led to the introduction of a great variety of new methods. With the emergence of numerous such methods, including approaches based on generative models, one-class classification, and reconstruction, there is a growing need to bring methods of this field into a systematic and unified perspective. In this review, we aim to identify the common underlying principles and the assumptions that are often made implicitly by various methods. In particular, we draw connections between classic “shallow” and novel deep approaches and show how this relation might cross-fertilize or extend both directions. We further provide an empirical assessment of major existing methods that are enriched by the use of recent explainability techniques and present specific worked-through examples together with practical advice. Finally, we outline critical open challenges and identify specific paths for future research in AD. Lukas Ruff, Jacob R. Kauffmann, Robert A. Vandermeulen, Grégoire Montavon, Wojciech Samek, Marius Kloft, Thomas G. Dietterich, Klaus-Robert Müller |
Proc. IEEE | 5 |
| 2021 | Explaining Deep Neural Networks and Beyond: A Review of Methods and ApplicationsabstractWith the broader and highly successful usage of machine learning (ML) in industry and the sciences, there has been a growing demand for explainable artificial intelligence (XAI). Interpretability and explanation methods for gaining a better understanding of the problem-solving abilities and strategies of nonlinear ML, in particular, deep neural networks, are, therefore, receiving increased attention. In this work, we aim to: 1) provide a timely overview of this active emerging field, with a focus on “post hoc” explanations, and explain its theoretical foundations; 2) put interpretability algorithms to a test both from a theory and comparative evaluation perspective using extensive simulations; 3) outline best practice aspects, i.e., how to best include interpretation methods into the standard usage of ML; and 4) demonstrate successful usage of XAI in a representative selection of application scenarios. Finally, we discuss challenges and possible future directions of this exciting foundational field of ML. Wojciech Samek, Grégoire Montavon, Sebastian Lapuschkin, Christopher J. Anders, Klaus-Robert Müller |
Proc. IEEE | 1 |
| 2021 | Pruning by explaining: A novel criterion for deep neural network pruningabstractThe success of convolutional neural networks (CNNs) in various applications is accompanied by a significant increase in computation and parameter storage costs. Recent efforts to reduce these overheads involve pruning and compressing the weights of various layers while at the same time aiming to not sacrifice performance. In this paper, we propose a novel criterion for CNN pruning inspired by neural network interpretability: The most relevant units, i.e. weights or filters, are automatically found using their relevance scores obtained from concepts of explainable AI (XAI). By exploring this idea, we connect the lines of interpretability and model compression research. We show that our proposed method can efficiently prune CNN models in transfer-learning setups in which networks pre-trained on large corpora are adapted to specialized tasks. The method is evaluated on a broad range of computer vision datasets. Notably, our novel criterion is not only competitive or better compared to state-of-the-art pruning criteria when successive retraining is performed, but clearly outperforms these previous criteria in the resource-constrained application scenario in which the data of the task to be transferred to is very scarce and one chooses to refrain from fine-tuning. Our method is able to compress the model iteratively while maintaining or even improving accuracy. At the same time, it has a computational cost in the order of gradient computation and is comparatively simple to apply without the need for tuning hyperparameters for pruning. Seul-Ki Yeom, Philipp Seegerer, Sebastian Lapuschkin, Alexander Binder, Simon Wiedemann, Klaus-Robert Müller, Wojciech Samek |
Pattern Recognit. | 7 |
| 2021 | Deep Learning for ECG Analysis: Benchmarks and Insights from PTB-XLabstractElectrocardiography (ECG) is a very common, non-invasive diagnostic procedure and its interpretation is increasingly supported by algorithms. The progress in the field of automatic ECG analysis has up to now been hampered by a lack of appropriate datasets for training as well as a lack of well-defined evaluation procedures to ensure comparability of different algorithms. To alleviate these issues, we put forward first benchmarking results for the recently published, freely accessible clinical 12-lead ECG dataset PTB-XL, covering a variety of tasks from different ECG statement prediction tasks to age and sex prediction. Among the investigated deep-learning-based timeseries classification algorithms, we find that convolutional neural networks, in particular resnet- and inception-based architectures, show the strongest performance across all tasks. We find consistent results on the ICBEB2018 challenge ECG dataset and discuss prospects of transfer learning using classifiers pretrained on PTB-XL. These benchmarking results are complemented by deeper insights into the classification algorithm in terms of hidden stratification, model uncertainty and an exploratory interpretability analysis, which provide connecting points for future research on the dataset. Our results emphasize the prospects of deep-learning-based algorithms in the field of ECG analysis, not only in terms of quantitative accuracy but also in terms of clinically equally important further quality metrics such as uncertainty quantification and interpretability. With this resource, we aim to establish the PTB-XL dataset as a resource for structured benchmarking of ECG analysis algorithms and encourage other researchers in the field to join these efforts. Nils Strodthoff, Patrick Wagner 0002, Tobias Schaeffter, Wojciech Samek |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | Clustered Federated Learning: Model-Agnostic Distributed Multitask Optimization Under Privacy ConstraintsabstractFederated learning (FL) is currently the most widely adopted framework for collaborative training of (deep) machine learning models under privacy constraints. Albeit its popularity, it has been observed that FL yields suboptimal results if the local clients' data distributions diverge. To address this issue, we present clustered FL (CFL), a novel federated multitask learning (FMTL) framework, which exploits geometric properties of the FL loss surface to group the client population into clusters with jointly trainable data distributions. In contrast to existing FMTL approaches, CFL does not require any modifications to the FL communication protocol to be made, is applicable to general nonconvex objectives (in particular, deep neural networks), does not require the number of clusters to be known a priori, and comes with strong mathematical guarantees on the clustering quality. CFL is flexible enough to handle client populations that vary over time and can be implemented in a privacy-preserving way. As clustering is only performed after FL has converged to a stationary point, CFL can be viewed as a postprocessing method that will always achieve greater or equal performance than conventional FL by allowing clients to arrive at more specialized models. We verify our theoretical analysis in experiments with deep convolutional and recurrent neural networks on commonly used FL data sets. Felix Sattler, Klaus-Robert Müller, Wojciech Samek |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Benign Examples: Imperceptible Changes Can Enhance Image Translation Performance
Vignesh Srinivasan, Klaus-Robert Müller, Wojciech Samek, Shinichi Nakajima |
AAAI | 3 |
| 2020 | On the Byzantine Robustness of Clustered Federated LearningabstractFederated Learning (FL) is currently the most widely adopted framework for collaborative training of (deep) machine learning models under privacy constraints. Albeit it's popularity, it has been observed that Federated Learning yields suboptimal results if the local clients' data distributions diverge. The recently proposed Clustered Federated Learning Framework addresses this issue, by separating the client population into different groups based on the pairwise cosine similarities between their parameter updates. In this work we investigate the application of CFL to byzantine settings, where a subset of clients behaves unpredictably or tries to disturb the joint training effort in an directed or undirected way. We perform experiments with deep neural networks on common Federated Learning datasets which demonstrate that CFL (without modifications) is able to reliably detect byzantine clients and remove them from training. Felix Sattler, Klaus-Robert Müller, Thomas Wiegand 0001, Wojciech Samek |
ICASSP | 4 |
| 2020 | Dependent Scalar Quantization For Neural Network CompressionabstractRecent approaches to compression of deep neural networks, like the emerging standard on compression of neural networks for multimedia content description and analysis (MPEG-7 part 17), apply scalar quantization and entropy coding of the quantization indexes. In this paper we present an advanced method for quantization of neural network parameters, which applies dependent scalar quantization (DQ) or trellis-coded quantization (TCQ), and an improved context modeling for the entropy coding of the quantization indexes. We show that the proposed method achieves 5.778% bitrate reduction and virtually no loss (0.37%) of network performance in average, compared to the baseline methods of the second test model (NCTM) of MPEG-7 part 17 for relevant working points. Paul Haase, Heiko Schwarz, Heiner Kirchhoffer, Simon Wiedemann, Talmaj Marinc, Arturo Marbán, Karsten Müller 0001, Wojciech Samek, Detlev Marpe, Thomas Wiegand 0001 |
ICIP | 8 |
| 2020 | Deepcabac: Plug & Play Compression of Neural Network Weights and Weight UpdatesabstractAn increasing number of distributed machine learning applications require efficient communication of neural network parameterizations. DeepCABAC, an algorithm in the current working draft of the emerging MPEG-7 part 17 standard for compression of neural networks for multimedia content description and analysis, has demonstrated high compression gains for a variety of neural network models. In this paper we propose a method for employing DeepCABAC in a Federated Learning scenario for the exchange of intermediate differential parameterizations. Furthermore, we discuss the efficiency of DeepCABAC when compressing trained neural networks. Our experiments on large neural networks show that in both scenarios, DeepCABAC achieves competitive compression rates, without degrading the network accuracy. David Neumann, Felix Sattler, Heiner Kirchhoffer, Simon Wiedemann, Karsten Müller 0001, Heiko Schwarz, Thomas Wiegand 0001, Detlev Marpe, Wojciech Samek |
ICIP | 9 |
| 2020 | Understanding Integrated Gradients with SmoothTaylor for Deep Neural Network AttributionabstractIntegrated Gradients as an attribution method for deep neural network models offers simple implementability. However, it suffers from noisiness of explanations which affects the ease of interpretability. The SmoothGrad technique is proposed to solve the noisiness issue and smoothen the attribution maps of any gradient-based attribution method. In this paper, we present SmoothTaylor as a novel theoretical concept bridging Integrated Gradients and SmoothGrad, from the Taylor's theorem perspective. We apply the methods to the image classification problem, using the ILSVRC2012 ImageNet object recognition dataset, and a couple of pretrained image models to generate attribution maps. These attribution maps are empirically evaluated using quantitative measures for sensitivity and noise level. We further propose adaptive noising to optimize for the noise scale hyperparameter value. From our experiments, we find that the SmoothTaylor approach together with adaptive noising is able to generate better quality saliency maps with lesser noise and higher sensitivity to the relevant points in the input space as compared to Integrated Gradients. Gary S. W. Goh, Sebastian Lapuschkin, Leander Weber, Wojciech Samek, Alexander Binder |
ICPR | 4 |
| 2020 | Explanation-Guided Training for Cross-Domain Few-Shot ClassificationabstractCross-domain few-shot classification task (CD-FSC) combines few-shot classification with the requirement to generalize across domains represented by datasets. This setup faces challenges originating from the limited labeled data in each class and, additionally, from the domain shift between training and test sets. In this paper, we introduce a novel training approach for existing FSC models. It leverages on the explanation scores, obtained from existing explanation methods when applied to the predictions of FSC models, computed for intermediate feature maps of the models. Firstly, we tailor the layer-wise relevance propagation (LRP) method to explain the predictions of FSC models. Secondly, we develop a model-agnostic explanation-guided training strategy that dynamically finds and emphasizes the features which are important for the predictions. Our contribution does not target a novel explanation method but lies in a novel application of explanations for the training phase. We show that explanation-guided training effectively improves the model generalization. We observe improved accuracy for three different FSC models: RelationNet, cross attention network, and a graph neural network-based formulation, on five few-shot learning datasets: miniImagenet, CUB, Cars, Places, and Plantae. The source code is available https://github.com/SunJiamei/few-shot-lrp-guided. Jiamei Sun, Sebastian Lapuschkin, Wojciech Samek, Yunqing Zhao, Ngai-Man Cheung, Alexander Binder |
ICPR | 3 |
| 2020 | Towards Best Practice in Explaining Neural Network Decisions with LRPabstractWithin the last decade, neural network based predictors have demonstrated impressive - and at times superhuman - capabilities. This performance is often paid for with an intransparent prediction process and thus has sparked numerous contributions in the novel field of explainable artificial intelligence (XAI). In this paper, we focus on a popular and widely used method of XAI, the Layer-wise Relevance Propagation (LRP). Since its initial proposition LRP has evolved as a method, and a best practice for applying the method has tacitly emerged, based however on humanly observed evidence alone. In this paper we investigate - and for the first time quantify - the effect of this current best practice on feedforward neural networks in a visual object detection setting. The results verify that the layer-dependent approach to LRP applied in recent literature better represents the model's reasoning, and at the same time increases the object localization and class discriminativity of LRP. Maximilian Kohlbrenner, Alexander Bauer 0001, Shinichi Nakajima, Alexander Binder, Wojciech Samek, Sebastian Lapuschkin |
IJCNN | 5 |
| 2020 | UDSMProt: universal deep sequence models for protein classificationabstractMOTIVATION: Inferring the properties of a protein from its amino acid sequence is one of the key problems in bioinformatics. Most state-of-the-art approaches for protein classification are tailored to single classification tasks and rely on handcrafted features, such as position-specific-scoring matrices from expensive database searches. We argue that this level of performance can be reached or even be surpassed by learning a task-agnostic representation once, using self-supervised language modeling, and transferring it to specific tasks by a simple fine-tuning step. RESULTS: We put forward a universal deep sequence model that is pre-trained on unlabeled protein sequences from Swiss-Prot and fine-tuned on protein classification tasks. We apply it to three prototypical tasks, namely enzyme class prediction, gene ontology prediction and remote homology and fold detection. The proposed method performs on par with state-of-the-art algorithms that were tailored to these specific tasks or, for two out of three tasks, even outperforms them. These results stress the possibility of inferring protein properties from the sequence alone and, on more general grounds, the prospects of modern natural language processing methods in omics. Moreover, we illustrate the prospects for explainable machine learning methods in this field by selected case studies. AVAILABILITY AND IMPLEMENTATION: Source code is available under https://github.com/nstrodt/UDSMProt. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nils Strodthoff, Patrick Wagner 0002, Markus Wenzel 0002, Wojciech Samek |
Bioinform. | 4 |
| 2020 | USMPep: universal sequence models for major histocompatibility complex binding affinity predictionabstractBACKGROUND: Immunotherapy is a promising route towards personalized cancer treatment. A key algorithmic challenge in this process is to decide if a given peptide (neoepitope) binds with the major histocompatibility complex (MHC). This is an active area of research and there are many MHC binding prediction algorithms that can predict the MHC binding affinity for a given peptide to a high degree of accuracy. However, most of the state-of-the-art approaches make use of complicated training and model selection procedures, are restricted to peptides of a certain length and/or rely on heuristics. RESULTS: We put forward USMPep, a simple recurrent neural network that reaches state-of-the-art approaches on MHC class I binding prediction with a single, generic architecture and even a single set of hyperparameters both on IEDB benchmark datasets and on the very recent HPV dataset. Moreover, the algorithm is competitive for a single model trained from scratch, while ensembling multiple regressors and language model pretraining can still slightly improve the performance. The direct application of the approach to MHC class II binding prediction shows a solid performance despite of limited training data. CONCLUSIONS: We demonstrate that competitive performance in MHC binding affinity prediction can be reached with a standard architecture and training procedure without relying on any heuristics. Johanna Vielhaben, Markus Wenzel 0002, Wojciech Samek, Nils Strodthoff |
BMC Bioinform. | 3 |
| 2020 | Accurate and robust neural networks for face morphing attack detectionabstractArtificial neural networks tend to use only what they need for a task. For example, to recognize a rooster, a network might only considers the rooster’s red comb and wattle and ignores the rest of the animal. This makes them vulnerable to attacks on their decision making process and can worsen their generality. Thus, this phenomenon has to be considered during the training of networks, especially in safety and security related applications. In this paper, we propose neural network training schemes, which are based on different alternations of the training data, to increase robustness and generality. Precisely, we limit the amount and position of information available to the neural network for the decision making process and study their effects on the accuracy, generality, and robustness against semantic and black box attacks for the particular example of face morphing attacks. In addition, we exploit layer-wise relevance propagation (LRP) to analyze the differences in the decision making process of the differently trained neural networks. A face morphing attack is an attack on a biometric facial recognition system, where the system is fooled to match two different individuals with the same synthetic face image. Such a synthetic image can be created by aligning and blending images of the two individuals that should be matched with this image. We train neural networks for face morphing attack detection using our proposed training schemes and show that they lead to an improvement of robustness against attacks on neural networks. Using LRP, we show that the improved training forces the networks to develop and use reliable models for all regions of the analyzed image. This redundancy in representation is of crucial importance to security related applications. Clemens Seibold, Wojciech Samek, Anna Hilsmann, Peter Eisert |
J. Inf. Secur. Appl. | 2 |
| 2020 | Robust and Communication-Efficient Federated Learning From Non-i.i.d. DataabstractFederated learning allows multiple parties to jointly train a deep learning model on their combined data, without any of the participants having to reveal their local data to a centralized server. This form of privacy-preserving collaborative learning, however, comes at the cost of a significant communication overhead during training. To address this problem, several compression methods have been proposed in the distributed training literature that can reduce the amount of required communication by up to three orders of magnitude. These existing methods, however, are only of limited utility in the federated learning setting, as they either only compress the upstream communication from the clients to the server (leaving the downstream communication uncompressed) or only perform well under idealized conditions, such as i.i.d. distribution of the client data, which typically cannot be found in federated learning. In this article, we propose sparse ternary compression (STC), a new compression framework that is specifically designed to meet the requirements of the federated learning environment. STC extends the existing compression technique of top- k gradient sparsification with a novel mechanism to enable downstream compression as well as ternarization and optimal Golomb encoding of the weight updates. Our experiments on four different learning tasks demonstrate that STC distinctively outperforms federated averaging in common federated learning scenarios. These results advocate for a paradigm shift in federated optimization toward high-frequency low-bitwidth communication, in particular in the bandwidth-constrained learning environments. Felix Sattler, Simon Wiedemann, Klaus-Robert Müller, Wojciech Samek |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Compact and Computationally Efficient Representation of Deep Neural NetworksabstractAt the core of any inference procedure, deep neural networks are dot product operations, which are the component that requires the highest computational resources. For instance, deep neural networks, such as VGG-16, require up to 15-G operations in order to perform the dot products present in a single forward pass, which results in significant energy consumption and thus limits their use in resource-limited environments, e.g., on embedded devices or smartphones. One common approach to reduce the complexity of the inference is to prune and quantize the weight matrices of the neural network. Usually, this results in matrices whose entropy values are low, as measured relative to the empirical probability mass distribution of its elements. In order to efficiently exploit such matrices, one usually relies on, inter alia, sparse matrix representations. However, most of these common matrix storage formats make strong statistical assumptions about the distribution of the elements; therefore, cannot efficiently represent the entire set of matrices that exhibit low-entropy statistics (thus, the entire set of compressed neural network weight matrices). In this paper, we address this issue and present new efficient representations for matrices with low-entropy statistics. Alike sparse matrix data structures, these formats exploit the statistical properties of the data in order to reduce the size and execution complexity. Moreover, we show that the proposed data structures can not only be regarded as a generalization of sparse formats but are also more energy and time efficient under practically relevant assumptions. Finally, we test the storage requirements and execution performance of the proposed formats on compressed neural networks and compare them to dense and sparse representations. We experimentally show that we are able to attain up to ×42 compression ratios, ×5 speed ups, and ×90 energy savings when we lossless convert the state-of-the-art networks, such as AlexNet, VGG-16, ResNet152, and DenseNet, into the new data structures and benchmark their respective dot product. Simon Wiedemann, Klaus-Robert Müller, Wojciech Samek |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Multi-Kernel Prediction Networks for Denoising of Burst ImagesabstractIn low light or short-exposure photography the image is often corrupted by noise. While longer exposure helps reduce the noise, it can produce blurry results due to the object and camera motion. The reconstruction of a noise-less image is an ill posed problem. Recent approaches for image denoising aim to predict kernels which are convolved with a set of successively taken images (burst) to obtain a clear image. We propose a deep neural network based approach called Multi-Kernel Prediction Networks (MKPN) for burst image denoising. MKPN predicts kernels of not just one size but of varying sizes and performs fusion of these different kernels resulting in one kernel per pixel. The advantages of our method are two fold: (a) the different sized kernels help in extracting different information from the image which results in better reconstruction and (b) kernel fusion assures retaining of the extracted information while maintaining computational efficiency. Experimental results reveal that MKPN outperforms state-of-the-art on our synthetic datasets with different noise levels. Talmaj Marinc, Vignesh Srinivasan, Serhan Gul, Cornelius Hellge, Wojciech Samek |
ICIP | 5 |
| 2019 | Sparse Binary Compression: Towards Distributed Deep Learning with minimal CommunicationabstractCurrently, progressively larger deep neural networks are trained on ever growing data corpora. In result, distributed training schemes are becoming increasingly relevant. A major issue in distributed training is the limited communication bandwidth between contributing nodes or prohibitive communication cost in general. To mitigate this problem we propose Sparse Binary Compression (SBC), a compression framework that allows for a drastic reduction of communication cost for distributed training. SBC combines existing techniques of communication delay and gradient sparsification with a novel binarization method and optimal weight update encoding to push compression gains to new limits. By doing so, our method also allows us to smoothly trade-off gradient sparsity and temporal sparsity to adapt to the requirements of the learning task. Our experiments show, that SBC can reduce the upstream communication on a variety of convolutional and recurrent neural network architectures by more than four orders of magnitude without significantly harming the convergence speed in terms of forward-backward passes. For instance, we can train ResNet50 on ImageNet in the same number of iterations to the baseline accuracy, using ×3531 less bits or train it to a 1% lower accuracy using ×37208 less bits. In the latter case, the total upstream communication required is cut from 125 terabytes to 3.35 gigabytes for every participating client. Felix Sattler, Simon Wiedemann, Klaus-Robert Müller, Wojciech Samek |
IJCNN | 4 |
| 2019 | Entropy-Constrained Training of Deep Neural NetworksabstractMotivated by the Minimum Description Length (MDL) principle, we first derive an expression for the entropy of a neural network which measures its complexity explicitly in terms of its bit-size. Then, we formalize the problem of neural network compression as an entropy-constrained optimization objective. This objective generalizes many of the currently proposed compression techniques in the literature, in that pruning or reducing the cardinality of the weight elements can be seen as special cases of entropy reduction methods. Furthermore, we derive a continuous relaxation of the objective, which allows us to minimize it using gradient-based optimization techniques. Finally, we show that we can reach compression results, which are competitive with those obtained using state-of-the-art techniques, on different network architectures and data sets, e.g. achieving×71 compression gains on a VGG-like architecture. Simon Wiedemann, Arturo Marbán, Klaus-Robert Müller, Wojciech Samek |
IJCNN | 4 |
| 2019 | DRAU: Dual Recurrent Attention Units for Visual Question Answering
Ahmed Osman 0003, Wojciech Samek |
Comput. Vis. Image Underst. | 2 |
| 2019 | iNNvestigate Neural Networks!abstractIn recent years, deep neural networks have revolutionized many application domains of machine learning and are key components of many critical decision or predictive processes. Therefore, it is crucial that domain specialists can understand and analyze actions and predictions, even of the most complex neural network architectures. Despite these arguments neural networks are often treated as black boxes. In the attempt to alleviate this shortcoming many analysis methods were proposed, yet the lack of reference implementations often makes a systematic comparison between the methods a major effort. The presented library innvestigate addresses this by providing a common interface and out-of-the-box implementation for many analysis methods, including the reference implementation for PatternNet and PatternAttribution as well as for LRP-methods. To demonstrate the versatility of innvestigate, we provide an analysis of image classifications for variety of state-of-the-art neural network architectures. Maximilian Alber, Sebastian Lapuschkin, Philipp Seegerer, Miriam Hägele, Kristof Schütt, Grégoire Montavon, Wojciech Samek, Klaus-Robert Müller, Sven Dähne, Pieter-Jan Kindermans |
J. Mach. Learn. Res. | 7 |
| 2019 | Enhanced Machine Learning Techniques for Early HARQ Feedback Prediction in 5GabstractWe investigate Early Hybrid Automatic Repeat reQuest (E-HARQ) feedback schemes enhanced by machine learning techniques as a path towards ultra-reliable and low-latency communication (URLLC). To this end, we propose machine learning methods to predict the outcome of the decoding process ahead of the end of the transmission. We discuss different input features and classification algorithms ranging from traditional methods to newly developed supervised autoencoders. These methods are evaluated based on their prospects of complying with the URLLC requirements of effective block error rates below 10-5at small latency overheads. We provide realistic performance estimates in a system model incorporating scheduling effects to demonstrate the feasibility of E-HARQ across different signal-to-noise ratios, subcode lengths, channel conditions and system loads, and show the benefit over regular HARQ and existing E-HARQ schemes without machine learning. Nils Strodthoff, Baris Göktepe, Thomas Schierl, Cornelius Hellge, Wojciech Samek |
IEEE J. Sel. Areas Commun. | 5 |
| 2018 | Transferring Information Between Neural NetworksabstractThis paper investigates techniques to transfer information between deep neural networks. We demonstrate that a student network, which has access to information computed by a teacher network on the training data, learns faster, can be less deep and requires less labeled examples to achieve a given performance level. For that we force the student to mimic the teacher by adding a penalty term to the student's objective. We evaluate different penalty terms: (1) mean squared error between the cost gradients, (2) the Jacobian of the pre-softmax layer, (3) its row-summed version, (4) the cost gradient differences to standard double backpropagation and (5) a targeted double backpropagation via gradient derived masks. The Jacobian method improves the accuracy proportional to the difference in training examples, in contrast to the cost gradient. If the difference in accuracy between teacher and student is large enough, we find an improvement from the Jacobian information, even if both had seen the same training data. This indicates that information transfer has a regularization effect. Christopher Ehmann, Wojciech Samek |
ICASSP | 2 |
| 2018 | Neural Network-Based Estimation of Distortion Sensitivity for Image Quality PredictionabstractDue to its computational simplicity, the PSNR is a popular and widely used image quality measure, although it correlates poorly with perceived visual quality. Distortion sensitivity, a reference image specific property, can be used to compensate for the lack of perceptual relevance of the PSNR. Based on the functional mapping between perceptual and computational quality a deep convolutional neural network is used to estimate patchwise distortion sensitivity. The local estimates are used for an imagewise perceptual adaptation of the PSNR. The performance of the proposed estimation approach is evaluated on the LIVE and TID2013 databases and shows comparable or superior performance as compared to benchmark image quality measures. Sebastian Bosse, Sören Becker 0001, Zacharias V. Fisches, Wojciech Samek, Thomas Wiegand 0001 |
ICIP | 4 |
| 2018 | Estimation of Interaction Forces in Robotic Surgery using a Semi-Supervised Deep Neural Network ModelabstractProviding force feedback as a feature in current Robot-Assisted Minimally Invasive Surgery systems still remains a challenge. In recent years, Vision-Based Force Sensing (VBFS) has emerged as a promising approach to address this problem. Existing methods have been developed in a Supervised Learning (SL) setting. Nonetheless, most of the video sequences related to robotic surgery are not provided with ground-truth force data, which can be easily acquired in a controlled environment. A powerful approach to process unlabeled video sequences and find a compact representation for each video frame relies on using an Unsupervised Learning (UL) method. Afterward, a model trained in an SL setting can take advantage of the available ground-truth force data. In the present work, UL and SL techniques are used to investigate a model in a Semi-Supervised Learning (SSL) framework, consisting of an encoder network and a Long-Short Term Memory (LSTM) network. First, a Convolutional Auto-Encoder (CAE) is trained to learn a compact representation for each RGB frame in a video sequence. To facilitate the reconstruction of high and low frequencies found in images, this CAE is optimized using an adversarial framework and a L1-loss, respectively. Thereafter, the encoder network of the CAE is serially connected with an LSTM network and trained jointly to minimize the difference between ground-truth and estimated force data. Datasets addressing the force estimation task are scarce. Therefore, the experiments have been validated in a custom dataset. The results suggest that the proposed approach is promising. Arturo Marbán, Vignesh Srinivasan, Wojciech Samek, Josep Fernández, Alicia Casals |
IROS | 3 |
| 2018 | On the Stimulation Frequency in SSVEP-based Image Quality AssessmentabstractSteady-state visual evoked potentials (SSVEP) are brain responses elicited by periodic visual stimuli. Recently it was shown that the use of SSVEP in quality studies allows for accurate psychophysiological assessment of perceived visual quality, but the influence of the stimulation frequency is still unclear. This paper studies experimentally the relation between the SNR of the neural signal and the stimulation frequency in an psychophysiological quality assessment setup. For various source images tested at different distortion magnitudes over the range of 6 different stimulation frequencies, we show physiologically plausible results that provide insights into the temporal dynamics of neural distortion processing. Our findings inform a rational choice of stimulation frequency in SSVEP-based image quality assessment studies. This potentially improves the experimental setup of future image quality assessment studies exploiting the SSVEP paradigm. Sebastian Bosse, Milena T. Bagdasarian, Wojciech Samek, Gabriel Curio, Thomas Wiegand 0001 |
QoMEX | 3 |
| 2018 | Assessing Perceived Image Quality Using Steady-State Visual Evoked Potentials and Spatio-Spectral DecompositionabstractSteady-state visual evoked potentials (SSVEPs) are neural responses, measurable using electroencephalography (EEG), that are directly linked to sensory processing of visual stimuli. In this paper, SSVEP is used to assess the perceived quality of texture images. The EEG-based assessment method is compared with conventional methods, and recorded EEG data are correlated to obtained mean opinion scores (MOSs). A dimensionality reduction technique for EEG data called spatio-spectral decomposition (SSD) is adapted for the SSVEP framework and used to extract physiologically meaningful and plausible neural components from the EEG recordings. It is shown that the use of SSD not only increases the correlation between neural features and MOS to r = -0.93, but also solves the problem of channel selection in an EEG-based image-quality assessment. Sebastian Bosse, Laura Acqualagna, Wojciech Samek, Anne Porbadnigk, Gabriel Curio, Benjamin Blankertz, Klaus-Robert Müller, Thomas Wiegand 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Deep Neural Networks for No-Reference and Full-Reference Image Quality AssessmentabstractWe present a deep neural network-based approach to image quality assessment (IQA). The network is trained end-to-end and comprises ten convolutional layers and five pooling layers for feature extraction, and two fully connected layers for regression, which makes it significantly deeper than related IQA models. Unique features of the proposed architecture are that: 1) with slight adaptations it can be used in a no-reference (NR) as well as in a full-reference (FR) IQA setting and 2) it allows for joint learning of local quality and local weights, i.e., relative importance of local quality to the global quality estimate, in an unified framework. Our approach is purely data-driven and does not rely on hand-crafted features or other types of prior domain knowledge about the human visual system or image statistics. We evaluate the proposed approach on the LIVE, CISQ, and TID2013 databases as well as the LIVE In the wild image quality challenge database and show superior performance to state-of-the-art NR and FR IQA methods. Finally, cross-database evaluation shows a high ability to generalize between different databases, indicating a high robustness of the learned features. Sebastian Bosse, Dominique Maniry, Klaus-Robert Müller, Thomas Wiegand 0001, Wojciech Samek |
IEEE Trans. Image Process. | 5 |
| 2017 | Interpretable human action recognition in compressed domainabstractCompressed domain human action recognition algorithms are extremely efficient, because they only require a partial decoding of the video bit stream. However, the question what exactly makes these algorithms decide for a particular action is still a mystery. In this paper, we present a general method, Layer-wise Relevance Propagation (LRP), to understand and interpret action recognition algorithms and apply it to a state-of-the-art compressed domain method based on Fisher vector encoding and SVM classification. By using LRP, the classifiers decisions are propagated back every step in the action recognition pipeline until the input is reached. This methodology allows to identify where and when the important (from the classifier's perspective) action happens in the video. To our knowledge, this is the first work to interpret a compressed domain action recognition algorithm. We evaluate our method on the HMDB51 dataset and show that in many cases a few significant frames contribute most towards the prediction of the video to a particular class. Vignesh Srinivasan, Sebastian Lapuschkin, Cornelius Hellge, Klaus-Robert Müller, Wojciech Samek |
ICASSP | 5 |
| 2017 | A perceptually relevant shearlet-based adaptation of the PSNRabstractAlthough being one of the simplest and most widely used image quality metrics (IQMs) the peak signal-to-noise ratio (PSNR) correlates only poorly with visual quality as perceived by humans. Based on an analysis of the non-linear mapping from PSNR to mean opinion scores (MOS) we identify a functional mapping parameter to adapt the PSNR perceptually meaningful. Neurophysiologically motivated, a shearlet-based correction is proposed for controlling this perceptual PSNR adaption. The performance of the proposed perceptually adapted PSNR is evaluated on the LIVE and TID2013 databases and shows to be superior or comparable to benchmark IQMs. Sebastian Bosse, Mischa Siekmann, Wojciech Samek, Thomas Wiegand 0001 |
ICIP | 3 |
| 2017 | Detection of Face Morphing Attacks by Deep Learning
Clemens Seibold, Wojciech Samek, Anna Hilsmann, Peter Eisert |
IWDW | 2 |
| 2017 | Explaining nonlinear classification decisions with deep Taylor decompositionabstractNonlinear methods such as Deep Neural Networks (DNNs) are the gold standard for various challenging machine learning problems such as image recognition. Although these methods perform impressively well, they have a significant disadvantage, the lack of transparency, limiting the interpretability of the solution and thus the scope of application in practice. Especially DNNs act as black boxes due to their multilayer nonlinear structure. In this paper we introduce a novel methodology for interpreting generic multilayer neural networks by decomposing the network classification decision into contributions of its input elements. Although our focus is on image classification, the method is applicable to a broad set of input data, learning tasks and network architectures. Our method called deep Taylor decomposition efficiently utilizes the structure of the network by backpropagating the explanations from the output to the input layer. We evaluate the proposed method empirically on the MNIST and ILSVRC data sets. Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, Klaus-Robert Müller |
Pattern Recognit. | 4 |
| 2017 | Evaluating the Visualization of What a Deep Neural Network Has LearnedabstractDeep neural networks (DNNs) have demonstrated impressive performance in complex machine learning tasks such as image classification or speech recognition. However, due to their multilayer nonlinear structure, they are not transparent, i.e., it is hard to grasp what makes them arrive at a particular classification or recognition decision, given a new unseen data sample. Recently, several approaches have been proposed enabling one to understand and interpret the reasoning embodied in a DNN for a single test image. These methods quantify the “importance” of individual pixels with respect to the classification decision and allow a visualization in terms of a heatmap in pixel/input space. While the usefulness of heatmaps can be judged subjectively by a human, an objective quality measure is missing. In this paper, we present a general methodology based on region perturbation for evaluating ordered collections of pixels such as heatmaps. We compare heatmaps computed by three different methods on the SUN397, ILSVRC2012, and MIT Places data sets. Our main result is that the recently proposed layer-wise relevance propagation algorithm qualitatively and quantitatively provides a better explanation of what made a DNN arrive at a particular classification decision than the sensitivity-based approach or the deconvolution method. We provide theoretical arguments to explain this result and discuss its practical implications. Finally, we investigate the use of heatmaps for unsupervised assessment of the neural network performance. Wojciech Samek, Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, Klaus-Robert Müller |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Analyzing Classifiers: Fisher Vectors and Deep Neural NetworksabstractFisher vector (FV) classifiers and Deep Neural Networks (DNNs) are popular and successful algorithms for solving image classification problems. However, both are generally considered 'black box' predictors as the non-linear transformations involved have so far prevented transparent and interpretable reasoning. Recently, a principled technique, Layer-wise Relevance Propagation (LRP), has been developed in order to better comprehend the inherent structured reasoning of complex nonlinear classification models such as Bag of Feature models or DNNs. In this paper we (1) extend the LRP framework also for Fisher vector classifiers and then use it as analysis tool to (2) quantify the importance of context for classification, (3) qualitatively compare DNNs against FV classifiers in terms of important image regions and (4) detect potential flaws and biases in data. All experiments are performed on the PASCAL VOC 2007 and ILSVRC 2012 data sets. Sebastian Lapuschkin, Alexander Binder, Grégoire Montavon, Klaus-Robert Müller, Wojciech Samek |
CVPR | 5 |
| 2016 | Layer-Wise Relevance Propagation for Neural Networks with Local Renormalization Layers
Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, Klaus-Robert Müller, Wojciech Samek |
ICANN (2) | 5 |
| 2016 | Controlling explanatory heatmap resolution and semantics via decomposition depthabstractWe present an application of the Layer-wise Relevance Propagation (LRP) algorithm to state of the art deep convolutional neural networks and Fisher Vector classifiers to compare the image perception and prediction strategies of both classifiers with the use of visualized heatmaps. Layer-wise Relevance Propagation (LRP) is a method to compute scores for individual components of an input image, denoting their contribution to the prediction of the classifier for one particular test point. We demonstrate the impact of different choices of decomposition cut-off points during the LRP-process, controlling the resolution and semantics of the heatmap on test images from the PASCAL VOC 2007 test data set. Sebastian Lapuschkin, Alexander Binder, Klaus-Robert Müller, Wojciech Samek |
ICIP | 4 |
| 2016 | Shearlet-based reduced reference image quality assessmentabstractThis paper proposes a reduced reference image quality assessment method using only a low number of features. It involves a shearlet decomposition, directional pooling of the obtained coefficient and extracts the scalewise statistical location parameter as a feature. The proposed method is tested and compared to similar approaches on the LIVE image database. On this database it outperforms the compared methods on five of seven distortion types and on the full testset with a linear correlation of = 0.89. Sebastian Bosse, Qiaobo Chen, Mischa Siekmann, Wojciech Samek, Thomas Wiegand 0001 |
ICIP | 4 |
| 2016 | A deep neural network for image quality assessmentabstractThis paper presents a no reference image (NR) quality assessment (IQA) method based on a deep convolutional neural network (CNN). The CNN takes unpreprocessed image patches as an input and estimates the quality without employing any domain knowledge. By that, features and natural scene statistics are learnt purely data driven and combined with pooling and regression in one framework. We evaluate the network on the LIVE database and achieve a linear Pearson correlation superior to state-of-the-art NR IQA methods. We also apply the network to the image forensics task of decoder-sided quantization parameter estimation and also here achieve correlations of r = 0.989. Sebastian Bosse, Dominique Maniry, Thomas Wiegand 0001, Wojciech Samek |
ICIP | 4 |
| 2016 | Quality assessment of image patches distorted by image compression using crowdsourcingabstractThree experiments addressing the assessment of perceived image quality in a patch-based manner are compared for HEVC compression artifacts. It is shown that image patches of a size small as 128×128 pixel are large enough to evaluate the perceived image quality in a Degradation Category Rating (DCR) setting. Ratings obtained with 128×128 pixel sized images patches and 512×512 pixel sized images of the same spatial statistics show a correlation of r=0.99. Based on this finding, image quality assessment of 128×128 pixel sized image patches degraded by HEVC compression is compared for controlled lab environment and uncontrolled crowdsourcing settings. Although we find high overall correlation between the quality ratings obtained in the two environments, observers tend to give worse ratings in the crowdsourcing setting and for conditions of higher quality a reduction of correlation is observed. These findings have implications for choosing controlled vs. uncontrolled viewing conditions for image quality assessment for real-life applications. Sebastian Bosse, Mischa Siekmann, Jennifer Rasch, Thomas Wiegand 0001, Wojciech Samek |
ICME | 5 |
| 2016 | Hybrid video object tracking in H.265/HEVC video streamsabstractIn this paper we propose a hybrid tracking method which detects moving objects in videos compressed according to H.265/HEVC standard. Our framework largely depends on motion vectors (MV) and block types obtained by partially decoding the video bit stream and occasionally uses pixel domain information to distinguish between two objects. The compressed domain method is based on a Markov Random Field (MRF) model that captures spatial and temporal coherence of the moving object and is updated on a frame-to-frame basis. The hybrid nature of our approach stems from the usage of a pixel domain method that extracts the color information from the fully-decoded I frames and is updated only after completion of each Group-of-Pictures (GOP). We test the tracking accuracy of our method using standard video sequences and show that our hybrid framework provides better tracking accuracy than a state-of-the-art MRF model. Serhan Gul, Jan Timo Meyer, Cornelius Hellge, Thomas Schierl, Wojciech Samek |
MMSP | 5 |
| 2016 | Neural network-based full-reference image quality assessmentabstractS.334-338 Sebastian Bosse, Dominique Maniry, Klaus-Robert Müller, Thomas Wiegand 0001, Wojciech Samek |
PCS | 5 |
| 2016 | Brain-Computer Interfacing for multimedia quality assessmentabstractThe assessment of perceived multimedia quality is a central research field in information and media technology. Conventionally, psychophysical techniques are used for determining the quality of multimedia signals. Recently, Brain-Computer Interfacing (BCI)-based methods have been proposed for the assessment of perceived multimedia signal quality. In this paper we give an overview over the shortcomings of conventional approaches, present the state-of-the art of BCI-based methods and discuss open questions and challenges relevant to the BCI community. Sebastian Bosse, Klaus-Robert Müller, Thomas Wiegand 0001, Wojciech Samek |
SMC | 4 |
| 2016 | Alternative CSP approaches for multimodal distributed BCI dataabstractBrain-Computer Interfaces (BCIs) are trained to distinguish between two (or more) mental states, e.g., left and right hand motor imagery, from the recorded brain signals. Common Spatial Patterns (CSP) is a popular method to optimally separate data from two motor imagery tasks under the assumption of an unimodal class distribution. In out of lab environments where users are distracted by additional noise sources this assumption may not hold. This paper systematically investigates BCI performance under such distractions and proposes two novel CSP variants, ensemble CSP and 2-step CSP, which can cope with multimodal class distributions. The proposed algorithms are evaluated using simulations and BCI data of 16 healthy participants performing motor imagery under 6 different types of distraction. Both methods are shown to significantly enhance the performance compared to the standard procedure. Stephanie Brandl, Klaus-Robert Müller, Wojciech Samek |
SMC | 3 |
| 2016 | The LRP Toolbox for Artificial Neural NetworksabstractThe Layer-wise Relevance Propagation (LRP) algorithm explains a classifier's prediction specific to a given data point by attributing relevance scores to important components of the input by using the topology of the learned model itself. With the LRP Toolbox we provide platform-agnostic implementations for explaining the predictions of pre-trained state of the art Caffe networks and stand-alone implementations for fully connected Neural Network models. The implementations for Matlab and python shall serve as a playing field to familiarize oneself with the LRP algorithm and are implemented with readability and transparency in mind. Models and data can be imported and exported using raw text formats, Matlab's .mat files and the .npy format for numpy or plain text. Sebastian Lapuschkin, Alexander Binder, Grégoire Montavon, Klaus-Robert Müller, Wojciech Samek |
J. Mach. Learn. Res. | 5 |
| 2015 | Multivariate Machine Learning Methods for Fusing Multimodal Functional Neuroimaging DataabstractMultimodal data are ubiquitous in engineering, communications, robotics, computer vision, or more generally speaking in industry and the sciences. All disciplines have developed their respective sets of analytic tools to fuse the information that is available in all measured modalities. In this paper, we provide a review of classical as well as recent machine learning methods (specifically factor models) for fusing information from functional neuroimaging techniques such as: LFP, EEG, MEG, fNIRS, and fMRI. Early and late fusion scenarios are distinguished, and appropriate factor models for the respective scenarios are presented along with example applications from selected multimodal neuroimaging studies. Further emphasis is given to the interpretability of the resulting model parameters, in particular by highlighting how factor models relate to physical models needed for source localization. The methods we discuss allow for the extraction of information from neural data, which ultimately contributes to 1) better neuroscientific understanding; 2) enhance diagnostic performance; and 3) discover neural signals of interest that correlate maximally with a given cognitive paradigm. While we clearly study the multimodal functional neuroimaging challenge, the discussed machine learning techniques have a wide applicability, i.e., in general data fusion, and may thus be informative to the general interested reader. Sven Dähne, Felix Bießmann, Wojciech Samek, Stefan Haufe, Dominique Goltz, Christopher Gundlach, Arno Villringer, Siamac Fazli, Klaus-Robert Müller |
Proc. IEEE | 3 |
| 2015 | Learning From More Than One Data Source: Data Fusion Techniques for Sensorimotor Rhythm-Based Brain-Computer InterfacesabstractBrain-computer interfaces (BCIs) are successfully used in scientific, therapeutic and other applications. Remaining challenges are among others a low signal-to-noise ratio of neural signals, lack of robustness for decoders in the presence of inter-trial and inter-subject variability, time constraints on the calibration phase and the use of BCIs outside a controlled lab environment. Recent advances in BCI research addressed these issues by novel combinations of complementary analysis as well as recording techniques, so called hybrid BCIs. In this paper, we review a number of data fusion techniques for BCI along with hybrid methods for BCI that have recently emerged. Our focus will be on sensorimotor rhythm-based BCIs. We will give an overview of the three main lines of research in this area, integration of complementary features of neural activation, integration of multiple previous sessions and of multiple subjects, and show how these techniques can be used to enhance modern BCI systems. Siamac Fazli, Sven Dähne, Wojciech Samek, Felix Bießmann, Klaus-Robert Müller |
Proc. IEEE | 3 |
| 2014 | Robust common spatial patterns by minimum divergence covariance estimatorabstractReliable estimation of covariance matrices from high-dimensional electroencephalographic recordings is crucial for a successful application of Brain-Computer Interface (BCI) systems. Artifactual trials and non-stationarity effects may have a large impact on the estimation quality and adversely affect the spatial filter computation and consequently the classification accuracy of the system. In this work we propose a novel robust estimator for covariance matrices that takes into account the trial structure of BCI experiments. Our estimator minimizes beta divergence between the empirical and a model Wishart distribution, thus allows to robustly average the estimated covariance matrices of different trials and downweight the influence of outlier trials. We evaluate this novel estimator on a data set with recordings from 80 subjects. Wojciech Samek, Motoaki Kawanabe |
ICASSP | 1 |
| 2014 | Robust Common Spatial Filters with a Maxmin ApproachabstractElectroencephalographic signals are known to be nonstationary and easily affected by artifacts; therefore, their analysis requires methods that can deal with noise. In this work, we present a way to robustify the popular common spatial patterns (CSP) algorithm under a maxmin approach. In contrast to standard CSP that maximizes the variance ratio between two conditions based on a single estimate of the class covariance matrices, we propose to robustly compute spatial filters by maximizing the minimum variance ratio within a prefixed set of covariance matrices called the tolerance set. We show that this kind of maxmin optimization makes CSP robust to outliers and reduces its tendency to overfit. We also present a data-driven approach to construct a tolerance set that captures the variability of the covariance matrices over time and shows its ability to reduce the nonstationarity of the extracted features and significantly improve classification accuracy. We test the spatial filters derived with this approach and compare them to standard CSP and a state-of-the-art method on a real-world brain-computer interface (BCI) data set in which we expect substantial fluctuations caused by environmental differences. Finally we investigate the advantages and limitations of the maxmin approach with simulations. Motoaki Kawanabe, Wojciech Samek, Klaus-Robert Müller, Carmen Vidaurre |
Neural Comput. | 2 |
| 2013 | Robust Spatial Filtering with Beta DivergenceabstractThe efficiency of Brain-Computer Interfaces (BCI) largely depends upon a reliable extraction of informative features from the high-dimensional EEG signal. A crucial step in this protocol is the computation of spatial filters. The Common Spatial Patterns (CSP) algorithm computes filters that maximize the difference in band power between two conditions, thus it is tailored to extract the relevant information in motor imagery experiments. However, CSP is highly sensitive to artifacts in the EEG data, i.e. few outliers may alter the estimate drastically and decrease classification performance. Inspired by concepts from the field of information geometry we propose a novel approach for robustifying CSP. More precisely, we formulate CSP as a divergence maximization problem and utilize the property of a particular type of divergence, namely beta divergence, for robustifying the estimation of spatial filters in the presence of artifacts in the data. We demonstrate the usefulness of our method on toy data and on EEG recordings from 80 subjects. Wojciech Samek, Duncan A. J. Blythe, Klaus-Robert Müller, Motoaki Kawanabe |
NIPS | 1 |
| 2013 | Enhanced representation and multi-task learning for image annotation
Alexander Binder, Wojciech Samek, Klaus-Robert Müller, Motoaki Kawanabe |
Comput. Vis. Image Underst. | 2 |
| 2011 | Multi-task Learning via Non-sparse Multiple Kernel Learning
Wojciech Samek, Alexander Binder, Motoaki Kawanabe |
CAIP (1) | 1 |
| 2011 | An Information Geometrical View of Stationary Subspace Analysis
Motoaki Kawanabe, Wojciech Samek, Paul von Bünau, Frank C. Meinecke |
ICANN (2) | 2 |
| 2011 | Stationary Common Spatial Patterns: Towards robust classification of non-stationary EEG signalsabstractBrain-Computer Interfaces (BCIs) allow a user to control a computer application by brain activity as acquired, e.g., by EEG. A standard step in a BCI system is to project the EEG signals to a low-dimensional subspace using Common Spatial Patterns (CSP). However, non-stationarities in the data can negatively affect the performance of CSP, i.e. variation of the signal properties within and across experimental sessions coming from electrode artefacts, alpha or muscular activity, or fatigue may result in suboptimal projection directions. We alleviate this problem by regularizing CSP towards stationary subspaces and show that this especially increases classification accuracy of people who are not able to control a BCI i.e. have more than 30% of error. These users very often show non-stationarities in their EEG signals. Wojciech Samek, Carmen Vidaurre, Motoaki Kawanabe |
ICASSP | 1 |
| 2011 | Multi-modal visual concept classification of images via Markov random walk over tagsabstractAutomatic annotation of images is a challenging task in computer vision because of “semantic gap” between highlevel visual concepts and image appearances. Therefore, user tags attached to images can provide further information to bridge the gap, even though they are partially uninformative and misleading. In this work, we investigate multi-modal visual concept classification based on visual features and user tags via kernel-based classifiers. An issue here is how to construct kernels between sets of tags. We deploy Markov random walks on graphs of key tags to incorporate co-occurrence between them. This procedure acts as a smoothing of tag based features. Our experimental result on the ImageCLEF2010 PhotoAnnotation benchmark shows that our proposed method outperforms the baseline relying solely on visual information and a recently published state-of-the-art approach. Motoaki Kawanabe, Alexander Binder, Christina Müller, Wojciech Samek |
WACV | 4 |
| 2010 | A Hybrid Supervised-Unsupervised Vocabulary Generation Algorithm for Visual Concept Recognition
Alexander Binder, Wojciech Samek, Christina Müller, Motoaki Kawanabe |
ACCV (3) | 2 |
| 2010 | Shrinking large visual vocabularies using multi-label agglomerative information bottleneckabstractThe quality of visual vocabularies is crucial for the performance of bag-of-words image classification methods. Several approaches have been developed for codebook construction, the most popular method is to cluster a set of image features (e.g. SIFT) by k-means. In this paper, we propose a two-step procedure which incorporates label information into the clustering process by efficiently generating a large and informative vocabulary using class-wise k-means and reducing its size by agglomerative information bottleneck (AIB). We introduce an extension of the AIB procedure for multi-label problems and show that this two-step approach improves the classification results while reducing computation time compared to the vanilla k-means. We analyse the reasons for the performance gain on the PASCAL VOC 2007 data set. Wojciech Samek, Alexander Binder, Motoaki Kawanabe |
ICIP | 1 |
| 2010 | Enhancing Image Classification with Class-wise Clustered VocabulariesabstractIn recent years bag-of-visual-words representations have gained increasing popularity in the field of image classification. Their performance highly relies on creating a good visual vocabulary from a set of image features (e.g. SIFT). For real-world photo archives such as Flicker, codebooks with larger than a few thousand words are desirable, which is infeasible by the standard k-means clustering. In this paper, we propose a two-step procedure which can generate more informative codebooks efficiently by class-wise k-means and a novel procedure for word selection. Our approach was compared favorably to the standard k-means procedure on the PASCAL VOC data sets. Wojciech Samek, Alexander Binder, Motoaki Kawanabe |
ICPR | 1 |