Bastian Pfeifer

dblp:157/9442 · DBLP profile ↗
← Back
17ranked-venue papers
11as first author
14since 2021 · last 2025
0000-0001-7035-9535ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 12 · 9 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Tree smoothing: Post-hoc regularization of tree ensembles for interpretable machine learning
abstract
Random Forests (RFs) are powerful ensemble learning algorithms that are widely used in various machine learning tasks. However, they tend to overfit noisy or irrelevant features, which can result in decreased generalization performance. Post-hoc regularization techniques aim to solve this problem by modifying the structure of the learned ensemble after training. We propose a novel post-hoc regularization via tree smoothing for classification tasks to leverage the reliable class distributions closer to the root node whilst reducing the impact of more specific and potentially noisy splits deeper in the tree. Our novel approach allows for a form of pruning that does not alter the general structure of the trees, adjusting the influence of nodes based on their proximity to the root node. We evaluated the performance of our method on various machine learning benchmark data sets and on cancer data from The Cancer Genome Atlas (TCGA). Our approach demonstrates competitive performance compared to the state-of-the-art and, in the majority of cases, and outperforms it in most cases in terms of prediction accuracy, generalization, and interpretability. • A novel post-regulation technique for Tree Ensembles called BBTS is introduced. • Interpretability is improved through posterior distributions in the leaf nodes. • This method allows the incorporation of domain knowledge through prior beliefs. • It was tested using ML benchmarks and real-world cancer data from the TCGA database. • BBTS has proven itself with the state-of-the-art and exceeds them on real-world cancer data.
Bastian Pfeifer, Arne Gevaert, Markus Loecher, Andreas Holzinger
Inf. Sci.1
2025 Explaining and visualizing black-box models through counterfactual paths
abstract
Abstract Explainable AI (XAI) is an increasingly important area of machine learning research, which aims to make black-box models transparent and interpretable. In this paper, we propose a novel approach to XAI that uses the so-called counterfactual paths for model-agnostic global explanations. The algorithm measures feature importance by identifying sequential permutations of features that most influence changes in model predictions. It is particularly suitable for generating explanations based on counterfactual paths in knowledge graphs incorporating domain knowledge. Counterfactual paths introduce an additional graph dimension to current XAI methods in both explaining and visualizing black-box models. Experiments with synthetic and bio-medical data demonstrate the practical applicability of our approach.
Bastian Pfeifer, Mateusz Krzyzinski, Hubert Baniecki, Andreas Holzinger, Przemyslaw Biecek
Pattern Anal. Appl.1
2024 Detection and quantification of introgression using Bayesian inference based on conjugate priors
abstract
SUMMARY: Introgression (the flow of genes between species) is a major force structuring the evolution of genomes, potentially providing raw material for adaptation. Here, we present a versatile Bayesian model selection approach for detecting and quantifying introgression, df-BF, that builds upon the recently published distance-based df statistic. Unlike df, df-BF accounts for the number of variant sites within a genomic region. The underlying model parameter of our df-BF method, here denoted as dfθ, accurately quantifies introgression, and the corresponding Bayes Factors (df-BF) enables weighing the strength of evidence for introgression. To ensure fast computation, we use conjugate priors with no need for computationally demanding MCMC iterations. We compare our method with other approaches including df, fd, Dp, and Patterson's D using a wide range of coalescent simulations. Furthermore, we showcase the applicability of df-BF and dfθ using whole-genome mosquito data. Finally, we integrate the new method into the powerful genomics R-package PopGenome. AVAILABILITY AND IMPLEMENTATION: The presented methods are implemented within the R-package PopGenome (https://github.com/pievos101/PopGenome) and the simulation as the application results can be reproduced from the source code available from a dedicated GitHub repository (https://github.com/pievos101/Introgression-Simulation).
Bastian Pfeifer, Durrell D. Kapan, Sereina A. Herzog
Bioinform.1
2024 Federated unsupervised random forest for privacy-preserving patient stratification
abstract
MOTIVATION: In the realm of precision medicine, effective patient stratification and disease subtyping demand innovative methodologies tailored for multi-omics data. Clustering techniques applied to multi-omics data have become instrumental in identifying distinct subgroups of patients, enabling a finer-grained understanding of disease variability. Meanwhile, clinical datasets are often small and must be aggregated from multiple hospitals. Online data sharing, however, is seen as a significant challenge due to privacy concerns, potentially impeding big data's role in medical advancements using machine learning. This work establishes a powerful framework for advancing precision medicine through unsupervised random forest-based clustering in combination with federated computing. RESULTS: We introduce a novel multi-omics clustering approach utilizing unsupervised random forests. The unsupervised nature of the random forest enables the determination of cluster-specific feature importance, unraveling key molecular contributors to distinct patient groups. Our methodology is designed for federated execution, a crucial aspect in the medical domain where privacy concerns are paramount. We have validated our approach on machine learning benchmark datasets as well as on cancer data from The Cancer Genome Atlas. Our method is competitive with the state-of-the-art in terms of disease subtyping, but at the same time substantially improves the cluster interpretability. Experiments indicate that local clustering performance can be improved through federated computing. AVAILABILITY AND IMPLEMENTATION: The proposed methods are available as an R-package (https://github.com/pievos101/uRF).
Bastian Pfeifer, Christel Sirocchi, Marcus D. Bloice, Markus Kreuzthaler, Martin Urschler
Bioinform.1
2024 Effective signal reconstruction from multiple ranked lists via convex optimization
abstract
Abstract The ranking of objects is widely used to rate their relative quality or relevance across multiple assessments. Beyond classical rank aggregation, it is of interest to estimate the usually unobservable latent signals that inform a consensus ranking. Under the only assumption of independent assessments, which can be incomplete, we introduce indirect inference via convex optimization in combination with computationally efficient Poisson Bootstrap. Two different objective functions are suggested, one linear and the other quadratic. The mathematical formulation of the signal estimation problem is based on pairwise comparisons of all objects with respect to their rank positions. Sets of constraints represent the order relations. The transitivity property of rank scales allows us to reduce substantially the number of constraints associated with the full set of object comparisons. The key idea is to globally reduce the errors induced by the rankers until optimal latent signals can be obtained. Its main advantage is low computational costs, even when handling $$n < < p$$ n < < p data problems. Exploratory tools can be developed based on the bootstrap signal estimates and standard errors. Simulation evidence, a comparison with the state-of-the-art rank centrality method, and two applications, one in higher education evaluation and the other in molecular cancer research, are presented.
Michael G. Schimek, Luca Vitale, Bastian Pfeifer, Michele La Rocca 0001
Data Min. Knowl. Discov.3
2024 Correction to: Effective signal reconstruction from multiple ranked lists via convex optimization
Michael G. Schimek, Luca Vitale, Bastian Pfeifer, Michele La Rocca 0001
Data Min. Knowl. Discov.3
2024 CLARUS: An interactive explainable AI platform for manual counterfactuals in graph neural networks
abstract
BACKGROUND: Lack of trust in artificial intelligence (AI) models in medicine is still the key blockage for the use of AI in clinical decision support systems (CDSS). Although AI models are already performing excellently in systems medicine, their black-box nature entails that patient-specific decisions are incomprehensible for the physician. Explainable AI (XAI) algorithms aim to "explain" to a human domain expert, which input features influenced a specific recommendation. However, in the clinical domain, these explanations must lead to some degree of causal understanding by a clinician. RESULTS: We developed the CLARUS platform, aiming to promote human understanding of graph neural network (GNN) predictions. CLARUS enables the visualisation of patient-specific networks, as well as, relevance values for genes and interactions, computed by XAI methods, such as GNNExplainer. This enables domain experts to gain deeper insights into the network and more importantly, the expert can interactively alter the patient-specific network based on the acquired understanding and initiate re-prediction or retraining. This interactivity allows us to ask manual counterfactual questions and analyse the effects on the GNN prediction. CONCLUSION: We present the first interactive XAI platform prototype, CLARUS, that allows not only the evaluation of specific human counterfactual questions based on user-defined alterations of patient networks and a re-prediction of the clinical outcome but also a retraining of the entire GNN after changing the underlying graph structures. The platform is currently hosted by the GWDG on https://rshiny.gwdg.de/apps/clarus/.
Jacqueline Michelle Metsch, Anna Saranti, Alessa Angerschmid, Bastian Pfeifer, Vanessa Klemt, Andreas Holzinger, Anne-Christin Hauschild
J. Biomed. Informatics4
2023 Human-in-the-Loop Integration with Domain-Knowledge Graphs for Explainable Federated Deep Learning
abstract
Abstract We explore the integration of domain knowledge graphs into Deep Learning for improved interpretability and explainability using Graph Neural Networks (GNNs). Specifically, a protein-protein interaction (PPI) network is masked over a deep neural network for classification, with patient-specific multi-modal genomic features enriched into the PPI graph’s nodes. Subnetworks that are relevant to the classification (referred to as “disease subnetworks”) are detected using explainable AI. Federated learning is enabled by dividing the knowledge graph into relevant subnetworks, constructing an ensemble classifier, and allowing domain experts to analyze and manipulate detected subnetworks using a developed user interface. Furthermore, the human-in-the-loop principle can be applied with the incorporation of experts, interacting through a sophisticated User Interface (UI) driven by Explainable Artificial Intelligence (xAI) methods, changing the datasets to create counterfactual explanations. The adapted datasets could influence the local model’s characteristics and thereby create a federated version that distils their diverse knowledge in a centralized scenario. This work demonstrates the feasibility of the presented strategies, which were originally envisaged in 2021 and most of it has now been materialized into actionable items. In this paper, we report on some lessons learned during this project.
Andreas Holzinger, Anna Saranti, Anne-Christin Hauschild, Jacqueline Michelle Metsch, Dominik Heider, Richard Röttger, Heimo Müller, Jan Baumbach, Bastian Pfeifer
CD-MAKE9
2023 Ensemble-GNN: federated ensemble learning with graph neural networks for disease module discovery and classification
abstract
SUMMARY: Federated learning enables collaboration in medicine, where data is scattered across multiple centers without the need to aggregate the data in a central cloud. While, in general, machine learning models can be applied to a wide range of data types, graph neural networks (GNNs) are particularly developed for graphs, which are very common in the biomedical domain. For instance, a patient can be represented by a protein-protein interaction (PPI) network where the nodes contain the patient-specific omics features. Here, we present our Ensemble-GNN software package, which can be used to deploy federated, ensemble-based GNNs in Python. Ensemble-GNN allows to quickly build predictive models utilizing PPI networks consisting of various node features such as gene expression and/or DNA methylation. We exemplary show the results from a public dataset of 981 patients and 8469 genes from the Cancer Genome Atlas (TCGA). AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/pievos101/Ensemble-GNN, and the data at Zenodo (DOI: 10.5281/zenodo.8305122).
Bastian Pfeifer, Hryhorii Chereda, Roman Martin, Anna Saranti, Sandra Clemens, Anne-Christin Hauschild, Tim Beißbarth, Andreas Holzinger, Dominik Heider
Bioinform.1
2023 Embedding-based terminology expansion via secondary use of large clinical real-world datasets
abstract
A log-likelihood based co-occurrence analysis of ∼1.9 million de-identified ICD-10 codes and related short textual problem list entries generated possible term candidates at a significance level of p<0.01. These top 10 term candidates, consisting of 1 to 5-grams, were used as seed terms for an embedding based nearest neighbor approach to fetch additional synonyms, hypernyms and hyponyms in the respective n-gram embedding spaces by leveraging two different language models. This was done to analyze the lexicality of the resulting term candidates and to compare the term classifications of both models. We found no difference in system performance during the processing of lexical and non-lexical content, i.e. abbreviations, acronyms, etc. Additionally, an application-oriented analysis of the SapBERT (Self-Alignment Pretraining for Biomedical Entity Representations) language model indicates suitable performance for the extraction of all term classifications such as synonyms, hypernyms, and hyponyms.
Amila Kugic, Bastian Pfeifer, Stefan Schulz 0001, Markus Kreuzthaler
J. Biomed. Informatics2
2023 Parea: Multi-view ensemble clustering for cancer subtype discovery
abstract
Multi-view clustering methods are essential for the stratification of patients into sub-groups of similar molecular characteristics. In recent years, a wide range of methods have been developed for this purpose. However, due to the high diversity of cancer-related data, a single method may not perform sufficiently well in all cases. We present Parea, a multi-view hierarchical ensemble clustering approach for disease subtype discovery. We demonstrate its performance on several machine learning benchmark datasets. We apply and validate our methodology on real-world multi-view patient data, comprising seven types of cancer. Parea outperforms the current state-of-the-art on six out of seven analysed cancer types. We have integrated the Parea method into our Python package Pyrea (https://github.com/mdbloice/Pyrea), which enables the effortless and flexible design of ensemble workflows while incorporating a wide range of fusion and clustering algorithms.
Bastian Pfeifer, Marcus D. Bloice, Michael G. Schimek
J. Biomed. Informatics1
2022 GNN-SubNet: disease subnetwork detection with explainable graph neural networks
abstract
MOTIVATION: The tremendous success of graphical neural networks (GNNs) already had a major impact on systems biology research. For example, GNNs are currently being used for drug target recognition in protein-drug interaction networks, as well as for cancer gene discovery and more. Important aspects whose practical relevance is often underestimated are comprehensibility, interpretability and explainability. RESULTS: In this work, we present a novel graph-based deep learning framework for disease subnetwork detection via explainable GNNs. Each patient is represented by the topology of a protein-protein interaction (PPI) network, and the nodes are enriched with multi-omics features from gene expression and DNA methylation. In addition, we propose a modification of the GNNexplainer that provides model-wide explanations for improved disease subnetwork detection. AVAILABILITY AND IMPLEMENTATION: The proposed methods and tools are implemented in the GNN-SubNet Python package, which we have made available on our GitHub for the international research community (https://github.com/pievos101/GNN-SubNet). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Bastian Pfeifer, Anna Saranti, Andreas Holzinger
Bioinform.1
2021 Integrative hierarchical ensemble clustering for improved disease subtype discovery
abstract
Multi-omics clustering methods are used for the stratification of patients into sub-groups of similar molecular characteristics. In recent years, a wide range of methods has been developed for this purpose. However, due to the high diversity of cancer-related data, a single method may not perform sufficiently well in all cases. Here, we propose a comprehensive framework for multi-omics hierarchical ensemble clustering. We provide a flexible environment that allows to build hierarchical clustering ensembles suitable for the available data and research goals. Survival analyses for data from The Cancer Genome Atlas (TCGA) indicate that our proposed ensembles provide more robust, and thus more reliable results than the state-of-the-art. We have implemented our architecture within the R-package HC-fused, which is freely available on Github.
Bastian Pfeifer, Andrei Voicu-Spineanu, Michael G. Schimek, Nikolaos Alachiotis 0001
BIBM1
2021 A hierarchical clustering and data fusion approach for disease subtype discovery
abstract
Recent advances in multi-omics clustering methods enable a more fine-tuned separation of cancer patients into clinical relevant clusters. These advancements have the potential to provide a deeper understanding of cancer progression and may facilitate the treatment of cancer patients. Here, we present a simple hierarchical clustering and data fusion approach, named HC-fused, for the detection of disease subtypes. Unlike other methods, the proposed approach naturally reports on the individual contribution of each single-omic to the data fusion process. We perform multi-view simulations with disjoint and disjunct cluster elements across the views to highlight fundamentally different data integration behavior of various state-of-the-art methods. HC-fused combines the strengths of some recently published methods and shows superior performance on real world cancer data from the TCGA (The Cancer Genome Atlas) database. An R implementation of our method is available on GitHub (pievos101/HC-fused).
Bastian Pfeifer, Michael G. Schimek
J. Biomed. Informatics1
2019 Estimates of introgression as a function of pairwise distances
abstract
Research over the last 10 years highlights the increasing importance of hybridization between species as a major force structuring the evolution of genomes and potentially providing raw material for adaptation by natural and/or sexual selection. Fueled by research in a few model systems where phenotypic hybrids are easily identified, research into hybridization and introgression (the flow of genes between species) has exploded with the advent of whole-genome sequencing and emerging methods to detect the signature of hybridization at the whole-genome or chromosome level. Amongst these are a general class of methods that utilize patterns of single-nucleotide polymorphisms (SNPs) across a tree as markers of hybridization. These methods have been applied to a variety of genomic systems ranging from butterflies to Neanderthals to detect introgression, however, when employed at a fine genomic scale these methods do not perform well to quantify introgression in small sample windows. We introduce a novel method to detect introgression by combining two widely used statistics: pairwise nucleotide diversity d xy and Patterson’s D . The resulting statistic, the distance fraction ( d f ), accounts for genetic distance across possible topologies and is designed to simultaneously detect and quantify introgression. We also relate our new method to the recently published f d and incorporate these statistics into the powerful genomics R-package PopGenome, freely available on GitHub ( pievos101/PopGenome ) and the Comprehensive R Archive Network (CRAN). The supplemental material contains a wide range of simulation studies and a detailed manual how to perform the statistics within the PopGenome framework. We present a new distance based statistic d f that avoids the pitfalls of Patterson’s D when applied to small genomic regions and accurately quantifies the fraction of introgression ( f ) for a wide range of simulation scenarios.
Bastian Pfeifer, Durrell D. Kapan
BMC Bioinform.1
2018 BlockFeST: Bayesian calculation of region-specific FST to detect local adaptation
abstract
Summary: The fixation index FST can be used to identify non-neutrally evolving loci from genome-scale SNP data across two or more populations. Recent years have seen the development of sophisticated approaches to estimate FST based on Markov-Chain Monte-Carlo simulations. Here, we present a vectorized R implementation of an extension of the widely used BayeScan software for codominant markers, adding the option to group individual SNPs into pre-defined blocks. A typical application of this new approach is the identification of genomic regions, genes, or gene sets containing SNPs that evolved under directional selection. Availability and implementation: The R implementation of our method, which builds on the powerful population genetics and genomics software PopGenome, is available freely from CRAN. Supplementary information: Supplementary data are available at Bioinformatics online.
Bastian Pfeifer, Martin J. Lercher
Bioinform.1
2015 WhopGenome: high-speed access to whole-genome variation and sequence data in R
abstract
SUMMARY: The statistical programming language R has become a de facto standard for the analysis of many types of biological data, and is well suited for the rapid development of new algorithms. However, variant call data from population-scale resequencing projects are typically too large to be read and processed efficiently with R's built-in I/O capabilities. WhopGenome can efficiently read whole-genome variation data stored in the widely used variant call format (VCF) file format into several R data types. VCF files can be accessed either on local hard drives or on remote servers. WhopGenome can associate variants with annotations such as those available from the UCSC genome browser, and can accelerate the reading process by filtering loci according to user-defined criteria. WhopGenome can also read other Tabix-indexed files and create indices to allow fast selective access to FASTA-formatted sequence files. AVAILABILITY AND IMPLEMENTATION: The WhopGenome R package is available on CRAN at http://cran.r-project.org/web/packages/WhopGenome/. A Bioconductor package has been submitted. CONTACT: [email protected].
Ulrich Wittelsbürger, Bastian Pfeifer, Martin J. Lercher
Bioinform.2