EDBT 2026 Demo / reviewers in the wild / expert
Enrico Glaab
dblp:77/8090
· DBLP profile ↗
15ranked-venue papers
8as first author
7since 2021 · last 2025
0000-0003-3977-7469ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 8 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Causal network analysis of omics data using prior knowledge databasesabstractIdentifying causal relationships in omics data is essential for understanding underlying biological processes. However, detecting these relationships remains challenging due to the complexity of molecular networks and observational data limitations. To guide researchers, we conducted a systematic literature review of data-driven causal omics analysis methods that use structured prior knowledge from regulatory and interaction databases. We grouped methods into three approaches based on the extent of prior knowledge integration: regulon-level (direct regulator-target links, straightforward interpretation, but with the risk of oversimplification), flow-level (multi-step propagation from regulators to targets, broader mechanism explanation, but lacking uncertainty modeling), and network-level (system-wide interactions and crosstalk, most comprehensive, but with increased computational complexity and requiring particularly careful interpretation). These methods have demonstrated utility across diverse applications, including identification of therapeutic targets in acute myeloid leukemia, elucidation of mechanisms in IgA nephropathy, and detection of regulatory perturbations in Alzheimer's disease. We discuss the strengths, limitations, and representative use cases of each approach, and address general limitations and outline future research directions. This review serves as a practical guide for the entire analysis process, from selecting prior knowledge databases (PKDBs) to choosing and applying causal analysis methods for different research questions. Gleb Svinin, Enrico Glaab |
Briefings Bioinform. | 2 |
| 2025 | Estimating sparse regression models in multi-task learning and transfer learning through adaptive penalisationabstractMETHOD: Here, we propose a simple two-stage procedure for sharing information between related high-dimensional prediction or classification problems. In both stages, we perform sparse regression separately for each problem. While this is done without prior information in the first stage, we use the coefficients from the first stage as prior information for the second stage. Specifically, we designed feature-specific and sign-specific adaptive weights to share information on feature selection, effect directions, and effect sizes between different problems. RESULTS: The proposed approach is applicable to multi-task learning as well as transfer learning. It provides sparse models (i.e. with few non-zero coefficients for each problem) that are easy to interpret. We show by simulation and application that it tends to select fewer features while achieving a similar predictive performance as compared to available methods. AVAILABILITY AND IMPLEMENTATION: An implementation is available in the R package "sparselink" (https://github.com/rauschenberger/sparselink, https://cran.r-project.org/package=sparselink). Armin Rauschenberger, Petr V. Nazarov, Enrico Glaab |
Bioinform. | 3 |
| 2024 | Bioinformatics approaches for studying molecular sex differences in complex diseasesabstractMany complex diseases exhibit pronounced sex differences that can affect both the initial risk of developing the disease, as well as clinical disease symptoms, molecular manifestations, disease progression, and the risk of developing comorbidities. Despite this, computational studies of molecular data for complex diseases often treat sex as a confounding variable, aiming to filter out sex-specific effects rather than attempting to interpret them. A more systematic, in-depth exploration of sex-specific disease mechanisms could significantly improve our understanding of pathological and protective processes with sex-dependent profiles. This survey discusses dedicated bioinformatics approaches for the study of molecular sex differences in complex diseases. It highlights that, beyond classical statistical methods, approaches are needed that integrate prior knowledge of relevant hormone signaling interactions, gene regulatory networks, and sex linkage of genes to provide a mechanistic interpretation of sex-dependent alterations in disease. The review examines and compares the advantages, pitfalls and limitations of various conventional statistical and systems-level mechanistic analyses for this purpose, including tailored pathway and network analysis techniques. Overall, this survey highlights the potential of specialized bioinformatics techniques to systematically investigate molecular sex differences in complex diseases, to inform biomarker signature modeling, and to guide more personalized treatment approaches. Rebecca Ting Jiin Loo, Mohamed Soudy, Francesco Nasta, Mirco Macchi, Enrico Glaab |
Briefings Bioinform. | 5 |
| 2023 | Penalized regression with multiple sources of prior effectsabstractMOTIVATION: In many high-dimensional prediction or classification tasks, complementary data on the features are available, e.g. prior biological knowledge on (epi)genetic markers. Here we consider tasks with numerical prior information that provide an insight into the importance (weight) and the direction (sign) of the feature effects, e.g. regression coefficients from previous studies. RESULTS: We propose an approach for integrating multiple sources of such prior information into penalized regression. If suitable co-data are available, this improves the predictive performance, as shown by simulation and application. AVAILABILITY AND IMPLEMENTATION: The proposed method is implemented in the R package transreg (https://github.com/lcsb-bds/transreg, https://cran.r-project.org/package=transreg). Armin Rauschenberger, Zied Landoulsi, Mark A. van de Wiel, Enrico Glaab |
Bioinform. | 4 |
| 2022 | Ten quick tips for biomarker discovery and validation analyses using machine learningabstractIntroductionAU : Pleaseconfirmthatallheadinglevelsarerepresentedcorrectly Ramón Díaz-Uriarte, Elisa Gómez de Lope, Rosalba Giugno, Holger Fröhlich, Petr V. Nazarov, Isabel A. Nepomuceno-Chamorro, Armin Rauschenberger, Enrico Glaab |
PLoS Comput. Biol. | 8 |
| 2021 | Predictive and interpretable models via the stacked elastic netabstractMOTIVATION: Machine learning in the biomedical sciences should ideally provide predictive and interpretable models. When predicting outcomes from clinical or molecular features, applied researchers often want to know which features have effects, whether these effects are positive or negative and how strong these effects are. Regression analysis includes this information in the coefficients but typically renders less predictive models than more advanced machine learning techniques. RESULTS: Here, we propose an interpretable meta-learning approach for high-dimensional regression. The elastic net provides a compromise between estimating weak effects for many features and strong effects for some features. It has a mixing parameter to weight between ridge and lasso regularization. Instead of selecting one weighting by tuning, we combine multiple weightings by stacking. We do this in a way that increases predictivity without sacrificing interpretability. AVAILABILITY AND IMPLEMENTATION: The R package starnet is available on GitHub (https://github.com/rauschenberger/starnet) and CRAN (https://CRAN.R-project.org/package=starnet). Armin Rauschenberger, Enrico Glaab, Mark A. van de Wiel |
Bioinform. | 2 |
| 2021 | Predicting correlated outcomes from molecular dataabstractMOTIVATION: Multivariate (multi-target) regression has the potential to outperform univariate (single-target) regression at predicting correlated outcomes, which frequently occur in biomedical and clinical research. Here we implement multivariate lasso and ridge regression using stacked generalization. RESULTS: Our flexible approach leads to predictive and interpretable models in high-dimensional settings, with a single estimate for each input-output effect. In the simulation, we compare the predictive performance of several state-of-the-art methods for multivariate regression. In the application, we use clinical and genomic data to predict multiple motor and non-motor symptoms in Parkinson's disease patients. We conclude that stacked multivariate regression, with our adaptations, is a competitive method for predicting correlated outcomes. AVAILABILITY AND IMPLEMENTATION: The R package joinet is available on GitHub (https://github.com/rauschenberger/joinet) and cran (https://cran.r-project.org/package=joinet). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Armin Rauschenberger, Enrico Glaab |
Bioinform. | 2 |
| 2016 | Building a virtual ligand screening pipeline using free software: a surveyabstractVirtual screening, the search for bioactive compounds via computational methods, provides a wide range of opportunities to speed up drug development and reduce the associated risks and costs. While virtual screening is already a standard practice in pharmaceutical companies, its applications in preclinical academic research still remain under-exploited, in spite of an increasing availability of dedicated free databases and software tools. In this survey, an overview of recent developments in this field is presented, focusing on free software and data repositories for screening as alternatives to their commercial counterparts, and outlining how available resources can be interlinked into a comprehensive virtual screening pipeline using typical academic computing facilities. Finally, to facilitate the set-up of corresponding pipelines, a downloadable software system is provided, using platform virtualization to integrate pre-installed screening tools and scripts for reproducible application across different operating systems. Enrico Glaab |
Briefings Bioinform. | 1 |
| 2016 | Using prior knowledge from cellular pathways and molecular networks for diagnostic specimen classificationabstractFor many complex diseases, an earlier and more reliable diagnosis is considered a key prerequisite for developing more effective therapies to prevent or delay disease progression. Classical statistical learning approaches for specimen classification using omics data, however, often cannot provide diagnostic models with sufficient accuracy and robustness for heterogeneous diseases like cancers or neurodegenerative disorders. In recent years, new approaches for building multivariate biomarker models on omics data have been proposed, which exploit prior biological knowledge from molecular networks and cellular pathways to address these limitations. This survey provides an overview of these recent developments and compares pathway- and network-based specimen classification approaches in terms of their utility for improving model robustness, accuracy and biological interpretability. Different routes to translate omics-based multifactorial biomarker models into clinical diagnostic tests are discussed, and a previous study is presented as example. Enrico Glaab |
Briefings Bioinform. | 1 |
| 2015 | RepExplore: addressing technical replicate variance in proteomics and metabolomics data analysisabstractUNLABELLED: High-throughput omics datasets often contain technical replicates included to account for technical sources of noise in the measurement process. Although summarizing these replicate measurements by using robust averages may help to reduce the influence of noise on downstream data analysis, the information on the variance across the replicate measurements is lost in the averaging process and therefore typically disregarded in subsequent statistical analyses.We introduce RepExplore, a web-service dedicated to exploit the information captured in the technical replicate variance to provide more reliable and informative differential expression and abundance statistics for omics datasets. The software builds on previously published statistical methods, which have been applied successfully to biomedical omics data but are difficult to use without prior experience in programming or scripting. RepExplore facilitates the analysis by providing a fully automated data processing and interactive ranking tables, whisker plot, heat map and principal component analysis visualizations to interpret omics data and derived statistics. AVAILABILITY AND IMPLEMENTATION: Freely available at http://www.repexplore.tk CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Enrico Glaab, Reinhard Schneider 0002 |
Bioinform. | 1 |
| 2012 | EnrichNet: network-based gene set enrichment analysisabstractMOTIVATION: Assessing functional associations between an experimentally derived gene or protein set of interest and a database of known gene/protein sets is a common task in the analysis of large-scale functional genomics data. For this purpose, a frequently used approach is to apply an over-representation-based enrichment analysis. However, this approach has four drawbacks: (i) it can only score functional associations of overlapping gene/proteins sets; (ii) it disregards genes with missing annotations; (iii) it does not take into account the network structure of physical interactions between the gene/protein sets of interest and (iv) tissue-specific gene/protein set associations cannot be recognized. RESULTS: To address these limitations, we introduce an integrative analysis approach and web-application called EnrichNet. It combines a novel graph-based statistic with an interactive sub-network visualization to accomplish two complementary goals: improving the prioritization of putative functional gene/protein set associations by exploiting information from molecular interaction networks and tissue-specific gene expression data and enabling a direct biological interpretation of the results. By using the approach to analyse sets of genes with known involvement in human diseases, new pathway associations are identified, reflecting a dense sub-network of interactions between their corresponding proteins. AVAILABILITY: EnrichNet is freely available at http://www.enrichnet.org. CONTACT: [email protected], [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics Online. Enrico Glaab, Anaïs Baudot, Natalio Krasnogor, Reinhard Schneider 0002, Alfonso Valencia |
Bioinform. | 1 |
| 2012 | PathVar: analysis of gene and protein expression variance in cellular pathways using microarray dataabstractSUMMARY: Finding significant differences between the expression levels of genes or proteins across diverse biological conditions is one of the primary goals in the analysis of functional genomics data. However, existing methods for identifying differentially expressed genes or sets of genes by comparing measures of the average expression across predefined sample groups do not detect differential variance in the expression levels across genes in cellular pathways. Since corresponding pathway deregulations occur frequently in microarray gene or protein expression data, we present a new dedicated web application, PathVar, to analyze these data sources. The software ranks pathway-representing gene/protein sets in terms of the differences of the variance in the within-pathway expression levels across different biological conditions. Apart from identifying new pathway deregulation patterns, the tool exploits these patterns by combining different machine learning methods to find clusters of similar samples and build sample classification models. AVAILABILITY: freely available at http://pathvar.embl.de CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Enrico Glaab, Reinhard Schneider 0002 |
Bioinform. | 1 |
| 2010 | TopoGSA: network topological gene set analysisabstractUNLABELLED: TopoGSA (Topology-based Gene Set Analysis) is a web-application dedicated to the computation and visualization of network topological properties for gene and protein sets in molecular interaction networks. Different topological characteristics, such as the centrality of nodes in the network or their tendency to form clusters, can be computed and compared with those of known cellular pathways and processes. AVAILABILITY: Freely available at http://www.infobiotics.net/topogsa. Enrico Glaab, Anaïs Baudot, Natalio Krasnogor, Alfonso Valencia |
Bioinform. | 1 |
| 2010 | Extending pathways and processes using molecular interaction networks to analyse cancer genome dataabstractBACKGROUND: Cellular processes and pathways, whose deregulation may contribute to the development of cancers, are often represented as cascades of proteins transmitting a signal from the cell surface to the nucleus. However, recent functional genomic experiments have identified thousands of interactions for the signalling canonical proteins, challenging the traditional view of pathways as independent functional entities. Combining information from pathway databases and interaction networks obtained from functional genomic experiments is therefore a promising strategy to obtain more robust pathway and process representations, facilitating the study of cancer-related pathways. RESULTS: We present a methodology for extending pre-defined protein sets representing cellular pathways and processes by mapping them onto a protein-protein interaction network, and extending them to include densely interconnected interaction partners. The added proteins display distinctive network topological features and molecular function annotations, and can be proposed as putative new components, and/or as regulators of the communication between the different cellular processes. Finally, these extended pathways and processes are used to analyse their enrichment in pancreatic mutated genes. Significant associations between mutated genes and certain processes are identified, enabling an analysis of the influence of previously non-annotated cancer mutated genes. CONCLUSIONS: The proposed method for extending cellular pathways helps to explain the functions of cancer mutated genes by exploiting the synergies of canonical knowledge and large-scale interaction data. Enrico Glaab, Anaïs Baudot, Natalio Krasnogor, Alfonso Valencia |
BMC Bioinform. | 1 |
| 2009 | ArrayMining: a modular web-application for microarray analysis combining ensemble and consensus methods with cross-study normalizationabstractBACKGROUND: Statistical analysis of DNA microarray data provides a valuable diagnostic tool for the investigation of genetic components of diseases. To take advantage of the multitude of available data sets and analysis methods, it is desirable to combine both different algorithms and data from different studies. Applying ensemble learning, consensus clustering and cross-study normalization methods for this purpose in an almost fully automated process and linking different analysis modules together under a single interface would simplify many microarray analysis tasks. RESULTS: We present ArrayMining.net, a web-application for microarray analysis that provides easy access to a wide choice of feature selection, clustering, prediction, gene set analysis and cross-study normalization methods. In contrast to other microarray-related web-tools, multiple algorithms and data sets for an analysis task can be combined using ensemble feature selection, ensemble prediction, consensus clustering and cross-platform data integration. By interlinking different analysis tools in a modular fashion, new exploratory routes become available, e.g. ensemble sample classification using features obtained from a gene set analysis and data from multiple studies. The analysis is further simplified by automatic parameter selection mechanisms and linkage to web tools and databases for functional annotation and literature mining. CONCLUSION: ArrayMining.net is a free web-application for microarray analysis combining a broad choice of algorithms based on ensemble and consensus methods, using automatic parameter selection and integration with annotation databases. Enrico Glaab, Jonathan M. Garibaldi, Natalio Krasnogor |
BMC Bioinform. | 1 |