EDBT 2026 Demo / reviewers in the wild / expert
Russell Bowler
dblp:139/1619 · also Russell P. Bowler
· DBLP profile ↗
15ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0003-4651-363XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BioNeuralNet: a graph neural network based Multi-Omics network data analysis toolabstractSUMMARY: Multi-omics data offer unprecedented insights into complex biological systems, yet their high dimensionality, sparsity, and intricate interactions pose significant analytical challenges. Network-based approaches have advanced multi-omics research by effectively capturing biologically relevant relationships among molecular features (e.g., genes, proteins, metabolites). While these methods are powerful for representing molecular interactions, there remains a need for tools specifically designed to effectively utilize these network representations across diverse downstream analyses. To fulfill this need, we introduce BioNeuralNet, a flexible and modular Python framework tailored for end-to-end network-based multi-omics data analysis. BioNeuralNet leverages Graph Neural Networks (GNNs) to learn biologically meaningful low-dimensional representations from multi-omics networks, converting these complex molecular networks into versatile embeddings. BioNeuralNet supports all major stages of multi-omics network analysis, including several network construction techniques, generation of low-dimensional representations, and a broad range of downstream analytical tasks. Its extensive utilities, including diverse GNN architectures, and compatibility with established Python packages (e.g., scikit-learn, PyTorch, NetworkX), enhance usability and facilitate quick adoption. BioNeuralNet is an open-source, user-friendly, and extensively documented framework designed to support flexible and reproducible multi-omics network analysis in precision medicine. AVAILABILITY AND IMPLEMENTATION: The BioNeuralNet library is available via The Python Package Index (PyPI). Source code, documentation, tutorials, and workflows are hosted at https://bioneuralnet.readthedocs.io. Code archived at https://doi.org/10.5281/zenodo.17503083. Vicente Ramos, Sundous Hussein, Mohamed Abdel-Hafiz, Arunangshu Sarkar, Weixuan Liu, Katerina J. Kechris, Russell Bowler, Leslie Lange, Farnoush Banaei Kashani |
Bioinform. | 7 |
| 2025 | A generalized higher-order correlation analysis framework for multi-omics network inferenceabstractMultiple -omics (genomics, proteomics, etc.) profiles are commonly generated to gain insight into a disease or physiological system. Constructing multi-omics networks with respect to the trait(s) of interest provides an opportunity to understand relationships between molecular features but integration is challenging due to multiple data sets with high dimensionality. One approach is to use canonical correlation to integrate one or two omics types and a single trait of interest. However, these types of methods may be limited due to (1) not accounting for higher-order correlations existing among features, (2) computational inefficiency when extending to more than two omics data when using a penalty term-based sparsity method, and (3) lack of flexibility for focusing on specific correlations (e.g., omics-to-phenotype correlation versus omics-to-omics correlations). In this work, we have developed a novel multi-omics network analysis pipeline called Sparse Generalized Tensor Canonical Correlation Analysis Network Inference (SGTCCA-Net) that can effectively overcome these limitations. We also introduce an implementation to improve the summarization of networks for downstream analyses. Simulation and real-data experiments demonstrate the effectiveness of our novel method for inferring omics networks and features of interest. Weixuan Liu, Katherine A. Pratte, Peter J. Castaldi, Craig P. Hersh, Russell Bowler, Farnoush Banaei Kashani, Katerina J. Kechris |
PLoS Comput. Biol. | 5 |
| 2024 | Learning from Multi-Omics Networks to Enhance Disease Prediction: An Optimized Network Embedding and Fusion ApproachabstractUnderstanding complex diseases hinges on a profound understanding of intricate biomolecular interactions unfolding within a complex, multidimensional landscape, challenging traditional methods to extract meaningful insights. While multi-omics networks capture the richness of biological data, providing a basis for predicting relationships between biomolecules and various phenotypic traits of complex diseases, their inherent complexity limits their predictive power. To address this challenge, we introduce a novel pipeline that leverages the power of Graph Neural Networks (GNNs) to extract and integrate meaningful information from multi-omics networks. By generating informative node embeddings and seamlessly incorporating them into the original subject-level data, our approach captures both local and global network dependencies, leading to substantial improvements in disease prediction accuracy. The proposed pipeline optimizes the embedding generation process for the specific prediction task, enabling the model to learn task-relevant representations. Through rigorous experimentation, we demonstrate the superior performance of our approach, surpassing existing methods by a substantial margin on nine real-world multi-omics datasets. With remarkable increases in accuracy ranging approximately from 8% to 10% over the best-performing baseline, particularly when the multi-omics networks are moderately dense, striking a balance between capturing complex relationships and avoiding excessive noise. Our findings underscore the potential of GNNs to significantly improve disease prediction by effectively extracting and representing knowledge embedded within multi-omics networks. Sundous Hussein, Vicente Ramos, Weixuan Liu, Katerina J. Kechris, Leslie Lange, Russell Bowler, Farnoush Banaei Kashani |
BIBM | 6 |
| 2024 | HIP: a method for high-dimensional multi-view data integration and prediction accounting for subgroup heterogeneityabstractEpidemiologic and genetic studies in many complex diseases suggest subgroup disparities (e.g. by sex, race) in disease course and patient outcomes. We consider this from the standpoint of integrative analysis where we combine information from different views (e.g. genomics, proteomics, clinical data). Existing integrative analysis methods ignore the heterogeneity in subgroups, and stacking the views and accounting for subgroup heterogeneity does not model the association among the views. We propose Heterogeneity in Integration and Prediction (HIP), a statistical approach for joint association and prediction that leverages the strengths in each view to identify molecular signatures that are shared by and specific to a subgroup. We apply HIP to proteomics and gene expression data pertaining to chronic obstructive pulmonary disease (COPD) to identify proteins and genes shared by, and unique to, males and females, contributing to the variation in COPD, measured by airway wall thickness. Our COPD findings have identified proteins, genes, and pathways that are common across and specific to males and females, some implicated in COPD, while others could lead to new insights into sex differences in COPD mechanisms. HIP accounts for subgroup heterogeneity in multi-view data, ranks variables based on importance, is applicable to univariate or multivariate continuous outcomes, and incorporates covariate adjustment. With the efficient algorithms implemented using PyTorch, this method has many potential scientific applications and could enhance multiomics research in health disparities. HIP is available at https://github.com/lasandrall/HIP, a video tutorial at https://youtu.be/O6E2OLmeMDo and a Shiny Application at https://multi-viewlearn.shinyapps.io/HIP_ShinyApp/ for users with limited programming experience. Jessica Butts, Leif Verace, Christine Wendt, Russell Bowler, Craig P. Hersh, Qi Long, Lynn E. Eberly, Sandra Safo |
Briefings Bioinform. | 4 |
| 2024 | PathIntegrate: Multivariate modelling approaches for pathway-based multi-omics data integrationabstractAs terabytes of multi-omics data are being generated, there is an ever-increasing need for methods facilitating the integration and interpretation of such data. Current multi-omics integration methods typically output lists, clusters, or subnetworks of molecules related to an outcome. Even with expert domain knowledge, discerning the biological processes involved is a time-consuming activity. Here we propose PathIntegrate, a method for integrating multi-omics datasets based on pathways, designed to exploit knowledge of biological systems and thus provide interpretable models for such studies. PathIntegrate employs single-sample pathway analysis to transform multi-omics datasets from the molecular to the pathway-level, and applies a predictive single-view or multi-view model to integrate the data. Model outputs include multi-omics pathways ranked by their contribution to the outcome prediction, the contribution of each omics layer, and the importance of each molecule in a pathway. Using semi-synthetic data we demonstrate the benefit of grouping molecules into pathways to detect signals in low signal-to-noise scenarios, as well as the ability of PathIntegrate to precisely identify important pathways at low effect sizes. Finally, using COPD and COVID-19 data we showcase how PathIntegrate enables convenient integration and interpretation of complex high-dimensional multi-omics datasets. PathIntegrate is available as an open-source Python package. Cecilia Wieder, Juliette Cooke, Clément Frainay, Nathalie Poupin, Russell Bowler, Fabien Jourdan, Katerina J. Kechris, Rachel P. J. Lai, Timothy M. D. Ebbels |
PLoS Comput. Biol. | 5 |
| 2023 | NetSHy: network summarization via a hybrid approach leveraging topological propertiesabstractMOTIVATION: Biological networks can provide a system-level understanding of underlying processes. In many contexts, networks have a high degree of modularity, i.e. they consist of subsets of nodes, often known as subnetworks or modules, which are highly interconnected and may perform separate functions. In order to perform subsequent analyses to investigate the association between the identified module and a variable of interest, a module summarization, that best explains the module's information and reduces dimensionality is often needed. Conventional approaches for obtaining network representation typically rely only on the profiles of the nodes within the network while disregarding the inherent network topological information. RESULTS: In this article, we propose NetSHy, a hybrid approach which is capable of reducing the dimension of a network while incorporating topological properties to aid the interpretation of the downstream analyses. In particular, NetSHy applies principal component analysis (PCA) on a combination of the node profiles and the well-known Laplacian matrix derived directly from the network similarity matrix to extract a summarization at a subject level. Simulation scenarios based on random and empirical networks at varying network sizes and sparsity levels show that NetSHy outperforms the conventional PCA approach applied directly on node profiles, in terms of recovering the true correlation with a phenotype of interest and maintaining a higher amount of explained variation in the data when networks are relatively sparse. The robustness of NetSHy is also demonstrated by a more consistent correlation with the observed phenotype as the sample size decreases. Lastly, a genome-wide association study is performed as an application of a downstream analysis, where NetSHy summarization scores on the biological networks identify more significant single nucleotide polymorphisms than the conventional network representation. AVAILABILITY AND IMPLEMENTATION: R code implementation of NetSHy is available at https://github.com/thaovu1/NetSHy. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Thao Vu, Elizabeth Litkowski, Weixuan Liu, Katherine A. Pratte, Leslie Lange, Russell Bowler, Farnoush Banaei Kashani, Katerina J. Kechris |
Bioinform. | 6 |
| 2023 | TreeKernel: interpretable kernel machine tests for interactions between -omics and clinical predictors with applications to metabolomics and COPD phenotypesabstractBACKGROUND: In this paper, we are interested in interactions between a high-dimensional -omics dataset and clinical covariates. The goal is to evaluate the relationship between a phenotype of interest and a high-dimensional omics pathway, where the effect of the omics data depends on subjects' clinical covariates (age, sex, smoking status, etc.). For instance, metabolic pathways can vary greatly between sexes which may also change the relationship between certain metabolic pathways and a clinical phenotype of interest. We propose partitioning the clinical covariate space and performing a kernel association test within those partitions. To illustrate this idea, we focus on hierarchical partitions of the clinical covariate space and kernel tests on metabolic pathways. RESULTS: We see that our proposed method outperforms competing methods in most simulation scenarios. It can identify different relationships among clinical groups with higher power in most scenarios while maintaining a proper Type I error rate. The simulation studies also show a robustness to the grouping structure within the clinical space. We also apply the method to the COPDGene study and find several clinically meaningful interactions between metabolic pathways, the clinical space, and lung function. CONCLUSION: TreeKernel provides a simple and interpretable process for testing for relationships between high-dimensional omics data and clinical outcomes in the presence of interactions within clinical cohorts. The method is broadly applicable to many studies. Charlie M. Carpenter, Lucas A. Gillenwater, Russell Bowler, Katerina J. Kechris, Debashis Ghosh |
BMC Bioinform. | 3 |
| 2022 | Effective Subject Representation based on Multi-omics Disease Networks using Graph EmbeddingabstractThe study of complex behavior of biological systems has become increasingly dependent on evolutionary network modeling. In particular, multi-omics networks capture interactions between biomolecules such as proteins and metabolites, providing a basis for predicting relationships between such biomolecules and various phenotypic traits of complex diseases. In this paper, we introduce an integrative framework that given a multi-omics network representing a cohort of subjects, learns expressive representations for network nodes, and combines the learned nodes representations with the biological profiles of individual subjects for enriched representation of the subjects. With extensive empirical evaluation using real-world multi-omics networks, we show that our proposed framework significantly outperforms existing and baseline methods in terms of subject representation accuracy, particularly when the multi-omics network representing the cohort is sparse and structured and therefore, more informative. Sundous Hussein, Thao Vu, Leslie Lange, Russell Bowler, Katerina J. Kechris, Farnoush Banaei Kashani |
BIBM | 4 |
| 2022 | Semi-supervised Embedding for Scalable and Accurate Time Series ClusteringabstractWhile time series data are abundant in numerous real world applications, large labeled time series datasets are scarce. Semi-supervised models, which leverage small amounts of labeled data along with a large set of unlabeled data, have been shown to significantly outperform unsupervised learning models that only rely on unlabeled data for time series clustering. However, existing semi-supervised time series clustering algorithms suffer from lack of scalability as they are limited to perform learning operations within the original data space. We propose a scalable and accurate autoencoder-based semi-supervised learning model for time series clustering in the embedded space. With this model, we also introduce multiple semi-supervised objective functions that leverage only a small number of labeled examples but significantly improve the quality of the autoencoder’s learned latent space for clustering. Our experiments on a variety of datasets show that our methods can often improve performance of a typical clustering method (namely, k-means). We demonstrate that our methods achieve a maximum average Adjusted Rand Index (ARI) of 0.897, a 140% increase over an unsupervised Convolutional Autoencoder (CAE) model. Finally, our proposed methods also achieve a maximum improvement of 44% over an existing semi-supervised model. Russell Bowler, Katerina J. Kechris, Farnoush Banaei Kashani |
IEEE Big Data | 2 |
| 2021 | PaIRKAT: A pathway integrated regression-based kernel association test with applications to metabolomics and COPD phenotypesabstractHigh-throughput data such as metabolomics, genomics, transcriptomics, and proteomics have become familiar data types within the "-omics" family. For this work, we focus on subsets that interact with one another and represent these "pathways" as graphs. Observed pathways often have disjoint components, i.e., nodes or sets of nodes (metabolites, etc.) not connected to any other within the pathway, which notably lessens testing power. In this paper we propose the Pathway Integrated Regression-based Kernel Association Test (PaIRKAT), a new kernel machine regression method for incorporating known pathway information into the semi-parametric kernel regression framework. This work extends previous kernel machine approaches. This paper also contributes an application of a graph kernel regularization method for overcoming disconnected pathways. By incorporating a regularized or "smoothed" graph into a score test, PaIRKAT can provide more powerful tests for associations between biological pathways and phenotypes of interest and will be helpful in identifying novel pathways for targeted clinical research. We evaluate this method through several simulation studies and an application to real metabolomics data from the COPDGene study. Our simulation studies illustrate the robustness of this method to incorrect and incomplete pathway knowledge, and the real data analysis shows meaningful improvements of testing power in pathways. PaIRKAT was developed for application to metabolomic pathway data, but the techniques are easily generalizable to other data sources with a graph-like structure. Charlie M. Carpenter, Lucas A. Gillenwater, Cameron Severn, Tusharkanti Ghosh, Russell Bowler, Katerina J. Kechris, Debashis Ghosh |
PLoS Comput. Biol. | 6 |
| 2021 | Improved prediction of smoking status via isoform-aware RNA-seq deep learning modelsabstractMost predictive models based on gene expression data do not leverage information related to gene splicing, despite the fact that splicing is a fundamental feature of eukaryotic gene expression. Cigarette smoking is an important environmental risk factor for many diseases, and it has profound effects on gene expression. Using smoking status as a prediction target, we developed deep neural network predictive models using gene, exon, and isoform level quantifications from RNA sequencing data in 2,557 subjects in the COPDGene Study. We observed that models using exon and isoform quantifications clearly outperformed gene-level models when using data from 5 genes from a previously published prediction model. Whereas the test set performance of the previously published model was 0.82 in the original publication, our exon-based models including an exon-to-isoform mapping layer achieved a test set AUC (area under the receiver operating characteristic) of 0.88, which improved to an AUC of 0.94 using exon quantifications from a larger set of genes. Isoform variability is an important source of latent information in RNA-seq data that can be used to improve clinical prediction models. Zifeng Wang 0002, Aria Masoomi, Zhonghui Xu, Adel Boueiz, Sool Lee, Russell Bowler, Michael H. Cho, Edwin K. Silverman, Craig P. Hersh, Jennifer G. Dy, Peter J. Castaldi |
PLoS Comput. Biol. | 7 |
| 2019 | The Utility of Shapelets for Analyzing Physical Activity of COPD Patients and non-COPD controlsabstractPhysical activity is an attractive endpoint for novel therapies in Chronic Obstructive Pulmonary Disease (COPD). However, a deep understanding about COPD physical activity patterns and disease severity is lacking. In this research, we study the physical activity patterns for 184 individuals with and without COPD from a single center in the COPDGene cohort. These subjects participated in a 3-week observational study wearing wrist-worn accelerometers for collecting physical activity data. Our exploratory data analysis finds using the whole range of activity data is insufficient for patient clustering. Alternatively, we use shapelets, small and local sub-sequences, to better capture patients' behaviors in different groups. We develop a length-bound heuristic algorithm for choosing the subset that has the best clustering result. The study shows the potentials of using shapelets for helping providers in assessing COPD patients' status. Nicholas Locantore, Matthew Allinder, Divya Mohan, Russell Bowler |
BIBM | 6 |
| 2017 | The discordant method: a novel approach for differential correlationabstractBioinformatics (2016) 32(5), 690–696. doi:10.1093/bioinformatics/btv633 The authors of the above paper wish to inform readers that there was a typographical error in Equations 2 and 6 of the published paper. The corrected equations are given below. The paper has now been corrected online. Charlotte Siska, Russell Bowler, Katerina J. Kechris |
Bioinform. | 2 |
| 2016 | The discordant method: a novel approach for differential correlationabstractMOTIVATION: Current differential correlation methods are designed to determine molecular feature pairs that have the largest magnitude of difference between correlation coefficients. These methods do not easily capture molecular feature pairs that experience no correlation in one group but correlation in another, which may reflect certain types of biological interactions. We have developed a tool, the Discordant method, which categorizes the correlation types for each group to make this possible. RESULTS: We compare the Discordant method to existing approaches using simulations and two biological datasets with different types of -omics data. In contrast to other methods, Discordant identifies phenotype-related features at a similar or higher rate while maintaining reasonable computational tractability and usability. AVAILABILITY AND IMPLEMENTATION: R code and sample data are available at https://github.com/siskac/discordant CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Charlotte Siska, Russell Bowler, Katerina J. Kechris |
Bioinform. | 2 |
| 2014 | MSPrep - Summarization, normalization and diagnostics for processing of mass spectrometry-based metabolomic dataabstractAbstract Motivation: Although R packages exist for the pre-processing of metabolomic data, they currently do not incorporate additional analysis steps of summarization, filtering and normalization of aligned data. We developed the MSPrep R package to complement other packages by providing these additional steps, implementing a selection of popular normalization algorithms and generating diagnostics to help guide investigators in their analyses. Availability: http://www.sourceforge.net/projects/msprep Contact: [email protected] Supplementary Information: Supplementary materials are available at Bioinformatics online. Grant Hughes, Charmion Cruickshank-Quinn, Richard Reisdorph, Sharon Lutz, Irina Petrache, Nichole Reisdorph, Russell Bowler, Katerina J. Kechris |
Bioinform. | 7 |