Kimberly Glass

dblp:155/1274 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0003-4394-5779ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Differential causal networks highlight sex-based differences in human tissues
abstract
Sex differences appear in healthy and pathological conditions and may influence sex-specific therapeutic responses. Understanding such differences is a key activity for developing precision medicine strategies. This study investigates sex differences in gene expression across 40 human tissues by applying a Differential Causal Network (DCN) analysis using data from the Genotype-Tissue Expression project. We identified sex-based DCNs that highlight distinct molecular mechanisms influencing both health and disease in men and women. For example, in pancreas tissue, genes associated with immune system show significant differences in their regulatory patterns between sexes, demonstrating a possible different response to diseases such as diabetes mellitus and cancer. Our findings provide valuable information on the biological underpinnings of sex differences, offering potential pathways for the development of precision medicine strategies.
Annamaria Defilippo, Kimberly Glass, Federico Manuel Giorgi, Tamer Kahveci, Pierangelo Veltri, Pietro H. Guzzi
Briefings Bioinform.2
2024 BONOBO: Bayesian Optimized Sample-Specific Networks Obtained by Omics Data
Enakshi Saha, Viola Fanfani, Panagiotis Mandros, Marouen Ben Guebila, Jonas Fischer, Katherine H. Shutta, Kimberly Glass, Dawn L. DeMeo, Camila Miranda Lopes-Ramos, John Quackenbush
RECOMB7
2024 Partial correlation network analysis identifies coordinated gene expression within a regional cluster of COPD genome-wide association signals
abstract
Chronic obstructive pulmonary disease (COPD) is a complex disease influenced by well-established environmental exposures (most notably, cigarette smoking) and incompletely defined genetic factors. The chromosome 4q region harbors multiple genetic risk loci for COPD, including signals near HHIP, FAM13A, GSTCD, TET2, and BTC. Leveraging RNA-Seq data from lung tissue in COPD cases and controls, we estimated the co-expression network for genes in the 4q region bounded by HHIP and BTC (~70MB), through partial correlations informed by protein-protein interactions. We identified several co-expressed gene pairs based on partial correlations, including NPNT-HHIP, BTC-NPNT and FAM13A-TET2, which were replicated in independent lung tissue cohorts. Upon clustering the co-expression network, we observed that four genes previously associated to COPD: BTC, HHIP, NPNT and PPM1K appeared in the same network community. Finally, we discovered a sub-network of genes differentially co-expressed between COPD vs controls (including FAM13A, PPA2, PPM1K and TET2). Many of these genes were previously implicated in cell-based knock-out experiments, including the knocking out of SPP1 which belongs to the same genomic region and could be a potential local key regulatory gene. These analyses identify chromosome 4q as a region enriched for COPD genetic susceptibility and differential co-expression.
Michele Gentili, Kimberly Glass, Enrico Maiorino, Brian D. Hobbs, Zhonghui Xu, Peter J. Castaldi, Michael H. Cho, Craig P. Hersh, Dandi Qiao, Jarrett D. Morrow, Vincent Carey, John Platig, Edwin K. Silverman
PLoS Comput. Biol.2
2023 Adjustment of spurious correlations in co-expression measurements from RNA-Sequencing data
abstract
MOTIVATION: Gene co-expression measurements are widely used in computational biology to identify coordinated expression patterns across a group of samples. Coordinated expression of genes may indicate that they are controlled by the same transcriptional regulatory program, or involved in common biological processes. Gene co-expression is generally estimated from RNA-Sequencing data, which are commonly normalized to remove technical variability. Here, we demonstrate that certain normalization methods, in particular quantile-based methods, can introduce false-positive associations between genes. These false-positive associations can consequently hamper downstream co-expression network analysis. Quantile-based normalization can, however, be extremely powerful. In particular, when preprocessing large-scale heterogeneous data, quantile-based normalization methods such as smooth quantile normalization can be applied to remove technical variability while maintaining global differences in expression for samples with different biological attributes. RESULTS: We developed SNAIL (Smooth-quantile Normalization Adaptation for the Inference of co-expression Links), a normalization method based on smooth quantile normalization specifically designed for modeling of co-expression measurements. We show that SNAIL avoids formation of false-positive associations in co-expression as well as in downstream network analyses. Using SNAIL, one can avoid arbitrary gene filtering and retain associations to genes that only express in small subgroups of samples. This highlights the method's potential future impact on network modeling and other association-based approaches in large-scale heterogeneous data. AVAILABILITY AND IMPLEMENTATION: The implementation of the SNAIL algorithm and code to reproduce the analyses described in this work can be found in the GitHub repository https://github.com/kuijjerlab/PySNAIL.
Ping-Han Hsieh, Camila Miranda Lopes-Ramos, Manuela Zucknick, Geir Kjetil Sandve, Kimberly Glass, Marieke L. Kuijjer
Bioinform.5
2021 Gene Regulatory Network Inference as Relaxed Graph Matching
abstract
Bipartite network inference is a ubiquitous problem across disciplines. One important example in the field molecular biology is gene regulatory network inference. Gene regulatory networks are an instrumental tool aiding in the discovery of the molecular mechanisms driving diverse diseases, including cancer. However, only noisy observations of the projections of these regulatory networks are typically assayed. In an effort to better estimate regulatory networks from their noisy projections, we formulate a non-convex but analytically tractable optimization problem called OTTER. This problem can be interpreted as relaxed graph matching between the two projections of the bipartite network. OTTER's solutions can be derived explicitly and inspire a spectral algorithm, for which we provide network recovery guarantees. We also provide an alternative approach based on gradient descent that is more robust to noise compared to the spectral algorithm. Interestingly, this gradient descent approach resembles the message passing equations of an established gene regulatory network inference method, PANDA. Using three cancer-related data sets, we show that OTTER outperforms state-of-the-art inference methods in predicting transcription factor binding to gene regulatory regions. To encourage new graph matching applications to this problem, we have made all networks and validation data publicly available.
Deborah A. Weighill, Marouen Ben Guebila, Camila Miranda Lopes-Ramos, Kimberly Glass, John Quackenbush, John Platig, Rebekka Burkholz
AAAI4
2020 PUMA: PANDA Using MicroRNA Associations
abstract
MOTIVATION: Conventional methods to analyze genomic data do not make use of the interplay between multiple factors, such as between microRNAs (miRNAs) and the messenger RNA (mRNA) transcripts they regulate, and thereby often fail to identify the cellular processes that are unique to specific tissues. We developed PUMA (PANDA Using MicroRNA Associations), a computational tool that uses message passing to integrate a prior network of miRNA target predictions with target gene co-expression information to model genome-wide gene regulation by miRNAs. We applied PUMA to 38 tissues from the Genotype-Tissue Expression project, integrating RNA-Seq data with two different miRNA target predictions priors, built on predictions from TargetScan and miRanda, respectively. We found that while target predictions obtained from these two different resources are considerably different, PUMA captures similar tissue-specific miRNA-target regulatory interactions in the different network models. Furthermore, the tissue-specific functions of miRNAs we identified based on regulatory profiles (available at: https://kuijjer.shinyapps.io/puma_gtex/) are highly similar between networks modeled on the two target prediction resources. This indicates that PUMA consistently captures important tissue-specific miRNA regulatory processes. In addition, using PUMA we identified miRNAs regulating important tissue-specific processes that, when mutated, may result in disease development in the same tissue. AVAILABILITY AND IMPLEMENTATION: PUMA is available in C++, MATLAB and Python on GitHub (https://github.com/kuijjerlab and https://netzoo.github.io/). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marieke L. Kuijjer, Maud Fagny, Alessandro Marin, John Quackenbush, Kimberly Glass
Bioinform.5
2019 Precision VISSTA: Bring-Your-Own-Device (BYOD) mHealth Data for Precision Health
Arlene E. Chung, Kimberly Glass, Jacob Leisey-Bartsch, Lucas K. Mentch, Nils Gehlenborg, David Gotz
AMIA2
2019 Precision VISSTA: Machine Learning Prediction and Inference for Bring-Your-Own-Device (BYOD) mHealth Data
Tim Coleman, Lucas K. Mentch, Kimberly Glass, David Gotz, Nils Gehlenborg, Arlene E. Chung
AMIA3
2017 Estimating gene regulatory networks with pandaR
abstract
Abstract PANDA (Passing Attributes between Networks for Data Assimilation) is a gene regulatory network inference method that begins with a model of transcription factor–target gene interactions and uses message passing to update the network model given available transcriptomic and protein–protein interaction data. PANDA is used to estimate networks for each experimental group and the network models are then compared between groups to explore transcriptional processes that distinguish the groups. We present pandaR (bioconductor.org/packages/pandaR), a Bioconductor package that implements PANDA and provides a framework for exploratory data analysis on gene regulatory networks. Availability and Implementation: PandaR is provided as a Bioconductor R Package and is available at bioconductor.org/packages/pandaR.
Daniel Schlauch, Joseph N. Paulson, Albert Young, Kimberly Glass, John Quackenbush
Bioinform.4
2017 Biomarker correlation network in colorectal carcinoma by tumor anatomic location
abstract
BACKGROUND: Colorectal carcinoma evolves through a multitude of molecular events including somatic mutations, epigenetic alterations, and aberrant protein expression, influenced by host immune reactions. One way to interrogate the complex carcinogenic process and interactions between aberrant events is to model a biomarker correlation network. Such a network analysis integrates multidimensional tumor biomarker data to identify key molecular events and pathways that are central to an underlying biological process. Due to embryological, physiological, and microbial differences, proximal and distal colorectal cancers have distinct sets of molecular pathological signatures. Given these differences, we hypothesized that a biomarker correlation network might vary by tumor location. RESULTS: We performed network analyses of 54 biomarkers, including major mutational events, microsatellite instability (MSI), epigenetic features, protein expression status, and immune reactions using data from 1380 colorectal cancer cases: 690 cases with proximal colon cancer and 690 cases with distal colorectal cancer matched by age and sex. Edges were defined by statistically significant correlations between biomarkers using Spearman correlation analyses. We found that the proximal colon cancer network formed a denser network (total number of edges, n = 173) than the distal colorectal cancer network (n = 95) (P < 0.0001 in permutation tests). The value of the average clustering coefficient was 0.50 in the proximal colon cancer network and 0.30 in the distal colorectal cancer network, indicating the greater clustering tendency of the proximal colon cancer network. In particular, MSI was a key hub, highly connected with other biomarkers in proximal colon cancer, but not in distal colorectal cancer. Among patients with non-MSI-high cancer, BRAF mutation status emerged as a distinct marker with higher connectivity in the network of proximal colon cancer, but not in distal colorectal cancer. CONCLUSION: In proximal colon cancer, tumor biomarkers tended to be correlated with each other, and MSI and BRAF mutation functioned as key molecular characteristics during the carcinogenesis. Our findings highlight the importance of considering multiple correlated pathways for therapeutic targets especially in proximal colon cancer.
Reiko Nishihara, Kimberly Glass, Kosuke Mima, Tsuyoshi Hamada, Jonathan A. Nowak, Zhi Rong Qian, Peter Kraft, Edward L. Giovannucci, Charles S. Fuchs, Andrew T. Chan, John Quackenbush, Shuji Ogino, Jukka-Pekka Onnela
BMC Bioinform.2
2017 Tissue-aware RNA-Seq processing and normalization for heterogeneous and sparse data
abstract
BACKGROUND: Although ultrahigh-throughput RNA-Sequencing has become the dominant technology for genome-wide transcriptional profiling, the vast majority of RNA-Seq studies typically profile only tens of samples, and most analytical pipelines are optimized for these smaller studies. However, projects are generating ever-larger data sets comprising RNA-Seq data from hundreds or thousands of samples, often collected at multiple centers and from diverse tissues. These complex data sets present significant analytical challenges due to batch and tissue effects, but provide the opportunity to revisit the assumptions and methods that we use to preprocess, normalize, and filter RNA-Seq data - critical first steps for any subsequent analysis. RESULTS: We find that analysis of large RNA-Seq data sets requires both careful quality control and the need to account for sparsity due to the heterogeneity intrinsic in multi-group studies. We developed Yet Another RNA Normalization software pipeline (YARN), that includes quality control and preprocessing, gene filtering, and normalization steps designed to facilitate downstream analysis of large, heterogeneous RNA-Seq data sets and we demonstrate its use with data from the Genotype-Tissue Expression (GTEx) project. CONCLUSIONS: An R package instantiating YARN is available at http://bioconductor.org/packages/yarn .
Joseph N. Paulson, Cho-Yi Chen, Camila Miranda Lopes-Ramos, Marieke L. Kuijjer, John Platig, Abhijeet R. Sonawane, Maud Fagny, Kimberly Glass, John Quackenbush
BMC Bioinform.8
2016 PyPanda: a Python package for gene regulatory network reconstruction
abstract
PANDA (Passing Attributes between Networks for Data Assimilation) is a gene regulatory network inference method that uses message-passing to integrate multiple sources of 'omics data. PANDA was originally coded in C ++. In this application note we describe PyPanda, the Python version of PANDA. PyPanda runs considerably faster than the C ++ version and includes additional features for network analysis. AVAILABILITY AND IMPLEMENTATION: The open source PyPanda Python package is freely available at http://github.com/davidvi/pypanda CONTACT: [email protected] or [email protected].
David G. P. van IJzendoorn, Kimberly Glass, John Quackenbush, Marieke L. Kuijjer
Bioinform.2
2015 A network model for angiogenesis in ovarian cancer
abstract
BACKGROUND: We recently identified two robust ovarian cancer subtypes, defined by the expression of genes involved in angiogenesis, with significant differences in clinical outcome. To identify potential regulatory mechanisms that distinguish the subtypes we applied PANDA, a method that uses an integrative approach to model information flow in gene regulatory networks. RESULTS: We find distinct differences between networks that are active in the angiogenic and non-angiogenic subtypes, largely defined by a set of key transcription factors that, although previously reported to play a role in angiogenesis, are not strongly differentially-expressed between the subtypes. Our network analysis indicates that these factors are involved in the activation (or repression) of different genes in the two subtypes, resulting in differential expression of their network targets. Mechanisms mediating differences between subtypes include a previously unrecognized pro-angiogenic role for increased genome-wide DNA methylation and complex patterns of combinatorial regulation. CONCLUSIONS: The models we develop require a shift in our interpretation of the driving factors in biological networks away from the genes themselves and toward their interactions. The observed regulatory changes between subtypes suggest therapeutic interventions that may help in the treatment of ovarian cancer.
Kimberly Glass, John Quackenbush, Dimitrios Spentzos, Benjamin Haibe-Kains, Guo-Cheng Yuan
BMC Bioinform.1
2015 Finding New Order in Biological Functions from the Network Structure of Gene Annotations
abstract
The Gene Ontology (GO) provides biologists with a controlled terminology that describes how genes are associated with functions and how functional terms are related to one another. These term-term relationships encode how scientists conceive the organization of biological functions, and they take the form of a directed acyclic graph (DAG). Here, we propose that the network structure of gene-term annotations made using GO can be employed to establish an alternative approach for grouping functional terms that captures intrinsic functional relationships that are not evident in the hierarchical structure established in the GO DAG. Instead of relying on an externally defined organization for biological functions, our approach connects biological functions together if they are performed by the same genes, as indicated in a compendium of gene annotation data from numerous different sources. We show that grouping terms by this alternate scheme provides a new framework with which to describe and predict the functions of experimentally identified sets of genes.
Kimberly Glass, Michelle Girvan
PLoS Comput. Biol.1