VLDB 2026 Research / reviewers in the wild / expert
Zuoheng Wang
dblp:123/5975
· DBLP profile ↗
11ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-7251-3687ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 11 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing patient representation learning with inferred family pedigrees improves disease risk predictionabstractBACKGROUND: Machine learning and deep learning are powerful tools for analyzing electronic health records (EHRs) in healthcare research. Although family health history has been recognized as a major predictor for a wide spectrum of diseases, research has so far adopted a limited view of family relations, essentially treating patients as independent samples in the analysis. METHODS: To address this gap, we present ALIGATEHR, which models inferred family relations in a graph attention network augmented with an attention-based medical ontology representation, thus accounting for the complex influence of genetics, shared environmental exposures, and disease dependencies. RESULTS: Taking disease risk prediction as a use case, we demonstrate that explicitly modeling family relations significantly improves predictions across the disease spectrum. We then show how ALIGATEHR's attention mechanism, which links patients' disease risk to their relatives' clinical profiles, successfully captures genetic aspects of diseases using longitudinal EHR diagnosis data. Finally, we use ALIGATEHR to successfully distinguish the 2 main inflammatory bowel disease subtypes with highly shared risk factors and symptoms (Crohn's disease and ulcerative colitis). CONCLUSION: Overall, our results highlight that family relations should not be overlooked in EHR research and illustrate ALIGATEHR's great potential for enhancing patient representation learning for predictive and interpretable modeling of EHRs. Xiayuan Huang, Jatin Arora 0010, Abdullah Mesut Erzurumluoglu, Stephen A. Stanhope, Daniel Lam, Pierre Khoueiry, Jan N. Jensen, James Cai, Nathan Lawless, Jan Kriegl, Zhihao Ding, Johann de Jong, Zuoheng Wang |
J. Am. Medical Informatics Assoc. | 14 |
| 2024 | Transformer with convolution and graph-node co-embedding: An accurate and interpretable vision backbone for predicting gene expressions from local histopathological imageabstractInferring gene expressions from histopathological images has long been a fascinating yet challenging task, primarily due to the substantial disparities between the two modality. Existing strategies using local or global features of histological images are suffering model complexity, GPU consumption, low interpretability, insufficient encoding of local features, and over-smooth prediction of gene expressions among neighboring sites. In this paper, we develop TCGN (Transformer with Convolution and Graph-Node co-embedding method) for gene expression estimation from H&E-stained pathological slide images. TCGN comprises a combination of convolutional layers, transformer encoders, and graph neural networks, and is the first to integrate these blocks in a general and interpretable computer vision backbone. Notably, TCGN uniquely operates with just a single spot image as input for histopathological image analysis, simplifying the process while maintaining interpretability. We validate TCGN on three publicly available spatial transcriptomic datasets. TCGN consistently exhibited the best performance (with median PCC 0.232). TCGN offers superior accuracy while keeping parameters to a minimum (just 86.241 million), and it consumes minimal memory, allowing it to run smoothly even on personal computers. Moreover, TCGN can be extended to handle bulk RNA-seq data while providing the interpretability. Enhancing the accuracy of omics information prediction from pathological images not only establishes a connection between genotype and phenotype, enabling the prediction of costly-to-measure biomarkers from affordable histopathological images, but also lays the groundwork for future multi-modal data modeling. Our results confirm that TCGN is a powerful tool for inferring gene expressions from histopathological images in precision health applications. Yan Kong, Ronghan Li, Zuoheng Wang, Hui Lu 0004 |
Medical Image Anal. | 4 |
| 2023 | iDESC: identifying differential expression in single-cell RNA sequencing data with multiple subjectsabstractBACKGROUND: Single-cell RNA sequencing (scRNA-seq) technology has enabled assessment of transcriptome-wide changes at single-cell resolution. Due to the heterogeneity in environmental exposure and genetic background across subjects, subject effect contributes to the major source of variation in scRNA-seq data with multiple subjects, which severely confounds cell type specific differential expression (DE) analysis. Moreover, dropout events are prevalent in scRNA-seq data, leading to excessive number of zeroes in the data, which further aggravates the challenge in DE analysis. RESULTS: We developed iDESC to detect cell type specific DE genes between two groups of subjects in scRNA-seq data. iDESC uses a zero-inflated negative binomial mixed model to consider both subject effect and dropouts. The prevalence of dropout events (dropout rate) was demonstrated to be dependent on gene expression level, which is modeled by pooling information across genes. Subject effect is modeled as a random effect in the log-mean of the negative binomial component. We evaluated and compared the performance of iDESC with eleven existing DE analysis methods. Using simulated data, we demonstrated that iDESC had well-controlled type I error and higher power compared to the existing methods. Applications of those methods with well-controlled type I error to three real scRNA-seq datasets from the same tissue and disease showed that the results of iDESC achieved the best consistency between datasets and the best disease relevance. CONCLUSIONS: iDESC was able to achieve more accurate and robust DE analysis results by separating subject effect from disease effect with consideration of dropouts to identify DE genes, suggesting the importance of considering subject effect and dropouts in the DE analysis of scRNA-seq data with multiple subjects. Taylor Sterling Adams, Ningya Wang, Jonas Schupp, Weimiao Wu, John E. McDonough, Geoffrey Lowell Chupp, Naftali Kaminski, Zuoheng Wang, Xiting Yan |
BMC Bioinform. | 10 |
| 2023 | Correction: iDESC: identifying differential expression in single-cell RNA sequencing data with multiple subjects
Taylor Sterling Adams, Ningya Wang, Jonas Schupp, Weimiao Wu, John E. McDonough, Geoffrey Lowell Chupp, Naftali Kaminski, Zuoheng Wang, Xiting Yan |
BMC Bioinform. | 10 |
| 2021 | G2S3: A gene graph-based imputation method for single-cell RNA sequencing dataabstractSingle-cell RNA sequencing technology provides an opportunity to study gene expression at single-cell resolution. However, prevalent dropout events result in high data sparsity and noise that may obscure downstream analyses in single-cell transcriptomic studies. We propose a new method, G2S3, that imputes dropouts by borrowing information from adjacent genes in a sparse gene graph learned from gene expression profiles across cells. We applied G2S3 and ten existing imputation methods to eight single-cell transcriptomic datasets and compared their performance. Our results demonstrated that G2S3 has superior overall performance in recovering gene expression, identifying cell subtypes, reconstructing cell trajectories, identifying differentially expressed genes, and recovering gene regulatory and correlation relationships. Moreover, G2S3 is computationally efficient for imputation in large-scale single-cell transcriptomic datasets. Weimiao Wu, Qile Dai, Xiting Yan, Zuoheng Wang |
PLoS Comput. Biol. | 5 |
| 2019 | Identification of trans-eQTLs using mediation analysis with multiple mediatorsabstractBACKGROUND: Mapping expression quantitative trait loci (eQTLs) has provided insight into gene regulation. Compared to cis-eQTLs, the regulatory mechanisms of trans-eQTLs are less known. Previous studies suggest that trans-eQTLs may regulate expression of remote genes by altering the expression of nearby genes. Trans-association has been studied in the mediation analysis with a single mediator. However, prior applications with one mediator are prone to model misspecification due to correlations between genes. Motivated from the observation that trans-eQTLs are more likely to associate with more than one cis-gene than randomly selected SNPs in the GTEx dataset, we developed a computational method to identify trans-eQTLs that are mediated by multiple mediators. RESULTS: We proposed two hypothesis tests for testing the total mediation effect (TME) and the component-wise mediation effects (CME), respectively. We demonstrated in simulation studies that the type I error rates were controlled in both tests despite model misspecification. The TME test was more powerful than the CME test when the two mediation effects are in the same direction, while the CME test was more powerful than the TME test when the two mediation effects are in opposite direction. Multiple mediator analysis had increased power to detect mediated trans-eQTLs, especially in large samples. In the HapMap3 data, we identified 11 mediated trans-eQTLs that were not detected by the single mediator analysis in the combined samples of African populations. Moreover, the mediated trans-eQTLs in the HapMap3 samples are more likely to be trait-associated SNPs. In terms of computation, although there is no limit in the number of mediators in our model, analysis takes more time when adding additional mediators. In the analysis of the HapMap3 samples, we included at most 5 cis-gene mediators. Majority of the trios we considered have one or two mediators. CONCLUSIONS: Trans-eQTLs are more likely to associate with multiple cis-genes than randomly selected SNPs. Mediation analysis with multiple mediators improves power of identification of mediated trans-eQTLs, especially in large samples. Nayang Shan, Zuoheng Wang, Lin Hou 0003 |
BMC Bioinform. | 2 |
| 2014 | A model for family-based case-control studies of genetic imprinting and epistasisabstractGenetic imprinting, or called the parent-of-origin effect, has been recognized to play an important role in the formation and pathogenesis of human diseases. Although the epigenetic mechanisms that establish genetic imprinting have been a focus of many genetic studies, our knowledge about the number of imprinting genes and their chromosomal locations and interactions with other genes is still scarce, limiting precise inference of the genetic architecture of complex diseases. In this article, we present a statistical model for testing and estimating the effects of genetic imprinting on complex diseases using a commonly used case-control design with family structure. For each subject sampled from a case and control population, we not only genotype its own single nucleotide polymorphisms (SNPs) but also collect its parents' genotypes. By tracing the transmission pattern of SNP alleles from parental to offspring generation, the model allows the characterization of genetic imprinting effects based on Pearson tests of a 2 × 2 contingency table. The model is expanded to test the interactions between imprinting effects and additive, dominant and epistatic effects in a complex web of genetic interactions. Statistical properties of the model are investigated, and its practical usefulness is validated by a real data analysis. The model will provide a useful tool for genome-wide association studies aimed to elucidate the picture of genetic control over complex human diseases. Xin Li 0025, Yihan Sui, Jianxin Wang 0004, Yongci Li, Zhenwu Lin, John Hegarty, Walter A. Koltun, Zuoheng Wang, Rongling Wu |
Briefings Bioinform. | 9 |
| 2014 | A case-control design for testing and estimating epigenetic effects on complex diseasesabstractEpigenetic modifications may play an important role in the formation and progression of complex diseases through the regulation of gene expression. The systematic identification of epigenetic variants that contribute to human diseases can be made possible using genome-wide association studies (GWAS), although epigenetic effects are currently not included in commonly used case-control designs for GWAS. Here, we show that epigenetic modifications can be integrated into a case-control setting by dissolving the overall genetic effect into its different components, additive, dominant and epigenetic. We describe a general procedure for testing and estimating the significance of each component based on a conventional chi-squared test approach. Simulation studies were performed to investigate the power and false-positive rate of this procedure, providing recommendations for its practical use. The integration of epigenetic variants into GWAS can potentially improve our understanding of how genetic, environmental and stochastic factors interact with epialleles to construct the genetic architecture of complex diseases. Yihan Sui, Weimiao Wu, Zhong Wang 0001, Jianxin Wang 0004, Zuoheng Wang, Rongling Wu |
Briefings Bioinform. | 5 |
| 2014 | Towards a comprehensive picture of the genetic landscape of complex traitsabstractThe formation of phenotypic traits, such as biomass production, tumor volume and viral abundance, undergoes a complex process in which interactions between genes and developmental stimuli take place at each level of biological organization from cells to organisms. Traditional studies emphasize the impact of genes by directly linking DNA-based markers with static phenotypic values. Functional mapping, derived to detect genes that control developmental processes using growth equations, has proven powerful for addressing questions about the roles of genes in development. By treating phenotypic formation as a cohesive system using differential equations, a different approach-systems mapping-dissects the system into interconnected elements and then map genes that determine a web of interactions among these elements, facilitating our understanding of the genetic machineries for phenotypic development. Here, we argue that genetic mapping can play a more important role in studying the genotype-phenotype relationship by filling the gaps in the biochemical and regulatory process from DNA to end-point phenotype. We describe a new framework, named network mapping, to study the genetic architecture of complex traits by integrating the regulatory networks that cause a high-order phenotype. Network mapping makes use of a system of differential equations to quantify the rule by which transcriptional, proteomic and metabolomic components interact with each other to organize into a functional whole. The synthesis of functional mapping, systems mapping and network mapping provides a novel avenue to decipher a comprehensive picture of the genetic landscape of complex phenotypes that underlie economically and biomedically important traits. Zhong Wang 0001, Yaqun Wang, Ningtao Wang, Jianxin Wang 0004, Zuoheng Wang, C. Eduardo Vallejos, Rongling Wu |
Briefings Bioinform. | 5 |
| 2013 | A quantitative model of transcriptional differentiation driving host-pathogen interactionsabstractDespite our expanding knowledge about the biochemistry of gene regulation involved in host-pathogen interactions, a quantitative understanding of this process at a transcriptional level is still limited. We devise and assess a computational framework that can address this question. This framework is founded on a mixture model-based likelihood, equipped with functionality to cluster genes per dynamic and functional changes of gene expression within an interconnected system composed of the host and pathogen. If genes from the host and pathogen are clustered in the same group due to a similar pattern of dynamic profiles, they are likely to be reciprocally co-evolving. If genes from the two organisms are clustered in different groups, this means that they experience strong host-pathogen interactions. The framework can test the rates of change for individual gene clusters during pathogenic infection and quantify their impacts on host-pathogen interactions. The framework was validated by a pathological study of poplar leaves infected by fungal Marssonina brunnea in which co-evolving and interactive genes that determine poplar-fungus interactions are identified. The new framework should find its wide application to studying host-pathogen interactions for any other interconnected systems. Zhong Wang 0001, Jianxin Wang 0004, Yaqun Wang, Ningtao Wang, Zuoheng Wang, Xiaohua Su, Mingxiu Wang, Shougong Zhang, Minren Huang, Rongling Wu |
Briefings Bioinform. | 6 |
| 2012 | A quantitative genetic and epigenetic model of complex traitsabstractBACKGROUND: Despite our increasing recognition of the mechanisms that specify and propagate epigenetic states of gene expression, the pattern of how epigenetic modifications contribute to the overall genetic variation of a phenotypic trait remains largely elusive. RESULTS: We construct a quantitative model to explore the effect of epigenetic modifications that occur at specific rates on the genome. This model, derived from, but beyond, the traditional quantitative genetic theory that is founded on Mendel's laws, allows questions concerning the prevalence and importance of epigenetic variation to be incorporated and addressed. CONCLUSIONS: It provides a new avenue for bringing chromatin inheritance into the realm of complex traits, facilitating our understanding of the means by which phenotypic variation is generated. Zhong Wang 0001, Zuoheng Wang, Jianxin Wang 0004, Yihan Sui, Jian Zhang 0106, Duanping Liao, Rongling Wu |
BMC Bioinform. | 2 |