VLDB 2026 Research / reviewers in the wild / expert
Jie Zhang 0010
dblp:84/6889-10
· DBLP profile ↗
14ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0001-6939-7905ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Identification of high-risk cells in single-cell spatially resolved transcriptomics data using Diagnostic Evidence GAuge of Single-cells with spatial smoothingabstractSUMMARY: The examination of high-risk cells and regions in tissue samples from spatially resolved transcriptomics platforms offers meaningful insights into specific disease processes. For existing methods, while cell types or clusters can be identified and associated with disease attributes, individual cells are unable to be associated in the same manner. METHOD: Diagnostic Evidence Gauge of Single-Cells and Spatial Transcriptomics (DEGAS) solves the above problem by employing latent representations of gene expression data and domain adaptation to transfer disease attributes from patients to individual cells from single-cell RNA sequencing datasets. In this research, we present and evaluate DEGAS's versatility in adapting to data arising from various single-cell spatially resolved transcriptomics (scSRT) platforms. DEGAS successfully identified high-risk cells and regions in liver hepatocellular carcinoma and skin cutaneous melanoma, which were validated through known markers. Additionally, DEGAS was applied to our newly generated Type II Diabetes Xenium dataset, revealing high-risk cells within the tissue samples. AVAILABILITY AND IMPLEMENTATION: The DEGAS software can be accessed at https://github.com/tsteelejohnson91/DEGAS. For the updated smoothing functions and associated codes, visit https://github.com/dchatter04/DEGAS-Spatial-Smoothing, which is archived at https://doi.org/10.5281/zenodo.18510221. Sources for the datasets reviewed are detailed in their respective sections. A description of some datasets, along with extra tables and figures, is provided in the Supplementary Materials file. Our newly generated Xenium data for Type II Diabetes can be found at https://doi.org/10.7303/syn68699752. Debolina Chatterjee, Justin L. Couetil, Kun Huang 0001, Chao Chen 0012, Jie Zhang 0010, Michael A Kalwat, Travis S. Johnson |
Bioinform. | 6 |
| 2025 | MERGE: Multi-faceted Hierarchical Graph-based GNN for Gene Expression Prediction from Whole Slide Histopathology ImagesabstractRecent advances in Spatial Transcriptomics (ST) pair histology images with spatially resolved gene expression profiles, enabling predictions of gene expression across different tissue locations based on image patches. This opens up new possibilities for enhancing whole slide image (WSI) prediction tasks with localized gene expression. However, existing methods fail to fully leverage the interactions between different tissue locations, which are crucial for accurate joint prediction. To address this, we introduce MERGE (Multi-faceted hiErarchical gRaph for Gene Expressions), which combines a multi-faceted hierarchical graph construction strategy with graph neural networks (GNN) to improve gene expression predictions from WSIs. By clustering tissue image patches based on both spatial and morphological features, and incorporating intra- and inter-cluster edges, our approach fosters interactions between distant tissue locations during GNN learning. As an additional contribution, we evaluate different data smoothing techniques that are necessary to mitigate artifacts in ST data, often caused by technical imperfections. We advocate for adopting gene-aware smoothing methods that are more biologically justified. Experimental results on gene expression prediction show that our GNN method outperforms state-of-the-art techniques across multiple metrics. Aniruddha Ganguly, Debolina Chatterjee, Jie Zhang 0010, Alisa Yurovsky, Travis Steele Johnson, Chao Chen 0012 |
CVPR | 4 |
| 2022 | SPCS: a spatial and pattern combined smoothing method for spatial transcriptomic expressionabstractHigh-dimensional, localized ribonucleic acid (RNA) sequencing is now possible owing to recent developments in spatial transcriptomics (ST). ST is based on highly multiplexed sequence analysis and uses barcodes to match the sequenced reads to their respective tissue locations. ST expression data suffer from high noise and dropout events; however, smoothing techniques have the promise to improve the data interpretability prior to performing downstream analyses. Single-cell RNA sequencing (scRNA-seq) data similarly suffer from these limitations, and smoothing methods developed for scRNA-seq can only utilize associations in transcriptome space (also known as one-factor smoothing methods). Since they do not account for spatial relationships, these one-factor smoothing methods cannot take full advantage of ST data. In this study, we present a novel two-factor smoothing technique, spatial and pattern combined smoothing (SPCS), that employs the k-nearest neighbor (kNN) technique to utilize information from transcriptome and spatial relationships. By performing SPCS on multiple ST slides from pancreatic ductal adenocarcinoma (PDAC), dorsolateral prefrontal cortex (DLPFC) and simulated high-grade serous ovarian cancer (HGSOC) datasets, smoothed ST slides have better separability, partition accuracy and biological interpretability than the ones smoothed by preexisting one-factor methods. Source code of SPCS is provided in Github (https://github.com/Usos/SPCS). Yusong Liu, Tongxin Wang, Ben Duggan, Michael F. Sharpnack, Kun Huang 0001, Jie Zhang 0010, Xiufen Ye, Travis S. Johnson |
Briefings Bioinform. | 6 |
| 2021 | Transfer Learning via Optimal Transportation for Integrative Cancer Patient StratificationabstractThe Stratification of early-stage cancer patients for the prediction of clinical outcome is a challenging task since cancer is associated with various molecular aberrations. A single biomarker often cannot provide sufficient information to stratify early-stage patients effectively. Understanding the complex mechanism behind cancer development calls for exploiting biomarkers from multiple modalities of data such as histopathology images and genomic data. The integrative analysis of these biomarkers sheds light on cancer diagnosis, subtyping, and prognosis. Another difficulty is that labels for early-stage cancer patients are scarce and not reliable enough for predicting survival times. Given the fact that different cancer types share some commonalities, we explore if the knowledge learned from one cancer type can be utilized to improve prognosis accuracy for another cancer type. We propose a novel unsupervised multi-view transfer learning algorithm to simultaneously analyze multiple biomarkers in different cancer types. We integrate multiple views using non-negative matrix factorization and formulate the transfer learning model based on the Optimal Transport theory to align features of different cancer types. We evaluate the stratification performance on three early-stage cancers from the Cancer Genome Atlas (TCGA) project. Comparing with other benchmark methods, our framework achieves superior accuracy for patient outcome prediction. Wei Shao 0005, Jie Zhang 0010, Kun Huang 0001 |
IJCAI | 3 |
| 2021 | TPSC: a module detection method based on topology potential and spectral clustering in weighted networks and its application in gene co-expression module discoveryabstractBACKGROUND: Gene co-expression networks are widely studied in the biomedical field, with algorithms such as WGCNA and lmQCM having been developed to detect co-expressed modules. However, these algorithms have limitations such as insufficient granularity and unbalanced module size, which prevent full acquisition of knowledge from data mining. In addition, it is difficult to incorporate prior knowledge in current co-expression module detection algorithms. RESULTS: In this paper, we propose a novel module detection algorithm based on topology potential and spectral clustering algorithm to detect co-expressed modules in gene co-expression networks. By testing on TCGA data, our novel method can provide more complete coverage of genes, more balanced module size and finer granularity than current methods in detecting modules with significant overall survival difference. In addition, the proposed algorithm can identify modules by incorporating prior knowledge. CONCLUSION: In summary, we developed a method to obtain as much as possible information from networks with increased input coverage and the ability to detect more size-balanced and granular modules. In addition, our method can integrate data from different sources. Our proposed method performs better than current methods with complete coverage of input genes and finer granularity. Moreover, this method is designed not only for gene co-expression networks but can also be applied to any general fully connected weighted network. Yusong Liu, Xiufen Ye, Christina Y. Yu, Wei Shao 0005, Weixing Feng, Jie Zhang 0010, Kun Huang 0001 |
BMC Bioinform. | 7 |
| 2021 | Weakly Supervised Deep Ordinal Cox Model for Survival Prediction From Whole-Slide Pathological ImagesabstractWhole-Slide Histopathology Image (WSI) is generally considered the gold standard for cancer diagnosis and prognosis. Given the large inter-operator variation among pathologists, there is an imperative need to develop machine learning models based on WSIs for consistently predicting patient prognosis. The existing WSI-based prediction methods do not utilize the ordinal ranking loss to train the prognosis model, and thus cannot model the strong ordinal information among different patients in an efficient way. Another challenge is that a WSI is of large size (e.g., 100,000-by-100,000 pixels) with heterogeneous patterns but often only annotated with a single WSI-level label, which further complicates the training process. To address these challenges, we consider the ordinal characteristic of the survival process by adding a ranking-based regularization term on the Cox model and propose a weakly supervised deep ordinal Cox model (BDOCOX) for survival prediction from WSIs. Here, we generate amounts of bags from WSIs, and each bag is comprised of the image patches representing the heterogeneous patterns of WSIs, which is assumed to match the WSI-level labels for training the proposed model. The effectiveness of the proposed method is well validated by theoretical analysis as well as the prognosis and patient stratification results on three cancer datasets from The Cancer Genome Atlas (TCGA). Wei Shao 0005, Tongxin Wang, Zhi Han, Jie Zhang 0010, Kun Huang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Multi-task multi-modal learning for joint diagnosis and prognosis of human cancers
Wei Shao 0005, Tongxin Wang, Liang Sun 0009, Tianhan Dong, Zhi Han, Jie Zhang 0010, Daoqiang Zhang, Kun Huang 0001 |
Medical Image Anal. | 7 |
| 2020 | Integrative Analysis of Pathological Images and Multi-Dimensional Genomic Data for Early-Stage Cancer PrognosisabstractThe integrative analysis of histopathological images and genomic data has received increasing attention for studying the complex mechanisms of driving cancers. However, most image-genomic studies have been restricted to combining histopathological images with the single modality of genomic data (e.g., mRNA transcription or genetic mutation), and thus neglect the fact that the molecular architecture of cancer is manifested at multiple levels, including genetic, epigenetic, transcriptional, and post-transcriptional events. To address this issue, we propose a novel ordinal multi-modal feature selection (OMMFS) framework that can simultaneously identify important features from both pathological images and multi-modal genomic data (i.e., mRNA transcription, copy number variation, and DNA methylation data) for the prognosis of cancer patients. Our model is based on a generalized sparse canonical correlation analysis framework, by which we also take advantage of the ordinal survival information among different patients for survival outcome prediction. We evaluate our method on three early-stage cancer datasets derived from The Cancer Genome Atlas (TCGA) project, and the experimental results demonstrated that both the selected image and multi-modal genomic markers are strongly correlated with survival enabling effective stratification of patients with distinct survival than the comparing methods, which is often difficult for early-stage cancer patients. Wei Shao 0005, Kun Huang 0001, Zhi Han, Jun Cheng 0006, Tongxin Wang, Liang Sun 0009, Zixiao Lu, Jie Zhang 0010, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 9 |
| 2019 | LAmbDA: label ambiguous domain adaptation dataset integration reduces batch effects and improves subtype detectionabstractMOTIVATION: Rapid advances in single cell RNA sequencing (scRNA-seq) have produced higher-resolution cellular subtypes in multiple tissues and species. Methods are increasingly needed across datasets and species to (i) remove systematic biases, (ii) model multiple datasets with ambiguous labels and (iii) classify cells and map cell type labels. However, most methods only address one of these problems on broad cell types or simulated data using a single model type. It is also important to address higher-resolution cellular subtypes, subtype labels from multiple datasets, models trained on multiple datasets simultaneously and generalizability beyond a single model type. RESULTS: We developed a species- and dataset-independent transfer learning framework (LAmbDA) to train models on multiple datasets (even from different species) and applied our framework on simulated, pancreas and brain scRNA-seq experiments. These models mapped corresponding cell types between datasets with inconsistent cell subtype labels while simultaneously reducing batch effects. We achieved high accuracy in labeling cellular subtypes (weighted accuracy simulated 1 datasets: 90%; simulated 2 datasets: 94%; pancreas datasets: 88% and brain datasets: 66%) using LAmbDA Feedforward 1 Layer Neural Network with bagging. This method achieved higher weighted accuracy in labeling cellular subtypes than two other state-of-the-art methods, scmap and CaSTLe in brain (66% versus 60% and 32%). Furthermore, it achieved better performance in correctly predicting ambiguous cellular subtype labels across datasets in 88% of test cases compared with CaSTLe (63%), scmap (50%) and MetaNeighbor (50%). LAmbDA is model- and dataset-independent and generalizable to diverse data types representing an advance in biocomputing. AVAILABILITY AND IMPLEMENTATION: github.com/tsteelejohnson91/LAmbDA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Travis S. Johnson, Tongxin Wang, Christina Y. Yu, Yatong Han, Kun Huang 0001, Jie Zhang 0010 |
Bioinform. | 9 |
| 2019 | Generalized gene co-expression analysis via subspace clustering using low-rank representationabstractBACKGROUND: Gene Co-expression Network Analysis (GCNA) helps identify gene modules with potential biological functions and has become a popular method in bioinformatics and biomedical research. However, most current GCNA algorithms use correlation to build gene co-expression networks and identify modules with highly correlated genes. There is a need to look beyond correlation and identify gene modules using other similarity measures for finding novel biologically meaningful modules. RESULTS: We propose a new generalized gene co-expression analysis algorithm via subspace clustering that can identify biologically meaningful gene co-expression modules with genes that are not all highly correlated. We use low-rank representation to construct gene co-expression networks and local maximal quasi-clique merger to identify gene co-expression modules. We applied our method on three large microarray datasets and a single-cell RNA sequencing dataset. We demonstrate that our method can identify gene modules with different biological functions than current GCNA methods and find gene modules with prognostic values. CONCLUSIONS: The presented method takes advantage of subspace clustering to generate gene co-expression networks rather than using correlation as the similarity measure between genes. Our generalized GCNA method can provide new insights from gene expression datasets and serve as a complement to current GCNA algorithms. Tongxin Wang, Jie Zhang 0010, Kun Huang 0001 |
BMC Bioinform. | 2 |
| 2018 | Genetic Mutations Associated with Histopathology Changes in Kidney Cancer
Jun Cheng 0006, Zhi Han, Qianjin Feng 0002, Jie Zhang 0010, Kun Huang 0001 |
AMIA | 5 |
| 2012 | Weighted Frequent Gene Co-expression Network Mining to Identify Genes Involved in Genome StabilityabstractGene co-expression network analysis is an effective method for predicting gene functions and disease biomarkers. However, few studies have systematically identified co-expressed genes involved in the molecular origin and development of various types of tumors. In this study, we used a network mining algorithm to identify tightly connected gene co-expression networks that are frequently present in microarray datasets from 33 types of cancer which were derived from 16 organs/tissues. We compared the results with networks found in multiple normal tissue types and discovered 18 tightly connected frequent networks in cancers, with highly enriched functions on cancer-related activities. Most networks identified also formed physically interacting networks. In contrast, only 6 networks were found in normal tissues, which were highly enriched for housekeeping functions. The largest cancer network contained many genes with genome stability maintenance functions. We tested 13 selected genes from this network for their involvement in genome maintenance using two cell-based assays. Among them, 10 were shown to be involved in either homology-directed DNA repair or centrosome duplication control including the well-known cancer marker MKI67. Our results suggest that the commonly recognized characteristics of cancers are supported by highly coordinated transcriptomic activities. This study also demonstrated that the co-expression network directed approach provides a powerful tool for understanding cancer physiology, predicting new gene functions, as well as providing new target candidates for cancer therapeutics. Jie Zhang 0010, Kewei Lu, Yang Xiang 0007, Muhtadi Islam, Shweta Kotian, Zeina Kais, Cindy Lee, Mansi Arora, Hui-wen Liu, Jeffrey D. Parvin, Kun Huang 0001 |
PLoS Comput. Biol. | 1 |
| 2010 | Multi-dimensional discovery of biomarker and phenotype complexesabstractBACKGROUND: Given the rapid growth of translational research and personalized healthcare paradigms, the ability to relate and reason upon networks of bio-molecular and phenotypic variables at various levels of granularity in order to diagnose, stage and plan treatments for disease states is highly desirable. Numerous techniques exist that can be used to develop networks of co-expressed or otherwise related genes and clinical features. Such techniques can also be used to create formalized knowledge collections based upon the information incumbent to ontologies and domain literature. However, reports of integrative approaches that bridge such networks to create systems-level models of disease or wellness are notably lacking in the contemporary literature. RESULTS: In response to the preceding gap in knowledge and practice, we report upon a prototypical series of experiments that utilize multi-modal approaches to network induction. These experiments are intended to elicit meaningful and significant biomarker-phenotype complexes spanning multiple levels of granularity. This work has been performed in the experimental context of a large-scale clinical and basic science data repository maintained by the National Cancer Institute (NCI) funded Chronic Lymphocytic Leukemia Research Consortium. CONCLUSIONS: Our results indicate that it is computationally tractable to link orthogonal networks of genes, clinical features, and conceptual knowledge to create multi-dimensional models of interrelated biomarkers and phenotypes. Further, our results indicate that such systems-level models contain interrelated bio-molecular and clinical markers capable of supporting hypothesis discovery and testing. Based on such findings, we propose a conceptual model intended to inform the cross-linkage of the results of such methods. This model has as its aim the identification of novel and knowledge-anchored biomarker-phenotype complexes. Philip R. O. Payne, Kun Huang 0001, Kristin Keen-Circle, Abhisek Kundu, Jie Zhang 0010, Tara Borlawsky |
BMC Bioinform. | 5 |
| 2010 | Using gene co-expression network analysis to predict biomarkers for chronic lymphocytic leukemiaabstractBACKGROUND: Chronic lymphocytic leukemia (CLL) is the most common adult leukemia. It is a highly heterogeneous disease, and can be divided roughly into indolent and progressive stages based on classic clinical markers. Immunoglobin heavy chain variable region (IgVH) mutational status was found to be associated with patient survival outcome, and biomarkers linked to the IgVH status has been a focus in the CLL prognosis research field. However, biomarkers highly correlated with IgVH mutational status which can accurately predict the survival outcome are yet to be discovered. RESULTS: In this paper, we investigate the use of gene co-expression network analysis to identify potential biomarkers for CLL. Specifically we focused on the co-expression network involving ZAP70, a well characterized biomarker for CLL. We selected 23 microarray datasets corresponding to multiple types of cancer from the Gene Expression Omnibus (GEO) and used the frequent network mining algorithm CODENSE to identify highly connected gene co-expression networks spanning the entire genome, then evaluated the genes in the co-expression network in which ZAP70 is involved. We then applied a set of feature selection methods to further select genes which are capable of predicting IgVH mutation status from the ZAP70 co-expression network. CONCLUSIONS: We have identified a set of genes that are potential CLL prognostic biomarkers IL2RB, CD8A, CD247, LAG3 and KLRK1, which can predict CLL patient IgVH mutational status with high accuracies. Their prognostic capabilities were cross-validated by applying these biomarker candidates to classify patients into different outcome groups using a CLL microarray datasets with clinical information. Jie Zhang 0010, Yang Xiang 0007, Liya Ding 0001, Kristin Keen-Circle, Tara Borlawsky, Hatice Gulcin Ozer, Ruoming Jin, Philip R. O. Payne, Kun Huang 0001 |
BMC Bioinform. | 1 |