VLDB 2026 Research / reviewers in the wild / expert
Xuekui Zhang
dblp:133/2814
· DBLP profile ↗
17ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0003-4728-2343ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | scGDCF: Graphical Deep Clustering With Fused Common Information for Single-Cell RNA-Seq DataabstractUnsupervised deep clustering plays a crucial role in analyzing single-cell RNA sequencing data (scRNA-seq) as it helps to identify potential cell types. However, most existing clustering methods face challenges in effectively fusing common information between feature and topological structure information, and they may not perform well on the sparse data, which are common in single-cell analysis. To address these challenges, we propose a novel Graphical Deep Clustering with Fused Common Information method for scRNA-seq data, named scGDCF. This method can accurately segregate different cell types even in large and sparse scRNA-seq datasets. In scGDCF, we first introduce a sparse feature representation method that utilizes an adversarial loss function to address the sparsity problem in scRNA-seq data and improve the performance of the discriminator. Next, we design a mutual information extracting operator to deeply mine and fuse the common information from feature and topological structure data, thereby improving the clustering performance. Furthermore, we incorporate the varying degrees of contribution from different neighbor nodes and information sources. To handle this, we promote a dual adaptive attention mechanism that operates at both global and local levels. Finally, experiments on seven real-world datasets and two simulated datasets show scGDCF outperforms 17 state-of-the-art methods. We further extend the clustering results for visualization, analysis of gene differential expression and enrichment, showing scGDCF provides novel insights into cell developmental lineages and preserved inter-cluster distances. Kunyu Li, Yongfeng Dong, Ziyu Ren, Jiaxue Zhang, Yushan Hu, Xuekui Zhang |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2026 | MOH: A Novel Multilayer Multi-Omics Heterogeneous Graph for Single-Cell ClusteringabstractCell clustering is crucial in single-cell multi-omics research for identifying distinct cellular populations. Although there has been progress in integrating multi-omics data for clustering, combining more than two types of omics data remains challenging due to the diversity and heterogeneity of these datasets. Traditional approaches typically use heterogeneous graphs that integrate only two types of omics data, constructing graphs with genes and cells as nodes and a single type of edge representing their relationships. However, this method has limitations as it overlooks cell-cell interactions and struggles to capture complex cellular dynamics. Additionally, the graph structure must be redesigned whenever new omics data are introduced, limiting the scalability of these models. To address these issues, we introduce MOH, a novel single-cell clustering algorithm based on a multilayer multi-omics heterogeneous graph. MOH integrates three key single-cell omics types: scRNA-seq, scATAC-seq, and spatial transcriptomics. It constructs a multilayer heterogeneous graph to simultaneously extract and enhance representations from all three omics layers, incorporating both intra-layer and inter-layer edges to capture association and similarity relationships. This enriched representation leads to an accurate clustering results. Extensive experiments show that MOH outperforms six state-of-the-art methods on unsupervised clustering metrics, offering a precise and comprehensive analysis with consistent improvements across all evaluation criteria. Moreover, downstream analyses validate the results, revealing novel biological insights into immune disorder complications in cancer, cancer drug repurposing, and new signaling pathways, which merit further investigation and validation. Yushan Hu, Xiaowen Cao 0002, Yongfeng Dong, Xuekui Zhang |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Decoupled Graph Neural Networks with Hybrid Data Augmentation
Zongyang Li, Xuekui Zhang |
ICIC (21) | 4 |
| 2025 | scSorterDL: a deep neural network-enhanced ensemble LDAs for single cell classificationsabstractThe emergence of single-cell RNA sequencing (scRNA-seq) technology has transformed our understanding of cellular diversity, yet it presents notable challenges for cell type annotation due to data's high dimensionality and sparsity. To tackle these issues, we present scSorterDL, an innovative approach that combines penalized Linear Discriminant Analysis (pLDA), swarm learning, and deep neural networks (DNNs) to improve cell type classification. In scSorterDL, we generate numerous random subsets of the data and apply pLDA models to each subset to capture varied data aspects. The model outputs are then consolidated using a DNN that identifies complex relationships among the pLDA scores, enhancing classification accuracy by considering interactions that simpler methods might overlook. Utilizing GPU computing for both swarm learning and deep learning, scSorterDL adeptly manages large datasets and high-dimensional gene expression data. We tested scSorterDL on 13 real scRNA-seq datasets from diverse species, tissues, and platforms, as well as on 20 pairs of cross-platform datasets. Our method surpassed nine current cell annotation tools in both accuracy and robustness, indicating exceptional performance in both cross-validation and cross-platform contexts. These findings underscore the potential of scSorterDL as an effective and adaptable tool for automated cell type annotation in scRNA-seq research. The code is available on GitHub: https://github.com/kellen8hao/scSorterDL. Kailun Bai, Belaid Moa, Xiaojian Shao, Xuekui Zhang |
Briefings Bioinform. | 4 |
| 2025 | HiCat: a semi-supervised approach for cell type annotationabstractExisting cell type annotation methods face significant hurdles: supervised approaches often fail to differentiate between novel cell types not present in reference data, while unsupervised techniques can suffer from cluster impurity and difficulties in robustly distinguishing multiple distinct unknown cell populations. This critical gap motivated the development of HiCat, a semi-supervised pipeline specifically designed to overcome these limitations. HiCat is a semi-supervised pipeline that integrates both approaches, leveraging reference (labeled) and query (unlabeled) genomic data to simultaneously enhance annotation accuracy for known cell types and improve the discovery and differentiation of novel ones. HiCat follows a structured pipeline: (1) removing batch effects and generate a low-dimensional embedding; (2) nonlinear dimensionality reduction for capturing key patterns; (3) unsupervised clustering for proposing novel cell type candidates; (4) merging multi-resolution features from previous steps into a condensed feature space; (5) training a classifier on reference data for supervised annotation; and (6) resolving inconsistencies between supervised predictions and unsupervised clusters to finalize annotations, particularly for unseen types. Performance was evaluated across 10 public genomic datasets and perform a case study on a molecular cell atlas of the human lung. HiCat demonstrated superior performance in both known cell type classification and novel cell type identification. In benchmark evaluations, HiCat consistently outperformed existing methods, critically excelling in identifying and distinguishing multiple novel cell types. HiCat presents a robust framework for scRNA-seq cell annotation, improving classification accuracy and novel type identification. In addition, it provides a scalable and transferable solution for biomedical research, directly addressing key challenges in automated cell annotation. Chang Bi, Kailun Bai, Xuekui Zhang |
Briefings Bioinform. | 3 |
| 2025 | scAGCI: an anchor graph-based method for cell clustering from integrated scRNA-seq and scATAC-seq dataabstractSingle-cell multi-omics clustering confronts noise and heterogeneity barriers. Current multi-view anchor graph approaches, though successful in noise reduction, inadequately model higher order feature interactions. To address this issue, we propose scAGCI, a cell clustering method based on anchor graphs that integrates both scRNA-seq and scATAC-seq data. Our method captures specific and shared anchor graphs representing the properties of omics data in the process of dynamic anchor unification, and mines high-order shared information to complete the omics representation. Subsequently, clustering results are obtained by integrating the specific and shared omics representation. Benchmarking against 13 state-of-the-art methods confirms scAGCI's superior clustering performance and computational efficiency in cell-type identification and subtype resolution. The method preserves biologically meaningful omics patterns, as evidenced by marker gene enrichment and functional analyses, establishing it as a robust tool for elucidating cellular heterogeneity in single-cell multi-omics data. Jiaxue Zhang, Yushan Hu, Xiaowen Cao 0002, Yongfeng Dong, Xuekui Zhang |
Briefings Bioinform. | 7 |
| 2025 | Novel machine learning model for predicting cancer drugs' susceptibilities and discovering novel treatmentsabstractBACKGROUND AND OBJECTIVE: Timely treatment is crucial for cancer patients, so it's important to administer the appropriate treatment as soon as possible. Because individuals can respond differently to a given drug due to their unique genomic profiles, we aim to use their genomic information to predict how various drugs will affect them and determine the best course of treatment. METHODS: We present Kernelized Residual Stacking (KRS), a new multi-task learning approach, and use it to predict the responses to anti-cancer drugs based on genomic data. We demonstrate the superior predictive performance of KRS, outperforming popular competitors, by utilizing the Genomics of Drug Sensitivity in Cancer (GDSC) study and the Cancer Cell Line Encyclopedia (CCLE) study. Downstream analysis of feature genes selected by KRS is conducted to discover novel therapies. RESULTS: We used two genomic studies to show that KRS outperforms a few popular competitors in predicting drugs' susceptibilities. Through downstream analysis of feature genes selected by KRS, we found that the PI3K-Akt pathway could alter drugs' susceptibilities, and its expression correlated positively with the hub gene ERBB2. We discovered eight novel small molecules based on these feature genes, which could be developed into novel combination therapies with anti-cancer drugs. CONCLUSIONS: KRS outperforms competitors in prediction performance and selects feature genes highly correlated with drugs' susceptibilities. Novel biological results are found by investigating KRS's feature genes. Xiaowen Cao 0002, Yushan Hu, Junhua Gu, Xuekui Zhang |
J. Biomed. Informatics | 9 |
| 2024 | Diagnostics of viral infections using high-throughput genome sequencing dataabstractPlant viral infections cause significant economic losses, totalling $350 billion USD in 2021. With no treatment for virus-infected plants, accurate and efficient diagnosis is crucial to preventing and controlling these diseases. High-throughput sequencing (HTS) enables cost-efficient identification of known and unknown viruses. However, existing diagnostic pipelines face challenges. First, many methods depend on subjectively chosen parameter values, undermining their robustness across various data sources. Second, artifacts (e.g. false peaks) in the mapped sequence data can lead to incorrect diagnostic results. While some methods require manual or subjective verification to address these artifacts, others overlook them entirely, affecting the overall method performance and leading to imprecise or labour-intensive outcomes. To address these challenges, we introduce IIMI, a new automated analysis pipeline using machine learning to diagnose infections from 1583 plant viruses with HTS data. It adopts a data-driven approach for parameter selection, reducing subjectivity, and automatically filters out regions affected by artifacts, thus improving accuracy. Testing with in-house and published data shows IIMI's superiority over existing methods. Besides a prediction model, IIMI also provides resources on plant virus genomes, including annotations of regions prone to artifacts. The method is available as an R package (iimi) on CRAN and will integrate with the web application www.virtool.ca, enhancing accessibility and user convenience. Haochen Ning, Ian Boyes, Ibrahim Numanagic, Michael Rott, Xuekui Zhang |
Briefings Bioinform. | 6 |
| 2024 | A Semi-parametric Estimation of Personalized Dose-response Function Using Instrumental VariablesabstractIn the application of instrumental variable analysis that conducts causal inference in the presence of unmeasured confounding, invalid instrumental variables and weak instrumental variables often exist which complicate the analysis. In this paper, we propose a model-free dimension reduction procedure to select the invalid instrumental variables and refine them into lower-dimensional linear combinations. The procedure also combines the weak instrumental variables into a few stronger instrumental variables that best condense their information. We then introduce the personalized dose-response function that incorporates the subject's personal characteristics into the conventional dose-response function, and use the reduced data from dimension reduction to propose a novel and easily implementable nonparametric estimator of this function. The proposed approach is suitable for both discrete and continuous treatment variables, and is robust to the dimensionality of data. Its effectiveness is illustrated by the simulation studies and the data analysis of ADNI-DoD study, where the causal relationship between depression and dementia is investigated. Yeying Zhu, Xuekui Zhang |
J. Mach. Learn. Res. | 3 |
| 2024 | Essential Number of Principal Components and Nearly Training-Free Model for Spectral AnalysisabstractLearning-enabled spectroscopic analysis, promising for automated real-time analysis of chemicals, is facing several challenges. First, a typical machine learning model requires a large number of training samples that physical systems can not provide. Second, it requires the testing samples to be in range with the training samples, which often is not the case in the real world. Further, a spectroscopy device is limited by its memory size, computing power, and battery capacity. That requires highly efficient learning models for on-site analysis. In this paper, by analyzing multi-gas mixtures and multi-molecule suspensions, we first show that orders of magnitude reduction of data dimension can be achieved as the number of principal components that need to be retained is the same as the independent constituents in the mixture. From this principle, we designed highly compact models in which the essential principal components can be directly extracted from the interrelations between the individual chemical properties and principal components; and only a few training samples are required. Our model can predict the constituent concentrations that have not been seen in the training dataset and provide estimations of measurement noises. This approach can be extended as an effectively standardized method for principle component extraction. Yifeng Bie, Shuai You, Xuekui Zhang, Tao Lu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | ADSP: An adaptive sample pooling strategy for diagnostic testing
Xuekui Zhang, Xiaolin Huang |
J. Biomed. Informatics | 1 |
| 2022 | Handling Sparse Longitudinal Data with Irregular Missing Data - Analysis of Fecal Coliform Bacteria DataabstractFecal coliform bacteria are commonly used as an indicator to reflect the fecal contamination level in the water. To manage a healthy and thriving aquaculture industry, the Canadian Shellfish Sanitation Program (CSSP) was established in 1948. As a part of the CSSP mandate, fecal coliform bacteria levels in shellfish growing habitat have been monitored at nearly 15, 000 shellfish harvesting sites across the six coastal provinces of Canada over 40 years (1980–2019). The irregular sparseness in the measurement data presented a critical challenge for reliable analysis of fecal contamination patterns along Canada's coastline. This paper illustrates a preprocessing approach to handle the irregular sparseness in the measurement data of fecal coliform bacteria levels and demonstrates the effectiveness of data filtering, pooling, binning, partitioning, and the application of functional principal component analysis. We managed to transform the irregularly sparse measurements at a site into surrogate variables without missing data that represent the differences in the site's contamination amplitude and seasonal variation from the average. The surrogate variables will be used for downstream analyses, such as associating contamination with climate change. Shuai You, Xiaolin Huang, Youlian Pan, Xuekui Zhang |
CIBCB | 5 |
| 2022 | cSurvival: a web resource for biomarker interactions in cancer outcomes and in cell linesabstractSurvival analysis is a technique for identifying prognostic biomarkers and genetic vulnerabilities in cancer studies. Large-scale consortium-based projects have profiled >11 000 adult and >4000 pediatric tumor cases with clinical outcomes and multiomics approaches. This provides a resource for investigating molecular-level cancer etiologies using clinical correlations. Although cancers often arise from multiple genetic vulnerabilities and have deregulated gene sets (GSs), existing survival analysis protocols can report only on individual genes. Additionally, there is no systematic method to connect clinical outcomes with experimental (cell line) data. To address these gaps, we developed cSurvival (https://tau.cmmt.ubc.ca/cSurvival). cSurvival provides a user-adjustable analytical pipeline with a curated, integrated database and offers three main advances: (i) joint analysis with two genomic predictors to identify interacting biomarkers, including new algorithms to identify optimal cutoffs for two continuous predictors; (ii) survival analysis not only at the gene, but also the GS level; and (iii) integration of clinical and experimental cell line studies to generate synergistic biological insights. To demonstrate these advances, we report three case studies. We confirmed findings of autophagy-dependent survival in colorectal cancers and of synergistic negative effects between high expression of SLC7A11 and SLC2A1 on outcomes in several cancers. We further used cSurvival to identify high expression of the Nrf2-antioxidant response element pathway as a main indicator for lung cancer prognosis and for cellular resistance to oxidative stress-inducing drugs. Altogether, these analyses demonstrate cSurvival's ability to support biomarker prognosis and interaction analysis via gene- and GS-level approaches and to integrate clinical and experimental biomedical studies. Xuanjin Cheng, Yongxing Liu, Andrew Gordon Robertson, Xuekui Zhang, Steven J. M. Jones, Stefan Taubert |
Briefings Bioinform. | 6 |
| 2022 | Integrative COVID-19 biological network inference with probabilistic core decompositionabstractThe severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) is responsible for millions of deaths around the world. To help contribute to the understanding of crucial knowledge and to further generate new hypotheses relevant to SARS-CoV-2 and human protein interactions, we make use of the information abundant Biomine probabilistic database and extend the experimentally identified SARS-CoV-2-human protein-protein interaction (PPI) network in silico. We generate an extended network by integrating information from the Biomine database, the PPI network and other experimentally validated results. To generate novel hypotheses, we focus on the high-connectivity sub-communities that overlap most with the integrated experimentally validated results in the extended network. Therefore, we propose a new data analysis pipeline that can efficiently compute core decomposition on the extended network and identify dense subgraphs. We then evaluate the identified dense subgraph and the generated hypotheses in three contexts: literature validation for uncovered virus targeting genes and proteins, gene function enrichment analysis on subgraphs and literature support on drug repurposing for identified tissues and diseases related to COVID-19. The major types of the generated hypotheses are proteins with their encoding genes and we rank them by sorting their connections to the integrated experimentally validated nodes. In addition, we compile a comprehensive list of novel genes, and proteins potentially related to COVID-19, as well as novel diseases which might be comorbidities. Together with the generated hypotheses, our results provide novel knowledge relevant to COVID-19 for further validation. Yang Guo 0002, Fatemeh Esfahani, Xiaojian Shao, S. Venkatesh 0001, Alex Thomo, Xuekui Zhang |
Briefings Bioinform. | 7 |
| 2021 | Multi-stage graph peeling algorithm for probabilistic core decompositionabstractMining dense subgraphs where vertices connect closely with each other is a common task when analyzing graphs. A very popular notion in subgraph analysis is core decomposition. Recently, Esfahani et al. presented a probabilistic core decomposition algorithm based on graph peeling and Central Limit Theorem (CLT) that is capable of handling very large graphs. Their proposed peeling algorithm (PA) starts from the lowest degree vertices and recursively deletes these vertices, assigning core numbers, and updating the degree of neighbour vertices until it reached the maximum core. However, in many applications, particularly in biology, more valuable information can be obtained from dense sub-communities and we are not interested in small cores where vertices do not interact much with others. To make the previous PA focus more on dense subgraphs, we propose a multi-stage graph peeling algorithm (M-PA) that has a two-stage data screening procedure added before the previous PA. After removing vertices from the graph based on the user-defined thresholds, we can reduce the graph complexity largely and without affecting the vertices in subgraphs that we are interested in. We show that M-PA is more efficient than the previous PA and with the properly set filtering threshold, can produce very similar if not identical dense subgraphs to the previous PA (in terms of graph density and clustering coefficient). Yang Guo 0002, Xuekui Zhang, Fatemeh Esfahani, S. Venkatesh 0001, Alex Thomo |
ASONAM | 2 |
| 2020 | Simultaneous prediction of multiple outcomes using revised stacking algorithmsabstractMOTIVATION: HIV is difficult to treat because its virus mutates at a high rate and mutated viruses easily develop resistance to existing drugs. If the relationships between mutations and drug resistances can be determined from historical data, patients can be provided personalized treatment according to their own mutation information. The HIV Drug Resistance Database was built to investigate the relationships. Our goal is to build a model using data in this database, which simultaneously predicts the resistance of multiple drugs using mutation information from sequences of viruses for any new patient. RESULTS: We propose two variations of a stacking algorithm which borrow information among multiple prediction tasks to improve multivariate prediction performance. The most attractive feature of our proposed methods is the flexibility with which complex multivariate prediction models can be constructed using any univariate prediction models. Using cross-validation studies, we show that our proposed methods outperform other popular multivariate prediction methods. AVAILABILITY AND IMPLEMENTATION: An R package is being developed. In the meantime, R code can be requested by email. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mary Lesperance, Xuekui Zhang |
Bioinform. | 3 |
| 2013 | PING 2.0: an R/Bioconductor package for nucleosome positioning using next-generation sequencing dataabstractSUMMARY: MNase-Seq and ChIP-Seq have evolved as popular techniques to study chromatin and histone modification. Although many tools have been developed to identify enriched regions, software tools for nucleosome positioning are still limited. We introduce a flexible and powerful open-source R package, PING 2.0, for nucleosome positioning using MNase-Seq data or MNase- or sonicated- ChIP-Seq data combined with either single-end or paired-end sequencing. PING uses a model-based approach, which enables nucleosome predictions even in the presence of low read counts. We illustrate PING using two paired-end datasets from Saccharomyces cerevisiae and compare its performance with nucleR and ChIPseqR. AVAILABILITY: PING 2.0 is available from the Bioconductor website at http://bioconductor.org. It can run on Linux, Mac and Windows. Sangsoon Woo, Xuekui Zhang, Renan Sauteraud, François Robert 0001, Raphael Gottardo |
Bioinform. | 2 |