Xuhua Yan

dblp:308/8230 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-3183-3342ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2026 DIMIR: Deep Incomplete Multi-view Information Recovery for Breast Cancer Subtype Classification
Wei Lan 0001, Yinghao Liu, Xuhua Yan, Qingfeng Chen, Liangliang Liu 0001, Min Li 0007, Yi Pan 0001
ISBRA (1)3
2026 SRLST: a unified multimodal representation learning framework for spatial transcriptomics analysis
abstract
MOTIVATION: Spatial transcriptomics (ST) enables molecular profiling within native tissue architecture, yet accurate delineation of spatial domains in ST data is challenging, as it demands the coordinated integration of transcriptomic, spatial, and tissue histological information. RESULTS: We present SRLST, an unsupervised representation learning framework that holistically harmonize these three complementary data modalities to precisely uncover tissue organization. SRLST employs a dual-graph variational autoencoding strategy to jointly model spatial proximity and morphological relations, fusing these with gene-expression embeddings into a unified latent space. Across distinct experimental datasets, SRLST consistently outperforms existing methods in delineating cortical organization, identifying small discontinuous tissue compartments, and capturing complex intratumor heterogeneity. AVAILABILITY AND IMPLEMENTATION: The code implementation of the SRLST algorithm is available at https://github.com/lanbiolab/SRLST.
Wei Lan 0001, Tongsheng Ling, Guohang He, Xuhua Yan, Ruiqing Zheng, Min Li 0007, Shirui Pan, Yi Pan 0001
Bioinform.5
2023 Bubble: a fast single-cell RNA-seq imputation using an autoencoder constrained by bulk RNA-seq data
abstract
Single-cell RNA-sequencing technology (scRNA-seq) brings research to single-cell resolution. However, a major drawback of scRNA-seq is large sparsity, i.e. expressed genes with no reads due to technical noise or limited sequence depth during the scRNA-seq protocol. This phenomenon is also called 'dropout' events, which likely affect downstream analyses such as differential expression analysis, the clustering and visualization of cell subpopulations, cellular trajectory inference, etc. Therefore, there is a need to develop a method to identify and impute these dropout events. We propose Bubble, which first identifies dropout events from all zeros based on expression rate and coefficient of variation of genes within cell subpopulation, and then leverages an autoencoder constrained by bulk RNA-seq data to only impute those values. Unlike other deep learning-based imputation methods, Bubble fuses the matched bulk RNA-seq data as a constraint to reduce the introduction of false positive signals. Using simulated and several real scRNA-seq datasets, we demonstrate that Bubble enhances the recovery of missing values, gene-to-gene and cell-to-cell correlations, and reduces the introduction of false positive signals. Regarding some crucial downstream analyses of scRNA-seq data, Bubble facilitates the identification of differentially expressed genes, improves the performance of clustering and visualization, and aids the construction of cellular trajectory. More importantly, Bubble provides fast and scalable imputation with minimal memory usage.
Xuhua Yan, Ruiqing Zheng, Min Li 0007
Briefings Bioinform.2
2023 scNCL: transferring labels from scRNA-seq to scATAC-seq data with neighborhood contrastive regularization
abstract
MOTIVATION: scATAC-seq has enabled chromatin accessibility landscape profiling at the single-cell level, providing opportunities for determining cell-type-specific regulation codes. However, high dimension, extreme sparsity, and large scale of scATAC-seq data have posed great challenges to cell-type identification. Thus, there has been a growing interest in leveraging the well-annotated scRNA-seq data to help annotate scATAC-seq data. However, substantial computational obstacles remain to transfer information from scRNA-seq to scATAC-seq, especially for their heterogeneous features. RESULTS: We propose a new transfer learning method, scNCL, which utilizes prior knowledge and contrastive learning to tackle the problem of heterogeneous features. Briefly, scNCL transforms scATAC-seq features into gene activity matrix based on prior knowledge. Since feature transformation can cause information loss, scNCL introduces neighborhood contrastive learning to preserve the neighborhood structure of scATAC-seq cells in raw feature space. To learn transferable latent features, scNCL uses a feature projection loss and an alignment loss to harmonize embeddings between scRNA-seq and scATAC-seq. Experiments on various datasets demonstrated that scNCL not only realizes accurate and robust label transfer for common types, but also achieves reliable detection of novel types. scNCL is also computationally efficient and scalable to million-scale datasets. Moreover, we prove scNCL can help refine cell-type annotations in existing scATAC-seq atlases. AVAILABILITY AND IMPLEMENTATION: The source code and data used in this paper can be found in https://github.com/CSUBioGroup/scNCL-release.
Xuhua Yan, Ruiqing Zheng, Jinmiao Chen, Min Li 0007
Bioinform.1
2023 CLAIRE: contrastive learning-based batch correction framework for better balance between batch mixing and preservation of cellular heterogeneity
abstract
MOTIVATION: Integration of growing single-cell RNA sequencing datasets helps better understand cellular identity and function. The major challenge for integration is removing batch effects while preserving biological heterogeneities. Advances in contrastive learning have inspired several contrastive learning-based batch correction methods. However, existing contrastive-learning-based methods exhibit noticeable ad hoc trade-off between batch mixing and preservation of cellular heterogeneities (mix-heterogeneity trade-off). Therefore, a deliberate mix-heterogeneity trade-off is expected to yield considerable improvements in scRNA-seq dataset integration. RESULTS: We develop a novel contrastive learning-based batch correction framework, CIAIRE, which achieves superior mix-heterogeneity trade-off. The key contributions of CLAIRE are proposal of two complementary strategies: construction strategy and refinement strategy, to improve the appropriateness of positive pairs. Construction strategy dynamically generates positive pairs by augmenting inter-batch mutual nearest neighbors (MNN) with intra-batch k-nearest neighbors (KNN), which improves the coverage of positive pairs for the whole distribution of shared cell types between batches. Refinement strategy aims to automatically reduce the potential false positive pairs from the construction strategy, which resorts to the memory effect of deep neural networks. We demonstrate that CLAIRE possesses superior mix-heterogeneity trade-off over existing contrastive learning-based methods. Benchmark results on six real datasets also show that CLAIRE achieves the best integration performance against eight state-of-the-art methods. Finally, comprehensive experiments are conducted to validate the effectiveness of CLAIRE. AVAILABILITY AND IMPLEMENTATION: The source code and data used in this study can be found in https://github.com/CSUBioGroup/CLAIRE-release. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xuhua Yan, Ruiqing Zheng, Fang-Xiang Wu, Min Li 0007
Bioinform.1
2022 GLOBE: a contrastive learning-based framework for integrating single-cell transcriptome datasets
abstract
Integration of single-cell transcriptome datasets from multiple sources plays an important role in investigating complex biological systems. The key to integration of transcriptome datasets is batch effect removal. Recent methods attempt to apply a contrastive learning strategy to correct batch effects. Despite their encouraging performance, the optimal contrastive learning framework for batch effect removal is still under exploration. We develop an improved contrastive learning-based batch correction framework, GLOBE. GLOBE defines adaptive translation transformations for each cell to guarantee the stability of approximating batch effects. To enhance the consistency of representations alignment, GLOBE utilizes a loss function that is both hardness-aware and consistency-aware to learn batch effect-invariant representations. Moreover, GLOBE computes batch-corrected gene matrix in a transparent approach to support diverse downstream analysis. Benchmarking results on a wide spectrum of datasets show that GLOBE outperforms other state-of-the-art methods in terms of robust batch mixing and superior conservation of biological signals. We further apply GLOBE to integrate two developing mouse neocortex datasets and show GLOBE succeeds in removing batch effects while preserving the contiguous structure of cells in raw data. Finally, a comprehensive study is conducted to validate the effectiveness of GLOBE.
Xuhua Yan, Ruiqing Zheng, Min Li 0007
Briefings Bioinform.1
2021 DeepCI: a deep learning based clustering method for single cell RNA-seq data
abstract
Single cell RNA sequencing enables researchers to analyze cellular heterogeneity at high resolution. In the cellular heterogeneity analysis, unsupervised clustering has been a common and powerful way to identify cell types. Nevertheless, the high dropout rate and high dimension of scRNA-seq data make it still a challenging task. In this study, we proposed DeepCI, a deep neural network based single cell clustering method, which simultaneously accomplishes low-dimensional representation learning and clustering with implicit imputation of scRNA-seq data. Tested on real datasets, DeepCI obtained overall better clustering and visualization performance than several state-of-the-art approaches.
Zhenlan Liang, Ruiqing Zheng, Xuhua Yan, Min Li 0007
BIBM4
2021 MKG: a mutual information based method to infer single cell gene regulatory network
abstract
GRN is the core of all living organisms that can explain how genes and their products interact at different levels. To infer the potential GRNs from gene expression data remains a great challenge in bioinformatics. Recently, with the development of single cell RNA sequencing technology, inferring cell specific GRNs involving in cell differentiation or cell function becomes a hot topic. Although there are some methods proposed to accomplish the task, it is still less than ideal because of the additional noises of pseudo time and high dropouts in datasets. Therefore, we propose a time-delayed mutual information based method, named MKG. MKG handles the above problems by partitioning the whole trajectory into several time windows then takes the average expression value of cells in each window as the representative cell. To further reduce the impact of dropouts, the mixed KSG estimator is applied to quantify the high-order time-delayed mutual information between pairs of genes. According to the experimental results on multiple simulated and real datasets, MKG has better performance and stability compared with other state-of-the-art algorithms.
Yanping Zeng, Xuhua Yan, Zhenlan Liang, Ruiqing Zheng, Min Li 0007
BIBM2