EDBT 2026 Demo / reviewers in the wild / expert
Xiaoshu Zhu
dblp:62/7808
· DBLP profile ↗
22ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0002-7696-9112ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 6 first-author · 14 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | scAQUA: A Quintuplet Constraint-Based Batch Effect Correction Method for Single-Cell Multi-omics Data
Tiezheng Qiao, Huiyong Zhang, Xiaoshu Zhu |
ISBRA (1) | 4 |
| 2026 | Essential Proteins Prediction Using Features Synergy Model and GO Pure CentralityabstractEssential proteins are a crucial component of living organisms, and their absence will lead to cell death or reproductive arrest. Discovering these proteins can propel advancements in synthetic biology and facilitate the development of novel antibiotics and therapies for various diseases. However, current computational methods suffer from two major drawbacks that hinder their discovery rate: one is the significant noise in protein-protein interaction (PPI) data, and the other is the inadequate consideration of feature relationships. To enhance identification capabilities, this study proposes a novel essential protein prediction method, Feature Synergy Method (FSM), which leverages a features synergy model and GO pure centrality. The FSM is described as follows:Firstly, based on the principle of co-expression, gene expression data are integrated with the original PPI network to construct a pure PPI network (PPIN). Subsequently, GO annotation data are employed to calculate GO_sim weights for the interactions within the original PPI network, forming a GS_PIN. The PPIN and GS_PIN are then fused to establish the GS_PPIN, which helps mitigate the impact of noise in PPI data. Secondly, a new centrality measure, GO pure centrality (GPC), is designed based on this GO similarity-weighted pure PPI network. Thirdly, an evolutionary conservation score (ECS) is extracted from subcellular localization and orthologous proteins data. Fourthly, after analyzing the relationship between GPC and ECS, a novel fusion model, the features synergy model, is developed to integrate GPC and ECS, ultimately leading to the proposal of the new essential protein prediction method, FSM. To validate the performance of FSM, six computational methods (PeC, WDC, ION, NCCO, E_POC, and JDC) and six centrality measures (NC, IC, EC, SC, CC, and DC) were evaluated on three distinct yeast datasets. The results demonstrate that FSM achieves a higher essential protein identification rate. Similarly, GPC identifies more essential proteins compared to the six centrality-based approaches (NC, IC, EC, SC, CC, and DC). Xinlong Luo 0002, Gaoshi Li, Zhipeng Hu, Jingli Wu, Wei Peng 0004, Jiafei Liu 0001, Xiaoshu Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2026 | A Deep Learning Framework for Identifying Essential Proteins Based on Vision TransformerabstractEssential proteins are fundamental to the reproduction and survival of cells, and if they are killed, the cells will stop reproducing or die. Many computational methods of identifying essential proteins are proposed which fuse a large number of features from multi-omics data. Some of them extract features from subcellular localization data by subjectively selecting certain subcellular locations. Meanwhile, there is still room to improve the identification rate of essential proteins. In this paper, a new deep learning framework for identifying essential proteins based on Vision Transformer is proposed, named EPViT. Firstly, topological features are extracted from the protein-protein interaction network. Secondly, a feature matrix is designed from the subcellular localization information without subjective factor. Then, the two classes of features are fused into a new feature matrix by outer product operation. Finally, the new feature matrix is input into the Vision Transformer model to discover essential proteins. The results show that EPViT has the highest recognition rate among the comparison experiments on yeast data. Gaoshi Li, Jingli Wu, Wei Peng 0004, Jiafei Liu 0001, Xiaoshu Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2025 | scSAGE: Combining Sparse Representation and Gene-Guided Clustering for Batch Effect Correction in scRNA-Seq DataabstractLarge-scale single-cell RNA sequencing (scRNAseq) datasets are often generated across multiple experimental batches, introducing batch effects that hinder downstream analysis and biological interpretation. We propose scSAGE, a novel batch effect correction framework that integrates sparse representation learning with gene-guided clustering. The method employs a sparsity-constrained autoencoder to extract lowdimensional embeddings that preserve biological signals while reducing technical noise. It further constructs a clustering similarity matrix based on gene expression to refine cell grouping. To enhance alignment across batches, a triplet constraint learning module is applied to pull together similar cells and push apart dissimilar ones. We evaluated scSAGE against six state-of-theart methods on six real scRNA-seq datasets. Experimental results show that scSAGE consistently achieved the highest performance in clustering accuracy metrics such as adjusted Rand index and normalized mutual information, improved batch mixing, and successfully identified rare pancreatic cell populations that were overlooked in uncorrected data. These findings demonstrate that scSAGE offers a robust and effective solution for constructing accurate and integrated single-cell atlases across batches. Xiaoshu Zhu, Tiezheng Qiao, Letong Chen, Zhenzhong Zeng |
BIBM | 1 |
| 2025 | ViDSG: A Hybrid Algorithm Integrating Statistical and Semantic Features via Dual-Channels for Identifying Prokaryotic and Eukaryotic Viruses
JianPeng Zhang, Changna Qian, Xiaoshu Zhu |
ISBRA (1) | 6 |
| 2025 | A Hybrid LSTM-CNN Algorithm for Swine Posture Recognition Using Multimodal Sensor Data
Xiaoshu Zhu, Changna Qian |
ISNN | 1 |
| 2025 | Using Multi-Feature Weak Consensus Model to Discover Essential ProteinsabstractEssential proteins play an essential role in cell survival and replication. Currently, more and more computational methods are developed to identify essential proteins, which overcome the time-consuming, costly and inefficient shortcomings with biological experimental methods. In order to improve the recognition rate, some new methods by fusing multiple features are developed, but they seldom consider the connection among features. After analyzing a large number of methods based on multi-feature fusion, a phenomenon among features is found, called weak consensus, then a weak consensus model to fuse these features is proposed in this paper. After analyzing the relationship between a protein and its neighbors in protein-protein interaction networks, a new centrality, namely neighborhood aggregation centrality(NAC) is developed in this paper. Then, a Max-Min strategy is used to integrate NAC with Pearson correlation coefficient and Jaccard similarity coefficient based on gene expression data to obtain local importance score. In addition, orthologous feature score is used to measure proteins conservation. Finally, by using the weak consensus model to fuse orthologous feature score with local importance score, a new method WOL is proposed in this paper. Then experiments are performed on S.cerevisiae data. The results show that compared with WDC, PeC, ION, JDC, NCCO and E_POC, WOL has a higher recognition rate. Zhipeng Hu, Gaoshi Li, Xinlong Luo 0002, Jiafei Liu 0001, Jingli Wu, Wei Peng 0004, Xiaoshu Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2024 | A Hybrid Algorithm based on Autoencoder and Text Convolutional Network for Integrating scATAC-seq and scRNA-seq DataabstractIntegrating single-cell RNA-seq (scRNA-seq) data and single-cell ATAC-seq (scATAC-seq) data provides a more comprehensive view of cellular heterogeneity. However, the high sparsity in scATAC-seq data presents significant challenges for cell type identification, while scRNA-seq offers richer gene expression information for more accurate annotation. We propose scEDI, a method that integrates scRNA-seq and scATAC- seq data using autoencoders and text convolutional networks. In scEDI, both data types are processed through an autoencoder to create a shared low-dimensional space, followed by a text convolutional network for label transfer. We evaluated scEDI against three state-of-the-art methods on datasets from the adult mouse cerebral cortex and human peripheral blood monocytes. Experimental results demonstrate that scEDI improves label transfer accuracy, particularly in small datasets. Xiaoshu Zhu, Wei Lan 0001 |
BIBM | 1 |
| 2024 | AGImpute: imputation of scRNA-seq data based on a hybrid GAN with dropouts identificationabstractMOTIVATION: Dropout events bring challenges in analyzing single-cell RNA sequencing data as they introduce noise and distort the true distributions of gene expression profiles. Recent studies focus on estimating dropout probability and imputing dropout events by leveraging information from similar cells or genes. However, the number of dropout events differs in different cells, due to the complex factors, such as different sequencing protocols, cell types, and batch effects. The dropout event differences are not fully considered in assessing the similarities between cells and genes, which compromises the reliability of downstream analysis. RESULTS: This work proposes a hybrid Generative Adversarial Network with dropouts identification to impute single-cell RNA sequencing data, named AGImpute. First, the numbers of dropout events in different cells in scRNA-seq data are differentially estimated by using a dynamic threshold estimation strategy. Next, the identified dropout events are imputed by a hybrid deep learning model, combining Autoencoder with a Generative Adversarial Network. To validate the efficiency of the AGImpute, it is compared with seven state-of-the-art dropout imputation methods on two simulated datasets and seven real single-cell RNA sequencing datasets. The results show that AGImpute imputes the least number of dropout events than other methods. Moreover, AGImpute enhances the performance of downstream analysis, including clustering performance, identifying cell-specific marker genes, and inferring trajectory in the time-course dataset. AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/xszhu-lab/AGImpute. Xiaoshu Zhu, Shuang Meng, Gaoshi Li, Jianxin Wang 0001, Xiaoqing Peng |
Bioinform. | 1 |
| 2024 | Identification of Cancer Driver Genes based on Dynamic Incentive ModelabstractCancer is a complex genomic mutation disease, and identifying cancer driver genes promotes the development of targeted drugs and personalized therapies. The current computational method takes less consideration of the relationship among features and the effect of noise in protein-protein interaction(PPI) data, resulting in a low recognition rate. In this paper, we propose a cancer driver genes identification method based on dynamic incentive model, DIM. This method firstly constructs a hypergraph to reduce the impact of false positive data in PPI. Then, the importance of genes in each hyperedge in hypergraph is considered from three perspectives, network and functional score(NFS) is proposed. By analyzing the relation among features, the dynamic incentive model is proposed to fuse NFS, the differential expression score of mRNA and the differential expression score of miRNA. DIM is compared with some classical methods on breast cancer, lung cancer, prostate cancer, and pan-cancer datasets. The results show that DIM has the best performance on statistical evaluation indicators, functional consistency and the partial area under the ROC curve, and has good cross-cancer capability. Zhipeng Hu, Gaoshi Li, Xinlong Luo 0002, Wei Peng 0004, Jiafei Liu 0001, Xiaoshu Zhu, Jingli Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2023 | Essential proteins identification based on weak consensus model and neighborhood aggregation centralityabstractEssential proteins play an essential role in cell survival and replication. Currently, more and more computational methods are developed to identify essential proteins, which overcome the time-consuming, costly and inefficient shortcomings with biological experimental methods. In order to improve the recognition rate, some new methods by fusing multiple features are developed, but they seldom consider the connection among features. After analyzing a large number of methods based on multi-feature fusion, a weak consensus model to fuse features is proposed in this paper. Then, this paper uses the weak consensus model to fuse protein-protein interaction network, gene expression data, and orthologous data, thus proposing a new method, WOL. Then experiments are performed on one S.cerevisiae dataset. The results show that compared with WDC, PeC, ION, JDC, NCCO and E_POC, WOL has a higher recognition rate. Zhipeng Hu, Gaoshi Li, Jingli Wu, Xinlong Luo 0002, Jiafei Liu 0001, Wei Peng 0004, Xiaoshu Zhu |
BIBM | 7 |
| 2023 | Gene selection and clustering of single-cell data based on Fisher score and genetic algorithm
Junhong Feng, Jie Zhang 0072, Xiaoshu Zhu |
J. Supercomput. | 3 |
| 2022 | scIAC: clustering scATAC-seq data based on Student's t-distribution similarity imputation and denoising autoencoderabstractAssay of single cell transposase-accessible chromatin with high-throughput sequencing (scATAC-seq) have enabled massively profiling of the chromatin accessibility landscape at the single-cell level. The essential step in analyzing scATAC-seq data is to cluster the cells into different clusters and utilize the clustering information in the subsequent downstream analysis. However, there are some challenges in the clustering analysis of scATAC-seq data. For example, scATAC-seq data are often high-dimensional and extremely sparse, as well as featuring high loss rate or noise. In this study, we proposed the scIAC to address these challenges of scATACseq data. In particular, scIAC combines the Student’s t-distribution similarity imputation and the denoising autoencoder based on the Zero-inflated Negative Binomial (ZINB) distribution. The Student’s t-distribution similarity imputation is used to solve the problem of high sparsity and high loss rate. The denoising autoencoder is employ to extract features which are useful for clustering and to reduce data noises. In addition, the self-training soft K-means and pairwise constraints are utilized in the clustering phase to enhance clustering performance. The experimental validation on several datasets shows that the proposed method performed better than other state-of-the-art methods. In conclusion, scIAC is an effective method to accurately cluster and identify cell types in scATAC-seq data. Wei Lan 0001, Jin Ye 0003, Xiaoshu Zhu, Qingfeng Chen, Yi Pan 0001 |
BIBM | 4 |
| 2022 | STgcor: A Distribution-Based Correlation Measurement Method for Spatial Transcriptome Data
Xiaoshu Zhu, Liyuan Pang, Wei Lan 0001, Shuang Meng, Xiaoqing Peng |
ISBRA | 1 |
| 2021 | SCOTCluster: Deep Clustering with Optimal Transport for Large-scale Single-cell RNA-seq DataabstractSingle-cell RNA sequencing (scRNA-seq) presents cell heterogeneity in a high resolution to explore cell development. The high dimension and the high noise in scRNAseq data bring some computational challenges. By introducing optimal transmission regularization, we proposed a novel deep clustering method, called SCOTCluster, which accurately learned low-dimensional representations. In SCOTCluster, a joint training strategy was designed by integrating AutoEncoder and soft k-means. Notably, to improve simultaneously the accuracy and robustness, the optimal transmission was introduced in the objective function of soft k-means, and entropy regularization and Sinkhorn iterative algorithm were performed to constraint the cluster size. To test the performance, we compared SCOTCluster with five state-of the-art methods on 16 real large-scale scRNA-seq datasets. The experimental results showed that SCOTCluster improved the training stability and clustering performance. Faning Long, Xiaoqing Peng, Jianxin Wang 0001, Xiaoshu Zhu |
BIBM | 5 |
| 2021 | ScDA: A Denoising AutoEncoder Based Dimensionality Reduction for Single-cell RNA-seq Data
Xiaoshu Zhu, Yongchang Lin, Jianxin Wang 0001, Xiaoqing Peng |
ISBRA | 1 |
| 2019 | A Global Similarity Learning for Clustering of Single-Cell RNA-Seq DataabstractSingle-cell RNA-seq (scRNA-seq) data analysis is a powerful tool for biological researches. Similarity plays an important role in clustering scRNA-seq data. Existing similarity measurements are mainly based on local distance information that is calculated between directly connected node pairs, or shared nearest neighbours' information, without considering the global information. Therefore, these similarity measurements may be not very accurate based on the insufficient information. Based on multi-kernel indices in a global feature space and path-based similarity, we proposed a new similarity measurement for single-cell clustering, called multi-kernel and path-based global similarity (MPGS). In MPGS, global information was incorporated by a new feature space from Spearman correlation coefficient, and a global similarity matrix calculated by multi-kernel. A path-based similarity metric was designed to expand the relevant node range. Based on this similaritiy, a modified Louvain community detection method was applied to cluster the scRNA-seq data, named MPGS-Louvain. To validate the performance of MPGS, the clustering performances of several clustering methods combined with different similarity measurements were compared. To demonstrate the performance of MPGS-Louvain, we compared MPGS-Louvain and five scRNA-seq clustering methods on twenty scRNA-seq datasets. The experimental results showed that MPGS outperformed other similarity measurements, and MPGS-Louvain achieved better performance on these datasets. It can be observed that MPGS provided a new insight to improve the accuracy of clustering scRNA-seq data by considering the global information in similarity measurement. MPGS-Louvain automatically detected clusters accurately without prior knowledge. Xiaoshu Zhu, Lilu Guo, Yunpei Xu, Hong-Dong Li, Xingyu Liao, Fang-Xiang Wu, Xiaoqing Peng |
BIBM | 1 |
| 2019 | Finding community of brain networks based on artificial bee colony with uniform design
Jie Zhang 0072, Xiaoshu Zhu, Junhong Feng, Yifang Yang |
Multim. Tools Appl. | 2 |
| 2017 | A multi-objective biclustering algorithm based on fuzzy mathematics
Xiaoshu Zhu, Miao Xie, Jianxin Wang 0001 |
Neurocomputing | 1 |
| 2017 | Multi-objective differential evolution with dynamic covariance matrix learning for multi-objective optimization problems with variable linkages
Qiaoyong Jiang, Lei Wang 0030, Jiatang Cheng, Xiaoshu Zhu, Wei Li 0068, Yanyan Lin, Guolin Yu, Xinhong Hei 0001, Jinwei Zhao |
Knowl. Based Syst. | 4 |
| 2017 | A novel chaos optimization algorithm
Junhong Feng, Jie Zhang 0072, Xiaoshu Zhu, Wenwu Lian |
Multim. Tools Appl. | 3 |
| 2016 | Efficient kNN classification algorithm for big data
Zhenyun Deng, Xiaoshu Zhu, Debo Cheng, Ming Zong, Shichao Zhang 0001 |
Neurocomputing | 2 |