EDBT 2026 Demo / reviewers in the wild / expert
Jianping Zhao 0001
dblp:39/10301-1 · also Jian-Ping Zhao 0001, Jian-ping Zhao 0001
· DBLP profile ↗
19ranked-venue papers
1as first author
19since 2021 · last 2026
0000-0002-8486-744XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 1 first-author · 19 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AdPrST:An Adversarial Graph Deep Learning Pre-Clustering Framework for Deciphering Spatiotemporal Structures in Spatially Resolved TranscriptomicsabstractSpatially Resolved Transcriptomics (SRT) has revolutionized our understanding of gene expression within tissue microenvironments, yet accurately deciphering spatiotemporal structures-encompassing spatial domain identification, trajectory inference, and pseudo-spatiotemporal map construction-in complex tissues remains a formidable challenge. AdPrST begins with a pre-clustering process on gene expression data to establish initial domain groupings. It then constructs dual-view graph structures using K-Nearest Neighbors (KNN) for local similarities and r-radius for broader spatial contexts. Through adversarial self-supervised contrast, leveraging Wasserstein distance-based GANs and contrastive learning, AdPrST generates robust low-dimensional embeddings for each view. These embeddings are fused via a dot-product attention mechanism, guided by pre-clustering labels, to achieve accurate spatial domain identification. Benchmarking across multiple datasets demonstrated AdPrST's superior performance over state-of-the-art methods, highlighting its potential to advance spatial transcriptomics research by elucidating spatial functional patterns and developmental trajectories. In particular, AdPrST excels in inferencing spatiotemporal structures, reconstructing developmental sequences and temporal features in complex tissues. Shensi Huang, Jianping Zhao 0001, Junfeng Xia |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | MultiPep-DLCL: recognition of multifunctional therapeutic peptides through deep learning with label-sequence contrastive learningabstractIdentifying multifunctional therapeutic peptides (MFTP) is an important yet complex challenge in the realm of peptide recognition. Unlike monofunctional peptides, MFTP classification requires discerning fine-grained labeling information associated with amino acids, making it more intricate. Existing methods often ignore the nuanced semantics of these labels and fail to fully explore the interplay between peptide sequences and their labels. To address these issues, we propose a multilabel classification method named MultiPep-DLCL. This method uses a deep learning-based model architecture to translate peptide sequences into sequence features by learning the local and global dependencies of multifunctional therapeutic peptide sequences. Additionally, the Label-Sequence Fusion Transformer is employed to efficiently learn high-quality label embeddings by mining effective information from peptide sequences. Finally, the correspondence between sequence features and label embeddings is strengthened through label-sequence contrastive learning. To tackle dataset imbalance, MultiPep-DLCL integrates a multilabel focal dice loss function alongside the traditional cross-entropy loss function. Experimental results demonstrate that the MultiPep-DLCL significantly outperforms existing methods in MFTP recognition. Henghui Fan, Jianping Zhao 0001, Xiaomei Yang, Junfeng Xia |
Briefings Bioinform. | 3 |
| 2025 | scSDNE: A semi-supervised method for inferring cell-cell interactions based on graph embeddingabstractAs a fundamental characteristic of multicellular organisms, cell-cell communication is achieved through ligand-receptor (L-R) interactions, enabling the exchange of information and revealing the diversity of biological processes and cellular functions. To gain a comprehensive understanding of these complex interaction mechanisms, we constructed a manually curated L-R interaction database and developed a semi-supervised graph embedding model called scSDNE for inferring cell-cell interactions mediated by L-R interactions. scSDNE model utilizes the power of deep learning to map genes from interacting cells into a shared latent space, allowing for a nuanced representation of their relationships. Leveraging the prior information provided by database, scSDNE can infer significant L-R pairs involved in intercellular communication. Experiments on real single-cell RNA sequencing (scRNA-seq) datasets demonstrate that our method detects interactions with a high degree of reliability compared with other methods. More importantly, the model integrates gene regulation information within cells to enhance the accuracy and biological interpretability of the inferences. Our method provides a more comprehensive view of cell-cell interactions, offering new insights into complex intercellular communication. Chenchen Jia, Jianping Zhao 0001, Junfeng Xia, Chun-Hou Zheng 0001 |
PLoS Comput. Biol. | 3 |
| 2025 | scDMSC: Deep Multi-View Subspace Clustering for Single-Cell Multi-Omics DataabstractSingle-cell multi-omics sequencing technology comprehensively considers various molecular features to reveal the complexity of cells information. The clustering analysis of multi-omics data provides new insight into cellular heterogeneity. However, multi-omics data are characterized by high dimensionality, sparsity, and heterogeneity. Here, we propose an unsupervised clustering algorithm based on deep multi-view subspace learning, called scDMSC. This approach coordinates the heterogeneity of omics data through weighted reconstruction and employs deep subspace learning to identify shared latent features, elucidating the correlations among the omics. Our algorithm was rigorously tested across multiple real and simulated datasets, outperforming existing single-cell multi-omics integration methods and standard single-cell transcriptomics clustering tools in terms of both precision and scalability. Furthermore, differential expression and modality interpretability analyses in downstream applications highlight the model's capacity in uncovering biological mechanisms. Zile Wang, Fengyu Lei, Jianping Zhao 0001, Junfeng Xia |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | scHyper: reconstructing cell-cell communication through hypergraph neural networksabstractCell-cell communications is crucial for the regulation of cellular life and the establishment of cellular relationships. Most approaches of inferring intercellular communications from single-cell RNA sequencing (scRNA-seq) data lack a comprehensive global network view of multilayered communications. In this context, we propose scHyper, a new method that can infer intercellular communications from a global network perspective and identify the potential impact of all cells, ligand, and receptor expression on the communication score. scHyper designed a new way to represent tripartite relationships, by extracting a heterogeneous hypergraph that includes the source (ligand expression), the target (receptor expression), and the relevant ligand-receptor (L-R) pairs. scHyper is based on hypergraph representation learning, which measures the degree of match between the intrinsic attributes (static embeddings) of nodes and their observed behaviors (dynamic embeddings) in the context (hyperedges), quantifies the probability of forming hyperedges, and thus reconstructs the cell-cell communication score. Additionally, to effectively mine the key mechanisms of signal transmission, we collect a rich dataset of multisubunit complex L-R pairs and propose a nonparametric test to determine significant intercellular communications. Comparing with other tools indicates that scHyper exhibits superior performance and functionality. Experimental results on the human tumor microenvironment and immune cells demonstrate that scHyper offers reliable and unique capabilities for analyzing intercellular communication networks. Therefore, we introduced an effective strategy that can build high-order interaction patterns, surpassing the limitations of most methods that can only handle low-order interactions, thus more accurately interpreting the complexity of intercellular communications. Wenying Li, Jianping Zhao 0001, Junfeng Xia |
Briefings Bioinform. | 3 |
| 2024 | scVSC: Deep Variational Subspace Clustering for Single-Cell Transcriptome DataabstractSingle-cell RNA sequencing (scRNA-seq) is a potent advancement for analyzing gene expression at the individual cell level, allowing for the identification of cellular heterogeneity and subpopulations. However, it suffers from technical limitations that result in sparse and heterogeneous data. Here, we propose scVSC, an unsupervised clustering algorithm built on deep representation neural networks. The method incorporates the variational inference into the subspace model, which imposes regularization constraints on the latent space and further prevents overfitting. In a series of experiments across multiple datasets, scVSC outperforms existing state-of-the-art unsupervised and semi-supervised clustering tools regarding clustering accuracy and running efficiency. Moreover, the study indicates that scVSC could visually reveal the state of trajectory differentiation, accurately identify differentially expressed genes, and further discover biologically critical pathways. Zile Wang, Jianping Zhao 0001, Junfeng Xia, Chun-Hou Zheng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | scGMAAE: Gaussian mixture adversarial autoencoders for diversification analysis of scRNA-seq dataabstractThe progress of single-cell RNA sequencing (scRNA-seq) has led to a large number of scRNA-seq data, which are widely used in biomedical research. The noise in the raw data and tens of thousands of genes pose a challenge to capture the real structure and effective information of scRNA-seq data. Most of the existing single-cell analysis methods assume that the low-dimensional embedding of the raw data belongs to a Gaussian distribution or a low-dimensional nonlinear space without any prior information, which limits the flexibility and controllability of the model to a great extent. In addition, many existing methods need high computational cost, which makes them difficult to be used to deal with large-scale datasets. Here, we design and develop a depth generation model named Gaussian mixture adversarial autoencoders (scGMAAE), assuming that the low-dimensional embedding of different types of cells follows different Gaussian distributions, integrating Bayesian variational inference and adversarial training, as to give the interpretable latent representation of complex data and discover the statistical distribution of different types of cells. The scGMAAE is provided with good controllability, interpretability and scalability. Therefore, it can process large-scale datasets in a short time and give competitive results. scGMAAE outperforms existing methods in several ways, including dimensionality reduction visualization, cell clustering, differential expression analysis and batch effect removal. Importantly, compared with most deep learning methods, scGMAAE requires less iterations to generate the best results. Jianping Zhao 0001, Chun-Hou Zheng 0001, Yansen Su |
Briefings Bioinform. | 2 |
| 2023 | scSemiAAE: a semi-supervised clustering model for single-cell RNA-seq dataabstractBACKGROUND: Single-cell RNA sequencing (scRNA-seq) strives to capture cellular diversity with higher resolution than bulk RNA sequencing. Clustering analysis is critical to transcriptome research as it allows for further identification and discovery of new cell types. Unsupervised clustering cannot integrate prior knowledge where relevant information is widely available. Purely unsupervised clustering algorithms may not yield biologically interpretable clusters when confronted with the high dimensionality of scRNA-seq data and frequent dropout events, which makes identification of cell types more challenging. RESULTS: We propose scSemiAAE, a semi-supervised clustering model for scRNA sequence analysis using deep generative neural networks. Specifically, scSemiAAE carefully designs a ZINB adversarial autoencoder-based architecture that inherently integrates adversarial training and semi-supervised modules in the latent space. In a series of experiments on scRNA-seq datasets spanning thousands to tens of thousands of cells, scSemiAAE can significantly improve clustering performance compared to dozens of unsupervised and semi-supervised algorithms, promoting clustering and interpretability of downstream analyses. CONCLUSION: scSemiAAE is a Python-based algorithm implemented on the VSCode platform that provides efficient visualization, clustering, and cell type assignment for scRNA-seq data. The tool is available from https://github.com/WHang98/scSemiAAE . Zile Wang, Jianping Zhao 0001, Chun-Hou Zheng 0001 |
BMC Bioinform. | 3 |
| 2023 | An Integrated Method Based on Wasserstein Distance and Graph for Cancer Subtype DiscoveryabstractDue to the complexity of cancer pathogenesis at different omics levels, it is necessary to find a comprehensive method to accurately distinguish and find cancer subtypes for cancer treatment. In this paper, we proposed a new cancer multi-omics subtype identification method, which is based on variational autoencoder measured by Wasserstein distance and graph autoencoder (WVGMO). This method depends on two foremost models. The first model is a variational autoencoder measured by Wasserstein distance (WVAE), which is used to extract potential spatial information of each omic data type. The second model is the graph autoencoder (GAE) with the second-order proximity. It has the capability to retain the topological structure information and feature information of the multi-omics data. And then, the identification of cancer subtypes via k-means clustering. Extensive experiments were conducted on seven different cancers based on four omics data from TCGA. The results show that WVGMO provides equivalent or even better results than the most of advanced synthesis methods. Jianping Zhao 0001, Chun-Hou Zheng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | PACVP: Prediction of Anti-Coronavirus Peptides Using a Stacking Learning Strategy With Effective Feature RepresentationabstractDue to the global outbreak of COVID-19 and its variants, antiviral peptides with anti-coronavirus activity (ACVPs) represent a promising new drug candidate for the treatment of coronavirus infection. At present, several computational tools have been developed to identify ACVPs, but the overall prediction performance is still not enough to meet the actual therapeutic application. In this study, we constructed an efficient and reliable prediction model PACVP (Prediction of Anti-CoronaVirus Peptides) for identifying ACVPs based on effective feature representation and a two-layer stacking learning framework. In the first layer, we use nine feature encoding methods with different feature representation angles to characterize the rich sequence information and fuse them into a feature matrix. Secondly, data normalization and unbalanced data processing are carried out. Next, 12 baseline models are constructed by combining three feature selection methods and four machine learning classification algorithms. In the second layer, we input the optimal probability features into the logistic regression algorithm (LR) to train the final model PACVP. The experiments show that PACVP achieves favorable prediction performance on independent test dataset, with ACC of 0.9208 and AUC of 0.9465. We hope that PACVP will become a useful method for identifying, annotating and characterizing novel ACVPs. Shouzhi Chen, Yanhong Liao, Jianping Zhao 0001, Yannan Bin, Chun-Hou Zheng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | NeuroPred-CLQ: incorporating deep temporal convolutional networks and multi-head attention mechanism to predict neuropeptidesabstractNeuropeptides (NPs) are a particular class of informative substances in the immune system and physiological regulation. They play a crucial role in regulating physiological functions in various biological growth and developmental stages. In addition, NPs are crucial for developing new drugs for the treatment of neurological diseases. With the development of molecular biology techniques, some data-driven tools have emerged to predict NPs. However, it is necessary to improve the predictive performance of these tools for NPs. In this study, we developed a deep learning model (NeuroPred-CLQ) based on the temporal convolutional network (TCN) and multi-head attention mechanism to identify NPs effectively and translate the internal relationships of peptide sequences into numerical features by the Word2vec algorithm. The experimental results show that NeuroPred-CLQ learns data information effectively, achieving 93.6% accuracy and 98.8% AUC on the independent test set. The model has better performance in identifying NPs than the state-of-the-art predictors. Visualization of features using t-distribution random neighbor embedding shows that the NeuroPred-CLQ can clearly distinguish the positive NPs from the negative ones. We believe the NeuroPred-CLQ can facilitate drug development and clinical trial studies to treat neurological disorders. Shouzhi Chen, Jianping Zhao 0001, Yannan Bin, Chun-Hou Zheng 0001 |
Briefings Bioinform. | 3 |
| 2022 | scCNC: a method based on capsule network for clustering scRNA-seq dataabstractMOTIVATION: A large number of studies have shown that clustering is a crucial step in scRNA-seq analysis. Most existing methods are based on unsupervised learning without the prior exploitation of any domain knowledge, which does not utilize available gold-standard labels. When confronted by the high dimensionality and general dropout events of scRNA-seq data, purely unsupervised clustering methods may not produce biologically interpretable clusters, which complicate cell type assignment. RESULTS: In this article, we propose a semi-supervised clustering method based on a capsule network named scCNC that integrates domain knowledge into the clustering step. Significantly, we also propose a Semi-supervised Greedy Iterative Training method used to train the whole network. Experiments on some real scRNA-seq datasets show that scCNC can significantly improve clustering performance and facilitate downstream analyses. AVAILABILITY AND IMPLEMENTATION: The source code of scCNC is freely available at https://github.com/WHY-17/scCNC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jianping Zhao 0001, Chun-Hou Zheng 0001, Yansen Su |
Bioinform. | 2 |
| 2022 | iEnhancer-DCLA: using the original sequence to identify enhancers and their strength based on a deep learning frameworkabstractEnhancers are small regions of DNA that bind to proteins, which enhance the transcription of genes. The enhancer may be located upstream or downstream of the gene. It is not necessarily close to the gene to be acted on, because the entanglement structure of chromatin allows the positions far apart in the sequence to have the opportunity to contact each other. Therefore, identifying enhancers and their strength is a complex and challenging task. In this article, a new prediction method based on deep learning is proposed to identify enhancers and enhancer strength, called iEnhancer-DCLA. Firstly, we use word2vec to convert k-mers into number vectors to construct an input matrix. Secondly, we use convolutional neural network and bidirectional long short-term memory network to extract sequence features, and finally use the attention mechanism to extract relatively important features. In the task of predicting enhancers and their strengths, this method has improved to a certain extent in most evaluation indexes. In summary, we believe that this method provides new ideas in the analysis of enhancers. Meng Liao, Jianping Zhao 0001, Chun-Hou Zheng 0001 |
BMC Bioinform. | 2 |
| 2022 | scDSSC: Deep Sparse Subspace Clustering for scRNA-seq DataabstractSingle cell RNA sequencing (scRNA-seq) enables researchers to characterize transcriptomic profiles at the single-cell resolution with increasingly high throughput. Clustering is a crucial step in single cell analysis. Clustering analysis of transcriptome profiled by scRNA-seq can reveal the heterogeneity and diversity of cells. However, single cell study still remains great challenges due to its high noise and dimension. Subspace clustering aims at discovering the intrinsic structure of data in unsupervised fashion. In this paper, we propose a deep sparse subspace clustering method scDSSC combining noise reduction and dimensionality reduction for scRNA-seq data, which simultaneously learns feature representation and clustering via explicit modelling of scRNA-seq data generation. Experiments on a variety of scRNA-seq datasets from thousands to tens of thousands of cells have shown that scDSSC can significantly improve clustering performance and facilitate the interpretability of clustering and downstream analysis. Compared to some popular scRNA-deq analysis methods, scDSSC outperformed state-of-the-art methods under various clustering performance metrics. Jianping Zhao 0001, Chun-Hou Zheng 0001, Yansen Su |
PLoS Comput. Biol. | 2 |
| 2022 | scCDG: A Method Based on DAE and GCN for scRNA-Seq Data AnalysisabstractIdentifying cell types is one of the main goals of single-cell RNA sequencing (scRNA-seq) analysis, and clustering is a common method for this item. However, the massive amount of data and the excess noise level bring challenge for single cell clustering. To address this challenge, in this paper, we introduced a novel method named single-cell clustering based on denoising autoencoder and graph convolution network (scCDG), which consists of two core models. The first model is a denoising autoencoder (DAE) used to fit the data distribution for data denoising. The second model is a graph autoencoder using graph convolution network (GCN), which projects the data into a low-dimensional space (compressed) preserving topological structure information and feature information in scRNA-seq data simultaneously. Extensive analysis on seven real scRNA-seq datasets demonstrate that scCDG outperforms state-of-the-art methods in some research sub-fields, including single cell clustering, visualization of transcriptome landscape, and trajectory inference. Jianping Zhao 0001, Yansen Su, Chun-Hou Zheng 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | SNEMO: Spectral Clustering Based on the Neighborhood for Multi-omics Data
Jianping Zhao 0001, Chun-Hou Zheng 0001 |
ICIC (3) | 2 |
| 2021 | SHDC: A Method of Similarity Measurement Using Heat Kernel Based on Denoising for Clustering scRNA-seq Data
Jianping Zhao 0001, Chun-Hou Zheng 0001 |
ICIC (3) | 1 |
| 2021 | Identification of driver genes based on gene mutational effects and network centralityabstractBACKGROUND: As one of the deadliest diseases in the world, cancer is driven by a few somatic mutations that disrupt the normal growth of cells, and leads to abnormal proliferation and tumor development. The vast majority of somatic mutations did not affect the occurrence and development of cancer; thus, identifying the mutations responsible for tumor occurrence and development is one of the main targets of current cancer treatments. RESULTS: To effectively identify driver genes, we adopted a semi-local centrality measure and gene mutation effect function to assess the effect of gene mutations on changes in gene expression patterns. Firstly, we calculated the mutation score for each gene. Secondly, we identified differentially expressed genes (DEGs) in the cohort by comparing the expression profiles of tumor samples and normal samples, and then constructed a local network for each mutation gene using DEGs and mutant genes according to the protein-protein interaction network. Finally, we calculated the score of each mutant gene according to the objective function. The top-ranking mutant genes were selected as driver genes. We name the proposed method as mutations effect and network centrality. CONCLUSIONS: Four types of cancer data in The Cancer Genome Atlas were tested. The experimental data proved that our method was superior to the existing network-centric method, as it was able to quickly and easily identify driver genes and rare driver factors. Yun-Yun Tang, Pi-Jing Wei, Jianping Zhao 0001, Junfeng Xia, Chun-Hou Zheng 0001 |
BMC Bioinform. | 3 |
| 2021 | Clustering of cancer data based on Stiefel manifold for multiple viewsabstractBACKGROUND: In recent years, various sequencing techniques have been used to collect biomedical omics datasets. It is usually possible to obtain multiple types of omics data from a single patient sample. Clustering of omics data plays an indispensable role in biological and medical research, and it is helpful to reveal data structures from multiple collections. Nevertheless, clustering of omics data consists of many challenges. The primary challenges in omics data analysis come from high dimension of data and small size of sample. Therefore, it is difficult to find a suitable integration method for structural analysis of multiple datasets. RESULTS: In this paper, a multi-view clustering based on Stiefel manifold method (MCSM) is proposed. The MCSM method comprises three core steps. Firstly, we established a binary optimization model for the simultaneous clustering problem. Secondly, we solved the optimization problem by linear search algorithm based on Stiefel manifold. Finally, we integrated the clustering results obtained from three omics by using k-nearest neighbor method. We applied this approach to four cancer datasets on TCGA. The result shows that our method is superior to several state-of-art methods, which depends on the hypothesis that the underlying omics cluster class is the same. CONCLUSION: Particularly, our approach has better performance than compared approaches when the underlying clusters are inconsistent. For patients with different subtypes, both consistent and differential clusters can be identified at the same time. Jianping Zhao 0001, Chun-Hou Zheng 0001 |
BMC Bioinform. | 2 |