EDBT 2026 Demo / reviewers in the wild / expert
Jiancheng Zhong
dblp:145/7954
· DBLP profile ↗
17ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0001-6334-9304ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 8 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MACE: A Multi-scale Attention Convolutional Evaluation Model for Top-Down Mass Spectrometry Isotopic Envelopes
Jiancheng Zhong, Maoqi Yuan, Shaokai Wang |
ISBRA (2) | 1 |
| 2025 | ScFold: a GNN-based model for efficient inverse folding of short-chain proteins via spatial reductionabstractIn the realm of protein design, the efficient construction of protein sequences that accurately fold into predefined structures has become an important area of research. Although advancements have been made in the study of long-chain proteins, the design of short-chain proteins requires equal consideration. The structural information inherent in short and single chains is typically less comprehensive than that of full-length chains, which can negatively impact their performance. To address this challenge, we introduce ScFold, a novel model that incorporates an innovative node module. This module utilizes spatial dimensionality reduction and positional encoding mechanisms to enhance the extraction of structural features. Experimental results indicate that ScFold achieves a recovery rate of 52.22$\%$ on the CATH4.2 dataset, demonstrating notable efficacy for short-chain proteins, with a recovery rate of 41.6$\%$. Additionally, ScFold further exhibits enhanced recovery rates of 59.32$\%$ and 61.59$\%$ on the TS50 and TS500 datasets, respectively, demonstrating its effectiveness across diverse protein types. Additionally, we performed protein length stratification on the TS500 and CATH4.2 datasets and tested ScFold on length-specific sub-datasets. The results confirm the model's superiority in handling short-chain proteins. Finally, we selected several protein sequence groups from the CATH4.2 dataset for structural visualization analysis and provided comparisons between the model-generated sequences and the target sequences. Jiancheng Zhong, Zhiwei Zou, Shaokai Wang |
Briefings Bioinform. | 1 |
| 2024 | Multi-source Data-driven Drug Repositioning based on Deep Self-attention NetworkabstractDrug-target interaction (DTI) prediction is the core of drug repositioning, which can promote rapid screening of candidate drugs and shorten drug development time. Methods based on graph neural networks (GNN) have attracted widespread attention and shown certain advantages in DTI prediction. However, extracting deep features from graph structure information in heterogeneous bioinformatic network (HBIN) still faces challenges. In this work, we propose a novel deep self-attention model called BCHNDTI. First, the model integrates the biological information of drugs and targets to construct HBIN, and preprocesses graph information by using the path integral thermonuclear (PIHK) method. Then, we use Graph Convolutional Network (GCN) and Multi-Head Attention to obtain representations of HBIN features and input them into a novel self-attention network, which can be used to learn deep feature representations of drugs and targets. Finally, we use the XGBoost classifier to predict DTIs. Compared with several recent DTI prediction methods, the results indicated that BCHNDTI performs better on the benchmark dataset for identifying DTIs. The case study further confirmed its ability to predict potential drug-target interactions. Qiu Xiao, Tuo Xiong, Yide Yang, Jiancheng Zhong |
BIBM | 6 |
| 2024 | PrSMBooster: Improving the Accuracy of Top-Down Proteoform Characterization Using Deep Learning Rescoring Models
Jiancheng Zhong, Maoqi Yuan, Shaokai Wang |
ISBRA (3) | 1 |
| 2022 | Predicting miRNA-disease associations via multi-channel graph convolutional networksabstractExtensive research evidence shows that variation and dysregulation of microRNAs(miRNAs) are important causes of disease, and therefore the study of miRNA-disease associations has important theoretical and applied implications in the field of human disease research and treatment. Based on the time and cost of validating miRNA-disease associations in traditional medicine clinical experiments, using multiple biological datasets to predict potential miRNA-disease associations (MDAs) has become a hot topic in the field of biological research in recent years. This paper develops a novel model of MDA-RGCN based on a multi-channel graph convolutional network and graph attention for MDAs prediction. Based on graph theory, this study treats MDAs prediction as a node classification task. To learn the topology and various interactions between feature graph nodes of various strengths, we employ two independent graph attention networks, which increases training efficiency and accuracy. In order to learn information that is shared by both graphs, we employ a GCN with a shared weight matrix simultaneously. Comprehensive experiments reveal that the prediction performance of MDA-RGCN excels other more sophisticated models for MDAs prediction. Furthermore, we further confirmed the predictive ability of MDA-RGCN to identify potential disease-related miRNAs by selecting two human diseases for case study. HaoRan Zheng, Qiu Xiao, Jiancheng Zhong |
BIBM | 3 |
| 2022 | Prediction of miRNA and Disease Association based on Graph Convolution Network using Latent Feature Vector by Positive SamplesabstractThe traditional approach of wet biological experiments tends to reveal whether there is an association between specific miRNA molecules and diseases, resulting in a lack of reliable negative samples of miRNA-disease associations in existing databases. To deal with the problem that many current computational methods treat unknown miRNA-disease associations in benchmark datasets as negative samples directly, we propose a graph convolutional neural network model, named DNMFGCN-MDA, based on feature extractions of positive samples. Firstly, by only using a dynamic matrix with positive samples, we extracted the potential feature vectors U and V of miRNAs and diseases in low-dimensional space. Then, we combined known miRNA-disease associations with potential feature vectors U and V to construct heterogeneous graph neural networks. Finally, we adopted graph convolutional neural networks to learn the structural features of the heterogeneous graph network and predicted the potential miRNA-disease associations by using linkage prediction. We use a 5-fold cross-validation experiment to evaluate the performance of our mode, and it turns out that our model has an average AUC of 95.23% and 96.07% on the HMDD v2.0 and HMDD v3.2 datasets, respectively, and achieves better-associated performance than that of the comparison method. Jiancheng Zhong, Jiedong Kang, Xingran Song, Qiu Xiao, Yi Pan 0001 |
BIBM | 1 |
| 2022 | Prediction of Drug-Disease Relationship on Heterogeneous Networks Based on Graph Convolution
Jiancheng Zhong, Pan Cui, Zuohang Qu, Liuping Wang, Qiu Xiao, Yihong Zhu |
ISBRA | 1 |
| 2022 | Drug response prediction using graph representation learning and Laplacian feature selectionabstractBACKGROUND: Knowing the responses of a patient to drugs is essential to make personalized medicine practical. Since the current clinical drug response experiments are time-consuming and expensive, utilizing human genomic information and drug molecular characteristics to predict drug responses is of urgent importance. Although a variety of computational drug response prediction methods have been proposed, their effectiveness is still not satisfying. RESULTS: In this study, we propose a method called LGRDRP (Learning Graph Representation for Drug Response Prediction) to predict cell line-drug responses. At first, LGRDRP constructs a heterogeneous network integrating multiple kinds of information: cell line miRNA expression profiles, drug chemical structure similarity, gene-gene interaction, cell line-gene interaction and known cell line-drug responses. Then, for each cell line, learning graph representation and Laplacian feature selection are combined to obtain network topology features related to the cell line. The learning graph representation method learns network topology structure features, and the Laplacian feature selection method further selects out some most important ones from them. Finally, LGRDRP trains an SVM model to predict drug responses based on the selected features of the known cell line-drug responses. Our five-fold cross-validation results show that LGRDRP is significantly superior to the art-of-the-state methods in the measures of the average area under the receiver operating characteristics curve, the average area under the precision-recall curve and the recall rate of top-k predicted sensitive cell lines. CONCLUSIONS: Our results demonstrated that the usage of multiple types of information about cell lines and drugs, the learning graph representation method, and the Laplacian feature selection is useful to the improvement of performance in predicting drug responses. We believe that such an approach would be easily extended to similar problems such as miRNA-disease relationship inference. Minzhu Xie, Xiaowen Lei, Jiancheng Zhong, Jianxing Ouyang, Guijing Li |
BMC Bioinform. | 3 |
| 2021 | Proteoform characterization based on top-down mass spectrometryabstractProteins are dominant executors of living processes. Compared to genetic variations, changes in the molecular structure and state of a protein (i.e. proteoforms) are more directly related to pathological changes in diseases. Characterizing proteoforms involves identifying and locating primary structure alterations (PSAs) in proteoforms, which is of practical importance for the advancement of the medical profession. With the development of mass spectrometry (MS) technology, the characterization of proteoforms based on top-down MS technology has become possible. This type of method is relatively new and faces many challenges. Since the proteoform identification is the most important process in characterizing proteoforms, we comprehensively review the existing proteoform identification methods in this study. Before identifying proteoforms, the spectra need to be preprocessed, and protein sequence databases can be filtered to speed up the identification. Therefore, we also summarize some popular deconvolution algorithms, various filtering algorithms for improving the proteoform identification performance and various scoring methods for localizing proteoforms. Moreover, commonly used methods were evaluated and compared in this review. We believe our review could help researchers better understand the current state of the development in this field and design new efficient algorithms for the proteoform characterization. Jiancheng Zhong, Yusui Sun, Minzhu Xie, Wei Peng 0004, Chushu Zhang, Fang-Xiang Wu, Jianxin Wang 0001 |
Briefings Bioinform. | 1 |
| 2021 | A novel essential protein identification method based on PPI networks and gene expression dataabstractBACKGROUND: Some proposed methods for identifying essential proteins have better results by using biological information. Gene expression data is generally used to identify essential proteins. However, gene expression data is prone to fluctuations, which may affect the accuracy of essential protein identification. Therefore, we propose an essential protein identification method based on gene expression and the PPI network data to calculate the similarity of "active" and "inactive" state of gene expression in a cluster of the PPI network. Our experiments show that the method can improve the accuracy in predicting essential proteins. RESULTS: In this paper, we propose a new measure named JDC, which is based on the PPI network data and gene expression data. The JDC method offers a dynamic threshold method to binarize gene expression data. After that, it combines the degree centrality and Jaccard similarity index to calculate the JDC score for each protein in the PPI network. We benchmark the JDC method on four organisms respectively, and evaluate our method by using ROC analysis, modular analysis, jackknife analysis, overlapping analysis, top analysis, and accuracy analysis. The results show that the performance of JDC is better than DC, IC, EC, SC, BC, CC, NC, PeC, and WDC. We compare JDC with both NF-PIN and TS-PIN methods, which predict essential proteins through active PPI networks constructed from dynamic gene expression. CONCLUSIONS: We demonstrate that the new centrality measure, JDC, is more efficient than state-of-the-art prediction methods with same input. The main ideas behind JDC are as follows: (1) Essential proteins are generally densely connected clusters in the PPI network. (2) Binarizing gene expression data can screen out fluctuations in gene expression profiles. (3) The essentiality of the protein depends on the similarity of "active" and "inactive" state of gene expression in a cluster of the PPI network. Jiancheng Zhong, Wei Peng 0004, Minzhu Xie, Yusui Sun, Qiang Tang 0014, Qiu Xiao, Jiahong Yang 0001 |
BMC Bioinform. | 1 |
| 2019 | Identifying Human Essential Genes by Network Embedding Protein-Protein Interaction Network
Wei Dai 0012, Wei Peng 0004, Jiancheng Zhong, Yongjiang Li |
ISBRA | 4 |
| 2019 | Improving Identification of Essential Proteins by a Novel Ensemble Method
Wei Dai 0012, Xia Li 0004, Wei Peng 0004, Jurong Song, Jiancheng Zhong, Jianxin Wang 0001 |
ISBRA | 5 |
| 2017 | Protein Inference from the Integration of Tandem MS Data and Interactome NetworksabstractSince proteins are digested into a mixture of peptides in the preprocessing step of tandem mass spectrometry (MS), it is difficult to determine which specific protein a shared peptide belongs to. In recent studies, besides tandem MS data and peptide identification information, some other information is exploited to infer proteins. Different from the methods which first use only tandem MS data to infer proteins and then use network information to refine them, this study proposes a protein inference method named TMSIN, which uses interactome networks directly. As two interacting proteins should co-exist, it is reasonable to assume that if one of the interacting proteins is confidently inferred in a sample, its interacting partners should have a high probability in the same sample, too. Therefore, we can use the neighborhood information of a protein in an interactome network to adjust the probability that the shared peptide belongs to the protein. In TMSIN, a multi-weighted graph is constructed by incorporating the bipartite graph with interactome network information, where the bipartite graph is built with the peptide identification information. Based on multi-weighted graphs, TMSIN adopts an iterative workflow to infer proteins. At each iterative step, the probability that a shared peptide belongs to a specific protein is calculated by using the Bayes' law based on the neighbor protein support scores of each protein which are mapped by the shared peptides. We carried out experiments on yeast data and human data to evaluate the performance of TMSIN in terms of ROC, q-value, and accuracy. The experimental results show that AUC scores yielded by TMSIN are 0.742 and 0.874 in yeast dataset and human dataset, respectively, and TMSIN yields the maximum number of true positives when q-value less than or equal to 0.05. The overlap analysis shows that TMSIN is an effective complementary approach for protein inference. Jiancheng Zhong, Jianxin Wang 0001, Zhen Zhang 0024, Min Li 0007, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2016 | Sprites: detection of deletions from sequencing data by re-aligning split readsabstractMOTIVATION: Advances of next generation sequencing technologies and availability of short read data enable the detection of structural variations (SVs). Deletions, an important type of SVs, have been suggested in association with genetic diseases. There are three types of deletions: blunt deletions, deletions with microhomologies and deletions with microsinsertions. The last two types are very common in the human genome, but they pose difficulty for the detection. Furthermore, finding deletions from sequencing data remains challenging. It is highly appealing to develop sensitive and accurate methods to detect deletions from sequencing data, especially deletions with microhomology and deletions with microinsertion. RESULTS: We present a novel method called Sprites (SPlit Read re-alIgnment To dEtect Structural variants) which finds deletions from sequencing data. It aligns a whole soft-clipping read rather than its clipped part to the target sequence, a segment of the reference which is determined by spanning reads, in order to find the longest prefix or suffix of the read that has a match in the target sequence. This alignment aims to solve the problem of deletions with microhomologies and deletions with microinsertions. Using both simulated and real data we show that Sprites performs better on detecting deletions compared with other current methods in terms of F-score. AVAILABILITY AND IMPLEMENTATION: Sprites is open source software and freely available at https://github.com/zhangzhen/sprites CONTACT: [email protected] data: Supplementary data are available at Bioinformatics online. Zhen Zhang 0024, Jianxin Wang 0001, Jiancheng Zhong, Jun Wang 0153, Fang-Xiang Wu, Yi Pan 0001 |
Bioinform. | 5 |
| 2015 | An efficient method to identify essential proteins for different species by integrating protein subcellular localization informationabstractEssential proteins are indispensable to maintain life activities in living organisms, and play important roles in the studies of pathology, synthetic biology, and drug design. Many computational methods are employed to identify essential proteins from Protein-protein Interaction Networks (PINs). In this paper, considering the different importance of protein-protein interactions which take place in different subcellular compartments, a Compartment Importance Centrality (CIC) method is proposed to detect essential proteins by integrating protein subcellular localization information. The experiments were carried on four species (Saccharomyces cerevisiae, Homo sapiens, Mus musculus and Drosophila melanogaster), and the performance of CIC was compared with other centrality methods, including the centrality methods solely based on topology and the ones combining both topology and other biological knowledge. The results show that CIC method has better performance to predict essential protein on four species. Furthermore, different from methods which overfits with the features of essential proteins of one species and may perform poor for other species, CIC has a wide applicable scope to identify essential proteins for different species. Xiaoqing Peng, Jianxin Wang 0001, Jiancheng Zhong, Yi Pan 0001 |
BIBM | 3 |
| 2015 | ClusterViz: A Cytoscape APP for Cluster Analysis of Biological NetworkabstractCluster analysis of biological networks is one of the most important approaches for identifying functional modules and predicting protein functions. Furthermore, visualization of clustering results is crucial to uncover the structure of biological networks. In this paper, ClusterViz, an APP of Cytoscape 3 for cluster analysis and visualization, has been developed. In order to reduce complexity and enable extendibility for ClusterViz, we designed the architecture of ClusterViz based on the framework of Open Services Gateway Initiative. According to the architecture, the implementation of ClusterViz is partitioned into three modules including interface of ClusterViz, clustering algorithms and visualization and export. ClusterViz fascinates the comparison of the results of different algorithms to do further related analysis. Three commonly used clustering algorithms, FAG-EC, EAGLE and MCODE, are included in the current version. Due to adopting the abstract interface of algorithms in module of the clustering algorithms, more clustering algorithms can be included for the future use. To illustrate usability of ClusterViz, we provided three examples with detailed steps from the important scientific articles, which show that our tool has helped several research teams do their research work on the mechanism of the biological networks. Jianxin Wang 0001, Jiancheng Zhong, Gang Chen 0010, Min Li 0007, Fang-Xiang Wu, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2014 | Predicting Essential Proteins Basedon Weighted Degree CentralityabstractEssential proteins are vital for an organism's viability under a variety of conditions. There are many experimental and computational methods developed to identify essential proteins. Computational prediction of essential proteins based on the global protein-protein interaction (PPI) network is severely restricted because of the insufficiency of the PPI data, but fortunately the gene expression profiles help to make up the deficiency. In this work, Pearson correlation coefficient (PCC) is used to bridge the gap between PPI and gene expression data. Based on PCC and edge clustering coefficient (ECC), a new centrality measure, i.e., the weighted degree centrality (WDC), is developed to achieve the reliable prediction of essential proteins. WDC is employed to identify essential proteins in the yeast PPI and e-Coli networks in order to estimate its performance. For comparison, other prediction technologies are also performed to identify essential proteins. Some evaluation methods are used to analyze the results from various prediction approaches. The prediction results and comparative analyses are shown in the paper. Furthermore, the parameter λ in the method WDC will be analyzed in detail and an optimal λ value will be found. Based on the optimal λ value, the differentiation of WDC and another prediction method PeC is discussed. The analyses prove that WDC outperforms other methods including DC, BC, CC, SC, EC, IC, NC, and PeC. At the same time, the analyses also mean that it is an effective way to predict essential proteins by means of integrating different data sources. Xiwei Tang, Jianxin Wang 0001, Jiancheng Zhong, Yi Pan 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |