Jingpu Zhang

dblp:211/4362 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Multi-Modality and Multi-Grained Transformer for Accurate Radiology Report Generation
Hongzhao Li, Liangzhi Zhang, Xiangrong Zhong, Jingpu Zhang, Shupan Li
ICIC (27)4
2022 Predicting circRNA-drug sensitivity associations via graph attention auto-encoder
abstract
BACKGROUND: Circular RNAs (circRNAs) play essential roles in cancer development and therapy resistance. Many studies have shown that circRNA is closely related to human health. The expression of circRNAs also affects the sensitivity of cells to drugs, thereby significantly affecting the efficacy of drugs. However, traditional biological experiments are time-consuming and expensive to validate drug-related circRNAs. Therefore, it is an important and urgent task to develop an effective computational method for predicting unknown circRNA-drug associations. RESULTS: In this work, we propose a computational framework (GATECDA) based on graph attention auto-encoder to predict circRNA-drug sensitivity associations. In GATECDA, we leverage multiple databases, containing the sequences of host genes of circRNAs, the structure of drugs, and circRNA-drug sensitivity associations. Based on the data, GATECDA employs Graph attention auto-encoder (GATE) to extract the low-dimensional representation of circRNA/drug, effectively retaining critical information in sparse high-dimensional features and realizing the effective fusion of nodes' neighborhood information. Experimental results indicate that GATECDA achieves an average AUC of 89.18% under 10-fold cross-validation. Case studies further show the excellent performance of GATECDA. CONCLUSIONS: Many experimental results and case studies show that our proposed GATECDA method can effectively predict the circRNA-drug sensitivity associations.
Lei Deng 0002, Yurong Qian, Jingpu Zhang
BMC Bioinform.4
2021 In silico drug repositioning based on the integration of chemical, genomic and pharmacological spaces
abstract
BACKGROUND: Drug repositioning refers to the identification of new indications for existing drugs. Drug-based inference methods for drug repositioning apply some unique features of drugs for new indication prediction. Complementary information is provided by these different features. It is therefore necessary to integrate these features for more accurate in silico drug repositioning. RESULTS: In this study, we collect 3 different types of drug features (i.e., chemical, genomic and pharmacological spaces) from public databases. Similarities between drugs are separately calculated based on each of the features. We further develop a fusion method to combine the 3 similarity measurements. We test the inference abilities of the 4 similarity datasets in drug repositioning under the guilt-by-association principle. Leave-one-out cross-validations show the integrated similarity measurement IntegratedSim receives the best prediction performance, with the highest AUC value of 0.8451 and the highest AUPR value of 0.2201. Case studies demonstrate IntegratedSim produces the largest numbers of confirmed predictions in most cases. Moreover, we compare our integration method with 3 other similarity-fusion methods using the datasets in our study. Cross-validation results suggest our method improves the prediction accuracy in terms of AUC and AUPR values. CONCLUSIONS: Our study suggests that the 3 drug features used in our manuscript are valuable information for drug repositioning. The comparative results indicate that integration of the 3 drug features would improve drug-disease association prediction. Our study provides a strategy for the fusion of different drug features for in silico drug repositioning.
Hailin Chen, Zuping Zhang 0001, Jingpu Zhang
BMC Bioinform.3
2021 MSCFS: inferring circRNA functional similarity based on multiple data sources
abstract
BACKGROUND: More and more evidence shows that circRNA plays an important role in various biological processes and human health. Therefore, inferring the circRNA's potential functions and obtaining circRNA functional similarity has become more and more significant. However, there is no effective approach to explore the functional similarity of circRNAs. METHODS: In this paper, we propose a new approach, called MSCFS, to calculate the functional similarity of circRNA by integrating multiple data sources. We combine circRNA-disease association, circRNA-gene-Gene Ontology association, and circRNA sequence information to explore the functional similarity of circRNA. Firstly, we employ different learning representation methods from three data sources to establish three circRNA functional similarity networks. Then we integrate the three networks to obtain the final circRNA functional similarity. RESULTS: We utilize circRNA-miRNA association similarity and circRNA co-expression similarity to evaluate the performance of MSCFS. The results show a positive correlation with miRNA association ([Formula: see text]) and circRNA co-expression similarity ([Formula: see text]). Finally, we construct a circRNA functional similarity network and perform case analysis. The result shows our method can be applied to infer new potential functions of circRNA and other associations. CONCLUSIONS: MSCFS combines multiple data sources related to circRNA functions. Correlation analysis and case analyses prove that MSCFS is a useful method to explore circRNA functional similarity.
Liang Shu, Xinxu Yuan, Jingpu Zhang, Lei Deng 0002
BMC Bioinform.4
2021 LDAH2V: Exploring Meta-Paths Across Multiple Networks for lncRNA-Disease Association Prediction
abstract
Accumulating evidence has demonstrated dysfunctions of long non-coding RNAs (lncRNAs) are involved in various complex human diseases. However, even today, the relationships between lncRNAs and diseases remain unknown in most cases. Developing effective computational approaches to identify potential lncRNA-disease associations has become a hot topic. Existing network-based approaches are usually focused on the intrinsic features of lncRNAs and diseases but ignore the heterogeneous information of biological networks. Considering the limitations in previous methods, we propose LDAH2V, an efficient computational framework for predicting potential lncRNA-disease associations. LDAH2V uses the HIN2Vec to calculate the meta-path and feature vector for each lncRNA-disease pair in the heterogeneous information network (HIN), which consists of lncRNA similarity network, disease similarity network, miRNA similarity network, and the associations between them. Then, a Gradient Boosting Tree (GBT) classifier to predict lncRNA-disease associations is built with the feature vectors. The results show that LDAH2V performs significantly better than the four existing state-of-the-art methods and gains an AUC of 0.97 in the 10-fold cross-validation test. Furthermore, case studies of colon cancer and ovarian cancer-related lncRNAs have been confirmed in related databases and medical literature.
Lei Deng 0002, Jingpu Zhang
IEEE ACM Trans. Comput. Biol. Bioinform.3
2020 Predict the Protein-protein Interaction between Virus and Host through Hybrid Deep Neural Network
abstract
Viral infection has been considered as a threat to human health for many years, where protein-protein interactions (PPIs) between viruses and hosts is involved. Researching the PPI between the virus and the host is conducive to understanding the mechanism of virus infection and the development of new drugs. Currently, most of the existing studies based on sequence only focus on extracting sequence features from original amino acid sequences, whereas the redundancy and noise of the features are neglected.In this paper, we employed Ll-regularized logistic regression to obtain efficacious sequence features related to PPIs without losing accuracy and generalization. A hybrid deep learning framework which combines convolutional neural network together with a long short term memory network to extract more hidden high-level features was designed to extract more latent features. As it is demonstrated in experiments results, the proposed framework is superior to the current advanced framework in both benchmark data and independent testing and is promising for identifying virus-host interactions.
Lei Deng 0002, Jiaojiao Zhao, Jingpu Zhang
BIBM3
2020 DeepciRGO: functional prediction of circular RNAs through hierarchical deep neural networks using heterogeneous network features
abstract
BACKGROUND: Circular RNAs (circRNAs) are special noncoding RNA molecules with closed loop structures. Compared with the traditional linear RNA, circRNA is more stable and not easily degraded. Many studies have shown that circRNAs are involved in the regulation of various diseases and cancers. Determining the functions of circRNAs in mammalian cells is of great significance for revealing their mechanism of action in physiological and pathological processes, diagnosis and treatment of diseases. However, determining the functions of circRNAs on a large scale is a challenging task because of the high experimental costs. RESULTS: In this paper, we present a hierarchical deep learning model, DeepciRGO, which can effectively predict gene ontology functions of circRNAs. We build a heterogeneous network containing circRNA co-expressions, protein-protein interactions and protein-circRNA interactions. The topology features of proteins and circRNAs are calculated using a novel representation learning approach HIN2Vec across the heterogeneous network. Then, a deep multi-label hierarchical classification model is trained with the topology features to predict the biological process function in the gene ontology for each circRNA. In particular, we manually curated a benchmark dataset containing 185 GO annotations for 62 circRNAs, namely, circRNA2GO-62. The DeepciRGO achieves promising performance on the circRNA2GO-62 dataset with a maximum F-measure of 0.412, a recall score of 0.400, and an accuracy of 0.425, which are significantly better than other state-of-the-art RNA function prediction methods. In addition, we demonstrate the considerable potential of integrating multiple interactions and association networks. CONCLUSIONS: DeepciRGO will be a useful tool for accurately annotating circRNAs. The experimental results show that integrating multi-source data can help to improve the predictive performance of DeepciRGO. Moreover, The model also can combine RNA structure and sequence information to further optimize predictive performance.
Lei Deng 0002, Jingpu Zhang
BMC Bioinform.4
2019 DKCirc2GO: Predicting Gene Ontology of circRNAs Using Dual KATZ Approach
abstract
Circular RNAs(circRNAs) have been demonstrated to play significant biological roles in many human biological processes such as competing endogenous RNAs or miRNA sponges, regulating gene transcription, translating proteins and others. Inferring the functions of circRNAs is an important strategy for understanding disease pathogenesis at the molecular level. However, most circRNAs have not been functionally characterized. Developing effective computational approaches to identify potential functions of circRNA has become a hot topic. In this paper, we propose an integrated model, DKCirc2GO, to infer the gene ontology (GO) functions of circRNAs by integrating multiple data sources, including the expression profiles of circRNAs, circRNA-protein associations, protein-protein interactions (PPI), protein-GO associations, GO-GO semantic similarity. The C2P global network is constructed by integrating three heterogeneous networks: circRNA-circRNA similarity network, circRNA-protein association network and protein-protein interaction network. The P2G global network is constructed by integrating three heterogeneous networks: protein-protein interaction network, protein-GO association network and GO-GO semantic similarity network. The KATZ measure is then employed respectively to calculate similarities in the two global network. The circRNAs-protein-GO(cpg) association network is constructed based on the KATZ similarity scores. Finally, we annotate circRNAs with Gene Ontology (GO) terms of their neighboring protein-GO terms. The experimental results show that DKCirc2GO has a significantly better performance than existing state-of-the-art methods in terms of F-max.
Lei Deng 0002, Jingpu Zhang
BIBM4
2019 A Protein Complex Identification Algorithm Based on essential protein
abstract
In this paper, we develop a novel protein complex identification algorithm from the angle of point (SCMA). First, we select the protein with high degree and high connection strength as seed to form the preliminary cores. Then, the preliminary cores are filtered according to the density of complex core to obtain the unique core. Finally, the protein complexes are generated by identifying attachment proteins through second order connection strength for each core. We compare our algorithm by using OS (matching degree), F-measure, Coverage rate, P-value metrics with the existing algorithms. The results show that our algorithm is effective, it can identify protein complexes more accurately.
Junmin Zhao, Jingpu Zhang
BIBM2
2019 Integrating Multiple Heterogeneous Networks for Novel LncRNA-Disease Association Inference
abstract
Accumulating experimental evidence has indicated that long non-coding RNAs (lncRNAs) are critical for the regulation of cellular biological processes implicated in many human diseases. However, only relatively few experimentally supported lncRNA-disease associations have been reported. Developing effective computational methods to infer lncRNA-disease associations is becoming increasingly important. Current network-based algorithms typically use a network representation to identify novel associations between lncRNAs and diseases. But these methods are concentrated on specific entities of interest (lncRNAs and diseases) and they do not allow to consider networks with more than two types of entities. Considering the limitations in previous computational methods, we develop a new global network-based framework, LncRDNetFlow, to prioritize disease-related lncRNAs. LncRDNetFlow utilizes a flow propagation algorithm to integrate multiple networks based on a variety of biological information including lncRNA similarity, protein-protein interactions, disease similarity, and the associations between them to infer lncRNA-disease associations. We show that LncRDNetFlow performs significantly better than the existing state-of-the-art approaches in cross-validation. To further validate the reproducibility of the performance, we use the proposed method to identify the related lncRNAs for ovarian cancer, glioma, and cervical cancer. The results are encouraging. Many predicted lncRNAs in the top list have been verified by the biological studies.
Jingpu Zhang, Zuping Zhang 0001, Zhigang Chen 0001, Lei Deng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.1
2019 KATZLGO: Large-Scale Prediction of LncRNA Functions by Using the KATZ Measure Based on Multiple Networks
abstract
Aggregating evidences have shown that long non-coding RNAs (lncRNAs) generally play key roles in cellular biological processes such as epigenetic regulation, gene expression regulation at transcriptional and post-transcriptional levels, cell differentiation, and others. However, most lncRNAs have not been functionally characterized. There is an urgent need to develop computational approaches for function annotation of increasing available lncRNAs. In this article, we propose a global network-based method, KATZLGO, to predict the functions of human lncRNAs at large scale. A global network is constructed by integrating three heterogeneous networks: lncRNA-lncRNA similarity network, lncRNA-protein association network, and protein-protein interaction network. The KATZ measure is then employed to calculate similarities between lncRNAs and proteins in the global network. We annotate lncRNAs with Gene Ontology (GO) terms of their neighboring protein-coding genes based on the KATZ similarity scores. The performance of KATZLGO is evaluated on a manually annotated lncRNA benchmark and a protein-coding gene benchmark with known function annotations. KATZLGO significantly outperforms state-of-the-art computational method both in maximum F-measure and coverage. Furthermore, we apply KATZLGO to predict functions of human lncRNAs and successfully map 12,318 human lncRNA genes to GO terms.
Zuping Zhang 0001, Jingpu Zhang, Yongjun Tang, Lei Deng 0002
IEEE ACM Trans. Comput. Biol. Bioinform.2
2018 Exploring Disease Similarity by Integrating Multiple Data Sources
Lei Deng 0002, Danyi Ye, Junmin Zhao, Jingpu Zhang
BIBM4
2018 Ontological function annotation of long non-coding RNAs through hierarchical multi-label classification
abstract
Motivation: Long non-coding RNAs (lncRNAs) are an enormous collection of functional non-coding RNAs. Over the past decades, a large number of novel lncRNA genes have been identified. However, most of the lncRNAs remain function uncharacterized at present. Computational approaches provide a new insight to understand the potential functional implications of lncRNAs. Results: Considering that each lncRNA may have multiple functions and a function may be further specialized into sub-functions, here we describe NeuraNetL2GO, a computational ontological function prediction approach for lncRNAs using hierarchical multi-label classification strategy based on multiple neural networks. The neural networks are incrementally trained level by level, each performing the prediction of gene ontology (GO) terms belonging to a given level. In NeuraNetL2GO, we use topological features of the lncRNA similarity network as the input of the neural networks and employ the output results to annotate the lncRNAs. We show that NeuraNetL2GO achieves the best performance and the overall advantage in maximum F-measure and coverage on the manually annotated lncRNA2GO-55 dataset compared to other state-of-the-art methods. Availability and implementation: The source code and data are available at http://denglab.org/NeuraNetL2GO/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Jingpu Zhang, Zuping Zhang 0001, Lei Deng 0002
Bioinform.1
2017 BiRWLGO: A global network-based strategy for lncRNA function annotation using bi-random walk
abstract
A large number of long non-coding RNAs (lncRNAs) have been identified over the past decades. Accumulating evidence proves that lncRNAs play key roles in various biological processes. However, the majority of the lncRNAs have not been functionally characterized. The annotation of lncRNA functions has become an area of focus in the fields of biology and bioinformatics. In this paper, we develop a global network-based strategy, BiRWLGO, to predict probable functions for lncRNAs at large scale. In BiRWLGO, we first build a global network consisting of three networks: lncRNA-lncRNA similarity network, lncRNA-protein interaction network and protein-protein interaction network. Then the bi-random walk algorithm is applied to explore similarities between lncRNAs and proteins. The functions of a query lncRNA can be obtained according to the Gene Ontology (GO) terms of its neighboring proteins. We compare the performance of BiRWLGO with other state-of-the-art approaches on a manually annotated lncRNA benchmark with known GO terms. As a result, BiRWLGO achieves the best predictive performance in terms of both maximum F-measure (Fmax) and coverage. Moreover, we demonstrate that integrating the protein-protein interactions can help improve the predictive performance of lncRNA functions.
Jingpu Zhang, Shuai Zou, Lei Deng 0002
BIBM1