Fuhao Zhang

dblp:75/7816 · DBLP profile ↗
← Back
26ranked-venue papers
6as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 20 · 4 first-author · 17 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 EssLM-MoE: A mixture-of-experts-enhanced framework for protein essentiality prediction using fused protein language models
Min Zeng 0004, Qianpei Liu, Wenkang Wang, Fuhao Zhang, Fei Guo 0001, Min Li 0007
Neurocomputing6
2025 piRNA-Disease Association Prediction via the Fusion of BERT and Graph Neural Network
abstract
PIWI-interacting RNAs (piRNAs) are key regulators in multiple disease processes, yet predicting their associations with diseases remains difficult due to data sparsity and feature complexity. To address this challenge, a multimodal deep learning framework, PDFBG, was proposed to formulate piRNA-disease association prediction as a three-class classification problem (positive, negative, and unrelated). PDFBG constructed a network using the secondary structure of piRNA and employed the BERT language model to generate feature embeddings for each nucleotide. These embed dings are then fused with topologi-cal structure information via a Graph Convolutional Network (GCN). In terms of diseases, this study constructed disease fea-ture embeddings based on MeSH semantic similarity, Gaussian interaction profile (GIP) kernel and disease-gene associations. Finally, an MLP classifier was used for end - to-end prediction. Experiments on benchmark datasets shown that PDFBG outper-forms state-of-the-art methods, achieving an accuracy of 0.9649 and AUC of 0.9761. Ablation studies confirmed the critical role of RNA-BERT and structural fusion. This work provided a robust and interpretable tool for exploring complex piRNA- disease relationships, while establishing a new paradigm for piRNA research. The code and dataset can be obtained from https://github.com/He-EugeneIPDFBG.
Yuchen Zhang 0003, Yuliang Pan, Fuhao Zhang
BIBM5
2025 DP-GPT: GPT-Driven Gene Text Feature Embedding Fused with Gene Expression Data for Depression Prediction
abstract
In recent years, with the improvement of living standards, the prevalence of depression has been steadily increasing, making it a growing public health concern. Gene expression data can reveal links between genes and diseases. Studies have shown that gene expression in depression patients differs significantly from healthy individuals, offering potential for early detection. However, existing methods often depend on selecting differentially expressed genes, which may overlook important signals from other genes and are vulnerable to batch effects, limiting model generalization. To address these limitations, we propose DP-GPT, a GPT-driven gene text feature embedding framework fused with gene expression data for depression prediction. DP-GPT integrates gene expression data with features extracted by GPT. Specifically, gene names and summaries are obtained from the NCBI Gene database, then embedded using GPT to generate feature vectors. These are fused with sample gene expression data and fed into a classifier for prediction. Extensive experiments show that DP-GPT achieves superior performance in depression prediction. The source code can be obtained from https://github.com/CSUBioGroup/DP-GPT.
Min Zeng 0004, Junyu Gao 0004, Qianpei Liu, Fuhao Zhang, Ruiqing Zheng, Min Li 0007
BIBM5
2025 Accurate Residue-Level Prediction of Linear Interacting Peptides with Auxiliary Multi-Scale Anchor Detection
abstract
Intrinsically disordered regions (IDRs) play essential roles in cellular signaling and regulation, primarily mediating interactions through short peptide motifs. Linear Interacting Peptides (LIPs) are a recently defined class of bindingassociated IDRs that undergo disorder-to-order transitions upon binding. LIPs encompass several well-characterized subclasses, including molecular recognition features (MoRFs) and short linear motifs (SLiMs). However, many current predictors focus exclusively on MoRFs prediction and exhibit limited effectiveness in identifying the broader class of LIPs. To address this limitation, we propose LipPredictor, the first deep learning framework specifically designed for LIP prediction. The model incorporates protein language model embeddings as input to a shared feature extraction module, which is composed of convolutional neural networks and a multi-head attention layer. It further adopts a dual-branch architecture that couples residue-level classification with auxiliary multi-scale anchor detection, enabling accurate identification of LIPs across diverse segment lengths through joint training. These results demonstrate its robustness and effectiveness in predicting LIPs. The source code can be obtained from https://github.com/Chenxi-Xia/LipPredictor.
Fuhao Zhang, Chenxi Xia, Mingxin Dong, Min Zeng 0004, Jian Zhang 0020, Min Li 0007
BIBM1
2025 PMA-Net: Parallel Mixed Attention Network for Predicting Intracranial Aneurysm Rupture Risk
abstract
Intracranial aneurysm (IA) is a life-threatening condition with high morbidity and mortality rates. Since preventive treatment of IA also carries inherent risks, accurate rupture risk prediction is crucial for optimizing clinical decisionmaking. However, current methods for IA rupture risk prediction remain limited in accuracy, generalization, and interpretability. To address the above issues, this study proposes PMA-Net, a novel framework based on a parallel mixed attention mechanism, to efficiently and accurately predict IA rupture risk from Computed Tomography Angiography (CTA) images. By jointly analyzing aneurysms and their surrounding regions, the model captures interdependent imaging features associated with rupture risk. A multi-branch attention module extracts both coarse-and fine-grained features, while clinical information is integrated to enhance prediction performance. In addition, interpretable visualization methods were employed to enhance the model interpretability. Experimental results show that the proposed method achieves superior accuracy, improving by 2.22 % and 1.08 % on internal and external test sets, reaching 92.22 % and 84.78 %, respectively.
Fuhao Zhang, Jingfeng Jiang, Nan Mu
ICTAI1
2025 DDLB: Using the Protein Language Model and Hierarchical Architecture to Improve Disordered Lipid-Binding Residues Prediction
Chaojin Wu, Fuhao Zhang, Pengzhen Jia, Min Zeng 0004, Min Li 0007
ISBRA (1)2
2025 Joint Geometric Self-Attention and Boundary-Aware Search for High-Precision Intracranial Aneurysm Mesh Segmentation
Fuhao Zhang, Ling Wang 0005, Dapeng Chen, Jinshan Tang, Jingfeng Jiang, Nan Mu
SMC2
2025 A Transformer-Based Dual-Branch Mesh Convolutional Neural Network for Aortic Dissection Segmentation
abstract
Aortic dissection (AD) is a life-threatening condition caused by a tear in the aortic intima, allowing blood to enter the vessel wall and form a false lumen. Due to its high mortality rate, timely diagnosis and precise treatment are critical. Clinical diagnosis and treatment of AD rely heavily on accurate 3D vascular image segmentation. To address existing methods’ low segmentation accuracy and insufficient geometric detail preservation, this paper proposes a Transformer-based Dual-Branch Mesh Segmentation Network (TD-MSeg) for AD. This network employs a mesh-based self-attention mechanism to retain vascular geometric details while adopting a dual-branch decoder to effectively fuse features and model long-range dependencies. Specifically, TD-MSeg incorporates three key components: a Hierarchical Mesh Transformer (HMT) module that enhances feature modeling of critical anatomical structures (e.g., intimal tears), a dual-branch decoder that facilitates collaborative optimization of multi-scale local and global features, and a mesh label refinement module that uses a wide-path exploration algorithm to eliminate deformation artifacts and improve spatial label continuity. Moreover, experiments on two AD mesh segmentation datasets demonstrate that the proposed TD-MSeg achieves a 6% improvement in accuracy compared to traditional models and significantly enhances the recognition of complex vascular structures, thereby providing high-precision 3D reconstruction support for endovascular surgical planning.
Fuhao Zhang, Ling Wang 0005, Dapeng Chen, Jinshan Tang, Jingfeng Jiang, Nan Mu
SMC2
2025 CREATE: a novel attention-based framework for efficient classification of transposable elements
abstract
Transposable elements (TEs) are DNA sequences that can move within a genome. They constitute a substantial portion of the eukaryotic genome and play essential roles in gene regulation and genome evolution. Accurate classification of these repetitive elements is crucial for investigating their potential impact on the genome. Over the past few decades, several alignment-based tools have been developed to annotate TE types. While these methods rely heavily on prior knowledge and are often computationally expensive, machine learning-based approaches have been proposed to overcome these limitations. However, most of these approaches fail to capture the multiscale features of TEs, resulting in suboptimal performance. Here, we propose a novel framework called CREATE, which simultaneously integrates the global pattern distribution and the local sequence profile of TEs using Convolutional neural networks and Recurrent neural nEtworks with an Attention mechanism for efficient TE classification. Due to the hierarchical structure of TE groups, we trained nine classifiers corresponding to parent nodes within the class hierarchy. We further applied a top-down hierarchical classification strategy to achieve a more complete classification of unknown TEs. Comprehensive experiments demonstrate that CREATE outperforms existing TE-type annotation methods and achieves superior performance in hierarchical classification tasks. In conclusion, CREATE exhibits great potential for improving the accuracy of TE annotation. The source code and demo data are available at https://github.com/yangqi-cs/CREATE.
Yingfu Wu, Meihong Gao, Fuhao Zhang, Xingyu Liao, Xuequn Shang 0001
Briefings Bioinform.6
2025 Leveraging protein language models for cross-variant CRISPR/Cas9 sgRNA activity prediction
abstract
MOTIVATION: Accurate prediction of single-guide RNA (sgRNA) activity is crucial for optimizing the CRISPR/Cas9 gene-editing system, as it directly influences the efficiency and accuracy of genome modifications. However, existing prediction methods mainly rely on large-scale experimental data of a single Cas9 variant to construct Cas9 protein (variants)-specific sgRNA activity prediction models, which limits their generalization ability and prediction performance across different Cas9 protein (variants), as well as their scalability to the continuously discovered new variants. RESULTS: In this study, we proposed PLM-CRISPR, a novel deep learning-based model that leverages protein language models to capture Cas9 protein (variants) representations for cross-variant sgRNA activity prediction. PLM-CRISPR uses tailored feature extraction modules for both sgRNA and protein sequences, incorporating a cross-variant training strategy and a dynamic feature fusion mechanism to effectively model their interactions. Extensive experiments demonstrate that PLM-CRISPR outperforms existing methods across datasets spanning seven Cas9 protein (variants) in three real-world scenarios, demonstrating its superior performance in handling data-scarce situations, including cases with few or no samples for novel variants. Comparative analyses with traditional machine learning and deep learning models further confirm the effectiveness of PLM-CRISPR. Additionally, motif analysis reveals that PLM-CRISPR accurately identifies high-activity sgRNA sequence patterns across diverse Cas9 protein (variants). Overall, PLM-CRISPR provides a robust, scalable, and generalizable solution for sgRNA activity prediction across diverse Cas9 protein (variants). AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/CSUBioGroup/PLM-CRISPR.
Yalin Hou, Ruiqing Zheng, Fuhao Zhang, Fei Guo 0001, Min Li 0007, Min Zeng 0004
Bioinform.4
2025 Neural dynamic fluid reconstruction technique for four-dimensional imaging of combustion flame based on deep learning
Fuhao Zhang, Zhiyin Ma, Can Gao, Gang Xun, Qingchun Lei
Eng. Appl. Artif. Intell.1
2024 ComLMEss: Combining multiple protein language models enables accurate essential protein prediction
abstract
Accurately predicting essential proteins is vital for comprehending organism survival, aiding in drug discovery, and informing strategies for treating diseases. While previous computational methods for essential protein prediction have predominantly focused on network-based approaches, recent advancements have seen rapid development in sequence-based prediction methods. However, existing sequence-based prediction methods tend to focus only on sequence-level features, ignoring other biological information at diverse levels. To make use of the diverse information across various biological levels, in this study, we introduce ComLMEss, a novel deep learning framework that combines three protein language models. ComLMEss integrates ProtTrans, ESMFold and OntoProtein, which contain different levels of biological information include protein sequence, conservation, structural, and functional information. ComLMEss employs convolutional neural networks and transformer structure to refine and contextualize the representations from three language models, enabling accurate and robust predictions. Experimental results demonstrate that ComLMEss consistently outperforms existing methods. Ablation studies confirm that the effectiveness of combining different language models focus on different biological information. All results underscore the potential of ComLMEss in essential protein prediction. The source code can be obtained at https://github.com/CSUBioGroup/ComLMEss.
Fuhao Zhang, Ruiqing Zheng, Fei Guo 0001, Min Li 0007, Min Zeng 0004
BIBM3
2024 A comprehensive review of protein-centric predictors for biomolecular interactions: from proteins to nucleic acids and beyond
abstract
Proteins interact with diverse ligands to perform a large number of biological functions, such as gene expression and signal transduction. Accurate identification of these protein-ligand interactions is crucial to the understanding of molecular mechanisms and the development of new drugs. However, traditional biological experiments are time-consuming and expensive. With the development of high-throughput technologies, an increasing amount of protein data is available. In the past decades, many computational methods have been developed to predict protein-ligand interactions. Here, we review a comprehensive set of over 160 protein-ligand interaction predictors, which cover protein-protein, protein-nucleic acid, protein-peptide and protein-other ligands (nucleotide, heme, ion) interactions. We have carried out a comprehensive analysis of the above four types of predictors from several significant perspectives, including their inputs, feature profiles, models, availability, etc. The current methods primarily rely on protein sequences, especially utilizing evolutionary information. The significant improvement in predictions is attributed to deep learning methods. Additionally, sequence-based pretrained models and structure-based approaches are emerging as new trends.
Pengzhen Jia, Fuhao Zhang, Chaojin Wu, Min Li 0007
Briefings Bioinform.2
2024 A comprehensive computational benchmark for evaluating deep learning-based protein function prediction approaches
abstract
Proteins play an important role in life activities and are the basic units for performing functions. Accurately annotating functions to proteins is crucial for understanding the intricate mechanisms of life and developing effective treatments for complex diseases. Traditional biological experiments struggle to keep pace with the growing number of known proteins. With the development of high-throughput sequencing technology, a wide variety of biological data provides the possibility to accurately predict protein functions by computational methods. Consequently, many computational methods have been proposed. Due to the diversity of application scenarios, it is necessary to conduct a comprehensive evaluation of these computational methods to determine the suitability of each algorithm for specific cases. In this study, we present a comprehensive benchmark, BeProf, to process data and evaluate representative computational methods. We first collect the latest datasets and analyze the data characteristics. Then, we investigate and summarize 17 state-of-the-art computational methods. Finally, we propose a novel comprehensive evaluation metric, design eight application scenarios and evaluate the performance of existing methods on these scenarios. Based on the evaluation, we provide practical recommendations for different scenarios, enabling users to select the most suitable method for their specific needs. All of these servers can be obtained from https://csuligroup.com/BEPROF and https://github.com/CSUBioGroup/BEPROF.
Wenkang Wang, Yunyan Shuai, Qiurong Yang, Fuhao Zhang, Min Zeng 0004, Min Li 0007
Briefings Bioinform.4
2023 DeepCellEss: cell line-specific essential protein prediction with attention-based interpretable deep learning
abstract
MOTIVATION: Protein essentiality is usually accepted to be a conditional trait and strongly affected by cellular environments. However, existing computational methods often do not take such characteristics into account, preferring to incorporate all available data and train a general model for all cell lines. In addition, the lack of model interpretability limits further exploration and analysis of essential protein predictions. RESULTS: In this study, we proposed DeepCellEss, a sequence-based interpretable deep learning framework for cell line-specific essential protein predictions. DeepCellEss utilizes a convolutional neural network and bidirectional long short-term memory to learn short- and long-range latent information from protein sequences. Further, a multi-head self-attention mechanism is used to provide residue-level model interpretability. For model construction, we collected extremely large-scale benchmark datasets across 323 cell lines. Extensive computational experiments demonstrate that DeepCellEss yields effective prediction performance for different cell lines and outperforms existing sequence-based methods as well as network-based centrality measures. Finally, we conducted some case studies to illustrate the necessity of considering specific cell lines and the superiority of DeepCellEss. We believe that DeepCellEss can serve as a useful tool for predicting essential proteins across different cell lines. AVAILABILITY AND IMPLEMENTATION: The DeepCellEss web server is available at http://csuligroup.com:8000/DeepCellEss. The source code and data underlying this study can be obtained from https://github.com/CSUBioGroup/DeepCellEss. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Min Zeng 0004, Fuhao Zhang, Fang-Xiang Wu, Min Li 0007
Bioinform.3
2023 A Deep Learning Framework for Predicting Protein Functions With Co-Occurrence of GO Terms
abstract
The understanding of protein functions is critical to many biological problems such as the development of new drugs and new crops. To reduce the huge gap between the increase of protein sequences and annotations of protein functions, many methods have been proposed to deal with this problem. These methods use Gene Ontology (GO) to classify the functions of proteins and consider one GO term as a class label. However, they ignore the co-occurrence of GO terms that is helpful for protein function prediction. We propose a new deep learning model, named DeepPFP-CO, which uses Graph Convolutional Network (GCN) to explore and capture the co-occurrence of GO terms to improve the protein function prediction performance. In this way, we can further deduce the protein functions by fusing the predicted propensity of the center function and its co-occurrence functions. We use Fmax and AUPR to evaluate the performance of DeepPFP-CO and compare DeepPFP-CO with state-of-the-art methods such as DeepGOPlus and DeepGOA. The computational results show that DeepPFP-CO outperforms DeepGOPlus and other methods. Moreover, we further analyze our model at the protein level. The results have demonstrated that DeepPFP-CO improves the performance of protein function prediction. DeepPFP-CO is available at https://csuligroup.com/DeepPFP/.
Min Li 0007, Fuhao Zhang, Min Zeng 0004, Yaohang Li
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 KiPT: Knowledge-injected Prompt Tuning for Event Detection
abstract
Event detection aims to detect events from the text by identifying and classifying event triggers (the most representative words). Most of the existing works rely heavily on complex downstream networks and require sufficient training data. Thus, those models may be structurally redundant and perform poorly when data is scarce. Prompt-based models are easy to build and are promising for few-shot tasks. However, current prompt-based methods may suffer from low precision because they have not introduced event-related semantic knowledge (e.g., part of speech, semantic correlation, etc.). To address these problems, this paper proposes a Knowledge-injected Prompt Tuning (KiPT) model. Specifically, the event detection task is formulated into a condition generation task. Then, knowledge-injected prompts are constructed using external knowledge bases, and a prompt tuning strategy is leveraged to optimize the prompts. Extensive experiments indicate that KiPT outperforms strong baselines, especially in few-shot scenarios.
Haochen Li 0001, Tong Mo, Hongcheng Fan, Fuhao Zhang, Weiping Li 0002
COLING6
2022 DeepLncLoc: a deep learning framework for long non-coding RNA subcellular localization prediction based on subsequence embedding
abstract
Long non-coding RNAs (lncRNAs) are a class of RNA molecules with more than 200 nucleotides. A growing amount of evidence reveals that subcellular localization of lncRNAs can provide valuable insights into their biological functions. Existing computational methods for predicting lncRNA subcellular localization use k-mer features to encode lncRNA sequences. However, the sequence order information is lost by using only k-mer features. We proposed a deep learning framework, DeepLncLoc, to predict lncRNA subcellular localization. In DeepLncLoc, we introduced a new subsequence embedding method that keeps the order information of lncRNA sequences. The subsequence embedding method first divides a sequence into some consecutive subsequences and then extracts the patterns of each subsequence, last combines these patterns to obtain a complete representation of the lncRNA sequence. After that, a text convolutional neural network is employed to learn high-level features and perform the prediction task. Compared with traditional machine learning models, popular representation methods and existing predictors, DeepLncLoc achieved better performance, which shows that DeepLncLoc could effectively predict lncRNA subcellular localization. Our study not only presented a novel computational model for predicting lncRNA subcellular localization but also introduced a new subsequence embedding method which is expected to be applied in other sequence-based prediction tasks. The DeepLncLoc web server is freely accessible at http://bioinformatics.csu.edu.cn/DeepLncLoc/, and source code and datasets can be downloaded from https://github.com/CSUBioGroup/DeepLncLoc.
Min Zeng 0004, Yifan Wu 0008, Chengqian Lu, Fuhao Zhang, Fang-Xiang Wu, Min Li 0007
Briefings Bioinform.4
2022 DeepDISOBind: accurate prediction of RNA-, DNA- and protein-binding intrinsically disordered residues with deep multi-task learning
abstract
Proteins with intrinsically disordered regions (IDRs) are common among eukaryotes. Many IDRs interact with nucleic acids and proteins. Annotation of these interactions is supported by computational predictors, but to date, only one tool that predicts interactions with nucleic acids was released, and recent assessments demonstrate that current predictors offer modest levels of accuracy. We have developed DeepDISOBind, an innovative deep multi-task architecture that accurately predicts deoxyribonucleic acid (DNA)-, ribonucleic acid (RNA)- and protein-binding IDRs from protein sequences. DeepDISOBind relies on an information-rich sequence profile that is processed by an innovative multi-task deep neural network, where subsequent layers are gradually specialized to predict interactions with specific partner types. The common input layer links to a layer that differentiates protein- and nucleic acid-binding, which further links to layers that discriminate between DNA and RNA interactions. Empirical tests show that this multi-task design provides statistically significant gains in predictive quality across the three partner types when compared to a single-task design and a representative selection of the existing methods that cover both disorder- and structure-trained tools. Analysis of the predictions on the human proteome reveals that DeepDISOBind predictions can be encoded into protein-level propensities that accurately predict DNA- and RNA-binding proteins and protein hubs. DeepDISOBind is available at https://www.csuligroup.com/DeepDISOBind/.
Fuhao Zhang, Bi Zhao, Min Li 0007, Lukasz A. Kurgan
Briefings Bioinform.1
2021 DeepPPF: A deep learning framework for predicting protein family
Shehu Mohammed Yusuf, Fuhao Zhang, Min Zeng 0004, Min Li 0007
Neurocomputing2
2021 A Deep Learning Framework for Gene Ontology Annotations With Sequence- and Network-Based Information
abstract
Knowledge of protein functions plays an important role in biology and medicine. With the rapid development of high-throughput technologies, a huge number of proteins have been discovered. However, there are a great number of proteins without functional annotations. A protein usually has multiple functions and some functions or biological processes require interactions of a plurality of proteins. Additionally, Gene Ontology provides a useful classification for protein functions and contains more than 40,000 terms. We propose a deep learning framework called DeepGOA to predict protein functions with protein sequences and protein-protein interaction (PPI) networks. For protein sequences, we extract two types of information: sequence semantic information and subsequence-based features. We use the word2vec technique to numerically represent protein sequences, and utilize a Bi-directional Long and Short Time Memory (Bi-LSTM) and multi-scale convolutional neural network (multi-scale CNN) to obtain the global and local semantic features of protein sequences, respectively. Additionally, we use the InterPro tool to scan protein sequences for extracting subsequence-based information, such as domains and motifs. Then, the information is plugged into a neural network to generate high-quality features. For the PPI network, the Deepwalk algorithm is applied to generate its embedding information of PPI. Then the two types of features are concatenated together to predict protein functions. To evaluate the performance of DeepGOA, several different evaluation methods and metrics are utilized. The experimental results show that DeepGOA outperforms DeepGO and BLAST.
Fuhao Zhang, Hong Song 0004, Min Zeng 0004, Fang-Xiang Wu, Yaohang Li, Yi Pan 0001, Min Li 0007
IEEE ACM Trans. Comput. Biol. Bioinform.1
2021 Deep Matrix Factorization Improves Prediction of Human CircRNA-Disease Associations
abstract
In recent years, more and more evidence indicates that circular RNAs (circRNAs) with covalently closed loop play various roles in biological processes. Dysregulation and mutation of circRNAs may be implicated in diseases. Due to its stable structure and resistance to degradation, circRNAs provide great potential to be diagnostic biomarkers. Therefore, predicting circRNA-disease associations is helpful in disease diagnosis. However, there are few experimentally validated associations between circRNAs and diseases. Although several computational methods have been proposed, precisely representing underlying features and grasping the complex structures of data are still challenging. In this paper, we design a new method, called DMFCDA (Deep Matrix Factorization CircRNA-Disease Association), to infer potential circRNA-disease associations. DMFCDA takes both explicit and implicit feedback into account. Then, it uses a projection layer to automatically learn latent representations of circRNAs and diseases. With multi-layer neural networks, DMFCDA can model the non-linear associations to grasp the complex structure of data. We assess the performance of DMFCDA using leave-one cross-validation and 5-fold cross-validation on two datasets. Computational results show that DMFCDA efficiently infers circRNA-disease associations according to AUC values, the percentage of precisely retrieved associations in various top ranks, and statistical comparison. We also conduct case studies to evaluate DMFCDA. All results show that DMFCDA provides accurate predictions.
Chengqian Lu, Min Zeng 0004, Fuhao Zhang, Fang-Xiang Wu, Min Li 0007, Jianxin Wang 0001
IEEE J. Biomed. Health Informatics3
2020 Protein-protein interaction site prediction through combining local and global features with deep neural networks
abstract
MOTIVATION: Protein-protein interactions (PPIs) play important roles in many biological processes. Conventional biological experiments for identifying PPI sites are costly and time-consuming. Thus, many computational approaches have been proposed to predict PPI sites. Existing computational methods usually use local contextual features to predict PPI sites. Actually, global features of protein sequences are critical for PPI site prediction. RESULTS: A new end-to-end deep learning framework, named DeepPPISP, through combining local contextual and global sequence features, is proposed for PPI site prediction. For local contextual features, we use a sliding window to capture features of neighbors of a target amino acid as in previous studies. For global sequence features, a text convolutional neural network is applied to extract features from the whole protein sequence. Then the local contextual and global sequence features are combined to predict PPI sites. By integrating local contextual and global sequence features, DeepPPISP achieves the state-of-the-art performance, which is better than the other competing methods. In order to investigate if global sequence features are helpful in our deep learning model, we remove or change some components in DeepPPISP. Detailed analyses show that global sequence features play important roles in DeepPPISP. AVAILABILITY AND IMPLEMENTATION: The DeepPPISP web server is available at http://bioinformatics.csu.edu.cn/PPISP/. The source code can be obtained from https://github.com/CSUBioGroup/DeepPPISP. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Min Zeng 0004, Fuhao Zhang, Fang-Xiang Wu, Yaohang Li, Jianxin Wang 0001, Min Li 0007
Bioinform.2
2020 PROBselect: accurate prediction of protein-binding residues from proteins sequences via dynamic predictor selection
abstract
MOTIVATION: Knowledge of protein-binding residues (PBRs) improves our understanding of protein-protein interactions, contributes to the prediction of protein functions and facilitates protein-protein docking calculations. While many sequence-based predictors of PBRs were published, they offer modest levels of predictive performance and most of them cross-predict residues that interact with other partners. One unexplored option to improve the predictive quality is to design consensus predictors that combine results produced by multiple methods. RESULTS: We empirically investigate predictive performance of a representative set of nine predictors of PBRs. We report substantial differences in predictive quality when these methods are used to predict individual proteins, which contrast with the dataset-level benchmarks that are currently used to assess and compare these methods. Our analysis provides new insights for the cross-prediction concern, dissects complementarity between predictors and demonstrates that predictive performance of the top methods depends on unique characteristics of the input protein sequence. Using these insights, we developed PROBselect, first-of-its-kind consensus predictor of PBRs. Our design is based on the dynamic predictor selection at the protein level, where the selection relies on regression-based models that accurately estimate predictive performance of selected predictors directly from the sequence. Empirical assessment using a low-similarity test dataset shows that PROBselect provides significantly improved predictive quality when compared with the current predictors and conventional consensuses that combine residue-level predictions. Moreover, PROBselect informs the users about the expected predictive quality for the prediction generated from a given input protein. AVAILABILITY AND IMPLEMENTATION: PROBselect is available at http://bioinformatics.csu.edu.cn/PROBselect/home/index. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fuhao Zhang, Jian Zhang 0020, Min Zeng 0004, Min Li 0007, Lukasz A. Kurgan
Bioinform.1
2020 Deep convolutional neural network for automatically segmenting acute ischemic stroke lesion in multi-modality MRI
Liangliang Liu 0001, Shaowu Chen, Fuhao Zhang, Fang-Xiang Wu, Yi Pan 0001, Jianxin Wang 0001
Neural Comput. Appl.3
2019 LncRNA-disease association prediction through combining linear and non-linear features with matrix factorization and deep learning techniques
abstract
Long non-coding RNAs (lncRNAs) are the foundation for understanding mechanisms of many human diseases. Considering the limited number of known experimentally verified associations between lncRNAs and diseases, it is appealing to develop accurate and effective computational methods to identify lncRNA-disease associations. Conventional matrix factorization-based methods cannot model complicated associations between lncRNAs and diseases. In this study, we propose a novel computational framework, through combining linear and non-linear features, which is used for lncRNA-disease association prediction. In our model, a conventional matrix factorization method is applied to extract linear features between lncRNAs and diseases. Deep learning techniques (fully connected layers) are applied to extract nonlinear features between lncRNAs and diseases. Finally, linear and non-linear features are fused to improve predictive performance. Compared to previous studies, our model can take advantages of the combination of linear and non-linear features between lncRNAs and diseases, and thus can effectively identify potential lncRNA-disease associations. The results show that our method achieves state-of-the-art performance in the leave-one-out cross-validation. The source codes of our method can be found at https://github.com/CSUBioGroup/DMFLDA2.
Min Zeng 0004, Chengqian Lu, Fuhao Zhang, Zhangli Lu, Fang-Xiang Wu, Yaohang Li, Min Li 0007
BIBM3