Junfeng Xia

dblp:31/7388 · also Jun-Feng Xia · DBLP profile ↗
← Back
76ranked-venue papers
2as first author
39since 2021 · last 2026
0000-0003-3024-1705ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 73 · 1 first-author · 37 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HybridSeqNet: A Deep Learning Framework for Blood Pressure Estimation
Fei Wang 0095, Feiyu Yu, Xiujuan Lei, Fang-Xiang Wu, Yansen Su, Chun-Hou Zheng 0001, Junfeng Xia
ICIC (29)8
2026 Pathogenicity prediction for noncanonical splice-altering variants based on multimodal feature fusion
abstract
Splice-altering variants (SAVs) are the second most prevalent class of pathogenic genetic variants and are strongly associated with the occurrence and development of various diseases. However, current computational tools exhibit limited predictive capability beyond canonical GT-AG splice sites, making accurate assessment of noncanonical SAV pathogenicity a considerable challenge. To address this limitation, we developed MOSAIC (multimodal feature fusion for noncanonical splice-altering variants pathogenicity prediction), a deep learning framework designed for precise assessment of noncanonical SAV pathogenicity. MOSAIC integrates long-range contextual signals derived from a pretrained DNA language model, local sequence features captured from multi-scale convolutional neural networks, and functional annotations. By employing a transformer encoder and a gated fusion module, the model adaptively integrates these multimodal features. Benchmarking across multiple independent datasets demonstrated that MOSAIC consistently outperforms existing state-of-the-art methods, such as CADD and SpliceAI. It remains highly accurate and robust when evaluated on rare variants, gene-independent contexts, and the largest subset where all comparative methods yielded outputs. Furthermore, feature importance analysis revealed that long-range dependencies in DNA sequences and transformer-based integration were critical contributors to model performance. Interpretability analyses indicated that MOSAIC could identify key regulatory sequence motifs associated with transcription factors and RNA-binding proteins, offering mechanistic insight into how noncanonical SAVs disrupt splicing regulation and contribute to pathogenic processes. Overall, MOSAIC offers an accurate and interpretable framework for predicting the pathogenicity of noncanonical SAVs, thereby serving as a dependable computational tool for genetic diagnostics and precision medicine applications. MOSAIC source code and data are available at https://github.com/Lilab-genomics/MOSAIC.
Xingpeng Zhou, Xiongjian Luo, Yansen Su, Chun-Hou Zheng 0001, Junfeng Xia
Briefings Bioinform.9
2026 scSCCNIA: similarity matrix based contrastive clustering with neighbor information aggregation for single-cell RNA sequencing data
abstract
The development of single-cell RNA sequencing (scRNA-seq) technology provides unprecedented opportunities for elucidating cell heterogeneity and gene expression. Identifying and discovering cell types through cell clustering is a crucial step in analyzing scRNA-seq data. However, the high-dimensionality nature and frequent dropout events of the data raise great challenges for cell clustering. Here, we propose a novel contrastive clustering framework called scSCCNIA (Similarity-matrix-based Contrastive Clustering with Neighbor Information Aggregation), for the accurate identification of cell clusters from scRNA-seq data. scSCCNIA adopts a Laplacian filter to conduct neighbor information aggregation, constructs different graph views by using special un-shared parameters Siamese encoders for data augmentation, and learns the latent low-dimensional embedding representations via similarity-matrix-based contrastive learning. Comparative analyses of multiple scRNA-seq datasets from different platforms and with varying cell numbers demonstrate that scSCCNIA outperforms existing methods in terms of cell clustering and marker gene identification. Furthermore, scSCCNIA reveals the heterogeneity and functional specificity of various cell types through Gene Ontology terms and Kyoto Encyclopedia of Genes and Genomes enrichment analyses. Overall, scSCCNIA is an effective algorithm for learning latent features from scRNA-seq data, enhancing cell type identification accuracy and facilitating downstream analyses of scRNA-seq data.
Jing Wang 0057, Junfeng Xia, Yansen Su, Chun-Hou Zheng 0001
Briefings Bioinform.2
2026 AdPrST:An Adversarial Graph Deep Learning Pre-Clustering Framework for Deciphering Spatiotemporal Structures in Spatially Resolved Transcriptomics
abstract
Spatially Resolved Transcriptomics (SRT) has revolutionized our understanding of gene expression within tissue microenvironments, yet accurately deciphering spatiotemporal structures-encompassing spatial domain identification, trajectory inference, and pseudo-spatiotemporal map construction-in complex tissues remains a formidable challenge. AdPrST begins with a pre-clustering process on gene expression data to establish initial domain groupings. It then constructs dual-view graph structures using K-Nearest Neighbors (KNN) for local similarities and r-radius for broader spatial contexts. Through adversarial self-supervised contrast, leveraging Wasserstein distance-based GANs and contrastive learning, AdPrST generates robust low-dimensional embeddings for each view. These embeddings are fused via a dot-product attention mechanism, guided by pre-clustering labels, to achieve accurate spatial domain identification. Benchmarking across multiple datasets demonstrated AdPrST's superior performance over state-of-the-art methods, highlighting its potential to advance spatial transcriptomics research by elucidating spatial functional patterns and developmental trajectories. In particular, AdPrST excels in inferencing spatiotemporal structures, reconstructing developmental sequences and temporal features in complex tissues.
Shensi Huang, Jianping Zhao 0001, Junfeng Xia
IEEE Trans. Comput. Biol. Bioinform.5
2026 scMSAC Assigns Single-Cell Multi-Omics Data at the Multi-Modal Cluster via Subgraph Attention Autoencoder
abstract
Single-cell multi-omics sequencing represents an advanced technology capable of simultaneously measuring multiple omics data from the same cell. The joint clustering of single-cell multi-omics sequencing data enables a comprehensive depiction of cell states and uncovers intricate molecular mechanisms, holding immense significance in fields such as oncology, neurology, and developmental biology. However, the disparities in feature spaces across different omics layers and data noise present substantial challenges for achieving accurate clustering. To tackle these challenges, we introduce a novel clustering method for single-cell multi-omics data, termed scMSAC, which is grounded in a denoising subgraph attention autoencoder. The proposed method employs a weighted nearest neighbor graph strategy to ascertain the weights of multi-omics data, subsequently generating a similarity graph that holistically encapsulates intercellular connections through the weighted amalgamation of diverse omics perspectives. The scMSAC model captures the topological features of cells through the subgraph attention autoencoder, constructing relationships among cells. For the omics features extracted by the subgraph attention autoencoder, scMSAC incorporates an SCA (Spatial Channel Attention) mechanism for feature fusion to reduce the differences in feature spaces of different omics and achieve better clustering performance. Comparative experiments with various existing methods demonstrate that scMSAC has excellent clustering performance and performs well in detecting rare cell types and differential expression analysis.
Jing Wang 0057, Weijie Cai, Dayu Tan, Yun Ding, Junfeng Xia, Yansen Su, Chun-Hou Zheng 0001
IEEE Trans. Comput. Biol. Bioinform.6
2026 Large-Scale Multimodality via Dual-Path Cooperative Feature Fusion Strategy for Medical Image Segmentation
abstract
Convolutional Neural Networks struggle with long-range dependencies modeling in medical image segmentation, and traditional Transformer models rely on Multi-Layer Perceptron (MLP) for channel information mixing, with performance issues as data dimensions increase. These issues prompt a reassessment of the model's design to enhance segmentation performance and effectively capture long-range dependencies. Consequently, this study presents the Kadformer, a novel network optimized for fine-grained multi-organ segmentation. The Kadformer model adopts an innovative U-shaped network architecture, which enhances the extraction of spatial and channel features in the encoder through the KAN-Enhanced Multi-Dimensional Attention (KMA) mechanism, effectively compensating for information loss during downsampling. We design a Dynamic Path Selection (DPS) strategy to mitigate the feature extraction discrepancies encountered by the linear attention mechanism when processing category-sparse and category-dense images while enhancing feature discrimination through long-range sequential modeling Mamba. Furthermore, we construct the Data Interaction (DAI) module to guide the dual-path encoder's channel and spatial information filtering and effectively integrate the semantically inconsistent features between the KMA and DPS modules. Our approach achieves more than 30% parameter reduction compared to state-of-the-art methods. In addition, the Kadformer network outperforms existing segmentation methods on six public datasets, demonstrating excellent performance. The code has been made available on GitHub: https://github.com/wxc9927/Kadformer.
Dayu Tan, Xingcheng Wang, Yansen Su, Junfeng Xia, Chun-Hou Zheng 0001, Weimin Zhong
IEEE Trans. Medical Imaging4
2025 DRExplainer: Quantifiable interpretability in drug response prediction with directed graph convolutional network
Tao Xu 0011, Zhiwei Xiong, Junfeng Xia
Artif. Intell. Medicine6
2025 MultiPep-DLCL: recognition of multifunctional therapeutic peptides through deep learning with label-sequence contrastive learning
abstract
Identifying multifunctional therapeutic peptides (MFTP) is an important yet complex challenge in the realm of peptide recognition. Unlike monofunctional peptides, MFTP classification requires discerning fine-grained labeling information associated with amino acids, making it more intricate. Existing methods often ignore the nuanced semantics of these labels and fail to fully explore the interplay between peptide sequences and their labels. To address these issues, we propose a multilabel classification method named MultiPep-DLCL. This method uses a deep learning-based model architecture to translate peptide sequences into sequence features by learning the local and global dependencies of multifunctional therapeutic peptide sequences. Additionally, the Label-Sequence Fusion Transformer is employed to efficiently learn high-quality label embeddings by mining effective information from peptide sequences. Finally, the correspondence between sequence features and label embeddings is strengthened through label-sequence contrastive learning. To tackle dataset imbalance, MultiPep-DLCL integrates a multilabel focal dice loss function alongside the traditional cross-entropy loss function. Experimental results demonstrate that the MultiPep-DLCL significantly outperforms existing methods in MFTP recognition.
Henghui Fan, Jianping Zhao 0001, Xiaomei Yang, Junfeng Xia
Briefings Bioinform.5
2025 Multi-view clustering for single-cell RNA-seq data based on graph fusion
abstract
Single-cell RNA sequencing (scRNA-seq) provides transcriptome profiling of individual cells, allowing for in-depth studies of cell heterogeneity at cell resolution. While cell clustering lays the basic foundation of scRNA-seq data analysis, the high-dimensionality and frequent dropout events of the data raise great challenges. Although plenty of dedicated clustering methods have been proposed, they often fail to fully explore the underlying data structure. Here, we introduce scMCGF, a new multi-view clustering algorithm based on graph fusion. It utilizes multi-view data generated from transcriptomic data to learn the consistent and complementary information across different view, ultimately constructing a unified graph matrix for robust cell clustering. Specifically, scMCGF utilizes two-dimensional-reduction methods (principal component analysis and diffusion maps) to capture both linear and non-linear characteristics of the data. Additionally, it calculates a cell-pathway score matrix to incorporate pathway-level information. These three features, along with the pre-processed gene expression data, form the multi-view data. scMCGF iteratively refines the structure of similarity graphs of each view through adaptive learning and learns a unified graph matrix by weighting and fusing the individual similarity graph matrix. The final clustering results are obtained by applying the rank constraint on the Laplacian matrix of the unified graph matrix. Experiments results of 13 real data sets reveal that scMCGF outperforms eight state-of-the-art methods in clustering accuracy and robustness. Furthermore, biological analysis validates that the clustering results of scMCGF provide a reliable foundation for downstream investigations.
Jing Wang 0057, Junfeng Xia, Dayu Tan, Yunjie Ma, Yansen Su, Chun-Hou Zheng 0001
Briefings Bioinform.2
2025 Ensemble learning-based predictor for driver synonymous mutation with sequence representation
abstract
Synonymous mutations, once considered neutral, are now understood to have significant implications for a variety of diseases, particularly cancer. It is indispensable to identify these driver synonymous mutations in human cancers, yet current methods are constrained by data limitations. In this study, we initially investigate the impact of sequence-based features, including DNA shape, physicochemical properties and one-hot encoding of nucleotides, and deep learning-derived features from pre-trained chemical molecule language models based on BERT. Subsequently, we propose EPEL, an effect predictor for synonymous mutations employing ensemble learning. EPEL combines five tree-based models and optimizes feature selection to enhance predictive accuracy. Notably, the incorporation of DNA shape features and deep learning-derived features from chemical molecule represents a pioneering effect in assessing the impact of synonymous mutations in cancer. Compared to existing state-of-the-art methods, EPEL demonstrates superior performance on the independent test dataset. Furthermore, our analysis reveals a significant correlation between effect scores and patient outcomes across various cancer types. Interestingly, while deep learning methods have shown promise in other fields, their DNA sequence representations do not significantly enhance the identification of driver synonymous mutations in this study. Overall, we anticipate that EPEL will facilitate researchers to more precisely target driver synonymous mutations. EPEL is designed with flexibility, allowing users to retrain the prediction model and generate effect scores for synonymous mutations in human cancers. A user-friendly web server for EPEL is available at http://ahmu.EPEL.bio/.
Chuanmei Bi, Junfeng Xia
PLoS Comput. Biol.3
2025 scSDNE: A semi-supervised method for inferring cell-cell interactions based on graph embedding
abstract
As a fundamental characteristic of multicellular organisms, cell-cell communication is achieved through ligand-receptor (L-R) interactions, enabling the exchange of information and revealing the diversity of biological processes and cellular functions. To gain a comprehensive understanding of these complex interaction mechanisms, we constructed a manually curated L-R interaction database and developed a semi-supervised graph embedding model called scSDNE for inferring cell-cell interactions mediated by L-R interactions. scSDNE model utilizes the power of deep learning to map genes from interacting cells into a shared latent space, allowing for a nuanced representation of their relationships. Leveraging the prior information provided by database, scSDNE can infer significant L-R pairs involved in intercellular communication. Experiments on real single-cell RNA sequencing (scRNA-seq) datasets demonstrate that our method detects interactions with a high degree of reliability compared with other methods. More importantly, the model integrates gene regulation information within cells to enhance the accuracy and biological interpretability of the inferences. Our method provides a more comprehensive view of cell-cell interactions, offering new insights into complex intercellular communication.
Chenchen Jia, Jianping Zhao 0001, Junfeng Xia, Chun-Hou Zheng 0001
PLoS Comput. Biol.4
2025 scDMSC: Deep Multi-View Subspace Clustering for Single-Cell Multi-Omics Data
abstract
Single-cell multi-omics sequencing technology comprehensively considers various molecular features to reveal the complexity of cells information. The clustering analysis of multi-omics data provides new insight into cellular heterogeneity. However, multi-omics data are characterized by high dimensionality, sparsity, and heterogeneity. Here, we propose an unsupervised clustering algorithm based on deep multi-view subspace learning, called scDMSC. This approach coordinates the heterogeneity of omics data through weighted reconstruction and employs deep subspace learning to identify shared latent features, elucidating the correlations among the omics. Our algorithm was rigorously tested across multiple real and simulated datasets, outperforming existing single-cell multi-omics integration methods and standard single-cell transcriptomics clustering tools in terms of both precision and scalability. Furthermore, differential expression and modality interpretability analyses in downstream applications highlight the model's capacity in uncovering biological mechanisms.
Zile Wang, Fengyu Lei, Jianping Zhao 0001, Junfeng Xia
IEEE J. Biomed. Health Informatics5
2025 ExplainMIX: Explaining Drug Response Prediction in Directed Graph Neural Networks With Multi-Omics Fusion
abstract
The intricacies of cancer present formidable challenges in achieving effective treatments. Despite extensive research in computational methods for drug response prediction, achieving personalized treatment insights remains challenging. Emerging solutions combine multiple omics data, leveraging graph neural networks to integrate molecular interactions into the reasoning process. However, effectively modeling and harnessing this information, as well as gaining the trust of clinical professionals remain complex. This paper introduces ExplainMIX, a pioneering approach that utilizes directed graph neural networks to predict drug responses with interpretability. ExplainMIX adeptly captures intricate structures and features within directed heterogeneous graphs, leveraging diverse data modalities such as genomics, proteomics, and metabolomics. ExplainMIX goes beyond prediction by generating transparent and interpretable explanations. Incorporating edge-level, meta-path, and graph structure information, it provides meaningful insights into factors influencing drug response, supporting clinicians and researchers in the development of targeted therapies. Empirical results validate the efficacy of ExplainMIX in prediction and interpretation tasks by constructing a quantitative evaluation ground truth. This approach aims to contribute to precision medicine research by addressing challenges in interpretable personalized drug response prediction within the landscape of cancer.
Ying Xiang, Junfeng Xia
IEEE J. Biomed. Health Informatics4
2024 DeepFGRN: inference of gene regulatory network with regulation type based on directed graph embedding
abstract
The inference of gene regulatory networks (GRNs) from gene expression profiles has been a key issue in systems biology, prompting many researchers to develop diverse computational methods. However, most of these methods do not reconstruct directed GRNs with regulatory types because of the lack of benchmark datasets or defects in the computational methods. Here, we collect benchmark datasets and propose a deep learning-based model, DeepFGRN, for reconstructing fine gene regulatory networks (FGRNs) with both regulation types and directions. In addition, the GRNs of real species are always large graphs with direction and high sparsity, which impede the advancement of GRN inference. Therefore, DeepFGRN builds a node bidirectional representation module to capture the directed graph embedding representation of the GRN. Specifically, the source and target generators are designed to learn the low-dimensional dense embedding of the source and target neighbors of a gene, respectively. An adversarial learning strategy is applied to iteratively learn the real neighbors of each gene. In addition, because the expression profiles of genes with regulatory associations are correlative, a correlation analysis module is designed. Specifically, this module not only fully extracts gene expression features, but also captures the correlation between regulators and target genes. Experimental results show that DeepFGRN has a competitive capability for both GRN and FGRN inference. Potential biomarkers and therapeutic drugs for breast cancer, liver cancer, lung cancer and coronavirus disease 2019 are identified based on the candidate FGRNs, providing a possible opportunity to advance our knowledge of disease treatments.
Yansen Su, Junfeng Xia, Yun Ding, Chun-Hou Zheng 0001, Pi-Jing Wei
Briefings Bioinform.3
2024 scHyper: reconstructing cell-cell communication through hypergraph neural networks
abstract
Cell-cell communications is crucial for the regulation of cellular life and the establishment of cellular relationships. Most approaches of inferring intercellular communications from single-cell RNA sequencing (scRNA-seq) data lack a comprehensive global network view of multilayered communications. In this context, we propose scHyper, a new method that can infer intercellular communications from a global network perspective and identify the potential impact of all cells, ligand, and receptor expression on the communication score. scHyper designed a new way to represent tripartite relationships, by extracting a heterogeneous hypergraph that includes the source (ligand expression), the target (receptor expression), and the relevant ligand-receptor (L-R) pairs. scHyper is based on hypergraph representation learning, which measures the degree of match between the intrinsic attributes (static embeddings) of nodes and their observed behaviors (dynamic embeddings) in the context (hyperedges), quantifies the probability of forming hyperedges, and thus reconstructs the cell-cell communication score. Additionally, to effectively mine the key mechanisms of signal transmission, we collect a rich dataset of multisubunit complex L-R pairs and propose a nonparametric test to determine significant intercellular communications. Comparing with other tools indicates that scHyper exhibits superior performance and functionality. Experimental results on the human tumor microenvironment and immune cells demonstrate that scHyper offers reliable and unique capabilities for analyzing intercellular communication networks. Therefore, we introduced an effective strategy that can build high-order interaction patterns, surpassing the limitations of most methods that can only handle low-order interactions, thus more accurately interpreting the complexity of intercellular communications.
Wenying Li, Jianping Zhao 0001, Junfeng Xia
Briefings Bioinform.4
2024 scVSC: Deep Variational Subspace Clustering for Single-Cell Transcriptome Data
abstract
Single-cell RNA sequencing (scRNA-seq) is a potent advancement for analyzing gene expression at the individual cell level, allowing for the identification of cellular heterogeneity and subpopulations. However, it suffers from technical limitations that result in sparse and heterogeneous data. Here, we propose scVSC, an unsupervised clustering algorithm built on deep representation neural networks. The method incorporates the variational inference into the subspace model, which imposes regularization constraints on the latent space and further prevents overfitting. In a series of experiments across multiple datasets, scVSC outperforms existing state-of-the-art unsupervised and semi-supervised clustering tools regarding clustering accuracy and running efficiency. Moreover, the study indicates that scVSC could visually reveal the state of trajectory differentiation, accurately identify differentially expressed genes, and further discover biologically critical pathways.
Zile Wang, Jianping Zhao 0001, Junfeng Xia, Chun-Hou Zheng 0001
IEEE ACM Trans. Comput. Biol. Bioinform.4
2024 Effect Predictor of Driver Synonymous Mutations Based on Multi-Feature Fusion and Iterative Feature Representation Learning
abstract
Accurate identification of driver mutations is crucial in genetic studies of human cancers. While numerous cancer driver missense mutations have been identified, research into potential cancer drivers for synonymous mutations has shown limited success to date. Here, we developed a novel machine learning framework, epSMic, for predicting cancer driver synonymous mutations. epSMic employs an iterative feature representation scheme that facilitates the learning of discriminative features from various sequential models in a supervised iterative mode. We constructed the benchmark datasets and encoded the embedding sequence, physicochemical property, and basic information such as conservation and splicing feature. The evaluation results on benchmark test datasets demonstrate that epSMic outperforms existing methods, making it a valuable tool for researchers in identifying functional synonymous mutations in cancer. We hope epSMic can enable researchers to concentrate on synonymous mutations that have a functional impact on cancer.
Chuanmei Bi, Mengkun Ren, Junfeng Xia
IEEE J. Biomed. Health Informatics7
2024 A Novel Skip-Connection Strategy by Fusing Spatial and Channel Wise Features for Multi-Region Medical Image Segmentation
abstract
Recent methods often introduce attention mechanisms into the skip connections of U-shaped networks to capture features. However, these methods usually overlook spatial information extraction in skip connections and exhibit inefficiency in capturing spatial and channel information. This issue prompts us to reevaluate the design of the skip-connection mechanism and propose a new deep-learning network called the Fusing Spatial and Channel Attention Network, abbreviated as FSCA-Net. FSCA-Net is a novel U-shaped network architecture that utilizes the Parallel Attention Transformer (PAT) to enhance the extraction of spatial and channel features in the skip-connection mechanism, further compensating for downsampling losses. We design the Cross-Attention Bridge Layer (CAB) to mitigate excessive feature and resolution loss when downsampling to the lowest level, ensuring meaningful information fusion during upsampling at the lowest level. Finally, we construct the Dual-Path Channel Attention (DPCA) module to guide channel and spatial information filtering for Transformer features, eliminating ambiguities with decoder features and better concatenating features with semantic inconsistencies between the Transformer and the U-Net decoder. FSCA-Net is designed explicitly for fine-grained segmentation tasks of multiple organs and regions. Our approach achieves over 48% reduction in FLOPs and over 32% reduction in parameters compared to the state-of-the-art method. Moreover, FSCA-Net outperforms existing segmentation methods on seven public datasets, demonstrating exceptional performance.
Dayu Tan, Junfeng Xia, Yansen Su, Chun-Hou Zheng 0001
IEEE J. Biomed. Health Informatics4
2023 scDCCA: deep contrastive clustering for single-cell RNA-seq data based on auto-encoder network
abstract
The advances in single-cell ribonucleic acid sequencing (scRNA-seq) allow researchers to explore cellular heterogeneity and human diseases at cell resolution. Cell clustering is a prerequisite in scRNA-seq analysis since it can recognize cell identities. However, the high dimensionality, noises and significant sparsity of scRNA-seq data have made it a big challenge. Although many methods have emerged, they still fail to fully explore the intrinsic properties of cells and the relationship among cells, which seriously affects the downstream clustering performance. Here, we propose a new deep contrastive clustering algorithm called scDCCA. It integrates a denoising auto-encoder and a dual contrastive learning module into a deep clustering framework to extract valuable features and realize cell clustering. Specifically, to better characterize and learn data representations robustly, scDCCA utilizes a denoising Zero-Inflated Negative Binomial model-based auto-encoder to extract low-dimensional features. Meanwhile, scDCCA incorporates a dual contrastive learning module to capture the pairwise proximity of cells. By increasing the similarities between positive pairs and the differences between negative ones, the contrasts at both the instance and the cluster level help the model learn more discriminative features and achieve better cell segregation. Furthermore, scDCCA joins feature learning with clustering, which realizes representation learning and cell clustering in an end-to-end manner. Experimental results of 14 real datasets validate that scDCCA outperforms eight state-of-the-art methods in terms of accuracy, generalizability, scalability and efficiency. Cell visualization and biological analysis demonstrate that scDCCA significantly improves clustering and facilitates downstream analysis for scRNA-seq data. The code is available at https://github.com/WJ319/scDCCA.
Jing Wang 0057, Junfeng Xia, Yansen Su, Chun-Hou Zheng 0001
Briefings Bioinform.2
2023 Deleterious synonymous mutation identification based on selective ensemble strategy
abstract
Although previous studies have revealed that synonymous mutations contribute to various human diseases, distinguishing deleterious synonymous mutations from benign ones is still a challenge in medical genomics. Recently, computational tools have been introduced to predict the harmfulness of synonymous mutations. However, most of these computational tools rely on balanced training sets without considering abundant negative samples that could result in deficient performance. In this study, we propose a computational model that uses a selective ensemble to predict deleterious synonymous mutations (seDSM). We construct several candidate base classifiers for the ensemble using balanced training subsets randomly sampled from the imbalanced benchmark training sets. The diversity measures of the base classifiers are calculated by the pairwise diversity metrics, and the classifiers with the highest diversities are selected for integration using soft voting for synonymous mutation prediction. We also design two strategies for filling in missing values in the imbalanced dataset and constructing models using different pairwise diversity metrics. The experimental results show that a selective ensemble based on double fault with the ensemble strategy EKNNI for filling in missing values is the most effective scheme. Finally, using 40-dimensional biology features, we propose a novel model based on a selective ensemble for predicting deleterious synonymous mutations (seDSM). seDSM outperformed other state-of-the-art methods on the independent test sets according to multiple evaluation indicators, indicating that it has an outstanding predictive performance for deleterious synonymous mutations. We hope that seDSM will be useful for studying deleterious synonymous mutations and advancing our understanding of synonymous mutations. The source code of seDSM is freely accessible at https://github.com/xialab-ahu/seDSM.git.
Lihong Yu, Chun-Hou Zheng 0001, Wenguang Yin, Junfeng Xia
Briefings Bioinform.6
2023 Deep learning-based multi-functional therapeutic peptides prediction with a multi-label focal dice loss function
abstract
MOTIVATION: With the great number of peptide sequences produced in the postgenomic era, it is highly desirable to identify the various functions of therapeutic peptides quickly. Furthermore, it is a great challenge to predict accurate multi-functional therapeutic peptides (MFTP) via sequence-based computational tools. RESULTS: Here, we propose a novel multi-label-based method, named ETFC, to predict 21 categories of therapeutic peptides. The method utilizes a deep learning-based model architecture, which consists of four blocks: embedding, text convolutional neural network, feed-forward network, and classification blocks. This method also adopts an imbalanced learning strategy with a novel multi-label focal dice loss function. multi-label focal dice loss is applied in the ETFC method to solve the inherent imbalance problem in the multi-label dataset and achieve competitive performance. The experimental results state that the ETFC method is significantly better than the existing methods for MFTP prediction. With the established framework, we use the teacher-student-based knowledge distillation to obtain the attention weight from the self-attention mechanism in the MFTP prediction and quantify their contributions toward each of the investigated activities. AVAILABILITY AND IMPLEMENTATION: The source code and dataset are available via: https://github.com/xialab-ahu/ETFC.
Henghui Fan, Wenhui Yan, Yannan Bin, Junfeng Xia
Bioinform.6
2023 PhaGAA: an integrated web server platform for phage genome annotation and analysis
abstract
MOTIVATION: Phage genome annotation plays a key role in the design of phage therapy. To date, there have been various genome annotation tools for phages, but most of these tools focus on mono-functional annotation and have complex operational processes. Accordingly, comprehensive and user-friendly platforms for phage genome annotation are needed. RESULTS: Here, we propose PhaGAA, an online integrated platform for phage genome annotation and analysis. By incorporating several annotation tools, PhaGAA is constructed to annotate the prophage genome at DNA and protein levels and provide the analytical results. Furthermore, PhaGAA could mine and annotate phage genomes from bacterial genome or metagenome. In summary, PhaGAA will be a useful resource for experimental biologists and help advance the phage synthetic biology in basic and application research. AVAILABILITY AND IMPLEMENTATION: PhaGAA is freely available at http://phage.xialab.info/.
Qingrui Liu, Jiliang Xu, Junyin Zhang, Minfeng Xiao, Yannan Bin, Junfeng Xia
Bioinform.9
2023 CNNGRN: A Convolutional Neural Network-Based Method for Gene Regulatory Network Inference From Bulk Time-Series Expression Data
abstract
Gene regulatory networks (GRNs) participate in many biological processes, and reconstructing them plays an important role in systems biology. Although many advanced methods have been proposed for GRN reconstruction, their predictive performance is far from the ideal standard, so it is urgent to design a more effective method to reconstruct GRN. Moreover, most methods only consider the gene expression data, ignoring the network structure information contained in GRN. In this study, we propose a supervised model named CNNGRN, which infers GRN from bulk time-series expression data via convolutional neural network (CNN) model, with a more informative feature. Bulk time series gene expression data imply the intricate regulatory associations between genes, and the network structure feature of ground-truth GRN contains rich neighbor information. Hence, CNNGRN integrates the above two features as model inputs. In addition, CNN is adopted to extract intricate features of genes and infer the potential associations between regulators and target genes. Moreover, feature importance visualization experiments are implemented to seek the key features. Experimental results show that CNNGRN achieved competitive performance on benchmark datasets compared to the state-of-the-art computational methods. Finally, hub genes identified based on CNNGRN have been confirmed to be involved in biological processes through literature.
Jin Tang 0001, Junfeng Xia, Chun-Hou Zheng 0001, Pi-Jing Wei
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 frDSM: An Ensemble Predictor With Effective Feature Representation for Deleterious Synonymous Mutation in Human Genome
abstract
With the discovery of causality between synonymous mutations and diseases, it has become increasingly important to identify deleterious synonymous mutations for better understanding of their functional mechanisms. Although several machine learning methods have been proposed to solve the task, an effective feature representation method that can make use of the inner difference and relevance between deleterious and benign synonymous mutations is still challenging considering the vast number of synonymous mutations in human genome. In this work, we developed a robust and accurate predictor called frDSM for deleterious synonymous mutation prediction using logistic regression. More specifically, we introduced an effective feature representation learning method which exploits multiple feature descriptors from different perspectives including functional scores obtained from previously computational methods, evolutionary conservation, splicing and sequence feature descriptors, and these features descriptors were input into the 76 XGBoost classifiers to obtain the predictive probabilities values. These probabilities were concatenated to generate the 76-dimension new feature vector, and feature selection method was used to remove redundant and irrelevant features. Experimental results show that frDSM enables robust and accurate prediction than the competing prediction methods with 31 optimal features, which demonstrated the effectiveness of the feature representation learning method. frDSM is freely available at http://frdsm.xialab.info.
Jianhui Sun, Chun-Hou Zheng 0001, Junfeng Xia
IEEE ACM Trans. Comput. Biol. Bioinform.5
2022 An Ensemble Framework Integrating Whole Slide Pathological Images and miRNA Data to Predict Radiosensitivity of Breast Cancer Patients
Wenhui Yan, Mengmeng Han, Junfeng Xia, Yannan Bin
ICIC (2)6
2022 Identifying multi-functional bioactive peptide functions using multi-label deep learning
abstract
The bioactive peptide has wide functions, such as lowering blood glucose levels and reducing inflammation. Meanwhile, computational methods such as machine learning are becoming more and more important for peptide functions prediction. Most of the previous studies concentrate on the single-functional bioactive peptides prediction. However, the number of multi-functional peptides is on the increase; therefore, novel computational methods are needed. In this study, we develop a method MLBP (Multi-Label deep learning approach for determining the multi-functionalities of Bioactive Peptides), which can predict multiple functions including anti-cancer, anti-diabetic, anti-hypertensive, anti-inflammatory and anti-microbial simultaneously. MLBP model takes the peptide sequence vector as input to replace the biological and physiochemical features used in other peptides predictors. Using the embedding layer, the dense continuous feature vector is learnt from the sequence vector. Then, we extract convolution features from the feature vector through the convolutional neural network layer and combine with the bidirectional gated recurrent unit layer to improve the prediction performance. The 5-fold cross-validation experiments are conducted on the training dataset, and the results show that Accuracy and Absolute true are 0.695 and 0.685, respectively. On the test dataset, Accuracy and Absolute true of MLBP are 0.709 and 0.697, with 5.0 and 4.7% higher than those of the suboptimum method, respectively. The results indicate MLBP has superior prediction performance on the multi-functional peptides identification. MLBP is available at https://github.com/xialab-ahu/MLBP and http://bioinfo.ahu.edu.cn/MLBP/.
Wending Tang, Ruyu Dai, Wenhui Yan, Yannan Bin, En-Hua Xia, Junfeng Xia
Briefings Bioinform.7
2022 scHFC: a hybrid fuzzy clustering method for single-cell RNA-seq data optimized by natural computation
abstract
Rapid development of single-cell RNA sequencing (scRNA-seq) technology has allowed researchers to explore biological phenomena at the cellular scale. Clustering is a crucial and helpful step for researchers to study the heterogeneity of cell. Although many clustering methods have been proposed, massive dropout events and the curse of dimensionality in scRNA-seq data make it still difficult to analysis because they reduce the accuracy of clustering methods, leading to misidentification of cell types. In this work, we propose the scHFC, which is a hybrid fuzzy clustering method optimized by natural computation based on Fuzzy C Mean (FCM) and Gath-Geva (GG) algorithms. Specifically, principal component analysis algorithm is utilized to reduce the dimensions of scRNA-seq data after it is preprocessed. Then, FCM algorithm optimized by simulated annealing algorithm and genetic algorithm is applied to cluster the data to output a membership matrix, which represents the initial clustering result and is taken as the input for GG algorithm to get the final clustering results. We also develop a cluster number estimation method called multi-index comprehensive estimation, which can estimate the cluster numbers well by combining four clustering effectiveness indexes. The performance of the scHFC method is evaluated on 17 scRNA-seq datasets, and compared with six state-of-the-art methods. Experimental results validate the better performance of our scHFC method in terms of clustering accuracy and stability of algorithm. In short, scHFC is an effective method to cluster cells for scRNA-seq data, and it presents great potential for downstream analysis of scRNA-seq data. The source code is available at https://github.com/WJ319/scHFC.
Jing Wang 0057, Junfeng Xia, Dayu Tan, Rongxin Lin, Yansen Su, Chun-Hou Zheng 0001
Briefings Bioinform.2
2022 PrMFTP: Multi-functional therapeutic peptides prediction based on multi-head self-attention mechanism and class weight optimization
abstract
Prediction of therapeutic peptide is a significant step for the discovery of promising therapeutic drugs. Most of the existing studies have focused on the mono-functional therapeutic peptide prediction. However, the number of multi-functional therapeutic peptides (MFTP) is growing rapidly, which requires new computational schemes to be proposed to facilitate MFTP discovery. In this study, based on multi-head self-attention mechanism and class weight optimization algorithm, we propose a novel model called PrMFTP for MFTP prediction. PrMFTP exploits multi-scale convolutional neural network, bi-directional long short-term memory, and multi-head self-attention mechanisms to fully extract and learn informative features of peptide sequence to predict MFTP. In addition, we design a class weight optimization scheme to address the problem of label imbalanced data. Comprehensive evaluation demonstrate that PrMFTP is superior to other state-of-the-art computational methods for predicting MFTP. We provide a user-friendly web server of PrMFTP, which is available at http://bioinfo.ahu.edu.cn/PrMFTP.
Wenhui Yan, Wending Tang, Yannan Bin, Junfeng Xia
PLoS Comput. Biol.5
2022 Extra Trees Method for Predicting LncRNA-Disease Association Based On Multi-Layer Graph Embedding Aggregation
abstract
Lots of experimental studies have revealed the significant associations between lncRNAs and diseases. Identifying accurate associations will provide a new perspective for disease therapy. Calculation-based methods have been developed to solve these problems, but these methods have some limitations. In this paper, we proposed an accurate method, named MLGCNET, to discover potential lncRNA-disease associations. Firstly, we reconstructed similarity networks for both lncRNAs and diseases using top k similar information, and constructed a lncRNA-disease heterogeneous network (LDN). Then, we applied Multi-Layer Graph Convolutional Network on LDN to obtain latent feature representations of nodes. Finally, the Extra Trees was used to calculate the probability of association between disease and lncRNA. The results of extensive 5-fold cross-validation experiments show that MLGCNET has superior prediction performance compared to the state-of-the-art methods. Case studies confirm the performance of our model on specific diseases. All the experiment results prove the effectiveness and practicality of MLGCNET in predicting potential lncRNA-disease associations.
Qing-Wen Wu, Junfeng Xia, Jiancheng Ni 0001, Chun-Hou Zheng 0001, Yansen Su
IEEE ACM Trans. Comput. Biol. Bioinform.3
2022 An Ensemble Framework for Improving the Prediction of Deleterious Synonymous Mutation
abstract
In recent years, the association between synonymous mutations (SMs) and human diseases has been uncovered in many studies. It is a challenge for identifying deleterious SMs in the field of medical genomics. Although there are several computational methods proposed in the past years, the precise prediction of deleterious SMs is still challenging. In this work, we proposed a predictor named as EnDSM, which is an accurate method based on the ensemble framework. We explored multimodal features across four groups including functional score, conservation, splicing, and sequence features, and we then trained eight conceptually different machine learning classifiers for each of them, resulting in 32 base classification models. We further selected four base models referring to their prediction performance and the predictive probabilities of these base classification models were subsequently used as the input feature vectors of logistic regression classifier to construct the ensemble learning model. The results suggested that EnDSM achieved better performance comparing with other state-of-the-art predictors on the training and independent test datasets. We anticipate that our ensemble predictor EnDSM will become a valuable tool for deleterious SM prediction.The EnDSM server interface along with the benchmarking data sets are freely available athttp://bioinfo.ahu.edu.cn/EnDSM.
Jie Gui, Chun-Hou Zheng 0001, Junfeng Xia
IEEE Trans. Circuits Syst. Video Technol.7
2022 DPProm: A Two-Layer Predictor for Identifying Promoters and Their Types on Phage Genome Using Deep Learning
abstract
With the number of phage genomes increasing, it is urgent to develop new bioinformatics methods for phage genome annotation. Promoter, a DNA region, is important for gene transcriptional regulation. In the era of post-genomics, the availability of data makes it possible to establish computational models for promoter identification with robustness. In this work, we introduce DPProm, a two-layer model composed of DPProm-1L and DPProm-2L, to predict promoters and their types for phages. On the first layer, as a dual-channel deep neural network ensemble method fusing multi-view features (sequence feature and handcrafted feature), the model DPProm-1L is proposed to identify whether a DNA sequence is a promoter or non-promoter. The sequence feature is extracted with convolutional neural network (CNN). And the handcrafted feature is the combination of free energy, GC content, cumulative skew, and Z curve features. On the second layer, DPProm-2L based on CNN is trained to predict the promoters' types (host or phage). For the realization of prediction on the whole genomes, the model DPProm, combines with a novel sequence data processing workflow, which contains sliding window and merging sequences modules. Experimental results show that DPProm outperforms the state-of-the-art methods, and decreases the false positive rate effectively on whole genome prediction. Furthermore, we provide a user-friendly web at http://bioinfo.ahu.edu.cn/DPProm. We expect that DPProm can serve as a useful tool for identification of promoters and their types.
Junyin Zhang, Minfeng Xiao, Junfeng Xia, Yannan Bin
IEEE J. Biomed. Health Informatics6
2021 usDSM: a novel method for deleterious synonymous mutation prediction using undersampling scheme
abstract
Although synonymous mutations do not alter the encoded amino acids, they may impact protein function by interfering with the regulation of RNA splicing or altering transcript splicing. New progress on next-generation sequencing technologies has put the exploration of synonymous mutations at the forefront of precision medicine. Several approaches have been proposed for predicting the deleterious synonymous mutations specifically, but their performance is limited by imbalance of the positive and negative samples. In this study, we firstly expanded the number of samples greatly from various data sources and compared six undersampling strategies to solve the problem of the imbalanced datasets. The results suggested that cluster centroid is the most effective scheme. Secondly, we presented a computational model, undersampling scheme based method for deleterious synonymous mutation (usDSM) prediction, using 14-dimensional biology features and random forest classifier to detect the deleterious synonymous mutation. The results on the test datasets indicated that the proposed usDSM model can attain superior performance in comparison with other state-of-the-art machine learning methods. Lastly, we found that the deep learning model did not play a substantial role in deleterious synonymous mutation prediction through a lot of experiments, although it achieves superior results in other fields. In conclusion, we hope our work will contribute to the future development of computational methods for a more accurate prediction of the deleterious effect of human synonymous mutation. The web server of usDSM is freely accessible at http://usdsm.xialab.info/.
Chun-Hou Zheng 0001, Junfeng Xia
Briefings Bioinform.6
2021 Erratum: usDSM: a novel method for deleterious synonymous mutation prediction using undersampling scheme
abstract
When this paper was originally published online, the lower part of Figure 2 was missing, and the web server name in Data availability Section was listed incorrectly. In addition, four values in Table 3 should have been shown in bold. The paper has been corrected online.
Chun-Hou Zheng 0001, Junfeng Xia
Briefings Bioinform.6
2021 GAERF: predicting lncRNA-disease associations by graph auto-encoder and random forest
abstract
Predicting disease-related long non-coding RNAs (lncRNAs) is beneficial to finding of new biomarkers for prevention, diagnosis and treatment of complex human diseases. In this paper, we proposed a machine learning techniques-based classification approach to identify disease-related lncRNAs by graph auto-encoder (GAE) and random forest (RF) (GAERF). First, we combined the relationship of lncRNA, miRNA and disease into a heterogeneous network. Then, low-dimensional representation vectors of nodes were learned from the network by GAE, which reduce the dimension and heterogeneity of biological data. Taking these feature vectors as input, we trained a RF classifier to predict new lncRNA-disease associations (LDAs). Related experiment results show that the proposed method for the representation of lncRNA-disease characterizes them accurately. GAERF achieves superior performance owing to the ensemble learning method, outperforming other methods significantly. Moreover, case studies further demonstrated that GAERF is an effective method to predict LDAs.
Qing-Wen Wu, Junfeng Xia, Jiancheng Ni 0001, Chun-Hou Zheng 0001
Briefings Bioinform.2
2021 PredCID: prediction of driver frameshift indels in human cancer
abstract
The discrimination of driver from passenger mutations has been a hot topic in the field of cancer biology. Although recent advances have improved the identification of driver mutations in cancer genomic research, there is no computational method specific for the cancer frameshift indels (insertions or/and deletions) yet. In addition, existing pathogenic frameshift indel predictors may suffer from plenty of missing values because of different choices of transcripts during the variant annotation processes. In this study, we proposed a computational model, called PredCID (Predictor for Cancer driver frameshift InDels), for accurately predicting cancer driver frameshift indels. Gene, DNA, transcript and protein level features are combined together and selected for classification with eXtreme Gradient Boosting classifier. Benchmarking results on the cross-validation dataset and independent dataset showed that PredCID achieves better and robust performance compared with existing noncancer-specific methods in distinguishing cancer driver frameshift indels from passengers and is therefore a valuable method for deeper understanding of frameshift indels in human cancer. PredCID is freely available for academic research at http://bioinfo.ahu.edu.cn:8080/PredCID.
Xinlu Chu, Junfeng Xia
Briefings Bioinform.3
2021 Identification of driver genes based on gene mutational effects and network centrality
abstract
BACKGROUND: As one of the deadliest diseases in the world, cancer is driven by a few somatic mutations that disrupt the normal growth of cells, and leads to abnormal proliferation and tumor development. The vast majority of somatic mutations did not affect the occurrence and development of cancer; thus, identifying the mutations responsible for tumor occurrence and development is one of the main targets of current cancer treatments. RESULTS: To effectively identify driver genes, we adopted a semi-local centrality measure and gene mutation effect function to assess the effect of gene mutations on changes in gene expression patterns. Firstly, we calculated the mutation score for each gene. Secondly, we identified differentially expressed genes (DEGs) in the cohort by comparing the expression profiles of tumor samples and normal samples, and then constructed a local network for each mutation gene using DEGs and mutant genes according to the protein-protein interaction network. Finally, we calculated the score of each mutant gene according to the objective function. The top-ranking mutant genes were selected as driver genes. We name the proposed method as mutations effect and network centrality. CONCLUSIONS: Four types of cancer data in The Cancer Genome Atlas were tested. The experimental data proved that our method was superior to the existing network-centric method, as it was able to quickly and easily identify driver genes and rare driver factors.
Yun-Yun Tang, Pi-Jing Wei, Jianping Zhao 0001, Junfeng Xia, Chun-Hou Zheng 0001
BMC Bioinform.4
2021 An improved DNA-binding hot spot residues prediction method by exploring interfacial neighbor properties
abstract
BACKGROUND: DNA-binding hot spots are dominant and fundamental residues that contribute most of the binding free energy yet accounting for a small portion of protein-DNA interfaces. As experimental methods for identifying hot spots are time-consuming and costly, high-efficiency computational approaches are emerging as alternative pathways to experimental methods. RESULTS: Herein, we present a new computational method, termed inpPDH, for hot spot prediction. To improve the prediction performance, we extract hybrid features which incorporate traditional features and new interfacial neighbor properties. To remove redundant and irrelevant features, feature selection is employed using a two-step feature selection strategy. Finally, a subset of 7 optimal features are chosen to construct the predictor using support vector machine. The results on the benchmark dataset show that this proposed method yields significantly better prediction accuracy than those previously published methods in the literature. Moreover, a user-friendly web server for inpPDH is well established and is freely available at http://bioinfo.ahu.edu.cn/inpPDH . CONCLUSIONS: We have developed an accurate improved prediction model, inpPDH, for hot spot residues in protein-DNA binding interfaces by given the structure of a protein-DNA complex. Moreover, we identify a comprehensive and useful feature subset including the proposed interfacial neighbor features that has an important strength for identifying hot spot residues. Our results indicate that these features are more effective than the conventional features considered previously, and that the combination of interfacial neighbor features and traditional features may support the creation of a discriminative feature set for efficient prediction of hot spot residues in protein-DNA complexes.
Menglu Li, Yannan Bin, Junfeng Xia
BMC Bioinform.8
2021 Double matrix completion for circRNA-disease association prediction
abstract
BACKGROUND: Circular RNAs (circRNAs) are a class of single-stranded RNA molecules with a closed-loop structure. A growing body of research has shown that circRNAs are closely related to the development of diseases. Because biological experiments to verify circRNA-disease associations are time-consuming and wasteful of resources, it is necessary to propose a reliable computational method to predict the potential candidate circRNA-disease associations for biological experiments to make them more efficient. RESULTS: In this paper, we propose a double matrix completion method (DMCCDA) for predicting potential circRNA-disease associations. First, we constructed a similarity matrix of circRNA and disease according to circRNA sequence information and semantic disease information. We also built a Gauss interaction profile similarity matrix for circRNA and disease based on experimentally verified circRNA-disease associations. Then, the corresponding circRNA sequence similarity and semantic similarity of disease are used to update the association matrix from the perspective of circRNA and disease, respectively, by matrix multiplication. Finally, from the perspective of circRNA and disease, matrix completion is used to update the matrix block, which is formed by splicing the association matrix obtained in the previous step with the corresponding Gaussian similarity matrix. Compared with other approaches, the model of DMCCDA has a relatively good result in leave-one-out cross-validation and five-fold cross-validation. Additionally, the results of the case studies illustrate the effectiveness of the DMCCDA model. CONCLUSION: The results show that our method works well for recommending the potential circRNAs for a disease for biological experiments.
Zong-Lan Zuo, Pi-Jing Wei, Junfeng Xia, Chun-Hou Zheng 0001
BMC Bioinform.4
2021 A Deep Learning-Based Method for Identification of Bacteriophage-Host Interaction
abstract
Multi-drug resistance (MDR) has become one of the greatest threats to human health worldwide, and novel treatment methods of infections caused by MDR bacteria are urgently needed. Phage therapy is a promising alternative to solve this problem, to which the key is correctly matching target pathogenic bacteria with the corresponding therapeutic phage. Deep learning is powerful for mining complex patterns to generate accurate predictions. In this study, we develop PredPHI (Predicting Phage-Host Interactions), a deep learning-based tool capable of predicting the host of phages from sequence data. We collect >3000 phage-host pairs along with their protein sequences from PhagesDB and GenBank databases and extract a set of features. Then we select high-quality negative samples based on the K-Means clustering method and construct a balanced training set. Finally, we employ a deep convolutional neural network to build the predictive model. The results indicate that PredPHI can achieve a predictive performance of 81 percent in terms of the area under the receiver operating characteristic curve on the test set, and the clustering-based method is significantly more robust than that based on randomly selecting negative samples. These results highlight that PredPHI is a useful and accurate tool for identifying phage-host interactions from sequence data.
Menglu Li, Yanan Wang 0003, Fuyi Li, Yun Zhao 0004, Yannan Bin, Alexander Ian Smith, Geoffrey I. Webb, Jian Li 0052, Jiangning Song, Junfeng Xia
IEEE ACM Trans. Comput. Biol. Bioinform.12
2020 Comparison and integration of computational methods for deleterious synonymous mutation prediction
abstract
Synonymous mutations do not change the encoded amino acids but may alter the structure or function of an mRNA in ways that impact gene function. Advances in next generation sequencing technologies have detected numerous synonymous mutations in the human genome. Several computational models have been proposed to predict deleterious synonymous mutations, which have greatly facilitated the development of this important field. Consequently, there is an urgent need to assess the state-of-the-art computational methods for deleterious synonymous mutation prediction to further advance the existing methodologies and to improve performance. In this regard, we systematically compared a total of 10 computational methods (including specific method for deleterious synonymous mutation and general method for single nucleotide mutation) in terms of the algorithms used, calculated features, performance evaluation and software usability. In addition, we constructed two carefully curated independent test datasets and accordingly assessed the robustness and scalability of these different computational methods for the identification of deleterious synonymous mutations. In an effort to improve predictive performance, we established an ensemble model, named Prediction of Deleterious Synonymous Mutation (PrDSM), which averages the ratings generated by the three most accurate predictors. Our benchmark tests demonstrated that the ensemble model PrDSM outperformed the reviewed tools for the prediction of deleterious synonymous mutations. Using the ensemble model, we developed an accessible online predictor, PrDSM, available at http://bioinfo.ahu.edu.cn:8080/PrDSM/. We hope that this comprehensive survey and the proposed strategy for building more accurate models can serve as a useful guide for inspiring future developments of computational methods for deleterious synonymous mutation prediction.
Menglu Li, Bo Zhang 0001, Yuhua Yang, Chun-Hou Zheng 0001, Junfeng Xia
Briefings Bioinform.7
2020 dbCPM: a manually curated database for exploring the cancer passenger mutations
abstract
While recently emergent driver mutation data sets are available for developing computational methods to predict cancer mutation effects, benchmark sets focusing on passenger mutations are largely missing. Here, we developed a comprehensive literature-based database of Cancer Passenger Mutations (dbCPM), which contains 941 experimentally supported and 978 putative passenger mutations derived from a manual curation of the literature. Using the missense mutation data, the largest group in the dbCPM, we explored patterns of missense passenger mutations by comparing them with the missense driver mutations and assessed the performance of four cancer-focused mutation effect predictors. We found that the missense passenger mutations showed significant differences with drivers at multiple levels, and several appeared in both the passenger and driver categories, showing pleiotropic functions depending on the tumor context. Although all the predictors displayed good true positive rates, their true negative rates were relatively low due to the lack of negative training samples with experimental evidence, which suggests that a suitable negative data set for developing a more robust methodology is needed. We hope that the dbCPM will be a benchmark data set for improving and evaluating prediction algorithms and serve as a valuable resource for the cancer research community. dbCPM is freely available online at http://bioinfo.ahu.edu.cn:8080/dbCPM.
Junfeng Xia
Briefings Bioinform.3
2020 A feature-based approach to predict hot spots in protein-DNA binding interfaces
abstract
DNA-binding hot spot residues of proteins are dominant and fundamental interface residues that contribute most of the binding free energy of protein-DNA interfaces. As experimental methods for identifying hot spots are expensive and time consuming, computational approaches are urgently required in predicting hot spots on a large scale. In this work, we systematically assessed a wide variety of 114 features from a combination of the protein sequence, structure, network and solvent accessible information and their combinations along with various feature selection strategies for hot spot prediction. We then trained and compared four commonly used machine learning models, namely, support vector machine (SVM), random forest, Naïve Bayes and k-nearest neighbor, for the identification of hot spots using 10-fold cross-validation and the independent test set. Our results show that (1) features based on the solvent accessible surface area have significant effect on hot spot prediction; (2) different but complementary features generally enhance the prediction performance; and (3) SVM outperforms other machine learning methods on both training and independent test sets. In an effort to improve predictive performance, we developed a feature-based method, namely, PrPDH (Prediction of Protein-DNA binding Hot spots), for the prediction of hot spots in protein-DNA binding interfaces using SVM based on the selected 10 optimal features. Comparative results on benchmark data sets indicate that our predictor is able to achieve generally better performance in predicting hot spots compared to the state-of-the-art predictors. A user-friendly web server for PrPDH is well established and is freely available at http://bioinfo.ahu.edu.cn:8080/PrPDH.
Chun-Hou Zheng 0001, Junfeng Xia
Briefings Bioinform.4
2020 Prediction of hot spots in protein-DNA binding interfaces based on supervised isometric feature mapping and extreme gradient boosting
abstract
BACKGROUND: Identification of hot spots in protein-DNA interfaces provides crucial information for the research on protein-DNA interaction and drug design. As experimental methods for determining hot spots are time-consuming, labor-intensive and expensive, there is a need for developing reliable computational method to predict hot spots on a large scale. RESULTS: Here, we proposed a new method named sxPDH based on supervised isometric feature mapping (S-ISOMAP) and extreme gradient boosting (XGBoost) to predict hot spots in protein-DNA complexes. We obtained 114 features from a combination of the protein sequence, structure, network and solvent accessible information, and systematically assessed various feature selection methods and feature dimensionality reduction methods based on manifold learning. The results show that the S-ISOMAP method is superior to other feature selection or manifold learning methods. XGBoost was then used to develop hot spots prediction model sxPDH based on the three dimensionality-reduced features obtained from S-ISOMAP. CONCLUSION: Our method sxPDH boosts prediction performance using S-ISOMAP and XGBoost. The AUC of the model is 0.773, and the F1 score is 0.713. Experimental results on benchmark dataset indicate that sxPDH can achieve generally better performance in predicting hot spots compared to the state-of-the-art methods.
Yannan Bin, Junfeng Xia
BMC Bioinform.5
2019 Improved Inductive Matrix Completion Method for Predicting MicroRNA-Disease Associations
Junfeng Xia, Jing Wang 0057, Chun-Hou Zheng 0001
ICIC (2)2
2019 Discovering Driver Mutation Profiles in Cancer with a Local Centrality Score
Ying Hui, Pi-Jing Wei, Junfeng Xia, Jing Wang 0057, Chun-Hou Zheng 0001
ICIC (2)3
2019 Sequence-Based Prediction of Hot Spots in Protein-RNA Complexes Using an Ensemble Approach
Junfeng Xia
ICIC (1)3
2019 dbCID: a manually curated resource for exploring the driver indels in human cancer
abstract
While recent advances in next-generation sequencing technologies have enabled the creation of a multitude of databases in cancer genomic research, there is no comprehensive database focusing on the annotation of driver indels (insertions and deletions) yet. Therefore, we have developed the database of Cancer driver InDels (dbCID), which is a collection of known coding indels that likely to be engaged in cancer development, progression or therapy. dbCID contains experimentally supported and putative driver indels derived from manual curation of literature and is freely available online at http://bioinfo.ahu.edu.cn:8080/dbCID. Using the data deposited in dbCID, we summarized features of driver indels in four levels (gene, DNA, transcript and protein) through comparing with putative neutral indels. We found that most of the genes containing driver indels in dbCID are known cancer genes playing a role in tumorigenesis. Contrary to the expectation, the sequences affected by driver frameshift indels are not larger than those by neutral ones. In addition, the frameshift and inframe driver indels prefer to disrupt high-conservative regions both in DNA sequences and protein domains. Finally, we developed a computational method for discriminating cancer driver from neutral frameshift indels based on the deposited data in dbCID. The proposed method outperformed other widely used non-cancer-specific predictors on an external test set, which demonstrated the usefulness of the data deposited in dbCID. We hope dbCID will be a benchmark for improving and evaluating prediction algorithms, and the characteristics summarized here may assist with investigating the mechanism of indel-cancer association.
Junfeng Xia
Briefings Bioinform.5
2019 The 2017 Network Tools and Applications in Biology (NETTAB) workshop: aims, topics and outcomes
abstract
The 17th International NETTAB workshop was held in Palermo, Italy, on October 16-18, 2017. The special topic for the meeting was "Methods, tools and platforms for Personalised Medicine in the Big Data Era", but the traditional topics of the meeting series were also included in the event. About 40 scientific contributions were presented, including four keynote lectures, five guest lectures, and many oral communications and posters. Also, three tutorials were organised before and after the workshop. Full papers from some of the best works presented in Palermo were submitted for this Supplement of BMC Bioinformatics. Here, we provide an overview of meeting aims and scope. We also shortly introduce selected papers that have been accepted for publication in this Supplement, for a complete presentation of the outcomes of the meeting.
Paolo Romano 0001, Arnaud Céol, Andreas Dräger, Antonino Fiannaca, Rosalba Giugno, Massimo La Rosa, Luciano Milanesi, Ulrich Pfeffer, Riccardo Rizzo, Soo-Yong Shin, Junfeng Xia, Alfonso Urso
BMC Bioinform.11
2018 Further Evidence for Role of Promoter Polymorphisms in TNF Gene in Alzheimer's Disease
Yannan Bin, Ling Shu, Qizhi Zhu, Huanhuan Zhu, Junfeng Xia
ICIC (2)5
2018 Nucleotide-Based Significance of Somatic Synonymous Mutations for Pan-Cancer
Yannan Bin, Qizhi Zhu, Pengbo Wen, Junfeng Xia
ICIC (2)5
2018 Computational Prediction of Driver Missense Mutations in Melanoma
Junfeng Xia, Yannan Bin, Di Zhang 0006
ICIC (2)4
2017 Investigating Alzheimer's Disease Candidate Genes Based on Combined Network Using Subnetwork Extraction Algorithms
Di Zhang 0006, Yannan Bin, Junfeng Xia
ICIC (2)6
2017 Cancer Subtype Discovery Based on Integrative Model of Multigenomic Data
abstract
One major goal of large-scale cancer omics study is to understand molecular mechanisms of cancer and find new biomedical targets. To deal with the high-dimensional multidimensional cancer omics data (DNA methylation, mRNA expression, etc.), which can be used to discover new insight on identifying cancer subtypes, clustering methods are usually used to find an effective low-dimensional subspace of the original data and then cluster cancer samples in the reduced subspace. However, due to data-type diversity and big data volume, few methods can integrate these data and map them into an effective low-dimensional subspace. In this paper, we develop a dimension-reduction and data-integration method for indentifying cancer subtypes, named Scluster. First, Scluster, respectively, projects the different original data into the principal subspaces by an adaptive sparse reduced-rank regression method. Then, a fused patient-by-patient network is obtained for these subgroups through a scaled exponential similarity kernel method. Finally, candidate cancer subtypes are identified using spectral clustering method. We demonstrate the efficiency of our Scluster method using three cancers by jointly analyzing mRNA expression, miRNA expression, and DNA methylation data. The evaluation results and analyses show that Scluster is effective for predicting survival and identifies novel cancer subtypes of large-scale multi-omics data.
Shu-Guang Ge, Junfeng Xia, Wen Sha, Chun-Hou Zheng 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2016 Cancer genes discovery based on integtating transcriptomic data and the impact of gene length
abstract
In this paper, we presented a network-based method, named DriverFinder, by filtering frequently mutated genes just because of their large size, and comparing tumor expression with normal expression data to obtain gene expression outliers which are more likely to be cancer genes. Then greedy algorithm was applied to prioritize candidate driver genes. The proposed method can not only indentify frequently mutated genes, but also novel and infrequently mutated driver genes.
Pi-Jing Wei, Di Zhang 0006, Chun-Hou Zheng 0001, Junfeng Xia
BIBM4
2016 Srrr-cluster: Using Sparse Reduced-Rank Regression to Optimize iCluster
Shu-Guang Ge, Junfeng Xia, Pi-Jing Wei, Chun-Hou Zheng 0001
ICIC (3)2
2016 dbDSM: a manually curated database for deleterious synonymous mutations
abstract
MOTIVATION: Synonymous mutations (SMs), which changed the sequence of a gene without directly altering the amino acid sequence of the encoded protein, were thought to have no functional consequences for a long time. They are often assumed to be neutral in models of mutation and selection and were completely ignored in many studies. However, accumulating experimental evidence has demonstrated that these mutations exert their impact on gene functions via splicing accuracy, mRNA stability, translation fidelity, protein folding and expression, and some of these mutations are implicated in human diseases. To the best of our knowledge, there is still no database specially focusing on disease-related SMs. RESULTS: We have developed a new database called dbDSM (database of Deleterious Synonymous Mutation), a continually updated database that collects, curates and manages available human disease-related SM data obtained from published literature. In the current release, dbDSM collects 1936 SM-disease association entries, including 1289 SMs and 443 human diseases from ClinVar, GRASP, GWAS Catalog, GWASdb, PolymiRTS database, PubMed database and Web of Knowledge. Additionally, we provided users a link to download all the data in the dbDSM and a link to submit novel data into the database. We hope dbDSM will be a useful resource for investigating the roles of SMs in human disease. AVAILABILITY AND IMPLEMENTATION: dbDSM is freely available online at http://bioinfo.ahu.edu.cn:8080/dbDSM/index.jsp with all major browser supported. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pengbo Wen, Junfeng Xia
Bioinform.3
2016 CINOEDV: a co-information based method for detecting and visualizing n-order epistatic interactions
abstract
BACKGROUND: Detecting and visualizing nonlinear interaction effects of single nucleotide polymorphisms (SNPs) or epistatic interactions are important topics in bioinformatics since they play an important role in unraveling the mystery of "missing heritability". However, related studies are almost limited to pairwise epistatic interactions due to their methodological and computational challenges. RESULTS: We develop CINOEDV (Co-Information based N-Order Epistasis Detector and Visualizer) for the detection and visualization of epistatic interactions of their orders from 1 to n (n ≥ 2). CINOEDV is composed of two stages, namely, detecting stage and visualizing stage. In detecting stage, co-information based measures are employed to quantify association effects of n-order SNP combinations to the phenotype, and two types of search strategies are introduced to identify n-order epistatic interactions: an exhaustive search and a particle swarm optimization based search. In visualizing stage, all detected n-order epistatic interactions are used to construct a hypergraph, where a real vertex represents the main effect of a SNP and a virtual vertex denotes the interaction effect of an n-order epistatic interaction. By deeply analyzing the constructed hypergraph, some hidden clues for better understanding the underlying genetic architecture of complex diseases could be revealed. CONCLUSIONS: Experiments of CINOEDV and its comparison with existing state-of-the-art methods are performed on both simulation data sets and a real data set of age-related macular degeneration. Results demonstrate that CINOEDV is promising in detecting and visualizing n-order epistatic interactions. CINOEDV is implemented in R and is freely available from R CRAN: http://cran.r-project.org and https://sourceforge.net/projects/cinoedv/files/ .
Junliang Shang, Yingxia Sun, Jin-Xing Liu 0001, Junfeng Xia, Chun-Hou Zheng 0001
BMC Bioinform.4
2016 LNDriver: identifying driver genes by integrating mutation and expression data based on gene-gene interaction network
abstract
BACKGROUND: Cancer is a complex disease which is characterized by the accumulation of genetic alterations during the patient's lifetime. With the development of the next-generation sequencing technology, multiple omics data, such as cancer genomic, epigenomic and transcriptomic data etc., can be measured from each individual. Correspondingly, one of the key challenges is to pinpoint functional driver mutations or pathways, which contributes to tumorigenesis, from millions of functional neutral passenger mutations. RESULTS: In this paper, in order to identify driver genes effectively, we applied a generalized additive model to mutation profiles to filter genes with long length and constructed a new gene-gene interaction network. Then we integrated the mutation data and expression data into the gene-gene interaction network. Lastly, greedy algorithm was used to prioritize candidate driver genes from the integrated data. We named the proposed method Length-Net-Driver (LNDriver). CONCLUSIONS: Experiments on three TCGA datasets, i.e., head and neck squamous cell carcinoma, kidney renal clear cell carcinoma and thyroid carcinoma, demonstrated that the proposed method was effective. Also, it can identify not only frequently mutated drivers, but also rare candidate driver genes.
Pi-Jing Wei, Di Zhang 0006, Junfeng Xia, Chun-Hou Zheng 0001
BMC Bioinform.3
2016 A Sequence-Based Dynamic Ensemble Learning System for Protein Ligand-Binding Site Prediction
abstract
BACKGROUND: Proteins have the fundamental ability to selectively bind to other molecules and perform specific functions through such interactions, such as protein-ligand binding. Accurate prediction of protein residues that physically bind to ligands is important for drug design and protein docking studies. Most of the successful protein-ligand binding predictions were based on known structures. However, structural information is not largely available in practice due to the huge gap between the number of known protein sequences and that of experimentally solved structures. RESULTS: This paper proposes a dynamic ensemble approach to identify protein-ligand binding residues by using sequence information only. To avoid problems resulting from highly imbalanced samples between the ligand-binding sites and non ligand-binding sites, we constructed several balanced data sets and we trained a random forest classifier for each of them. We dynamically selected a subset of classifiers according to the similarity between the target protein and the proteins in the training data set. The combination of the predictions of the classifier subset to each query protein target yielded the final predictions. The ensemble of these classifiers formed a sequence-based predictor to identify protein-ligand binding sites. CONCLUSIONS: Experimental results on two Critical Assessment of protein Structure Prediction datasets and the ccPDB dataset demonstrated that of our proposed method compared favorably with the state-of-the-art. AVAILABILITY: http://www2.ahu.edu.cn/pchen/web/LigandDSES.htm.
Peng Chen 0001, Jun Zhang 0011, Xin Gao 0001, Jinyan Li 0001, Junfeng Xia, Bing Wang 0004
IEEE ACM Trans. Comput. Biol. Bioinform.6
2015 Identification of Colorectal Cancer Candidate Genes Based on Subnetwork Extraction Algorithm
Haitao Li 0004, Chun-Hou Zheng 0001, Junfeng Xia
ICIC (3)5
2015 Multi-objective Optimization Method for Identifying Mutated Driver Pathways in Cancer
Junfeng Xia, Yan Zhang 0106, Chun-Hou Zheng 0001
ICIC (2)2
2015 Prediction of Clinical Drug Response Based on Differential Gene Expression Levels
Junfeng Xia
ICIC (2)3
2015 Discovery of Ovarian Cancer Candidate Genes Using Protein Interaction Information
Di Zhang 0006, Qingbao Wang, Rongrong Zhu, Haitao Li 0004, Chun-Hou Zheng 0001, Junfeng Xia
ICIC (2)6
2014 Comparative Assessment of Data Sets of Protein Interaction Hot Spots Used in the Computational Method
Yun-Qiang Di, Changchang Wang, Junfeng Xia
ICIC (3)5
2014 Tumor Clustering Using Independent Component Analysis and Adaptive Affinity Propagation
Fen Ye, Junfeng Xia, Yanwen Chong, Yan Zhang 0106, Chun-Hou Zheng 0001
ICIC (3)2
2014 Potential Driver Genes Regulated by OncomiRNA Are Associated with Druggability in Pan-Negative Melanoma
Di Zhang 0006, Junfeng Xia
ICIC (3)2
2013 Prediction of cytochrome P450 inhibition using ensemble of extreme learning machine
abstract
Adverse side effects of drug-drug interactions induced by human cytochrome P450 (CYP) inhibition play crucial roles in drug discovery. It is urgent and challenging to develop computational methods to efficiently and accurately predict the inhibitive effect of a compound against a specific CYP isoform. In this work we present a novel EELM (ensemble of extreme learning machine) model to predict CYP inhibition. Particularly, extreme learning machine (ELM) and fingerprint descriptors are firstly used to build the weak learning machines. And then EELM is constructed by combining the outputs of each individual ELM using majority voting strategy. Experimental results demonstrate that the proposed method yields good results compared with the existing methods.
Yun-Qiang Di, Chun-Hou Zheng 0001, Junfeng Xia
BIBM4
2013 Differential coexpression analysis in gene modules level and its application to type 2 diabetes
abstract
More and more studies have shown many complex diseases are contributed jointly by alterations of numerous genes. In this paper, we propose a gene differential coexpression analysis algorithm in the level of gene sets and apply the algorithm to a publicly available type 2 diabetes (T2D) expression dataset. The experimental results on simulated data show that the new approach performed well. Moreover, we apply the new approach to clinical data, many additional discoveries can be found through our method.
Lin Yuan 0001, Wen Sha, Jun Zhang 0011, Chun-Hou Zheng 0001, Junfeng Xia
BIBM5
2013 Application of next generation sequencing to human gene fusion detection: computational tools, features and perspectives
abstract
Gene fusions are important genomic events in human cancer because their fusion gene products can drive the development of cancer and thus are potential prognostic tools or therapeutic targets in anti-cancer treatment. Major advancements have been made in computational approaches for fusion gene discovery over the past 3 years due to improvements and widespread applications of high-throughput next generation sequencing (NGS) technologies. To identify fusions from NGS data, existing methods typically leverage the strengths of both sequencing technologies and computational strategies. In this article, we review the NGS and computational features of existing methods for fusion gene detection and suggest directions for future development.
Qingguo Wang, Junfeng Xia, Peilin Jia, William Pao, Zhongming Zhao
Briefings Bioinform.2
2013 Network analysis of gene fusions in human cancer
abstract
Background Gene fusions are hybrid genes formed when two discrete genes are incorrectly joined together. Gene fusions are found to play roles in tumorigenesis. For example, the fusion gene BCR-ABL translates into an abnormal tyrosine kinase that accelerates development of chronic myelogenous leukemia [1]. A network is a relational representation of nodes (e.g., genes) with edges, and is a useful approach to explore biological interactions among many related nodes. Network analysis of gene fusions in cancer would aid the exploration of gene fusion occurrence and association with tumorigenesis. Hoglund et al [2] performed an initial investigation of gene fusions network after collecting 291 tumorigenesis related gene fusions from the Mitelman database in 2006. Since then, gene fusion data has exponentially increased. There is no current and comprehensive cancer-related gene fusion network to assist in targeting cancer-associated genes.
Morgan Harrell, Junfeng Xia, Zhongming Zhao
BMC Bioinform.2
2013 Prediction of protein-protein interactions from amino acid sequences with ensemble extreme learning machines and principal component analysis
abstract
BACKGROUND: Protein-protein interactions (PPIs) play crucial roles in the execution of various cellular processes and form the basis of biological mechanisms. Although large amount of PPIs data for different species has been generated by high-throughput experimental techniques, current PPI pairs obtained with experimental methods cover only a fraction of the complete PPI networks, and further, the experimental methods for identifying PPIs are both time-consuming and expensive. Hence, it is urgent and challenging to develop automated computational methods to efficiently and accurately predict PPIs. RESULTS: We present here a novel hierarchical PCA-EELM (principal component analysis-ensemble extreme learning machine) model to predict protein-protein interactions only using the information of protein sequences. In the proposed method, 11188 protein pairs retrieved from the DIP database were encoded into feature vectors by using four kinds of protein sequences information. Focusing on dimension reduction, an effective feature extraction method PCA was then employed to construct the most discriminative new feature set. Finally, multiple extreme learning machines were trained and then aggregated into a consensus classifier by majority voting. The ensembling of extreme learning machine removes the dependence of results on initial random weights and improves the prediction performance. CONCLUSIONS: When performed on the PPI data of Saccharomyces cerevisiae, the proposed method achieved 87.00% prediction accuracy with 86.15% sensitivity at the precision of 87.59%. Extensive experiments are performed to compare our method with state-of-the-art techniques Support Vector Machine (SVM). Experimental results demonstrate that proposed PCA-EELM outperforms the SVM method by 5-fold cross-validation. Besides, PCA-EELM performs faster than PCA-SVM based method. Consequently, the proposed approach can be considered as a new promising and powerful tools for predicting PPI with excellent performance and less time.
Zhu-Hong You, Ying-Ke Lei, Lin Zhu 0008, Junfeng Xia, Bing Wang 0004
BMC Bioinform.4
2011 Do MicroRNAs Preferentially Target the Genes with Low DNA Methylation Level at the Promoter Region?
Zhixi Su, Junfeng Xia, Zhongming Zhao
ICIC (3)2
2010 Prediction of Protein-Protein Interaction Sites by Using Autocorrelation Descriptor and Support Vector Machine
Xiao-Ming Ren, Junfeng Xia
ICIC (2)2
2010 APIS: accurate prediction of hot spots in protein interfaces by combining protrusion index with solvent accessibility
abstract
BACKGROUND: It is well known that most of the binding free energy of protein interaction is contributed by a few key hot spot residues. These residues are crucial for understanding the function of proteins and studying their interactions. Experimental hot spots detection methods such as alanine scanning mutagenesis are not applicable on a large scale since they are time consuming and expensive. Therefore, reliable and efficient computational methods for identifying hot spots are greatly desired and urgently required. RESULTS: In this work, we introduce an efficient approach that uses support vector machine (SVM) to predict hot spot residues in protein interfaces. We systematically investigate a wide variety of 62 features from a combination of protein sequence and structure information. Then, to remove redundant and irrelevant features and improve the prediction performance, feature selection is employed using the F-score method. Based on the selected features, nine individual-feature based predictors are developed to identify hot spots using SVMs. Furthermore, a new ensemble classifier, namely APIS (A combined model based on Protrusion Index and Solvent accessibility), is developed to further improve the prediction accuracy. The results on two benchmark datasets, ASEdb and BID, show that this proposed method yields significantly better prediction accuracy than those previously published in the literature. In addition, we also demonstrate the predictive power of our proposed method by modelling two protein complexes: the calmodulin/myosin light chain kinase complex and the heat shock locus gene products U and V complex, which indicate that our method can identify more hot spots in these two complexes compared with other state-of-the-art methods. CONCLUSION: We have developed an accurate prediction model for hot spot residues, given the structure of a protein complex. A major contribution of this study is to propose several new features based on the protrusion index of amino acid residues, which has been shown to significantly improve the prediction performance of hot spots. Moreover, we identify a compact and useful feature subset that has an important implication for identifying hot spot residues. Our results indicate that these features are more effective than the conventional evolutionary conservation, pairwise residue potentials and other traditional features considered previously, and that the combination of our and traditional features may support the creation of a discriminative feature set for efficient prediction of hot spot residues. The data and source code are available on web site http://home.ustc.edu.cn/~jfxia/hotspot.html.
Junfeng Xia, Xing-Ming Zhao, Jiangning Song, De-Shuang Huang
BMC Bioinform.1
2009 A New Source and Receiver Localization Method with Erroneous Receiver Positions
Ying-Ke Lei, Junfeng Xia
ICIC (2)2
2007 Inferring Strengths of Protein-Protein Interaction Using Artificial Neural Network
abstract
Many computational methods have been proposed for inference of protein-protein interactions as protein-protein interaction plays an important role in many cellular processes. One of methods is to infer protein-protein interactions based on domain-domain interactions, and the preliminary results have represented their feasibility. In this paper, we use the neural networks for predicting the strengths of protein interaction. This method is capable of exploring all possible interactions between domains and make predictions based on all the domains. Compared to expectation-maximization method and association method, the experimental results show that the proposed schemes can infer strengths of protein-protein interactions with better performances.
Junfeng Xia, Bing Wang 0004, De-Shuang Huang
IJCNN1