Wei Peng 0004

dblp:16/5560-4 · DBLP profile ↗
← Back
58ranked-venue papers
25as first author
44since 2021 · last 2026
0000-0002-9572-951XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 47 · 24 first-author · 33 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 BioMOE-CDG: Pretrained Biological Sequence Embedding-Guided MoE for Cancer Driver Gene Prediction
Wei Dai 0012, Wei Peng 0004, Xiaodong Fu, Li Liu 0032
ISBRA (1)4
2026 Identifying Spatial Domains via Hierarchical Fusion of Multi-scale Biological Priors
Zhihao Ping, Wei Peng 0004, Wei Dai 0012, Xiaodong Fu, Wei Lan 0001, Li Liu 0032
ISBRA (1)2
2026 Cross-modal Orthogonal Adaptive Contrastive Learning for robust chest radiology report generation
Deng Zhu, Jiaman Ding, Wei Peng 0004, Li Liu 0032
Eng. Appl. Artif. Intell.5
2026 Extract-before-mix: Multi-domain topology-aware recovery for one-shot federated clustering under local differential privacy
Xiaodong Fu, Li Liu 0032, Jiaman Ding, Wei Peng 0004
Neurocomputing5
2026 Causal gradient intervention for debiased and evidence-grounded medical visual question answering
Ziyuan Yang 0001, Jiaman Ding, Wei Peng 0004
Medical Image Anal.5
2026 Essential Proteins Prediction Using Features Synergy Model and GO Pure Centrality
abstract
Essential proteins are a crucial component of living organisms, and their absence will lead to cell death or reproductive arrest. Discovering these proteins can propel advancements in synthetic biology and facilitate the development of novel antibiotics and therapies for various diseases. However, current computational methods suffer from two major drawbacks that hinder their discovery rate: one is the significant noise in protein-protein interaction (PPI) data, and the other is the inadequate consideration of feature relationships. To enhance identification capabilities, this study proposes a novel essential protein prediction method, Feature Synergy Method (FSM), which leverages a features synergy model and GO pure centrality. The FSM is described as follows:Firstly, based on the principle of co-expression, gene expression data are integrated with the original PPI network to construct a pure PPI network (PPIN). Subsequently, GO annotation data are employed to calculate GO_sim weights for the interactions within the original PPI network, forming a GS_PIN. The PPIN and GS_PIN are then fused to establish the GS_PPIN, which helps mitigate the impact of noise in PPI data. Secondly, a new centrality measure, GO pure centrality (GPC), is designed based on this GO similarity-weighted pure PPI network. Thirdly, an evolutionary conservation score (ECS) is extracted from subcellular localization and orthologous proteins data. Fourthly, after analyzing the relationship between GPC and ECS, a novel fusion model, the features synergy model, is developed to integrate GPC and ECS, ultimately leading to the proposal of the new essential protein prediction method, FSM. To validate the performance of FSM, six computational methods (PeC, WDC, ION, NCCO, E_POC, and JDC) and six centrality measures (NC, IC, EC, SC, CC, and DC) were evaluated on three distinct yeast datasets. The results demonstrate that FSM achieves a higher essential protein identification rate. Similarly, GPC identifies more essential proteins compared to the six centrality-based approaches (NC, IC, EC, SC, CC, and DC).
Xinlong Luo 0002, Gaoshi Li, Zhipeng Hu, Jingli Wu, Wei Peng 0004, Jiafei Liu 0001, Xiaoshu Zhu
IEEE Trans. Comput. Biol. Bioinform.5
2026 A Deep Learning Framework for Identifying Essential Proteins Based on Vision Transformer
abstract
Essential proteins are fundamental to the reproduction and survival of cells, and if they are killed, the cells will stop reproducing or die. Many computational methods of identifying essential proteins are proposed which fuse a large number of features from multi-omics data. Some of them extract features from subcellular localization data by subjectively selecting certain subcellular locations. Meanwhile, there is still room to improve the identification rate of essential proteins. In this paper, a new deep learning framework for identifying essential proteins based on Vision Transformer is proposed, named EPViT. Firstly, topological features are extracted from the protein-protein interaction network. Secondly, a feature matrix is designed from the subcellular localization information without subjective factor. Then, the two classes of features are fused into a new feature matrix by outer product operation. Finally, the new feature matrix is input into the Vision Transformer model to discover essential proteins. The results show that EPViT has the highest recognition rate among the comparison experiments on yeast data.
Gaoshi Li, Jingli Wu, Wei Peng 0004, Jiafei Liu 0001, Xiaoshu Zhu
IEEE Trans. Comput. Biol. Bioinform.5
2026 Beyond Static Knowledge: Dynamic Context-Aware Cross-Modal Contrastive Learning for Medical Visual Question Answering
abstract
Medical Visual Question Answering (Med-VQA) aims to analyze medical images and accurately respond to natural language queries, thereby optimizing clinical workflows and improving diagnostic and therapeutic outcomes. Although medical images contain rich visual information, the corresponding textual queries frequently lack sufficient descriptive content. This imbalance of information and modality differences leads to significant semantic bias. Furthermore, existing approaches integrate external medical knowledge to enhance model performance, they primarily rely on static knowledge that lacks dynamic adaptation to specific input samples, leading to redundant information and noise interference. To address these challenges, we propose a Contextual Knowledge-Aware Dynamic Perception for the Cross-Modal Reasoning and Alignment (CKRA) Model. To mitigate knowledge redundancy, CKRA employs a dynamic perception mechanism that leverages semantic cues from the query to selectively filter relevant medical knowledge specific to the current sample's context. To alleviate cross-modal semantic bias, CKRA bridges the distance between visual and linguistic features through knowledge-image contrastive learning, optimizing knowledge feature representation and directing the model's attention to key image regions. Further, we design a dual-stream guided attention network that facilitates cross-modal interaction and alignment across multiple dimensions. Experimental results show that the proposed CKRA model outperforms the state-of-the-art method on SLAKE and VQA-RAD datasets. In addition, ablation studies validate the effectiveness of each module, while Grad-CAM maps further demonstrate the feasibility of CKRA for medical visual questioning tasks. The source code and weights of the model are available at https://github.com/cloneiq/CKRA-MedVQA.
Xupeng Feng, Wei Peng 0004, Xiaobing Yang
IEEE Trans. Medical Imaging4
2026 Garment-Aware Neural Radiance Fields for Generalizable 3D Human Digitization
abstract
High-quality garment representation is both a challenge and a key factor in constructing generalized 3D humans from a single-view image. Existing techniques often perform poorly when handling complex garments, primarily due to two critical challenges: (1) Single-view images lack complete information about the garments, limiting the completeness and realism of the reconstruction results. (2) The model’s generalization ability is insufficient, resulting in significant inconsistencies in garment texture and structure when rendered from different viewpoints, which severely impacts the quality of novel view images. To improve the quality of novel view images, we propose a three-stage garment-aware Neural Radiance Field (NeRF) method for generalizable 3D human digitization. To supplement the missing garment information in single-view images, the first garment prior awareness stage focuses on extracting prior knowledge of the garment’s shape, pose deformations, and style. To comprehensively eliminate ambiguities in rendered images across different viewpoints, we then introduce a set of prior-aware feature learning in the second stage to represent garment’s global texture, geometry, and fine details. Additionally, a garment-aware NeRF module with fusion and decoder is designed in the third stage to effectively fuse these prior features, and thus our model can render novel view clothed human and generate high-quality results. Experimental results on RenderPeople, Thuman, and HuMMan datasets demonstrate that our method achieves superior performance and robust generalization in garment representation over the existing methods, especially for synthesizing novel view images of garments without the human body.
Li Liu 0032, Xiaodong Fu, Wei Peng 0004
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Tissue-Aware Prototype Learning Model for Predicting Anticancer Drug Response in Patients
abstract
Cancer treatments often yield different results from patient to patient due to the genomic heterogeneity of tumors. Accurately predicting a patient's response to an anticancer drug is challenging, especially when using traditional machine learning models trained on cell lines and applied to patient data. These models struggle with domain shift (out-of-distribution data due to differences between cell line and patient data), loss of tissue specificity, and data imbalance across domains. To address these issues, we developed the Tissue-Aware and Prototype Learning for Drug Response Prediction (TAPL-DRP) model. This model predicts anticancer drug responses using a two-stage process. At the stage of tissue-aware cross-domain feature extraction, we integrate a Variational Autoencoder (VAE) and a Generative Adversarial Network (GAN) to extract features from both cell line and patient data. This process incorporates tissue prototypes, InfoNCE loss, and class-balance loss to ensure the features are not only domain-invariant but also biologically meaningful and tissue-specific. At the drug response prediction stage, the model combines these extracted features with molecular drug graph features. It then uses the tissue prototypes to guide the training of the classifier, which further improves the prediction accuracy. We tested TAPL-DRP on the TCGA clinical dataset and the PDTC in vitro dataset. The results show that our model significantly outperforms other methods in key metrics like AUC, AUPRC, ACC, and MCC. This demonstrates its effectiveness in handling domain shift and maintaining tissue specificity. Further analysis confirmed that the tissue prototypes, mutual information loss, and class-balance strategies are all crucial components of the model. In summary, TAPL-DRP offers an effective and precise solution for predicting anticancer drug responses in personalized medicine.The source code is available at https://github.com/weiba/TAPL-DRP.
Wei Peng 0004, Wei Dai 0012, Xiaodong Fu, Li Liu 0032, Ning Yu 0004
BIBM2
2025 A Survival Prediction Model Integrating Hierarchical Pathological Image and Pathway Features
Wei Peng 0004, Wei Dai 0012, Xiaodong Fu, Li Liu 0032
ISBRA (2)2
2025 Multi-domains personalized local differential privacy frequency estimation mechanism for utility optimization
abstract
Local Differential Privacy (LDP) has garnered considerable attention in recent years because it does not rely on trusted third parties and has low interactivity and high operational efficiency. However, current LDP frequency estimation mechanisms aggregate data using different privacy budgets within the same domain of attribute values, overlooking the aggregation requirements across different domains of attribute values. This limits the potential for enhancing the data utility under fixed privacy budgets and meeting user preferences in multiple domains of attribute values and privacy budgets. To address this issue, we define a Multi-Domains Personalized Local Differential Privacy (MDPLDP) model that allows users to freely choose domains of attribute values and privacy budgets according to their privacy preferences. Furthermore, based on the MDPLDP model, two new frequency estimation mechanisms are proposed: MDPLDP-Generalized Randomized Response and MDPLDP-basic Randomized Aggregatable Privacy-Preserving Ordinal Response. These mechanisms support cross-domains data aggregation and optimize data utility by adjusting the domains of attribute values and increasing privacy budgets. Theoretical analysis reveals that these new mechanisms have lower estimation errors than the traditional LDP mechanisms. Experiments on real and synthetic datasets demonstrate that the proposed mechanisms effectively reduce estimation errors and enhance the utility of data-frequency estimation.
Xiaodong Fu, Li Liu 0032, Jiaman Ding, Wei Peng 0004, Lianyin Jia
Comput. Secur.5
2025 Instance-Category Feature Representation and Association Learning for Multi-Human Parsing
Lanqing Ye, Li Liu 0032, Xiaodong Fu, Wei Peng 0004
IET Image Process.5
2025 Fusion of brain imaging genetic data for alzheimer's disease diagnosis and causal factors identification using multi-stream attention mechanisms and graph convolutional networks
Wei Peng 0004, Yanhan Ma, Chunshan Li, Wei Dai 0012, Xiaodong Fu, Li Liu 0032, Jin Liu 0012
Neural Networks1
2025 Using Multi-Feature Weak Consensus Model to Discover Essential Proteins
abstract
Essential proteins play an essential role in cell survival and replication. Currently, more and more computational methods are developed to identify essential proteins, which overcome the time-consuming, costly and inefficient shortcomings with biological experimental methods. In order to improve the recognition rate, some new methods by fusing multiple features are developed, but they seldom consider the connection among features. After analyzing a large number of methods based on multi-feature fusion, a phenomenon among features is found, called weak consensus, then a weak consensus model to fuse these features is proposed in this paper. After analyzing the relationship between a protein and its neighbors in protein-protein interaction networks, a new centrality, namely neighborhood aggregation centrality(NAC) is developed in this paper. Then, a Max-Min strategy is used to integrate NAC with Pearson correlation coefficient and Jaccard similarity coefficient based on gene expression data to obtain local importance score. In addition, orthologous feature score is used to measure proteins conservation. Finally, by using the weak consensus model to fuse orthologous feature score with local importance score, a new method WOL is proposed in this paper. Then experiments are performed on S.cerevisiae data. The results show that compared with WDC, PeC, ION, JDC, NCCO and E_POC, WOL has a higher recognition rate.
Zhipeng Hu, Gaoshi Li, Xinlong Luo 0002, Jiafei Liu 0001, Jingli Wu, Wei Peng 0004, Xiaoshu Zhu
IEEE Trans. Comput. Biol. Bioinform.6
2025 Guest Editorial: Introduction to the Special Section on Bioinformatics Research and Applications
Wei Peng 0004, Zhipeng Cai 0001, Alex Zelikovsky
IEEE Trans. Comput. Biol. Bioinform.1
2025 Predicting Anti-Cancer Drug Response Based on Hypergraph Representation Learning
abstract
Accurate prediction of drug responses is critical for advancing personalized cancer therapies. Although current graph neural network (GNN)-based approaches predominantly focus on pairwise interactions between cell lines and drugs, they often neglect the potential of higher-order interactions. In this study, we present HRLCDR, a novel computational framework that utilizes Hypergraph Representation Learning to predict Cancer Drug Responses. HRLCDR begins by constructing hypergraphs for both cell lines and drugs and then processes through low-pass and high-pass hypergraph convolutions, allowing the model to extract both common and different features from the complex higher-order interactions between cell lines and drugs. After that, HRLCDR constructs a heterogeneous graph using known cell line responses to drugs. Parallel heterogeneous graph convolution operations are then employed to extract primary interaction features between cell lines and drugs from these associations. Finally, HRLCDR integrates the features learned from both the hypergraphs and the heterogeneous graph, predicting drug response via Classifiers. We evaluated HRLCDR's performance on two major cancer drug response datasets: the Cancer Drug Sensitivity Data (GDSC) and the Cancer Cell Line Encyclopedia (CCLE). The results demonstrate that HRLCDR outperforms current state-of-the-art methods, underscoring its potential to enhance the accuracy and reliability of cancer drug response predictions.
Wei Peng 0004, Jiangzhen Lin, Wei Dai 0012, Xiaodong Fu, Li Liu 0032, Ning Yu 0004
IEEE Trans. Comput. Biol. Bioinform.1
2025 The Large Language Models on Biomedical Data Analysis: A Survey
abstract
With the rapid development of Large Language Model (LLM) technology, it has become an indispensable force in biomedical data analysis research. However, biomedical researchers currently have limited knowledge about LLM. Therefore, there is an urgent need for a summary of LLM applications in biomedical data analysis. Herein, we propose this review by summarizing the latest research work on LLM in biomedicine. In this review, LLM techniques are first outlined. We then discuss biomedical datasets and frameworks for biomedical data analysis, followed by a detailed analysis of LLM applications in genomics, proteomics, transcriptomics, radiomics, single-cell analysis, medical texts and drug discovery. Finally, the challenges of LLM in biomedical data analysis are discussed. In summary, this review is intended for researchers interested in LLM technology and aims to help them understand and apply LLM in biomedical data analysis research.
Wei Lan 0001, Zhentao Tang, Qingfeng Chen, Wei Peng 0004, Yi-Ping Phoebe Chen, Yi Pan 0001
IEEE J. Biomed. Health Informatics5
2025 Predicting Clinical Anticancer Drug Response of Patients by Using Domain Alignment and Prototypical Learning
abstract
Anticancer drug response prediction is crucial in developing personalized treatment plans for cancer patients. However, High-quality patient anticancer drug response data are scarce and cell line data and patient data have different distributions, models trained solely on cell line data perform poorly. Some existing methods predict anticancer drug response by transferring knowledge from the cell line domain to the patient domain using transfer learning. However, the robustness of these classifiers is affected by anomalies in the cell line data, and they do not utilize the knowledge in the unlabeled target domain data. To this end, we proposed a model called DAPL to predict patient responses to anticancer drugs. The model extracts domain-invariant features from cell lines and patients by constructing multiple VAEs and extracts drug features using GNNs. These features are then combined for prototypical learning to train a classifier, resulting in better predictions of patient anticancer drug response. We used the cell line datasets CCLE and GDSC as source domains and the patient datasets TCGA and PDTC as target domains and conducted experiments. The results indicate that DAPL shows excellent performance in predicting patient anticancer drug response compared to other state-of-the-art methods.
Wei Peng 0004, Chuyue Chen, Wei Dai 0012, Ning Yu 0004, Jianxin Wang 0001
IEEE J. Biomed. Health Informatics1
2025 Hierarchical Graph Representation Learning With Multi-Granularity Features for Anti-Cancer Drug Response Prediction
abstract
Patients with the same type of cancer often respond differently to identical drug treatments due to unique genomic traits. Accurately predicting a patient's response to drug is crucial in guiding treatment decisions, alleviating patient suffering, and improving cancer prognosis. Current computational methods utilize deep learning models trained on extensive drug screening data to predict anti-cancer drug responses based on features of cell lines and drugs. However, the interaction between cell lines and drugs is a complex biological process involving interactions across various levels, from internal cellular and drug structures to the external interactions among different molecules.To address this complexity, we propose a novel Hierarchical graph representation Learning with Multi-Granularity features (HLMG) algorithm for predicting anti-cancer drug responses. The HLMG algorithm combines features at two granularities: the overall gene expression and pathway substructures of cell lines, and the overall molecular fingerprints and substructures of drugs. Subsequently, it constructs a heterogeneous graph including cell lines, drugs, known cell line-drug responses, and the associations between similar cell lines and similar drugs. Through a graph convolutional network model, the HLMG learns the final cell line and drug representations by aggregating features of their multi-level neighbor in the heterogeneous graph. The multi-level neighbors consist of the node self, directly related drugs/cell lines, and indirectly related similar drugs/cell lines. Finally, a linear correlation coefficient decoder is employed to reconstruct the cell line-drug correlation matrix to predict anti-cancer drug responses. Our model was tested on the Genomics of Drug Sensitivity in Cancer (GDSC) and the Cancer Cell Line Encyclopedia (CCLE) databases. Results indicate that HLMG outperforms other state-of-the-art methods in accurately predicting anti-cancer drug responses.
Wei Peng 0004, Jiangzhen Lin, Wei Dai 0012, Ning Yu 0004, Jianxin Wang 0001
IEEE J. Biomed. Health Informatics1
2025 Unified 3D Gaussian splatting for motion and defocus blur reconstruction
abstract
This paper proposes a unified 3D Gaussian splatting framework consisting of three key components for motion and defocus blur reconstruction. First, a dual-blur perception module is designed to generate pixel-wise masks and predict the types of motion and defocus blur, guiding structural feature extraction. Second, a blur-aware Gaussian splatting integrates blur-aware features into the splatting process for accurate modeling of the global and local scene structure. Third, an Unoptimized Gaussian Ratio (UGR)-opacity joint optimization strategy is proposed to refine under-optimized regions, improving reconstruction accuracy under complex blur conditions. Experiments on a newly constructed motion and defocus blur dataset demonstrate the effectiveness of the proposed method for novel view synthesis. Compared with state-of-the-art methods, our framework achieves improvements of 0.28 dB, 2.46% and 39.88% on PSNR, SSIM, and LPIPS, respectively. For deblurring tasks, it achieves improvements of 0.36 dB, 3.24% and 28.96% on the same metrics. These results highlight the robustness and effectiveness of this approach. Additional visual results and video renderings are available on our project webpage: https://sunbeam-217.github.io/Dual-blur-reconstruction/ .
Li Liu 0032, Jing Duan, Xiaodong Fu, Wei Peng 0004
Vis. Informatics4
2024 Sparse Attention-based Hierarchical Node Representation for Spatial Domain Identification
abstract
Using deep learning models on spatial transcriptomics data to identify the spatial domain is crucial for uncovering the spatial distribution of cells and gene expression patterns within tissues, essential for understanding complex biological processes and disease mechanisms. Existing methods for spatial domain partitioning often rely on predefined adjacency relationships at a single scale, overlooking the hierarchical structure and functional characteristics of biological tissues. In this paper, we propose SpaNFM, a novel method that leverages sparse attention-based hierarchical node representation and multi-view contrastive learning for spatial domain identification in spatial transcriptomics data. The SpaNFM first treats each spot as a node and constructs two views using different data augmentation techniques based on tissue image information, gene expression profiles, and spatial coordinates of cells. Subsequently, SpaNFM utilizes a sparse attention-based hierarchical node fusion module to generate coarse-grained node representations. This fine-to-coarse hierarchical structure integrates complementary information from multi-granularity node features and reduces model complexity due to the decreased node size. The model parameters are updated using gene expression reconstruction loss and contrastive loss on the coarse-grained node representations from the two views. Finally, the learned node features are subjected to downstream clustering using the Leiden algorithm. We tested SpaNFM on the human dorsolateral prefrontal cortex dataset. The results demonstrate that SpaNFM outperforms other state-of-the-art methods in most cases. The data and code are available at: https://github.com/weiba/SpaNFM
Wei Peng 0004, Zhihao Ping, Wei Dai 0012, Xiaodong Fu, Li Liu 0032, Ning Yu 0004
BIBM1
2024 Patient Anticancer Drug Response Prediction Based on Single-Cell Deconvolution
Wei Peng 0004, Chuyue Chen, Wei Dai 0012
ISBRA (3)1
2024 Hypergraph Representation Learning for Cancer Drug Response Prediction
Wei Peng 0004, Jiangzhen Lin, Wei Dai 0012, Xiaodong Fu, Li Liu 0032
ISBRA (2)1
2024 DGCL: A Contrastive Learning Method for Predicting Cancer Driver Genes Based on Graph Diffusion
Wei Peng 0004, Zhengnan Zhou, Wei Dai 0012, Xinping Xu, Xiaodong Fu, Li Liu 0032
ISBRA (2)1
2024 MHCLMDA: multihypergraph contrastive learning for miRNA-disease association prediction
abstract
The correct prediction of disease-associated miRNAs plays an essential role in disease prevention and treatment. Current computational methods to predict disease-associated miRNAs construct different miRNA views and disease views based on various miRNA properties and disease properties and then integrate the multiviews to predict the relationship between miRNAs and diseases. However, most existing methods ignore the information interaction among the views and the consistency of miRNA features (disease features) across multiple views. This study proposes a computational method based on multiple hypergraph contrastive learning (MHCLMDA) to predict miRNA-disease associations. MHCLMDA first constructs multiple miRNA hypergraphs and disease hypergraphs based on various miRNA similarities and disease similarities and performs hypergraph convolution on each hypergraph to capture higher order interactions between nodes, followed by hypergraph contrastive learning to learn the consistent miRNA feature representation and disease feature representation under different views. Then, a variational auto-encoder is employed to extract the miRNA and disease features in known miRNA-disease association relationships. Finally, MHCLMDA fuses the miRNA and disease features from different views to predict miRNA-disease associations. The parameters of the model are optimized in an end-to-end way. We applied MHCLMDA to the prediction of human miRNA-disease association. The experimental results show that our method performs better than several other state-of-the-art methods in terms of the area under the receiver operating characteristic curve and the area under the precision-recall curve.
Wei Peng 0004, Zhichen He, Wei Dai 0012, Wei Lan 0001
Briefings Bioinform.1
2024 Identification of Cancer Driver Genes based on Dynamic Incentive Model
abstract
Cancer is a complex genomic mutation disease, and identifying cancer driver genes promotes the development of targeted drugs and personalized therapies. The current computational method takes less consideration of the relationship among features and the effect of noise in protein-protein interaction(PPI) data, resulting in a low recognition rate. In this paper, we propose a cancer driver genes identification method based on dynamic incentive model, DIM. This method firstly constructs a hypergraph to reduce the impact of false positive data in PPI. Then, the importance of genes in each hyperedge in hypergraph is considered from three perspectives, network and functional score(NFS) is proposed. By analyzing the relation among features, the dynamic incentive model is proposed to fuse NFS, the differential expression score of mRNA and the differential expression score of miRNA. DIM is compared with some classical methods on breast cancer, lung cancer, prostate cancer, and pan-cancer datasets. The results show that DIM has the best performance on statistical evaluation indicators, functional consistency and the partial area under the ROC curve, and has good cross-cancer capability.
Zhipeng Hu, Gaoshi Li, Xinlong Luo 0002, Wei Peng 0004, Jiafei Liu 0001, Xiaoshu Zhu, Jingli Wu
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 Essential proteins identification based on weak consensus model and neighborhood aggregation centrality
abstract
Essential proteins play an essential role in cell survival and replication. Currently, more and more computational methods are developed to identify essential proteins, which overcome the time-consuming, costly and inefficient shortcomings with biological experimental methods. In order to improve the recognition rate, some new methods by fusing multiple features are developed, but they seldom consider the connection among features. After analyzing a large number of methods based on multi-feature fusion, a weak consensus model to fuse features is proposed in this paper. Then, this paper uses the weak consensus model to fuse protein-protein interaction network, gene expression data, and orthologous data, thus proposing a new method, WOL. Then experiments are performed on one S.cerevisiae dataset. The results show that compared with WDC, PeC, ION, JDC, NCCO and E_POC, WOL has a higher recognition rate.
Zhipeng Hu, Gaoshi Li, Jingli Wu, Xinlong Luo 0002, Jiafei Liu 0001, Wei Peng 0004, Xiaoshu Zhu
BIBM6
2023 A multi-view comparative learning method for spatial transcriptomics data clustering
abstract
Clustering individual cells or spots based on their gene expression profiles in a spatial context is a powerful approach to uncovering the underlying biological diversity and relationships among cells. The intricate information within spatial transcriptomics data demands sophisticated algorithms that effectively integrate gene expression, cell position, and tissue image data for accurate cell or spot clustering. This work proposes a Multi-View Comparative Learning method for clustering Spatial Transcriptomics data (MVCLST). MVCLST first builds on two data views using gene expression profiles, cell space coordinates, and image features. Then it employs four different encoders to capture the common and private features of the two views. The model employs a contrastive learning loss to encourage effective interaction between the two views and ensure feature consistency. The shared and private features from both views are fused using corresponding decoders. Finally, the model employs the Leiden algorithm for downstream clustering of the learned features. We test the MVCLST method on a human dorsolateral prefrontal cortex dataset. The results show that MVCLST outperforms other state-of-the-art methods in most cases. Additionally, the clusters identified by MVCLST align closely with manual annotations and established neuroscience definitions.
Wei Peng 0004, Wei Dai 0012, Xiaodong Fu, Li Liu 0032, Ning Yu 0004
BIBM1
2023 Identifying cancer driver genes based on multi-view heterogeneous graph convolutional network and self-attention mechanism
abstract
BACKGROUND: Correctly identifying the driver genes that promote cell growth can significantly assist drug design, cancer diagnosis and treatment. The recent large-scale cancer genomics projects have revealed multi-omics data from thousands of cancer patients, which requires to design effective models to unlock the hidden knowledge within the valuable data and discover cancer drivers contributing to tumorigenesis. RESULTS: In this work, we propose a graph convolution network-based method called MRNGCN that integrates multiple gene relationship networks to identify cancer driver genes. First, we constructed three gene relationship networks, including the gene-gene, gene-outlying gene and gene-miRNA networks. Then, genes learnt feature presentations from the three networks through three sharing-parameter heterogeneous graph convolution network (HGCN) models with the self-attention mechanism. After that, these gene features pass a convolution layer to generate fused features. Finally, we utilized the fused features and the original feature to optimize the model by minimizing the node and link prediction losses. Meanwhile, we combined the fused features, the original features and the three features learned from every network through a logistic regression model to predict cancer driver genes. CONCLUSIONS: We applied the MRNGCN to predict pan-cancer and cancer type-specific driver genes. Experimental results show that our model performs well in terms of the area under the ROC curve (AUC) and the area under the precision-recall curve (AUPRC) compared to state-of-the-art methods. Ablation experimental results show that our model successfully improved the cancer driver identification by integrating multiple gene relationship networks.
Wei Peng 0004, Wei Dai 0012, Ning Yu 0004
BMC Bioinform.1
2023 Feature Representation for High-resolution Clothed Human Reconstruction
abstract
Abstract Detailed and accurate feature representation is essential for high‐resolution reconstruction of clothed human. Herein we introduce a unified feature representation for clothed human reconstruction, which can adapt to changeable posture and various clothing details. The whole method can be divided into two parts: the human shape feature representation and the details feature representation. Specifically, we firstly combine the voxel feature learned from semantic voxel with the pixel feature from input image as an implicit representation for human shape. Then, the details feature mixed with the clothed layer feature and the normal feature is used to guide the multi‐layer perceptron to capture geometric surface details. The key difference from existing methods is that we use the clothing semantics to infer clothed layer information, and further restore the layer details with geometric height. We qualitative and quantitative experience results demonstrate that proposed method outperforms existing methods in terms of handling limb swing and clothing details. Our method provides a new solution for clothed human reconstruction with high‐resolution details (style, wrinkles and clothed layers), and has good potential in three‐dimensional virtual try‐on and digital characters.
Juncheng Pu, Li Liu 0032, Xiaodong Fu, Zhuo Su 0001, Wei Peng 0004
Comput. Graph. Forum6
2023 Crowded pose-guided multi-task learning for instance-level human parsing
Li Liu 0032, Xiaodong Fu, Wei Peng 0004
Mach. Vis. Appl.5
2023 Predicting miRNA-Disease Associations From miRNA-Gene-Disease Heterogeneous Network With Multi-Relational Graph Convolutional Network Model
abstract
MiRNAs are reported to be linked to the pathogenesis of human complex diseases. Disease-related miRNAs may serve as novel bio-marks and drug targets. This work focuses on designing a multi-relational Graph Convolutional Network model to predict miRNA-disease associations (HGCNMDA) from a Heterogeneous network. HGCNMDA introduces a gene layer to construct a miRNA-gene-disease heterogeneous network. We refine the features of nodes into initial and inductive features so that the direct and indirect associations between diseases and miRNA can be considered simultaneously. Then HGCNMDA learns feature embeddings for miRNAs and disease through a multi-relational graph convolutional network model that can assign appropriate weights to different types of edges in the heterogeneous network. Finally, the miRNA-disease associations were decoded by the inner product between miRNA and disease feature embeddings. We apply our model to predict human miRNA-disease associations. The HGCNMDA is superior to the other state-of-the-art models in identifying missing miRNA-disease associations and also performs well on recommending related miRNAs/diseases to new diseases/ miRNAs. The codes are available at https://github.com/weiba/HGCNMDA.
Wei Peng 0004, Zicheng Che, Wei Dai 0012, Shoulin Wei, Wei Lan 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2023 Multi-View Feature Aggregation for Predicting Microbe-Disease Association
abstract
Microbes play a crucial role in human health and disease. Figuring out the relationship between microbes and diseases leads to significant potential applications in disease treatments. It is an urgent need to devise robust and effective computational methods for identifying disease-related microbes. This work proposes a Multi-View Feature Aggregation (MVFA) scheme that integrates the linear and nonlinear features to identify disease-related microbes. We introduce a non-negative matrix tri-factorization (NMTF) model to extract linear features for diseases and microbes. Then we learn another type of linear feature by utilizing a bi-random walk model. The nonlinear feature is obtained by inputting the two kinds of linear features into a capsule neural network. These three types of features describe the associations between diseases and microbes from different views. Finally, considering the complementary of these features, we leverage a logistic regression model to combine the NMTF model predictions, bi-random walk model predictions, and the capsule neural network predictions to obtain the final microbe-disease pair scores. We apply our method to predict human microbe-disease associations on two datasets. Experimental results show that our multi-view model outperforms the state-of-the-art models in recovering missing microbe-disease associations and predicting associations for new microbes. The ablation study shows that aggregating multi-view linear and nonlinear features can improve the prediction performance. Case studies on two diseases, i.e. Type 1 diabetes and Liver cirrhosis, further validate our method effectiveness.
Wei Peng 0004, Wei Dai 0012, Tielin Chen, Yi Pan 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2022 Identification of personalized driver genes for individuals using graph convolution network
abstract
The correct identification of the driver genes that lead to cancer development is essential for understanding the mechanisms of cancer and developing drugs to treat it. Currently, most computational methods for identifying cancer driver genes are based on a cohort of patients. However, due to the heterogeneity of cancers, patients diagnosed with the same cancers may have different genomic characteristics and present varied clinical symptoms. It requires devising effective methods to identify personalized cancer driver genes in an individual. This work developed a novel method to predict personalized cancer driver genes of a single sample based on graph convolution networks, namely pDriverGCN. pDriverGCN constructed a mutant gene-sample heterogeneous network according to the known driver genes of samples. Then it employed two separate graph convolution network models to learn feature representations for genes and samples by gathering the features of themselves and their neighbors. Finally, pDriverGCN used the feature representations to reconstruct the association matrix between genes and samples through a linear correlation coefficient decoder. We apply our model to identify personalized driver genes of samples on the TCGA datasets. The experimental results show that our model outperforms state-of-the-art methods being evaluated at both population and individual levels.
Wei Peng 0004, Piaofang Yu, Wei Dai 0012, Xiaodong Fu, Li Liu 0032, Yi Pan 0001
BIBM1
2022 Collusion Attack Analysis and Detection of DPoS Consensus Mechanism
Xinxin Qi, Xiaodong Fu, Fei Dai 0002, Li Liu 0032, Jiaman Ding, Wei Peng 0004
BlockSys7
2022 Improving cancer driver gene identification using multi-task learning on graph convolutional network
abstract
Cancer is thought to be caused by the accumulation of driver genetic mutations. Therefore, identifying cancer driver genes plays a crucial role in understanding the molecular mechanism of cancer and developing precision therapies and biomarkers. In this work, we propose a Multi-Task learning method, called MTGCN, based on the Graph Convolutional Network to identify cancer driver genes. First, we augment gene features by introducing their features on the protein-protein interaction (PPI) network. After that, the multi-task learning framework propagates and aggregates nodes and graph features from input to next layer to learn node embedding features, simultaneously optimizing the node prediction task and the link prediction task. Finally, we use a Bayesian task weight learner to balance the two tasks automatically. The outputs of MTGCN assign each gene a probability of being a cancer driver gene. Our method and the other four existing methods are applied to predict cancer drivers for pan-cancer and some single cancer types. The experimental results show that our model shows outstanding performance compared with the state-of-the-art methods in terms of the area under the Receiver Operating Characteristic (ROC) curves and the area under the precision-recall curves. The MTGCN is freely available via https://github.com/weiba/MTGCN.
Wei Peng 0004, Wei Dai 0012, Tielin Chen
Briefings Bioinform.1
2022 Predicting cancer drug response using parallel heterogeneous graph convolutional networks with neighborhood interactions
abstract
MOTIVATION: Due to cancer heterogeneity, the therapeutic effect may not be the same when a cohort of patients of the same cancer type receive the same treatment. The anticancer drug response prediction may help develop personalized therapy regimens to increase survival and reduce patients' expenses. Recently, graph neural network-based methods have aroused widespread interest and achieved impressive results on the drug response prediction task. However, most of them apply graph convolution to process cell line-drug bipartite graphs while ignoring the intrinsic differences between cell lines and drug nodes. Moreover, most of these methods aggregate node-wise neighbor features but fail to consider the element-wise interaction between cell lines and drugs. RESULTS: This work proposes a neighborhood interaction (NI)-based heterogeneous graph convolution network method, namely NIHGCN, for anticancer drug response prediction in an end-to-end way. Firstly, it constructs a heterogeneous network consisting of drugs, cell lines and the known drug response information. Cell line gene expression and drug molecular fingerprints are linearly transformed and input as node attributes into an interaction model. The interaction module consists of a parallel graph convolution network layer and a NI layer, which aggregates node-level features from their neighbors through graph convolution operation and considers the element-level of interactions with their neighbors in the NI layer. Finally, the drug response predictions are made by calculating the linear correlation coefficients of feature representations of cell lines and drugs. We have conducted extensive experiments to assess the effectiveness of our model on Cancer Drug Sensitivity Data (GDSC) and Cancer Cell Line Encyclopedia (CCLE) datasets. It has achieved the best performance compared with the state-of-the-art algorithms, especially in predicting drug responses for new cell lines, new drugs and targeted drugs. Furthermore, our model that was well trained on the GDSC dataset can be successfully applied to predict samples of PDX and TCGA, which verified the transferability of our model from cell line in vitro to the datasets in vivo. AVAILABILITY AND IMPLEMENTATION: The source code can be obtained from https://github.com/weiba/NIHGCN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wei Peng 0004, Hancheng Liu, Wei Dai 0012, Ning Yu 0004, Jianxin Wang 0001
Bioinform.1
2022 GANLDA: Graph attention network for lncRNA-disease associations prediction
Wei Lan 0001, Ximin Wu, Qingfeng Chen, Wei Peng 0004, Jianxin Wang 0001, Yi-Ping Phoebe Chen
Neurocomputing4
2022 Identifying Cancer Patient Subgroups by Finding Co-Modules From the Driver Mutation Profiles and Downstream Gene Expression Profiles
abstract
Nowadays, the heterogeneous characteristics of cancer patients throw a big challenge to precision medicine and targeted therapy. Identifying cancer subtypes shed new light on effective personalized cancer medicine, future therapeutic strategies and minimizing treatment-related costs. Recently, there are many clustering methods have been proposed in categorizing cancer patients. Although these methods obtained a certain achievement in cancer subtype identification, they still fail to fully use the prior known biological information in the model designing process to improve precision and efficiency. It is acknowledged that the driver gene always regulates its downstream genes in the network to perform a certain function. By analyzing the known clinic cancer subtype data, we found some special co-pathways between the driver genes and the downstream genes in the cancer patients of the same subgroup. Hence, we proposed a novel model named DDCMNMF(Driver and Downstream gene Co-Module Assisted Multiple Non-negative Matrix Factorization model) that first for cancer subtypes by identifying co-modules of driver genes and downstream genes. We applied our model on lung and breast cancer datasets and compared it with the other four state-of-the-art models. The final results show that our model could identify the cancer subtypes with high compactness and separateness and achieve a high degree of consistency with the known cancer subtypes. The survival time analysis further proves the significant clinical characteristic of identified cancer subgroups by our model. Availability and implementation: It is available at https://github.com/weiba/DDCMNMF/.
Junrong Song, Wei Peng 0004
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 Predicting Drug Response Based on Multi-Omics Fusion and Graph Convolution
abstract
Different cancer patients may respond differently to cancer treatment due to the heterogeneity of cancer. It is an urgent task to develop an efficient computational method to identify drug responses in different cell lines, which guides us to design personalized therapy for an individual patient. Hence, we propose an end-to-end algorithm, namely MOFGCN, to predict drug response in cell lines based on Multi-Omics Fusion and Graph Convolution Network. MOFGCN first fuses multiple omics data to calculate the cell line similarity and then constructs a heterogeneous network by combining the cell line similarity, drug similarity, and the known cell line-drug associations. Secondly, it learns the latent features for cancer cell lines and drugs by performing graph convolution operations on the heterogeneous network. Finally, MOFGCN applies the linear correlation coefficient to reconstruct the cancer cell line-drug correlation matrix to predict drug sensitivity. To our knowledge, this is the first attempt to combine graph convolutional neural network and linear correlation coefficient for this significant task. We performed extensive evaluation experiments on the Genomics of Drug Sensitivity in Cancer (GDSC) and Cancer Cell Line Encyclopedia (CCLE) databases to validate MOFGCN's performance. The experimental results show that MOFGCN is superior to the state-of-the-art algorithms in predicting missing drug responses. It also leads to higher performance in predicting drug responses for new cell lines, new drugs, and targeted drugs.
Wei Peng 0004, Tielin Chen, Wei Dai 0012
IEEE J. Biomed. Health Informatics1
2021 A Heterogeneous Graph Convolutional Network-Based Deep Learning Model to Identify miRNA-Disease Association
Zicheng Che, Wei Peng 0004, Wei Dai 0012, Shoulin Wei, Wei Lan 0001
ISBRA2
2021 Proteoform characterization based on top-down mass spectrometry
abstract
Proteins are dominant executors of living processes. Compared to genetic variations, changes in the molecular structure and state of a protein (i.e. proteoforms) are more directly related to pathological changes in diseases. Characterizing proteoforms involves identifying and locating primary structure alterations (PSAs) in proteoforms, which is of practical importance for the advancement of the medical profession. With the development of mass spectrometry (MS) technology, the characterization of proteoforms based on top-down MS technology has become possible. This type of method is relatively new and faces many challenges. Since the proteoform identification is the most important process in characterizing proteoforms, we comprehensively review the existing proteoform identification methods in this study. Before identifying proteoforms, the spectra need to be preprocessed, and protein sequence databases can be filtered to speed up the identification. Therefore, we also summarize some popular deconvolution algorithms, various filtering algorithms for improving the proteoform identification performance and various scoring methods for localizing proteoforms. Moreover, commonly used methods were evaluated and compared in this review. We believe our review could help researchers better understand the current state of the development in this field and design new efficient algorithms for the proteoform characterization.
Jiancheng Zhong, Yusui Sun, Minzhu Xie, Wei Peng 0004, Chushu Zhang, Fang-Xiang Wu, Jianxin Wang 0001
Briefings Bioinform.4
2021 A novel essential protein identification method based on PPI networks and gene expression data
abstract
BACKGROUND: Some proposed methods for identifying essential proteins have better results by using biological information. Gene expression data is generally used to identify essential proteins. However, gene expression data is prone to fluctuations, which may affect the accuracy of essential protein identification. Therefore, we propose an essential protein identification method based on gene expression and the PPI network data to calculate the similarity of "active" and "inactive" state of gene expression in a cluster of the PPI network. Our experiments show that the method can improve the accuracy in predicting essential proteins. RESULTS: In this paper, we propose a new measure named JDC, which is based on the PPI network data and gene expression data. The JDC method offers a dynamic threshold method to binarize gene expression data. After that, it combines the degree centrality and Jaccard similarity index to calculate the JDC score for each protein in the PPI network. We benchmark the JDC method on four organisms respectively, and evaluate our method by using ROC analysis, modular analysis, jackknife analysis, overlapping analysis, top analysis, and accuracy analysis. The results show that the performance of JDC is better than DC, IC, EC, SC, BC, CC, NC, PeC, and WDC. We compare JDC with both NF-PIN and TS-PIN methods, which predict essential proteins through active PPI networks constructed from dynamic gene expression. CONCLUSIONS: We demonstrate that the new centrality measure, JDC, is more efficient than state-of-the-art prediction methods with same input. The main ideas behind JDC are as follows: (1) Essential proteins are generally densely connected clusters in the PPI network. (2) Binarizing gene expression data can screen out fluctuations in gene expression profiles. (3) The essentiality of the protein depends on the similarity of "active" and "inactive" state of gene expression in a cluster of the PPI network.
Jiancheng Zhong, Wei Peng 0004, Minzhu Xie, Yusui Sun, Qiang Tang 0014, Qiu Xiao, Jiahong Yang 0001
BMC Bioinform.3
2020 A multi-view approach for predicting microbedisease associations by fusing the linear and nonlinear features
abstract
Microbes play a crucial role in human health and disease. Understanding the relationship between microbes and diseases is conducive to the treatment and diagnosis of diseases. Recently, many computational methods have been proposed to predict disease-microbe associations. However, most of the existing methods only consider a single model and explore the disease-microbe associations from a single view. To improve the prediction accuracy, we propose a novel multi-view approach that fuses the linear and nonlinear features to predict new potential associations between diseases and microbes. We first design a non-negative matrix tri-factorization method to extract the linear features of diseases and microbes. We input the linear features from the non-negative matrix tri-factorization model and bi-random walk model into a capsule neural network to obtain the diseases and microbes' nonlinear features. Finally, we leverage a logistic regression model to combine the non-negative matrix tri-factorization model predictions, bi-random walk model predictions and the capsule neural network predictions to obtain the final association scores between microbes and diseases. We apply our method to predict human microbedisease associations. Experimental results show that our fusion model outperforms the non-negative matrix tri-factorization model, bi-random walk model and other existing models.
Wei Dai 0012, Wei Peng 0004, Yi Pan 0001
BIBM3
2020 An Entropy-Based Method for Identifying Mutual Exclusive Driver Genes in Cancer
abstract
Cancer in essence is a complex genomic alteration disease which is caused by the somatic mutations during the lifetime. According to previous researches, the first step to overcome cancer is to identify driver genes which can promote carcinogenesis. However, it is still a big challenge to precisely and efficiently extract the cancer related driver genes because the nature of cancer is heterogeneous and there exists tremendously irrelevant passenger mutations which have no function impact on the cancer's development. In this work, we proposed a novel entropy-based method namely EntroRank to identify driver genes by integrating the subcellular localization information and mutual exclusive of variation frequency into the network. EntroRank can take into full consideration different properties of driver genes. Considering the modularity of driver genes, the mutated genes in the network were first clustered into different subgroups according to their located compartments. After that, the structural entropy of the gene in the subgroup was employed to measure its indispensability. Considering mutual exclusive property between driver genes in the modules, relative entropy was utilized to measure the degree of mutual exclusive between two mutated genes in terms of their variation frequency. We applied our method to three different cancers including lung, prostate, and breast cancer. The results show our method not only detect the well-known important drivers but also prioritiz the rare unknown driver genes. Besides, EntroRank can identify driver genes having mutual exclusive property. Compared with other existing methods, our method achieves a better performance for most of cancer types in terms of Precision, Recall, and Fscore.
Junrong Song, Wei Peng 0004
IEEE ACM Trans. Comput. Biol. Bioinform.2
2019 Predicting protein functions through non-negative matrix factorization regularized by protein-protein interaction network and gene functional information
abstract
Protein function prediction is necessary for understanding life and is valuable for application on drug design and health care. It is still a big challenge to predict protein function correctly by integrating multi-biological information. In this work, we propose a novel non-Negative Matrix Factorization (NMF) based method regularized by PPI network and GO similarity network, namely PONMF, for protein function prediction, which decomposes the known GO-protein association matrix into two low-rank matrixes for proteins and GO terms. In the course of factorization, PONMF incorporates the label matrix factorization term, an additional network regularization term and a GO similarity regularization term into the objective function. Finally, the potential protein functions are predicted by referring to the product of the two low-rank matrixes. PONMF not only successfully integrates diverse biological information to predict protein functions, but also naturally partitions proteins into different modules and infers functions from the proteins in the same modules. Our methods as well as the other two state-of-the-art methods (UBiRW and NMFGO) are applied to predict functions for protein of S. cerevisiae and H. sapiens. The prediction results show that PONMF outperforms the other two existing methods.
Wei Peng 0004, Wei Dai 0012, Jielin Du, Wei Lan 0001
BIBM1
2019 Identifying Human Essential Genes by Network Embedding Protein-Protein Interaction Network
Wei Dai 0012, Wei Peng 0004, Jiancheng Zhong, Yongjiang Li
ISBRA3
2019 Improving Identification of Essential Proteins by a Novel Ensemble Method
Wei Dai 0012, Xia Li 0004, Wei Peng 0004, Jurong Song, Jiancheng Zhong, Jianxin Wang 0001
ISBRA3
2019 A random walk-based method to identify driver genes by integrating the subcellular localization and variation frequency into bipartite graph
abstract
BACKGROUND: Cancer as a worldwide problem is driven by genomic alterations. With the advent of high-throughput sequencing technology, a huge amount of genomic data generates at every second which offer many valuable cancer information and meanwhile throw a big challenge to those investigators. As the major characteristic of cancer is heterogeneity and most of alterations are supposed to be useless passenger mutations that make no contribution to the cancer progress. Hence, how to dig out driver genes that have effect on a selective growth advantage in tumor cells from those tremendously and noisily data is still an urgent task. RESULTS: Considering previous network-based method ignoring some important biological properties of driver genes and the low reliability of gene interactive network, we proposed a random walk method named as Subdyquency that integrates the information of subcellular localization, variation frequency and its interaction with other dysregulated genes to improve the prediction accuracy of driver genes. We applied our model to three different cancers: lung, prostate and breast cancer. The results show our model can not only identify the well-known important driver genes but also prioritize the rare unknown driver genes. Besides, compared with other existing methods, our method can improve the precision, recall and fscore to a higher level for most of cancer types. CONCLUSIONS: The final results imply that driver genes are those prone to have higher variation frequency and impact more dysregulated genes in the common significant compartment. AVAILABILITY: The source code can be obtained at https://github.com/weiba/Subdyquency .
Junrong Song, Wei Peng 0004
BMC Bioinform.2
2017 Protein-protein interactions: detection, reliability assessment and applications
abstract
Protein-protein interactions (PPIs) participate in all important biological processes in living organisms, such as catalyzing metabolic reactions, DNA replication, DNA transcription, responding to stimuli and transporting molecules from one location to another. To reveal the function mechanisms in cells, it is important to identify PPIs that take place in the living organism. A large number of PPIs have been discovered by high-throughput experiments and computational methods. However, false-positive PPIs have been introduced too. Therefore, to obtain reliable PPIs, many computational methods have been proposed. Generally, these methods can be classified into two categories. One category includes the methods that are designed to determine new reliable PPIs. The other one is designed to assess the reliability of existing PPIs and filter out the unreliable ones. In this article, we review the two kinds of methods for detecting reliable PPIs, and then focus on evaluating the performance of some of these typical methods. Later on, we also enumerate several PPI network-based applications with taking a reliability assessment of the PPI data into consideration. Finally, we will discuss the challenges for obtaining reliable PPIs and future directions of the construction of reliable PPI networks. Our research will provide readers some guidance for choosing appropriate methods and features for obtaining reliable PPIs.
Xiaoqing Peng, Jianxin Wang 0001, Wei Peng 0004, Fang-Xiang Wu, Yi Pan 0001
Briefings Bioinform.3
2017 Predicting Protein Functions by Using Unbalanced Random Walk Algorithm on Three Biological Networks
abstract
With the gap between the sequence data and their functional annotations becomes increasing wider, many computational methods have been proposed to annotate functions for unknown proteins. However, designing effective methods to make good use of various biological resources is still a big challenge for researchers due to function diversity of proteins. In this work, we propose a new method named ThrRW, which takes several steps of random walking on three different biological networks: protein interaction network (PIN), domain co-occurrence network (DCN), and functional interrelationship network (FIN), respectively, so as to infer functional information from neighbors in the corresponding networks. With respect to the topological and structural differences of the three networks, the number of walking steps in the three networks will be different. In the course of working, the functional information will be transferred from one network to another according to the associations between the nodes in different networks. The results of experiment on S. cerevisiae data show that our method achieves better prediction performance not only than the methods that consider both PIN data and GO term similarities, but also than the methods using both PIN data and protein domain information, which verifies the effectiveness of our method on integrating multiple biological data sources.
Wei Peng 0004, Min Li 0007, Lusheng Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2016 Predicting microRNA-disease associations by walking on four biological networks
abstract
MicroRNA(miRNA) plays an important role in regulating the expression of target mRNAs. The deregulation of microRNAs appears to associate with various diseases. Recently, researchers focus on making use of various biological properties to identify the associations between microRNAs and diseases so as to provide helpful information for disease therapies. Accumulate evidences have shown that the inter- and intra-relationships of microRNAs, diseases, environment factors and genes contribute to correctly detect candidate microRNA-disease associations. However, there lack of methods that can comprehensively make use of the advantage of these relationships. In this work, we construct four separate biological networks, that are microRNA functional similarity network(MFN), disease semantic similarity network(DSN), environmental factor chemical structure similarity network(ESN) and gene-gene functional similarity network( GSN). After that, an unbalanced four random walking method, namely FourRW is implemented on the four networks, which not only can flexibly infer information from different levels of neighbors in the four networks, but also realizes the information transfer between different networks. The results of experiment show that our method achieves better prediction performance than the other state-of-the-art methods.
Wei Peng 0004, Wei Lan 0001, Jianxin Wang 0001, Yi Pan 0001
BIBM1
2016 Predicting MicroRNA-Disease Associations by Random Walking on Multiple Networks
Wei Peng 0004, Wei Lan 0001, Zeng Yu 0001, Jianxin Wang 0001, Yi Pan 0001
ISBRA1
2015 UDoNC: An Algorithm for Identifying Essential Proteins Based on Protein Domains and Protein-Protein Interaction Networks
abstract
Prediction of essential proteins which are crucial to an organism's survival is important for disease analysis and drug design, as well as the understanding of cellular life. The majority of prediction methods infer the possibility of proteins to be essential by using the network topology. However, these methods are limited to the completeness of available protein-protein interaction (PPI) data and depend on the network accuracy. To overcome these limitations, some computational methods have been proposed. However, seldom of them solve this problem by taking consideration of protein domains. In this work, we first analyze the correlation between the essentiality of proteins and their domain features based on data of 13 species. We find that the proteins containing more protein domain types which rarely occur in other proteins tend to be essential. Accordingly, we propose a new prediction method, named UDoNC, by combining the domain features of proteins with their topological properties in PPI network. In UDoNC, the essentiality of proteins is decided by the number and the frequency of their protein domain types, as well as the essentiality of their adjacent edges measured by edge clustering coefficient. The experimental results on S. cerevisiae data show that UDoNC outperforms other existing methods in terms of area under the curve (AUC). Additionally, UDoNC can also perform well in predicting essential proteins on data of E. coli.
Wei Peng 0004, Jianxin Wang 0001, Yingjiao Cheng, Fang-Xiang Wu, Yi Pan 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2015 Identification of Protein Complexes Using Weighted PageRank-Nibble Algorithm and Core-Attachment Structure
abstract
Protein complexes play a significant role in understanding the underlying mechanism of most cellular functions. Recently, many researchers have explored computational methods to identify protein complexes from protein-protein interaction (PPI) networks. One group of researchers focus on detecting local dense subgraphs which correspond to protein complexes by considering local neighbors. The drawback of this kind of approach is that the global information of the networks is ignored. Some methods such as Markov Clustering algorithm (MCL), PageRank-Nibble are proposed to find protein complexes based on random walk technique which can exploit the global structure of networks. However, these methods ignore the inherent core-attachment structure of protein complexes and treat adjacent node equally. In this paper, we design a weighted PageRank-Nibble algorithm which assigns each adjacent node with different probability, and propose a novel method named WPNCA to detect protein complex from PPI networks by using weighted PageRank-Nibble algorithm and core-attachment structure. Firstly, WPNCA partitions the PPI networks into multiple dense clusters by using weighted PageRank-Nibble algorithm. Then the cores of these clusters are detected and the rest of proteins in the clusters will be selected as attachments to form the final predicted protein complexes. The experiments on yeast data show that WPNCA outperforms the existing methods in terms of both accuracy and p-value. The software for WPNCA is available at "http://netlab.csu.edu.cn/bioinfomatics/weipeng/WPNCA/download.html".
Wei Peng 0004, Jianxin Wang 0001, Lusheng Wang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2013 A dividing-and-matching algorithm to detect conserved protein complexes via local network alignment
abstract
Local network alignment is an effective way to comparatively analyze a pair of protein-protein interaction (PPI) networks so as to identify the common subnetworks (conserved protein complexes) across species, which helps us better understand the structure, function and evolution of biological cells. In this work, we propose a new dividing-and-matching method named by DAMAlign to detect conserved protein complexes via local network alignment. DAMAlign firstly partitions one of PPI network into subnetworks and then these subnetworks are mapped to the other PPI network to find common connected components. In the course of finding common connected components, DAMAlign adopts a lenient criteria that is we locally extend a pair of homologous proteins if there exists at least one path of length not larger than 2 to connect one of node in the homologous protein pair in its corresponding network. We implement network alignment between S. cerevisiae and D. melanogaster. The experimental results show that DAMAlign outperforms other existing methods in recovering known protein complexes. Moreover, the conserved protein complexes that are detected by DAMAlign from different PPI networks are also functional similar in terms of their GO semantic similarity.
Wei Peng 0004, Jianxin Wang 0001, Fang-Xiang Wu
BIBM1
2013 Identifying essential proteins based on protein domains in protein-protein interaction networks
abstract
Prediction of essential proteins which are crucial to an organism survival is important for disease analysis and drug design, as well as the understanding of cellular life. The majority of prediction methods infer the possibility of proteins to be essential by using the network topology. However, these methods are limited to the complementation of available protein-protein interaction (PPI) data and depend on the network accuracy. To overcome these limitation, some computational methods have been proposed while seldom of them solve this problem by taking consideration of protein domains. In this work, we firstly analyze the correlation between the essentiality of proteins and their domain features based on data of 13 species. We find that the proteins containing more protein domain types which rarely occur in other proteins tend to be essential. Accordingly we propose a new prediction method, named UDoNC, by combining the domain features of proteins with their topological properties in PPI network. In UDoNC, the essentiality of proteins is decided by the number and the frequency of their protein domain types, as well as the essentiality of their adjacent edges measured by edge clustering coefficient. The experimental results on S. cerevisiae data show that UDoNC outperforms other existing methods in terms of area under the curve (AUC).
Jianxin Wang 0001, Wei Peng 0004, Yingjiao Chen, Yi Pan 0001
BIBM2