Lun Hu

dblp:163/5983 · DBLP profile ↗
← Back
94ranked-venue papers
16as first author
74since 2021 · last 2026
0000-0002-1591-8549ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 75 · 11 first-author · 60 since 2021Artificial intelligence and machine learning · 13 · 4 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MMC-Det: Structure-Preserving and Morphology-Aware Detection for Organoids in Bright-Field Microscopy
Xi Zhou 0007, Jun Zhang 0003, Le Tong, Lun Hu, Pengwei Hu 0001
ICIC (30)7
2026 From Self-supervised Pre-training to Confounder-Aware Refinement: A New Approach for miRNA-Drug Association Prediction
Runzhou Tang, Xi Zhou 0007, Jun Zhang 0003, Lun Hu, Pengwei Hu 0001
ICIC (30)5
2026 Spatial-spectral fusion enables drug repositioning by capturing indirect and long-range associations in biological networks
Lei Wang 0121, Runzhou Tang, Zhi-an Huang, Feng Tan 0002, Lun Hu, Zhu-Hong You, Pengwei Hu 0001
Bioinform.7
2026 DeShiftNet: a deformable-shifted cross-attention network for lightweight and robust organoid image segmentation
abstract
BACKGROUND: Organoid image segmentation is essential for quantitative analysis in disease modeling and drug screening, yet remains highly challenging due to substantial morphological variability and blurred boundaries in organoid images. Existing approaches often struggle to achieve a favorable balance between segmentation accuracy and computational efficiency. RESULTS: In this paper, DeShiftNet, a lightweight segmentation framework, is proposed to extract discriminative features with high accuracy while maintaining low computational overhead. The model incorporates a deformable-shifted encoding strategy that adaptively samples local structures. It also includes a cross-attention-guided decoder for selective multi-scale feature alignment. Furthermore, a deformable multi-scale contextual refinement module enhances boundary coherence and contextual consistency. Extensive experiments on the multi-type OrganoID dataset show that DeShiftNet achieves competitive performance compared with recent segmentation models, while maintaining only 1.78M parameters and 2.65 GFLOPs. Notably, DeShiftNet achieves a Dice score of 0.961 on the Lung subset. CONCLUSION: These results indicate its potential practical value for efficient organoid segmentation in high-throughput experimental workflows.
Le Tong, Tao Shu, Xinru Zhuang, Jingrui Bai, Lun Hu, Feng Tan 0002, Zhu-Hong You, Pengwei Hu 0001
BMC Bioinform.5
2026 Robust Data-Driven Control of Heterogeneous Multi-Agent Systems and Its Application in Autonomous Vehicle and Drone Collaboration
abstract
This paper addresses the output synchronization problem in heterogeneous leader-follower multi-agent systems with unknown follower dynamics. A data-driven control strategy is proposed, wherein a tracking controller is designed based on finite-horizon input-output data. The data is further employed to solve the output regulator equation, thereby facilitating synchronization of the followers’ outputs with that of the leader. To account for bounded noise interference, the matrix S-lemma is employed to reformulate the synchronization conditions into a tractable convex optimization problem. Furthermore, it is analytically shown that the synchronization error remains bounded and is explicitly constrained by the solution of the optimization problem. The proposed methodology is substantiated through comprehensive numerical simulations and experimental validations, demonstrating its effectiveness and robustness.
Xiufeng Zhang, Shuli Tan, Lun Hu
IEEE Trans Autom. Sci. Eng.5
2026 Fuzzy Mixture-of-Experts Aggregation for Organoid Identification With Multiscale State Space Features
abstract
Accurate and automated identification of organoids from bright-field images is essential for enabling high-throughput drug screening and precision medicine. Organoids, as 3-D in vitro cellular models, closely recapitulate the functional and structural characteristics of their tissue or organ of origin, presenting an unprecedented opportunity for biomedical research. However, the complexity of bright-field microscopy images, including heterogeneous backgrounds and diverse organoid morphologies, poses significant challenges for existing computational methods, often hindering robust feature extraction and high-throughput analysis. To address these issues at the intersection of computational vision and organoid biology, we propose FEMSSorg, a novel organoid recognition framework designed to adaptively aggregate multiscale scan-selected state space features through a fuzzy mixture-of-experts (FuzzyMoE) scoring mechanism. FEMSSorg introduces a fuzzy expert soft routing mechanism (fuzzy route), implemented via Fuzzy C-Means-based soft routing assignments, forming a new class of fuzzy MoE that leverages fuzzy expert clustering scores to dynamically integrate local (LocalSS) and global (GlobalSS) state space features. This approach enables effective balancing of global pixel dependencies and local texture information, thereby substantially reducing background interference and image noise in bright-field images and improving the accuracy of organoid identification. Furthermore, we incorporate a Dual Downsampling Adaptive Pooling Feature Fusion module, which combines original backbone features with parallel downsampled features and utilizes content-aware pooling for adaptive multilevel and multiscale feature fusion. Experimental results on multiclass organoid bright-field image datasets demonstrate that FEMSSorg achieves state-of-the-art performance in both organoid detection and morphological texture classification, highlighting its value as a robust computational tool for advancing real-time, high-throughput organoid research.
Pengwei Hu 0001, Thomas Herget, Feng Tan 0002, Jun Zhang 0003, Lun Hu, Zhu-Hong You, Xin Luo 0001
IEEE Trans. Fuzzy Syst.8
2026 LLM-DDI: Leveraging Large Language Models for Drug-Drug Interaction Prediction on Biomedical Knowledge Graph
abstract
Drug-drug interaction (DDI) refers to the interaction relationships between drugs. Discovering new DDIs is crucial for advancing drug development and enhancing clinical treatments. Given the significant progress achieved through graph neural networks (GNNs), network-based models have become a prevalent approach for tackling this challenge. However, current network-based approaches are incapable of seamlessly integrating a wide range of information. Motivated by this discovery, we propose a novel model, namely LLM-DDI, which aims to comprehensively tackle DDI prediction tasks by integrating various information of molecules in the BKG. LLM-DDI initially incorporates the generative pre-trained transformer (GPT) model to generate embeddings for each molecule within the biomedical knowledge graph (BKG). These embeddings encompass diverse types of information pertaining to each molecule. Subsequently, LLM-DDI utilizes a message-passing GNN framework to enhance the learning of molecular representations with the embeddings derived from GPT as input. LLM-DDI governs the propagation of information within the BKG by semantic relationships. These semantic relationships determine how information flows and is exchanged between different entities in the BKG. Finally, LLM-DDI leverages the learned drug representations to predict potential DDIs. Experiments show the effectiveness of LLM-DDI, as it achieves the best performance on two real-world datasets, providing valuable guidance for drug development and clinical treatment.
Dongxu Li 0002, Yue Yang 0035, Ziwen Cui, Hengchuang Yin, Pengwei Hu 0001, Lun Hu
IEEE J. Biomed. Health Informatics6
2026 Multi-View Contrastive Learning for Drug-Drug Interaction Event Prediction
abstract
Drug-drug interactions (DDIs) represent a critical challenge in pharmacology, often leading to adverse effects and compromised therapeutic efficacy. Accurate prediction of DDI events, which involve not only identifying interacting drug pairs but also characterizing the specific nature and context of their interactions, is essential for drug safety and personalized medicine. In this study, we propose a novel Multi-view Contrastive Learning framework, namely MCL-DDI, for DDI Event Prediction by leveraging multi-view representations of drugs to enhance predictive performance. MCL-DDI integrates molecular structures and network features, capturing complementary information about drug properties and interactions. By employing contrastive learning, we align and unify drug representations across these diverse views, enabling the framework to distinguish complex interaction patterns. Extensive experiments on benchmark datasets demonstrate that MCL-DDI outperforms state-of-the-art methods in terms of predictive accuracy. Furthermore, case studies highlight the model's ability to identify clinically relevant DDIs, offering practical insights for drug development and risk assessment. Our work establishes a robust and accurate paradigm for DDI event prediction, paving the way for safer and more effective pharmacological interventions.
Dongxu Li 0002, Feifan Zhao, Yue Yang 0035, Ziwen Cui, Pengwei Hu 0001, Lun Hu
IEEE J. Biomed. Health Informatics6
2026 Graph-Based Prediction of miRNA-Drug Associations With Multisource Information and Metapath Enhancement Matrices
abstract
Recent studies have demonstrated that miRNA expression dysregulation is closely related to the occurrence of various diseases; thus, miRNA-based drug development strategies have received increasing research interest. Most existing computational methods focus on the attribute information of individual nodes and are limited to the direct associations between nodes, thereby ignoring the complex associations inherent in the network. This limitation may lead to the loss of key potential information, which impacts the prediction accuracy. To address these issues, we propose a multisource information fusion and metapath enhancement matrix based graph autoencoder (MSMP-GAE) to predict the potential associations between miRNAs and drugs. The proposed MSMP-GAE model comprises a metapath instance extraction module, a metapath feature enhanced encoder module, a weighted feature fusion module, and a graph autoencoder. First, we construct an miRNA-drug heterogeneous network using experimentally validated miRNA-drug interactions and integrate various miRNA and drug features into an initial feature matrix to comprehensively represent their intrinsic property information. Then, we extract metapath instances from the interaction network, generate multiple metapath enhancement matrices, and fuse them with the initial feature matrix to generate high-quality node feature embeddings. Finally, we employ the graph autoencoder for fivefold cross-validation on a public dataset and test it on an independent test set. Experimental results demonstrate that the proposed MSMP-GAE model obtained an area under the curve (AUC) and AUPR values of 98.61% and 98.23%, respectively, which is considerably better than the several state-of-the-art methods. This highlights the importance of the higher-order complex associations between nodes in the miRNA-drug association (MDA) prediction task and provides a new method and approach to advance MDA prediction.
Ming-Yang Wu, Pengwei Hu 0001, Zhu-Hong You, Jun Zhang 0003, Lun Hu, Xin Luo 0001
IEEE J. Biomed. Health Informatics5
2025 Improving Cancer Gene Identification via Mixture-of-Experts-Based Graph Representation Learning
abstract
Accurately identifying cancer driver genes is crucial for understanding tumorigenesis and advancing precision oncology. However, integrating multi-omics data within complex biological networks remains challenging, particularly when it comes to capturing diverse structural information and leveraging the distinct signals from different omics modalities. While graphbased methods have demonstrated high accuracy in cancer gene identification, they might overlook the heterogeneity between omics features. To address this limitation, we propose CGI-MoE, a Mixture-of-Experts-inspired graph representation learning framework that incorporates omics-feature-specific expert modules based on graph transformers, together with adaptive gating and subgraph aggregation mechanisms. CGI-MoE extracts both local and global structural encodings for each node by sampling multiple subgraphs, enabling the model to capture comprehensive and robust network features. Each expert module focuses on a specific omics modality, and their outputs are fused by a lightweight gating network that dynamically weighs their contributions. When evaluated on both homogeneous and heterogeneous benchmark datasets, CGI-MoE achieves state-of-the-art performance in terms of accuracy, AUC, and AUPR, consistently surpassing existing methods. Ablation studies further highlight the critical roles of each component in achieving robust performance. Using the trained models, CGI-MoE predicted 46 novel cancer gene candidates from all unlabeled genes, demonstrating its potential for novel discovery and for deepening our understanding of cancer development. The code is available at https://github.com/moomight/CGI-MoE.
Ying Chang, Yue Yang 0035, Dongxu Li 0002, Ziwen Cui, Hengchuang Yin, Pengwei Hu 0001, Lun Hu
BIBM7
2025 GA-RLImpute: A Genetic Algorithm with Reinforcement Learning for Imputation of Medical Tabular Data
abstract
Existing imputation methods typically apply a uniform, one-size-fits-all strategy, which is ill-suited for medical data where features exhibit highly heterogeneous marginal distributions, leading to biased results and degraded downstream performance. To address this, we propose GA-RLImpute, a framework integrating Genetic Algorithms (GA) and Reinforcement Learning (RL). Its GA module optimizes featurespecific imputation strategies, identifying the best method for each column via a multi-objective fitness function that balances reconstruction accuracy with task performance. Concurrently, the RL module enhances adaptability by dynamically tuning key GA hyperparameters, such as mutation and crossover rates. Comprehensive experiments show GA-RLImpute consistently outperforms baselines across classification, regression, clustering and data quality evaluation tasks, outperforming the next-best method by 6.34% in multi-class classification. GA-RLImpute is thus a robust and flexible solution for complex missing data scenarios in healthcare.
Xiwen Yang, Deng Xun, Lu You, Lun Hu
BIBM5
2025 MHGN: A Model for Drug Discovery Based on Molecular Hierarchical Graph Network
abstract
Understanding molecular structures and their hierarchical relationships with biological entities is crucial for precise drug modeling and discovery. In this study, we proposed the Molecular Hierarchical Graph Network (MHGN), a unified framework that connects the drug molecular graph structure constructed based on atom and bond information with the drug heterogeneous interaction graph. MHGN first encodes detailed chemical features-including atom type, valence, chirality, bond type, aromaticity, and spatial topology-into expressive molecular graph representations. We then constructed a heterogeneous interaction graph for drug discovery to enable bidirectional information exchange between chemical information within drug molecules and non-drug entities, thereby linking atomic-level knowledge with biological information between target or miRNAs. Beyond molecular graph representations, MHGN also leverages the latest molecular visual modality, integrating 3D molecular visual information through a late-fusion strategy to enhance representation completeness. Experimental results on drug-target interaction (DTI) and drug-miRNA interaction (DMI) prediction tasks demonstrate that MHGN outperforms recent state-of-the-art baselines across ROC-AUC, AUPR, and F1-score metrics, highlighting the superiority of MHGN in enabling atom-level information exchange. Furthermore, MHGN generalizes effectively to molecular property regression, achieving an RMSE of 0.6633 and an$R^{2}$of 0.8735 on the ESOL solubility dataset. In summary, MHGN provides a powerful graph-based model for data-driven drug discovery, enabling efficient information exchange between drugs and target or miRNAs at the atomic level.
Zimai Zhang, Yujie Qi, Lun Hu
BIBM5
2025 Dual-Channel MiRNA Drug Resistance Prediction Model Based on Multimodal Feature Alignment
Runzhou Tang, Zimai Zhang, Jun Zhang 0003, Lun Hu, Xi Zhou 0007, Pengwei Hu 0001
ICIC (26)4
2025 Adversarial Domain Adaptation for Accurately Predicting the Associations between Herbal Compounds and Target Proteins
abstract
Traditional Chinese medicine (TCM), as a treasure of traditional Chinese medical heritage, has developed a unique therapeutic system through millennia of practice. Its multi-compound, multi-target mechanism demonstrates distinct advantages in disease treatment. However, under the modern medical framework, the modernization of TCM research faces dual challenges: on one hand, the complex interaction network between herbal compounds and target proteins has not yet been systematically elucidated, and traditional methods based on in vitro experiments or statistical correlations suffer from long experimental cycles and high failure rates; on the other hand, compared to chemical drugs with well-established large-scale annotated databases, the association data between herbal compounds and targets exhibit significant sparsity, making it difficult for existing methods to effectively model cross-domain knowledge transfer due to mismatched data dimensions. To address this scientific challenge, this study proposes a cross-domain association prediction framework, DGCPI, which integrates graph neural networks and adversarial learning. To begin with, DGCPI constructs two heterogeneous graph networks for modelling chemical drug-target (source domain) and herbal compound-target (taget domain) associations respectively. It then applies a graph convolutional network to capture source-domain-specific topological features in the chemical drug-target graph. After that, an adversarial discriminator is introduced to dynamically align the embedding distributions between the source and target domains, overcoming the dependence of traditional methods on the assumption of homomorphic data. Experimental results show that only with a small number of TCM samples involved in the training, the model significantly outperforms baseline models on two TCM datasets.
Yantong Qiao, Lun Hu, Jun Zhang 0003, Pengwei Hu 0001, Xin Luo 0001
SMC2
2025 Identifying novel therapeutic targets of natural compounds in traditional Chinese medicine herbs with hypergraph representation learning
abstract
Traditional Chinese medicine (TCM), with its roots in centuries of clinical practice, has established itself as an effective therapeutic system that involves a diverse range of herbal plants. Despite its proven efficacy, the intricate relationships between herbal multi-component preparations and multi-target therapies present challenges for systematic study, thereby limiting its broader application in managing chronic diseases. In this work, we aim to identify novel therapeutic targets of natural compounds found in TCM herbs by leveraging advanced hypergraph representation learning techniques. Following the multi-component, multi-target pharmacological mechanisms, we first construct two hypergraphs to represent herb-compound and disease-target interactions, respectively. The connection between these hypergraphs is established through compound-target associations. A convolutional operator is then employed to capture the high-order correlations between compound (or target) nodes and herb (or disease) hyperedges within each hypergraph. Furthermore, we incorporate the PageRank algorithm and a multi-head attention mechanism to enhance the representation capabilities of node embeddings. By integrating these methods, our model is able to accurately identify novel therapeutic targets of natural compounds in TCM herbs in an end-to-end manner. Extensive experiments conducted on three benchmark datasets demonstrate the superior performance of our model when compared with several state-of-the-art approaches. Furthermore, case studies on two natural compounds, coumarin and progesterone, reveal that 7 and 8 out of the Top-10 identified targets, respectively, have been validated through literature review. These results highlight the effectiveness of our model in discovering new therapeutic targets for natural compounds in TCM.
Yantong Qiao, Lun Hu, Jun Zhang 0003, Pengwei Hu 0001, Xin Luo 0001
Briefings Bioinform.2
2025 Regulation-aware graph learning for drug repositioning over heterogeneous biological network
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Dongxu Li 0002, Pengwei Hu 0001, Zhu-Hong You, Xin Luo 0001, Lun Hu
Inf. Sci.9
2025 A bijective inference network for interpretable identification of RNA N6-methyladenosine modification sites
Yue Yang 0035, Dongxu Li 0002, Xiao-Rui Su 0001, Zhi Zeng 0001, Pengwei Hu 0001, Lun Hu
Pattern Recognit.7
2025 DeepHIV: A Sequence-Based Deep Learning Model for Predicting HIV-1 Protease Cleavage Sites
abstract
Human immunodeficiency virus type 1 (HIV-1) is one of the main causative agents of acquired immunodeficiency syndrome (AIDS), and effectively identifying HIV-1 protease cleavage sites (PCSs) is of great importance for the design of new anti-AIDS inhibitors. Computational prediction of HIV-1 PCSs can be used to discover new cleavable substrates, and further facilitates the understanding of substrate specificity. A novel deep learning model, namely DeepHIV, is designed to predict HIV-1 PCSs from substrate sequence information alone. In particular, DeepHIV first applies a convolutional neural network combined with an attention mechanism to capture the rich contextual information of position-specific amino acids in the substrate sequences, thus improving the quality of features learned for substrates. Considering the imbalance observed between cleavable and uncleavable substrates, a biased support vector machine is adopted as the classifier of DeepHIV to complete the prediction task. Experimental results demonstrate that DeepHIV outperforms several state-of-the-art prediction methods across all benchmark datasets and evaluation metrics. Hence, DeepHIV is an accurate and robust tool to predict HIV-1 PCSs. Moreover, the promising predictive performance of DeepHIV also reveals that our deep learning model is capable of fully leveraging the sequence information to effectively learn the latent features of substrates.
Dongxu Li 0002, Zhenfeng Li, Bo-Wei Zhao, Xiao-Rui Su 0001, Lun Hu
IEEE Trans. Comput. Biol. Bioinform.6
2025 Knowledge Graph Neural Network With Spatial-Aware Capsule for Drug-Drug Interaction Prediction
abstract
Uncovering novel drug-drug interactions (DDIs) plays a pivotal role in advancing drug development and improving clinical treatment. The outstanding effectiveness of graph neural networks (GNNs) has garnered significant interest in the field of DDI prediction. Consequently, there has been a notable surge in the development of network-based computational approaches for predicting DDIs. However, current approaches face limitations in capturing the spatial relationships between neighboring nodes and their higher-level features during the aggregation of neighbor representations. To address this issue, this study introduces a novel model, KGCNN, designed to comprehensively tackle DDI prediction tasks by considering spatial relationships between molecules within the biomedical knowledge graph (BKG). KGCNN is built upon a message-passing GNN framework, consisting of propagation and aggregation. In the context of the BKG, KGCNN governs the propagation of information based on semantic relationships, which determine the flow and exchange of information between different molecules. In contrast to traditional linear aggregators, KGCNN introduces a spatial-aware capsule aggregator, which effectively captures the spatial relationships among neighboring molecules and their higher-level features within the graph structure. The ultimate goal is to leverage these learned drug representations to predict potential DDIs. To evaluate the effectiveness of KGCNN, it undergoes testing on two datasets. Extensive experimental results demonstrate its superiority in DDI predictions and quantified performance.
Xiao-Rui Su 0001, Bo-Wei Zhao, Jun Zhang 0003, Pengwei Hu 0001, Zhu-Hong You, Lun Hu
IEEE J. Biomed. Health Informatics7
2025 Toward Multilabel Classification for Multiple Disease Prediction Using Gut Microbiota Profiles
abstract
Advancements in high-throughput technologies have yielded large-scale human gut microbiota profiles, sparking considerable interest in exploring the relationship between the gut microbiome and complex human diseases. Through extracting and integrating knowledge from complex microbiome data, existing machine learning (ML)-based studies have demonstrated their effectiveness in the precise identification of high-risk individuals. However, these approaches struggle to address the heterogeneity and sparsity of microbial features and explore the intrinsic relatedness among human diseases. In this work, we reframe human gut microbiome-based disease detection as a multilabel classification (MLC) problem and integrate a range of innovative techniques within the proposed MLC framework, aptly named GutMLC. Specifically, the entity semantic similarity as priori knowledge is incorporated into multilabel feature selection and loss functions by capturing the shared attributes and inherent associations among diseases and microbes. To tackle the issue of label imbalance, both within and between labels, we adapt the focal loss (FL) function for MLC using debiased inverse weighting. Extensive experiment results consistently demonstrate the competitive performance of GutMLC in comparison with commonly used MLC and single-label classification (SLC) algorithms. This work seeks to unlock the potential of gut microbiota as robust biomarkers for multiple disease prediction.
Zhi-an Huang, Pengwei Hu 0001, Lun Hu, Zhu-Hong You, Kay Chen Tan
IEEE Trans. Neural Networks Learn. Syst.3
2025 Link-Based Attributed Graph Clustering via Approximate Generative Bayesian Learning
abstract
To understand the mechanisms of complex systems, attributed graphs (AGs) are recognized as a valuable model by their capability of describing nontrivial topological structures and rich node contents, and their emergence raises new challenges on the task of graph clustering. Although a variety of computational algorithms have been proposed to perform accurate clustering analysis on AGs, most of them are incapable of inferring the cluster labels of nodes through links, thus falling short of explaining node behaviors on how to formulate overlapping clusters. Moreover, the vast amount of links considerably decreases the computation efficiency if they are explicitly taken into account for AG clustering. To overcome this problem, we present a novel variational Bayesian learning model, which avoids generating a complete AG by only simulating the generative process of its skeleton with the prior knowledge on the cluster labels of links. When addressing the inference problem, we develop an efficient algorithm, namely, LCAAG, for determining the optimal cluster labels of nodes by estimating local community structures of links. The convergence of LCAAG has been proved theoretically. Compared with several state-of-the-art algorithms, LCAAG has demonstrated its promising performance in terms of both accuracy and scalability on five different scaled benchmark datasets. The source code and datasets are available at https://github.com/shallowdreamoon/LCAAG.git.
Yue Yang 0035, Lun Hu, Dongxu Li 0002, Pengwei Hu 0001, Xin Luo 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2025 FMvPCI: A Multiview Fusion Neural Network for Identifying Protein Complex via Fuzzy Clustering
abstract
Protein complexes play a crucial role in regulating various biological processes that govern cell activities. Numerous computational algorithms have been proposed to identify protein complexes from protein-protein interaction (PPI) networks. However, many of these algorithms face limitations in effectively leveraging multiview biological information of proteins, restricting their ability to capture the intricate characteristics of protein complexes in PPI networks. While deep learning-based algorithms have significantly advanced the identification of protein complexes, they often integrate graph representation learning techniques into traditional clustering algorithms without explicitly capturing the dependency between protein embeddings and resulting complexes. To address these issues, we present a multiview fusion neural network, named FMvPCI, for protein complex identification via fuzzy clustering. In FMvPCI, we introduce a novel multiview graph convolution encoder to effectively manipulate and fuse the biological information of proteins from different perspectives. Subsequently, the optimization of FMvPCI incorporates our expectations about protein complexes through the concept of fuzzy clustering. This approach unifies the embeddings of proteins and their cluster memberships within a coherent framework. Leveraging a heuristic search strategy, FMvPCI can discover overlapping protein complexes based on the cluster memberships of proteins. A series of experiments on five different PPI networks collected from two species have been conducted to evaluate the performance of FMvPCI by comparing it with state-of-the-art identification algorithms, and the results demonstrate the superior performance of FMvPCI by significantly improving the identification accuracy for protein complexes.
Yue Yang 0035, Lun Hu, Dongxu Li 0002, Pengwei Hu 0001, Xin Luo 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2024 Dual-Channel Learning Framework for Drug-Drug Interaction Prediction via Relation-Aware Heterogeneous Graph Transformer
abstract
Identifying novel drug-drug interactions (DDIs) is a crucial task in pharmacology, as the interference between pharmacological substances can pose serious medical risks. In recent years, several network-based techniques have emerged for predicting DDIs. However, they primarily focus on local structures within DDI-related networks, often overlooking the significance of indirect connections between pairwise drug nodes from a global perspective. Additionally, effectively handling heterogeneous information present in both biomedical knowledge graphs and drug molecular graphs remains a challenge for improved performance of DDI prediction. To address these limitations, we propose a Transformer-based relatIon-aware Graph rEpresentation leaRning framework (TIGER) for DDI prediction. TIGER leverages the Transformer architecture to effectively exploit the structure of heterogeneous graph, which allows it direct learning of long dependencies and high-order structures. Furthermore, TIGER incorporates a relation-aware self-attention mechanism, capturing a diverse range of semantic relations that exist between pairs of nodes in heterogeneous graph. In addition to these advancements, TIGER enhances predictive accuracy by modeling DDI prediction task using a dual-channel network, where drug molecular graph and biomedical knowledge graph are fed into two respective channels. By incorporating embeddings obtained at graph and node levels, TIGER can benefit from structural properties of drugs as well as rich contextual information provided by biomedical knowledge graph. Extensive experiments conducted on three real-world datasets demonstrate the effectiveness of TIGER in DDI prediction. Furthermore, case studies highlight its ability to provide a deeper understanding of underlying mechanisms of DDIs.
Xiao-Rui Su 0001, Pengwei Hu 0001, Zhu-Hong You, Philip S. Yu, Lun Hu
AAAI5
2024 MSSOrg: a multi-scale SSM-based model for organoid location and classification
abstract
Organoids, 3D in vitro cultures replicating their tissue or organ of origin, hold great potential for disease research, drug testing, and regenerative medicine. However, current analysis methods primarily rely on staining techniques, which limit high-throughput analysis and require manual adjustments. Automated and non-invasive methods remain underexplored, and most rely on small-scale datasets that fail to capture the complexity of organoids, which vary significantly in shape and size and have more intricate backgrounds than cells. To address these limitations, we developed MSSOrg, a multiscale framework for organoid localization and classification, integrating Local and Global State Space Models (SSMs) to balance pixel dependencies and texture features. MSSOrg mitigates background interference and image noise in bright-field images, achieving superior performance in both organoid detection and morphological classification. This innovation enables real-time, large-scale organoid analysis, providing a powerful tool for drug screening and disease research, ultimately accelerating progress in these fields.
Zhibin Li 0006, Mei Gao, Shaolin Liang, Zhipeng Shang, Yafang Wei, Lun Hu, Pengwei Hu 0001
BIBM9
2024 A Multi-view Nested Contrastive Learning Framework for Predicting Drug-Drug Interaction Events
abstract
Exploring drug-drug interactions (DDIs) is crucial for avoiding unknown physicochemical incompatibilities between coadministered drugs. While most studies concentrate on detecting the presence or absence of DDIs, they often overlook the diversity of DDI event types that can significantly enhance drug research and guide scientific drug use. To address this limitation, we propose MNCLDDI, a multi-view nested contrastive learning model designed for the precise prediction of DDI events. MN-CLDDI begins by employing a relational graph convolutional network to capture the various explicit relationships between drugs within a multi-relational DDI graph. This is followed by a transformer framework combined with a convolutional neural network (CNN) to learn the biological features of drugs from their Smiles information. The model then integrates these two feature types into a novel multi-view nested contrastive learning framework, thereby improving the expressiveness of drug embeddings from multiple biological perspectives. Experimental results on two real-world datasets demonstrate that MNCLDDI outperforms state-of-the-art models in predicting DDI events. Moreover, our case studies reveal that considering the multi-view features of drugs simultaneously enables MNCLDDI to predict DDI events with greater accuracy and from a more comprehensive perspective, offering valuable insights into the study of DDI events.
Dongxu Li 0002, Yue Yang 0035, Pengwei Hu 0001, Lun Hu
BIBM5
2024 DNMDA: Deep Non-negative Matrix Factorization with Multi-level Integration for MiRNA-Drug Interaction Prediction
abstract
Numerous studies have demonstrated that the interaction between miRNAs and drugs plays a pivotal role in regulating gene expression and cellular function. Therefore, predicting these interactions is crucial for the development of novel drugs and personalized therapies. Existing methods for predicting miRNA-drug interactions often fail to leverage the full spectrum of molecular and biological features and overlook complex high-dimensional patterns. Deep non-negative matrix factorization (DNMF) addresses these limitations by extracting higher-level representations, thereby enhancing prediction accuracy and robustness. Building on this, this paper proposes a model called DNMDA. In this model, we integrate multiple similarity networks for both miRNAs and drugs and then extract their features through three key modules. Moreover, autoencoders are used to combine various feature sets, allowing for the capture of complementary information and enhancing the model’s capacity for making precise and reliable predictions. The resulting features are consolidated into a unified feature vector for each miRNA-drug pair. Ultimately, these feature vectors and their associated labels are provided to the classifier for training. To verify the predictions, a five-fold cross-validation was conducted. The five-fold cross-validation demonstrated a clear advantage in DNMDA’s metrics, underscoring its reliability in predicting potential miRNA-drug interactions. This claim is further supported by the predictive results section in the paper, offering concrete evidence of DNMDA’s efficacy in this field.
Yujie Qi, Zhu-Hong You, Zimai Zhang, Lun Hu, Xi Zhou 0007, Pengwei Hu 0001
BIBM6
2024 Predicting the Associations between Herbal Compounds and Target Proteins via Hypergraph Convolutional Network
abstract
Traditional Chinese medicine (TCM) has a long history of herbal treatments for diseases, but our understanding to the interactions between herbal compounds and target proteins remains considerably incomplete due to the complexity of multi-compound, multi-target mechanisms. Most of existing prediction models fall short of considering such mechanisms, limiting their applicability in discovering novel compound-target interactions(CTIs). To address this problem, we propose DHGCTI, a novel Hypergraph Convolutional Network for improving performance on the task of CTI prediction. In the context of hypergraph, herbs and diseases are regarded as the hyperedges of herbal compounds and target proteins respectively. An end-to-end prediction model is then specifically designed to identify CTIs by using a hypergraph convolutional network. Extensive testing shows experimental results demonstrate that DHGCTI consistently outperforms baseline models, highlighting its superior performance. Moreover, we conducted case studies on acacetin and luteolin, and 6 and 7 of the top ten scoring targets were validated by literature, respectively. This highlights the advantages of our model in exploring new targets of natural compounds in traditional Chinese medicine.
Yantong Qiao, Yafang Wei, Pengwei Hu 0001, Lun Hu
BIBM5
2024 Knowledge-guided Protein Complex Identification with Fuzzy-based Graph Representation Learning
abstract
Protein complexes are essential in regulating various cellular processes. A number of computational algorithms have been developed to identify protein complexes from protein-protein interaction (PPI) networks, but they are limited in their ability to effectively leverage diverse biological knowledge of proteins. Additionally, while deep learning-based algorithms perform well in identifying protein complexes, they fail to explicitly capture the dependency between protein embeddings and resulting complexes. To address these challenges, this paper proposes a knowledge-guided protein complex identification algorithm with fuzzy-based graph representation learning, named KPCI-FGRL. In particular, a fuzzy-based graph representation learning framework is developed by KPCI-FGRL to manipulate and fuse network structure with multi-view biological knowledge of proteins. During the training phase of KPCI-FGRL, besides employing self-supervised loss to improve the cohesion of the complexes, we also specifically incorporate the expectation about protein complexes based on fuzzy clustering concept, and thus the dependency between protein embeddings and complexes can be coupled. Furthermore, KPCI-FGRL is capable of achieving the identification of overlapping protein complexes through a heuristic search strategy upon fuzzy memberships of proteins. Extensive experimental results on four different PPI networks collected from two species demonstrate that KPCI-FGRL significantly outperforms several state-of-the-art protein complex identification algorithms.
Yue Yang 0035, Dongxu Li 0002, Pengwei Hu 0001, Lun Hu
BIBM5
2024 Dual-channel hypergraph convolutional network for predicting herb-disease associations
abstract
Herbs applicability in disease treatment has been verified through experiences over thousands of years. The understanding of herb-disease associations (HDAs) is yet far from complete due to the complicated mechanism inherent in multi-target and multi-component (MTMC) botanical therapeutics. Most of the existing prediction models fail to incorporate the MTMC mechanism. To overcome this problem, we propose a novel dual-channel hypergraph convolutional network, namely HGHDA, for HDA prediction. Technically, HGHDA first adopts an autoencoder to project components and target protein onto a low-dimensional latent space so as to obtain their embeddings by preserving similarity characteristics in their original feature spaces. To model the high-order relations between herbs and their components, we design a channel in HGHDA to encode a hypergraph that describes the high-order patterns of herb-component relations via hypergraph convolution. The other channel in HGHDA is also established in the same way to model the high-order relations between diseases and target proteins. The embeddings of drugs and diseases are then aggregated through our dual-channel network to obtain the prediction results with a scoring function. To evaluate the performance of HGHDA, a series of extensive experiments have been conducted on two benchmark datasets, and the results demonstrate the superiority of HGHDA over the state-of-the-art algorithms proposed for HDA prediction. Besides, our case study on Chuan Xiong and Astragalus membranaceus is a strong indicator to verify the effectiveness of HGHDA, as seven and eight out of the top 10 diseases predicted by HGHDA for Chuan-Xiong and Astragalus-membranaceus, respectively, have been reported in literature.
Lun Hu, Menglong Zhang, Pengwei Hu 0001, Jun Zhang 0003, Xueying Lu, Xiangrui Jiang, Yupeng Ma
Briefings Bioinform.1
2024 Fuzzy-Based Deep Attributed Graph Clustering
abstract
Attributed graph (AG) clustering is a fundamental, yet challenging, task for studying underlying network structures. Recently, a variety of graph representation learning models has been proposed to effectively infer the node embeddings, which are then incorporated into conventional clustering techniques to identify meaningful clusters. While these models tend to preserve node proximities, which reflect the similarity between nodes in both structural and attribute dimensions, for representation learning, they generally overlook the crucial dependencies between node embeddings and the resulting clusters. To overcome this problem, we propose a novel fuzzy-based deep AG clustering model, namely FDAGC, which is capable of achieving the task in a purely unsupervised and end-to-end manner without additionally incorporating conventional clustering techniques. In particular, FDAGC first encodes network structures and node attributes into a compact representation with graph convolution. A reconstruction error is then estimated to minimize the information loss during network message-passing. Besides, we utilize a self-monitoring training strategy to optimize node embeddings, thus improving the cluster cohesion by guiding them toward cluster centers. In the training phase, our expectations about resulting clusters are explicitly incorporated into the optimization of FDAGC via the concept of fuzzy clustering, thus leading to more accurate clustering by coupling the dependency between graph representation learning and AG clustering. Extensive experiments have demonstrated the superior performance of FDAGC in terms of several evaluation metrics, such as accuracy, normalized mutual information, F1-score and adjusted rand index, on six real-world AGs with different scales.
Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Pengwei Hu 0001, Jun Zhang 0003, Lun Hu
IEEE Trans. Fuzzy Syst.7
2024 Discovering Consensus Regions for Interpretable Identification of RNA N6-Methyladenosine Modification Sites via Graph Contrastive Clustering
abstract
As a pivotal post-transcriptional modification of RNA, N6-methyladenosine (m6A) has a substantial influence on gene expression modulation and cellular fate determination. Although a variety of computational models have been developed to accurately identify potential m6A modification sites, few of them are capable of interpreting the identification process with insights gained from consensus knowledge. To overcome this problem, we propose a deep learning model, namely M6A-DCR, by discovering consensus regions for interpretable identification of m6A modification sites. In particular, M6A-DCR first constructs an instance graph for each RNA sequence by integrating specific positions and types of nucleotides. The discovery of consensus regions is then formulated as a graph clustering problem in light of aggregating all instance graphs. After that, M6A-DCR adopts a motif-aware graph reconstruction optimization process to learn high-quality embeddings of input RNA sequences, thus achieving the identification of m6A modification sites in an end-to-end manner. Experimental results demonstrate the superior performance of M6A-DCR by comparing it with several state-of-the-art identification models. The consideration of consensus regions empowers our model to make interpretable predictions at the motif level. The analysis of cross validation through different species and tissues further verifies the consistency between the identification results of M6A-DCR and the evolutionary relationships among species.
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Xi Zhou 0007, Lun Hu
IEEE J. Biomed. Health Informatics7
2024 Motif-Aware miRNA-Disease Association Prediction via Hierarchical Attention Network
abstract
As post-transcriptional regulators of gene expression, micro-ribonucleic acids (miRNAs) are regarded as potential biomarkers for a variety of diseases. Hence, the prediction of miRNA-disease associations (MDAs) is of great significance for an in-depth understanding of disease pathogenesis and progression. Existing prediction models are mainly concentrated on incorporating different sources of biological information to perform the MDA prediction task while failing to consider the fully potential utility of MDA network information at the motif-level. To overcome this problem, we propose a novel motif-aware MDA prediction model, namely MotifMDA, by fusing a variety of high- and low-order structural information. In particular, we first design several motifs of interest considering their ability to characterize how miRNAs are associated with diseases through different network structural patterns. Then, MotifMDA adopts a two-layer hierarchical attention to identify novel MDAs. Specifically, the first attention layer learns high-order motif preferences based on their occurrences in the given MDA network, while the second one learns the final embeddings of miRNAs and diseases through coupling high- and low-order preferences. Experimental results on two benchmark datasets have demonstrated the superior performance of MotifMDA over several state-of-the-art prediction models. This strongly indicates that accurate MDA prediction can be achieved by relying solely on MDA network information. Furthermore, our case studies indicate that the incorporation of motif-level structure information allows MotifMDA to discover novel MDAs from different perspectives.
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Zhu-Hong You, Lun Hu
IEEE J. Biomed. Health Informatics9
2023 CAMPEOD: A Cross Attention-Based Multi-Scale Patch Embedding Organoid Detection Model
abstract
In medical research, organoids, which exhibit structural and functional resemblances to authentic organs, offer a substantial avenue for delving into aspects encompassing physiology, pathophysiology, diseases, and pharmaceutical screening. Critical insights into drug responsiveness are often gleaned from these organoids' dimensional and configurational disparities. However, conventional detection methodologies reliant upon fluorescent labeling engender potential hazards to organoid integrity, thereby impinging upon their intrinsic dynamic attributes. Traditional bounding-box detection methodologies fall short in encapsulating intricate morphological particulars, and certain deep-learning approaches grapple with the intricate task of capturing multi-scale data, particularly when tasked with discerning organoid structures characterized by marked shape and size heterogeneities. In a bid to surmount these constraints, our study introduces CAMPEOD, an innovative framework that synergistically amalgamates multi-scale attributes derived from organoid specimens, employing cross-attention mechanisms. This novel approach effectively obviates superfluous background interference and image noise, thereby endowing an automated, finely-tuned dissection of organoid samples. Significantly, this segmentation process ensures congruence with authentic organoid quantities and morphological characteristics. By facilitating comprehensive scrutiny of microscopy images of organoid samples on a large scale, CAMPEOD assumes considerable implications for the realm of pharmaceutical screening and ailment emulation.
Lun Hu, Zhu-Hong You, Pengwei Hu 0001
BIBM2
2023 Making the Implicit Explicit: Depression Detection in Web across Posted Texts and Images
abstract
The utilization of web social media for depression detection has been proven effective in recent years since the multimedia signal on web can reflect users’ emotions, feelings, and personality traits in advance. However, most earlier studies simply used users’ submitted words or user profiles to predict depression risk. The implicit information accessible in users’ posted images, which can be effective in depression detection, still remains unexplored. In this paper, an implicit and explicit multi-modal feature fusion (IEMFF) model is proposed for depression detection. We successfully make the implicit information inherent in users’ posted images explicit and further incorporate such explicit features with the textual features directly extracted from user-posted texts. A multi-modal feature fusion approach is applied for depression detection. Extensive experiments have been conducted on public Twitter datasets. Experimental results show that our approach has achieved state-of-the-art performance for depression detection.
Pengwei Hu 0001, Chenhao Lin, Jiajia Li 0004, Feng Tan 0002, Xue Han 0018, Xi Zhou 0007, Lun Hu
BIBM7
2023 MLGL: Model-free Lesion Generation and Learning for Diabetic Retinopathy Diagnosis
abstract
The approaches based on deep learning have achieved remarkable success in diabetic retinopathy detection. Due to the accountability in medical diagnosis, the interpretability of computer-aided diagnosis has recently been investigated. However, few existing approaches make full use of the explainable evidence to improve the diagnosis accuracy. In this paper, we propose a Model-free Lesion Generation and Learning (MLGL) framework to study the interpretability of diabetic retinopathy detection. We first generate visual explanations for diabetic retinopathy diagnosis using the proposed Gated Multi-layer Saliency Map (GMSM) module, which locates the accurate region of lesions by combining multi-layer heatmaps. Then we use the GMSM to extract the lesion patches and conduct the adaptive lesion transfer, iteratively generating new retinal fundus images with lesions. Especially, in this process, no additional generative models are trained. Finally, we merge the generated and original retinal fundus images for the model's training to learn robust lesion features. Overall, our method provides accurate explainable evidence and further addresses the data imbalance problem in diabetic retinopathy detection. The experimental results on four public datasets demonstrate the efficiency of our approach.
Jiajia Li 0004, Chenhao Lin, Feng Tan 0002, Lun Hu, Pengwei Hu 0001
BIBM6
2023 Learning RNA sequence patterns to interpretably identify m6A modification sites
abstract
N6-methyladenosine (m6A) regulates RNA post-transcriptional modification and translation processes, thereby regulating gene expression and cell fate. Hence, accurate identification of potential m6A modification sites is a key step to further reveal their biological functions and understand multiple biological processes such as gene regulation and epigenetic variation. Many computational methods have been developed to address this challenge. However, fewer studies have focused on an interpretable process of m6A modification site identification. Here, we propose an interpretable end-to-end predictor, called M6AInter, which learns the RNA sequence patterns related to modification sites through contrastive learning frameworks to achieve accurate identification of m6A modification sites. Specifically, M6AInter first utilizes chaos game representation theory and one-hot encoding to initialize the position and type information of nucleotides, respectively. On this basis, M6AInter extracts the position and type correlations shared by RNA sequences, and predicts the common sequence patterns by utilizing a graph contrastive clustering framework. These motifs and patterns are involved in describing the associations between RNA sequences and obtaining their low-dimensional representations. Finally, through a designed bias fusion block, these representations are combined with the frequency information of nucleotides to realize the identification of m6A modification sites. Extensive experimental results show that our model can accurately identify modified RNA sequences and can adaptively locate sequential regions associated with m6A modification sites on RNA sequences. Importantly, by exploring the role of these patterns in the identification tasks, M6AInter provides interpretable predictions and analysis at the sequence level.
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Pengwei Hu 0001, Lun Hu
BIBM6
2023 A Novel Graph Representation Learning Model for Drug Repositioning Using Graph Transition Probability Matrix Over Heterogenous Information Networks
Dongxu Li 0002, Bo-Wei Zhao, Xiao-Rui Su 0001, Zhu-Hong You, Pengwei Hu 0001, Lun Hu
ICIC (3)8
2023 TransOrga: End-To-End Multi-modal Transformer-Based Organoid Segmentation
Jiajia Li 0004, Zhu-Hong You, Lun Hu, Pengwei Hu 0001, Feng Tan 0002
ICIC (3)7
2023 A Deep Learning Approach Incorporating Data Missing Mechanism in Predicting Acute Kidney Injury in ICU
Zhengbo Zhang, Lei Zha, Fengcong, Xiao-Rui Su 0001, Bo-Wei Zhao, Lun Hu, Pengwei Hu 0001
ICIC (3)8
2023 Multi-level Subgraph Representation Learning for Drug-Disease Association Prediction Over Heterogeneous Biological Information Network
Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Dongxu Li 0002, Pengwei Hu 0001, Zhu-Hong You, Lun Hu
ICIC (3)7
2023 Incorporating higher order network structures to improve miRNA-disease association prediction based on functional modularity
abstract
As microRNAs (miRNAs) are involved in many essential biological processes, their abnormal expressions can serve as biomarkers and prognostic indicators to prevent the development of complex diseases, thus providing accurate early detection and prognostic evaluation. Although a number of computational methods have been proposed to predict miRNA-disease associations (MDAs) for further experimental verification, their performance is limited primarily by the inadequacy of exploiting lower order patterns characterizing known MDAs to identify missing ones from MDA networks. Hence, in this work, we present a novel prediction model, namely HiSCMDA, by incorporating higher order network structures for improved performance of MDA prediction. To this end, HiSCMDA first integrates miRNA similarity network, disease similarity network and MDA network to preserve the advantages of all these networks. After that, it identifies overlapping functional modules from the integrated network by predefining several higher order connectivity patterns of interest. Last, a path-based scoring function is designed to infer potential MDAs based on network paths across related functional modules. HiSCMDA yields the best performance across all datasets and evaluation metrics in the cross-validation and independent validation experiments. Furthermore, in the case studies, 49 and 50 out of the top 50 miRNAs, respectively, predicted for colon neoplasms and lung neoplasms have been validated by well-established databases. Experimental results show that rich higher order organizational structures exposed in the MDA network gain new insight into the MDA prediction based on higher order connectivity patterns.
Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Shengwu Xiong 0001, Lun Hu
Briefings Bioinform.6
2023 iGRLDTI: an improved graph representation learning method for predicting drug-target interactions over heterogeneous biological information network
abstract
MOTIVATION: The task of predicting drug-target interactions (DTIs) plays a significant role in facilitating the development of novel drug discovery. Compared with laboratory-based approaches, computational methods proposed for DTI prediction are preferred due to their high-efficiency and low-cost advantages. Recently, much attention has been attracted to apply different graph neural network (GNN) models to discover underlying DTIs from heterogeneous biological information network (HBIN). Although GNN-based prediction methods achieve better performance, they are prone to encounter the over-smoothing simulation when learning the latent representations of drugs and targets with their rich neighborhood information in HBIN, and thereby reduce the discriminative ability in DTI prediction. RESULTS: In this work, an improved graph representation learning method, namely iGRLDTI, is proposed to address the above issue by better capturing more discriminative representations of drugs and targets in a latent feature space. Specifically, iGRLDTI first constructs an HBIN by integrating the biological knowledge of drugs and targets with their interactions. After that, it adopts a node-dependent local smoothing strategy to adaptively decide the propagation depth of each biomolecule in HBIN, thus significantly alleviating over-smoothing by enhancing the discriminative ability of feature representations of drugs and targets. Finally, a Gradient Boosting Decision Tree classifier is used by iGRLDTI to predict novel DTIs. Experimental results demonstrate that iGRLDTI yields better performance that several state-of-the-art computational methods on the benchmark dataset. Besides, our case study indicates that iGRLDTI can successfully identify novel DTIs with more distinguishable features of drugs and targets. AVAILABILITY AND IMPLEMENTATION: Python codes and dataset are available at https://github.com/stevejobws/iGRLDTI/.
Bo-Wei Zhao, Xiao-Rui Su 0001, Pengwei Hu 0001, Zhu-Hong You, Lun Hu
Bioinform.6
2023 Biocaiv: an integrative webserver for motif-based clustering analysis and interactive visualization of biological networks
abstract
BACKGROUND: As an important task in bioinformatics, clustering analysis plays a critical role in understanding the functional mechanisms of many complex biological systems, which can be modeled as biological networks. The purpose of clustering analysis in biological networks is to identify functional modules of interest, but there is a lack of online clustering tools that visualize biological networks and provide in-depth biological analysis for discovered clusters. RESULTS: Here we present BioCAIV, a novel webserver dedicated to maximize its accessibility and applicability on the clustering analysis of biological networks. This, together with its user-friendly interface, assists biological researchers to perform an accurate clustering analysis for biological networks and identify functionally significant modules for further assessment. CONCLUSIONS: BioCAIV is an efficient clustering analysis webserver designed for a variety of biological networks. BioCAIV is freely available without registration requirements at http://bioinformatics.tianshanzw.cn:8888/BioCAIV/ .
Dongxu Li 0002, Bo-Wei Zhao, Xiao-Rui Su 0001, Jun Zhang 0003, Pengwei Hu 0001, Lun Hu
BMC Bioinform.8
2023 A survey of transformer-based multimodal pre-trained modals
Xue Han 0018, Junlan Feng, Chao Deng 0002, Hui Su, Lun Hu, Pengwei Hu 0001
Neurocomputing8
2023 Artificial intelligence accelerates multi-modal biomedical process: A Survey
Jiajia Li 0004, Xue Han 0018, Feng Tan 0002, Xi Zhou 0007, Lun Hu, Pengwei Hu 0001
Neurocomputing10
2023 Knowledge graph embedding for profiling the interaction between transcription factors and their target genes
abstract
Interactions between transcription factor and target gene form the main part of gene regulation network in human, which are still complicating factors in biological research. Specifically, for nearly half of those interactions recorded in established database, their interaction types are yet to be confirmed. Although several computational methods exist to predict gene interactions and their type, there is still no method available to predict them solely based on topology information. To this end, we proposed here a graph-based prediction model called KGE-TGI and trained in a multi-task learning manner on a knowledge graph that we specially constructed for this problem. The KGE-TGI model relies on topology information rather than being driven by gene expression data. In this paper, we formulate the task of predicting interaction types of transcript factor and target genes as a multi-label classification problem for link types on a heterogeneous graph, coupled with solving another link prediction problem that is inherently related. We constructed a ground truth dataset as benchmark and evaluated the proposed method on it. As a result of the 5-fold cross experiments, the proposed method achieved average AUC values of 0.9654 and 0.9339 in the tasks of link prediction and link type classification, respectively. In addition, the results of a series of comparison experiments also prove that the introduction of knowledge information significantly benefits to the prediction and that our methodology achieve state-of-the-art performance in this problem.
Yang-Han Wu, Jianqiang Li 0001, Zhu-Hong You, Pengwei Hu 0001, Lun Hu, Victor C. M. Leung, Zhihua Du
PLoS Comput. Biol.6
2023 Predicting Protein-Protein Interactions Using Sequence and Network Information via Variational Graph Autoencoder
abstract
Protein-protein interactions (PPIs) play a critical role in the proteomics study, and a variety of computational algorithms have been developed to predict PPIs. Though effective, their performance is constrained by high false-positive and false-negative rates observed in PPI data. To overcome this problem, a novel PPI prediction algorithm, namely PASNVGA, is proposed in this work by combining the sequence and network information of proteins via variational graph autoencoder. To do so, PASNVGA first applies different strategies to extract the features of proteins from their sequence and network information, and obtains a more compact form of these features using principal component analysis. In addition, PASNVGA designs a scoring function to measure the higher-order connectivity between proteins and so as to obtain a higher-order adjacency matrix. With all these features and adjacency matrices, PASNVGA trains a variational graph autoencoder model to further learn the integrated embeddings of proteins. The prediction task is then completed by using a simple feedforward neural network. Extensive experiments have been conducted on five PPI datasets collected from different species. Compared with several state-of-the-art algorithms, PASNVGA has been demonstrated as a promising PPI prediction algorithm.
Xin Luo 0001, Pengwei Hu 0001, Lun Hu
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 PPISB: A Novel Network-Based Algorithm of Predicting Protein-Protein Interactions With Mixed Membership Stochastic Blockmodel
abstract
Protein-protein interactions (PPIs) play an essential role for most of biological processes in cells. Many computational algorithms have thus been proposed to predict PPIs. However, most of them heavily rest on the biological information of proteins while ignoring the latent structural features of proteins presented in a PPI network. In this paper, we propose an efficient network-based prediction algorithm, namely PPISB, based on a mixed membership stochastic blockmodel. By simulating the generative process of a PPI network, PPISB is able to capture the latent community structures. The inference procedure adopted by PPISB further optimizes the membership distributions of proteins over different complexes. After that, a distance measure is designed to compute the similarity between two proteins in terms of their likelihoods of being in the same complex, thus verifying whether they interact with each other or not. To evaluate the performance of PPISB, a series of extensive experiments have been conducted with five PPI networks collected from different species and the results demonstrate that PPISB has a promising performance when applied to predict PPIs in terms of several evaluation metrics. Hence, we reason that PPISB is preferred over state-of-the-art network-based prediction algorithms especially for predicting potential PPIs.
Wen Yang 0019, Yue Yang 0035, Jun Zhang 0003, Lusheng Wang 0001, Lun Hu
IEEE ACM Trans. Comput. Biol. Bioinform.7
2023 FCAN-MOPSO: An Improved Fuzzy-Based Graph Clustering Algorithm for Complex Networks With Multiobjective Particle Swarm Optimization
abstract
Performing an accurate clustering analysis is of great significance for us to understand the behavior of complex networks, and a variety of graph clustering algorithms have, thus, been proposed to do so by taking into account network topology and node attributes. Among them, fuzzy clustering algorithm for complex networks (FCAN) is an established fuzzy clustering algorithm that optimizes the memberships of nodes based on dense structures and content relevance. This article proposes an improved fuzzy-based graph clustering algorithm, namely FCAN-multi objective particle swarm optimization (MOPSO) that retains all the benefits associated with FCAN while achieving significantly increased convergence rate using multiobjective particle swarm optimization (MOPSO). To do so, FCAN-MOPSO first modifies the original optimization model of FCAN by adopting an instance-frequency-weighted regularization, which enhances the ability of FCAN-MOPSO to handle the imbalance observed in the distribution of fuzzy memberships of nodes. After that, FCAN-MOPSO decomposes its optimization problem into a set of suboptimization problems. Following the MOPSO framework, FCAN-MOPSO develops an effective solution to reach a consensus optimization among them by balancing the global exploration and local exploitation abilities of particles. A theoretical analysis is provided to prove the global convergence of FCAN-MOPSO. Extensive experiments have been conducted to evaluate the performance of FCAN-MOPSO on five real-world complex networks with different scale, and experimental results demonstrate that when compared with state-of-the-art clustering algorithms, FCAN-MOPSO achieves a better accuracy performance with improved convergence. Hence, FCAN-MOPSO is a promising graph clustering algorithm to precisely and efficiently discover clusters in complex networks.
Lun Hu, Yue Yang 0035, Zehai Tang, Xin Luo 0001
IEEE Trans. Fuzzy Syst.1
2023 PPAEDTI: Personalized Propagation Auto-Encoder Model for Predicting Drug-Target Interactions
abstract
Identifying protein targets for drugs establishes an indispensable knowledge foundation for drug repurposing and drug development. Though expensive and time-consuming, vitro trials are widely employed to discover drug targets, and the existing relevant computational algorithms still cannot satisfy the demand for real application in drug R&D with regards to the prediction accuracy and performance efficiency, which are urgently needed to be improved. To this end, we propose here the PPAEDTI model, which uses the graph personalized propagation technique to predict drug-target interactions from the known interaction network. To evaluate the prediction performance, six benchmark datasets were used for testing with some state-of-the-art methods compared. As a result, using the 5-fold cross-validation, the proposed PPAEDTI model achieves average AUCs>90% on 5 collected datasets. We also manually checked the top-20 prediction list for 2 proteins (hsa:775 and hsa:779) and a kind of drug (D00618), and successfully confirmed 18, 17, and 20 items from the public datasets, respectively. The experimental results indicate that, given known drug-target interactions, the PPAEDTI model can provide accurate predictions for the new ones, which is anticipated to serve as a useful tool for pharmacology research. Using the proposed model that was trained with the collected datasets, we have built a computational platform that is accessible at http://120.77.11.78/PPAEDTI/ and corresponding codes and datasets are also released.
Yue-Chao Li, Zhu-Hong You, Lei Wang 0121, Leon Wong, Lun Hu, Pengwei Hu 0001
IEEE J. Biomed. Health Informatics6
2023 Predicting Drug-Target Interactions Over Heterogeneous Information Network
abstract
Identifying Drug-Target Interactions (DTIs) is a critical step in studying pathogenesis and drug development. Due to the fact that conventional experimental methods usually suffer from high costs and low efficiency, various computational methods have been proposed to detect potential DTIs by extracting features from the biological information of drugs and their target proteins. Though effective, most of them fall short of considering the topological structure of the DTI network, which provides a global view to discover novel DTIs. In this paper, a network-based computational method, namely LG-DTI, is proposed to accurately predict DTIs over a heterogeneous information network. For drugs and target proteins, LG-DTI first learns not only their local representations from drug molecular structures and protein sequences, but also their global representations by using a semi-supervised heterogeneous network embedding method. These two kinds of representations consist of the final representations of drugs and target proteins, which are then incorporated into a Random Forest classifier to complete the task of DTI prediction. The performance of LG-DTI has been evaluated on two independent datasets and also compared with several state-of-the-art methods. Experimental results show the superior performance of LG-DTI. Moreover, our case study indicates that LG-DTI can be a valuable tool for identifying novel DTIs.
Xiao-Rui Su 0001, Pengwei Hu 0001, Zhu-Hong You, Lun Hu
IEEE J. Biomed. Health Informatics5
2022 Cost and Care Insight: An Interactive and Scalable Hierarchical Learning System for Identifying Cost Saving Opportunities
David Koepke, Bibo Hao, Jing Mei, Xu Min, Rachna Gupta, Rajashree Joshi, Fiona McNaughton, Bo-Wei Zhao, Lun Hu, Pengwei Hu 0001
ICIC (1)11
2022 Predicting Drug-Disease Associations via Meta-path Representation Learning based on Heterogeneous Information Net works
Menglong Zhang, Bo-Wei Zhao, Lun Hu, Zhu-Hong You
ICIC (2)3
2022 MRLDTI: A Meta-path-Based Representation Learning Model for Drug-Target Interaction Prediction
Bo-Wei Zhao, Lun Hu, Pengwei Hu 0001, Zhu-Hong You, Xiao-Rui Su 0001, Dongxu Li 0002, Ping Zhang 0027
ICIC (2)2
2022 A Novel Fuzzy-Based MOPSO Algorithm for Identifying Clusters From Complex Networks
abstract
Many complicated systems can be modeled as complex networks, and a variety of graph clustering algorithms have been proposed to perform accurate clustering analysis for better understanding system behaviors. However, most of them suffer the disadvantage of slow convergence. In this paper, we incorporate multi-objective particle swarm optimization (MOPSO) into a well-established fuzzy clustering algorithm, i.e., FCAN, and propose an improved Fuzzy-based Graph Clustering Algorithm, namely IMFCAN, which retains all the benefits gained with FCAN while achieving significantly fast convergence rate. Specially, IMFCAN enhances the ability of handling the imbalance observed in the distribution of fuzzy membership of nodes by introducing an instance-frequency-weighted regularization (IR) scheme. After that, IMFCAN develops an effective solution to reach a consensus optimization among them by balancing global exploration and local exploitation abilities of particles. Experimental results on four practical datasets demonstrate that IMFCAN performs better than several state-of-the-art clustering algorithm in terms of accuracy and convergence. Hence, IMFCAN is a promising algorithm for addressing the clustering analysis of complex networks.
Yue Yang 0035, Xiao-Rui Su 0001, Bo-Wei Zhao, Lun Hu
ICTAI5
2022 GraphTGI: an attention-based graph embedding model for predicting TF-target gene interactions
abstract
MOTIVATION: Interaction between transcription factor (TF) and its target genes establishes the knowledge foundation for biological researches in transcriptional regulation, the number of which is, however, still limited by biological techniques. Existing computational methods relevant to the prediction of TF-target interactions are mostly proposed for predicting binding sites, rather than directly predicting the interactions. To this end, we propose here a graph attention-based autoencoder model to predict TF-target gene interactions using the information of the known TF-target gene interaction network combined with two sequential and chemical gene characters, considering that the unobserved interactions between transcription factors and target genes can be predicted by learning the pattern of the known ones. To the best of our knowledge, the proposed model is the first attempt to solve this problem by learning patterns from the known TF-target gene interaction network. RESULTS: In this paper, we formulate the prediction task of TF-target gene interactions as a link prediction problem on a complex knowledge graph and propose a deep learning model called GraphTGI, which is composed of a graph attention-based encoder and a bilinear decoder. We evaluated the prediction performance of the proposed method on a real dataset, and the experimental results show that the proposed model yields outstanding performance with an average AUC value of 0.8864 +/- 0.0057 in the 5-fold cross-validation. It is anticipated that the GraphTGI model can effectively and efficiently predict TF-target gene interactions on a large scale. AVAILABILITY: Python code and the datasets used in our studies are made available at https://github.com/YanghanWu/GraphTGI.
Zhihua Du, Yang-Han Wu, Jie Chen 0027, Gui-Qing Pan, Lun Hu, Zhu-Hong You, Jianqiang Li 0001
Briefings Bioinform.6
2022 A deep learning method for repurposing antiviral drugs against new viruses via multi-view nonnegative matrix factorization and its application to SARS-CoV-2
abstract
The outbreak of COVID-19 caused by SARS-coronavirus (CoV)-2 has made millions of deaths since 2019. Although a variety of computational methods have been proposed to repurpose drugs for treating SARS-CoV-2 infections, it is still a challenging task for new viruses, as there are no verified virus-drug associations (VDAs) between them and existing drugs. To efficiently solve the cold-start problem posed by new viruses, a novel constrained multi-view nonnegative matrix factorization (CMNMF) model is designed by jointly utilizing multiple sources of biological information. With the CMNMF model, the similarities of drugs and viruses can be preserved from their own perspectives when they are projected onto a unified latent feature space. Based on the CMNMF model, we propose a deep learning method, namely VDA-DLCMNMF, for repurposing drugs against new viruses. VDA-DLCMNMF first initializes the node representations of drugs and viruses with their corresponding latent feature vectors to avoid a random initialization and then applies graph convolutional network to optimize their representations. Given an arbitrary drug, its probability of being associated with a new virus is computed according to their representations. To evaluate the performance of VDA-DLCMNMF, we have conducted a series of experiments on three VDA datasets created for SARS-CoV-2. Experimental results demonstrate that the promising prediction accuracy of VDA-DLCMNMF. Moreover, incorporating the CMNMF model into deep learning gains new insight into the drug repurposing for SARS-CoV-2, as the results of molecular docking experiments reveal that four antiviral drugs identified by VDA-DLCMNMF have the potential ability to treat SARS-CoV-2 infections.
Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Lei Wang 0121, Bo-Wei Zhao
Briefings Bioinform.2
2022 Attention-based Knowledge Graph Representation Learning for Predicting Drug-drug Interactions
abstract
Drug-drug interactions (DDIs) are known as the main cause of life-threatening adverse events, and their identification is a key task in drug development. Existing computational algorithms mainly solve this problem by using advanced representation learning techniques. Though effective, few of them are capable of performing their tasks on biomedical knowledge graphs (KGs) that provide more detailed information about drug attributes and drug-related triple facts. In this work, an attention-based KG representation learning framework, namely DDKG, is proposed to fully utilize the information of KGs for improved performance of DDI prediction. In particular, DDKG first initializes the representations of drugs with their embeddings derived from drug attributes with an encoder-decoder layer, and then learns the representations of drugs by recursively propagating and aggregating first-order neighboring information along top-ranked network paths determined by neighboring node embeddings and triple facts. Last, DDKG estimates the probability of being interacting for pairwise drugs with their representations in an end-to-end manner. To evaluate the effectiveness of DDKG, extensive experiments have been conducted on two practical datasets with different sizes, and the results demonstrate that DDKG is superior to state-of-the-art algorithms on the DDI prediction task in terms of different evaluation metrics across all datasets.
Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Bo-Wei Zhao
Briefings Bioinform.2
2022 iGRLCDA: identifying circRNA-disease association based on graph representation learning
abstract
While the technologies of ribonucleic acid-sequence (RNA-seq) and transcript assembly analysis have continued to improve, a novel topology of RNA transcript was uncovered in the last decade and is called circular RNA (circRNA). Recently, researchers have revealed that they compete with messenger RNA (mRNA) and long noncoding for combining with microRNA in gene regulation. Therefore, circRNA was assumed to be associated with complex disease and discovering the relationship between them would contribute to medical research. However, the work of identifying the association between circRNA and disease in vitro takes a long time and usually without direction. During these years, more and more associations were verified by experiments. Hence, we proposed a computational method named identifying circRNA-disease association based on graph representation learning (iGRLCDA) for the prediction of the potential association of circRNA and disease, which utilized a deep learning model of graph convolution network (GCN) and graph factorization (GF). In detail, iGRLCDA first derived the hidden feature of known associations between circRNA and disease using the Gaussian interaction profile (GIP) kernel combined with disease semantic information to form a numeric descriptor. After that, it further used the deep learning model of GCN and GF to extract hidden features from the descriptor. Finally, the random forest classifier is introduced to identify the potential circRNA-disease association. The five-fold cross-validation of iGRLCDA shows strong competitiveness in comparison with other excellent prediction models at the gold standard data and achieved an average area under the receiver operating characteristic curve of 0.9289 and an area under the precision-recall curve of 0.9377. On reviewing the prediction results from the relevant literature, 22 of the top 30 predicted circRNA-disease associations were noted in recent published papers. These exceptional results make us believe that iGRLCDA can provide reliable circRNA-disease associations for medical research and reduce the blindness of wet-lab experiments.
Lei Wang 0121, Zhu-Hong You, Lun Hu, Bo-Wei Zhao, Zhengwei Li 0001, Yang-Ming Li
Briefings Bioinform.4
2022 HINGRL: predicting drug-disease associations with graph representation learning on heterogeneous information networks
abstract
Identifying new indications for drugs plays an essential role at many phases of drug research and development. Computational methods are regarded as an effective way to associate drugs with new indications. However, most of them complete their tasks by constructing a variety of heterogeneous networks without considering the biological knowledge of drugs and diseases, which are believed to be useful for improving the accuracy of drug repositioning. To this end, a novel heterogeneous information network (HIN) based model, namely HINGRL, is proposed to precisely identify new indications for drugs based on graph representation learning techniques. More specifically, HINGRL first constructs a HIN by integrating drug-disease, drug-protein and protein-disease biological networks with the biological knowledge of drugs and diseases. Then, different representation strategies are applied to learn the features of nodes in the HIN from the topological and biological perspectives. Finally, HINGRL adopts a Random Forest classifier to predict unknown drug-disease associations based on the integrated features of drugs and diseases obtained in the previous step. Experimental results demonstrate that HINGRL achieves the best performance on two real datasets when compared with state-of-the-art models. Besides, our case studies indicate that the simultaneous consideration of network topology and biological knowledge of drugs and diseases allows HINGRL to precisely predict drug-disease associations from a more comprehensive perspective. The promising performance of HINGRL also reveals that the utilization of rich heterogeneous information provides an alternative view for HINGRL to identify novel drug-disease associations especially for new diseases.
Bo-Wei Zhao, Lun Hu, Zhu-Hong You, Lei Wang 0121, Xiao-Rui Su 0001
Briefings Bioinform.2
2022 A geometric deep learning framework for drug repositioning over heterogeneous information networks
abstract
Drug repositioning (DR) is a promising strategy to discover new indicators of approved drugs with artificial intelligence techniques, thus improving traditional drug discovery and development. However, most of DR computational methods fall short of taking into account the non-Euclidean nature of biomedical network data. To overcome this problem, a deep learning framework, namely DDAGDL, is proposed to predict drug-drug associations (DDAs) by using geometric deep learning (GDL) over heterogeneous information network (HIN). Incorporating complex biological information into the topological structure of HIN, DDAGDL effectively learns the smoothed representations of drugs and diseases with an attention mechanism. Experiment results demonstrate the superior performance of DDAGDL on three real-world datasets under 10-fold cross-validation when compared with state-of-the-art DR methods in terms of several evaluation metrics. Our case studies and molecular docking experiments indicate that DDAGDL is a promising DR tool that gains new insights into exploiting the geometric prior knowledge for improved efficacy.
Bo-Wei Zhao, Xiao-Rui Su 0001, Pengwei Hu 0001, Yu-Peng Ma, Xi Zhou 0007, Lun Hu
Briefings Bioinform.6
2022 Effectively predicting HIV-1 protease cleavage sites by using an ensemble learning approach
abstract
BACKGROUND: The site information of substrates that can be cleaved by human immunodeficiency virus 1 proteases (HIV-1 PRs) is of great significance for designing effective inhibitors against HIV-1 viruses. A variety of machine learning-based algorithms have been developed to predict HIV-1 PR cleavage sites by extracting relevant features from substrate sequences. However, only relying on the sequence information is not sufficient to ensure a promising performance due to the uncertainty in the way of separating the datasets used for training and testing. Moreover, the existence of noisy data, i.e., false positive and false negative cleavage sites, could negatively influence the accuracy performance. RESULTS: In this work, an ensemble learning algorithm for predicting HIV-1 PR cleavage sites, namely EM-HIV, is proposed by training a set of weak learners, i.e., biased support vector machine classifiers, with the asymmetric bagging strategy. By doing so, the impact of data imbalance and noisy data can thus be alleviated. Besides, in order to make full use of substrate sequences, the features used by EM-HIV are collected from three different coding schemes, including amino acid identities, chemical properties and variable-length coevolutionary patterns, for the purpose of constructing more relevant feature vectors of octamers. Experiment results on three independent benchmark datasets demonstrate that EM-HIV outperforms state-of-the-art prediction algorithm in terms of several evaluation metrics. Hence, EM-HIV can be regarded as a useful tool to accurately predict HIV-1 PR cleavage sites.
Lun Hu, Zhenfeng Li, Zehai Tang, Xi Zhou 0007, Pengwei Hu 0001
BMC Bioinform.1
2022 Multi-view heterogeneous molecular network representation learning for protein-protein interaction prediction
abstract
BACKGROUND: Protein-protein interaction (PPI) plays an important role in regulating cells and signals. Despite the ongoing efforts of the bioassay group, continued incomplete data limits our ability to understand the molecular roots of human disease. Therefore, it is urgent to develop a computational method to predict PPIs from the perspective of molecular system. METHODS: In this paper, a highly efficient computational model, MTV-PPI, is proposed for PPI prediction based on a heterogeneous molecular network by learning inter-view protein sequences and intra-view interactions between molecules simultaneously. On the one hand, the inter-view feature is extracted from the protein sequence by k-mer method. On the other hand, we use a popular embedding method LINE to encode the heterogeneous molecular network to obtain the intra-view feature. Thus, the protein representation used in MTV-PPI is constructed by the aggregation of its inter-view feature and intra-view feature. Finally, random forest is integrated to predict potential PPIs. RESULTS: To prove the effectiveness of MTV-PPI, we conduct extensive experiments on a collected heterogeneous molecular network with the accuracy of 86.55%, sensitivity of 82.49%, precision of 89.79%, AUC of 0.9301 and AUPR of 0.9308. Further comparison experiments are performed with various protein representations and classifiers to indicate the effectiveness of MTV-PPI in predicting PPIs based on a complex network. CONCLUSION: The achieved experimental results illustrate that MTV-PPI is a promising tool for PPI prediction, which may provide a new perspective for the future interactions prediction researches based on heterogeneous molecular network.
Xiao-Rui Su 0001, Lun Hu, Zhu-Hong You, Pengwei Hu 0001, Bo-Wei Zhao
BMC Bioinform.2
2022 RLFDDA: a meta-path based graph representation learning model for drug-disease association prediction
abstract
BACKGROUND: Drug repositioning is a very important task that provides critical information for exploring the potential efficacy of drugs. Yet developing computational models that can effectively predict drug-disease associations (DDAs) is still a challenging task. Previous studies suggest that the accuracy of DDA prediction can be improved by integrating different types of biological features. But how to conduct an effective integration remains a challenging problem for accurately discovering new indications for approved drugs. METHODS: In this paper, we propose a novel meta-path based graph representation learning model, namely RLFDDA, to predict potential DDAs on heterogeneous biological networks. RLFDDA first calculates drug-drug similarities and disease-disease similarities as the intrinsic biological features of drugs and diseases. A heterogeneous network is then constructed by integrating DDAs, disease-protein associations and drug-protein associations. With such a network, RLFDDA adopts a meta-path random walk model to learn the latent representations of drugs and diseases, which are concatenated to construct joint representations of drug-disease associations. As the last step, we employ the random forest classifier to predict potential DDAs with their joint representations. RESULTS: To demonstrate the effectiveness of RLFDDA, we have conducted a series of experiments on two benchmark datasets by following a ten-fold cross-validation scheme. The results show that RLFDDA yields the best performance in terms of AUC and F1-score when compared with several state-of-the-art DDAs prediction models. We have also conducted a case study on two common diseases, i.e., paclitaxel and lung tumors, and found that 7 out of top-10 diseases and 8 out of top-10 drugs have already been validated for paclitaxel and lung tumors respectively with literature evidence. Hence, the promising performance of RLFDDA may provide a new perspective for novel DDAs discovery over heterogeneous networks.
Menglong Zhang, Bo-Wei Zhao, Xiao-Rui Su 0001, Yue Yang 0035, Lun Hu
BMC Bioinform.6
2022 An Algorithm of Inductively Identifying Clusters From Attributed Graphs
abstract
Attributed graphs are widely used to represent network data where the attribute information of nodes is available. To address the problem of identifying clusters in attributed graphs, most of existing solutions are developed simply based on certain particular assumptions related to the characteristics of clusters of interest. However, it is yet unknown whether such assumed characteristics are consistent with attributed graphs. To overcome this issue, we innovatively introduce an inductive clustering algorithm that tends to address the clustering problem for attributed graphs without any assumption made on the clusters. To do so, we first process the attribute information to obtain pairwise attribute values that significantly frequently co-occur in adjacent nodes as we believe that they have potential ability to represent the characteristics of a given attributed graph. For two adjacent nodes, their likelihood of being grouped in the same cluster can be weighted by their ability to characterize the graph. Then based on these verifed characteristics instead of assumed ones, a depth-first search strategy is applied to perform the clustering task. Moreover, we are able to classify clusters such that their significance can be indicated. The experimental results demonstrate the performance and usefulness of our algorithm.
Lun Hu, Shicheng Yang, Xin Luo 0001, MengChu Zhou
IEEE Trans. Big Data1
2022 Identifying Protein Complexes From Protein-Protein Interaction Networks Based on Fuzzy Clustering and GO Semantic Information
abstract
Protein complexes are of great significance to provide valuable insights into the mechanisms of biological processes of proteins. A variety of computational algorithms have thus been proposed to identify protein complexes in a protein-protein interaction network. However, few of them can perform their tasks by taking into account both network topology and protein attribute information in a unified fuzzy-based clustering framework. Since proteins in the same complex are similar in terms of their attribute information and the consideration of fuzzy clustering can also make it possible for us to identify overlapping complexes, we target to propose such a novel fuzzy-based clustering framework, namely FCAN-PCI, for an improved identification accuracy. To do so, the semantic similarity between the attribute information of proteins is calculated and we then integrate it into a well-established fuzzy clustering model together with the network topology. After that, a momentum method is adopted to accelerate the clustering procedure. FCAN-PCI finally applies a heuristical search strategy to identify overlapping protein complexes. A series of extensive experiments have been conducted to evaluate the performance of FCAN-PCI by comparing it with state-of-the-art identification algorithms and the results demonstrate the promising performance of FCAN-PCI.
Xiangyu Pan, Lun Hu, Pengwei Hu 0001, Zhu-Hong You
IEEE ACM Trans. Comput. Biol. Bioinform.2
2022 A Fast Fuzzy Clustering Algorithm for Complex Networks via a Generalized Momentum Method
abstract
Complex networks have been widely adopted to represent a variety of complicated systems. Given a complex network, it is of great significance to perform accurate clustering for better understanding its intrinsic organization. To this end, a fuzzy-based clustering algorithm, i.e., FCAN, has been developed. Though effective, FCAN suffers from the disadvantage of slow convergence, which in return constrains its efficiency. To address this issue, this article proposes a fast fuzzy clustering algorithm, namely, F$^2$CAN, which incorporates a generalized momentum method into FCAN. Its fast convergence is rigorous justified in theory. Empirical studies on five datasets from real applications demonstrate that F$^2$CAN achieves a better performance when compared with FCAN and several state-of-the-art clustering algorithms in terms of convergence rate and clustering accuracy simultaneously. Hence, F$^2$CAN has potential for addressing the clustering analysis of large-scale complex networks emerging from industrial applications.
Lun Hu, Xiangyu Pan, Zehai Tang, Xin Luo 0001
IEEE Trans. Fuzzy Syst.1
2022 Generalized Nesterov's Acceleration-Incorporated, Non-Negative and Adaptive Latent Factor Analysis
abstract
A non-negative latent factor (NLF) model with a single latent factor-dependent, non-negative and multiplicative update (SLF-NMU) algorithm is frequently adopted to extract useful knowledge from non-negative data represented by high-dimensional and sparse (HiDS) matrices arising from various service-oriented applications. However, its convergence rate is slow. To address this issue, this study proposes aGeneralized Nesterov's acceleration-incorporated,Non-negative andAdaptiveLatentFactor (GNALF) model. It results from a) incorporating a generalized Nesterov's accelerated gradient (NAG) method into an SLF-NMU algorithm, thereby achieving anNAG-incorporated andelement-orientednon-negative (NEN) algorithm to perform efficient parameter update; and b) making its regularization and acceleration parameters self-adaptive via incorporating the principle of a particle swarm optimization algorithm into the training process, thereby implementing a highly adaptive and practical model. Empirical studies on six large sparse matrices from different recommendation service applications show that a GNALF model achieves very high convergence rate without the need of hyper-parameter tuning, making its computational efficiency significantly higher than state-of-the-art models. Meanwhile, such efficiency gain does not result in accuracy loss, since its prediction accuracy is comparable with its peers. Hence, it can better serve practical service applications with real-time demands.
Xin Luo 0001, Zhigang Liu 0006, Lun Hu, MengChu Zhou
IEEE Trans. Serv. Comput.4
2021 OnSum: Extractive Single Document Summarization Using Ordered Neuron LSTM
Xue Han 0018, Lun Hu, Pengwei Hu 0001
ICIC (2)4
2021 An Ensemble Learning Algorithm for Predicting HIV-1 Protease Cleavage Sites
Zhenfeng Li, Pengwei Hu 0001, Lun Hu
ICIC (3)3
2021 A Multi-graph Deep Learning Model for Predicting Drug-Disease Associations
Bo-Wei Zhao, Zhu-Hong You, Lun Hu, Leon Wong, Ping Zhang 0027
ICIC (3)3
2021 Predicting Large-scale Protein-protein Interactions by Extracting Coevolutionary Patterns with MapReduce Paradigm
abstract
Protein-protein interactions are of great significance for us to understand the functional mechanisms of proteins. With the rapid development of high-throughput genomic technology, the amount of protein-protein interaction data has become so big that most of existing prediction algorithms are no longer applicable. To address this problem, we develop a distributed framework by reimplementing one of state-of-the-art algorithms, i.e., CoFex, by using MapReduce. In particular, we adopt a novel tree-based data structure to reduce the heavy memory consumption cased by the huge sequence information of proteins. After that, the procedure of CoFex is modified by following the paradigm of MapReduce such that the prediction task can be completed in a distributed manner, thus fulfilling the demanding requirements of large-scale protein-protein interaction prediction. A series of experiments have been conducted to evaluate the performance of the proposed distributed framework in terms of both efficiency and effectiveness. Experimental results demonstrate that the proposed framework can considerably improve the efficiency of CoFex by achieving more than two-orders-of-magnitude improvement in computational efficiency while retaining a comparable level of accuracy.
Lun Hu, Bo-Wei Zhao, Shicheng Yang, Xin Luo 0001, MengChu Zhou
SMC1
2021 A survey on computational models for predicting protein-protein interactions
abstract
Proteins interact with each other to play critical roles in many biological processes in cells. Although promising, laboratory experiments usually suffer from the disadvantages of being time-consuming and labor-intensive. The results obtained are often not robust and considerably uncertain. Due recently to advances in high-throughput technologies, a large amount of proteomics data has been collected and this presents a significant opportunity and also a challenge to develop computational models to predict protein-protein interactions (PPIs) based on these data. In this paper, we present a comprehensive survey of the recent efforts that have been made towards the development of effective computational models for PPI prediction. The survey introduces the algorithms that can be used to learn computational models for predicting PPIs, and it classifies these models into different categories. To understand their relative merits, the paper discusses different validation schemes and metrics to evaluate the prediction performance. Biological databases that are commonly used in different experiments for performance comparison are also described and their use in a series of extensive experiments to compare different prediction models are discussed. Finally, we present some open issues in PPI prediction for future work. We explain how the performance of PPI prediction can be improved if these issues are effectively tackled.
Lun Hu, Pengwei Hu 0001, Zhu-Hong You
Briefings Bioinform.1
2021 HiSCF: leveraging higher-order structures for clustering analysis in biological networks
abstract
MOTIVATION: Clustering analysis in a biological network is to group biological entities into functional modules, thus providing valuable insight into the understanding of complex biological systems. Existing clustering techniques make use of lower-order connectivity patterns at the level of individual biological entities and their connections, but few of them can take into account of higher-order connectivity patterns at the level of small network motifs. RESULTS: Here, we present a novel clustering framework, namely HiSCF, to identify functional modules based on the higher-order structure information available in a biological network. Taking advantage of higher-order Markov stochastic process, HiSCF is able to perform the clustering analysis by exploiting a variety of network motifs. When compared with several state-of-the-art clustering models, HiSCF yields the best performance for two practical clustering applications, i.e. protein complex identification and gene co-expression module detection, in terms of accuracy. The promising performance of HiSCF demonstrates that the consideration of higher-order network motifs gains new insight into the analysis of biological networks, such as the identification of overlapping protein complexes and the inference of new signaling pathways, and also reveals the rich higher-order organizational structures presented in biological networks. AVAILABILITY AND IMPLEMENTATION: HiSCF is available at https://github.com/allenv5/HiSCF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lun Hu, Jun Zhang 0003, Xiangyu Pan, Hong Yan 0001, Zhu-Hong You
Bioinform.1
2020 A Highly Efficient Biomolecular Network Representation Model for Predicting Drug-Disease Associations
Hanjing Jiang, Zhu-Hong You, Lun Hu, Zhen-Hao Guo, Leon Wong
ICIC (3)3
2020 Embracing Disease Progression with a Learning System for Real World Evidence Discovery
Zefang Tang, Lun Hu, Xu Min, Jing Mei, Kenney Ng, Shaochun Li, Pengwei Hu 0001, Zhu-Hong You
ICIC (2)2
2020 A Novel Stochastic Block Model for Network-Based Prediction of Protein-Protein Interactions
Pengwei Hu 0001, Lun Hu
ICIC (2)3
2020 The identification of variable-length coevolutionary patterns for predicting HIV-1 protease cleavage sites
abstract
The substrate specificity of human immunodeficiency virus 1 (HIV-1) plays an essential role in designing HIV-1 inhibitors for therapy purpose. Hence, to predict the existence of cleavage sites in HIV-1 protease, a variety of computational algorithms have been developed by following the homogeneous information in substrate sequences. However, few of them can fully exploit such information, as they are not capable of identifying variable-length coevolutionary patterns. To overcome this limitation, we propose a novel algorithm with which variable-length coevolutionary patterns can be identified. Based on these patterns, we compose the feature vector for each of substrates and train the SVM classifier to the purpose of predicting HIV1 protease cleavage sites. Experimental results show that the use of variable-length coevolutionary patterns can improve the prediction performance in terms of AUC and PR-AUC analysis.
Zhenfeng Li, Lun Hu
SMC2
2020 Energy-aware task scheduling with time constraint for heterogeneous cloud datacenters
abstract
Summary Energy optimization with time constraint has become a timely and significant challenge for the datacenters. In this paper, a hardware and software collaborative optimization strategy is implemented to minimize the energy cost while satisfying the time constraint of the datacenters. In the hardware aspect, a DVFS‐capable CPU/GPU/FPGA heterogeneous computing infrastructure is built. This infrastructure can adjust its hardware characteristics dynamically in terms of the software run‐time contexts so that the applications can be executed efficiently with less time and lower energy cost. In the software aspect, a deadline‐aware energy‐efficient task scheduling algorithm based on the Q‐learning approach is investigated. This algorithm can adjust its searching directions smartly in terms of the environment feedback so that it can achieve better optimization performance comparing with the traditional genetic algorithm. However, its convergence time is long due to the large amount of training work, making it inappropriate to be applied in the large‐scale datacenters. To ease this problem, we proposed another new algorithm named Rapid Local Convolution Optimization (RLCO) and combine it with the Q‐learning algorithm. By doing this, the convergence time of the Q‐learning mechanism can be decreased significantly. We conducted both the simulation and real‐world experiments to evaluate the performance of our approaches, and the results proved the proposed algorithm running on the DVFS‐capable heterogeneous hardware architecture could decrease the energy cost of the datacenter significantly even if the datacenter is in large scale.
Xing Liu 0002, Panwen Liu, Lun Hu, Chengming Zou, Zhangyu Cheng
Concurr. Comput. Pract. Exp.3
2020 Incorporating the Coevolving Information of Substrates in Predicting HIV-1 Protease Cleavage Sites
abstract
Human immunodeficiency virus 1 (HIV-1) protease (PR) plays a crucial role in the maturation of the virus. The study of substrate specificity of HIV-1 PR as a new endeavor strives to increase our ability to understand how HIV-1 PR recognizes its various cleavage sites. To predict HIV-1 PR cleavage sites, most of the existing approaches have been developed solely based on the homogeneity of substrate sequence information with supervised classification techniques. Although efficient, these approaches are found to be restricted to the ability of explaining their results and probably provide few insights into the mechanisms by which HIV-1 PR cleaves the substrates in a site-specific manner. In this work, a coevolutionary pattern-based prediction model for HIV-1 PR cleavage sites, namely EvoCleave, is proposed by integrating the coevolving information obtained from substrate sequences with a linear SVM classifier. The experiment results showed that EvoCleave yielded a very promising performance in terms of ROC analysis and f-measure. We also prospectively assessed the biological significance of coevolutionary patterns by applying them to study three fundamental issues of HIV-1 PR cleavage site. The analysis results demonstrated that the coevolutionary patterns offered valuable insights into the understanding of substrate specificity of HIV-1 PR.
Lun Hu, Pengwei Hu 0001, Xin Luo 0001, Zhu-Hong You
IEEE ACM Trans. Comput. Biol. Bioinform.1
2020 Non-Negativity Constrained Missing Data Estimation for High-Dimensional and Sparse Matrices from Industrial Applications
abstract
High-dimensional and sparse (HiDS) matrices are commonly seen in big-data-related industrial applications like recommender systems. Latent factor (LF) models have proven to be accurate and efficient in extracting hidden knowledge from them. However, they mostly fail to fulfill the non-negativity constraints that describe the non-negative nature of many industrial data. Moreover, existing models suffer from slow convergence rate. An alternating-direction-method of multipliers-based non-negative LF (AMNLF) model decomposes the task of non-negative LF analysis on an HiDS matrix into small subtasks, where each task is solved based on the latest solutions to the previously solved ones, thereby achieving fast convergence and high prediction accuracy for its missing data. This paper theoretically analyzes the characteristics of an AMNLF model, and presents detailed empirical studies regarding its performance on nine HiDS matrices from industrial applications currently in use. Therefore, its capability of addressing HiDS matrices is justified in both theory and practice.
Xin Luo 0001, MengChu Zhou, Shuai Li 0002, Lun Hu, Mingsheng Shang 0001
IEEE Trans. Cybern.4
2020 A Variational Bayesian Framework for Cluster Analysis in a Complex Network
abstract
A complex network is a network with non-trivial topological structures. It contains not just topological information but also attribute information available in the rich content of nodes. Concerning the task of cluster analysis in a complex network, model-based algorithms are preferred over distance-based ones, as they avoid designing specific distance measures. However, their models are only applicable to complex networks where the attribute information is composed of attributes in binary form. To overcome this disadvantage, we introduce a three-layer node-attribute-value hierarchical structure to describe the attribute information in a flexible and interpretable manner. Then, a new Bayesian model is proposed to simulate the generative process of a complex network. In this model, the attribute information is generated by following the hierarchical structure while the links between pairwise nodes are generated by a stochastic blockmodel. To solve the corresponding inference problem, we develop a variational Bayesian algorithm called TARA, which allows us to identify functionally meaningful clusters through an iterative procedure. Our extensive experiment results show that TARA can be an effective algorithm for cluster analysis in a complex network. Moreover, the parallelized version of TARA makes it possible to perform efficiently at its tasks when applied to large complex networks.
Lun Hu, Keith C. C. Chan, Shengwu Xiong 0001
IEEE Trans. Knowl. Data Eng.1
2019 Learning from Deep Representations of Multiple Networks for Predicting Drug-Target Interactions
Pengwei Hu 0001, Zhu-Hong You, Shaochun Li, Keith C. C. Chan, Henry Leung 0001, Lun Hu
ICIC (2)7
2019 A Gated Recurrent Unit Model for Drug Repositioning by Combining Comprehensive Similarity Measures and Gaussian Interaction Profile Kernel
Zhu-Hong You, Liping Li 0003, Lun Hu, Leon Wong
ICIC (2)6
2019 A fast algorithm to identify coevolutionary patterns from protein sequences based on tree-based data structure
abstract
Knowing how proteins interact with each other are crucial to for us to understand the functional mechanisms of proteins. It is for this reason that the CoFex has been developed in attempts to predict protein-protein interactions (PPIs) computationally. However, the procedure of obtaining coevolutionary patterns adopted by CoFex is inefficient especially for large-scale prediction of PPIs, as it needs to traverse the entire protein sequence dataset once when computing the number of co-occurrences for each candidate of coevolutionary patterns. Hence, to improve the efficiency of CoFex, we propose a novel tree-based data structure, namely CF-Tree, to integrate with CoFex so that the running time of CoFex can be reduced by only traversing the sequence dataset once. The experiment results show that CF-Tree is a promising tree-based data structure to identify the coevolutionary patterns more efficiently from the sequence information of proteins in different species.
Lun Hu, Shicheng Yang
SMC1
2019 Efficiently Detecting Protein Complexes from Protein Interaction Networks via Alternating Direction Method of Multipliers
abstract
Protein complexes are crucial in improving our understanding of the mechanisms employed by proteins. Various computational algorithms have thus been proposed to detect protein complexes from protein interaction networks. However, given massive protein interactome data obtained by high-throughput technologies, existing algorithms, especially those with additionally consideration of biological information of proteins, either have low efficiency in performing their tasks or suffer from limited effectiveness. For addressing this issue, this work proposes to detect protein complexes from a protein interaction network with high efficiency and effectiveness. To do so, the original detection task is first formulated into an optimization problem according to the intuitive properties of protein complexes. After that, the framework of alternating direction method of multipliers is applied to decompose this optimization problem into several subtasks, which can be subsequently solved in a separate and parallel manner. An algorithm for implementing this solution is then developed. Experimental results on five large protein interaction networks demonstrated that compared to state-of-the-art protein complex detection algorithms, our algorithm outperformed them in terms of both effectiveness and efficiency. Moreover, as number of parallel processes increases, one can expect an even higher computational efficiency for the proposed algorithm with no compromise on effectiveness.
Lun Hu, Xing Liu 0002, Shengwu Xiong 0001, Xin Luo 0001
IEEE ACM Trans. Comput. Biol. Bioinform.1
2018 A gene-phenotype relationship extraction pipeline from the biomedical literature using a representation learning approach
abstract
Motivation: The fundamental challenge of modern genetic analysis is to establish gene-phenotype correlations that are often found in the large-scale publications. Because lexical features of gene are relatively regular in text, the main challenge of these relation extraction is phenotype recognition. Due to phenotypic descriptions are often study- or author-specific, few lexicon can be used to effectively identify the entire phenotypic expressions in text, especially for plants. Results: We have proposed a pipeline for extracting phenotype, gene and their relations from biomedical literature. Combined with abbreviation revision and sentence template extraction, we improved the unsupervised word-embedding-to-sentence-embedding cascaded approach as representation learning to recognize the various broad phenotypic information in literature. In addition, the dictionary- and rule-based method was applied for gene recognition. Finally, we integrated one of famous information extraction system OLLIE to identify gene-phenotype relations. To demonstrate the applicability of the pipeline, we established two types of comparison experiment using model organism Arabidopsis thaliana. In the comparison of state-of-the-art baselines, our approach obtained the best performance (F1-Measure of 66.83%). We also applied the pipeline to 481 full-articles from TAIR gene-phenotype manual relationship dataset to prove the validity. The results showed that our proposed pipeline can cover 70.94% of the original dataset and add 373 new relations to expand it. Availability and implementation: The source code is available at http://www.wutbiolab.cn: 82/Gene-Phenotype-Relation-Extraction-Pipeline.zip. Supplementary information: Supplementary data are available at Bioinformatics online.
Wenhui Xing, Junsheng Qi, Lin Li 0001, Yuhua Fu, Shengwu Xiong 0001, Lun Hu
Bioinform.8
2017 Bar charts detection and analysis in biomedical literature of PubMed Central
Xiaohan Yu 0001, Yangjing Gan, Tujin Zhu, Shengwu Xiong 0001, Lun Hu
AMIA7
2017 Discovering second-order sub-structure associations in drug molecules for side-effect prediction
abstract
Possible drug side-effects (SEs) are usually verified by many years of repeated clinical trials. Despite the effort, some drugs are still expected to cause adverse reactions in some patients. To better predict drug SEs without having to go through the laborious processes of testing and re-testing, machine learning (ML) techniques are more and more used to uncovered patterns in drug data for such purpose. Most existing such techniques are black-box techniques. Since correlations between sub-structures involving multiple variables may exist, these techniques may not always work well. For ML techniques to be effective, they should be accurate, efficient and the patterns they discover should be interpretable. Towards these goals, we have developed a second-order association discovering (SOAD) algorithm for SE prediction. Given a set of drug data for training, the SOAD algorithm can discover SO associations between multiple drug sub-structures and multiple SEs in drug data for the purpose of predicting the SEs. SOAD performs its tasks by first making use of a residual measure to test the significance of occurrence of a chemical sub-structure within a drug and the SE of the drug. Once an association is established between a sub-structure and a SE, we test if two or more such sub-structures are significantly associated with a SE. Based on such second-order associations, we derive from them a set of “informative” SO patterns so that the SEs of new unseen drugs can be predicted based on the frequency of appearance of such patterns. To ensure interpretability of the SE discovery process, we make use of the Bayesian to predict if certain SO relationship in a drug may be related to a certain side-effect. Based on the experimental results, SOAD is found to be very promising.
Pengwei Hu 0001, Keith C. C. Chan, Lun Hu, Henry Leung 0001
BIBM3
2017 Identifying overlapping protein complexes in yeast protein interaction network via fuzzy clustering
abstract
The problem of identifying protein complexes is of great significance for studying the protein mechanisms in different cellular systems. It is for this reason that many computational approaches have been proposed to solve the problem. Yet few of them have endeavored to discover overlapping protein complexes, which are crucial to improve the accuracy performance. Hence, in this paper, we explore the feasibility of making use of a fuzzy clustering approach to identify overlapping protein complexes in a natural manner. To do so, we first formulate the identification problem as an optimization problem by following certain intuitions and then develop an algorithm to solve it so that the memberships of each protein to different protein complexes can be optimized to eventually infer the protein complexes of interest. The experimental results on several yeast protein interaction networks show that our algorithm is promising in terms of accuracy.
Lun Hu, Shengwu Xiong 0001
FUZZ-IEEE1
2017 Extracting Coevolutionary Features from Protein Sequences for Predicting Protein-Protein Interactions
abstract
Knowing the ways proteins interact with each other are crucial to our understanding of the functional mechanisms of proteins. It is for this reason that different approaches have been developed in attempts to predict protein-protein interactions (PPIs) computationally. Among them, the sequence-based approaches are preferred to the others as they do not require any information about protein properties to perform their tasks. Instead, most sequence-based approaches make use of feature extraction methods to extract features directly from protein sequences so that for each protein sequence, we can construct a feature vector. The feature vectors of every pair of proteins are then concatenated to form two classes of interacting and non-interacting proteins. The prediction of whether or not two proteins interact with each other is then formulated as a classification problem. How accurate PPI predictions can be made therefore depends on how good the features are that can be extracted from the protein sequences to allow interacting or non-interacting to be best distinguished. To do so, instead of extracting such features from individual protein sequences independently of the other protein in the same pair, we propose to jointly consider features from both sequences in a protein pair during the feature extraction process through using a novel coevolutionary feature extraction approach called CoFex. Coevolutionary features extracted by CoFex refer to the covariations found at coevolving positions. Based on the presence and absence of these coevolutionary features in the sequences of two proteins, feature vectors can be composed for pairs of proteins rather than individual proteins. The experiment results show that CoFex is a promising feature extraction approach and can improve the performance of PPI prediction.
Lun Hu, Keith C. C. Chan
IEEE ACM Trans. Comput. Biol. Bioinform.1
2016 InDel marker detection by integration of multiple softwares using machine learning techniques
abstract
BACKGROUND: In the biological experiments of soybean species, molecular markers are widely used to verify the soybean genome or construct its genetic map. Among a variety of molecular markers, insertions and deletions (InDels) are preferred with the advantages of wide distribution and high density at the whole-genome level. Hence, the problem of detecting InDels based on next-generation sequencing data is of great importance for the design of InDel markers. To tackle it, this paper integrated machine learning techniques with existing software and developed two algorithms for InDel detection, one is the best F-score method (BF-M) and the other is the Support Vector Machine (SVM) method (SVM-M), which is based on the classical SVM model. RESULTS: The experimental results show that the performance of BF-M was promising as indicated by the high precision and recall scores, whereas SVM-M yielded the best performance in terms of recall and F-score. Moreover, based on the InDel markers detected by SVM-M from soybeans that were collected from 56 different regions, highly polymorphic loci were selected to construct an InDel marker database for soybean. CONCLUSIONS: Compared to existing software tools, the two algorithms proposed in this work produced substantially higher precision and recall scores, and remained stable in various types of genomic regions. Moreover, based on SVM-M, we have constructed a database for soybean InDel markers and published it for academic research.
Jianqiu Yang, Xinyi Shi, Lun Hu, Daipeng Luo, Shengwu Xiong 0001, Fanjing Kong, Baohui Liu
BMC Bioinform.3
2016 Fuzzy Clustering in a Complex Network Based on Content Relevance and Link Structures
abstract
Many real-world problems can be represented as complex networks with nodes representing different objects and links between nodes representing relationships between objects. As different attributes can be considered as associating with different objects, other than nontrivial link structures, complex networks also contain rich content information, and it can be a big challenge to find interesting clusters in such networks by fully exploiting the knowledge of both content and link information in them. Although some attempts have been made to tackle this clustering problem, few of them have considered the feasibility of identifying clusters in complex networks using a fuzzy-based clustering approach. We believe that, if the degree of membership to a cluster that a node belongs to can be considered, we will be able to better identify clusters in complex networks, as we may be able to identify overlapping clusters. In this paper, we, therefore, propose a fuzzy-based clustering algorithm for this task. The algorithm, which we call Fuzzy Clustering Algorithm for Complex Networks (FCAN), can discover clusters by taking into consideration both link and content information. It does so by first processing the content information by introducing a measure to quantify the relevance of contents between each pair of nodes within the network. It then proceeds to leverage the link information in the clustering process by considering a measure of cluster density. Based on these measures, FCAN identifies fuzzy clusters that are more densely connected and more highly relevant in their contents to optimize the degrees of memberships of each node belonging to different clusters. The performance of FCAN has been evaluated with several synthetic and real datasets involving those of document classification and social community detection. The results show that, in terms of accuracy, computation efficiency, and scalability, FCAN can be a very promising approach.
Lun Hu, Keith C. C. Chan
IEEE Trans. Fuzzy Syst.1
2015 A density-based clustering approach for identifying overlapping protein complexes with functional preferences
abstract
BACKGROUND: Identifying protein complexes is an essential task for understanding the mechanisms of proteins in cells. Many computational approaches have thus been developed to identify protein complexes in protein-protein interaction (PPI) networks. Regarding the information that can be adopted by computational approaches to identify protein complexes, in addition to the graph topology of PPI network, the consideration of functional information of proteins has been becoming popular recently. Relevant approaches perform their tasks by relying on the idea that proteins in the same protein complex may be associated with similar functional information. However, we note from our previous researches that for most protein complexes their proteins are only similar in specific subsets of categories of functional information instead of the entire set. Hence, if the preference of each functional category can also be taken into account when identifying protein complexes, the accuracy will be improved. RESULTS: To implement the idea, we first introduce a preference vector for each of proteins to quantitatively indicate the preference of each functional category when deciding the protein complex this protein belongs to. Integrating functional preferences of proteins and the graph topology of PPI network, we formulate the problem of identifying protein complexes into a constrained optimization problem, and we propose the approach DCAFP to address it. For performance evaluation, we have conducted extensive experiments with several PPI networks from the species of Saccharomyces cerevisiae and Human and also compared DCAFP with state-of-the-art approaches in the identification of protein complexes. The experimental results show that considering the integration of functional preferences and dense structures improved the performance of identifying protein complexes, as DCAFP outperformed the other approaches for most of PPI networks based on the assessments of independent measures of f-measure, Accuracy and Maximum Matching Rate. Furthermore, the function enrichment experiments indicated that DCAFP identified more protein complexes with functional significance when compared with approaches, such as PCIA, that also utilize the functional information. CONCLUSIONS: According to the promising performance of DCAFP, the integration of functional preferences and dense structures has made it possible to identify protein complexes more accurately and significantly.
Lun Hu, Keith C. C. Chan
BMC Bioinform.1