VLDB 2026 Research / reviewers in the wild / expert
Zhijian Huang 0001
dblp:95/2725-1
· DBLP profile ↗
28ranked-venue papers
7as first author
28since 2021 · last 2026
0009-0000-1949-167XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 26 · 7 first-author · 26 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Structure-Property Alignment for Data-Efficient Molecular Generation and EditingabstractProperty-constrained molecular generation and editing are crucial in AI-driven drug discovery but remain hindered by two factors: (i) capturing the complex relationships between molecular structures and multiple properties remains challenging, and (ii) the narrow coverage and incomplete annotations of molecular properties weaken the effectiveness of property-based models. To tackle these limitations, we propose HSPAG, a data-efficient framework featuring hierarchical structure–property alignment. By treating SMILES and molecular properties as complementary modalities, the model learns their relationships at atom, substructure, and whole-molecule levels. Moreover, we select representative samples through scaffold clustering and hard samples via an auxiliary variational auto-encoder (VAE), substantially reducing the required pre-training data. In addition, we incorporate a property relevance-aware masking mechanism and diversified perturbation strategies to enhance generation quality under sparse annotations. Experiments demonstrate that HSPAG captures fine-grained structure–property relationships and supports controllable generation under multiple property constraints. Two real-world case studies further validate the editing capabilities of HSPAG. Ziyu Fan, Zhijian Huang 0001, Yahan Li, Yunliang Wang, Zeyu Zhong, Shuhong Liu, Shuning Yang, Shangqian Wu, Min Wu 0008, Lei Deng 0002 |
AAAI | 2 |
| 2026 | DeepLMI: deep feature mining with a globally enhanced graph convolutional network for robust lncRNA-miRNA interaction predictionabstractMOTIVATION: Interactions between long noncoding RNAs (lncRNAs) and microRNAs (miRNAs) play pivotal roles in gene regulation and disease progression, notably through mechanisms such as competitive miRNA sponging. Accurate identification of lncRNA-miRNA interactions is therefore essential for understanding disease mechanisms and discovering therapeutic targets. However, current knowledge is largely derived from labor-intensive and costly biological experiments, underscoring the need for reliable computational approaches. RESULTS: We propose DeepLMI, a novel deep learning framework for lncRNA-miRNA interaction prediction that integrates deep feature mining with a globally enhanced graph convolutional network. To effectively capture the distinct properties of lncRNAs and miRNAs, DeepLMI employs specialized feature extraction modules: for lncRNAs, we combine sequence pretraining with self-attention mechanisms to learn multiscale semantic representations; for miRNAs, we fuse heterogeneous features through a graph convolutional encoder. To further address the sparsity and structural complexity of known RNA interaction networks, we design a Global-Enhanced Graph Convolutional Network that jointly models local neighborhood information and global topological signals. The embeddings learned for lncRNAs and miRNAs are then integrated to infer interaction probabilities. Extensive experiments across multiple datasets and evaluation settings demonstrate that DeepLMI consistently outperforms existing state-of-the-art methods and exhibits strong robustness, highlighting its potential as a valuable tool for RNA interaction analysis and disease research. AVAILABILITY AND IMPLEMENTATION: The codes and data are publicly available at https://github.com/Hhhzj-7/DeepLMI. Zhijian Huang 0001, Xianshu Wang, Junheng Wang, Yuanpeng Zhang 0004, Min Wu 0008, Lei Deng 0002 |
Bioinform. | 1 |
| 2026 | MM-GHiNet: A multimodal graph hierarchical network for brain disease diagnosis
Zhijian Huang 0001, Lei Deng 0002 |
Expert Syst. Appl. | 3 |
| 2025 | DSSA: Dual-Stream Synthetic Accessibility Framework for Organic CompoundsabstractSynthetic accessibility prediction remains a key bottleneck in AI-driven drug discovery, as a large proportion of computationally generated molecules prove infeasible to synthesize. Existing approaches often struggle to distinguish structurally similar compounds with divergent synthetic profiles, limiting their usefulness in practical design pipelines. We present DSSA (Dual-Stream Synthetic Accessibility), a novel architecture that integrates Graph Attention Networks for molecular topology with a Bidirectional GRU for sequential SMILES representations through transformer-based cross-modal fusion. DSSA effectively captures both local structural complexity and global sequential patterns, enabling robust generalization across diverse molecular classes. Cross-modal attention analysis reveals that the model dynamically adapts to molecular complexity, with graph-dominant attention highlighting stereochemical constraints that sequenceonly models overlook. Ablation studies further confirm that cross-modal fusion is essential for achieving balanced structural and sequential reasoning. Collectively, DSSA bridges the gap between computational molecular generation and real-world synthetic feasibility, offering a reliable foundation for data-driven molecular design. Web tool: http://dssa.denglab.org/; code/data: https://github.com/Q-Aljanabi/DSSA. Qahtan Adnan Aljanabi, Zhijian Huang 0001, Gebremedhin Assefa Girmay, Zhengkang Wang, Yuanpeng Zhang 0004, Lei Deng 0002 |
BIBM | 2 |
| 2025 | ASRSMA: Atomic-Scale and Structure-Based Modeling for RNA-Small Molecule Binding Affinity Prediction via Contrastive PretrainingabstractRNA is intricately involved in aberrant cellular functions and a wide range of disease processes, playing pivotal roles in gene regulation, viral replication, and innate immunity. Consequently, it has emerged as a highly promising therapeutic target. To accelerate the discovery of such drugs, it is essential to develop an effective computational method for predicting RNA-small molecule affinity. Therefore, we propose ASRSMA as an atomiclevel structure-aware model with contrastive pre-training for RNA-small molecule binding affinity prediction. ASRSMA represents RNA and small molecules at atomic resolution, enabling the capture of fine-grained structural features. It incorporates an atom-pair encoding module and an inter-molecular interaction module to model intra- and inter-molecular interactions. In addition, a self-supervised contrastive learning strategy is employed during pre-training, maximising the similarity between different views of the same complex while minimising similarities across different complexes, thereby yielding deep representations with strong generalisation capacity. Experimental results demonstrate that ASRSMA significantly outperforms state-of-the-art baseline models in predicting RNA-small molecule binding affinities. Yahan Li, Zhijian Huang 0001, Yucheng Wang 0001, Min Wu 0008, Lei Deng 0002 |
BIBM | 2 |
| 2025 | GroupTransUNet: Group Transformer UNet for Medical Image SegmentationabstractAccurate medical image segmentation is a pivotal task in medical image analysis, serving as the foundation for extracting critical clinical information and directly influencing the precision of disease identification, diagnosis, and treatment decision-making. Although Convolutional Neural Networks (CNNs) and Transformers have demonstrated remarkable performance in medical image segmentation, their sensitivity to variations in target size and morphology remains insufficient, resulting in limited capability of single-scale features to comprehensively capture multi-target details. Current approaches also face bottlenecks, including restricted global context modeling capabilities, high computational complexity, and inefficient multilevel feature fusion. To address these challenges in medical image segmentation, this study proposes an innovative architecture named GroupTransUNet. The model implements dual enhancements based on the UNet framework: (i) A novel Next-Generation Transformer Block (NGTB) is designed for the bottleneck layer, which organically integrates Efficient Multi-Head Self-Attention (E-MHSA) with Convolution-Enhanced Multi-Layer Perceptron (MLP) through a complementary mechanism to simultaneously enhance global semantic understanding and local detail characterization. (ii) The Grouped Feature Fusion Module (GFFM) is introduced in skip connections, employing grouped convolutions and multi-dilation rate strategies to construct multi-scale contextual receptive fields while preserving detailed integrity, thereby significantly improving feature fusion efficiency. Experimental results on the Synapse and ACDC datasets validate the effectiveness of our proposed GroupTransUNet. The source code can be obtained at https://anonymous.4open.science/r/GroupTransUNetA3D6. Yunliang Wang, Ziyu Fan, Zhijian Huang 0001, Jinmiao Song, Lei Deng 0002 |
BIBM | 3 |
| 2025 | MVFDSP: A Multi-View Fusion Framework for Drug Side-Effect Frequency PredictionabstractAccurate prediction of drug side effect frequencies is critical for drug safety evaluation and clinical decision-making. Current methods primarily emphasize the associations between drugs and side effects, yet they often neglect the underlying structural and semantic features of both, which limits further advancements in prediction accuracy. In this study, we propose a novel multi-view fusion framework, MVFDSP, which integrates pre-trained molecular representation of 1D and 2D views with graph-based side effect information for side effect frequency prediction. Firstly, we obtain both 1D and 2D molecular representations from the pretrained molecular language model, and combine them using an adaptive fusion strategy. Subsequently, we construct a similarity network based on the side effect frequency matrix using K-Nearest Neighbors (KNN), and incorporate semantic embeddings derived from the terminology system of MedDRA to construct a side effect information graph. A multi-head graph attention network is then employed to capture the multi-dimensional information within this graph, allowing the model to attend to diverse aspects of the semantic and structural relationships among side effects. The final frequency prediction matrix is derived from the inner product between the learned drug and side effect embeddings. Experimental results on the SIDER 4.1 dataset demonstrate that MVFDSP outperforms existing methods, highlighting its effectiveness in capturing complex relationships of drugs and side effects. The code and data are available at https://github.com/Sonder-Echo/MVFDSP. Zhengkang Wang, Zhijian Huang 0001, Yurong Qian, Yuanpeng Zhang 0004, Yahan Li, Qahtan Adnan Aljanabi, Jinmiao Song, Lei Deng 0002 |
BIBM | 2 |
| 2025 | MVRBind: multi-view learning for RNA-small molecule binding site predictionabstractRNA plays a critical role in cellular processes, and its dysregulation is linked to many diseases, positioning RNA-targeted drugs as an important area of research. Accurate prediction of RNA-small molecule binding sites is crucial for advancing RNA-targeted therapies. Although deep learning has shown promise in this area, challenges remain in integrating and processing multi-dimensional data, such as RNA sequences and structural features, particularly given the inherent flexibility of RNA structures. In this study, we present MVRBind, a multi-view graph convolutional network designed to predict RNA-small molecule binding sites. MVRBind generates feature representations of RNA nucleotides across different structural levels. To effectively integrate these features, we developed a multi-view feature fusion module that constructs graphs based on RNA's primary, secondary, and tertiary structural views, enabling the model to capture diverse aspects of RNA structure. In addition, we fuse embeddings from multi-scale to obtain a comprehensive representation of RNA nucleotides, which is then used to predict RNA-small molecule binding sites. Extensive experiments demonstrate that MVRBind consistently outperforms baseline methods in various experimental settings. Our MVRBind shows exceptional performance in predicting binding sites for both the holo and apo forms of RNA, even when RNA adopts multiple conformations. These results suggest that MVRBind offers a robust model for structure-based RNA analysis, contributing toward accurate prediction and analysis of RNA-small molecule binding sites. All datasets and resource codes are available at https://github.com/cschen-y/MVRBind. Zhijian Huang 0001, Yucheng Wang 0001, Yahan Li, Yaw Sing Tan, Lei Deng 0002, Min Wu 0008 |
Briefings Bioinform. | 2 |
| 2025 | DeepHeteroCDA: circRNA-drug sensitivity associations prediction via multi-scale heterogeneous network and graph attention mechanismabstractDrug sensitivity is essential for identifying effective treatments. Meanwhile, circular RNA (circRNA) has potential in disease research and therapy. Uncovering the associations between circRNAs and cellular drug sensitivity is crucial for understanding drug response and resistance mechanisms. In this study, we proposed DeepHeteroCDA, a novel circRNA-drug sensitivity association prediction method based on multi-scale heterogeneous network and graph attention mechanism. We first constructed a heterogeneous graph based on drug-drug similarity, circRNA-circRNA similarity, and known circRNA-drug sensitivity associations. Then, we embedded the 2D structure of drugs into the circRNA-drug sensitivity heterogeneous graph and use graph convolutional networks (GCN) to extract fine-grained embeddings of drug. Finally, by simultaneously updating graph attention network for processing heterogeneous networks and GCN for processing drug structures, we constructed a multi-scale heterogeneous network and use a fully connected layer to predict the circRNA-drug sensitivity associations. Extensive experimental results highlight the superior of DeepHeteroCDA. The visualization experiment shows that DeepHeteroCDA can effectively extract the association information. The case studies demonstrated the effectiveness of our model in identifying potential circRNA-drug sensitivity associations. The source code and dataset are available at https://github.com/Hhhzj-7/DeepHeteroCDA. Zhijian Huang 0001, Xiaojun Xiao, Ziyu Fan, Yuanpeng Zhang 0004, Lei Deng 0002 |
Briefings Bioinform. | 1 |
| 2025 | Contrastive hypergraph collaborative filtering for transfer RNA-disease association predictionabstractTransfer RNAs (tRNAs) play critical roles in the process of protein synthesis by decoding messenger RNA codons into amino acids, which is essential for cellular function across various biological pathways and for maintaining metabolic homeostasis. Available evidence implicates that tRNAs are involved in the progression of diverse diseases, underscoring the importance of accurately predicting tRNA-disease associations to understand disease mechanisms and support precision medicine. However, existing methods often struggle with the complexity and heterogeneity inherent in these associations. To address these challenges, we introduce contrastive hypergraph collaborative filtering (CoHGCL), a prediction framework that integrates hypergraph contrastive learning with collaborative filtering. CoHGCL employs graph attention networks to capture local structural features and random walk with restart algorithms to encode global topological patterns. Subsequently, a node-level contrastive learning mechanism alternates between standard graph and hypergraph representations to enhance multiview feature embeddings. These enriched representations are integrated by a collaborative filtering approach through the utilization of generalized matrix factorization for modeling linear associations and multilayer perceptrons for capturing nonlinear interactions. Extensive experimental results on five-fold cross-validation demonstrate that CoHGCL achieves superior performance compared to existing methods, with an area under the receiver operating characteristic curve of 0.9623, area under the precision-recall curve of 0.9430, outperforming all baselines across all metrics. Furthermore, case studies further confirm CoHGCL's effectiveness in discovering novel and biologically meaningful tRNA-disease associations. The source code and datasets are publicly available at https://github.com/Ouyang-cmd/CoHGCL. Tianxiang Ouyang, Yuanpeng Zhang 0004, Zhijian Huang 0001, Lei Deng 0002 |
Briefings Bioinform. | 3 |
| 2025 | Precise prediction of hotspot residues in protein-RNA complexes using graph attention networks and pretrained protein language modelsabstractMOTIVATION: Protein-RNA interactions play a pivotal role in biological processes and disease mechanisms, with hotspot residues being critical for targeted drug design. Traditional experimental methods for identifying hotspot residues are often inefficient and expensive. Moreover, many existing prediction methods rely heavily on high-resolution structural data, which may not always be available. Consequently, there is an urgent need for an accurate and efficient sequence-based computational approach for predicting hotspot residues in protein-RNA complexes. RESULTS: In this study, we introduce DeepHotResi, a sequence-based computational method designed to predict hotspot residues in protein-RNA complexes. DeepHotResi leverages a pretrained protein language model to predict protein structure and generate an amino acid contact map. To enhance feature representation, DeepHotResi integrates the Squeeze-and-Excitation (SE) module, which processes diverse amino acid-level features. Next, it constructs an amino acid feature network from the contact map and SE-module-derived features. Finally, DeepHotResi employs a graph attention network to model hotspot residue prediction as a graph node classification task. Experimental results demonstrate that DeepHotResi outperforms state-of-the-art methods, effectively identifying hotspot residues in protein-RNA complexes with superior accuracy on the test set. AVAILABILITY AND IMPLEMENTATION: The source code and dataset are available at https://github.com/Q1DT/DeepHotResi. Zhijian Huang 0001, Yuanpeng Zhang 0004, Ziyu Fan, Yuting Kong, Lei Deng 0002 |
Bioinform. | 3 |
| 2025 | HLN-DDI: hierarchical molecular representation learning with co-attention mechanism for drug-drug interaction predictionabstractBACKGROUND: Accurate identification of drug-drug interactions (DDIs) is critical in pharmacology, as DDIs can either enhance therapeutic efficacy or trigger adverse reactions when multiple medications are administered concurrently. Traditional methods for identifying DDIs are labor-intensive and time-consuming, prompting the development of computational alternatives. However, existing computational approaches frequently encounter challenges related to interpretability and struggle to effectively capture the complex, multi-level structures inherent in drug molecules. Specifically, they often fail to adequately analyze substructural components and neglect interactions across hierarchical structural levels, resulting in incomplete molecular representations. RESULTS: In this study, we propose a Hierarchical Learning Network with a co-attention mechanism tailored to molecular structure representation for predicting DDIs, named HLN-DDI. The proposed method advances existing approaches by explicitly encoding motif-level structures and capturing hierarchical molecular representations at atom-level, motif-level, and whole-molecule scales. These hierarchical representations are integrated using a co-attention mechanism and combined with interaction-type information to enhance predictive performance. Comprehensive evaluations demonstrate that HLN-DDI significantly outperforms state-of-the-art methods across multiple benchmark datasets, achieving over 98% accuracy under transductive scenarios and surpassing 99% on various evaluation metrics. Moreover, HLN-DDI achieves a notable accuracy improvement of 2.75% in predicting DDIs involving unseen drugs. Practical assessments with real-world DDI scenarios further validate the efficacy and utility of our proposed model. CONCLUSION: By leveraging hierarchical molecular structures and employing a co-attention mechanism to effectively integrate multi-level representations, HLN-DDI generates comprehensive and precise drug representations, leading to substantially improved predictions of potential drug-drug interactions. Lei Deng 0002, Zhijian Huang 0001 |
BMC Bioinform. | 3 |
| 2025 | CircGO: Predicting circRNA Functions Through Self-Supervised Learning of Heterogeneous NetworksabstractCircular RNAs (circRNAs), a class of non-coding RNAs characterized by their covalently closed loop structures, play active roles in diverse physiological processes through interactions with biological macromolecules. Despite the growing discovery of circRNAs enabled by high-throughput technologies, their functional annotations remain largely unexplored. This highlights the need for automated batch annotation methods to unveil the functional roles of circRNAs. In this study, we present a novel approach for predicting Gene Ontology (GO) functions associated with circRNA by leveraging self-supervised pre-training on circRNA-protein heterogeneous network. First, we construct the heterogeneous network by combining circRNA co-expression data, circRNA-protein association data, and protein-protein interaction (PPI) data. Second, we initialize the features and pseudo-labels for nodes using three graph processing methods including walking, aggregation and clustering. The initialized node features and pseudo labels, combined with protein GO annotations, are employed for heterogeneous graph pre-training. During the pre-training, the node features are learned using a heterogeneous graph attention network and the pseudo-labels are updated using the label propagation algorithm (LPA) with an attention mechanism. Finally, the initial node features are combined with those learned during pre-training to predict circRNA GO terms. Evaluation results on the independent test set reveal the superior performance of our method compared with existing approaches. Furthermore, our analysis underscores the importance of network structure and initialization strategies, highlighting the potential benefits of incorporating additional heterogeneous information and association networks. Zhijian Huang 0001, Rongtao Zheng, Min Wu 0008, Lei Deng 0002 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2025 | Enhancing Predictions of Drug Solubility Through Multidimensional Structural Characterization ExploitationabstractSolubility is not only a significant physical property of molecules but also a vital factor in small-molecule drug development. Determining drug solubility demands stringent equipment, controlled environments, and substantial human and material resources. The accurate prediction of drug solubility using computational methods has long been a goal for researchers. In this study, we introduce MSCSol, a solubility prediction model that integrates multidimensional molecular structure information. We incorporate a graph neural network with geometric vector perceptrons (GVP-GNN) to encode 3D molecular structures, representing spatial arrangement and orientation of atoms, as well as atomic sequences and interactions. We also employ Selective Kernel Convolution combined with Global and Local attention mechanisms to capture molecular features context at different scales. Additionally, various descriptors are calculated to enrich the molecular representation. For the 2D and 3D structural data of molecules, we design different data augmentation strategies to enhance generalization ability and prevent the model from learning irrelevant information. Extensive experiments on benchmark and independent datasets demonstrate MSCSol's superior performance. Ablation studies further confirm the effectiveness of different modules. Interpretability analysis highlights the importance of various atomic groups and substructures for solubility and verifies that our model effectively captures functional molecular structures and higher-order knowledge. Ziyu Fan, Zhijian Huang 0001, Lei Deng 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | AGCLNDA: Enhancing the Prediction of ncRNA-Drug Resistance Association Using Adaptive Graph Contrastive LearningabstractNon-coding RNAs (ncRNAs), which do not encode proteins, have been implicated in chemotherapy resistance in cancer treatment. Given the high costs and time requirements of traditional biological experiments, there is an increasing need for computational models to predict ncRNA-drug resistance associations. In this study, we introduce AGCLNDA, an adaptive contrastive learning method designed to uncover these associations. AGCLNDA begins by constructing a bipartite graph from existing ncRNA-drug resistance data. It then utilizes a light graph convolutional network (LightGCN) to learn vector representations for both ncRNAs and drugs. The method assesses resistance association scores through the inner product of these vectors. To tackle data sparsity and noise, AGCLNDA incorporates learnable augmented view generators and denoised view generators, which provide contrastive views for enhanced data augmentation. Comparative experiments demonstrate that AGCLNDA outperforms five other advanced methods. Case studies further validate AGCLNDA as an effective tool for predicting ncRNA-drug resistance associations. Yanhao Fan, Che Zhang, Zhijian Huang 0001, Lei Deng 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Enhancing miRNA-Disease Prediction with Attention-Based Multi-View Contrastive LearningabstractMicroRNAs (miRNAs) play a critical role in gene expression regulation, and their dysregulation is implicated in diseases such as cancer and cardiovascular disorders. Identifying miRNA-disease associations is therefore crucial for advancing prevention and treatment strategies for these conditions. We introduce MCLAMDA (Multi-view Contrastive Learning with Attention Mechanism for miRNA-Disease Association Prediction), a novel method inspired by contrastive learning and attention mechanisms. MCLAMDA constructs diverse similarity networks for miRNAs and diseases and employs both topological and feature contrastive learning to capture comprehensive node information. An attention mechanism further enhances the model by weighting and aggregating node features based on the relative importance of their neighboring nodes. Our approach outperforms five state-of-the-art methods in predictive accuracy. Case studies provide additional support for MCLAMDA’s robustness and reliable predictive performance. The code and data for MCLAMDA are available at https://github.com/Ouyang-cmd/MCLAMDA. Tianxiang Ouyang, Zhijian Huang 0001, Lei Deng 0002 |
BIBM | 2 |
| 2024 | DTA-Net: Dual-Task Attention Network for Medical Image Segmentation and ClassificationabstractMedical image segmentation and classification are fundamental tasks in computer-aided diagnosis, where accurate segmentation plays a key role in identifying disease-related features and regions of interest, thus aiding subsequent classification. In this paper, we propose a novel Dual-Task Attention Network (DTA-Net), which simultaneously generates high-quality segmentation masks and performs image classification by incorporating Electronic Medical Records (EMR). The DTA-Net architecture introduces an innovative Shared Proxy Attention Module (SPAM), which leverages a shared mapping function to effectively encode the attention weights of both queries and keys for spatial and channel attention. This dual attention mechanism facilitates complementary learning and interdependence between spatial and channel features. Notably, the proxy projection mechanism within SPAM significantly reduces the computational complexity of spatial attention to a linear level. We validated our approach on three datasets: MSD Prostate for prostate segmentation, the RICORD dataset for lung segmentation, and the iCTCF COVID-19 dataset for both segmentation and classification tasks. Experimental results demonstrate that the proposed network achieves promising performance across all datasets. The source code is available at: https://github.com/HaifengQi/DTA-Net. Haifeng Qi, Zhijian Huang 0001, Lei Deng 0002 |
BIBM | 3 |
| 2024 | GATBind: Accurate protein-RNA binding sites prediction via graph attention networks with pre-trained language modelabstractProtein-RNA interactions play a pivotal role in various biological processes, making them essential for discovering novel therapeutic targets. Understanding these interactions is crucial for identifying potential drug targets and designing effective therapeutics. However, traditional experimental approaches are costly and time-consuming, highlighting the need for accurate and efficient computational methods to predict RNA-binding sites. In this study, we propose a novel method called GATBind, designed to predict RNA-binding sites on proteins. GATBind introduces a local environment-aware module that captures the local structural information surrounding target residues. Additionally, we incorporate sequence and structural features, including PSSM, HMM, DSSP, and sequence embeddings generated by the pre-trained protein language model ESMFold-2. These features are integrated and processed through a Graph Attention Network (GAT) module to predict RNA-binding sites accurately. Experimental results demonstrate that GATBind outperforms existing state-of-the-art methods, and case studies further validate its effectiveness in predicting RNA-binding sites. Shangqian Wu, Zhijian Huang 0001, Lei Deng 0002 |
BIBM | 4 |
| 2024 | SGCLDGA: unveiling drug-gene associations through simple graph contrastive learningabstractDrug repurposing offers a viable strategy for discovering new drugs and therapeutic targets through the analysis of drug-gene interactions. However, traditional experimental methods are plagued by their costliness and inefficiency. Despite graph convolutional network (GCN)-based models' state-of-the-art performance in prediction, their reliance on supervised learning makes them vulnerable to data sparsity, a common challenge in drug discovery, further complicating model development. In this study, we propose SGCLDGA, a novel computational model leveraging graph neural networks and contrastive learning to predict unknown drug-gene associations. SGCLDGA employs GCNs to extract vector representations of drugs and genes from the original bipartite graph. Subsequently, singular value decomposition (SVD) is employed to enhance the graph and generate multiple views. The model performs contrastive learning across these views, optimizing vector representations through a contrastive loss function to better distinguish positive and negative samples. The final step involves utilizing inner product calculations to determine association scores between drugs and genes. Experimental results on the DGIdb4.0 dataset demonstrate SGCLDGA's superior performance compared with six state-of-the-art methods. Ablation studies and case analyses validate the significance of contrastive learning and SVD, highlighting SGCLDGA's potential in discovering new drug-gene associations. The code and dataset for SGCLDGA are freely available at https://github.com/one-melon/SGCLDGA. Yanhao Fan, Che Zhang, Zhijian Huang 0001, Jiameng Xue, Lei Deng 0002 |
Briefings Bioinform. | 4 |
| 2024 | MolMVC: Enhancing molecular representations for drug-related tasks through multi-view contrastive learningabstractMOTIVATION: Effective molecular representation is critical in drug development. The complex nature of molecules demands comprehensive multi-view representations, considering 1D, 2D, and 3D aspects, to capture diverse perspectives. Obtaining representations that encompass these varied structures is crucial for a holistic understanding of molecules in drug-related contexts. RESULTS: In this study, we introduce an innovative multi-view contrastive learning framework for molecular representation, denoted as MolMVC. Initially, we use a Transformer encoder to capture 1D sequence information and a Graph Transformer to encode the intricate 2D and 3D structural details of molecules. Our approach incorporates a novel attention-guided augmentation scheme, leveraging prior knowledge to create positive samples tailored to different molecular data views. To align multi-view molecular positive samples effectively in latent space, we introduce an adaptive multi-view contrastive loss (AMCLoss). In particular, we calculate AMCLoss at various levels within the model to effectively capture the hierarchical nature of the molecular information. Eventually, we pre-train the encoders via minimizing AMCLoss to obtain the molecular representation, which can be used for various down-stream tasks. In our experiments, we evaluate the performance of our MolMVC on multiple tasks, including molecular property prediction (MPP), drug-target binding affinity (DTA) prediction and cancer drug response (CDR) prediction. The results demonstrate that the molecular representation learned by our MolMVC can enhance the predictive accuracy on these tasks and also reduce the computational costs. Furthermore, we showcase MolMVC's efficacy in drug repositioning across a spectrum of drug-related applications. AVAILABILITY AND IMPLEMENTATION: The code and pre-trained model are publicly available at https://github.com/Hhhzj-7/MolMVC. Zhijian Huang 0001, Ziyu Fan, Min Wu 0008, Lei Deng 0002 |
Bioinform. | 1 |
| 2024 | DeepRSMA: a cross-fusion-based deep learning method for RNA-small molecule binding affinity predictionabstractMOTIVATION: RNA is implicated in numerous aberrant cellular functions and disease progressions, highlighting the crucial importance of RNA-targeted drugs. To accelerate the discovery of such drugs, it is essential to develop an effective computational method for predicting RNA-small molecule affinity (RSMA). Recently, deep learning-based computational methods have been promising due to their powerful nonlinear modeling ability. However, the leveraging of advanced deep learning methods to mine the diverse information of RNAs, small molecules, and their interaction still remains a great challenge. RESULTS: In this study, we present DeepRSMA, an innovative cross-attention-based deep learning method for RSMA prediction. To effectively capture fine-grained features from RNA and small molecules, we developed nucleotide-level and atomic-level feature extraction modules for RNA and small molecules, respectively. Additionally, we incorporated both sequence and graph views into these modules to capture features from multiple perspectives. Moreover, a transformer-based cross-fusion module is introduced to learn the general patterns of interactions between RNAs and small molecules. To achieve effective RSMA prediction, we integrated the RNA and small molecule representations from the feature extraction and cross-fusion modules. Our results show that DeepRSMA outperforms baseline methods in multiple test settings. The interpretability analysis and the case study on spinal muscular atrophy demonstrate that DeepRSMA has the potential to guide RNA-targeted drug design. AVAILABILITY AND IMPLEMENTATION: The codes and data are publicly available at https://github.com/Hhhzj-7/DeepRSMA. Zhijian Huang 0001, Yucheng Wang 0001, Yaw Sing Tan, Lei Deng 0002, Min Wu 0008 |
Bioinform. | 1 |
| 2023 | Enhancing Protein Solubility Prediction through Pre-trained Language Models and Graph Convolutional Neural NetworksabstractAchieving optimal protein solubility is pivotal for efficient high-throughput purification, especially in industrial settings. However, conventional experimental techniques for assessing protein solubility in such contexts are not only costly but also time-intensive. Currently, numerous methods are available for predicting protein solubility, yet their effectiveness remains limited. Most of these approaches are predominantly sequence-based, failing to harness the invaluable structural insights inherent in proteins. Addressing these limitations, we introduce PPSol, an innovative protein solubility prediction methodology. Operating on protein sequences, PPSol employs ESM2 to predict protein contact maps, forming the basis for constructing protein graphs. Subsequently, well-established techniques are employed to predict protein feature representations as node features, including the utilization of the Position-Specific Scoring Matrix (PSSM). The resulting graph is fed into a graph convolutional neural network (GCN), enabling the acquisition of spatial structural information from proteins. Concurrently, ESM2-generated features undergo dimensional reduction via fully connected layers, integrating into every layer of the GCN for precise protein solubility prediction. Our approach excels through the fusion of pre-trained protein language models and GCNs, surpassing existing methodologies. Notably, PPSol attains state-of-the-art performance, showcasing a remarkable 2.8% enhancement in AUROC performance com-pared to prior strategies. Yurong Qian, Zhijian Huang 0001, Xiaojun Xiao, Lei Deng 0002 |
BIBM | 3 |
| 2023 | TGC-ARG: Predicting Antibiotic Resistance through Transformer-based Modeling and Contrastive LearningabstractThe escalating severity of antibiotic resistance poses substantial challenges across diverse sectors, encompassing everyday life, agriculture, and clinical medical interventions. Conventional methods for investigating antibiotic resistance genes (ARGs), such as culture-based techniques and whole-genome sequencing, often suffer from demands of time, labor, and limited accuracy. Moreover, the fragmented nature of existing datasets hampers a comprehensive analysis of antibiotic resistance gene sequences. In this study, we introduce an innovative computational framework known as TGC-ARG, designed to predict potential ARGs. TGC-ARG harnesses protein sequences as input, retrieves protein structures through SCRATCH-1D, and employs a feature extraction module to deduce feature representations for both protein sequences and structures. Subsequently, we integrate a siamese network to establish a contrastive learning paradigm, thus augmenting the model’s representational capabilities. The resultant sequence embeddings and structure embeddings are merged and directed into a Multilayer Perceptron (MLP) for predicting ARG presence. To assess the performance, we curate a pioneering publicly available dataset named ARSS (Antibiotic Resistance Sequence Statistics). Our extensive comparative experimental outcomes underscore the superiority of our approach over the current state-of-the-art (SOTA) methodology. Furthermore, through comprehensive case analyses, we demonstrate the efficacy of our approach in predicting potential ARGs. The dataset and source code are accessible at https://github.com/angel1gel/TGC-ARG. Yihan Dong, Zhijian Huang 0001, Lei Deng 0002 |
BIBM | 3 |
| 2023 | Predicting Associations between circRNAs and Drug Sensitivity using Heterogeneous Graphs and Graph Attention NetworksabstractDiscovering associations between circular RNAs (circRNAs) and cellular drug sensitivity is essential for understanding drug efficacy and therapeutic resistance. Traditional experimental methods to verify such associations are costly and time-consuming. Thus, the development of efficient computational methods for predicting circRNA-drug associations is crucial. In this study, we introduce a novel computational predictor called HETACDA, aimed at predicting potential circRNA-drug sensitivity associations. HETACDA constructs a heterogeneous graph network, incorporating the characteristic structure graph of drugs and circRNAs, along with the circRNA-drug sensitivity topology graph. By employing a graph convolutional network, the drug embedding vector is computed from the molecular structure of the drug. Through a graph attention mechanism, HETACDA assigns distinct attention weights to nodes in order to emphasize the contribution of various neighborhood nodes to the central node. Subsequently, an association score between circRNA-drug sensitivity is predicted using a three-layer fully connected neural network. Extensive experimental comparisons against several state-of-the-art methods highlight the effectiveness of our proposed framework. The availability of our source code and datasets on GitHub (https://github.com/xiaoxiaojun131421/HETACDA) facilitates replication and further research in this area. Xiaojun Xiao, Yurong Qian, Zhijian Huang 0001, Rongtao Zheng, Lei Deng 0002 |
BIBM | 3 |
| 2023 | PTDA-SWGCL: Predicting tRNA-Disease Associations using Supplementarily Weighted Graph Contrastive LearningabstracttRNAs play a pivotal role in protein synthesis by transporting amino acids to the ribosome according to mRNA instructions. These molecules are essential regulators in various biological processes, and their dysregulation is closely linked to human diseases. Predicting associations between tRNAs and diseases is valuable for uncovering biomarkers that aid in disease prevention, detection, prognosis, diagnosis, and treatment. However, experimental validation of such associations is resource-intensive, necessitating the development of robust computational methods. In this study, we propose PTDA-SWGCL, a novel model for predicting potential tRNA-disease associations. PTDA-SWGCL integrates tRNA and disease similarity information derived from Gaussian kernel similarity, sequence similarity, and semantic similarity. It initializes tRNA and disease embeddings using this similarity information and refines them through supplementarily weight and graph comparison learning training on the tRNA-disease association graph. The final association pair prediction is obtained by the inner product of the tRNA and disease embeddings. Experimental results demonstrate that PTDA-SWGCL outperforms state-of-the-art methods. Case studies confirm its effectiveness in predicting tRNA-disease associations. The code and data are available at https://github.com/ZYPssss/PTDA-SWGCL. Yuanpeng Zhang 0004, Yurong Qian, Xiaojun Xiao, Zhijian Huang 0001, Lei Deng 0002 |
BIBM | 5 |
| 2023 | Large-scale predicting protein functions through heterogeneous feature fusionabstractAs the volume of protein sequence and structure data grows rapidly, the functions of the overwhelming majority of proteins cannot be experimentally determined. Automated annotation of protein function at a large scale is becoming increasingly important. Existing computational prediction methods are typically based on expanding the relatively small number of experimentally determined functions to large collections of proteins with various clues, including sequence homology, protein-protein interaction, gene co-expression, etc. Although there has been some progress in protein function prediction in recent years, the development of accurate and reliable solutions still has a long way to go. Here we exploit AlphaFold predicted three-dimensional structural information, together with other non-structural clues, to develop a large-scale approach termed PredGO to annotate Gene Ontology (GO) functions for proteins. We use a pre-trained language model, geometric vector perceptrons and attention mechanisms to extract heterogeneous features of proteins and fuse these features for function prediction. The computational results demonstrate that the proposed method outperforms other state-of-the-art approaches for predicting GO functions of proteins in terms of both coverage and accuracy. The improvement of coverage is because the number of structures predicted by AlphaFold is greatly increased, and on the other hand, PredGO can extensively use non-structural information for functional prediction. Moreover, we show that over 205 000 ($\sim $100%) entries in UniProt for human are annotated by PredGO, over 186 000 ($\sim $90%) of which are based on predicted structure. The webserver and database are available at http://predgo.denglab.org/. Rongtao Zheng, Zhijian Huang 0001, Lei Deng 0002 |
Briefings Bioinform. | 2 |
| 2023 | DeepCoVDR: deep transfer learning with graph transformer and cross-attention for predicting COVID-19 drug responseabstractMOTIVATION: The coronavirus disease 2019 (COVID-19) remains a global public health emergency. Although people, especially those with underlying health conditions, could benefit from several approved COVID-19 therapeutics, the development of effective antiviral COVID-19 drugs is still a very urgent problem. Accurate and robust drug response prediction to a new chemical compound is critical for discovering safe and effective COVID-19 therapeutics. RESULTS: In this study, we propose DeepCoVDR, a novel COVID-19 drug response prediction method based on deep transfer learning with graph transformer and cross-attention. First, we adopt a graph transformer and feed-forward neural network to mine the drug and cell line information. Then, we use a cross-attention module that calculates the interaction between the drug and cell line. After that, DeepCoVDR combines drug and cell line representation and their interaction features to predict drug response. To solve the problem of SARS-CoV-2 data scarcity, we apply transfer learning and use the SARS-CoV-2 dataset to fine-tune the model pretrained on the cancer dataset. The experiments of regression and classification show that DeepCoVDR outperforms baseline methods. We also evaluate DeepCoVDR on the cancer dataset, and the results indicate that our approach has high performance compared with other state-of-the-art methods. Moreover, we use DeepCoVDR to predict COVID-19 drugs from FDA-approved drugs and demonstrate the effectiveness of DeepCoVDR in identifying novel COVID-19 drugs. AVAILABILITY AND IMPLEMENTATION: https://github.com/Hhhzj-7/DeepCoVDR. Zhijian Huang 0001, Lei Deng 0002 |
Bioinform. | 1 |
| 2022 | DeepFusionGO: Protein function prediction by fusing heterogeneous features through deep learningabstractExploring the functions of proteins is crucial for explaining cellular mechanisms, treating diseases, and developing new drugs. Due to experimental limitations, large-scale identification of protein function remains a challenging task in cell biology. Here we propose DeepFusionGo, a novel protein function prediction method that adopts a graph representation learning approach (GraphSAGE) to extract features from heterogeneous data sources. First, we generate embeddings from protein sequences using the pre-trained protein language model and InterPro domains with scaling gradient. Then we integrate these two embeddings with adaptive feature weights to the PPI graph and use GraphSAGE to generate the representation vector. Finally, we build the classification model to predict protein function based on the concatenated feature vector. The experimental results show that DeepFusionGO outperforms existing state-of-the-art methods, including sequence-based DeepGOPLUS, and PPI-based DeepGraphGO. DeepFusionGO also performs well in difficult protein function prediction. We demonstrate that selecting an appropriate protein features fusion method can improve the prediction performance, and using the PPI network and the protein representation vector obtained from the protein language model through the GraphSAGE algorithm is an effective way to mine potential functional clues. The source code and data sets are available at: https://github.com/Hhhzj-7/DeepFusionGO. Zhijian Huang 0001, Rongtao Zheng, Lei Deng 0002 |
BIBM | 1 |