Shugang Zhang

dblp:188/6629 · DBLP profile ↗
← Back
15ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0002-9774-9709ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 13 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Artificial intelligence-enabled multi-scale virtual cell: perspective, challenges, and opportunities
abstract
As the fundamental unit of life, cells coordinate biological activities through the interaction between microscopic molecular mechanisms and macroscopic tissue organization. Traditional research studies, experiments, and biochemical analyses, give rise to important insights, although they are restricted in spatiotemporal resolution and processing power, thereby precluding the understanding of dynamic cross-scale biological events . Breakthroughs in artificial intelligence (AI) have given birth to the AI virtual cell (AIVC) as a new way to do research. By integrating multi-omics data and mixing methods from multidisciplinary models, AIVC establishes a digital twin system to simulate cell functions and behaviors. AIVC still faces a number of pressing challenges that need to be addressed in its current development stage. In this review, we are proposing a unified definition and technical framework for AIVC and analyze in detail the cross-scale coupling mechanisms of the "gene-protein-pathway-cell" hierarchy. Furthermore, we decompose the technical construction framework of AIVC from cross-scale representation engineering, functional submodule design, and multi-component dynamic regulation mechanisms. Additionally, we summarize the existing models and datasets in the field to provide reference resources for researchers. Finally, we deeply discuss the challenges faced by AIVC, such as data heterogeneity and model interpretability, and aim to accelerate the research progress in the AIVC field while driving the life sciences to shift from observational analysis to a paradigm that integrates predictability and innovation. Despite being in the early stage, AIVC is a trending topic that has garnered widespread interest. This review aims to integrate existing models, datasets, and technical ideas to provide a unified framework for field development.
Huasen Jiang, Xiangpeng Bi, Wenjian Ma, Haibo Ni, Zhiqiang Wei 0002, Pin Sun, Henggui Zhang, Shugang Zhang
Briefings Bioinform.9
2026 Protein interaction pattern recognition using heterogeneous semantics mining and hierarchical graph representation
Xiangpeng Bi, Wenjian Ma, Huasen Jiang, Weigang Lu 0002, Jie Nie, Zhiqiang Wei 0002, Shugang Zhang
Pattern Recognit.7
2026 Geometric Deep Learning for Protein-Ligand Affinity Prediction With Hybrid Message Passing Strategies
abstract
Accurate prediction of protein-ligand affinity (PLA) is critical for drug discovery. Recent deep learning approaches have adopted data-driven models for PLA prediction by learning intrinsic patterns from one-dimensional (1D) sequential or two-dimensional (2D) graph representations of proteins and ligands. However, these low-dimensional methods overlook the three-dimensional (3D) geometric features, which are hypothesized to be critical in binding interaction. To address the above problem, we present a Geometric deep learning approach with Hybrid message passing strategies--HybridGeo, for protein-ligand affinity prediction. We adopt dual-view graph learning to model the intra- and inter-molecular atomic interactions and propose to aggregate the spatial information with hybrid strategies. In addition, to fully model the inter-residue dependency upon message aggregation, we adopt a geometric graph transformer on the residue-scale graph of protein pockets. Extensive experiments on the PDBbind dataset show that HybridGeo achieves state-of-the-art performance with a Root Mean Square Error (RMSE) of 1.172. HybridGeo also achieves the best among all baseline models on three external test sets, showcasing good generalizability and robustness. Through systematic ablation experiments, we validated the effectiveness of the proposed modules, and further demonstrated the superior performance of HybridGeo in predicting the binding affinity of macrocyclic compound complexes through case studies. Visualization analysis further indicates the biological interpretability of the model predictions.
Jiaren Li, Huasen Jiang, Wenjian Ma, Xiangpeng Bi, Weigang Lu 0002, Fei Yang 0003, Zhiqiang Wei 0002, Shugang Zhang
IEEE J. Biomed. Health Informatics10
2026 SSPPI: Cross-Modality Enhanced Protein-Protein Interaction Prediction From Sequence and Structure Perspectives
abstract
Recent advances have shown great promise in mining multimodal protein knowledge for better protein-protein interaction (PPI) prediction by enriching the representation of proteins. However, existing solutions lack a comprehensive consideration of both local patterns and global dependencies in proteins, hindering the full exploitation of modal information. Additionally, the inherent disparities between modalities are often disregarded, which may lead to inferior modality complementarity effects. To address these issues, we propose a cross-modality enhanced PPI prediction method from the perspectives of protein sequence and structure modalities, namely SSPPI. In this framework, our main contribution is that we integrate both sequence and structural modalities of proteins and employ an alignment and fusion method between modalities to further generate more comprehensive protein representations for PPI prediction. Specifically, we design two modal representation modules (Convformer and Graphormer) tailored for protein sequence and structure modalities, respectively, to enhance the quality of modal representation. Subsequently, we introduce a Cross-modality enhancer module to achieve alignment and fusion between modalities, thereby generating more informative modal joint representations. Finally, we devise a cross-protein fusion (CPF) module to model residue interaction processes between proteins, thereby enriching the joint representation of protein pairs. Extensive experimentation on four benchmark datasets demonstrates that our proposed model surpasses all current state-of-the-art (SOTA) methods. The source codes are publicly available at the following link https://github.com/bixiangpeng/SSPPI/.
Xiangpeng Bi, Wenjian Ma, Huasen Jiang, Weigang Lu 0002, Zhiqiang Wei 0002, Shugang Zhang
IEEE Trans. Neural Networks Learn. Syst.6
2025 SuperCardioTox: Drug-Induced Cardiotoxicity Prediction Based on Self-Supervised Learning
abstract
In this study, we propose a pre-trained graph neural network-based model, called SuperCardioTox, for predicting hERG cardiotoxicity. The model employs a high-quality dataset derived from the CHEMBL database to quantify the inhibitory capacity of compounds through pIC50 values, and uses the GraphMAE graph autoencoder to learn the molecular structure characterisation of drugs. Its functionality is based on the masking of node features and their reconstruction, thereby enabling the learning of the intrinsic laws governing graph data. Furthermore, a drug toxicity prediction model that integrates graph neural network (GNN) and multilayer perceptron (MLP) has been developed to achieve accurate drug toxicity prediction by extracting deep features of graph data. The objective of this study is to enhance the precision of drug design and toxicity prediction through the utilisation of extensive unlabelled molecular data, thereby providing robust support for novel drug development and ensuring the safety of patients' medication. The experimental results demonstrate that the model demonstrates robust performance in drug toxicity prediction, thereby validating its utility and generalisability.
Xiangpeng Bi, Zhiqiang Wei 0002, Pin Sun, Shugang Zhang
BIBM6
2025 Dual-protein embedding-based graph model with dynamic attention for interaction prediction
abstract
Protein-protein interactions (PPIs) are fundamental to biological processes, yet experimental determination of PPIs remains costly and labor-intensive. While computational methods have emerged as promising alternatives, sequence-based approaches face critical challenges: (1) effectively capturing long-range dependencies and critical biochemical patterns in variable-length sequences, and (2) balancing computational efficiency with sensitivity to subtle residue-level interactions. Here, we present Dual Protein Embedding-based Graph Model (DPEG), which leverages dynamic graph attention networks to enable robust sequence-driven PPI prediction. Unlike structure-dependent methods, DPEG operates solely on sequence data, bypassing the need for structural or domain annotations. Specifically, we employ ESM-2 to transform sequences into residue-level graphs, preserving evolutionary and physicochemical context. To address variable sequence lengths, we design a module that can represent protein sequences of arbitrary lengths as graph networks at the amino acid level. Further, a gated attention mechanism is introduced to adaptively refining residue representations. Finally, a dynamic attention mechanism prioritizes functionally critical motifs within the graph. Evaluated on four diverse PPI datasets spanning different species and interaction types, DPEG achieves state-of-the-art performance and demonstrates strong cross-dataset generalizability. By integrating deep sequence semantics with graph-based interaction modeling, DPEG advances sequence-only PPI prediction, offering a scalable and biologically plausible framework for proteome-wide studies.
Shunpeng Pang, Mingjian Jiang, Shugang Zhang, Zhen Li 0024, Li Guo 0012
Briefings Bioinform.3
2025 SuperEdgeGO: Edge-supervised graph representation learning for enhanced protein function prediction
abstract
Understanding the functions of proteins is of great importance for deciphering the mechanisms of life activities. To date, there have been over 200 million known proteins, but only 0.2% of them have well-annotated functional terms. By measuring the contacts among residues, proteins can be described as graphs so that the graph leaning approaches can be applied to learn protein representations. However, existing graph-based methods put efforts in enriching the residue node information and did not fully exploit the edge information, which leads to suboptimal representations considering the strong association of residue contacts to protein structures and to the functions. In this article, we propose SuperEdgeGO, which introduces the supervision of edges in protein graphs to learn a better graph representation for protein function prediction. Different from common graph convolution methods that uses edge information in a plain or unsupervised way, we introduce a supervised attention to encode the residue contacts explicitly into the protein representation. Comprehensive experiments demonstrate that SuperEdgeGO achieves state-of-the-art performance on all three categories of protein functions. Additional ablation analysis further proves the effectiveness of the devised edge supervision strategy. The implementation of edge supervision in SuperEdgeGO resulted in enhanced graph representations for protein function prediction, as demonstrated by its superior performance across all the evaluated categories. This superior performance was confirmed through ablation analysis, which validated the effectiveness of the edge supervision strategy. This strategy has a broad application prospect in the study of protein function and related fields.
Shugang Zhang, Yuntong Li, Wenjian Ma, Xiangpeng Bi, Huasen Jiang, Zhiqiang Wei 0002
PLoS Comput. Biol.1
2025 HiSIF-DTA: A Hierarchical Semantic Information Fusion Framework for Drug-Target Affinity Prediction
abstract
Accurately identifying drug-target affinity (DTA) plays a significant role in promoting drug discovery and has attracted increasing attention in recent years. Exploring appropriate protein representation methods and increasing the abundance of protein information is critical in enhancing the accuracy of DTA prediction. Recently, numerous deep learning-based models have been proposed to utilize the sequential or structural features of target proteins. However, these models capture only the low-order semantics that exist in a single protein, while the high-order semantics abundant in biological networks are largely ignored. In this article, we propose HiSIF-DTA-a hierarchical semantic information fusion framework for DTA prediction. In this framework, a hierarchical protein graph is constructed that includes not only contact maps as low-order structural semantics but also protein-protein interaction (PPI) networks as high-order functional semantics. Particularly, two distinct hierarchical fusion strategies (i.e., Top-down and Bottom-Up) are designed to integrate the different protein semantics, therefore contributing to a richer protein representation. Comprehensive experimental results demonstrate that HiSIF-DTA outperforms current state -of-the-art methods for prediction on the benchmark datasets of the DTA task. Further validation on binary tasks and visualization analysis demonstrates the generalization and interpretation abilities of the proposed method.
Xiangpeng Bi, Shugang Zhang, Wenjian Ma, Huasen Jiang, Zhiqiang Wei 0002
IEEE J. Biomed. Health Informatics2
2025 MHAN-DTA: A Multiscale Hybrid Attention Network for Drug-Target Affinity Prediction
abstract
Drug-target affinity prediction is a key challenge in the drug discovery process. Recent advances have demonstrated the great potential of deep learning in predicting affinities; however, existing approaches learn the representation of drug-target complex insufficiently, leading to suboptimal performance. Here, we propose a Multiscale Hybrid Attention Network for the Drug-Target Affinity prediction, named MHAN-DTA, which aims to address the problem of insufficient feature mining thereby improving the prediction performance. To empower the model with global perception ability, a pocket-oriented feature aggregation and extraction module is developed based on self-attention mechanisms, together with a hierarchical strategy applied to the target proteins. We further introduce a cross-modal fusion module and a cross-entity interaction module for mining the multiscale intra-molecular and inter-molecular features within the binding sites. Comprehensive evaluations on four benchmark test sets, including an internal and three external benchmark datasets, demonstrate that the proposed approach achieves superior and robust performance.
Jiaren Li, Xiangpeng Bi, Wenjian Ma, Huasen Jiang, Shanglong Liu, Zhiqiang Wei 0002, Shugang Zhang
IEEE J. Biomed. Health Informatics8
2024 CollaPPI: A Collaborative Learning Framework for Predicting Protein-Protein Interactions
abstract
Exploring protein-protein interaction (PPI) is of paramount importance for elucidating the intrinsic mechanism of various biological processes. Nevertheless, experimental determination of PPI can be both time-consuming and expensive, motivating the exploration of data-driven deep learning technologies as a viable, efficient, and accurate alternative. Nonetheless, most current deep learning-based methods regarded a pair of proteins to be predicted for possible interaction as two separate entities when extracting PPI features, thus neglecting the knowledge sharing among the collaborative protein and the target protein. Aiming at the above issue, a collaborative learning framework CollaPPI was proposed in this study, where two kinds of collaboration, i.e., protein-level collaboration and task-level collaboration, were incorporated to achieve not only the knowledge-sharing between a pair of proteins, but also the complementation of such shared knowledge between biological domains closely related to PPI (i.e., protein function, and subcellular location). Evaluation results demonstrated that CollaPPI obtained superior performance compared to state-of-the-art methods on two PPI benchmarks. Besides, evaluation results of CollaPPI on the additional PPI type prediction task further proved its excellent generalization ability.
Wenjian Ma, Xiangpeng Bi, Huasen Jiang, Shugang Zhang, Zhiqiang Wei 0002
IEEE J. Biomed. Health Informatics4
2023 A Personalized Learning Path Recommendation Method for Learning Objects with Diverse Coverage Levels
Tengju Li, Shugang Zhang, Fei Yang 0003, Weigang Lu 0002
AIED3
2023 Predicting Drug-Target Affinity by Learning Protein Knowledge From Biological Networks
abstract
Predicting drug-target affinity (DTA) is a crucial step in the process of drug discovery. Efficient and accurate prediction of DTA would greatly reduce the time and economic cost of new drug development, which has encouraged the emergence of a large number of deep learning-based DTA prediction methods. In terms of the representation of target proteins, current methods can be classified into 1D sequence- and 2D-protein graph-based methods. However, both two approaches focused only on the inherent properties of the target protein, but neglected the broad prior knowledge regarding protein interactions that have been clearly elucidated in past decades. Aiming at the above issue, this work presents an end-to-end DTA prediction method named MSF-DTA (Multi-Source Feature Fusion-based Drug-Target Affinity). The contributions can be summarized as follows. First, MSF-DTA adopts a novel "neighboring feature"-based protein representation. Instead of utilizing only the inherent features of a target protein, MSF-DTA gathers additional information for the target protein from its biologically related "neighboring" proteins in PPI (i.e., protein-protein interaction) and SSN (i.e., sequence similarity) networks to get prior knowledge. Second, the representation was learned using an advanced graph pre-training framework, VGAE, which could not only gather node features but also learn topological connections, therefore contributing to a richer protein representation and benefiting the downstream DTA prediction task. This study provides new perspective for the DTA prediction task, and evaluation results demonstrated that MSF-DTA obtained superior performances compared to current state-of-the-art methods.
Wenjian Ma, Shugang Zhang, Zhen Li 0024, Mingjian Jiang, Nianfan Guo, Yuanfei Li, Xiangpeng Bi, Huasen Jiang, Zhiqiang Wei 0002
IEEE J. Biomed. Health Informatics2
2022 Molecular substructure tree generative model for de novo drug design
abstract
Deep learning shortens the cycle of the drug discovery for its success in extracting features of molecules and proteins. Generating new molecules with deep learning methods could enlarge the molecule space and obtain molecules with specific properties. However, it is also a challenging task considering that the connections between atoms are constrained by chemical rules. Aiming at generating and optimizing new valid molecules, this article proposed Molecular Substructure Tree Generative Model, in which the molecule is generated by adding substructure gradually. The proposed model is based on the Variational Auto-Encoder architecture, which uses the encoder to map molecules to the latent vector space, and then builds an autoregressive generative model as a decoder to generate new molecules from Gaussian distribution. At the same time, for the molecular optimization task, a molecular optimization model based on CycleGAN was constructed. Experiments showed that the model could generate valid and novel molecules, and the optimized model effectively improves the molecular properties.
Tao Song 0001, Shugang Zhang, Mingjian Jiang, Zhiqiang Wei 0002, Zhen Li 0024
Briefings Bioinform.3
2020 Mechanisms Underlying Sulfur Dioxide Pollution Induced Ventricular Arrhythmia: A Simulation Study
abstract
Air pollution has been long recognized as a hazardous factor for the human cardiovascular system. Sulfur dioxide (SO2) is a common ambient air pollutant that is able to cause detrimental effects on hearts. Though the cardiotoxicity effects by sulfur dioxide were well documented in epidemiological reports, however, the underlying mechanisms remain unclear owing to the technical limitations that exist in traditional experimental measures. In this article, we developed a multi-scale virtual ventricular tissue, which incorporated electrophysiological activities from subcellular to tissue levels and could provide comprehensive records and insightful mechanisms of the SO2induced ventricular arrhythmias. Based on the available cellular and molecular experimental data, our findings provide a rationale at tissue level in support of epidemiologic studies pointing to the deleterious effects of SO2pollution on cardiac function.
Shugang Zhang, Weigang Lu 0002, Zhen Li 0024, Mingjian Jiang, Zhiqiang Wei 0002, Henggui Zhang
BIBM1
2016 How to record the amount of exercise automatically? A general real-time recognition and counting approach for repetitive activities
abstract
Exercise is considered as an effective mean against overweight and obesity-related diseases. In this paper, a real-time activity recognition and counting approach is proposed to evaluate amount of exercise only using a wearable smart watch. First, accelerometer and gyroscope data are collected to extract efficient features. Then Support Vector Machine classifiers are trained to recognize nine common exercise activities in real time. In order to measure the frequency of repetitive activity, a general activity counting algorithm based on gyroscope is proposed which is applicable for different types of activity. Various activities can be counted uninterruptedly using the proposed general method without frequently changing algorithms. Through experiments, it is demonstrated that the extracted features are efficient for real time exercise activity recognition. Moreover, our comparative experiments have shown that our counting approach is more accurate than other products on the market.
Shugang Zhang, Zhen Li 0024, Jie Nie, Lei Huang 0010, Zhiqiang Wei 0002
BIBM1