VLDB 2026 Research / reviewers in the wild / expert
Pengyong Li
dblp:281/8336
· DBLP profile ↗
17ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing cross-context generalization in drug perturbation prediction with a multimodal conditional diffusion frameworkabstractMOTIVATION: Predicting drug-induced transcriptional perturbations is critical for precision medicine, yet existing models fail to capture multimodal biological context, limiting generalization across unseen drugs and cell lines. RESULTS: We present PertDiff, a conditional diffusion framework that integrates control gene expression, LLM-derived cell semantics, and pretrained molecular graph representations to predict transcriptome-wide perturbations. PertDiff outperforms state-of-the-art baselines in prediction accuracy and generalizes robustly across drugs and cell lines. It further demonstrates translational utility through accurate drug sensitivity prediction, therapeutic repurposing for pancreatic cancer, and concordance with real-world clinical treatment outcomes, establishing it as a biologically grounded transcriptomic modeling tool. AVAILABILITY: The source code and data are available at https://github.com/Panda-myj/PertDiff and https://doi.org/10.5281/zenodo.18427848. Yanjie Ma, Pengyong Li, Liang Yu 0002 |
Bioinform. | 4 |
| 2026 | SHE-SFL: An efficient and privacy-preserving heterogeneous federated split learning architecture based on homomorphic encryption
Jiaqi Xia, Pengyong Li |
Future Gener. Comput. Syst. | 3 |
| 2025 | ZeroGEN: leveraging language models for zero-shot ligand design from protein sequencesabstractMOTIVATION: Deep generative methods based on language models have the capability to generate new data that resemble a given distribution and have begun to gain traction in ligand design. However, existing models face significant challenges when it comes to generating ligands for unseen targets, a scenario known as zero-shot learning. The ability to effectively generate ligands for novel targets is crucial for accelerating drug discovery and expanding the applicability of ligand design. Therefore, there is a pressing need to develop robust deep generative frameworks that can operate efficiently in zero-shot scenarios. RESULTS: In this study, we introduce ZeroGEN, a novel zero-shot deep generative framework based on protein sequences. ZeroGEN analyzes extensive data on protein-ligand inter-relationships and incorporates contrastive learning to align known protein-ligand features, thereby enhancing the model's understanding of potential interactions between proteins and ligands. Additionally, ZeroGEN employs self-distillation to filter the initially generated data, retaining only the ligands deemed reliable by the model. It also implements data augmentation techniques to aid the model in identifying ligands that match unseen targets. Experimental results demonstrate that ZeroGEN successfully generates ligands for unseen targets with strong affinity and desirable drug-like properties. Furthermore, visualizations of molecular docking and attention matrices reveal that ZeroGEN can autonomously focus on key residues of proteins, underscoring its capability to understand and generate effective ligands for novel targets. AVAILABILITY AND IMPLEMENTATION: The source code and data of this work is freely available in the https://github.com/viko-3/ZeroGEN. Yangyang Chen 0006, Pengyong Li, Xiangxiang Zeng, Lei Xu 0002 |
Bioinform. | 3 |
| 2025 | RNA language model and graph attention network for RNA and small molecule binding sites predictionabstractMOTIVATION: The structural complexities enable RNA to serve as a versatile molecular scaffold capable of binding small molecules with high specificity. Understanding these interactions is essential for elucidating RNA's role in disease mechanisms and developing RNA-targeted therapeutics. However, predicting RNA-small molecule binding sites remains a significant challenge due to their conformational flexibility, structural diversity, and the limited availability of high-resolution structural data. RESULTS: In this study, we propose RLsite, a novel computational framework integrating pre-trained RNA language models with graph attention networks (GAT) to predict small-molecule binding sites on RNA. Our method effectively captures both sequential and structural features of RNA by leveraging large-scale RNA sequence data to learn intrinsic patterns and processing graph-based RNA structures to highlight key topological and spatial features. Compared to existing methods, RLsite demonstrates superior accuracy, generalizability, and biological relevance, achieving a Precision of 0.749, a Recall of 0.654, an MCC of 0.474, and an AUC of 0.828 on the public test set, which significantly outperforms the previous models, such as CapBind (an AUC of 0.770), MultiModRLBP (an AUC of 0.780), and RNABind (an AUC of 0.471). Notably, a case study of the PreQ1 riboswitch has achieved strong predictive performance (AUC = 0.97, Recall = 0.9), and its predicted binding sites have been confirmed experimentally. These results underscore our method as a potentially powerful tool for RNA-targeted drug discovery and advancing our understanding of RNA-ligand interactions. AVAILABILITY AND IMPLEMENTATION: The resource codes and data can be accessed at https://github.com/SaisaiSun/RLsite. Saisai Sun, Jianyi Yang 0002, Lin Gao 0006, Pengyong Li |
Bioinform. | 4 |
| 2025 | SFML: A personalized, efficient, and privacy-preserving collaborative traffic classification architecture based on split learning and mutual learning
Jiaqi Xia, Meng Wu 0003, Pengyong Li |
Future Gener. Comput. Syst. | 3 |
| 2024 | Improving drug response prediction via integrating gene relationships with deep learningabstractPredicting the drug response of cancer cell lines is crucial for advancing personalized cancer treatment, yet remains challenging due to tumor heterogeneity and individual diversity. In this study, we present a deep learning-based framework named Deep neural network Integrating Prior Knowledge (DIPK) (DIPK), which adopts self-supervised techniques to integrate multiple valuable information, including gene interaction relationships, gene expression profiles and molecular topologies, to enhance prediction accuracy and robustness. We demonstrated the superior performance of DIPK compared to existing methods on both known and novel cells and drugs, underscoring the importance of gene interaction relationships in drug response prediction. In addition, DIPK extends its applicability to single-cell RNA sequencing data, showcasing its capability for single-cell-level response prediction and cell identification. Further, we assess the applicability of DIPK on clinical data. DIPK accurately predicted a higher response to paclitaxel in the pathological complete response (pCR) group compared to the residual disease group, affirming the better response of the pCR group to the chemotherapy compound. We believe that the integration of DIPK into clinical decision-making processes has the potential to enhance individualized treatment strategies for cancer patients. Pengyong Li, Zhengxiang Jiang, Tianxiao Liu |
Briefings Bioinform. | 1 |
| 2024 | DeepDR: a deep learning library for drug response predictionabstractSUMMARY: Accurate drug response prediction is critical to advancing precision medicine and drug discovery. Recent advances in deep learning (DL) have shown promise in predicting drug response; however, the lack of convenient tools to support such modeling limits their widespread application. To address this, we introduce DeepDR, the first DL library specifically developed for drug response prediction. DeepDR simplifies the process by automating drug and cell featurization, model construction, training, and inference, all achievable with brief programming. The library incorporates three types of drug features along with nine drug encoders, four types of cell features along with nine cell encoders, and two fusion modules, enabling the implementation of up to 135 DL models for drug response prediction. We also explored benchmarking performance with DeepDR, and the optimal models are available on a user-friendly visual interface. AVAILABILITY AND IMPLEMENTATION: DeepDR can be installed from PyPI (https://pypi.org/project/deepdr). The source code and experimental data are available on GitHub (https://github.com/user15632/DeepDR). Zhengxiang Jiang, Pengyong Li |
Bioinform. | 2 |
| 2024 | Secure architecture for Industrial Edge of Things(IEoT): A hierarchical perspective
Pengyong Li, Jiaqi Xia, Qian Wang 0028, Meng Wu 0003 |
Comput. Networks | 1 |
| 2024 | MuLDOM: Forecasting Multivariate Anomalies on Edge Devices in IIoT Using Multibranch LSTM and Differential Overfitting Mitigation ModelabstractIn the Industrial Internet of Things (IIoT) environment, there is a multitude of heterogeneous industrial edge devices (IEDs) from various sources. Real-time monitoring and precise prediction of its operational status are typically essential. However, existing deep learning-based models often encounter overfitting issues due to complex parameter configurations. Furthermore, ensuring the comprehensive performance of anomaly event forecasts for IEDs has emerged as a pressing issue requiring resolution to accommodate a wider range of practical applications. In this article, we introduce a novel multibranch long short term memory and differential overfitting mitigation scheme (MuLDOM). This scheme is designed to achieve two primary objectives: 1) to extract features and denoise multivariate time series adaptively and 2) to implement the differential overfitting mitigation algorithm for the first time, thereby enabling robust intelligent anomaly detection and forecast (IADF). Expanding on this framework, we provide detailed information on the development of an online prediction scoring mechanism based on multivariate time series data. This mechanism aims to enhance the efficiency of quantitatively estimating the spatial and temporal characteristics associated with IEDs. We conducted extensive experiments on four publicly available industrial data sets and compared our approach with nine recent baseline methods. The results indicate that our method surpasses the recent state-of-the-art methods, validating its effectiveness. These findings underscore its significant potential for real-world applications. Pengyong Li, Meng Wu 0003, Jiaqi Xia, Qian Wang 0028 |
IEEE Internet Things J. | 1 |
| 2024 | PT-ADP: A personalized privacy-preserving federated learning scheme based on transaction mechanism
Jiaqi Xia, Pengyong Li, Yiming Mao 0010 |
Inf. Sci. | 2 |
| 2023 | Improving drug-target affinity prediction via feature fusion and knowledge distillationabstractRapid and accurate prediction of drug-target affinity can accelerate and improve the drug discovery process. Recent studies show that deep learning models may have the potential to provide fast and accurate drug-target affinity prediction. However, the existing deep learning models still have their own disadvantages that make it difficult to complete the task satisfactorily. Complex-based models rely heavily on the time-consuming docking process, and complex-free models lacks interpretability. In this study, we introduced a novel knowledge-distillation insights drug-target affinity prediction model with feature fusion inputs to make fast, accurate and explainable predictions. We benchmarked the model on public affinity prediction and virtual screening dataset. The results show that it outperformed previous state-of-the-art models and achieved comparable performance to previous complex-based models. Finally, we study the interpretability of this model through visualization and find it can provide meaningful explanations for pairwise interaction. We believe this model can further improve the drug-target affinity prediction for its higher accuracy and reliable interpretability. Ruiqiang Lu, Jun Wang 0123, Pengyong Li, Shuoyan Tan, Yiting Pan, Huanxiang Liu, Peng Gao 0015, Guo Tong Xie |
Briefings Bioinform. | 3 |
| 2022 | HCL: Improving Graph Representation with Hierarchical Contrastive Learning
Jun Wang 0123, Weixun Li, Changyu Hou, Yixuan Qiao, Pengyong Li, Peng Gao 0015, Guo Tong Xie |
ISWC | 7 |
| 2021 | Pairwise Half-graph Discrimination: A Simple Graph-level Self-supervised Strategy for Pre-training Graph Neural NetworksabstractSelf-supervised learning has gradually emerged as a powerful technique for graph representation learning. However, transferable, generalizable, and robust representation learning on graph data still remains a challenge for pre-training graph neural networks. In this paper, we propose a simple and effective self-supervised pre-training strategy, named Pairwise Half-graph Discrimination (PHD), that explicitly pre-trains a graph neural network at graph-level. PHD is designed as a simple binary classification task to discriminate whether two half-graphs come from the same source. Experiments demonstrate that the PHD is an effective pre-training strategy that offers comparable or superior performance on 13 graph classification tasks compared with state-of-the-art strategies, and achieves notable improvements when combined with node-level strategies. Moreover, the visualization of learned representation revealed that PHD strategy indeed empowers the model to learn graph-level knowledge like the molecular scaffold. These results have established PHD as a powerful and effective self-supervised learning strategy in graph-level representation learning. Pengyong Li, Jun Wang 0123, Ziliang Li, Yixuan Qiao, Xianggen Liu, Peng Gao 0015, Sen Song, Guo Tong Xie |
IJCAI | 1 |
| 2021 | TrimNet: learning molecular representation from triplet messages for biomedicineabstractMOTIVATION: Computational methods accelerate drug discovery and play an important role in biomedicine, such as molecular property prediction and compound-protein interaction (CPI) identification. A key challenge is to learn useful molecular representation. In the early years, molecular properties are mainly calculated by quantum mechanics or predicted by traditional machine learning methods, which requires expert knowledge and is often labor-intensive. Nowadays, graph neural networks have received significant attention because of the powerful ability to learn representation from graph data. Nevertheless, current graph-based methods have some limitations that need to be addressed, such as large-scale parameters and insufficient bond information extraction. RESULTS: In this study, we proposed a graph-based approach and employed a novel triplet message mechanism to learn molecular representation efficiently, named triplet message networks (TrimNet). We show that TrimNet can accurately complete multiple molecular representation learning tasks with significant parameter reduction, including the quantum properties, bioactivity, physiology and CPI prediction. In the experiments, TrimNet outperforms the previous state-of-the-art method by a significant margin on various datasets. Besides the few parameters and high prediction accuracy, TrimNet could focus on the atoms essential to the target properties, providing a clear interpretation of the prediction tasks. These advantages have established TrimNet as a powerful and useful computational tool in solving the challenging problem of molecular representation learning. AVAILABILITY: The quantum and drug datasets are available on the website of MoleculeNet: http://moleculenet.ai. The source code is available in GitHub: https://github.com/yvquanli/trimnet. CONTACT: [email protected], [email protected]. Pengyong Li, Chang-Yu Hsieh, Shengyu Zhang 0002, Xianggen Liu, Huanxiang Liu, Sen Song |
Briefings Bioinform. | 1 |
| 2021 | An effective self-supervised framework for learning expressive molecular global representations to drug discoveryabstractHow to produce expressive molecular representations is a fundamental challenge in artificial intelligence-driven drug discovery. Graph neural network (GNN) has emerged as a powerful technique for modeling molecular data. However, previous supervised approaches usually suffer from the scarcity of labeled data and poor generalization capability. Here, we propose a novel molecular pre-training graph-based deep learning framework, named MPG, that learns molecular representations from large-scale unlabeled molecules. In MPG, we proposed a powerful GNN for modelling molecular graph named MolGNet, and designed an effective self-supervised strategy for pre-training the model at both the node and graph-level. After pre-training on 11 million unlabeled molecules, we revealed that MolGNet can capture valuable chemical insights to produce interpretable representation. The pre-trained MolGNet can be fine-tuned with just one additional output layer to create state-of-the-art models for a wide range of drug discovery tasks, including molecular properties prediction, drug-drug interaction and drug-target interaction, on 14 benchmark datasets. The pre-trained MolGNet in MPG has the potential to become an advanced molecular encoder in the drug discovery pipeline. Pengyong Li, Jun Wang 0123, Yixuan Qiao, Yihuan Yu, Peng Gao 0015, Guo Tong Xie, Sen Song |
Briefings Bioinform. | 1 |
| 2021 | Simulated annealing for optimization of graphs and sequences
Xianggen Liu, Pengyong Li, Fandong Meng, Hao Zhou 0012, Huasong Zhong, Jie Zhou 0016, Lili Mou, Sen Song |
Neurocomputing | 2 |
| 2021 | Deep geometric representations for modeling effects of mutations on protein-protein binding affinityabstractModeling the impact of amino acid mutations on protein-protein interaction plays a crucial role in protein engineering and drug design. In this study, we develop GeoPPI, a novel structure-based deep-learning framework to predict the change of binding affinity upon mutations. Based on the three-dimensional structure of a protein, GeoPPI first learns a geometric representation that encodes topology features of the protein structure via a self-supervised learning scheme. These representations are then used as features for training gradient-boosting trees to predict the changes of protein-protein binding affinity upon mutations. We find that GeoPPI is able to learn meaningful features that characterize interactions between atoms in protein structures. In addition, through extensive experiments, we show that GeoPPI achieves new state-of-the-art performance in predicting the binding affinity changes upon both single- and multi-point mutations on six benchmark datasets. Moreover, we show that GeoPPI can accurately estimate the difference of binding affinities between a few recently identified SARS-CoV-2 antibodies and the receptor-binding domain (RBD) of the S protein. These results demonstrate the potential of GeoPPI as a powerful and useful computational tool in protein design and engineering. Our code and datasets are available at: https://github.com/Liuxg16/GeoPPI. Xianggen Liu, Yunan Luo, Pengyong Li, Sen Song, Jian Peng 0001 |
PLoS Comput. Biol. | 3 |