EDBT 2026 Demo / reviewers in the wild / expert
Shaokai Wang
dblp:166/7762
· DBLP profile ↗
26ranked-venue papers
12as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 7 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MACE: A Multi-scale Attention Convolutional Evaluation Model for Top-Down Mass Spectrometry Isotopic Envelopes
Jiancheng Zhong, Maoqi Yuan, Shaokai Wang |
ISBRA (2) | 4 |
| 2026 | MMFF-DDI: A Multi-Modal Fusion Framework for Drug-Drug Interaction Event Prediction With Contrastive LearningabstractAccurately predicting drug-drug interaction events (DDIEs) is critical for optimizing combination therapies and ensuring drug safety. However, existing methods typically rely on either handcrafted molecular fingerprints or static embeddings from pretrained models, which limits their ability to jointly capture local chemical substructures and three-dimensional geometric features. To overcome these limitations, we propose MMFF-DDI, a multi-modal fusion framework based on contrastive learning for drug-drug interaction event (DDIE) prediction. MMFF-DDI extracts drug representations from three modalities-Morgan fingerprints, canonical SMILES, and 3D molecular graphs-using an attention-augmented autoencoder, a MolFormer encoder, and an Equivariant Graph Neural Network (EGNN), respectively. Furthermore, a contrastive multi-modal integration submodule is designed to transform multi-modal representation learning from a concatenation-based paradigm to an alignment-based paradigm, thereby achieving cross-modal consistency and complementary feature fusion. Experimental results show that MMFF-DDI outperforms the best competitive method (MRGCDDI) in predicting DDIE involving existing drugs, achieving improvements of 7.87% and 7.99% in Macro-F1 and Macro-precision, respectively. Furthermore, MMFF-DDI outperforms the best competitive method (DSN-DDI) in predicting DDIEs involving new drugs, achieving improvements of 8.06% and 12.79% in Macro-F1 and Macro-precision, respectively. Visualization experiments and case studies validate its practical applicability and superior predictive performance. Guihua Duan, Shaokai Wang |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | ScFold: a GNN-based model for efficient inverse folding of short-chain proteins via spatial reductionabstractIn the realm of protein design, the efficient construction of protein sequences that accurately fold into predefined structures has become an important area of research. Although advancements have been made in the study of long-chain proteins, the design of short-chain proteins requires equal consideration. The structural information inherent in short and single chains is typically less comprehensive than that of full-length chains, which can negatively impact their performance. To address this challenge, we introduce ScFold, a novel model that incorporates an innovative node module. This module utilizes spatial dimensionality reduction and positional encoding mechanisms to enhance the extraction of structural features. Experimental results indicate that ScFold achieves a recovery rate of 52.22$\%$ on the CATH4.2 dataset, demonstrating notable efficacy for short-chain proteins, with a recovery rate of 41.6$\%$. Additionally, ScFold further exhibits enhanced recovery rates of 59.32$\%$ and 61.59$\%$ on the TS50 and TS500 datasets, respectively, demonstrating its effectiveness across diverse protein types. Additionally, we performed protein length stratification on the TS500 and CATH4.2 datasets and tested ScFold on length-specific sub-datasets. The results confirm the model's superiority in handling short-chain proteins. Finally, we selected several protein sequence groups from the CATH4.2 dataset for structural visualization analysis and provided comparisons between the model-generated sequences and the target sequences. Jiancheng Zhong, Zhiwei Zou, Shaokai Wang |
Briefings Bioinform. | 4 |
| 2025 | Application of machine learning in drug side effect prediction: databases, methods, and challengesabstractAbstract Drug side effects have become paramount concerns in drug safety research, ranking as the fourth leading cause of mortality following cardiovascular diseases, cancer, and infectious diseases. Simultaneously, the widespread use of multiple prescription and over-the-counter medications by many patients in their daily lives has heightened the occurrence of side effects resulting from Drug-Drug Interactions (DDIs). Traditionally, assessments of drug side effects relied on resource-intensive and time-consuming laboratory experiments. However, recent advancements in bioinformatics and the rapid evolution of artificial intelligence technology have led to the accumulation of extensive biomedical data. Based on this foundation, researchers have developed diverse machine learning methods for discovering and detecting drug side effects. This paper provides a comprehensive overview of recent advancements in predicting drug side effects, encompassing the entire spectrum from biological data acquisition to the development of sophisticated machine learning models. The review commences by elucidating widely recognized datasets and Web servers relevant to the field of drug side effect prediction. Subsequently, The study delves into machine learning methods customized for binary, multi-class, and multi-label classification tasks associated with drug side effects. These methods are applied to a variety of representative computational models designed for identifying side effects induced by single drugs and DDIs. Finally, the review outlines the challenges encountered in predicting drug side effects using machine learning approaches and concludes by illuminating important future research directions in this dynamic field. Chenliang Xie, Shaokai Wang |
Frontiers Comput. Sci. | 5 |
| 2025 | NeoMS: Mass Spectrometry-Based Method for Uncovering Mutated MHC-I NeoantigensabstractMajor Histocompatibility Complex (MHC) molecules play a critical role in the immune system by presenting peptides on the cell surface for recognition by T-cells. Tumor cells often produce MHC peptides with amino acid mutations, known as neoantigens, which evade T-cell recognition, leading to rapid tumor growth. In immunotherapies such as TCR-T and CAR-T, identifying these mutated MHC peptide sequences is crucial. Current mass spectrometry-based peptide identification methods primarily rely on database searching, which fails to detect mutated peptides not present in human databases. In this paper, we propose a novel workflow called NeoMS, designed to efficiently identify both non-mutated and mutated MHC-I peptides from mass spectrometry data. NeoMS utilizes a tagging algorithm to generate an expanded sequence database that includes potential mutated proteins for each sample. Furthermore, it employs a machine learning-based scoring function for each peptide-spectrum match (PSM) to maximize search sensitivity. Finally, a rigorous target-decoy approach is implemented to control the false discovery rates (FDR) of the peptides with and without mutations separately. Experimental results for regular peptides demonstrate that NeoMS outperforms four benchmark methods. For mutated peptides, NeoMS successfully identifies hundreds of high-quality mutated peptides in a melanoma-associated sample, with their validity confirmed by further studies. Shaokai Wang, Bin Ma 0002 |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2025 | Anti-Cancer Peptides Identification and Activity Type Classification With Protein Sequence Pre-TrainingabstractCancer remains a significant global health challenge, responsible for millions of deaths annually. Addressing this issue necessitates the discovery of novel anti-cancer drugs. Anti-cancer peptides (ACPs), with their unique ability to selectively target cancer cells, offer new hope in discovering low side-effect anti-cancer drugs. However, the process of discovering novel ACPs is both time-consuming and costly. Therefore, there is an urgent need for a computational method that can predict whether a given peptide is an ACP and classify its specific functional types. In this paper, we introduce DUO-ACP, a model serving dual roles in ACP prediction: identification and functional type classification. DUO-ACP employs two embedding modules to acquire knowledge about global protein features and local ACP characteristics, complemented by a prediction module. When assessed on two publicly available datasets for each task, DUO-ACP surpasses all existing methods, achieving outstanding results: an ACP identification accuracy of 89.5% and a Macro-averaged AUC of 88.6% in ACP functional type classification. We further interpret the contribution of each part of our model, including the two types of embeddings as well as ensemble learning. On a new curated dataset, the prediction results of DUO-ACP closely match existing literature, highlighting DUO-ACP's generalization capabilities on previously unseen data and displaying the potential capability of discovering novel ACP. Shaokai Wang, Bin Ma 0002 |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | AGPred: An End-to-End Deep Learning Model for Predicting Drug Approvals in Clinical Trials Based on Molecular FeaturesabstractOne of the major challenges in drug development is maintaining acceptable levels of efficacy and safety throughout the various stages of clinical trials and successfully bringing the drug to market. However, clinical trials are time-consuming and expensive. While there are computational methods designed to predict the likelihood of a drug passing clinical trials and reaching the market, these methods heavily rely on manual feature engineering and cannot automatically learn drug molecular representations, resulting in relatively low model performance. In this study, we propose AGPred, an attention-based deep Graph Neural Network (GNN) designed to predict drug approval rates in clinical trials accurately. Unlike the few existing studies on drug approval prediction, which only use predicted targets of compounds, our novel approach employs a GNN module to extract high-potential features of compounds based on their molecular graphs. Additionally, a cross-attention-based fusion module is utilized to learn molecular fingerprint features, enhancing the model's representation of chemical structures. Meanwhile, AGPred integrates the physicochemical properties of drugs to provide a comprehensive description of the molecules. Experimental results indicate that AGPred outperforms four state-of-the-art models on both benchmark and independent datasets. The study also includes several ablation experiments and visual analyses to demonstrate the effectiveness of our method in predicting drug approval during clinical trials. Chenliang Xie, Shaokai Wang |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | Novel Fine-Tuning Strategy on Pre-trained Protein Model Enhances ACP Functional Type Classification
Shaokai Wang, Bin Ma 0002 |
ISBRA (1) | 1 |
| 2024 | PrSMBooster: Improving the Accuracy of Top-Down Proteoform Characterization Using Deep Learning Rescoring Models
Jiancheng Zhong, Maoqi Yuan, Shaokai Wang |
ISBRA (3) | 4 |
| 2023 | Deep learning boosted amyloidosis diagnosisabstractAmyloid light chain (AL) amyloidosis is a disorder characterized by the deposition of antibody light chains in organs. Early and accurate diagnosis of AL amyloidosis is crucial for timely implementation of appropriate treatment strategies. However, existing computational methods for predicting AL amyloidosis often heavily rely on manually extracted features and their performance is less than satisfactory. In this study, we introduce DeepAL, a deep learning-based approach designed to predict AL amyloidosis with high precision. DeepAL utilizes a pre-trained model to extract light chain features and is then fine-tuned with AL amyloidosis knowledge. On two benchmark datasets, DeepAL achieved impressive results with area under the ROC curves (AUCs) of 0.9072 and 0.8919, outperforming previous approaches. Our ablation study shows the use of the pre-trained model can significantly improve identification performance. The code is available at https://github.com/waterlooms/DeepAL. Shaokai Wang, Bin Ma 0002 |
BIBM | 1 |
| 2023 | NeoMS: Identification of Novel MHC-I Peptides with Tandem Mass Spectrometry
Shaokai Wang, Bin Ma 0002 |
ISBRA | 1 |
| 2023 | Prior Knowledge Constrained Adaptive Graph Framework for Partial Label LearningabstractPartial label learning (PLL) aims to learn a robust multi-class classifier from the ambiguous data, where each instance is given with several candidate labels, among which only one label is real. Most existing methods usually cope with such problem by utilizing a feature similarity graph to conduct label disambiguation. However, these methods construct the feature graph by only employing original features, while the influences of latent outliers and the contributions of label space are regrettably ignored. To tackle these issues, in this article, we propose aPrior KnOwledge ConsTrainedAdaptiveGraph FramEwork (POTAGE) for partial label learning, which utilizes an adaptive graph fused with label information to accurately describe the instance relationship and guide the desired model training. Compared with the feature-induced fixed graph, the adaptive graph is deemed to be more robust and accurate to reveal the intrinsic manifold structure within the data, and the embedding label information is expected to effectively alleviate the label ambiguities and enlarge the gap of label confidences between two instances from different classes. Extensive experiments demonstrate that POTAGE achieves state-of-the-art performance. Gengyu Lyu, Songhe Feng, Shaokai Wang, Zhen Yang 0004 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2022 | Partial label learning with noisy side information
Shaokai Wang, Mingxuan Xia, Gengyu Lyu, Songhe Feng |
Appl. Intell. | 1 |
| 2022 | SADeepcry: a deep learning framework for protein crystallization propensity prediction using self-attention and auto-encoder networksabstractThe X-ray diffraction (XRD) technique based on crystallography is the main experimental method to analyze the three-dimensional structure of proteins. The production process of protein crystals on which the XRD technique relies has undergone multiple experimental steps, which requires a lot of manpower and material resources. In addition, studies have shown that not all proteins can form crystals under experimental conditions, and the success rate of the final crystallization of proteins is only <10%. Although some protein crystallization predictors have been developed, not many tools capable of predicting multi-stage protein crystallization propensity are available and the accuracy of these tools is not satisfactory. In this paper, we propose a novel deep learning framework, named SADeepcry, for predicting protein crystallization propensity. The framework can be used to estimate the three steps (protein material production, purification and crystallization) in protein crystallization experiments and the success rate of the final protein crystallization. SADeepcry uses the optimized self-attention and auto-encoder modules to extract sequence, structure and physicochemical features from the proteins. Compared with other state-of-the-art protein crystallization propensity prediction models, SADeepcry can obtain more complex global spatial long-distance dependence of protein sequence information. Our computational results show that SADeepcry has increased Matthews correlation coefficient and area under the curve, by 100.3% and 13.4%, respectively, over the DCFCrystal method on the benchmark dataset. The codes of SADeepcry are available at https://github.com/zhc940702/SADeepcry. Shaokai Wang |
Briefings Bioinform. | 1 |
| 2022 | A similarity-based deep learning approach for determining the frequencies of drug side effectsabstractThe side effects of drugs present growing concern attention in the healthcare system. Accurately identifying the side effects of drugs is very important for drug development and risk assessment. Some computational models have been developed to predict the potential side effects of drugs and provided satisfactory performance. However, most existing methods can only predict whether side effects will occur and cannot determine the frequency of side effects. Although a few existing methods can predict the frequency of drug side effects, they strongly depend on the known drug-side effect relationships. Therefore, they cannot be applied to new drugs without known side effect frequency information. In this paper, we develop a novel similarity-based deep learning method, named SDPred, for determining the frequencies of drug side effects. Compared with the existing state-of-the-art models, SDPred integrates rich features and can be applied to predict the side effect frequencies of new drugs without any known drug-side effect association or frequency information. To our knowledge, this is the first work that can predict the side effect frequencies of new drugs in the population. The comparison results indicate that SDPred is much superior to all previously reported models. In addition, some case studies also demonstrate the effectiveness of our proposed method in practical applications. The SDPred software and data are freely available at https://github.com/zhc940702/SDPred, https://zenodo.org/record/5112573 and https://hub.docker.com/r/zhc940702/sdpred. Shaokai Wang, Kai Zheng 0020, Qichang Zhao, Feng Zhu 0004, Jianxin Wang 0001 |
Briefings Bioinform. | 2 |
| 2020 | MultiGuideScan: a multi-processing tool for designing CRISPR guide RNA librariesabstractSUMMARY: The recent advance in genome engineering technologies based on CRISPR/Cas9 system is enabling people to systematically understand genomic functions. A short RNA string (the CRISPR guide RNA) can guide the Cas9 endonuclease to specific locations in complex genomes to cut DNA double-strands. The CRISPR guide RNA is essential for gene editing systems. Recently, the GuideScan software is developed to design CRISPR guide RNA libraries, which can be used for genome editing of coding and non-coding genomic regions effectively. However, GuideScan is a serial program and computationally expensive for designing CRISPR guide RNA libraries from large genomes. Here, we present an efficient guide RNA library designing tool (MultiGuideScan) by implementing multiple processes of GuideScan. MultiGuideScan speeds up the guide RNA library designing about 9-12 times on a 32-process mode comparing to GuideScan. MultiGuideScan makes it possible to design guide RNA libraries from large genomes. AVAILABILITY AND IMPLEMENTATION: MULTIGUIDESCAN IS AVAILABLE AT GITHUB: https://github.com/bioinfomaticsCSU/MultiGuideScan. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tao Li 0033, Shaokai Wang, Feng Luo 0001, Fang-Xiang Wu, Jianxin Wang 0001 |
Bioinform. | 2 |
| 2018 | A secure data collection scheme based on compressive sensing in wireless sensor networks
Shaokai Wang, Kehua Guo, Jianxin Wang 0001 |
Ad Hoc Networks | 2 |
| 2018 | Computational drug repositioning using low-rank matrix approximation and randomized algorithmsabstractMotivation: Computational drug repositioning is an important and efficient approach towards identifying novel treatments for diseases in drug discovery. The emergence of large-scale, heterogeneous biological and biomedical datasets has provided an unprecedented opportunity for developing computational drug repositioning methods. The drug repositioning problem can be modeled as a recommendation system that recommends novel treatments based on known drug-disease associations. The formulation under this recommendation system is matrix completion, assuming that the hidden factors contributing to drug-disease associations are highly correlated and thus the corresponding data matrix is low-rank. Under this assumption, the matrix completion algorithm fills out the unknown entries in the drug-disease matrix by constructing a low-rank matrix approximation, where new drug-disease associations having not been validated can be screened. Results: In this work, we propose a drug repositioning recommendation system (DRRS) to predict novel drug indications by integrating related data sources and validated information of drugs and diseases. Firstly, we construct a heterogeneous drug-disease interaction network by integrating drug-drug, disease-disease and drug-disease networks. The heterogeneous network is represented by a large drug-disease adjacency matrix, whose entries include drug pairs, disease pairs, known drug-disease interaction pairs and unknown drug-disease pairs. Then, we adopt a fast Singular Value Thresholding (SVT) algorithm to complete the drug-disease adjacency matrix with predicted scores for unknown drug-disease pairs. The comprehensive experimental results show that DRRS improves the prediction accuracy compared with the other state-of-the-art approaches. In addition, case studies for several selected drugs further demonstrate the practical usefulness of the proposed method. Availability and implementation: http://bioinformatics.csu.edu.cn/resources/softs/DrugRepositioning/DRRS/index.html. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Huimin Luo, Min Li 0007, Shaokai Wang, Yaohang Li, Jianxin Wang 0001 |
Bioinform. | 3 |
| 2018 | Multi-attribute and relational learning via hypergraph regularized generative model
Shaokai Wang, Xutao Li 0003, Yunming Ye, Xiaohui Huang 0003, Yan Li 0040 |
Neurocomputing | 1 |
| 2018 | Parameterized counting matching and packing: A family of hard problems that admit FPTRAS
Yunlong Liu 0001, Shaokai Wang, Jianxin Wang 0001 |
Theor. Comput. Sci. | 2 |
| 2017 | A generative model with hypergraph regularizers for protein function predictionabstractHeterogeneous data sources and multi-label are two important characteristics of protein function prediction. They describe protein data from two different aspects. However, it is of considerable challenge to integrate multiple data sources and multi-label simultaneously for predicting protein functions, especially when there are only a limited number of labeled proteins. In this paper, we propose a generative model with hypergraph regularizers algorithm, called GMHR, for predicting proteins with multiple functions. The GMHR algorithm integrates all data sources that are available, including protein attribute features, interaction networks, label correlations, and unlabeled data. Experimental results on the real-world datasets predicting the functions of proteins demonstrate the superiority of our proposed method compared with the state-of-the-art baselines. Shaokai Wang, Xutao Li 0003, Yunming Ye, Yan Li 0040, Xiaohui Huang 0003, Xiaolin Du |
IJCNN | 1 |
| 2017 | Multi-view learning via multiple graph regularized generative model
Shaokai Wang, Ke Wang 0068, Xutao Li 0003, Yunming Ye, Raymond Y. K. Lau, Xiaolin Du |
Knowl. Based Syst. | 1 |
| 2017 | Semi-supervised Collective Classification in Multi-attribute Network Data
Shaokai Wang, Yunming Ye, Xutao Li 0003, Xiaohui Huang 0003, Raymond Y. K. Lau |
Neural Process. Lett. | 1 |
| 2016 | Time series k-means: A new k-means type smooth subspace clustering for time series data
Xiaohui Huang 0003, Yunming Ye, Liyan Xiong, Raymond Y. K. Lau, Nan Jiang 0013, Shaokai Wang |
Inf. Sci. | 6 |
| 2016 | Clustering time-stamped data using multiple nonnegative matrices factorization
Xiaohui Huang 0003, Yunming Ye, Liyan Xiong, Shaokai Wang, Xiaofei Yang 0002 |
Knowl. Based Syst. | 4 |
| 2015 | A Generative Model with Ensemble Manifold Regularization for Multi-view Clustering
Shaokai Wang, Yunming Ye, Raymond Y. K. Lau |
ICIC (3) | 1 |