VLDB 2026 Research / reviewers in the wild / expert
Shutao Chen
dblp:89/7838
· DBLP profile ↗
13ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-6188-2631ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Cheap Lunch: Synthetic Annotation With Reduced Human Effort for Medical Text Mining
Shutao Chen, Piek Vossen |
LREC | 1 |
| 2026 | PepLM-GNN: A graph neural network framework leveraging pre-trained language models for peptide-protein binding predictionabstractMOTIVATION: The precise prediction of peptide-protein interaction (PepPI) is a core support for promoting breakthroughs in peptide drug research, as well as understanding the regulatory mechanisms of biomolecules. Researchers have developed several computational methods to predict PepPI. However, existing computational methods also have significant limitations. At the level of data feature characterisation, the problem of PepPI does not conform to the Euclidean axioms, making it difficult for conventional prediction methods to effectively measure the underlying correlations between peptides and proteins. At the level of model generalisation performance, existing approaches are often hampered by insufficient generalisation ability, as manifested by their markedly degraded performance in cold start scenarios involving novel peptides, novel proteins, and novel binding pairs. RESULTS: In this study, we propose a computing framework, PepLM-GNN, that integrates a pre-trained language ProtT5 model with a hybrid graph network for accurate identification of PepPI. This model constructs a graph by using ProtT5-extracted semantic context features of peptides and proteins to form heterogeneous nodes, with edges connecting interacting peptide-protein pairs. The hybrid graph network Graph Convolutional Networks (GCN) provides the comprehensive information of the peptide and protein sequences, while employing the Graph Isomorphism Network (GIN) to capture the global interactions between them. Specifically, the GCN aggregates both the semantic context information of node sequences and local neighbourhood information, effectively representing non-Euclidean data. To capture the global associations, we adopt a GIN strategy to optimize the cross-node feature interaction and transfer process, thereby enhancing the generalisation performance of addressing the cold start scenario. Compared with the existing advanced methods, PepLM-GNN demonstrated highly accurate performance and robustness in predicting the PepPI. We further demonstrated the capabilities of PepLM-GNN in virtual peptide drug screening, which is expected to facilitate the discovery of peptide drugs and the elucidation of protein functions. Ke Yan 0003, Meijing Li, Shutao Chen, Bin Liu 0014 |
PLoS Comput. Biol. | 3 |
| 2026 | DSCA-HLAII: A dual-stream cross-attention model for predicting peptide-HLA class II interaction and presentationabstractMOTIVATION: The interaction between peptides and human leukocyte antigen class II (HLA-II) molecules plays a pivotal role in adaptive immune responses, as HLA-II mediates the recognition of exogenous antigens and initiates T cell activation through peptide presentation. Accurate prediction of peptide-HLA-II binding serves as a cornerstone for deciphering cellular immune responses, and is essential for guiding the optimization of antibody therapeutics. Researchers have developed several computational approaches to identify peptide-HLA-II interaction and presentation. However, most computational approaches exhibit inconsistent predictive performance, poor generalization ability and limited biological interpretability. RESULTS: In this study, we present DSCA-HLAII, a novel predictive framework for peptide-HLA-II interactions and presentation based on a dual-stream cross-attention architecture. The framework proposes a dual-stream cross-attention (DSCA) mechanism to integrate pre-trained semantic embedding ESMC with sequence-level ONE-HOT features. The DSCA mechanism effectively models the interaction dynamics between peptides and HLA-II molecules, enabling the precise identification of key binding sites. Experimental results demonstrate that DSCA-HLAII consistently surpasses existing state-of-the-art approaches, demonstrating high accuracy and robustness in predicting peptide-HLA-II interactions and presentation. We further demonstrate the capability of DSCA-HLAII for predicting peptide binding cores and assessing antibody immunogenicity, which is expected to advance artificial intelligence-based peptide drug discovery. Shutao Chen, Alexey K. Shaytan, Youyu Wang |
PLoS Comput. Biol. | 3 |
| 2025 | FusionEncoder: identification of intrinsically disordered regions based on multi-feature fusionabstractMOTIVATION: Intrinsic disorder regions (IDRs) play a significant role in diverse biological processes and are widely distributed in proteins. Thus, accurately predicting these regions is essential for analyzing protein structure and function. Amino acid feature extraction servers as a foundational process in the development of computational predictive models. Existing methods typically rely on traditional biological features (e.g. PSSM) or use pre-trained protein language models (PPLMs) to capture sequence semantic information, often resorting to straightforward feature concatenation. However, these approaches fail to capture the multi-semantic interactions between traditional biological features and PPLMs-based features. RESULTS: In this study, we propose a method named FusionEncoder designed for the integration of traditional biological and PPLMs-based features of the protein. FusionEncoder is a fusion network built on a variant of long short-term memory (LSTM). We consider traditional biological features and PPLMs-based features to be two types of semantic inputs within a "multi-semantic" space. Traditional features are input into the cell state of the LSTM, while PPLMs-based features are fed into the input part. A fusion cell is then utilized to fuse these two types of features. This strategy leverages the capability of LSTM to encode long sequences, enhancing context-aware semantic learning of amino acid sequences. Finally, a transformer-based encoder layer is employed to predict the IDRs. Evaluation on four independent test datasets indicate that FusionEncoder obviously improves the accuracy of amino acid feature representation and achieves superior performance compared to the other existing methods. AVAILABILITY AND IMPLEMENTATION: To facilitate accessibility for experimental researchers, a user-friendly and publicly available webserver for the FusionEncoder predictor has been deployed at http://bliulab.net/FusionEncoder/. FusionEncoder is expected to serve as a valuable tool for the accurate identification of IDRs. Sicen Liu, Shutao Chen, Bin Liu 0014 |
Bioinform. | 2 |
| 2025 | Accurate prediction of toxicity peptide and its function using multi-view tensor learning and latent semantic learning frameworkabstractMOTIVATION: Therapeutic peptide is an important ingredient in the treatment of various diseases and drug discovery. The toxicity of peptides is one of the major challenges in peptide drug therapy. With the abundance of therapeutic peptides generated in the post-genomics era, it is a challenge to promptly identify toxicity peptides using computational methods. Although several efforts have been made, few algorithms are designed to identify whether a query peptide exhibits toxicity. Considering the varied levels of biological activities, the toxicity peptides should be further classified into multi-functional peptides. RESULTS: This study introduces a two-level predictor, ToxPre-2L, developed using the multi-view tensor learning and latent semantic learning framework. The proposed method utilized multi-label learning with feature induced labels to avoid the redundancy of information from each view. Then the multi-view tensor learning was employed to establish the latent semantic information among different views, while low-rank constraint learning was leveraged to exploit the correlation information among multi-labels. Finally, we constructed an updated toxicity peptide benchmark dataset to assess the effectiveness of the proposed method. Experimental results demonstrated that ToxPre-2L achieves a better performance than alternative computational methods in the prediction of toxicity peptides and their multi-functional types. AVAILABILITY AND IMPLEMENTATION: The source code and data of ToxPre-2L can be accessed at http://bliulab.net/ToxPre-2L. Ke Yan 0003, Shutao Chen, Bin Liu 0014, Hao Wu 0066 |
Bioinform. | 2 |
| 2025 | Protein Language Pragmatic Analysis and Progressive Transfer Learning for Profiling Peptide-Protein InteractionsabstractProtein complex structural data are growing at an unprecedented pace, but its complexity and diversity pose significant challenges for protein function research. Although deep learning models have been widely used to capture the syntactic structure, word semantics, or semantic meanings of polypeptide and protein sequences, these models often overlook the complex contextual information of sequences. Here, we propose interpretable interaction deep learning (IIDL)-peptide-protein interaction (PepPI), a deep learning model designed to tackle these challenges using data-driven and interpretable pragmatic analysis to profile PepPIs. IIDL-PepPI constructs bidirectional attention modules to represent the contextual information of peptides and proteins, enabling pragmatic analysis. It then adopts a progressive transfer learning framework to simultaneously predict PepPIs and identify binding residues for specific interactions, providing a solution for multilevel in-depth profiling. We validate the performance and robustness of IIDL-PepPI in accurately predicting peptide-protein binary interactions and identifying binding residues compared with the state-of-the-art methods. We further demonstrate the capability of IIDL-PepPI in peptide virtual drug screening and binding affinity assessment, which is expected to advance artificial intelligence-based peptide drug discovery and protein function elucidation. Shutao Chen, Ke Yan 0003, Xuelong Li 0001, Bin Liu 0014 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | TPpred-SC: multi-functional therapeutic peptide prediction based on multi-label supervised contrastive learning
Ke Yan 0003, Hongwu Lv, Jiangyi Shao, Shutao Chen, Bin Liu 0014 |
Sci. China Inf. Sci. | 4 |
| 2023 | FGFICA: Independent Component Analysis of Fusion Genomic Features for Mining Epi-Transcriptome Profiling DataabstractA. Shutao Chen, Lin Zhang 0015, Xiangzhi Chen 0003, Hui Liu 0024 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | FBCwPlaid: A Functional Biclustering Analysis of Epi-Transcriptome Profiling Data Via a Weighted Plaid ModelabstractRecent studies have shown that in-depth studies on epi-transcriptomic patterns of N6-methyladenosine (m6A) may help understand its complex functions and co-regulatory mechanisms. Since most biclustering algorithms are developed in scenarios of gene expression analysis, which does not share the same characteristics with m6A methylation profile, we propose a weighted Plaid biclustering model (FBCwPlaid) based on the Lagrange multiplier method to discover the potential functional patterns. Each pattern is achieved by minimizing approximation error between FBCwPlaid predicted value and real data. To address the issue that site expression level determines methylation level confidence, it uses RNA expression levels of each site as weights to make lower expressed sites less confident. FBCwPlaid also allows overlapping biclusters, indicating some sites may participate in multiple biological functions. FBCwPlaid was then applied on MeRIP-Seq data of 69,446 methylation sites under 32 experimental conditions, each of which represented a stimulus to a particular cell line or environment. Finally, three patterns were discovered, and further pathway analysis and enzyme specificity test showed that sites involved in each pattern are highly relevant to m6A methyltransferases. Further detailed analyses showed that some patterns are condition-specific, indicating that some specific sites’ methylation profiles may occur in specific cell lines or conditions. Shutao Chen, Lin Zhang 0015, Jia Meng 0001, Hui Liu 0024 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | BDBB: A Novel Beta-Distribution-Based Biclustering Algorithm for Revealing Local Co-Methylation Patterns in Epi-Transcriptome Profiling DataabstractN6-methyladenosine (m6A) has been shown to play crucial roles in RNA metabolism, physiology, and pathological processes. However, the specific regulatory mechanisms of most methylation sites remain uncharted due to the complexity of life processes. Biological experimental methods are costly to solve this problem, and computational methods are relatively lacking. The discovery of local co-methylation patterns (LCPs) of m6A epi-transcriptome data can benefit to solve the above problems. Based on this, we propose a novel biclustering algorithm based on the beta distribution (BDBB), which realizes the mining of LCPs of m6A epi-transcriptome data. BDBB employs the Gibbs sampling method to complete parameter estimation. In the process of modeling, LCPs are recognized as sharp beta distributions compared to the background distribution. Simulation study showed BDBB can extract all the three actual LCPs implanted in the background data and the overlap conditions between them with considerable accuracy (almost close to 100%). On MeRIP-Seq data of 69,446 methylation sites under 32 experimental conditions from 10 human cell lines, BDBB unveiled two LCPs, and Gene Ontology (GO) enrichment analysis showed that they were enriched in histone modification and embryo development, etc. important biological processes respectively. The GOE_Score scoring indicated that the biclustering results of BDBB in the m6A epi-transcriptome data are more biologically meaningful than the results of other biclustering algorithms. Zhaoyang Liu 0002, Yuteng Xiao, Hongsheng Yin 0001, Shutao Chen, Kaijian Xia, Lin Zhang 0015 |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | EDLm6APred: ensemble deep learning approach for mRNA m6A site predictionabstractAbstract Background As a common and abundant RNA methylation modification, N6-methyladenosine (m6A) is widely spread in various species' transcriptomes, and it is closely related to the occurrence and development of various life processes and diseases. Thus, accurate identification of m6A methylation sites has become a hot topic. Most biological methods rely on high-throughput sequencing technology, which places great demands on the sequencing library preparation and data analysis. Thus, various machine learning methods have been proposed to extract various types of features based on sequences, then occupied conventional classifiers, such as SVM, RF, etc., for m6A methylation site identification. However, the identification performance relies heavily on the extracted features, which still need to be improved. Results This paper mainly studies feature extraction and classification of m6A methylation sites in a natural language processing way, which manages to organically integrate the feature extraction and classification simultaneously, with consideration of upstream and downstream information of m6A sites. One-hot, RNA word embedding, and Word2vec are adopted to depict sites from the perspectives of the base as well as its upstream and downstream sequence. The BiLSTM model, a well-known sequence model, was then constructed to discriminate the sequences with potential m6A sites. Since the above-mentioned three feature extraction methods focus on different perspectives of m6A sites, an ensemble deep learning predictor (EDLm6APred) was finally constructed for m6A site prediction. Experimental results on human and mouse data sets show that EDLm6APred outperforms the other single ones, indicating that base, upstream, and downstream information are all essential for m6A site detection. Compared with the existing m6A methylation site prediction models without genomic features, EDLm6APred obtains 86.6% of the area under receiver operating curve on the human data sets, indicating the effectiveness of sequential modeling on RNA. To maximize user convenience, a webserver was developed as an implementation of EDLm6APred and made publicly available at www.xjtlu.edu.cn/biologicalsciences/EDLm6APred . Conclusions Our proposed EDLm6APred method is a reliable predictor for m6A methylation sites. Lin Zhang 0015, Gangshen Li, Xiuyu Li, Shutao Chen, Hui Liu 0024 |
BMC Bioinform. | 5 |
| 2020 | REW-ISA: unveiling local functional blocks in epi-transcriptome profiling data via an RNA expression-weighted iterative signature algorithmabstractAbstract Background Recent studies have shown that N6-methyladenosine (m6A) plays a critical role in numbers of biological processes and complex human diseases. However, the regulatory mechanisms of most methylation sites remain uncharted. Thus, in-depth study of the epi-transcriptomic patterns of m6A may provide insights into its complex functional and regulatory mechanisms. Results Due to the high economic and time cost of wet experimental methods, revealing methylation patterns through computational models has become a more preferable way, and drawn more and more attention. Considering the theoretical basics and applications of conventional clustering methods, an RNA Expression Weighted Iterative Signature Algorithm (REW-ISA) is proposed to find potential local functional blocks (LFBs) based on MeRIP-Seq data, where sites are hyper-methylated or hypo-methylated simultaneously across the specific conditions. REW-ISA adopts RNA expression levels of each site as weights to make sites of lower expression level less significant. It starts from random sets of sites, then follows iterative search strategies by thresholds of rows and columns to find the LFBs in m6A methylation profile. Its application on MeRIP-Seq data of 69,446 methylation sites under 32 experimental conditions unveiled 6 LFBs, which achieve higher enrichment scores than ISA. Pathway analysis and enzyme specificity test showed that sites remained in LFBs are highly relevant to the m6A methyltransferase, such as METTL3, METTL14, WTAP and KIAA1429. Further detailed analyses for each LFB even showed that some LFBs are condition-specific, indicating that methylation profiles of some specific sites may be condition relevant. Conclusions REW-ISA finds potential local functional patterns presented in m6A profiles, where sites are co-methylated under specific conditions. Lin Zhang 0015, Shutao Chen, Jia Meng 0001, Hui Liu 0024 |
BMC Bioinform. | 2 |
| 2019 | Effects of Base-Frequency and Spectral Envelope on Deep-Learning Speech Separation and Recognition Models
Jun Hui, Shutao Chen, Richard H. Y. So |
INTERSPEECH | 3 |