EDBT 2026 Demo / reviewers in the wild / expert
Xiaoyu Wang 0016
dblp:401/0651
· DBLP profile ↗
14ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Comprehensive review and assessment of multi-species splicing variant prediction: task-specific deep learning models and genomic foundation modelsabstractAlternative splicing generates transcriptomic and proteomic diversity essential for eukaryotic complexity, yet genetic variants disrupting the splicing code underlie numerous human diseases. Deep learning (DL) models and genomic foundation models (GFMs) have achieved outstanding accuracy for predicting splicing variant effects in humans. However, their transferability to non-human species remains poorly understood, limiting applications in agricultural genomics, comparative biology, and non-model organism research, where experimentally validated variant datasets are limited or lacking. In this study, we comprehensively reviewed 35 computational approaches in terms of their architectural characteristics for splicing site and variant prediction and analysis. We systematically benchmarked the performance of 10 representative models for splicing variant prediction across human, rat, pig, and chicken, including four task-specific DL models and six GFMs, using our manually assembled benchmark datasets. Our benchmarking results revealed a substantial cross-species performance decrease (~21%-33% in the area under the receiver operating characteristic curve - AUROC) using task-specific models from human to non-human species datasets. We then applied a supervised adaptation to frozen GFM embeddings (DNABERT-2, Evo 2, Genos) by adding a lightweight classifier (i.e. a multi-layer perceptron) and reduced the cross-species performance decrease for rat and pig (8.56%-23.84% in AUROC), while performance on chicken was very close to human (decline within 1%, even exceeding by 0.52% when using the Evo 2 embedding). We proposed several directions to improve the prediction performance of splicing variants, including feature representation transfer and multi-modal fusion integrating global context, universal embeddings, and species-aware conditioning. We hope our comprehensive review and performance benchmarking can provide useful computational insights for further advancement of splicing variant prediction. Yinuo Sun, Xiaoyu Wang 0016, Yuheng Jia, Seiya Imoto, Fuyi Li, Chen Li 0021, Jiangning Song |
Briefings Bioinform. | 2 |
| 2025 | CoPRA: Bridging Cross-domain Pretrained Sequence Models with Complex Structures for Protein-RNA Binding Affinity PredictionabstractAccurately measuring protein-RNA binding affinity is crucial in many biological processes and drug design. Previous computational methods for protein-RNA binding affinity prediction rely on either sequence or structure features, unable to capture the binding mechanisms comprehensively. The recent emerging pre-trained language models trained on massive unsupervised sequences of protein and RNA have shown strong representation ability for various in-domain downstream tasks, including binding site prediction. However, applying different-domain language models collaboratively for complex-level tasks remains unexplored. In this paper, we propose CoPRA to bridge pre-trained language models from different biological domains via Complex structure for Protein-RNA binding Affinity prediction. We demonstrate for the first time that cross-biological modal language models can collaborate to improve binding affinity prediction. We propose a Co-Former to combine the cross-modal sequence and structure information and a bi-scope pre-training strategy for improving Co-Former's interaction understanding. Meanwhile, we build the largest protein-RNA binding affinity dataset PRA310 for performance evaluation. We also test our model on a public dataset for mutation effect prediction. CoPRA reaches state-of-the-art performance on all the datasets. We provide extensive analyses and verify that CoPRA can (1) accurately predict the protein-RNA binding affinity; (2) understand the binding affinity change caused by mutations; and (3) benefit from scaling data and model size. Xiaohong Liu 0007, Tong Pan, Jing Xu 0008, Xiaoyu Wang 0016, Wuyang Lan, Jiangning Song, Ting Chen 0006 |
AAAI | 5 |
| 2025 | Multimodal geometric learning for antimicrobial peptide identification by leveraging alphafold2-predicted structures and surface featuresabstractAntimicrobial peptides (AMPs) are short peptides that play critical roles in diverse biological processes and exhibit functional activities against target organisms. While numerous methods have demonstrated the effectiveness of deep neural networks for AMP identification using sequence features; nevertheless, higher-level peptide characteristics-such as 3D structure and geometric surface features-have not been comprehensively explored. To address this gap, we introduce the SSFGM-Model (Sequence, Structure, Surface, Graph, and Geometric-based Model), a novel framework that integrates multiple feature types to enhance AMP identification. The model represents each peptide sequence as a graph, where nodes are characterized by amino acid features derived from ProteinBERT, ESM-2, and One-hot embeddings. Graph convolutional networks and an attention mechanism are employed to capture high-order structural and sequential relationships. Additionally, surface geometry and physicochemical properties are processed using a geometric neural network. Finally, a feature fusion strategy combines the outputs from these subnetworks to enable robust AMP identification. Extensive benchmarking experiments demonstrate that the SSFGM-Model outperforms current state-of-the-art methods. An ablation study further confirms the critical role of sequence, structural, and surface features in AMP identification. The key contribution of this work is the innovative integration of multiple levels of peptide characteristics and the combination of geometric and graph neural networks. This approach provides a more comprehensive understanding of the sequence-structure-function relationship of peptides, paving the way for more accurate AMP prediction. The SSFGM-Model has a significant potential for applications in the discovery and design of novel AMP-based therapeutics. The source code is publicly available at https://github.com/ggcameronnogg/SSFGM-Model. Zehua Sun, Jing Xu 0008, Zhikang Wang, Xiaoyu Wang 0016, Shanshan Li 0008, Yuming Guo 0001, Hsin Hui Shen, Jiangning Song |
Briefings Bioinform. | 6 |
| 2025 | MORE: a multi-omics data-driven hypergraph integration network for biomedical data classification and biomarker identificationabstractHigh-throughput sequencing methods have brought about a huge change in omics-based biomedical study. Integrating various omics data is possibly useful for identifying some correlations across data modalities, thus improving our understanding of the underlying biological mechanisms and complexity. Nevertheless, most existing graph-based feature extraction methods overlook the complementary information and correlations across modalities. Moreover, these methods tend to treat the features of each omics modality equally, which contradicts current biological principles. To solve these challenges, we introduce a novel approach for integrating multi-omics data termed Multi-Omics hypeRgraph integration nEtwork (MORE). MORE initially constructs a comprehensive hyperedge group by extensively investigating the informative correlations within and across modalities. Subsequently, the multi-omics hypergraph encoding module is employed to learn the enriched omics-specific information. Afterward, the multi-omics self-attention mechanism is then utilized to adaptatively aggregate valuable correlations across modalities for representation learning and making the final prediction. We assess MORE's performance on datasets characterized by message RNA (mRNA) expression, Deoxyribonucleic Acid (DNA) methylation, and microRNA (miRNA) expression for Alzheimer's disease, invasive breast carcinoma, and glioblastoma. The results from three classification tasks highlight the competitive advantage of MORE in contrast with current state-of-the-art (SOTA) methods. Moreover, the results also show that MORE has the capability to identify a greater variety of disease-related biomarkers compared to existing methods, highlighting its advantages in biomedical data mining and interpretation. Overall, MORE can be investigated as a valuable tool for facilitating multi-omics analysis and novel biomarker discovery. Our code and data can be publicly accessed at https://github.com/Wangyuhanxx/MORE. Zhikang Wang, Xiaoyu Wang 0016, Jiangning Song, Dongjun Yu, Fang Ge |
Briefings Bioinform. | 4 |
| 2024 | Deep learning approaches for non-coding genetic variant effect prediction: current progress and future prospectsabstractRecent advancements in high-throughput sequencing technologies have significantly enhanced our ability to unravel the intricacies of gene regulatory processes. A critical challenge in this endeavor is the identification of variant effects, a key factor in comprehending the mechanisms underlying gene regulation. Non-coding variants, constituting over 90% of all variants, have garnered increasing attention in recent years. The exploration of gene variant impacts and regulatory mechanisms has spurred the development of various deep learning approaches, providing new insights into the global regulatory landscape through the analysis of extensive genetic data. Here, we provide a comprehensive overview of the development of the non-coding variants models based on bulk and single-cell sequencing data and their model-based interpretation and downstream tasks. This review delineates the popular sequencing technologies for epigenetic profiling and deep learning approaches for discerning the effects of non-coding variants. Additionally, we summarize the limitations of current approaches in variant effect prediction research and outline opportunities for improvement. We anticipate that our study will offer a practical and useful guide for the bioinformatic community to further advance the unraveling of genetic variant effects. Xiaoyu Wang 0016, Fuyi Li, Seiya Imoto, Hsin-Hui Shen, Shanshan Li 0008, Yuming Guo 0001, Jian Yang 0005, Jiangning Song |
Briefings Bioinform. | 1 |
| 2024 | MLSNet: a deep learning model for predicting transcription factor binding sitesabstractAccurate prediction of transcription factor binding sites (TFBSs) is essential for understanding gene regulation mechanisms and the etiology of diseases. Despite numerous advances in deep learning for predicting TFBSs, their performance can still be enhanced. In this study, we propose MLSNet, a novel deep learning architecture designed specifically to predict TFBSs. MLSNet innovatively integrates multisize convolutional fusion with long short-term memory (LSTM) networks to effectively capture DNA-sparse higher-order sequence features. Further, MLSNet incorporates super token attention and Bi-LSTM to systematically extract and integrate higher-order DNA shape features. Experimental results on 165 ChIP-seq (chromatin immunoprecipitation followed by sequencing) datasets indicate that MLSNet consistently outperforms several state-of-the-art algorithms in the prediction of TFBSs. Specifically, MLSNet reports average metrics: 0.8306 for ACC, 0.8992 for AUROC, and 0.9035 for AUPRC, surpassing the second-best methods by 1.82%, 1.68%, and 1.54%, respectively. This research delineates the effectiveness of combining multi-size convolutional layers with LSTM and DNA shape-based features in enhancing predictive accuracy. Moreover, this study comprehensively assesses the variability in model performance across different cell lines and transcription factors. The source code of MLSNet is available at https://github.com/minghaidea/MLSNet. Yuchuan Zhang, Zhikang Wang, Fang Ge, Xiaoyu Wang 0016, Shanshan Li 0008, Yuming Guo 0001, Jiangning Song, Dongjun Yu |
Briefings Bioinform. | 4 |
| 2023 | SMG: self-supervised masked graph learning for cancer gene identificationabstractCancer genomics is dedicated to elucidating the genes and pathways that contribute to cancer progression and development. Identifying cancer genes (CGs) associated with the initiation and progression of cancer is critical for characterization of molecular-level mechanism in cancer research. In recent years, the growing availability of high-throughput molecular data and advancements in deep learning technologies has enabled the modelling of complex interactions and topological information within genomic data. Nevertheless, because of the limited labelled data, pinpointing CGs from a multitude of potential mutations remains an exceptionally challenging task. To address this, we propose a novel deep learning framework, termed self-supervised masked graph learning (SMG), which comprises SMG reconstruction (pretext task) and task-specific fine-tuning (downstream task). In the pretext task, the nodes of multi-omic featured protein-protein interaction (PPI) networks are randomly substituted with a defined mask token. The PPI networks are then reconstructed using the graph neural network (GNN)-based autoencoder, which explores the node correlations in a self-prediction manner. In the downstream tasks, the pre-trained GNN encoder embeds the input networks into feature graphs, whereas a task-specific layer proceeds with the final prediction. To assess the performance of the proposed SMG method, benchmarking experiments are performed on three node-level tasks (identification of CGs, essential genes and healthy driver genes) and one graph-level task (identification of disease subnetwork) across eight PPI networks. Benchmarking experiments and performance comparison with existing state-of-the-art methods demonstrate the superiority of SMG on multi-omic feature engineering. Yan Cui 0008, Zhikang Wang, Xiaoyu Wang 0016, Ying Zhang 0053, Tong Pan, Shanshan Li 0008, Yuming Guo 0001, Tatsuya Akutsu, Jiangning Song |
Briefings Bioinform. | 3 |
| 2023 | TIMER is a Siamese neural network-based framework for identifying both general and species-specific bacterial promotersabstractBACKGROUND: Promoters are DNA regions that initiate the transcription of specific genes near the transcription start sites. In bacteria, promoters are recognized by RNA polymerases and associated sigma factors. Effective promoter recognition is essential for synthesizing the gene-encoded products by bacteria to grow and adapt to different environmental conditions. A variety of machine learning-based predictors for bacterial promoters have been developed; however, most of them were designed specifically for a particular species. To date, only a few predictors are available for identifying general bacterial promoters with limited predictive performance. RESULTS: In this study, we developed TIMER, a Siamese neural network-based approach for identifying both general and species-specific bacterial promoters. Specifically, TIMER uses DNA sequences as the input and employs three Siamese neural networks with the attention layers to train and optimize the models for a total of 13 species-specific and general bacterial promoters. Extensive 10-fold cross-validation and independent tests demonstrated that TIMER achieves a competitive performance and outperforms several existing methods on both general and species-specific promoter prediction. As an implementation of the proposed method, the web server of TIMER is publicly accessible at http://web.unimelb-bioinfortools.cloud.edu.au/TIMER/. Yan Zhu 0006, Fuyi Li, Xiaoyu Wang 0016, Lachlan James M. Coin, Geoffrey I. Webb, Jiangning Song, Cangzhi Jia |
Briefings Bioinform. | 4 |
| 2023 | Targeting tumor heterogeneity: multiplex-detection-based multiple instance learning for whole slide image classificationabstractMOTIVATION: Multiple instance learning (MIL) is a powerful technique to classify whole slide images (WSIs) for diagnostic pathology. The key challenge of MIL on WSI classification is to discover the critical instances that trigger the bag label. However, tumor heterogeneity significantly hinders the algorithm's performance. RESULTS: Here, we propose a novel multiplex-detection-based multiple instance learning (MDMIL) which targets tumor heterogeneity by multiplex detection strategy and feature constraints among samples. Specifically, the internal query generated after the probability distribution analysis and the variational query optimized throughout the training process are utilized to detect potential instances in the form of internal and external assistance, respectively. The multiplex detection strategy significantly improves the instance-mining capacity of the deep neural network. Meanwhile, a memory-based contrastive loss is proposed to reach consistency on various phenotypes in the feature space. The novel network and loss function jointly achieve high robustness towards tumor heterogeneity. We conduct experiments on three computational pathology datasets, e.g. CAMELYON16, TCGA-NSCLC, and TCGA-RCC. Benchmarking experiments on the three datasets illustrate that our proposed MDMIL approach achieves superior performance over several existing state-of-the-art methods. AVAILABILITY AND IMPLEMENTATION: MDMIL is available for academic purposes at https://github.com/ZacharyWang-007/MDMIL. Zhikang Wang, Yue Bi, Tong Pan, Xiaoyu Wang 0016, Chris Bain, Richard Bassed, Seiya Imoto, Jianhua Yao 0001, Roger J. Daly, Jiangning Song |
Bioinform. | 4 |
| 2022 | Positive-unlabeled learning in bioinformatics and computational biology: a brief reviewabstractConventional supervised binary classification algorithms have been widely applied to address significant research questions using biological and biomedical data. This classification scheme requires two fully labeled classes of data (e.g. positive and negative samples) to train a classification model. However, in many bioinformatics applications, labeling data is laborious, and the negative samples might be potentially mislabeled due to the limited sensitivity of the experimental equipment. The positive unlabeled (PU) learning scheme was therefore proposed to enable the classifier to learn directly from limited positive samples and a large number of unlabeled samples (i.e. a mixture of positive or negative samples). To date, several PU learning algorithms have been developed to address various biological questions, such as sequence identification, functional site characterization and interaction prediction. In this paper, we revisit a collection of 29 state-of-the-art PU learning bioinformatic applications to address various biological questions. Various important aspects are extensively discussed, including PU learning methodology, biological application, classifier design and evaluation strategy. We also comment on the existing issues of PU learning and offer our perspectives for the future development of PU learning applications. We anticipate that our work serves as an instrumental guideline for a better understanding of the PU learning framework in bioinformatics and further developing next-generation PU learning frameworks for critical biological applications. Fuyi Li, Shuangyu Dong, André Leier, Meiya Han, Jing Xu 0008, Xiaoyu Wang 0016, Shirui Pan, Cangzhi Jia, Yang Zhang 0010, Geoffrey I. Webb, Lachlan James M. Coin, Chen Li 0021, Jiangning Song |
Briefings Bioinform. | 7 |
| 2022 | RBP-TSTL is a two-stage transfer learning framework for genome-scale prediction of RNA-binding proteinsabstractRNA binding proteins (RBPs) are critical for the post-transcriptional control of RNAs and play vital roles in a myriad of biological processes, such as RNA localization and gene regulation. Therefore, computational methods that are capable of accurately identifying RBPs are highly desirable and have important implications for biomedical and biotechnological applications. Here, we propose a two-stage deep transfer learning-based framework, termed RBP-TSTL, for accurate prediction of RBPs. In the first stage, the knowledge from the self-supervised pre-trained model was extracted as feature embeddings and used to represent the protein sequences, while in the second stage, a customized deep learning model was initialized based on an annotated pre-training RBPs dataset before being fine-tuned on each corresponding target species dataset. This two-stage transfer learning framework can enable the RBP-TSTL model to be effectively trained to learn and improve the prediction performance. Extensive performance benchmarking of the RBP-TSTL models trained using the features generated by the self-supervised pre-trained model and other models trained using hand-crafting encoding features demonstrated the effectiveness of the proposed two-stage knowledge transfer strategy based on the self-supervised pre-trained models. Using the best-performing RBP-TSTL models, we further conducted genome-scale RBP predictions for Homo sapiens, Arabidopsis thaliana, Escherichia coli, and Salmonella and established a computational compendium containing all the predicted putative RBPs candidates. We anticipate that the proposed RBP-TSTL approach will be explored as a useful tool for the characterization of RNA-binding proteins and exploration of their sequence-structure-function relationships. Xinxin Peng, Xiaoyu Wang 0016, Yuming Guo 0001, ZongYuan Ge, Fuyi Li, Xin Gao 0001, Jiangning Song |
Briefings Bioinform. | 2 |
| 2022 | ASPIRER: a new computational approach for identifying non-classical secreted proteins based on deep learningabstractProtein secretion has a pivotal role in many biological processes and is particularly important for intercellular communication, from the cytoplasm to the host or external environment. Gram-positive bacteria can secrete proteins through multiple secretion pathways. The non-classical secretion pathway has recently received increasing attention among these secretion pathways, but its exact mechanism remains unclear. Non-classical secreted proteins (NCSPs) are a class of secreted proteins lacking signal peptides and motifs. Several NCSP predictors have been proposed to identify NCSPs and most of them employed the whole amino acid sequence of NCSPs to construct the model. However, the sequence length of different proteins varies greatly. In addition, not all regions of the protein are equally important and some local regions are not relevant to the secretion. The functional regions of the protein, particularly in the N- and C-terminal regions, contain important determinants for secretion. In this study, we propose a new hybrid deep learning-based framework, referred to as ASPIRER, which improves the prediction of NCSPs from amino acid sequences. More specifically, it combines a whole sequence-based XGBoost model and an N-terminal sequence-based convolutional neural network model; 5-fold cross-validation and independent tests demonstrate that ASPIRER achieves superior performance than existing state-of-the-art approaches. The source code and curated datasets of ASPIRER are publicly available at https://github.com/yanwu20/ASPIRER/. ASPIRER is anticipated to be a useful tool for improved prediction of novel putative NCSPs from sequences information and prioritization of candidate proteins for follow-up experimental validation. Xiaoyu Wang 0016, Fuyi Li, Jing Xu 0008, Jia Rong, Geoffrey I. Webb, ZongYuan Ge, Jian Li 0052, Jiangning Song |
Briefings Bioinform. | 1 |
| 2022 | PreAcrs: a machine learning framework for identifying anti-CRISPR proteinsabstractBACKGROUND: Anti-CRISPR proteins are potent modulators that inhibit the CRISPR-Cas immunity system and have huge potential in gene editing and gene therapy as a genome-editing tool. Extensive studies have shown that anti-CRISPR proteins are essential for modifying endogenous genes, promoting the RNA-guided binding and cleavage of DNA or RNA substrates. In recent years, identifying and characterizing anti-CRISPR proteins has become a hot and significant research topic in bioinformatics. However, as most anti-CRISPR proteins fall short in sharing similarities to those currently known, traditional screening methods are time-consuming and inefficient. Machine learning methods could fill this gap with powerful predictive capability and provide a new perspective for anti-CRISPR protein identification. RESULTS: Here, we present a novel machine learning ensemble predictor, called PreAcrs, to identify anti-CRISPR proteins from protein sequences directly. Three features and eight different machine learning algorithms were used to train PreAcrs. PreAcrs outperformed other existing methods and significantly improved the prediction accuracy for identifying anti-CRISPR proteins. CONCLUSIONS: In summary, the PreAcrs predictor achieved a competitive performance for predicting new anti-CRISPR proteins in terms of accuracy and robustness. We anticipate PreAcrs will be a valuable tool for researchers to speed up the research process. The source code is available at: https://github.com/Lyn-666/anti_CRISPR.git . Xiaoyu Wang 0016, Fuyi Li, Jiangning Song |
BMC Bioinform. | 2 |
| 2021 | Leveraging the attention mechanism to improve the identification of DNA N6-methyladenine sitesabstractDNA N6-methyladenine is an important type of DNA modification that plays important roles in multiple biological processes. Despite the recent progress in developing DNA 6mA site prediction methods, several challenges remain to be addressed. For example, although the hand-crafted features are interpretable, they contain redundant information that may bias the model training and have a negative impact on the trained model. Furthermore, although deep learning (DL)-based models can perform feature extraction and classification automatically, they lack the interpretability of the crucial features learned by those models. As such, considerable research efforts have been focused on achieving the trade-off between the interpretability and straightforwardness of DL neural networks. In this study, we develop two new DL-based models for improving the prediction of N6-methyladenine sites, termed LA6mA and AL6mA, which use bidirectional long short-term memory to respectively capture the long-range information and self-attention mechanism to extract the key position information from DNA sequences. The performance of the two proposed methods is benchmarked and evaluated on the two model organisms Arabidopsis thaliana and Drosophila melanogaster. On the two benchmark datasets, LA6mA achieves an area under the receiver operating characteristic curve (AUROC) value of 0.962 and 0.966, whereas AL6mA achieves an AUROC value of 0.945 and 0.941, respectively. Moreover, an in-depth analysis of the attention matrix is conducted to interpret the important information, which is hidden in the sequence and relevant for 6mA site prediction. The two novel pipelines developed for DNA 6mA site prediction in this work will facilitate a better understanding of the underlying principle of DL-based DNA methylation site prediction and its future applications. Ying Zhang 0053, Yan Liu 0038, Jian Xu 0009, Xiaoyu Wang 0016, Xinxin Peng, Jiangning Song, Dongjun Yu |
Briefings Bioinform. | 4 |