VLDB 2026 Research / reviewers in the wild / expert
Yuming Guo 0001
dblp:91/9543-1
· DBLP profile ↗
12ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0002-1766-6592ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 9 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multimodal geometric learning for antimicrobial peptide identification by leveraging alphafold2-predicted structures and surface featuresabstractAntimicrobial peptides (AMPs) are short peptides that play critical roles in diverse biological processes and exhibit functional activities against target organisms. While numerous methods have demonstrated the effectiveness of deep neural networks for AMP identification using sequence features; nevertheless, higher-level peptide characteristics-such as 3D structure and geometric surface features-have not been comprehensively explored. To address this gap, we introduce the SSFGM-Model (Sequence, Structure, Surface, Graph, and Geometric-based Model), a novel framework that integrates multiple feature types to enhance AMP identification. The model represents each peptide sequence as a graph, where nodes are characterized by amino acid features derived from ProteinBERT, ESM-2, and One-hot embeddings. Graph convolutional networks and an attention mechanism are employed to capture high-order structural and sequential relationships. Additionally, surface geometry and physicochemical properties are processed using a geometric neural network. Finally, a feature fusion strategy combines the outputs from these subnetworks to enable robust AMP identification. Extensive benchmarking experiments demonstrate that the SSFGM-Model outperforms current state-of-the-art methods. An ablation study further confirms the critical role of sequence, structural, and surface features in AMP identification. The key contribution of this work is the innovative integration of multiple levels of peptide characteristics and the combination of geometric and graph neural networks. This approach provides a more comprehensive understanding of the sequence-structure-function relationship of peptides, paving the way for more accurate AMP prediction. The SSFGM-Model has a significant potential for applications in the discovery and design of novel AMP-based therapeutics. The source code is publicly available at https://github.com/ggcameronnogg/SSFGM-Model. Zehua Sun, Jing Xu 0008, Zhikang Wang, Xiaoyu Wang 0016, Shanshan Li 0008, Yuming Guo 0001, Hsin Hui Shen, Jiangning Song |
Briefings Bioinform. | 8 |
| 2025 | Supervised contrastive learning enhances MHC-II peptide binding affinity prediction
Long-Chen Shen, Yan Liu 0038, Zi Liu, Zhikang Wang, Yuming Guo 0001, Jamie Rossjohn, Jiangning Song, Dongjun Yu |
Expert Syst. Appl. | 6 |
| 2025 | Integrating Graph Convolutional Networks for Missing Gene Expression ImputationabstractSingle-cell RNA sequencing (scRNA-seq) techniques are emerging to revolutionize modern biomedical sciences by providing a detailed landscape of individual cells. However, these methods often lack crucial spatial localization information. To address this gap, spatial transcriptomic technologies have developed, enabling gene expression profiling while mapping cells spatial information. Yet, the gene throughput in spatial transcriptomic technologies makes it challenging to characterize whole-transcriptome-level data for single cells in space. In this context, approaches for predicting the spatial distribution of genes are still under development. Here, we present GCNgene, a novel method to predict the spatial distribution of the undetected RNA transcripts, through integrating spatial and scRNA-seq datasets. GCNgene leverages a graph convolutional network to embed spatial transcriptomics data and then applies a learned rule to reconstruct gene expression by combining the reference single-cell data with the calculated cell-type proportions. Ultimately, this learned paradigm enables accurate predictions of gene expression levels. Ying Zhang 0053, Hong-Jin Yu, Zihao Yan, Tong Pan, Yan Liu 0038, Shanshan Li 0008, Yuming Guo 0001, Jiangning Song, Dongjun Yu |
IEEE Trans. Comput. Biol. Bioinform. | 8 |
| 2024 | Deep learning approaches for non-coding genetic variant effect prediction: current progress and future prospectsabstractRecent advancements in high-throughput sequencing technologies have significantly enhanced our ability to unravel the intricacies of gene regulatory processes. A critical challenge in this endeavor is the identification of variant effects, a key factor in comprehending the mechanisms underlying gene regulation. Non-coding variants, constituting over 90% of all variants, have garnered increasing attention in recent years. The exploration of gene variant impacts and regulatory mechanisms has spurred the development of various deep learning approaches, providing new insights into the global regulatory landscape through the analysis of extensive genetic data. Here, we provide a comprehensive overview of the development of the non-coding variants models based on bulk and single-cell sequencing data and their model-based interpretation and downstream tasks. This review delineates the popular sequencing technologies for epigenetic profiling and deep learning approaches for discerning the effects of non-coding variants. Additionally, we summarize the limitations of current approaches in variant effect prediction research and outline opportunities for improvement. We anticipate that our study will offer a practical and useful guide for the bioinformatic community to further advance the unraveling of genetic variant effects. Xiaoyu Wang 0016, Fuyi Li, Seiya Imoto, Hsin-Hui Shen, Shanshan Li 0008, Yuming Guo 0001, Jian Yang 0005, Jiangning Song |
Briefings Bioinform. | 7 |
| 2024 | MLSNet: a deep learning model for predicting transcription factor binding sitesabstractAccurate prediction of transcription factor binding sites (TFBSs) is essential for understanding gene regulation mechanisms and the etiology of diseases. Despite numerous advances in deep learning for predicting TFBSs, their performance can still be enhanced. In this study, we propose MLSNet, a novel deep learning architecture designed specifically to predict TFBSs. MLSNet innovatively integrates multisize convolutional fusion with long short-term memory (LSTM) networks to effectively capture DNA-sparse higher-order sequence features. Further, MLSNet incorporates super token attention and Bi-LSTM to systematically extract and integrate higher-order DNA shape features. Experimental results on 165 ChIP-seq (chromatin immunoprecipitation followed by sequencing) datasets indicate that MLSNet consistently outperforms several state-of-the-art algorithms in the prediction of TFBSs. Specifically, MLSNet reports average metrics: 0.8306 for ACC, 0.8992 for AUROC, and 0.9035 for AUPRC, surpassing the second-best methods by 1.82%, 1.68%, and 1.54%, respectively. This research delineates the effectiveness of combining multi-size convolutional layers with LSTM and DNA shape-based features in enhancing predictive accuracy. Moreover, this study comprehensively assesses the variability in model performance across different cell lines and transcription factors. The source code of MLSNet is available at https://github.com/minghaidea/MLSNet. Yuchuan Zhang, Zhikang Wang, Fang Ge, Xiaoyu Wang 0016, Shanshan Li 0008, Yuming Guo 0001, Jiangning Song, Dongjun Yu |
Briefings Bioinform. | 7 |
| 2023 | SMG: self-supervised masked graph learning for cancer gene identificationabstractCancer genomics is dedicated to elucidating the genes and pathways that contribute to cancer progression and development. Identifying cancer genes (CGs) associated with the initiation and progression of cancer is critical for characterization of molecular-level mechanism in cancer research. In recent years, the growing availability of high-throughput molecular data and advancements in deep learning technologies has enabled the modelling of complex interactions and topological information within genomic data. Nevertheless, because of the limited labelled data, pinpointing CGs from a multitude of potential mutations remains an exceptionally challenging task. To address this, we propose a novel deep learning framework, termed self-supervised masked graph learning (SMG), which comprises SMG reconstruction (pretext task) and task-specific fine-tuning (downstream task). In the pretext task, the nodes of multi-omic featured protein-protein interaction (PPI) networks are randomly substituted with a defined mask token. The PPI networks are then reconstructed using the graph neural network (GNN)-based autoencoder, which explores the node correlations in a self-prediction manner. In the downstream tasks, the pre-trained GNN encoder embeds the input networks into feature graphs, whereas a task-specific layer proceeds with the final prediction. To assess the performance of the proposed SMG method, benchmarking experiments are performed on three node-level tasks (identification of CGs, essential genes and healthy driver genes) and one graph-level task (identification of disease subnetwork) across eight PPI networks. Benchmarking experiments and performance comparison with existing state-of-the-art methods demonstrate the superiority of SMG on multi-omic feature engineering. Yan Cui 0008, Zhikang Wang, Xiaoyu Wang 0016, Ying Zhang 0053, Tong Pan, Shanshan Li 0008, Yuming Guo 0001, Tatsuya Akutsu, Jiangning Song |
Briefings Bioinform. | 9 |
| 2023 | MULGA, a unified multi-view graph autoencoder-based approach for identifying drug-protein interaction and drug repositioningabstractMOTIVATION: Identifying drug-protein interactions (DPIs) is a critical step in drug repositioning, which allows reuse of approved drugs that may be effective for treating a different disease and thereby alleviates the challenges of new drug development. Despite the fact that a great variety of computational approaches for DPI prediction have been proposed, key challenges, such as extendable and unbiased similarity calculation, heterogeneous information utilization, and reliable negative sample selection, remain to be addressed. RESULTS: To address these issues, we propose a novel, unified multi-view graph autoencoder framework, termed MULGA, for both DPI and drug repositioning predictions. MULGA is featured by: (i) a multi-view learning technique to effectively learn authentic drug affinity and target affinity matrices; (ii) a graph autoencoder to infer missing DPI interactions; and (iii) a new "guilty-by-association"-based negative sampling approach for selecting highly reliable non-DPIs. Benchmark experiments demonstrate that MULGA outperforms state-of-the-art methods in DPI prediction and the ablation studies verify the effectiveness of each proposed component. Importantly, we highlight the top drugs shortlisted by MULGA that target the spike glycoprotein of severe acute respiratory syndrome coronavirus 2 (SAR-CoV-2), offering additional insights into and potentially useful treatment option for COVID-19. Together with the availability of datasets and source codes, we envision that MULGA can be explored as a useful tool for DPI prediction and drug repositioning. AVAILABILITY AND IMPLEMENTATION: MULGA is publicly available for academic purposes at https://github.com/jianiM/MULGA/. Jiani Ma, Chen Li 0021, Zhikang Wang, Shanshan Li 0008, Yuming Guo 0001, Lin Zhang 0015, Hui Liu 0024, Xin Gao 0001, Jiangning Song |
Bioinform. | 6 |
| 2023 | Interpretable prediction models for widespread m6A RNA modification across cell lines and tissuesabstractMOTIVATION: RNA N6-methyladenosine (m6A) in Homo sapiens plays vital roles in a variety of biological functions. Precise identification of m6A modifications is thus essential to elucidation of their biological functions and underlying molecular-level mechanisms. Currently available high-throughput single-nucleotide-resolution m6A modification data considerably accelerated the identification of RNA modification sites through the development of data-driven computational methods. Nevertheless, existing methods have limitations in terms of the coverage of single-nucleotide-resolution cell lines and have poor capability in model interpretations, thereby having limited applicability. RESULTS: In this study, we present CLSM6A, comprising a set of deep learning-based models designed for predicting single-nucleotide-resolution m6A RNA modification sites across eight different cell lines and three tissues. Extensive benchmarking experiments are conducted on well-curated datasets and accordingly, CLSM6A achieves superior performance than current state-of-the-art methods. Furthermore, CLSM6A is capable of interpreting the prediction decision-making process by excavating critical motifs activated by filters and pinpointing highly concerned positions in both forward and backward propagations. CLSM6A exhibits better portability on similar cross-cell line/tissue datasets, reveals a strong association between highly activated motifs and high-impact motifs, and demonstrates complementary attributes of different interpretation strategies. AVAILABILITY AND IMPLEMENTATION: The webserver is available at http://csbio.njust.edu.cn/bioinf/clsm6a. The datasets and code are available at https://github.com/zhangying-njust/CLSM6A/. Ying Zhang 0053, Zhikang Wang, Shanshan Li 0008, Yuming Guo 0001, Jiangning Song, Dongjun Yu |
Bioinform. | 5 |
| 2023 | Selective Feature Bagging of one-class classifiers for novelty detection in high-dimensional data
Guanglei Meng, Tiankuo Meng, Yingnan Wang, Yuming Guo 0001, Zhihua Qiao, Zhizhong Mao |
Eng. Appl. Artif. Intell. | 7 |
| 2022 | Clarion is a multi-label problem transformation method for identifying mRNA subcellular localizationsabstractSubcellular localization of messenger RNAs (mRNAs) plays a key role in the spatial regulation of gene activity. The functions of mRNAs have been shown to be closely linked with their localizations. As such, understanding of the subcellular localizations of mRNAs can help elucidate gene regulatory networks. Despite several computational methods that have been developed to predict mRNA localizations within cells, there is still much room for improvement in predictive performance, especially for the multiple-location prediction. In this study, we proposed a novel multi-label multi-class predictor, termed Clarion, for mRNA subcellular localization prediction. Clarion was developed based on a manually curated benchmark dataset and leveraged the weighted series method for multi-label transformation. Extensive benchmarking tests demonstrated Clarion achieved competitive predictive performance and the weighted series method plays a crucial role in securing superior performance of Clarion. In addition, the independent test results indicate that Clarion outperformed the state-of-the-art methods and can secure accuracy of 81.47, 91.29, 79.77, 92.10, 89.15, 83.74, 80.74, 79.23 and 84.74% for chromatin, cytoplasm, cytosol, exosome, membrane, nucleolus, nucleoplasm, nucleus and ribosome, respectively. The webserver and local stand-alone tool of Clarion is freely available at http://monash.bioweb.cloud.edu.au/Clarion/. Yue Bi, Fuyi Li, Zhikang Wang, Tong Pan, Yuming Guo 0001, Geoffrey I. Webb, Jianhua Yao 0001, Cangzhi Jia, Jiangning Song |
Briefings Bioinform. | 6 |
| 2022 | RBP-TSTL is a two-stage transfer learning framework for genome-scale prediction of RNA-binding proteinsabstractRNA binding proteins (RBPs) are critical for the post-transcriptional control of RNAs and play vital roles in a myriad of biological processes, such as RNA localization and gene regulation. Therefore, computational methods that are capable of accurately identifying RBPs are highly desirable and have important implications for biomedical and biotechnological applications. Here, we propose a two-stage deep transfer learning-based framework, termed RBP-TSTL, for accurate prediction of RBPs. In the first stage, the knowledge from the self-supervised pre-trained model was extracted as feature embeddings and used to represent the protein sequences, while in the second stage, a customized deep learning model was initialized based on an annotated pre-training RBPs dataset before being fine-tuned on each corresponding target species dataset. This two-stage transfer learning framework can enable the RBP-TSTL model to be effectively trained to learn and improve the prediction performance. Extensive performance benchmarking of the RBP-TSTL models trained using the features generated by the self-supervised pre-trained model and other models trained using hand-crafting encoding features demonstrated the effectiveness of the proposed two-stage knowledge transfer strategy based on the self-supervised pre-trained models. Using the best-performing RBP-TSTL models, we further conducted genome-scale RBP predictions for Homo sapiens, Arabidopsis thaliana, Escherichia coli, and Salmonella and established a computational compendium containing all the predicted putative RBPs candidates. We anticipate that the proposed RBP-TSTL approach will be explored as a useful tool for the characterization of RNA-binding proteins and exploration of their sequence-structure-function relationships. Xinxin Peng, Xiaoyu Wang 0016, Yuming Guo 0001, ZongYuan Ge, Fuyi Li, Xin Gao 0001, Jiangning Song |
Briefings Bioinform. | 3 |
| 2022 | Boosting the prediction of molten steel temperature in ladle furnace with a dynamic outlier ensemble
Guanglei Meng, Zhihua Qiao, Yuming Guo 0001, Zhizhong Mao |
Eng. Appl. Artif. Intell. | 5 |