Yansu Wang

dblp:302/1144 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MuFGPS: enhancing liquid-liquid phase separation protein prediction through multi-level features and ensemble learning
abstract
Liquid-liquid phase separation (LLPS) is a key mechanism driving the assembly of membrane-less organelles and is increasingly recognized for its involvement in essential cellular functions and various diseases. However, existing computational approaches largely rely on sequence-level descriptors and often fail to explicitly incorporate structural topology information, limiting their ability to capture the complex determinants of LLPS behavior. Accurate identification of LLPS-capable proteins remains challenging due to their sequence diversity and complex structural determinants. Here, we present MuFGPS (Multi-level Feature Graph-based Predictor for Phase-Separating proteins), a predictive framework integrating sequence-derived physicochemical features, Define Secondary Structure of Proteins-annotated secondary structures, and graph-based structural embeddings from AlphaFold residue contact maps via a multi-head Graph Attention Network. Class imbalance is addressed using Synthetic Minority Oversampling Technique (SMOTE), and classification is performed through a stacking ensemble of Random Forest, XGBoost, and LightGBM. Benchmarks against six representative methods demonstrate that MuFGPS achieves superior performance across all metrics, with notable gains in F1-score and matthews correlation coefficient (MCC). Ablation analyses confirm the synergistic contributions of structural features and ensemble learning to accuracy and robustness. MuFGPS offers a scalable and high-accuracy framework for proteome-wide LLPS protein prediction.
Lei Xian, Quan Zou 0001, Ren Qi, Mengting Niu, Yansu Wang
Briefings Bioinform.5
2026 HMA-GCA: hybrid manifold augmentation and gated cross-attention for circRNA-miRNA interaction prediction
abstract
MOTIVATION: Circular RNAs (circRNAs) interact with microRNAs (miRNAs) to regulate gene expression and influence disease progression. However, traditional models tend to overlook the significant contributions of certain features when dealing with diverse sequence information, resulting in the inability to capture some deep topological structures and thus leaving room for improvement in prediction performance. RESULTS: We propose HMA-GCA, a novel framework that integrates hybrid manifold augmentation and gated cross-attention for CMI prediction. The model first constructs multi-scale descriptors by combining sequence-derived features (K-mer, CTD, Doc2Vec) and topological features (Role2Vec, node degree, neighborhood proximity). It then applies PCA for global linear projection and UMAP for local nonlinear manifold learning, enhancing feature representations while preserving intrinsic data geometry. A channel-wise gated cross-attention mechanism dynamically controls the injection of miRNA information into circRNA representations. Extensive experiments on three benchmark datasets show that HMA-GCA consistently outperforms state-of-the-art methods across multiple metrics. To ensure interpretability, we conducted SHAP analysis to quantify the contribution of each feature type, revealing that sequence-derived features and topological similarities are the most influential. Ablation studies confirm the necessity of each module, while case studies demonstrate that top-ranked predictions are supported by literature evidence. Overall, HMA-GCA not only achieves state-of-the-art predictive performance but also provides interpretable insights into the molecular features. AVAILABILITY AND IMPLEMENTATION: The source code and data are freely available at https://github.com/Lixunwind/Prediction-circ-mi-by-Gate.git. The implementation is based on Python and the required dependencies are listed in the repository.
Yunzhou Hu, Yansu Wang, Yifeng Bai, Lei Xu 0047, Quan Zou 0001, Chunyu Wang 0002, Mengting Niu
Bioinform.2
2026 EAGP: an efficient generative augmentation framework for phage protein classification under severe class imbalance
abstract
MOTIVATION: The accurate classification of phage proteins is critical for advancing bacteriophage research. Despite the proliferation of machine learning approaches in this domain, the persistent issue of data imbalance continues to hinder performance, particularly for rare protein sequences. Previous attempts to address this by re-weighting minority classes have faced limitations due to insufficient feature extraction capabilities. RESULTS: In this paper, we introduce EAGP, a novel approach that integrates a generative model-functionally equivalent to a WGAN yet tailored for one-dimensional data-with the Evolutionary Scale Modeling (ESM) protein large language model for robust feature extraction. EAGP exhibits exceptional performance in binary classification and protein function annotation tasks. Crucially, our method not only improves overall classification efficacy but also significantly alleviates the performance degradation typically observed in minority classes. AVAILABILITY AND IMPLEMENTATION: The data and code underlying this article are available in GitHub at https://github.com/Innerly/EAGP and have been archived on Zenodo at https://doi.org/10.5281/zenodo.19928069.
Jiaru Li, Yansu Wang, Quan Zou 0001, Hongling Zhu
Bioinform.3
2025 UPDMIA: Unified Processing of Diverse Medical Imaging Data via a Multi-Channel EfficientNet Architecture
abstract
Pulmonary medical image analysis has become increasingly challenged by the overwhelming volume and heterogeneity of CT, MRI, PET, X-ray and histopathology data, leading to time-consuming workflows and error-prone manual interpretation. To address this, we propose UPDMIA, a unified deep-learning pipeline built on a structurally enhanced EfficientNet-B7 backbone, coupled with an intelligent modality-adaptive classification head and a comprehensive multimodal data augmentation suite. All input images are preprocessed to$224 \times 224$resolution, normalized, and subjected to probabilistic augmentation (flips, rotations, cropping, Gaussian blur, color jitter, random erasing). Modality detection then dynamically routes features through lightweight, task-specific classifiers. Experimental evaluation across five modalities demonstrates strong performance: CT (F1 =$96.99 \%)$, MRI$(\mathrm{F} 1=87.40 \%)$, PET$(\mathrm{F} 1=93.35 \%)$, X-ray ($\mathrm{F} 1= 95.49 \%$) and histological slides ($F 1=97.60 \%$), with overall precision and recall above 88 % for all domains. This unified approach matches or exceeds specialized single-modality models while simplifying deployment.
Binyun Yang, Quan Zou 0001, Yansu Wang
BIBM3
2025 metaTP: a meta-transcriptome data analysis pipeline with integrated automated workflows
abstract
BACKGROUND: The accessibility of sequencing technologies has enabled meta-transcriptomic studies to provide a deeper understanding of microbial ecology at the transcriptional level. Analyzing omics data involves multiple steps that require the use of various bioinformatics tools. With the increasing availability of public microbiome datasets, conducting meta-analyses can reveal new insights into microbiome activity. However, the reproducibility of data is often compromised due to variations in processing methods for sample omics data. Therefore, it is essential to develop efficient analytical workflows that ensure repeatability, reproducibility, and the traceability of results in microbiome research. RESULTS: We developed metaTP, a pipeline that integrates bioinformatics tools for analyzing meta-transcriptomic data comprehensively. The pipeline includes quality control, non-coding RNA removal, transcript expression quantification, differential gene expression analysis, functional annotation, and co-expression network analysis. To quantify mRNA expression, we rely on reference indexes built using protein-coding sequences, which help overcome the limitations of database analysis. Additionally, metaTP provides a function for calculating the topological properties of gene co-expression networks, offering an intuitive explanation for correlated gene sets in high-dimensional datasets. The use of metaTP is anticipated to support researchers in addressing microbiota-related biological inquiries and improving the accessibility and interpretation of microbiota RNA-Seq data. CONCLUSIONS: We have created a conda package to integrate the tools into our pipeline, making it a flexible and versatile tool for handling meta-transcriptomic sequencing data. The metaTP pipeline is freely available at: https://github.com/nanbei45/metaTP .
Limuxuan He, Quan Zou 0001, Yansu Wang
BMC Bioinform.3
2024 Adversarial regularized autoencoder graph neural network for microbe-disease associations prediction
abstract
BACKGROUND: Microorganisms inhabit various regions of the human body and significantly contribute to numerous diseases. Predicting the associations between microbes and diseases is crucial for understanding pathogenic mechanisms and informing prevention and treatment strategies. Biological experiments to determine these associations are time-consuming and costly. Therefore, integrating deep learning with biological networks can efficiently identify potential microbe-disease associations on a large scale. METHODS: We propose an adversarial regularized autoencoder graph neural network algorithm, named Stacked Adversarial Regularization for Microbe-Disease Associations Prediction (SARMDA), for predicting associations between microbes and diseases. First, we integrate topological structural similarity and functional similarity metrics of microbes and diseases to construct a heterogeneous network. Then, utilizing an autoencoder based on GraphSAGE, we learn both the topological and attribute representations of nodes within the constructed network. Finally, we introduce an adversarial regularized autoencoder graph neural network embedding model to address the inherent limitations of traditional GraphSAGE autoencoders in capturing global information. RESULTS: Under the five-fold cross-validation on microbe-disease pairs, SARMDA was compared with eight advanced methods using the Human Microbe-Disease Association Database (HMDAD) and Disbiome databases. The best area under the ROC curve (AUC) achieved by SARMDA on HMDAD was 0.9891$\pm$0.0057, and the best area under the precision-recall curve (AUPR) was 0.9902$\pm$0.0128. On the Disbiome dataset, the AUC was 0.9328$\pm$0.0072, and the best AUPR was 0.9233$\pm$0.0089, outperforming the other eight MDAs prediction methods. Furthermore, the effectiveness of our model was demonstrated through a detailed analysis of asthma and inflammatory bowel disease cases.
Limuxuan He, Quan Zou 0001, Shuang Cheng, Yansu Wang
Briefings Bioinform.5
2023 One step multi-view spectral clustering via joint adaptive graph learning and matrix factorization
Wenqi Yang, Yansu Wang, Chang Tang, Hengjian Tong, Ao Wei
Neurocomputing2
2023 Recall DNA methylation levels at low coverage sites using a CNN model in WGBS
abstract
DNA methylation is an important regulator of gene transcription. WGBS is the gold-standard approach for base-pair resolution quantitative of DNA methylation. It requires high sequencing depth. Many CpG sites with insufficient coverage in the WGBS data, resulting in inaccurate DNA methylation levels of individual sites. Many state-of-arts computation methods were proposed to predict the missing value. However, many methods required either other omics datasets or other cross-sample data. And most of them only predicted the state of DNA methylation. In this study, we proposed the RcWGBS, which can impute the missing (or low coverage) values from the DNA methylation levels on the adjacent sides. Deep learning techniques were employed for the accurate prediction. The WGBS datasets of H1-hESC and GM12878 were down-sampled. The average difference between the DNA methylation level at 12× depth predicted by RcWGBS and that at >50× depth in the H1-hESC and GM2878 cells are less than 0.03 and 0.01, respectively. RcWGBS performed better than METHimpute even though the sequencing depth was as low as 12×. Our work would help to process methylation data of low sequencing depth. It is beneficial for researchers to save sequencing costs and improve data utilization through computational methods.
Ximei Luo, Yansu Wang, Quan Zou 0001, Lei Xu 0047
PLoS Comput. Biol.2
2022 Effector-GAN: prediction of fungal effector proteins based on pretrained deep representation learning methods and generative adversarial networks
abstract
MOTIVATION: Phytopathogenic fungi secrete effector proteins to subvert host defenses and facilitate infection. Systematic analysis and prediction of candidate fungal effector proteins are crucial for experimental validation and biological control of plant disease. However, two problems are still considered intractable to be solved in fungal effector prediction: one is the high-level diversity in effector sequences that increases the difficulty of protein feature learning, and the other is the class imbalance between effector and non-effector samples in the training dataset. RESULTS: In our study, pretrained deep representation learning methods are presented to represent multiple characteristics of sequences for predicting fungal effectors and generative adversarial networks are adapted to create synthetic feature samples to address the data imbalance problem. Compared with the state-of-the-art fungal effector prediction methods, Effector-GAN shows an overall improvement in accuracy in the independent test set. AVAILABILITY AND IMPLEMENTATION: Effector-GAN offers a user-friendly interface to inspect potential fungal effector proteins (http://lab.malab.cn/~wys/webserver/Effector-GAN). The Python script can be downloaded from http://lab.malab.cn/~wys/gitlab/effector-gan. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yansu Wang, Ximei Luo, Quan Zou 0001
Bioinform.1
2021 Machine learning for phytopathology: from the molecular scale towards the network scale
abstract
With the increasing volume of high-throughput sequencing data from a variety of omics techniques in the field of plant-pathogen interactions, sorting, retrieving, processing and visualizing biological information have become a great challenge. Within the explosion of data, machine learning offers powerful tools to process these complex omics data by various algorithms, such as Bayesian reasoning, support vector machine and random forest. Here, we introduce the basic frameworks of machine learning in dissecting plant-pathogen interactions and discuss the applications and advances of machine learning in plant-pathogen interactions from molecular to network biology, including the prediction of pathogen effectors, plant disease resistance protein monitoring and the discovery of protein-protein networks. The aim of this review is to provide a summary of advances in plant defense and pathogen infection and to indicate the important developments of machine learning in phytopathology.
Yansu Wang, Murong Zhou, Quan Zou 0001, Lei Xu 0047
Briefings Bioinform.1