Yan Wang 0028

dblp:59/2227-28 · DBLP profile ↗
← Back
63ranked-venue papers
7as first author
40since 2021 · last 2026
0000-0002-4751-0708ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 37 · 29 since 2021Artificial intelligence and machine learning · 18 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 HpMiX: A Disease ceRNA biomarker prediction framework driven by graph topology-constrained Mixup and hypergraph residual enhancement
Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Fengfeng Zhou
Neural Networks3
2026 A pre-trained language model-based cross-modal fusion framework for predicting miRNA-drug resistance and sensitivity associations
abstract
MicroRNAs (miRNAs) are pivotal regulators of drug resistance and sensitivity in cancer cells, functioning as tumor suppressors or oncogenes that modulate the cellular response to anticancer drugs. While experimental identification of miRNA-mediated drug resistance and sensitivity is both costly and laborious, computational methods present a promising alternative. Recent advances in pre-trained language models (PLMs) offer new opportunities to leverage large-scale unlabeled biomolecular data for enhanced relationship prediction. In this study, we introduce PLMF-MDA, a PLM-based cross-modal fusion model designed to predict miRNA-drug resistance (MDR) and miRNA-drug sensitivity (MDS) associations. PLMF-MDA integrates miRNA and drug multimodal embeddings derived from PLMs and intrinsic feature extractors, and employs a cross-modal attention fusion module to adaptively capture key interactions between modalities. To evaluate the performance of the approach, we manually constructed two benchmark datasets. Experimental results demonstrate that the PLMF-MDA achieves superior prediction performance. Furthermore, case studies on anticancer drug docetaxel and gefitinib demonstrate its potential in discovering novel MDR (MDS) associations. All data and source code are available on GitHub: https://github.com/sheng-n/PLMF-MDA.
Nan Sheng, Yun-Zhi Liu, Wenju Hou, Lan Huang 0002, Yan Wang 0028
PLoS Comput. Biol.6
2026 Structure and Semantics Aware Multi-View Contrastive Learning for Predicting Association Among lncRNAs, miRNAs and Diseases
abstract
Exploring associations among long non-coding RNAs (lncRNAs), microRNAs (miRNAs), and diseases is crucial for biomarker discovery and precision medicine. Existing computational methods are hindered by sparse known associations and the complexity of biological networks. To address this challenge, we propose SSMVCL (Structure- and Semantic-aware Multi-View Contrastive Learning), a unified framework for predicting lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs), and lncRNA-miRNA interactions (LMIs). SSMVCL constructs a heterogeneous bioinformatics network from multi-source biological data and learns representations from two complementary views: a structure-aware view for local topology and a semantic-aware view using biologically meaningful meta-paths to capture high-order relationships. A cross-view contrastive alignment module with adaptive negative sampling enforces consistency between views and enhances discriminative capability. On two benchmark datasets, SSMVCL achieves state-of-the-art performance: for Dataset2, AUC/AUPR of 0.9736/0.9716 (LDA), 0.9364/0.9309 (MDA), and 0.9297/0.9234 (LMI) Case studies on gastric and prostate cancers further validated robustness and translational potential by identifying supported associations.
Lan Huang 0002, Yujuan Zhang, Yan Wang 0028, Nan Sheng
IEEE J. Biomed. Health Informatics5
2026 A Dynamic Multi-Scale Hypergraph Learning Framework Driven by Features and Structures for ceRNA-Disease Association Prediction
abstract
Competitive endogenous RNA (ceRNA) networks are pivotal for uncovering disease molecular mechanisms. Graph representation learning is a cornerstone for modeling biological regulatory networks and predicting disease-related biomarkers. However, current methods face challenges: traditional graph neural network (GNN) rely on low-order graph structures, which struggle to capture high-order molecular interactions, resulting in topological information loss; shallow GNN fail to model long-range dependencies, while deep architectures suffer from over-smoothing, limiting complex regulatory expression; static embeddings overlook dynamic molecular interactions, reducing biomarker accuracy. These limitations highlight the need for advanced graph learning frameworks. To address these challenges, we propose DMHLF, a Dynamic Multi-scale Hypergraph Learning Framework for predicting disease-associated ceRNA biomarkers. The framework first integrates multiple regulatory relationships among miRNAs, lncRNAs, circRNAs, mRNAs, and diseases to construct disease-specific ceRNA regulatory networks, capturing local and global regulatory patterns through multi-Hop hyperedges. Subsequently, we devise a Hypergraph-Weighted Dynamic Random Walk (HEDRW) method to dynamically extract node meta-embeddings that encode high-order regulatory information. Concurrently, we extend Eigen-GNN spectral analysis to hypergraph structures, incorporating a residual-enhanced hypergraph neural network to preserve the global topological properties of shallow hypergraphs. Finally, a cross-scale attention mechanism aligns and fuses multi-scale features to generate high-quality node embeddings for disease-ceRNA association prediction. Experiments on diverse datasets demonstrate that DMHLF significantly outperforms existing methods. Case study further validates the framework's efficacy in identifying disease-related ceRNA biomarkers, providing a reliable predictive tool for biomedical research.
Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Fengfeng Zhou, Yu-Qing Li
IEEE J. Biomed. Health Informatics3
2026 Reliable Multimodal Cancer Survival Prediction With Confidence-Aware Risk Modeling
abstract
Multimodal survival methods that integrate histology whole-slide images and transcriptomic profiles hold significant promise for understanding patient prognostication and guiding personalized treatment strategies. However, existing approaches primarily focus on improving predictive performance through multimodal information fusion, often neglecting the reliability estimation of the prediction results and the inherent alignment noise across modalities. Thus, we propose ReCaSP, a novel and reliable cancer survival prediction framework that effectively integrates histology and transcriptomics data via multimodal alignment and fusion, providing the auxiliary confidence levels for survival predictions through a confidence-aware risk modeling mechanism. Specifically, our approach incorporates a fine-grained risk classifier that models risk labels jointly over both multiple time intervals and censorship status, utilizing evidential deep learning to yield fine-grained risk predictions accompanied by confidence scores. Additionally, to mitigate the inherent noise in multimodal data alignment, we introduce a cross-attention alignment module that effectively aligns histology data with transcriptomics data prior to multimodal fusion, thereby facilitating cross-modal interaction learning. Extensive experiments on five datasets demonstrate that ReCaSP significantly outperforms state-of-the-art methods, achieving a 4.58% improvement in the overall C-Index.
Xuping Xie, Qixing Yang, Lan Huang 0002, Fengfeng Zhou, Yan Wang 0028
IEEE J. Biomed. Health Informatics7
2025 Multimodal Hypothetical Summary for Retrieval-based Multi-image Question Answering
abstract
Retrieval-based multi-image question answering (QA) task involves retrieving multiple question-related images and synthesizing these images to generate an answer. Conventional "retrieve-then-answer" pipelines often suffer from cascading errors because the training objective of QA fails to optimize the retrieval stage. To address this issue, we propose a novel method to effectively introduce and reference retrieved information into the QA. Given the image set to be retrieved, we employ a multimodal large language model (visual perspective) and a large language model (textual perspective) to obtain multimodal hypothetical summary in question-form and description-form. By combining visual and textual perspectives, MHyS captures image content more specifically and replaces real images in retrieval, which eliminates the modality gap by transforming into text-to-text retrieval and helps improve retrieval. To more advantageously introduce retrieval with QA, we employ contrastive learning to align queries (questions) with MHyS. Moreover, we propose a coarse-to-fine strategy for calculating both sentence-level and word-level similarity scores, to further enhance retrieval and filter out irrelevant details. Our approach achieves a 3.7% absolute improvement over state-of-the-art methods on RETVQA and a 14.5% improvement over CLIP. Comprehensive experiments and detailed ablation studies demonstrate the superiority of our method.
Peize Li, Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Yan Wang 0028
AAAI5
2025 GSAM-MRI: Frequency-Based Domain Randomization for Generalized MR Image Segmentation with Segment Anything Model
abstract
Magnetic resonance imaging (MRI) data segmentation plays a critical role in clinical diagnosis and treatment planning. However, the performance of deep learning-based segmentation models is often hindered by domain shifts caused by variations in imaging factors. To address this challenge, we propose GSAM-MRI, a Generalized Segment Anything Model for robust MRI segmentation in the scenario of single-source domain generalization (SDG). GSAM-MRI integrates multiple components to enhance domain generalization: (1) Frequency-based Domain Randomization module that simulates inter-site variability by perturbing the frequency domain; (2) Domain Adversarial Block that promotes domain-invariant feature learning through adversarial training; (3) General Embedding Generator that fuses multi-scale hierarchical features to produce dense prompt embeddings; Additionally, a hybrid loss function is employed for output consistency. Experiments on prostate segmentation and white matter hyperintensity segmentation tasks demonstrate that GSAM-MRI consistently outperforms state-of-the-art SDG methods and baseline models, achieving superior generalization across unseen domains.
Lan Huang 0002, Yinglu Sun, Xinfei Wang 0001, Qixing Yang, Xuping Xie, Wenju Hou, Chunjie Guo, Yan Wang 0028
BIBM9
2025 PepLand: a large-scale pre-trained peptide representation model for a comprehensive landscape of both canonical and non-canonical amino acids
abstract
The recent interest in peptides incorporating non-canonical amino acids has surged within the scientific community, driven by their enhanced stability and resistance to proteolytic degradation. These so-called non-canonical peptides offer significant potential for modifying biological, pharmacological, and physiochemical characteristics in both native and synthetic contexts. Despite their advantages, there remains a notable gap in the availability of an efficient pre-trained model capable of effectively capturing feature representations from such intricate peptide sequences. This study herein introduces PepLand, a novel pre-training framework designed for the comprehensive representation and analysis of peptides, encompassing both canonical and non-canonical amino acids. PepLand leverages a general-purpose multi-view heterogeneous graph neural network to unveil the subtle structural representations of peptides. Our empirical evaluations demonstrate PepLand's proficiency in a range of peptide property prediction tasks, including cell penetrability, solubility, and protein-peptide binding affinity. These rigorous assessments affirm PepLand's superior capability in discerning critical representations of peptides with both canonical and non-canonical amino acids, and provide a robust foundation for transformative advances in peptide-focused pharmaceutical research. We have made the entire source code and datasets available at http://www.healthinformaticslab.org/supp/resources.php or https://github.com/zhangruochi/PepLand.
Ruochi Zhang, Chang Liu 0082, Yuting Xiu, Ningning Chen, Yu Wang 0225, Yan Wang 0028, Xin Gao 0001, Fengfeng Zhou
Briefings Bioinform.9
2025 HATZFS predicts pancreatic cancer driver biomarkers by hierarchical reinforcement learning and zero-forcing set
Wenju Hou, Nan Sheng, Chunman Zuo, Yan Wang 0028
Expert Syst. Appl.5
2025 Supervised contrastive knowledge graph learning for ncRNA-disease association prediction
Yan Wang 0028, Xuping Xie, Nan Sheng, Lan Huang 0002, Chunman Zuo
Expert Syst. Appl.1
2025 PCG-CAM: Enhanced class activation map using principal components of gradients and its applications in brain MRI
Lan Huang 0002, Yangguang Shao, Wenju Hou, Yan Wang 0028, Nan Sheng, Yinglu Sun, Yao Wang 0010
Inf. Sci.5
2025 Multi-view fusion based on graph convolutional network with attention mechanism for predicting miRNA related to drugs
abstract
MicroRNAs (miRNAs) play crucial roles in cancer progression, invasion, and response to treatment, particularly in regulating anticancer drug resistance and sensitivity. Identifying potential human miRNA-drug associations (MDAs) that manifest as resistance or sensitivity relationships offers valuable insights for cancer treatment and drug development. With the growing availability of biological data, computational methods have emerged as powerful tools to complement experimental approaches. However, limited attention has been paid to computational prediction of MDAs. Furthermore, existing approaches typically rely on known MDA information, overlooking the valuable insights available from multi-source data related to miRNAs and drugs. In this study, we present a multi-view fusion-based graph convolutional network with attention mechanism (MGCNA) to predict miRNA-associated drug resistance/sensitivity. Specifically, MGCNA integrates macro- and micro- level information of miRNAs and drugs to construct multi-view node features from different perspectives. The proposed multi-view graph convolutional network (GCN) encoder obtains miRNA and disease features from different views and learns adaptive importance weights of the embedding using an attention mechanism. Extensive experiments on manually curated benchmark datasets demonstrate that MGCNA outperforms existing baseline methods. Case studies of two common drugs further establish MGCNA's effectiveness in discovering novel MDAs.
Nan Sheng, Yun-Zhi Liu, Lei Wang 0121, Lan Huang 0002, Yan Wang 0028
PLoS Comput. Biol.6
2025 Self-Supervised Contrastive Learning on Attribute and Topology Graphs for Predicting Relationships Among lncRNAs, miRNAs and Diseases
abstract
Exploring associations between long non-coding RNAs (lncRNAs), microRNAs (miRNAs) and diseases is crucial for disease prevention, diagnosis and treatment. While determining these relationships experimentally is resource-intensive and time-consuming, computational methods have emerged as an attractive way. However, existing computational methods tend to focus on single tasks, neglecting the benefits of leveraging multiple biomolecular interactions and domain-specific knowledge for multi-task prediction. Furthermore, the scarcity of labeled data for lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs) and lncRNA-miRNA interactions (LMIs) poses challenges for comprehensive node embedding learning. This paper proposes a multi-task prediction model (called SSCLMD) that employs self-supervised contrastive learning on attribute and topology graphs to identify potential LDAs, MDAs and LMIs. Firstly, domain knowledge of lncRNAs, miRNAs and diseases as well as their interactions are exploited to construct attribute graph and topology graph, respectively. Then, the nodes are encoded in the attribute and topology spaces to extract the specific and common feature. Meanwhile, the attention mechanism is performed to adaptively fuse the embedding from different views. SSCLMD incorporates contrastive self-supervised learning as a regularize to guide node embedding learning in both attribute and topology space without relying on labels. Severing as a regularize in multi-task learning paradigm, it to improves the model.s generalization capabilities. Extensive experiments on 2 manually curated datasets demonstrate that SSCLMD significantly outperforms baseline methods in LDA, MDA and LMI prediction tasks. Case studies on both old and new datasets further supported SSCLMD's ability to uncover novel disease-related lncRNAs and miRNAs.
Lan Huang 0002, Nan Sheng, Lei Wang 0121, Wenju Hou, Yan Wang 0028
IEEE J. Biomed. Health Informatics7
2024 Object Attribute Matters in Visual Question Answering
abstract
Visual question answering is a multimodal task that requires the joint comprehension of visual and textual information. However, integrating visual and textual semantics solely through attention layers is insufficient to comprehensively understand and align information from both modalities. Intuitively, object attributes can naturally serve as a bridge to unify them, which has been overlooked in previous research. In this paper, we propose a novel VQA approach from the perspective of utilizing object attribute, aiming to achieve better object-level visual-language alignment and multimodal scene understanding. Specifically, we design an attribute fusion module and a contrastive knowledge distillation module. The attribute fusion module constructs a multimodal graph neural network to fuse attributes and visual features through message passing. The enhanced object-level visual features contribute to solving fine-grained problem like counting-question. The better object-level visual-language alignment aids in understanding multimodal scenes, thereby improving the model's robustness. Furthermore, to augment scene understanding and the out-of-distribution performance, the contrastive knowledge distillation module introduces a series of implicit knowledge. We distill knowledge into attributes through contrastive loss, which further strengthens the representation learning of attribute features and facilitates visual-linguistic alignment. Intensive experiments on six datasets, COCO-QA, VQAv2, VQA-CPv2, VQA-CPv1, VQAvs and TDIUC, show the superiority of the proposed method.
Peize Li, Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Yan Wang 0028
AAAI5
2024 2.5D ASF-UNet: Adjacent Slice Spatial Feature Fusion Model for WMH Segmentation from 3D MR Brain Image
abstract
Segmenting brain white matter hyperintensities (WMH) from 3D Magnetic Resonance (MR) images is crucial for the diagnosis, treatment, and prognosis of Multiple Sclerosis (MS). Unlike common 2D images, this task is more challenging and time-consuming. Classical deep learning methods for 3D image segmentation face two main challenges: 1) Pure 3D networks have more parameters and are prone to overfitting. 2) When using 2D networks to segment slices of 3D images, the lack of 3D structural information results in suboptimal segmentation after reconstruction. To address these difficulties, we propose the 2.5D ASF-UNet, which employs the 2.5D workflow and uses adjacent slices as the input for 2D segmentation network ASF-UNet. In ASF-UNet, separate down-sampling paths are used for the adjacent slices, and the Local Spatial Attention Module (LSAM) is designed to more effectively integrate 3D spatial information into the 2D network. Additionally, the Conv_Spectral_Block (CSB) is designed to extract and integrate local and global features. It allows the model to capture global spatial structures while preserving detailed information. Experimental results on MICCAI MSSEG 2016 and Local MS datasets show that 2.5D ASF-UNet achieve better segmentation performance than other deep learning methods.
Lan Huang 0002, Yinglu Sun, Chunjie Guo, Yan Wang 0028
BIBM5
2024 A multi-task prediction method based on neighborhood structure embedding and signed graph representation learning to infer the relationship between circRNA, miRNA, and cancer
abstract
MOTIVATION: Research shows that competing endogenous RNA is widely involved in gene regulation in cells, and identifying the association between circular RNA (circRNA), microRNA (miRNA), and cancer can provide new hope for disease diagnosis, treatment, and prognosis. However, affected by reductionism, previous studies regarded the prediction of circRNA-miRNA interaction, circRNA-cancer association, and miRNA-cancer association as separate studies. Currently, few models are capable of simultaneously predicting these three associations. RESULTS: Inspired by holism, we propose a multi-task prediction method based on neighborhood structure embedding and signed graph representation learning, CMCSG, to infer the relationship between circRNA, miRNA, and cancer. Our method aims to extract feature descriptors of all molecules from the circRNA-miRNA-cancer regulatory network using known types of association information to predict unknown types of molecular associations. Specifically, we first constructed the circRNA-miRNA-cancer association network (CMCN), which is constructed based on the experimentally verified biomedical entity regulatory network; next, we combine topological structure embedding methods to extract feature representations in CMCN from local and global perspectives, and use denoising autoencoder for enhancement; then, combined with balance theory and state theory, molecular features are extracted from the point of social relations through the propagation and aggregation of signed graph attention network; finally, the GBDT classifier is used to predict the association of molecules. The results show that CMCSG can effectively predict the relationship between circRNA, miRNA, and cancer. Additionally, the case studies also demonstrate that CMCSG is capable of accurately identifying biomarkers across various types of cancer. The data and source code can be found at https://github.com/1axin/CMCSG.
Lan Huang 0002, Xinfei Wang 0001, Yan Wang 0028, Renchu Guan, Nan Sheng, Xuping Xie, Lei Wang 0121
Briefings Bioinform.3
2024 Multi-view learning framework for predicting unknown types of cancer markers via directed graph neural networks fitting regulatory networks
abstract
The discovery of diagnostic and therapeutic biomarkers for complex diseases, especially cancer, has always been a central and long-term challenge in molecular association prediction research, offering promising avenues for advancing the understanding of complex diseases. To this end, researchers have developed various network-based prediction techniques targeting specific molecular associations. However, limitations imposed by reductionism and network representation learning have led existing studies to narrowly focus on high prediction efficiency within single association type, thereby glossing over the discovery of unknown types of associations. Additionally, effectively utilizing network structure to fit the interaction properties of regulatory networks and combining specific case biomarker validations remains an unresolved issue in cancer biomarker prediction methods. To overcome these limitations, we propose a multi-view learning framework, CeRVE, based on directed graph neural networks (DGNN) for predicting unknown type cancer biomarkers. CeRVE effectively extracts and integrates subgraph information through multi-view feature learning. Subsequently, CeRVE utilizes DGNN to simulate the entire regulatory network, propagating node attribute features and extracting various interaction relationships between molecules. Furthermore, CeRVE constructed a comparative analysis matrix of three cancers and adjacent normal tissues through The Cancer Genome Atlas and identified multiple types of potential cancer biomarkers through differential expression analysis of mRNA, microRNA, and long noncoding RNA. Computational testing of multiple types of biomarkers for 72 cancers demonstrates that CeRVE exhibits superior performance in cancer biomarker prediction, providing a powerful tool and insightful approach for AI-assisted disease biomarker discovery.
Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Nan Sheng, Xuping Xie, Wenju Hou
Briefings Bioinform.3
2024 A multichannel graph neural network based on multisimilarity modality hypergraph contrastive learning for predicting unknown types of cancer biomarkers
abstract
Identifying potential cancer biomarkers is a key task in biomedical research, providing a promising avenue for the diagnosis and treatment of human tumors and cancers. In recent years, several machine learning-based RNA-disease association prediction techniques have emerged. However, they primarily focus on modeling relationships of a single type, overlooking the importance of gaining insights into molecular behaviors from a complete regulatory network perspective and discovering biomarkers of unknown types. Furthermore, effectively handling local and global topological structural information of nodes in biological molecular regulatory graphs remains a challenge to improving biomarker prediction performance. To address these limitations, we propose a multichannel graph neural network based on multisimilarity modality hypergraph contrastive learning (MML-MGNN) for predicting unknown types of cancer biomarkers. MML-MGNN leverages multisimilarity modality hypergraph contrastive learning to delve into local associations in the regulatory network, learning diverse insights into the topological structures of multiple types of similarities, and then globally modeling the multisimilarity modalities through a multichannel graph autoencoder. By combining representations obtained from local-level associations and global-level regulatory graphs, MML-MGNN can acquire molecular feature descriptors benefiting from multitype association properties and the complete regulatory network. Experimental results on predicting three different types of cancer biomarkers demonstrate the outstanding performance of MML-MGNN. Furthermore, a case study on gastric cancer underscores the outstanding ability of MML-MGNN to gain deeper insights into molecular mechanisms in regulatory networks and prominent potential in cancer biomarker prediction.
Xinfei Wang 0001, Lan Huang 0002, Yan Wang 0028, Renchu Guan, Zhu-Hong You, Nan Sheng, Xuping Xie, Qixing Yang
Briefings Bioinform.3
2024 MolFeSCue: enhancing molecular property prediction in data-limited and imbalanced contexts using few-shot and contrastive learning
abstract
MOTIVATION: Predicting molecular properties is a pivotal task in various scientific domains, including drug discovery, material science, and computational chemistry. This problem is often hindered by the lack of annotated data and imbalanced class distributions, which pose significant challenges in developing accurate and robust predictive models. RESULTS: This study tackles these issues by employing pretrained molecular models within a few-shot learning framework. A novel dynamic contrastive loss function is utilized to further improve model performance in the situation of class imbalance. The proposed MolFeSCue framework not only facilitates rapid generalization from minimal samples, but also employs a contrastive loss function to extract meaningful molecular representations from imbalanced datasets. Extensive evaluations and comparisons of MolFeSCue and state-of-the-art algorithms have been conducted on multiple benchmark datasets, and the experimental data demonstrate our algorithm's effectiveness in molecular representations and its broad applicability across various pretrained models. Our findings underscore MolFeSCues potential to accelerate advancements in drug discovery. AVAILABILITY AND IMPLEMENTATION: We have made all the source code utilized in this study publicly accessible via GitHub at http://www.healthinformaticslab.org/supp/ or https://github.com/zhangruochi/MolFeSCue. The code (MolFeSCue-v1-00) is also available as the supplementary file of this paper.
Ruochi Zhang, Chang Liu 0082, Yan Wang 0028, Lan Huang 0002, Fengfeng Zhou
Bioinform.5
2024 BEROLECMI: a novel prediction method to infer circRNA-miRNA interaction from the role definition of molecular attributes and biological networks
abstract
Circular RNA (CircRNA)-microRNA (miRNA) interaction (CMI) is an important model for the regulation of biological processes by non-coding RNA (ncRNA), which provides a new perspective for the study of human complex diseases. However, the existing CMI prediction models mainly rely on the nearest neighbor structure in the biological network, ignoring the molecular network topology, so it is difficult to improve the prediction performance. In this paper, we proposed a new CMI prediction method, BEROLECMI, which uses molecular sequence attributes, molecular self-similarity, and biological network topology to define the specific role feature representation for molecules to infer the new CMI. BEROLECMI effectively makes up for the lack of network topology in the CMI prediction model and achieves the highest prediction performance in three commonly used data sets. In the case study, 14 of the 15 pairs of unknown CMIs were correctly predicted.
Xinfei Wang 0001, Zhu-Hong You, Yan Wang 0028, Lan Huang 0002, Yan Qiao 0002, Lei Wang 0121, Zhengwei Li 0001
BMC Bioinform.4
2024 MMGAT: a graph attention network framework for ATAC-seq motifs finding
abstract
BACKGROUND: Motif finding in Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) data is essential to reveal the intricacies of transcription factor binding sites (TFBSs) and their pivotal roles in gene regulation. Deep learning technologies including convolutional neural networks (CNNs) and graph neural networks (GNNs), have achieved success in finding ATAC-seq motifs. However, CNN-based methods are limited by the fixed width of the convolutional kernel, which makes it difficult to find multiple transcription factor binding sites with different lengths. GNN-based methods has the limitation of using the edge weight information directly, makes it difficult to aggregate the neighboring nodes' information more efficiently when representing node embedding. RESULTS: To address this challenge, we developed a novel graph attention network framework named MMGAT, which employs an attention mechanism to adjust the attention coefficients among different nodes. And then MMGAT finds multiple ATAC-seq motifs based on the attention coefficients of sequence nodes and k-mer nodes as well as the coexisting probability of k-mers. Our approach achieved better performance on the human ATAC-seq datasets compared to existing tools, as evidenced the highest scores on the precision, recall, F1_score, ACC, AUC, and PRC metrics, as well as finding 389 higher quality motifs. To validate the performance of MMGAT in predicting TFBSs and finding motifs on more datasets, we enlarged the number of the human ATAC-seq datasets to 180 and newly integrated 80 mouse ATAC-seq datasets for multi-species experimental validation. Specifically on the mouse ATAC-seq dataset, MMGAT also achieved the highest scores on six metrics and found 356 higher-quality motifs. To facilitate researchers in utilizing MMGAT, we have also developed a user-friendly web server named MMGAT-S that hosts the MMGAT method and ATAC-seq motif finding results. CONCLUSIONS: The advanced methodology MMGAT provides a robust tool for finding ATAC-seq motifs, and the comprehensive server MMGAT-S makes a significant contribution to genomics research. The open-source code of MMGAT can be found at https://github.com/xiaotianr/MMGAT , and MMGAT-S is freely available at https://www.mmgraphws.com/MMGAT-S/ .
Wenju Hou, Lan Huang 0002, Nan Sheng, Qixing Yang, Shuangquan Zhang, Yan Wang 0028
BMC Bioinform.8
2024 Deep learning model for human-intuitive shoeprint reconstruction
Yan Wang 0028, Di Wang 0004, Wei Pang 0001, Daixi Li, You Zhou 0008, Dong Xu 0002, Sami Ur Rahman, Amin ur Rahman, Ahmed Ameen Fateh, Peiwu Qin
Expert Syst. Appl.2
2024 FairCare: Adversarial training of a heterogeneous graph neural network with attention mechanism to learn fair representations of electronic health records
Yan Wang 0028, Ruochi Zhang, Qiong Zhou, Shengde Zhang, Yusi Fan, Lan Huang 0002, Fengfeng Zhou
Inf. Process. Manag.1
2024 A Survey of Deep Learning for Detecting miRNA- Disease Associations: Databases, Computational Methods, Challenges, and Future Directions
abstract
MicroRNAs (miRNAs) are an important class of non-coding RNAs that play an essential role in the occurrence and development of various diseases. Identifying the potential miRNA-disease associations (MDAs) can be beneficial in understanding disease pathogenesis. Traditional laboratory experiments are expensive and time-consuming. Computational models have enabled systematic large-scale prediction of potential MDAs, greatly improving the research efficiency. With recent advances in deep learning, it has become an attractive and powerful technique for uncovering novel MDAs. Consequently, numerous MDA prediction methods based on deep learning have emerged. In this review, we first summarize publicly available databases related to miRNAs and diseases for MDA prediction. Next, we outline commonly used miRNA and disease similarity calculation and integration methods. Then, we comprehensively review the 48 existing deep learning-based MDA computation methods, categorizing them into classical deep learning and graph neural network-based techniques. Subsequently, we investigate the evaluation methods and metrics that are frequently used to assess MDA prediction performance. Finally, we discuss the performance trends of different computational methods, point out some problems in current research, and propose 9 potential future research directions. Data resources and recent advances in MDA prediction methods are summarized in the GitHub repository https://github.com/sheng-n/DL-miRNA-disease-association-methods.
Nan Sheng, Xuping Xie, Yan Wang 0028, Lan Huang 0002, Shuangquan Zhang
IEEE ACM Trans. Comput. Biol. Bioinform.3
2024 Cross-modality Multiple Relations Learning for Knowledge-based Visual Question Answering
abstract
Knowledge-based visual question answering not only needs to answer the questions based on images but also incorporates external knowledge to study reasoning in the joint space of vision and language. To bridge the gap between visual content and semantic cues, it is important to capture the question-related and semantics-rich vision-language connections. Most existing solutions model simple intra-modality relation or represent cross-modality relation using a single vector, which makes it difficult to effectively model complex connections between visual features and question features. Thus, we propose a cross-modality multiple relations learning model, aiming to better enrich cross-modality representations and construct advanced multi-modality knowledge triplets. First, we design a simple yet effective method to generate multiple relations that represent the rich cross-modality relations. The various cross-modality relations link the textual question to the related visual objects. These multi-modality triplets efficiently align the visual objects and corresponding textual answers. Second, to encourage multiple relations to better align with different semantic relations, we further formulate a novel global-local loss. The global loss enables the visual objects and corresponding textual answers close to each other through cross-modality relations in the vision-language space, and the local loss better preserves semantic diversity among multiple relations. Experimental results on the Outside Knowledge VQA and Knowledge-Routed Visual Question Reasoning datasets demonstrate that our model outperforms the state-of-the-art methods.
Yan Wang 0028, Peize Li, Qingyi Si, Hanwen Zhang 0010, Wenyu Zang, Zheng Lin 0001, Peng Fu 0008
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Contrastive self-supervised graph convolutional network for detecting the relationship among lncRNAs, miRNAs, and diseases
abstract
Inferring potential relationships among long non-coding RNAs (lncRNAs), microRNAs (miRNAs), and diseases play a crucial role in investigation of disease aetiology and pathogenesis. Due to the high cost of laboratory experiments, there is a practical requirement to develop appropriate computational methods that promise to accelerate the experimental screening process for potential lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs), and lncRNA-miRNA interactions (LMIs). However, most existing methods are applied to predict LDAs, MDAs, and LMIs in specific domains, neglecting the important benefits of integrating multiple sources data and limiting the ability of transferring models to other tasks. Furthermore, with the high sparsity of LDA, MDA, and LMI data, it is difficult for many computational models to exploit enough knowledge to learn the comprehensive patterns of node embedding. In this study, inspired by the recent success of graph contrastive learning, we develop a Contrastive Self-supervised Graph convolutional network to identify potential LDAs, MDAs, and LMIs (called CSGLMD). CSGLMD combines supervised learning and self-supervised learning to fully capture node features. Specifically, CSGLMD primarily leverages the rich association and similarity relationships among lncRNA, miRNA, and disease to construct a lncRNA-miRNA-disease heterogeneous graph (LMDHG) that contains three types of biological entities. It can effectively embed multi-source biological data and assist the model extension to other prediction tasks. In addition, we consider applying a label instantiation mechanism to make the LMDHG better adapt graph neural network structures and control the strength of similarity relationships between the same biological entities. Secondly, CSGLMD implements graph convolutional network (GCN) as encoder to extract node embedding features from the LMDHG, and utilizes a multi-relational modelling decoder to predict LDAs, MDAs, or LMIs. Finally, we designed a contrastive self-supervised learning task that guides the learning of node embeddings without relying on labels, and acts as a regularize in a multi-task learning paradigm to enhance the generalization ability of the model. Extensive results on two datasets (from the old and new versions of the database, respectively) show that CSGLMD significantly outperforms 12 state-of-the-art methods (5 LDA prediction and 7 MDA prediction) in predicting disease-associated lncRNAs and miRNAs. Case studies on old and new datasets can further demonstrate the capability of CSGLMD to discover disease-related new candidate lncRNAs and miRNAs. The source data and code for the proposed model are publicly available on https://github.com/sheng-n/CSGLMD.
Nan Sheng, Lan Huang 0002, Yan Wang 0028, Huiyan Sun, Xuping Xie
BIBM3
2023 Multi-task prediction-based graph contrastive learning for inferring the relationship among lncRNAs, miRNAs and diseases
abstract
MOTIVATION: Identifying the relationships among long non-coding RNAs (lncRNAs), microRNAs (miRNAs) and diseases is highly valuable for diagnosing, preventing, treating and prognosing diseases. The development of effective computational prediction methods can reduce experimental costs. While numerous methods have been proposed, they often to treat the prediction of lncRNA-disease associations (LDAs), miRNA-disease associations (MDAs) and lncRNA-miRNA interactions (LMIs) as separate task. Models capable of predicting all three relationships simultaneously remain relatively scarce. Our aim is to perform multi-task predictions, which not only construct a unified framework, but also facilitate mutual complementarity of information among lncRNAs, miRNAs and diseases. RESULTS: In this work, we propose a novel unsupervised embedding method called graph contrastive learning for multi-task prediction (GCLMTP). Our approach aims to predict LDAs, MDAs and LMIs by simultaneously extracting embedding representations of lncRNAs, miRNAs and diseases. To achieve this, we first construct a triple-layer lncRNA-miRNA-disease heterogeneous graph (LMDHG) that integrates the complex relationships between these entities based on their similarities and correlations. Next, we employ an unsupervised embedding model based on graph contrastive learning to extract potential topological feature of lncRNAs, miRNAs and diseases from the LMDHG. The graph contrastive learning leverages graph convolutional network architectures to maximize the mutual information between patch representations and corresponding high-level summaries of the LMDHG. Subsequently, for the three prediction tasks, multiple classifiers are explored to predict LDA, MDA and LMI scores. Comprehensive experiments are conducted on two datasets (from older and newer versions of the database, respectively). The results show that GCLMTP outperforms other state-of-the-art methods for the disease-related lncRNA and miRNA prediction tasks. Additionally, case studies on two datasets further demonstrate the ability of GCLMTP to accurately discover new associations. To ensure reproducibility of this work, we have made the datasets and source code publicly available at https://github.com/sheng-n/GCLMTP.
Nan Sheng, Yan Wang 0028, Lan Huang 0002, Yangkun Cao, Xuping Xie
Briefings Bioinform.2
2023 Predicting miRNA-disease associations based on PPMI and attention network
abstract
BACKGROUND: With the development of biotechnology and the accumulation of theories, many studies have found that microRNAs (miRNAs) play an important role in various diseases. Uncovering the potential associations between miRNAs and diseases is helpful to better understand the pathogenesis of complex diseases. However, traditional biological experiments are expensive and time-consuming. Therefore, it is necessary to develop more efficient computational methods for exploring underlying disease-related miRNAs. RESULTS: In this paper, we present a new computational method based on positive point-wise mutual information (PPMI) and attention network to predict miRNA-disease associations (MDAs), called PATMDA. Firstly, we construct the heterogeneous MDA network and multiple similarity networks of miRNAs and diseases. Secondly, we respectively perform random walk with restart and PPMI on different similarity network views to get multi-order proximity features and then obtain high-order proximity representations of miRNAs and diseases by applying the convolutional neural network to fuse the learned proximity features. Then, we design an attention network with neural aggregation to integrate the representations of a node and its heterogeneous neighbor nodes according to the MDA network. Finally, an inner product decoder is adopted to calculate the relationship scores between miRNAs and diseases. CONCLUSIONS: PATMDA achieves superior performance over the six state-of-the-art methods with the area under the receiver operating characteristic curve of 0.933 and 0.946 on the HMDD v2.0 and HMDD v3.2 datasets, respectively. The case studies further demonstrate the validity of PATMDA for discovering novel disease-associated miRNAs.
Xuping Xie, Yan Wang 0028, Nan Sheng
BMC Bioinform.2
2023 A Survey of Computational Methods and Databases for lncRNA-MiRNA Interaction Prediction
abstract
Long non-coding RNAs (lncRNAs) and microRNAs (miRNAs) are two prevalent non-coding RNAs in current research. They play critical regulatory roles in the life processes of animals and plants. Studies have shown that lncRNAs can interact with miRNAs to participate in post-transcriptional regulatory processes, mainly involved in regulating cancer development, metastatic progression, and drug resistance. Additionally, these interactions have significant effects on plant growth, development, and responses to biotic and abiotic stresses. Deciphering the potential relationships between lncRNAs and miRNAs may provide new insights into our understanding of the biological functions of lncRNAs and miRNAs, and the pathogenesis of complex diseases. In contrast, gathering information on lncRNA-miRNA interactions (LMIs) through biological experiments is expensive and time-consuming. With the accumulation of multi-omics data, computational models are extremely attractive in systematically exploring potential LMIs. To the best of our knowledge, this is the first comprehensive review of computational methods for identifying LMIs. Specifically, we first summarized the available public databases for predicting animal and plant LMIs. Second, we comprehensively reviewed the computational methods for predicting LMIs and classified them into two categories, including network-based methods and sequence-based methods. Third, we analyzed the standard evaluation methods and metrics used in LMI prediction. Finally, we pointed out some problems in the current study and discuss future research directions. Relevant databases and the latest advances in LMI prediction are summarized in a GitHub repository https://github.com/sheng-n/lncRNA-miRNA-interaction-methods, and we'll keep it updated.
Nan Sheng, Lan Huang 0002, Yangkun Cao, Xuping Xie, Yan Wang 0028
IEEE ACM Trans. Comput. Biol. Bioinform.6
2022 RDDriver: A novel method based on multi-layer heterogeneous transcriptional regulation network for identifying pancreatic cancer biomarker
abstract
Pancreatic cancer is a malignant cancer with rapid progression and poor prognosis. The use of transcriptional data can be effective in finding cancer biomarkers. Most of the existing network-based methods do not study RNA, rely on prior or mutation information, or can only achieve classification tasks. In this paper, wep ropose a method combining Relational Graph Convolutional Network and Deep Q-Network called RDDriver to identify pancreatic cancer biomarkers based on multi-layer heterogeneous transcriptional regulation network. We first construct a regulation network containing three types RNAs. Then, Relational Graph Convolutional Network is used to learn the node representation. Finally, we combine the idea of Deep Q-Network to prioritize each RNA with the network controllability theory. We train RDDriver on three simulated small networks, and calculate the average score after applying the model parameters to the regulation networks. To demonstrate the effectiveness of the method, we compare RDDriver with other eight methods based on the approximate cancer drivers benchmark RNAs.
Yan Wang 0028, Nan Sheng, Yuan Tian 0016
BIBM2
2022 Artificial intelligence in clinical research of cancers
abstract
Several factors, including advances in computational algorithms, the availability of high-performance computing hardware, and the assembly of large community-based databases, have led to the extensive application of Artificial Intelligence (AI) in the biomedical domain for nearly 20 years. AI algorithms have attained expert-level performance in cancer research. However, only a few AI-based applications have been approved for use in the real world. Whether AI will eventually be capable of replacing medical experts has been a hot topic. In this article, we first summarize the cancer research status using AI in the past two decades, including the consensus on the procedure of AI based on an ideal paradigm and current efforts of the expertise and domain knowledge. Next, the available data of AI process in the biomedical domain are surveyed. Then, we review the methods and applications of AI in cancer clinical research categorized by the data types including radiographic imaging, cancer genome, medical records, drug information and biomedical literatures. At last, we discuss challenges in moving AI from theoretical research to real-world cancer research applications and the perspectives toward the future realization of AI participating cancer treatment.
Dan Shao, Yinfei Dai, Nianfeng Li, Xuqing Cao, Zhuqing Rong, Lan Huang 0002, Yan Wang 0028
Briefings Bioinform.9
2022 Multi-channel graph attention autoencoders for disease-related lncRNAs prediction
abstract
MOTIVATION: Predicting disease-related long non-coding RNAs (lncRNAs) can be used as the biomarkers for disease diagnosis and treatment. The development of effective computational prediction approaches to predict lncRNA-disease associations (LDAs) can provide insights into the pathogenesis of complex human diseases and reduce experimental costs. However, few of the existing methods use microRNA (miRNA) information and consider the complex relationship between inter-graph and intra-graph in complex-graph for assisting prediction. RESULTS: In this paper, the relationships between the same types of nodes and different types of nodes in complex-graph are introduced. We propose a multi-channel graph attention autoencoder model to predict LDAs, called MGATE. First, an lncRNA-miRNA-disease complex-graph is established based on the similarity and correlation among lncRNA, miRNA and diseases to integrate the complex association among them. Secondly, in order to fully extract the comprehensive information of the nodes, we use graph autoencoder networks to learn multiple representations from complex-graph, inter-graph and intra-graph. Thirdly, a graph-level attention mechanism integration module is adopted to adaptively merge the three representations, and a combined training strategy is performed to optimize the whole model to ensure the complementary and consistency among the multi-graph embedding representations. Finally, multiple classifiers are explored, and Random Forest is used to predict the association score between lncRNA and disease. Experimental results on the public dataset show that the area under receiver operating characteristic curve and area under precision-recall curve of MGATE are 0.964 and 0.413, respectively. MGATE performance significantly outperformed seven state-of-the-art methods. Furthermore, the case studies of three cancers further demonstrate the ability of MGATE to identify potential disease-correlated candidate lncRNAs. The source code and supplementary data are available at https://github.com/sheng-n/MGATE. CONTACT: [email protected], [email protected].
Nan Sheng, Lan Huang 0002, Yan Wang 0028, Ping Xuan, Yangkun Cao
Briefings Bioinform.3
2022 Assessing deep learning methods in cis-regulatory motif finding based on genomic sequencing data
abstract
Identifying cis-regulatory motifs from genomic sequencing data (e.g. ChIP-seq and CLIP-seq) is crucial in identifying transcription factor (TF) binding sites and inferring gene regulatory mechanisms for any organism. Since 2015, deep learning (DL) methods have been widely applied to identify TF binding sites and predict motif patterns, with the strengths of offering a scalable, flexible and unified computational approach for highly accurate predictions. As far as we know, 20 DL methods have been developed. However, without a clear and systematic assessment, users will struggle to choose the most appropriate tool for their specific studies. In this manuscript, we evaluated 20 DL methods for cis-regulatory motif prediction using 690 ENCODE ChIP-seq, 126 cancer ChIP-seq and 55 RNA CLIP-seq data. Four metrics were investigated, including the accuracy of motif finding, the performance of DNA/RNA sequence classification, algorithm scalability and tool usability. The assessment results demonstrated the high complementarity of the existing DL methods. It was determined that the most suitable model should primarily depend on the data size and type and the method's outputs.
Shuangquan Zhang, Anjun Ma, Dong Xu 0002, Qin Ma 0003, Yan Wang 0028
Briefings Bioinform.6
2022 MMGraph: a multiple motif predictor based on graph neural network and coexisting probability for ATAC-seq data
abstract
MOTIVATION: Transcription factor binding sites (TFBSs) prediction is a crucial step in revealing functions of transcription factors from high-throughput sequencing data. Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) provides insight on TFBSs and nucleosome positioning by probing open chromatic, which can simultaneously reveal multiple TFBSs compare to traditional technologies. The existing tools based on convolutional neural network (CNN) only find the fixed length of TFBSs from ATAC-seq data. Graph neural network (GNN) can be considered as the extension of CNN, which has great potential in finding multiple TFBSs with different lengths from ATAC-seq data. RESULTS: We develop a motif predictor called MMGraph based on three-layer GNN and coexisting probability of k-mers for finding multiple motifs from ATAC-seq data. The results of the experiment which has been conducted on 88 ATAC-seq datasets indicate that MMGraph has achieved the best performance on area of eight metrics radar score of 2.31 and could find 207 higher-quality multiple motifs than other existing tools. AVAILABILITY AND IMPLEMENTATION: MMGraph is wrapped in Python package, which is available at https://github.com/zhangsq06/MMGraph.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shuangquan Zhang, Lili Yang 0004, Nan Sheng, Anjun Ma, Yan Wang 0028
Bioinform.7
2022 Restorable-inpainting: A novel deep learning approach for shoeprint restoration
Yan Wang 0028, Di Wang 0004, Wei Pang 0001, Kangping Wang, Daixi Li, You Zhou 0008, Dong Xu 0002
Inf. Sci.2
2021 An AutoEncoder-Based Matrix Factorization Approach to Estimating Cell Proportion from Bulk Tumor RNA-seq Data
abstract
The deconvolution of infiltrating immune cells and stromal cells from complex tumor tissues is significant for studying the impact of these cells on tumor development, as well as assisting cancer therapies. Integrating cell-specific marker genes from several references, we put forward an AutoEncoder-based matrix factorization method to estimate the cell proportion in heterogeneous samples. The proportion predicted by our method achieved over 0.95 Pearson correlation coefficient (PCC) with the ground truth. Moreover, through analyzing association between cell proportion of tumor tissues and clinical information of the tumor patients, we found that the proportion of cancer-associated fibroblast (CAF) gradually increased with the progress of tumor, and the proportion of B cell in tissues was significantly related to the five-year survival rate of tumor patients.
Yingze Xu, Yan Wang 0028, Xuping Xie, Huiyan Sun
BIBM2
2021 Research and Application of Reinforcement Learning Recommendation Method for Taobao
abstract
Nowadays, many e-commerce companies are using reinforcement learning recommendation methods to maximize long-term benefits. Alibaba Group and Nanjing University build “Virtual Taobao”, a Taobao simulator. In this paper, we proposed TTD3 based on TD3 and trained it in Virtual Taobao. There are three important improvements in TTD3's training process. First, the current actor-network and target actor-network will predict two candidate actions for Virtual Taobao's current state, and the action with a larger value evaluated by the current critic-network is selected as the final execution action. Second, the Ornstein-Uhlenbeck (OU) process is used as the exploration noise to improve the agent's ability to explore Virtual Taobao. Third, prioritized experience replay is adopted to improve sampling efficiency. TTD3 achieves the highest average CTR of about 0.85 in Virtual Taobao which is superior to TD3 as well as DPPO, SAC, and DDPG used by Virtual Taobao's author.
Lan Huang 0002, Yan Wang 0028, Xuping Xie
ISCC3
2021 The bioinformatics toolbox for circRNA discovery and analysis
abstract
Circular RNAs (circRNAs) are a unique class of RNA molecule identified more than 40 years ago which are produced by a covalent linkage via back-splicing of linear RNA. Recent advances in sequencing technologies and bioinformatics tools have led directly to an ever-expanding field of types and biological functions of circRNAs. In parallel with technological developments, practical applications of circRNAs have arisen including their utilization as biomarkers of human disease. Currently, circRNA-associated bioinformatics tools can support projects including circRNA annotation, circRNA identification and network analysis of competing endogenous RNA (ceRNA). In this review, we collected about 100 circRNA-associated bioinformatics tools and summarized their current attributes and capabilities. We also performed network analysis and text mining on circRNA tool publications in order to reveal trends in their ongoing development.
Liang Chen 0021, Changliang Wang, Huiyan Sun, Juexin Wang, Yanchun Liang 0001, Yan Wang 0028, Garry Wong
Briefings Bioinform.6
2021 Human body-fluid proteome: quantitative profiling and computational prediction
abstract
Empowered by the advancement of high-throughput bio technologies, recent research on body-fluid proteomes has led to the discoveries of numerous novel disease biomarkers and therapeutic drugs. In the meantime, a tremendous progress in disclosing the body-fluid proteomes was made, resulting in a collection of over 15 000 different proteins detected in major human body fluids. However, common challenges remain with current proteomics technologies about how to effectively handle the large variety of protein modifications in those fluids. To this end, computational effort utilizing statistical and machine-learning approaches has shown early successes in identifying biomarker proteins in specific human diseases. In this article, we first summarized the experimental progresses using a combination of conventional and high-throughput technologies, along with the major discoveries, and focused on current research status of 16 types of body-fluid proteins. Next, the emerging computational work on protein prediction based on support vector machine, ranking algorithm, and protein-protein interaction network were also surveyed, followed by algorithm and application discussion. At last, we discuss additional critical concerns about these topics and close the review by providing future perspectives especially toward the realization of clinical disease biomarker discovery.
Lan Huang 0002, Dan Shao, Yan Wang 0028, Xueteng Cui, Juan Cui
Briefings Bioinform.3
2021 DeepSec: a deep learning framework for secreted protein discovery in human body fluids
abstract
MOTIVATION: Human proteins that are secreted into different body fluids from various cells and tissues can be promising disease indicators. Modern proteomics research empowered by both qualitative and quantitative profiling techniques has made great progress in protein discovery in various human fluids. However, due to the large number of proteins and diverse modifications present in the fluids, as well as the existing technical limits of major proteomics platforms (e.g. mass spectrometry), large discrepancies are often generated from different experimental studies. As a result, a comprehensive proteomics landscape across major human fluids are not well determined. RESULTS: To bridge this gap, we have developed a deep learning framework, named DeepSec, to identify secreted proteins in 12 types of human body fluids. DeepSec adopts an end-to-end sequence-based approach, where a Convolutional Neural Network is built to learn the abstract sequence features followed by a Bidirectional Gated Recurrent Unit with fully connected layer for protein classification. DeepSec has demonstrated promising performances with average area under the ROC curves of 0.85-0.94 on testing datasets in each type of fluids, which outperforms existing state-of-the-art methods available mostly on blood proteins. As an illustration of how to apply DeepSec in biomarker discovery research, we conducted a case study on kidney cancer by using genomics data from the cancer genome atlas and have identified 104 possible marker proteins. AVAILABILITY: DeepSec is available at https://bmbl.bmi.osumc.edu/deepsec/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Dan Shao, Lan Huang 0002, Yan Wang 0028, Xueteng Cui, Yao Wang 0010, Qin Ma 0003, Juan Cui
Bioinform.3
2019 Computational Prediction of Human Body-Fluid Protein
abstract
Research on body fluid proteomes has led to the discoveries of context-dependent proteomics profiles and numerous novel disease biomarkers. Common challenges remain with current proteomics technologies about how to effectively handle the large variety of protein modifications in those fluids. To this end, computational efforts have shown early successes in identifying biomarker proteins in specific human diseases. In this article, we first reviewed published computational methods on this topic and then presented a new database system and a novel deep learning-based method for predicting body-fluid proteins in 16 types of human fluids. The results show that our new system outperforms existing methods and provides a highly promising tool to facilitate clinical proteomic discovery.
Dan Shao, Lan Huang 0002, Yan Wang 0028, Xueteng Cui, Yao Wang 0010
BIBM3
2018 ε-Distance Weighted Support Vector Regression
Ge Ou, Yan Wang 0028, Lan Huang 0002, Wei Pang 0001, George Macleod Coghill
PAKDD (1)2
2016 Partitioning Clustering Based on Support Vector Ranking
Qing Peng, Yan Wang 0028, Ge Ou, Yuan Tian 0016, Lan Huang 0002, Wei Pang 0001
ADMA2
2016 PUEPro: A Computational Pipeline for Prediction of Urine Excretory Proteins
Yan Wang 0028, Wei Du 0002, Yanchun Liang 0001, Xin Chen 0113, Chi Zhang 0021, Wei Pang 0001, Ying Xu 0001
ADMA1
2016 A method for exploring implicit concept relatedness in biomedical knowledge network
abstract
BACKGROUND: Biomedical information and knowledge, structural and non-structural, stored in different repositories can be semantically connected to form a hybrid knowledge network. How to compute relatedness between concepts and discover valuable but implicit information or knowledge from it effectively and efficiently is of paramount importance for precision medicine, and a major challenge facing the biomedical research community. RESULTS: In this study, a hybrid biomedical knowledge network is constructed by linking concepts across multiple biomedical ontologies as well as non-structural biomedical knowledge sources. To discover implicit relatedness between concepts in ontologies for which potentially valuable relationships (implicit knowledge) may exist, we developed a Multi-Ontology Relatedness Model (MORM) within the knowledge network, for which a relatedness network (RN) is defined and computed across multiple ontologies using a formal inference mechanism of set-theoretic operations. Semantic constraints are designed and implemented to prune the search space of the relatedness network. CONCLUSIONS: Experiments to test examples of several biomedical applications have been carried out, and the evaluation of the results showed an encouraging potential of the proposed approach to biomedical knowledge discovery.
Tian Bai 0002, Leiguang Gong, Yan Wang 0028, Casimir A. Kulikowski, Lan Huang 0002
BMC Bioinform.4
2014 Essential protein identification based on essential protein-protein interaction prediction by integrated edge weights
abstract
Essential proteins are crucial to cellular survival and development. Traditionally, essential proteins are identified by knock-out experiments, which are expensive and often fatal to the target organisms. Regarding this, an important approach to essential protein identification is through computational prediction. In this research, we present a novel computational method, Integrated Edge Weights (IEW), to innovatively predict proteins' essentiality based on essential protein-protein interactions. The experimental results on all three organisms: Saccharomyces cere-visiae (Yeast), Escherichia coli (E. coli), and Caenorhabditis ele-gans (C. elegans) show that IEW achieves better performance than the state-of-the-art methods in terms of precision-recall. Furthermore, we have demonstrated that the highly-ranked protein-protein interactions predicted by our approach tend to be biologically significant in Yeast, E. coli, and C. elegans protein-protein interaction (PPI) networks.
Yuexu Jiang, Yan Wang 0028, Wei Pang 0001, Liang Chen 0021, Huiyan Sun, Yanchun Liang 0001, Enrico Blanzieri
BIBM2
2014 Computational Prediction of Human Saliva-Secreted Proteins
Chunguang Zhou, Zhongbo Cao, Wei Du 0002, Yan Wang 0028
ISBRA6
2013 Effective and stable feature selection method based on filter for gene signature identification in paired microarray data
abstract
A huge amount of microarray datasets are produced with big number of genes and small samples. Feature selection methods have become a very sharp tool to select the gene signatures from the whole gene set. In recent years, researchers are concerned much about the datasets containing samples of cancer as well as corresponding control tissues. However, few feature selection methods consider the effect of paired samples. In this article, we propose a new feature selection method for paired microarray datasets based on the original paired t-test approach. We apply on the paired datasets across six common cancer types. Through comparison with some widely used methods on the performance of prediction power, stability of gene lists and functional stability, our method shows excellent performance. The proposed method has good effectiveness, stability and consistency, which enables the method to be applicative to feature selection for paired microarray expression data analysis.
Zhongbo Cao, Yan Wang 0028, Wei Du 0002, Yanchun Liang 0001
BIBM2
2010 PMirP: A pre-microRNA prediction method based on structure-sequence hybrid features
Dongyu Zhao, Yan Wang 0028, Xiaohu Shi, Liupu Wang, Dong Xu 0002, Yanchun Liang 0001
Artif. Intell. Medicine2
2009 Immune Particle Swarm Optimization for Support Vector Regression on Forest Fire Prediction
Yan Wang 0028, Juexin Wang, Wei Du 0002, Chuncai Wang, Yanchun Liang 0001, Chunguang Zhou, Lan Huang 0002
ISNN (2)1
2009 Methods for labeling error detection in microarrays based on the effect of data perturbation on the regression model
abstract
MOTIVATION: Mislabeled samples often appear in gene expression profile because of the similarity of different sub-type of disease and the subjective misdiagnosis. The mislabeled samples deteriorate supervised learning procedures. The LOOE-sensitivity algorithm is an approach for mislabeled sample detection for microarray based on data perturbation. However, the failure of measuring the perturbing effect makes the LOOE-sensitivity algorithm a poor performance. The purpose of this article is to design a novel detection method for mislabeled samples of microarray, which could take advantage of the measuring effect of data perturbations. RESULTS: To measure the effect of data perturbation, we define an index named perturbing influence value (PIV), based on the support vector machine (SVM) regression model. The Column Algorithm (CAPIV), Row Algorithm (RAPIV) and progressive Row Algorithm (PRAPIV) based on the PIV value are proposed to detect the mislabeled samples. Experimental results obtained by using six artificial datasets and five microarray datasets demonstrate that all proposed methods in this article are superior to LOOE-sensitivity. Moreover, compared with the simple SVM and CL-stability, the PRAPIV algorithm shows an increase in precision and high recall. AVAILABILITY: The program and source code (in JAVA) are publicly available at http://ccst.jlu.edu.cn/CSBG/PIVS/index.htm
Chunguo Wu, Enrico Blanzieri, You Zhou 0008, Yan Wang 0028, Wei Du 0002, Yanchun Liang 0001
Bioinform.5
2008 Pep-3D-Search: a method for B-cell epitope prediction based on mimotope analysis
abstract
BACKGROUND: The prediction of conformational B-cell epitopes is one of the most important goals in immunoinformatics. The solution to this problem, even if approximate, would help in designing experiments to precisely map the residues of interaction between an antigen and an antibody. Consequently, this area of research has received considerable attention from immunologists, structural biologists and computational biologists. Phage-displayed random peptide libraries are powerful tools used to obtain mimotopes that are selected by binding to a given monoclonal antibody (mAb) in a similar way to the native epitope. These mimotopes can be considered as functional epitope mimics. Mimotope analysis based methods can predict not only linear but also conformational epitopes and this has been the focus of much research in recent years. Though some algorithms based on mimotope analysis have been proposed, the precise localization of the interaction site mimicked by the mimotopes is still a challenging task. RESULTS: In this study, we propose a method for B-cell epitope prediction based on mimotope analysis called Pep-3D-Search. Given the 3D structure of an antigen and a set of mimotopes (or a motif sequence derived from the set of mimotopes), Pep-3D-Search can be used in two modes: mimotope or motif. To evaluate the performance of Pep-3D-Search to predict epitopes from a set of mimotopes, 10 epitopes defined by crystallography were compared with the predicted results from a Pep-3D-Search: the average Matthews correlation coefficient (MCC), sensitivity and precision were 0.1758, 0.3642 and 0.6948. Compared with other available prediction algorithms, Pep-3D-Search showed comparable MCC, specificity and precision, and could provide novel, rational results. To verify the capability of Pep-3D-Search to align a motif sequence to a 3D structure for predicting epitopes, 6 test cases were used. The predictive performance of Pep-3D-Search was demonstrated to be superior to that of other similar programs. Furthermore, a set of test cases with different lengths of sequences was constructed to examine Pep-3D-Search's capability in searching sequences on a 3D structure. The experimental results demonstrated the excellent search capability of Pep-3D-Search, especially when the length of the query sequence becomes longer; the iteration numbers of Pep-3D-Search to precisely localize the target paths did not obviously increase. This means that Pep-3D-Search has the potential to quickly localize the epitope regions mimicked by longer mimotopes. CONCLUSION: Our Pep-3D-Search provides a powerful approach for localizing the surface region mimicked by the mimotopes. As a publicly available tool, Pep-3D-Search can be utilized and conveniently evaluated, and it can also be used to complement other existing tools. The data sets and open source code used to obtain the results in this paper are available on-line and as supplementary material. More detailed materials may be accessed at (http://kyc.nenu.edu.cn/Pep3DSearch/).
Yanxin Huang, Yongli Bao, Shu Yan Guo, Yan Wang 0028, Chunguang Zhou, Yuxin Li 0001
BMC Bioinform.4
2007 Operon Prediction Using Neural Network Based on Multiple Information of Log-Likelihoods
Wei Du 0002, Yan Wang 0028, Fangxun Sun, Chunguang Zhou, Chengquan Hu, Yanchun Liang 0001
ISNN (1)2
2007 A Novel Method for Prediction of Protein Domain Using Distance-Based Maximal Entropy
Shu-Xue Zou, Yanxin Huang, Yan Wang 0028, Chengquan Hu, Yanchun Liang 0001, Chunguang Zhou
ISNN (2)3
2007 A multi-approaches-guided genetic algorithm with application to operon prediction
Yan Wang 0028, Wei Du 0002, Fangxun Sun, Chunguang Zhou, Yanchun Liang 0001
Artif. Intell. Medicine2
2007 A novel quantum swarm evolutionary algorithm and its applications
Yan Wang 0028, Xiaoyue Feng, Yanxin Huang, Dongbing Pu, Yanchun Liang 0001, Chunguang Zhou
Neurocomputing1
2006 Prediction of Protein Domains from Sequence Information Using Support Vector Machines
Shu-Xue Zou, Yanxin Huang, Yan Wang 0028, Chunguang Zhou
ISNN (2)3
2006 Self-adaptive Two-Phase Support Vector Clustering for Multi-Relational Data Mining
Ping Ling, Yan Wang 0028, Chunguang Zhou
PAKDD2
2005 Two-Phase Support Vector Clustering for Multi-Relational Data Mining
abstract
A novel two-phase support vector clustering (TPSVC) algorithm is proposed in this paper, which is implemented in multi-relational data mining (MRDM). Based on the designed kernel which is incorporated with MRDM environment, TPSVC provides an appreciate description of cluster contours using support vectors at the first step and then a support vector machine (SVM) classification procedure is employed to further extract the information of cluster central zones. The algorithm does the cluster assignment according to desired definition of affinity without suffering the expensive operations of adjacent matrix computation used in traditional support vector clustering (SVC). Experimental results indicate that the designed kernel can capture the features of relational schema and TPSVC is of fine clustering performance
Ping Ling, Yan Wang 0028, Chunguang Zhou
CW2
2005 A Fuzzy Neural Network System Based on Generalized Class Cover and Particle Swarm Optimization
Yanxin Huang, Yan Wang 0028, Zhezhou Yu, Chunguang Zhou
ICIC (2)2
2004 Training Minimal Uncertainty Neural Networks by Bayesian Theorem and Particle Swarm Optimization
Yan Wang 0028, Chunguang Zhou, Yanxin Huang, Xiaoyue Feng
ICONIP1
2004 A Rough-Set-Based Fuzzy-Neural-Network System for Taste Signal Identification
Yanxin Huang, Chunguang Zhou, Shu-Xue Zou, Yan Wang 0028, Yanchun Liang 0001
ISNN (2)4
2004 A Hybrid Algorithm for Combining Forecasting Based on AFTER-PSO
Xiaoyue Feng, Yanchun Liang 0001, Heow Pueh Lee, Chunguang Zhou, Yan Wang 0028
PRICAI6