VLDB 2026 Research / reviewers in the wild / expert
Chunyu Wang 0002
dblp:63/7235-2
· DBLP profile ↗
56ranked-venue papers
1as first author
39since 2021 · last 2026
0000-0002-2965-9920ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 39 · 1 first-author · 25 since 2021Artificial intelligence and machine learning · 11 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Image-Text Knowledge Modeling for Unsupervised Multi-Scenario Person Re-IdentificationabstractWe propose unsupervised multi-scenario (UMS) person re-identification (ReID) as a new task that expands ReID across diverse scenarios (cross-resolution, clothing change, etc.) within a single coherent framework. To tackle UMS-ReID, we introduce image-text knowledge modeling (ITKM) -- a three-stage framework that effectively exploits the representational power of vision-language models. We start with a pre-trained CLIP model with an image encoder and a text encoder. In Stage I, we introduce a scenario embedding in the image encoder and fine-tune the encoder to adaptively leverage knowledge from multiple scenarios. In Stage II, we optimize a set of learned text embeddings to associate with pseudo-labels from Stage I and introduce a multi-scenario separation loss to increase the divergence between inter-scenario text representations. In Stage III, we first introduce cluster-level and instance-level heterogeneous matching modules to obtain reliable heterogeneous positive pairs (e.g., a visible image and an infrared image of the same person) within each scenario. Next, we propose a dynamic text representation update strategy to maintain consistency between text and image supervision signals. Experimental results across multiple scenarios demonstrate the superiority and generalizability of ITKM; it not only outperforms existing scenario-specific methods but also enhances overall performance by integrating knowledge from multiple scenarios. Zhiqi Pang, Lingling Zhao, Yang Liu 0006, Chunyu Wang 0002, Gaurav Sharma 0001 |
AAAI | 4 |
| 2026 | Identification and characterization of lncRNA-stemness-immune regulatory patternsabstractLong noncoding RNAs (lncRNAs) play critical roles in regulating stemness signature genes (SSGs) and tumor immunity, thereby shaping the tumor microenvironment and antitumor immune responses. Increasing evidence suggests that cancer stem cell traits are closely associated with immune evasion and therapeutic resistance, underscoring the need to systematically characterize the pan-cancer interplay among SSGs, lncRNAs, and tumor immunity. Here, we developed an integrative analytical framework that combines network-based modeling with Bayesian network inference to identify core regulatory triplets (STEM-LncCRTs), each consisting of an lncRNA, an SSG, and an immune gene. We demonstrate that specific stemness-related lncRNAs can distinguish cancer subtypes, and that common stemness-related lncRNAs correlate significantly with immune cell infiltration. Notably, the ATAD5/PRR11-AS1/SKP2 triplet exhibits favorable prognostic potential across multiple cancers and consistently outperforms individual gene markers in predicting 1-, 3-, and 5-year overall survival. Furthermore, using four machine learning algorithms across three independent immunotherapy cohorts, we validate the predictive value of STEM-LncCRTs for immune checkpoint inhibitor response. Importantly, integrating STEM-LncCRTs with tumor mutation burden further improves predictive accuracy. Collectively, this study provides a systems-level view of stemness-related lncRNA regulation in tumor immunity and offers practical biomarkers for predicting immunotherapy efficacy. Zhipeng Qian, Chunlong Zhang, Guohua Wang 0001, Chunyu Wang 0002, Yang Li 0130 |
Briefings Bioinform. | 5 |
| 2026 | HMA-GCA: hybrid manifold augmentation and gated cross-attention for circRNA-miRNA interaction predictionabstractMOTIVATION: Circular RNAs (circRNAs) interact with microRNAs (miRNAs) to regulate gene expression and influence disease progression. However, traditional models tend to overlook the significant contributions of certain features when dealing with diverse sequence information, resulting in the inability to capture some deep topological structures and thus leaving room for improvement in prediction performance. RESULTS: We propose HMA-GCA, a novel framework that integrates hybrid manifold augmentation and gated cross-attention for CMI prediction. The model first constructs multi-scale descriptors by combining sequence-derived features (K-mer, CTD, Doc2Vec) and topological features (Role2Vec, node degree, neighborhood proximity). It then applies PCA for global linear projection and UMAP for local nonlinear manifold learning, enhancing feature representations while preserving intrinsic data geometry. A channel-wise gated cross-attention mechanism dynamically controls the injection of miRNA information into circRNA representations. Extensive experiments on three benchmark datasets show that HMA-GCA consistently outperforms state-of-the-art methods across multiple metrics. To ensure interpretability, we conducted SHAP analysis to quantify the contribution of each feature type, revealing that sequence-derived features and topological similarities are the most influential. Ablation studies confirm the necessity of each module, while case studies demonstrate that top-ranked predictions are supported by literature evidence. Overall, HMA-GCA not only achieves state-of-the-art predictive performance but also provides interpretable insights into the molecular features. AVAILABILITY AND IMPLEMENTATION: The source code and data are freely available at https://github.com/Lixunwind/Prediction-circ-mi-by-Gate.git. The implementation is based on Python and the required dependencies are listed in the repository. Yunzhou Hu, Yansu Wang, Yifeng Bai, Lei Xu 0047, Quan Zou 0001, Chunyu Wang 0002, Mengting Niu |
Bioinform. | 7 |
| 2026 | CFGSCDSA: Predicting circRNA-drug sensitivity associations based on collaborative feature learning and graph structure learningabstractMOTIVATION: The expression of circular RNAs (circRNAs) has been shown to be strongly correlated with drug sensitivity in human cells. However, experimental validation using wet-lab techniques is costly and inefficient, leaving a substantial portion of circRNA-drug sensitivity associations undiscovered. Therefore, improving the prediction efficiency of circRNA and sensitivity associations remains critical. METHODS: Here, we describe a method that integrates collaborative feature learning and graph structure learning to predict associations between circRNAs and drug sensitivity (CFGSCDSA). Specifically, collaborative learning integrated heterogeneous features from diverse data sources, thereby addressing the issue of data sparsity. Furthermore, graph structure learning with a confidence-guided pseudo-labeling strategy was employed to mitigate the detrimental effect of excessive negative samples. Results: Experimental evaluation revealed that CFGSCDSA attained superior performance compared to all competing models. Moreover, case studies provided further evidence of its capability to accurately predict both novel associations and new drug-related links. Quan Zou 0001, Chunyu Wang 0002, Mengting Niu |
PLoS Comput. Biol. | 3 |
| 2025 | Identity-Clothing Similarity Modeling for Unsupervised Clothing Change Person Re-IdentificationabstractClothing change person re-identification (CC-ReID) aims to match different images of the same person, even when the clothing varies across images. To reduce manual labeling costs, existing unsupervised CC-ReID methods employ clustering algorithms to generate pseudo-labels. However, they often fail to assign the same pseudo-label to two images with the same identity but different clothing—referred to as a clothing change positive pair—thus hindering clothing-invariant feature learning. To address this issue, we propose the identity-clothing similarity modeling (ICSM) framework. To effectively connect clothing change positive pairs, ICSM first performs clothing-aware learning to leverage all discriminative information, including clothing, to obtain compact clusters. It then extracts cluster-level identity and clothing features and performs inter-cluster similarity estimation to identify clothing change positive clusters, reliable negative clusters, and hard negative clusters for each compact cluster. During optimization, we design an adaptive version of existing optimization methods to enhance similarities of clothing change positive pairs, while also introducing text semantics as a supervisory signal to further promote clothing invariance. Extensive experimental results across multiple datasets validate the effectiveness of the proposed framework, demonstrating its superiority over existing unsupervised methods and its competitiveness with some supervised approaches. Zhiqi Pang, Junjie Wang 0005, Lingling Zhao, Chunyu Wang 0002 |
CVPR | 4 |
| 2025 | Augmented and Softened Matching for Unsupervised Visible-Infrared Person Re-Identification
Zhiqi Pang, Chunyu Wang 0002, Lingling Zhao, Junjie Wang 0005 |
ICCV | 2 |
| 2025 | LVLM-Driven Attribute-Aware Modeling for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) aims to match visible and infrared images of the same individual. Supervised VI-ReID (SVI-ReID) methods have achieved promising performance under the guidance of manually annotated identity labels. However, the substantial annotation cost severely limits their scalability in real-world applications. As a result, unsupervised VI-ReID (UVI-ReID) methods have attracted increasing attention. These methods typically rely on pseudo-labels generated by clustering and matching algorithms to replace manual annotations. Nevertheless, the quality of pseudo-labels is often difficult to guarantee, and low-quality pseudo-labels can significantly hinder model performance improvements. To address these challenges, we explore the use of attribute arrays extracted by a large vision-language model (LVLM) to enhance VI-ReID, and propose a novel LVLM-driven attribute-aware modeling (LVLM-AAM) approach. Specifically, we first design an attribute-aware reliable labeling strategy, which refines intra-modality clustering results based on image-level attributes and improves inter-modality matching by grouping clusters according to cluster-level attributes. Next, we develop an explicit-implicit attribute fusion module, which integrates explicit and implicit attributes to obtain more fine-grained identity-related text features. Finally, we introduce an attribute-aware contrastive learning module, which jointly leverages static and dynamic text features to promote modality-invariant feature learning. Extensive experiments conducted on VI-ReID datasets validate the effectiveness of the proposed LVLM-AAM and its individual components. LVLM-AAM not only significantly outperforms existing unsupervised methods but also surpasses several supervised methods. Zhiqi Pang, Lingling Zhao, Junjie Wang 0005, Chunyu Wang 0002 |
NeurIPS | 4 |
| 2025 | Cross-RNA transferable sequence representation learning for lncRNA m6A site detection via novel deep domain separation networksabstractN6-methyladenosine (m6A) is a key epitranscriptomic marker enriched in long noncoding RNAs (lncRNAs) that is closely involved in complex disease mechanisms. Although accurate detection of m6A sites in lncRNAs is essential for understanding disease mechanisms, the development of effective computational predictors remains challenging due to the limited number of annotated sites. Moreover, most existing predictors are specifically designed for messenger RNAs (mRNAs) based on abundant mRNA-specific knowledge, yet they exhibit limited generalizability to lncRNAs. Given the similarities between mRNAs and lncRNAs, a transferable framework capable of leveraging their shared features is critical for advancing m6A site prediction in lncRNAs. To address this challenge, we propose DSNm6A, a deep learning framework that learns cross-RNA transferable sequence representations for effective lncRNA m6A site detection. To comprehensively capture patterns and signals of m6A sites, lncRNA and mRNA sequences are first encoded from complementary multiple facets, including One-Hot encoding, nucleotide physicochemical properties and cumulative frequency, and position-specific propensity. Based on these sequence encodings, a domain separation network integrating CNN, Bi-LSTM, and BERT modules is then employed to explicitly disentangle domain-invariant features shared between mRNAs and lncRNAs from their domain-specific counterparts. The shared features are finally fed into a fully connected layer for accurate lncRNA m6A sites prediction. Cross-validation and independent test results demonstrate that DSNm6A consistently outperforms existing methods across nearly all performance metrics, attributed to its superior capacity to learn transferable m6A-related features across RNA types. In addition, DSNm6A exhibits strong robustness and generalization across species. Zhixia Teng, Chunyu Wang 0002, Guohua Wang 0001 |
Briefings Bioinform. | 4 |
| 2025 | Long Noncoding RNA function prediction via multiview cross-contrastive learning combined with multiscale semantic adaptive optimizationabstractUnderstanding long noncoding RNA (lncRNA) function is essential for revealing molecular mechanisms and developing effective therapies for complex diseases, as lncRNAs play important regulatory roles in many disease-related biological processes. However, existing lncRNA function predictors struggle to extract discriminative features from multimodal omics data and to model the semantic and topological structure of the gene ontology (GO), which severely limits their ability to achieve biologically meaningful and functionally informative predictions. To address these challenges, we propose a novel framework for lncRNA function prediction, namely MiCLSAO. Firstly, MiCLSAO utilizes multiview cross-contrastive learning with attention mechanisms to extract highly discriminative lncRNA features from diverse omics similarity networks. Secondly, graph convolutional networks are applied to learn initial features of GO terms, while multiscale topological and semantic relationships are incorporated to adaptively refine term representations. Finally, an lncRNA function predictor is developed by dynamically integrating the representations of lncRNAs and GO terms using a Kolmogorov-Arnold network. Extensive experiments demonstrate that MiCLSAO consistently outperforms state-of-the-art methods across multiple metrics, with significant capability to recover known functions and uncover novel ones. Moreover, MiCLSAO demonstrates remarkable practical utility and potential value by providing more functionally informative annotations for lncRNAs. Zhixia Teng, Qingqi Li, Chunyu Wang 0002, Guohua Wang 0001 |
Briefings Bioinform. | 4 |
| 2025 | Predicting circRNA-disease associations with shared units and multi-channel attention mechanismsabstractMOTIVATION: Circular RNAs (circRNAs) have been identified as key players in the progression of several diseases; however, their roles have not yet been determined because of the high financial burden of biological studies. This highlights the urgent need to develop efficient computational models that can predict circRNA-disease associations, offering an alternative approach to overcome the limitations of expensive experimental studies. Although multi-view learning methods have been widely adopted, most approaches fail to fully exploit the latent information across views, while simultaneously overlooking the fact that different views contribute to varying degrees of significance. RESULTS: This study presents a method that combines multi-view shared units and multichannel attention mechanisms to predict circRNA-disease associations (MSMCDA). MSMCDA first constructs similarity and meta-path networks for circRNAs and diseases by introducing shared units to facilitate interactive learning across distinct network features. Subsequently, multichannel attention mechanisms were used to optimize the weights within similarity networks. Finally, contrastive learning strengthened the similarity features. Experiments on five public datasets demonstrated that MSMCDA significantly outperformed other baseline methods. Additionally, case studies on colorectal cancer, gastric cancer, and nonsmall cell lung cancer confirmed the effectiveness of MSMCDA in uncovering new associations. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/zhangxue2115/MSMCDA.git. Quan Zou 0001, Mengting Niu, Chunyu Wang 0002 |
Bioinform. | 4 |
| 2025 | Image-text semantic learning for unsupervised cross-resolution person re-identification
Fuqi Liu, Zhiqi Pang, Chunyu Wang 0002 |
Expert Syst. Appl. | 3 |
| 2025 | Computational approaches for circRNA-disease association prediction: a reviewabstractAbstract Circular RNA (circRNA) is a covalently closed RNA molecule formed by back splicing. The role of circRNAs in posttranscriptional gene regulation provides new insights into several types of cancer and neurological diseases. CircRNAs are associated with multiple diseases and are emerging biomarkers in cancer diagnosis and treatment. The associations prediction is one of the current research hotspots in the field of bioinformatics. Although research on circRNAs has made great progress, the traditional biological method of verifying circRNA-disease associations is still a great challenge because it is a difficult task and requires much time. Fortunately, advances in computational methods have made considerable progress in circRNA research. This review comprehensively discussed the functions and databases related to circRNA, and then focused on summarizing the calculation model of related predictions, detailed the mainstream algorithm into 4 categories, and analyzed the advantages and limitations of the 4 categories. This not only helps researchers to have overall understanding of circRNA, but also helps researchers have a detailed understanding of the past algorithms, guide new research directions and research purposes to solve the shortcomings of previous research. Mengting Niu, Yaojia Chen, Chunyu Wang 0002, Quan Zou 0001, Lei Xu 0047 |
Frontiers Comput. Sci. | 3 |
| 2025 | MVHGCN: Predicting circRNA-disease associations with multi-view heterogeneous graph convolutional neural networksabstractCircular RNA, a class of RNA molecules gaining widespread attentions, has been widely recognized as a potential biomarker for many diseases. In recent years, significant progress has been made in the study of the associations between circRNA and diseases. However, traditional experimental methods are often inefficient and costly, making computational models an effective alternative. Nevertheless, existing computational methods still face challenges such as data sparsity and the difficulty of confirming negative samples, which limits the accuracy of predictions. To address these challenges, a novel computational method, namely MVHGCN, is proposed based on multi-view and graph convolutional networks to predict potential associations between circRNA and diseases. MVHGCN first constructs a heterogeneous graph and generates feature descriptors by integrating multiple databases. Then it extracts different connection views of circRNA and diseases through meta-paths, maximizing the utilization of known association information, and aggregates deep feature information through graph convolutional networks. Finally, a MLP is used to predict the association scores. The experimental results show that MVHGCN significantly outperforms existing methods on benchmark datasets by 5-fold cross-validation. This research provides an effective new approach to studying the associations between circRNAs and diseases, capable of alleviating the problem of data sparsity and accurately identifying potential associations. Yan Miao, Chunyu Wang 0002, Zhenyuan Sun, Guohua Wang 0001 |
PLoS Comput. Biol. | 3 |
| 2025 | Joint Augmentation and Part Learning for Unsupervised Clothing Change Person Re-IdentificationabstractClothing change person re-identification (CC-ReID) is a crucial task in intelligent surveillance, aiming to match images of the same person wearing different clothing. Promising performance in existing CC-ReID methods is achieved at the cost of labor-intensive manual annotation of identity labels. While some researchers have explored unsupervised CC-ReID, these methods still depend on additional deep learning models for preprocessing. To eliminate the need for additional models and improve performance, we propose a joint augmentation and part learning (JAPL) framework that obtains clothing change positive pairs in an unsupervised fashion by synergistically combining augmentation-based invariant learning (AugIL) and part-based invariant learning (ParIL). AugIL first constructs clothing change pseudo-positive pairs and then encourages the model to focus on clothing-invariant information by enhancing feature consistency between the pseudo-positive pairs. ParIL beneficially encourages high similarity between inter-cluster clothing change positive pair using part images and a prediction sharpening loss. PartIL also introduces a soft consistency loss that promotes clothing-invariant feature learning by encouraging consistency of class vectors between the real features actually used for CC-ReID and the part features. Experimental results on multiple ReID datasets demonstrate that the proposed JAPL not only surpasses existing unsupervised methods but also achieves competitive performance compared to some supervised CC-ReID methods. Zhiqi Pang, Lingling Zhao, Yang Liu 0006, Gaurav Sharma 0001, Chunyu Wang 0002 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Robust Labeling and Invariance Modeling for Unsupervised Cross-Resolution Person Re-IdentificationabstractCross-resolution person re-identification (CR-ReID) aims to match low-resolution (LR) and high-resolution (HR) images of the same individual. To reduce the cost of manual annotation, existing unsupervised CR-ReID methods typically rely on cross-resolution fusion to obtain pseudo-labels and resolution-invariant features. However, the fusion process requires two encoders and a fusion module, which significantly increases computational complexity and reduces efficiency. To address this issue, we propose a robust labeling and invariance modeling (RLIM) framework, which utilizes a single encoder to tackle the unsupervised CR-ReID problem. To obtain pseudo-labels robust to resolution gaps, we develop cross-resolution robust labeling (CRL), which utilizes two clustering criteria to encourage cross-resolution positive pairs to cluster together and exploit the reliable relationships between images. We also introduce random texture augmentation (TexA) to enhance the model's robustness to noisy textures related to artifacts and backgrounds by randomly adjusting texture strength. During the optimization process, we introduce the resolution-cluster consistency loss, which promotes resolution-invariant feature learning by aligning inter-resolution distances with intra-cluster distances. Experimental results on multiple datasets demonstrate that RLIM not only surpasses existing unsupervised methods, but also achieves performance close to some supervised CR-ReID methods. Code is available at https://github.com/zqpang/RLIM. Zhiqi Pang, Lingling Zhao, Yang Liu 0006, Chunyu Wang 0002, Gaurav Sharma 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Predicting circRNA-Drug Resistance Associations Based on a Multimodal Graph Representation Learning FrameworkabstractCircular RNA (circRNA) is a class of noncoding RNA that is highly conserved and exhibit exceptional stability. Due to its function as a microRNA sponge, circRNA has gained significant attention as an essential biomarker and potential drug target in the pathogenesis of several cancers. Although many circRNAs have been identified to play a role in cancer resistance, traditional methods are time-consuming and expensive. In this context, computational methods offer a promising way to facilitate the discovery process. However, most existing prediction models focus on the association between circRNAs and drug resistance, without considering the corresponding disease-related information in the circRNA-drug resistance association. Incorporating disease-related information into the prediction of circRNA-drug resistance associations could potentially improve the efficiency and speed of discovering and developing circRNA-targeting drugs. We propose a computational framework, named GraphCDD, for predicting the association between circRNA and drug resistance. Our model utilizes data from three sources, namely circRNA, disease, and drug, to construct three similarity networks that represent the features of circRNA, disease, and drug, respectively. We utilize a multimodal graph neural network to acquire efficient representations of circRNAs, diseases, and drugs by integrating various types of information, and establish a predictive model. The experimental results have validated the effectiveness of our model and provided a promising method in predicting potential associations between circRNA and drug resistance. Qiguo Dai, Xianhai Yu, Xiaodong Duan, Chunyu Wang 0002 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Multimodal Contrastive Learning for Protein-Protein Interaction Inhibitor PredictionabstractProtein-protein interactions (PPIs) are crucial for various cellular activities and disease development, and modulating PPIs using small molecule inhibitors (PPIIs) has gradually become a promising therapeutic strategy. Recently, researchers have proposed several machine learning methods to screen PPIIs, but most of the works focused on unimodal representations of molecules or combining multimodal features in a naive splicing manner. Meanwhile, current research progress is being slowed by the lack of large-scale PPII datasets. To address these issues, we propose MCLPPII, a unified multimodal contrastive learning framework for PPII prediction. MCLPPII extracts comprehensive molecular information from four modalities and effectively combines them through an adaptive feature fusion method. Furthermore, we propose a three-stage training strategy to enhance the PPII prediction capability of MCLPPII by leveraging self-supervised pre-training on a large unlabeled dataset. We evaluate MCLPPII on nine PPI targets and two downstream tasks, including the PPI inhibitor identification task and potency prediction task. Experimental results show that MCLPPII achieves competitive performance. The source code and datasets are freely available at https://github.com/1zzt/MCLPPII. Zitong Zhang 0002, Zhixian Wang, Lingling Zhao, Junjie Wang 0005, Chunyu Wang 0002 |
BIBM | 5 |
| 2024 | Dual-Resolution Fusion Modeling for Unsupervised Cross-Resolution Person Re-IdentificationabstractCross-resolution person re-identification (CR-ReID) aims to match images of the same person with different resolutions in different scenarios. Existing CR-ReID methods achieve promising performance by relying on large-scale manually annotated identity labels. However, acquiring manual labels requires considerable human effort, greatly limiting the flexibility of existing CR-ReID methods. To address this issue, we propose a dual-resolution fusion modeling (DRFM) framework to tackle the CR-ReID problem in an unsupervised manner. Firstly, we design a cross-resolution pseudo-label generation (CPG) method, which initially clusters high-resolution images and then obtains reliable identity pseudo-labels by fusing class vectors in both resolution spaces. Subsequently, we develop a cross-resolution feature fusion (CRFF) module to fuse features from both high-resolution and low-resolution spaces. The fusion features have the potential to serve as a new form of resolution-invariant features. Finally, we introduce cross-resolution contrastive loss and probability sharpening loss in DRFM to facilitate resolution-invariant learning and effectively utilize ambiguous samples for optimization. Experimental results on multiple CR-ReID datasets demonstrate that the proposed DRFM not only outperforms existing unsupervised methods but also approaches the performance of early supervised methods. Zhiqi Pang, Lingling Zhao, Chunyu Wang 0002 |
ACM Multimedia | 3 |
| 2024 | Identification, characterization and expression analysis of circRNA encoded by SARS-CoV-1 and SARS-CoV-2abstractVirus-encoded circular RNA (circRNA) participates in the immune response to viral infection, affects the human immune system, and can be used as a target for precision therapy and tumor biomarker. The coronaviruses SARS-CoV-1 and SARS-CoV-2 (SARS-CoV-1/2) that have emerged in recent years are highly contagious and have high mortality rates. In coronaviruses, little is known about the circRNA encoded by the SARS-CoV-1/2. Therefore, this study explores whether SARS-CoV-1/2 encodes circRNA and characteristics and functions of circRNA. Based on RNA-seq data of SARS-CoV-1 and SARS-CoV-2 infections, we used circRNA identification tools (circRNA_finder, find_circ and CIRI2) to identify circRNAs. The number of circRNAs encoded by SARS-CoV-1 and SARS-CoV-2 was identified as 151 and 470, respectively. It can be found that SARS-CoV-2 shows more prominent circRNA encoding ability than SARS-CoV-1. Expression analysis showed that only a few circRNAs encoded by SARS-CoV-1/2 showed high expression levels, and the positive strand produced more abundant circRNAs. Then, based on the identified SARS-CoV-1/2-encoded circRNAs, we performed circRNA identification and characterization using the previously developed CirRNAPL. Finally, target gene prediction and functional enrichment analysis were performed. It was found that viral circRNA is closely related to cancer and has a potential role in regulating host cell functions. This study studied the characteristics and functions of viral circRNA encoded by coronavirus SARS-CoV-1/2, providing a valuable resource for further research on the function and molecular mechanism of coronavirus circRNA. Mengting Niu, Chunyu Wang 0002, Yaojia Chen, Quan Zou 0001, Lei Xu 0047 |
Briefings Bioinform. | 2 |
| 2024 | CircSI-SSL: circRNA-binding site identification based on self-supervised learningabstractMOTIVATION: In recent years, circular RNAs (circRNAs), the particular form of RNA with a closed-loop structure, have attracted widespread attention due to their physiological significance (they can directly bind proteins), leading to the development of numerous protein site identification algorithms. Unfortunately, these studies are supervised and require the vast majority of labeled samples in training to produce superior performance. But the acquisition of sample labels requires a large number of biological experiments and is difficult to obtain. RESULTS: To resolve this matter that a great deal of tags need to be trained in the circRNA-binding site prediction task, a self-supervised learning binding site identification algorithm named CircSI-SSL is proposed in this article. According to the survey, this is unprecedented in the research field. Specifically, CircSI-SSL initially combines multiple feature coding schemes and employs RNA_Transformer for cross-view sequence prediction (self-supervised task) to learn mutual information from the multi-view data, and then fine-tuning with only a few sample labels. Comprehensive experiments on six widely used circRNA datasets indicate that our CircSI-SSL algorithm achieves excellent performance in comparison to previous algorithms, even in the extreme case where the ratio of training data to test data is 1:9. In addition, the transplantation experiment of six linRNA datasets without network modification and hyperparameter adjustment shows that CircSI-SSL has good scalability. In summary, the prediction algorithm based on self-supervised learning proposed in this article is expected to replace previous supervised algorithms and has more extensive application value. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/cc646201081/CircSI-SSL. Chunyu Wang 0002, Shuhong Yang, Quan Zou 0001 |
Bioinform. | 2 |
| 2024 | A MLP-Mixer and mixture of expert model for remaining useful life prediction of lithium-ion batteriesabstractAbstract Accurately predicting the Remaining Useful Life (RUL) of lithium-ion batteries is crucial for battery management systems. Deep learning-based methods have been shown to be effective in predicting RUL by leveraging battery capacity time series data. However, the representation learning of features such as long-distance sequence dependencies and mutations in capacity time series still needs to be improved. To address this challenge, this paper proposes a novel deep learning model, the MLP-Mixer and Mixture of Expert (MMMe) model, for RUL prediction. The MMMe model leverages the Gated Recurrent Unit and Multi-Head Attention mechanism to encode the sequential data of battery capacity to capture the temporal features and a re-zero MLP-Mixer model to capture the high-level features. Additionally, we devise an ensemble predictor based on a Mixture-of-Experts (MoE) architecture to generate reliable RUL predictions. The experimental results on public datasets demonstrate that our proposed model significantly outperforms other existing methods, providing more reliable and precise RUL predictions while also accurately tracking the capacity degradation process. Our code and dataset are available at the website of github. Lingling Zhao, Shitao Song, Pengyan Wang, Chunyu Wang 0002, Junjie Wang 0005, Maozu Guo 0001 |
Frontiers Comput. Sci. | 4 |
| 2024 | MIMR: Modality-Invariance Modeling and Refinement for unsupervised visible-infrared person re-identification
Zhiqi Pang, Chunyu Wang 0002, Honghu Pan, Lingling Zhao, Junjie Wang 0005, Maozu Guo 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Clothing-invariant contrastive learning for unsupervised person re-identification
Zhiqi Pang, Lingling Zhao, Chunyu Wang 0002 |
Neural Networks | 3 |
| 2024 | AutoEdge-CCP: A novel approach for predicting cancer-associated circRNAs and drugs based on automated edge embeddingabstractThe unique expression patterns of circRNAs linked to the advancement and prognosis of cancer underscore their considerable potential as valuable biomarkers. Repurposing existing drugs for new indications can significantly reduce the cost of cancer treatment. Computational prediction of circRNA-cancer and drug-cancer relationships is crucial for precise cancer therapy. However, prior computational methods fail to analyze the interaction between circRNAs, drugs, and cancer at the systematic level. It is essential to propose a method that uncover more valuable information for achieving cancer-centered multi-association prediction. In this paper, we present a novel computational method, AutoEdge-CCP, to unveil cancer-associated circRNAs and drugs. We abstract the complex relationships between circRNAs, drugs, and cancer into a multi-source heterogeneous network. In this network, each molecule is represented by two types information, one is the intrinsic attribute information of molecular features, and the other is the link information explicitly modeled by autoGNN, which searches information from both intra-layer and inter-layer of message passing neural network. The significant performance on multi-scenario applications and case studies establishes AutoEdge-CCP as a potent and promising association prediction tool. Yaojia Chen, Jiacheng Wang 0009, Chunyu Wang 0002, Quan Zou 0001 |
PLoS Comput. Biol. | 3 |
| 2024 | Drug-Target Binding Affinity Prediction in a Continuous Latent Space Using Variational AutoencodersabstractAccurate prediction of Drug-Target binding Affinity (DTA) is a daunting yet pivotal task in the sphere of drug discovery. Over the years, a plethora of deep learning-based DTA models have emerged, rendering promising results in predicting the binding affinities between drugs and their target proteins. However, in contrast to the conventional approach of modeling binding affinity in vector spaces, we propose a more nuanced modeling process in a continuous space to account for the diversity of input samples. Initially, the drug is encoded using the Simplified Molecular Input Line Entry System (SMILES), while the target sequences are characterized via a pretrained language model. Subsequently, highly correlative information is extracted utilizing residual gated convolutional neural networks. In a departure from existing deep learning-based models, our model learns the hidden representations of the drugs and targets jointly. Instead of employing two vectors, our hidden representations consist of two Gaussian distributions. To validate the effectiveness of our proposal, we conducted evaluations on commonly utilized benchmark datasets. The experimental outcomes corroborated that our method surpasses the state-of-the-art vectorial representation methods in terms of performance. This approach, therefore, offers potential enhancements in the precision of DTA predictions, potentially contributing to more efficient drug discovery processes. Lingling Zhao, Yan Zhu 0006, Naifeng Wen, Chunyu Wang 0002, Junjie Wang 0005, Yongfeng Yuan |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | Cross-Modality Hierarchical Clustering and Refinement for Unsupervised Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a challenging cross-modality image retrieval task. Compared to visible modality person re-identification that handles only the intra-modality discrepancy, VI-ReID suffers from an additional modality gap. Most existing VI-ReID methods achieve promising accuracy in a supervised setting, but the high annotation cost limits their scalability to real-world scenarios. Although a few unsupervised VI-ReID methods already exist, they typically rely on intra-modality initialization and cross-modality instance selection, despite the additional computational time required for intra-modality initialization. In this paper, we study the fully unsupervised VI-ReID problem and propose a novel cross-modality hierarchical clustering and refinement (CHCR) method by promoting modality-invariant feature learning and improving the reliability of pseudo-labels. Unlike conventional VI-ReID methods, CHCR does not rely on any manual identity annotation and intra-modality initialization. First, we design a simple and effective cross-modality clustering baseline that clusters between modalities. Then, to provide sufficient inter-modality positive sample pairs for modality-invariant feature learning, we propose a cross-modality hierarchical clustering algorithm to promote the clustering of inter-modality positive samples into the same cluster. In addition, we develop an inter-channel pseudo-label refinement algorithm to eliminate unreliable pseudo-labels by checking the clustering results of three channels in the visible modality. Extensive experiments demonstrate that CHCR outperforms state-of-the-art unsupervised methods and achieves performance competitive with many supervised methods. Zhiqi Pang, Chunyu Wang 0002, Lingling Zhao, Yang Liu 0006, Gaurav Sharma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Inter-Modality Similarity Learning for Unsupervised Multi-Modality Person Re-IdentificationabstractRGB (visible), near-infrared (NI), and thermal infrared (TI) imaging modalities are commonly combined for round-the-clock surveillance. We introduce a novel unsupervised multi-modality person re-identification (MM-ReID) task, which, based on an individual’s image in any one modality, seeks to identify matches in the other two modalities. Compared to prior MM-ReID problem formulations, unsupervised MM-ReID significantly reduces labeling cost and imaging constraints. To address the unsupervised MM-ReID task, we propose a novel inter-modality similarity learning (IMSL) framework consisting of four synergistic interconnected modules: modality mean clustering (MMC), multi-modality reliability estimation (MMRE), shape-based mutual reinforcement (SMR), and modality-aware invariant learning (MIL). MMC iterates with SMR and MIL in a mutually beneficial manner to provide pseudo-labels that are robust to modality gap. MMRE normalizes sample weights, mitigating the impact of noisy labels in the multi-modality setting. SMR emphasizes shape information to implicitly enhance the model’s robustness to the modality gap and is additionally guided by pseudo-labels provided by MMC to attend to identity-related details. MIL explicitly encourages learning of modality-invariant and identity-related features via contrastive feedback for the MMC module. Extensive experimental results on the multi-modality and cross-modality datasets demonstrate that IMSL provides substantial performance gains over existing methods. Code is made available at https://github.com/zqpang/IMSL. Zhiqi Pang, Lingling Zhao, Yang Liu 0006, Gaurav Sharma 0001, Chunyu Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | MAMLCDA: A Meta-Learning Model for Predicting circRNA-Disease Association Based on MAML Combined With CNNabstractCircular RNAs (circRNAs) exist in vivo and are a class of noncoding RNA molecules. They have a single-stranded, closed, annular structure. Many studies have shown that circRNAs and diseases are linked. Therefore, it is critical to build a reliable and accurate predictor to find the circRNA-disease association. In this paper, we presented a meta-learning model named MAMLCDA to identify the circRNA-disease association, which is based on model-agnostic meta-learning (MAML) combined with CNN classification. Specifically, similarities between diseases and circRNAs are extracted and integrated to characterize their relationships, and k-means is used to cluster majority samples and select a certain number of samples from each cluster to obtain the same number of negative samples as the positive samples. To further reduce the dimension of the features and save operation time, we applied probabilistic principal component analysis (PPCA) to compact the integrated circRNA and disease similarity network feature vectors. The feature vectors are converted into images. At this time, the prediction problem is transformed into the 2-way 1-shot problem of the image and input into the model with MAML as the meta-learner and CNN as the base-learner. Comparison results of five-fold cross-validation on two benchmark datasets illustrate that MAMLCDA outperforms several state-of-the-art approaches with the best accuracies of 95.33% and 98%. Therefore, MAMLCDA can help to understand the pathogenesis of complex diseases at the circRNA level. Yuanyi Tian, Quan Zou 0001, Chunyu Wang 0002, Cangzhi Jia |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | A Hierarchical Graph Neural Network Framework for Predicting Protein-Protein Interaction Modulators With Functional Group Information and Hypergraph StructureabstractAccurate prediction of small molecule modulators targeting protein-protein interactions (PPIMs) remains a significant challenge in drug discovery. Existing machine learning-based models rely on manual feature engineering, which is tedious and task-specific. Recently, deep learning models based on graph neural networks have made remarkable progress in molecular representation learning. However, many graph-based approaches ignore molecular hierarchical structure modeling guided by domain knowledge. In chemistry, the functional groups of a molecule determine its interaction with specific targets. Therefore, we propose a hierarchical graph neural network framework (called HiGPPIM) for predicting PPIMs by integrating atom-level and functional group-level features of molecules. HiGPPIM constructs atom-level and functional group-level graphs based on chemical knowledge and learns graph representations using graph attention networks. Furthermore, a hypergraph attention network is designed in HiGPPIM to aggregate and transform two-level graph information. We evaluate the performance of HiGPPIM on eight PPI families and two prediction tasks, namely PPIM identification and potency prediction. Experimental results demonstrate that HiGPPIM achieves state-of-the-art performance on both tasks and that using functional group information to guide PPIM prediction is effective. Zitong Zhang 0002, Lingling Zhao, Junjie Wang 0005, Chunyu Wang 0002 |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Continuous Prompt for Chemical Language Model Aided Anticancer Synergistic Drug Combination PredictionabstractIdentifying synergistic drug combinations is paramount significance in addressing complex diseases while reducing the risks of toxicities and other adverse effects. Although a plethora of computational methods have been proposed in this domain, most of them are underpinned only by physicochemical or biological features. Recently, Chemical Language Models (CLMs) are shown to be capable of learning better representations that hold utility across diverse tasks, from molecular property prediction, de novo drug design, drug-target interaction, and more. In this study, we proposed CLMSyn, a continuous prompt for CLM aided synergistic drug combinations prediction. Unlike existing works employ CLMs for downstream tasks, we adopt the prompt learning to fine-tune CLM, that is, only train small-scale prompt while keeping CLM fixed. Furthermore, we harness the the multi-head attention mechanism to fuse the learned vector from the CLM, chemical descriptors and gene expression of cell line. A comprehensive array of experiments conducted on a benchmark dataset, encompassing four distinct synergy types, substantiates the superior performance of CLMSyn when contrasted against existing state-of-the-art methods. These empirical findings provide compelling evidence attesting to the efficacy of CLMSyn as a potent instrumentality in expediting the identification of pioneering combination therapies. Guannan Geng, Lingling Zhao, Chunyu Wang 0002, Junjie Wang 0005 |
IEEE Big Data | 3 |
| 2023 | DataDTA: a multi-feature and dual-interaction aggregation framework for drug-target binding affinity predictionabstractMOTIVATION: Accurate prediction of drug-target binding affinity (DTA) is crucial for drug discovery. The increase in the publication of large-scale DTA datasets enables the development of various computational methods for DTA prediction. Numerous deep learning-based methods have been proposed to predict affinities, some of which only utilize original sequence information or complex structures, but the effective combination of various information and protein-binding pockets have not been fully mined. Therefore, a new method that integrates available key information is urgently needed to predict DTA and accelerate the drug discovery process. RESULTS: In this study, we propose a novel deep learning-based predictor termed DataDTA to estimate the affinities of drug-target pairs. DataDTA utilizes descriptors of predicted pockets and sequences of proteins, as well as low-dimensional molecular features and SMILES strings of compounds as inputs. Specifically, the pockets were predicted from the three-dimensional structure of proteins and their descriptors were extracted as the partial input features for DTA prediction. The molecular representation of compounds based on algebraic graph features was collected to supplement the input information of targets. Furthermore, to ensure effective learning of multiscale interaction features, a dual-interaction aggregation neural network strategy was developed. DataDTA was compared with state-of-the-art methods on different datasets, and the results showed that DataDTA is a reliable prediction tool for affinities estimation. Specifically, the concordance index (CI) of DataDTA is 0.806 and the Pearson correlation coefficient (R) value is 0.814 on the test dataset, which is higher than other methods. AVAILABILITY AND IMPLEMENTATION: The codes and datasets of DataDTA are available at https://github.com/YanZhu06/DataDTA. Yan Zhu 0006, Lingling Zhao, Naifeng Wen, Junjie Wang 0005, Chunyu Wang 0002 |
Bioinform. | 5 |
| 2023 | Reliability modeling and contrastive learning for unsupervised person re-identification
Zhiqi Pang, Chunyu Wang 0002, Junjie Wang 0005, Lingling Zhao |
Knowl. Based Syst. | 2 |
| 2023 | Camera Invariant Feature Learning for Unsupervised Person Re-IdentificationabstractFully unsupervised person re-identification (ReID) methods aim to learn discriminative features without using labeled ReID data. Because these methods are easily affected by camera discrepancies, similar studies have typically designed optimization methods to enable the model to learn camera-invariant features. However, they often ignore the impact of camera discrepancies on clustering results. Specifically, camera discrepancies will reduce the intra-class camera diversity and promote the generation of noise labels. To solve the above problems, we propose a unified unsupervised learning framework: camera invariant feature learning (CIFL) framework. First, we designed a novel DBSCAN-NN algorithm in the CIFL framework that improves the intra-class camera diversity by forcibly merging samples from different cameras. Then, we designed feature ensemble clustering that improves the accuracy of the pseudo-labels by clustering feature ensembles. In addition, we designed an optimization method for camera discrepancies: stochastic pulled loss. With the stochastic pulled loss, the ReID model is forced to learn camera-invariant features. We verified the effectiveness and generalization of CIFL on four ReID datasets (Market-1501, DukeMTMC-reID, MSMT17 and CUHK03-NP). The experimental results show that CIFL not only outperforms the existing fully unsupervised methods but also is superior to the unsupervised domain adaptation methods. Zhiqi Pang, Lingling Zhao, Qiuyang Liu, Chunyu Wang 0002 |
IEEE Trans. Multim. | 4 |
| 2022 | Learning representations for gene ontology terms by jointly encoding graph structure and textual node descriptorsabstractMeasuring the semantic similarity between Gene Ontology (GO) terms is a fundamental step in numerous functional bioinformatics applications. To fully exploit the metadata of GO terms, word embedding-based methods have been proposed recently to map GO terms to low-dimensional feature vectors. However, these representation methods commonly overlook the key information hidden in the whole GO structure and the relationship between GO terms. In this paper, we propose a novel representation model for GO terms, named GT2Vec, which jointly considers the GO graph structure obtained by graph contrastive learning and the semantic description of GO terms based on BERT encoders. Our method is evaluated on a protein similarity task on a collection of benchmark datasets. The experimental results demonstrate the effectiveness of using a joint encoding graph structure and textual node descriptors to learn vector representations for GO terms. Lingling Zhao, Huiting Sun, Xinyi Cao, Naifeng Wen, Junjie Wang 0005, Chunyu Wang 0002 |
Briefings Bioinform. | 6 |
| 2022 | GMNN2CD: identification of circRNA-disease associations based on variational inference and graph Markov neural networksabstractMOTIVATION: With the analysis of the characteristic and function of circular RNAs (circRNAs), people have realized that they play a critical role in the diseases. Exploring the relationship between circRNAs and diseases is of far-reaching significance for searching the etiopathogenesis and treatment of diseases. Nevertheless, it is inefficient to learn new associations only through biotechnology. RESULTS: Consequently, we present a computational method, GMNN2CD, which employs a graph Markov neural network (GMNN) algorithm to predict unknown circRNA-disease associations. First, used verified associations, we calculate semantic similarity and Gaussian interactive profile kernel similarity (GIPs) of the disease and the GIPs of circRNA and then merge them to form a unified descriptor. After that, GMNN2CD uses a fusion feature variational map autoencoder to learn deep features and uses a label propagation map autoencoder to propagate tags based on known associations. Based on variational inference, GMNN alternate training enhances the ability of GMNN2CD to obtain high-efficiency high-dimensional features from low-dimensional representations. Finally, 5-fold cross-validation of five benchmark datasets shows that GMNN2CD is superior to the state-of-the-art methods. Furthermore, case studies have shown that GMNN2CD can detect potential associations. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/nmt315320/GMNN2CD.git. Mengting Niu, Quan Zou 0001, Chunyu Wang 0002 |
Bioinform. | 3 |
| 2021 | SeqGO-CPA: Improving Compound-Protein Binding Affinity Prediction with Sequence Information and Gene Ontology KnowledgeabstractThe compound-protein binding affinity (CPA) pre-diction is vital for drug discovery and drug repurposing. Deep learning methods have been developed to model the complicated relationship between CPA and the sequences or structures of proteins and molecules. This study proposes a novel deep learning method, SeqGO-CPA, integrating protein function knowledge represented by Gene Ontology (GO) annotations in the CPA prediction. To capture the semantic information of GO annotations, a fine-tuned natural language processing model for biomedical domains is utilized to encode the set of GO terms. Meanwhile, based on the observation that CPA often occurs in sub-structures, our method uses the tokenization algorithm to learn sub-structure information of proteins and compounds from a large number of unlabeled sequences. Further, a deep neural network architecture involving the jointly-feature representation and a highway block is developed to enhance the CPA prediction ability. The proposed model was evaluated on two public benchmark datasets in both standard cross-validation and blinding split settings. The experimental results demonstrate our method outperforms the deep learning-based baselines, meanwhile the incorporating of GO information further improves the prediction performance. Chunyu Wang 0002, Yan Zhu 0006, Naifeng Wen, Lingling Zhao, Junjie Wang 0005 |
BIBM | 1 |
| 2021 | Pathogenic gene prediction based on network embeddingabstractIn disease research, the study of gene-disease correlation has always been an important topic. With the emergence of large-scale connected data sets in biology, we use known correlations between the entities, which may be from different sets, to build a biological heterogeneous network and propose a new network embedded representation algorithm to calculate the correlation between disease and genes, using the correlation score to predict pathogenic genes. Then, we conduct several experiments to compare our method to other state-of-the-art methods. The results reveal that our method achieves better performance than the traditional methods. Yang Liu 0006, Chunyu Wang 0002, Maozu Guo 0001 |
Briefings Bioinform. | 4 |
| 2021 | Critical downstream analysis steps for single-cell RNA sequencing dataabstractSingle-cell RNA sequencing (scRNA-seq) has enabled us to study biological questions at the single-cell level. Currently, many analysis tools are available to better utilize these relatively noisy data. In this review, we summarize the most widely used methods for critical downstream analysis steps (i.e. clustering, trajectory inference, cell-type annotation and integrating datasets). The advantages and limitations are comprehensively discussed, and we provide suggestions for choosing proper methods in different situations. We hope this paper will be useful for scRNA-seq data analysts and bioinformatics tool developers. Feifei Cui, Chen Lin 0001, Lingling Zhao, Chunyu Wang 0002, Quan Zou 0001 |
Briefings Bioinform. | 5 |
| 2021 | Goals and approaches for each processing step for single-cell RNA sequencing dataabstractSingle-cell RNA sequencing (scRNA-seq) has enabled researchers to study gene expression at the cellular level. However, due to the extremely low levels of transcripts in a single cell and technical losses during reverse transcription, gene expression at a single-cell resolution is usually noisy and highly dimensional; thus, statistical analyses of single-cell data are a challenge. Although many scRNA-seq data analysis tools are currently available, a gold standard pipeline is not available for all datasets. Therefore, a general understanding of bioinformatics and associated computational issues would facilitate the selection of appropriate tools for a given set of data. In this review, we provide an overview of the goals and most popular computational analysis tools for the quality control, normalization, imputation, feature selection and dimension reduction of scRNA-seq data. Feifei Cui, Chunyu Wang 0002, Lingling Zhao, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2020 | OntoSem: an Ontology Semantic Representation Methodology for Biomedical DomainabstractOntologies are essential description tools for biomedical concepts and entities, supporting biomedical fundamental research such as semantic similarity analysis, protein-protein interaction prediction and so on. An increasing amount of ontology-like domain knowledge is published in scientific publications, meanwhile, advanced natural language processing (NLP) techniques have been widespread to extract information from text resources automatically, both of which facilitate the exploration of the semantic representation of biomedical ontologies. We propose a novel distributional semantic representation methodology based on the combination of two pre-trained and domain-specific word embedding tools, the non-contextualized Word2Vec and the context-dependent NCBI-blueBERT, to enhance the encoding ability for biomedical ontologies. Furthermore, we utilize a randomly initialized bidirectional LSTM to project the obtained word vector sequence to a fixed-length sentence vector, facilitating a flexible and uniform way for the computation of downstream tasks. We evaluate our method in two categories of tasks: the similarity access of ontology terms, and the ontology annotation-based protein-protein interaction classification. Experimental results demonstrate that our method provides encouraging results compared to the baselines in all tests. Our approach offers promising opportunities for representing ontologies semantics and in turn characterizing entities including proteins in biomedical research. Lingling Zhao, Junjie Wang 0005, Liang Cheng 0006, Chunyu Wang 0002 |
BIBM | 4 |
| 2020 | Gene Ontology aided Compound Protein Binding Affinity Prediction Using BERT EncodingabstractThe drug-target binding affinity(DTA) indicates the strength of the drug-target interaction; therefore, predicting DTA by computational approaches can considerably benefit drug discovery by narrowing down the searching space and pruning those drug-target pairs with low binding affinity scores. In the computational methods, feature representation of proteins is one of the most important parts due to its strong influence on the following regression task. This paper introduces the BERT-based language representation to embed the gene ontology annotations, combined with the raw sequence to characterize a protein by fusing its physical structure and human knowledge. We exploit CNN network stacked over full connected layers to learn the prediction of DTA scores in a supervised manner. This framework enhances the feature representation ability, leading to the improvement of the DTA prediction precision. The evaluation on the Davis and KIBA datasets compared to the state-of-the-art baselines demonstrates our feature representation's superiority. Lingling Zhao, Peijin Xie, Lingfeng Hao, Chunyu Wang 0002 |
BIBM | 5 |
| 2018 | Identification and prioritization of differentially expressed genes for time-series gene expression data
Linlin Xing, Maozu Guo 0001, Chunyu Wang 0002 |
Frontiers Comput. Sci. | 4 |
| 2017 | Refine gene functional similarity network based on interaction networksabstractBACKGROUND: In recent years, biological interaction networks have become the basis of some essential study and achieved success in many applications. Some typical networks such as protein-protein interaction networks have already been investigated systematically. However, little work has been available for the construction of gene functional similarity networks so far. In this research, we will try to build a high reliable gene functional similarity network to promote its further application. RESULTS: Here, we propose a novel method to construct and refine the gene functional similarity network. It mainly contains three steps. First, we establish an integrated gene functional similarity networks based on different functional similarity calculation methods. Then, we construct a referenced gene-gene association network based on the protein-protein interaction networks. At last, we refine the spurious edges in the integrated gene functional similarity network with the help of the referenced gene-gene association network. Experiment results indicate that the refined gene functional similarity network (RGFSN) exhibits a scale-free, small world and modular architecture, with its degrees fit best to power law distribution. In addition, we conduct protein complex prediction experiment for human based on RGFSN and achieve an outstanding result, which implies it has high reliability and wide application significance. CONCLUSIONS: Our efforts are insightful for constructing and refining gene functional similarity networks, which can be applied to build other high quality biological networks. Zhen Tian 0004, Maozu Guo 0001, Chunyu Wang 0002 |
BMC Bioinform. | 3 |
| 2016 | Constructing an integrated gene similarity network for the identification of disease genesabstractDiscovering novel genes that are involved in human diseases is a challenging task. In recent years, several computational approaches have been proposed to prioritize candidate disease genes. Most of these methods are mainly based on protein-protein interaction (PPI) networks. However, since these PPI networks contain false positives and only cover less half of known human genes, their reliability and coverage are both very low. Therefore, it is highly necessary to fuse multiple genomic data to construct a reliable gene similarity network and then infer disease genes on the whole genomic scale. Here, we proposed a novel method, named RWRB, to infer causal genes of interested disease. First, we construct five individual gene (protein) similarity networks based on multiple genomic data of human genes. Then, an integrated gene similarity network (IGSN) is reconstructed based on similarity network fusion (SNF) method. Finally, we employ the random walk with restart algorithm on the phenotype-gene bilayer network, which combines phenotype similarity network, IGSN as well as the phenotype-gene association network, to prioritize candidate disease genes. We investigate the effectiveness of RWRB through leave-one-out cross-validation methods in inferring phenotype-gene relationships. Results show that RWRB is more accurate than state-of-the-art methods on most evaluation metrics. Further analysis shows that the success of RWRB is benefited from IGSN which has a wider coverage and higher reliability comparing with current PPI networks. Zhen Tian 0004, Maozu Guo 0001, Chunyu Wang 0002, Linlin Xing, Lei Wang 0085, Yin Zhang 0009 |
BIBM | 3 |
| 2016 | Reconstructing gene regulatory network based on candidate auto selection methodabstractThe reconstruction of gene regulatory network (GRN) is a great challenge in systems biology and bioinformatics, and methods based on Bayesian network (BN) draw most of attention because of its inherent probability characteristics. As NP-hard problems, most of the BN methods often adopt the heuristic search, but they are time-consuming for biological networks with a large number of nodes. To solve this problem, this paper presents a Candidate Auto Selection algorithm (CAS) based on mutual information and breakpoint detection to limit the search space in order to accelerate the learning process. The proposed algorithm automatically restricts the neighbors of each node to a small set of candidates before structure learning. Then based on CAS algorithm, we propose a globally optimal greedy search method (CAS+G), which focuses on finding the high-scoring network structure, and a local learning method (CAS+L), which focuses on faster learning the structure with small loss of quality. Results show that the proposed CAS algorithm can effectively identify the neighbor nodes of each node. In the experiments, the CAS+G method outperforms the state-of-the-art method on simulation data for inferring GRNs, and the CAS+L method is significantly faster than the state-of-the-art method with little loss of accuracy. Hence, the CAS based algorithms are more suitable for GRN inference. Linlin Xing, Maozu Guo 0001, Chunyu Wang 0002, Lei Wang 0085, Yin Zhang 0009 |
BIBM | 4 |
| 2016 | SGFSC: speeding the gene functional similarity calculation based on hash tablesabstractBACKGROUND: In recent years, many measures of gene functional similarity have been proposed and widely used in all kinds of essential research. These methods are mainly divided into two categories: pairwise approaches and group-wise approaches. However, a common problem with these methods is their time consumption, especially when measuring the gene functional similarities of a large number of gene pairs. The problem of computational efficiency for pairwise approaches is even more prominent because they are dependent on the combination of semantic similarity. Therefore, the efficient measurement of gene functional similarity remains a challenging problem. RESULTS: To speed current gene functional similarity calculation methods, a novel two-step computing strategy is proposed: (1) establish a hash table for each method to store essential information obtained from the Gene Ontology (GO) graph and (2) measure gene functional similarity based on the corresponding hash table. There is no need to traverse the GO graph repeatedly for each method with the help of the hash table. The analysis of time complexity shows that the computational efficiency of these methods is significantly improved. We also implement a novel Speeding Gene Functional Similarity Calculation tool, namely SGFSC, which is bundled with seven typical measures using our proposed strategy. Further experiments show the great advantage of SGFSC in measuring gene functional similarity on the whole genomic scale. CONCLUSIONS: The proposed strategy is successful in speeding current gene functional similarity calculation methods. SGFSC is an efficient tool that is freely available at http://nclab.hit.edu.cn/SGFSC . The source code of SGFSC can be downloaded from http://pan.baidu.com/s/1dFFmvpZ . Zhen Tian 0004, Chunyu Wang 0002, Maozu Guo 0001, Zhixia Teng |
BMC Bioinform. | 2 |
| 2016 | MiRTDL: A Deep Learning Approach for miRNA Target PredictionabstractMicroRNAs (miRNAs) regulate genes that are associated with various diseases. To better understand miRNAs, the miRNA regulatory mechanism needs to be investigated and the real targets identified. Here, we present miRTDL, a new miRNA target prediction algorithm based on convolutional neural network (CNN). The CNN automatically extracts essential information from the input data rather than completely relying on the input dataset generated artificially when the precise miRNA target mechanisms are poorly known. In this work, the constraint relaxing method is first used to construct a balanced training dataset to avoid inaccurate predictions caused by the existing unbalanced dataset. The miRTDL is then applied to 1,606 experimentally validated miRNA target pairs. Finally, the results show that our miRTDL outperforms the existing target prediction algorithms and achieves significantly higher sensitivity, specificity and accuracy of 88.43, 96.44, and 89.98 percent, respectively. We also investigate the miRNA target mechanism, and the results show that the complementation features are more important than the others. Shuang Cheng, Maozu Guo 0001, Chunyu Wang 0002, Yang Liu 0006, Xuejian Wu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2014 | Identification of functional miRNA regulatory modules and their associations via dynamic miRNA regulatory functionabstractMicroRNAs (miRNAs) are small non-coding RNAs which cause target genes degradation or translational inhibition. Constructing functional miRNAs regulatory module can be a significant step towards the discovery of their regulatory roles in various development programs. In this paper, we present a Correlated Correspondence Regulatory Module model which builds on modified Correlated Topic Model (CTM). We apply the proposed method to the expression profiles of miRNAs and genes on 89 human cancer samples. The approach computationally predicts miRNA-gene interactions according to the negative or positive correlation relationship between miRNA and gene expression data and identifies functional miRNA regulatory modules from which we can infer multiple and dynamic miRNA function according to the known elements, the result shows consistency with published literature and database. Furthermore, a miRNA regulatory network is constructed in order to study the associations among various regulatory modules, we can detect evolution of miRNA function in biological process according these associations, which solve restriction of traditional methods that only focus on static miRNA function in single regulatory module. Online services can be accessed at the website (http://nclab.hit.edu.cn/CCRM). Shuang Cheng, Maozu Guo 0001, Chunyu Wang 0002, Yang Liu 0006 |
BIBM | 3 |
| 2014 | Inferring the soybean (Glycine max) microRNA functional network based on target gene networkabstractMOTIVATION: The rapid accumulation of microRNAs (miRNAs) and experimental evidence for miRNA interactions has ushered in a new area of miRNA research that focuses on network more than individual miRNA interaction, which provides a systematic view of the whole microRNome. So it is a challenge to infer miRNA functional interactions on a system-wide level and further draw a miRNA functional network (miRFN). A few studies have focused on the well-studied human species; however, these methods can neither be extended to other non-model organisms nor take fully into account the information embedded in miRNA-target and target-target interactions. Thus, it is important to develop appropriate methods for inferring the miRNA network of non-model species, such as soybean (Glycine max), without such extensive miRNA-phenotype associated data as miRNA-disease associations in human. RESULTS: Here we propose a new method to measure the functional similarity of miRNAs considering both the site accessibility and the interactive context of target genes in functional gene networks. We further construct the miRFNs of soybean, which is the first study on soybean miRNAs on the network level and the core methods can be easily extended to other species. We found that miRFNs of soybean exhibit a scale-free, small world and modular architecture, with their degrees fit best to power-law and exponential distribution. We also showed that miRNA with high degree tends to interact with those of low degree, which reveals the disassortativity and modularity of miRFNs. Our efforts in this study will be useful to further reveal the soybean miRNA-miRNA and miRNA-gene interactive mechanism on a systematic level. AVAILABILITY AND IMPLEMENTATION: A web tool for information retrieval and analysis of soybean miRFNs and the relevant target functional gene networks can be accessed at SoymiRNet: http://nclab.hit.edu.cn/SoymiRNet. Yungang Xu, Maozu Guo 0001, Chunyu Wang 0002, Yang Liu 0006 |
Bioinform. | 4 |
| 2014 | CPL: Detecting Protein Complexes by Propagating Labels on Protein-Protein Interaction Network
Qiguo Dai, Maozu Guo 0001, Zhixia Teng, Chunyu Wang 0002 |
J. Comput. Sci. Technol. | 5 |
| 2013 | Effective constructing training sets for object detectionabstractThis paper addresses the problem of building up effective training sets at minimal labeling cost for object detection. This problem occurs in the situation that the part-based detector is trained on a group of positive examples with bounding box labels, but the images selected by uniform sampling do not reflect the desired training distribution and need additional labeling cost in order to obtain enough positive examples. We study the active training process in which some object windows are sampled from a pool of unlabeled candidate windows, and then their corresponding bounding annotations are queried. We derive an effective training set by selecting a group of most uncertain object windows according to the current detector. Our approach has been empirically demonstrated on the object detection task of PASCAL VOC dataset. The experiment results show that our proposed algorithm outperforms common uniform sampling within the same labeling cost. Weining Wu, Yang Liu 0006, Wei Zeng 0006, Maozu Guo 0001, Chunyu Wang 0002 |
ICIP | 5 |
| 2013 | Measuring gene functional similarity based on group-wise comparison of GO termsabstractMOTIVATION: Compared with sequence and structure similarity, functional similarity is more informative for understanding the biological roles and functions of genes. Many important applications in computational molecular biology require functional similarity, such as gene clustering, protein function prediction, protein interaction evaluation and disease gene prioritization. Gene Ontology (GO) is now widely used as the basis for measuring gene functional similarity. Some existing methods combined semantic similarity scores of single term pairs to estimate gene functional similarity, whereas others compared terms in groups to measure it. However, these methods may make error-prone judgments about gene functional similarity. It remains a challenge that measuring gene functional similarity reliably. RESULT: We propose a novel method called SORA to measure gene functional similarity in GO context. First of all, SORA computes the information content (IC) of a term making use of semantic specificity and coverage. Second, SORA measures the IC of a term set by means of combining inherited and extended IC of the terms based on the structure of GO. Finally, SORA estimates gene functional similarity using the IC overlap ratio of term sets. SORA is evaluated against five state-of-the-art methods in the file on the public platform for collaborative evaluation of GO-based semantic similarity measure. The carefully comparisons show SORA is superior to other methods in general. Further analysis suggests that it primarily benefits from the structure of GO, which implies expressive information about gene function. SORA offers an effective and reliable way to compare gene function. AVAILABILITY: The web service of SORA is freely available at http://nclab.hit.edu.cn/SORA/ Zhixia Teng, Maozu Guo 0001, Qiguo Dai, Chunyu Wang 0002, Ping Xuan |
Bioinform. | 5 |
| 2013 | Lnetwork: an efficient and effective method for constructing phylogenetic networksabstractMOTIVATION: The evolutionary history of species is traditionally represented with a rooted phylogenetic tree. Each tree comprises a set of clusters, i.e. subsets of the species that are descended from a common ancestor. When rooted phylogenetic trees are built from several different datasets (e.g. from different genes), the clusters are often conflicting. These conflicting clusters cannot be expressed as a simple phylogenetic tree; however, they can be expressed in a phylogenetic network. Phylogenetic networks are a generalization of phylogenetic trees that can account for processes such as hybridization, horizontal gene transfer and recombination, which are difficult to represent in standard tree-like models of evolutionary histories. There is currently a large body of research aimed at developing appropriate methods for constructing phylogenetic networks from cluster sets. The Cass algorithm can construct a much simpler network than other available methods, but is extremely slow for large datasets or for datasets that need lots of reticulate nodes. The networks constructed by Cass are also greatly dependent on the order of input data, i.e. it generally derives different phylogenetic networks for the same dataset when different input orders are used. RESULTS: In this study, we introduce an improved Cass algorithm, Lnetwork, which can construct a phylogenetic network for a given set of clusters. We show that Lnetwork is significantly faster than Cass and effectively weakens the influence of input data order. Moreover, we show that Lnetwork can construct a much simpler network than most of the other available methods. AVAILABILITY: Lnetwork has been built as a Java software package and is freely available at http://nclab.hit.edu.cn/∼wangjuan/Lnetwork/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Juan Wang 0011, Maozu Guo 0001, Yang Liu 0006, Chunyu Wang 0002, Linlin Xing, Kai Che |
Bioinform. | 5 |
| 2013 | A probabilistic model of active learning with multiple noisy oracles
Weining Wu, Yang Liu 0006, Maozu Guo 0001, Chunyu Wang 0002 |
Neurocomputing | 4 |
| 2009 | CGTS: a site-clustering graph based tagSNP selection algorithm in genotype dataabstractBACKGROUND: Recent studies have shown genetic variation is the basis of the genome-wide disease association research. However, due to the high cost on genotyping large number of single nucleotide polymorphisms (SNPs), it is essential to choose a small subset of informative SNPs (tagSNPs), which are able to capture most variation in a population, to represent the rest SNPs. Several methods have been proposed to find the minimum set of tagSNPs, but most of them still have some disadvantages such as information loss and block-partition limit. RESULTS: This paper proposes a new hybrid method named CGTS which combines the ideas of the clustering and the graph algorithms to select tagSNPs on genotype data. This method aims to maximize the number of the discarding nontagSNPs in the given set. CGTS integrates the information of the LD association and the genotype diversity using the site graphs, discards redundant SNPs using the algorithm based on these graph structures. The clustering algorithm is used to reduce the running time of CGTS. The efficiency of the algorithm and quality of solutions are evaluated on biological data and the comparisons with three popular selecting methods are shown in the paper. CONCLUSION: Our theoretical analysis and experimental results show that our algorithm CGTS is not only more efficient than other methods but also can be get higher accuracy in tagSNP selection. Jun Wang 0035, Maozu Guo 0001, Chunyu Wang 0002 |
BMC Bioinform. | 3 |
| 2009 | A hybrid clustering and graph based algorithm for tagSNP selection
Maozu Guo 0001, Jun Wang 0035, Chunyu Wang 0002, Yang Liu 0006 |
Soft Comput. | 3 |