VLDB 2026 Research / reviewers in the wild / expert
Zhaowei Wang 0005
dblp:120/1278-5
· DBLP profile ↗
9ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-3703-287XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Twin cross contrastive learning with multi-modality fusion for drug-target affinity prediction
Linna Zhang, Zhaowei Wang 0005, Wuhao Liu, Xiaodong Duan, Qiguo Dai |
Artif. Intell. Medicine | 2 |
| 2025 | Attention-augmented multi-domain cooperative graph representation learning for molecular interaction prediction
Zhaowei Wang 0005, Jun Meng, Qiguo Dai, Xiaohui Lin 0002, Yushi Luan |
Neural Networks | 1 |
| 2024 | DeepPepPI: A deep cross-dependent framework with information sharing mechanism for predicting plant peptide-protein interactions
Zhaowei Wang 0005, Jun Meng, Qiguo Dai, Shihao Xia, Ruirui Yang, Yushi Luan |
Expert Syst. Appl. | 1 |
| 2023 | TGAAL: Combining Transformer-based GAN and active learning to identify the coding potential of sORFs in plant lncRNAsabstractSome small open reading frames (sORFs) in plant long non-coding RNAs (lncRNAs) are capable of encoding small peptides, which play key roles in the growth and development of organisms. Therefore, it is particularly important to identify the coding potential of sORFs in plant lncRNAs. However, existing methods often ignore the differences in length distribution between coding sORFs (csORFs) and non-coding sORFs (non-csORFs), which may lead to incorrect identification of csORFs. To address this issue, we propose a novel method to identify the coding potential of sORFs in plant lncRNAs, named Transformer Generative Adversarial Active Learning (TGAAL), which combines Transformer-based Generative Adversarial Network (TGAN) and active learning based on KL-topk sampling strategy. TGAN can generate sORF sequences in a specific length interval, which have the same class as the input sORFs. Meanwhile, using active learning based on KL-topk sampling strategy, samples with high confidence can be selected for data augmentation. 5-fold cross-validation shows that KL-topk sampling strategy significantly improves the prediction performance compared with commonly adopted sampling strategies. The experimental results show that TGAAL significantly outperforms existing methods in identifying the coding potential of sORFs in Arabidopsis thaliana, reaching 0.7761, 0.7906 and 0.7529 unweighted average recall in three sORF length intervals, respectively. Jun Meng, Shihao Xia, Zhaowei Wang 0005, Yushi Luan |
BIBM | 5 |
| 2023 | A multi-granularity information-enhanced pre-training method for predicting the coding potential of sORFs in plant lncRNAsabstractSmall open reading frames (sORFs) are nucleotide sequences that may be translated into small peptides. Recently, increasing studies have demonstrated that peptides encoded by sORFs in plant long noncoding RNAs (lncRNAs) play a vital role in growth regulation and disease treatment. To accelerate the discovery of lncRNA-encoded peptides, it is essential to predict translatable sORFs in lncRNAs (lncRNA-sORFs) by computational methods. As only a few translatable plant lncRNA-sORFs have been discovered to date, there is a lack of effective methods for characterizing the coding potential of lncRNA-sORFs in data-scarce scenarios. Therefore, a novel method for plant lncRNA-sORFs coding potential prediction using the pre-trained bidirectional encoder representations from transformer (LSCPP-BERT) is proposed. Firstly, the BERT model is trained to extract multi-granularity context information from large-scale unlabeled lncRNA-sORFs through two pre-training tasks. Then, the pre-trained model can be fine-tuned with two additional linear layers for classification. The LSCPP-BERT is featured by a self-supervised pre-training scheme and multi-granularity context information, aiming to enhance the representational power of the network. In addition, an extra pre-training task called contextual relation of lncRNA-sORFs prediction (CRSP) is presented to extract sentence-level information. Experiment results show that the accuracy of LSCPP-BERT is increased by 8.14% compared with state-of-the-art methods. We hope that the proposed method can serve as a reliable tool for the prediction of coding lncRNA-sORFs, thereby further contributing to drug development and agronomical applications. Shihao Xia, Jun Meng, Zhaowei Wang 0005, Zhaojing Qin, Yushi Luan |
BIBM | 3 |
| 2022 | GraphCDA: a hybrid graph representation learning framework based on GCN and GAT for predicting disease-associated circRNAsabstractMOTIVATION: CircularRNA (circRNA) is a class of noncoding RNA with high conservation and stability, which is considered as an important disease biomarker and drug target. Accumulating pieces of evidence have indicated that circRNA plays a crucial role in the pathogenesis and progression of many complex diseases. As the biological experiments are time-consuming and labor-intensive, developing an accurate computational prediction method has become indispensable to identify disease-related circRNAs. RESULTS: We presented a hybrid graph representation learning framework, named GraphCDA, for predicting the potential circRNA-disease associations. Firstly, the circRNA-circRNA similarity network and disease-disease similarity network were constructed to characterize the relationships of circRNAs and diseases, respectively. Secondly, a hybrid graph embedding model combining Graph Convolutional Networks and Graph Attention Networks was introduced to learn the feature representations of circRNAs and diseases simultaneously. Finally, the learned representations were concatenated and employed to build the prediction model for identifying the circRNA-disease associations. A series of experimental results demonstrated that GraphCDA outperformed other state-of-the-art methods on several public databases. Moreover, GraphCDA could achieve good performance when only using a small number of known circRNA-disease associations as the training set. Besides, case studies conducted on several human diseases further confirmed the prediction capability of GraphCDA for predicting potential disease-related circRNAs. In conclusion, extensive experimental results indicated that GraphCDA could serve as a reliable tool for exploring the regulatory role of circRNAs in complex diseases. Qiguo Dai, Zhaowei Wang 0005, Xiaodong Duan, Maozu Guo 0001 |
Briefings Bioinform. | 3 |
| 2022 | Predicting miRNA-disease associations using an ensemble learning framework with resampling methodabstractMOTIVATION: Accumulating evidences have indicated that microRNA (miRNA) plays a crucial role in the pathogenesis and progression of various complex diseases. Inferring disease-associated miRNAs is significant to explore the etiology, diagnosis and treatment of human diseases. As the biological experiments are time-consuming and labor-intensive, developing effective computational methods has become indispensable to identify associations between miRNAs and diseases. RESULTS: We present an Ensemble learning framework with Resampling method for MiRNA-Disease Association (ERMDA) prediction to discover potential disease-related miRNAs. Firstly, the resampling strategy is proposed for building multiple different balanced training subsets to address the challenge of sample imbalance within the database. Then, ERMDA extracts miRNA and disease feature representations by integrating miRNA-miRNA similarities, disease-disease similarities and experimentally verified miRNA-disease association information. Next, the feature selection approach is applied to reduce the redundant information and increase the diversity among these subsets. Lastly, ERMDA constructs an individual learner on each subset to yield primitive outcomes, and the soft voting method is introduced for making the final decision based on the prediction results of individual learners. A series of experimental results demonstrates that ERMDA outperforms other state-of-the-art methods on both balanced and unbalanced testing sets. Besides, case studies conducted on the three human diseases further confirm the ERMDA's prediction capability for identifying potential disease-related miRNAs. In conclusion, these experimental results demonstrate that our method can serve as an effective and reliable tool for researchers to explore the regulatory role of miRNAs in complex diseases. Qiguo Dai, Zhaowei Wang 0005, Xiaodong Duan, Jinmiao Song, Maozu Guo 0001 |
Briefings Bioinform. | 2 |
| 2022 | Predicting RBP Binding Sites of RNA With High-Order Encoding Features and CNN-BLSTM Hybrid ModelabstractRNA binding protein (RBP) is extensively involved in various cellular regulatory processes through the interaction with RNAs. Capturing the RBP binding preferences is fundamental for revealing the pathogenesis of complex diseases. Many experimental detection techniques are still time-consuming and labor-intensive, therefore, it is indispensable to develop a computational method with convincing accuracy. In this study, we proposed a CNN-BLSTM hybrid deep learning framework, named DeepDW, for predicting the RBP binding sites on RNAs with high-order encoding features of RNA sequence and secondary structure. The high-order encoding strategy was used to characterize the dependencies among adjacency nucleotides. For CNN-BLSTM hybrid model, DeepDW first employed two 1-D convolutional neural networks (CNNs) for learning the local features from high-order encoded matrices of RNA sequence and structure separately, and then applied two bidirectional long short-term memory networks (BLSTMs) to capture the global information in a higher level. Moreover, a series of experiments were carried out on 31 public datasets to evaluate our proposed framework, and DeepDW achieved superior performance than the state-of-the-art methods. The results indicated that the combination of high-order encoding method and CNN-BLSTM hybrid model had advantages in identifying RBP-RNA binding sites. Zhaowei Wang 0005, Qiguo Dai, Jinmiao Song, Xiaodong Duan, Hongpeng Yang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | A Stacked Ensemble Learning Framework with Heterogeneous Feature Combinations for Predicting ncRNA-Protein InteractionabstractThe interaction between ncRNA and protein is a kind of crucial molecular activities in a cell. Developing computational methods to predict ncRNA-protein interactions has attracted increasing attentions in recent years. In this work, a novel stacked ensemble learning framework is presented for predicting ncRNA-protein interaction based on heterogeneous feature combinations, named HFC-RPI. Firstly, the compositional features of k-mer with different orders were extracted from the primary sequence and secondary structure of RNA and protein respectively. Secondly, we trained a set of base learners using a variety of heterogeneous combinations of the extracted features respectively. Thirdly, the prediction results of these base learners were employed to train the stacked learner, which output the final prediction result at the higher layer in HFC-RPI. Moreover, in order to improve the generalization of HFC-RPI, when training the base learners, a cross-validation based method was applied. Extensive experimental results showed that the proposed learning framework HFC-RPI was effective and feasible for predicting the interaction of ncRNA and protein. By comparing with state-of-the-art methods, HFC-RPI was superior to them on most performance evaluation metrics. Qiguo Dai, Zhaowei Wang 0005, Jinmiao Song, Xiaodong Duan, Maozu Guo 0001, Zhen Tian 0004 |
BIBM | 2 |