VLDB 2026 Research / reviewers in the wild / expert
Xin Wang 0124
dblp:10/5630-124
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0001-8341-2323ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | GRIT: An Accurate and Efficient Graph Stream Summarization for Temporal QueryabstractGraph stream summarization refers to the technique used to process graph streams-unbounded sequences of edges-by constructing compressed representations that support approximate queries on both graph topology and temporal information in computing power networks. However, existing methods struggle to achieve accurate and efficient temporal queries due to two key limitations: (1) inefficient integration of temporal information, leading to high latency in both edge processing and query execution; and (2) redundant multilayer structures that accumulate errors, significantly reducing query accuracy. In this paper, we propose GRIT, an accurate and efficient Graph stReam summarIzation for Temporal query. GRIT introduces a new structure FlatIndex, which organizes temporal information in a flattened form, playing a critical role in minimizing error accumulation and ensuring accurate temporal queries. To further enhance edge processing efficiency, we introduce a lazy update strategy, which updates only a single element in the FlatIndex upon edge insertion, significantly reducing insertion latency. Moreover, our greedy-based decomposition (GBD) algorithm decomposes the target query range into the minimal number of intervals corresponding to the FlatIndex, enabling efficient execution of temporal queries over arbitrary time ranges. Extensive experiments on five real-world datasets demonstrate that GRIT improves query accuracy by 2-3 orders of magnitude, while reducing query latency by 1-2 orders of magnitude and increasing throughput by 7-13 times compared to state-of-the-art methods. Jingxian Hu, Guozhang Sun, Xin Wang 0124, Yuhai Zhao, Yuan Li 0008, Xingwei Wang 0001 |
CIKM | 3 |
| 2025 | GPU-Powered Evolutionary Auxiliary Multitasking for Fast SNP Interaction DetectionabstractIdentifying complex interactions among millions of single nucleotide polymorphisms (SNPs) is a key challenge in Genome-Wide Association Studies (GWAS), offering crucial insights into the genetic architecture of complex diseases. Evolutionary algorithm (EA)-based methods have gained significant attention for their global search capabilities, controllable runtime, and multi-objective optimization potential. However, when applied to high-dimensional GWAS datasets, many existing EA-based methods encounter challenges such as getting trapped in local optima and facing high computational demands. To address these issues, the evolutionary multitasking (EMT) paradigm presents a promising solution, enhancing population diversity and convergence speed through collaborative, cross-task knowledge sharing. Furthermore, the multi-tasking framework and EA can be seamlessly deployed across multiple Graphics Processing Units (GPUs), leveraging their high parallelism and aggregated memory bandwidth. Therefore, we introduce a GPU-powered evolutionary auxiliary multitasking algorithm (GEAMT) for fast SNP interaction detection. GEAMT first constructs a main task along with several low-dimensional auxiliary tasks to redefine the original task. The main task explores the entire search space, while the auxiliary tasks search distinct subspaces to enhance local optimization capabilities. In each iteration, the auxiliary tasks transfer high-quality information to the main task via an information transfer mechanism. Subsequently, an auxiliary task update strategy based on feature regrouping is employed to switch the search subspaces of the auxiliary tasks. The final results are derived from the Pareto-optimal solutions of the main task. Implemented across multiple GPUs, GEAMT achieves notable scalability and efficiency. Comprehensive experiments on both synthetic and real-world datasets demonstrate that GEAMT can significantly enhance search accuracy and speed up the search process. Ying Yin 0001, Xin Wang 0124, Changyong Yu, Yuhai Zhao |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2023 | miProBERT: identification of microRNA promoters based on the pre-trained model BERTabstractAccurate prediction of promoter regions driving miRNA gene expression has become a major challenge due to the lack of annotation information for pri-miRNA transcripts. This defect hinders our understanding of miRNA-mediated regulatory networks. Some algorithms have been designed during the past decade to detect miRNA promoters. However, these methods rely on biosignal data such as CpG islands and still need to be improved. Here, we propose miProBERT, a BERT-based model for predicting promoters directly from gene sequences without using any structural or biological signals. According to our information, it is the first time a BERT-based model has been employed to identify miRNA promoters. We use the pre-trained model DNABERT, fine-tune the pre-trained model on the gene promoter dataset so that the model includes information about the richer biological properties of promoter sequences in its representation, and then systematically scan the upstream regions of each intergenic miRNA using the fine-tuned model. About, 665 miRNA promoters are found. The innovative use of a random substitution strategy to construct a negative dataset improves the discriminative ability of the model and further reduces the false positive rate (FPR) to as low as 0.0421. On independent datasets, miProBERT outperformed other gene promoter prediction methods. With comparison on 33 experimentally validated miRNA promoter datasets, miProBERT significantly outperformed previously developed miRNA promoter prediction programs with 78.13% precision and 75.76% recall. We further verify the predicted promoter regions by analyzing conservation, CpG content and histone marks. The effectiveness and robustness of miProBERT are highlighted. Xin Wang 0124, Xin Gao 0001, Guohua Wang 0001 |
Briefings Bioinform. | 1 |
| 2023 | MicroRNA Promoter Identification in Human With a Three-level Prediction MethodabstractThe accurate annotation of miRNA promoters is critical for the mechanistic understanding of miRNA gene regulation. Various computational methods have been developed for the prediction of miRNA promoters solely employing a single classifier. Most of these computational methods extract either sequence features or one-sided signal features, and the accuracy and reliability of predictions need to be improved. To address these issues, we present miPTP, a three-level prediction method that combines SVM, RF, and correlation coefficients. It is capable of identifying miRNA promoters based on both DNA sequence and ChIP-Seq data (RPol II). By sequentially integrating these two types of information sources with the three methods selected, miPTP can identify miRNA promoters with higher accuracy and sensitivity compared to specific existing methods. Finally, the reliability of miPTP is validated by examining the conservation, CpG content, and activating histone marks in the identified miRNA promoters. Xin Wang 0124, Jie Li 0055, Guohua Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | StackCirRNAPred: computational classification of long circRNA from other lncRNA based on stacking strategyabstractBACKGROUND: CircRNAs are essential for the regulation of post-transcriptional gene expression, including as miRNA sponges, and play an important role in disease development. Some computational tools have been proposed recently to predict circRNA, since only one classifier is used, there is still much that can be done to improve the performance. RESULTS: StackCirRNAPred was proposed, the computational classification of long circRNA from other lncRNA based on stacking strategy. In order to cope with the potential problem that a single feature might not be able to distinguish circRNA well from other lncRNA, we first extracted features from different sources, including nucleic acid composition, sequence spatial features and physicochemical properties, Alu and tandem repeats. We innovatively apply the stacking strategy to integrate the more advantageous classifiers of RF, LightGBM, XGBoost. This allows the model to incorporate these features more flexibly. StackCirRNAPred was found to be significantly better than other tools, with precision, accuracy, F1, recall and MCC of 0.843, 0.833, 0.831, 0.819 and 0.666 respectively. We tested it directly on the mouse dataset. StackCirRNAPred was still significantly better than other methods, with precision, accuracy, F1, recall and MCC of 0.837, 0.839, 0.839, 0.841, 0.677. CONCLUSIONS: We proposed StackCirRNAPred based on stacking strategy to distinguish long circRNAs from other lncRNAs. With the test results demonstrating the validity and robustness of StackCirRNAPred, we hope StackCirRNAPred will complement existing circRNA prediction methods and is helpful in down-stream research. Xin Wang 0124, Yadong Liu 0001, Jie Li 0055, Guohua Wang 0001 |
BMC Bioinform. | 1 |
| 2021 | The stacking strategy-based hybrid framework for identifying non-coding RNAsabstractWith the development of next-generation sequencing technology, a large number of transcripts need to be analyzed, and it has been a challenge to distinguish non-coding ribonucleic acid (RNAs) (ncRNAs) from coding RNAs. And for non-model organisms, due to the lack of transcriptional data, many existing methods cannot identify them. Therefore, in addition to using deoxyribonucleic acid-based and RNA-based features, we also proposed a hybrid framework based on the stacking strategy to identify ncRNAs, and we innovatively added eight features based on predicted peptides. The proposed framework was based on stacking two-layer classifier which combined random forest (RF), LightGBM, XGBoost and logistic regression (LR) models. We used this framework to build two types of models. For cross-species ncRNAs identification model, we tested it on six different species: human, mouse, zebrafish, fruit fly, worm and Arabidopsis. Compared with other tools, our model was the best in datasets of Arabidopsis, worm and zebrafish with the accuracy of 98.36%, 99.65% and 94.12%. For performance metrics analysis, the datasets of the six species were considered as a whole set, and the sensitivity, accuracy, precision and F1 values of our model were the best. For the plant-specific ncRNAs identification model, the average values of the six metrics of the two experiments were all greater than 95%, which demonstrated it can be used to identify ncRNAs in plants. The above indicates that the hybrid framework we designed is universal between animals and plants and has significant advantages in the identification of cross-species ncRNAs. Xin Wang 0124, Guohua Wang 0001 |
Briefings Bioinform. | 1 |