VLDB 2026 Research / reviewers in the wild / expert
Xubin Wang 0001
dblp:303/0383-1
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2026
0000-0001-6217-1305ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 67% Reinforcement learning · 33% | |
| Databases, data mining, and information retrieval
1 paper |
Data mining · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › in-context learning
demonstration selection |
0.9 | 1 | 2025 | Demonstration Selection for In-Context Learning via Reinforcement Learning · ICML 2025 |
Natural language and speech › Language models and text generation
in-context learning |
0.9 | 1 | 2025 | Demonstration Selection for In-Context Learning via Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning
policy optimization |
0.9 | 1 | 2025 | Demonstration Selection for In-Context Learning via Reinforcement Learning · ICML 2025 |
Data mining › dimensionality reduction
feature selection |
0.8 | 1 | 2024 | MEL: Efficient Multi-Task Evolutionary Learning for High-Dimensional Feature Selection · IEEE Trans. Knowl. Data Eng. 2024 |
Data mining › dimensionality reduction › feature selection
high-dimensional feature selection |
0.8 | 1 | 2024 | MEL: Efficient Multi-Task Evolutionary Learning for High-Dimensional Feature Selection · IEEE Trans. Knowl. Data Eng. 2024 |
Methods — techniques the papers use, named apart from their topics
q-learning · 0.9chain-of-thought · 0.9PPO · 0.9particle swarm optimization · 0.8multi-task learning · 0.8evolutionary computation · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A niching archive-assisted evolutionary algorithm for multimodal feature selection in high-dimensional data classification
Yunhe Wang 0002, Zhengyu Du, Zeming Zhou, Xubin Wang 0001, Shengxiang Yang |
Knowl. Based Syst. | 4 |
| 2025 | Enhancing Text Annotation Through Rationale-Driven Collaborative Few-Shot PromptingabstractThe traditional data annotation process is often labor-intensive, time-consuming, and susceptible to human bias, which complicates the management of increasingly complex datasets. This study explores the potential of large language models (LLMs) as automated data annotators to improve efficiency and consistency in annotation tasks. By employing rationale-driven collaborative few-shot prompting techniques, we aim to improve the performance of LLMs in text annotation. We conduct a rigorous evaluation of five LLMs across four benchmark datasets, comparing seven distinct methodologies. Our results demonstrate that collaborative methods consistently outperform traditional few-shot techniques and other baseline approaches, particularly in complex annotation tasks. Our work provides valuable insights and a robust framework for leveraging collaborative learning methods to tackle challenging text annotation tasks. Jianfei Wu, Xubin Wang 0001, Weijia Jia 0001 |
ICASSP | 2 |
| 2025 | Demonstration Selection for In-Context Learning via Reinforcement LearningabstractDiversity in demonstration selection is critical for enhancing model generalization by enabling broader coverage of structures and concepts. Constructing appropriate demonstration sets remains a key research challenge. This paper introduces the Relevance-Diversity Enhanced Selection (RDES), an innovative approach that leverages reinforcement learning (RL) frameworks to optimize the selection of diverse reference demonstrations for tasks amenable to in-context learning (ICL), particularly text classification and reasoning, in few-shot prompting scenarios. RDES employs frameworks like Q-learning and a PPO-based variant to dynamically identify demonstrations that maximize both diversity (quantified by label distribution) and relevance to the task objective. This strategy ensures a balanced representation of reference data, leading to improved accuracy and generalization. Through extensive experiments on multiple benchmark datasets, including diverse reasoning tasks, and involving 14 closed-source and open-source LLMs, we demonstrate that RDES significantly enhances performance compared to ten established baselines. Our evaluation includes analysis of performance across varying numbers of demonstrations on selected datasets. Furthermore, we investigate incorporating Chain-of-Thought (CoT) reasoning, which further boosts predictive performance. The results highlight the potential of RL for adaptive demonstration selection and addressing challenges in ICL. Xubin Wang 0001, Jianfei Wu, Deyu Cai, Weijia Jia 0001 |
ICML | 1 |
| 2024 | Evolving pathway activation from cancer gene expression data using nature-inspired ensemble optimization
Xubin Wang 0001, Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
Expert Syst. Appl. | 1 |
| 2024 | Exhaustive Exploitation of Nature-Inspired Computation for Cancer Screening in an Ensemble MannerabstractAccurate screening of cancer types is crucial for effective cancer detection and precise treatment selection. However, the association between gene expression profiles and tumors is often limited to a small number of biomarker genes. While computational methods using nature-inspired algorithms have shown promise in selecting predictive genes, existing techniques are limited by inefficient search and poor generalization across diverse datasets. This study presents a framework termed Evolutionary Optimized Diverse Ensemble Learning (EODE) to improve ensemble learning for cancer classification from gene expression data. The EODE methodology combines an intelligent grey wolf optimization algorithm for selective feature space reduction, guided random injection modeling for ensemble diversity enhancement, and subset model optimization for synergistic classifier combinations. Extensive experiments were conducted across 35 gene expression benchmark datasets encompassing varied cancer types. Results demonstrated that EODE obtained significantly improved screening accuracy over individual and conventionally aggregated models. The integrated optimization of advanced feature selection, directed specialized modeling, and cooperative classifier ensembles helps address key challenges in current nature-inspired approaches. This provides an effective framework for robust and generalized ensemble learning with gene expression biomarkers. Xubin Wang 0001, Yunhe Wang 0002, Zhiqiang Ma 0003, Ka-Chun Wong, Xiangtao Li |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | MEL: Efficient Multi-Task Evolutionary Learning for High-Dimensional Feature SelectionabstractFeature selection is a crucial step in data mining to enhance model performance by reducing data dimensionality. However, the increasing dimensionality of collected data exacerbates the challenge known as the “curse of dimensionality”, where computation grows exponentially with the number of dimensions. To tackle this issue, evolutionary computational (EC) approaches have gained popularity due to their simplicity and applicability. Unfortunately, the diverse designs of EC methods result in varying abilities to handle different data, often underutilizing and not sharing information effectively. In this article, we propose a novel approach called PSO-based Multi-task Evolutionary Learning (MEL) that leverages multi-task learning to address these challenges. By incorporating information sharing between different feature selection tasks, MEL achieves enhanced learning ability and efficiency. We evaluate the effectiveness of MEL through extensive experiments on 22 high-dimensional datasets. Comparing against 24 EC approaches, our method exhibits strong competitiveness. In addition, we have open-sourced our code on GitHub. Xubin Wang 0001, Haojiong Shangguan, Shangrui Wu, Weijia Jia 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | A Feature Weighting Particle Swarm Optimization Method to Identify Biomarker GenesabstractThe discovery of biomarker genes from gene expression data is a hot topic for understanding the mechanisms underlying disease etiology. However, while the collection of high-dimensional gene expression data has been made possible by the adoption of technologies such as DNA microarray, it also poses challenges for the identification of key disease-causing genes due to its high-dimensional nature. To address this problem, we propose a feature weighting particle swarm optimization method (FWPSO) for efficiently identifying biomarker genes from high-dimensional microarray data. Specifically, there are two significant phases in FWPSO: 1) Feature Weighting Phase: Features will be discriminated into relevant and irrelevant based on the evolutionary performance of individuals in the PSO population in each generation, and features will be assigned weights based on this. 2) Feature Selection Phase: By focusing the search on a feature set that have been determined to be relevant based on the results of the previous phase, the PSO population will improve the efficiency of removing redundant features and discovering the most related genes. Both phases work together and operate in synergy to achieve the optimized results. The experimental results on four microarray datasets shows that FWPSO not only reduces the number of feature dimensions to a large extent, but also achieves higher classification accuracy compared to other methods, demonstrating the effectiveness of our method. Our implementation of FWPSO is available at https://github.com/wangxb96/FWPSO. Xubin Wang 0001, Weijia Jia 0001 |
BIBM | 1 |
| 2022 | A self-adaptive weighted differential evolution approach for large-scale feature selection
Xubin Wang 0001, Yunhe Wang 0002, Ka-Chun Wong, Xiangtao Li |
Knowl. Based Syst. | 1 |