VLDB 2026 Research / reviewers in the wild / expert
Xiangchen Wu
dblp:37/1097
· DBLP profile ↗
8ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Solving the Min-Max Multiple Traveling Salesmen Problem via Learning-Based Path Generation and Optimal SplittingabstractThis study addresses the Min-Max Multiple Traveling Salesmen Problem (m3-TSP), which aims to coordinate tours for multiple salesmen such that the length of the longest tour is minimized. Due to its NP-hard nature, exact solvers become impractical under the assumption that P ≠ NP. As a result, learning-based approaches have gained traction for their ability to rapidly generate high-quality approximate solutions. Among these, two-stage methods combine learning-based components with classical solvers, simplifying the learning objective. However, this decoupling often disrupts consistent optimization, potentially degrading solution quality. To address this issue, we propose a novel two-stage framework named Generate-and-Split (GaS), which integrates reinforcement learning (RL) with an optimal splitting algorithm in a joint training process. The splitting algorithm offers near-linear scalability with respect to the number of cities and guarantees optimal splitting in Euclidean space for any given path. To facilitate the joint optimization of the RL component with the algorithm, we adopt an LSTM-enhanced model architecture to address partial observability. Extensive experiments show that the proposed GaS framework significantly outperforms existing learning-based approaches in both solution quality and transferability. Xiangchen Wu, Liang Wang 0006, Hao Hu 0001, XianPing Tao, Linghao Zhang |
ECAI | 2 |
| 2025 | EarlyPR: Early Prediction of Potential Pull-Requests from ForksabstractIn this work, we propose the EarlyPR framework that identifies and predicts potential pull-request (PR) contri-butions from an open source software (OSS) project's forks, which can potentially improve the efficiency of the fork-and-pull based development in OSS projects by supporting early warning of duplicated and rejected contributions, and detection of lost contributions. Unlike traditional, PR-based studies that rely on the descriptions and contents of PRs provided by their creators, which are only available after the PRs are created, EarlyPR makes predictions before the creation of PRs by mining the forks' commit history. EarlyPR's task is challenging because of the explosive number of commit subsets in a fork's commit history that may form PRs, and the absence of resulting, real PR-related information. To tackle the challenges, we adopt the state-of-the-art, Transformer-based architecture to extract rich statistical and content information from the forks and their commits to support the prediction of potential PR contributions. And to make the algorithms scalable, we devise a TemporalFilter to find candidate PRs by mimicking the real-world processes of picking subsets of commits from a fork's commit history when creating PRs. Experimental results on real-world OSS project data suggest that EarlyPR is effective in predicting PRs, which are essentially sets of commits selected from forks to compose these PRs. Experimental results obtained using real-world OSS projects' and their forks' data suggest that EarlyPR is effective by achieving a hitting rate of 0.790 and a missing rate of 0.367 by matching the predicted and real PRs under a stringent criterion of IoU > 0.5. We further demonstrate that we can forecast the merging of PRs based on EarlyPR's predictions with an accuracy of 70.8%. In summary, the proposed approach can potentially improve the efficiency of the fork-and-pull based OSS development by making accurate and early predictions of PR contributions from the distributed, and often independently, developed forks. Xiangchen Wu, Liang Wang 0006, XianPing Tao |
SANER | 1 |
| 2025 | An entropy-based measure of fork diversity and its correlations with open source software projects' received contributions
Xiangchen Wu, Liang Wang 0006, Baihui Sang, Jierui Zhang, XianPing Tao |
Empir. Softw. Eng. | 1 |
| 2024 | FSL-DSC: A Hybrid Pap Smear Cervical Cancer Image Classification Framework Using Few-shot Learning with Depthwise Separable ConvolutionsabstractCervical cancer poses a significant threat to the health of women worldwide. Cervical cytopathology screening is an effective method for diagnosing cervical cancer. However, manual screening is time-consuming and prone to errors. The advent of automatic Computer-Aided Diagnosis (CAD) systems based on deep learning addresses this problem, however training these models requires large amounts of labeled data, which may not always be available. This paper proposes a Few-Shot Learning (FSL) framework called FSL-DSC to perform cervical cell classification tasks on small dataset. FSL-DSC first proposes inner loop learning and outer loop learning for individual tasks and overall parameter updates respectively, then a depthwise separable module is designed to further enhance the performance of the model. Among three repeated experiments, the FSL-DSC framework achieves an average accuracy of 83.34%, which shows the effectiveness and potential of the proposed FSL-DSC in the field of cervical image classification and few-shot tasks. Xiangchen Wu, Changzhong Li, Hongzan Sun, Tao Jiang 0014, Marcin Grzegorzek, Chen Li 0022 |
IEEE Big Data | 2 |
| 2024 | Step detection in complex walking environments based on continuous wavelet transform
Xiangchen Wu, Xiaoqin Zeng, Xiaoxiang Lu, Keman Zhang |
Multim. Tools Appl. | 1 |
| 2023 | Fork Entropy: Assessing the Diversity of Open Source Software Projects' ForksabstractOn open source software (OSS) platforms such as GitHub, forking and accepting pull-requests is an important approach for OSS projects to receive contributions, especially from external contributors who cannot directly commit into the source repositories. Having a large number of forks is often considered as an indicator of a project being popular. While extensive studies have been conducted to understand the reasons of forking, communications between forks, features and impacts of forks, there are few quantitative measures that can provide a simple yet informative way to gain insights about an OSS project's forks besides their count. Inspired by studies on biodiversity and OSS team diversity, in this paper, we propose an approach to measure the diversity of an OSS project's forks (i.e., its fork population). We devise a novel fork entropy metric based on Rao's quadratic entropy to measure such diversity according to the forks' modifications to project files. With properties including symmetry, continuity, and monotonicity, the proposed fork entropy metric is effective in quantifying the diversity of a project's fork population. To further examine the usefulness of the proposed metric, we conduct empirical studies with data retrieved from fifty projects on GitHub. We observe significant correlations between a project's fork entropy and different outcome variables including the project's external productivity measured by the number of external contributors' commits, acceptance rate of external contributors' pull-requests, and the number of reported bugs. We also observe significant interactions between fork entropy and other factors such as the number of forks. The results suggest that fork entropy effectively enriches our understanding of OSS projects' forks beyond the simple number of forks, and can potentially support further research and applications. Liang Wang 0006, Xiangchen Wu, Baihui Sang, Jierui Zhang, XianPing Tao |
ASE | 3 |
| 2022 | Auto-Encoding GAN for Reducing Mode Collapse and Enhancing Feature RepresentationabstractGenerative Adversarial Nets (GAN) has been a popular research topic in processing of images, speech, texts, and videos, and many other fields.However, GAN still has some drawbacks such as unstable training and mode collapse.To address these challenges, this paper proposes an auto-encoding GAN, which is composed of a set of generators, a discriminator, an encoder and a decoder.A set of generators is responsible for learning different modes, accelerating the convergence of the model and preventing model collapse.The discriminator is used to distinguish between real samples and generated ones.In order to improve feature representation of the encoder and prevent multiple generators from covering a certain mode, an approach consisting of three phases is proposed accordingly.First, a clustering algorithm is presented to perceive the distribution of real and generated samples.Then, cluster center matching is utilized to keep consistency of the distribution of real and generated samples.Finally, the encoder and decoder are jointly optimized by the generated and real samples.Therefore, the encoder can map the generated and real samples to the embedding space so as to encode distinguishable features, and the decoder can distinguish from which generator the generated samples come and from which mode the real samples come.Experiments are conducted on image datasets to verify effectiveness of the auto-encoding GAN for reducing mode collapse and enhancing feature representation. Xiaoxiang Lu, Yang Zou 0001, Xiaoqin Zeng, Xiangchen Wu, Pengfei Qiu |
SEKE | 4 |
| 2022 | CVM-Cervix: A hybrid cervical Pap-smear image classification framework using CNN, visual transformer and multilayer perceptron
Wanli Liu, Chen Li 0022, Ning Xu 0012, Tao Jiang 0014, Md Mamunur Rahaman, Hongzan Sun, Xiangchen Wu, Changhao Sun, Yu-Dong Yao, Marcin Grzegorzek |
Pattern Recognit. | 7 |