VLDB 2026 Research / reviewers in the wild / expert
Qingjie Wei
dblp:224/7198
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0007-7097-1071ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Participants sample generation based on association rules and data imputation in vertical federated learning
Xin Liu 0092, Hangxuan He, Weihao Tan, Qingjie Wei |
Expert Syst. Appl. | 6 |
| 2025 | FedPSG-CAG: Generating Participants Samples Based on Correlated Attributes Generation and Vertical Federated Imputation
Xin Liu 0092, Hangxuan He, Qingjie Wei |
IEEE Big Data | 5 |
| 2025 | Multi-Project Just-in-Time Software Defect Prediction Based on Multi-Task Learning for Mobile ApplicationsabstractIn the rapid development of mobile applications, frequent code commits pose significant challenges for quality assurance. Just-in- Time Software Defect Prediction (JIT -SDP) helps at the commit level but often struggles due to insufficient labeled data, particularly in newer applications. To address this issue, we introduce JMFM, a novel approach leveraging Multi-Task Learning (MTL), Fuzzy C-Means (FCM) clustering, and Multi-Head Attention (MHA) for JIT-SDP. JMFM integrates multiple projects for training under the MTL framework, treating each project as a distinct task and enabling cross-project learning. In JMFM, FCM clustering determines the membership of each data sample to various clusters, which is then used in the MHA module to compute the weights of similarity between samples. By integrating the weighted sum of other samples' data, each sample is augmented with additional information for shared learning. Simultaneously, each project is trained in a task-specific layer to retain its unique features. We calculate the joint loss by giving greater weight to projects with fewer samples to ensure that they are not overshadowed by larger projects. Experiments on 15 Android mobile applications show that JMFM outperforms existing models on metrics such as Fl, MCC and AUC especially for projects with scarce data. Yuxin Ke, Qingjie Wei |
ICST | 4 |
| 2024 | Multi-Party Federated Recommendation Based on Semi-Supervised LearningabstractLeveraging multi-party data to provide recommendations remains a challenge, particularly when the party in need of recommendation services possesses only positive samples while other parties just have unlabeled data. To address UDD-PU learning problem, this paper proposes an algorithm VFPU, Vertical Federated learning with Positive and Unlabeled data. VFPU conducts random sampling repeatedly from the multi-party unlabeled data, treating sampled data as negative ones. It hence forms multiple training datasets with balanced positive and negative samples, and multiple testing datasets with those unsampled data. For each training dataset, VFPU trains a base estimator adapted for the vertical federated learning framework iteratively. We use the trained base estimator to generate forecast scores for each sample in the testing dataset. Based on the sum of scores and their frequency of occurrence in the testing datasets, we calculate the probability of being positive for each unlabeled sample. Those with top probabilities are regarded as reliable positive samples. They are then added to the positive samples and subsequently removed from the unlabeled data. This process of sampling, training, and selecting positive samples is iterated repeatedly. Experimental results demonstrated that VFPU performed comparably to its non-federated counterparts and outperformed other federated semi-supervised learning methods. Xin Liu 0092, Jiuluan Lv, Qingjie Wei, Hangxuan He |
IEEE Trans. Big Data | 4 |
| 2021 | A K-means Improved CTGAN Oversampling Method for Data Imbalance ProblemabstractCTGAN is a tabular data synthesis method for privacy preservation, which is used in this paper for data imbalance problem. This paper proposes a method for dealing with imbalanced data sets that combines K-means clustering and CTGAN to address the imbalanced distribution of minority class examples that result from oversampling with CTGAN. By conducting experiments with the LightGBM algorithm on home loan and online shopping datasets, it is demonstrated that the CTGAN method achieves superior learning results in f1-score and G-mean metrics compared to the interpolation-based oversampling technique represented by SMOTE. The preceding results indicate that by applying the method described in this paper to handle an imbalanced dataset, one can obtain a dataset with more examples, a more uniform distribution, and less overfitting while still satisfying the original dataset's probability distribution. Chunsheng An, Jingtong Sun, Qingjie Wei |
QRS | 4 |