EDBT 2026 Demo / reviewers in the wild / expert
Wenjun Zhang 0012
dblp:46/3359-12
· DBLP profile ↗
20ranked-venue papers
5as first author
20since 2021 · last 2025
0000-0002-7269-0376ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TLLC: Transfer Learning-based Label Completion for CrowdsourcingabstractLabel completion serves as a preprocessing approach to handling the sparse crowdsourced label matrix problem, significantly boosting the effectiveness of the downstream label aggregation. In recent advances, worker modeling has been proved to be a powerful strategy to further improve the performance of label completion. However, in real-world scenarios, workers typically annotate only a few instances, leading to insufficient worker modeling and thus limiting the improvement of label completion. To address this issue, we propose a novel transfer learning-based label completion (TLLC) method. Specifically, we first identify all high-confidence instances from the whole crowdsourced data as a source domain and use it to pretrain a Siamese network. The abundant annotated instances in the source domain provide essential knowledge for worker modeling. Then, we transfer the pretrained network to the target domain with the instances annotated by each worker separately, ensuring worker modeling captures unique characteristics of each worker. Finally, we leverage the new embeddings learned by the transferred network to complete each worker’s missing labels. Extensive experiments on several widely used real-world datasets demonstrate the effectiveness of TLLC. Our codes and datasets are available at https://github.com/jiangliangxiao/TLLC. Wenjun Zhang 0012, Liangxiao Jiang, Chaoqun Li 0001 |
ICML | 1 |
| 2025 | Instance Correlation Graph-based Naive BayesabstractDue to its simplicity, effectiveness and robustness, naive Bayes (NB) has continued to be one of the top 10 data mining algorithms. To improve its performance, a large number of improved algorithms have been proposed in the last few decades. However, in addition to Gaussian naive Bayes (GNB), there is little work on numerical attributes. At the same time, none of them takes into account the correlations among instances. To fill this gap, we propose a novel algorithm called instance correlation graph-based naive Bayes (ICGNB). Specifically, it first uses original attributes to construct an instance correlation graph (ICG) to represent the correlations among instances. Then, it employs a variational graph auto-encoder (VGAE) to generate new attributes from the constructed ICG and uses them to augment original attributes. Finally, it weights each augmented attribute to alleviate the attribute redundancy and builds GNB on the weighted attributes. The experimental results on tens of datasets show that ICGNB significantly outperforms its deserved competitors.Our codes and datasets are available at https://github.com/jiangliangxiao/ICGNB. Liangxiao Jiang, Wenjun Zhang 0012, Liangjun Yu, Huan Zhang 0007 |
ICML | 3 |
| 2025 | Label Distribution Propagation-based Label Completion for CrowdsourcingabstractIn real-world crowdsourcing scenarios, most workers often annotate a few instances only, which results in a significantly sparse crowdsourced label matrix and subsequently harms the performance of label integration algorithms. Recent work called worker similarity-based label completion (WSLC) has been proven to be an effective algorithm to addressing this issue. However, WSLC considers solely the correlation of the labels annotated by different workers on per individual instance while totally ignoring the correlation of the labels annotated by different workers among similar instances. To fill this gap, we propose a novel label distribution propagation-based label completion (LDPLC) algorithm. At first, we use worker similarity weighted majority voting to initialize a label distribution for each missing label. Then, we design a label distribution propagation algorithm to enable each missing label of each instance to iteratively absorb its neighbors’ label distributions. Finally, we complete each missing label based on its converged label distribution. Experimental results on both real-world and simulated crowdsourced datasets show that LDPLC significantly outperforms WSLC in enhancing the performance of label integration algorithms. Our codes and datasets are available at https://github.com/jiangliangxiao/LDPLC. Liangxiao Jiang, Wenjun Zhang 0012, Chaoqun Li 0001 |
ICML | 3 |
| 2025 | ELDP: Enhanced Label Distribution Propagation for CrowdsourcingabstractIn crowdsourcing scenarios, we can obtain multiple noisy labels for an instance from crowd workers and then aggregate these labels to infer the unknown true label of this instance. Due to the lack of expertise of workers, obtained labels usually contain a degree of noise. Existing studies usually focus on the crowdsourcing scenarios with low noise ratios but rarely focus on the crowdsourcing scenarios with high noise ratios. In this paper, we focus on the crowdsourcing scenarios with high noise ratios and propose a novel label aggregation algorithm called enhanced label distribution propagation (ELDP). First, ELDP harnesses an internal worker weighting method to estimate the weights of workers and then performs the first label distribution enhancement. Then, for instances not covered in the first enhancement, ELDP performs the second enhancement using a class membership estimation method based on the intra-cluster distance. Finally, ELDP propagates enhanced label distributions from accurately enhanced instances to inaccurately enhanced instances. Experimental results on both simulated and real-world crowdsourced datasets show that ELDP significantly outperforms all the other state-of-the-art label aggregation algorithms. Wenjun Zhang 0012, Liangxiao Jiang, Chaoqun Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Worker Similarity-Based Label Completion for CrowdsourcingabstractIn real-world crowdsourcing scenarios, it is a common phenomenon that each worker only annotates a few instances, resulting in a significantly sparse crowdsourcing label matrix. Consequently, only a small number of workers influence the inferred integrated label of each instance, which may weaken the performance of label integration algorithms. To address this problem, we propose a novel label completion algorithm called Worker Similarity-based Label Completion (WSLC). WSLC is grounded on the assumption that workers with similar cognitive abilities will annotate similar labels on the same instances. Specifically, we first construct a data set for each worker that includes all instances annotated by this worker and learn a feature vector for each worker. Then, we define a metric based on cosine similarity to estimate worker similarity based on the learned feature vectors. Finally, we complete the labels for each worker on unannotated instances based on the worker similarity and the annotations of similar workers. The experimental results on one real-world and 34 simulated crowdsourced data sets consistently show that WSLC effectively addresses the problem of the sparse crowdsourcing label matrix and enhances the integration accuracies of label integration algorithms. Liangxiao Jiang, Wenjun Zhang 0012, Chaoqun Li 0001 |
IEEE Trans. Big Data | 3 |
| 2025 | Dual-View Learning from CrowdsabstractCrowdsourcing services provide a fast and cheap way to obtain substantial labeled data by employing crowd workers on the Internet. In crowdsourcing learning, two-stage methods have been widely used, which first infer the integrated label for each instance and then build a learning model using instances with their integrated labels. However, existing two-stage methods mainly focus on how to infer more accurate integrated labels, after that, most of them directly regard the integrated labels as class labels to build a learning model, which loses the detailed worker labeling information in multiple noisy labels and thus results in sub-optimal model accuracy. To solve this problem, in this study, we take the multiple noisy labels of each instance as its attribute value vector to construct another view in addition to the original attribute view, and propose a novel two-stage method called dual-view learning from crowds (DVLFC). In DVLFC, we first pick out workers with sufficient number of labels and augment the multiple noisy label set for each instance, then we build a supervised learning model in each view and at last we fuse their class-membership probabilities to get the final classification result. Extensive experiments on both real-world and artificial crowdsourced datasets prove the effectiveness of DVLFC. Huan Zhang 0007, Liangxiao Jiang, Wenjun Zhang 0012, Geoffrey I. Webb |
ACM Trans. Knowl. Discov. Data | 3 |
| 2025 | Label Consistency-Based Ground Truth Inference for CrowdsourcingabstractIn crowdsourcing scenarios, we can obtain each instance's multiple noisy labels from different crowd workers and then infer its unknown ground truth via a ground truth inference method. However, to the best of our knowledge, the existing ground truth inference methods always attempt to aggregate multiple noisy labels into a single consensus label as the ground truth. In this article, we aim to explore a new strategy, i.e., label selection, which directly selects the label of the highest quality worker as the ground truth. To this end, we propose a label consistency-based ground truth inference (LCGTI) method. In LCGTI, we argue that high-quality workers should have a low bias with other workers in labeling the same instances and a low variance with themselves in labeling similar instances. To estimate the bias, we calculate the label consistency of different workers on the same instances. To estimate the variance, we calculate the label consistency of the same worker on similar instances. Finally, we combine these two components to calculate the labeling quality of each worker on the inferred instance and perform label selection instead of label aggregation to achieve inference. The experimental results on 34 simulated and two real-world datasets show that LCGTI significantly outperforms all the other state-of-the-art label aggregation-based ground truth inference methods. Liangxiao Jiang, Wenjun Zhang 0012 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Probabilistic Matrix Factorization-based Three-stage Label Completion for CrowdsourcingabstractCrowdsourcing provides a cost-effective solution to the problem of obtaining large annotated datasets. In real-world crowdsourcing scenarios, most workers often annotate a few instances only, which results in a significantly sparse crowdsourcing label matrix and subsequently harms the performance of label integration algorithms. Probabilistic matrix factorization (PMF) has been proven to be an effective method for crowdsourcing label completion. However, its low-quality input and output labels limit its performance. To improve its performance, this paper proposes a PMF-based three-stage label completion (PMF-TLC) method. In the first stage, we design a label confidence-based strategy to estimate the quality of each raw label of each worker. Then we flip those low-quality labels in the original crowdsourcing label matrix. In the second stage, we conduct PMF on the flipped label matrix and obtain the completed label matrix with soft labels. In the third stage, we design a between-class margin-based filter to delete those low-quality soft labels in the completed label matrix. Then we convert the remaining high-quality soft labels to hard (logic) labels and obtain the final processed label matrix. Extensive experimental results on real-world and simulated crowdsourced datasets show that PMF-TLC can significantly improve label integration algorithms' performance. Boyi Yang, Liangxiao Jiang, Wenjun Zhang 0012 |
ICDM | 3 |
| 2024 | IWBVT: Instance Weighting-based Bias-Variance Trade-off for CrowdsourcingabstractIn recent years, a large number of algorithms for label integration and noise correction have been proposed to infer the unknown true labels of instances in crowdsourcing. They have made great advances in improving the label quality of crowdsourced datasets. However, due to the presence of intractable instances, these algorithms are usually not as significant in improving the model quality as they are in improving the label quality. To improve the model quality, this paper proposes an instance weighting-based bias-variance trade-off (IWBVT) approach. IWBVT at first proposes a novel instance weighting method based on the complementary set and entropy, which mitigates the impact of intractable instances and thus makes the bias and variance of trained models closer to the unknown true results. Then, IWBVT performs probabilistic loss regressions based on the bias-variance decomposition, which achieves the bias-variance trade-off and thus reduces the generalization error of trained models. Experimental results indicate that IWBVT can serve as a universal post-processing approach to significantly improving the model quality of existing state-of-the-art label integration algorithms and noise correction algorithms. Wenjun Zhang 0012, Liangxiao Jiang, Chaoqun Li 0001 |
NeurIPS | 1 |
| 2024 | KFNN: K-Free Nearest Neighbor For CrowdsourcingabstractTo reduce annotation costs, it is common in crowdsourcing to collect only a few noisy labels from different crowd workers for each instance. However, the limited noisy labels restrict the performance of label integration algorithms in inferring the unknown true label for the instance. Recent works have shown that leveraging neighbor instances can help alleviate this problem. Yet, these works all assume that each instance has the same neighborhood size, which defies common sense. To address this gap, we propose a novel label integration algorithm called K-free nearest neighbor (KFNN). In KFNN, the neighborhood size of each instance is automatically determined based on its attributes and noisy labels. Specifically, KFNN initially estimates a Mahalanobis distance distribution from the attribute space to model the relationship between each instance and all classes. This distance distribution is then utilized to enhance the multiple noisy label distribution of each instance. Subsequently, a Kalman filter is designed to mitigate the impact of noise incurred by neighbor instances. Finally, KFNN determines the optimal neighborhood size by the max-margin learning. Extensive experimental results demonstrate that KFNN significantly outperforms all the other state-of-the-art algorithms and exhibits greater robustness in various crowdsourcing scenarios. Wenjun Zhang 0012, Liangxiao Jiang, Chaoqun Li 0001 |
NeurIPS | 1 |
| 2024 | Label Consistency-based Worker Filtering for CrowdsourcingabstractIn crowdsourcing scenarios, we can obtain multiple noisy labels from different crowd workers on the Internet for each instance and then infer its unknown true label via a label integration method. However, noisy labels often have a serious negative impact on label integration. In this case, most existing works always focus on designing more complex label integration methods to infer unknown true labels more accurately from multiple noisy labels, but little attention has been paid to another perspective, i.e., purifying noisy labels before label integration. In this paper, we aim to purify noisy labels for existing label integration methods and propose a label consistency-based worker filtering (LCWF) algorithm. In LCWF, we consider that if all low-quality workers are filtered out and only high-quality workers remain, the label consistency should be high. Therefore, we utilize label consistency to filter out low-quality workers. Firstly, we directly transform the worker filtering problem into a discrete optimization problem and utilize label consistency to define the fitness function for this problem. Then, we search for the optimal solution to this problem by a genetic algorithm. Finally, we filter out all labels from low-quality workers according to the optimal solution we obtained. Experimental results on simulated and real-world datasets demonstrate that LCWF can effectively purify noisy labels and improve the integration accuracy of existing label integration methods. Liangxiao Jiang, Chaoqun Li 0001, Wenjun Zhang 0012 |
UAI | 4 |
| 2024 | Learning from Crowds with Dual-View K-Nearest NeighborabstractIn crowdsourcing scenarios, we can obtain multiple noisy labels from different crowd workers for each instance and then infer its integrated label via label integration. To achieve better performance, some recently published label integration methods have attempted to exploit the multiple noisy labels of inferred instances’ nearest neighbors via the K-nearest neighbor (KNN) algorithm. However, the used KNN algorithm searches inferred instances’ nearest neighbors only relying on the defined distance functions in the original attribute view and totally ignoring the valuable information hidden in the multiple noisy labels, which limits their performance. Motivated by multi-view learning, we define the multiple noisy labels as another label view of instances and propose to search inferred instances’ nearest neighbors using the joint information from both the original attribute view and the multiple noisy label view. To this end, we propose a novel label integration method called dual-view K-nearest neighbor (DVKNN). In DVKNN, we first define a new distance function to search the K-nearest neighbors of an inferred instance. Then, we define a fine-grained weight for each noisy label from each neighbor. Finally, we perform weighted majority voting (WMV) on all these noisy labels to obtain the integrated label of the inferred instance. Extensive experiments validate the effectiveness and rationality of DVKNN. Liangxiao Jiang, Wenjun Zhang 0012 |
UAI | 4 |
| 2024 | FNNWV: farthest-nearest neighbor-based weighted voting for class-imbalanced crowdsourcing
Wenjun Zhang 0012, Liangxiao Jiang, Chaoqun Li 0001 |
Sci. China Inf. Sci. | 1 |
| 2024 | Label distribution similarity-based noise correction for crowdsourcing
Lijuan Ren, Liangxiao Jiang, Wenjun Zhang 0012, Chaoqun Li 0001 |
Frontiers Comput. Sci. | 3 |
| 2024 | Worker similarity-based noise correction for crowdsourcing
Yufei Hu, Liangxiao Jiang, Wenjun Zhang 0012 |
Inf. Syst. | 3 |
| 2024 | Weighted Adversarial Learning From CrowdsabstractCrowdsourcing services provide a fast and cheap way to annotate instances by employing crowd workers on the Internet. As a result, many learning from crowds (LFC) methods have been proposed in recent years. However, almost all these methods assume that all instances are benign, which makes them vulnerable to adversarial attacks. To improve the model's robustness to adversarial attacks, adversarial LFC (A-LFC) has attracted remarkable attention. A-LFC iteratively updates the true labels' estimations and the trained model by adversarial learning and uses the model's predictions in turn to help estimate the true labels. In A-LFC, to the best of our knowledge, the true labels' estimations and the model's predictions are inaccurate and the stopping condition for iterations is rough, which limit the model's performance. To further improve A-LFC, this paper proposes weighted A-LFC (WA-LFC). To reduce the impact of misinformation in the true labels' estimations and model's predictions, our method weights instances for adversarial learning and weights the model's predictions for estimating the true labels. Our method iteratively updates these weights and uses instance-weighted cross-entropy loss to decide when the iterative process should be stopped. Experiments on three real-world datasets show that our method substantially improves the trained model's performance. On average, the test accuracy of our method outperforms that of A-LFC by 10.06% and 27.55% in white-box and black-box attack settings, respectively. Liangxiao Jiang, Wenjun Zhang 0012, Chaoqun Li 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Instance Weighting-Based Noise Correction for Crowdsourcing
Liangxiao Jiang, Wenjun Zhang 0012 |
ICIC (4) | 3 |
| 2023 | Three-way decision-based noise correction for crowdsourcing
Liangxiao Jiang, Wenjun Zhang 0012, Chaoqun Li 0001 |
Int. J. Approx. Reason. | 3 |
| 2023 | Dual-View Noise Correction for CrowdsourcingabstractIn crowdsourcing scenarios, each instance obtains a multiple noisy label set from different crowd workers on the Internet and then gets its integrated label via label integration. Although the label integration algorithms are often effective, a certain level of noise still remains in the integrated labels. To reduce the impact of noise on label quality, many noise correction algorithms have been proposed in recent years. However, most of them are hard to fully utilize the joint information of the original attribute view and multiple noisy label view. Motivated by multi-view learning, in this article, we propose a dual-view noise correction (DVNC) algorithm. Benefiting from the complementary and consensus principle, DVNC can fully utilize the joint information of two views and thus enhance the effect of noise correction. In DVNC, we first construct the original attribute view and multiple noisy label view for each instance and, respectively, train a classifier on each view. Then, we use the trained two classifiers to filter each noise instance and thus obtain a clean set and noise set. Finally, we train a classifier on the clean set to correct each instance in the noise set via reclassifying it as the class with the maximum posterior probability. The experimental results on simulated and real-world data sets indicate that DVNC significantly outperforms all the other noise correction algorithms used for comparison. Liangxiao Jiang, Wenjun Zhang 0012 |
IEEE Internet Things J. | 3 |
| 2023 | Multi-View Attribute Weighted Naive BayesabstractNaive Bayes (NB) continues to be one of the top 10 data mining algorithms due to its simplicity, efficiency and efficacy. Numerous enhancements have been proposed to weaken its attribute conditional independence assumption. However, all of them only focus on the raw attribute view, which is hard to reflect all the data characteristics in real-world applications. To portray data characteristics more comprehensively, in this study, we construct two label views from the raw attributes and propose a novel model called multi-view attribute weighted naive Bayes (MAWNB). In MAWNB, we first build multiple super-parent one-dependence estimators (SPODEs) as well as random trees (RTs), then we utilize each of them to classify each training instance in turn and use all their predicted class labels to construct two label views. Next, to avoid attribute redundancy, we optimize the weight of each attribute value for each class by minimizing the negative conditional log-likelihood (CLL) in each view. Finally, the estimated class-membership probabilities by three views are fused to predict the class label for each test instance. Extensive experiments show that MAWNB significantly outperforms NB and all the other existing state-of-the-art competitors. Huan Zhang 0007, Liangxiao Jiang, Wenjun Zhang 0012, Chaoqun Li 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |