VLDB 2026 Research / reviewers in the wild / expert
Fuchao Yang
dblp:336/2241
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0002-5209-7153ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Learning paradigms · 65% Trustworthy machine learning · 17% Language models and text generation · 10% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 100% |
Topics — the 11 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Learning paradigms › weakly supervised learning
partial label learning |
2.4 | 3 | 2025 | Mixed Blessing: Class-Wise Embedding guided Instance-Dependent Partial Label Learning · KDD (1) 2025 Partial Label Clustering · IJCAI 2025 Partial Label Learning with Dissimilarity Propagation guided Candidate Label Shrinkage · NeurIPS 2023 |
Machine learning › Learning paradigms
weakly supervised learning |
2.4 | 3 | 2025 | Partial Label Clustering · IJCAI 2025 Noise Separation guided Candidate Label Reconstruction for Noisy Partial Label Learning · ICLR 2025 Partial Label Learning with Dissimilarity Propagation guided Candidate Label Shrinkage · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
1.7 | 2 | 2025 | Mixed Blessing: Class-Wise Embedding guided Instance-Dependent Partial Label Learning · KDD (1) 2025 Noise Separation guided Candidate Label Reconstruction for Noisy Partial Label Learning · ICLR 2025 |
Machine learning › Learning paradigms › weakly supervised learning › partial label learning
instance-dependent partial label learning |
0.9 | 1 | 2025 | Mixed Blessing: Class-Wise Embedding guided Instance-Dependent Partial Label Learning · KDD (1) 2025 |
Machine learning › Learning paradigms › weakly supervised learning › partial label learning
noisy partial label learning |
0.9 | 1 | 2025 | Noise Separation guided Candidate Label Reconstruction for Noisy Partial Label Learning · ICLR 2025 |
Data mining
clustering |
0.9 | 1 | 2025 | Partial Label Clustering · IJCAI 2025 |
Data mining › clustering
constrained clustering |
0.9 | 1 | 2025 | Partial Label Clustering · IJCAI 2025 |
Data mining › predictive modeling › classification
multi-label classification |
0.8 | 1 | 2024 | Noisy Label Removal for Partial Multi-Label Learning · KDD 2024 |
Data mining › predictive modeling › classification › multi-label classification
partial multi-label learning |
0.8 | 1 | 2024 | Noisy Label Removal for Partial Multi-Label Learning · KDD 2024 |
Machine learning › Representation and self-supervised learning › representation learning
embedding learning |
0.3 | 1 | 2025 | Mixed Blessing: Class-Wise Embedding guided Instance-Dependent Partial Label Learning · KDD (1) 2025 |
Machine learning › Learning paradigms › weakly supervised learning
label disambiguation |
0.3 | 1 | 2025 | Partial Label Clustering · IJCAI 2025 |
Methods — techniques the papers use, named apart from their topics
pairwise constraint propagation · 1.7dual-graph learning · 1.7adversarial prior · 1.7optical compression · 1.0prototype learning · 0.9noise separation · 0.9generalization error bound · 0.9contrastive learning · 0.9discrete optimization · 0.8alternative optimization · 0.8inexact augmented lagrange multiplier · 0.7constrained regression · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AgentOCR: Reimagining Agent History via Optical Self-CompressionabstractLang Feng, Fuchao Yang, Feng Chen, Xin Cheng, Haiyang Xu, Zhenglin Wan, Ming Yan, Bo An. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Lang Feng 0002, Fuchao Yang, Xin Cheng 0007, Haiyang Xu 0001, Zhenglin Wan, Ming Yan 0008, Bo An 0001 |
ACL (1) | 2 |
| 2025 | Noise Separation guided Candidate Label Reconstruction for Noisy Partial Label LearningabstractPartial label learning is a weakly supervised learning problem in which an instance is annotated with a set of candidate labels, among which only one is the correct label. However, in practice the correct label is not always in the candidate label set, leading to the noisy partial label learning (NPLL) problem. In this paper, we theoretically prove that the generalization error of the classifier constructed under NPLL paradigm is bounded by the noise rate and the average length of the candidate label set. Motivated by the theoretical guide, we propose a novel NPLL framework that can separate the noisy samples from the normal samples to reduce the noise rate and reconstruct the shorter candidate label sets for both of them. Extensive experiments on multiple benchmark datasets confirm the efficacy of the proposed method in addressing NPLL. For example, on CIFAR100 dataset with severe noise, our method improves the classification accuracy of the state-of-the-art one by 11.57%. The code is available at: https://github.com/pruirui/PLRC. Xiaorui Peng, Yuheng Jia, Fuchao Yang, Ran Wang 0001, Min-Ling Zhang |
ICLR | 3 |
| 2025 | Partial Label ClusteringabstractPartial label learning (PLL) is a significant weakly supervised learning framework, where each training example corresponds to a set of candidate labels and only one label is the ground-truth label. For the first time, this paper investigates the partial label clustering problem, which takes advantage of the limited available partial labels to improve the clustering performance. Specifically, we first construct a weight matrix of examples based on their relationships in the feature space and disambiguate the candidate labels to estimate the ground-truth label based on the weight matrix. Then, we construct a set of must-link and cannot-link constraints based on the disambiguation results. Moreover, we propagate the initial must-link and cannot-link constraints based on an adversarial prior promoted dual-graph learning approach. Finally, we integrate weight matrix construction, label disambiguation, and pairwise constraints propagation into a joint model to achieve mutual enhancement. We also theoretically prove that a better disambiguated label matrix can help improve clustering performance. Comprehensive experiments demonstrate our method realizes superior performance when comparing with state-of-the-art constrained clustering methods, and outperforms PLL and semi-supervised PLL methods when only limited samples are annotated. The code and appendix are publicly available at https://github.com/xyt-ml/PLC. Yutong Xie 0014, Fuchao Yang, Yuheng Jia |
IJCAI | 2 |
| 2025 | Mixed Blessing: Class-Wise Embedding guided Instance-Dependent Partial Label LearningabstractIn partial label learning (PLL), every sample is associated with a candidate label set comprising the ground-truth label and several noisy labels. The conventional PLL assumes the noisy labels are randomly generated (instance-independent), while in practical scenarios, the noisy labels are always instance-dependent and are highly related to the sample features, leading to the instance-dependent partial label learning (IDPLL) problem. Instance-dependent noisy label is a double-edged sword. On one side, it may promote model training as the noisy labels can depict the sample to some extent. On the other side, it brings high label ambiguity as the noisy labels are quite undistinguishable from the ground-truth label. To leverage the nuances of IDPLL effectively, for the first time we create class-wise embeddings for each sample, which allow us to explore the relationship of instance-dependent noisy labels, i.e., the class-wise embeddings in the candidate label set should have high similarity, while the class-wise embeddings between the candidate label set and the non-candidate label set should have high dissimilarity. Moreover, to reduce the high label ambiguity, we introduce the concept of class prototypes containing global feature information to disambiguate the candidate label set. Extensive experimental comparisons with twelve methods on six benchmark data sets, including four fine-grained data sets, demonstrate the effectiveness of the proposed method. The code implementation is publicly available at https://github.com/Yangfc-ML/CEL. Fuchao Yang, Jianhong Cheng, Hui Liu 0032, Yongqiang Dong, Yuheng Jia, Junhui Hou |
KDD (1) | 1 |
| 2024 | Noisy Label Removal for Partial Multi-Label LearningabstractThis paper addresses the problem of partial multi-label learning (PML), a challenging weakly supervised learning framework, where each sample is associated with a candidate label set comprising both ground-true labels and noisy labels. We theoretically reveal that an increased number of noisy labels in the candidate label set leads to an enlarged generalization error bound, consequently degrading the classification performance. Accordingly, the key to solving PML lies in accurately removing the noisy labels within the candidate label set. To achieve this objective, we leverage prior knowledge about the noisy labels in PML, which suggests that they only exist within the candidate label set and possess binary values. Specifically, we propose a constrained regression model to learn a PML classifier and select the noisy labels. The constraints of the model strictly enforce the location and value of the noisy labels. Simultaneously, the supervision information provided by the candidate label set is unreliable due to the presence of noisy labels. In contrast, the non-candidate labels of a sample precisely indicate the classes to which the sample does not belong. To aid in the selection of noisy labels, we construct a competitive classifier based on the non-candidate labels. The PML classifier and the competitive classifier form a competitive relationship, encouraging mutual learning. We formulate the proposed model as a discrete optimization problem to effectively remove the noisy labels, and we solve it using an alternative algorithm. Extensive experiments conducted on 6 real-world partial multi-label data sets and 7 synthetic data sets, employing various evaluation metrics, demonstrate that our method significantly outperforms state-of-the-art PML methods. The code implementation is publicly available at https://github.com/Yangfc-ML/NLR. Fuchao Yang, Yuheng Jia, Hui Liu 0032, Yongqiang Dong, Junhui Hou |
KDD | 1 |
| 2023 | Partial Label Learning with Dissimilarity Propagation guided Candidate Label ShrinkageabstractIn partial label learning (PLL), each sample is associated with a group of candidate labels, among which only one label is correct. The key of PLL is to disambiguate the candidate label set to find the ground-truth label. To this end, we first construct a constrained regression model to capture the confidence of the candidate labels, and multiply the label confidence matrix by its transpose to build a second-order similarity matrix, whose elements indicate the pairwise similarity relationships of samples globally. Then we develop a semantic dissimilarity matrix by considering the complement of the intersection of the candidate label set, and further propagate the initial dissimilarity relationships to the whole data set by leveraging the local geometric structure of samples. The similarity and dissimilarity matrices form an adversarial relationship, which is further utilized to shrink the solution space of the label confidence matrix and promote the dissimilarity matrix. We finally extend the proposed model to a kernel version to exploit the non-linear structure of samples and solve the proposed model by the inexact augmented Lagrange multiplier method. By exploiting the adversarial prior, the proposed method can significantly outperform
state-of-the-art PLL algorithms when evaluated on 10 artificial and 7 real-world partial label data sets. We also prove the effectiveness of our method with some theoretical guarantees. The code is publicly available at https://github.com/Yangfc-ML/DPCLS. Yuheng Jia, Fuchao Yang, Yongqiang Dong |
NeurIPS | 2 |