EDBT 2026 Demo / reviewers in the wild / expert
Yuzhou Cao
dblp:256/5052
· DBLP profile ↗
18ranked-venue papers
8as first author
17since 2021 · last 2025
0000-0002-3034-7967ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-author · 15 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Establishing Linear Surrogate Regret Bounds for Convex Smooth Losses via Convolutional Fenchel-Young LossesabstractSurrogate regret bounds, also known as excess risk bounds,
bridge the gap between the convergence rates of surrogate and target losses. The regret transfer is lossless if the surrogate regret bound is linear.
While convex smooth surrogate losses are appealing in particular due to the efficient estimation and optimization,
the existence of a trade-off between the loss smoothness and linear regret bound has been believed in the community.
Under this scenario, the better optimization and estimation properties of convex smooth surrogate losses may inevitably deteriorate after undergoing the regret transfer onto a target loss.
We overcome this dilemma for arbitrary discrete target losses
by constructing a convex smooth surrogate loss,
which entails a linear surrogate regret bound composed with a tailored
prediction link.
The construction is based on Fenchel--Young losses generated by the *convolutional negentropy*,
which are equivalent to the infimal convolution of a generalized negentropy and the target Bayes risk.
Consequently, the infimal convolution enables us to derive a smooth loss while maintaining the surrogate regret bound linear.
We additionally benefit from the infimal convolution to have a consistent estimator of the underlying class probability.
Our results are overall a novel demonstration of how convex analysis penetrates into optimization and statistical efficiency in risk minimization. Yuzhou Cao, Lei Feng 0006, Bo An 0001 |
NeurIPS | 1 |
| 2024 | Consistent Hierarchical Classification with A Generalized MetricabstractIn multi-class hierarchical classification, a natural evaluation metric is the tree distance loss that takes the value of two labels’ distance on the pre-defined tree hierarchy. This metric is motivated by that its Bayes optimal solution is the deepest label on the tree whose induced superclass (subtree rooted at it) includes the true label with probability at least $\frac{1}{2}$. However, it can hardly handle the risk sensitivity of different tasks since its accuracy requirement for induced superclasses is fixed at $\frac{1}{2}$. In this paper, we first introduce a new evaluation metric that generalizes the tree distance loss, whose solution’s accuracy constraint $\frac{1+c}{2}$ can be controlled by a penalty value $c$ tailored for different tasks: a higher c indicates the emphasis on prediction’s accuracy and a lower one indicates that on specificity. Then, we propose a novel class of consistent surrogate losses based on an intuitive presentation of our generalized metric and its regret, which can be compatible with various binary losses. Finally, we theoretically derive the regret transfer bounds for our proposed surrogates and empirically validate their usefulness on benchmark datasets. Yuzhou Cao, Lei Feng 0006, Bo An 0001 |
AISTATS | 1 |
| 2024 | Mitigating Underfitting in Learning to Defer with Consistent LossesabstractLearning to defer (L2D) allows the classifier to defer its prediction to an expert for safer predictions, by balancing the system’s accuracy and extra costs incurred by consulting the expert. Various loss functions have been proposed for L2D, but they were shown to cause the underfitting of trained classifiers when extra consulting costs exist, resulting in degraded performance. In this paper, we propose a novel loss formulation that can mitigate the underfitting issue while remaining the statistical consistency. We first show that our formulation can avoid a common characteristic shared by most existing losses, which has been shown to be a cause of underfitting, and show that it can be combined with the representative losses for L2D to enhance their performance and yield consistent losses. We further study the regret transfer bounds of the proposed losses and experimentally validate its improvements over existing methods. Shuqi Liu 0002, Yuzhou Cao, Qiaozhen Zhang, Lei Feng 0006, Bo An 0001 |
AISTATS | 2 |
| 2024 | Consistent Multi-Class Classification from Multiple Unlabeled DatasetsabstractWeakly supervised learning aims to construct effective predictive models from imperfectly labeled data. The recent trend of weakly supervised learning has focused on how to learn an accurate classifier from completely unlabeled data, given little supervised information such as class priors. In this paper, we consider a newly proposed weakly supervised learning problem called multi-class classification from multiple unlabeled datasets, where only multiple sets of unlabeled data and their class priors (i.e., the proportions of each class) are provided for training the classifier. To solve this problem, we first propose a classifier-consistent method (CCM) based on a probability transition matrix. However, CCM cannot guarantee risk consistency and lacks of purified supervision information during training. Therefore, we further propose a risk-consistent method (RCM) that progressively purifies supervision information during training by importance weighting. We provide comprehensive theoretical analyses for our methods to demonstrate the statistical consistency. Experimental results on multiple benchmark datasets and various prior matrices demonstrate the superiority of our proposed methods. Zixi Wei, Senlin Shu, Yuzhou Cao, Hongxin Wei, Bo An 0001, Lei Feng 0006 |
ICLR | 3 |
| 2024 | On the Vulnerability of Adversarially Trained Models Against Two-faced AttacksabstractAdversarial robustness is an important standard for measuring the quality of learned models, and adversarial training is an effective strategy for improving the adversarial robustness of models. In this paper, we disclose that adversarially trained models are vulnerable to two-faced attacks, where slight perturbations in input features are crafted to make the model exhibit a false sense of robustness in the verification phase. Such a threat is significantly important as it can mislead our evaluation of the adversarial robustness of models, which could cause unpredictable security issues when deploying substandard models in reality. More seriously, this threat seems to be pervasive and tricky: we find that many types of models suffer from this threat, and models with higher adversarial robustness tend to be more vulnerable. Furthermore, we provide the first attempt to formulate this threat, disclose its relationships with adversarial risk, and try to circumvent it via a simple countermeasure. These findings serve as a crucial reminder for practitioners to exercise caution in the verification phase, urging them to refrain from blindly trusting the exhibited adversarial robustness of models. Lue Tao, Yuzhou Cao, Tao Xiang 0001, Bo An 0001, Lei Feng 0006 |
ICLR | 3 |
| 2024 | Exploiting Human-AI Dependence for Learning to DeferabstractThe learning to defer (L2D) framework allows models to defer their decisions to human experts. For L2D, the Bayes optimality is the basic requirement of theoretical guarantees for the design of consistent surrogate loss functions, which requires the minimizer (i.e., learned classifier) by the surrogate loss to be the Bayes optimality. However, we find that the original form of Bayes optimality fails to consider the dependence between the model and the expert, and such a dependence could be further exploited to design a better consistent loss for L2D. In this paper, we provide a new formulation for the Bayes optimality called dependent Bayes optimality, which reveals the dependence pattern in determining whether to defer. Based on the dependent Bayes optimality, we further present a deferral principle for L2D. Following the guidance of the deferral principle, we propose a novel consistent surrogate loss. Comprehensive experimental results on both synthetic and real-world datasets demonstrate the superiority of our proposed method. Zixi Wei, Yuzhou Cao, Lei Feng 0006 |
ICML | 2 |
| 2023 | Consistent Complementary-Label Learning via Order-Preserving LossesabstractIn contrast to ordinary supervised classification tasks that require massive data with high-quality labels, complementary-label learning (CLL) deals with the weakly-supervised learning scenario where each instance is equipped with a complementary label, which specifies a class the instance does not belong to. However, most of the existing statistically consistent CLL methods suffer from overfitting intrinsically, due to the negative empirical risk issue. In this paper, we aim to propose overfitting-resistant and theoretically grounded methods for CLL. Considering the unique property of the distribution of complementarily labeled samples, we provide a risk estimator via order-preserving losses, which are naturally non-negative and thus can avoid overfitting caused by negative terms in risk estimators. Moreover, we provide classifier-consistency analysis and statistical guarantee for this estimator. Furthermore, we provide a weighed version of the proposed risk estimator to further enhance its generalization ability and prove its statistical consistency. Experiments on benchmark datasets demonstrate the effectiveness of our proposed methods. Shuqi Liu 0002, Yuzhou Cao, Qiaozhen Zhang, Lei Feng 0006, Bo An 0001 |
AISTATS | 2 |
| 2023 | Weakly Supervised Regression with Interval TargetsabstractThis paper investigates an interesting weakly supervised regression setting called regression with interval targets (RIT). Although some of the previous methods on relevant regression settings can be adapted to RIT, they are not statistically consistent, and thus their empirical performance is not guaranteed. In this paper, we provide a thorough study on RIT. First, we proposed a novel statistical model to describe the data generation process for RIT and demonstrate its validity. Second, we analyze a simple selecting method for RIT, which selects a particular value in the interval as the target value to train the model. Third, we propose a statistically consistent limiting method for RIT to train the model by limiting the predictions to the interval. We further derive an estimation error bound for our limiting method. Finally, extensive experiments on various datasets demonstrate the effectiveness of our proposed method. Xin Cheng 0007, Yuzhou Cao, Ximing Li 0002, Bo An 0001, Lei Feng 0006 |
ICML | 2 |
| 2023 | In Defense of Softmax Parametrization for Calibrated and Consistent Learning to DeferabstractEnabling machine learning classifiers to defer their decision to a downstream expert when the expert is more accurate will ensure improved safety and performance. This objective can be achieved with the learning-to-defer framework which aims to jointly learn how to classify and how to defer to the expert. In recent studies, it has been theoretically shown that popular estimators for learning to defer parameterized with softmax provide unbounded estimates for the likelihood of deferring which makes them uncalibrated. However, it remains unknown whether this is due to the widely used softmax parameterization and if we can find a softmax-based estimator that is both statistically consistent and possesses a valid probability estimator. In this work, we first show that the cause of the miscalibrated and unbounded estimator in prior literature is due to the symmetric nature of the surrogate losses used and not due to softmax. We then propose a novel statistically consistent asymmetric softmax-based surrogate loss that can produce valid estimates without the issue of unboundedness. We further analyze the non-asymptotic properties of our proposed method and empirically validate its performance and calibration on benchmark datasets. Yuzhou Cao, Hussein Mozannar, Lei Feng 0006, Hongxin Wei, Bo An 0001 |
NeurIPS | 1 |
| 2023 | Regression with Cost-based RejectionabstractLearning with rejection is an important framework that can refrain from making predictions to avoid critical mispredictions by balancing between prediction and rejection. Previous studies on cost-based rejection only focused on the classification setting, which cannot handle the continuous and infinite target space in the regression setting. In this paper, we investigate a novel regression problem called regression with cost-based rejection, where the model can reject to make predictions on some examples given certain rejection costs. To solve this problem, we first formulate the expected risk for this problem and then derive the Bayes optimal solution, which shows that the optimal model should reject to make predictions on the examples whose variance is larger than the rejection cost when the mean squared error is used as the evaluation metric. Furthermore, we propose to train the model by a surrogate loss function that considers rejection as binary classification and we provide conditions for the model consistency, which implies that the Bayes optimal solution can be recovered by our proposed surrogate loss. Extensive experiments demonstrate the effectiveness of our proposed method. Xin Cheng 0007, Yuzhou Cao, Haobo Wang 0001, Hongxin Wei, Bo An 0001, Lei Feng 0006 |
NeurIPS | 2 |
| 2023 | On the Importance of Feature Separability in Predicting Out-Of-Distribution ErrorabstractEstimating the generalization performance is practically challenging on out-of-distribution (OOD) data without ground-truth labels. While previous methods emphasize the connection between distribution difference and OOD accuracy, we show that a large domain gap not necessarily leads to a low test accuracy. In this paper, we investigate this problem from the perspective of feature separability empirically and theoretically. Specifically, we propose a dataset-level score based upon feature dispersion to estimate the test accuracy under distribution shift. Our method is inspired by desirable properties of features in representation learning: high inter-class dispersion and high intra-class compactness. Our analysis shows that inter-class dispersion is strongly correlated with the model accuracy, while intra-class compactness does not reflect the generalization performance on OOD data. Extensive experiments demonstrate the superiority of our method in both prediction performance and computational efficiency. Renchunzi Xie, Hongxin Wei, Lei Feng 0006, Yuzhou Cao, Bo An 0001 |
NeurIPS | 4 |
| 2023 | Average Transmission Rate and Energy Efficiency Optimization in UAV-assisted IoTabstractInternet of Things (IoT) has gradually been applied to various fields, including industries and agriculture, and plays an increasingly important role in society. However, the limited coverage of terrestrial IoT network restricts the communication performance of IoT devices, making the network inefficient. Unmanned aerial vehicles (UAVs) have the potential to be an efficient solution to improve the communication efficiency of the terrestrial IoT devices. Thus, we formulate a UAV-assisted data collection multi-objective optimization problem (UAVDCMOP) to jointly maximize the average transmission rate, minimize the total time of UAVs, and minimize the average energy consumed by UAVs via determining the optimal positions of UAVs. To this end, we propose an improved multi-objective grey wolf-based optimization (IMOGWO) algorithm with chaotic mapping initialization operator and inversion opposition generation operator, making it suitable for optimizing the formulated UAVDCMOP. Simulation results demonstrate that the proposed approach contributes to enhance the system average transmission rate and energy efficiency, and it has superior performance compared to other approaches. Yuzhou Cao, Aimin Wang 0001, Geng Sun 0001, Lingling Liu |
WCNC | 1 |
| 2023 | Multiple-Instance Learning From Unlabeled Bags With Pairwise SimilarityabstractInmultiple-instance learning(MIL), each training example is represented by a bag of instances. A training bag is either negative if it contains no positive instances or positive if it has at least one positive instance. Previous MIL methods generally assume that training bags are fully labeled. However, the exact labels of training examples may not be accessible, due to security, confidentiality, and privacy concerns. Fortunately, it could be easier for us to access the pairwise similarity between two bags (indicating whether two bags share the same label or not) and unlabeled bags, as we do not need to know the underlying label of each bag. In this paper, we provide the first attempt to investigate MIL from only similar-dissimilar-unlabeled bags. To solve this new MIL problem, we first propose a strong baseline method that trains an instance-level classifier by employing an unlabeled-unlabeled learning strategy. Then, we also propose to train a bag-level classifier based on a convex formulation and theoretically derive a generalization error bound for this method. Comprehensive experimental results show that our instance-level classifier works well, while our bag-level classifier even has better performance. Lei Feng 0006, Senlin Shu, Yuzhou Cao, Lue Tao, Hongxin Wei, Tao Xiang 0001, Bo An 0001, Gang Niu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Generalizing Consistent Multi-Class Classification with Rejection to be Compatible with Arbitrary Lossesabstract\emph{Classification with rejection} (CwR) refrains from making a prediction to avoid critical misclassification when encountering test samples that are difficult to classify. Though previous methods for CwR have been provided with theoretical guarantees, they are only compatible with certain loss functions, making them not flexible enough when the loss needs to be changed with the dataset in practice. In this paper, we derive a novel formulation for CwR that can be equipped with arbitrary loss functions while maintaining the theoretical guarantees. First, we show that $K$-class CwR is equivalent to a $(K\!+\!1)$-class classification problem on the original data distribution with an augmented class, and propose an empirical risk minimization formulation to solve this problem with an estimation error bound. Then, we find necessary and sufficient conditions for the learning \emph{consistency} of the surrogates constructed on our proposed formulation equipped with any classification-calibrated multi-class losses, where consistency means the surrogate risk minimization implies the target risk minimization for CwR. Finally, experiments on benchmark datasets validate the effectiveness of our proposed method. Yuzhou Cao, Tianchi Cai, Lei Feng 0006, Lihong Gu, Jinjie Gu, Bo An 0001, Gang Niu 0001, Masashi Sugiyama |
NeurIPS | 1 |
| 2022 | Multi-complementary and unlabeled learning for arbitrary losses and models
Yuzhou Cao, Shuqi Liu 0002, Yitian Xu |
Pattern Recognit. | 1 |
| 2021 | Learning from Similarity-Confidence DataabstractWeakly supervised learning has drawn considerable attention recently to reduce the expensive time and labor consumption of labeling massive data. In this paper, we investigate a novel weakly supervised learning problem of learning from similarity-confidence (Sconf) data, where only unlabeled data pairs equipped with confidence that illustrates their degree of similarity (two examples are similar if they belong to the same class) are needed for training a discriminative binary classifier. We propose an unbiased estimator of the classification risk that can be calculated from only Sconf data and show that the estimation error bound achieves the optimal convergence rate. To alleviate potential overfitting when flexible models are used, we further employ a risk correction scheme on the proposed risk estimator. Experimental results demonstrate the effectiveness of the proposed methods. Yuzhou Cao, Lei Feng 0006, Yitian Xu, Bo An 0001, Gang Niu 0001, Masashi Sugiyama |
ICML | 1 |
| 2021 | Multiple-Instance Learning from Similar and Dissimilar BagsabstractMultiple-instance learning (MIL) is an important weakly supervised binary classification problem, where training instances are arranged in bags, and each bag is assigned a positive or negative label. Most of the previous studies for MIL assume that training bags are fully labeled. However, in some real-world scenarios, it could be difficult to collect fully labeled bags, due to the expensive time and labor consumption of the labeling task. Fortunately, it could be much easier for us to collect similar and dissimilar bags (indicating whether two bags share the same label or not), because we do not need to figure out the underlying label of each bag in this case. Therefore, in this paper, we for the first time investigate MIL from only similar and dissimilar bags. To solve this new MIL problem, we propose a convex formulation to train a bag-level classifier based on empirical risk minimization and theoretically derive a generalization error bound. In addition, we also propose a strong baseline for this new MIL problem, which aims to train an instance-level classifier by minimizing the instance-level empirical risk. Extensive experimental results clearly demonstrate that our proposed baseline works well, while our proposed convex formulation is even better. Lei Feng 0006, Senlin Shu, Yuzhou Cao, Lue Tao, Hongxin Wei, Tao Xiang 0001, Bo An 0001, Gang Niu 0001 |
KDD | 3 |
| 2020 | Multi-variable estimation-based safe screening rule for small sphere and large margin support vector machine
Yuzhou Cao, Yitian Xu, Junling Du |
Knowl. Based Syst. | 1 |