Nan Lu 0001

dblp:14/4965-1 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0003-2233-3984ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Trustworthy machine learning · 24% Learning theory · 21% Transfer learning and domain adaptation · 17%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
distribution shift
1.122023
Generalizing Importance Weighting to A Universal Solver for Distribution Shift Problems · NeurIPS 2023
Rethinking Importance Weighting for Deep Learning under Distribution Shift · NeurIPS 2020
Machine learning › Transfer learning and domain adaptation › instance weighting
importance weighting
1.122023
Generalizing Importance Weighting to A Universal Solver for Distribution Shift Problems · NeurIPS 2023
Rethinking Importance Weighting for Deep Learning under Distribution Shift · NeurIPS 2020
Machine learning › Learning paradigms
weakly supervised learning
1.022021
Binary Classification from Multiple Unlabeled Datasets via Surrogate Set Classification · ICML 2021
Pointwise Binary Classification with Pairwise Confidence Comparisons · ICML 2021
Machine learning › Learning theory › classification
binary classification
0.922021
Binary Classification from Multiple Unlabeled Datasets via Surrogate Set Classification · ICML 2021
On the Minimal Supervision for Training Any Binary Classifier from Only Unlabeled Data · ICLR (Poster) 2019
Machine learning › Efficient and distributed learning
federated learning
0.612022
Federated Learning from Only Unlabeled Data with Class-conditional-sharing Clients · ICLR 2022
Machine learning › Trustworthy machine learning
pairwise classification
0.512021
Pointwise Binary Classification with Pairwise Confidence Comparisons · ICML 2021
Machine learning › Learning theory › statistical estimation › risk estimation
unbiased risk estimator
0.512021
Pointwise Binary Classification with Pairwise Confidence Comparisons · ICML 2021
Machine learning › Representation and self-supervised learning
unlabeled data
0.412019
On the Minimal Supervision for Training Any Binary Classifier from Only Unlabeled Data · ICLR (Poster) 2019

Methods — techniques the papers use, named apart from their topics

risk decomposition · 0.7one-class support vector machine · 0.7surrogate set classification · 0.5noisy label learning · 0.5linear-fractional transformation · 0.5correction function · 0.5consistency regularization · 0.5importance weighting · 0.4feature extraction · 0.4end-to-end training · 0.4
YearPublicationVenuePosition
2025 Learning from Ambiguous Data with Hard Labels
abstract
Real-world data often contains intrinsic ambiguity that the common single-hard-label annotation paradigm ignores. Standard training using ambiguous data with these hard labels may produce overly confident models and thus leading to poor generalization. In this paper, we propose a novel framework called Quantized Label Learning (QLL) to alleviate this issue. First, we formulate QLL as learning from (very) ambiguous data with hard labels: ideally, each ambiguous instance should be associated with a ground-truth soft-label distribution describing its corresponding probabilistic weight in each class, however, this is usually not accessible; in practice, we can only observe a quantized label, i.e., a hard label sampled (quantized) from the corresponding ground-truth soft-label distribution, of each instance, which can be seen as a biased approximation of the ground-truth soft-label. Second, we propose a Class-wise Positive-Unlabeled (CPU) risk estimator that allows us to train accurate classifiers from only ambiguous data with quantized labels. Third, to simulate ambiguous datasets with quantized labels in the real world, we design a mixing-based ambiguous data generation procedure for empirical evaluation. Experiments demonstrate that our CPU method can significantly improve model generalization performance and outperform the baselines.
Zeke Xie, Nan Lu 0001, Lichen Bai, Shuo Yang 0006, Mingming Sun 0001, Ping Li 0001
ICASSP3
2023 Generalizing Importance Weighting to A Universal Solver for Distribution Shift Problems
abstract
Distribution shift (DS) may have two levels: the distribution itself changes, and the support (i.e., the set where the probability density is non-zero) also changes. When considering the support change between the training and test distributions, there can be four cases: (i) they exactly match; (ii) the training support is wider (and thus covers the test support); (iii) the test support is wider; (iv) they partially overlap. Existing methods are good at cases (i) and (ii), while cases (iii) and (iv) are more common nowadays but still under-explored. In this paper, we generalize importance weighting (IW), a golden solver for cases (i) and (ii), to a universal solver for all cases. Specifically, we first investigate why IW might fail in cases (iii) and (iv); based on the findings, we propose generalized IW (GIW) that could handle cases (iii) and (iv) and would reduce to IW in cases (i) and (ii). In GIW, the test support is split into an in-training (IT) part and an out-of-training (OOT) part, and the expected risk is decomposed into a weighted classification term over the IT part and a standard classification term over the OOT part, which guarantees the risk consistency of GIW. Then, the implementation of GIW consists of three components: (a) the split of validation data is carried out by the one-class support vector machine, (b) the first term of the empirical risk can be handled by any IW algorithm given training data and IT validation data, and (c) the second term just involves OOT validation data. Experiments demonstrate that GIW is a universal solver for DS problems, outperforming IW methods in cases (iii) and (iv).
Tongtong Fang, Nan Lu 0001, Gang Niu 0001, Masashi Sugiyama
NeurIPS2
2022 Multi-class Classification from Multiple Unlabeled Datasets with Partial Risk Regularization
Yuting Tang, Nan Lu 0001, Masashi Sugiyama
ACML2
2022 Federated Learning from Only Unlabeled Data with Class-conditional-sharing Clients
Nan Lu 0001, Zhao Wang 0006, Gang Niu 0001, Qi Dou 0001, Masashi Sugiyama
ICLR1
2021 Pointwise Binary Classification with Pairwise Confidence Comparisons
abstract
To alleviate the data requirement for training effective binary classifiers in binary classification, many weakly supervised learning settings have been proposed. Among them, some consider using pairwise but not pointwise labels, when pointwise labels are not accessible due to privacy, confidentiality, or security reasons. However, as a pairwise label denotes whether or not two data points share a pointwise label, it cannot be easily collected if either point is equally likely to be positive or negative. Thus, in this paper, we propose a novel setting called pairwise comparison (Pcomp) classification, where we have only pairs of unlabeled data that we know one is more likely to be positive than the other. Firstly, we give a Pcomp data generation process, derive an unbiased risk estimator (URE) with theoretical guarantee, and further improve URE using correction functions. Secondly, we link Pcomp classification to noisy-label learning to develop a progressive URE and improve it by imposing consistency regularization. Finally, we demonstrate by experiments the effectiveness of our methods, which suggests Pcomp is a valuable and practically useful type of pairwise supervision besides the pairwise label.
Lei Feng 0006, Senlin Shu, Nan Lu 0001, Bo Han 0003, Miao Xu 0001, Gang Niu 0001, Bo An 0001, Masashi Sugiyama
ICML3
2021 Binary Classification from Multiple Unlabeled Datasets via Surrogate Set Classification
abstract
To cope with high annotation costs, training a classifier only from weakly supervised data has attracted a great deal of attention these days. Among various approaches, strengthening supervision from completely unsupervised classification is a promising direction, which typically employs class priors as the only supervision and trains a binary classifier from unlabeled (U) datasets. While existing risk-consistent methods are theoretically grounded with high flexibility, they can learn only from two U sets. In this paper, we propose a new approach for binary classification from $m$ U-sets for $m\ge2$. Our key idea is to consider an auxiliary classification task called surrogate set classification (SSC), which is aimed at predicting from which U set each observed sample is drawn. SSC can be solved by a standard (multi-class) classification method, and we use the SSC solution to obtain the final binary classifier through a certain linear-fractional transformation. We built our method in a flexible and efficient end-to-end deep learning framework and prove it to be classifier-consistent. Through experiments, we demonstrate the superiority of our proposed method over state-of-the-art methods.
Nan Lu 0001, Shida Lei, Gang Niu 0001, Issei Sato, Masashi Sugiyama
ICML1
2020 A One-step Approach to Covariate Shift Adaptation
abstract
A default assumption in many machine learning scenarios is that the training and test samples are drawn from the same probability distribution. However, such an assumption is often violated in the real world due to non-stationarity of the environment or bias in sample selection. In this work, we consider a prevalent setting called covariate shift, where the input distribution differs between the training and test stages while the conditional distribution of the output given the input remains unchanged. Most of the existing methods for covariate shift adaptation are two-step approaches, which first calculate the importance weights and then conduct importance-weighted empirical risk minimization. In this paper, we propose a novel one-step approach that jointly learns the predictive model and the associated weights in one optimization by minimizing an upper bound of the test risk. We theoretically analyze the proposed method and provide a generalization error bound. We also empirically demonstrate the effectiveness of the proposed method.
Ikko Yamane, Nan Lu 0001, Masashi Sugiyama
ACML3
2020 Mitigating Overfitting in Supervised Classification from Two Unlabeled Datasets: A Consistent Risk Correction Approach
abstract
The recently proposed unlabeled-unlabeled (UU) classification method allows us to train a binary classifier only from two unlabeled datasets with different class priors. Since this method is based on the empirical risk minimization, it works as if it is a supervised classification method, compatible with any model and optimizer. However, this method sometimes suffers from severe overfitting, which we would like to prevent in this paper. Our empirical finding in applying the original UU method is that overfitting often co-occurs with the empirical risk going negative, which is not legitimate. Therefore, we propose to wrap the terms that cause a negative empirical risk by certain correction functions. Then, we prove the consistency of the corrected risk estimator and derive an estimation error bound for the corrected risk minimizer. Experiments show that our proposal can successfully mitigate overfitting of the UU method and significantly improve the classification accuracy.
Nan Lu 0001, Gang Niu 0001, Masashi Sugiyama
AISTATS1
2020 Rethinking Importance Weighting for Deep Learning under Distribution Shift
abstract
Under distribution shift (DS) where the training data distribution differs from the test one, a powerful technique is importance weighting (IW) which handles DS in two separate steps: weight estimation (WE) estimates the test-over-training density ratio and weighted classification (WC) trains the classifier from weighted training data. However, IW cannot work well on complex data, since WE is incompatible with deep learning. In this paper, we rethink IW and theoretically show it suffers from a circular dependency: we need not only WE for WC, but also WC for WE where a trained deep classifier is used as the feature extractor (FE). To cut off the dependency, we try to pretrain FE from unweighted training data, which leads to biased FE. To overcome the bias, we propose an end-to-end solution dynamic IW that iterates between WE and WC and combines them in a seamless manner, and hence our WE can also enjoy deep networks and stochastic optimizers indirectly. Experiments with two representative types of DS on three popular datasets show that our dynamic IW compares favorably with state-of-the-art methods.
Tongtong Fang, Nan Lu 0001, Gang Niu 0001, Masashi Sugiyama
NeurIPS2
2019 On the Minimal Supervision for Training Any Binary Classifier from Only Unlabeled Data
Nan Lu 0001, Gang Niu 0001, Aditya Krishna Menon, Masashi Sugiyama
ICLR (Poster)1