VLDB 2026 Research / reviewers in the wild / expert
Nontawat Charoenphakdee
dblp:227/3074
· DBLP profile ↗
14ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-0214-4943ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Trustworthy machine learning · 40% Learning theory · 21% Reinforcement learning · 11% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification |
0.9 | 2 | 2021 | Classification with Rejection Based on Cost-sensitive Classification · ICML 2021 On the Calibration of Multiclass Classification with Rejection · NeurIPS 2019 |
Machine learning › Learning theory › classification › classification error analysis
bayes error estimation |
0.7 | 1 | 2023 | Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary Classification · ICLR 2023 |
Machine learning › Trustworthy machine learning › calibration
classifier calibration |
0.5 | 1 | 2021 | On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective · CVPR 2021 |
Machine learning › Learning paradigms › cost-sensitive learning
cost-sensitive classification |
0.5 | 1 | 2021 | Classification with Rejection Based on Cost-sensitive Classification · ICML 2021 |
Machine learning › Deep learning architectures and training › loss function design
focal loss |
0.5 | 1 | 2021 | On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective · CVPR 2021 |
Machine learning › Trustworthy machine learning
interpretability |
0.5 | 1 | 2021 | On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective · CVPR 2021 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation |
0.4 | 1 | 2020 | Learning from Aggregate Observations · NeurIPS 2020 |
Machine learning › Learning paradigms
weakly supervised learning |
0.4 | 1 | 2020 | Learning from Aggregate Observations · NeurIPS 2020 |
Machine learning › Trustworthy machine learning
calibration |
0.4 | 1 | 2019 | On the Calibration of Multiclass Classification with Rejection · NeurIPS 2019 |
Machine learning › Learning theory › loss function › surrogate loss
classification calibration |
0.4 | 1 | 2019 | On Symmetric Losses for Learning from Corrupted Labels · ICML 2019 |
Machine learning › Learning theory
discrepancy measure |
0.4 | 1 | 2019 | Unsupervised Domain Adaptation Based on Source-Guided Discrepancy · AAAI 2019 |
Machine learning › Learning theory › generalization bounds
domain adaptation bound |
0.4 | 1 | 2019 | Unsupervised Domain Adaptation Based on Source-Guided Discrepancy · AAAI 2019 |
Machine learning › Learning theory
generalization bounds |
0.4 | 1 | 2019 | Unsupervised Domain Adaptation Based on Source-Guided Discrepancy · AAAI 2019 |
Machine learning › Reinforcement learning › imitation learning › generative imitation learning
generative adversarial imitation learning |
0.4 | 1 | 2019 | Imitation Learning from Imperfect Demonstration · ICML 2019 |
Machine learning › Reinforcement learning
imitation learning |
0.4 | 1 | 2019 | Imitation Learning from Imperfect Demonstration · ICML 2019 |
Machine learning › Reinforcement learning › imitation learning
learning from imperfect demonstrations |
0.4 | 1 | 2019 | Imitation Learning from Imperfect Demonstration · ICML 2019 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.4 | 1 | 2019 | On Symmetric Losses for Learning from Corrupted Labels · ICML 2019 |
Machine learning › Trustworthy machine learning
robustness |
0.4 | 1 | 2019 | On the Calibration of Multiclass Classification with Rejection · NeurIPS 2019 |
Machine learning › Trustworthy machine learning
uncertainty and calibration |
0.4 | 1 | 2019 | On the Calibration of Multiclass Classification with Rejection · NeurIPS 2019 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.4 | 1 | 2019 | Unsupervised Domain Adaptation Based on Source-Guided Discrepancy · AAAI 2019 |
Natural language and speech › Information extraction and text analysis › text classification
weakly supervised text classification |
0.4 | 1 | 2019 | Learning Only from Relevant Keywords and Unlabeled Documents · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Language models and text generation › large language model › large language model adaptation
pre-trained language model fine-tuning |
0.1 | 1 | 2019 | Learning Only from Relevant Keywords and Unlabeled Documents · EMNLP/IJCNLP (1) 2019 |
Methods — techniques the papers use, named apart from their topics
confidence scores · 0.8binary classification · 0.7bayes error estimation · 0.7proper scoring rules · 0.5ensemble learning · 0.5cost-sensitive learning · 0.5classification calibration · 0.5maximum likelihood estimation · 0.4consistency analysis · 0.4finite-sample convergence · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Calm Composite Losses: Being Improper Yet Proper CompositeabstractStrict proper losses are fundamental loss functions inducing classifiers capable of estimating class probabilities. While practitioners have devised many loss functions, their properness is often unverified. In this paper, we identify several losses as improper, calling into question the validity of class probability estimates derived from their simplex-projected outputs. Nevertheless, we show that these losses are strictly proper composite with appropriate link functions, allowing predictions to be mapped into true class probabilities. We invent the calmness condition, which we prove suffices to identify that a loss has a strictly proper composite representation, and provide the general form of the inverse link. To further understand proper composite losses, we explore proper composite losses through the framework of property elicitation, revealing a connection between inverse link functions and Bregman projections. Numerical simulations are provided to demonstrate the behavior of proper composite losses and the effectiveness of the inverse link function. Han Bao 0002, Nontawat Charoenphakdee |
AISTATS | 2 |
| 2023 | Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary Classification
Takashi Ishida 0001, Ikko Yamane, Nontawat Charoenphakdee, Gang Niu 0001, Masashi Sugiyama |
ICLR | 3 |
| 2021 | Robust Imitation Learning from Noisy DemonstrationsabstractRobust learning from noisy demonstrations is a practical but highly challenging problem in imitation learning. In this paper, we first theoretically show that robust imitation learning can be achieved by optimizing a classification risk with a symmetric loss. Based on this theoretical finding, we then propose a new imitation learning method that optimizes the classification risk by effectively combining pseudo-labeling with co-training. Unlike existing methods, our method does not require additional labels or strict assumptions about noise distributions. Experimental results on continuous-control benchmarks show that our method is more robust compared to state-of-the-art methods. Voot Tangkaratt, Nontawat Charoenphakdee, Masashi Sugiyama |
AISTATS | 2 |
| 2021 | On Focal Loss for Class-Posterior Probability Estimation: A Theoretical PerspectiveabstractThe focal loss has demonstrated its effectiveness in many real-world applications such as object detection and image classification, but its theoretical understanding has been limited so far. In this paper, we first prove that the focal loss is classification-calibrated, i.e., its minimizer surely yields the Bayes-optimal classifier and thus the use of the focal loss in classification can be theoretically justified. However, we also prove a negative fact that the focal loss is not strictly proper, i.e., the confidence score of the classifier obtained by focal loss minimization does not match the true class-posterior probability. This may cause the trained classifier to give an unreliable confidence score, which can be harmful in critical applications. To mitigate this problem, we prove that there exists a particular closed-form transformation that can recover the true class-posterior probability from the outputs of the focal risk minimizer. Our experiments show that our proposed transformation successfully improves the quality of class-posterior probability estimation and improves the calibration of the trained classifier, while preserving the same prediction accuracy. Nontawat Charoenphakdee, Jayakorn Vongkulbhisal, Nuttapong Chairatanakul, Masashi Sugiyama |
CVPR | 1 |
| 2021 | Classification with Rejection Based on Cost-sensitive ClassificationabstractThe goal of classification with rejection is to avoid risky misclassification in error-critical applications such as medical diagnosis and product inspection. In this paper, based on the relationship between classification with rejection and cost-sensitive classification, we propose a novel method of classification with rejection by learning an ensemble of cost-sensitive classifiers, which satisfies all the following properties: (i) it can avoid estimating class-posterior probabilities, resulting in improved classification accuracy. (ii) it allows a flexible choice of losses including non-convex ones, (iii) it does not require complicated modifications when using different losses, (iv) it is applicable to both binary and multiclass cases, and (v) it is theoretically justifiable for any classification-calibrated loss. Experimental results demonstrate the usefulness of our proposed approach in clean-labeled, noisy-labeled, and positive-unlabeled classification. Nontawat Charoenphakdee, Zhenghang Cui, Yivan Zhang, Masashi Sugiyama |
ICML | 1 |
| 2021 | Semisupervised Ordinal Regression Based on Empirical Risk MinimizationabstractOrdinal regression is aimed at predicting an ordinal class label. In this letter, we consider its semisupervised formulation, in which we have unlabeled data along with ordinal-labeled data to train an ordinal regressor. There are several metrics to evaluate the performance of ordinal regression, such as the mean absolute error, mean zero-one error, and mean squared error. However, the existing studies do not take the evaluation metric into account, restrict model choice, and have no theoretical guarantee. To overcome these problems, we propose a novel generic framework for semisupervised ordinal regression based on the empirical risk minimization principle that is applicable to optimizing all of the metrics mentioned above. In addition, our framework has flexible choices of models, surrogate losses, and optimization algorithms without the common geometric assumption on unlabeled data such as the cluster assumption or manifold assumption. We provide an estimation error bound to show that our risk estimator is consistent. Finally, we conduct experiments to show the usefulness of our framework. Taira Tsuchiya, Nontawat Charoenphakdee, Issei Sato, Masashi Sugiyama |
Neural Comput. | 2 |
| 2020 | Learning from Aggregate ObservationsabstractWe study the problem of learning from aggregate observations where supervision signals are given to sets of instances instead of individual instances, while the goal is still to predict labels of unseen individuals. A well-known example is multiple instance learning (MIL). In this paper, we extend MIL beyond binary classification to other problems such as multiclass classification and regression. We present a general probabilistic framework that accommodates a variety of aggregate observations, e.g., pairwise similarity/triplet comparison for classification and mean/difference/rank observation for regression. Simple maximum likelihood solutions can be applied to various differentiable models such as deep neural networks and gradient boosting machines. Moreover, we develop the concept of consistency up to an equivalence relation to characterize our estimator and show that it has nice convergence properties under mild assumptions. Experiments on three problem settings --- classification via triplet comparison and regression via mean/rank observation indicate the effectiveness of the proposed method. Yivan Zhang, Nontawat Charoenphakdee, Zhenguo Wu, Masashi Sugiyama |
NeurIPS | 2 |
| 2020 | Classification from Triplet Comparison DataabstractLearning from triplet comparison data has been extensively studied in the context of metric learning, where we want to learn a distance metric between two instances, and ordinal embedding, where we want to learn an embedding in a Euclidean space of the given instances that preserve the comparison order as much as possible. Unlike fully labeled data, triplet comparison data can be collected in a more accurate and human-friendly way. Although learning from triplet comparison data has been considered in many applications, an important fundamental question of whether we can learn a classifier only from triplet comparison data without all the labels has remained unanswered. In this letter, we give a positive answer to this important question by proposing an unbiased estimator for the classification risk under the empirical risk minimization framework. Since the proposed method is based on the empirical risk minimization framework, it inherently has the advantage that any surrogate loss function and any model, including neural networks, can be easily applied. Furthermore, we theoretically establish an estimation error bound for the proposed empirical risk minimizer. Finally, we provide experimental results to show that our method empirically works well and outperforms various baseline methods. Zhenghang Cui, Nontawat Charoenphakdee, Issei Sato, Masashi Sugiyama |
Neural Comput. | 2 |
| 2019 | Unsupervised Domain Adaptation Based on Source-Guided DiscrepancyabstractUnsupervised domain adaptation is the problem setting where data generating distributions in the source and target domains are different and labels in the target domain are unavailable. An important question in unsupervised domain adaptation is how to measure the difference between the source and target domains. Existing discrepancy measures for unsupervised domain adaptation either require high computation costs or have no theoretical guarantee. To mitigate these problems, this paper proposes a novel discrepancy measure called source-guided discrepancy (S-disc), which exploits labels in the source domain unlike the existing ones. As a consequence, S-disc can be computed efficiently with a finitesample convergence guarantee. In addition, it is shown that S-disc can provide a tighter generalization error bound than the one based on an existing discrepancy measure. Finally, experimental results demonstrate the advantages of S-disc over the existing discrepancy measures. Seiichi Kuroki, Nontawat Charoenphakdee, Han Bao 0002, Junya Honda, Issei Sato, Masashi Sugiyama |
AAAI | 2 |
| 2019 | Learning Only from Relevant Keywords and Unlabeled DocumentsabstractNontawat Charoenphakdee, Jongyeong Lee, Yiping Jin, Dittaya Wanvarie, Masashi Sugiyama. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Nontawat Charoenphakdee, Jongyeong Lee, Yiping Jin, Dittaya Wanvarie, Masashi Sugiyama |
EMNLP/IJCNLP (1) | 1 |
| 2019 | On Symmetric Losses for Learning from Corrupted LabelsabstractThis paper aims to provide a better understanding of a symmetric loss. First, we emphasize that using a symmetric loss is advantageous in the balanced error rate (BER) minimization and area under the receiver operating characteristic curve (AUC) maximization from corrupted labels. Second, we prove general theoretical properties of symmetric losses, including a classification-calibration condition, excess risk bound, conditional risk minimizer, and AUC-consistency condition. Third, since all nonnegative symmetric losses are non-convex, we propose a convex barrier hinge loss that benefits significantly from the symmetric condition, although it is not symmetric everywhere. Finally, we conduct experiments to validate the relevance of the symmetric condition. Nontawat Charoenphakdee, Jongyeong Lee, Masashi Sugiyama |
ICML | 1 |
| 2019 | Imitation Learning from Imperfect DemonstrationabstractImitation learning (IL) aims to learn an optimal policy from demonstrations. However, such demonstrations are often imperfect since collecting optimal ones is costly. To effectively learn from imperfect demonstrations, we propose a novel approach that utilizes confidence scores, which describe the quality of demonstrations. More specifically, we propose two confidence-based IL methods, namely two-step importance weighting IL (2IWIL) and generative adversarial IL with imperfect demonstration and confidence (IC-GAIL). We show that confidence scores given only to a small portion of sub-optimal demonstrations significantly improve the performance of IL both theoretically and empirically. Yueh-Hua Wu, Nontawat Charoenphakdee, Han Bao 0002, Voot Tangkaratt, Masashi Sugiyama |
ICML | 2 |
| 2019 | On the Calibration of Multiclass Classification with RejectionabstractWe investigate the problem of multiclass classification with rejection, where a classifier can choose not to make a prediction to avoid critical misclassification. First, we consider an approach based on simultaneous training of a classifier and a rejector, which achieves the state-of-the-art performance in the binary case. We analyze this approach for the multiclass case and derive a general condition for calibration to the Bayes-optimal solution, which suggests that calibration is hard to achieve by general loss functions unlike the binary case. Next, we consider another traditional approach based on confidence scores, in which the existing work focuses on a specific class of losses. We propose rejection criteria for more general losses for this approach and guarantee calibration to the Bayes-optimal solution. Finally, we conduct experiments to validate the relevance of our theoretical findings. Chenri Ni, Nontawat Charoenphakdee, Junya Honda, Masashi Sugiyama |
NeurIPS | 2 |
| 2019 | Positive-Unlabeled Classification under Class Prior Shift and Asymmetric ErrorabstractBottlenecks of binary classification from positive and unlabeled data (PU classification) are the requirements that given unlabeled patterns are drawn from the same distribution as the test distribution, and the penalty of the false positive error is identical to the false negative error. However, such requirements are often not fulfilled in practice. In this paper, we generalize PU classification to the class prior shift and asymmetric error scenarios. Based on the analysis of the Bayes optimal classifier, we show that given a test class prior, PU classification under class prior shift is equivalent to PU classification with asymmetric error. Then, we propose two different frameworks to handle these problems, namely, a risk minimization framework and density ratio estimation framework. Finally, we demonstrate the effectiveness of the proposed frameworks through experiments. Nontawat Charoenphakdee, Masashi Sugiyama |
SDM | 1 |