Nontawat Charoenphakdee

dblp:227/3074 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-0214-4943ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Trustworthy machine learning · 40% Learning theory · 21% Reinforcement learning · 11%

Topics — the 22 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › uncertainty estimation
selective classification
0.922021
Classification with Rejection Based on Cost-sensitive Classification · ICML 2021
On the Calibration of Multiclass Classification with Rejection · NeurIPS 2019
Machine learning › Learning theory › classification › classification error analysis
bayes error estimation
0.712023
Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary Classification · ICLR 2023
Machine learning › Trustworthy machine learning › calibration
classifier calibration
0.512021
On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective · CVPR 2021
Machine learning › Learning paradigms › cost-sensitive learning
cost-sensitive classification
0.512021
Classification with Rejection Based on Cost-sensitive Classification · ICML 2021
Machine learning › Deep learning architectures and training › loss function design
focal loss
0.512021
On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective · CVPR 2021
Machine learning › Trustworthy machine learning
interpretability
0.512021
On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective · CVPR 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
maximum likelihood estimation
0.412020
Learning from Aggregate Observations · NeurIPS 2020
Machine learning › Learning paradigms
weakly supervised learning
0.412020
Learning from Aggregate Observations · NeurIPS 2020
Machine learning › Trustworthy machine learning
calibration
0.412019
On the Calibration of Multiclass Classification with Rejection · NeurIPS 2019
Machine learning › Learning theory › loss function › surrogate loss
classification calibration
0.412019
On Symmetric Losses for Learning from Corrupted Labels · ICML 2019
Machine learning › Learning theory
discrepancy measure
0.412019
Unsupervised Domain Adaptation Based on Source-Guided Discrepancy · AAAI 2019
Machine learning › Learning theory › generalization bounds
domain adaptation bound
0.412019
Unsupervised Domain Adaptation Based on Source-Guided Discrepancy · AAAI 2019
Machine learning › Learning theory
generalization bounds
0.412019
Unsupervised Domain Adaptation Based on Source-Guided Discrepancy · AAAI 2019
Machine learning › Reinforcement learning › imitation learning › generative imitation learning
generative adversarial imitation learning
0.412019
Imitation Learning from Imperfect Demonstration · ICML 2019
Machine learning › Reinforcement learning
imitation learning
0.412019
Imitation Learning from Imperfect Demonstration · ICML 2019
Machine learning › Reinforcement learning › imitation learning
learning from imperfect demonstrations
0.412019
Imitation Learning from Imperfect Demonstration · ICML 2019
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.412019
On Symmetric Losses for Learning from Corrupted Labels · ICML 2019
Machine learning › Trustworthy machine learning
robustness
0.412019
On the Calibration of Multiclass Classification with Rejection · NeurIPS 2019
Machine learning › Trustworthy machine learning
uncertainty and calibration
0.412019
On the Calibration of Multiclass Classification with Rejection · NeurIPS 2019
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.412019
Unsupervised Domain Adaptation Based on Source-Guided Discrepancy · AAAI 2019
Natural language and speech › Information extraction and text analysis › text classification
weakly supervised text classification
0.412019
Learning Only from Relevant Keywords and Unlabeled Documents · EMNLP/IJCNLP (1) 2019
Natural language and speech › Language models and text generation › large language model › large language model adaptation
pre-trained language model fine-tuning
0.112019
Learning Only from Relevant Keywords and Unlabeled Documents · EMNLP/IJCNLP (1) 2019

Methods — techniques the papers use, named apart from their topics

confidence scores · 0.8binary classification · 0.7bayes error estimation · 0.7proper scoring rules · 0.5ensemble learning · 0.5cost-sensitive learning · 0.5classification calibration · 0.5maximum likelihood estimation · 0.4consistency analysis · 0.4finite-sample convergence · 0.4
YearPublicationVenuePosition
2025 Calm Composite Losses: Being Improper Yet Proper Composite
abstract
Strict proper losses are fundamental loss functions inducing classifiers capable of estimating class probabilities. While practitioners have devised many loss functions, their properness is often unverified. In this paper, we identify several losses as improper, calling into question the validity of class probability estimates derived from their simplex-projected outputs. Nevertheless, we show that these losses are strictly proper composite with appropriate link functions, allowing predictions to be mapped into true class probabilities. We invent the calmness condition, which we prove suffices to identify that a loss has a strictly proper composite representation, and provide the general form of the inverse link. To further understand proper composite losses, we explore proper composite losses through the framework of property elicitation, revealing a connection between inverse link functions and Bregman projections. Numerical simulations are provided to demonstrate the behavior of proper composite losses and the effectiveness of the inverse link function.
Han Bao 0002, Nontawat Charoenphakdee
AISTATS2
2023 Is the Performance of My Deep Network Too Good to Be True? A Direct Approach to Estimating the Bayes Error in Binary Classification
Takashi Ishida 0001, Ikko Yamane, Nontawat Charoenphakdee, Gang Niu 0001, Masashi Sugiyama
ICLR3
2021 Robust Imitation Learning from Noisy Demonstrations
abstract
Robust learning from noisy demonstrations is a practical but highly challenging problem in imitation learning. In this paper, we first theoretically show that robust imitation learning can be achieved by optimizing a classification risk with a symmetric loss. Based on this theoretical finding, we then propose a new imitation learning method that optimizes the classification risk by effectively combining pseudo-labeling with co-training. Unlike existing methods, our method does not require additional labels or strict assumptions about noise distributions. Experimental results on continuous-control benchmarks show that our method is more robust compared to state-of-the-art methods.
Voot Tangkaratt, Nontawat Charoenphakdee, Masashi Sugiyama
AISTATS2
2021 On Focal Loss for Class-Posterior Probability Estimation: A Theoretical Perspective
abstract
The focal loss has demonstrated its effectiveness in many real-world applications such as object detection and image classification, but its theoretical understanding has been limited so far. In this paper, we first prove that the focal loss is classification-calibrated, i.e., its minimizer surely yields the Bayes-optimal classifier and thus the use of the focal loss in classification can be theoretically justified. However, we also prove a negative fact that the focal loss is not strictly proper, i.e., the confidence score of the classifier obtained by focal loss minimization does not match the true class-posterior probability. This may cause the trained classifier to give an unreliable confidence score, which can be harmful in critical applications. To mitigate this problem, we prove that there exists a particular closed-form transformation that can recover the true class-posterior probability from the outputs of the focal risk minimizer. Our experiments show that our proposed transformation successfully improves the quality of class-posterior probability estimation and improves the calibration of the trained classifier, while preserving the same prediction accuracy.
Nontawat Charoenphakdee, Jayakorn Vongkulbhisal, Nuttapong Chairatanakul, Masashi Sugiyama
CVPR1
2021 Classification with Rejection Based on Cost-sensitive Classification
abstract
The goal of classification with rejection is to avoid risky misclassification in error-critical applications such as medical diagnosis and product inspection. In this paper, based on the relationship between classification with rejection and cost-sensitive classification, we propose a novel method of classification with rejection by learning an ensemble of cost-sensitive classifiers, which satisfies all the following properties: (i) it can avoid estimating class-posterior probabilities, resulting in improved classification accuracy. (ii) it allows a flexible choice of losses including non-convex ones, (iii) it does not require complicated modifications when using different losses, (iv) it is applicable to both binary and multiclass cases, and (v) it is theoretically justifiable for any classification-calibrated loss. Experimental results demonstrate the usefulness of our proposed approach in clean-labeled, noisy-labeled, and positive-unlabeled classification.
Nontawat Charoenphakdee, Zhenghang Cui, Yivan Zhang, Masashi Sugiyama
ICML1
2021 Semisupervised Ordinal Regression Based on Empirical Risk Minimization
abstract
Ordinal regression is aimed at predicting an ordinal class label. In this letter, we consider its semisupervised formulation, in which we have unlabeled data along with ordinal-labeled data to train an ordinal regressor. There are several metrics to evaluate the performance of ordinal regression, such as the mean absolute error, mean zero-one error, and mean squared error. However, the existing studies do not take the evaluation metric into account, restrict model choice, and have no theoretical guarantee. To overcome these problems, we propose a novel generic framework for semisupervised ordinal regression based on the empirical risk minimization principle that is applicable to optimizing all of the metrics mentioned above. In addition, our framework has flexible choices of models, surrogate losses, and optimization algorithms without the common geometric assumption on unlabeled data such as the cluster assumption or manifold assumption. We provide an estimation error bound to show that our risk estimator is consistent. Finally, we conduct experiments to show the usefulness of our framework.
Taira Tsuchiya, Nontawat Charoenphakdee, Issei Sato, Masashi Sugiyama
Neural Comput.2
2020 Learning from Aggregate Observations
abstract
We study the problem of learning from aggregate observations where supervision signals are given to sets of instances instead of individual instances, while the goal is still to predict labels of unseen individuals. A well-known example is multiple instance learning (MIL). In this paper, we extend MIL beyond binary classification to other problems such as multiclass classification and regression. We present a general probabilistic framework that accommodates a variety of aggregate observations, e.g., pairwise similarity/triplet comparison for classification and mean/difference/rank observation for regression. Simple maximum likelihood solutions can be applied to various differentiable models such as deep neural networks and gradient boosting machines. Moreover, we develop the concept of consistency up to an equivalence relation to characterize our estimator and show that it has nice convergence properties under mild assumptions. Experiments on three problem settings --- classification via triplet comparison and regression via mean/rank observation indicate the effectiveness of the proposed method.
Yivan Zhang, Nontawat Charoenphakdee, Zhenguo Wu, Masashi Sugiyama
NeurIPS2
2020 Classification from Triplet Comparison Data
abstract
Learning from triplet comparison data has been extensively studied in the context of metric learning, where we want to learn a distance metric between two instances, and ordinal embedding, where we want to learn an embedding in a Euclidean space of the given instances that preserve the comparison order as much as possible. Unlike fully labeled data, triplet comparison data can be collected in a more accurate and human-friendly way. Although learning from triplet comparison data has been considered in many applications, an important fundamental question of whether we can learn a classifier only from triplet comparison data without all the labels has remained unanswered. In this letter, we give a positive answer to this important question by proposing an unbiased estimator for the classification risk under the empirical risk minimization framework. Since the proposed method is based on the empirical risk minimization framework, it inherently has the advantage that any surrogate loss function and any model, including neural networks, can be easily applied. Furthermore, we theoretically establish an estimation error bound for the proposed empirical risk minimizer. Finally, we provide experimental results to show that our method empirically works well and outperforms various baseline methods.
Zhenghang Cui, Nontawat Charoenphakdee, Issei Sato, Masashi Sugiyama
Neural Comput.2
2019 Unsupervised Domain Adaptation Based on Source-Guided Discrepancy
abstract
Unsupervised domain adaptation is the problem setting where data generating distributions in the source and target domains are different and labels in the target domain are unavailable. An important question in unsupervised domain adaptation is how to measure the difference between the source and target domains. Existing discrepancy measures for unsupervised domain adaptation either require high computation costs or have no theoretical guarantee. To mitigate these problems, this paper proposes a novel discrepancy measure called source-guided discrepancy (S-disc), which exploits labels in the source domain unlike the existing ones. As a consequence, S-disc can be computed efficiently with a finitesample convergence guarantee. In addition, it is shown that S-disc can provide a tighter generalization error bound than the one based on an existing discrepancy measure. Finally, experimental results demonstrate the advantages of S-disc over the existing discrepancy measures.
Seiichi Kuroki, Nontawat Charoenphakdee, Han Bao 0002, Junya Honda, Issei Sato, Masashi Sugiyama
AAAI2
2019 Learning Only from Relevant Keywords and Unlabeled Documents
abstract
Nontawat Charoenphakdee, Jongyeong Lee, Yiping Jin, Dittaya Wanvarie, Masashi Sugiyama. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Nontawat Charoenphakdee, Jongyeong Lee, Yiping Jin, Dittaya Wanvarie, Masashi Sugiyama
EMNLP/IJCNLP (1)1
2019 On Symmetric Losses for Learning from Corrupted Labels
abstract
This paper aims to provide a better understanding of a symmetric loss. First, we emphasize that using a symmetric loss is advantageous in the balanced error rate (BER) minimization and area under the receiver operating characteristic curve (AUC) maximization from corrupted labels. Second, we prove general theoretical properties of symmetric losses, including a classification-calibration condition, excess risk bound, conditional risk minimizer, and AUC-consistency condition. Third, since all nonnegative symmetric losses are non-convex, we propose a convex barrier hinge loss that benefits significantly from the symmetric condition, although it is not symmetric everywhere. Finally, we conduct experiments to validate the relevance of the symmetric condition.
Nontawat Charoenphakdee, Jongyeong Lee, Masashi Sugiyama
ICML1
2019 Imitation Learning from Imperfect Demonstration
abstract
Imitation learning (IL) aims to learn an optimal policy from demonstrations. However, such demonstrations are often imperfect since collecting optimal ones is costly. To effectively learn from imperfect demonstrations, we propose a novel approach that utilizes confidence scores, which describe the quality of demonstrations. More specifically, we propose two confidence-based IL methods, namely two-step importance weighting IL (2IWIL) and generative adversarial IL with imperfect demonstration and confidence (IC-GAIL). We show that confidence scores given only to a small portion of sub-optimal demonstrations significantly improve the performance of IL both theoretically and empirically.
Yueh-Hua Wu, Nontawat Charoenphakdee, Han Bao 0002, Voot Tangkaratt, Masashi Sugiyama
ICML2
2019 On the Calibration of Multiclass Classification with Rejection
abstract
We investigate the problem of multiclass classification with rejection, where a classifier can choose not to make a prediction to avoid critical misclassification. First, we consider an approach based on simultaneous training of a classifier and a rejector, which achieves the state-of-the-art performance in the binary case. We analyze this approach for the multiclass case and derive a general condition for calibration to the Bayes-optimal solution, which suggests that calibration is hard to achieve by general loss functions unlike the binary case. Next, we consider another traditional approach based on confidence scores, in which the existing work focuses on a specific class of losses. We propose rejection criteria for more general losses for this approach and guarantee calibration to the Bayes-optimal solution. Finally, we conduct experiments to validate the relevance of our theoretical findings.
Chenri Ni, Nontawat Charoenphakdee, Junya Honda, Masashi Sugiyama
NeurIPS2
2019 Positive-Unlabeled Classification under Class Prior Shift and Asymmetric Error
abstract
Bottlenecks of binary classification from positive and unlabeled data (PU classification) are the requirements that given unlabeled patterns are drawn from the same distribution as the test distribution, and the penalty of the false positive error is identical to the false negative error. However, such requirements are often not fulfilled in practice. In this paper, we generalize PU classification to the class prior shift and asymmetric error scenarios. Based on the analysis of the Bayes optimal classifier, we show that given a test class prior, PU classification under class prior shift is equivalent to PU classification with asymmetric error. Then, we propose two different frameworks to handle these problems, namely, a risk minimization framework and density ratio estimation framework. Finally, we demonstrate the effectiveness of the proposed frameworks through experiments.
Nontawat Charoenphakdee, Masashi Sugiyama
SDM1