Kunwoong Kim

dblp:296/1715 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0000-2750-2463ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Data mining · 100%
Artificial intelligence
4 papers
Trustworthy machine learning · 87% Generative modeling · 13%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining
clustering
1.922026
Fair Model-based Clustering · AAAI 2026
Fair Clustering via Alignment · ICML 2025
Data mining › clustering › constrained clustering
fair clustering
1.922026
Fair Model-based Clustering · AAAI 2026
Fair Clustering via Alignment · ICML 2025
Machine learning › Trustworthy machine learning
fairness
1.422025
Fair Representation Learning for Continuous Sensitive Attributes Using Expectation of Integral Probability Metrics · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Learning fair representation with a parametric integral probability metric · ICML 2022
Machine learning › Trustworthy machine learning › fairness
fair representation learning
1.422025
Fair Representation Learning for Continuous Sensitive Attributes Using Expectation of Integral Probability Metrics · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Learning fair representation with a parametric integral probability metric · ICML 2022
Data mining › clustering › model-based clustering
mixture model clustering
1.012026
Fair Model-based Clustering · AAAI 2026
Data mining › clustering
model-based clustering
1.012026
Fair Model-based Clustering · AAAI 2026
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.812024
IOFM: Using the Interpolation Technique on the Over-Fitted Models to Identify Clean-Annotated Samples · AAAI 2024
Machine learning › Generative modeling
likelihood-based outlier detection
0.812024
ODIM: Outlier Detection via Likelihood of Under-Fitted Generative Models · ICML 2024
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
memorization effect
0.812024
IOFM: Using the Interpolation Technique on the Over-Fitted Models to Identify Clean-Annotated Samples · AAAI 2024
Data mining
anomaly detection
0.812024
ODIM: Outlier Detection via Likelihood of Under-Fitted Generative Models · ICML 2024
Data mining › anomaly detection
outlier detection
0.812024
ODIM: Outlier Detection via Likelihood of Under-Fitted Generative Models · ICML 2024
Data mining › anomaly detection › outlier detection
unsupervised outlier detection
0.812024
ODIM: Outlier Detection via Likelihood of Under-Fitted Generative Models · ICML 2024
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.612022
Learning fair representation with a parametric integral probability metric · ICML 2022

Methods — techniques the papers use, named apart from their topics

under-fitted deep generative models · 1.5inlier-memorization effect · 1.5integral probability metric · 1.4mini-batch learning · 1.0expectation-maximization · 1.0maximum mean discrepancy · 0.9joint probability distribution alignment · 0.9alternating optimization · 0.9adversarial learning · 0.9over-fitted deep neural network · 0.8interpolation · 0.8fine-tuning · 0.8parametric discriminator · 0.6adversarial training · 0.6
YearPublicationVenuePosition
2026 Fair Model-based Clustering
abstract
The goal of fair clustering is to find clusters such that the proportion of sensitive attributes (e.g., gender, race, etc) in each cluster is similar to the proportion of the entire data. Various fair clustering algorithms have been proposed, which modify standard K-means clustering to satisfy a given fairness constraint. A critical limitation of several existing fair clustering algorithms is that the number of parameters to be learned is proportional to the sample size because the cluster assignment of each datum should be optimized simultaneously with the cluster center, and thus scaling up the algorithms is difficult. In this paper, we propose a new fair clustering algorithm based on finite mixture model called Fair Model-based Clustering (FMC). A main advantage of FMC is that the number of learnable parameters is independent to the sample size and thus can be scaled up easily. In particular, a mini-batch learning is possible to obtain clusters that are approximately fair. Moreover, FMC can be applied to non-metric data (e.g., categorical data) as long as the likelihood is well-defined. Theoretical and empirical justifications of the superiority of the proposed algorithm are provided.
Jinwon Park, Kunwoong Kim, Jihu Lee, Yongdai Kim
AAAI2
2025 Fair Clustering via Alignment
abstract
Algorithmic fairness in clustering aims to balance the proportions of instances assigned to each cluster with respect to a given sensitive attribute. While recently developed fair clustering algorithms optimize clustering objectives under specific fairness constraints, their inherent complexity or approximation often results in suboptimal clustering utility or numerical instability in practice. To resolve these limitations, we propose a new fair clustering algorithm based on a novel decomposition of the fair $K$-means clustering objective function. The proposed algorithm, called Fair Clustering via Alignment (FCA), operates by alternately (i) finding a joint probability distribution to align the data from different protected groups, and (ii) optimizing cluster centers in the aligned space. A key advantage of FCA is that it theoretically guarantees approximately optimal clustering utility for any given fairness level without complex constraints, thereby enabling high-utility fair clustering in practice. Experiments show that FCA outperforms existing methods by (i) attaining a superior trade-off between fairness level and clustering utility, and (ii) achieving near-perfect fairness without numerical instability.
Kunwoong Kim, Jihu Lee, Sangchul Park, Yongdai Kim
ICML1
2025 Fair Representation Learning for Continuous Sensitive Attributes Using Expectation of Integral Probability Metrics
abstract
AI fairness, also known as algorithmic fairness, aims to ensure that algorithms operate without bias or discrimination towards any individual or group. Among various AI algorithms, the Fair Representation Learning (FRL) approach has gained significant interest in recent years. However, existing FRL algorithms have a limitation: they are primarily designed for categorical sensitive attributes and thus cannot be applied to continuous sensitive attributes, such as age or income. In this paper, we propose an FRL algorithm for continuous sensitive attributes. First, we introduce a measure called the Expectation of Integral Probability Metrics (EIPM) to assess the fairness level of representation space for continuous sensitive attributes. We demonstrate that if the distribution of the representation has a low EIPM value, then any prediction head constructed on the top of the representation become fair, regardless of the selection of the prediction head. Furthermore, EIPM possesses a distinguished advantage in that it can be accurately estimated using our proposed estimator with finite samples. Based on these properties, we propose a new FRL algorithm called Fair Representation using EIPM with MMD (FREM). Experimental evidences show that FREM outperforms other baseline methods.
Insung Kong, Kunwoong Kim, Yongdai Kim
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 IOFM: Using the Interpolation Technique on the Over-Fitted Models to Identify Clean-Annotated Samples
abstract
Most recent state-of-the-art algorithms for handling noisy label problems are based on the memorization effect, which is a phenomenon that deep neural networks (DNNs) memorize clean data before noisy ones. While the memorization effect can be a powerful tool, there are several cases where memorization effect does not occur. Examples are imbalanced class distributions and heavy contamination on labels. To address this limitation, we introduce a whole new approach called the interpolation with the over-fitted model (IOFM), which leverages over-fitted deep neural networks. The IOFM utilizes a new finding of over-fitted DNNs: for a given training sample, its neighborhoods chosen from the feature space are distributed differently on the original input space depending on the cleanness of the target sample. The IOFM has notable features in two aspects: 1) it yields superior results even when the training data are imbalanced or heavily noisy, 2) since we utilize over-fitted deep neural networks, a fine-tuning procedure to select the optimal training epoch, which is an essential yet sensitive factor for the success of the memorization effect, is not required, and thus, the IOFM can be used for non-experts. Through extensive experiments, we show that our method can serve as a promising alternative to existing solutions dealing with noisy labels, offering improved performance even in challenging situations.
Yongchan Choi, Kunwoong Kim, Ilsang Ohn, Yongdai Kim
AAAI3
2024 ODIM: Outlier Detection via Likelihood of Under-Fitted Generative Models
abstract
The unsupervised outlier detection (UOD) problem refers to a task to identify inliers given training data which contain outliers as well as inliers, without any labeled information about inliers and outliers. It has been widely recognized that using fully-trained likelihood-based deep generative models (DGMs) often results in poor performance in distinguishing inliers from outliers. In this study, we claim that the likelihood itself could serve as powerful evidence for identifying inliers in UOD tasks, provided that DGMs are carefully under-fitted. Our approach begins with a novel observation called the inlier-memorization (IM) effect--when training a deep generative model with data including outliers, the model initially memorizes inliers before outliers. Based on this finding, we develop a new method called the outlier detection via the IM effect (ODIM). Remarkably, the ODIM requires only a few updates, making it computationally efficient--at least tens of times faster than other deep-learning-based algorithms. Also, the ODIM filters out outliers excellently, regardless of the data type, including tabular, image, and text data. To validate the superiority and efficiency of our method, we provide extensive empirical analyses on close to 60 datasets.
Jaesung Hwang, Jongjin Lee, Kunwoong Kim, Yongdai Kim
ICML4
2022 Learning fair representation with a parametric integral probability metric
abstract
As they have a vital effect on social decision-making, AI algorithms should be not only accurate but also fair. Among various algorithms for fairness AI, learning fair representation (LFR), whose goal is to find a fair representation with respect to sensitive variables such as gender and race, has received much attention. For LFR, the adversarial training scheme is popularly employed as is done in the generative adversarial network type algorithms. The choice of a discriminator, however, is done heuristically without justification. In this paper, we propose a new adversarial training scheme for LFR, where the integral probability metric (IPM) with a specific parametric family of discriminators is used. The most notable result of the proposed LFR algorithm is its theoretical guarantee about the fairness of the final prediction model, which has not been considered yet. That is, we derive theoretical relations between the fairness of representation and the fairness of the prediction model built on the top of the representation (i.e., using the representation as the input). Moreover, by numerical experiments, we show that our proposed LFR algorithm is computationally lighter and more stable, and the final prediction model is competitive or superior to other LFR algorithms using more complex discriminators.
Kunwoong Kim, Insung Kong, Ilsang Ohn, Yongdai Kim
ICML2
2022 SLIDE: A surrogate fairness constraint to ensure fairness consistency
Kunwoong Kim, Ilsang Ohn, Sara Kim, Yongdai Kim
Neural Networks1