VLDB 2026 Research / reviewers in the wild / expert
Kunwoong Kim
dblp:296/1715
· DBLP profile ↗
7ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0000-2750-2463ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Data mining · 100% | |
| Artificial intelligence
4 papers |
Trustworthy machine learning · 87% Generative modeling · 13% |
Topics — the 13 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Data mining
clustering |
1.9 | 2 | 2026 | Fair Model-based Clustering · AAAI 2026 Fair Clustering via Alignment · ICML 2025 |
Data mining › clustering › constrained clustering
fair clustering |
1.9 | 2 | 2026 | Fair Model-based Clustering · AAAI 2026 Fair Clustering via Alignment · ICML 2025 |
Machine learning › Trustworthy machine learning
fairness |
1.4 | 2 | 2025 | Fair Representation Learning for Continuous Sensitive Attributes Using Expectation of Integral Probability Metrics · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Learning fair representation with a parametric integral probability metric · ICML 2022 |
Machine learning › Trustworthy machine learning › fairness
fair representation learning |
1.4 | 2 | 2025 | Fair Representation Learning for Continuous Sensitive Attributes Using Expectation of Integral Probability Metrics · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Learning fair representation with a parametric integral probability metric · ICML 2022 |
Data mining › clustering › model-based clustering
mixture model clustering |
1.0 | 1 | 2026 | Fair Model-based Clustering · AAAI 2026 |
Data mining › clustering
model-based clustering |
1.0 | 1 | 2026 | Fair Model-based Clustering · AAAI 2026 |
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels |
0.8 | 1 | 2024 | IOFM: Using the Interpolation Technique on the Over-Fitted Models to Identify Clean-Annotated Samples · AAAI 2024 |
Machine learning › Generative modeling
likelihood-based outlier detection |
0.8 | 1 | 2024 | ODIM: Outlier Detection via Likelihood of Under-Fitted Generative Models · ICML 2024 |
Machine learning › Trustworthy machine learning › robustness › learning with noisy labels
memorization effect |
0.8 | 1 | 2024 | IOFM: Using the Interpolation Technique on the Over-Fitted Models to Identify Clean-Annotated Samples · AAAI 2024 |
Data mining
anomaly detection |
0.8 | 1 | 2024 | ODIM: Outlier Detection via Likelihood of Under-Fitted Generative Models · ICML 2024 |
Data mining › anomaly detection
outlier detection |
0.8 | 1 | 2024 | ODIM: Outlier Detection via Likelihood of Under-Fitted Generative Models · ICML 2024 |
Data mining › anomaly detection › outlier detection
unsupervised outlier detection |
0.8 | 1 | 2024 | ODIM: Outlier Detection via Likelihood of Under-Fitted Generative Models · ICML 2024 |
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training |
0.6 | 1 | 2022 | Learning fair representation with a parametric integral probability metric · ICML 2022 |
Methods — techniques the papers use, named apart from their topics
under-fitted deep generative models · 1.5inlier-memorization effect · 1.5integral probability metric · 1.4mini-batch learning · 1.0expectation-maximization · 1.0maximum mean discrepancy · 0.9joint probability distribution alignment · 0.9alternating optimization · 0.9adversarial learning · 0.9over-fitted deep neural network · 0.8interpolation · 0.8fine-tuning · 0.8parametric discriminator · 0.6adversarial training · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fair Model-based ClusteringabstractThe goal of fair clustering is to find clusters such that the proportion of sensitive attributes (e.g., gender, race, etc) in each cluster is similar to the proportion of the entire data. Various fair clustering algorithms have been proposed, which modify standard K-means clustering to satisfy a given fairness constraint. A critical limitation of several existing fair clustering algorithms is that the number of parameters to be learned is proportional to the sample size because the cluster assignment of each datum should be optimized simultaneously with the cluster center, and thus scaling up the algorithms is difficult. In this paper, we propose a new fair clustering algorithm based on finite mixture model called Fair Model-based Clustering (FMC). A main advantage of FMC is that the number of learnable parameters is independent to the sample size and thus can be scaled up easily. In particular, a mini-batch learning is possible to obtain clusters that are approximately fair. Moreover, FMC can be applied to non-metric data (e.g., categorical data) as long as the likelihood is well-defined. Theoretical and empirical justifications of the superiority of the proposed algorithm are provided. Jinwon Park, Kunwoong Kim, Jihu Lee, Yongdai Kim |
AAAI | 2 |
| 2025 | Fair Clustering via AlignmentabstractAlgorithmic fairness in clustering aims to balance the proportions of instances assigned to each cluster with respect to a given sensitive attribute.
While recently developed fair clustering algorithms optimize clustering objectives under specific fairness constraints, their inherent complexity or approximation often results in suboptimal clustering utility or numerical instability in practice.
To resolve these limitations, we propose a new fair clustering algorithm based on a novel decomposition of the fair $K$-means clustering objective function.
The proposed algorithm, called Fair Clustering via Alignment (FCA), operates by alternately (i) finding a joint probability distribution to align the data from different protected groups, and (ii) optimizing cluster centers in the aligned space.
A key advantage of FCA is that it theoretically guarantees approximately optimal clustering utility for any given fairness level without complex constraints, thereby enabling high-utility fair clustering in practice.
Experiments show that FCA outperforms existing methods by (i) attaining a superior trade-off between fairness level and clustering utility, and (ii) achieving near-perfect fairness without numerical instability. Kunwoong Kim, Jihu Lee, Sangchul Park, Yongdai Kim |
ICML | 1 |
| 2025 | Fair Representation Learning for Continuous Sensitive Attributes Using Expectation of Integral Probability MetricsabstractAI fairness, also known as algorithmic fairness, aims to ensure that algorithms operate without bias or discrimination towards any individual or group. Among various AI algorithms, the Fair Representation Learning (FRL) approach has gained significant interest in recent years. However, existing FRL algorithms have a limitation: they are primarily designed for categorical sensitive attributes and thus cannot be applied to continuous sensitive attributes, such as age or income. In this paper, we propose an FRL algorithm for continuous sensitive attributes. First, we introduce a measure called the Expectation of Integral Probability Metrics (EIPM) to assess the fairness level of representation space for continuous sensitive attributes. We demonstrate that if the distribution of the representation has a low EIPM value, then any prediction head constructed on the top of the representation become fair, regardless of the selection of the prediction head. Furthermore, EIPM possesses a distinguished advantage in that it can be accurately estimated using our proposed estimator with finite samples. Based on these properties, we propose a new FRL algorithm called Fair Representation using EIPM with MMD (FREM). Experimental evidences show that FREM outperforms other baseline methods. Insung Kong, Kunwoong Kim, Yongdai Kim |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | IOFM: Using the Interpolation Technique on the Over-Fitted Models to Identify Clean-Annotated SamplesabstractMost recent state-of-the-art algorithms for handling noisy label problems are based on the memorization effect, which is a phenomenon that deep neural networks (DNNs) memorize clean data before noisy ones. While the memorization effect can be a powerful tool, there are several cases where memorization effect does not occur. Examples are imbalanced class distributions and heavy contamination on labels. To address this limitation, we introduce a whole new approach called the interpolation with the over-fitted model (IOFM), which leverages over-fitted deep neural networks. The IOFM utilizes a new finding of over-fitted DNNs: for a given training sample, its neighborhoods chosen from the feature space are distributed differently on the original input space depending on the cleanness of the target sample. The IOFM has notable features in two aspects: 1) it yields superior results even when the training data are imbalanced or heavily noisy, 2) since we utilize over-fitted deep neural networks, a fine-tuning procedure to select the optimal training epoch, which is an essential yet sensitive factor for the success of the memorization effect, is not required, and thus, the IOFM can be used for non-experts. Through extensive experiments, we show that our method can serve as a promising alternative to existing solutions dealing with noisy labels, offering improved performance even in challenging situations. Yongchan Choi, Kunwoong Kim, Ilsang Ohn, Yongdai Kim |
AAAI | 3 |
| 2024 | ODIM: Outlier Detection via Likelihood of Under-Fitted Generative ModelsabstractThe unsupervised outlier detection (UOD) problem refers to a task to identify inliers given training data which contain outliers as well as inliers, without any labeled information about inliers and outliers. It has been widely recognized that using fully-trained likelihood-based deep generative models (DGMs) often results in poor performance in distinguishing inliers from outliers. In this study, we claim that the likelihood itself could serve as powerful evidence for identifying inliers in UOD tasks, provided that DGMs are carefully under-fitted. Our approach begins with a novel observation called the inlier-memorization (IM) effect--when training a deep generative model with data including outliers, the model initially memorizes inliers before outliers. Based on this finding, we develop a new method called the outlier detection via the IM effect (ODIM). Remarkably, the ODIM requires only a few updates, making it computationally efficient--at least tens of times faster than other deep-learning-based algorithms. Also, the ODIM filters out outliers excellently, regardless of the data type, including tabular, image, and text data. To validate the superiority and efficiency of our method, we provide extensive empirical analyses on close to 60 datasets. Jaesung Hwang, Jongjin Lee, Kunwoong Kim, Yongdai Kim |
ICML | 4 |
| 2022 | Learning fair representation with a parametric integral probability metricabstractAs they have a vital effect on social decision-making, AI algorithms should be not only accurate but also fair. Among various algorithms for fairness AI, learning fair representation (LFR), whose goal is to find a fair representation with respect to sensitive variables such as gender and race, has received much attention. For LFR, the adversarial training scheme is popularly employed as is done in the generative adversarial network type algorithms. The choice of a discriminator, however, is done heuristically without justification. In this paper, we propose a new adversarial training scheme for LFR, where the integral probability metric (IPM) with a specific parametric family of discriminators is used. The most notable result of the proposed LFR algorithm is its theoretical guarantee about the fairness of the final prediction model, which has not been considered yet. That is, we derive theoretical relations between the fairness of representation and the fairness of the prediction model built on the top of the representation (i.e., using the representation as the input). Moreover, by numerical experiments, we show that our proposed LFR algorithm is computationally lighter and more stable, and the final prediction model is competitive or superior to other LFR algorithms using more complex discriminators. Kunwoong Kim, Insung Kong, Ilsang Ohn, Yongdai Kim |
ICML | 2 |
| 2022 | SLIDE: A surrogate fairness constraint to ensure fairness consistency
Kunwoong Kim, Ilsang Ohn, Sara Kim, Yongdai Kim |
Neural Networks | 1 |