EDBT 2026 Demo / reviewers in the wild / expert
Yeongmin Kim
dblp:121/8864
· DBLP profile ↗
16ranked-venue papers
6as first author
16since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 5 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less ComputeabstractDespite their remarkable performance, modern Diffusion Transformers (DiTs) are hindered by substantial resource requirements during inference, stemming from the fixed and large amount of compute needed for each denoising step. In this work, we revisit the conventional static paradigm that allocates a fixed compute budget per denoising iteration and propose a dynamic strategy instead. Our simple and sample-efficient framework enables pre-trained DiT models to be converted into flexible ones — dubbed FlexiDiT— allowing them to process inputs at varying compute budgets. We demonstrate how a single flexible model can generate images without any drop in quality, while reducing the required FLOPs by more than 40% compared to their static counterparts, for both class-conditioned and text-conditioned image generation. Our method is general and agnostic to input and conditioning modalities. We show how our approach can be readily extended for video generation, where FlexiDiT models generate samples with up to 75% less compute without compromising performance. Sotiris Anagnostidis, Gregor Bachmann, Yeongmin Kim, Jonas Kohler, Markos Georgopoulos, Artsiom Sanakoyeu, Yuming Du, Albert Pumarola, Ali K. Thabet, Edgar Schönfeld |
CVPR | 3 |
| 2025 | Autoregressive Distillation of Diffusion TransformersabstractDiffusion models with transformer architectures have demonstrated promising capabilities in generating high-fidelity images and scalability for high resolution. However, iterative sampling process required for synthesis is very resource-intensive. A line of work has focused on distilling solutions to probability flow ODEs into few-step student models. Nevertheless, existing methods have been limited by their reliance on the most recent denoised samples as input, rendering them susceptible to exposure bias. To address this limitation, we propose AutoRegressive Distillation (ARD), a novel approach that leverages the historical trajectory of the ODE to predict future steps. ARD offers two key benefits: 1) it mitigates exposure bias by utilizing a predicted historical trajectory that is less susceptible to accumulated errors, and 2) it leverages the previous history of the ODE trajectory as a more effective source of coarse-grained information. ARD modifies the teacher transformer architecture by adding token-wise time embedding to mark each input from the trajectory history and employs a block-wise causal attention mask for training. Furthermore, incorporating historical inputs only in lower transformer layers enhances performance and efficiency. We validate the effectiveness of ARD in a class-conditioned generation on ImageNet and T2I synthesis. Our model achieves a 5× reduction in FID degradation compared to the baseline methods while requiring only 1.1% extra FLOPs on ImageNet-256. Moreover, ARD reaches FID of 1.84 on ImageNet-256 in merely 4 steps and outperforms the publicly available 1024p text-to-image distilled models in prompt adherence score with a minimal drop in FID compared to the teacher. Project page: https://github.com/alsdudrla10/ARD. Yeongmin Kim, Sotiris Anagnostidis, Yuming Du, Edgar Schönfeld, Jonas Kohler, Markos Georgopoulos, Albert Pumarola, Ali K. Thabet, Artsiom Sanakoyeu |
CVPR | 1 |
| 2025 | Diffusion Bridge AutoEncoders for Unsupervised Representation LearningabstractDiffusion-based representation learning has achieved substantial attention due to its promising capabilities in latent representation and sample generation. Recent studies have employed an auxiliary encoder to identify a corresponding representation from data and to adjust the dimensionality of a latent variable $\mathbf{z}$. Meanwhile, this auxiliary structure invokes an *information split problem*; the information of each data instance $\mathbf{x}_0$ is divided into diffusion endpoint $\mathbf{x}_T$ and encoded $\mathbf{z}$ because there exist two inference paths starting from the data. The latent variable modeled by diffusion endpoint $\mathbf{x}_T$ has some disadvantages. The diffusion endpoint $\mathbf{x}_T$ is computationally expensive to obtain and inflexible in dimensionality. To address this problem, we introduce Diffusion Bridge AuteEncoders (DBAE), which enables $\mathbf{z}$-dependent endpoint $\mathbf{x}_T$ inference through a feed-forward architecture. This structure creates an information bottleneck at $\mathbf{z}$, so $\mathbf{x}_T$ becomes dependent on $\mathbf{z}$ in its generation. This results in $\mathbf{z}$ holding the full information of data. We propose an objective function for DBAE to enable both reconstruction and generative modeling, with their theoretical justification. Empirical evidence supports the effectiveness of the intended design in DBAE, which notably enhances downstream inference quality, reconstruction, and disentanglement. Additionally, DBAE generates high-fidelity samples in the unconditional generation. Our code is
available at https://github.com/aailab-kaist/DBAE. Yeongmin Kim, Kwanghyeon Lee, Minsang Park, Byeonghu Na, Il-Chul Moon |
ICLR | 1 |
| 2025 | Global Context-aware Representation Learning for Spatially Resolved TranscriptomicsabstractSpatially Resolved Transcriptomics (SRT) is a cutting-edge technique that captures the spatial context of cells within tissues, enabling the study of complex biological networks. Recent graph-based methods leverage both gene expression and spatial information to identify relevant spatial domains. However, these approaches fall short in obtaining meaningful spot representations, especially for spots near spatial domain boundaries, as they heavily emphasize adjacent spots that have minimal feature differences from an anchor node. To address this, we propose Spotscape, a novel framework that introduces the Similarity Telescope module to capture global relationships between multiple spots. Additionally, we propose a similarity scaling strategy to regulate the distances between intra- and inter-slice spots, facilitating effective multi-slice integration. Extensive experiments demonstrate the superiority of Spotscape in various downstream tasks, including single-slice and multi-slice scenarios. Yunhak Oh, Junseok Lee 0002, Yeongmin Kim, Sangwoo Seo, Namkyeong Lee, Chanyoung Park 0001 |
ICML | 3 |
| 2025 | Preference Optimization by Estimating the Ratio of the Data DistributionabstractDirect preference optimization (DPO) is widely used as a simple and stable method for aligning large language models (LLMs) with human preferences. This paper investigates a generalized DPO loss that enables a policy model to match the target policy from a likelihood ratio estimation perspective. The ratio of the target policy provides a unique identification of the policy distribution without relying on reward models or partition functions. This allows the generalized loss to retain both simplicity and theoretical guarantees, which prior work such as $f$-PO fails to achieve simultaneously. We propose \textit{Bregman preference optimization} (BPO), a generalized framework for ratio matching that provides a family of objective functions achieving target policy optimality. BPO subsumes DPO as a special case and offers tractable forms for all instances, allowing implementation with a few lines of code. We further develop scaled Basu's power divergence (SBA), a gradient scaling method that can be used for BPO instances. The BPO framework complements other DPO variants and is applicable to target policies defined by these variants. In experiments, unlike other probabilistic loss extensions such as $f$-DPO or $f$-PO, which exhibits a trade-off between generation fidelity and diversity, instances of BPO improve both win rate and entropy compared with DPO. When applied to Llama-3-8B-Instruct, BPO achieves state-of-the-art performance among Llama-3-8B backbones, with a 55.9\% length-controlled win rate on AlpacaEval2. Project page: https://github.com/aailab-kaist/BPO. Yeongmin Kim, HeeSun Bae, Byeonghu Na, Il-Chul Moon |
NeurIPS | 1 |
| 2025 | Disentangling Hyperedges through the Lens of Category TheoryabstractDespite the promising results of disentangled representation learning in discovering latent patterns in graph-structured data, few studies have explored disentanglement for hypergraph-structured data.
Integrating hyperedge disentanglement into hypergraph neural networks enables models to leverage hidden hyperedge semantics, such as unannotated relations between nodes, that are associated with labels.
This paper presents an analysis of hyperedge disentanglement from a category-theoretical perspective and proposes a novel criterion for disentanglement derived from the naturality condition.
Our proof-of-concept model experimentally showed the potential of the proposed criterion by successfully capturing functional relations of genes (nodes) in genetic pathways (hyperedges). Yoonho Lee 0002, Junseok Lee 0002, Sangwoo Seo, Sungwon Kim 0002, Yeongmin Kim, Chanyoung Park 0001 |
NeurIPS | 5 |
| 2024 | Reward-based Input Construction for Cross-document Relation ExtractionabstractRelation extraction (RE) is a fundamental task in natural language processing, aiming to identify relations between target entities in text.While many RE methods are designed for a single sentence or document, cross-document RE has emerged to address relations across multiple long documents.Given the nature of long documents in cross-document RE, extracting document embeddings is challenging due to the length constraints of pre-trained language models.Therefore, we propose REward-based Input Construction (REIC), the first learningbased sentence selector for cross-document RE.REIC extracts sentences based on relational evidence, enabling the RE module to effectively infer relations.Since supervision of evidence sentences is generally unavailable, we train REIC using reinforcement learning with RE prediction scores as rewards.Experimental results demonstrate the superiority of our method over heuristic methods for different RE structures and backbones in crossdocument RE. Byeonghu Na, Suhyeon Jo, Yeongmin Kim, Il-Chul Moon |
ACL (1) | 3 |
| 2024 | Early Prediction of Depressive Episodes in Mood Disorders Using Circadian Rhythm Indicators and Deep LearningabstractThe early prediction of depressive mood episodes is crucial for effective intervention in patients with Major Depressive Disorder (MDD) and Bipolar Disorder (BD). This study explores a predictive framework leveraging digital phenotypic data collected from smartphones and smartwatches, with a focus on circadian rhythm indicators such as Dim Light Melatonin Onset (DLMO). Using data from 164 participants within the Mood Disorder Cohort Research Consortium in Korea, time-series features related to sleep, heart rate, activity levels, and light exposure were processed to predict mood episodes seven days in advance. Deep learning models, including LSTM, GRU, and an LSTM-GRU hybrid, were applied to analyze this data, with the GRU model achieving the highest recall (0.767) and the LSTM model displaying superior robustness across metrics. SHAP value analysis of DLMO-related variables further underscored the association between circadian rhythm disruptions and depressive episodes, with delayed wake-up times relative to ideal schedules linked to increased depressive symptoms. Our findings demonstrate the feasibility of using digital phenotypes for early detection of mood episodes. These results highlight the potential of automated monitoring systems in clinical practice, which enable proactive intervention strategies through continuous, objective monitoring of patient conditions. Byeongsu Kim, Minsu Chae, Yihyun Kim, Seokjin Kong, Yeongmin Kim, Taewon Jung, Jaegwon Jeong, SooHyun Park, Chul-Hyun Cho, Ji Won Yeom, Taek Lee, Heon-Jeong Lee, Hwa-Min Lee |
BIBM | 5 |
| 2024 | Training Unbiased Diffusion Models From Biased DatasetabstractWith significant advancements in diffusion models, addressing the potential risks of dataset bias becomes increasingly important. Since generated outputs directly suffer from dataset bias, mitigating latent bias becomes a key factor in improving sample quality and proportion. This paper proposes time-dependent importance reweighting to mitigate the bias for the diffusion models. We demonstrate that the time-dependent density ratio becomes more precise than previous approaches, thereby minimizing error propagation in generative learning. While directly applying it to score-matching is intractable, we discover that using the time-dependent density ratio both for reweighting and score correction can lead to a tractable form of the objective function to regenerate the unbiased data density. Furthermore, we theoretically establish a connection with traditional score-matching, and we demonstrate its convergence to an unbiased distribution. The experimental evidence supports the usefulness of the proposed method, which outperforms baselines including time-independent importance reweighting on CIFAR-10, CIFAR-100, FFHQ, and CelebA with various bias settings. Our code is available at https://github.com/alsdudrla10/TIW-DSM. Yeongmin Kim, Byeonghu Na, Minsang Park, JoonHo Jang, Wanmo Kang, Il-Chul Moon |
ICLR | 1 |
| 2024 | Label-Noise Robust Diffusion ModelsabstractConditional diffusion models have shown remarkable performance in various generative tasks, but training them requires large-scale datasets that often contain noise in conditional inputs, a.k.a. noisy labels. This noise leads to condition mismatch and quality degradation of generated data. This paper proposes Transition-aware weighted Denoising Score Matching (TDSM) for training conditional diffusion models with noisy labels, which is the first study in the line of diffusion models. The TDSM objective contains a weighted sum of score networks, incorporating instance-wise and time-dependent label transition probabilities. We introduce a transition-aware weight estimator, which leverages a time-dependent noisy-label classifier distinctively customized to the diffusion process. Through experiments across various datasets and noisy label settings, TDSM improves the quality of generated samples aligned with given conditions. Furthermore, our method improves generation performance even on prevalent benchmark datasets, which implies the potential noisy labels and their risk of generative model learning. Finally, we show the improved performance of TDSM on top of conventional noisy label corrections, which empirically proving its contribution as a part of label-noise robust generative models. Our code is available at: https://github.com/byeonghu-na/tdsm. Byeonghu Na, Yeongmin Kim, HeeSun Bae, Jung Hyun Lee, Se Jung Kwon, Wanmo Kang, Il-Chul Moon |
ICLR | 2 |
| 2024 | Diffusion Rejection SamplingabstractRecent advances in powerful pre-trained diffusion models encourage the development of methods to improve the sampling performance under well-trained diffusion models. This paper introduces Diffusion Rejection Sampling (DiffRS), which uses a rejection sampling scheme that aligns the sampling transition kernels with the true ones at each timestep. The proposed method can be viewed as a mechanism that evaluates the quality of samples at each intermediate timestep and refines them with varying effort depending on the sample. Theoretical analysis shows that DiffRS can achieve a tighter bound on sampling error compared to pre-trained models. Empirical results demonstrate the state-of-the-art performance of DiffRS on the benchmark datasets and the effectiveness of DiffRS for fast diffusion samplers and large-scale text-to-image diffusion models. Our code is available at https://github.com/aailabkaist/DiffRS. Byeonghu Na, Yeongmin Kim, Minsang Park, Donghyeok Shin, Wanmo Kang, Il-Chul Moon |
ICML | 2 |
| 2024 | Single-cell RNA sequencing data imputation using bi-level feature propagationabstractSingle-cell RNA sequencing (scRNA-seq) enables the exploration of cellular heterogeneity by analyzing gene expression profiles in complex tissues. However, scRNA-seq data often suffer from technical noise, dropout events and sparsity, hindering downstream analyses. Although existing works attempt to mitigate these issues by utilizing graph structures for data denoising, they involve the risk of propagating noise and fall short of fully leveraging the inherent data relationships, relying mainly on one of cell-cell or gene-gene associations and graphs constructed by initial noisy data. To this end, this study presents single-cell bilevel feature propagation (scBFP), two-step graph-based feature propagation method. It initially imputes zero values using non-zero values, ensuring that the imputation process does not affect the non-zero values due to dropout. Subsequently, it denoises the entire dataset by leveraging gene-gene and cell-cell relationships in the respective steps. Extensive experimental results on scRNA-seq data demonstrate the effectiveness of scBFP in various downstream tasks, uncovering valuable biological insights. Junseok Lee 0002, Sukwon Yun, Yeongmin Kim, Tianlong Chen 0001, Manolis Kellis, Chanyoung Park 0001 |
Briefings Bioinform. | 3 |
| 2023 | SAAL: Sharpness-Aware Active LearningabstractWhile deep neural networks play significant roles in many research areas, they are also prone to overfitting problems under limited data instances. To overcome overfitting, this paper introduces the first active learning method to incorporate the sharpness of loss space into the acquisition function. Specifically, our proposed method, Sharpness-Aware Active Learning (SAAL), constructs its acquisition function by selecting unlabeled instances whose perturbed loss becomes maximum. Unlike the Sharpness-Aware learning with fully-labeled datasets, we design a pseudo-labeling mechanism to anticipate the perturbed loss w.r.t. the ground-truth label, which we provide the theoretical bound for the optimization. We conduct experiments on various benchmark datasets for vision-based tasks in image classification, object detection, and domain adaptive semantic segmentation. The experimental results confirm that SAAL outperforms the baselines by selecting instances that have the potentially maximal perturbation on the loss. The code is available at https://github.com/YoonyeongKim/SAAL. Yoon-Yeong Kim, Youngjae Cho 0002, JoonHo Jang, Byeonghu Na, Yeongmin Kim, Kyungwoo Song, Wanmo Kang, Il-Chul Moon |
ICML | 5 |
| 2023 | Refining Generative Process with Discriminator Guidance in Score-based Diffusion ModelsabstractThe proposed method, Discriminator Guidance, aims to improve sample generation of pre-trained diffusion models. The approach introduces a discriminator that gives explicit supervision to a denoising sample path whether it is realistic or not. Unlike GANs, our approach does not require joint training of score and discriminator networks. Instead, we train the discriminator after score training, making discriminator training stable and fast to converge. In sample generation, we add an auxiliary term to the pre-trained score to deceive the discriminator. This term corrects the model score to the data score at the optimal discriminator, which implies that the discriminator helps better score estimation in a complementary way. Using our algorithm, we achive state-of-the-art results on ImageNet 256x256 with FID 1.83 and recall 0.64, similar to the validation data’s FID (1.68) and recall (0.66). We release the code at https://github.com/alsdudrla10/DG. Yeongmin Kim, Se Jung Kwon, Wanmo Kang, Il-Chul Moon |
ICML | 2 |
| 2022 | Evaluation and validation of an artificial neural network for predicting the performance of a desiccant coated heat exchangerabstractThis study presents an artificial neural network (ANN) model to predict the performance of a desiccant coated heat exchanger (DCHE) that is employed for dehumidification applications. The performance of DCHE were evaluated in terms of the water vapor removal capacity and coefficient of performance (COP). Different air-side parameters and waterside conditions are were used as input for the ANN. DCHE was developed by coating a finned tube heat exchanger with adsorbent powder. The finned tube heat exchanger has a dimension of 200 mm x 150 mm x 22 mm, fin thickness of 0.1 mm and spacing 1.5 mm, and 4 tube passes with tube diameter of 9.5 mm. From the previous experimental study, 146 data samples were employed for training the ANN model. MATLAB code was developed to study feed forward and back propagation. 75% of the experimental data was used to train the model and the remaining 15% and 15% were used to validate and test the model respectively. The results revealed that the maximum discrepancy between the ANN and experimental date for water vapor removal rate and COP were 0.05 and 0.01, respectively. Yeongmin Kim, Yoon Jung Ko, Seung Jin Oh |
IEEE Big Data | 1 |
| 2021 | Predict Sequential Credit Card Delinquency with VaDE-Seq2SeqabstractFor successful debt collection, it is important for credit card companies to judge the users’ capability of debt repayment. This has been assessed by domain experts in the past, but as the amount of data increases, there has been a rising demand for a more effective decision-making methods. Several machine learning algorithms have been proposed to pursue interpretation and high performance. We newly propose Variational Deep Embedding with Sequence to Sequence (VaDE-Seq2Seq), based on a deep neural network. By adding the VaDE structure to the encoder, the model properly reflects information on cluster assignments in latent space, and the model explains decision-making by tracking the cluster assignments. Most delinquency prediction studies predict only the next time step, whereas our model predicts the future sequence. It is a strength of our model because sequence prediction is difficult, but more practical. The model was tested with the data of 10,000 users from a Korean credit card company, and VaDE-Seq2Seq outperforms the other baseline models in terms of performance. In addition, we observe the history in the latent cluster assignment that was clearly distinguished between non-delinquency users and delinquency users. Yeongmin Kim, Youngjae Cho 0002, Hanbit Lee, Il-Chul Moon |
SMC | 1 |