VLDB 2026 Research / reviewers in the wild / expert
Kwan-Young Kim
dblp:120/3921 · also Kwanyoung Kim
· DBLP profile ↗
9ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0001-7508-7145ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 6 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward the Frontiers of Reliable Diffusion Sampling via Adversarial Sinkhorn Attention GuidanceabstractDiffusion models have demonstrated strong generative performance when using guidance methods such as classifier-free guidance (CFG), which enhance output quality by modifying the sampling trajectory. These methods typically improve a target output by intentionally degrading another, often the unconditional output, using heuristic perturbation functions such as identity mixing or blurred conditions. However, these approaches lack a principled foundation and rely on manually designed distortions. In this work, we propose Adversarial Sinkhorn Attention Guidance (ASAG), a novel method that reinterprets attention scores in diffusion models through the lens of optimal transport and intentionally increases the transport cost to disrupt unreliable attention flows. Instead of naively corrupting the attention mechanism, ASAG injects an adversarial cost within self-attention layers to reduce pixel-wise similarity between queries and keys. This deliberate degradation weakens misleading attention alignments and leads to improved conditional and unconditional sample quality. ASAG shows consistent improvements in text-to-image diffusion, and enhances controllability and fidelity in downstream applications such as IP-Adapter and ControlNet. The method is lightweight, plug-and-play, and improves reliability without requiring any model retraining. Kwan-Young Kim |
AAAI | 1 |
| 2025 | PLADIS: Pushing the Limits of Attention in Diffusion Models at Inference Time by Leveraging SparsityabstractDiffusion models have shown impressive results in generating high-quality conditional samples using guidance techniques such as Classifier-Free Guidance (CFG). However, existing methods often require additional training or neural function evaluations (NFEs), making them incompatible with guidance-distilled models. Also, they rely on heuristic approaches that need identifying target layers. In this work, we propose a novel and efficient method, termed PLADIS, which boosts pre-trained models (U-Net/Transformer) by leveraging sparse attention. Specifically, we extrapolate query-key correlations using softmax and its sparse counterpart in the cross-attention layer during inference, without requiring extra training or NFEs. By leveraging the noise robustness of sparse attention, our PLADIS unleashes the latent potential of text-to-image diffusion models, enabling them to excel in areas where they once struggled with newfound effectiveness. It integrates seamlessly with guidance techniques, including guidance-distilled models. Extensive experiments show notable improvements in text alignment and human preference, offering a highly efficient and universally applicable solution. See Our project page : https://cubeyoung.github.io/pladis-proejct/ Kwan-Young Kim, Byeongsu Sim |
ICCV | 1 |
| 2025 | End-to-end breast cancer radiotherapy planning via LMMs with consistency embeddingabstractRecent advances in AI foundation models have significant potential for lightening the clinical workload by mimicking the comprehensive and multi-faceted approaches used by medical professionals. In the field of radiation oncology, the integration of multiple modalities holds great importance, so the opportunity of foundational model is abundant. Inspired by this, here we present RO-LMM, a multi-purpose, comprehensive large multimodal model (LMM) tailored for the field of radiation oncology. This model effectively manages a series of tasks within the clinical workflow, including clinical context summarization, radiotherapy strategy suggestion, and plan-guided target volume segmentation by leveraging the capabilities of LMM. In particular, to perform consecutive clinical tasks without error accumulation, we present a novel Consistency Embedding Fine-Tuning (CEFTune) technique, which boosts LMM's robustness to noisy inputs while preserving the consistency of handling clean inputs. We further extend this concept to LMM-driven segmentation framework, leading to a novel Consistency Embedding Segmentation (CESEG) techniques. Experimental results including multi-center validation confirm that our RO-LMM with CEFTune and CESEG results in promising performance for multiple clinical tasks with generalization capabilities. Kwan-Young Kim, Yujin Oh, Sangjoon Park, Hwa Kyung Byun, Joongyo Lee, Yong Bae Kim, Jong Chul Ye |
Medical Image Anal. | 1 |
| 2024 | OTSeg: Multi-Prompt Sinkhorn Attention for Zero-Shot Semantic Segmentation
Kwan-Young Kim, Yujin Oh, Jong Chul Ye |
ECCV (77) | 1 |
| 2024 | Noise2one: One-Shot Image Denoising with Local Implicit LearningabstractRecently, the self-supervised learning paradigm, involving pretraining and fine-tuning large-scale models for downstream tasks, has shown promise in computer vision. Inspired by this, here we introduce Noise2One, a simple and effective image denoising method building upon this paradigm. Noise2One leverages self-supervised learning to pre-train a model on noisy images by learning a score function. Subsequently, fine-tuning on a single clean image enables denoising noisy images. Unlike Noise2Score, which relies on Tweedie’s formula, our method introduces a lightweight decoder based on Local Implicit Image Function (LIIF) for per-pixel noise adjustment and clean image reconstruction. This approach is more versatile, accommodating various noise models, including real-world noise. Our extensive experiments on benchmark datasets demonstrate that Noise2One achieves state-of-the-art denoising performance with only a 2.5% (+0.04M) increase in parameters compared to existing methods. Kwan-Young Kim, Jong Chul Ye |
ICASSP | 1 |
| 2024 | Unpaired Image-to-Image Translation via Neural Schrödinger BridgeabstractDiffusion models are a powerful class of generative models which simulate stochastic differential equations (SDEs) to generate data from noise. While diffusion models have achieved remarkable progress, they have limitations in unpaired image-to-image (I2I) translation tasks due to the Gaussian prior assumption. Schrödinger Bridge (SB), which learns an SDE to translate between two arbitrary distributions, have risen as an attractive solution to this problem. Yet, to our best knowledge, none of SB models so far have been successful at unpaired translation between high-resolution images. In this work, we propose Unpaired Neural Schrödinger Bridge (UNSB), which expresses the SB problem as a sequence of adversarial learning problems. This allows us to incorporate advanced discriminators and regularization to learn a SB between unpaired data. We show that UNSB is scalable and successfully solves various unpaired I2I translation tasks. Code: \url{https://github.com/cyclomon/UNSB} Gihyun Kwon, Kwan-Young Kim, Jong Chul Ye |
ICLR | 3 |
| 2022 | Noise Distribution Adaptive Self-Supervised Image Denoising using Tweedie Distribution and Score MatchingabstractTweedie distributions are a special case of exponential dispersion models, which are often used in classical statistics as distributions for generalized linear models. Here, we show that Tweedie distributions also play key roles in modern deep learning era, leading to a distribution adaptive self-supervised image denoising formula without clean reference images. Specifically, by combining with the recent Noise2Score self-supervised image denoising approach and the saddle point approximation of Tweedie distribution, we provide a general closed-form denoising formula that can be used for large classes of noise distributions without ever knowing the underlying noise distribution. Similar to the original Noise2Score, the new approach is composed of two successive steps: score matching using perturbed noisy images, followed by a closed form image denoising formula via distribution-independent Tweedie's formula. In addition, we reveal a systematic algorithm to estimate the noise model and noise parameters for a given noisy image data set. Through extensive experiments, we demonstrate that the proposed method can accurately estimate noise models and parameters, and provide the state-of-the-art self-supervised image denoising performance in the benchmark dataset and real-world dataset. Kwan-Young Kim, Taesung Kwon, Jong Chul Ye |
CVPR | 1 |
| 2021 | Task-Aware Variational Adversarial Active LearningabstractOften, labeling large amount of data is challenging due to high labeling cost limiting the application domain of deep learning techniques. Active learning (AL) tackles this by querying the most informative samples to be annotated among unlabeled pool. Two promising directions for AL that have been recently explored are task-agnostic approach to select data points that are far from the current labeled pool and task-aware approach that relies on the perspective of task model. Unfortunately, the former does not exploit structures from tasks and the latter does not seem to well-utilize overall data distribution. Here, we propose task-aware variational adversarial AL (TA-VAAL) that modifies task-agnostic VAAL, that considered data distribution of both label and unlabeled pools, by relaxing task learning loss prediction to ranking loss prediction and by using ranking conditional generative adversarial network to embed normalized ranking loss information on VAAL. Our proposed TA-VAAL outperforms state-of-the-arts on various benchmark datasets for classifications with balanced / imbalanced labels as well as semantic segmentation and its task-aware and task-agnostic AL properties were confirmed with our in-depth analyses. Kwan-Young Kim, Dongwon Park, Kwang In Kim, Se Young Chun |
CVPR | 1 |
| 2021 | Noise2Score: Tweedie's Approach to Self-Supervised Image Denoising without Clean ImagesabstractRecently, there has been extensive research interest in training deep networks to denoise images without clean reference.However, the representative approaches such as Noise2Noise, Noise2Void, Stein's unbiased risk estimator (SURE), etc. seem to differ from one another and it is difficult to find the coherent mathematical structure. To address this, here we present a novel approach, called Noise2Score, which reveals a missing link in order to unite these seemingly different approaches.Specifically, we show that image denoising problems without clean images can be addressed by finding the mode of the posterior distribution and that the Tweedie's formula offers an explicit solution through the score function (i.e. the gradient of loglikelihood). Our method then uses the recent finding that the score function can be stably estimated from the noisy images using the amortized residual denoising autoencoder, the method of which is closely related to Noise2Noise or Nose2Void. Our Noise2Score approach is so universal that the same network training can be used to remove noises from images that are corrupted by any exponential family distributions and noise parameters. Using extensive experiments with Gaussian, Poisson, and Gamma noises, we show that Noise2Score significantly outperforms the state-of-the-art self-supervised denoising methods in the benchmark data set such as (C)BSD68, Set12, and Kodak, etc. Kwan-Young Kim, Jong Chul Ye |
NeurIPS | 1 |