EDBT 2026 Demo / reviewers in the wild / expert
Jeong-gi Kwak
dblp:278/3610
· DBLP profile ↗
11ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-7513-616XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-OffabstractVirtual try-on aims to synthesize a realistic image of a person wearing a target garment, but accurately modeling garment–body correspondence remains a persistent challenge, especially under pose and appearance variation. In this paper, we propose Voost—a unified and scalable framework that jointly learns virtual try-on and try-off with a single diffusion transformer. By modeling both tasks jointly, Voost enables each garment-person pair to supervise both directions and supports flexible conditioning over generation direction and garment category—enhancing garment–body relational reasoning without task-specific networks, auxiliary losses, or additional labels. In addition, we introduce two inference-time techniques: attention temperature scaling for robustness to resolution or mask variation, and self-corrective sampling that leverages bidirectional consistency between tasks. Extensive experiments demonstrate that Voost achieves state-of-the-art results on both try-on and try-off benchmarks, consistently outperforming strong baselines in alignment accuracy, visual fidelity, and generalization. Jeong-gi Kwak |
SIGGRAPH Asia | 2 |
| 2024 | ViVid-1-to-3: Novel View Synthesis with Video Diffusion ModelsabstractGenerating novel views of an object from a single image is a challenging task. It requires an understanding of the underlying 3D structure of the object from an image and ren-dering high-quality, spatially consistent new views. While recent methods for view synthesis based on diffusion have shown great progress, achieving consistency among various view estimates and at the same time abiding by the desired camera pose remains a critical problem yet to be solved. In this work, we demonstrate a strikingly simple method, where we utilize a pre-trained video diffusion model to solve this problem. Our key idea is that synthesizing a novel view could be reformulated as synthesizing a video of a cam-era going around the object of interest-a scanning video-which then allows us to leverage the powerful priors that a video diffusion model would have learned. Thus, to perform novel-view synthesis, we create a smooth camera trajectory to the target view that we wish to render, and denoise using both a view-conditioned diffusion model and a video diffusion model. By doing so, we obtain a highly consistent novel view synthesis, outperforming the state of the art. Jeong-gi Kwak, Erqun Dong, Yuhe Jin, Hanseok Ko, Shweta Mahajan, Kwang Moo Yi |
CVPR | 1 |
| 2024 | Towards Multi-Domain Face Landmark Detection with Synthetic Data from Diffusion ModelabstractRecently, deep learning-based facial landmark detection for in-the-wild faces has achieved significant improvement. However, there are still challenges in face landmark detection in other domains (e.g. cartoon, caricature, etc). This is due to the scarcity of extensively annotated training data. To tackle this concern, we design a two-stage training approach that effectively leverages limited datasets and the pre-trained diffusion model to obtain aligned pairs of landmarks and face in multiple domains. In the first stage, we train a landmark-conditioned face generation model on a large dataset of real faces. In the second stage, we fine-tune the above model on a small dataset of image-landmark pairs with text prompts for controlling the domain. Our new designs enable our method to generate high-quality synthetic paired datasets from multiple domains while preserving the alignment between landmarks and facial features. Finally, we fine-tuned a pre-trained face landmark detection model on the synthetic dataset to achieve multi-domain face landmark detection. Our qualitative and quantitative results demonstrate that our method outperforms existing methods on multi-domain face landmark detection. Yuanming Li, Gwantae Kim, Jeong-gi Kwak, Bonhwa Ku, Hanseok Ko |
ICASSP | 3 |
| 2024 | Towards high-fidelity facial UV map generation in real-world
Yuanming Li, Jeong-gi Kwak, Bonhwa Ku, David K. Han, Hanseok Ko |
Pattern Recognit. Lett. | 2 |
| 2022 | Injecting 3D Perception of Controllable NeRF-GAN into StyleGAN for Editable Portrait Image Synthesis
Jeong-gi Kwak, Yuanming Li, Dongsik Yoon, David K. Han, Hanseok Ko |
ECCV (17) | 1 |
| 2022 | DIFAI: Diverse Facial Inpainting using StyleGAN InversionabstractImage inpainting is an old problem in computer vision that restores occluded regions and completes damaged images. In the case of facial image inpainting, most of the methods generate only one result for each masked image, even though there are other reasonable possibilities. To prevent any potential biases and unnatural constraints stemming from generating only one image, we propose a novel framework for diverse facial inpainting exploiting the embedding space of StyleGAN. Our framework employs pSp encoder and SeFa algorithm to identify semantic components of the StyleGAN embeddings and feed them into our proposed SPARN decoder that adopts region normalization for plausible inpainting. We demonstrate that our proposed method outperforms several state-of-the-art methods. Dongsik Yoon, Jeong-gi Kwak, Yuanming Li, David K. Han, Hanseok Ko |
ICIP | 2 |
| 2022 | Efficient Dynamic Filter For Robust and Low Computational Feature ExtractionabstractThe unseen noise signal is difficult to anticipate, and various approaches have been developed to address this issue. In our earlier work, we proposed a lightweight dynamic filter by splitting the filter into kernel and spatial parts. This small footprint model showed robust results in an unseen noisy environment. However, a simple pooling process for dividing the feature would limit the performance. In this paper, we propose an efficient dynamic filter to enhance the performance of the existing dynamic filter. Instead of the simple feature mean, we separate the input features as non-overlapping chunks, and separable convolutions take place for each feature direction. We also propose a dynamic filter based attention pooling method. These methods are applied to the kernel part in our previous work, and experiments are carried out for keyword spotting and speaker verification. We confirm that our proposed method performs better in unseen environments than the recently developed models. Jeong-gi Kwak, Hanseok Ko |
SLT | 2 |
| 2021 | Adverse Weather Image Translation with Asymmetric and Uncertainty-aware GAN
Jeong-gi Kwak, Youngsaeng Jin, Yuanming Li, Dongsik Yoon, Hanseok Ko |
BMVC | 1 |
| 2021 | Adaptive Content Feature Enhancement GAN for Multimodal Selfie to Anime Translation
Yuanming Li, Jeong-gi Kwak, Dongsik Yoon, Youngsaeng Jin, David K. Han, Hanseok Ko |
BMVC | 2 |
| 2021 | Reference Guided Image Inpainting using Facial Attributes
Dongsik Yoon, Youngsaeng Jin, Jeong-gi Kwak, Yuanming Li, David K. Han, Hanseok Ko |
BMVC | 3 |
| 2020 | CAFE-GAN: Arbitrary Face Attribute Editing with Complementary Attention Feature
Jeong-gi Kwak, David K. Han, Hanseok Ko |
ECCV (14) | 1 |