Seunghwan Choi

dblp:92/8604 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021
YearPublicationVenuePosition
2025 FLEX: Expert-level False-Less EXecution Metric for Text-to-SQL Benchmark
abstract
Heegyu Kim, Jeon Taeyang, SeungHwan Choi, Seungtaek Choi, Hyunsouk Cho. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Heegyu Kim, Taeyang Jeon, Seunghwan Choi, Seungtaek Choi, Hyunsouk Cho
NAACL (Long Papers)3
2025 Disentangling Subject-Irrelevant Elements in Personalized Text-to-Image Diffusion via Filtered Self-Distillation
abstract
Recent research has unveiled the development of customizing large-scale text-to-image models. These models bind a unique subject desired by a user to a specific token, using the token to generate the subject in various contexts. However, models from previous studies also bind elements unrelated to the subject's identity, such as common backgrounds or poses in the reference images. This often leads to conflicts between the token and the context of text prompts during inference, causing the model to fail to generate both the subject and the prompted context. In this work, we approach this issue from a data scarcity perspective and propose to augment the number of reference images through a novel self-distillation framework. Our framework selects high-quality samples from images generated by a teacher model and uses them in student training. Our framework can be applied to any models that suffer from the conflicts, and we demonstrate that our framework most effectively resolves the issue through comprehensive evaluations.
Seunghwan Choi, Jooyeol Yun, Jeonghoon Park, Jaegul Choo
WACV1
2023 Shortcut-V2V: Compression Framework for Video-to-Video Translation based on Temporal Redundancy Reduction
abstract
Video-to-video translation aims to generate video frames of a target domain from an input video. Despite its usefulness, the existing networks require enormous computations, necessitating their model compression for wide use. While there exist compression methods that improve computational efficiency in various image/video tasks, a generally-applicable compression method for video-to-video translation has not been studied much. In response, we present Shortcut-V2V, a general-purpose compression framework for video-to-video translation. Shortcut-V2V avoids full inference for every neighboring video frame by approximating the intermediate features of a current frame from those of the previous frame. Moreover, in our framework, a newly-proposed block called AdaBD adaptively blends and deforms features of neighboring frames, which makes more accurate predictions of the intermediate features possible. We conduct quantitative and qualitative evaluations using well-known video-to-video translation models on various tasks to demonstrate the general applicability of our framework. The results show that Shortcut-V2V achieves comparable performance compared to the original video-to-video translation model while saving 3.2-5.7× computational cost and 7.8-44× memory at test time. Our code and videos are available at https://shortcut-v2v.github.io/.
Chaeyeon Chung, Yeojeong Park, Seunghwan Choi, Munkhsoyol Ganbat, Jaegul Choo
ICCV3
2023 Learning to Generate Semantic Layouts for Higher Text-Image Correspondence in Text-to-Image Synthesis
abstract
Existing text-to-image generation approaches have set high standards for photorealism and text-image correspondence, largely benefiting from web-scale text-image datasets, which can include up to 5 billion pairs. However, text-to-image generation models trained on domain-specific datasets, such as urban scenes, medical images, and faces, still suffer from low text-image correspondence due to the lack of text-image pairs. Additionally, collecting billions of text-image pairs for a specific domain can be time-consuming and costly. Thus, ensuring high text-image correspondence without relying on web-scale text-image datasets remains a challenging task. In this paper, we present a novel approach for enhancing text-image correspondence by leveraging available semantic layouts. Specifically, we propose a Gaussian-categorical diffusion process that simultaneously generates both images and corresponding layout pairs. Our experiments reveal that we can guide text-to-image generation models to be aware of the semantics of different image regions, by training the model to generate semantic labels for each pixel. We demonstrate that our approach achieves higher text-image correspondence compared to existing text-to-image generation approaches in the Multi-Modal CelebA-HQ and the Cityscapes dataset, where text-image pairs are scarce. Codes are available at https://pmh9960.github.io/research/GCDP.
Minho Park 0003, Jooyeol Yun, Seunghwan Choi, Jaegul Choo
ICCV3
2022 High-Resolution Virtual Try-On with Misalignment and Occlusion-Handled Conditions
Gyojung Gu, Sunghyun Park 0005, Seunghwan Choi, Jaegul Choo
ECCV (17)4
2021 HairFIT: Pose-invariant Hairstyle Transfer via Flow-based Hair Alignment and Semantic-region-aware Inpainting
Chaeyeon Chung, Hyelin Nam, Seunghwan Choi, Gyojung Gu, Sunghyun Park 0005, Jaegul Choo
BMVC4
2021 VITON-HD: High-Resolution Virtual Try-On via Misalignment-Aware Normalization
abstract
The task of image-based virtual try-on aims to transfer a target clothing item onto the corresponding region of a person, which is commonly tackled by fitting the item to the desired body part and fusing the warped item with the person. While an increasing number of studies have been conducted, the resolution of synthesized images is still limited to low (e.g., 256×192), which acts as the critical limitation against satisfying online consumers. We argue that the limitation stems from several challenges: as the resolution increases, the artifacts in the misaligned areas between the warped clothes and the desired clothing regions become noticeable in the final results; the architectures used in existing methods have low performance in generating high-quality body parts and maintaining the texture sharpness of the clothes. To address the challenges, we propose a novel virtual try-on method called VITON-HD that successfully synthesizes 1024×768 virtual try-on images. Specifically, we first prepare the segmentation map to guide our virtual try-on synthesis, and then roughly fit the target clothing item to a given person’s body. Next, we propose ALIgnment-Aware Segment (ALIAS) normalization and ALIAS generator to handle the misaligned areas and preserve the details of 1024×768 inputs. Through rigorous comparison with existing methods, we demonstrate that VITON-HD highly surpasses the baselines in terms of synthesized image quality both qualitatively and quantitatively.
Seunghwan Choi, Sunghyun Park 0005, Minsoo Lee, Jaegul Choo
CVPR1
2010 An Improved CIR-Based STR Scheme for MISO Mode in DVB-T2 System
abstract
We propose an improved CIR(Channel Impulse Response)-based fine STR(Symbol Timing Recovery) scheme for multi-input single-output (MISO) transmission mode of DVB-T2 system. At first, we present an efficient method to estimate the CIRs of MISO transmission. Then, the effective method for resolving the ambiguity of CIR is proposed by categorizing the ambiguity effect under the assumption of a false coarse STO(Symbol Timing Offset) and by changing the starting point of FFT window successively. Simulation results show that the proposed fine STR achieves the fine synchronization without an inter-symbol interference(ISI) in the large delay channels causing the ambiguity effect of CIR.
Seunghwan Choi, JongSeob Baek, Jong-Soo Seo
VTC Fall1