Chaehun Shin

dblp:287/9294 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-5299-9918ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Style-Friendly SNR Sampler for Style-Driven Generation
abstract
Recent text-to-image diffusion models generate high-quality images but struggle to learn new styles, which limits the personalized content creation. In response, style-driven generation has become a popular task, wherein users supply reference images capturing the target style, complemented by text prompts that specify stylistic cues. Fine-tuning is a common approach, yet it often blindly utilizes pre-training configurations without modification, especially for noise schedules defined in terms of signal-to-noise ratio (SNR), which determines the amount of image information available at each denoising step. We discover that stylistic features predominantly emerge at low SNR range, leading current fine-tuning methods using regular noise schedules to exhibit suboptimal style alignment. We propose the Style-friendly SNR sampler, which focuses the fine-tuning on low SNR range where stylistic features emerge. We demonstrate improved generation of novel styles that cannot be described solely with a text prompt, enabling high-fidelity personalized content creation.
Jooyoung Choi 0001, Chaehun Shin, Yeongtak Oh, Heeseung Kim, Jungbeom Lee, Sungroh Yoon
WACV2
2026 Exploring the transformer-based and diffusion-based models for single image deblurring
Seunghwan Park, Chaehun Shin, Jaihyun Lew, Sungroh Yoon
J. Vis. Commun. Image Represent.2
2025 Disentangled Motion Modeling for Video Frame Interpolation
abstract
Video Frame Interpolation (VFI) aims to synthesize intermediate frames between existing frames to enhance visual smoothness and quality. Beyond the conventional methods based on the reconstruction loss, recent works have employed generative models for improved perceptual quality. However, they require complex training and large computational costs for pixel space modeling. In this paper, we introduce disentangled Motion Modeling (MoMo), a diffusion-based approach for VFI that enhances visual quality by focusing on intermediate motion modeling. We propose a disentangled two-stage training process. In the initial stage, frame synthesis and flow models are trained to generate accurate frames and flows optimal for synthesis. In the subsequent stage, we introduce a motion diffusion model, which incorporates our novel U-Net architecture specifically designed for optical flow, to generate bi-directional flows between frames. By learning the simpler low-frequency representation of motions, MoMo achieves superior perceptual quality with reduced computational demands compared to the generative modeling methods on the pixel space. MoMo surpasses state-of-the-art methods in perceptual metrics across various benchmarks, demonstrating its efficacy and efficiency in VFI.
Jaihyun Lew, Jooyoung Choi 0001, Chaehun Shin, Dahuin Jung, Sungroh Yoon
AAAI3
2025 Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
abstract
Subject-driven text-to-image generation aims to produce images of a new subject within a desired context by accurately capturing both the visual characteristics of the subject and the semantic content of a text prompt. Traditional methods rely on time- and resource-intensive fine-tuning for subject alignment, while recent zero-shot approaches leverage on-the-fly image prompting, often sacrificing subject alignment. In this paper, we introduce Diptych Prompting, a novel zero-shot approach that reinterprets as an inpainting task with precise subject alignment by leveraging the emergent property of diptych generation in large-scale text-to-image models. Diptych Prompting arranges an incomplete diptych with the reference image in the left panel, and performs text-conditioned inpainting on the right panel. We further prevent unwanted content leakage by removing the background in the reference image and improve finegrained details in the generated subject by enhancing attention weights between the panels during inpainting. Experimental results confirm that our approach significantly outperforms zero-shot image prompting methods, resulting in images that are visually preferred by users. Additionally, our method supports not only subject-driven generation but also stylized image generation and subject-driven image editing, demonstrating versatility across diverse image generation applications.
Chaehun Shin, Jooyoung Choi 0001, Heeseung Kim, Sungroh Yoon
CVPR1
2025 Diffusion-Stego: Training-free diffusion generative steganography via message projection
Daegyu Kim, Chaehun Shin, Jooyoung Choi 0001, Dahuin Jung, Sungroh Yoon
Inf. Sci.2
2024 ControlDreamer: Blending Geometry and Style in Text-to-3D
Yeongtak Oh, Jooyoung Choi 0001, Yongsung Kim, Minjun Park, Chaehun Shin, Sungroh Yoon
BMVC5
2023 Edit-A-Video: Single Video Editing with Object-Aware Consistency
Chaehun Shin, Heeseung Kim, Che Hyun Lee, Sang-gil Lee, Sungroh Yoon
ACML1
2022 Perception Prioritized Training of Diffusion Models
abstract
Diffusion models learn to restore noisy data, which is corrupted with different levels of noise, by optimizing the weighted sum of the corresponding loss terms, i.e., denoising score matching loss. In this paper, we show that restoring data corrupted with certain noise levels offers a proper pretext task for the model to learn rich visual concepts. We propose to prioritize such noise levels over other levels during training, by redesigning the weighting scheme of the objective function. We show that our simple redesign of the weighting scheme significantly improves the performance of diffusion models regardless of the datasets, architectures, and sampling strategies.
Jooyoung Choi 0001, Jungbeom Lee, Chaehun Shin, Sungwon Kim 0001, Sungroh Yoon
CVPR3
2022 PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior
Sang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan 0003, Chang Liu 0030, Tao Qin 0001, Wei Chen 0034, Sungroh Yoon, Tie-Yan Liu
ICLR3
2021 BBAM: Bounding Box Attribution Map for Weakly Supervised Semantic and Instance Segmentation
abstract
Weakly supervised segmentation methods using bounding box annotations focus on obtaining a pixel-level mask from each box containing an object. Existing methods typically depend on a class-agnostic mask generator, which operates on the low-level information intrinsic to an image. In this work, we utilize higher-level information from the behavior of a trained object detector, by seeking the smallest areas of the image from which the object detector produces almost the same result as it does from the whole image. These areas constitute a bounding-box attribution map (BBAM), which identifies the target object in its bounding box and thus serves as pseudo ground-truth for weakly supervised semantic and instance segmentation. This approach significantly outperforms recent comparable techniques on both the PASCAL VOC and MS COCO benchmarks in weakly supervised semantic and instance segmentation. In addition, we provide a detailed analysis of our method, offering deeper insight into the behavior of the BBAM.
Jungbeom Lee, Jihun Yi, Chaehun Shin, Sungroh Yoon
CVPR3
2021 Personalized Face Authentication Based On Few-Shot Meta-Learning
abstract
Existing face authentication methods rely heavily on the general feature extractor trained with face verification algorithms without considering the target identity. No consideration of the target identity causes inefficiency in that the face authentication works on a specific target identity. Personalized face authentication systems that integrate target identity information is demanded to yield an efficient and superior performance than current methods. With the help of fast adaptation in few-shot meta-learning, we propose a novel personalized Face Authentication based on few-shot Meta-Learning (FAML). FAML introduces a unique person-specific binary classification task and trains an initialization model that can evolve into the personalized face authentication model for any target identity with few data and a small number of iterations. We evaluate the FAML on two public datasets, and it outperforms the baselines by a large margin and demonstrates the practicability as a security application.
Chaehun Shin, Jangho Lee 0002, Byunggook Na, Sungroh Yoon
ICIP1