VLDB 2026 Research / reviewers in the wild / expert
Jooyoung Choi 0001
dblp:170/0982-1
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0009-0009-3862-0639ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Style-Friendly SNR Sampler for Style-Driven GenerationabstractRecent text-to-image diffusion models generate high-quality images but struggle to learn new styles, which limits the personalized content creation. In response, style-driven generation has become a popular task, wherein users supply reference images capturing the target style, complemented by text prompts that specify stylistic cues. Fine-tuning is a common approach, yet it often blindly utilizes pre-training configurations without modification, especially for noise schedules defined in terms of signal-to-noise ratio (SNR), which determines the amount of image information available at each denoising step. We discover that stylistic features predominantly emerge at low SNR range, leading current fine-tuning methods using regular noise schedules to exhibit suboptimal style alignment. We propose the Style-friendly SNR sampler, which focuses the fine-tuning on low SNR range where stylistic features emerge. We demonstrate improved generation of novel styles that cannot be described solely with a text prompt, enabling high-fidelity personalized content creation. Jooyoung Choi 0001, Chaehun Shin, Yeongtak Oh, Heeseung Kim, Jungbeom Lee, Sungroh Yoon |
WACV | 1 |
| 2026 | GDoFS: Gaussian DoF Separation for Plausible 3D Geometry in Sparse-View 3DGSabstractWhile learning-based Multi-View Stereo (MVS) excels in sparse-view reconstruction, refining its output with 3D Gaussian Splatting (3DGS) remains challenging. Excessive positional degrees of freedom (DoFs) in Gaussians often cause instability and geometric artifacts, sometimes distorting geometry to represent texture patterns. To address this issue, we propose GDoFS (Gaussian DoF Separation), a strategy that divides positional DoFs into two categories—image-plane-parallel and ray-aligned—based on their uncertainty. For each category, GDoFS introduces tailored optimization techniques, including bounded offsets for low-uncertainty DoFs and a visibility-guided loss for ray-aligned DoFs. Experiments on standard benchmarks demonstrate that GDoFS effectively mitigates geometric artifacts and produces reconstructions that are both visually coherent and structurally accurate. Yongsung Kim, Jooyoung Choi 0001, Sungroh Yoon |
WACV | 2 |
| 2026 | Guiding What Not to Generate: Automated Negative Prompting for Text-Image AlignmentabstractDespite substantial progress in text–to–image generation, achieving precise text–image alignment remains challenging, particularly for prompts with rich compositional structure or imaginative elements. To address this, we introduce Negative Prompting for Image Correction (NPC), an automated pipeline that improves alignment by identifying and applying negative prompts that suppress unintended content. We begin by analyzing cross-attention patterns to explain why both targeted negatives—those directly tied to the prompt’s alignment error—and untargeted negatives—tokens unrelated to the prompt but present in the generated image—can enhance alignment. To discover useful negatives, NPC generates candidate prompts using a verifier–captioner–proposer framework and ranks them with a salient text-space score, enabling effective selection without requiring additional image synthesis. On GenEval++ and Imagine-Bench, NPC outperforms strong baselines, achieving 0.571 vs. 0.371 on GenEval++ and the best overall performance on Imagine-Bench. By guiding what not to generate, NPC provides a principled, fully automated route to stronger text–image alignment in diffusion models. Code is released at https://github.com/wiarae/NPC. Sangha Park, Eunji Kim 0002, Yeongtak Oh, Jooyoung Choi 0001, Sungroh Yoon |
WACV | 4 |
| 2026 | DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer StrategyabstractDespite recent text-to-image models achieving high-fidelity text rendering, they still struggle with long or multiple texts due to diluted global attention. We propose DC-Text, a training-free visual text generation method that adopts a divide-and-conquer strategy, leveraging the reliable short-text generation of Multi-Modal Diffusion Transformers. Our method first decomposes a prompt by extracting and dividing the target text, then assigns each to a designated region. To accurately render each segment within their regions while preserving overall image coherence, we introduce two attention masks—Text-Focus and Context-Expansion—applied sequentially during denoising. Additionally, Localized Noise Initialization further improves text accuracy and region alignment without increasing computational cost. Extensive experiments on single- and multi-sentence benchmarks show that DCText achieves the best text accuracy without compromising image quality while also delivering the lowest generation latency. Jaewoo Song, Jooyoung Choi 0001, Kanghyun Baek, Sangyub Lee 0005, Daemin Park, Sungroh Yoon |
WACV | 2 |
| 2025 | Disentangled Motion Modeling for Video Frame InterpolationabstractVideo Frame Interpolation (VFI) aims to synthesize intermediate frames between existing frames to enhance visual smoothness and quality. Beyond the conventional methods based on the reconstruction loss, recent works have employed generative models for improved perceptual quality. However, they require complex training and large computational costs for pixel space modeling. In this paper, we introduce disentangled Motion Modeling (MoMo), a diffusion-based approach for VFI that enhances visual quality by focusing on intermediate motion modeling. We propose a disentangled two-stage training process. In the initial stage, frame synthesis and flow models are trained to generate accurate frames and flows optimal for synthesis. In the subsequent stage, we introduce a motion diffusion model, which incorporates our novel U-Net architecture specifically designed for optical flow, to generate bi-directional flows between frames. By learning the simpler low-frequency representation of motions, MoMo achieves superior perceptual quality with reduced computational demands compared to the generative modeling methods on the pixel space. MoMo surpasses state-of-the-art methods in perceptual metrics across various benchmarks, demonstrating its efficacy and efficiency in VFI. Jaihyun Lew, Jooyoung Choi 0001, Chaehun Shin, Dahuin Jung, Sungroh Yoon |
AAAI | 2 |
| 2025 | Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image GeneratorabstractSubject-driven text-to-image generation aims to produce images of a new subject within a desired context by accurately capturing both the visual characteristics of the subject and the semantic content of a text prompt. Traditional methods rely on time- and resource-intensive fine-tuning for subject alignment, while recent zero-shot approaches leverage on-the-fly image prompting, often sacrificing subject alignment. In this paper, we introduce Diptych Prompting, a novel zero-shot approach that reinterprets as an inpainting task with precise subject alignment by leveraging the emergent property of diptych generation in large-scale text-to-image models. Diptych Prompting arranges an incomplete diptych with the reference image in the left panel, and performs text-conditioned inpainting on the right panel. We further prevent unwanted content leakage by removing the background in the reference image and improve finegrained details in the generated subject by enhancing attention weights between the panels during inpainting. Experimental results confirm that our approach significantly outperforms zero-shot image prompting methods, resulting in images that are visually preferred by users. Additionally, our method supports not only subject-driven generation but also stylized image generation and subject-driven image editing, demonstrating versatility across diverse image generation applications. Chaehun Shin, Jooyoung Choi 0001, Heeseung Kim, Sungroh Yoon |
CVPR | 2 |
| 2025 | DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual InspectionabstractDeveloping effective visual inspection models remains challenging due to the scarcity of defect data. While image generation models have been used to synthesize defect images, producing highly realistic defects remains difficult. We propose DefectFill, a novel method for realistic defect generation that requires only a few reference defect images. It leverages a fine-tuned inpainting diffusion model, optimized with our custom loss functions incorporating defect, object, and attention terms. It enables precise capture of detailed, localized defect features and their seamless integration into defect-free objects. Additionally, our Low-Fidelity Selection method further enhances the defect sample quality. Experiments show that DefectFill generates high-quality defect images, enabling visual inspection models to achieve state- of-the-art performance on the MVTec AD dataset. Jaewoo Song, Daemin Park, Kanghyun Baek, Sangyub Lee 0005, Jooyoung Choi 0001, Eunji Kim 0002, Sungroh Yoon |
CVPR | 5 |
| 2025 | NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple SpeakersabstractWe present NanoVoice, a personalized text-to-speech model that efficiently constructs voice adapters for multiple speakers simultaneously. NanoVoice introduces a batch-wise speaker adaptation technique capable of fine-tuning multiple references in parallel, significantly reducing training time. Beyond building separate adapters for each speaker, we also propose a parameter sharing technique that reduces the number of parameters used for speaker adaptation. By incorporating a novel trainable scale matrix, NanoVoice mitigates potential performance degradation during parameter sharing. NanoVoice achieves performance comparable to the baselines, while training 4 times faster and using 45 percent fewer parameters for speaker adaptation with 40 reference voices. Extensive ablation studies and analysis further validate the efficiency of our model. Nohil Park, Heeseung Kim, Che Hyun Lee, Jooyoung Choi 0001, Jiheum Yeom, Sungroh Yoon |
ICASSP | 4 |
| 2025 | VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via AutoguidanceabstractWhen applying parameter-efficient finetuning via LoRA onto speaker adaptive text-to-speech models, adaptation performance may decline compared to full-finetuned counterparts, especially for out-of-domain speakers. Here, we propose VoiceGuider, a parameter-efficient speaker adaptive text-to-speech system reinforced with autoguidance to enhance the speaker adaptation performance, reducing the gap against full-finetuned models. We carefully explore various ways of strengthening autoguidance, ultimately finding the optimal strategy. VoiceGuider as a result shows robust adaptation performance especially on extreme out-of-domain speech data. We provide audible samples in our demo page. Jiheum Yeom, Heeseung Kim, Jooyoung Choi 0001, Che Hyun Lee, Nohil Park, Sungroh Yoon |
ICASSP | 3 |
| 2025 | Diffusion-Stego: Training-free diffusion generative steganography via message projection
Daegyu Kim, Chaehun Shin, Jooyoung Choi 0001, Dahuin Jung, Sungroh Yoon |
Inf. Sci. | 3 |
| 2024 | ControlDreamer: Blending Geometry and Style in Text-to-3D
Yeongtak Oh, Jooyoung Choi 0001, Yongsung Kim, Minjun Park, Chaehun Shin, Sungroh Yoon |
BMVC | 2 |
| 2024 | Efficient Diffusion-Driven Corruption Editor for Test-Time Adaptation
Yeongtak Oh, Jonghyun Lee 0004, Jooyoung Choi 0001, Dahuin Jung, Uiwon Hwang, Sungroh Yoon |
ECCV (52) | 3 |
| 2022 | Perception Prioritized Training of Diffusion ModelsabstractDiffusion models learn to restore noisy data, which is corrupted with different levels of noise, by optimizing the weighted sum of the corresponding loss terms, i.e., denoising score matching loss. In this paper, we show that restoring data corrupted with certain noise levels offers a proper pretext task for the model to learn rich visual concepts. We propose to prioritize such noise levels over other levels during training, by redesigning the weighting scheme of the objective function. We show that our simple redesign of the weighting scheme significantly improves the performance of diffusion models regardless of the datasets, architectures, and sampling strategies. Jooyoung Choi 0001, Jungbeom Lee, Chaehun Shin, Sungwon Kim 0001, Sungroh Yoon |
CVPR | 1 |
| 2021 | ILVR: Conditioning Method for Denoising Diffusion Probabilistic ModelsabstractDenoising diffusion probabilistic models (DDPM) have shown remarkable performance in unconditional image generation. However, due to the stochasticity of the generative process in DDPM, it is challenging to generate images with the desired semantics. In this work, we propose Iterative Latent Variable Refinement (ILVR), a method to guide the generative process in DDPM to generate high-quality images based on a given reference image. Here, the refinement of the generative process in DDPM enables a single DDPM to sample images from various sets directed by the reference image. The proposed ILVR method generates high-quality images while controlling the generation. The controllability of our method allows adaptation of a single DDPM without any additional learning in various image generation tasks, such as generation from various downsampling factors, multi-domain image translation, paint-to-image, and editing with scribbles. Jooyoung Choi 0001, Sungwon Kim 0001, Yonghyun Jeong, Youngjune Gwon, Sungroh Yoon |
ICCV | 1 |
| 2021 | Toward Spatially Unbiased Generative ModelsabstractRecent image generation models show remarkable generation performance. However, they mirror strong location preference in datasets, which we call spatial bias. Therefore, generators render poor samples at unseen locations and scales. We argue that the generators rely on their implicit positional encoding to render spatial content. From our observations, the generator’s implicit positional encoding is translation-variant, making the generator spatially biased. To address this issue, we propose injecting explicit positional encoding at each scale of the generator. By learning the spatially unbiased generator, we facilitate the robust use of generators in multiple tasks, such as GAN inversion, multi-scale generation, generation of arbitrary sizes and aspect ratios. Furthermore, we show that our method can also be applied to denoising diffusion probabilistic models. Our code is available at: https://github.com/jychoill8/toward_spatial_unbiased. Jooyoung Choi 0001, Jungbeom Lee, Yonghyun Jeong, Sungroh Yoon |
ICCV | 1 |
| 2021 | Reducing Information Bottleneck for Weakly Supervised Semantic SegmentationabstractWeakly supervised semantic segmentation produces pixel-level localization from class labels; however, a classifier trained on such labels is likely to focus on a small discriminative region of the target object. We interpret this phenomenon using the information bottleneck principle: the final layer of a deep neural network, activated by the sigmoid or softmax activation functions, causes an information bottleneck, and as a result, only a subset of the task-relevant information is passed on to the output. We first support this argument through a simulated toy experiment and then propose a method to reduce the information bottleneck by removing the last activation function. In addition, we introduce a new pooling method that further encourages the transmission of information from non-discriminative regions to the classification. Our experimental evaluations demonstrate that this simple modification significantly improves the quality of localization maps on both the PASCAL VOC 2012 and MS COCO 2014 datasets, exhibiting a new state-of-the-art performance for weakly supervised semantic segmentation. Jungbeom Lee, Jooyoung Choi 0001, Jisoo Mok, Sungroh Yoon |
NeurIPS | 2 |