Sangkyung Kwak

dblp:345/0067 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0001-9145-5876ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 41% Transfer learning and domain adaptation · 26% Trustworthy machine learning · 18%
Computer graphics and multimedia
3 papers
Visual content generation and editing · 62% Multimedia analysis and retrieval · 38%

Topics — the 15 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
3.342025
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think · ICLR 2025
Controllable Human Image Generation with Personalized Multi-Garments · CVPR 2025
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion models · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation
fine-tuning
1.622025
StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment · IJCAI 2025
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion models · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
text-to-image generation
1.622025
Controllable Human Image Generation with Personalized Multi-Garments · CVPR 2025
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion models · NeurIPS 2024
Visual content generation and editing
virtual try-on
1.122025
Controllable Human Image Generation with Personalized Multi-Garments · CVPR 2025
Improving Diffusion Models for Authentic Virtual Try-on in the Wild · ECCV (86) 2024
Machine learning › Representation and self-supervised learning › representation matching
feature alignment
0.912025
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think · ICLR 2025
Machine learning › Transfer learning and domain adaptation › fine-tuning
robust fine-tuning
0.912025
StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment · IJCAI 2025
Machine learning › Trustworthy machine learning
robustness
0.912025
StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment · IJCAI 2025
Machine learning › Trustworthy machine learning › robustness
spurious correlation
0.912025
StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment · IJCAI 2025
Machine learning › Generative modeling › image generation
conditional image generation
0.812024
Improving Diffusion Models for Authentic Virtual Try-on in the Wild · ECCV (86) 2024
Machine learning › Trustworthy machine learning
consistency optimization
0.812024
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion models · NeurIPS 2024
Machine learning › Transfer learning and domain adaptation › model adaptation
model customization
0.812024
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion models · NeurIPS 2024
Machine learning › Efficient and distributed learning
model merging
0.812024
Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion models · NeurIPS 2024
Computer vision › Vision and language › vision-language model
CLIP
0.312025
StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment · IJCAI 2025
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot models
0.312025
StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment · IJCAI 2025
Interaction techniques and input › spatial interaction › navigation
video navigation
0.212023
Beyond Instructions: A Taxonomy of Information Types in How-to Videos · CHI 2023

Methods — techniques the papers use, named apart from their topics

diffusion model · 3.3synthetic data generation · 1.7perceptual similarity filtering · 1.7regularization · 1.7user study · 1.3open coding · 1.3representation alignment · 0.9language model · 0.9classifier-free guidance · 0.9fine-tuning objectives · 0.8direct consistency optimization · 0.8
YearPublicationVenuePosition
2025 Controllable Human Image Generation with Personalized Multi-Garments
abstract
We present BootComp, a novel framework based on text-to-image diffusion models for controllable human image generation with multiple reference garments. Here, the main bottleneck is data acquisition for training: collecting a large-scale dataset of high-quality reference garment images per human subject is quite challenging, i.e., ideally, one needs to manually gather every single garment photograph worn by each human. To address this, we propose a data generation pipeline to construct a large synthetic dataset, consisting of human and multiple-garment pairs, by introducing a model to extract any reference garment images from each human image. To ensure data quality, we also propose a filtering strategy to remove undesirable generated data based on measuring perceptual similarities between the garment presented in human image and extracted garment. Finally, by utilizing the constructed synthetic dataset, we train a diffusion model having two parallel denoising paths that use multiple garment images as conditions to generate human images while preserving their fine-grained details. We further show the wide-applicability of our framework by adapting it to different types of reference-based generation in the fashion domain, including virtual try-on, and controllable human image generation with other conditions, e.g., pose, face, etc.
Yisol Choi, Sangkyung Kwak, Sihyun Yu, Hyungwon Choi, Jinwoo Shin
CVPR2
2025 Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
abstract
Recent studies have shown that the denoising process in (generative) diffusion models can induce meaningful (discriminative) representations inside the model, though the quality of these representations still lags behind those learned through recent self-supervised learning methods. We argue that one main bottleneck in training large-scale diffusion models for generation lies in effectively learning these representations. Moreover, training can be made easier by incorporating high-quality external visual representations, rather than relying solely on the diffusion models to learn them independently. We study this by introducing a straightforward regularization called REPresentation Alignment (REPA), which aligns the projections of noisy input hidden states in denoising networks with clean image representations obtained from external, pretrained visual encoders. The results are striking: our simple strategy yields significant improvements in both training efficiency and generation quality when applied to popular diffusion and flow-based transformers, such as DiTs and SiTs. For instance, our method can speed up SiT training by over 17.5$\times$, matching the performance (without classifier-free guidance) of a SiT-XL model trained for 7M steps in less than 400K steps. In terms of final generation quality, our approach achieves state-of-the-art results of FID=1.42 using classifier-free guidance with the guidance interval.
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, Saining Xie
ICLR2
2025 StarFT: Robust Fine-tuning of Zero-shot Models via Spuriosity Alignment
abstract
Learning robust representations from data often requires scale, which has led to the success of recent zero-shot models such as CLIP. However, the obtained robustness can easily be deteriorated when these models are fine-tuned on other downstream tasks (e.g., of smaller scales). Previous works often interpret this phenomenon in the context of domain shift, developing fine-tuning methods that aim to preserve the original domain as much as possible. However, in a different context, fine-tuned models with limited data are also prone to learning features that are spurious to humans, such as background or texture. In this paper, we propose StarFT (Spurious Textual Alignment Regularization), a novel framework for fine-tuning zero-shot models to enhance robustness by preventing them from learning spuriosity. We introduce a regularization that aligns the output distribution for spuriosity-injected labels with the original zero-shot model, ensuring that the model is not induced to extract irrelevant features further from these descriptions. We leverage recent language models to get such spuriosity-injected labels by generating alternative textual descriptions that highlight potentially confounding features. Extensive experiments validate the robust generalization of StarFT and its emerging properties: zero-shot group robustness and improved zero-shot classification. Notably, StarFT boosts both worst-group and average accuracy by 14.30% and 3.02%, respectively, in the Waterbirds group shift scenario, where other robust fine-tuning baselines show even degraded performance.
Jongheon Jeong, Sangkyung Kwak, Kyungmin Lee, Jinwoo Shin
IJCAI3
2024 Improving Diffusion Models for Authentic Virtual Try-on in the Wild
Yisol Choi, Sangkyung Kwak, Kyungmin Lee, Hyungwon Choi, Jinwoo Shin
ECCV (86)2
2024 Direct Consistency Optimization for Robust Customization of Text-to-Image Diffusion models
abstract
Text-to-image (T2I) diffusion models, when fine-tuned on a few personal images, can generate visuals with a high degree of consistency. However, such fine-tuned models are not robust; they often fail to compose with concepts of pretrained model or other fine-tuned models. To address this, we propose a novel fine-tuning objective, dubbed Direct Consistency Optimization, which controls the deviation between fine-tuning and pretrained models to retain the pretrained knowledge during fine-tuning. Through extensive experiments on subject and style customization, we demonstrate that our method positions itself on a superior Pareto frontier between subject (or style) consistency and image-text alignment over all previous baselines; it not only outperforms regular fine-tuning objective in image-text alignment, but also shows higher fidelity to the reference images than the method that fine-tunes with additional prior dataset. More importantly, the models fine-tuned with our method can be merged without interference, allowing us to generate custom subjects in a custom style by composing separately customized subject and style models. Notably, we show that our approach achieves better prompt fidelity and subject fidelity than those post-optimized for merging regular fine-tuned models.
Kyungmin Lee, Sangkyung Kwak, Kihyuk Sohn, Jinwoo Shin
NeurIPS2
2023 Beyond Instructions: A Taxonomy of Information Types in How-to Videos
abstract
How-to videos are rich in information—they not only give instructions but also provide justifications or descriptions. People seek different information to meet their needs, and identifying different types of information present in the video can improve access to the desired knowledge. Thus, we present a taxonomy of information types in how-to videos. Through an iterative open coding of 4k sentences in 48 videos, 21 information types under 8 categories emerged. The taxonomy represents diverse information types that instructors provide beyond instructions. We first show how our taxonomy can serve as an analytical framework for video navigation systems. Then, we demonstrate through a user study (n=9) how type-based navigation helps participants locate the information they needed. Finally, we discuss how the taxonomy enables a wide range of video-related tasks, such as video authoring, viewing, and analysis. To allow researchers to build upon our taxonomy, we release a dataset of 120 videos containing 9.9k sentences labeled using the taxonomy.
Saelyne Yang, Sangkyung Kwak, Juhoon Lee, Juho Kim 0001
CHI2