Yeji Song

dblp:271/0101 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-5436-5801ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Targeted Data Protection for Diffusion Model by Matching Training Trajectory
abstract
Recent advancements in diffusion models have made fine-tuning text-to-image models for personalization increasingly accessible, but have also raised significant concerns regarding unauthorized data usage and privacy infringement. Current protection methods are limited to passively degrading image quality, failing to achieve stable control. While Targeted Data Protection (TDP) offers a promising paradigm for active redirection toward user-specified target concepts, existing TDP attempts suffer from poor controllability due to snapshot-matching approaches that fail to account for complete learning dynamics. We introduce TAFAP (Trajectory Alignment via Fine-tuning with Adversarial Perturbations), the first method to successfully achieve effective TDP by controlling the entire training trajectory. Unlike snapshot-based methods whose protective influence is easily diluted as training progresses, TAFAP employs trajectory-matching inspired by dataset distillation to enforce persistent, verifiable transformations throughout fine-tuning. We validate our method through extensive experiments, demonstrating the first successful targeted transformation in diffusion models with simultaneous control over both identity and visual patterns. TAFAP significantly outperforms existing TDP attempts, achieving robust redirection toward target concepts while maintaining high image quality. This work enables verifiable safeguards and provides a new framework for controlling and tracing alterations in diffusion model outputs.
Hojun Lee 0002, Mijin Koo, Yeji Song, Nojun Kwak
AAAI3
2025 Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
abstract
In a surge of text-to-image (T2I) models and their customization methods that generate new images of a user-provided subject, current works focus on alleviating the costs incurred by a lengthy per-subject optimization. These zero-shot customization methods encode the image of a specified subject into a visual embedding which is then utilized alongside the textual embedding for diffusion guidance. The visual embedding incorporates intrinsic information about the subject, while the textual embedding provides a new context. However, the existing methods often 1) generate images with the same pose as an input image, and 2) exhibit deterioration in the subject's identity when facing a pose variation prompt. We first pin down the problem and show that redundant pose information in the visual embedding interferes with the pose indication in the textual embedding. Conversely, the textual embedding also harms the subject's identity which is tightly entangled with the pose in the visual embedding. As a remedy, we propose text-orthogonal visual embedding which effectively harmonizes with the given textual embedding. We also adopt the visual-only embedding and inject the subject's clear features utilizing a self-attention swap. Our method is both effective and robust, offering highly flexible zero-shot generation while effectively maintaining the subject's identity.
Yeji Song, Jimyeong Kim, Wonhark Park, Wonsik Shin, Wonjong Rhee, Nojun Kwak
AAAI1
2025 ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
abstract
Rectified Flow text-to-image models surpass diffusion models in image quality and text alignment, but adapting ReFlow for real-image editing remains challenging. We propose a new real-image editing method for ReFlow by analyzing the intermediate representations of multimodal transformer blocks and identifying three key features. To extract these features from real images with sufficient structural preservation, we leverage mid-step latent, which is inverted only up to the mid-step. We then adapt attention during injection to improve editability and enhance alignment to the target text. Our method is training-free, requires no user-provided mask, and can be applied even without a source prompt. Extensive experiments on two benchmarks with nine baselines demonstrate its superior performance over prior methods, further validated by human evaluations confirming a strong user preference for our approach.
Jimyeong Kim, Jungwon Park, Yeji Song, Nojun Kwak, Wonjong Rhee
ICCV3
2025 Differential privacy in statistical queries for synthetic trajectories generated by generative adversarial networks
abstract
With the widespread adoption of smartphones and the rapid advancement of information and communication technologies, the use of Location-Based Services (LBS) has significantly increased across various domains. Consequently, the collection and utilisation of user trajectory data are also growing rapidly. While such data can provide valuable insights for personalised services and other analyses, it inherently contains sensitive location information, posing serious privacy risks if used without proper anonymization. Previous studies have attempted to mitigate privacy concerns by applying Differential Privacy (DP) to prefix tree structures for statistical analysis. However, these approaches often suffer from diminished data utility due to the excessive noise required by DP mechanisms. To address this issue, we propose a two-stage trajectory privacy framework. In the first stage, we employ a Category Auxiliary Classifier-Generative Adversarial Network (CAC-GAN) to generate synthetic trajectory data that preserves the statistical characteristics of the original data, thereby providing primary privacy protection. In the second stage, we apply a prefix tree-based DP algorithm to the synthetic data, offering enhanced privacy during statistical analysis and query processing. Experimental results demonstrate that the proposed CAC-GAN method achieves approximately 53% improvement in both data utility and anonymity compared to existing methods. Furthermore, relative error analysis across various ϵ values confirms that our two-stage protection scheme maintains superior statistical accuracy. This study presents a novel methodology that effectively balances trajectory data privacy and utility.
Jihwan Shin, Yeji Song, Minsoo Jang, Taewhi Lee, Dong-Hyuk Im
Connect. Sci.2
2024 SAVE: Protagonist Diversification with Structure Agnostic Video Editing
Yeji Song, Wonsik Shin, Junsoo Lee 0002, Jeesoo Kim, Nojun Kwak
ECCV (80)1
2023 Except-Condition Generative Adversarial Network for Generating Trajectory Data
Yeji Song, Jihwan Shin, Taewhi Lee, Dong-Hyuk Im
DEXA (2)1
2023 Finding Efficient Pruned Network via Refined Gradients for Pruned Weights
abstract
With the growth of deep neural networks (DNN), the number of DNN parameters has drastically increased. This makes DNN models hard to be deployed on resource-limited embedded systems. To alleviate this problem, dynamic pruning methods have emerged, which try to find diverse sparsity patterns during training by utilizing Straight-Through-Estimator (STE) to approximate gradients of pruned weights. STE can help the pruned weights revive in the process of finding dynamic sparsity patterns. However, using these coarse gradients causes training instability and performance degradation owing to the unreliable gradient signal of the STE approximation. In this work, to tackle this issue, we introduce refined gradients to update the pruned weights by forming dual forwarding paths from two sets (pruned and unpruned) of weights. We propose a novel Dynamic Collective Intelligence Learning (DCIL) which makes use of the learning synergy between the collective intelligence of both weight sets. We verify the usefulness of the refined gradients by showing enhancements in the training stability and the model performance on the CIFAR and ImageNet datasets. DCIL outperforms various previously proposed pruning schemes including other dynamic pruning methods with enhanced stability during training. The code is provided in Github.
Jangho Kim, Jayeon Yoo, Yeji Song, KiYoon Yoo, Nojun Kwak
ACM Multimedia3
2022 Towards Efficient Neural Scene Graphs by Learning Consistency Fields
Yeji Song, Chaerin Kong, Seoyoung Lee 0001, Nojun Kwak, Joonseok Lee
BMVC1
2021 Part-Aware Data Augmentation for 3D Object Detection in Point Cloud
abstract
Data augmentation has greatly contributed to improving the performance in image recognition tasks, and a lot of related studies have been conducted. However, data augmentation on 3D point cloud data has not been much explored. 3D label has more sophisticated and rich structural information than the 2D label, so it enables more diverse and effective data augmentation. In this paper, we propose part-aware data augmentation (PA-AUG) that can better utilize rich information of 3D label to enhance the performance of 3D object detectors. PA-AUG divides objects into partitions and stochastically applies five augmentation methods to each local region. It is compatible with existing point cloud data augmentation methods and can be used universally regardless of the detector’s architecture. PA-AUG has improved the performance of state-of-the-art 3D object detector for all classes of the KITTI dataset and has the equivalent effect of increasing the train data by about 2.5×. We also show that PA-AUG not only increases performance for a given dataset but also is robust to corrupted data. The code is available at https://github.com/sky77764/pa-aug.pytorch
Jaeseok Choi, Yeji Song, Nojun Kwak
IROS2