Kyungjune Baek

dblp:223/5659 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Rethinking Direct Preference Optimization in Diffusion Models
abstract
Aligning text-to-image (T2I) diffusion models with human preferences has emerged as a critical research challenge. While Direct Preference Optimization (DPO) has established a foundation for preference learning in large language models (LLMs), its extension to diffusion models remains limited in alignment performance. In this work, we propose an enhanced version of Diffusion-DPO by introducing a stable reference model update strategy. This strategy facilitates the exploration of better alignment solutions while maintaining training stability. Moreover, we design a timestep-aware optimization strategy that further boosts performance by addressing preference learning imbalance across timesteps. Through the synergistic combination of our exploration and timestep-aware optimization, our method significantly improves the alignment performance of Diffusion-DPO on human preference evaluation benchmarks, achieving state-of-the-art results.
Junyong Kang, Seohyun Lim, Kyungjune Baek, Hyunjung Shim
AAAI3
2026 MomentMix Augmentation with Length-Aware DETR for Temporally Robust Moment Retrieval
abstract
Video Moment Retrieval (MR) aims to localize moments within a video based on a given natural language query. Given the prevalent use of platforms like YouTube for information retrieval, the demand for MR techniques is significantly growing. Recent DETR-based models have made notable advances in performance but still struggle with accurately localizing short moments. Through data analysis, we identified limited feature diversity in short moments, which motivated the development of MomentMix. MomentMix generates new short-moment samples by employing two augmentation strategies: ForegroundMix and BackgroundMix, each enhancing the ability to understand the query-relevant and irrelevant frames, respectively. Additionally, our analysis of prediction bias revealed that short moments particularly struggle with accurately predicting their center positions and length of moments. To address this, we propose a Length-Aware Decoder, which conditions length through a novel bipartite matching process. Our extensive studies demonstrate the efficacy of our length-aware approach, especially in localizing short moments, leading to improved overall performance. Our method surpasses state-of-the-art DETR-based methods on benchmark datasets, achieving the highest R1 and mAP on QVHighlights and the highest [email protected] on TACoS and Charades-STA (such as a 9.62% gain in [email protected] and an 16.9% gain in mAP average for QVHighlights). The code is available at https://github.com/sjpark5800/LA-DETR.
Seojeong Park, Jiho Choi, Kyungjune Baek, Hyunjung Shim
WACV3
2026 Probing intrinsic bias: Internal attention feature analysis for social bias evaluation in diffusion models
Hyeongmin Lee, Kyungjune Baek
Neurocomputing2
2026 Toward stable world models: Measuring and addressing world instability in generative environments
Soonwoo Kwon, Hyojun Go, Kyungjune Baek
Pattern Recognit.4
2025 Temporal Smoothness-Aware Rate-Distortion Optimized 4D Gaussian Splatting
abstract
Dynamic 4D Gaussian Splatting (4DGS) effectively extends the high-speed rendering capabilities of 3D Gaussian Splatting (3DGS) to represent volumetric videos. However, the large number of Gaussians, substantial temporal redundancies, and especially the absence of an entropy-aware compression framework result in large storage requirements. Consequently, this poses significant challenges for practical deployment, efficient edge-device processing, and data transmission. In this paper, we introduce a novel end-to-end RD-optimized compression framework tailored for 4DGS, aiming to enable flexible, high-fidelity rendering across varied computational platforms. Leveraging Fully Explicit Dynamic Gaussian Splatting (Ex4DGS), one of the state-of-the-art 4DGS methods, as our baseline, we start from the existing 3DGS compression methods for compatibility while effectively addressing additional challenges introduced by the temporal axis. In particular, instead of storing motion trajectories independently per point, we employ a wavelet transform to reflect the real-world smoothness prior, significantly enhancing storage efficiency. This approach yields significantly improved compression ratios and provides a user-controlled balance between compression efficiency and rendering quality. Extensive experiments demonstrate the effectiveness of our method, achieving up to 91$\times$ compression compared to the original Ex4DGS model while maintaining high visual fidelity. These results highlight the applicability of our framework for real-time dynamic scene rendering in diverse scenarios, from resource-constrained edge devices to high-performance environments. The source code is available at https://github.com/HyeongminLEE/RD4DGS.
Hyeongmin Lee, Kyungjune Baek
NeurIPS2
2022 Commonality in Natural Images Rescues GANs: Pretraining GANs with Generic and Privacy-free Synthetic Data
abstract
Transfer learning for GANs successfully improves generation performance under low-shot regimes. However, existing studies show that the pretrained model using a single benchmark dataset is not generalized to various target datasets. More importantly, the pretrained model can be vulnerable to copyright or privacy risks as membership inference attack advances. To resolve both issues, we propose an effective and unbiased data synthesizer, namely Primitives - PS, inspired by the generic characteristics of natural images. Specifically, we utilize 1) the generic statistics on the frequency magnitude spectrum, 2) the elementary shape (i.e., image composition via elementary shapes) for representing the structure information, and 3) the existence of saliency as prior. Since our synthesizer only considers the generic properties of natural images, the single model pretrained on our dataset can be consistently transferred to various target datasets, and even outperforms the previous methods pretrained with the natural images in terms of Fréchet inception distance. Extensive analysis, ablation study, and evaluations demonstrate that each component of our data synthesizer is effective, and provide insights on the desirable nature of the pretrained model for the transferability of GANs.
Kyungjune Baek, Hyunjung Shim
CVPR1
2022 Learning from Better Supervision: Self-distillation for Learning with Noisy Labels
abstract
The remarkable performance of deep neural networks heavily rely on large-scale datasets with high-quality annotations. Since the data collection process such as web crawling naturally involves unreliable supervision (i.e., noisy label), handling samples with noisy labels has been actively studied. Existing methods in learning with noisy labels (LNL) 1) develop the sampling strategy for filtering out the noisy labels or 2) devise the robust loss function against noisy labels. As a result of these efforts, existing LNL models achieve impressive performance, recording a higher accuracy than the ratio of the clean samples in the dataset. Based on this observation, we propose a self-distillation framework to utilize the prediction of existing LNL models and further improve the performance via rectified distillation; hard pseudo label and feature distillation. Our rectified distillation can be easily applied to existing LNL models, thus we can enjoy their state-of-the-art performances. From extensive evaluations, we confirm that our model is effective on both synthetic and real noisy datasets with state-of-the-art performances on four benchmark datasets.
Kyungjune Baek, Hyunjung Shim
ICPR1
2022 Logit Mixing Training for More Reliable and Accurate Prediction
abstract
When a person solves the multi-choice problem, she considers not only what is the answer but also what is not the answer. Knowing what choice is not the answer and utilizing the relationships between choices, she can improve the prediction accuracy. Inspired by this human reasoning process, we propose a new training strategy to fully utilize inter-class relationships, namely LogitMix. Our strategy is combined with recent data augmentation techniques, e.g., Mixup, Manifold Mixup, CutMix, and PuzzleMix. Then, we suggest using a mixed logit, i.e., a mixture of two logits, as an auxiliary training objective. Since the logit can preserve both positive and negative inter-class relationships, it can impose a network to learn the probability of wrong answers correctly. Our extensive experimental results on the image- and language-based tasks demonstrate that LogitMix achieves state-of-the-art performance among recent data augmentation techniques regarding calibration error and prediction accuracy.
Duhyeon Bang, Kyungjune Baek, Yunho Jeon, Jin-Hwa Kim, Jongwuk Lee, Hyunjung Shim
IJCAI2
2021 Rethinking the Truly Unsupervised Image-to-Image Translation
abstract
Every recent image-to-image translation model inherently requires either image-level (i.e. input-output pairs) or set-level (i.e. domain labels) supervision. However, even set-level supervision can be a severe bottleneck for data collection in practice. In this paper, we tackle image-to-image translation in a fully unsupervised setting, i.e., neither paired images nor domain labels. To this end, we propose a truly unsupervised image-to-image translation model (TUNIT) that simultaneously learns to separate image domains and translates input images into the estimated domains. Experimental results show that our model achieves comparable or even better performance than the set-level supervised model trained with full labels, generalizes well on various datasets, and is robust against the choice of hyperparameters (e.g. the preset number of pseudo domains). Furthermore, TUNIT can be easily extended to semi-supervised learning with a few labeled data.
Kyungjune Baek, Yunjey Choi, Youngjung Uh, Jaejun Yoo 0001, Hyunjung Shim
ICCV1
2021 GridMix: Strong regularization through local context mapping
Kyungjune Baek, Duhyeon Bang, Hyunjung Shim
Pattern Recognit.1
2020 PsyNet: Self-Supervised Approach to Object Localization Using Point Symmetric Transformation
abstract
Existing co-localization techniques significantly lose performance over weakly or fully supervised methods in accuracy and inference time. In this paper, we overcome common drawbacks of co-localization techniques by utilizing self-supervised learning approach. The major technical contributions of the proposed method are two-fold. 1) We devise a new geometric transformation, namely point symmetric transformation and utilize its parameters as an artificial label for self-supervised learning. This new transformation can also play the role of region-drop based regularization. 2) We suggest a heat map extraction method for computing the heat map from the network trained by self-supervision, namely class-agnostic activation mapping. It is done by computing the spatial attention map. Based on extensive evaluations, we observe that the proposed method records new state-of-the-art performance in three fine-grained datasets for unsupervised object localization. Moreover, we show that the idea of the proposed method can be adopted in a modified manner to solve the weakly supervised object localization task. As a result, we outperform the current state-of-the-art technique in weakly supervised object localization by a significant gap.
Kyungjune Baek, Minhyun Lee, Hyunjung Shim
AAAI1
2018 Editable Generative Adversarial Networks: Generating and Editing Faces Simultaneously
Kyungjune Baek, Duhyeon Bang, Hyunjung Shim
ACCV (1)1