Jaihoon Kim

dblp:355/1743 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 68% 3D vision · 32%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 65% Geometric modeling and processing · 35%
Network and information security
1 paper
Privacy and data protection · 100%

Topics — the 7 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.622025
StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary Spaces · ICLR 2025
SyncTweedies: A General Generative Framework Based on Synchronized Diffusions · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
score distillation sampling
0.912025
StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary Spaces · ICLR 2025
Machine learning › Generative modeling › image generation › data-efficient image generation
zero-shot image generation
0.912025
StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary Spaces · ICLR 2025
Geometric modeling and processing › mesh processing
mesh texturing
0.912025
StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary Spaces · ICLR 2025
Visual content generation and editing › image generation › multi-view image generation
view-consistent generation
0.812024
SyncTweedies: A General Generative Framework Based on Synchronized Diffusions · NeurIPS 2024
Computer vision › 3D vision › visual localization
privacy-preserving visual localization
0.712023
Paired-Point Lifting for Enhanced Privacy-Preserving Visual Localization · CVPR 2023
Computer vision › 3D vision
visual localization
0.712023
Paired-Point Lifting for Enhanced Privacy-Preserving Visual Localization · CVPR 2023

Methods — techniques the papers use, named apart from their topics

diffusion synchronization · 3.3score distillation sampling · 1.7diffusion model · 1.7tweedie's formula · 1.5paired-point lifting · 1.33d line cloud construction · 1.3
YearPublicationVenuePosition
2026 Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models
abstract
Classifier-Free Guidance (CFG) is a fundamental technique in training conditional diffusion models. The common practice for CFG-based training is to use a single network to learn both conditional and unconditional noise prediction, with a small dropout rate for conditioning. However, we observe that the joint learning of unconditional noise with limited bandwidth in training results in poor priors for the unconditional case. More importantly, these poor unconditional noise predictions become a serious reason for degrading the quality of conditional generation. Inspired by the fact that most CFG-based conditional models are trained by fine-tuning a base model with better unconditional generation, we first show that simply replacing the unconditional noise in CFG with that predicted by the base model can significantly improve conditional generation. Furthermore, we show that a diffusion model other than the one the fine-tuned model was trained on can be used for unconditional noise replacement. We experimentally verify our claim with a range of CFG-based conditional models for both image and video generation, including Zero-1-to-3, Versatile Diffusion, DiT, DynamiCrafter, and InstructPix2Pix.
Prin Phunyaphibarn, Phillip Y. Lee, Jaihoon Kim, Minhyuk Sung
WACV3
2025 StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary Spaces
abstract
We propose a zero-shot method for generating images in arbitrary spaces (e.g., a sphere for 360◦ panoramas and a mesh surface for texture) using a pretrained image diffusion model. The zero-shot generation of various visual content using a pretrained image diffusion model has been explored mainly in two directions. First, Diffusion Synchronization–performing reverse diffusion processes jointly across different projected spaces while synchronizing them in the target space–generates high-quality outputs when enough conditioning is provided, but it struggles in its absence. Second, Score Distillation Sampling–gradually updating the target space data through gradient descent–results in better coherence but often lacks detail. In this paper, we reveal for the first time the interconnection between these two methods while highlighting their differences. To this end, we propose StochSync, a novel approach that combines the strengths of both, enabling effective performance with weak conditioning. Our experiments demonstrate that StochSync provides the best performance in 360◦ panorama generation (where image conditioning is not given), outperforming previous finetuning-based methods, and also delivers comparable results in 3D mesh texturing (where depth conditioning is provided) with previous methods.
Kyeongmin Yeo, Jaihoon Kim, Minhyuk Sung
ICLR2
2025 Moment- and Power-Spectrum-Based Gaussianity Regularization for Text-to-Image Models
abstract
We propose a novel regularization loss that enforces standard Gaussianity, encouraging samples to align with a standard Gaussian distribution. This facilitates a range of downstream tasks involving optimization in the latent space of text-to-image models. We treat elements of a high-dimensional sample as one-dimensional standard Gaussian variables and define a composite loss that combines moment-based regularization in the spatial domain with power spectrum-based regularization in the spectral domain. Since the expected values of moments and power spectrum distributions are analytically known, the loss promotes conformity to these properties. To ensure permutation invariance, the losses are applied to randomly permuted inputs. Notably, existing Gaussianity-based regularizations fall within our unified framework: some correspond to moment losses of specific orders, while the previous covariance-matching loss is equivalent to our spectral loss but incurs higher time complexity due to its spatial-domain computation. We showcase the application of our regularization in generative modeling for test-time reward alignment with a text-to-image model, specifically to enhance aesthetics and text alignment. Our regularization outperforms previous Gaussianity regularization, effectively prevents reward hacking and accelerates convergence.
Jisung Hwang, Jaihoon Kim, Minhyuk Sung
NeurIPS2
2025 Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing
abstract
We propose an inference-time scaling approach for pretrained flow models. Recently, inference-time scaling has gained significant attention in LLMs and diffusion models, improving sample quality or better aligning outputs with user preferences by leveraging additional computation. For diffusion models, particle sampling has allowed more efficient scaling due to the stochasticity at intermediate denoising steps. On the contrary, while flow models have gained popularity as an alternative to diffusion models--offering faster generation and high-quality outputs--efficient inference-time scaling methods used for diffusion models cannot be directly applied due to their deterministic generative process. To enable efficient inference-time scaling for flow models, we propose three key ideas: 1) SDE-based generation, enabling particle sampling in flow models, 2) Interpolant conversion, broadening the search space and enhancing sample diversity, and 3) Rollover Budget Forcing (RBF), an adaptive allocation of computational resources across timesteps to maximize budget utilization. Our experiments show that SDE-based generation and variance-preserving (VP) interpolant-based generation, improves the performance of particle sampling methods for inference-time scaling in flow models. Additionally, we demonstrate that RBF with VP-SDE achieves the best performance, outperforming all previous inference-time scaling approaches.
Jaihoon Kim, Taehoon Yoon, Jisung Hwang, Minhyuk Sung
NeurIPS1
2024 SyncTweedies: A General Generative Framework Based on Synchronized Diffusions
abstract
We introduce a general diffusion synchronization framework for generating diverse visual content, including ambiguous images, panorama images, 3D mesh textures, and 3D Gaussian splats textures, using a pretrained image diffusion model. We first present an analysis of various scenarios for synchronizing multiple diffusion processes through a canonical space. Based on the analysis, we introduce a synchronized diffusion method, SyncTweedies, which averages the outputs of Tweedie’s formula while conducting denoising in multiple instance spaces. Compared to previous work that achieves synchronization through finetuning, SyncTweedies is a zero-shot method that does not require any finetuning, preserving the rich prior of diffusion models trained on Internet-scale image datasets without overfitting to specific domains. We verify that SyncTweedies offers the broadest applicability to diverse applications and superior performance compared to the previous state-of-the-art for each application. Our project page is at https://synctweedies.github.io.
Jaihoon Kim, Juil Koo, Kyeongmin Yeo, Minhyuk Sung
NeurIPS1
2023 Paired-Point Lifting for Enhanced Privacy-Preserving Visual Localization
abstract
Visual localization refers to the process of recovering camera pose from input image relative to a known scene, forming a cornerstone of numerous vision and robotics systems. While many algorithms utilize sparse 3D point cloud of the scene obtained via structure-from-motion (SfM) for localization, recent studies have raised privacy concerns by successfully revealing high-fidelity appearance of the scene from such sparse 3D representation. One prominent approach for bypassing this attack was to lift 3D points to randomly oriented 3D lines thereby hiding scene geometry, but latest work have shown such random line cloud has a critical statistical flaw that can be exploited to break through protection. In this work, we present an alternative lightweight strategy called Paired-Point Lifting (PPL) for constructing 3D line clouds. Instead of drawing one randomly oriented line per 3D point, PPL splits 3D points into pairs and joins each pair to form 3D lines. This seemingly simple strategy yields 3 benefits, i) new ambiguity in feature selection, ii) increased line cloud sparsity and iii) nontrivial distribution of 3D lines, all of which contributes to enhanced protection against privacy attacks. Extensive experimental results demonstrate the strength of PPL in concealing scene details without compromising localization accuracy, unlocking the true potential of 3D line clouds.
Chunghwan Lee, Jaihoon Kim, Chanhyuk Yun, Je Hyeong Hong
CVPR2