Suhyeon Lee 0004

dblp:342/2820 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2025
0009-0004-0840-0161ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 60% 3D vision · 18% Vision and language · 13%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 100%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 20 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.232024
Don't Play Favorites: Minority Guidance for Diffusion Models · ICLR 2024
Decomposed Diffusion Sampler for Accelerating Large-Scale Inverse Problems · ICLR 2024
Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion Models · ICCV 2023
Computer vision › 3D vision › novel view synthesis
multi-view video generation
0.912025
Reangle-A-Video: 4D Video Generation as Video-to-Video Translation · ICCV 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.912025
Reangle-A-Video: 4D Video Generation as Video-to-Video Translation · ICCV 2025
Machine learning › Generative modeling › diffusion model
diffusion sampling
0.812024
Decomposed Diffusion Sampler for Accelerating Large-Scale Inverse Problems · ICLR 2024
Machine learning › Generative modeling › diffusion model
guided sampling
0.812024
Don't Play Favorites: Minority Guidance for Diffusion Models · ICLR 2024
Natural language and speech › Language models and text generation
instruction tuning
0.812024
LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation · ICLR 2024
Machine learning › Generative modeling › diffusion model
inverse problem solving
0.812024
Decomposed Diffusion Sampler for Accelerating Large-Scale Inverse Problems · ICLR 2024
Computer vision › Vision and language › vision-language model › domain-specific vision-language model
medical vision-language model
0.812024
LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation · ICLR 2024
Machine learning › Generative modeling › synthetic data generation
minority sample generation
0.812024
Don't Play Favorites: Minority Guidance for Diffusion Models · ICLR 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.812024
LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation · ICLR 2024
Computer vision › 3D vision
3d reconstruction
0.712023
Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion Models · ICCV 2023
Machine learning › Generative modeling › diffusion model › diffusion prior
diffusion prior for 3d reconstruction
0.712023
Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion Models · ICCV 2023
Computer vision › 3D vision
medical image reconstruction
0.712023
Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion Models · ICCV 2023
Machine learning › Generative modeling
image generation
0.212024
LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation · ICLR 2024
Machine learning › Generative modeling › image generation
medical image synthesis
0.212024
LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation · ICLR 2024
Medical and health informatics › medical imaging › x-ray imaging
chest x-ray
0.212024
LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation · ICLR 2024
Medical and health informatics › medical imaging
medical image analysis
0.212024
LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation · ICLR 2024
Medical and health informatics
medical image reconstruction
0.212024
Decomposed Diffusion Sampler for Accelerating Large-Scale Inverse Problems · ICLR 2024
Image and video processing
compressive sensing
0.212023
Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion Models · ICCV 2023
Image and video processing
super-resolution
0.212023
Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion Models · ICCV 2023

Methods — techniques the papers use, named apart from their topics

vision-language alignment · 1.5tweedie's formula · 1.5krylov subspace method · 1.5instruction tuning · 1.5conjugate gradient · 1.5self-supervised fine-tuning · 0.9cross-view consistency guidance · 0.9tweedie's denoising formula · 0.8minority guidance · 0.8VQGAN · 0.8VQ-GAN · 0.8perpendicular 2d priors · 0.7diffusion model · 0.7
YearPublicationVenuePosition
2025 PlaceSim: An LLM-based Interactive Platform for Human Behavior Simulation in Physical Facilities
abstract
Physical facility design faces a fundamental cold-start problem: predicting human behavior in non-existent spaces. Traditional surveys and observational studies create gaps between stated preferences and actual usage, while existing simulation tools require significant technical expertise, limiting accessibility. We introduce PlaceSim, a web-based platform leveraging Large Language Models (LLMs) to simulate realistic human behavior in facilities through a zero-code interface. PlaceSim employs a Persona-Environment-Scenario (P.E.S.) framework that structures LLM reasoning through context-aware AI personas with transparent decision-making processes. The platform provides interactive facility design, AI-driven persona generation, live simulation with reasoning visualization, and what-if analysis for scenario comparison. Evaluated on 18 months of real-world apartment facility data (789,238 usage records from 8,435 residents), our zero-shot approach achieves Jensen-Shannon Divergence scores as low as 0.006, outperforming both supervised learning methods and existing LLM-based tools like SocioVerse without requiring training data. PlaceSim establishes new benchmarks for spatial behavior prediction while providing immediate, actionable insights for architects, urban planners, and facility managers through systematic simulation. The platform is available at https://simulation-viewer.vercel.app/
Suhyeon Lee 0004, Youngjun Yu, Donghyuk Shin, Rita Singh
CIKM1
2025 Reangle-A-Video: 4D Video Generation as Video-to-Video Translation
abstract
We introduce Reangle-A-Video, a unified framework for generating synchronized multi-view videos from a single input video. Unlike mainstream approaches that train multi-view video diffusion models on large-scale 4D datasets, our method reframes the multi-view video generation task as video-to-videos translation, leveraging publicly available image and video diffusion priors. In essence, Reangle-A-Video operates in two stages. (1) Multi-View Motion Learning: An image-to-video diffusion transformer is synchronously fine-tuned in a self-supervised manner to distill view-invariant motion from a set of warped videos. (2) Multi-View Consistent Image-to-Images Translation: The first frame of the input video is warped and inpainted into various camera perspectives under an inference-time cross-view consistency guidance using DUSt3R, generating multi-view consistent starting images. Extensive experiments on static view transport and dynamic camera control show that Reangle-A-Video surpasses existing methods, establishing a new solution for multi-view video generation. We will publicly release our code and data. Project page: https://hyeonho99.github.io/reangle-a-video/
Hyeonho Jeong, Suhyeon Lee 0004, Jong Chul Ye
ICCV2
2024 LLM-CXR: Instruction-Finetuned LLM for CXR Image Understanding and Generation
abstract
Following the impressive development of LLMs, vision-language alignment in LLMs is actively being researched to enable multimodal reasoning and visual input/output. This direction of research is particularly relevant to medical imaging because accurate medical image analysis and generation consist of a combination of reasoning based on visual features and prior knowledge. Many recent works have focused on training adapter networks that serve as an information bridge between image processing (encoding or generating) networks and LLMs; but presumably, in order to achieve maximum reasoning potential of LLMs on visual information as well, visual and language features should be allowed to interact more freely. This is especially important in the medical domain because understanding and generating medical images such as chest X-rays (CXR) require not only accurate visual and language-based reasoning but also a more intimate mapping between the two modalities. Thus, taking inspiration from previous work on the transformer and VQ-GAN combination for bidirectional image and text generation, we build upon this approach and develop a method for instruction-tuning an LLM pre-trained only on text to gain vision-language capabilities for medical images. Specifically, we leverage a pretrained LLM’s existing question-answering and instruction-following abilities to teach it to understand visual inputs by instructing it to answer questions about image inputs and, symmetrically, output both text and image responses appropriate to a given query by tuning the LLM with diverse tasks that encompass image-based text-generation and text-based image-generation. We show that our LLM-CXR trained in this approach shows better image-text alignment in both CXR understanding and generation tasks while being smaller in size compared to previously developed models that perform a narrower range of tasks.
Suhyeon Lee 0004, Won Jun Kim, Jinho Chang, Jong Chul Ye
ICLR1
2024 Decomposed Diffusion Sampler for Accelerating Large-Scale Inverse Problems
abstract
Krylov subspace, which is generated by multiplying a given vector by the matrix of a linear transformation and its successive powers, has been extensively studied in classical optimization literature to design algorithms that converge quickly for large linear inverse problems. For example, the conjugate gradient method (CG), one of the most popular Krylov subspace methods, is based on the idea of minimizing the residual error in the Krylov subspace. However, with the recent advancement of high-performance diffusion solvers for inverse problems, it is not clear how classical wisdom can be synergistically combined with modern diffusion models. In this study, we propose a novel and efficient diffusion sampling strategy that synergistically combines the diffusion sampling and Krylov subspace methods. Specifically, we prove that if the tangent space at a denoised sample by Tweedie's formula forms a Krylov subspace, then the CG initialized with the denoised data ensures the data consistency update to remain in the tangent space. This negates the need to compute the manifold-constrained gradient (MCG), leading to a more efficient diffusion sampling method. Our method is applicable regardless of the parametrization and setting (i.e., VE, VP). Notably, we achieve state-of-the-art reconstruction quality on challenging real-world medical inverse imaging problems, including multi-coil MRI reconstruction and 3D CT reconstruction. Moreover, our proposed method achieves more than 80 times faster inference time than the previous state-of-the-art method. Code is available at https://github.com/HJ-harry/DDS
Hyungjin Chung, Suhyeon Lee 0004, Jong Chul Ye
ICLR2
2024 Don't Play Favorites: Minority Guidance for Diffusion Models
abstract
We explore the problem of generating minority samples using diffusion models. The minority samples are instances that lie on low-density regions of a data manifold. Generating a sufficient number of such minority instances is important, since they often contain some unique attributes of the data. However, the conventional generation process of the diffusion models mostly yields majority samples (that lie on high-density regions of the manifold) due to their high likelihoods, making themselves ineffective and time-consuming for the minority generating task. In this work, we present a novel framework that can make the generation process of the diffusion models focus on the minority samples. We first highlight that Tweedie's denoising formula yields favorable results for majority samples. The observation motivates us to introduce a metric that describes the uniqueness of a given sample. To address the inherent preference of the diffusion models w.r.t. the majority samples, we further develop *minority guidance*, a sampling technique that can guide the generation process toward regions with desired likelihood levels. Experiments on benchmark real datasets demonstrate that our minority guidance can greatly improve the capability of generating high-quality minority samples over existing generative samplers. We showcase that the performance benefit of our framework persists even in demanding real-world scenarios such as medical imaging, further underscoring the practical significance of our work. Code is available at https://github.com/soobin-um/minority-guidance.
Soobin Um, Suhyeon Lee 0004, Jong Chul Ye
ICLR2
2023 Improving 3D Imaging with Pre-Trained Perpendicular 2D Diffusion Models
abstract
Diffusion models have become a popular approach for image generation and reconstruction due to their numerous advantages. However, most diffusion-based inverse problem-solving methods only deal with 2D images, and even recently published 3D methods do not fully exploit the 3D distribution prior. To address this, we propose a novel approach using two perpendicular pre-trained 2D diffusion models to solve the 3D inverse problem. By modeling the 3D data distribution as a product of 2D distributions sliced in different directions, our method effectively addresses the curse of dimensionality. Our experimental results demonstrate that our method is highly effective for 3D medical image reconstruction tasks, including MRI Z-axis super-resolution, compressed sensing MRI, and sparse-view CT. Our method can generate high-quality voxel volumes suitable for medical applications. The code is available at https://github.com/hyn2028/tpdm
Suhyeon Lee 0004, Hyungjin Chung, Jonghyuk Park 0006, Wi-Sun Ryu, Jong Chul Ye
ICCV1