VLDB 2026 Research / reviewers in the wild / expert
Min-Jung Kim 0001
dblp:13/1439-1
· DBLP profile ↗
8ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0003-3799-8225ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Spatiotemporal Skip Guidance for Enhanced Video Diffusion SamplingabstractDiffusion models have emerged as a powerful tool for generating high-quality images, videos, and 3D content. While sampling guidance techniques like CFG improve quality, they reduce diversity and motion. Autoguidance mitigates these issues but demands extra weak model training, limiting its practicality for large-scale models. In this work, we introduce Spatiotemporal Skip Guidance (STG), a simple training-free sampling guidance method for enhancing transformer-based video diffusion models. STG employs an implicit weak model via self-perturbation, avoiding the need for external models or additional training. By selectively skipping spatiotemporal layers, STG produces an aligned, degraded version of the original model to boost sample quality without compromising diversity or dynamic degree. Our contributions include: (1) introducing STG as an efficient, high-performing guidance technique for video diffusion models, (2) eliminating the need for auxiliary models by simulating a weak model through layer skipping, and (3) ensuring quality-enhanced guidance without compromising sample diversity or dynamics unlike CFG. For additional results, visit https://junhahyung.github.io/STGuidance/. Junha Hyung, Susung Hong, Min-Jung Kim 0001, Jaegul Choo |
CVPR | 4 |
| 2025 | SurFhead: Affine Rig Blending for Geometrically Accurate 2D Gaussian Surfel Head AvatarsabstractRecent advancements in head avatar rendering using Gaussian primitives have achieved significantly high-fidelity results. Although precise head geometry is crucial for applications like mesh reconstruction and relighting, current methods struggle to capture intricate geometric details and render unseen poses due to their reliance on similarity transformations, which cannot handle stretch and shear transforms essential for detailed deformations of geometry. To address this, we propose SurFhead, a novel method that reconstructs riggable head geometry from RGB videos using 2D Gaussian surfels, which offer well-defined geometric properties, such as precise depth from fixed ray intersections and normals derived from their surface orientation, making them advantageous over 3D counterparts. SurFhead ensures high-fidelity rendering of both normals and images, even in extreme poses, by leveraging classical mesh-based deformation transfer and affine transformation interpolation. SurFhead introduces precise geometric deformation and blends surfels through polar decomposition of transformations, including those affecting normals. Our key contribution lies in bridging classical graphics techniques, such as mesh-based deformation, with modern Gaussian primitives, achieving state-of-the-art geometry reconstruction and rendering quality. Unlike previous avatar rendering approaches, SurFhead enables efficient reconstruction driven by Gaussian primitives while preserving high-fidelity geometry. Taewoong Kang, Marcel C. Bühler, Min-Jung Kim 0001, Sungwon Hwang, Junha Hyung, Hyojin Jang, Jaegul Choo |
ICLR | 4 |
| 2024 | VEGS: View Extrapolation of Urban Scenes in 3D Gaussian Splatting Using Learned Priors
Sungwon Hwang, Min-Jung Kim 0001, Taewoong Kang, Jayeon Kang, Jaegul Choo |
ECCV (55) | 2 |
| 2024 | TCAN: Animating Human Images with Temporally Consistent Pose Guidance Using Diffusion Models
Jeongho Kim 0007, Min-Jung Kim 0001, Junsoo Lee 0002, Jaegul Choo |
ECCV (59) | 2 |
| 2024 | LensNeRF: Rethinking Volume Rendering based on Thin-Lens Camera ModelabstractRecent advances in Neural Radiance Field (NeRF) show promising results in rendering realistic novel view images. However, NeRF and its variants assume that input images are captured using a pinhole camera and that subjects in images are always all-in-focus by tacit agreement. In this paper, we propose aperture-aware NeRF optimization and rendering methods using a thin-lens model (dubbed LensNeRF), which allows defocus images of any aperture size as input and output. To generalize a pinhole camera model to a thin-lens camera model in NeRF framework, we define multiple rays originating from the aperture area, solving world-to-pixel scale ambiguity. Also, we propose in-focus loss that assigns the given pixel color to points on the focus plane to alleviate the color ambiguity caused by the use of multiple rays. For the rigorous evaluation of the proposed method, we collect a real forward-facing dataset with different F-numbers for each viewpoint. Experimental results demonstrate that our method successfully fuses an aperture-size adjustable thin-lens camera model into the NeRF architecture, showing favorable qualitative and quantitative results compared to baseline models. The dataset will be made available. Min-Jung Kim 0001, Gyojung Gu, Jaegul Choo |
WACV | 1 |
| 2023 | FaceCLIPNeRF: Text-driven 3D Face Manipulation using Deformable Neural Radiance FieldsabstractAs recent advances in Neural Radiance Fields (NeRF) have enabled high-fidelity 3D face reconstruction and novel view synthesis, its manipulation also became an essential task in 3D vision. However, existing manipulation methods require extensive human labor, such as a user-provided semantic mask and manual attribute search unsuitable for non-expert users. Instead, our approach is designed to require a single text to manipulate a face reconstructed with NeRF. To do so, we first train a scene manipulator, a latent code-conditional deformable NeRF, over a dynamic scene to control a face deformation using the latent code. However, representing a scene deformation with a single latent code is unfavorable for compositing local deformations observed in different instances. As so, our proposed Position-conditional Anchor Compositor (PAC) learns to represent a manipulated scene with spatially varying latent codes. Their renderings with the scene manipulator are then optimized to yield high cosine similarity to a target text in CLIP embedding space for text-driven manipulation. To the best of our knowledge, our approach is the first to address the text-driven manipulation of a face reconstructed with NeRF. Extensive results, comparisons, and ablation studies demonstrate the effectiveness of our approach. Sungwon Hwang, Junha Hyung, Min-Jung Kim 0001, Jaegul Choo |
ICCV | 4 |
| 2022 | VISOLO: Grid-Based Space-Time Aggregation for Efficient Online Video Instance SegmentationabstractFor online video instance segmentation (VIS), fully utilizing the information from previous frames in an efficient manner is essential for real-time applications. Most previous methods follow a two-stage approach requiring additional computations such as RPN and RoIAlign, and do not fully exploit the available information in the video for all subtasks in VIS. In this paper, we propose a novel single-stage framework for online VIS built based on the grid structured feature representation. The grid-based features allow us to employ fully convolutional networks for real-time processing, and also to easily reuse and share features within different components. We also introduce cooperatively operating modules that aggregate information from available frames, in order to enrich the features for all subtasks in VIS. Our design fully takes advantage of previous information in a grid form for all tasks in VIS in an efficient way, and we achieved the new state-of-the-art accuracy (38.6 AP and 36.9 AP) and speed (40.0 FPS) on YouTube-VIS 2019 and 2021 datasets among online VIS methods. The code is available at https://github.com/SuHoHan95/VISOLQ Su Ho Han, Sukjun Hwang, Seoung Wug Oh, Yeonchool Park, Min-Jung Kim 0001, Seon Joo Kim |
CVPR | 6 |
| 2014 | Cost-aware depth map estimation for Lytro cameraabstractSince commercial light field cameras became available, the light field camera has aroused much interest from computer vision and image processing communities due to its versatile functions. Most of its special features are based on an estimated depth map, so reliable depth estimation is a crucial step. However, estimating depth on real light field cameras is a challenging problem due to noise and short baselines among sub-aperture images. We propose a depth map estimation method for light field cameras by exploiting correspondence and focus cues. We aggregate costs among all the sub-aperture images on cost volume to alleviate noise effects. With efficiency of the cost volume, cost-aware depth estimation is quickly achieved by discrete-continuous optimization. In addition, we analyze each property of correspondence and focus cues and utilize them to select reliable anchor points. A well reconstructed initial depth map from the anchors is shown to enhance convergence. We show our method outperforms the state-of-the-art methods by validating it on real datasets acquired with a Lytro camera. Min-Jung Kim 0001, Tae-Hyun Oh, In-So Kweon |
ICIP | 1 |