Min-Gyu Park

dblp:122/3594 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
3D vision · 69% Trustworthy machine learning · 13% Generative modeling · 12%
Computer graphics and multimedia
4 papers
Visual content generation and editing · 68% Computational photography and imaging · 18% Geometric modeling and processing · 7%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
3d human reconstruction
1.422024
CanonicalFusion: Generating Drivable 3D Human Avatars from Multiple Images · ECCV (23) 2024
High-fidelity 3D Human Digitization from Single 2K Resolution Images · CVPR 2023
Computer vision › 3D vision › 3d reconstruction › object reconstruction
3d head reconstruction
0.912025
WarpHE4D: Dense 4D Head Map Toward Full Head Reconstruction · ICCV 2025
Visual content generation and editing › avatar generation
3d avatar creation
0.912025
Text2Avatar: Articulated 3D Avatar Creation With Text Instructions · IEEE Trans. Multim. 2025
Visual content generation and editing › 3d content creation
text-to-3d avatar generation
0.912025
Text2Avatar: Articulated 3D Avatar Creation With Text Instructions · IEEE Trans. Multim. 2025
Visual content generation and editing › avatar generation
3d human avatar generation
0.812024
CanonicalFusion: Generating Drivable 3D Human Avatars from Multiple Images · ECCV (23) 2024
Computational photography and imaging
depth estimation
0.712023
High-fidelity 3D Human Digitization from Single 2K Resolution Images · CVPR 2023
Machine learning › Trustworthy machine learning › uncertainty estimation
confidence estimation
0.622019
Learning and Selecting Confidence Measures for Robust Stereo Matching · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Leveraging stereo matching with learning-based confidence measures · CVPR 2015
Computer vision › 3D vision › stereo vision
stereo matching
0.622019
Learning and Selecting Confidence Measures for Robust Stereo Matching · IEEE Trans. Pattern Anal. Mach. Intell. 2019
Leveraging stereo matching with learning-based confidence measures · CVPR 2015
Machine learning › Generative modeling
generative adversarial network
0.412019
Predicting Future Frames Using Retrospective Cycle GAN · CVPR 2019
Machine learning › Generative modeling
video generation
0.412019
Predicting Future Frames Using Retrospective Cycle GAN · CVPR 2019
Computer vision › Video understanding and tracking
video prediction
0.412019
Predicting Future Frames Using Retrospective Cycle GAN · CVPR 2019
Computer vision › 3D vision
3d reconstruction
0.312017
Joint Layout Estimation and Global Multi-view Registration for Indoor Reconstruction · ICCV 2017
Computer vision › 3D vision › 3d scene reconstruction
indoor scene reconstruction
0.312017
Joint Layout Estimation and Global Multi-view Registration for Indoor Reconstruction · ICCV 2017
Computer vision › 3D vision › geometric estimation
registration
0.312017
Joint Layout Estimation and Global Multi-view Registration for Indoor Reconstruction · ICCV 2017
Computer vision › 3D vision › 3d scene understanding
room layout estimation
0.312017
Joint Layout Estimation and Global Multi-view Registration for Indoor Reconstruction · ICCV 2017
Computer vision › 3D vision
neural radiance field
0.312025
Text2Avatar: Articulated 3D Avatar Creation With Text Instructions · IEEE Trans. Multim. 2025
Virtual and augmented reality › avatar
avatar animation
0.312025
Text2Avatar: Articulated 3D Avatar Creation With Text Instructions · IEEE Trans. Multim. 2025
Geometric modeling and processing
deformable models
0.312025
WarpHE4D: Dense 4D Head Map Toward Full Head Reconstruction · ICCV 2025
Machine learning › Trustworthy machine learning
uncertainty estimation
0.212015
Leveraging stereo matching with learning-based confidence measures · CVPR 2015

Methods — techniques the papers use, named apart from their topics

warping · 1.7poisson reconstruction · 1.7linear blend skinning · 1.7NeRF · 1.74d representation · 1.7neural radiance field · 1.5part-wise image-to-normal network · 1.3multi-resolution depth network · 1.3SMPL · 1.3retrospective cycle consistency · 0.4
YearPublicationVenuePosition
2026 Towards robust 3D human reconstruction with uncertainty-aware low-rank adaptation
Inho Chang, Ju-Mi Kang, Yong-Hoon Kwon, Ju Hong Yoon, Min-Gyu Park
Comput. Vis. Image Underst.5
2025 WarpHE4D: Dense 4D Head Map Toward Full Head Reconstruction
Jong Seob Yun, Yong-Hoon Kwon, Min-Gyu Park, Ju-Mi Kang, Min-Ho Lee, Inho Chang, Ju Hong Yoon, Kuk-Jin Yoon
ICCV3
2025 Text2Avatar: Articulated 3D Avatar Creation With Text Instructions
abstract
We propose a framework for creating articulated human avatars, editing their styles, and animating the human avatars from three different types of text instructions. The three types of instructions, identity, edit, and action, are fed into three models that generate, edit, and animate human avatars. Specifically, the proposed framework takes identity instruction and multi-view pose condition images to generate the images of a human using the avatar generation model. Then, the avatar can be edited with text instructions by changing the style of the images generated. We apply the Neural Radiance Field (NeRF) and Poisson reconstruction to extract a human mesh model from images and assign linear blend skinning (LBS) weights to the vertices. Finally, the action instructions can animate human avatars, where we use the off-the-shelf method to generate the motions from text instructions. Notably, our proposed method adapts the appearance of hundreds of different individuals to construct a conditionally editable avatar-generated model, allowing easy creation of 3D avatars using text instructions. We demonstrate high-fidelity 3D animatable avatar creation with text instructions on various datasets and highlight a superior performance of the proposed method compared to the previous studies.
Yong-Hoon Kwon, Ju Hong Yoon, Min-Gyu Park
IEEE Trans. Multim.3
2024 CanonicalFusion: Generating Drivable 3D Human Avatars from Multiple Images
Jisu Shin 0002, Junmyeong Lee, Seongmin Lee 0009, Min-Gyu Park, Ju-Mi Kang, Ju Hong Yoon, Hae-Gon Jeon
ECCV (23)4
2023 High-fidelity 3D Human Digitization from Single 2K Resolution Images
abstract
High-quality 3D human body reconstruction requires high-fidelity and large-scale training data and appropriate network design that effectively exploits the high-resolution input images. To tackle these problems, we propose a simple yet effective 3D human digitization method called 2K2K, which constructs a large-scale 2K human dataset and infers 3D human models from 2K resolution images. The proposed method separately recovers the global shape of a human and its details. The low-resolution depth network predicts the global structure from a low-resolution image, and the part-wise image-to-normal network predicts the details of the 3D human body structure. The high-resolution depth network merges the global 3D shape and the detailed structures to infer the high-resolution front and back side depth maps. Finally, an off-the-shelf mesh generator reconstructs the full 3D human model, which are available at https://github.com/SangHunHan92/2K2K. In addition, we also provide 2,050 3D human models, including texture maps, 3D joints, and SMPL parameters for research purposes. In experiments, we demonstrate competitive performance over the recent works on various datasets.
Sang-Hun Han, Min-Gyu Park, Ju Hong Yoon, Ju-Mi Kang, Young-Jae Park, Hae-Gon Jeon
CVPR2
2019 Learning Depth from Endoscopic Images
abstract
We propose an unsupervised approach to predict depth maps from images captured by a wireless endoscopic capsule. Recent advances in deep learning have shown that accurate depth maps can be predicted from a single image, where the deep network is trained via unsupervised or self-supervised learning by using monocular video sequences or stereo image pairs. However, directly applying these techniques to endoscopic images does not yield satisfactory results owing to the inherent difficulties of the wireless capsule imaging such as dim lighting and low-resolution of images, which are different from normal imaging conditions. For that reason, we exploit the environmental characteristics of endoscopic images - there is no external light source except ones attached to the capsule. Based on this condition, we propose the direct attenuation model-based depth map prediction scheme to guide depth prediction and to add meaningful cues to the loss function. We experimentally verify the proposed method with various endoscopic images.
Ju Hong Yoon, Min-Gyu Park, Youngbae Hwang, Kuk-Jin Yoon
3DV2
2019 Predicting Future Frames Using Retrospective Cycle GAN
abstract
Recent advances in deep learning have significantly improved the performance of video prediction, however, top-performing algorithms start to generate blurry predictions as they attempt to predict farther future frames. In this paper, we propose a unified generative adversarial network for predicting accurate and temporally consistent future frames over time, even in a challenging environment. The key idea is to train a single generator that can predict both future and past frames while enforcing the consistency of bi-directional prediction using the retrospective cycle constraints. Moreover, we employ two discriminators not only to identify fake frames but also to distinguish fake contained image sequences from the real sequence. The latter discriminator, the sequence discriminator, plays a crucial role in predicting temporally consistent future frames. We experimentally verify the proposed framework using various real-world videos captured by car-mounted cameras, surveillance cameras, and arbitrary devices with state-of-the-art methods.
Yong-Hoon Kwon, Min-Gyu Park
CVPR2
2019 As-planar-as-possible depth map estimation
Min-Gyu Park, Kuk-Jin Yoon
Comput. Vis. Image Underst.1
2019 Learning and Selecting Confidence Measures for Robust Stereo Matching
abstract
We present a robust approach for computing disparity maps with a supervised learning-based confidence prediction. This approach takes into consideration following features. First, we analyze the characteristics of various confidence measures in the random forest framework to select effective confidence measures depending on the characteristics of the training data and matching strategies, such as similarity measures and parameters. We then train a random forest using the selected confidence measures to improve the efficiency of confidence prediction and to build a better prediction model. Second, we present a confidence-based matching cost modulation scheme, based on predicted confidence values, to improve the robustness and accuracy of the (semi-) global stereo matching algorithms. Finally, we apply the proposed modulation scheme to popularly used algorithms to make them robust against unexpected difficulties that could occur in an uncontrolled environment using challenging outdoor datasets. The proposed confidence measure selection and cost modulation schemes are experimentally verified from various perspectives using the KITTI and Middlebury datasets.
Min-Gyu Park, Kuk-Jin Yoon
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Memory-Efficient Parametric Semiglobal Matching
abstract
Accurate stereo matching for depth extraction requires a large memory space, which restricts its use in resource-limited systems. The problem is aggravated by the recent trend of applications requiring significantly high pixel resolution and disparity levels. To alleviate the high memory requirement, we propose to represent the aggregation costs as a Gaussian mixture model (GMM) function. Only a set of GMM parameters is stored and used instead of all the costs for each pixel. We also propose GMM parameter update-based aggregation along multiple paths. To preserve the accuracy of the disparity map, we employ a depth confidence measure and propose an update rule for the slanted surface of an object. Experimental results over the KITTI dataset show that the proposed method reduces the memory requirement to less than 5% of that of semiglobal matching, while the accuracy is maintained at the level of state-of-the-art semiglobal and local methods.
Yeongmin Lee, Min-Gyu Park, Youngbae Hwang, Youngsoo Shin, Chong-Min Kyung
IEEE Signal Process. Lett.2
2017 Joint Layout Estimation and Global Multi-view Registration for Indoor Reconstruction
abstract
In this paper, we propose a novel method to jointly solve scene layout estimation and global registration problems for accurate indoor 3D reconstruction. Given a sequence of range data, we first build a set of scene fragments using KinectFusion and register them through pose graph optimization. Afterwards, we alternate between layout estimation and layout-based global registration processes in iterative fashion to complement each other. We extract the scene layout through hierarchical agglomerative clustering and energy-based multi-model fitting in consideration of noisy measurements. Having the estimated scene layout in one hand, we register all the range data through the global iterative closest point algorithm where the positions of 3D points that belong to the layout such as walls and a ceiling are constrained to be close to the layout. We experimentally verify the proposed method with the publicly available synthetic and real-world datasets in both quantitative and qualitative ways.
Jeong-Kyun Lee, Jae-Won Yea, Min-Gyu Park, Kuk-Jin Yoon
ICCV3
2017 Learning to detect dynamic feature points
abstract
The detection of dynamic points on a moving platform is an important task to avoid a potential collision. However, it is difficult to detect dynamic points using only two frames, especially when various input data such as ego-motion, disparity map, and optical flow are noisy for computing the motion of points. In this paper, we propose a supervised learning-based approach to detect dynamic points in consideration of noisy input data. First of all, to consider depth ambiguity that proportionally increases according to the distance to the ego-vehicle, we divide the XZ-plane (bird-eye view) into several subregions. Then, we train a random forest for each subregion by constructing motion vectors computed based on two motion metrics. Here, in order to reduce errors of the input data, the motion vectors are filtered based on a pairwise planarity check and then filtered motion vectors are used for training. In the experiments, the proposed method is verified by comparing the detection performance with that of previous approaches on the KITTI dataset.
Min-Gyu Park, Ju Hong Yoon, Jonghee Park, Jeong-Kyun Lee, Kuk-Jin Yoon
Intelligent Vehicles Symposium1
2015 Leveraging stereo matching with learning-based confidence measures
abstract
We propose a new approach to associate supervised learning-based confidence prediction with the stereo matching problem. First of all, we analyze the characteristics of various confidence measures in the regression forest framework to select effective confidence measures using training data. We then train regression forests again to predict the correctness (confidence) of a match by using selected confidence measures. In addition, we present a confidence-based matching cost modulation scheme based on the predicted correctness for improving the robustness and accuracy of various stereo matching algorithms. We apply the proposed scheme to the semi-global matching algorithm to make it robust under unexpected difficulties that can occur in outdoor environments. We verify the proposed confidence measure selection and cost modulation methods through extensive experimentation with various aspects using KITTI and challenging outdoor datasets.
Min-Gyu Park, Kuk-Jin Yoon
CVPR1
2014 Dynamic Point Clustering with Line Constraints for Moving Object Detection in DAS
abstract
In this letter, we propose a robust dynamic point clustering method for detecting moving objects in stereo image sequences, which is essential for collision detection in driver assistance system. If multiple objects with similar motions are located in close proximity, dynamic points from different moving objects may be clustered together when using the position and velocity as clustering criteria. To solve this problem, we apply a geometric constraint between dynamic points using line segments. Based on this constraint, we propose a variable K-nearest neighbor clustering method and three cost functions that are defined between line segments and points. The proposed method is verified experimentally in terms of its accuracy, and comparisons are also made with conventional methods that only utilize the positions and velocities of dynamic points.
Jonghee Park, Ju Hong Yoon, Min-Gyu Park, Kuk-Jin Yoon
IEEE Signal Process. Lett.3
2012 Efficient Point Feature Tracking based on Self-aware Distance Transform
Min-Gyu Park, Kuk-Jin Yoon
BMVC1