VLDB 2026 Research / reviewers in the wild / expert
Ju Yong Chang
dblp:93/2243
· DBLP profile ↗
24ranked-venue papers
11as first author
7since 2021 · last 2026
0000-0003-3710-7314ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 9 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bidirectional regression for monocular 6DoF head pose estimation and reference system alignment
Sungho Chun, Boeun Kim, Hyung Jin Chang, Ju Yong Chang |
Pattern Recognit. | 4 |
| 2025 | PersonaBooth: Personalized Text-to-Motion GenerationabstractThis paper introduces Motion Personalization, a new task that generates personalized motions aligned with text descriptions using several basic motions containing Persona. To support this novel task, we introduce a new large-scale motion dataset called PerMo (PersonaMotion), which captures the unique personas of multiple actors. We also propose a multi-modal finetuning method of a pretrained motion diffusion model called PersonaBooth. PersonaBooth addresses two main challenges: i) A significant distribution gap between the persona-focused PerMo dataset and the pretraining datasets, which lack persona-specific data, and ii) the difficulty of capturing a consistent persona from the motions vary in content (action type). To tackle the dataset distribution gap, we introduce a persona token to accept new persona features and perform multi-modal adaptation for both text and visuals during finetuning. To capture a consistent persona, we incorporate a contrastive learning technique to enhance intra-cohesion among samples with the same persona. Furthermore, we introduce a context-aware fusion mechanism to maximize the integration of persona cues from multiple input motions. PersonaBooth outperforms state-of-the-art motion style transfer methods, establishing a new benchmark for motion personalization. Boeun Kim, Hea In Jeong, JungHoon Sung, Yihua Cheng, Jeongmin Lee 0007, Ju Yong Chang, Sang-Il Choi, Younggeun Choi 0001, Saim Shin, Hyung Jin Chang |
CVPR | 6 |
| 2025 | Stable Diffusion-Based Approach for Human De-OcclusionabstractHumans can infer the missing parts of an occluded object by leveraging prior knowledge and visible cues. However, enabling deep learning models to accurately predict such occluded regions remains a challenging task. De-occlusion addresses this problem by reconstructing both the mask and RGB appearance. In this work, we focus on human de-occlusion, specifically targeting the recovery of occluded body structures and appearances. Our approach decomposes the task into two stages: mask completion and RGB completion. The first stage leverages a diffusion-based human body prior to provide a comprehensive representation of body structure, combined with occluded joint heatmaps that offer explicit spatial cues about missing regions. The reconstructed amodal mask then serves as a conditioning input for the second stage, guiding the model on which areas require RGB reconstruction. To further enhance RGB generation, we incorporate human-specific textual features derived using a visual question answering (VQA) model and encoded via a CLIP encoder. RGB completion is performed using Stable Diffusion, with decoder fine-tuning applied to mitigate pixel-level degradation in visible regions---a known limitation of prior diffusion-based de-occlusion methods caused by latent space transformations. Our method effectively reconstructs human appearances even under severe occlusions and consistently outperforms existing methods in both mask and RGB completion. Moreover, the de-occluded images generated by our approach can improve the performance of downstream human-centric tasks, such as 2D pose estimation and 3D human reconstruction. The code will be made publicly available. Seung Young Noh, Ju Yong Chang |
ACM Multimedia | 2 |
| 2024 | 6DoF Head Pose Estimation Through Explicit Bidirectional Interaction with Face Geometry
Sungho Chun, Ju Yong Chang |
ECCV (33) | 2 |
| 2023 | Representation Learning of Vertex Heatmaps for 3D Human Mesh Reconstruction From Multi-View ImagesabstractThis study addresses the problem of 3D human mesh reconstruction from multi-view images. Recently, approaches that directly estimate the skinned multi-person linear model (SMPL)-based human mesh vertices based on volumetric heatmap representation from input images have shown good performance. We show that representation learning of vertex heatmaps using an autoencoder helps improve the performance of such approaches. Vertex heatmap autoencoder (VHA) learns the manifold of plausible human meshes in the form of latent codes using AMASS, which is a large-scale motion capture dataset. Body code predictor (BCP) utilizes the learned body prior from VHA for human mesh reconstruction from multi-view images through latent code-based supervision and transfer of pretrained weights. According to experiments on Human3.6M and LightStage datasets, the proposed method outperforms previous methods and achieves state-of-the-art human mesh reconstruction performance. Sungho Chun, Sungbum Park, Ju Yong Chang |
ICIP | 3 |
| 2023 | Learnable Human Mesh Triangulation for 3D Human Pose and Shape EstimationabstractCompared to joint position, the accuracy of joint rotation and shape estimation has received relatively little attention in the skinned multi-person linear model (SMPL)-based human mesh reconstruction from multi-view images. The work in this field is broadly classified into two categories. The first approach performs joint estimation and then produces SMPL parameters by fitting SMPL to resultant joints. The second approach regresses SMPL parameters directly from the input images through a convolutional neural network (CNN)-based model. However, these approaches suffer from the lack of information for resolving the ambiguity of joint rotation and shape reconstruction and the difficulty of network learning. To solve the aforementioned problems, we propose a two-stage method. The proposed method first estimates the coordinates of mesh vertices through a CNN-based model from input images, and acquires SMPL parameters by fitting the SMPL model to the estimated vertices. Estimated mesh vertices provide sufficient information for determining joint rotation and shape, and are easier to learn than SMPL parameters. According to experiments using Human3.6M and MPI-INF-3DHP datasets, the proposed method significantly outperforms the previous works in terms of joint rotation and shape estimation, and achieves competitive performance in terms of joint location estimation. Sungho Chun, Sungbum Park, Ju Yong Chang |
WACV | 3 |
| 2021 | Beyond Static Features for Temporally Consistent 3D Human Pose and Shape From a VideoabstractDespite the recent success of single image-based 3D human pose and shape estimation methods, recovering temporally consistent and smooth 3D human motion from a video is still challenging. Several video-based methods have been proposed; however, they fail to resolve the single image-based methods’ temporal inconsistency issue due to a strong dependency on a static feature of the current frame. In this regard, we present a temporally consistent mesh recovery system (TCMR). It effectively focuses on the past and future frames’ temporal information without being dominated by the current static feature. Our TCMR significantly outperforms previous video-based methods in temporal consistency with better per-frame 3D pose and shape accuracy. We also release the codes. Hongsuk Choi, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee |
CVPR | 3 |
| 2019 | PoseFix: Model-Agnostic General Human Pose Refinement NetworkabstractMulti-person pose estimation from a 2D image is an essential technique for human behavior understanding. In this paper, we propose a human pose refinement network that estimates a refined pose from a tuple of an input image and input pose. The pose refinement was performed mainly through an end-to-end trainable multi-stage architecture in previous methods. However, they are highly dependent on pose estimation models and require careful model design. By contrast, we propose a model-agnostic pose refinement method. According to a recent study, state-of-the-art 2D human pose estimation methods have similar error distributions. We use this error statistics as prior information to generate synthetic poses and use the synthesized poses to train our model. In the testing stage, pose estimation results of any other methods can be input to the proposed method. Moreover, the proposed model does not require code or knowledge about other methods, which allows it to be easily used in the post-processing step. We show that the proposed approach achieves better performance than the conventional multi-stage refinement models and consistently improves the performance of various state-of-the-art pose estimation methods on the commonly used benchmark. The code is available in. Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee |
CVPR | 2 |
| 2019 | Camera Distance-Aware Top-Down Approach for 3D Multi-Person Pose Estimation From a Single RGB ImageabstractAlthough significant improvement has been achieved recently in 3D human pose estimation, most of the previous methods only treat a single-person case. In this work, we firstly propose a fully learning-based, camera distance-aware top-down approach for 3D multi-person pose estimation from a single RGB image. The pipeline of the proposed system consists of human detection, absolute 3D human root localization, and root-relative 3D single-person pose estimation modules. Our system achieves comparable results with the state-of-the-art 3D single-person pose estimation models without any ground truth information and significantly outperforms previous 3D multi-person pose estimation methods on publicly available datasets. The code is available in1,2. Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee |
ICCV | 2 |
| 2018 | V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation From a Single Depth MapabstractMost of the existing deep learning-based methods for 3D hand and human pose estimation from a single depth map are based on a common framework that takes a 2D depth map and directly regresses the 3D coordinates of keypoints, such as hand or human body joints, via 2D convolutional neural networks (CNNs). The first weakness of this approach is the presence of perspective distortion in the 2D depth map. While the depth map is intrinsically 3D data, many previous methods treat depth maps as 2D images that can distort the shape of the actual object through projection from 3D to 2D space. This compels the network to perform perspective distortion-invariant estimation. The second weakness of the conventional approach is that directly regressing 3D coordinates from a 2D image is a highly nonlinear mapping, which causes difficulty in the learning procedure. To overcome these weaknesses, we firstly cast the 3D hand and human pose estimation problem from a single depth map into a voxel-to-voxel prediction that uses a 3D voxelized grid and estimates the per-voxel likelihood for each keypoint. We design our model as a 3D CNN that provides accurate estimates while running in real-time. Our system outperforms previous methods in almost all publicly available 3D hand and human pose estimation datasets and placed first in the HANDS 2017 frame-based 3D hand pose estimation challenge. The code is available in1. Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee |
CVPR | 2 |
| 2018 | Depth-Based 3D Hand Pose Estimation: From Current Achievements to Future GoalsabstractIn this paper, we strive to answer two questions: What is the current state of 3D hand pose estimation from depth images? And, what are the next challenges that need to be tackled? Following the successful Hands In the Million Challenge (HIM2017), we investigate the top 10 state-of-the-art methods on three tasks: single frame 3D pose estimation, 3D hand tracking, and hand pose estimation during object interaction. We analyze the performance of different CNN structures with regard to hand shape, joint visibility, view point and articulation distributions. Our findings include: (1) isolated 3D hand pose estimation achieves low mean errors (10 mm) in the view point range of [70, 120] degrees, but it is far from being solved for extreme view points; (2) 3D volumetric representations outperform 2D CNNs, better capturing the spatial structure of the depth data; (3) Discriminative methods still generalize poorly to unseen hand shapes; (4) While joint occlusions pose a challenge for most methods, explicit modeling of structure constraints can significantly narrow the gap between errors on visible and occluded joints. Shanxin Yuan, Guillermo Garcia-Hernando, Björn Stenger, Gyeongsik Moon, Ju Yong Chang, Kyoung Mu Lee, Pavlo Molchanov 0001, Jan Kautz, Sina Honari, Liuhao Ge, Junsong Yuan 0001, Xinghao Chen 0001, Guijin Wang, Fan Yang 0032, Kai Akiyama, Yang Wu 0001, Qingfu Wan, Meysam Madadi, Sergio Escalera, Shile Li, Dongheui Lee, Iasonas Oikonomidis, Antonis A. Argyros, Tae-Kyun Kim 0001 |
CVPR | 5 |
| 2018 | 2D-3D pose consistency-based conditional random fields for 3D human pose estimation
Ju Yong Chang, Kyoung Mu Lee |
Comput. Vis. Image Underst. | 1 |
| 2016 | Nonparametric Feature Matching Based Conditional Random Fields for Gesture Recognition from Multi-Modal VideoabstractWe present a new gesture recognition method that is based on the conditional random field (CRF) model using multiple feature matching. Our approach solves the labeling problem, determining gesture categories and their temporal ranges at the same time. A generative probabilistic model is formalized and probability densities are nonparametrically estimated by matching input features with a training dataset. In addition to the conventional skeletal joint-based features, the appearance information near the active hand in an RGB image is exploited to capture the detailed motion of fingers. The estimated likelihood function is then used as the unary term for our CRF model. The smoothness term is also incorporated to enforce the temporal coherence of our solution. Frame-wise recognition results can then be obtained by applying an efficient dynamic programming technique. To estimate the parameters of the proposed CRF model, we incorporate the structured support vector machine (SSVM) framework that can perform efficient structured learning by using large-scale datasets. Experimental results demonstrate that our method provides effective gesture recognition results for challenging real gesture datasets. By scoring 0.8563 in the mean Jaccard index, our method has obtained the state-of-the-art results for the gesture recognition track of the 2014 ChaLearn Looking at People (LAP) Challenge. Ju Yong Chang |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Large margin learning of hierarchical semantic similarity for image classification
Ju Yong Chang, Kyoung Mu Lee |
Comput. Vis. Image Underst. | 1 |
| 2014 | A hand posture recognition system utilizing frequency difference of infrared lightabstractHand gesture is one of the most effective methods to perform interactions between humans and also between humans and computers. However, currently existing depth cameras do not provide sufficient resolution and precision for effectively recognizing hand postures in distance (>2 meters). Existing researches tried to solve the limitation by using a combination of depth information and color information. However, they all could not have stable performance, because the color information is naturally affected by visible light condition. In this paper, we introduce a hardware system and an algorithm to recognize hand postures of a distant user while guaranteeing its performance even in the dark. Specifically, by utilizing infrared(IR) lights and their frequency difference, our system simultaneously gathers a depth map from Kinect and a high resolution IR image of a scene from an additional IR camera without any interference. The system analyzes the IR image of a hand using histogram of oriented gradients and support vector machine. In addition, the recognition system has a technique to compensate errors of hand position estimation unavoidable in any hand detection algorithms. As a result, from the experiment on real-time data, the proposed system classifies seven different hand postures with an average precision rate of 92.17% and the precision rate is maintained in the dark (<5 lux) with an average precision rate of 93.28%. Soonchan Park, Moonwook Ryu, Ju Yong Chang |
VRST | 3 |
| 2012 | Learning object relationships via graph-based context modelabstractIn this paper, we propose a novel framework for modeling image-dependent contextual relationships using graph-based context model. This approach enables us to selectively utilize the contextual relationships suitable for an input query image. We introduce a context link view of contextual knowledge, where the relationship between a pair of annotated regions is represented as a context link on a similarity graph of regions. Link analysis techniques are used to estimate the pairwise context scores of all pairs of unlabeled regions in the input image. Our system integrates the learned context scores into a Markov Random Field (MRF) framework in the form of pairwise cost and infers the semantic segmentation result by MRF optimization. Experimental results on object class segmentation show that the proposed graph-based context model outperforms the current state-of-the-art methods. Heesoo Myeong, Ju Yong Chang, Kyoung Mu Lee |
CVPR | 2 |
| 2011 | GPU-friendly multi-view stereo reconstruction using surfel representation and graph cuts
Ju Yong Chang, Haesol Park, In Kyu Park, Kyoung Mu Lee, Sang Uk Lee |
Comput. Vis. Image Underst. | 1 |
| 2009 | 3D pose estimation and segmentation using specular cuesabstractWe present a system for fast model-based segmentation and 3D pose estimation of specular objects using appearance based specular features. We use observed (a) specular reflection and (b) specular flow as cues, which are matched against similar cues generated from a CAD model of the object in various poses. We avoid estimating 3D geometry or depths, which is difficult and unreliable for specular scenes. In the first method, the environment map of the scene is utilized to generate a database containing synthesized specular reflections of the object for densely sampled 3D poses. This database is compared with captured images of the scene at run time to locate and estimate the 3D pose of the object. In the second method, specular flows are generated for dense 3D poses as illumination invariant features and are matched to the specular flow of the scene. We incorporate several practical heuristics such as use of saturated/highlight pixels for fast matching and normal selection to minimize the effects of inter-reflections and cluttered backgrounds. Despite its simplicity, our approach is effective in scenes with multiple specular objects, partial occlusions, inter-reflections, cluttered backgrounds and changes in ambient illumination. Experimental results demonstrate the effectiveness of our method for various synthetic and real objects. Ju Yong Chang, Ramesh Raskar, Amit K. Agrawal |
CVPR | 1 |
| 2008 | Shape from shading using graph cuts
Ju Yong Chang, Kyoung Mu Lee, Sang Uk Lee |
Pattern Recognit. | 1 |
| 2007 | Multiview normal field integration using level set methodsabstractIn this paper, we propose a new method to integrate multiview normal fields using level sets. In contrast with conventional normal integration algorithms used in shape from shading and photometric stereo that reconstruct a 2.5D surface using a single-view normal field, our algorithm can combine multiview normal fields simultaneously and recover the full 3D shape of a target object. We formulate this multiview normal integration problem by an energy minimization framework and find an optimal solution in a least square sense using a variational technique. A level set method is applied to solve the resultant geometric PDE that minimizes the proposed error functional. It is shown that the resultant flow is composed of the well known mean curvature and flux maximizing flows. In particular, we apply the proposed algorithm to the problem of 3D shape modelling in a multiview photometric stereo setting. Experimental results for various synthetic data show the validity of our approach. Ju Yong Chang, Kyoung Mu Lee, Sang Uk Lee |
CVPR | 1 |
| 2007 | Stereo matching using iterative reliable disparity map expansion in the color-spatial-disparity space
Ju Yong Chang, Kyoung Mu Lee, Sang Uk Lee |
Pattern Recognit. | 1 |
| 2006 | Stereo Matching Using Iterated Graph Cuts and Mean Shift Filtering
Ju Yong Chang, Kyoung Mu Lee, Sang Uk Lee |
ACCV (1) | 1 |
| 2006 | A New Stereo Matching Model Using Visibility Constraint Based on Disparity Consistency
Ju Yong Chang, Kyoung Mu Lee, Sang Uk Lee |
ACIVS | 1 |
| 2003 | Shape from shading using graph cutsabstractThis paper describes a new semiglobal method for SFS (shape-from-shading) using graph cuts. The new algorithm combines the local method proposed by Lee and Rosenfeld (1985) and the global method using energy minimization technique. By employing a new global energy minimization formulation, the convex/concave ambiguity problem of the Lee and Rosenfeld method can be resolved efficiently. A new combinatorial optimization technique, graph cuts method is used for the minimization of the proposed energy functional. Experimental results on a variety of synthetic and real-world images show that the proposed algorithm reconstructs the 3-D shape of objects very efficiently. Ju Yong Chang, Kyoung Mu Lee, Sang Uk Lee |
ICIP (1) | 1 |