Hyung-Gun Chi

dblp:270/5700 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
18since 2021 · last 2025
0000-0001-5454-3404ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 12 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CARING-AI: Towards Authoring Context-aware Augmented Reality INstruction through Generative Artificial Intelligence
Jingyu Shi, Rahul Jain 0018, Seunggeun Chi, Hyungjun Doh, Hyung-Gun Chi, Alexander J. Quinn, Karthik Ramani
CHI5
2025 DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
Hyung-Gun Chi, Zakaria Aldeneh, Tatiana Likhomanenko, Ognjen Rudovic, Takuya Higuchi, Shinji Watanabe 0001, Ahmed Hussen Abdelaziz
INTERSPEECH1
2025 Adaptive Knowledge Distillation for Device-Directed Speech Detection
Hyung-Gun Chi, Florian Pesce, Wonil Chang, Ognjen Rudovic, Arturo Argueta, Vineet Garg, Ahmed Hussen Abdelaziz
INTERSPEECH1
2025 InfoGCN++: Learning Representation by Predicting the Future for Online Skeleton-Based Action Recognition
abstract
Skeleton-based action recognition has made significant advancements recently, with models like InfoGCN showcasing remarkable accuracy. However, these models exhibit a key limitation: they necessitate complete action observation prior to classification, which constrains their applicability in real-time situations such as surveillance and robotic systems. To overcome this barrier, we introduce InfoGCN++, an innovative extension of InfoGCN, explicitly developed for online skeleton-based action recognition. InfoGCN++ augments the abilities of the original InfoGCN model by allowing real-time categorization of action types, independent of the observation sequence's length. It transcends conventional approaches by learning from current and anticipated future movements, thereby creating a more thorough representation of the entire sequence. Our approach to prediction is managed as an extrapolation issue, grounded on observed actions. To enable this, InfoGCN++ incorporates Neural Ordinary Differential Equations, a concept that lets it effectively model the continuous evolution of hidden states. Following rigorous evaluations on three skeleton-based action recognition benchmarks, InfoGCN++ demonstrates exceptional performance in online action recognition. It consistently equals or exceeds existing techniques, highlighting its significant potential to reshape the landscape of real-time action recognition applications. Consequently, this work represents a major leap forward from InfoGCN, pushing the limits of what's possible in online, skeleton-based action recognition.
Seunggeun Chi, Hyung-Gun Chi, Qixing Huang, Karthik Ramani
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Higher-order Relational Reasoning for Pedestrian Trajectory Prediction
abstract
Social relations have substantial impacts on the potential trajectories of each individual. Modeling these dynamics has been a central solution for more precise and accurate trajectory forecasting. However, previous works ignore the importance of ‘social depth’, meaning the influences flowing from different degrees of social relations. In this work, we propose HighGraph, a graph-based pedestrian relational reasoning method that captures the higherorder dynamics of social interactions. First, we construct a collision-aware relation graph based on the agents' observed trajectories. Upon this graph structure, we build our core module that aggregates the agent features from diverse social distances. As a result, the network is able to model complex social relations, thereby yielding more accurate and socially acceptable trajectories. Our High-Graph is a plug-and-play module that can be easily applied to any current trajectory predictors. Extensive experiments with ETH/UCY and SDD datasets demonstrate that our HighGraph noticeably improves the previous state-of-the-art baselines both quantitatively and qualitatively.
Sungjune Kim, Hyung-Gun Chi, Hyerin Lim, Karthik Ramani, Jinkyu Kim 0001, Sangpil Kim
CVPR2
2024 M2D2M: Multi-Motion Generation from Text with Discrete Diffusion Models
Seunggeun Chi, Hyung-Gun Chi, Hengbo Ma, Nakul Agarwal, Faizan Siddiqui 0001, Karthik Ramani, Kwonjoon Lee
ECCV (14)2
2024 Enhanced Motion Forecasting with Visual Relation Reasoning
Sungjune Kim, Hadam Baek, Seunggwan Lee, Hyung-Gun Chi, Hyerin Lim, Jinkyu Kim 0001, Sangpil Kim
ECCV (56)4
2024 VisionTrap: Vision-Augmented Trajectory Prediction Guided by Textual Descriptions
Seokha Moon, Hyun Woo, Hongbeen Park, Haeji Jung, Reza Mahjourian, Hyung-Gun Chi, Hyerin Lim, Sangpil Kim, Jinkyu Kim 0001
ECCV (6)6
2024 Multi-Modal Representation Learning with Tactile Data
abstract
Advancements in embodied language models like PALM-E and RT-2 have significantly enhanced language-conditioned robotic manipulation. However, these advances remain predominantly focused on vision and language, often overlooking the pivotal role of tactile feedback which is advantageous in contact-rich interactions. Our research introduces a novel approach that synergizes tactile information with vision and language. We present the Multi-Modal Wand (MMWand) dataset enriched with linguistic descriptions and tactile data. By integrating tactile feedback, we aim to bridge the divide between human linguistic understanding and robotic sensory interpretation. Our multi-modal representation model is trained on these datasets by employing the multi-modal embedding alignment principle from ImageBind which has shown promising results, emphasizing the potential of tactile data in robotic applications. The validation of our approach in downstream robotics tasks, such as texture-based object classification, cross-modality retrieval, and the dense reward function for visuomotor control, attests to its effectiveness. Our contributions underscore the importance of tactile feedback in multi-modal robotic learning and its potential to enhance robotic tasks. The MMWand dataset is publicly available at https://hyung-gun.me/mmwand/.
Hyung-Gun Chi, Jose A. Barreiros, Jean Mercat, Karthik Ramani, Thomas Kollar
IROS1
2024 Novel approach for fast structured light framework using deep learning
Won-Hoe Kim, Bongjoong Kim, Hyung-Gun Chi, Jae-Sang Hyun
Image Vis. Comput.3
2024 Robust sound-guided image manipulation
Hyung-Gun Chi, Gyeongrok Oh, Wonmin Byeon, Sang Ho Yoon, Hyunje Park, Wonjun Cho, Jinkyu Kim 0001, Sangpil Kim
Neural Networks2
2023 Functional Hand Type Prior for 3D Hand Pose Estimation and Action Recognition from Egocentric View Monocular Videos
Wonseok Roh, Wonjeong Ryoo, Jakyung Lee, Gyeongrok Oh, Sooyeon Hwang, Hyung-Gun Chi, Sangpil Kim
BMVC7
2023 AdamsFormer for Spatial Action Localization in the Future
abstract
Predicting future action locations is vital for applications like human-robot collaboration. While some computer vision tasks have made progress in predicting human actions, accurately localizing these actions in future frames remains an area with room for improvement. We introduce a new task called spatial action localization in the future (SALF), which aims to predict action locations in both observed and future frames. SALF is challenging because it requires understanding the underlying physics of video observations to predict future action locations accurately. To address SALF, we use the concept of NeuralODE, which models the latent dynamics of sequential data by solving ordinary differential equations (ODE) with neural networks. We propose a novel architecture, AdamsFormer, which extends observed frame features to future time horizons by modeling continuous temporal dynamics through ODE solving. Specifically, we employ the Adams method, a multi-step approach that efficiently uses information from previous steps without discarding it. Our extensive experiments on UCF101-24 and JHMDB-21 datasets demonstrate that our proposed model outperforms existing long-range temporal modeling methods by a significant margin in terms of frame-mAP.
Hyung-Gun Chi, Kwonjoon Lee, Nakul Agarwal, Yi Xu 0005, Karthik Ramani, Chiho Choi
CVPR1
2023 Uncovering the Missing Pattern: Unified Framework Towards Trajectory Imputation and Prediction
abstract
Trajectory prediction is a crucial undertaking in understanding entity movement or human behavior from observed sequences. However, current methods often assume that the observed sequences are complete while ignoring the potential for missing values caused by object occlusion, scope limitation, sensor failure, etc. This limitation inevitably hinders the accuracy of trajectory prediction. To address this issue, our paper presents a unified framework, the Graph-based Conditional Variational Recurrent Neural Network (GC-VRNN), which can perform trajectory imputation and prediction simultaneously. Specifically, we introduce a novel Multi-Space Graph Neural Network (MS-GNN) that can extract spatial features from incomplete observations and leverage missing patterns. Additionally, we employ a Conditional VRNN with a specifically designed Temporal Decay (TD) module to capture temporal dependencies and temporal missing patterns in incomplete trajectories. The inclusion of the TD module allows for valuable information to be conveyed through the temporal flow. We also curate and benchmark three practical datasets for the joint problem of trajectory imputation and prediction. Extensive experiments verify the exceptional performance of our proposed method. As far as we know, this is the first work to address the lack of benchmarks and techniques for trajectory imputation and prediction in a unified manner.
Yi Xu 0005, Armin Bazarjani, Hyung-Gun Chi, Chiho Choi, Yun Fu 0001
CVPR3
2023 Pose Relation Transformer Refine Occlusions for Human Pose Estimation
abstract
Accurately estimating the human pose is an essential task for many applications in robotics. However, existing pose estimation methods suffer from poor performance when occlusion occurs. Recent advances in NLP have been very successful in predicting the missing words conditioned on visible words. We draw upon the sentence completion analogy in NLP to guide our model to address occlusions in the pose estimation problem. We propose a novel approach that can mitigate the effect of occlusions motivated by the sentence completion task of NLP. In an analogous manner, we designed our model to reconstruct occluded joints given the visible joints utilizing joint correlations by capturing the implicit joint connectivity through the attention mechanism. In this work, we propose a POse Relation Transformer (PORT) that captures the global context of the pose using self-attention and a local context by aggregating adjacent joint features. To supervise PORT in learning joint correlations, we guide PORT to reconstruct randomly masked joints, which we call Masked Joint Modeling (MJM). PORT trained with MJM adds to existing keypoint detection methods and successfully refines occlusions. Notably, PORT is a model-agnostic plug-and-play module for pose refinement under occlusion that can be plugged into any keypoint detector with substantially low computational costs. We conducted extensive experiments to demonstrate the advantage of PORT mitigating the occlusion on the hand and body pose PORT improves the pose estimation accuracy of existing human pose estimation methods by up to 16% with only 5% of additional parameters. The code is publicly available at https://github.com/stnoah1/PORT.
Hyung-Gun Chi, Seunggeun Chi, Stanley Chan, Karthik Ramani
ICRA1
2023 Simplification of 3D CAD Model in Voxel Form for Mechanical Parts Using Generative Adversarial Networks
Hyunoh Lee, Soonjo Kwon, Karthik Ramani, Hyung-Gun Chi, Duhwan Mun
Comput. Aided Des.5
2022 InfoGCN: Representation Learning for Human Skeleton-based Action Recognition
abstract
Human skeleton-based action recognition offers a valuable means to understand the intricacies of human behavior because it can handle the complex relationships between physical constraints and intention. Although several studies have focused on encoding a skeleton, less attention has been paid to embed this information into the latent representations of human action. InfoGCN proposes a learning framework for action recognition combining a novel learning objective and an encoding method. First, we design an information bottleneck-based learning objective to guide the model to learn informative but compact latent representations. To provide discriminative information for classifying action, we introduce attention-based graph convolution that captures the context-dependent intrinsic topology of human action. In addition, we present a multi-modal representation of the skeleton using the relative position of joints, designed to provide complementary spatial information for joints. InfoGcn11Code is available at github.com/stnoahl/infogcn surpasses the known state-of-the-art on multiple skeleton-based action recognition benchmarks with the accuracy of 93.0% on NTU RGB+D 60 cross-subject split, 89.8% on NTU RGB+D 120 cross-subject split, and 97.0% on NW-UCLA.
Hyung-Gun Chi, Myoung Hoon Ha, Seunggeun Chi, Sang Wan Lee, Qixing Huang, Karthik Ramani
CVPR1
2021 Object Synthesis by Learning Part Geometry with Surface and Volumetric Representations
Sangpil Kim, Hyung-Gun Chi, Karthik Ramani
Comput. Aided Des.2
2020 First-Person View Hand Segmentation of Multi-Modal Hand Activity Video Dataset
Sangpil Kim, Hyung-Gun Chi, Xiao Hu 0004, Anirudh Vegesana, Karthik Ramani
BMVC2
2020 A Large-Scale Annotated Mechanical Components Benchmark for Classification and Retrieval Tasks with Deep Neural Networks
Sangpil Kim, Hyung-Gun Chi, Xiao Hu 0004, Qixing Huang, Karthik Ramani
ECCV (18)2
2020 Latent transformations neural network for object view synthesis
Sangpil Kim, Nick Winovich, Hyung-Gun Chi, Guang Lin 0001, Karthik Ramani
Vis. Comput.3