Jinhyung Park

dblp:97/10525 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
13since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 6 first-author · 8 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Leveraging Temporal Cues for Semi-Supervised Multi-View 3D Object Detection
abstract
While recent advancements in camera-based 3D object detection demonstrate remarkable performance, they require thousands or even millions of human-annotated frames. This requirement significantly inhibits their deployment in various locations and sensor configurations. To address this gap, we propose a performant semi-supervised framework that leverages unlabeled RGB-only driving sequences - data easily collected with cost-effective RGB cameras - to significantly improve temporal, camera-only 3D detectors. We observe that the standard semi-supervised pseudo-labeling paradigm underperforms in this temporal, camera-only setting due to poor 3D localization of pseudo-labels. To address this, we train a single 3D detector to handle RGB sequences both forward and backward in time, then ensemble both its forwards and backwards pseudo-labels for semi-supervised learning. We further improve the pseudo-label quality by leveraging 3D object tracking to infill missing detections and by eschewing simple confidence thresholding in favor of using the auxiliary 2D detection head to filter 3D predictions. Finally, to enable the backbone to learn directly from the unlabeled data itself, we introduce an object-query conditioned masked reconstruction objective. Our framework demonstrates remarkable performance improvement on large-scale autonomous driving datasets nuScenes and nuPlan.
Jinhyung Park, Navyata Sanghvi, Hiroki Adachi, Yoshihisa Shibata, Shawn Hunt, Shinya Tanaka, Hironobu Fujiyoshi, Kris Makoto Kitani
CVPR1
2025 ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human Modeling
abstract
Parametric body models offer expressive 3D representation of humans across a wide range of poses, shapes, and facial expressions, typically derived by learning a basis over registered 3D meshes. However, existing human mesh modeling approaches struggle to capture detailed variations across diverse body poses and shapes, largely due to limited training data diversity and restrictive modeling assumptions. Moreover, the common paradigm first optimizes the external body surface using a linear basis, then regresses internal skeletal joints from surface vertices. This approach introduces problematic dependencies between internal skeleton and outer soft tissue, limiting direct control over body height and bone lengths. To address these issues, we present ATLAS, a high-fidelity body model learned from 600k high-resolution scans captured using 240 synchronized cameras. Unlike previous methods, we explicitly decouple the shape and skeleton bases by grounding our mesh representation in the human skeleton. This decoupling enables enhanced shape expressivity, fine-grained customization of body attributes, and keypoint fitting independent of external soft-tissue characteristics. ATLAS outperforms existing methods by fitting unseen subjects in diverse poses more accurately, and quantitative evaluations show that our non-linear pose correctives more effectively capture complex poses compared to linear models.
Jinhyung Park, Javier Romero 0002, Shunsuke Saito, Fabian Prada, Takaaki Shiratori, Federica Bogo, Shoou-I Yu, Kris Makoto Kitani, Rawal Khirodkar
ICCV1
2025 F-RDW: Redirected Walking With Forecasting Future Position
abstract
In order to serve better VR experiences to users, existing predictive methods of Redirected Walking (RDW) exploit future information to reduce the number of reset occurrences. However, such methods often impose a precondition during deployment, either in the virtual environment's layout or the user's walking direction, which constrains its universal applications. To tackle this challenge, we propose a mechanism F-RDW that is twofold: (1) forecasts the future information of a user in the virtual space without any assumptions by using the conventional method, and (2) fuse this information while maneuvering existing RDW methods. The backbone of the first step is an LSTM-based model that ingests the user's spatial and eye-tracking data to predict the user's future position in the virtual space, and the following step feeds those predicted values into existing RDW methods (such as MPCRed, S2C, TAPF, and ARC) while respecting their internal mechanism in applicable ways. The results of our simulation test and user study demonstrate the significance of future information when using RDW in small physical spaces or complex environments. We prove that the proposed mechanism significantly reduces the number of resets and increases the traveled distance between resets, hence augmenting the redirection performance of all RDW methods explored in this work.
Sang-Bin Jeon, Jaeho Jung, Jinhyung Park, In-Kwon Lee
IEEE Trans. Vis. Comput. Graph.3
2024 Flexible Depth Completion for Sparse and Varying Point Densities
abstract
While recent depth completion methods have achieved remarkable results filling in relatively dense depth maps (e.g., projected 64-line LiDAR on KITTI or 500 sampled points on NYUv2) with RGB guidance, their performance on very sparse input (e.g., 4-line LiDAR or 32 depth point measurements) is unverified. These sparser regimes present new challenges, as a 4-line LiDAR increases the distance between pixels without depth and their nearest depth point sixfold from 5 pixels to 30 pixels compared to 64 lines. Ob-serving that existing methods struggle with sparse and variable distribution depth maps, we propose an Affinity-Based Shift Correction (ASC) module that iteratively aligns depth predictions to input depth based on predicted affinities between image pixels and depth points. Our framework enables each depth point to adaptively influence and improve predictions across the image, leading to largely improved results for fewer-line, fewer-point, and variable sparsity settings. Further, we show improved performance in domain transfer from KITTI to nuScenes andfrom random sampling to irregular point distributions. Our correction module can easily be added to any depth completion or RGB-only depth estimation model, notably allowing the latter to perform both completion and estimation with a single model.
Jinhyung Park, Yu-Jhe Li, Kris Makoto Kitani
CVPR1
2023 Azimuth Super-Resolution for FMCW Radar in Autonomous Driving
abstract
We tackle the task of Azimuth (angular dimension) super-resolution for Frequency Modulated Continuous Wave (FMCW) multiple-input multiple-output (MIMO) radar. FMCW MIMO radar is widely used in autonomous driving alongside Lidar and RGB cameras. However, compared to Lidar, MIMO radar is usually of low resolution due to hardware size restrictions. For example, achieving 1° azimuth resolution requires at least 100 receivers, but a single MIMO device usually supports at most 12 receivers. Having limitations on the number of receivers is problematic since a high-resolution measurement of azimuth angle is essential for estimating the location and velocity of objects. To improve the azimuth resolution of MIMO radar, we propose a light, yet efficient, Analog-to-Digital super-resolution model (ADC-SR) that predicts or hallucinates additional radar signals using signals from only a few receivers. Compared with the baseline models that are applied to processed radar Range-Azimuth-Doppler (RAD) maps, we show that our ADC-SR method that processes raw ADC signals achieves comparable performance with 98% (50 times) fewer parameters. We also propose a hybrid super-resolution model (Hybrid-SR) combining our ADC-SR with a standard RAD super-resolution model, and show that performance can be improved by a large margin. Experiments on our Pitt-Radar dataset and the RADIal dataset validate the importance of leveraging raw radar ADC signals. To assess the value of our super-resolution model for autonomous driving, we also perform object detection on the results of our super-resolution model and find that our super-resolution model improves detection performance by around 4% in mAP. The Pitt-Radar and the code will be released at the link.
Yu-Jhe Li, Shawn Hunt, Jinhyung Park, Matthew O'Toole, Kris Makoto Kitani
CVPR3
2023 Time Will Tell: New Outlooks and A Baseline for Temporal Multi-View 3D Object Detection
Jinhyung Park, Chenfeng Xu, Shijia Yang, Kurt Keutzer, Kris Makoto Kitani, Masayoshi Tomizuka
ICLR1
2023 "To be or Not to be Me?": Exploration of Self-Similar Effects of Avatars on Social Virtual Reality Experiences
abstract
The growing interest in the self-similarity effect of avatars in virtual reality (VR) has spurred the creation of realistic avatars that closely mirror their users. However, despite extensive research on the self-similarity effect in single-user VR environments, our understanding of its impact in social VR settings remains underdeveloped. This shortfall exists despite the unique socio-psychological phenomena arising from the illusion of embodiment that could potentially alter these effects. To fill this gap, this paper provides an in-depth empirical investigation of how avatars' self-similarity influences social VR experiences. Our research uncovers several notable findings: 1) A high level of avatar self-similarity boosts users' sense of embodiment and social presence but has minimal effects on the overall presence and even slightly hinders immersion. These results are driven by increased self-awareness. 2) Among various factors that contribute to the self-similarity of avatars, voice stands out as a significant influencer of social VR experiences, surpassing other representational factors. 3) The impact of avatar self-similarity shows negligible differences between male and female users. Based on these findings, we discuss the pros and cons of incorporating self-similarity into social VR avatars. Our study serves as a foundation for further research in this field.
Jinhyung Park, In-Kwon Lee
IEEE Trans. Vis. Comput. Graph.2
2022 Modality-Agnostic Learning for Radar-Lidar Fusion in Vehicle Detection
abstract
Fusion of multiple sensor modalities such as camera, Lidar, and Radar, which are commonly found on autonomous vehicles, not only allows for accurate detection but also robustifies perception against adverse weather conditions and individual sensor failures. Due to inherent sensor characteristics, Radar performs well under extreme weather conditions (snow, rain, fog) that significantly degrade camera and Lidar. Recently, a few works have developed vehicle detection methods fusing Lidar and Radar signals, i.e., MVD-Net. However, these models are typically developed under the assumption that the models always have access to two error-free sensor streams. If one of the sensors is unavailable or missing, the model may fail catastrophically. To mitigate this problem, we propose the Self-Training Multimodal Vehicle Detection Network (ST-MVDNet) which leverages a Teacher-Student mutual learning framework and a simulated sensor noise model used in strong data augmentation for Lidar and Radar. We show that by (1) enforcing output consistency between a Teacher network and a Student network and by (2) introducing missing modalities (strong augmentations) during training, our learned model breaks away from the error-free sensor assumption. This consistency enforcement enables the Student model to handle missing data properly and improve the Teacher model by updating it with the Student model's exponential moving average. Our experiments demonstrate that our proposed learning framework for multi-modal detection is able to better handle missing sensor data during inference. Furthermore, our method achieves new state-of-the-art performance (5% gain) on the Oxford Radar Robotcar dataset under various evaluation settings.
Yu-Jhe Li, Jinhyung Park, Matthew O'Toole, Kris Makoto Kitani
CVPR2
2022 DetMatch: Two Teachers are Better than One for Joint 2D and 3D Semi-Supervised Object Detection
Jinhyung Park, Chenfeng Xu, Yiyang Zhou, Masayoshi Tomizuka
ECCV (10)1
2022 Infinite Virtual Space Exploration Using Space Tiling and Perceivable Reset at Fixed Positions
abstract
A simultaneous walking experience in virtual and real spaces can provide a high sense of presence. However, users may face challenges when walking within a large virtual space while walking in a small and complex real space. Several methods such as Redirected Walking (RDW) and Substitutional Reality (SR) have been proposed as different approaches to this problem. However, the users must “reset” their movement direction at unpredictable moments to avoid collision in a small and complex real space when using subtle RDW that does not maintain the correspondence between virtual and real space. Contrarily, exploration through the SR has a limitation in that the VR scene is restricted to a controlled area. In this paper, we propose Reset at Fixed Positions (RFP), a method that combines RDW with the advantage of the SR and matches walkable real space with walkable virtual space. To utilize RFP, we defined Guaranteed Space Block (GSB), a unit space that constitutes a walkable virtual space. This space is obtained through the point reflection of the GSB utilizing the reset position within the GSB. RFPs can be implemented by two methods: Generating Virtual Space Using RFP (G-RFP) and Implementing Given Virtual Space Using RFP (I-RFP). G-RFP can create an infinitely large virtual space for exploration. On the other hand, I-RFP can conFigure a given virtual environment to make users walk. We observed that G-RFP provides higher presence, immersion and a higher mean distance traveled between resets compared to the existing RDW method in a complex real space through a user study. In addition, exploration through I-RFP provided a higher immersion, a comparable presence, and a similar number of resets.
SoonUk Kwon, Sang-Bin Jeon, June-Young Hwang, Yong-Hun Cho, Jinhyung Park, In-Kwon Lee
ISMAR5
2022 Dynamic optimal space partitioning for redirected walking in multi-user environment
abstract
In multi-user Redirected Walking (RDW), the space subdivision method divides a shared physical space into sub-spaces and allocates a sub-space to each user. While this approach has the advantage of precluding any collisions between users, the conventional space subdivision method suffers from frequent boundary resets due to the reduction of available space per user. To address this challenge, in this study, we propose a space subdivision method called Optimal Space Partitioning (OSP) that dynamically divides the shared physical space in real-time. By exploiting spatial information of the physical and virtual environment, OSP predicts the movement of users and divides the shared physical space into optimal sub-spaces separated with shutters. Our OSP framework is trained using deep reinforcement learning to allocate optimal sub-space to each user and provide optimal steering. Our experiments demonstrate that OSP provides higher sense of immersion to users by minimizing the total number of reset counts, while preserving the advantage of the existing space subdivision strategy: ensuring better safety to users by completely eliminating the possibility of any collisions between users beforehand. Our project is available at https://github.com/AppleParfait/OSP-Archive.
Sang-Bin Jeon, SoonUk Kwon, June-Young Hwang, Yong-Hun Cho, Jinhyung Park, In-Kwon Lee
ACM Trans. Graph.6
2021 Multi-Modality Task Cascade for 3D Object Detection
Jinhyung Park, Xinshuo Weng, Yunze Man, Kris Makoto Kitani
BMVC1
2021 Crack Detection and Refinement Via Deep Reinforcement Learning
abstract
Detecting small cracks in concrete is difficult due to the complexity and thinness of cracking patterns, which requires the development of refined vision-based segmentation algorithms that can accurately characterize the details of crack defects. While existing methods are good at generally outlining cracks, due to inherent differences in shape distributions between common objects and cracks, their predictions often have disconnected segments and inaccuracy along boundaries. To this end, we develop a refinement framework using reinforcement learning (RL) that can better recognize details specific to cracks. Our method uses an RL agent to iteratively improve per-pixel crack predictions of a general segmentation model. We find that in addition to connecting gaps in predictions, the RL agent is also able to detect cracks that are missed in the original predictions. It does so by using the originally detected regions as crack priors to branch out from. Refining outputs of a commonly used per-pixel segmentation model, our method outperforms the current state-of-the-art approaches for crack segmentation. Our experiments also demonstrate that our method generalizes well to a similar task of vessel segmentation.
Jinhyung Park, Yu-Jhe Li, Kris Makoto Kitani
ICIP1
2013 netShip: a networked virtual platform for large-scale heterogeneous distributed embedded systems
abstract
From a single SoC to a network of embedded devices communicating with a backend cloud-computing server, emerging classes of embedded systems feature an increasing number of heterogeneous components that operate concurrently in a distributed environment. As the scale and complexity of these systems continues to grow, there is a critical need for scalable and efficient simulators. We propose a networked virtual platform as a scalable environment for modeling and simulation. The goal is to support the development and optimization of embedded computing applications by handling heterogeneity at the chip, node, and network level. To illustrate the properties of our approach, we present two very different case studies: the design of an Open MPI scheduler for a heterogeneous distributed embedded system and the development of an application for crowd estimation through the analysis of pictures uploaded from mobile phones.
YoungHoon Jung, Jinhyung Park, Michele Petracca, Luca P. Carloni
DAC2