Hyunsoo Park

dblp:316/0117 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 100%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computational science and engineering › model simulation › atomistic simulation
machine learning interatomic potential
0.912025
MLIP Arena: Advancing Fairness and Transparency in Machine Learning Interatomic Potentials via an Open, Accessible Benchmark Platform · NeurIPS 2025
Performance modeling and evaluation
benchmarking
0.912025
MLIP Arena: Advancing Fairness and Transparency in Machine Learning Interatomic Potentials via an Open, Accessible Benchmark Platform · NeurIPS 2025
Computer vision › 3D vision
3d reconstruction
0.812024
HandNeRF: Learning to Reconstruct Hand-Object Interaction Scene from a Single RGB Image · ICRA 2024
Computer vision › 3D vision › 3d human reconstruction
hand-object interaction reconstruction
0.812024
HandNeRF: Learning to Reconstruct Hand-Object Interaction Scene from a Single RGB Image · ICRA 2024
Computer vision › 3D vision
3d human pose estimation
0.612022
PoseKernelLifter: Metric Lifting of 3D Human Pose using Sound · CVPR 2022
Computer vision › 3D vision › 3d reconstruction
multimodal 3d reconstruction
0.612022
PoseKernelLifter: Metric Lifting of 3D Human Pose using Sound · CVPR 2022
Computer vision › 3D vision
neural radiance field
0.212024
HandNeRF: Learning to Reconstruct Hand-Object Interaction Scene from a Single RGB Image · ICRA 2024

Methods — techniques the papers use, named apart from their topics

force field evaluation · 1.7density functional theory · 1.7neural radiance field · 0.8implicit function learning · 0.8transfer function · 0.6audio-visual fusion · 0.63D CNN · 0.6
YearPublicationVenuePosition
2025 MLIP Arena: Advancing Fairness and Transparency in Machine Learning Interatomic Potentials via an Open, Accessible Benchmark Platform
abstract
Machine learning interatomic potentials (MLIPs) have revolutionized molecular and materials modeling, but existing benchmarks suffer from data leakage, limited transferability, and an over-reliance on error-based metrics tied to specific density functional theory (DFT) references. We introduce MLIP Arena, a benchmark platform that evaluates force field performance based on physics awareness, chemical reactivity, stability under extreme conditions, and predictive capabilities for thermodynamic properties and physical phenomena. By moving beyond static DFT references and revealing the important failure modes of current foundation MLIPs in real-world settings, MLIP Arena provides a reproducible framework to guide the next-generation MLIP development toward improved predictive accuracy and runtime efficiency while maintaining physical consistency. The Python package and online leaderboard are available at https://github.com/atomind-ai/mlip-arena.
Yuan Chiang, Tobias Kreiman, Christine Zhang, Matthew C. Kuner, Elizabeth Weaver, Ishan Amin, Hyunsoo Park, Yunsung Lim, Jihan Kim, Daryl Chrzan, Aron Walsh, Samuel M. Blau, Mark Asta, Aditi S. Krishnapriyan
NeurIPS7
2024 HandNeRF: Learning to Reconstruct Hand-Object Interaction Scene from a Single RGB Image
abstract
This paper presents a method to learn hand-object interaction prior for reconstructing a 3D hand-object scene from a single RGB image. The inference as well as training-data generation for 3D hand-object scene reconstruction is challenging due to the depth ambiguity of a single image and occlusions by the hand and object. We turn this challenge into an opportunity by utilizing the hand shape to constrain the possible relative configuration of the hand and object geometry. We design a generalizable implicit function, HandNeRF, that explicitly encodes the correlation of the 3D hand shape features and 2D object features to predict the hand and object scene geometry. With experiments on real-world datasets, we show that HandNeRF can reconstruct hand-object scenes of novel grasp configurations more accurately than comparable methods. Moreover, we demonstrate that object reconstruction from HandNeRF ensures more accurate execution of downstream tasks, such as grasping for robotic hand-over.
Hongsuk Choi, Nikhil Chavan Dafle, Jiacheng Yuan, Volkan Isler, Hyunsoo Park
ICRA5
2024 Impact of interactive learning elements on personal learning performance in immersive virtual reality for construction safety training
abstract
To improve construction safety through proactive prevention strategies, it is essential to leverage virtual reality (VR) training programs that foster active learning through dynamic interaction between VR systems and learners. Addressing this need, this study proposed an interactive immersive VR (IVR)-based construction safety training framework, incorporating four interactive learning elements (ILEs): immediate feedback, basic interaction with objects, assembling objects, and knowledge testing. This study divided sixty trainees into two groups: one experienced the proposed novel interactive IVR-based training approach, while the other underwent traditional, non-interactive IVR-based training. To accurately evaluate the impact of the four ILEs on individual learning performance, this study conducted t-test analyses and utilized machine learning-based SHAP (Shapley Additive exPlanations) analysis to compare the results between both groups. The findings indicated that the proposed method significantly enhanced active learning, positively influencing the sensory and knowledge domains of trainees more effectively than the existing methods. Notably, 'ILE-(1). Immediate feedback' and 'ILE-(2). Basic interaction with objects' were identified as key factors in improving personal learning outcomes during training. This study contributes significantly to the field of management in engineering, demonstrating that interactive IVR-based training can offer a superior learning experience and enhance safety performance, thereby reducing accident rates and creating safer construction sites.
Seungwon Seo, Hyunsoo Park, Choongwan Koo
Expert Syst. Appl.2
2022 PoseKernelLifter: Metric Lifting of 3D Human Pose using Sound
abstract
Reconstructing the 3D pose of a person in metric scale from a single view image is a geometrically ill-posed problem. For example, we can not measure the exact distance of a person to the camera from a single view image without additional scene assumptions (e.g., known height). Existing learning based approaches circumvent this issue by reconstructing the 3D pose up to scale. However, there are many applications such as virtual telepresence, robotics, and augmented reality that require metric scale reconstruction. In this paper, we show that audio signals recorded along with an image, provide complementary information to reconstruct the metric 3D pose of the person. The key insight is that as the audio signals traverse across the 3D space, their interactions with the body provide metric information about the body's pose. Based on this insight, we introduce a time-invariant transfer function called pose kernel-the impulse response of audio signals induced by the body pose. The main properties of the pose kernel are that (1) its envelope highly correlates with 3D pose, (2) the time response corresponds to arrival time, indicating the metric distance to the microphone, and (3) it is invariant to changes in the scene geometry configurations. Therefore, it is readily generalizable to unseen scenes. We design a multistage 3D CNN that fuses audio and visual signals and learns to reconstruct 3D pose in a metric scale. We show that our multi-modal method produces accurate metric reconstruction in realworld scenes, which is not possible with state-of-the-art lifting approaches including parametric mesh regression and depth regression.
Zhijian Yang, Xiaoran Fan, Volkan Isler, Hyunsoo Park
CVPR4
2022 Self-supervised Wide Baseline Visual Servoing via 3D Equivariance
abstract
One of the challenging input settings for visual servoing is when the initial and goal camera views are far apart. Such settings are difficult because the wide baseline can cause drastic changes in object appearance and cause occlusions. This paper presents a novel self-supervised visual servoing method for wide baseline images which does not require 3D ground truth supervision. Existing approaches that regress absolute camera pose with respect to an object require 3D ground truth data of the object in the forms of 3D bounding boxes or meshes. We learn a coherent visual representation by leveraging a geometric property called 3D equivariance—the representation is transformed in a predictable way as a function of 3D transformation. To ensure that the feature-space is faithful to the underlying geodesic space, a geodesic preserving constraint is applied in conjunction with the equivariance. We design a Siamese network that can effectively enforce these two geometric properties without requiring 3D supervision. With the learned model, the relative transformation can be inferred simply by following the gradient in the learned space and used as feedback for closed-loop visual servoing. Our method is evaluated on objects from the YCB dataset, showing meaningful outperformance on a visual servoing task, or object alignment task with respect to state-of-the-art approaches that use 3D supervision. Ours yields more than 35% average distance error reduction and more than 90% success rate with 3cm error tolerance.
Jinwook Huh, Jungseok Hong, Suveer Garg, Hyunsoo Park, Volkan Isler
IROS4