VLDB 2026 Research / reviewers in the wild / expert
Dongseok Yang
dblp:153/0198
· DBLP profile ↗
10ranked-venue papers
6as first author
7since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-authorArtificial intelligence and machine learning · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DivaTrack: Diverse Bodies and Motions from Acceleration-Enhanced Three-Point TrackersabstractAbstract Full‐body avatar presence is important for immersive social and environmental interactions in digital reality. However, current devices only provide three six degrees of freedom (DOF) poses from the headset and two controllers (i.e. three‐point trackers). Because it is a highly under‐constrained problem, inferring full‐body pose from these inputs is challenging, especially when supporting the full range of body proportions and use cases represented by the general population. In this paper, we propose a deep learning framework, DivaTrack, which outperforms existing methods when applied to diverse body sizes and activities. We augment the sparse three‐point inputs with linear accelerations from Inertial Measurement Units (IMU) to improve foot contact prediction. We then condition the otherwise ambiguous lower‐body pose with the predictions of foot contact and upper‐body pose in a two‐stage model. We further stabilize the inferred full‐body pose in a wide range of configurations by learning to blend predictions that are computed in two reference frames, each of which is designed for different types of motions. We demonstrate the effectiveness of our design on a large dataset that captures 22 subjects performing challenging locomotion for three‐point tracking, including lunges, hula‐hooping, and sitting. As shown in a live demo using the Meta VR headset and Xsens IMUs, our method runs in real‐time while accurately tracking a user's motion when they perform a diverse set of movements. Dongseok Yang, Jiho Kang, Lingni Ma, Joseph D. Greer, Yuting Ye, Sung-Hee Lee |
Comput. Graph. Forum | 1 |
| 2024 | ELMO: Enhanced Real-time LiDAR Motion Capture through UpsamplingabstractThis paper introduces ELMO, a real-time upsampling motion capture framework designed for a single LiDAR sensor. Modeled as a conditional autoregressive transformer-based upsampling motion generator, ELMO achieves 60 fps motion capture from a 20 fps LiDAR point cloud sequence. The key feature of ELMO is the coupling of the self-attention mechanism with thoughtfully designed embedding modules for motion and point clouds, significantly elevating the motion quality. To facilitate accurate motion capture, we develop a one-time skeleton calibration model capable of predicting user skeleton off-sets from a single-frame point cloud. Additionally, we introduce a novel data augmentation technique utilizing a LiDAR simulator, which enhances global root tracking to improve environmental understanding. To demonstrate the effectiveness of our method, we compare ELMO with state-of-the-art methods in both image-based and point cloud-based motion capture. We further conduct an ablation study to validate our design principles. ELMO's fast inference time makes it well-suited for real-time applications, exemplified in our demo video featuring live streaming and interactive gaming scenarios. Furthermore, we contribute a high-quality LiDAR-mocap synchronized dataset comprising 20 different subjects performing a range of motions, which can serve as a valuable resource for future research. The dataset and evaluation code are available at https://movin3d.github.io/ELMO_SIGASIA2024/ Deok-Kyeong Jang, Dongseok Yang, Deok-Yun Jang, Byeoli Choi, Sung-Hee Lee |
ACM Trans. Graph. | 2 |
| 2024 | Visual Guidance for User Placement in Avatar-Mediated Telepresence Between Dissimilar SpacesabstractRapid advances in technology gradually realize immersive mixed-reality (MR) telepresence between distant spaces. This paper presents a novel visual guidance system for avatar-mediated telepresence, directing users to optimal placements that facilitate the clear transfer of gaze and pointing contexts through remote avatars in dissimilar spaces, where the spatial relationship between the remote avatar and the interaction targets may differ from that of the local user. Representing the spatial relationship between the user/avatar and interaction targets with angle-based interaction features, we assign recommendation scores of sampled local placements as their maximum feature similarity with remote placements. These scores are visualized as color-coded 2D sectors to inform the users of better placements for interaction with selected targets. In addition, virtual objects of the remote space are overlapped with the local space for the user to better understand the recommendations. We examine whether the proposed score measure agrees with the actual user perception of the partner's interaction context and find a score threshold for recommendation through user experiments in virtual reality (VR). A subsequent user study in VR investigates the effectiveness and perceptual overload of different combinations of visualizations. Finally, we conduct a user study in an MR telepresence scenario to evaluate the effectiveness of our method in real-world applications. Dongseok Yang, Jiho Kang, Taehei Kim, Sung-Hee Lee |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2023 | Real-time Retargeting of Deictic Motion to Virtual Avatars for Augmented Reality TelepresenceabstractAvatar-mediated augmented reality telepresence aims to enable distant users to collaborate remotely through avatars. When two spaces involved in telepresence are dissimilar, with different object sizes and arrangements, the avatar movement must be adjusted to convey the user’s intention rather than directly following their motion, which poses a significant challenge. In this paper, we propose a novel neural network-based framework for real-time retargeting of users’ deictic motions (pointing at and touching objects) to virtual avatars in dissimilar environments. Our framework translates the user’s deictic motion, acquired from a sparse set of tracking signals, to the virtual avatar’s deictic motion for a corresponding remote object in real-time. One of the main features of our framework is that a single trained network can generate natural deictic motions for various sizes of users. To this end, our network includes two sub-networks: AngleNet and MotionNet. AngleNet maps the angular state of the user’s motion into a latent representation, which is subsequently converted by MotionNet into the avatar’s pose, considering the user’s scale. We validate the effectiveness of our method in terms of deictic intention preservation and movement naturalness through quantitative comparison with alternative approaches. Additionally, we demonstrate the utility of our approach through several AR telepresence scenarios. Jiho Kang, Dongseok Yang, Taehei Kim, Yewon Lee 0001, Sung-Hee Lee |
ISMAR | 2 |
| 2023 | MOVIN: Real-time Motion Capture using a Single LiDARabstractAbstract Recent advancements in technology have brought forth new forms of interactive applications, such as the social metaverse, where end users interact with each other through their virtual avatars. In such applications, precise full‐body tracking is essential for an immersive experience and a sense of embodiment with the virtual avatar. However, current motion capture systems are not easily accessible to end users due to their high cost, the requirement for special skills to operate them, or the discomfort associated with wearable devices. In this paper, we present MOVIN, the data‐driven generative method for real‐time motion capture with global tracking, using a single LiDAR sensor. Our autoregressive conditional variational autoencoder (CVAE) model learns the distribution of pose variations conditioned on the given 3D point cloud from LiDAR. As a central factor for high‐accuracy motion capture, we propose a novel feature encoder to learn the correlation between the historical 3D point cloud data and global, local pose features, resulting in effective learning of the pose prior. Global pose features include root translation, rotation, and foot contacts, while local features comprise joint positions and rotations. Subsequently, a pose generator takes into account the sampled latent variable along with the features from the previous frame to generate a plausible current pose. Our framework accurately predicts the performer's 3D global information and local joint details while effectively considering temporally coherent movements across frames. We demonstrate the effectiveness of our architecture through quantitative and qualitative evaluations, comparing it against state‐of‐the‐art methods. Additionally, we implement a real‐time application to showcase our method in real‐world scenarios. MOVIN dataset is available at https://movin3d.github.io/movin_pg2023/https://movin3d.github.io/movin_pg2023/">https://movin3d.github.io/movin_pg2023/ . Deok-Kyeong Jang, Dongseok Yang, Deok-Yun Jang, Byeoli Choi, Taeil Jin, Sung-Hee Lee |
Comput. Graph. Forum | 2 |
| 2022 | Placement Retargeting of Virtual Avatars to Dissimilar Indoor EnvironmentsabstractRapidly developing technologies are realizing a 3D telepresence, in which geographically separated users can interact with each other through their virtual avatars. In this article, we present novel methods to determine the avatar's position in an indoor space to preserve the semantics of the user's position in a dissimilar indoor space with different space configurations and furniture layouts. To this end, we first perform a user survey on the preferred avatar placements for various indoor configurations and user placements, and identify a set of related attributes, including interpersonal relation, visual attention, pose, and spatial characteristics, and quantify these attributes with a set of features. By using the obtained dataset and identified features, we train a neural network that predicts the similarity between two placements. Next, we develop an avatar placement method that preserves the semantics of the placement of the remote user in a different space as much as possible. We show the effectiveness of our methods by implementing a prototype AR-based telepresence system and user evaluations. Leonard Yoon, Dongseok Yang, Choongho Chung, Sung-Hee Lee |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | LoBSTr: Real-time Lower-body Pose Prediction from Sparse Upper-body Tracking SignalsabstractAbstract With the popularization of games and VR/AR devices, there is a growing need for capturing human motion with a sparse set of tracking data. In this paper, we introduce a deep neural network (DNN) based method for real‐time prediction of the lower‐body pose only from the tracking signals of the upper‐body joints. Specifically, our Gated Recurrent Unit (GRU)‐based recurrent architecture predicts the lower‐body pose and feet contact states from a past sequence of tracking signals of the head, hands, and pelvis. A major feature of our method is that the input signal is represented by the velocity of tracking signals. We show that the velocity representation better models the correlation between the upper‐body and lower‐body motions and increases the robustness against the diverse scales and proportions of the user body than position‐orientation representations. In addition, to remove foot‐skating and floating artifacts, our network predicts feet contact state, which is used to post‐process the lower‐body pose with inverse kinematics to preserve the contact. Our network is lightweight so as to run in real‐time applications. We show the effectiveness of our method through several quantitative evaluations against other architectures and input representations with respect to wild tracking data obtained from commercial VR devices. Dongseok Yang, Sung-Hee Lee |
Comput. Graph. Forum | 1 |
| 2020 | TapSix: A Palm-Worn Glove with a Low-Cost Camera Sensor that Turns a Tactile Surface into a Six-Key Chorded Keyboard by Detection Finger TapsabstractTapSix is a one-handed wearable keyboard that enables typing in situations where using a keyboard is not possible. It detects finger taps on six virtual keys on a tactile surface while users can type without paying visual attention to their fingers. Its unique palm-worn design provides the stable view of all five fingers for the low-cost camera sensor, even when there is unusual motion in the upper limb. The captured image is processed using the proposed algorithm that both robustly detects finger taps on keys using only geometric features and measures the distance between each finger and the tactile surface. A new letter-to-tap mapping that fully utilizes the six keys in light of learnability, anatomical comfort, and algorithm accuracy is also proposed. We demonstrate the utility of TapSix in a virtual reality environment and evaluate the algorithm’s accuracy, typing performance, and user acceptance by comparing it with three commercial virtual reality interfaces. Dongseok Yang, Younggeun Choi 0001 |
Int. J. Hum. Comput. Interact. | 1 |
| 2018 | Synthetic Hands Generator for RGB Hand TrackingabstractIn this study, we addressed the challenging problem of 2D hand-pose tracking based on an RGB-only sequence by using a hand data generator. For training various deep networks on hand-pose tracking, we propose a synthetic hand generator based on an application. Our generator could be combined with a kinematic hand model to generalize well to unseen data. In addition, it is robust to occlusions and varying camera viewpoints and leads to anatomically smooth hand motions. Our generator also allows to set the range of each property and add objects (hand models and backgrounds) easily to the application. This greatly diversifies the architecture and improves performance of hand pose tracking. We evaluated our generator by comparing with other public hand datasets and propose a novel annotation technique for accurate 2D (3D) hand labeling and joint angles even in case of partial occlusions. We demonstrate that the dataset generated through our generator outperforms other public datasets with only challenging RGB. Dongseok Yang, BackSan Moon, Haneurl Kim, Younggeun Choi 0001 |
TENCON | 1 |
| 2014 | Early childhood education by hand gesture recognition using a smartphone based robotabstractWe propose a light and fast hand gesture recognition method using geometric feature for a smartphone based robot and apply it to early childhood mathematics education. The feature of hand gesture is defined by the number of extrema in the plot for the distances between the center point of hand and the outer points of hand from active contour model or snakes. The region of interest (ROI) is continuously updated by Continuously Adaptive Mean Shift Algorithm (CamShift) algorithm, and the snake model is used to make the outer points sequential efficiently. A mathematics learning application for an Android OS smartphone based robot is developed using the hand gesture recognition algorithm. The experiment with Korean children (5-6 years of age) is conducted to evaluate if hand gesture based HRI could promote their mathematics learning. The result suggests that the idea of hand gesture based HRI for early childhood education is feasible and that children can learn mathematics by hand gesture based interaction with a robot. Dongseok Yang, Jong-Kuk Lim, Younggeun Choi 0001 |
RO-MAN | 1 |