VLDB 2026 Research / reviewers in the wild / expert
Hail Song
dblp:363/9958
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0006-4008-196XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Streamlined Facial Data Collection Based on Utterance and Emotional Data for Human-to-Avatar ReconstructionabstractThis study explores a streamlined facial data collection method for conversational contexts, addressing the limitations of existing approaches that often require extensive datasets and prioritize technical metrics over user perception and experience. We systematically investigate which facial expression data are essential for reconstructing photorealistic avatars and how they can be captured efficiently. Our research employs a two-phase methodology to identify efficient facial data collection strategies and evaluate their effectiveness. In the first phase, we conduct facial data acquisition and evaluate reconstruction performance using utterance data and emotional data. In the second phase, we carry out a comprehensive user evaluation comparing three progressive conditions: utterance only, utterance and emotional data, and a control condition involving extensive data. Findings from 24 participants engaged in simulated face-to-face conversations reveal that targeted utterance and emotional data achieve comparable levels of perceived realism, naturalness, and telepresence, while reducing training time and data usage when compared to the extensive data collection approach. These results demonstrate that targeted data inputs can enable efficient avatar face reconstruction, offering practical guidelines for real-time applications such as AR/VR telepresence and highlighting the trade-off between data quantity and perceived quality. Seoyoung Kang, Seokhwan Yang, Hail Song, Boram Yoon, Kangsoo Kim, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2026 | SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB SequencesabstractWe introduce SceneLinker, a novel framework that generates compositional 3D scenes via semantic scene graph from RGB sequences. To adaptively experience Mixed Reality (MR) content based on each user's space, it is essential to generate a 3D scene that reflects the real-world layout by compactly capturing the semantic cues of the surroundings. Prior works struggled to fully capture the contextual relationship between objects or mainly focused on synthesizing diverse shapes, making it challenging to generate 3D scenes aligned with object arrangements. We address these challenges by designing a graph network with cross-check feature attention for scene graph prediction and constructing a graph-variational autoencoder (graph-VAE), which consists of a joint shape and layout block for 3D scene generation. Experiments on the 3RScan/3DSSG and SG-FRONT datasets demonstrate that our approach outperforms state-of-the-art methods in both quantitative and qualitative evaluations, even in complex indoor environments and under challenging scene graph constraints. Our work enables users to generate consistent 3D spaces from their physical environments via scene graphs, allowing them to create spatial MR content. Project page is https://scenelinker2026.github.io. Seokyoung Kim 0002, Dooyoung Kim 0001, Woojin Cho 0002, Hail Song, Suji Kang, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | VRGaussianAvatar: Integrating 3D Gaussian Avatars into VRabstractWe present VRGaussianAvatar, an integrated system that enables real-time full-body 3D Gaussian Splatting (3DGS) avatars in virtual reality using only head-mounted display (HMD) tracking signals. The system adopts a parallel pipeline with a VR Frontend and a GA Backend. The VR Frontend uses inverse kinematics to estimate full-body pose and streams the resulting pose along with stereo camera parameters to the backend. The GA Backend stereoscopically renders a 3DGS avatar reconstructed from a single image. To improve stereo rendering efficiency, we introduce Binocular Batching, which jointly processes left and right eye views in a single batched pass to reduce redundant computation and support high-resolution VR displays. We evaluate VRGaussianAvatar with quantitative performance tests and a within-subject user study against image- and video-based mesh avatar baselines. Results show that VRGaussianAvatar sustains interactive VR performance and yields higher perceived appearance similarity, embodiment, and plausibility. Project page and source code are available at https://vrgaussianavatar.github.io. Hail Song, Boram Yoon, Seokhwan Yang, Seoyoung Kang, Hyunjeong Kim, Henning Metzmacher, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2026 | OFERA: Blendshape-Driven 3D Gaussian Control for Occluded Facial Expression to Realistic Avatars in VRabstractWe propose OFERA, a novel framework for real-time expression control of photorealistic Gaussian head avatars for VR headset users. Existing approaches attempt to recover occluded facial expressions using additional sensors or internal cameras, but sensor-based methods increase device weight and discomfort, while camera-based methods raise privacy concerns and suffer from limited access to raw data. To overcome these limitations, we leverage the blendshape signals provided by commercial VR headsets as expression inputs. Our framework consists of three key components: (1) Blendshape Distribution Alignment (BDA), which applies linear regression to align the headset-provided blendshape distribution to a canonical input space; (2) an Expression Parameter Mapper (EPM) that maps the aligned blendshape signals into an expression parameter space for controlling Gaussian head avatars; and (3) a Mapper-integrated Avatar (MiA) that incorporates EPM into the avatar learning process to ensure distributional consistency. Furthermore, OFERA establishes an end-to-end pipeline that senses and maps expressions, updates Gaussian avatars, and renders them in real-time within VR environments. We show that EPM outperforms existing mapping methods on quantitative metrics, and we demonstrate through a user study that the full OFERA framework enhances expression fidelity while preserving avatar realism. By enabling real-time and photorealistic avatar expression control, OFERA significantly improves telepresence in VR communication. A project page is available at https://ysshwan147.github.io/projects/ofera/. Seokhwan Yang, Boram Yoon, Seoyoung Kang, Hail Song, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | The Influence of Emotion-based Prioritized Facial Expressions on Social Presence in Avatar-mediated Remote CommunicationabstractIn avatar-mediated remote communication, avatars’ facial expressions can be dynamically adjusted according to each user’s computational and device constraints, highlighting the importance of varied expressions and their impact on user perception. However, there is a lack of research on how variations in avatar facial expressions, especially when simplified, influence user perception, particularly in terms of social presence. To address this, we examine the impact of various facial expression combinations on social presence in avatar-mediated communication scenarios, ranging from informative speeches to emotional conversations. Our approach involves prioritizing avatar facial blendshape combinations using two main approaches: (1) commonly activated expressions that reflect the active facial movements observed during casual conversations, and (2) emotion-based expressions derived from Facial Action Coding System (FACS). These combinations were compared against minimal baseline and full blendshape conditions through a comprehensive study involving 32 participants. Our findings reveal that emotion-based condition achieves comparable levels of social presence and communication quality to the full condition, in both informative speeches and emotional conversations. This highlights the effectiveness of prioritizing emotion-based expressions and adopting a streamlined approach to avatar facial control. By focusing on emotional expressions while optimizing resources, this approach shows potential for enhancing the avatar-mediated communication experience, accommodating the diverse users’ contexts. Seoyoung Kang, Hail Song, Boram Yoon, Kangsoo Kim, Woontack Woo |
ISMAR | 2 |
| 2023 | RC-SMPL: Real-time Cumulative SMPL-based Avatar Body GenerationabstractWe present a novel method for avatar body generation that cumulatively updates the texture and normal map in real-time. Multiple images or videos have been broadly adopted to create detailed 3D human models that capture more realistic user identities in both Augmented Reality (AR) and Virtual Reality (VR) environments. However, this approach has a higher spatiotemporal cost because it requires a complex camera setup and extensive computational resources. For lightweight reconstruction of personalized avatar bodies, we design a system that progressively captures the texture and normal values using a single RGBD camera to generate the widely-accepted 3D parametric body model, SMPL-X. Quantitatively, our system maintains real-time performance while delivering reconstruction quality comparable to the state-of-the-art method. Moreover, user studies reveal the benefits of real-time avatar creation and its applicability in various collaborative scenarios. By enabling the production of high-fidelity avatars at a lower cost, our method provides more general way to create personalized avatar in AR/VR applications, thereby fostering more expressive self-representation in the metaverse. Hail Song, Boram Yoon, Woojin Cho 0002, Woontack Woo |
ISMAR | 1 |