VLDB 2026 Research / reviewers in the wild / expert
Woontack Woo
dblp:w/WoontackWoo
· DBLP profile ↗
117ranked-venue papers
5as first author
47since 2021 · last 2026
0000-0002-5501-4421ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 80 · 4 first-author · 37 since 2021Human-computer interaction and ubiquitous computing · 66 · 1 first-author · 25 since 2021Artificial intelligence and machine learning · 6 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | What Are You Really Asking For? A Comparative 5W1H Analysis of Learner Questioning in CPR Training with IVAs in Screen-based and Augmented Reality EnvironmentsabstractQuestion-asking is one of the key indicators of cognitive engagement. However, understanding how the distinct psychological affordances of presentation media shape learners' spoken inquiries with embodied Intelligent Virtual Agents (IVAs) remains limited. To systematically examine this process, we propose a 5W1H-based framework for analyzing learner questions. Using this framework, we conducted a user study comparing an Augmented Reality-based IVA (AR-IVA) deployed in the physical environment with a screen-based IVA (Video-IVA) during cardiopulmonary resuscitation (CPR) instruction. Results showed that the AR-IVA elicited higher spatial and social presence and promoted more frequent and longer questions focused on clarification and understanding. In contrast, the Video-IVA encouraged questions regarding procedural refinement. Presence acted as a selective filter, shaping the timing and topic of questions rather than as a universal mediator. These effects were significantly moderated by learners' motivational and strategic characteristics toward learning. Based on these findings, we propose design implications for IVA-supported learning systems. Hyerim Park, Jinseok Hong, Heejeong Ko, Woontack Woo |
CHI | 4 |
| 2026 | A Unified Hand and Gesture Tracking via Offloading Framework for Object-mediated Interaction in Wearable ARabstractWe propose a novel object-mediated hand interaction system that enables real-time operation with everyday objects on wearable augmented reality (AR) devices. Despite recent advances, both commercial and academic hand interaction techniques remain constrained, typically requiring external hardware or depending exclusively on bare-hand gestures. Motivated by these constraints, we developed an offloading framework that integrates a high-fidelity transformer-based 3D hand reconstruction model with a dynamic gesture recognition network powered by gated recurrent units (GRU). This architecture ensures stable and accurate gesture recognition even during interaction with physical objects. To evaluate its quantitative performance, we collected a custom dataset based on a predefined gesture set, achieving 93.0% accuracy in 5-fold cross-validation. The complete system implemented on Microsoft HoloLens 2 operates at a real-time framerate, and we further analyze the latency of each step in our framework. Through this interaction paradigm, users can experience immersive and intuitive AR in everyday environments with minimal disruption to natural action behavior. Our projects are available at https://github.com/kaist-uvrlab/UnifiedHOInteraction. Woojin Cho 0002, Taewook Ha, Taejun Son, Woontack Woo |
VR | 4 |
| 2026 | Task Breakpoint Generation using Origin-Centric Graph in Virtual Reality Recordings for Adaptive PlaybackabstractWe propose a method for generating task breakpoints based on an Origin-Centric Graph (OCG) to segment goal-oriented activity recordings into task units for adaptive playback in Virtual Reality (VR) environments. With the development of Augmented Reality (AR)/VR head-mounted displays (HMDs), research on adaptive tutorials and authoring tools has become active, but existing task segmentation methods mainly rely on manual annotation or are restricted to 2D video which limits their applicability to 3D VR contexts. In our approach, assembly scenarios with clearly defined task boundaries are recorded using a structured spatio-temporal scene graph (STSG), and the OCG is employed to track changes in the central object and the formation of new groups, thereby generating task breakpoints automatically. A user study collected user-perceived task breakpoints to establish ground truth (GT), and comparison with the algorithm-detected breakpoints demonstrated high agreement and confirmed accuracy in supporting adaptive playback. The proposed task segmentation method provides a foundation for dynamically adjusting VR playback according to user proficiency and progress, with potential for extension into automatic timeline segmentation systems for diverse VR recordings. Selin Choi, Dooyoung Kim 0001, Taewook Ha, Seonji Kim, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2026 | Int3DNet: Scene-Motion Cross Attention Network for 3D Intention Prediction in Mixed RealityabstractWe propose Int3DNet, a scene-aware network that predicts 3D intention areas directly from scene geometry and head-hand motion cues, enabling robust human intention prediction without explicit object-level perception. In Mixed Reality (MR), intention prediction is critical as it enables the system to anticipate user actions and respond proactively, reducing interaction delays and ensuring seamless user experiences. Our method employs a cross attention fusion of sparse motion cues and scene point clouds, offering a novel approach that directly interprets the user's spatial intention within the scene. We evaluated Int3DNet on MoGaze and CIRCLE datasets, which are public datasets for full-body human-scene interactions, showing consistent performance across time horizons of up to 1500 ms and outperforming the baselines, even in diverse and unseen scenes. Moreover, we demonstrate the usability of proposed method through a demonstration of efficient visual question answering (VQA) based on intention areas. Int3DNet provides reliable 3D intention areas derived from head-hand motion and scene geometry, thus enabling seamless interaction between humans and MR systems through proactive processing of intention areas. Taewook Ha, Woojin Cho 0002, Dooyoung Kim 0001, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | Streamlined Facial Data Collection Based on Utterance and Emotional Data for Human-to-Avatar ReconstructionabstractThis study explores a streamlined facial data collection method for conversational contexts, addressing the limitations of existing approaches that often require extensive datasets and prioritize technical metrics over user perception and experience. We systematically investigate which facial expression data are essential for reconstructing photorealistic avatars and how they can be captured efficiently. Our research employs a two-phase methodology to identify efficient facial data collection strategies and evaluate their effectiveness. In the first phase, we conduct facial data acquisition and evaluate reconstruction performance using utterance data and emotional data. In the second phase, we carry out a comprehensive user evaluation comparing three progressive conditions: utterance only, utterance and emotional data, and a control condition involving extensive data. Findings from 24 participants engaged in simulated face-to-face conversations reveal that targeted utterance and emotional data achieve comparable levels of perceived realism, naturalness, and telepresence, while reducing training time and data usage when compared to the extensive data collection approach. These results demonstrate that targeted data inputs can enable efficient avatar face reconstruction, offering practical guidelines for real-time applications such as AR/VR telepresence and highlighting the trade-off between data quantity and perceived quality. Seoyoung Kang, Seokhwan Yang, Hail Song, Boram Yoon, Kangsoo Kim, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2026 | SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB SequencesabstractWe introduce SceneLinker, a novel framework that generates compositional 3D scenes via semantic scene graph from RGB sequences. To adaptively experience Mixed Reality (MR) content based on each user's space, it is essential to generate a 3D scene that reflects the real-world layout by compactly capturing the semantic cues of the surroundings. Prior works struggled to fully capture the contextual relationship between objects or mainly focused on synthesizing diverse shapes, making it challenging to generate 3D scenes aligned with object arrangements. We address these challenges by designing a graph network with cross-check feature attention for scene graph prediction and constructing a graph-variational autoencoder (graph-VAE), which consists of a joint shape and layout block for 3D scene generation. Experiments on the 3RScan/3DSSG and SG-FRONT datasets demonstrate that our approach outperforms state-of-the-art methods in both quantitative and qualitative evaluations, even in complex indoor environments and under challenging scene graph constraints. Our work enables users to generate consistent 3D spaces from their physical environments via scene graphs, allowing them to create spatial MR content. Project page is https://scenelinker2026.github.io. Seokyoung Kim 0002, Dooyoung Kim 0001, Woojin Cho 0002, Hail Song, Suji Kang, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2026 | Spatial Affordance-Aware Affine Transformation Between Heterogeneous Spaces for Mixed Reality Remote CollaborationabstractWe propose a spatial affordance-aware affine transformation method between heterogeneous spaces for continuous multi-object matching in shared Mixed Reality (MR) spaces. While previous redirection and spatial mapping approaches utilize physical objects and walkable areas, a critical gap remains in enabling continuous mapping between dissimilar physical environments that supports both precise object alignment and seamless locomotion in a shared space. Our method structurally segments heterogeneous spaces into interaction zones and constructs affine patches based on object adjacency and facing configuration, enabling continuous correspondence. We evaluate our method using a dataset of paired dissimilar spaces and demonstrate that, unlike conventional grid-based methods, our approach achieves broader spatial alignment and richer object matching. The results show that our method can serve as an effective mapping framework for shared environments requiring semantic continuity and structural coherence across diverse real-world spaces. Seonji Kim, Dooyoung Kim 0001, Selin Choi, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | Event-Based Referred Vibrotactile Feedback for Bare-Hand XR InteractionabstractWe propose event-based vibrotactile feedback as an effective approach for enhancing bare-hand XR interaction, delivered through commodity smartwatches that users already wear. While controllers and haptic gloves provide rich tactile feedback, they introduce additional hardware that users must carry and wear throughout the day. Importantly, they interpose hardware between the user and the virtual world, occupying the hands, constraining finger motion, and undermining truly unobstructed bare-hand interaction. Our proposed approach delivers short discrete pulses at key manipulation moments such as object contact and state changes, providing tactile confirmation at the wrist without requiring additional hardware or disrupting finger movement. Two user studies with 26 participants demonstrate that this referred feedback significantly enhances user experience at either wrist placement, with 94% finding it helpful and particular benefits during complex tasks like knob rotation where haptic cues reduced visual attention demands. These findings establish design guidelines for integrating accessible haptics into everyday XR through personal devices, supporting broader adoption of natural hand interaction. Hyunseo Seo, Hyunjin Lee 0005, Minju Baeck, Hui-Shyong Yeo, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2026 | ForceCtrl: Hand-Raycasting With User-Defined Pinch Force for Control-Display Gain ApplicationabstractWe present ForceCtrl, a novel 3D hand raycasting technique that enhances pointing precision based on control-display (CD) gain controlled with user-defined pinch force. We introduce a target-agnostic approach for refining raycasting precision, overcoming limitations in human motor accuracy. User-defined pinch force, detected with surface electromyography (sEMG), enables users to easily activate or deactivate CD gain during interaction. We propose three CD gain strategies and compare them through target selection and placement tasks. Our system reduces selection errors, placement jitters, and user workload, especially for distant targets in high-difficulty tasks. These results highlight the effectiveness of applying CD gain to hand raycasting and demonstrate the potential of user-defined pinch force as a robust input modality for precise hand interaction in AR/VR. Seoyoung Oh, Junghoon Seo, Boram Yoon, Sang Ho Yoon, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2026 | VRGaussianAvatar: Integrating 3D Gaussian Avatars into VRabstractWe present VRGaussianAvatar, an integrated system that enables real-time full-body 3D Gaussian Splatting (3DGS) avatars in virtual reality using only head-mounted display (HMD) tracking signals. The system adopts a parallel pipeline with a VR Frontend and a GA Backend. The VR Frontend uses inverse kinematics to estimate full-body pose and streams the resulting pose along with stereo camera parameters to the backend. The GA Backend stereoscopically renders a 3DGS avatar reconstructed from a single image. To improve stereo rendering efficiency, we introduce Binocular Batching, which jointly processes left and right eye views in a single batched pass to reduce redundant computation and support high-resolution VR displays. We evaluate VRGaussianAvatar with quantitative performance tests and a within-subject user study against image- and video-based mesh avatar baselines. Results show that VRGaussianAvatar sustains interactive VR performance and yields higher perceived appearance similarity, embodiment, and plausibility. Project page and source code are available at https://vrgaussianavatar.github.io. Hail Song, Boram Yoon, Seokhwan Yang, Seoyoung Kang, Hyunjeong Kim, Henning Metzmacher, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2026 | OFERA: Blendshape-Driven 3D Gaussian Control for Occluded Facial Expression to Realistic Avatars in VRabstractWe propose OFERA, a novel framework for real-time expression control of photorealistic Gaussian head avatars for VR headset users. Existing approaches attempt to recover occluded facial expressions using additional sensors or internal cameras, but sensor-based methods increase device weight and discomfort, while camera-based methods raise privacy concerns and suffer from limited access to raw data. To overcome these limitations, we leverage the blendshape signals provided by commercial VR headsets as expression inputs. Our framework consists of three key components: (1) Blendshape Distribution Alignment (BDA), which applies linear regression to align the headset-provided blendshape distribution to a canonical input space; (2) an Expression Parameter Mapper (EPM) that maps the aligned blendshape signals into an expression parameter space for controlling Gaussian head avatars; and (3) a Mapper-integrated Avatar (MiA) that incorporates EPM into the avatar learning process to ensure distributional consistency. Furthermore, OFERA establishes an end-to-end pipeline that senses and maps expressions, updates Gaussian avatars, and renders them in real-time within VR environments. We show that EPM outperforms existing mapping methods on quantitative metrics, and we demonstrate through a user study that the full OFERA framework enhances expression fidelity while preserving avatar realism. By enabling real-time and photorealistic avatar expression control, OFERA significantly improves telepresence in VR communication. A project page is available at https://ysshwan147.github.io/projects/ofera/. Seokhwan Yang, Boram Yoon, Seoyoung Kang, Hail Song, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical StudyabstractThe remarkable reasoning and generalization capabilities of Large Language Models (LLMs) have paved the way for their expanding applications in embodied AI, robotics, and other real-world tasks. To effectively support these applications, grounding in spatial and temporal understanding in multimodal environments is essential. To this end, recent works have leveraged scene graphs, a structured representation that encodes entities, attributes, and their relationships in a scene. However, a comprehensive evaluation of LLMs’ ability to utilize scene graphs remains limited. In this work, we introduce Text-Scene Graph (TSG) Bench, a benchmark designed to systematically assess LLMs’ ability to (1) understand scene graphs and (2) generate them from textual narratives. With TSG Bench we evaluate 11 LLMs and reveal that, while models perform well on scene graph understanding, they struggle with scene graph generation, particularly for complex narratives. Our analysis indicates that these models fail to effectively decompose discrete scenes from a complex narrative, leading to a bottleneck when generating scene graphs. These findings underscore the need for improved methodologies in scene graph generation and provide valuable insights for future research. The demonstration of our benchmark is available at https://tsg-bench.netlify.app. Additionally, our code and evaluation data are publicly available at https://github.com/docworlds/tsg-bench. Dongil Yang, Minjin Kim, Sunghwan Kim 0005, Beong-woo Kwak, Minjun Park, Jinseok Hong, Woontack Woo, Jinyoung Yeo |
ACL (1) | 7 |
| 2025 | AReading with Smartphones: Understanding the Trade-offs between Enhanced Legibility and Display Switching Costs in Hybrid AR Interfaces
Sunyoung Bang, Hyunjin Lee 0005, Seoyoung Oh, Woontack Woo |
CHI | 4 |
| 2025 | Gender Congruence and Social Context in Xr: Effects on Partner Preference, Warmth, Competence, and UncanninessabstractAs immersive virtual environments become more prevalent, avatars serve as critical social interfaces. This study explores how combinations of visual appearance, vocal characteristics, and informed identity influence users' initial impressions and partner preferences in four distinct XR scenarios: physical, intellectual, social, and romantic. A within-subject experiment with 40 participants assessed perceived warmth, competence, uncanniness, and selection preferences across diverse avatar configurations. Results indicate that vocal cues had a particularly strong impact on social perception, often shaping feelings of approachability and clarity in communication. While some cue alignments enhanced perceived social comfort and engagement, inconsistencies across gender-related cues occasionally led to increased perceptions of uncanniness, especially in emotionally sensitive contexts. These findings highlight the importance of designing avatars that thoughtfully adapt to different interaction contexts, supporting inclusive and responsive user experiences in social XR platforms. Hyeongil Nam, Seoyoung Kang, Isaac Cho, Woontack Woo, Kangsoo Kim |
ISMAR | 5 |
| 2025 | Holistic quantified-self for context-aware wearable augmented reality
Eunhwa Song, Taewook Ha, Hyunjin Lee 0005, Woontack Woo |
Int. J. Hum. Comput. Stud. | 5 |
| 2025 | Visuo-Tactile Feedback with Hand Outline Styles for Modulating Affective Roughness PerceptionabstractWe propose a visuo-tactile feedback method that combines virtual hand visualization and fingertip vibrations to modulate affective roughness perception in VR. While prior work has focused on object-based textures and vibrotactile feedback, the role of visual feedback on virtual hands remains underexplored. Our approach introduces affective visual cues including line shape, motion, and color applied to hand outlines, and examines their influence on both affective responses (arousal, valence) and perceived roughness. Results show that sharp contours enhanced perceived roughness, increased arousal, and reduced valence, intensifying the emotional impact of haptic feedback. In contrast, color affected valence only, with red consistently lowering emotional positivity. These effects were especially noticeable at lower haptic intensities, where visual cues extended affective modulation into mid-level perceptual ranges. Overall, the findings highlight how integrating expressive visual cues with tactile feedback can enrich affective rendering and offer flexible emotional tuning in immersive VR interactions. Minju Baeck, Yoonseok Shin, Dooyoung Kim 0001, Hyunjin Lee 0005, Sang Ho Yoon, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | How Collaboration Context and Personality Traits Shape the Social Norms of Human-to-Avatar Identity RepresentationabstractAs avatars have evolved from simple digital representations into extensions of our identities, they offer unprecedented opportunities for self-expression and customization beyond the physical world limitations. While virtual platforms foster new forms of identity exploration, social norms still play a crucial role in defining what is considered appropriate in these environments. In this study, we surveyed 150 participants to investigate social norms surrounding avatar modifications, examining how perspectives, contexts, and personality traits influence attitudes toward appropriateness. Our findings reveal that avatar modifications are generally viewed as more appropriate when considered from a partner's perspective, especially for changeable attributes. However, these modifications are perceived as less acceptable in professional settings such as workplaces. Additionally, individuals with high self-monitoring tendencies tend to be more resistant to changes, while those scoring higher on Machiavellianism are more accepting of changes, particularly regarding unchangeable attributes and emotional expressions. These findings provide valuable insights for platform developers and designers, highlighting the importance of implementing context-aware customization options that balance core identity elements with personality-driven preferences, thereby enhancing user experiences while respecting social norms. Seoyoung Kang, Boram Yoon, Kangsoo Kim, Jonathan Gratch, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | Viewpoint-Tolerant Depth Perception for Shared Extended Space Experience on Wall-Sized DisplayabstractWe proposed viewpoint-tolerant shared depth perception without individual tracking by leveraging human cognitive compensation in universally 3D rendered images on a wall-sized display. While traditional 3D-perception-enabled display systems have primarily focused on single-user scenarios-adapting rendering based on head and eye tracking-the use of wall-sized displays to extend spatial experiences and support perceptually coherent multi-user interactions remains underexplored. We investigated the effects of virtual depths (dv) and absolute viewing distance (da) on human cognitive compensation factors (perceived distance difference, viewing angle threshold, and perceived presence) to construct the wall display-based eXtended Reality (XR) space. Results show that participants experienced a compelling depth perception even from off-center angles of 23°-37°, and largely increasing virtual depth worsens depth perception and presence factors, highlighting the importance of balancing extended depth of virtual space and viewing distance from the wall-sized display. Drawing on these findings, wall-sized displays in venues such as museums, galleries, and classrooms can evolve beyond 2D information sharing to offer immersive, spatially extended group experiences without individualized tracking or wearables. Dooyoung Kim 0001, Jinseok Hong, Heejeong Ko, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Effects of AI-Powered Embodied Avatars on Communication Quality and Social Connection in Asynchronous Virtual MeetingsabstractImmersive technologies such as virtual and augmented reality (VR/AR) allow remote users to meet and interact in a shared virtual space using embodied virtual avatars, creating a sense of co-presence. However, asynchronous communication-essential in many real-world contexts-remains underexplored in these environments. Traditional playback-based systems lack interactivity and often fail to preserve critical contextual cues necessary for effective asynchronous communication. In this paper, we introduce AVAGENTs, AI-powered virtual avatars that replicate users' verbal and nonverbal cues from recordings of past meetings. Avagents can interpret meeting context and generate appropriate responses to questions posed by asynchronous viewers. Through a user study (N = 30), we evaluated Avagents against a traditional playback method and a voice-based AI assistant across two asynchronous meeting scenarios: analytic reasoning and affective resonance. Results showed that Avagents enhance the asynchronous communication experience by increasing social presence, sense of belonging, emotional intimacy, and other user perceptions. We discuss the findings and their implications for designing effective AI-driven asynchronous communication tools in VR/AR environments. Hyeongil Nam, Muskan Sarvesh, Seoyoung Kang, Woontack Woo, Kangsoo Kim |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Comfortable Mobility vs. Attractive Scenery: The Key to Augmenting Narrative Worlds in Outdoor Locative Augmented Reality StorytellingabstractWe investigate how path context, encompassing both comfort and attractiveness, shapes user experiences in outdoor locative storytelling using Augmented Reality (AR). Addressing a research gap that predominantly concentrates on indoor settings or narrative backdrops, our user-focused research delves into the interplay between perceived path context and locative AR storytelling on routes with diverse walkability levels. We examine the correlation and causation between narrative engagement, spatial presence, perceived workload, and perceived path context. Our findings show that on paths with reasonable path walkability, attractive elements positively influence the narrative experience. However, even in environments with assured narrative walkability, inappropriate safety elements can divert user attention to mobility, hindering the integration of real-world features into the narrative. These results carry significant implications for path creation in outdoor locative AR storytelling, underscoring the importance of ensuring comfort and maintaining a balance between comfort and attractiveness to enrich the outdoor AR storytelling experience. Hyerim Park, Aram Min, Hyunjin Lee 0005, Maryam Shakeri, Ikbeom Jeon, Woontack Woo |
CHI | 6 |
| 2024 | Investigating the Design of Augmented Narrative Spaces Through Virtual-Real Connections: A Systematic Literature ReviewabstractAugmented Reality (AR) is regarded as an innovative storytelling medium that presents novel experiences by layering a virtual narrative space over a real 3D space. However, understanding of how the virtual narrative space and the real space are connected with one another in the design of augmented narrative spaces has been limited. For this, we conducted a systematic literature review of 64 articles featuring AR storytelling applications and systems in HCI, AR, and MR research. We investigated how virtual narrative spaces have been paired, functionalized, placed, and registered in relation to the real spaces they target. Based on these connections, we identified eight dominant types of augmented narrative spaces that are primarily categorized by whether they virtually narrativize reality or realize the virtual narrative. We discuss our findings to propose design recommendations on how virtual-real connections can be incorporated into a more structured approach to AR storytelling. Jae-eun Shin, Hayun Kim, Hyerim Park, Woontack Woo |
CHI | 4 |
| 2024 | Dense Hand-Object (HO) GraspNet with Full Grasping Taxonomy and Dynamics
Woojin Cho 0002, Minjae Yi, Taeyun Woo, Taewook Ha, Hyokeun Lee, Je-Hwan Ryu, Woontack Woo, Tae-Kyun Kim 0001 |
ECCV (82) | 10 |
| 2024 | Crowd Data-driven Artwork Placement in Virtual Exhibitions for Visitor Density Distribution PlanningabstractWe propose a novel crowd data-driven optimization approach for artwork placement in virtual exhibitions. With the emerging concept of Metaverse, a multitude of users can engage with content contemporaneously in virtual exhibitions. Yet, few studies have suggested a method to resolve crowd density concentration in multiuser Mixed Reality (MR) and Virtual Reality (VR) environments. In this study, our approach leverages crowd data engaged with artworks to predict optimal placement of artworks to distribute crowd density in virtual exhibitions, prior to exhibition planning. To investigate the requirement and validity of our approach, we conducted focus group interviews and an artwork relocation experiment as preliminary studies. In the generation of solution scenes for optimal placement, our optimizer adaptively scrutinizes placeable areas with considerations of crowd density distribution, scene rationality, and artwork similarity. Through a performance comparison analysis between optimization results, we confirmed that the optimizer successfully fulfilled the intended objectives with respect to the design considerations, resolving practical scenarios in exhibition planning. Jinseok Hong, Taewook Ha, Hyerim Park, Hayun Kim, Woontack Woo |
ISMAR | 5 |
| 2024 | Gender Differences in Perceiving Avatar Face and Interpersonal Distance: Exploring Realism and Social Presence in Mixed RealityabstractUnderstanding gender differences in facial and spatial recognition is crucial for enhancing avatar-mediated communication. However, there remains a gap in understanding how participant gender influences perceptions of avatar facial expressions and spatial dynamics in Mixed Reality communication. Therefore, our study investigates how avatar non-verbal cues interact with gender differences to affect user experience and understanding in MR environments. To examine these complex relationships, we conducted a user study comparing the effects of various avatar facial expressions (Full, Mouth-Only, and Emotion-based) and interpersonal distances (Closer vs. Farther) on facial animation realism and social presence, with a focus on gender-balanced participant groups. Our findings revealed that female participants were particularly sensitive to the avatar’s proximity and facial expressions, reporting significantly higher perceptions of facial animation realism, copresence, message understanding, and affective understanding at farther distances compared to male participants. They also perceived higher copresence and message understanding when exposed to emotion-based facial expressions, as opposed to a mouth-only condition-a distinction not observed among male participants. Based on our findings, we advocate for avatar design strategies that accommodate gender differences in perception and preference, potentially through customizable levels of expressiveness to cater to diverse user needs and contexts. Seoyoung Kang, Boram Yoon, Kangsoo Kim, Woontack Woo |
ISMAR | 5 |
| 2024 | The Influence of Emotion-based Prioritized Facial Expressions on Social Presence in Avatar-mediated Remote CommunicationabstractIn avatar-mediated remote communication, avatars’ facial expressions can be dynamically adjusted according to each user’s computational and device constraints, highlighting the importance of varied expressions and their impact on user perception. However, there is a lack of research on how variations in avatar facial expressions, especially when simplified, influence user perception, particularly in terms of social presence. To address this, we examine the impact of various facial expression combinations on social presence in avatar-mediated communication scenarios, ranging from informative speeches to emotional conversations. Our approach involves prioritizing avatar facial blendshape combinations using two main approaches: (1) commonly activated expressions that reflect the active facial movements observed during casual conversations, and (2) emotion-based expressions derived from Facial Action Coding System (FACS). These combinations were compared against minimal baseline and full blendshape conditions through a comprehensive study involving 32 participants. Our findings reveal that emotion-based condition achieves comparable levels of social presence and communication quality to the full condition, in both informative speeches and emotional conversations. This highlights the effectiveness of prioritizing emotion-based expressions and adopting a streamlined approach to avatar facial control. By focusing on emotional expressions while optimizing resources, this approach shows potential for enhancing the avatar-mediated communication experience, accommodating the diverse users’ contexts. Seoyoung Kang, Hail Song, Boram Yoon, Kangsoo Kim, Woontack Woo |
ISMAR | 5 |
| 2024 | Spatial Affordance-aware Interactable Subspace Allocation for Mixed Reality TelepresenceabstractTo enable remote Virtual Reality (VR) and Augmented Reality (AR) clients to collaborate as if they were in the same space during Mixed Reality (MR) telepresence, it is essential to overcome spatial heterogeneity and generate a unified shared collaborative environment by integrating remote spaces into a target host space. Especially when multiple remote users connect, a large shared space is necessary for people to maintain their personal space while collaborating, but the existing simple intersection method leads to the creation of narrow shared spaces as the number of remote spaces increases. To robustly align to the host space even as the number of remote spaces increases, we propose a spatial affordance-aware interactable subspace allocation algorithm. The key concept of our approach is to consider the perceivable and interactable areas separately, where every user views the same mutual space, but each remote user has a different interactable subspace, considering their location and spatial affordance. We conducted an evaluation with 900 space combinations, varying the number of remote spaces as two, four, and six, and results show our method outperformed in securing wide interactable mutual space and instantiating users compared to the other spatial matching methods. Our work enables multiple clients from diverse remote locations to access the AR host’s space, allowing them to interact directly with the table, wall, or floor by aligning their physical subspaces within a connected mutual space. Dooyoung Kim 0001, Seonji Kim, Selin Choi, Woontack Woo |
ISMAR | 4 |
| 2024 | Whirling Interface: Hand-based Motion Matching Selection for Small Target on XR DisplaysabstractWe introduce “Whirling Interface,” a selection method for XR displays using bare-hand motion matching gestures as an input technique. We extend the motion matching input method, by introducing different input states to provide visual feedback and guidance to the users. Using the wrist joint as the primary input modality, our technique reduces user fatigue and improves performance while selecting small and distant targets. In a study with 16 participants, we compared the whirling interface with a standard ray casting method using hand gestures. The results demonstrate that the Whirling Interface consistently achieves high success rates, especially for distant targets, averaging 95.58% with a completion time of 5.58 seconds. Notably, it requires a smaller camera sensing field of view of only 21.45° horizontally and 24.7° vertically. Participants reported lower workloads on distant conditions and expressed a higher preference for the Whirling Interface in general. These findings suggest that the Whirling Interface could be a useful alternative input method for XR displays with a small camera sensing FOV or when interacting with small targets. Seoyoung Oh, Minju Baeck, Hui-Shyong Yeo, Hyungil Kim, Thad Starner, Woontack Woo |
ISMAR | 7 |
| 2024 | Object Cluster Registration of Dissimilar Rooms Using Geometric Spatial Affordance Graph to Generate Shared Virtual SpacesabstractWe propose Object Cluster Registration (OCR) using Geometric Spatial Affordance Graph (GSAG) to support user interaction with multiple objects in a shared space generated from two dissimilar rooms. Previous research on generating a shared virtual space from dissimilar real spaces has only reflected the information of individual objects and aimed at maximizing the area of the shared space. This led to limited interactions relying on the singular affordances of objects, neglecting to consider the usability and effectiveness of the generated shared spaces. The proposed OCR with GSAG, which considers the relationship between objects based on facing formation, extracts optimal object cluster pairs to align dissimilar rooms in generating shared virtual spaces. In an evaluation study involving 100 multi-cluster space pairs, applying OCR using GSAG showed greater effectiveness in preserving object correlations compared to cases where OCR was not used. Furthermore, the size of the shared space did not significantly differ between the two methods. This suggests that factoring in the relationship between objects does not compromise the objective of maximizing the shared virtual space. The proposed method is expected to serve as a foundation for generating shared virtual spaces that are more user-oriented and efficient by facilitating a wider range of collaborative activities for remote users in dissimilar real spaces with varied configurations. Seonji Kim, Dooyoung Kim 0001, Jae-eun Shin, Woontack Woo |
VR | 4 |
| 2023 | How Space is Told: Linking Trajectory, Narrative, and Intent in Augmented Reality Storytelling for Cultural Heritage SitesabstractWe report on a qualitative study in which 22 participants created Augmented Reality (AR) stories for outdoor cultural heritage sites. As storytelling is a crucial strategy for AR content aimed at providing meaningful experiences, the emphasis has been on what storytelling does, rather than how it is done, the end user’s needs prioritized over the author’s. To address this imbalance, we identify how recurring patterns in the spatial trajectories and narrative compositions of AR stories for cultural heritage sites are linked to the author’s intent and creative process: While authors tend to bind story arcs tightly to confined trajectories for narrative delivery, the need for spatial exploration results in thematic content mapped loosely onto encompassing trajectories. Based on our analysis, we present design recommendations for site-specific AR storytelling tools that can support authors in delivering their intent while leveraging the placeness of cultural heritage sites as a creative resource. Jae-eun Shin, Woontack Woo |
CHI | 2 |
| 2023 | OmniSense: Exploring Novel Input Sensing and Interaction Techniques on Mobile Device with an Omni-Directional CameraabstractAn omni-directional (360°) camera captures the entire viewing sphere surrounding its optical center. Such cameras are growing in use to create highly immersive content and viewing experiences. When such a camera is held by a user, the view includes the user’s hand grip, finger, body pose, face, and the surrounding environment, providing a complete understanding of the visual world and context around it. This capability opens up numerous possibilities for rich mobile input sensing. In OmniSense, we explore the broad input design space for mobile devices with a built-in omni-directional camera and broadly categorize them into three sensing pillars: i) near device ii) around device and iii) surrounding device. In addition we explore potential use cases and applications that leverage these sensing capabilities to solve user needs. Following this, we develop a working system to put these concepts into action, by leveraging these sensing capabilities to enable potential use cases and applications. We studied the system in a technical evaluation and a preliminary user study to gain initial feedback and insights. Collectively these techniques illustrate how a single, omni-purpose sensor on a mobile device affords many compelling ways to enable expressive input, while also affording a broad range of novel applications that improve user experience during mobile interaction. Hui-Shyong Yeo, Erwin Wu, Daehwa Kim, Hyungil Kim, Seoyoung Oh, Luna Takagi, Woontack Woo, Hideki Koike, Aaron J. Quigley |
CHI | 8 |
| 2023 | Edge-Centric Space Rescaling with Redirected Walking for Dissimilar Physical-Virtual Space RegistrationabstractWe propose a novel space-rescaling technique for registering dissimilar physical-virtual spaces by utilizing the effects of adjusting physical space with redirected walking. Achieving a seamless and immersive Virtual Reality (VR) experience requires overcoming the spatial heterogeneities between the physical and virtual spaces and accurately aligning the VR environment with the user’s tracked physical space. However, existing space-matching algorithms that rely on one-to-one scale mapping are inadequate when dealing with highly dissimilar physical and virtual spaces, and redirected walking controllers could not utilize basic geometric information from physical space in the virtual space due to coordinate distortion. To address these issues, we apply relative translation gains to partitioned space grids based on the main interactable object’s edge, which enables space-adaptive modification effects of physical space without coordinate distortion. Our evaluation results demonstrate the effectiveness of our algorithm in aligning the main object’s edge, surface, and wall, as well as securing the largest registered area compared to alternative methods under all conditions. These findings can be used to create an immersive play area for VR content where users can receive passive feedback from the plane and edge in their physical environment. Dooyoung Kim 0001, Woontack Woo |
ISMAR | 2 |
| 2023 | RC-SMPL: Real-time Cumulative SMPL-based Avatar Body GenerationabstractWe present a novel method for avatar body generation that cumulatively updates the texture and normal map in real-time. Multiple images or videos have been broadly adopted to create detailed 3D human models that capture more realistic user identities in both Augmented Reality (AR) and Virtual Reality (VR) environments. However, this approach has a higher spatiotemporal cost because it requires a complex camera setup and extensive computational resources. For lightweight reconstruction of personalized avatar bodies, we design a system that progressively captures the texture and normal values using a single RGBD camera to generate the widely-accepted 3D parametric body model, SMPL-X. Quantitatively, our system maintains real-time performance while delivering reconstruction quality comparable to the state-of-the-art method. Moreover, user studies reveal the benefits of real-time avatar creation and its applicability in various collaborative scenarios. By enabling the production of high-fidelity avatars at a lower cost, our method provides more general way to create personalized avatar in AR/VR applications, thereby fostering more expressive self-representation in the metaverse. Hail Song, Boram Yoon, Woojin Cho 0002, Woontack Woo |
ISMAR | 4 |
| 2023 | Enhancing the Reading Experience on AR HMDs by Using Smartphones as Assistive DisplaysabstractThe reading experience on current augmented reality (AR) head mounted displays (HMDs) is often impeded by the devices' low perceived resolution, translucency, and small field of view, especially in situations involving lengthy text. Although many researchers have proposed methods to resolve this issue, the inherent characteristics prevent these displays from delivering a readability on par with that of more traditional displays. As a solution, we explore the use of smartphones as assistive displays to AR HMDs. To validate the feasibility of our approach, we conducted a user study in which we compared a smartphone-assisted hybrid interface against using the HMD only for two different text lengths. The results demonstrate that the hybrid interface yields a lower task load regardless of the text length, although it does not improve task performance. Furthermore, the hybrid interface provides a better experience regarding user comfort, visual fatigue, and perceived readability. Based on these results, we claim that joining the spatial output capabilities of the HMD with the high-resolution display of the smartphone is a viable solution for improving the reading experience in AR. Sunyoung Bang, Woontack Woo |
VR | 2 |
| 2023 | Exploring the Effects of Augmented Reality Notification Type and Placement in AR HMD while WalkingabstractAugmented reality (AR) helps users easily accept information when they are walking by providing virtual information in front of their eyes. However, it remains unclear how to present AR notifications considering the expected user reaction to interruption. Therefore, we investigated to confirm appropriate placement methods for each type by dividing it into notification types that are handled immediately (high) or that are performed later (low). We compared two coordinate systems (display-fixed and body-fixed) and three positions (top, right, and bottom) for the notification placement. We found significant effects of notification type and placement on how notifications are perceived during the AR notification experience. Using a display-fixed coordinate system responded faster for high notification types, whereas using a body-fixed coordinate system resulted in quick walking speed for low ones. As for the position, the high types had a higher notification performance at the bottom position, but the low types had enhanced walking performance at the right position. Based on the finding of our experiment, we suggest some recommendations for the future design of AR notification while walking. Hyunjin Lee 0005, Woontack Woo |
VR | 2 |
| 2023 | Seg&Struct: The Interplay Between Part Segmentation and Structure Inference for 3D Shape ParsingabstractWe propose Seg&Struct, a supervised learning framework leveraging the interplay between part segmentation and structure inference and demonstrating their synergy in an integrated framework. Both part segmentation and structure inference have been extensively studied in the recent deep learning literature, while the supervisions used for each task have not been fully exploited to assist the other task. Namely, structure inference has been typically conducted with an autoencoder that does not lever-age the point-to-part associations. Also, segmentation has been mostly performed without structural priors that tell the plausibility of the output segments. We present how these two tasks can be best combined while fully utilizing super-vision to improve performance. Our framework first decomposes a raw input shape into part segments using an off-the-shelf algorithm, whose outputs are then mapped to nodes in a part hierarchy, establishing point-to-part associations. Following this, ours predicts the structural information, e.g., part bounding boxes and part relationships. Lastly, the segmentation is rectified by examining the confusion of part boundaries using the structure-based part features. Our experimental results based on the StructureNet and PartNet demonstrate that the interplay between two tasks results in remarkable improvements in both tasks: 27.91% in structure inference and 0.5% in segmentation. Kaichun Mo, Minhyuk Sung, Woontack Woo |
WACV | 4 |
| 2023 | "Enjoy, but Moderately!": Designing a Social Companion Robot for Social Engagement and Behavior Moderation in Solitary Drinking ContextabstractSocially assistive robots can support people in making behavior changes by socially engaging in or moderating certain behaviors, such as physical exercise and snacking. However, there has not been much work on designing social robots that aim to support both social engagement and behavior moderation, i.e., offering social interactions for engaging in behaviors without over-engagement. This work explores how social robots can moderate alcohol consumption while socially engaging them in a solitary drinking context. As alcohol consumption can have benefits when done in moderation, this companion robot aims to guide the user toward moderate drinking by using social engagement (i.e., creating an enjoyable atmosphere) and drinking moderation (i.e., regulating the drinking pace). Our preliminary user study (n=20) reveals that the robot is perceived as a friendly companion, and its human-likeness is partly attributed to the robot's intervention. Most participants followed the robot's guidance and perceived it as an intelligent friend due to its social interactions and behavior tracking features. We discuss the benefit of physical interactions for social engagement, utilizing interaction rituals for enjoyable but moderate commensality, and ethical considerations in solitary drinking contexts. Yugyeong Jung, Gyuwon Jung, Sooyeon Jeong, Woontack Woo, Hwajung Hong, Uichin Lee |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2023 | Visualizing Hand Force with Wearable Muscle Sensing for Enhanced Mixed Reality Remote CollaborationabstractIn this paper, we present a prototype system for sharing a user's hand force in mixed reality (MR) remote collaboration on physical tasks, where hand force is estimated using wearable surface electromyography (sEMG) sensor. In a remote collaboration between a worker and an expert, hand activity plays a crucial role. However, the force exerted by the worker's hand has not been extensively investigated. Our sEMG-based system reliably captures the worker's hand force during physical tasks and conveys this information to the expert through hand force visualization, overlaid on the worker's view or on the worker's avatar. A user study was conducted to evaluate the impact of visualizing a worker's hand force on collaboration, employing three distinct visualization methods across two view modes. Our findings demonstrate that sensing and sharing hand force in MR remote collaboration improves the expert's awareness of the worker's task, significantly enhances the expert's perception of the collaborator's hand force and the weight of the interacting object, and promotes a heightened sense of social presence for the expert. Based on the findings, we provide design implications for future mixed reality remote collaboration systems that incorporate hand force sensing and visualization. Hyungil Kim, Boram Yoon, Seoyoung Oh, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | The Effects of Spatial Complexity on Narrative Experience in Space-Adaptive AR StorytellingabstractA critical yet unresolved challenge in designing space-adaptive narratives for Augmented Reality (AR) is to provide consistently immersive user experiences anywhere, regardless of physical features specific to a space. For this, we present a comprehensive analysis on a series of user studies investigating how the size, density, and layout of real indoor spaces affect users playing Fragments, a space-adaptive AR detective game. Based on the studies, we assert that moderate levels of traversability and visual complexity afforded in counteracting combinations of size and complexity are beneficial for narrative experience. To confirm our argument, we combined the experimental data of the studies (n=112) to compare how five different spatial complexity conditions impact narrative experience when applied to contrasting room sizes. Results show that whereas factors of narrative experience are rated significantly higher in relatively simple settings for a small space, they are less affected by complexity in a large space. Ultimately, we establish guidelines on the design and placement of space-adaptive augmentations in location-independent AR narratives to compensate for the lack or excess of affordances in various real spaces and enhance user experiences therein. Jae-eun Shin, Boram Yoon, Dooyoung Kim 0001, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Effects of Avatar Transparency on Social Presence in Task-Centric Mixed Reality Remote CollaborationabstractDespite the importance of avatar representation on user experience for Mixed Reality (MR) remote collaboration involving various device environments and large amounts of task-related information, studies on how controlling visual parameters for avatars can benefit users in such situations have been scarce. Thus, we conducted a user study comparing the effects of three avatars with different transparency levels (Nontransparent, Semi-transparent, and Near-transparent) on social presence for users in Augmented Reality (AR) and Virtual Reality (VR) during task-centric MR remote collaboration. Results show that avatars with a strong visual presence are not required in situations where accomplishing the collaborative task is prioritized over social interaction. However, AR users preferred more vivid avatars than VR users. Based on our findings, we suggest guidelines on how different levels of avatar transparency should be applied based on the context of the task and device type for MR remote collaboration. Boram Yoon, Jae-eun Shin, Hyungil Kim, Seoyoung Oh, Dooyoung Kim 0001, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | The Effects of Device and Spatial Layout on Social Presence During a Dynamic Remote Collaboration Task in Mixed RealityabstractThis paper evaluates factors of social presence during a dynamic remote collaboration task in a technologically asymmetric Mixed Reality (MR) setting for two spatial layouts. While active movement during MR remote collaboration is afforded by how the shared 3D space is mediated and configured, studies investigating the impact of these conditions on user experience have been scarce. In a between-group study $(\mathrm{n}=48)$, a host user in Augmented Reality (AR) and a remote user in Virtual Reality (VR), both wearing Head Mounted Displays (HMDs), simultaneously moved around the shared space to find and assemble parts of a Mars exploration rover together, one group in a Peripheral layout and the other in a Scattered layout with disparate levels of spatial affordance. Results show that while VR facilitates higher co-presence and spatial presence than AR through HMDs, the Peripheral layout enables users to pay more attention to one another than the Scattered. We analyze the results and derive implications aimed at bridging the AR-VR gap in social presence for dynamic MR remote collaboration through the adaptive placement of virtual content in shared spaces. Jae-eun Shin, Boram Yoon, Dooyoung Kim 0001, Hyungil Kim, Woontack Woo |
ISMAR | 5 |
| 2022 | Effects of Virtual Room Size and Objects on Relative Translation Gain Thresholds in Redirected WalkingabstractThis paper investigates how the size of virtual space and objects within it affect the threshold range of relative translation gains, a Redirected Walking (RDW) technique that scales the user’s movement in virtual space in different ratios for the width and depth. While previous studies assert that a virtual room’s size affects relative translation gain thresholds on account of the virtual horizon’s location, additional research is needed to explore this assumption through a structured approach to visual perception in Virtual Reality (VR). We estimate the relative translation gain thresholds in six spatial conditions configured by three room sizes and the presence of virtual objects (3 × 2), which were set according to differing Angles of Declination (AoDs) between eye-gaze and the forward-gaze. Results show that both size and virtual objects significantly affect the threshold range, it being greater in the large-sized condition and furnished condition. This indicates that the effect of relative translation gains can be further increased by constructing a perceived virtual movable space that is even larger than the adjusted virtual movable space and placing objects in it. Our study can be applied to adjust virtual spaces in synchronizing heterogeneous spaces without coordinate distortion where real and virtual objects can be leveraged to create realistic mutual spaces. Dooyoung Kim 0001, Jae-eun Shin, Boram Yoon, Jeongmi Lee, Woontack Woo |
VR | 6 |
| 2022 | Correction to: RealityBrush: an AR authoring system that captures and utilizes kinetic properties of everyday objects
Sanghwa Hong, Junki Kim, Taesoo Jang, Woontack Woo, Seongkook Heo, Byungjoo Lee |
Multim. Tools Appl. | 5 |
| 2021 | A User-Oriented Approach to Space-Adaptive Augmentation: The Effects of Spatial Affordance on Narrative Experience in an Augmented Reality Detective GameabstractSpace-adaptive algorithms aim to effectively align the virtual with the real to provide immersive user experiences for Augmented Reality(AR) content across various physical spaces. While such measures are reliant on real spatial features, efforts to understand those features from the user’s perspective and reflect them in designing adaptive augmented spaces have been lacking. For this, we compared factors of narrative experience in six spatial conditions during the gameplay of Fragments, a space-adaptive AR detective game. Configured by size and furniture layout, each condition afforded disparate degrees of traversability and visibility. Results show that whereas centered furniture clusters are suitable for higher presence in sufficiently large rooms, the same layout leads to lower narrative engagement. Based on our findings, we suggest guidelines that can enhance the effects of space adaptivity by considering how users perceive and navigate augmented space generated from different physical environments. Jae-eun Shin, Boram Yoon, Dooyoung Kim 0001, Woontack Woo |
CHI | 4 |
| 2021 | Adjusting Relative Translation Gains According to Space Size in Redirected Walking for Mixed Reality Mutual Space GenerationabstractWe propose the concept of relative translation gains, a novel Redirected Walking (RDW) method to create a mutual movable space between the Augmented Reality (AR) host's reference space and the Virtual Reality (VR) client's space. Previous RDW methods have focused on maximizing the movable space at the expense of aligning the coordinates between the AR and VR side, and could only be applied to collaborative scenarios involving sequential tasks. Our method solves these problems by adjusting the remote client's walking speed for each axis of a VR space to modify the movable area without coordinate distortion. We estimate the relative translation gain threshold, defined as the extent to which the walking speed can be altered without creating a perceived difference in distance. In order to reflect features of the reference space in generating the mutual space, we then examine how changing its size affects the threshold value. Our study showed that for remote clients connected to the larger reference space, relative translation gains can be increased to utilize a VR space bigger than their real space. Our method can be applied to create optimal mutual spaces for a wider variety of asymmetric Mixed Reality (MR) remote collaboration systems. Dooyoung Kim 0001, Jae-eun Shin, Jeongmi Lee, Woontack Woo |
VR | 4 |
| 2021 | Video Content Representation to Support the Hyper-reality Experience in Virtual RealityabstractMost research on providing location-based content in 3D interactive virtual reality has been limited to social media content. Few studies have suggested how to represent the video clip of movies or TV shows in virtual reality. This paper investigates a video content representation method to provide a hyper-reality experience of the narrative world in virtual reality. We reflect the time and place settings of the video content in virtual reality and have participants watch the video in four different virtual reality environments. We reveal that reflecting the story's environment settings to the virtual reality environment significantly improves the spatial presence and narratives engagement. We also confirm a positive correlation between spatial presence and narrative engagement, including subscales such as emotional engagement and narrative presence. Based on the study results, we discuss how to provide the hyper-reality experience in content-adaptive virtual reality. Hyerim Park, Woontack Woo |
VR | 2 |
| 2021 | RealityBrush: an AR authoring system that captures and utilizes kinetic properties of everyday objects
Sanghwa Hong, Junki Kim, Taesoo Jang, Woontack Woo, Seongkook Heo, Byungjoo Lee |
Multim. Tools Appl. | 5 |
| 2021 | Instant Panoramic Texture Mapping with Semantic Object Matching for Large-Scale Urban Scene ReproductionabstractThis paper proposes a novel panoramic texture mapping-based rendering system for real-time, photorealistic reproduction of large-scale urban scenes at a street level. Various image-based rendering (IBR) methods have recently been employed to synthesize high-quality novel views, although they require an excessive number of adjacent input images or detailed geometry just to render local views. While the development of global data, such as Google Street View, has accelerated interactive IBR techniques for urban scenes, such methods have hardly been aimed at high-quality street-level rendering. To provide users with free walk-through experiences in global urban streets, our system effectively covers large-scale scenes by using sparsely sampled panoramic street-view images and simplified scene models, which are easily obtainable from open databases. Our key concept is to extract semantic information from the given street-view images and to deploy it in proper intermediate steps of the suggested pipeline, which results in enhanced rendering accuracy and performance time. Furthermore, our method supports real-time semantic 3D inpainting to handle occluded and untextured areas, which appear often when the user's viewpoint dynamically changes. Experimental results validate the effectiveness of this method in comparison with the state-of-the-art approaches. We also present real-time demos in various urban streets. Ikbeom Jeon, Sung-Eui Yoon, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | Bare-hand Depth Inpainting for 3D Tracking of Hand Interacting with ObjectabstractWe propose a 3D hand tracking system using bare-hand depth inpainting from an RGB-depth image for a hand interacting with an object. The effectiveness of most existing hand-object tracking methods is impeded by the insufficiency of data, which do not include hand data occluded by the object, and their reliance on the information inferred from assuming the specific object type. We generate a sufficiently accurate bare-hand depth image from a hand interacting with an object using a conditional generative adversarial network, which is trained using the synthesized 2D silhouettes of the object to learn the morphology of the hand. We evaluate the proposed approach using a hierarchical particle filter-based hand tracker and prove that our approach utilizing the bare-hand tracker in the hand-object interaction dataset achieve state-of-the-art performance. The generalization of our work will enable visual-tactile interaction that is more natural in various wearable augmented reality applications. Woojin Cho 0002, Gabyong Park, Woontack Woo |
ISMAR | 3 |
| 2020 | 3D Hand Pose Estimation with a Single Infrared Camera via Domain Transfer LearningabstractPrevious methods successfully estimated 3D hand poses from unblurred depth images with slow and smooth hand motions. However, the performance drops when the depth images are contaminated by motion blur due to fast hand motion. In this paper, we exploit an infrared (IR) image input, which is weakly blurred under fast hand motion. The proposed method is based on domain transfer learning from depth to infrared images. Note we do not have IR images with hand skeletons, thus proposing self-supervision rather than direct supervision using the skeleton labels. We train a Hand Image Generator (HIG) and two Hand Pose Estimators (HPEs) on paired depth and infrared images via self-supervision using a consistency loss, guided by an existing HPE trained on paired depth and hand skeleton entries. The IR-based HPE is then refined on the weakly blurred infrared images. The qualitative and quantitative experiments demonstrate that the proposed method accurately estimates 3D hand poses under motion blur by fast hand motion, while existing depth-based methods fail. Our solution therefore supports fast 3D manipulation of virtual objects for augmented reality applications. Our model and dataset are publicly available for future research.1 Gabyong Park, Tae-Kyun Kim 0001, Woontack Woo |
ISMAR | 3 |
| 2020 | Evaluating Remote Virtual Hands Models on Social Presence in Hand-based 3D Remote CollaborationabstractThis study investigates the effects of a virtual hand representation on the user experience including social presence during hand-based 3D remote collaboration. Although a remote hand appearance is a critical parts of a hand-based telepresence, it has been rarely studied in comparison to studies on the self-embodiment of virtual hands in a 3D environment. Thus, we conducted a user study comparing the three virtual hands models (Skeleton, Low Polygon and Realistic) while performing a remote collaborative task based on the American Sign Language (ASL) using both Augmented Reality (AR) and Virtual Reality (VR) environments. We found that the realistic type was perceived as the most sense of being together, human-like, and trustable representation. The low polygon model could also convey a clear sign and moderate level of social presence. Although the system was configured asymmetrically in AR and VR, little difference in perception was found except for the participant's mental load and message understanding. We then discuss the results and suggest design implications for future hand-based 3D telepresence systems. Boram Yoon, Hyungil Kim, Seoyoung Oh, Woontack Woo |
ISMAR | 4 |
| 2020 | 3D Hand Tracking in the Presence of Excessive Motion BlurabstractWe present a sensor-fusion method that exploits a depth camera and a gyroscope to track the articulation of a hand in the presence of excessive motion blur. In case of slow and smooth hand motions, the existing methods estimate the hand pose fairly accurately and robustly, despite challenges due to the high dimensionality of the problem, self-occlusions, uniform appearance of hand parts, etc. However, the accuracy of hand pose estimation drops considerably for fast-moving hands because the depth image is severely distorted due to motion blur. Moreover, when hands move fast, the actual hand pose is far from the one estimated in the previous frame, therefore the assumption of temporal continuity on which tracking methods rely, is not valid. In this paper, we track fast-moving hands with the combination of a gyroscope and a depth camera. As a first step, we calibrate a depth camera and a gyroscope attached to a hand so as to identify their time and pose offsets. Following that, we fuse the rotation information of the calibrated gyroscope with model-based hierarchical particle filter tracking. A series of quantitative and qualitative experiments demonstrate that the proposed method performs more accurately and robustly in the presence of motion blur, when compared to state of the art algorithms, especially in the case of very fast hand rotations. Gabyong Park, Antonis A. Argyros, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | Physically-inspired Deep Light Estimation from a Homogeneous-Material Object for Mixed Reality LightingabstractIn mixed reality (MR), augmenting virtual objects consistently with real-world illumination is one of the key factors that provide a realistic and immersive user experience. For this purpose, we propose a novel deep learning-based method to estimate high dynamic range (HDR) illumination from a single RGB image of a reference object. To obtain illumination of a current scene, previous approaches inserted a special camera in that scene, which may interfere with user's immersion, or they analyzed reflected radiances from a passive light probe with a specific type of materials or a known shape. The proposed method does not require any additional gadgets or strong prior cues, and aims to predict illumination from a single image of an observed object with a wide range of homogeneous materials and shapes. To effectively solve this ill-posed inverse rendering problem, three sequential deep neural networks are employed based on a physically-inspired design. These networks perform end-to-end regression to gradually decrease dependency on the material and shape. To cover various conditions, the proposed networks are trained on a large synthetic dataset generated by physically-based rendering. Finally, the reconstructed HDR illumination enables realistic image-based lighting of virtual objects in MR. Experimental results demonstrate the effectiveness of this approach compared against state-of-the-art methods. The paper also suggests some interesting MR applications in indoor and outdoor scenes. Hunmin Park, Sung-Eui Yoon, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | Evaluating the Combination of Visual Communication Cues for HMD-based Mixed Reality Remote CollaborationabstractMany researchers have studied various visual communication cues (e.g. pointer, sketching, and hand gesture) in Mixed Reality remote collaboration systems for real-world tasks. However, the effect of combining them has not been so well explored. We studied the effect of these cues in four combinations: hand only, hand + pointer, hand + sketch, and hand + pointer + sketch, with three problem tasks: Lego, Tangram, and Origami. The study results showed that the participants completed the task significantly faster and felt a significantly higher level of usability when the sketch cue is added to the hand gesture cue, but not with adding the pointer cue. Participants also preferred the combinations including hand and sketch cues over the other combinations. However, using additional cues (pointer or sketch) increased the perceived mental effort and did not improve the feeling of co-presence. We discuss the implications of these results and future research directions. Seungwon Kim, Gun A. Lee, Weidong Huang 0001, Hayun Kim, Woontack Woo, Mark Billinghurst |
CHI | 5 |
| 2019 | SelfSync: exploring self-synchronous body-based hotword gestures for initiating interactionabstractSelfSync enables rapid, robust initiation of a gesture interface using synchronized movement of different body parts. SelfSync is the gestural equivalent of a hotword such as OK-Google in a speech interface and is enabled by the increasing trend where a user wears two or more wearables, such as a smartwatch, wireless earbuds, or a smartphone. In a user study comparing five potential SelfSync gestures in isolation, our system averages 96%, 98% and 88% for user dependent, user adapted, and user independent accuracy, respectively. For when the user has a phone in a pocket and a smart-watch, we suggest twisting the hand about the wrist while moving the leg with the phone in synchrony left and right. When the user has a head worn device and a smartwatch, we suggest twisting the hand while twisting the head left and right. Shaurye Aggarwal, Jason Wu 0001, Thad Starner, Woontack Woo |
UbiComp | 5 |
| 2019 | Is Any Room Really OK? The Effect of Room Size and Furniture on Presence, Narrative Engagement, and Usability During a Space-Adaptive Augmented Reality GameabstractOne of the main challenges in creating narrative-driven Augmented Reality (AR) content for Head Mounted Displays (HMDs) is to make them equally accessible and enjoyable in different types of indoor environments. However, little has been studied in regards to whether such content can indeed provide similar, if not the same, levels of experience across different spaces. To gain more understanding towards this issue, we examine the effect of room size and furniture on the player experience of Fragments, a space-adaptive, indoor AR crime-solving game created for the Microsoft HoloLens. The study compares factors of player experience in four types of spatial conditions: (1) Large Room - Fully Furnished; (2) Large Room - Scarcely Furnished; (3) Small Room - Fully Furnished; and (4) Small Room - Scarcely Furnished. Our results show that while large spaces facilitate a higher sense of presence and narrative engagement, fully-furnished rooms raise perceived workload. Based on our findings, we propose design suggestions that can support narrative-driven, space-adaptive indoor HMD-based AR content in delivering optimal experiences for various types of rooms. Jae-eun Shin, Hayun Kim, Callum Parker, Hyungil Kim, Seoyoung Oh, Woontack Woo |
ISMAR | 6 |
| 2019 | WRIST: Watch-Ring Interaction and Sensing Technique for Wrist Gestures and Macro-Micro PointingabstractTo better explore the incorporation of pointing and gesturing into ubiquitous computing, we introduce WRIST, an interaction and sensing technique that leverages the dexterity of human wrist motion. WRIST employs a sensor fusion approach which combines inertial measurement unit (IMU) data from a smartwatch and a smart ring. The relative orientation difference of the two devices is measured as the wrist rotation that is independent from arm rotation, which is also position and orientation invariant. Employing our test hardware, we demonstrate that WRIST affords and enables a number of novel yet simplistic interaction techniques, such as (i) macro-micro pointing without explicit mode switching and (ii) wrist gesture recognition when the hand is held in different orientations (e.g., raised or lowered). We report on two studies to evaluate the proposed techniques and we present a set of applications that demonstrate the benefits of WRIST. We conclude with a discussion of the limitations and highlight possible future pathways for research in pointing and gesturing with wearable devices. Hui-Shyong Yeo, Hyungil Kim, Aakar Gupta, Andrea Bianchi, Daniel Vogel 0001, Hideki Koike, Woontack Woo, Aaron J. Quigley |
MobileHCI | 8 |
| 2019 | Novel View Synthesis with Multiple 360 Images for Large-Scale 6-DOF Virtual Reality SystemabstractWe present a novel view synthesis method that allows users to experience a large-scale Six-Degree-of-Freedom (6-DOF) virtual environment. Our main contributions are the construction of a large-scale 6-DOF virtual environment using multiple 360 images as well as synthesis of a scene from novel viewpoints. Novel view synthesis from a single 360 image can give free viewpoint experience with full 6-DOF of head motion to players, but the moveable space is limited within a context of the image. We propose a novel view synthesis process that references multiple 360 images via reconstructing a large-scale real world based virtual data map and perform a weighted blending for interpolating multiple novel view images. Our results show that our approach provides a wider area of virtual environment as well as a smooth transition between each reference 360 images. Hochul Cho, Jangyoon Kim, Woontack Woo |
VR | 3 |
| 2019 | The Effect of Avatar Appearance on Social Presence in an Augmented Reality Remote CollaborationabstractThis paper investigates the effect of avatar appearance on Social Presence and users' perception in an Augmented Reality (AR) telep-resence system. Despite the development of various commercial 3D telepresence systems, there has been little evaluation and discussions about the appearance of the collaborator's avatars. We conducted two user studies comparing the effect of avatar appearances with three levels of body part visibility (head & hands, upper body, and whole body) and two different character styles (realistic and cartoon-like) on Social Presence while performing two different remote collaboration tasks. We found that a realistic whole body avatar was perceived as being the best for remote collaboration, but an upper body or cartoon style could be considered as a substitute depending on the collaboration context. We discuss these results and suggest guidelines for designing future avatar-mediated AR remote collaboration systems. Boram Yoon, Hyungil Kim, Gun A. Lee, Mark Billinghurst, Woontack Woo |
VR | 5 |
| 2018 | Hybrid 3D Hand Articulations Tracking Guided by Classification and Search Space AdaptationabstractWe propose a novel method for model-based 3D tracking of hand articulations that is effective even for fastmoving hand postures in depth images. A large number of augmented reality (AR) and virtual reality (VR) studies have used model-based approaches for estimating hand postures and tracking movements. However, these approaches exhibit limitations if the hand moves rapidly or into the camera's field of view. To overcome these problems, researchers attempted a hybrid strategy that uses multiple initializations for 3D tracking of articulations. However, this strategy also exhibits limitations. For example, in genetic optimization, the hypotheses generated from the previous solution may search for a solution in an incorrect search space in a fast-moving hand gesture. This problem also occurs if the search space selected from the results of a trained model does not cover the true solution although the tracked hand moves slowly. Our proposed method estimates the hand pose based on model-based tracking guided by classification and search space adaptation. From the classification by a convolutional neural network (CNN), a data-driven prior is included in the objective function and additional hypotheses are generated in particle swarm optimization (PSO). In addition, the search spaces of the two sets of the hypotheses, generated by the data-driven prior and the previous solution, are adaptively updated using the distribution of each set of the hypotheses. We demonstrated the effectiveness of the proposed method by applying it to an American Sign Language (ASL) dataset consisting of fast-moving hand postures. The experimental results demonstrate that the proposed algorithm exhibits more accurate tracking results compared to other state-of-the-art tracking algorithms. Gabyong Park, Woontack Woo |
ISMAR | 2 |
| 2017 | Ontology-based mobile augmented reality in cultural heritage sites: information modeling and user study
Hayun Kim, Tamás Matuszka, Jea In Kim, Jungwha Kim, Woontack Woo |
Multim. Tools Appl. | 5 |
| 2017 | Metaphoric Hand Gestures for Orientation-Aware VR Object Manipulation With an Egocentric ViewpointabstractWe present a novel natural user interface framework, called Meta-Gesture, for selecting and manipulating rotatable virtual reality (VR) objects in egocentric viewpoint. Meta-Gesture uses the gestures of holding and manipulating the tools of daily use. Specifically, the holding gesture is used to summon a virtual object into the palm, and the manipulating gesture to trigger the function of the summoned virtual tool. Our contributions are broadly threefold: 1) Meta-Gesture is the first to perform bare hand-gesture-based orientation-aware selection and manipulation of very small (nail-sized) VR objects, which has become possible by combining a stable 3-D palm pose estimator (publicly available) with the proposed static-dynamic (SD) gesture estimator; 2) the proposed novel SD random forest, as an SD gesture estimator can classify a 3-D static gesture and its action status hierarchically, in a single classifier; and 3) our novel voxel coding scheme, called layered shape pattern, which is configured by calculating the fill rate of point clouds (raw source of data) in each voxel on the top of the palm pose estimation, allows for dispensing with the need for preceding hand skeletal tracking or joint classification while defining a gesture. Experimental results show that the proposed method can deliver promising performance, even under frequent occlusions, during orientation-aware selection and manipulation of objects in VR space by wearing head-mounted display with an attached egocentric-depth camera (see the supplementary video available at: http://ieeexplore.ieee.org). Youngkyoon Jang, Ikbeom Jeon, Tae-Kyun Kim 0001, Woontack Woo |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2017 | TunnelSlice: Freehand Subspace Acquisition Using an Egocentric Tunnel for Wearable Augmented RealityabstractIn this paper, we propose TunnelSlice, which enables natural acquisition of subspace in an augmented scene from an egocentric view, even for scenarios involving ambiguous center objects or object occlusion. In wearable augmented reality (AR), approaching a three-dimensional (3-D) region including the objects of interest has become more important than approaching distant objects one by one. However, existing ray-based volumetric selection through a head worn display accompanies difficulties in defining a desired 3-D region due to obstacles by occlusion and depth perception. The proposed TunnelSlice effectively determines a cuboid transform, excluding unnecessary areas of a user-defined tunnel via two-handed pinch-based procedural slicing from an egocentric view. Through six scenarios involving central object status and different occlusion levels, we conducted a user study of TunnelSlice. Compared with two existing approaches, TunnelSlice was preferred by the subjects and showed greater stability for all scenarios, and outperformed the other approaches in a scenario involving strong occlusion without a central object. TunnelSlice is thus expected to serve as a key technology for spatial protocol and interaction using a subspace in wearable AR. Hyeongmook Lee, Seungtak Noh, Woontack Woo |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2016 | Mirror Mirror: An On-Body T-shirt Design SystemabstractVirtual fitting rooms equipped with magic mirrors let people evaluate fashion items without actually putting them on. The mirrors superimpose virtual clothes on the user's reflection. We contribute the Mirror Mirror system, which not only supports mixing and matching of existing fashion items, but also lets users design new items in front of the mirror and export designs to fabric printers. While much of the related work deals with interactive cloth simulation on live user data, we focus on collaborative design activities and explore various ways of designing on the body with a mirror. Daniel Saakes, Hui-Shyong Yeo, Seungtak Noh, Gyeol Han, Woontack Woo |
CHI | 5 |
| 2016 | Smooth eye movement interaction using EOG glassesabstractOrbits combines a visual display and an eye motion sensor to allow a user to select between options by tracking a cursor with the eyes as the cursor travels in a circular path around each option. Using an off-the-shelf Jins MEME pair of eyeglasses, we present a pilot study that suggests that the eye movement required for Orbits can be sensed using three electrodes: one in the nose bridge and one in each nose pad. For forced choice binary selection, we achieve a 2.6 bits per second (bps) input rate at 250ms per input. We also inntroduce Head Orbits, where the user fixates the eyes on a target and moves the head in synchrony with the orbiting target. Measuring only the relative movement of the eyes in relation to the head, this method achieves a maximum rate of 2.0 bps at 500ms per input. Finally, we combine the two techniques together with a gyro to create an interface with a maximum input rate of 5.0 bps. Murtaza Dhuliawala, Junichi Shimizu, Andreas Bulling, Kai Kunze, Thad Starner, Woontack Woo |
ICMI | 7 |
| 2015 | Physiological evidence for a dual process model of the social effects of emotion in computers
Ahyoung Choi, Celso de Melo, Peter Khooshabeh, Woontack Woo, Jonathan Gratch |
Int. J. Hum. Comput. Stud. | 4 |
| 2015 | 3D Finger CAPE: Clicking Action and Position Estimation under Self-Occlusions in Egocentric ViewpointabstractIn this paper we present a novel framework for simultaneous detection of click action and estimation of occluded fingertip positions from egocentric viewed single-depth image sequences. For the detection and estimation, a novel probabilistic inference based on knowledge priors of clicking motion and clicked position is presented. Based on the detection and estimation results, we were able to achieve a fine resolution level of a bare hand-based interaction with virtual objects in egocentric viewpoint. Our contributions include: (i) a rotation and translation invariant finger clicking action and position estimation using the combination of 2D image-based fingertip detection with 3D hand posture estimation in egocentric viewpoint. (ii) a novel spatio-temporal random forest, which performs the detection and estimation efficiently in a single framework. We also present (iii) a selection process utilizing the proposed clicking action detection and position estimation in an arm reachable AR/VR space, which does not require any additional device. Experimental results show that the proposed method delivers promising performance under frequent self-occlusions in the process of selecting objects in AR/VR space whilst wearing an egocentric-depth camera-attached HMD. Youngkyoon Jang, Seungtak Noh, Hyung Jin Chang, Tae-Kyun Kim 0001, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2014 | WeARHand: Head-worn, RGB-D camera-based, bare-hand user interface with visually enhanced depth perceptionabstractWe introduce WeARHand, which allows a user to manipulate virtual 3D objects with a bare hand in a wearable augmented reality (AR) environment. Our method uses no environmentally tethered tracking devices and localizes a pair of near-range and far-range RGB-D cameras mounted on a head-worn display and a moving bare hand in 3D space by exploiting depth input data. Depth perception is enhanced through egocentric visual feedback, including a semi-transparent proxy hand. We implement a virtual hand interaction technique and feedback approaches, and evaluate their performance and usability. The proposed method can apply to many 3D interaction scenarios using hands in a wearable AR environment, such as AR information browsing, maintenance, design, and games. Taejin Ha, Steven K. Feiner, Woontack Woo |
ISMAR | 3 |
| 2014 | Adaptive content recommendation for mobile users: Ordering recommendations using a hierarchical context model with granularityabstractRetrieving timely and relevant information on-site is an important task for mobile users. A context-aware system can understand a user’s information needs and thus select contents according to relevance. We propose a context-dependent search engine that represents user context in a knowledge-based context model, implemented in a hierarchical structure with granularity information. Search results are ordered based on semantic relevance computed as similarity between the current context and tags of search results. Compared against baseline algorithms, the proposed approach enhances precision by 22% and pooled recall by 17%. The use of size-based granularity to compute similarity makes the approach more robust against changes in the context model in comparison to graph-based methods, facilitating import of existing knowledge repositories and end-user defined vocabularies (folksonomies). The reasoning engine being light-weight, privacy protection is ensured, as all user information is processed locally on the user’s phone without requiring communication with an external server. Jonghyun Han, Hedda R. Schmidtke, Xing Xie 0001, Woontack Woo |
Pervasive Mob. Comput. | 4 |
| 2014 | Vision-based all-in-one solution for augmented reality and its storytelling applications
Kiyoung Kim, Nohyoung Park, Woontack Woo |
Vis. Comput. | 3 |
| 2013 | Unified Visual Perception Model for context-aware wearable ARabstractWe propose Unified Visual Perception Model (UVPM), which imitates the human visual perception process, for the stable object recognition necessarily required for augmented reality (AR) in the field. The proposed model is designed based on the theoretical bases in the field of cognitive informatics, brain research and psychological science. The proposed model consists of Working Memory (WM) in charge of low-level processing (in a bottomup manner), Long-Term Memory (LTM) and Short-Term Memory (STM), which are in charge of high-level processing (in a top-down manner). WM and LTM/STM are mutually complementary to increase recognition accuracies. By implementing the initial prototype of each boxes of the model, we could know that the proposed model works for stable object recognition. The proposed model is available to support context-aware AR with the optical see-through HMD. Youngkyoon Jang, Woontack Woo |
ISMAR | 2 |
| 2013 | IMAF: in situ indoor modeling and annotation framework on mobile phones
Gerhard Reitmayr, Woontack Woo |
Pers. Ubiquitous Comput. | 3 |
| 2012 | Real-time interactive modeling and scalable multiple object tracking for AR
Kiyoung Kim, Vincent Lepetit, Woontack Woo |
Comput. Graph. | 3 |
| 2012 | An interactive 3D movement path manipulation method in an augmented reality environmentabstractIn this paper, we evaluate a path editing method using a tangible user interface to generate and manipulate the movement path of a 3D object in an Augmented Reality (AR) scene. To generate the movement path, each translation point of a real 3D manipulation prop is examined to determine which point should be used as a control point for the path. Interpolation using splines is then used to reconstruct the path with a smooth line. A dynamic score-based selection method is also used to effectively select small and dense control points of the path. In an experimental evaluation, our method took the same time and generated a similar amount of errors as a more traditional approach, however the number of control points needed was significantly reduced. For control manipulation, the task completion time was quicker and there was less hand movement needed. Our method can be applied to drawing or curve editing methods in AR educational, gaming, and simulation applications. Taejin Ha, Mark Billinghurst, Woontack Woo |
Interact. Comput. | 3 |
| 2012 | Affective engagement to emotional facial expressions of embodied social agents in a decision-making gameabstractABSTRACT Previous research illustrates that people can be influenced by the emotional displays of computer‐generated agents. What is less clear is if these influences arise from cognitive or affective process (i.e., do people use agent displays as information or do they provoke user emotions). To unpack these processes, we examine the decisions and physiological reactions of participants (heart rate and electrodermal activity) when engaged in a decision task (prisoner's dilemma game) with emotionally expressive agents. Our results replicate findings that people's decisions are influenced by such emotional displays, but these influences differ depending on the extent to which these displays provoke an affective response. Specifically, we show that an individual difference known as electrodermal lability predicts the extent to whether people will engage affectively or strategically with such agents, thereby better predicting their decisions. We discuss implications for designing agent facial expressions to enhance social interaction between humans and agents. Copyright © 2012 John Wiley & Sons, Ltd. Ahyoung Choi, Celso de Melo, Woontack Woo, Jonathan Gratch |
Comput. Animat. Virtual Worlds | 3 |
| 2012 | Social itinerary recommendation from user-generated digital trails
Hyoseok Yoon, Yu Zheng 0004, Xing Xie 0001, Woontack Woo |
Pers. Ubiquitous Comput. | 4 |
| 2012 | Handling Motion-Blur in 3D Tracking and Rendering for Augmented RealityabstractThe contribution of this paper is two-fold. First, we show how to extend the ESM algorithm to handle motion blur in 3D object tracking. ESM is a powerful algorithm for template matching-based tracking, but it can fail under motion blur. We introduce an image formation model that explicitly consider the possibility of blur, and shows its results in a generalization of the original ESM algorithm. This allows to converge faster, more accurately and more robustly even under large amount of blur. Our second contribution is an efficient method for rendering the virtual objects under the estimated motion blur. It renders two images of the object under 3D perspective, and warps them to create many intermediate images. By fusing these images we obtain a final image for the virtual objects blurred consistently with the captured image. Because warping is much faster than 3D rendering, we can create realistically blurred images at a very low computational cost. Vincent Lepetit, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2011 | Interactive annotation on mobile phones for real and virtual space registrationabstractRegistration of real space and virtual information is a fundamental requirement for any augmented reality system. This paper presents an interactive method to quickly create a 3D room model and annotate locations within the room to provide registration anchors for virtual information. The method operates on a mobile phone and uses a visual rotation tracker to obtain orientation tracking for in-situ applications. The simple interaction allows non-expert users to create models of their environment and thus contribute marked-up representations to an online AR platform. Gerhard Reitmayr, Woontack Woo |
ISMAR | 3 |
| 2011 | Texture-less object tracking with online training using an RGB-D cameraabstractWe propose a texture-less object detection and 3D tracking method which automatically extracts on the fly the information it needs from color images and the corresponding depth maps. While texture-less 3D tracking is not new, it requires a prior CAD model, and real-time methods for detection still have to be developed for robust tracking. To detect the target, we propose to rely on a fast template-based method, which provides an initial estimate of its 3D pose, and we refine this estimate using the depth and image contours information. We automatically extract a 3D model for the target from the depth information. To this end, we developed methods to enhance the depth map and to stabilize the 3D pose estimation. We demonstrate our method on challenging sequences exhibiting partial occlusions and fast motions. Vincent Lepetit, Woontack Woo |
ISMAR | 3 |
| 2011 | Silhouette Segmentation in Multiple ViewsabstractIn this paper, we present a method for extracting consistent foreground regions when multiple views of a scene are available. We propose a framework that automatically identifies such regions in images under the assumption that, in each image, background and foreground regions present different color properties. To achieve this task, monocular color information is not sufficient and we exploit the spatial consistency constraint that several image projections of the same space region must satisfy. Combining the monocular color consistency constraint with multiview spatial constraints allows us to automatically and simultaneously segment the foreground and background regions in multiview images. In contrast to standard background subtraction methods, the proposed approach does not require a priori knowledge of the background nor user interaction. Experimental results under realistic scenarios demonstrate the effectiveness of the method for multiple camera set ups. Wonwoo Lee, Woontack Woo, Edmond Boyer |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Enhancing and evaluating users' social experience with a mobile phone guide applied to cultural heritage
Youngjung Suh, Choonsung Shin, Woontack Woo, Steven Dow, Blair MacIntyre |
Pers. Ubiquitous Comput. | 3 |
| 2011 | Video-Based In Situ Tagging on Mobile PhonesabstractWe propose a novel way to augment a real-world scene with minimal user intervention on a mobile phone; the user only has to point the phone camera to the desired location of the augmentation. Our method is valid for horizontal or vertical surfaces only, but this is not a restriction in practice in manmade environments, and it avoids going through any reconstruction of the 3-D scene, which is still a delicate process on a resource-limited system like a mobile phone. Our approach is inspired by recent work on perspective patch recognition, but we adapt it for better performances on mobile phones. We reduce user interaction with real scenes by exploiting the phone accelerometers to relax the need for fronto-parallel views. As a result, we can learn a planar target in situ from arbitrary viewpoints and augment it with virtual objects in real-time on a mobile phone. Wonwoo Lee, Vincent Lepetit, Woontack Woo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2011 | Extended Keyframe Detection with Stable Tracking for Multiple 3D Object TrackingabstractWe present a method that is able to track several 3D objects simultaneously, robustly, and accurately in real time. While many applications need to consider more than one object in practice, the existing methods for single object tracking do not scale well with the number of objects, and a proper way to deal with several objects is required. Our method combines object detection and tracking: frame-to-frame tracking is less computationally demanding but is prone to fail, while detection is more robust but slower. We show how to combine them to take the advantages of the two approaches and demonstrate our method on several real sequences. Vincent Lepetit, Woontack Woo |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2010 | Keyframe-based modeling and tracking of multiple 3D objectsabstractWe propose a real-time solution for modeling and tracking multiple 3D objects in unknown environments. Our contribution is two-fold: First, we show how to scale with the number of objects. This is done by combining recent techniques for image retrieval and online Structure from Motion, which can be run in parallel. As a result, tracking 40 objects in 3D can be done within 6 to 25 milliseconds per frame, even under difficult conditions for tracking. Second, we propose a method to let the user add new objects very quickly. The user simply has to select in an image a 2D region lying on the object. A 3D primitive is then fitted to the features within this region, and adjusted to create the object 3D model. In practice, this procedure takes less than a minute. Kiyoung Kim, Vincent Lepetit, Woontack Woo |
ISMAR | 3 |
| 2010 | Point-and-shoot for ubiquitous tagging on mobile phonesabstractWe propose a novel way to augment a real scene with minimalist user intervention on a mobile phone: The user only has to point the phone camera to the desired location of the augmentation. Our method is valid for vertical or horizontal surfaces only, but this is not a restriction in practice in man-made environments, and avoids to go through any reconstruction of the 3D scene, which is still a delicate process. Our approach is inspired by recent work on perspective patch recognition and we show how to modify it for better performances on mobile phones and how to exploit the phone accelerometers to relax the need for fronto-parallel views. In addition, our implementation allows to share the augmentations and the required data over peer-to-peer communication to build a shared AR space on mobile phones. Wonwoo Lee, Vincent Lepetit, Woontack Woo |
ISMAR | 4 |
| 2010 | Smart Itinerary Recommendation Based on User-Generated GPS Trajectories
Hyoseok Yoon, Yu Zheng 0004, Xing Xie 0001, Woontack Woo |
UIC | 4 |
| 2010 | Immersive modeling system (IMMS) for personal electronic products using a multi-modal interface
Yong-Gu Lee, Hyungjun Park, Woontack Woo, Jeha Ryu, Hong Kook Kim, Sung Wook Baik, Kwang Hee Ko, Han Kyun Choi, Sun-Uk Hwang, Duck Bong Kim, Hyun Soo Kim, Kwan H. Lee |
Comput. Aided Des. | 3 |
| 2010 | Daily Physiological Signal Monitoring System for Fostering Social Well-Being in Smart SpacesabstractWe propose socially acceptable physiological signal monitoring system consisting of a natural sensing interface and an intuitive information display. The natural sensing interface, BioPebble, is a simple, natural, and enjoyable stone-type interface embedded in daily objects such as remote controllers or mobile devices. The intuitive information display, Rainbow display, represents physiological signal information using an intuitive information visualization method for novice users. The information is transferred to the mobile personal station wirelessly and visualized using an appropriate mapping strategy considering the differences of individual users. From the experiment, we found that the sensing interface is more comfortable and involves more aesthetic feeling compared to disposable wired interfaces while retaining effectiveness in monitoring physiological signals regularly during a short period of time. In addition, visualized information in a rainbow-shaped display is easily understood by novice users but not by professionals and caregivers. From these observations, we expect that this monitoring system enables users to monitor their physiological status easily with the proposed interface and display. As a result, this system accelerates social well-being by helping users to regularly check their physiological status for precaution of abnormal condition. Ahyoung Choi, Woontack Woo |
Cybern. Syst. | 2 |
| 2010 | Toward Combining Automatic Resolution with Social Mediation for Resolving Multiuser ConflictsabstractIn spite of intensive effort to resolve conflicts between multiple users of context-aware applications in a smart space, there has been no practical solution for flexibly resolving them based on the situation of the users. In this paper, we propose a mixed resolution method to combine automatic resolution with social participation for resolving multiuser conflicts. For combining the two resolution approaches, various contexts such as preferences, priority, and types of applications are used to select an appropriate resolution method for the encountered conflict. Through an evaluation, we found that the performance of selection algorithms mainly depended on the number of users and the similarity between their preferences, and we derived an appropriate threshold for determining whether users were similar or not. With a user study of 3 applications in a smart-space test-bed, we observed that the combination of automatic resolution when preferences are similar and social mediation when preferences are different effectively resolved multiuser conflict even though social pressure played an important role, and the 3 different applications had different thresholds. Choonsung Shin, Anind K. Dey, Woontack Woo |
Cybern. Syst. | 3 |
| 2010 | Synthetic vision-based perceptual attention for augmented reality agentsabstractAbstract We describe our model for synthetic vision‐based perceptual attention for autonomous agents in augmented reality (AR) environments. Since virtual and physical objects coexist in their environment, such agents must adaptively perceive and attend to objects relevant to their goals. To enable agents to perceive their surroundings, our approach allows the agents to determine currently visible objects from the scene description of what virtual and physical objects are configured in the camera's viewing area. In our model, a degree of attention is assigned to each perceived object based on its similarity to target objects related to an agent's goals. The agent can thus focus on a reduced set of perceived objects with respect to the estimated degree of attention. Moreover, by continuously and smartly updating the perceptual memory, it eliminates the processing loads associated to previously observed objects. To demonstrate the effectiveness of our approach, we implemented an animated character that was overlaid over a miniature version of campus in real‐time and that attended to building blocks relevant to given tasks. Experiments showed that our model could reduce a character's perceptual load at any time, even when surroundings change. Copyright © 2010 John Wiley & Sons, Ltd. Sejin Oh, Woonhyuk Baek, Woontack Woo |
Comput. Animat. Virtual Worlds | 3 |
| 2010 | u-BabSang: a context-aware food recommendation system
Yoosoo Oh, Ahyoung Choi, Woontack Woo |
J. Supercomput. | 3 |
| 2010 | Scalable real-time planar targets tracking for digilog books
Kiyoung Kim, Vincent Lepetit, Woontack Woo |
Vis. Comput. | 3 |
| 2009 | ESM-Blur: Handling & rendering blur in 3D tracking and augmentationabstractThe contribution of this paper is two-fold. First, we show how to extend the ESM algorithm to handle motion blur in 3D object tracking. ESM is a powerful algorithm for template matching-based tracking, but it can fail under motion blur. We introduce an image formation model that explicitly considers the possibility of blur, and show it results in a generalization of the original ESM algorithm. This allows to converge faster, more accurately and more robustly even under large amount of blur. Our second contribution is an efficient method for rendering the virtual objects under the estimated motion blur. It renders two images of the object under 3D perspective, and warps them to create many intermediate images. By fusing these images we obtain a final image for the virtual objects blurred consistently with the captured image. Because warping is much faster that 3D rendering, we can create realistically blurred images at a very low computational cost. Vincent Lepetit, Woontack Woo |
ISMAR | 3 |
| 2009 | Marker-Less Tracking for Multi-layer Authoring in AR Books
Kiyoung Kim, Jonghee Park, Woontack Woo |
ICEC | 3 |
| 2009 | Multi-layer Based Authoring Tool for Digilog Book
Jonghee Park, Woontack Woo |
ICEC | 2 |
| 2008 | Mixed-initiative conflict resolution for context-aware applicationsabstractA number of technologies have contributed to automatically resolving resource conflicts between multiple users in a smart space. However, such systems eliminate the users' ability to perform this conflict resolution by themselves, which they actually prefer to do in certain circumstances. Since both resolution approaches have their merits, we propose a mixed-initiative conflict resolution system, which combines automatic conflict resolution with mediated, or user-driven, resolution by exploiting contextual information in context-aware applications. An evaluation of our system found that users prefer to use a mediated resolution approach when their preferences about outcome are very different from others', but have no preferred method when their preferences about outcome are similar to others'. Choonsung Shin, Anind K. Dey, Woontack Woo |
UbiComp | 3 |
| 2008 | Multiple 3D Object tracking for augmented realityabstractWe present a method that is able to track several 3D objects simultaneously, robustly, and accurately in real-time. While many applications need to consider more than one object in practice, the existing methods for single object tracking do not scale well with the number of objects, and a proper way to deal with several objects is required. Our method combines object detection and tracking: Frame-to-frame tracking is less computationally demanding but is prone to fail, while detection is more robust but slower. We show how to combine them to take the advantages of the two approaches, and demonstrate our method on several real sequences. Vincent Lepetit, Woontack Woo |
ISMAR | 3 |
| 2008 | User Identification with User's Stepping Pattern over the ubiFloorIIabstractIn this paper, we propose the UbiFloorII, a novel floor-based user identification system to recognize humans based on their stepping pattern, the arrays of the transitional footprints from heel-strike to toe-off. To obtain users' stepping pattern from their gait, we deployed photo interrupter sensors instead of switch sensors used in the UbiFloorI. We developed a software module to extract stepping pattern from users' gait. For user identification, we employed neural network trained with users' stepping samples. We achieved about 92% recognition accuracy using this floor-based approach. The UbiFloorII system may be used to automatically and transparently identify users in a home environment. Jaeseok Yun, Gregory D. Abowd, Jeha Ryu, Woontack Woo |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2007 | Identifying Foreground from Multiple Images
Wonwoo Lee, Woontack Woo, Edmond Boyer |
ACCV (2) | 2 |
| 2007 | Explanatory Style for Socially Interactive Agents
Sejin Oh, Jonathan Gratch, Woontack Woo |
ACII | 3 |
| 2007 | A Size-Based Qualitative Approach to the Representation of Spatial Granularity
Hedda R. Schmidtke, Woontack Woo |
IJCAI | 2 |
| 2006 | wear-UCAM: A Toolkit for Mobile User Interactions in Smart Environments
Dongpyo Hong, Youngjung Suh, Ahyoung Choi, Woontack Woo |
EUC | 5 |
| 2006 | Bare Hand Interface for Interaction in the Video See-Through HMD Based Wearable AR Environment
Taejin Ha, Woontack Woo |
ICEC | 2 |
| 2006 | A 3D Vision-Based Ambient User InterfaceabstractThis article proposes a 3-dimensional (3D) vision-based ambient user interface as an interaction metaphor that exploits a user's personal space and its dynamic gestures. In human-computer interaction, to provide natural interactions with a system, a user interface should not be a bulky or complicated device. In this regard, the proposed ambient user interface utilizes an invisible personal space to remove cumbersome devices where the invisible personal space is virtually augmented through exploiting 3D vision techniques. For natural interactions with the user's dynamic gestures, the user of interest is extracted from the image sequences by the proposed user segmentation method. This method can retrieve 3D information from the segmented user image through 3D vision techniques and a multiview camera. With the retrieved 3D information of the user, a set of 3D boxes (SpaceSensor) can be constructed and augmented around the user; then the user can interact with the system by touching the augmented SpaceSensor. In the user's dynamic gesture tracking, the computational complexity of SpaceSensor is relatively lower than that of conventional 2-dimensional vision-based gesture tracking techniques, because the touched positions of SpaceSensor are tracked. According to the experimental results, the proposed ambient user interface can be applied to various systems that require real-time user's dynamic gestures for their interactions both in real and virtual environments. Dongpyo Hong, Woontack Woo |
Int. J. Hum. Comput. Interact. | 2 |
| 2006 | 3-D Virtual Studio for Natural Inter-"Acting"abstractVirtual studios have long been used in commercial broadcasting. However, most virtual studios are based on "blue screen" technology, and its two-dimensional (2-D) nature restricts the user from making natural three-dimensional (3-D) interactions. Actors have to follow prewritten scripts and pretend as if directly interacting with the synthetic objects. This often creates an unnatural and seemingly uncoordinated output. In this paper, we introduce an improved virtual-studio framework to enable actors/users to interact in 3-D more naturally with the synthetic environment and objects. The proposed system uses a stereo camera to first construct a 3-D environment (for the actor to act in), a multiview camera to extract the image and 3-D information about the actor, and a real-time registration and rendering software for generating the final output. Synthetic 3-D objects can be easily inserted and rendered, in real time, together with the 3-D environment and video actor for natural 3-D interaction. The enabling of natural 3-D interaction would make more cinematic techniques possible including live and spontaneous acting. The proposed system is not limited to broadcast production, but can also be used for creating virtual/augmented-reality environments for training and entertainment Namgyu Kim, Woontack Woo, Gerard Jounghyun Kim, Chan-Mo Park |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2005 | Physiological Sensing and Feature Extraction for Emotion Recognition by Exploiting Acupuncture Spots
Ahyoung Choi, Woontack Woo |
ACII | 2 |
| 2005 | Collaborative billiARds: Towards the Ultimate Gaming Experience
Usman Sargaana, Hossein S. Farahani, Jong-Weon Lee 0002, Jeha Ryu, Woontack Woo |
ICEC | 5 |
| 2004 | Manipulating Multimedia Contents with Tangible Media Control System
Sejin Oh, Woontack Woo |
ICEC | 2 |
| 2003 | Image-based panoramic 3D virtual environment using rotating two multiview camerasabstractIn this paper, we propose a new method for generating an image-based 3D panoramic virtual environment (VE). The panoramic VE is generated using 3D depth information estimated from rotating two multiview cameras. Even though conventional 2D image-based mosaicking methods provide a wide view, they have limitations in providing a user with a navigation-enabled virtual environment. In order to resolve such obstacles, we first estimate the depth of the scene using two calibrated multiview cameras and then stitch 3D point clouds instead of images. By rotating two cameras using a turn-table it enables users to navigate the resulting 3D virtual environment with HMD. Sehwan Kim, Eun-Young Chang, Chung-Hyun Ahn, Woontack Woo |
ICIP (1) | 4 |
| 2002 | Three-dimensional movement tracking with asynchronous digital cameras for interactive systems
Se-Heon Kim, Woontack Woo |
VCIP | 2 |
| 2001 | Image retrieval using multi-scale color clusteringabstractA fundamental issue in content-based image retrieval is how to select image features that can represent image contents appropriately. A multi-scale color clustering algorithm based on human perceptual properties of color images is proposed for image retrieval. The multi-scale clustering algorithm is an unsupervised clustering method that utilizes the perceptual uniformity property in the (p,q) color space. The proposed color clustering algorithm produces a small set of representative color vectors for each image that capture color properties of the image, and a set of correlogram values that contain the spatial information of the image. Sehwan Kim, Woontack Woo, Yo-Sung Ho |
ICIP (1) | 2 |
| 2001 | Photorealistic interactive virtual environment generation using multiview cameras
Namgyu Kim, Woontack Woo, Makoto Tadenuma |
VCIP | 2 |
| 2001 | Sketch on dynamic gesture tracking and analysis exploiting vision-based 3D interface
Woontack Woo, Namgyu Kim, Karen Wong, Makoto Tadenuma |
VCIP | 1 |
| 2000 | MIDAS: MIC Interactive DAnce SystemabstractWe have been studying how to establish a method for extracting human emotion in order to express emotional images by utilizing multimedia, such as video and sound. Human body motion is the most basic essence in expressing human emotion. MIDAS (MIC Interactive DAnce System) is an application of the emotional information extracting/expressing method in a dance system. In this system, we applied R. Laban's (1879-1958) dance theory to extract physical characteristics of dance motion from real-time video sequences, and we mapped this information to categories of emotional image expression. In this way, we can relate the motion to an emotional representation using multimedia. Ryotaro Suzuki, Yuichi Iwadate, Masayuki Inoue, Woontack Woo |
SMC | 4 |
| 2000 | Stereo imaging using a camera with stereoscopic adapterabstractThe authors analyze the characteristics of the stereoscopic adapter, which is a cost-effective way to generate stereo video sequences with a camera. We also propose an efficient way to compensate for the inherent distortions. In general, stereo sequences can be captured using a pair of cameras but the resulting sequences tend to yield various well-known problems due to different characteristics of the pair of stereo cameras. Meanwhile, a camera with the stereoscopic adapter provides a natural way to capture and display stereoscopic video. It allows users to access all the functions built into the camera, e.g. zoom, auto-focus, auto-exposure, special effects, etc. The cost however is the reduced quality of the videos since the adapter allows capture of stereo video sequences in the field sequential format, i.e. left and right images in different scan lines, respectively. In addition, it generates size and color distortions due to the physical configuration of the mirror in the adapter. We analyze and compensate for such distortions to reduce possible errors in vision applications exploiting the stereo images. According to our preliminary study, the adapter with the proposed compensation scheme will pave the way for various low cost image based virtual reality applications at hand. Woontack Woo, Namgyu Kim, Yuichi Iwadate |
SMC | 1 |
| 2000 | Overlapped block disparity compensation with adaptive windows for stereo image codingabstractWe propose a modified overlapped block-matching (OBM) scheme for stereo image coding. OBM has been used in video coding but, to the best of our knowledge, it has not been applied to stereo image coding to date. In video coding, OBM has proven useful in reducing blocking artifacts (since multiple vectors can be used for each block), while also maintaining most of the advantages of fixed-size block matching. There are two main novelties in this work. First, we show that OBM techniques can be successfully applied to stereo image coding. Second, we take advantage of the smoothness properties typically found in disparity fields to further improve the performance of OBM in this particular application. Specifically, we note that practical OBM approaches use noniterative estimation techniques, which produce lower quality estimates than iterative methods. By introducing smoothness constraints into the noniterative DV computation, we improve the quality of the estimated disparity as compared to standard noniterative OBM approaches. In addition, we propose a disparity estimation/compensation approach using adaptive windows with variable shapes, which results in a reduction in complexity. We provide experimental results that show that our proposed hybrid OBM scheme achieves a PSNR gain (about 1.5-2 dB) as compared to a simple block-based scheme, with some slight PSNR gains (about 0.2-0.5 dB) in a reduced complexity, as compared to an approach based on standard OBM with half-pixel accuracy. Woontack Woo, Antonio Ortega |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1999 | Stereo Image Coding Using Hierarchical Mrf Model and Selective Overlapped Block Disparity CompensationabstractIn this paper, we propose a novel hierarchical disparity estimation/compensation (DE/DC) algorithm for stereo image coding. One way to limit the well-known drawbacks of block matching (e.g., inaccurate disparity and blocking artifacts in the decoded image) is to resort to variable size block matching (VSBM). However, VSMB may result in an inconsistent disparity estimation (i.e., such that the estimated disparity field does not correspond to the true disparity) especially as the subblock becomes small, thus leading to high frequency energy at block boundaries in the residue image. To address these problems, in this paper we propose a hybrid quadtree-based DE/DC scheme, where a Markov Random Field (MRF) model is used in combination with VSBM and selective overlapped disparity compensation to improve the disparity field consistency and lead to higher coding efficiency. Our experimental results demonstrate that the proposed block segmentation scheme achieves a higher PSNR and a more consistent disparity field, as compared to conventional VSBM schemes. Woontack Woo, Antonio Ortega, Yuichi Iwadate |
ICIP (2) | 1 |
| 1999 | Optimal blockwise dependent quantization for stereo image codingabstractResearch in coding of stereo images has focused mostly on the issue of disparity estimation to exploit the redundancy between the two images in a stereo pair, with less attention being devoted to the equally important problem of allocating bits between the two images. This bit allocation problem is complicated by the dependencies arising from using a prediction based on the quantized reference images. We address the problem of blockwise bit allocation for coding of stereo images and show how, given the special characteristics of the disparity field, one can achieve an optimal solution with reasonable complexity, whereas in similar problems in motion compensated video only approximate solutions are feasible. We present algorithms based on dynamic programming that provide the optimal blockwise bit allocation. Our experiments based on a modified JPEG coder show that the proposed scheme achieves higher mean peak signal-to-noise ratio over the two frames (0.2-0.5 dB improvements) as compared with blockwise independent quantization. We also propose a fast algorithm that provides most of the gain at a fraction of the complexity. Woontack Woo, Antonio Ortega |
IEEE Trans. Circuits Syst. Video Technol. | 1 |