Taehyun Rhee

dblp:07/6568 · also Taehyun James Rhee · DBLP profile ↗
← Back
33ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-6150-0637ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 6 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 13 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
YearPublicationVenuePosition
2026 SE360: Semantic Edit in 360° Panoramas via Hierarchical Data Construction
abstract
While instruction-based image editing is emerging, extending it to 360° panorama introduces additional challenges. Existing methods often produce implausible results in both equirectangular projections (ERP) and perspective views. To address these limitations, we propose SE360, a novel framework for multi-condition guided object editing in 360° panoramas. At its core is a novel coarse-to-fine autonomous data generation pipeline without manual intervention. This pipeline leverages a Vision-Language Model (VLM) and adaptive projection adjustment for hierarchical analysis, ensuring the holistic segmentation of objects and their physical context. The resulting data pairs are both semantically meaningful and geometrically consistent, even when sourced from unlabeled panoramas. Furthermore, we introduce a cost-effective, two-stage data refinement strategy to improve data realism and mitigate model overfitting to erasing artifacts. Based on the constructed dataset, we train a Transformer-based diffusion model to allow flexible object editing guided by text, mask, or reference image in 360° panoramas. Our experiments demonstrate that our method outperforms existing methods in both visual quality and semantic accuracy.
Haoyi Zhong, Andrew Chalmers, Taehyun Rhee
AAAI4
2026 Affective and Cognitive Feedback from a Robot for Human-attributed Failure Handling
abstract
Human–robot collaboration increasingly frames robots as teammates rather than tools, yet there is limited guidance on how robots should respond when failures are attributed to the human collaborator. We investigate how robot collaborators should respond to support collaboration experience after a human-attributed failure. In a 4 × 2 mixed factorial design (N = 60), participants completed a collaborative block-stacking task with either a humanoid robot (NAO) or a human collaborator under four scenarios: success, affective feedback, cognitive feedback, and no feedback. We measured collaboration experience in terms of teamwork quality, perceived copresence, and intimacy. Both affective and cognitive feedback improved these outcomes compared with no feedback: affective cues yielded the strongest socio-relational gains (copresence, intimacy), whereas cognitive cues more strongly enhanced perceived teamwork quality. These patterns were consistent across human–robot and human–human collaboration, indicating shared team-level expectations that extend beyond the individual actor. The results provide empirical evidence for socially adaptive robots that pair brief emotional reassurance with concrete guidance to support collaboration after human-attributed failures.
Myeongul Jung, Taehyun Rhee, Dooyong Kim, Kwanguk (Kenny) Kim
CHI4
2025 TouchWalker: Real-Time Avatar Locomotion from Touchscreen Finger Walking
abstract
We present TouchWalker, a real-time system for controlling fullbody avatar locomotion using finger-walking gestures on a touchscreen. The system comprises two main components: TouchWalker-MotionNet, a neural motion generator that synthesizes full-body avatar motion on a per-frame basis from temporally sparse twofinger input, and TouchWalker-UI, a compact touch interface that interprets user touch input to avatar-relative foot positions. Unlike prior systems that rely on symbolic gesture triggers or predefined motion sequences, TouchWalker uses its neural component to generate continuous, context-aware full-body motion on a per-frame basis-including airborne phases such as running, even without input during mid-air steps-enabling more expressive and immediate interaction. To ensure accurate alignment between finger contacts and avatar motion, it employs a MoE-GRU architecture with a dedicated foot-alignment loss. We evaluate TouchWalker in a user study comparing it to a virtual joystick baseline with predefined motion across diverse locomotion tasks. Results show that TouchWalker improves users' sense of embodiment, enjoyment, and immersion.
Geuntae Park, Jiwon Yi, Taehyun Rhee, Kwanguk (Kenny) Kim, Yoonsang Lee 0001
ISMAR3
2025 Parameter-Free Neural Lens Blur Rendering for High-Fidelity Composites
abstract
Consistent and natural camera lens blur is important for seamlessly blending 3D virtual objects into photographed real-scenes. Since lens blur typically varies with scene depth, the placement of virtual objects and their corresponding blur levels significantly affect the visual fidelity of mixed reality compositions. Existing pipelines often rely on camera parameters (e.g., focal length, focus distance, aperture size) and scene depth to compute the circle of confusion (CoC) for realistic lens blur rendering. However, such information is often unavailable to ordinary users, limiting the accessibility and generalizability of these methods. In this work, we propose a novel compositing approach that directly estimates the CoC map from RGB images, bypassing the need for scene depth or camera metadata. The CoC values for virtual objects are inferred through a linear relationship between its signed CoC map and depth, and realistic lens blur is rendered using a neural reblurring network. Our method provides flexible and practical solution for real-world applications. Experimental results demonstrate that our method achieves high-fidelity compositing with realistic defocus effects, outperforming state-of-the-art techniques in both qualitative and quantitative evaluations.
Lingyan Ruan, Bin Chen 0019, Taehyun Rhee
ISMAR3
2025 Interaction With Virtual Objects Using Human Pose and Shape Estimation
abstract
ABSTRACT In this article, we propose an AR system that facilitates a user's natural interaction with virtual objects in an augmented reality environment. The system consists of three modules: human pose and shape estimation, camera‐space calibration, and physics simulation. The first module estimates a user's 3D pose and shape from a single RGB video stream, thereby reducing the system setup cost and broadening potential applications. The camera‐space calibration module estimates the user's camera‐space position to align the user with the input RGB image. The physics simulation enables seamless and physically natural interaction with virtual objects. Two prototyping applications built upon the system prove an enhancement in the quality of interaction, fostering a more immersive and intuitive user experience.
Hong Son Nguyen, DaEun Cheong, Andrew Chalmers, MyoungGon Kim, Taehyun Rhee
Comput. Animat. Virtual Worlds5
2024 Full-Body Human De-lighting with Semi-supervised Learning
Joshua Weir, Junhong Zhao, Andrew Chalmers, Taehyun Rhee
ACCV (1)4
2024 Neural Radiance Fields for Dynamic View Synthesis Using Local Temporal Priors
Rongsen Chen, Junhong Zhao, Andrew Chalmers, Taehyun Rhee
CVM (1)5
2024 Avatar360: Emulating 6-DoF Perception in 360°Panoramas through Avatar-Assisted Navigation
abstract
360° images offer panoramic views of captured environments, placing users within an egocentric perspective. While users can freely rotate their viewpoint, they don’t experience 6-DoF navigation with translational movement. In this research, we introduce Avatar360, a novel method to elicit 6-DoF perception in 360° panoramas, using avatar-assisted navigation combined with an exocentric view of the 360° panorama. We seamlessly integrate a 3D avatar into 360° panoramas, allowing users to navigate a 3D virtual landscape congruent with the 360° background. By aligning the exocentric perspective of the 360° panorama with the avatar’s movements, we replicate a sensation of 6-DoF navigation in 360° panoramas. We explore mechanisms for simultaneous avatar and viewpoint controls, as well as procedures for transitions between spatially connected 360° panoramas. A user study was conducted to assess the perception of 6-DoF navigation in 360° panoramas via a 3D avatar, evaluating users’ sense of movement, disorientation, and presence. We also gained insight into perspective view controls and transition techniques between panoramas. Statistical analysis shows avatar-assisted navigation elicits a user’s sense of movement within 360° panoramas. Our results also provide guidelines for effective view control and transition strategies in avatar-assisted 360° navigation.
Andrew Chalmers, Faisal Zaman, Taehyun Rhee
VR3
2023 MRMAC: Mixed Reality Multi-user Asymmetric Collaboration
abstract
We present MRMAC, a Mixed Reality Multi-user Asymmetric Collaboration system that allows remote users to teleport virtually into a real-world collaboration space to communicate and collaborate with local users. Our system enables telepresence for remote users by live-streaming the physical environment of local users using a 360° camera while blending 3D virtual assets into the mixed-reality collaboration space. Our novel client-server architecture enables asymmetric collaboration for multiple AR and VR users and incorporates avatars, view controls, as well as synchronized low-latency audio, video, and asset streaming. We evaluated our implementation with two baseline conditions: conventional 2D and standard 360° videoconferencing. Results show that MRMAC outperformed both baselines in inducing a sense of presence, improving task performance, usability, and overall user preference, demonstrating its potential for immersive multi-user telecollaboration.
Faisal Zaman, Craig Anslow, Andrew Chalmers, Taehyun Rhee
ISMAR4
2023 Deep Learning-based Simulator Sickness Estimation from 3D Motion
abstract
This paper presents a novel solution for estimating simulator sickness in HMDs using machine learning and 3D motion data, informed by user-labeled simulator sickness data and user analysis. We conducted a novel VR user study, which decomposed motion data and used an instant dial-based sickness scoring mechanism. We were able to emulate typical VR usage and collect user simulator sickness scores. Our user analysis shows that translation and rotation differently impact user simulator sickness in HMDs. In addition, users’ demographic information and self-assessed simulator sickness susceptibility data are collected and show some indication of potential simulator sickness. Guided by the findings from the user study, we developed a novel deep learning-based solution to better estimate simulator sickness with decomposed 3D motion features and user profile information. The model was trained and tested using the 3D motion dataset with user-labeled simulator sickness and profiles collected from the user study. The results show higher estimation accuracy when using the 3D motion data compared with methods based on optical flow extracted from the recorded video, as well as improved accuracy when decomposing the motion data and incorporating user profile information.
Junhong Zhao, Kien T. P. Tran, Andrew Chalmers, Weng Khuan Hoh, Richard Yao, Arindam Dey 0001, James Wilmott, Mark Billinghurst, Robert W. Lindeman, Taehyun Rhee
ISMAR11
2023 Vicarious: Context-aware Viewpoints Selection for Mixed Reality Collaboration
abstract
Mixed-perspective, combining egocentric (first-person) and exocentric (third-person) viewpoints, have been shown to improve the collaborative experience in remote settings. Such experiences allow remote users to switch between different viewpoints to gain alternative perspectives of the remote space. However, existing systems lack seamless selection and transition between multiple perspectives that better fit the task at hand. To address this, we present a new approach called Vicarious, which simplifies and automates the selection between egocentric and exocentric viewpoints. Vicarious employs a context-aware method for dynamically switching or highlighting the optimal viewpoint based on user actions and the current context. To evaluate the effectiveness of the viewpoint selection method, we conducted a user study (n = 27) using an asymmetric AR-VR setup where users performed remote collaboration tasks under four distinct conditions: No-view, Manual, Guided, and Automatic selection. The results showed that Guided and Automatic viewpoint selection improved users’ understanding of the task space and task performance, and reduced cognitive load compared to Manual or No-view selection. The results also suggest that the asymmetric setup had minimal impact on spatial and social presence, except for differences in task load and preference. Based on these findings, we provide design implications for future research in mixed reality collaboration.
Faisal Zaman, Craig Anslow, Taehyun Rhee
VRST3
2023 Casual 6-DoF: Free-Viewpoint Panorama Using a Handheld 360° Camera
abstract
Six degrees-of-freedom (6-DoF) video provides telepresence by enabling users to move around in the captured scene with a wide field of regard. Compared to methods requiring sophisticated camera setups, the image-based rendering method based on photogrammetry can work with images captured with any poses, which is more suitable for casual users. However, existing image-based rendering methods are based on perspective images. When used to reconstruct 6-DoF views, it often requires capturing hundreds of images, making data capture a tedious and time-consuming process. In contrast to traditional perspective images, 360° images capture the entire surrounding view in a single shot, thus, providing a faster capturing process for 6-DoF view reconstruction. This article presents a novel method to provide 6-DoF experiences over a wide area using an unstructured collection of 360° panoramas captured by a conventional 360° camera. Our method consists of 360° data capturing, novel depth estimation to produce a high-quality spherical depth panorama, and high-fidelity free-viewpoint generation. We compared our method against state-of-the-art methods, using data captured in various environments. Our method shows better visual quality and robustness in the tested scenes.
Rongsen Chen, Simon Finnie, Andrew Chalmers, Taehyun Rhee
IEEE Trans. Vis. Comput. Graph.5
2022 Deep Portrait Delighting
Joshua Weir, Junhong Zhao, Andrew Chalmers, Taehyun Rhee
ECCV (16)4
2022 TeleFest: Augmented Virtual Teleportation for Live Concerts
abstract
We present TeleFest, a novel system for live-streaming mixed reality 360° videos to online streaming services. TeleFest allows a producer to control multiple cameras in real time, providing viewers with different locations for experiencing the concert, and an intermediate software stack allows virtual content to be overlaid with coherent illumination that matches the real-world footage. TeleFest was evaluated by livestreaming a concert to almost 2,000 online viewers, allowing them to watch the performance from the crowd, the stage, or via a catered experience controlled by a producer in real time that included camera switching and augmented content. The results of an online survey completed by virtual and physical attendees of the festival are presented, showing positive feedback for our setup and suggesting that the addition of virtual and immersive content to live events could lead to a more enjoyable experience for viewers.
Jacob Young, Stephen Thompson 0001, Holly Downer, Benjamin Allen, Nadia Pantidi, Lukas Stoecklein, Taehyun Rhee
IMX7
2022 Illumination Browser: An intuitive representation for radiance map databases
Andrew Chalmers, Todd E. Zickler, Taehyun Rhee
Comput. Graph.3
2021 Spectator View: Enabling Asymmetric Interaction between HMD Wearers and Spectators with a Large Display
abstract
In this paper, we present a system that allows a user with a head-mounted display (HMD) to communicate and collaborate with spectators outside of the headset. We evaluate its impact on task performance, immersion, and collaborative interaction. Our solution targets scenarios like live presentations or multi-user collaborative systems, where it is not convenient to develop a VR multiplayer experience and supply each user (and spectator) with an HMD. The spectator views the virtual world on a large-scale tiled video wall and is given the ability to control the orientation of their own virtual camera. This allows spectators to stay focused on the immersed user's point of view or freely look around the environment. To improve collaboration between users, we implemented a pointing system where a spectator can point at objects on the screen, which maps an indicator directly onto the objects in the virtual world. We conducted a user study to investigate the influence of rotational camera decoupling and pointing gestures in the context of HMD-immersed and non-immersed users utilizing a large-scale display. Our results indicate that camera decoupling and pointing positively impacts collaboration. A decoupled view is preferable in situations where both users need to indicate objects of interest in the scene, such as presentations and joint-task scenarios, as it requires a shared reference space. A coupled view, on the other hand, is preferable in synchronous interactions such as remote-assistant scenarios.
Finn Welsford-Ackroyd, Andrew Chalmers, Rafael Kuffner dos Anjos, Daniel Medeiros 0001, Taehyun Rhee
Proc. ACM Hum. Comput. Interact.6
2021 Shading Rig: Dynamic Art-directable Stylised Shading for 3D Characters
abstract
Despite the popularity of three-dimensional (3D) animation techniques, the style of 2D cel animation is seeing increased use in games and interactive applications. However, conventional 3D toon shading frequently requires manual editing to clean up undesired shadows or add stylistic details based on art direction. This editing is impractical for the frame-by-frame editing in cartoon feature film post-production. For interactive stylised media and games, post-production is unavailable due to real-time constraints, so art-direction must be preserved automatically. For these reasons, artists often resort to mesh and texture edits to mitigate undesired shadows typical of toon shaders. Such edits allow real-time rendering but are limited in resolution, animation quality and lack detail control for stylised shadow design. In our framework, artists build a “shading rig,” a collection of these edits, that allows artists to animate toon shading. Artists pre-animate the shading rig under changing lighting, to dynamically preserve artistic intent in a live application, without manual intervention. We show our method preserves continuous motion and shape interpolation, with fewer keyframes than previous work. Our shading shape interpolation is computationally cheaper than state-of-the-art image interpolation techniques. We achieve these improvements while preserving vector quality rendering, without resorting either to high texture resolution or mesh density.
Lohit Petikam, Ken Anjyo, Taehyun Rhee
ACM Trans. Graph.3
2021 Reconstructing Reflection Maps Using a Stacked-CNN for Mixed Reality Rendering
abstract
Corresponding lighting and reflectance between real and virtual objects is important for spatial presence in augmented and mixed reality (AR and MR) applications. We present a method to reconstruct real-world environmental lighting, encoded as a reflection map (RM), from a conventional photograph. To achieve this, we propose a stacked convolutional neural network (SCNN) that predicts high dynamic range (HDR) 360° RMs with varying roughness from a limited field of view, low dynamic range photograph. The SCNN is progressively trained from high to low roughness to predict RMs at varying roughness levels, where each roughness level corresponds to a virtual object's roughness (from diffuse to glossy) for rendering. The predicted RM provides high-fidelity rendering of virtual objects to match with the background photograph. We illustrate the use of our method with indoor and outdoor scenes trained on separate indoor/outdoor SCNNs showing plausible rendering and composition of virtual objects in AR/MR. We show that our method has improved quality over previous methods with a comparative user study and error metrics.
Andrew Chalmers, Junhong Zhao, Daniel Medeiros 0001, Taehyun Rhee
IEEE Trans. Vis. Comput. Graph.4
2021 Adaptive Light Estimation using Dynamic Filtering for Diverse Lighting Conditions
abstract
High dynamic range (HDR) panoramic environment maps are widely used to illuminate virtual objects to blend with real-world scenes. However, in common applications for augmented and mixed-reality (AR/MR), capturing 360° surroundings to obtain an HDR environment map is often not possible using consumer-level devices. We present a novel light estimation method to predict 360° HDR environment maps from a single photograph with a limited field-of-view (FOV). We introduce the Dynamic Lighting network (DLNet), a convolutional neural network that dynamically generates the convolution filters based on the input photograph sample to adaptively learn the lighting cues within each photograph. We propose novel Spherical Multi-Scale Dynamic (SMD) convolutional modules to dynamically generate sample-specific kernels for decoding features in the spherical domain to predict 360° environment maps. Using DLNet and data augmentations with respect to FOV, an exposure multiplier, and color temperature, our model shows the capability of estimating lighting under diverse input variations. Compared with prior work that fixes the network filters once trained, our method maintains lighting consistency across different exposure multipliers and color temperature, and maintains robust light estimation accuracy as FOV increases. The surrounding lighting information estimated by our method ensures coherent illumination of 3D objects blended with the input photograph, enabling high fidelity augmented and mixed reality supporting a wide range of environmental lighting conditions and device sensors.
Junhong Zhao, Andrew Chalmers, Taehyun Rhee
IEEE Trans. Vis. Comput. Graph.3
2020 Augmented Virtual Teleportation for High-Fidelity Telecollaboration
abstract
Telecollaboration involves the teleportation of a remote collaborator to another real-world environment where their partner is located. The fidelity of the environment plays an important role for allowing corresponding spatial references in remote collaboration. We present a novel asymmetric platform, Augmented Virtual Teleportation (AVT), which provides high-fidelity telepresence of a remote VR user (VR-Traveler) into a real-world collaboration space to interact with a local AR user (AR-Host). AVT uses a 360° video camera (360-camera) that captures and live-streams the omni-directional scenes over a network. The remote VR-Traveler watching the video in a VR headset experiences live presence and co-presence in the real-world collaboration space. The VR-Traveler's movements are captured and transmitted to a 3D avatar overlaid onto the 360-camera which can be seen in the AR-Host's display. The visual and audio cues for each collaborator are synchronized in the Mixed Reality Collaboration space (MRC-space), where they can interactively edit virtual objects and collaborate in the real environment using the real objects as a reference. High fidelity, real-time rendering of virtual objects and seamless blending into the real scene allows for unique mixed reality use-case scenarios. Our working prototype has been tested with a user study to evaluate spatial presence, co-presence, and user satisfaction during telecollaboration. Possible applications of AVT are identified and proposed to guide future usage.
Taehyun Rhee, Stephen Thompson 0001, Daniel Medeiros 0001, Rafael Kuffner dos Anjos, Andrew Chalmers
IEEE Trans. Vis. Comput. Graph.1
2019 Real-Time Mixed Reality Rendering for Underwater 360° Videos
abstract
We present a novel mixed reality (MR) rendering and composition solution that illuminates and blends virtual objects into underwater 360° videos (360-video) in real-time. Real-time underwater lighting (caustics, god rays, fog, and particulates) were developed to improve the overall lighting and blending quality. We also provide a MR toolkit, an interface to tune the parameters of the underwater lighting so the user can match the lighting observed in the 360-video. Our image based lighting provides automatic ambient and high frequency underwater lighting. This ensures that the virtual objects are lit and blend similarly to each frame of the video semi-automatically and in real-time. We conducted a user study by having participants rate our method based on the visual quality and presence using a five point Likert Scale. The results show that our underwater lighting is preferred over no underwater effects or using naive ambient lighting. We also have a few takeaways on what elements of our underwater lighting and interaction have a significant impact on visual quality and presence in underwater MR.
Stephen Thompson 0001, Andrew Chalmers, Taehyun Rhee
ISMAR3
2019 Real-time Underwater Caustics for Mixed Reality 360° Videos
abstract
We present a novel mixed reality (MR) rendering solution that illuminates and blends virtual objects into underwater 360° video with real-time underwater caustic effects. Image-based lighting is used in conjunction with underwater caustics to provide automatic ambient and high frequency underwater lighting. This ensures that the caustics and virtual objects are lit and blend into each frame of the video semi-automatically and in real-time. We provide an interactive interface with intuitive parameter controls to fine tune caustics to match with the background video.
Stephen Thompson 0001, Andrew Chalmers, Taehyun Rhee
VR3
2018 Visual Perception of Real World Depth Map Resolution for Mixed Reality Rendering
abstract
Compositing virtual objects into photographs with known real world geometry is a common task in mixed reality (MR) applications. This geometry enables rendering of global illumination effects, such as mutual lighting, shadowing, and occlusions between the background photograph and virtual objects. Obtaining high fidelity geometric representations of the real world can be a costly procedure, and is often approximated with depth data. However, it is not clear how much fidelity the depth data should have in order to maintain high visual quality in MR rendering. in this paper, we investigate the relationship between real world depth fidelity and visual quality in MR rendering. We do this by conducting a series of user experiments that measure how seamlessly virtual objects are blended with the background under varying depth resolutions. We independently evaluate the noticeability of multiple composition artifacts that occur with approximate depth. Perceptual thresholds in depth resolution are then obtained for each artifact. The findings can be used to inform trade-off decisions for optimising depth acquisition pipelines in MR applications.
Lohit Petikam, Andrew Chalmers, Taehyun Rhee
VR3
2018 Low-Rank Matrix Completion to Reconstruct Incomplete Rendering Images
abstract
Path tracing provides photo-realistic rendering in many applications but intermediate previsualization often suffers from distracting noise. Since the fundamental underlying problem is insufficient samples, we exploit the coherence of the visual signal to reconstruct missing samples, using a low-rank matrix completion framework. We present novel methods to construct low rank matrices for incomplete images including missing pixel, missing sub-pixel, and multi-frame scenarios. A convolutional neural network provides fast pre-completion for initialising missing values, and subsequent weighted nuclear norm minimisation (WNNM) with a parameter adjustment strategy (PAWNNM) efficiently recovers missing values even in high frequency details. The result shows better visual quality than recent methods including compressed sensing based reconstruction.
John P. Lewis, Taehyun Rhee
IEEE Trans. Vis. Comput. Graph.3
2017 MR360: Mixed Reality Rendering for 360° Panoramic Videos
abstract
This paper presents a novel immersive system called MR360 that provides interactive mixed reality (MR) experiences using a conventional low dynamic range (LDR) 360° panoramic video (360-video) shown in head mounted displays (HMDs). MR360 seamlessly composites 3D virtual objects into a live 360-video using the input panoramic video as the lighting source to illuminate the virtual objects. Image based lighting (IBL) is perceptually optimized to provide fast and believable results using the LDR 360-video as the lighting source. Regions of most salient lights in the input panoramic video are detected to optimize the number of lights used to cast perceptible shadows. Then, the areas of the detected lights adjust the penumbra of the shadow to provide realistic soft shadows. Finally, our real-time differential rendering synthesizes illumination of the virtual 3D objects into the 360-video. MR360 provides the illusion of interacting with objects in a video, which are actually 3D virtual objects seamlessly composited into the background of the 360-video. MR360 was implemented in a commercial game engine and tested using various 360-videos. Since our MR360 pipeline does not require any pre-computation, it can synthesize an interactive MR scene using a live 360-video stream while providing realistic high performance rendering suitable for HMDs.
Taehyun Rhee, Lohit Petikam, Benjamin Allen, Andrew Chalmers
IEEE Trans. Vis. Comput. Graph.1
2012 Split and merge approach for detecting multiple planes in a depth image
abstract
We propose a novel method for detecting multiple planar structures in a scene from a depth image, and estimating their parametric models in real time. To realize this goal we split an entire depth image into small patches. Initially, we assume that all patches are on the planar structure and estimate their parametric models using least square fitting. We select patches where our assumption is valid and the selected planar patches are iteratively merged and refined in case that they are in the same planar structure. The qualitative and quantitative experiments shows that the proposed method has clear benefits in terms of accuracy and processing time.
Seon-Min Rhee, Yong-Beom Lee, James Dokyoon Kim, Taehyun Rhee
ICIP4
2012 Time-of-flight sensor and color camera calibration for multi-view acquisition
Hyunjung Shim, Rolf Adelsberger, James Dokyoon Kim, Seon-Min Rhee, Taehyun Rhee, Jae-Young Sim, Markus Gross 0001, Chang-Yeong Kim
Vis. Comput.5
2011 Realtime human motion control with a small number of inertial sensors
abstract
This paper introduces an approach to performance animation that employs a small number of motion sensors to create an easy-to-use system for an interactive control of a full-body human character. Our key idea is to construct a series of online local dynamic models from a prerecorded motion database and utilize them to construct full-body human motion in a maximum a posteriori framework (MAP). We have demonstrated the effectiveness of our system by controlling a variety of human actions, such as boxing, golf swinging, and table tennis, in real time. Given an appropriate motion capture database, the results are comparable in quality to those obtained from a commercial motion capture system with a full set of motion sensors (e.g., XSens [2009]); however, our performance animation system is far less intrusive and expensive because it requires a small of motion sensors for full body control. We have also evaluated the performance of our system by leave-one-out-experiments and by comparing with two baseline algorithms.
Huajun Liu, Xiaolin K. Wei, Jinxiang Chai, Inwoo Ha, Taehyun Rhee
SI3D5
2011 Scan-Based Volume Animation Driven by Locally Adaptive Articulated Registrations
abstract
This paper describes a complete system to create anatomically accurate example-based volume deformation and animation of articulated body regions, starting from multiple in vivo volume scans of a specific individual. In order to solve the correspondence problem across volume scans, a template volume is registered to each sample. The wide range of pose variations is first approximated by volume blend deformation (VBD), providing proper initialization of the articulated subject in different poses. A novel registration method is presented to efficiently reduce the computation cost while avoiding strong local minima inherent in complex articulated body volume registration. The algorithm highly constrains the degrees of freedom and search space involved in the nonlinear optimization, using hierarchical volume structures and locally constrained deformation based on the biharmonic clamped spline. Our registration step establishes a correspondence across scans, allowing a data-driven deformation approach in the volume domain. The results provide an occlusion-free person-specific 3D human body model, asymptotically accurate inner tissue deformations, and realistic volume animation of articulated movements driven by standard joint control estimated from the actual skeleton. Our approach also addresses the practical issues arising in using scans from living subjects. The robustness of our algorithms is tested by their applications on the hand, probably the most complex articulated region in the body, and the knee, a frequent subject area for medical imaging due to injuries.
Taehyun Rhee, John P. Lewis, Ulrich Neumann, Krishna S. Nayak
IEEE Trans. Vis. Comput. Graph.1
2010 A probabilistic approach to realistic face synthesis
abstract
This paper presents a novel approach to face modeling for realistic synthesis, powered by a probabilistic face diffuse model and a generic face specular map. We first construct a probabilistic face diffuse model for estimating the albedo and the normals of a face from an unknown input image. Then, we introduce a generic face specular map for estimating the specularity of the face. Using the estimated albedo, normal and specular information, we can synthesize the face under arbitrary lighting and viewing directions realistically. Unlike many existing face modeling techniques, our approach can retain both the diffuse and specular properties of the face without involving an elaborating 3D matching procedure. Thanks to the compact representation and the effective inference scheme, our technique can be applied to many practical applications, such as face normalization, avatar creation and de-identification.
Hyunjung Shim, Inwoo Ha, Taehyun Rhee, James Dokyoon Kim, Chang-Yeong Kim
ICIP3
2007 Soft-Tissue Deformation for In Vivo Volume Animation
abstract
Articulated body animation with smooth skin deformation is an important topic in computer graphics. This paper presents a pipeline that extends articulated body deformation to the volume graphics domain. The pipeline consists of in-vivo volume scans, kinematic joint estimation, volumetric joint weight computation, soft-tissue volume deformation, and direct volume rendering. The result is a fully articulated body volume driven by intuitive joint control that respects rigid deformation of the bone structures and produces smooth deformations of both the skin surface and the interior soft tissue regions.
Taehyun Rhee, John P. Lewis, Ulrich Neumann, Krishna S. Nayak
PG1
2006 Human hand modeling from surface anatomy
abstract
The human hand is an important interface with complex shape and movement. In virtual reality and gaming applications the use of an individualized rather than generic hand representation can increase the sense of immersion and in some cases may lead to more effortless and accurate interaction with the virtual world. We present a method for constructing a person-specific model from a single canonically posed palm image of the hand without human guidance. Tensor voting is employed to extract the principal creases on the palmar surface. Joint locations are estimated using extracted features and analysis of surface anatomy. The skin geometry of a generic 3D hand model is deformed using radial basis functions guided by correspondences to the extracted surface anatomy and hand contours. The result is a 3D model of an individual's hand, with similar joint locations, contours, and skin texture.
Taehyun Rhee, Ulrich Neumann, John P. Lewis
SI3D1
2006 Real-Time Weighted Pose-Space Deformation on the GPU
abstract
Abstract WPSD (Weighted Pose Space Deformation) is an example based skinning method for articulated body animation. The per‐vertex computation required in WPSD can be parallelized in a SIMD (Single Instruction Multiple Data) manner and implemented on a GPU. While such vertex‐parallel computation is often done on the GPU vertex processors, further parallelism can potentially be obtained by using the fragment processors. In this paper, we develop a parallel deformation method using the GPU fragment processors. Joint weights for each vertex are automatically calculated from sample poses, thereby reducing manual effort and enhancing the quality of WPSD as well as SSD (Skeletal Subspace Deformation). We show sufficient speed‐up of SSD, PSD (Pose Space Deformation) and WPSD to make them suitable for real‐time applications. Categories and Subject Descriptors (according to ACM CCS): I.3.1 [Computer Graphics]: Hardware Architecture‐Parallel processing, I.3.5 [Computer Graphics]: Computational Geometry and Object Modeling‐Curve, surface, solid and object modeling, I.3.7 [Computer Graphics]: Three‐Dimensional Graphics and Realism‐Animation.
Taehyun Rhee, John P. Lewis, Ulrich Neumann
Comput. Graph. Forum1