VLDB 2026 Research / reviewers in the wild / expert
Takafumi Taketomi
dblp:95/1736
· DBLP profile ↗
65ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-5353-0895ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 57 · 6 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 23 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | KaoLRM: Repurposing Pre-Trained Large Reconstruction Models for Parametric 3D Face ReconstructionabstractWe propose KaoLRM to re-target the learned prior of the Large Reconstruction Model (LRM) for parametric 3D face reconstruction from single-view images. Parametric 3D Morphable Models (3DMMs) have been widely used for facial reconstruction due to their compact and interpretable parameterization, yet existing 3DMM regressors often exhibit poor consistency across varying viewpoints. To address this, we harness the pre-trained 3D prior of LRM and incorporate FLAME-based 2D Gaussian Splatting into LRM's rendering pipeline. Specifically, KaoLRM projects LRM's pre-trained triplane features into the FLAME parameter space to recover geometry, and models appearance via 2D Gaussian primitives that are tightly coupled to the FLAME mesh. The rich prior enables the FLAME regressor to be aware of the 3D structure, leading to accu-rate and robust reconstructions under self-occlusions and diverse viewpoints. Experiments on both controlled and in-the-wild benchmarks demonstrate that KaoLRM achieves superior reconstruction accuracy and cross-view consistency, while existing methods remain sensitive to viewpoint variations. The code is released at https://github.com/CyberAgentAILab/KaoLRM. Qingtian Zhu, Zhixiang Wang 0001, Yinqiang Zheng, Takafumi Taketomi |
3DV | 5 |
| 2025 | ShadowSG: Spherical Gaussian Illumination from ShadowsabstractThis work leverages shadow cues in a scene to infer the surrounding illumination of the shadow-casting object. Unlike prior works that optimize a discrete environment map, we model scene illumination using a mixture of spherical Gaussians (SGs). SG illumination provides more intuitive relations to shadow appearance and offers a more compact parameterization compared to discrete environment maps. To estimate SG parameters, we employ an SG-based, differentiable, closed-form rendering equation to explain the shading of the shadow plane and minimize a photometric loss between the rendered and observed shadow plane shading. Experiments on synthetic and real-world images under various surrounding illumination demonstrate that our method estimates illumination more accurately than approaches based on discrete environment maps. With our estimated lighting, consistent shadow effects are realized when blending virtual objects into real-world images11Code is available at https://github.com/CyberAgentAILab/ShadowSG.. Hiroshi Kawasaki, Takafumi Taketomi |
3DV | 4 |
| 2025 | FreeUV: Ground-Truth-Free Realistic Facial UV Texture Recovery via Cross-Assembly Inference StrategyabstractRecovering high-quality 3D facial textures from single-view 2D images is a challenging task, especially under the constraints of limited data and complex facial details such as wrinkles, makeup, and occlusions. In this paper, we introduce FreeUV, a novel ground-truth-free UV texture recovery framework that eliminates the need for annotated or synthetic UV data. FreeUV leverages a pre-trained stable diffusion model alongside a Cross-Assembly inference strategy to fulfill this objective. In FreeUV, separate networks are trained independently to focus on realistic appearance and structural consistency, and these networks are combined during inference to generate coherent textures. Our approach accurately captures intricate facial features and demonstrates robust performance across diverse poses and occlusions. Extensive experiments validate FreeUV’s effectiveness, with results surpassing state-of-the-art methods in both quantitative and qualitative metrics. Additionally, FreeUV enables new applications, including local editing, facial feature interpolation, and texture recovery from multi-view images. By reducing data requirements, FreeUV offers a scalable solution for generating high-fidelity 3D facial textures suitable for real-world scenarios. Xingchao Yang, Takafumi Taketomi, Yuki Endo 0001, Yoshihiro Kanamori |
CVPR | 2 |
| 2025 | Neural Multi-View Self-Calibrated Photometric Stereo without Photometric Stereo CuesabstractWe propose a neural inverse rendering approach that jointly reconstructs geometry, spatially varying reflectance, and lighting conditions from multi-view images captured under varying directional lighting. Unlike prior multi-view photometric stereo methods that require light calibration or intermediate cues such as per-view normal maps, our method jointly optimizes all scene parameters from raw images in a single stage. We represent both geometry and reflectance as neural implicit fields and apply shadow-aware volume rendering. A spatial network first predicts the signed distance and a reflectance latent code for each scene point. A reflectance network then estimates reflectance values conditioned on the latent code and angularly encoded surface normal, view, and light directions. The proposed method outperforms state-of-the-art normal-guided approaches in shape and lighting estimation accuracy, generalizes to view-unaligned multi-light images, and handles objects with challenging geometry and reflectance. Takafumi Taketomi |
ICCV | 2 |
| 2025 | TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion InterpolationabstractWe present TANGO, a framework for generating co-speech body-gesture videos. Given a few-minute, single-speaker reference video and target speech audio, TANGO produces high-fidelity videos with synchronized body gestures. TANGO builds on Gesture Video Reenactment (GVR), which splits and retrieves video clips using a directed graph structure - representing video frames as nodes and valid transitions as edges. We address two key limitations of GVR: audio-motion misalignment and visual artifacts in GAN-generated transition frames. In particular, i) we propose retrieving gestures using latent feature distance to improve cross-modal alignment. To ensure the latent features could effectively model the relationship between speech audio and gesture motion, we implement a hierarchical joint embedding space (AuMoClip); ii) we introduce the diffusion-based model to generate high-quality transition frames. Our diffusion model, Appearance Consistent Interpolation (ACInterp), is built upon AnimateAnyone and includes a reference motion module and homography background flow to preserve appearance consistency between generated and reference videos. By integrating these components into the graph-based retrieval framework, TANGO reliably produces realistic, audio-synchronized videos and outperforms all existing generative and retrieval methods. Our code, pretrained models, and datasets are publicly available at https://github.com/CyberAgentAILab/TANGO. Xingchao Yang, Tomoya Akiyama, Yuantian Huang, Qiaoge Li, Shigeru Kuriyama, Takafumi Taketomi |
ICLR | 7 |
| 2025 | DeMapGS: Simultaneous Mesh Deformation and Surface Attribute Mapping via Gaussian SplattingabstractWe propose DeMapGS, a structured Gaussian Splatting framework that jointly optimizes deformable surfaces and surface-attached 2D Gaussian splats. By anchoring splats to a deformable template mesh, our method overcomes topological inconsistencies and enhances editing flexibility, addressing limitations of prior Gaussian Splatting methods that treat points independently. The unified representation in our method supports extraction of high-fidelity diffuse, normal, and displacement maps, enabling the reconstructed mesh to inherit the photorealistic rendering quality of Gaussian Splatting. To support robust optimization, we introduce a gradient diffusion strategy that propagates supervision across the surface, along with an alternating 2D/3D rendering scheme to handle concave regions. Experiments demonstrate that DeMapGS achieves state-of-the-art mesh reconstruction quality and enables downstream applications for Gaussian splats such as editing and cross-object manipulation through a shared parametric surface. Shengze Zhong, Kenshi Takayama, Takafumi Taketomi, Takeshi Oishi |
SIGGRAPH Asia | 4 |
| 2025 | BeautyBank: Encoding Facial Makeup in Latent SpaceabstractThe advancement of makeup transfer, editing, and image encoding has demonstrated their effectiveness and superior quality. However, existing makeup works primarily focus on low-dimensional features such as color distributions and patterns, limiting their versatillity across a wide range of makeup applications. Futhermore, existing high-dimensional latent encoding methods mainly target global features such as structure and style, and are less effective for tasks that require detailed attention to local color and pattern features of makeup. To overcome these limitations, we propose BeautyBank, a novel makeup encoder that disentangles pattern features of bare and makeup faces. Our method encodes makeup features into a high-dimensional space, preserving essential details necessary for makeup reconstruction and broadening the scope of potential makeup research applications. We also propose a Progressive Makeup Tuning (PMT) strategy, specifically designed to enhance the preservation of detailed makeup features while preventing the inclusion of irrelevant attributes. We further explore novel makeup applications, including facial image generation with makeup injection and makeup similarity measure. Extensive empirical experiments validate that our method offers superior task adaptability and holds significant potential for widespread application in various makeup-related fields. Furthermore, to address the lack of large-scale, high-quality paired makeup datasets in the field, we constructed the Bare-Makeup Synthesis Dataset (BMS), comprising 324,000 pairs of 512×512 pixel images of bare and makeup-enhanced faces. Qianwen Lu, Xingchao Yang, Takafumi Taketomi |
WACV | 3 |
| 2024 | SuperNormal: Neural Surface Reconstruction via Multi-View Normal IntegrationabstractWe present SuperNormal, a fast, high-fidelity approach to multi-view 3D reconstruction using surface normal maps. With a few minutes, SuperNormal produces detailed surfaces on par with 3D scanners. We harness volume ren-dering to optimize a neural signed distance function (SDF) powered by multi-resolution hash encoding. To accelerate training, we propose directional finite difference and patch-based ray marching to approximate the SDF gradients nu-merically. While not compromising reconstruction quality, this strategy is nearly twice as efficient as analytical gra-dients and about three times faster than axis-aligned finite difference. Experiments on the benchmark dataset demon-strate the superiority of SuperNormal in efficiency and ac-curacy compared to existing multi-view photometric stereo methods. On our captured objects, SuperNormal produces more fine-grained geometry than recent neural 3D reconstruction methods. Our code is available at https://github.com/CyberAgentAILab/SuperNormal. Takafumi Taketomi |
CVPR | 2 |
| 2024 | Makeup Prior Models for 3D Facial Makeup Estimation and ApplicationsabstractIn this work, we introduce two types of makeup prior models to extend existing 3D face prior models: PCA-based and StyleGAN2-based priors. The PCA-based prior model is a linear model that is easy to construct and is computationally efficient. However, it retains only low-frequency information. Conversely, the StyleGAN2-based model can represent high-frequency information with relatively higher computational cost than the PCA-based model. Although there is a trade-off between the two models, both are applicable to 3D facial makeup estimation and related applications. By leveraging makeup prior models and designing a makeup consistency module, we effectively address the challenges that previous methods faced in robustly estimating makeup, particularly in the context of handling self-occluded faces. In experiments, we demonstrate that our approach reduces computational costs by several orders of magnitude, achieving speeds up to 180 times faster. In addition, by improving the accuracy of the estimated makeup, we confirm that our methods are highly advantageous for various 3D facial makeup applications such as 3D makeup face reconstruction, user-friendly makeup editing, makeup transfer, and interpolation. Xingchao Yang, Takafumi Taketomi, Yuki Endo 0001, Yoshihiro Kanamori |
CVPR | 2 |
| 2023 | BlendFace: Re-designing Identity Encoders for Face-SwappingabstractThe great advancements of generative adversarial networks and face recognition models in computer vision have made it possible to swap identities on images from single sources. Although a lot of studies seems to have proposed almost satisfactory solutions, we notice previous methods still suffer from an identity-attribute entanglement that causes undesired attributes swapping because widely used identity encoders, e.g., ArcFace, have some crucial attribute biases owing to their pretraining on face recognition tasks. To address this issue, we design BlendFace, a novel identity encoder for face-swapping. The key idea behind BlendFace is training face recognition models on blended images whose attributes are replaced with those of another mitigates inter-personal biases such as hairsyles. BlendFace feeds disentangled identity features into generators and guides generators properly as an identity loss function. Extensive experiments demonstrate that BlendFace improves the identity-attribute disentanglement in face-swapping models, maintaining a comparable quantitative performance to previous methods. The code and models are available at https://github.com/mapooon/BlendFace. Kaede Shiohara, Xingchao Yang, Takafumi Taketomi |
ICCV | 3 |
| 2023 | Garment Model Extraction from Clothed Mannequin ScanabstractAbstract Modelling garments with rich details require enormous time and expertise of artists. Recent works re‐construct garments through segmentation of clothed human scan. However, existing methods rely on certain human body templates and do not perform as well on loose garments such as skirts. This paper presents a two‐stage pipeline for extracting high‐fidelity garments from static scan data of clothed mannequins. Our key contribution is a novel method for tracking both tight and loose boundaries between garments and mannequin skin. Our algorithm enables the modelling of off‐the‐shelf clothing with fine details. It is independent of human template models and requires only minimal mannequin priors. The effectiveness of our method is validated through quantitative and qualitative comparison with the baseline method. The results demonstrate that our method can accurately extract both tight and loose garments within reasonable time. Qiqi Gao 0001, Takafumi Taketomi |
Comput. Graph. Forum | 2 |
| 2023 | Refinement of Hair Geometry by Strand IntegrationabstractAbstract Reconstructing 3D hair is challenging due to its complex micro‐scale geometry, and is of essential importance for the efficient creation of high‐fidelity virtual humans. Existing hair capture methods based on multi‐view stereo tend to generate results that are noisy and inaccurate. In this study, we propose a refinement method for hair geometry by incorporating the gradient of strands into the computation of their position. We formulate a gradient integration strategy for hair strands. We evaluate the performance of our method using a synthetic multi‐view dataset containing four hairstyles, and show that our refinement produces more accurate hair geometry. Furthermore, we tested our method with a real image input. Our method produces a plausible result. Our source code is publicly available at https://github.com/elerac/strand_integration . Ryota Maeda, Kenshi Takayama, Takafumi Taketomi |
Comput. Graph. Forum | 3 |
| 2023 | Makeup Extraction of 3D Representation via Illumination-Aware Image DecompositionabstractAbstract Facial makeup enriches the beauty of not only real humans but also virtual characters; therefore, makeup for 3D facial models is highly in demand in productions. However, painting directly on 3D faces and capturing real‐world makeup are costly, and extracting makeup from 2D images often struggles with shading effects and occlusions. This paper presents the first method for extracting makeup for 3D facial models from a single makeup portrait. Our method consists of the following three steps. First, we exploit the strong prior of 3D morphable models via regression‐based inverse rendering to extract coarse materials such as geometry and diffuse/specular albedos that are represented in the UV space. Second, we refine the coarse materials, which may have missing pixels due to occlusions. We apply inpainting and optimization. Finally, we extract the bare skin, makeup, and an alpha matte from the diffuse albedo. Our method offers various applications for not only 3D facial models but also 2D portrait images. The extracted makeup is well‐aligned in the UV space, from which we build a large‐scale makeup dataset and a parametric makeup model for 3D faces. Our disentangled materials also yield robust makeup transfer and illumination‐aware makeup interpolation/removal without a reference image. Xingchao Yang, Takafumi Taketomi, Yoshihiro Kanamori |
Comput. Graph. Forum | 2 |
| 2022 | Context-based style transfer of tokenized gesturesabstractAbstract Gestural animations in the amusement or entertainment field often require rich expressions; however, it is still challenging to synthesize characteristic gestures automatically. Although style transfer based on a neural network model is a potential solution, existing methods mainly focus on cyclic motions such as gaits and require re‐training in adding new motion styles. Moreover, their per‐pose transformation cannot consider the time‐dependent features, and therefore motion styles of different periods and timings are difficult to be transferred. This limitation is fatal for the gestural motions requiring complicated time alignment due to the variety of exaggerated or intentionally performed behaviors. This study introduces a context‐based style transfer of gestural motions with neural networks to ensure stable conversion even for exaggerated, dynamically complicated gestures. We present a model based on a vision transformer for transferring gestures' content and style features by time‐segmenting them to compose tokens in a latent space. We extend this model to yield the probability of swapping gestures' tokens for style‐transferring. A transformer model is suited to semantically consistent matching among gesture tokens, owing to the correlation with spoken words. The compact architecture of our network model requires only a small number of parameters and computational costs, which is suitable for real‐time applications with an ordinary device. We introduce loss functions provided by the restoration error of identically and cyclically transferred gesture tokens and the similarity losses of content and style evaluated by splicing features inside the transformer. This design of losses allows unsupervised and zero‐shot learning, by which the scalability for motion data is obtained. We comparatively evaluated our style transfer method, mainly focusing on expressive gestures using our dataset captured for various scenarios and styles by introducing new error metrics tailored for gestures. Our experiment showed the superiority of our method in numerical accuracy and stability of style transfer against the existing methods. Shigeru Kuriyama, Tomohiko Mukai, Takafumi Taketomi, Tomoyuki Mukasa |
Comput. Graph. Forum | 3 |
| 2022 | BareSkinNet: De-makeup and De-lighting via 3D Face ReconstructionabstractAbstract We propose BareSkinNet, a novel method that simultaneously removes makeup and lighting influences from the face image. Our method leverages a 3D morphable model and does not require a reference clean face image or a specified light condition. By combining the process of 3D face reconstruction, we can easily obtain 3D geometry and coarse 3D textures. Using this information, we can infer normalized 3D face texture maps (diffuse, normal, roughness, and specular) by an image‐translation network. Consequently, reconstructed 3D face textures without undesirable information will significantly benefit subsequent processes, such as re‐lighting or re‐makeup. In experiments, we show that BareSkinNet outperforms state‐of‐the‐art makeup removal methods. In addition, our method is remarkably helpful in removing makeup to generate consistent high‐fidelity texture maps, which makes it extendable to many realistic face generation applications. It can also automatically build graphic assets of face makeup images before and after with corresponding 3D data. This will assist artists in accelerating their work, such as 3D makeup avatar creation. Xingchao Yang, Takafumi Taketomi |
Comput. Graph. Forum | 2 |
| 2021 | 4DComplete: Non-Rigid Motion Estimation Beyond the Observable SurfaceabstractTracking non-rigidly deforming scenes using range sensors has numerous applications including computer vision, AR/VR, and robotics. However, due to occlusions and physical limitations of range sensors, existing methods only handle the visible surface, thus causing discontinuities and in-completeness in the motion field. To this end, we introduce 4DComplete, a novel data-driven approach that estimates the non-rigid motion for the unobserved geometry. 4DComplete takes as input a partial shape and motion observation, extracts 4D time-space embedding, and jointly infers the missing geometry and motion field using a sparse fully-convolutional network. For network training, we constructed a large-scale synthetic dataset called DeformingThings4D, which consists of 1,972 animation sequences spanning 31 different animals or humanoid categories with dense 4D annotation. Experiments show that 4DComplete 1) reconstructs high-resolution volumetric shape and motion field from a partial observation, 2) learns an entangled 4D feature representation that benefits both shape and motion estimation, 3) yields more accurate and natural deformation than classic non-rigid priors such as As-Rigid-As-Possible (ARAP) deformation, and 4) generalizes well to unseen objects in real-world sequences. Yang Li 0143, Hikari Takehara, Takafumi Taketomi, Matthias Nießner |
ICCV | 3 |
| 2021 | SlidAR+: Gravity-aware 3D object manipulation for handheld augmented realityabstractAccurately placing virtual objects in a scene is a challenging tasks in handheld augmented reality (HAR). To add and arrange virtual objects in HAR, users must manipulate 6 degrees of freedom (DoFs) of the virtual object, namely: position (3) and orientation (3). However, it is difficult to manipulate all DoFs with the two-dimensional display of the handheld device. We present SlidAR+, a method for controlling the position and orientation of objects in HAR. SlidAR+ is an extension of SlidAR [1], a technique that allows users to control the position of a virtual object by manipulating only 1 DoF. We use the direction of gravity as a constraint to improve the user’s control and reduce the time it takes to adjust the orientation. Upon comparing it with a state-of-the-art object manipulation method, using SlidAR+, user were able to complete the tasks faster under our expected conditions and were also preferred by most participants. Varunyu Fuvattanasilp, Yuichiro Fujimoto, Alexander Plopski, Takafumi Taketomi, Christian Sandor, Masayuki Kanbara, Hirokazu Kato 0001 |
Comput. Graph. | 4 |
| 2021 | Impact of facial contour compensation on self-recognition in face-swapping technology
Haruka Matsumura, Takafumi Taketomi, Hirokazu Kato 0001 |
Multim. Tools Appl. | 2 |
| 2021 | Robust Reflectance Estimation for Projection-Based Appearance Control in a Dynamic Light EnvironmentabstractWe present a novel method that robustly estimates the reflectance, even in an environment with dynamically changing light. To control the appearance of an object by using a projector-camera system, an appropriate estimate of the object's reflectance is vital to the creation of an appropriate projection image. Most conventional estimation methods assume static light conditions; however, in practice, the appearance is affected by both the reflectance and environmental light. In an environment with dynamically changing light, conventional reflectance estimation methods require calibration every time the conditions change. In contrast, our method requires no additional calibration because it simultaneously estimates both the reflectance and environmental light. Our method is based on the concept of creating two different light conditions by switching the projection at a rate higher than that perceived by the human eye and captures the images of a target object separately under each condition. The reflectance and environmental light are then simultaneously estimated by using the pair of images acquired under these two conditions. We implemented a projector-camera system that switches the projection on and off at 120 Hz. Experiments confirm the robustness of our method when changing the environmental light. Further, our method can robustly estimate the reflectance under practical indoor lighting conditions. Ryo Akiyama, Goshiro Yamamoto, Toshiyuki Amano, Takafumi Taketomi, Alexander Plopski, Christian Sandor, Hirokazu Kato 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2020 | Illusory light: Perceptual appearance control using a projection-induced illusionabstractWith projection mapping, we can control the appearance of real-world objects by adding illumination. A projector can be used to control light reflected from an object, where the reflected light depends not only on the projection but also on the reflectance and environmental light. Because the resulting colors are affected by the reflectance and environmental light, the presentable color range of a projector is limited. The purpose of this work is to broaden this limited range by focusing on the perceived colors. Although our eyes capture reflected light to perceive the colors of an object, the colors perceived by humans are not always the same as the actual colors, and there are often significant differences between them because of the human visual system. To overcome the limitations of a projector based on human perception, we intentionally generate this difference by inducing a visual illusion, namely, color constancy. In this work, we designed an algorithm to determine the projected colors for presenting the desired colors perceptually by employing a color constancy effect. In addition, we conducted a user study and confirmed that our algorithm can (1) create a misperception regarding the color of illumination, (2) broaden the presentable color range of a projector, and (3) shift the perceptual colors in the desirable direction. Ryo Akiyama, Goshiro Yamamoto, Toshiyuki Amano, Takafumi Taketomi, Alexander Plopski, Yuichiro Fujimoto, Masayuki Kanbara, Christian Sandor, Hirokazu Kato 0001 |
Comput. Graph. | 4 |
| 2019 | Perceptual Appearance Control by Projection-Induced IllusionabstractUsing projection mapping, we can control the appearance of realworld objects by projecting colored light onto them. Since a projector can only add illumination to the scene, limited color gamut can be presented through projection mapping. However, actual color and perceived color are not always the same, and there is often large difference between them. We intentionally generate this difference by inducing visual illusion for extending the controllable color gamut of a projector. In particular, we induce color constancy. We demonstrate changing object color with and without inducing color constancy. Audiences perceive the controlled colors with and without illusion as different color nevertheless these colors are physically completely same. Ryo Akiyama, Goshiro Yamamoto, Toshiyuki Amano, Takafumi Taketomi, Christian Sandor, Alexander Plopski, Hirokazu Kato 0001 |
VR | 4 |
| 2019 | Towards large scale high fidelity collaborative augmented reality
Damien Constantine Rompapas, Christian Sandor, Alexander Plopski, Daniel Saakes, Joon Gi Shin, Takafumi Taketomi, Hirokazu Kato 0001 |
Comput. Graph. | 6 |
| 2018 | Multimodal Augmented Reality - Augmenting Auditory-Tactile Feedback to Change the Perception of Thickness
Geert Lugtenberg, Wolfgang Hürst, Nina Rosa, Christian Sandor, Alexander Plopski, Takafumi Taketomi, Hirokazu Kato 0001 |
MMM (1) | 6 |
| 2018 | Light Projection-Induced Illusion for Controlling Object ColorabstractUsing projection mapping, we can control the appearance of realworld objects by projecting colored light onto them. Because a projector can only add illumination to the scene, only a limited color gamut can be presented through projection mapping. In this paper we describe how the controllable color gamut can be extended by accounting for human perception and visual illusions. In particular, we induce color constancy to control what color space observers will perceive. In this paper, we explain the concept of our approach, and show first results of our system. Ryo Akiyama, Goshiro Yamamoto, Toshiyuki Amano, Takafumi Taketomi, Alexander Plopski, Christian Sandor, Hirokazu Kato 0001 |
VR | 4 |
| 2018 | Transferability of Spatial Maps: Augmented Versus Virtual Reality TrainingabstractWork space simulations help trainees acquire skills necessary to perform their tasks efficiently without disrupting the workflow, forgetting important steps during a procedure, or the location of important information. This training can be conducted in Augmented and Virtual Reality (AR, VR) to enhance its effectiveness and speed. When the skills are transferred to the actual application, it is referred to as positive training transfer. However, thus far, it is unclear which training, AR or VR, achieves better results in terms of positive training transfer. We compare the effectiveness of AR and VR for spatial memory training in a control-room scenario, where users have to memorize the location of buttons and information displays in their surroundings. We conducted a within-subject study with 16 participants and evaluated the impact the training had on short-term and long-term memory. Results of our study show that VR outperformed AR when tested in the same medium after the training. In a memory transfer test conducted two days later AR outperformed VR. Our findings have implications on the design of future training scenarios and applications. Nicko R. Caluya, Alexander Plopski, Jayzon Flores Ty, Christian Sandor, Takafumi Taketomi, Hirokazu Kato 0001 |
VR | 5 |
| 2018 | Towards Situated Knee Trajectory Visualization for Self Analysis in CyclingabstractInflammation, stiffness, and swelling are frequently reported symptoms of patellar tendinitis among cyclists; making knee pain a consistently observed overuse injury in cycling. In this paper, we investigate the applicability of a knee trajectory visualization to self-analysis for increasing awareness of movement patterns leading to injuries. We briefly explain overuse injuries and patellar instability, describe the experiments we did with cyclists for gathering requirements, and finally illustrate an augmented reality concept. We also show two different types of visualizations with participant opinions; one being conventional and other being a video-based one and discuss how situated visualizations can be utilized for improving self awareness to injury causes. Oral Kaplan, Goshiro Yamamoto, Takafumi Taketomi, Yasuhide Yoshitake, Alexander Plopski, Christian Sandor |
VR | 3 |
| 2018 | Augmented Reality versus Virtual Reality for 3D Object ManipulationabstractVirtual Reality (VR) Head-Mounted Displays (HMDs) are on the verge of becoming commodity hardware available to the average user and feasible to use as a tool for 3D work. Some HMDs include front-facing cameras, enabling Augmented Reality (AR) functionality. Apart from avoiding collisions with the environment, interaction with virtual objects may also be affected by seeing the real environment. However, whether these effects are positive or negative has not yet been studied extensively. For most tasks it is unknown whether AR has any advantage over VR. In this work we present the results of a user study in which we compared user performance measured in task completion time on a 9 degrees of freedom object selection and transformation task performed either in AR or VR, both with a 3D input device and a mouse. Our results show faster task completion time in AR over VR. When using a 3D input device, a purely VR environment increased task completion time by 22.5 percent on average compared to AR ( ). Surprisingly, a similar effect occurred when using a mouse: users were about 17.3 percent slower in VR than in AR ( ). Mouse and 3D input device produced similar task completion times in each condition (AR or VR) respectively. We further found no differences in reported comfort. Max Krichenbauer, Goshiro Yamamoto, Takafumi Taketomi, Christian Sandor, Hirokazu Kato 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | Handheld Guides in Inspection Tasks: Augmented Reality versus PictureabstractInspection tasks focus on observation of the environment and are required in many industrial domains. Inspectors usually execute these tasks by using a guide such as a paper manual, and directly observing the environment. The effort required to match the information in a guide with the information in an environment and the constant gaze shifts required between the two can severely lower the work efficiency of inspector in performing his/her tasks. Augmented reality (AR) allows the information in a guide to be overlaid directly on an environment. This can decrease the amount of effort required for information matching, thus increasing work efficiency. AR guides on head-mounted displays (HMDs) have been shown to increase efficiency. Handheld AR (HAR) is not as efficient as HMD-AR in terms of manipulability, but is more practical and features better information input and sharing capabilities. In this study, we compared two handheld guides: an AR interface that shows 3D registered annotations, that is, annotations having a fixed 3D position in the AR environment, and a non-AR picture interface that displays non-registered annotations on static images. We focused on inspection tasks that involve high information density and require the user to move, as well as to perform several viewpoint alignments. The results of our comparative evaluation showed that use of the AR interface resulted in lower task completion times, fewer errors, fewer gaze shifts, and a lower subjective workload. We are the first to present findings of a comparative study of an HAR and a picture interface when used in tasks that require the user to move and execute viewpoint alignments, focusing only on direct observation. Our findings can be useful for AR practitioners and psychology researchers. Jarkko Polvi, Takafumi Taketomi, Atsunori Moteki, Toshiyuki Yoshitake, Toshiyuki Fukuoka, Goshiro Yamamoto, Christian Sandor, Hirokazu Kato 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2017 | Promoting Short-Term Gains in Physical Exercise Through Digital Media Creation
Oral Kaplan, Goshiro Yamamoto, Takafumi Taketomi, Yasuhide Yoshitake, Alexander Plopski, Christian Sandor, Hirokazu Kato 0001 |
ACE | 3 |
| 2017 | Evaluating the effect of positional head-tracking on task performance in 3D modeling user interfaces
Max Krichenbauer, Goshiro Yamamoto, Takafumi Taketomi, Christian Sandor, Hirokazu Kato 0001, Steven K. Feiner |
Comput. Graph. | 3 |
| 2016 | The COMPASS Framework for Digital Entertainment: Discussing Augmented Reality Activities for ScoutsabstractEntertainment is challenging to observe, especially with children, due to limited analytical tools. In response, we present a modified framework for entertainment computing, COMPASS -- COmbined Mental, PhysicAl, Social and Spatial factors, which we use to analyze augmented reality activities for cub scouts. Marc Ericson C. Santos, Damien Constantine Rompapas, Yoshinari Nishiki, Takafumi Taketomi, Goshiro Yamamoto, Christian Sandor, Hirokazu Kato 0001 |
ACE | 4 |
| 2016 | In-situ visualization of pedaling forces on cycling training videosabstractOver the last decades, visual representations of data has been a commonly used medium to bolster human cognition in performance evaluation of professional athletes. However, the current approaches to these visualizations still build upon the paper based principles of initial designs with solid backgrounds. Due to this situation, same visualizations usually fail to provide explicit information about the physical characteristics of the scenario that the data was captured, such as the form of athletes. In this work, we present a data visualization method which combines visual representations of cyclist's pedaling with correlated frames of indoor training videos. We designed a prototype system which allows us to superimpose various pedaling visualizations onto simultaneously captured training videos of cyclists. The results of user studies we conducted with twelve professional cyclists confirmed their interest in new possibilities emerging from intuitive data visualizations. We also received valuable feedback about the feasible benefits of our approach over traditional approaches, such as reduced cognitive overload in understanding visualizations. We conclude by discussing the future implementations and application areas of our approach and further need of adjusting it to distinct training scenarios. Oral Kaplan, Goshiro Yamamoto, Yasuhide Yoshitake, Takafumi Taketomi, Christian Sandor, Hirokazu Kato 0001 |
SMC | 4 |
| 2016 | Exploring the perception of co-location errors during tool interaction in visuo-haptic augmented realityabstractCo-located haptic feedback in mixed and augmented reality environments can improve realism and user performance, but it also requires careful system design and calibration. In this poster, we determine the thresholds for perceiving co-location errors through two psychophysics experiments in a typical fine-motor manipulation task. In these experiments we simulate the two fundamental ways of implementing VHAR systems: first, attaching a real tool; second, augmenting a virtual tool. We determined the just-noticeable co-location errors for position and orientation in both experiments and found that users are significantly more sensitive to co-location errors with virtual tools. Our overall findings are useful for designing visuo-haptic augmented reality workspaces and calibration procedures. Ulrich Eck, Liem Hoang, Christian Sandor, Goshiro Yamamoto, Takafumi Taketomi, Hirokazu Kato 0001, Hamid Laga |
VR | 5 |
| 2016 | SharpView: Improved clarity of defocussed content on optical see-through head-mounted displaysabstractA common factor among current generation optical see-through augmented reality systems is fixed focal distance to virtual content. In this work, we investigate the issue of focus blur, in particular, the blurring caused by simultaneously viewing virtual content and physical objects in the environment at differing focal distances. We examine the application of dynamic sharpening filters as a straight forward, system independent, means for mitigating this effect improving the clarity of defocused AR content. We assess the utility of this method, termed SharpView, by employing an adjustment experiment in which users actively apply varying amounts of sharpening to reduce the perception of blur in AR content. Our experimental results validate the ability of our SharpView model to improve the visual clarity of focus blurred content, with optimal performance at focal differences well suited for near field AR applications. Kohei Oshima, Kenneth R. Moser, Damien Constantine Rompapas, J. Edward Swan II, Sei Ikeda, Goshiro Yamamoto, Takafumi Taketomi, Christian Sandor, Hirokazu Kato 0001 |
VR | 7 |
| 2016 | SlidAR: A 3D positioning method for SLAM-based handheld augmented realityabstractHandheld Augmented Reality (HAR) has the potential to introduce Augmented Reality (AR) to large audiences due to the widespread use of suitable handheld devices. However, many of the current HAR systems are not considered very practical and they do not fully answer to the needs of the users. One of the challenging areas in HAR is the in-situ AR content creation where the correct and accurate positioning of virtual objects to the real world is fundamental. Due to the hardware limitations of handheld devices and possible restrictions in the environment, the correct 3D positioning of objects can be difficult to achieve we are unable to use AR markers or correctly map the 3D structure of the environment. We present SlidAR, a 3D positioning for Simultaneous Localization And Mapping (SLAM) based HAR systems. SlidAR utilizes 3D ray-casting and epipolar geometry for virtual object positioning. It does not require a perfect 3D reconstruction of the environment nor any virtual depth cues. We have conducted a user experiment to evaluate the efficiency of SlidAR method against an existing device-centric positioning method that we call HoldAR. Results showed that SlidAR was significantly faster, required significantly less device movement, and also got significantly better subjective evaluation from the test participants. SlidAR also had higher positioning accuracy, although not significantly. Jarkko Polvi, Takafumi Taketomi, Goshiro Yamamoto, Arindam Dey 0001, Christian Sandor, Hirokazu Kato 0001 |
Comput. Graph. | 2 |
| 2016 | Exploring legibility of augmented reality X-ray
Marc Ericson C. Santos, Igor de Souza Almeida, Goshiro Yamamoto, Takafumi Taketomi, Christian Sandor, Hirokazu Kato 0001 |
Multim. Tools Appl. | 4 |
| 2015 | Toward Guidelines for Designing Handheld Augmented Reality in Learning Support
Marc Ericson C. Santos, Takafumi Taketomi, Goshiro Yamamoto, Ma. Mercedes T. Rodrigo, Christian Sandor, Hirokazu Kato 0001 |
ICCE | 2 |
| 2015 | Focal length change compensation for monocular slamabstractIn this paper, we propose a method for handling focal length changes in the SLAM algorithm. Our method is designed as a pre-processing step to first estimate the change of the camera focal length, and then compensate for the zooming effects before running the actual SLAM algorithm. By using our method, camera zooming can be used in the existing SLAM algorithms with minor modifications. In the experiments, the effectiveness of the proposed method was quantitatively evaluated. The results indicate that the method can successfully deal with abrupt changes of the camera focal length. Takafumi Taketomi, Janne Heikkilä |
ICIP | 1 |
| 2015 | Pseudo Printed Fabrics through Projection MappingabstractProjection-based Augmented Reality commonly projects on rigid objects, while only few systems project on deformable objects. In this paper, we present Pseudo Printed Fabrics (PPF), which enables the projection on a deforming piece of cloth. This can be applied to previewing a cloth design while manipulating its shape. We support challenging manipulations, including heavy occlusions and stretching the cloth. In previous work, we developed a similar system, based on a novel marker pattern; PPF extends it in two important aspects. First, we improved performance by two orders of magnitudes to achieve interactive performance. Second, we developed a new interpolation algorithm to keep registration during challenging manipulations. We believe that PPF can be applied to domains including virtual-try on and fashion design. Yuichiro Fujimoto, Goshiro Yamamoto, Takafumi Taketomi, Christian Sandor, Hirokazu Kato 0001 |
ISMAR | 3 |
| 2015 | Towards Estimating Usability Ratings of Handheld Augmented Reality Using Accelerometer DataabstractUsability evaluations are important to the development of augmented reality systems. However, conducting large-scale longitudinal studies remains challenging because of the lack of inexpensive but appropriate methods. In response, we propose a method for implicitly estimating usability ratings based on readily available sensor logs. To demonstrate our idea, we explored the use of features of accelerometer data in estimating usability ratings in an annotation task. Results show that our implicit method corresponds with explicit usability ratings at 79% and 84%. These results should be investigated further in other use cases, with other sensor logs. Marc Ericson C. Santos, Takafumi Taketomi, Goshiro Yamamoto, Gudrun Klinker, Christian Sandor, Hirokazu Kato 0001 |
ISMAR | 2 |
| 2015 | Abecedary Tracking and Mapping: A Toolkit for Tracking CompetitionsabstractThis paper introduces a toolkit with camera calibration, monocular visual Simultaneous Localization and Mapping (vSLAM) and registration with a calibration marker. With the toolkit, users can perform the whole procedure of the ISMAR on-site tracking competition in 2015. Since the source code is designed to be well-structured and highly-readable, users can easily install and modify the toolkit. By providing the toolkit, we encourage beginners to learn tracking techniques and to participate in the competition. Hideaki Uchiyama, Takafumi Taketomi, Sei Ikeda, Joao Paulo Silva do Monte Lima |
ISMAR | 2 |
| 2015 | Zoom factor compensation for monocular SLAMabstractSLAM algorithms are widely used in augmented reality applications for registering virtual objects. Most SLAM algorithms estimate camera poses and 3D positions of feature points using known intrinsic camera parameters that are calibrated and fixed in advance. This assumption means that the algorithm does not allow changing the intrinsic camera parameters during runtime. We propose a method for handling focal length changes in the SLAM algorithm. Our method is designed as a pre-processing step for the SLAM algorithm input. In our method, the change of the focal length is estimated before the tracking process of the SLAM algorithm. Camera zooming effects in the input camera images are compensated for by using the estimated focal length change. By using our method, camera zooming can be used in the existing SLAM algorithms such as PTAM [4] with minor modifications. In the experiment, the effectiveness of the proposed method was quantitatively evaluated. The results indicate that the method can successfully deal with abrupt changes of the camera focal length. Takafumi Taketomi, Janne Heikkilä |
VR | 1 |
| 2014 | Evaluating Augmented Reality for Situated Vocabulary LearningabstractAugmented reality (AR) is an emerging technology for communicating learning contents. Several AR systems are designed for learning. However, studies that have investigated instructional strategies for applying AR are few. This investigation requires the implementation of prototypes that use state-of-the-art technology and sound learning theory. In this work, we implemented two prototypes for learning Filipino and German words by first developing a handheld AR platform. These prototypes demonstrate situated vocabulary learning. Using our AR system, students can learn words related to their current environment. We assessed the quality of these prototypes by conducting usability evaluations. For the theoretical grounding, we leveraged on multimedia learning theory to design the content. Through our handheld AR platform, we evaluated situated vocabulary learning by comparing our prototypes to a flash cards application. In the first evaluation, students scored significantly lower when using AR in an immediate post-test. However, this difference disappeared after taking into account the variability in usability scores via analysis of covariance. Taking account usability is fairer when comparing an emerging technology to traditional technology. Test scores were also not significantly different in a delayed post-test. In the second evaluation, although the post-test score and answering time of students did not differ, our results showed that they feel more satisfied and can keep their attention better when using AR. For the first time, we demonstrated situated vocabulary learning by using AR. Moreover, our preliminary study confirms the intuition that students can achieve the same score using AR, but with benefits such as ease in maintaining attention and increased satisfaction. Marc Ericson C. Santos, Arno in Wolde Lübke, Takafumi Taketomi, Goshiro Yamamoto, Ma. Mercedes T. Rodrigo, Christian Sandor, Hirokazu Kato 0001 |
ICCE | 3 |
| 2014 | Authoring Augmented Reality as Situated MultimediaabstractAugmented reality (AR) is an enabling technology for presenting information in relation to real objects or real environments. AR is situated multimedia or information that is positioned in authentic physical contexts. In this paper, we discuss how we address issues in creating AR content for educational settings. From the learning theory perspective, we explain that AR is a logical extension of multimedia learning theory. From the development perspective, we demonstrate how AR content can be created through our in situ authoring tool and our platform for handheld AR. Marc Ericson C. Santos, Jayzon Flores Ty, Arno in Wolde Lübke, Ma. Mercedes T. Rodrigo, Takafumi Taketomi, Goshiro Yamamoto, Christian Sandor, Hirokazu Kato 0001 |
ICCE | 5 |
| 2014 | Towards Augmented Reality user interfaces in 3D media productionabstractThe idea of using Augmented Reality (AR) user interfaces (UIs) to create 3D media content, such as 3D models for movies and games has been repeatedly suggested over the last decade. Even though the concept is intuitively compelling and recent technological advances have made such an application increasingly feasible, very little progress has been made towards an actual real-world application of AR in professional media production. To this day, no immersive 3D UI has been commonly used by professionals for 3D computer graphics (CG) content creation. In this paper, we are first to publish a requirements analysis for our target application in the professional domain. Based on a survey that we conducted with media professionals, the analysis of professional 3D CG software, and professional training tutorials, we identify these requirements and put them into the context of AR UIs. From these findings, we derive several interaction design principles that aim to address the challenges of real-world application of AR to the production pipeline. We implemented these in our own prototype system while receiving feedback from media professionals. The insights gained in the survey, requirements analysis, and user interface design are relevant for research and development aimed at creating production methods for 3D media production. Max Krichenbauer, Goshiro Yamamoto, Takafumi Taketomi, Christian Sandor, Hirokazu Kato 0001 |
ISMAR | 3 |
| 2014 | Towards augmented reality user interfaces in 3D media productionabstractFor this demo, we present an Augmented Reality (AR) User Interface (UI) for the 3D design software Autodesk Maya, aimed at professional media creation. A user wears a head-mounted display (HMD) and thin cotton gloves which allow him to interact with virtual 3D models in the work area. Additional viewers can see the video stream on a projector and thus share the users view. Both head and hand positions are tracked from the HMD video stream, and an inertial measurement unit (IMU) and conductive materials on the gloves allow interaction with virtual objects. This system is built using Autodesk Maya — a professional 3D software package commonly used in the media industry — and aims to fulfill the requirements of professional 3D design work which we identified in our paper of the same title. While still an early prototype, it was already tested with media professionals to evaluate our approach. Max Krichenbauer, Goshiro Yamamoto, Takafumi Taketomi, Christian Sandor, Hirokazu Kato 0001 |
ISMAR | 3 |
| 2014 | Geometrically-correct projection-based texture mapping onto a clothabstractWe demonstrate the geometrically-correct projection-based texture mapping onto a deformable object like a cloth. This system can be used to simulate design that involves change in shape, such as sheets of malleable material. The geometrically-correct projection-based texture mapping onto a cloth is conducted using the measurement of object's 3D shape and the detection of the retro-reflective marker on the object's surface. Rapid prototyping is used as an example application of this projection technique. Yuichiro Fujimoto, Jun Miyazaki, Takafumi Taketomi, Hirokazu Kato 0001, Bruce H. Thomas, Goshiro Yamamoto, Ross Smith 0001 |
VR | 3 |
| 2014 | A usability scale for handheld augmented realityabstractHandheld augmented reality (HAR) applications must be carefully designed and improved based on user feedback to sustain commercial use. However, no standard questionnaire considers perceptual and ergonomic issues found in HAR. We address this issue by creating a HAR Usability Scale (HARUS). Marc Ericson C. Santos, Takafumi Taketomi, Christian Sandor, Jarkko Polvi, Goshiro Yamamoto, Hirokazu Kato 0001 |
VRST | 2 |
| 2014 | Camera pose estimation under dynamic intrinsic parameter change for augmented realityabstractIn this paper, we propose a method for estimating the camera pose for an environment in which the intrinsic camera parameters change dynamically. In video see-through augmented reality (AR) technology, image-based methods for estimating the camera pose are used to superimpose virtual objects onto the real environment. In general, video see-through-based AR cannot change the image magnification that results from a change in the camera׳s field-of-view because of the difficulty of dealing with changes in the intrinsic camera parameters. To remove this limitation, we propose a novel method for simultaneously estimating the intrinsic and extrinsic camera parameters based on an energy minimization framework. Our method is composed of both online and offline stages. An intrinsic camera parameter change depending on the zoom values is calibrated in the offline stage. Intrinsic and extrinsic camera parameters are then estimated based on the energy minimization framework in the online stage. In our method, two energy terms are added to the conventional marker-based method to estimate the camera parameters: reprojection errors based on the epipolar constraint and the constraint of the continuity of zoom values. By using a novel energy function, our method can accurately estimate intrinsic and extrinsic camera parameters. We confirmed experimentally that the proposed method can achieve accurate camera parameter estimation during camera zooming. Takafumi Taketomi, Kazuya Okada, Goshiro Yamamoto, Jun Miyazaki, Hirokazu Kato 0001 |
Comput. Graph. | 1 |
| 2014 | Geometrically-Correct Projection-Based Texture Mapping onto a Deformable ObjectabstractProjection-based Augmented Reality commonly employs a rigid substrate as the projection surface and does not support scenarios where the substrate can be reshaped. This investigation presents a projection-based AR system that supports deformable substrates that can be bent, twisted or folded. We demonstrate a new invisible marker embedded into a deformable substrate and an algorithm that identifies deformations to project geometrically correct textures onto the deformable object. The geometrically correct projection-based texture mapping onto a deformable marker is conducted using the measurement of the 3D shape through the detection of the retro-reflective marker on the surface. In order to achieve accurate texture mapping, we propose a marker pattern that can be partially recognized and can be registered to an object’s surface. The outcome of this work addresses a fundamental vision recognition challenge that allows the underlying material to change shape and be recognized by the system. Our evaluation demonstrated the system achieved geometrically correct projection under extreme deformation conditions. We envisage the techniques presented are useful for domains including prototype development, design, entertainment and information based AR systems. Yuichiro Fujimoto, Ross Smith 0001, Takafumi Taketomi, Goshiro Yamamoto, Jun Miyazaki, Hirokazu Kato 0001, Bruce H. Thomas |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2013 | Augmented Reality X-Ray Interaction in K-12 Education: Theory, Student Perception and Teacher EvaluationabstractAugmented reality (AR) x-ray interaction is an enabling technology for providing students with virtual abstractions of the interior of an object. It provides students contextual visualization - the presentation of virtual information in the rich context of a real environment - thereby offering compelling experiences. According to experiential learning theory, such personal experiences are necessary for reaching different types of learners. AR x-ray is a novel interaction technique for education, thus, it is necessary to investigate how it affects the students' perception. We implemented AR x-ray using a state-of-the-art occlusion technique, and compared it to viewing 3D objects without occlusion. Results of two user studies (n=23 and n=47) show that there are no significant differences in realism, perception of depth, and visibility with occlusion and without occlusion, and that the current technique is usable for educational purposes. We also conducted interviews with both students (n=23) and teachers (n=12). Results indicate that AR x-ray is perceived to be useful for motivating and explaining to students. The teachers expressed willingness to adopt AR x-ray and to undergo training for using AR-based teaching materials. Marc Ericson C. Santos, Angie Chen, Mitsuaki Terawaki, Goshiro Yamamoto, Takafumi Taketomi, Jun Miyazaki, Hirokazu Kato 0001 |
ICALT | 5 |
| 2013 | Authoring Augmented Reality Learning Experiences as Learning ObjectsabstractEngineers and educators alike have prototyped a variety of augmented reality learning experiences (ARLEs). However, adapting ARLEs in educational practice would require an interdisciplinary approach that considers learning theory, pedagogy and instructional design. To address this requirement, we model ARLEs as learning objects by outlining the necessary components, and we propose a participatory design to demonstrate the authoring process of an augmented reality learning object (ARLO). ARLOs can be made useful in many scenarios if teachers are empowered to edit its context elements, content and instructional activity. Lastly, we point to the research questions entailed in modeling ARLEs as ARLOs. Marc Ericson C. Santos, Goshiro Yamamoto, Takafumi Taketomi, Jun Miyazaki, Hirokazu Kato 0001 |
ICALT | 3 |
| 2013 | Detection of 3D points on moving objects from point cloud data for 3D modeling of outdoor environmentsabstractA 3D modeling technique for an urban environment can be applied to several applications such as landscape simulations, navigational systems, and mixed reality systems. In this field, the target environment is first measured using several types of sensors (laser rangefinders, cameras, GPS sensors, and gyroscopes). A 3D model of the environment is then constructed based on the results of the 3D measurements. In this 3D modeling process, 3D points that exist on moving objects become obstacles or outliers to enable the construction of an accurate 3D model. To solve this problem, we propose a method for detecting 3D points on moving objects from 3D point cloud data based on photometric consistency and knowledge of the road environment. In our method, 3D points on moving objects are detected based on luminance variations obtained by projecting 3D points onto omnidirectional images. After detecting 3D the points based on evaluation value, the points are detected using prior information of the road environment. Tsunetake Kanatani, Hideyuki Kume, Takafumi Taketomi, Tomokazu Sato, Naokazu Yokoya |
ICIP | 3 |
| 2013 | Geometric registration for zoomable camera using epipolar constraint and pre-calibrated intrinsic camera parameter changeabstractIn general, video see-through based augmented reality (AR) cannot change the magnification of camera zooming parameter due to the difficulty of dealing with changes in intrinsic camera parameters. To realize the usage of camera zooming in AR, we propose a novel simultaneous intrinsic and extrinsic camera parameter estimation method based on an energy minimization framework. Our method is composed of the online and offline stages. An intrinsic camera parameter change depending on the zoom values is calibrated in the offline stage. Intrinsic and extrinsic camera parameters are then estimated based on the energy minimization framework in the online stage. In our method, two energy terms are added to the conventional marker-based camera parameter estimation method. One is reprojection errors based on the epipolar constraint. The other is the constraint of continuity of zoom values. By using a novel energy function, our method can estimate accurate intrinsic and extrinsic camera parameters. In an experiment, we confirmed that the proposed method can achieve accurate camera parameter estimation during camera zooming. Takafumi Taketomi, Kazuya Okada, Goshiro Yamamoto, Jun Miyazaki, Hirokazu Kato 0001 |
ISMAR | 1 |
| 2012 | Where Buddhism Encounters Entertainment Computing
Daisuke Uriu, Naohito Okude, Masahiko Inami, Takafumi Taketomi, Chihiro Sato |
Advances in Computer Entertainment | 4 |
| 2012 | Position estimation of near point light sources using a clear hollow sphere
Takahito Aoto 0002, Takafumi Taketomi, Tomokazu Sato, Yasuhiro Mukaigawa, Naokazu Yokoya |
ICPR | 2 |
| 2012 | Robust model-based tracking considering changes in the measurable DoF of the target object
Kenzo Kumagai, Marina Atsumi Oikawa, Takafumi Taketomi, Goshiro Yamamoto, Jun Miyazaki, Hirokazu Kato 0001 |
ICPR | 3 |
| 2012 | Fast and incremental indexing in effective and efficient XML element retrieval systemsabstractA method for fast and incremental indexing, with both effective and efficient query processing, is proposed for XML element retrieval. When frequent document updates occur on the Web, they must be handled to maintain the effectiveness of the search system. When new topics are added and document statistics change drastically, search accuracy is also reduced. We therefore consider a method not only for updating indices efficiently but also for processing queries effectively and efficiently. We construct indices for fast updating and propose a method for computing accurate term weights even under dynamically changing statistics. Experimental results show that our proposed system can handle document updates at low cost and search documents accurately even when their statistics change. Atsushi Keyaki, Jun Miyazaki, Kenji Hatano, Goshiro Yamamoto, Takafumi Taketomi, Hirokazu Kato 0001 |
iiWAS | 5 |
| 2012 | Relationship between features of augmented reality and user memorizationabstractThe objective of this study is to investigate the relationship between the features of augmented reality (AR) and human memorization ability. The basis of this relation is derived from the following features. The AR feature is that AR can provide information associated with specific locations in the real world. The feature of human memory is that humans can easily memorize information if the information is visually associated with specific locations. To investigate this relation, we conduct a pilot user study in which blocks are picked from some drawers. As a result, significant differences are found between a situation in which visual information is displayed at the location of each drawer in the real world and that in which textual information is displayed at an unrelated location. Yuichiro Fujimoto, Goshiro Yamamoto, Takafumi Taketomi, Jun Miyazaki, Hirokazu Kato 0001 |
ISMAR | 3 |
| 2012 | Augmented prototyping of 3D rigid curved surfacesabstractThis paper presents an application of Augmented Reality (AR) in Rapid Prototyping (RP) of non-textured rigid curved surfaces. By enhancing the prototypes with AR, evaluation of its design and aesthetic concepts in real-time becomes easier, saving time and production costs. In our application, no fiducial markers are required and the CAD model used to build the prototype is applied in an edge-based tracking system specially designed to deal with curved shapes. Results from a pilot user study comparing the use of a 3D software and the proposed application are also presented. Marina Atsumi Oikawa, Igor de Souza Almeida, Takafumi Taketomi, Goshiro Yamamoto, Jun Miyazaki, Hirokazu Kato 0001 |
ISMAR | 3 |
| 2012 | Interactive photomosaic system using GPUabstractA photomosaic is a type of decorative art made up from various other photographs. We present a method for quickly generating photomosaics and propose an interactive recursive photomosaic system. Users can operate the system by using a large display with a touch input function, which allows them to alter the appearance of the image dynamically. Makoto Fujisawa, Toshiyuki Amano, Takafumi Taketomi, Goshiro Yamamoto, Yuuki Uranishi, Jun Miyazaki |
ACM Multimedia | 3 |
| 2011 | Real-time and accurate extrinsic camera parameter estimation using feature landmark database for augmented reality
Takafumi Taketomi, Tomokazu Sato, Naokazu Yokoya |
Comput. Graph. | 1 |
| 2010 | Extrinsic Camera Parameter Estimation Using Video Images and GPS Considering GPS Positioning AccuracyabstractThis paper proposes a method for estimating extrinsic camera parameters using video images and position data acquired by GPS. In conventional methods, the accuracy of the estimated camera position largely depends on the accuracy of GPS positioning data because they assume that GPS position error is very small or normally distributed. However, the actual error of GPS positioning easily grows to the 10m level and the distribution of these errors is changed depending on satellite positions and conditions of the environment. In order to achieve more accurate camera positioning in outdoor environments, in this study, we have employed a simple assumption that true GPS position exists within a certain range from the observed GPS position and the size of the range depends on the GPS positioning accuracy. Concretely, the proposed method estimates camera parameters by minimizing an energy function that is defined by using the reprojection error and the penalty term for GPS positioning. Hideyuki Kume, Takafumi Taketomi, Tomokazu Sato, Naokazu Yokoya |
ICPR | 2 |
| 2008 | Real-time outdoor pre-visualization method for videographers - real-time geometric registration using point-based model - abstractThis paper describes a real-time pre-visualization method using augmented reality techniques for videographers. It enables them to test camera work without real actors in a real environment. As a substitute for real actors, virtual ones are superimposed on a live video in real-time according to a real camera motion and an illumination condition. The key technique of this method is real-time motion estimation of a camera, which can be applied to unknown complex environments including natural objects. In our method, geometric and photometric registration problems for such unknown environments are solved to realize the above visualization. A prototype system demonstrates availability of the pre-visualization method. Sei Ikeda, Takafumi Taketomi, Bunyo Okumura, Tomokazu Sato, Masayuki Kanbara, Naokazu Yokoya, Kunihiro Chihara |
ICME | 2 |
| 2008 | Real-time camera position and posture estimation using a feature landmark database with prioritiesabstractIn the field of computer vision, many kinds of camera parameter estimation methods have been proposed. As one of these methods, an extrinsic camera parameter estimation method that uses pre-constructed feature landmark database has been studied. In this method, extrinsic camera parameters of video images are estimated from correspondences between landmarks and image features. Although this method can work in a large outdoor environment, its computational cost in matching process is expensive and it cannot work in real-time. In this paper, to achieve real-time camera parameter estimation, the number of matching candidates are reduced by using priorities of landmarks that are determined from previously captured video sequences. Takafumi Taketomi, Tomokazu Sato, Naokazu Yokoya |
ICPR | 1 |