Hideaki Uchiyama

dblp:14/1212 · DBLP profile ↗
← Back
54ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0002-6119-1184ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 9 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 21 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 13 · 1 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 EMA: Effort Metric Attention for Anatomical Effort-Guided Human Motion Diffusion
Joshua Siy, Huakun Liu, Yutaro Hirao, Monica Perusquía-Hernández, Hideaki Uchiyama, Kiyoshi Kiyokawa
FG5
2026 Electrooculography-Based Detection of Refractive Vision Problems
abstract
Early detection of visual impairments remains a persistent challenge, especially due to the subtle and often unnoticed nature of early-stage symptoms. Recent works have attempted to transition clinical tests to home-based services or develop innovative diagnostic methods, but most approaches remain self-initiated and discrete. In this study, we focused on refractive disorders and explored the feasibility of using electrooculography (EOG) to detect changes in refractive power passively. Thirty-nine participants used optometry trial lenses to simulate different refractive conditions. Participants performed a series of visual tasks while their EOG signals were recorded. We trained classification models to predict simulated refractive power levels relative to baseline visual condition across multiple evaluation settings, including within-subject, temporal generalization, and across-subject scenarios. The findings reveal that refractive power classification models achieve a mean accuracy of $0.950 \pm 0.034$ in within-subject, within-condition scenarios. Within-subject models tested on data from a different time point showed highly variable performance. While some participants achieved promising results, overall accuracy remained low, with a mean of $0.159 \pm 0.285$. We employed three strategies to evaluate the across-subject models. Naive models performed poorly ($0.161 \pm 0.063$) and linear normalization provided limited improvement ($0.175 \pm 0.062$). However, the fine-tuning strategy substantially improved the model's performance ($0.785 \pm 0.123$). EOG signals contain useful information for refractive power classification, particularly in personalized contexts. However, generalizing across time and individuals remains challenging. Overall, this work offers valuable insights for advancing EOG-based systems aimed at passive, real-time monitoring of visual conditions.
Xin Wei 0007, Huakun Liu, Yutaro Hirao, Monica Perusquía-Hernández, Katsutoshi Masai, Hideaki Uchiyama, Kiyoshi Kiyokawa
IEEE J. Biomed. Health Informatics6
2026 Visual and Somatosensory Integration With Higher Sitting Posture Enhances the Sense of Standing and Self-Motion in Seated VR
abstract
Users are often seated in the real environment, while their virtual avatars either remain standing stationary or move in virtual reality (VR). This creates posture inconsistencies between the real and virtual embodiment representations. The relationship between posture consistency in locomotion techniques and sense of presence in VR is still unclear. This study investigates how visual and somatosensory integration affects the sense of standing (SoSt) and the sense of self-motion (SoSm) when the sitting posture is varied slightly, including highlighting the importance of sitting posture for locomotion design in VR. The degree and occurrence of SoSt and SoSm were assessed by subjective experiments, and it was found that higher sitting and lower sitting postures present higher SoSt and lower SoSm, respectively. Invocation of SoSt also influences postural perception. Perception of travel distance varied according to the posture condition when identical visual flow was presented. The findings suggest that visual and somatosensory integration related to posture enhances SoSt and SoSm, and a sitting posture with a higher seating position is recommended in seated VR locomotion design.
Daiki Hagimori, Naoya Isoyama, Monica Perusquía-Hernández, Shunsuke Yoshimoto, Hideaki Uchiyama, Nobuchika Sakata, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.5
2026 Tag-Along Virtual Windows Increase Perceived Resistance and Task Load in Augmented Reality
abstract
Augmented Reality (AR) can enhance accessibility by anchoring virtual windows to the user's body. Among common approaches, head-following windows help maintain floating virtual windows within the user's field of view. Previous studies have actively explored this new design space to improve user experience and efficiency. In contrast, this study focuses on the perceived resistance of head-following windows in AR, despite their lack of physical mass. We conducted a within-subject experiment with 24 participants, manipulating Follow-Up Delay, Window Size, and UI Type. We measured subjective resistance ratings, NASA-TLX (Raw TLX Scores), and the gaze-head angular offset. The results showed that both a certain level of Follow-Up Delay and the Tag-Along elicited significantly stronger perceived resistance as well as task load. Although Window Size alone did not show a significant effect on resistance ratings, we observed an interaction between the size and UI Type. These findings extend existing pseudo-haptics research by revealing the previously unexplored domain of resistance in head-based interactions with head-following virtual windows. We further provide design implications for head-following windows in AR.
Motoki Kagami, Yuta Kataoka, Yutaro Hirao, Monica Perusquía-Hernández, Satoshi Hashiguchi, Hideaki Uchiyama, Kiyoshi Kiyokawa, Shohei Mori
IEEE Trans. Vis. Comput. Graph.6
2025 UMotion: Uncertainty-driven Human Motion Estimation from Inertial and Ultra-wideband Units
abstract
Sparse wearable inertial measurement units (IMUs) have gained popularity for estimating 3D human motion. However, challenges such as pose ambiguity, data drift, and limited adaptability to diverse bodies persist. To address these issues, we propose UMotion, an uncertainty-driven, online fusing-all state estimation framework for 3D human shape and pose estimation, supported by six integrated, body-worn ultra-wideband (UWB) distance sensors with IMUs. UWB sensors measure inter-node distances to infer spatial relationships, aiding in resolving pose ambiguities and body shape variations when combined with anthropometric data. Unfortunately, IMUs are prone to drift, and UWB sensors are affected by body occlusions. Consequently, we develop a tightly coupled Unscented Kalman Filter (UKF) framework that fuses uncertainties from sensor data and estimated human motion based on individual body shape. The UKF iteratively refines IMU and UWB measurements by aligning them with uncertain human motion constraints in real-time, producing optimal estimates for each. Experiments on both synthetic and real-world datasets demonstrate the effectiveness of UMotion in stabilizing sensor data and the improvement over state of the art in pose accuracy. Code is available at: https://github.com/kk9six/umotion.
Huakun Liu, Hiroki Ota, Xin Wei 0007, Yutaro Hirao, Monica Perusquía-Hernández, Hideaki Uchiyama, Kiyoshi Kiyokawa
CVPR6
2025 Have a Seat: An Enhanced Reactive Alignment of a Single Target's Position and Angle from the User's Perspective in VR
abstract
Redirected Walking (RDW) techniques allow users to explore virtually infinite environments within constrained physical spaces. However, achieving precise alignment between physical and virtual targets remains a significant challenge. In particular, when both position and orientation of the targets need to align in order to get a proper haptic feedback like siting on a virtual chair. This paper introduces a revised version of the Reactive Alignment (REA) controller that simultaneously minimizes the Angular and Positional Distance Errors between a physical and a virtual target. The proposed method enhances spatial alignment and optimizes user navigation using a novel rotation gain control algorithm that takes angular misalignment into account. In addition, a new metric,$\Delta p$, is proposed to quantify the angular alignment, complementing the redefined Physical Distance Error (PDE) for positional accuracy. We implemented the algorithm on Oculus Quest head-mounted display and utilized the HMD's physical space tracking to locate the physical prop's location without the need of any external tracking. We also incorporated saccadic redirection by utilizing the HMD's eyetracking functionality to complement the revised REA approach. A user study demonstrates that the revised REA controller outperforms the original REA by reducing Physical Distance Error, angular error, and reset counts. It also enhanced user interaction with physical props by enabling users to successfully sit on a physical chair 60% of the time compared to 0% with the original REA when$\Delta p$is zero.
Habiba H. AbdelAziz, Yutaro Hirao, Monica Perusquía-Hernández, Hideaki Uchiyama, Kiyoshi Kiyokawa
ISMAR4
2025 Mind Your Vision: A Passive Multimodal Framework for Refractive Disorders Measurement Combining Electrooculography and Eye Tracking
abstract
Refractive errors are among the most common visual impairments globally, yet their diagnosis often relies on active user participation and clinical oversight. This study explores a passive method for estimating refractive power using two eye movement recording techniques: electrooculography (EOG) and video-based eye tracking. Using a publicly available dataset recorded under varying diopter conditions, we trained Long Short-Term Memory (LSTM) models to classify refractive power from unimodal (EOG or video-based eye tracking) and multimodal configurations. In the context of eye movement analysis, EOG captures fine-grained electrical signals, while video-based tracking provides rich features such as pupil dynamics and gaze behavior, making the two modalities complementary. We assess performance in both subject-dependent and subject-independent settings to evaluate model personalization and generalizability across individuals. Results show that the multimodal model consistently outperforms unimodal models, achieving the highest average accuracy in both settings: 96.568% in the subject-dependent scenario and 9.344% in the subject-independent scenario. Statistical comparisons in the subject-dependent setting confirmed that both unimodal and multimodal models significantly exceeded the chance level. Among them, the multimodal model significantly outperformed the EOG and eye-tracking models. The strong performance of subject-dependent models highlights the potential for developing personalized models tailored to the target user for refractive power monitoring. However, generalization remains limited, with classification accuracy only marginally above chance in the subject-independent evaluations. Our findings demonstrate both the potential and current limitations of eye movement data-based refractive error estimation, contributing to the development of continuous, non-invasive screening methods using EOG signals and eye-tracking data.
Xin Wei 0007, Huakun Liu, Yutaro Hirao, Monica Perusquía-Hernández, Katsutoshi Masai, Hideaki Uchiyama, Kiyoshi Kiyokawa
MUM6
2025 Perception-Driven Soft-Edge Occlusion for Optical See-Through Head-Mounted Displays
abstract
Systems with occlusion capabilities, such as those used in vision augmentation, image processing, and optical see-through head-mounted display (OST-HMD), have gained popularity. Achieving precise (hard-edge) occlusion in these systems is challenging, often requiring complex optical designs and bulky volumes. On the other hand, utilizing a single transparent liquid crystal display (LCD) is a simple approach to create occlusion masks. However, the generated mask will appear defocused (soft-edge) resulting in insufficient blocking or occlusion leakage. In our work, we delve into the perception of soft-edge occlusion by the human visual system and present a preference-based optimal expansion method that minimizes perceived occlusion leakage. In a user study involving 20 participants, we made a noteworthy observation that the human eye perceives a sharper edge blur of the occlusion mask when individuals see through it and gaze at a far distance, in contrast to the camera system's observation. Moreover, our study revealed significant individual differences in the perception of soft-edge masks in human vision when focusing. These differences may lead to varying degrees of demand for mask size among individuals. Our evaluation demonstrates that our method successfully accounts for individual differences and achieves optimal masking effects at arbitrary distances and pupil sizes.
Xiaodan Hu, Yan Zhang 0101, Alexander Plopski, Yuta Itoh 0001, Monica Perusquía-Hernández, Naoya Isoyama, Hideaki Uchiyama, Kiyoshi Kiyokawa
IEEE Trans. Vis. Comput. Graph.7
2025 Application of Transitional Mixed Reality Interfaces: A Co-Design Study with Flood-Prone Communities
abstract
Flood risk communication in disaster-prone communities often relies on traditional tools (e.g., paper and browser-based hazard/flood maps) that struggle to engage community stakeholders and reflect intuitive flood situations. In this paper, we applied the transitional mixed reality (MR) interface concept from pioneering work and extended it for flood risk communication scenarios through co-design with community stakeholders to help vulnerable residents understand flood risk and facilitate preparedness. Starting with an initial transitional MR prototype, we conducted three iterative workshops - each dedicated to device usability, visualization techniques, and interaction methods. We collaborated with diverse community stakeholders in flood-prone areas, collecting feedback to refine the system according to community needs. Our preliminary evaluation indicates that this co-designed system significantly improves user understanding and engagement compared to traditional tools, though some older residents faced usability challenges. We detailed this iterative co-design process, critical insights and design implications, offering our work as a practical case of mixed reality application in strengthening flood risk communication. We also discuss the system's potential to support community-driven collaboration in flood preparedness.
Zhiling Jie, Geert Lugtenberg, Armin Teubert, Makoto Fujisawa, Hideaki Uchiyama, Kiyoshi Kiyokawa, Isidro Butaslac, Taishi Sawabe, Hirokazu Kato 0001
IEEE Trans. Vis. Comput. Graph.6
2024 U2R: Underwater Ultrasonic Reflection Wave Dataset Toward Pose-Invariant Material Recognition
abstract
In underwater environments, the reflected ultrasonic waves from objects generally provide more than just information about their color and shape for object recognition. Previous studies have overlooked the influence of object pose on these wave components. It is crucial to investigate how these poses affect the reflected wave components because object poses can vary widely and are often unpredictable in real-world scenarios. In this work, we introduce a novel dataset comprising reflected wave components collected from objects made of various materials and observed from various angles. We also show the preliminary evaluations on the performance of machine learning-based material classification on object pose. Our results indicate that the accuracy is consistently high (≥ 91%) for known angles but significantly drops (< 60%) when dealing with unknown angles in most cases. Based on these evaluations, we suggest several directions for future research. Our dataset is available at https://github.com/Nyamotaro/U2R.
Mayuka Kono, Yutaro Hirao, Monica Perusquía-Hernández, Naoya Isoyama, Hideaki Uchiyama, Nobuchika Sakata, Jun Takamatsu, Kiyoshi Kiyokawa
ICASSP5
2024 First-Person Perspective Induces Stronger Feelings of Awe and Presence Compared to Third-Person Perspective in Virtual Reality
abstract
Awe is a complex emotion described as a perception of vastness and a need for accommodation to integrate new, overwhelming experiences. Virtual Reality (VR) has recently gained attention as a convenient means to facilitate experiences of awe. In VR, a first-person perspective might increase awe due to its immersive nature, while a third-person perspective might enhance the perception of vastness. However, the impact of VR perspectives on experiencing awe has not been thoroughly examined. We created two types of VR scenes: one with elements designed to induce high awe, such as a snowy mountain, and a low awe scene without such elements. We compared first-person and third-person perspectives in each scene. Forty-two participants explored the VR scenes, with their physiological responses captured by electrocardiogram (ECG) and face tracking (FT). Subsequently, participants self-reported their experience of awe (AWE-S) and presence (IPQ) within VR. The results revealed that the first-person perspective induced stronger feelings of awe and presence than the third-person perspective. The findings of this study provide useful guidelines for designing VR content that enhances emotional experiences.
Hiromu Otsubo, Alexander Marquardt, Melissa Steininger, Marvin Lehnort, Felix Dollack, Yutaro Hirao, Monica Perusquía-Hernández, Hideaki Uchiyama, Ernst Kruijff, Bernhard E. Riecke, Kiyoshi Kiyokawa
ICMI8
2024 Hap'n'Roll: A Scroll-inspired Device for Delivering Diverse Haptic Feedback with a Single Actuator
abstract
Hap’n’Roll is a wearable device that leverages the concept of a scroll to present, with a single motor, tactile sensations of various sizes, shapes, and textures. Hap’n’Roll is composed of two axes, a sheet, and one motor. By changing the number of sheet wraps, the thickness within the user’s hand can be adjusted. Additionally, using holes on the sheet to secure the fingertips, it can present a wide range of sizes and shapes. Unlike typical existing handheld shape-changing devices, Hap’n’Roll is not limited to cylindrical forms. Furthermore, by moving different materials attached on the sheet to the fingertips, it can also express different textures. A user study showed that Hap’n’Roll can convey at least three sizes (small, medium, and large) and four types of shapes (a cylinder, a rectangle, a cone, and a cup), with a shape and size identification accuracy of approx. 76.1%. The identification accuracy for shape alone was approx. 98.5%. Moreover, several applications were developed to showcase the effectiveness of Hap’n’Roll’s mechanism for various haptic feedback.
Hiroki Ota, Daiki Hagimori, Monica Perusquía-Hernández, Naoya Isoyama, Yutaro Hirao, Hideaki Uchiyama, Kiyoshi Kiyokawa
VR6
2024 Fast direct multi-person radiance fields from sparse input with dense pose priors
abstract
Volumetric radiance fields have been popular in reconstructing small-scale 3D scenes from multi-view images. With additional constraints such as person correspondences, reconstructing a large 3D scene with multiple persons becomes possible. However, existing methods fail for sparse input views or when person correspondences are unavailable. In such cases, the conventional depth image supervision may be insufficient because it only captures the relative position of each person with respect to the camera center. In this paper, we investigate an alternative approach by supervising the optimization framework with a dense pose prior that represents correspondences between the SMPL model and the input images. The core ideas of our approach consist in exploiting dense pose priors estimated from the input images to perform person segmentation and incorporating such priors into the learning of the radiance field. Our proposed dense pose supervision is view-independent, significantly speeding up computational time and improving 3D reconstruction accuracy , with less floaters and noise. We confirm the advantages of our proposed method with extensive evaluation in a subset of the publicly available CMU Panoptic dataset. When training with only five input views, our proposed method achieves an average improvement of 6.1% in PSNR , 3.5% in SSIM , 17.2% in LPIPS vgg , 19.3% in LPIPS alex , and 39.4% in training time.
Joao Paulo Silva do Monte Lima, Hideaki Uchiyama, Diego Thomas, Veronica Teichrieb
Comput. Graph.2
2023 Effects of Visual Presentation Near the Mouth on Cross-Modal Effects of Multisensory Flavor Perception and Ease of Eating
abstract
Various studies have suggested that altering the appearance of food can impact multisensory flavor perception. The cross-modal effect of such visual changes on gustation may allow for the presentation of food tastes that are difficult to express with simple combinations of taste stimuli. This cross-modal effect of visual changes on gustation holds potential for applications in gustatory displays. However, the current limitation of existing Head-Mounted Displays (HMDs) is their restricted vertical Field of View (FoV), which prohibits the display of images near the mouth while eating. This limitation may impede the cross-modal effect of visual changes on multisensory flavor perception. Additionally, the lack of visibility around the mouth area challenges the ease of eating. To address these issues, we design a Video See-Through (VST)-HMD with an expanded vertical FoV (approx. 100 [deg]). Using the HMD, we investigated how presenting visual information near the mouth affects the cross-modal effects of flavor perception and ease of eating. In our experiment, machine learning techniques were utilized to alter the appearance of food. However, the result showed no significant differences in the amount of cross-modal effects or the ease of eating between the groups with and without visual information near the mouth. As a discussion of this result, the participants may not direct their visual attention to the food when they put the food in their mouths. The experiment also examined whether visual changes alter the taste as well as the smell and texture of the food. The findings demonstrated that visual changes could present the smell and texture of the food following the modifications. This result was confirmed irrespective of the visibility near the mouth.
Kizashi Nakano, Monica Perusquía-Hernández, Naoya Isoyama, Hideaki Uchiyama, Kiyoshi Kiyokawa
ISMAR4
2023 A Two-layer Haptic Device for Presenting a Wide Range of Softness and Hardness Using a Pneumatic Balloon and a Mechanical Piston
abstract
Although a variety of haptic devices are used for virtual reality (VR) and augmented reality (AR) experiences, few can present a wide range of softness-hardness of the surface of the virtual objects. We propose a haptic device that can present a wide range of softness-hardness by using a two-layered structure consisting of a pneumatic balloon and a mechanical piston. Through a series of user studies, we confirmed that the prototype can present five levels of softness and three levels of hardness, and that the prototype device improves the VR experience in terms of realism, enjoyment, and comfort for virtual objects with a variety of softness/hardness.
Takuya Sasaki, Daiki Hagimori, Monica Perusquía-Hernández, Naoya Isoyama, Hideaki Uchiyama, Kiyoshi Kiyokawa, Yoshihiro Kuroda
RO-MAN5
2023 I'm Transforming! Effects of Visual Transitions to Change of Avatar on the Sense of Embodiment in AR
abstract
Virtual avatars are more and more often featured in Virtual Reality (VR) and Augmented Reality (AR) applications. When embodying a virtual avatar, one may desire to change of appearance over the course of the embodiment. However, switching suddenly from one appearance to another can break the continuity of the user experience and potentially impact the sense of embodiment (SoE), especially when the new appearance is very different. In this paper, we explore how applying smooth visual transitions at the moment of the change can help to maintain the SoE and benefit the general user experience. To address this, we implemented an AR system allowing users to embody a regular-shaped avatar that can be transformed into a muscular one through a visual effect. The avatar's transformation can be triggered either by the user through physical action (“active” transition), or automatically launched by the system (“passive” transition). We conducted a user study to evaluate the effects of these two types of transformations on the SoE by comparing them to control conditions where there was no visual feedback of the transformation. Our results show that changing the appearance of one's avatar with an active transition (with visual feedback), compared to a passive transition, helps to maintain the user's sense of agency, a component of the SoE. They also partially suggest that the Proteus effects experienced during the embodiment were enhanced by these transitions. Therefore, we conclude that visual effects controlled by the user when changing their avatar's appearance can benefit their experience by preserving the SoE and intensifying the Proteus effects.
Riku Otono, Adélaïde Genay, Monica Perusquía-Hernández, Naoya Isoyama, Hideaki Uchiyama, Martin Hachet, Anatole Lécuyer, Kiyoshi Kiyokawa
VR5
2022 Unsupervised Multi-view Multi-person 3D Pose Estimation Using Reprojection Error
Diógenes Wallis de França Silva, Joao Paulo Silva do Monte Lima, David Macedo, Cleber Zanchettin, Diego Thomas, Hideaki Uchiyama, Veronica Teichrieb
ICANN (3)6
2022 MOTSLAM: MOT-assisted monocular dynamic SLAM using single-view depth estimation
abstract
Visual SLAM systems targeting static scenes have been developed with satisfactory accuracy and robustness. Dynamic 3D object tracking has then become a significant capability in visual SLAM with the requirement of under-standing dynamic surroundings in various scenarios including autonomous driving, augmented and virtual reality. However, performing dynamic SLAM solely with monocular images remains a challenging problem due to the difficulty of asso-ciating dynamic features and estimating their positions. In this paper, we present MOTSLAM, a dynamic visual SLAM system with the monocular configuration that tracks both poses and bounding boxes of dynamic objects. MOTSLAM first performs multiple object tracking (MOT) with associated both 2D and 3D bounding box detection to create initial 3D objects. Then, neural-network-based monocular depth estimation is applied to fetch the depth of dynamic features. Finally, camera poses, object poses, and both static, as well as dynamic map points, are jointly optimized using a novel bundle adjustment. Our experiments on the KITTI dataset demonstrate that our system has reached best performance on both camera ego-motion and object tracking on monocular dynamic SLAM.
Hanwei Zhang 0003, Hideaki Uchiyama, Shintaro Ono, Hiroshi Kawasaki
IROS2
2022 An Object Synthesis Method to Enhance Visuo-Haptic Consistency
abstract
The sense of reality is enhanced by presenting appropriate haptic feedback when the user interacts with a virtual object in virtual reality (VR). To present appropriate feedback, we often use a real object that resembles the virtual one to manipulate. However, such a real object is not always available. The user may feel a visuohaptic inconsistency between real and virtual objects when their shapes are different. To alleviate such an inconsistency, we propose a novel object synthesis method that combines the shape of the real object that the user manipulates in reality and the shape of the virtual object which was to be presented to the user in VR. In other words, this synthesized object, a chimera object, is created by transforming the part of the virtual object that the user would touch into a shape that is similar to the corresponding part of the real object while maintaining the other parts of the virtual object intact. The expected haptic sensation from the appearance of our chimera object is more consistent with the one produced by the real object. Therefore, the visuo-haptic inconsistency is expected to alleviate in VR, compared to the original virtual object. In addition, we propose an interactive system to support the design of a chimera object for users. To investigate the effectiveness of our method, we conducted two user studies. The first experiment confirmed that our proposed chimera object helps enhance visuo-haptic consistency. The second experiment confirmed that our system was effective for chimera object creation by users with acceptable system usability.
Naoya Fukumoto, Naoya Isoyama, Hideaki Uchiyama, Nobuchika Sakata, Kiyoshi Kiyokawa
ISMAR3
2022 Recent advances in vision-based indoor navigation: A systematic literature review
Dawar Khan, Zhanglin Cheng, Hideaki Uchiyama, Sikandar Ali 0002, Muhammad Asshad, Kiyoshi Kiyokawa
Comput. Graph.3
2022 3D pedestrian localization using multiple cameras: a generalizable approach
Joao Paulo Silva do Monte Lima, Rafael Roberto, Lucas Silva Figueiredo, Francisco Simões, Diego Thomas, Hideaki Uchiyama, Veronica Teichrieb
Mach. Vis. Appl.6
2020 TetraTSDF: 3D Human Reconstruction From a Single Image With a Tetrahedral Outer Shell
abstract
Recovering the 3D shape of a person from its 2D appearance is ill-posed due to ambiguities. Nevertheless, with the help of convolutional neural networks (CNN) and prior knowledge on the 3D human body, it is possible to overcome such ambiguities to recover detailed 3D shapes of human bodies from single images. Current solutions, however, fail to reconstruct all the details of a person wearing loose clothes. This is because of either (a) huge memory requirement that cannot be maintained even on modern GPUs or (b) the compact 3D representation that cannot encode all the details. In this paper, we propose the tetrahedral outer shell volumetric truncated signed distance function (TetraTSDF) model for the human body, and its corresponding part connection network (PCN) for 3D human body shape regression. Our proposed model is compact, dense, accurate, and yet well suited for CNN-based regression task. Our proposed PCN allows us to learn the distribution of the TSDF in the tetrahedral volume from a single image in an end-to-end manner. Results show that our proposed method allows to reconstruct detailed shapes of humans wearing loose clothes from single RGB images.
Hayato Onizuka, Zehra Hayirci, Diego Thomas, Akihiro Sugimoto, Hideaki Uchiyama, Rin-Ichiro Taniguchi
CVPR5
2020 On-the-fly Extrinsic Calibration of Non-Overlapping in-Vehicle Cameras based on Visual SLAM under 90-degree Backing-up Parking
abstract
Calibration of relative poses between cameras is a challenging problem, known as extrinsic calibration, for nonoverlapping cameras that do not share the field of view. We propose a method for calibrating non-overlapping in-vehicle cameras placed at front, back, left and right positions by using visual SLAM(vSLAM). Our proposal is to calibrate the cameras during the motion of 90-degree backing-up parking on the fly, without using any dedicated calibration equipment. With this motion, the adjacent cameras are able to have the close field of view at different moments. The relative poses can be computed if the maps computed with vSLAM on each camera are merged by using the common structures. Therefore, we propose an efficient calibration framework with this feature. The proposed method is divided into three steps: map reconstruction with vSLAM on each camera, map merging for all the cameras, and extrinsic calibration. Especially, we propose to separately utilize the frames for vSLAM and the ones for the calibration so that the accuracy of vSLAM can be maximized for the calibration. In the evaluation, the calibration was performed in a practical environment to investigate the performance in comparison with the ground truth acquired by using a calibration equipment.
Kazuki Nishiguchi, Hideaki Uchiyama, Kazutaka Hayakawa, Jun Adachi, Diego Thomas, Atsushi Shimada 0001, Rin-Ichiro Taniguchi
IV2
2019 Mobile Photometric Stereo with Keypoint-Based SLAM for Dense 3D Reconstruction
abstract
The standard photometric stereo is a technique to densely reconstruct objects' surfaces using light variation under the assumption of a static camera with a moving light source. In this work, we use photometric stereo to reconstruct dense 3D scenes while moving the camera and the light altogether. In such non-static case, camera poses as well as correspondences between pixels of each frame to apply photometric stereo are required. ORB-SLAM is a technique that can be used to estimate camera poses. To retrieve correspondences, our idea is to start from a sparse 3D mesh obtained with ORB SLAM and then densify the mesh by a plane sweep method using a multi-view photometric consistency. By combining ORB-SLAM and photometric stereo, it is possible to reconstruct dense 3D scenes with a off-the-shelf smartphone and its embedded torchlight. Note that SLAM systems usually struggle with textureless object, which is effectively compensated by the photometric stereo in our method. Experiments are conducted to show that our proposed method gives better results than SLAM alone or COLMAP, especially for partially textureless surfaces.
Remy Maxence, Hideaki Uchiyama, Hiroshi Kawasaki, Diego Thomas, Vincent Nozick, Hideo Saito 0001
3DV2
2019 3D Positioning System Based on One-handed Thumb Interactions for 3D Annotation Placement
abstract
This paper presents a 3D positioning system based on one-handed thumb interactions for simple 3D annotation placement with a smart-phone. To place an annotation at a target point in the real environment, the 3D coordinate of the point is computed by interactively selecting the corresponding points in multiple views by users while performing SLAM. Generally, it is difficult for users to precisely select an intended pixel on the touchscreen. Therefore, we propose to compute the 3D coordinate from multiple observations with a robust estimator to have the tolerance to the inaccurate user inputs. In addition, we developed three pixel selection methods based on one-handed thumb interactions. A pixel is selected at the thumb position at a live view in FingAR, the position of a reticle marker at a live view in SnipAR, or that of a movable reticle marker at a freezed view in FreezAR. In the preliminary evaluation, we investigated the 3D positioning accuracy of each method.
So Tashiro, Hideaki Uchiyama, Diego Thomas, Rin-Ichiro Taniguchi
VR2
2019 Geometrical and statistical incremental semantic modeling on mobile devices
Rafael Alves Roberto, Joao Paulo Silva do Monte Lima, Hideaki Uchiyama, Veronica Teichrieb, Rin-Ichiro Taniguchi
Comput. Graph.3
2018 Transparent Random Dot Markers
abstract
This paper presents random dot markers (RDM) printed on transparent sheets as transparent fiducial markers. They are extremely unobstructive, and useful for developing novel user interfaces. However, the marker identification is required to be robust to observable back sides of the transparent sheets. To realize such markers, we propose a graph based framework for geometric feature based robust point matching for RDM. Instead of building one-to-one correspondences, we first build one-to-many correspondences using a 2D affinity matrix, and then globally optimize the matching assignment from the matrix. Especially, we incorporate pairwise relationship between neighboring points using local geometric descriptors into the matrix, and finally solve it with spectral matching. In the evaluation, we investigate the effectiveness of the global assignment from one-to-many correspondences, and finally show that our proposed method is enough robust to identifying overlapped markers.
Hideaki Uchiyama, Yuji Oyamada
ICPR1
2018 Two-step Transfer Learning for Semantic Plant Segmentation
Shunsuke Sakurai, Hideaki Uchiyama, Atsushi Shimada 0001, Daisaku Arita, Rin-Ichiro Taniguchi
ICPRAM2
2018 Live Structural Modeling Using RGB-D SLAM
abstract
This paper presents a method for localizing primitive shapes in a dense point cloud computed by the RGB-D SLAM system. To stably generate a shape map containing only primitive shapes, the primitive shape is incrementally modeled by fusing the shapes estimated at previous frames in the SLAM, so that an accurate shape can be finally generated. Specifically, the history of the fusing process is used to avoid the influence of error accumulation in the SLAM. The point cloud of the shape is then updated by fusing the points in all the previous frames into a single point cloud. In the experimental results, we show that metric primitive modeling in texture-less and unprepared environments can be achieved online.
Nicolas Olivier, Hideaki Uchiyama, Masashi Mishima, Diego Thomas, Rin-Ichiro Taniguchi, Rafael Alves Roberto, Joao Paulo Silva do Monte Lima, Veronica Teichrieb
ICRA2
2018 Deep Localization on Panoramic Images
abstract
Sensor pose estimation is an essential technology for various applications. For instance, it can be used not only to display immersive contents according user movements in Virtual Reality (VR) and but also to superimpose computer-generated objects onto images from a camera in Augmented Reality (AR). As a technical term definition, camera localization with respect to a pre-created map database is specifically referred to as image based localization, memory based localization, or camera relocalization.
Atsutoshi Hanasaki, Hideaki Uchiyama, Atsushi Shimada 0001, Rin-ichiro Taniquch
VR2
2018 Texture synthesis for stable planar tracking
abstract
We propose a texture synthesis method to enhance the trackability of a target planar object by embedding natural features into the object in the object design process. To transform an input object into an easy-to-track object in the design process, we extend an inpainting method for naturally embedding the features into the texture. First, a feature-less region in an input object is extracted based on feature distribution based segmentation. Then, the region is filled by using an inpainting method with a feature-rich region searched in an object database. By using context based region search, the inpainted region can be consistent in terms of the object context while improving the feature distribution.
Clément Glédel, Hideaki Uchiyama, Yuji Oyamada, Rin-Ichiro Taniguchi
VRST2
2018 Incremental Structural Modeling Based on Geometric and Statistical Analyses
abstract
Finding high-level semantic information from a point cloud is a challenging task, and it can be used in various applications. For instance, it is useful to compactly represent the scene structure and efficiently understand the scene context. This task is even more challenging when using a hand-held monocular visual SLAM system that outputs a noisy sparse point cloud. In order to tackle this issue, we propose an incremental primitive modeling method using both geometric and statistical analyses for such point cloud. The main idea is to select only reliably-modeled shapes by analyzing the geometric relationship between the point cloud and the estimated shapes. Besides that, a statistical evaluation is incorporated to filter wrongly-detected primitives in a noisy point cloud. As a result of this processing, our approach largely improved precision when compared with state of the art methods. We also show the impact of segmenting and representing a scene using primitives instead of a point cloud.
Rafael Alves Roberto, Joao Paulo Silva do Monte Lima, Hideaki Uchiyama, Clemens Arth, Veronica Teichrieb, Rin-Ichiro Taniguchi, Dieter Schmalstieg
WACV3
2017 Adaptive background model registration for moving cameras
Tsubasa Minematsu, Hideaki Uchiyama, Atsushi Shimada 0001, Hajime Nagahara, Rin-Ichiro Taniguchi
Pattern Recognit. Lett.2
2016 Real-Time Surface of Revolution Reconstruction on Dense SLAM
abstract
We present a fast and accurate method for reconstructing surfaces of revolution (SoR) on 3D data and its application to structural modeling of a cluttered scene in real-time. To estimate a SoR axis, we derive an approximately linear cost function for fast convergence. Also, we design a framework for reconstructing SoR on dense SLAM. In the experiment results, we show our method is accurate, robust to noise and runs in real-time.
Hideaki Uchiyama, Jean-Marie Normand, Guillaume Moreau, Hajime Nagahara, Rin-Ichiro Taniguchi
3DV2
2016 Design of a Low-false-positive Gesture for a Wearable Device
abstract
As smartwatches are becoming more widely used in society, gesture recognition, as an important aspect of interaction with smartwatches, is attracting attention. An accelerometer that is incorporated in a device is often used to recognize gestures. However, a gesture is often detected falsely when a similar pattern of action occurs in daily life. In this paper, we present a novel method of designing a new gesture that reduces false detection. We refer to such a gesture as a low-false-positive (LFP) gesture. The proposed method enables a gesture design system to suggest LFP motion gestures automatically. The user of the system can design LFP gestures more easily and quickly than what has been possible in previous work. Our method combines primitive gestures to create an LFP gesture. The combination of primitive gestures is recognized quickly and accurately by a random forest algorithm using our method. We experimentally demonstrate the good recognition performance of our method for a designed gesture with a high recognition rate and without false detection.
Ryo Kawahata, Atsushi Shimada 0001, Takayoshi Yamashita, Hideaki Uchiyama, Rin-Ichiro Taniguchi
ICPRAM4
2016 Depth-assisted rectification for real-time object detection and pose estimation
Joao Paulo Silva do Monte Lima, Francisco Simões, Hideaki Uchiyama, Veronica Teichrieb, Éric Marchand
Mach. Vis. Appl.3
2016 Pose Estimation for Augmented Reality: A Hands-On Survey
abstract
Augmented reality (AR) allows to seamlessly insert virtual objects in an image sequence. In order to accomplish this goal, it is important that synthetic elements are rendered and aligned in the scene in an accurate and visually acceptable way. The solution of this problem can be related to a pose estimation or, equivalently, a camera localization process. This paper aims at presenting a brief but almost self-contented introduction to the most important approaches dedicated to vision-based camera localization along with a survey of several extension proposed in the recent years. For most of the presented approaches, we also provide links to code of short examples. This should allow readers to easily bridge the gap between theoretical aspects and practical implementations.
Éric Marchand, Hideaki Uchiyama, Fabien Spindler
IEEE Trans. Vis. Comput. Graph.2
2015 Estimating Surface Normals with Depth Image Gradients for Fast and Accurate Registration
abstract
We present a fast registration framework with estimating surface normals from depth images. The key component in the framework is to utilize adjacent pixels and compute the normal at each pixel on a depth image by following three steps. First, image gradients on a depth image are computed with a 2D differential filtering. Next, two 3D gradient vectors are computed from horizontal and vertical depth image gradients. Finally, the normal vector is obtained from the cross product of the 3D gradient vectors. Since horizontal and vertical adjacent pixels at each pixel are considered composing a local 3D plane, the 3D gradient vectors are equivalent to tangent vectors of the plane. Compared with existing normal estimation based on fitting a plane to a point cloud, our depth image gradients based normal estimation is extremely faster because it needs only a few mathematical operations. We apply it to normal space sampling based 3D registration and validate the effectiveness of our registration framework by evaluating its accuracy and computational cost with a public dataset.
Yosuke Nakagawa, Hideaki Uchiyama, Hajime Nagahara, Rin-Ichiro Taniguchi
3DV2
2015 Adaptive search of background models for object detection in images taken by moving cameras
abstract
We propose a strategy of background subtraction for an image sequence captured by a moving camera. To adapt for camera motion, it is necessary to estimate the relation between consecutive frames in background subtraction. However, simple background subtraction using the relation between consecutive frames results in many false detections. We use re-projection error to handle this problem. The re-projection error has a low value in a background region. According to re-projection error, our method searches neighboring background models and tunes a threshold value for detection in order to reduce false detections. We evaluated the accuracy of detection of our method in experiments. Our method provided better detection than a method that does not search neighboring background models. Our method thus reduced the number of false detections.
Tsubasa Minematsu, Hideaki Uchiyama, Atsushi Shimada 0001, Hajime Nagahara, Rin-Ichiro Taniguchi
ICIP2
2015 Abecedary Tracking and Mapping: A Toolkit for Tracking Competitions
abstract
This paper introduces a toolkit with camera calibration, monocular visual Simultaneous Localization and Mapping (vSLAM) and registration with a calibration marker. With the toolkit, users can perform the whole procedure of the ISMAR on-site tracking competition in 2015. Since the source code is designed to be well-structured and highly-readable, users can easily install and modify the toolkit. By providing the toolkit, we encourage beginners to learn tracking techniques and to participate in the competition.
Hideaki Uchiyama, Takafumi Taketomi, Sei Ikeda, Joao Paulo Silva do Monte Lima
ISMAR1
2012 Texture-less planar object detection and pose estimation using Depth-Assisted Rectification of Contours
abstract
This paper presents a method named Depth-Assisted Rectification of Contours (DARC) for detection and pose estimation of texture-less planar objects using RGB-D cameras. It consists in matching contours extracted from the current image to previously acquired template contours. In order to achieve invariance to rotation, scale and perspective distortions, a rectified representation of the contours is obtained using the available depth information. DARC requires only a single RGB-D image of the planar objects in order to estimate their pose, opposed to some existing approaches that need to capture a number of views of the target object. It also does not require to generate warped versions of the templates, which is commonly needed by existing object detection techniques. It is shown that the DARC method runs in real-time and its detection and pose estimation quality are suitable for augmented reality applications.
Joao Paulo Silva do Monte Lima, Hideaki Uchiyama, Veronica Teichrieb, Éric Marchand
ISMAR2
2011 Toward augmenting everything: Detecting and tracking geometrical features on planar objects
abstract
This paper presents an approach for detecting and tracking various types of planar objects with geometrical features. We combine traditional keypoint detectors with Locally Likely Arrangement Hashing (LLAH) [21] for geometrical feature based keypoint matching. Because the stability of keypoint extraction affects the accuracy of the keypoint matching, we set the criteria of keypoint selection on keypoint response and the distance between keypoints. In order to produce robustness to scale changes, we build a non-uniform image pyramid according to keypoint distribution at each scale. In the experiments, we evaluate the applicability of traditional keypoint detectors with LLAH for the detection. We also compare our approach with SURF and finally demonstrate that it is possible to detect and track different types of textures including colorful pictures, binary fiducial markers and handwritings.
Hideaki Uchiyama, Éric Marchand
ISMAR1
2011 Deformable random dot markers
abstract
We extend planar fiducial markers using random dots [8] to nonrigidly deformable markers. Because the recognition and tracking of random dot markers are based on keypoint matching, we can estimate the deformation of the markers with nonrigid surface detection from keypoint correspondences. First, the initial pose of the markers is computed from a homography with RANSAC as a planar detection. Second, deformations are estimated from the minimization of a cost function for deformable surface fitting. We show augmentation results of 2D surface deformation recovery with several markers.
Hideaki Uchiyama, Éric Marchand
ISMAR1
2011 onNote: playing printed music scores as a musical instrument
abstract
This paper presents a novel musical performance system named onNote that directly utilizes printed music scores as a musical instrument. This system can make users believe that sound is indeed embedded on the music notes in the scores. The users can play music simply by placing, moving and touching the scores under a desk lamp equipped with a camera and a small projector. By varying the movement, the users can control the playing sound and the tempo of the music. To develop this system, we propose an image processing based framework for retrieving music from a music database by capturing printed music scores. From a captured image, we identify the scores by matching them with the reference music scores, and compute the position and pose of the scores with respect to the camera. By using this framework, we can develop novel types of musical interactions.
Yusuke Yamamoto, Hideaki Uchiyama, Yasuaki Kakehi
UIST2
2011 Random dot markers
abstract
This paper presents a novel approach for detecting and tracking markers with randomly scattered dots for augmented reality applications. Compared with traditional markers with square pattern, our random dot markers have several significant advantages for flexible marker design, robustness against occlusion and user interaction. The retrieval and tracking of these markers are based on geometric feature based keypoint matching and tracking. We experimentally demonstrate that the discriminative ability of forty random dots per marker is applicable for retrieving up to one thousand markers.
Hideaki Uchiyama, Hideo Saito 0001
VR1
2011 Random dot markers
abstract
We introduce a novel type of markers with randomly scattered dots for augmented reality applications. Compared with traditional square markers, our markers have several significant advantages for flexible marker design, robustness against occlusion and user interaction. Our markers do not need to have a black frame, and their shape is not limited to square because the retrieval and tracking of the markers are based on geometric feature based keypoint matching. In our demonstration, we show real-time simultaneous retrieval and tracking of the markers on a laptop.
Hideaki Uchiyama, Hideo Saito 0001
VR1
2010 An Augmented Reality Setup with an Omnidirectional Camera Based on Multiple Object Detection
abstract
We propose a novel augmented reality (AR) setup with an omni directional camera on a table top display. The table acts as a mirror on which real playing cards appear augmented with virtual elements. The omni directional camera captures and recognizes its surrounding based on a feature based image retrieval approach which achieves fast and scalable registration. It allows our system to superimpose virtual visual effects to the omni directional camera image. In our AR card game, users sit around a table top display and show a card to the other players. The system recognizes it and augments it with virtual elements in the omni directional image acting as a mirror. While playing the game, the users can interact with each other directly and through the display. Our setup is a new, simple, and natural approach to augmented reality. It opens new doors to traditional card games.
Tomoki Hayashi, Hideaki Uchiyama, Julien Pilet, Hideo Saito 0001
ICPR2
2010 Foldable augmented maps
abstract
This paper presents folded surface detection and tracking for augmented maps. For the detection, plane detection is iteratively applied to 2D correspondences between an input image and a reference plane because the folded surface is composed of multiple planes. In order to compute the exact folding line from the detected planes, the intersection line of the planes is computed from their positional relationship. After the detection is done, each plane is individually tracked by frame-by-frame descriptor update. For a natural augmentation on the folded surface, we overlay virtual geographic data on each detected plane. The user can interact with the geographic data by finger pointing because the finger tip of the user is also detected during the tracking. As scenario of use, some interactions on the folded surface are introduced. Experimental results show the accuracy and performance of folded surface detection for evaluating the effectiveness of our approach.
Sandy Martedi, Hideaki Uchiyama, Guillermo Enriquez, Hideo Saito 0001, Tsutomu Miyashita, Takenori Hara
ISMAR2
2010 Foldable augmented maps
abstract
This demonstration presents folded surface detection and tracking for augmented maps. We model the folded surface as multiple planes. To detect a folded surface, plane detection is iteratively applied to 2D correspondences between an input image and a reference plane. In order to compute the exact folding line from the detected planes, the intersection line of the planes is computed from their positional relationship. After the detection is done, each plane is individually tracked by frame-by-frame descriptor update. For a natural augmentation on the folded surface, we overlay virtual geographic data on each detected plane.
Sandy Martedi, Hideaki Uchiyama, Guillermo Enriquez, Hideo Saito 0001, Tsutomu Miyashita, Takenori Hara
ISMAR2
2010 An intermediate report of TrakMark WG - international voluntary activities on establishing benchmark test schemes for AR/MR geometric registration and tracking methods
abstract
In the study of AR/MR field, tracking and geometric registration methods are very important topics that are actively discussed. Especially, the study on tracking is flourishing and many algorithms are being proposed every year. With this trend in mind, we, the TrakMark WG, had proposed benchmark test schemes for geometric registration and tracking in AR/MR at ISMAR 2009 [1]. This paper is an intermediate report of the TrakMark WG, which describes its activities and the first proposal on benchmarking image sequences.
Fumihisa Shibata, Sei Ikeda, Takeshi Kurata, Hideaki Uchiyama
ISMAR4
2010 Clickable augmented documents
abstract
This paper presents an Augmented Reality (AR) system for physical text documents that enable users to click a document. In the system, we track the relative pose between a camera and a document to overlay some virtual contents on the document continuously. In addition, we compute the trajectory of a fingertip based on skin color detection for clicking interaction. By merging a document tracking and an interaction technique, we have developed a novel tangible document system. As an application, we develop an AR dictionary system that overlays the meaning and explanation of words by clicking on a document. In the experiment part, we present the accuracy of the clicking interaction and the robustness of our document tracking method against the occlusion.
Sandy Martedi, Hideaki Uchiyama, Hideo Saito 0001
MMSP2
2009 Augmenting text document by on-line learning of local arrangement of keypoints
abstract
We propose a technique for text document tracking over a large range of viewpoints. Since the popular SIFT or SURF descriptors typically fail on such documents, our method considers instead local arrangement of keypoints. We extends locally likely arrangement hashing (LLAH), which is limited to fronto-parallel images: We handle a large range of viewpoints by learning the behavior of keypoint patterns when the camera viewpoint changes. Our method starts tracking a document from a nearly frontal view. Then, it undergoes motion, and new configurations of keypoints appear. The database is incrementally updated to reflect these new observations, allowing the system to detect the document under the new viewpoint. We demonstrate the performance and robustness of our method by comparing it with the original LLAH.
Hideaki Uchiyama, Hideo Saito 0001
ISMAR1
2009 Rotated Image Based Photomosaic Using Combination of Principal Component Hashing
Hideaki Uchiyama, Hideo Saito 0001
PSIVT1
2006 Position Estimation of Solid Balls from Handy Camera for Pool Supporting System
Hideaki Uchiyama, Hideo Saito 0001
PSIVT1