Yoshihiro Watanabe

dblp:85/6634 · DBLP profile ↗
← Back
43ranked-venue papers
9as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 34 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 2 since 2021Systems, architecture and hardware · 8 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 Neural Inverse Rendering for High-Accuracy 3D Measurement of Moving Objects with Fewer Phase-Shifting Patterns
Yuki Urakawa, Yoshihiro Watanabe
ICCV2
2025 Perceptually-Aligned Dynamic Facial Projection Mapping by High-Speed Face-Tracking Method and Lens-Shift Co-Axial Setup
abstract
Dynamic Facial Projection Mapping (DFPM) overlays computer-generated images onto human faces to create immersive experiences that have been used in the makeup and entertainment industries. In this study, we propose two concepts to reduce the misalignment artifacts between projected images and target faces, which is a persistent challenge for DFPM. Our first concept is a high-speed face-tracking method that exploits temporal information. We first introduce a cropped-area-limited inter/extrapolation-based face detection framework, which allows parallel execution with facial landmark detection. We then propose a novel hybrid facial landmark detection method that combines fast Ensemble of Regression Trees (ERT)-based detections and an auxiliary detection. ERT-based detections rapidly produce results in 0.107 ms using temporal information with the support of auxiliary detection to recover from detection errors. To train the facial landmark detection method, we propose an innovative method for simulating high-frame-rate video annotations to address the lack of publicly available high-frame-rate annotated datasets. Our second concept is a lens-shift co-axial projector-camera setup that maintains a high optical alignment with only a 1.274-pixel error between 1 m and 2 m depth. This setup reduces misalignment by applying the same optical designs to the projector and camera without causing large misalignment as in conventional methods. Based on these concepts, we developed a novel high-speed DFPM system that achieves nearly perfect alignment with human visual perception.
Hao-Lun Peng, Kengo Sato, Soran Nakagawa, Yoshihiro Watanabe
IEEE Trans. Vis. Comput. Graph.4
2024 Improving Real-Time Near-Infrared Face Alignment With a Paired VIS-NIR Dataset and Data Augmentation Through Image-to-Image Translation
abstract
Real-time near-infrared (NIR) face alignment holds significant importance across various domains, such as security, healthcare, and augmented reality. However, existing face alignment techniques tailored for visible-light (VIS) encounter a decline in accuracy when applied in NIR settings. This decline stems from the domain discrepancy between VIS and NIR facial domains and the absence of meticulously annotated NIR facial data. To address this issue, we introduce a system and strategy for gathering paired VIS-NIR facial images and meticulously annotating precise landmarks. Our system facilitates streamlined dataset preparation by utilizing automatic annotation transfer from VIS images to their corresponding NIR counterparts. Following our devised approach, we constructed an inaugural dataset comprising high-frame-rate paired VIS-NIR facial images with landmark annotations. Additionally, to enhance the diversity of facial data, we augment our dataset through VIS-NIR image-to-image (img2img) translation using publicly available facial landmark datasets. Through the retraining of face alignment models and subsequent evaluations, our findings demonstrate a noteworthy enhancement in the accuracy of face alignment under NIR conditions using our dataset. Furthermore, the augmented dataset exhibits refined accuracy, particularly notable in the case of different individuals’ facial features.
Langning Miao, Ryo Kakimoto, Kaoru Ohishi, Yoshihiro Watanabe
ICIP4
2024 Projection Mapping with a Brightly Lit Surrounding Using a Mixed Light Field Approach
abstract
Projection mapping (PM) exhibits suboptimal performance in well-lit environments because of the interference caused by ambient light. This interference degrades the contrast of the projected images. Consequently, conventional methodologies restrict the application of PM to dimly lit settings, leading to an unnatural visual experience, as only the PM target is prominently illuminated. To overcome these limitations, we introduce an innovative approach that leverages a mixed light field, blending traditional PM with ray-controllable ambient lighting. This methodological combination, despite its simplicity, is effective because it ensures that the projector exclusively illuminates the PM target, preserving the optimal contrast. Precise control of ambient light rays is essential to prevent them from illuminating the PM target while adequately illuminating the surrounding environment. Furthermore, we propose the integration of a kaleidoscopic array with integral photography to generate dense light fields for ray-controllable ambient lighting. Additionally, we present an efficient binary-search-based calibration method tailored to this intricate optical system. Our optical simulations and the developed system collectively validate the effectiveness of our approach. Our results show that PM targets and ordinary objects coexist naturally in environments that are brightly lit as a result of our method, enhancing the overall visual experience.
Masahiko Yasui, Ryota Iwataki, Masatoshi Ishikawa, Yoshihiro Watanabe
IEEE Trans. Vis. Comput. Graph.4
2023 High-Frame-Rate Projection with Thousands of Frames Per Second Based on the Multi-Bit Superimposition Method
abstract
The growing need for high-frame-rate projectors in the fields of dynamic projection mapping (DPM) and three-dimensional (3D) displays has increased. Conventional methods allow for an increase in the frame rate to as much as 2,841 frames per second (fps) for 8-bit image projection, using digital light processing (DLP) technology when the minimum digital mirror device (DMD) control time is $44 \mu \mathrm{s}$. However, this rate needs to be further augmented to suit specific applications. In this study, we developed a novel high-frame-rate projection method, which divides the bit depth of an image among multiple projectors and simultaneously projects them in synchronization. The simultaneously projected bit images are superimposed such that a high-bit-depth image is generated within a reduced single-frame duration. Additionally, we devised an optimization process to determine the system parameters necessary for attaining maximum brightness. We constructed a prototype system utilizing two high-frame-rate projectors and validated the feasibility of using our system to project 8-bit images at a rate of 5,600 fps. Furthermore, the quality assessment of our projected image exhibited superior performance in comparison to a dithered image.
Soran Nakagawa, Yoshihiro Watanabe
ISMAR2
2023 Studying User Perceptible Misalignment in Simulated Dynamic Facial Projection Mapping
abstract
High-speed dynamic facial projection mapping (DFPM) is an advanced technology that aims to create perceptual changes in facial appearance by overlapping images based on facial position and shape. Compared to traditional monitor-based augmented reality systems, DFPM offers a higher level of immersion because users can directly observe digital content on their faces. However, DFPM suffers from misalignment issues owing to a slight temporal delay from sensing to projection, which reduces the level of immersion. To the best of our knowledge, no previous study has established the necessary latency requirements to avoid perceptible misalignment and achieve an immersive experience. Furthermore, conventional DFPM works followed latency requirements that were not reported for the DFPM scenario. Therefore, this study measured the latency that provided a just-noticeable difference (JND) in DFPM under different facial motion conditions, using the weighted up-down two-alternative forced-choice method. The results showed that user-perceptible misalignment was influenced by facial motion types and their velocities. Additionally, it was found that an average latency of 3.87 ms was necessary to avoid perceptible misalignment in the DFPM system when the translation speed was 0.5 m/s, which contradicts the commonly held belief regarding the required latency threshold.
Hao-Lun Peng, Shin'ya Nishida, Yoshihiro Watanabe
ISMAR3
2022 Dynamic Projection Mapping for Robust Sphere Posture Tracking Using Uniform/Biased Circumferential Markers
abstract
In spatial augmented reality, a widely dynamic projection mapping system has been developed as a novel approach to graphics presentation for widely moving objects in dynamic situations. However, this method necessitates a novel tracking marker design that is resistant to random and complex occlusion and out-of-focus blurring, which conventional markers have not achieved. This article presents a uniform circumferential marker that becomes an ellipse in perspective projection and expresses geometric information. It can track the relative posture of a dynamically moving sphere with high speed, high accuracy, and robustness owing to continuous contour lines, thereby supporting both wide-range movement in the depth direction and human interaction. Moreover, a biased circumferential marker is proposed to embed unique coding, where the absolute posture is decoded with a novel recognition algorithm. Moreover, rough initialization using the geometry of multiple ellipses is proposed for both markers to start the automatic and precise tracking. Real-time rotation visualization onto the surface of a moving sphere is made possible with the high-speed, widely dynamic projection mapping system. The tracking performance is demonstrated to exhibit sufficient basic tracking performance as well as robustness against blurring and occlusion compared to conventional dot-based markers.
Yuri Mikawa, Tomohiro Sueishi, Yoshihiro Watanabe, Masatoshi Ishikawa
IEEE Trans. Vis. Comput. Graph.3
2022 Dynamic Multi-projection Mapping Based on Parallel Intensity Control
abstract
Projection mapping using multiple projectors is promising for spatial augmented reality; however, it is difficult to apply it to dynamic scenes. This is because the conventional method decides all pixel intensities of multiple images simultaneously based on the global optimization method, and it is hard to reduce the latency from motion to projection. To mitigate this, we propose a novel method of controlling the intensity based on a pixel-parallel calculation for each projector in real-time with low latency. This parallel calculation leverages the insight that the projected pixels from different projectors in overlapping areas can be approximated independently if the pixel is sufficiently small relative to the surface structure. Additionally, our pixel-parallel calculation method allows a distributed system configuration, such that the number of projectors can be increased to form a network for high scalability. We demonstrate a seamless mapping into dynamic scenes at 360 fps with a 9.5-ms latency using ten cameras and four projectors.
Takashi Nomoto, Wanlong Li, Hao-Lun Peng, Yoshihiro Watanabe
IEEE Trans. Vis. Comput. Graph.4
2021 Feature-aided Bundle Adjustment Learning Framework for Self-supervised Monocular Visual Odometry
abstract
Bundle adjustment refines scene geometry and relative camera poses simultaneously via reprojection error, computed by a set of images from different viewpoints, which is the gold standard for visual odometry. However, deep learning methods have not been well exploited within this area of study. This paper introduces a self-supervised learning framework for monocular visual odometry, inside which depth maps, relative camera poses, and dense feature maps (with the same resolution as images) are estimated and used for photometric, geometric, and feature-metric losses in bundle adjustment. In this manner, we consider that the learning of neural networks can be geometrically constrained by multi-view geometry. Furthermore, bundle adjustment is only required during the training time, allowing the networks to benefit from bundle adjustment without any additional computation burden during the inference time. To stabilize the training process, we apply a two-stage strategy that yields promising results. Finally, we carefully select the neural network architectures to ensure efficiency, and experimental results demonstrate the success of our proposed approach in terms of visual odometry accuracy and high speed.
Weijun Mai, Yoshihiro Watanabe
IROS2
2021 Realistic 3D Swept-Volume Display with Hidden-Surface Removal Using Physical Materials
abstract
Conventional swept-volume displays can provide accurate physical cues for depth perception. However, the corresponding texture reproduction does not have high quality because such displays employ high-speed projectors with low bit-depth and low resolution. In this study, to address the limitation of swept-volume displays while retaining their advantages, a novel swept-volume three-dimensional (3D) display is proposed by incorporating physical materials as screens. Physical materials such as wool, felt, and so on are directly used for reproducing textures on a displayed 3D surface. Furthermore, we introduce the adaptive pattern generation based on real-time viewpoint tracking to perform the hidden-surface removal. Our algorithm leverages the ray-tracing concept and can run at high speed on GPU.
Ray Asahina, Takashi Nomoto, Takatoshi Yoshida, Yoshihiro Watanabe
VR4
2020 Projection Mapping System To A Widely Dynamic Sphere With Circumferential Markers
abstract
Image projection on spheres and their surroundings have been researched for all-around display and motion visualization. However, wide-range projection onto a dynamic sphere can suffer from tracking errors and projection latency. We propose a high-speed projection mapping system for dynamic spheres, and circumferential markers for sphere posture estimation. Marker detection based on ellipse appearance can easily start tracking, and is robust against interactive occlusion; therefore, the system enables projection mapping and rotational visualization for a dynamically moving sphere. We experimentally confirmed sufficient pose estimation accuracy of the circumferential markers compared with initialized dot markers, and sufficient tracking ability against initialization, occlusion, and depth-direction movement. Demonstrations showed accurate, rotation-visualized projection onto a dynamic sphere, which is useful for sports practice applications.
Yuri Mikawa, Tomohiro Sueishi, Yoshihiro Watanabe, Masatoshi Ishikawa
ICME3
2018 Portable Lumipen: Dynamic SAR in Your Hand
abstract
Based on the concept of spatial augmented reality (SAR), a number of applications have been developed by extending the type of target objects to include dynamically moving objects; however, the systems themselves still remain fixed. In this paper, we propose a new paradigm for dynamic SAR implemented with a mobile device and a practical architecture using a 3D-stacked vision chip and a high-speed optical gaze controller. We evaluated the prototype system that we developed, called “Portable Lumipen”, and confirmed that image capturing at 1000 fps, rapid processing and a 3 ms response time for projection were achieved on a mobile device. We also propose potential applications using this portable dynamic SAR system and demonstrated miniaturization of a high-speed visual feedback system.
Leo Miyashita, Tomohiro Yamazaki, Kenji Uehara, Yoshihiro Watanabe, Masatoshi Ishikawa
ICME4
2018 Multi-pattern Embedded Phase Shifting Using a High-Speed Projector for Fast and Accurate Dynamic 3D Measurement
abstract
The goal of this study is to achieve high-speed and high-accuracy 3D measurement of a moving object using phase shifting. Phase shifting is a multi-shot method that uses structured light and enables high-accuracy and highresolution measurement with few projections. However, this method requires projecting a multi-tone pattern, which causes the projection speed of the projector to become a bottleneck. This in turn makes performing high-speed measurements difficult. Furthermore, because phase shifting is a multi-shot method, reduced measurement accuracy for moving objects becomes another concern. In this study, we overcome these problems by using a newly developed highspeed projector. First, this projector enables high-speed 3D measurement, as it can project 8-bit tone images at frame rates of 1000fps. Second, proposed method improves the measurement accuracy of moving objects by enabling phase unwrapping, while minimizing the number of projected patterns through pattern embedding. Through experiments, we demonstrate that, by embedding a Gray-code pattern within an interval of 100ìs, we can measure a object moving at a speed of approximately 30-50cm/s with a frame rate of 500fps and an average error of less than 1mm.
Michika Maruyama, Satoshi Tabata, Yoshihiro Watanabe
WACV3
2018 MIDAS projection: markerless and modelless dynamic projection mapping for material representation
abstract
The visual appearance of an object can be disguised by projecting virtual shading as if overwriting the material. However, conventional projection-mapping methods depend on markers on a target or a model of the target shape, which limits the types of targets and the visual quality. In this paper, we focus on the fact that the shading of a virtual material in a virtual scene is mainly characterized by surface normals of the target, and we attempt to realize markerless and modelless projection mapping for material representation. In order to deal with various targets, including static, dynamic, rigid, soft, and fluid objects, without any interference with visible light, we measure surface normals in the infrared region in real time and project material shading with a novel high-speed texturing algorithm in screen space. Our system achieved 500-fps high-speed projection mapping of a uniform material and a tileable-textured material with millisecond-order latency, and it realized dynamic and flexible material representation for unknown objects. We also demonstrated advanced applications and showed the expressive shading performance of our technique.
Leo Miyashita, Yoshihiro Watanabe, Masatoshi Ishikawa
ACM Trans. Graph.2
2017 Extended Dot Cluster Marker for High-speed 3D Tracking in Dynamic Projection Mapping
abstract
The technique of Projection Mapping, which is useful for merging real-world geometry with an augmented appearance, is a promising core technology for augmented reality (AR). In recent years, dynamically changing environments, mainly a consequence of the growing demand for interactive user experiences, have contributed to a new style of AR applications. However, performance levels of current systems for realizing 3D effects, in terms of the tracking speed and projection ability, are insufficient to meet these demands. In this paper, we present a high-speed, occlusion-robust marker-based 3D tracking method achieved by only using a monocular monochrome image. The objective of our research is to develop an automatic marker design method for any 3D shape and an effective framework for stabilizing tracking at high throughput by extending the latest promising work based on a deformable dot cluster marker [46]. Furthermore, this tracking method was used in combination with a high-speed projector, both of which can achieve high throughput and low latency, on the order of milliseconds. This enabled us to realize a high-quality computational display capable of representing the material appearance of a dynamically moving target. The demonstration showed that the effect of a dynamically changing appearance with nearly imperceptible latency drastically enriches the sense of immersion in the recognition of augmented materials with the naked eye.
Yoshihiro Watanabe, Toshiyuki Kato, Masatoshi Ishikawa
ISMAR1
2017 Rapid blending of closed curves based on curvature flow
Masahiro Hirano, Yoshihiro Watanabe, Masatoshi Ishikawa
Comput. Aided Geom. Des.2
2017 Dynamic Projection Mapping onto Deforming Non-Rigid Surface Using Deformable Dot Cluster Marker
abstract
Dynamic projection mapping for moving objects has attracted much attention in recent years. However, conventional approaches have faced some issues, such as the target objects being limited to rigid objects, and the limited moving speed of the targets. In this paper, we focus on dynamic projection mapping onto rapidly deforming non-rigid surfaces with a speed sufficiently high that a human does not perceive any misalignment between the target object and the projected images. In order to achieve such projection mapping, we need a high-speed technique for tracking non-rigid surfaces, which is still a challenging problem in the field of computer vision. We propose the Deformable Dot Cluster Marker (DDCM), a novel fiducial marker for high-speed tracking of non-rigid surfaces using a high-frame-rate camera. The DDCM has three performance advantages. First, it can be detected even when it is strongly deformed. Second, it realizes robust tracking even in the presence of external and self occlusions. Third, it allows millisecond-order computational speed. Using DDCM and a high-speed projector, we realized dynamic projection mapping onto a deformed sheet of paper and a T-shirt with a speed sufficiently high that the projected images appeared to be printed on the objects.
Gaku Narita, Yoshihiro Watanabe, Masatoshi Ishikawa
IEEE Trans. Vis. Comput. Graph.2
2016 Occlusion-robust 3D sensing using aerial imaging
abstract
Conventional active 3D sensing systems do not work well when other objects get between the measurement target and the measurement equipment, occluding the line of sight. In this paper, we propose an active 3D sensing method that solves this occlusion problem by using a light field created using aerial imaging. In this light field, aerial luminous spots can be formed by focusing rays of light from multiple directions. Towards the occlusion problem, this configuration is effective, because even if some of the rays are occluded, the rays of other directions keep the spots. Our results showed that this method was able to measure the position and inclination of a target by using an aerial image of a single point light source and was robust against occlusions. In addition, we confirmed that multiple point light sources also worked well.
Masahiko Yasui, Yoshihiro Watanabe, Masatoshi Ishikawa
ICCP2
2016 Phyxel: Realistic Display of Shape and Appearance using Physical Objects with High-speed Pixelated Lighting
abstract
A computer display that is sufficiently realistic such that the difference between a presented image and a real object cannot be discerned is in high demand in a wide range of fields, such as entertainment, digital signage, and design industry. To achieve such a level of reality, it is essential to reproduce the three-dimensional (3D) shape and material appearances simultaneously; however, to date, developing a display that can satisfy both conditions has been difficult. To address this problem, we propose a system that places physical elements at desired locations to create a visual image that is perceivable by the naked eye. This configuration can be realized by exploiting characteristics of human visual perception. Humans perceive light modulation as perfectly steady light if the modulation rate is sufficiently high. Therefore, if high-speed spatially varying illumination is projected to the actuated physical elements possessing various appearances at the desired timing, a realistic visual image that can be transformed dynamically by simply modifying the lighting pattern can be obtained. We call the proposed display technology Phyxel. This paper describes the proposed configuration and required performance for Phyxel. We also demonstrate three applications: dynamic stop motion, a layered 3D display, and shape mixture.
Takatoshi Yoshida, Yoshihiro Watanabe, Masatoshi Ishikawa
UIST2
2016 ZoeMatrope: a system for physical material design
abstract
Reality is the most realistic representation. We introduce a material display called ZoeMatrope that can reproduce a variety of materials with high resolution, dynamic range and light field reproducibility by using compositing and animation principles used in a zoetrope and a thaumatrope. With ZoeMatrope, the quality of the material is equivalent to that of real objects and the range of expressible materials is diversified by overlaying a set of base materials in a linear combination. ZoeMatrope is also able to express spatially-varying materials, and even augmented materials such as materials with an alpha channel. In this paper, we propose a method for selecting the optimal material set and determining the weights of the linear combination to reproduce a wide range of target materials properly. We also demonstrate the effectiveness of this approach with the developed system and show the results for various materials.
Leo Miyashita, Kota Ishihara, Yoshihiro Watanabe, Masatoshi Ishikawa
ACM Trans. Graph.3
2015 Development of fast-response master-slave system using high-speed non-contact 3D sensing and high-speed robot hand
abstract
In this paper we focus on master-slave robot hand systems that can realize non-contact sensing and intuitive mapping between human hand motion and robot hand motion. Such a master-slave robot hand system can be effective from a viewpoint of usability. However, conventional systems are not able to adapt to dynamically changing environments because they have high latency from input to output. Therefore, we developed a fast-response master-slave robot hand system using a high-speed vision system and a high-speed robot hand. The latency of the proposed system is so small that humans cannot recognize it. The motion of a human hand is obtained with high-speed non-contact 3D sensing, and this motion is mapped to a high-speed robot hand, while taking account of structural differences between the human hand and the robot hand. We confirmed the effectiveness of our proposed system through experiments.
Yugo Katsuki, Yuji Yamakawa, Yoshihiro Watanabe, Masatoshi Ishikawa
IROS3
2015 High-speed image rotator for blur-canceling roll camera
abstract
We developed an optical high-speed image rotation controller and realized a high-speed roll camera that is able to cancel the rotational motion blur of a rotating target. This system is composed of a hollow motor, a Dove prism, and a high-speed camera and controls optical image rotation according to the target rotation by using high-speed image processing. This so-called optical lever formed of the Dove prism worked effectively also for a high-speed rotating target, and our prototype system shows that rotational motion blur was suppressed to 0.125 [°] at 1420 [r/min].
Leo Miyashita, Yoshihiro Watanabe, Masatoshi Ishikawa
IROS2
2015 High-speed 3D sensing with three-view geometry using a segmented pattern
abstract
High-speed vision technology in which not just image capturing and recording but also image processing are executed simultaneously at high frame rates, exceeding video rates (30 Hz), has recently been considered an important technology for various applications, such as robotics, and man-machine interfaces. However, the image processing performed in conventional high-speed vision systems is mainly based on two-dimensional pattern recognition. In order to extend the possibilities of this technology, here we focus on real-time three-dimensional sensing at the speeds achievable by high-speed vision systems. Although a related approach for high-speed 3D sensing can achieve a frame rate of over 200 fps, there are disadvantages, including the need for multiple captured frames during sensing, a limited measurement range and low resolution. Our proposed real-time 3D sensing system consists of a projector and two cameras. By projecting a well-designed segmented pattern and using three-viewpoint epipolar constraints, the proposed system can obtain 3D points at high speed. The developed system robustly obtained a 3D shape at 450-500 fps in real time.
Satoshi Tabata, Shohei Noguchi, Yoshihiro Watanabe, Masatoshi Ishikawa
IROS3
2015 Dynamic projection mapping onto a deformable object with occlusion based on high-speed tracking of dot marker array
abstract
In recent years, projection mapping has attracted much attention in a variety of fields. Generally, however, the objects in projection mapping are limited to rigid and static or quasistatic objects. Dynamic projection mapping onto a deformable object could remarkably expand the possibilities. In order to achieve such a projection mapping, it is necessary to recognize the deformation of the object even when it is occluded. However, it is still a challenging problem to achieve this task in real-time with low latency. In this paper, we propose an efficient, high-speed tracking method utilizing high-frame-rate imaging. Our method is able to track an array of dot markers arranged on a deformable object even when there is external occlusion caused by the user interaction and self-occlusion caused by the deformation of the object itself. Additionally, our method can be applied to a stretchable object. Dynamic projection mapping with our method showed robust and consistent display onto a sheet of paper and cloth with a tracking performance of about 0.2 ms per frame, with the result that the projected pattern appeared to be printed on the deformable object.
Gaku Narita, Yoshihiro Watanabe, Masatoshi Ishikawa
VRST2
2015 3D motion sensing of any object without prior knowledge
abstract
We propose a novel three-dimensional motion sensing method using lasers. Recently, object motion information is being used in various applications, and the types of targets that can be sensed continue to diversify. Nevertheless, conventional motion sensing systems have low universality because they require some devices to be mounted on the target, such as accelerometers and gyro sensors, or because they are based on cameras, which limits the types of targets that can be detected. Our method solves this problem and enables noncontact, high-speed, deterministic measurement of the velocity of a moving target without any prior knowledge about the target shape and texture, and can be applied to any unconstrained, unspecified target. These distinctive features are achieved by using a system consisting of a laser range finder, a laser Doppler velocimeter, and a beam controller, in addition to a robust 3D motion calculation method. The motion of the target is recovered from fragmentary physical information, such as the distance and speed of the target at the laser irradiation points. From the acquired laser information, our method can provide a numerically stable solution based on the generalized weighted Tikhonov regularization. Using this technique and a prototype system that we developed, we also demonstrated a number of applications, including motion capture, video game control, and 3D shape integration with everyday objects.
Leo Miyashita, Ryota Yonezawa, Yoshihiro Watanabe, Masatoshi Ishikawa
ACM Trans. Graph.3
2014 Rapid SVBRDF Measurement by Algebraic Solution Based on Adaptive Illumination
abstract
In this paper, we propose an algebraic solution for rapid SVBRDF measurement. The algebraic approach requires only a few reflectance samples to obtain the parameters described by the physically based Cook - Torrance model. This solution, however, also involves constraints concerning light and the normal direction in the acquisition process. To meet these constraints, we developed a system that changes the illumination according to the target 3D shape at high speed. As a result, the proposed method provides BRDF parameters at each texel without optimization and over-sampling. We demonstrated rapid measurement with real objects that do not have uniform reflectance and confirmed the validity of this approach by comparison with conventional methods.
Leo Miyashita, Yoshihiro Watanabe, Masatoshi Ishikawa
3DV2
2014 3D rectification of distorted document image based on tiled rectangle fragments
abstract
This paper presents an approach for document rectification using a quasi-isometric mapping derived from a novel model that represents a deformed document surface. The model is composed of rectangle fragments of the developed document plane. This was realized by introducing gap relaxation which permits unconnected fragments. Our experiments show that our rectification approach can be applied to different types of document deformation.
Masahiro Hirano, Yoshihiro Watanabe, Masatoshi Ishikawa
ICIP2
2014 Real-time 3D page tracking and book status recognition for high-speed book digitization based on adaptive capturing
abstract
In this paper, we propose a new book digitization system that can obtain high-resolution document images while flipping the pages automatically. The distinctive feature of our system is the adaptive capturing that has a crucial role in achieving high speed and high resolution. This adaptive capturing requires observing the state of the flipped pages at high speed and with high accuracy. In order to meet this requirement, we newly propose a method of obtaining the 3D shape of the book, tracking each page, and evaluating the state. In addition, we explain the details of the proposed high-speed book digitization system. We also report some experiments conducted to verify the performance of the developed system.
Shohei Noguchi, Masahiro Yamada, Yoshihiro Watanabe, Masatoshi Ishikawa
WACV3
2013 High-resolution surface reconstruction based on multi-level implicit surface from multiple range images
abstract
In this paper, we propose a method for improving the resolution of a 3D shape reconstructed from multiple range images acquired from a moving target. In our approach, the alignment and surface estimation problems are solved in a simultaneous estimation framework based on multi-level implicit surface. We present results of experiments for evaluating the reconstruction accuracy with different point cloud densities and noise levels which shows that our method achieved good performance even when the input range images were low-resolution and sparse.
Shohei Noguchi, Yoshihiro Watanabe, Masatoshi Ishikawa
ICIP2
2013 Automatic page turner machine for high-speed book digitization
abstract
In recent years, there has been an increasing demand to digitize a huge number of books. A promising new approach for meeting this demand, called Book Flipping Scanning, has been proposed. This is a new style of scanning in which all pages of a book are captured while a user continuously flips through the pages without stopping at each page. Although this new technology has had a tremendous impact in the field of book digitization, page turning is still done manually, which acts as a bottleneck in the development of high-speed book digitization. Against this background, this paper proposes a newly designed high-speed, high-precision book page turner machine. Our machine turns the pages in a contactless manner by utilizing the elastic force of the paper and an air blast. This design enables high-speed performance that is ten times faster than conventional approaches and, in addition, causes no obstruction in the digitization process. This paper reports the evaluation of the proposed machine using various types of paper with different qualities. Our machine achieved almost 100% success rate when turning pages at around 300 pages/min, showing that it is a promising technology for turning pages at high-speed and with high precision.
Yoshihiro Watanabe, Miho Tamei, Masahiro Yamada, Masatoshi Ishikawa
IROS1
2012 Reconstruction of 3D Surface and Restoration of Flat Document Image from Monocular Image Sequence
Hiroki Shibayama, Yoshihiro Watanabe, Masatoshi Ishikawa
ACCV (4)2
2012 Digitization of Deformed Documents Using a High-Speed Multi-camera Array
Yoshihiro Watanabe, Kotaro Itoyama, Masahiro Yamada, Masatoshi Ishikawa
ACCV (2)1
2011 Stereo 3D reconstruction using prior knowledge of indoor scenes
abstract
We propose a new method of indoor-scene stereo vision that uses probabilistic prior knowledge of indoor scenes in order to exploit the global structure of artificial objects. In our method, we assume three properties of the global structure - planarity, connectivity, and parallelism/orthogonality - and we formulate them in the framework of maximum a posteriori (MAP) estimation. To enable robust estimation, we employ a probability distribution that has both high peaks and wide flat tails. In experiments, we demonstrated that our approach can estimate shapes whose surfaces are not constrained by three orthogonal planes. Furthermore, comparing our results with those of a conventional method that assumes a locally smooth disparity map suggested that the proposed method can estimate more globally consistent shapes.
Kentaro Kofuji, Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ICRA2
2011 VolVision: high-speed capture in unconstrained camera motion
abstract
In this paper, we propose a novel concept called VolVision that encompasses using a camera to reconstruct 6DoF unconstrained motion. "VolVision" is designed to handle imagery falling, tossed or thrown cameras. And VolVision also allows users to reconstruct dynamic images and generate a 3D-mapped scene from image sequences. It could be used to model severe environments like valley and mountains that are normally not easily viewed by humans. We produced a prototype that embodies the concept above, and were able to reconstruct the camera's path, perform image mosaicing, and track 3D information of feature points in images.
Hideki Takeoka, Yushi Moko, Carson Reynolds, Takashi Komuro, Yoshihiro Watanabe, Masatoshi Ishikawa
SIGGRAPH Asia Sketches5
2011 Human gait estimation using a wearable camera
abstract
We focus on the growing need for a technology that can achieve motion capture in outdoor environments. The conventional approaches have relied mainly on fixed installed cameras. With this approach, however, it is difficult to capture motion in everyday surroundings. This paper describes a new method for motion estimation using a single wearable camera. We focused on walking motion. The key point is how the system can estimate the original walking state using limited information from a wearable sensor. This paper describes three aspects: the configuration of the sensing system, gait representation, and the gait estimation method.
Yoshihiro Watanabe, Tetsuo Hatanaka, Takashi Komuro, Masatoshi Ishikawa
WACV1
2010 Wide range image sensing using a thrown-up camera
abstract
In this paper, we propose a wide-range image sensing method using a camera thrown up into the air. By using camera thrown up in this way, we can get images that are otherwise difficult to obtain, such as those taken from overhead. As an example of wide-range image sensing, we integrated video images captured by a thrown-up camera using an image mosaicing technique. When rotation about the optical axis of the camera can be ignored, we can integrate images by mosaicing using a translational approximation, which preferentially pastes pixels around the image center. To obtain the information about the camera direction, a rotational approximation using the angles of incident light rays is required. We also propose use of a high-frame-rate camera (HFR camera) in order to acquire a large amount of information. A seamless large image was obtained by synthesizing the images captured by a thrown-up HFR camera. We found that high frame rates of around 1000 fps were necessary.
Toshitaka Kuwa, Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ICME2
2010 Estimation of Non-rigid Surface Deformation Using Developable Surface Model
abstract
There is a strong demand for a method of acquiring a non-rigid shape under deformation with high accuracy and high resolution. However, this is difficult to achieve because of performance limitations in measurement hardware. In this paper, we propose a model based method for estimating non-rigid deformation of a developable surface. The model is based on geometric characteristics of the surface, which are important in various applications. This method improves the accuracy of surface estimation and planar development from a low-resolution point cloud. Experiments using curved documents showed the effectiveness of the proposed method.
Yoshihiro Watanabe, Takashi Nakashima, Takashi Komuro, Masatoshi Ishikawa
ICPR1
2009 High-resolution shape reconstruction from multiple range images based on simultaneous estimation of surface and motion
abstract
Recognition of dynamic scenes based on shape information could be useful for various applications. In this study, we aimed at improving the resolution of three-dimensional (3D) data obtained from moving targets. We present a simple clean and robust method that jointly estimates motion parameters and a high-resolution 3D shape. Experimental results are provided to illustrate the performance of the proposed algorithm.
Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ICCV1
2008 High-S/N imaging of a moving object using a high-frame-rate camera
abstract
In this paper we propose a high-S/N imaging method involving combining many images captured with small blur using a video camera capable of high-frame-rate image capturing at 1000 frames/s. Use of a high-frame-rate camera makes the image change between frames small, enabling easy motion estimation, and makes it possible to use more light information, even when the exposure time is reduced to avoid blurring. To obtain a clear picture without misalignment due to motion parallax, it is necessary to determine both the motion and a depth map of the subject from noisy input images. We show results when applying the proposed algorithm to an image sequence captured by a high-frame-rate camera.
Takashi Komuro, Yoshihiro Watanabe, Masatoshi Ishikawa, Tadakuni Narabu
ICIP2
2008 Integration of time-sequential range images for reconstruction of a high-resolution 3D shape
abstract
The recognition of dynamic scenes using 3D shapes could provide useful approaches for various applications. However, the conventional 3D-shape sensing systems dedicated for such scenes have had problems in spatial resolution, though they have achieved high sampling rate in temporal domain. In order to solve this limits, we present a method that integrates time-sequential partial range images capturing moving targets to reconstruct a high-resolution range image. In the proposed method, multiple range images are set in the same coordinate system based on multi-frame simultaneous alignment. This paper also demonstrates the performance of the proposed method using some example rigid bodies.
Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ICPR1
2007 A High-Speed Vision System for Moment-Based Analysis of Numerous Objects
abstract
We describe a high-speed vision system for real-time applications, which is capable of processing visual information at a frame rate of 1 kfps, including both imaging and processing. Our system performs moment-based analysis of numerous objects. Moments are useful values providing information about geometric features and invariant features with respect to image-plane transformations. In addition, the simultaneous observation of numerous objects allows recognition of various complex phenomena. The proposed system achieves high-speed image processing by providing a dedicated massively parallel co-processor for moment extraction. The co-processor has a high-performance core based on a pixel-parallel and object-parallel calculation method. We constructed a prototype system and evaluated its performance. We present results obtained in actual operation.
Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ICIP (5)1
2007 955-fps Real-time Shape Measurement of a Moving/Deforming Object using High-speed Vision for Numerous-point Analysis
abstract
This paper describes real-time shape measurement using a newly developed high-speed vision system. Our proposed measurement system can observe a moving/deforming object at high frame rate and can acquire data in real-time. This is realized by using two-dimensional pattern projection and a high-speed vision system with a massively parallel co-processor for numerous-point analysis. We detail our proposed shape measurement system and present some results of evaluation experiments. The experimental results show the advantages of our system compared with conventional approaches.
Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ICRA1
2007 Design of a Massively Parallel Vision Processor based on Multi-SIMD Architecture
abstract
Increasing demands for robust image recognition systems require vision processors not only with enormous computational capacities but also with sufficient flexibility to handle highly complicated recognition tasks. We describe a multi-SIMD architecture and the design of a vision processor based on it for carrying out such difficult image recognition tasks. The proposed architecture consists of two SIMD parallel processing modules and a shared memory, allowing highly parallelized and flexible computation of complicated recognition tasks, which were difficult to process on a conventional massively parallel SIMD architecture. We designed a prototype vision processor for evaluation purposes and confirmed that the processor could be implemented in FPGA.
Kota Yamaguchi, Yoshihiro Watanabe, Takashi Komuro, Masatoshi Ishikawa
ISCAS2