Hiroshi Kawasaki

dblp:84/4575 · DBLP profile ↗
← Back
109ranked-venue papers
18as first author
31since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 91 · 15 first-author · 25 since 2021Artificial intelligence and machine learning · 53 · 11 first-author · 11 since 2021Systems, architecture and hardware · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 2Security and privacy · 1Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Neural 4D Scene Reconstruction with Multiple One-Shot Scanning Systems
abstract
Recently, 3D reconstruction from multiview stereo (MVS) has advanced significantly with the introduction of neural implicit representation methods, which estimate voxel densities or signed distance fields (SDFs) to describe the 3D structure of a scene. Although such neural-based methods typically require a large number of captured images to estimate dense volumetric information during training, developing systems that can recover the 3D shape of moving objects using only a small number of stationary cameras remains highly demanding and challenging. To address the issue of sparse views, various active lighting techniques have been proposed. However, the problem remains inherently difficult, particularly when attempting to capture the complete shape of an object with a wide baseline. In this paper, we propose a novel approach that combines active lighting with photometric stereo (PS) using neural representations. Additionally, we introduce a multiplexed illumination technique that captures the entire shape of an object in a single shot. Although this results in a low signal-to-noise ratio (SNR), our method also addresses this issue. The advantages of our technique are demonstrated through real-world experiments, showcasing its ability to capture a 4D scene.
Ryusuke Sagawa, Kota Nishihara, Takafumi Iwaguchi, Hiroshi Kawasaki
3DV4
2026 Learning Underwater Image Enhancement Iteratively Without Reference Images
abstract
Since high-fidelity reference images are difficult to obtain in real underwater scenes, most deep models trained by synthetic paired data cannot match real-world data exactly. In this paper, we propose an unsupervised training framework for underwater image enhancement (UIE) by leveraging an iterative training strategy and quantification of specific neural units. Specifically, to eliminate the heavy color cast and distortion in the underwater images, we decompose the unsupervised image enhancement as two targeted sub-tasks, namely colorization and color compensation. First, a diffusion model is introduced for colorization to correct the green and blue color casts. Then, to intensify the learning ability of balanced color information, we introduce an extra network branch and propose a quantification mechanism for color compensation. The extra branch encodes style information from normal images into the generative model, while the quantification mechanism identifies and adjusts neural units relevant to warm colors, improving the model’s ability to learn balanced color feature representations for robust generation. In the end, through iterative training, color cast and distortion are progressively reduced, leading to a gradual improvement in the quality of the generated images. Experimental results on various widely used underwater datasets demonstrate that our approach achieves excellent performance, even when compared to recent supervised methods.
Yi Tang 0008, Hiroshi Kawasaki, Takafumi Iwaguchi, Hiroshi Masui
AAAI2
2026 UnderWater SLAM with Laser-light sectioning method using ST-GAT
abstract
Multi-line laser ID assignment is crucial for underwater 3D reconstruction but fails when lines fragment. We reformulate this as a graph-based sequence labeling task and propose a novel two-stage hierarchical framework using Spatio-Temporal Graph Attention Networks (ST-GAT). Our method first reasons over a spatio-temporal graph of laser endpoints and intersections to handle local fragmentation, then elevates this to a global segment-level optimization with trajectory-constrained Viterbi decoding to ensure temporal consistency. This GNN-based approach eliminates the reliance on complete epipolar geometry. Experiments on real underwater datasets demonstrate superior reconstruction completeness and temporal stability, especially in challenging environments where traditional methods fail.
Heyang Gao, Kazuto Ichimaru, Takafumi Iwaguchi, Hiroshi Kawasaki
WACV4
2026 Multi-view stereo with multiple projectors for oneshot entire shape scan based on Neural SDF and DSSS demultiplexing
abstract
3D reconstruction has been widely studied and applied in various fields. Multi-view stereo (MVS) methods can recover dense geometry from multiple views, but often fail for texture-less objects due to unreliable feature matching. Active stereo with structured light (SL) addresses this limitation, however, when using multiple cameras and projectors for entire shape acquisition, overlapped SL patterns interfere with one another, leading to decoding failures. We propose a novel MVS framework based on neural signed distance field (Neural SDF) with multiple projectors that employs Direct Sequence Spread Spectrum (DSSS) to separate multiplexed patterns. This approach enables robust and accurate 3D shape reconstruction through Neural SDF optimization with a photometric loss that accounts for both the positions and the patterns of the projectors. We built a real scanning device that surrounds an object and captures its entire shape at once. Experiments on several static and a dynamics are conducted to validate the effectiveness of the proposed method.
Kota Nishihara, Ryo Furukawa 0001, Ryusuke Sagawa, Hiroshi Kawasaki
WACV4
2025 ShadowSG: Spherical Gaussian Illumination from Shadows
abstract
This work leverages shadow cues in a scene to infer the surrounding illumination of the shadow-casting object. Unlike prior works that optimize a discrete environment map, we model scene illumination using a mixture of spherical Gaussians (SGs). SG illumination provides more intuitive relations to shadow appearance and offers a more compact parameterization compared to discrete environment maps. To estimate SG parameters, we employ an SG-based, differentiable, closed-form rendering equation to explain the shading of the shadow plane and minimize a photometric loss between the rendered and observed shadow plane shading. Experiments on synthetic and real-world images under various surrounding illumination demonstrate that our method estimates illumination more accurately than approaches based on discrete environment maps. With our estimated lighting, consistent shadow effects are realized when blending virtual objects into real-world images11Code is available at https://github.com/CyberAgentAILab/ShadowSG..
Hiroshi Kawasaki, Takafumi Taketomi
3DV3
2025 Shape Reconstruction of Foreground and Background in Scenes with Translucent Objects Based on Coding Curves
abstract
Time-of-flight (ToF) cameras, widely used in commercial applications such as augmented reality and autonomous driving, measure depth by analyzing the flight time of emitted laser signals. Indirect ToF (I-ToF) cameras are particularly popular due to their high resolution, high frame rate, and affordability. However, they struggle with depth estimation in scenes containing translucent objects, as their fundamental assumption – single direct reflectance – breaks down due to complex light interactions. In this paper, we address depth estimation with translucent objects by leveraging coding curve (CC) distortions, which have recently been shown to mitigate multi-path interference (MPI). The CC is formulated as a function derived from multiple values of various types of temporal encoding by the photodetector at each pixel. Specifically, inspired by recent MPI solutions using CC, we sample both values and depth errors under translucent objects at fixed intervals to train a model that predicts foreground and background depth errors from the CC, ensuring accurate reconstruction of translucent scenes through error correction. Our approach is validated through real-world experiments, demonstrating its effectiveness in improving depth estimation in scenes with translucent objects.
Wenbin Luo, Takafumi Iwaguchi, Ryusuke Sagawa, Hiroshi Kawasaki
ICIP4
2025 Neural SDF for Shadow-Aware Unsupervised Structured Light
abstract
Among various active 3D measurement techniques, Structured Light (SL) is one of the most popular methods for its robustness and high accuracy. The ordinary SL system consists of a camera and a projector, and by projecting a pre-defined pattern, we can obtain pixel-to-pixel correspondences between the camera and the projector for triangulation. However, if we lack knowledge of the projected pattern for some reason, e.g., the projected pattern is not as expected due to lens distortion, inaccurate calibration, undesired optical phenomena like inter-reflection, and so on, the accuracy of conventional SL is severely degraded. As a remedy, we propose unsupervised structured light (USSL), which does not explicitly use prior knowledge of the pattern. Inspired by the fact that humans can recognize the scene structure illuminated by an unknown light source (e.g. rotating mirror ball), and some prior works have succeeded in novel-view-synthesis under unknown illumination conditions, we implement USSL on Neural Signed Distance Fields (Neural SDF) pipeline with implicit reflection module powered by a neural network. Additionally, since every SL method causes occlusion (shadow) by pattern projection, we must consider it for accurate shape reconstruction. To this end, we integrate shadow volume rendering into the proposed pipeline. Experiments with synthetic and real datasets are conducted to confirm the feasibility of the proposed method.
Kazuto Ichimaru, Diego Thomas, Takafumi Iwaguchi, Hiroshi Kawasaki
WACV4
2024 ActiveNeuS: Neural Signed Distance Fields for Active Stereo
abstract
3D-shape reconstruction in extreme environments, such as low illumination or scattering condition, has been an open problem and intensively researched. Active stereo is one of potential solution for such environments for its robustness and high accuracy. However, active stereo systems usually consist of specialized system configurations with complicated algorithms, which narrow their application. In this paper, we propose Neural Signed Distance Field for active stereo systems to enable implicit correspondence search and triangulation in generalized Structured Light. With our technique, textureless or equivalent surfaces by low light condition are successfully reconstructed even with a small number of captured images. Experiments were conducted to confirm that the proposed method could achieve state-of-the-art reconstruction quality under such severe condition. We also demonstrated that the proposed method worked in an underwater scenario.
Kazuto Ichimaru, Takaki Ikeda, Diego Thomas, Takafumi Iwaguchi, Hiroshi Kawasaki
3DV5
2024 Depth Reconstruction with Neural Signed Distance Fields in Structured Light Systems
abstract
We introduce a novel depth estimation technique for multi-frame structured light setups using neural implicit representations of 3D space. Our approach employs a neural signed distance field (SDF), trained through self-supervised differentiable rendering. Unlike passive vision, where joint estimation of radiance and geometry fields is necessary, we capitalize on known radiance fields from projected patterns in structured light systems. This enables isolated optimization of the geometry field, ensuring convergence and network efficacy with fixed device positioning. To enhance geometric fidelity, we incorporate an additional color loss based on object surfaces during training. Real-world experiments demonstrate our method’s superiority in geometric performance for few-shot scenarios, while achieving comparable results with increased pattern availability.
Rukun Qiao, Hiroshi Kawasaki, Hongbin Zha
3DV2
2024 Neural Active Structure-from-Motion in Dark and Textureless Environment
Kazuto Ichimaru, Diego Thomas, Takafumi Iwaguchi, Hiroshi Kawasaki
ACCV (10)4
2024 Multiple Active Stereo Systems Calibration Method Based on Neural SDF Using DSSS for Wide Area 3D Reconstruction
Kota Nishihara, Ryo Furukawa 0001, Ryusuke Sagawa, Hiroshi Kawasaki
ACCV (9)4
2024 A Practical Calibration Method for Cameras and Multiple Line-Lasers in Light Sectioning Systems for Underwater Environments
abstract
In recent years, the increasing demand for underwater 3D measurement for various applications has brought about challenges such as low accuracy in 3D shape acquisition and difficulties in localizing sensor positions. This paper introduces a robust calibration method for underwater 3D sensors, comprising line lasers and cameras, utilizing a physically accurate model. Specifically, our proposed line laser calibration method estimates laser plane parameters using two types of planar constraints, avoiding the need for the costly process of backward-projection of refraction for optimization. For camera calibration, we advocate a two-step approach incorporating a simple yet effective deep-learning-based marker detection algorithm to estimate parameters of refraction, representing a physically correct lens model. Through experiments, we validate the superior performance of our methods over previous approximation-based approaches, as demonstrated in simulations and actual experiments conducted in a swimming pool.
Takaki Ikeda, Takafumi Iwaguchi, Diego Thomas, Hiroshi Kawasaki
ICIP4
2024 Multi-Path Interference Mitigation For Indirect Time-of-Flight Camera By the Distortion of Coding Curve
abstract
The indirect time-of-flight camera measures depth based on the phase shift between modulated laser and its reflected light. However, when there are multi-paths due to interreflections, depth estimation is significantly affected as it assumes that only a single reflected light is observed. In this paper, we propose a method for mitigating multi-path interference utilizing coding curves derived from measurements of multiple sensor taps, which represent the contribution of global light component, i.e., multi-path. First, to obtain the coding curve, we offset the laser light to the sensor and measure the sensor tap values for different delays. We propose a data-driven method for mitigating MPI, utilizing a dataset where distorted coding curves are paired with MPI intensity. By using the measured coding curve as a query to obtain correction values from the dataset, we can mitigate the effect of MPI and obtain the correct depth. Unlike previous methods, our approach can solve the general MPI problem while imposing fewer restrictions on the type of reflections, the number of light paths, and the modulation frequency. We validate the effectiveness of our method through both simulation and real experiments.
Wenbin Luo, Takafumi Iwaguchi, Ryusuke Sagawa, Hiroshi Kawasaki
ICIP4
2024 Two-stage pose optimization algorithm using color information for underwater SLAM with light-sectioning-based 3D scanning method
abstract
The demand for 3D shape measurement of underwater scene is increasing in various applications. Especially, simultaneous localization and mapping (SLAM) technique utilizing remotely operated vehicle (ROV) attached with 3D sensors has been intensively researched. This paper focuses on solving pose optimization problem for underwater robots with camera/multiple-line-lasers setup, especially for the scene with some textures (color information). To this end, a two-stage pose optimization technique is proposed. In the first stage, due to the sparse nature of the reconstructed shape in the light-sectioning method consisting of several 3D curves, we bundle 10 to 20 consecutive frames to form a block shape, refining significant errors in the initial sensor poses using a novel bundle adjustment algorithm. In the second stage, remaining pose errors are corrected by a block-based matching algorithm utilizing iterative closest point (ICP) algorithm with color information. Through experiments in underwater environment with a real system, it was validated that the proposed method demonstrates superior performance compared to past underwater SLAM techniques.
Takaki Ikeda, Takafumi Iwaguchi, Diego Thomas, Hiroshi Kawasaki
IROS4
2024 Specular Object Reconstruction Behind Frosted Glass by Differentiable Rendering
abstract
This paper addresses the problem of reconstructing scenes behind optical diffusers, which is common in applications such as imaging through frosted glass. We propose a new approach that exploits specular reflection to capture sharp light distributions with a point light source, which can be used to detect reflections in low signal-to-noise scenarios. In this paper, we propose a rasterizer-based differentiable renderer to solve this problem by minimizing the difference between the captured and rendered images. Because our method can simultaneously optimize multiple observations for different light source positions, it is confirmed that ambiguities of the scene are efficiently eliminated by increasing the number of observations. Experiments show that the proposed method can reconstruct a scene with several mirror-like objects behind the diffuser in both simulated and real environments.
Takafumi Iwaguchi, Hiroyuki Kubo, Hiroshi Kawasaki
WACV3
2023 A Two-Step Approach for Interactive Animatable Avatars
Takumi Kitamura, Naoya Iwamoto, Hiroshi Kawasaki, Diego Thomas
CGI3
2023 Online Adaptive Disparity Estimation for Dynamic Scenes in Structured Light Systems
abstract
In recent years, deep neural networks have shown remarkable progress in dense disparity estimation from dynamic scenes in monocular structured light systems. However, their performance significantly drops when applied in unseen environments. To address this issue, self-supervised online adaptation has been proposed as a solution to bridge this performance gap. Unlike traditional fine-tuning processes, online adaptation performs test-time optimization to adapt networks to new domains. Therefore, achieving fast convergence during the adaptation process is critical for attaining satisfactory accuracy. In this paper, we propose an unsupervised loss function based on long sequential inputs. It ensures better gradient directions and faster convergence. Our loss function is designed using a multi-frame pattern flow, which comprises a set of sparse trajectories of the projected pattern along the sequence. We estimate the sparse pseudo ground truth with a confidence mask using a filter-based method, which guides the online adaptation process. Our proposed framework significantly improves the online adaptation speed and achieves superior performance on unseen data. The code is available on https://github.com/CodePointer/TIDENet.
Rukun Qiao, Hiroshi Kawasaki, Hongbin Zha
IROS2
2023 Underwater Image Enhancement by Transformer-based Diffusion Model with Non-uniform Sampling for Skip Strategy
abstract
In this paper, we present an approach to image enhancement with diffusion model in underwater scenes. Our method adapts conditional denoising diffusion probabilistic models to generate the corresponding enhanced images by using the underwater images and the Gaussian noise as the inputs. Additionally, in order to improve the efficiency of the reverse process in the diffusion model, we adopt two different ways. We firstly propose a lightweight transformer-based denoising network, which can effectively promote the time of network forward per iteration. On the other hand, we introduce a skip sampling strategy to reduce the number of iterations. Besides, based on the skip sampling strategy, we propose two different non-uniform sampling methods for the sequence of the time step, namely piecewise sampling and searching with the evolutionary algorithm. Both of them are effective and can further improve performance by using the same steps against the previous uniform sampling. In the end, we conduct a relative evaluation of the widely used underwater enhancement datasets between the recent state-of-the-art methods and the proposed approach. The experimental results prove that our approach can achieve both competitive performance and high efficiency. Our code is available at https://github.com/piggy2009/DM_underwater.
Yi Tang 0008, Hiroshi Kawasaki, Takafumi Iwaguchi
ACM Multimedia2
2023 Surface normal estimation from optimized and distributed light sources using DNN-based photometric stereo
abstract
Photometric stereo (PS) is a major technique to recover surface normal for each pixel. However, since it assumes Lambertian surface and directional light to estimate the value, a large number of images are usually required to avoid the effects of outliers and noise. In this paper, we propose a technique to reduce the number of images by using distributed light sources, where the patterns are optimized by a deep neural network (DNN). In addition, to efficiently realize the distributed light, we use an optical diffuser with a video projector, where the diffuser is illuminated by the projector from behind, the illuminated area on the diffuser works as if an arbitrary-shaped area light. To estimate the surface normal using the distributed light source, we propose a near-light photometric stereo (NLPS) using DNN. Since optimization of the pattern of distributed light is achieved by a differentiable renderer, it is connected with NLPS network, achieving end-to-end learning. The experiments are conducted to show the successful estimation of the surface normal by our method from a small number of images.
Takafumi Iwaguchi, Hiroshi Kawasaki
WACV2
2022 AutoEnhancer: Transformer on U-Net Architecture Search for Underwater Image Enhancement
Yi Tang 0008, Takafumi Iwaguchi, Hiroshi Kawasaki, Ryusuke Sagawa, Ryo Furukawa 0001
ACCV (3)3
2022 Robust Calibration-Marker and Laser-Line Detection For Underwater 3d Shape Reconstruction By Deep Neural Network
abstract
There are various demands for underwater 3D reconstruction, however, since most active stereo 3D reconstruction methods focus on the air environment, it is difficult to directly apply them to underwater due to the several critical reasons, such as refraction, water flow and severe attenuation. Typically, calibration-markers or laser-lines are strongly blurred and saturated by attenuation, which makes difficult to recover shape in the water. Another problem is that it is difficult to keep cameras, projectors and objects static in the water because of strong water flow, which prevents accurate calibration. In this paper, we propose a method to solve those problems by novel algorithm using deep neural network (DNN), epipolar constraint and specially designed devices. We also built a real system and tested it in the water, e.g., pool and sea. Experimental results confirmed the effectiveness of the proposed method. We also demonstrated real 3D scan in the sea.
Hanbin Wang, Takafumi Iwaguchi, Hiroshi Kawasaki
ICIP3
2022 Self-calibration of multiple-line-lasers based on coplanarity and Epipolar constraints for wide area shape scan using moving camera
abstract
High-precision three-dimensional scanning systems have been intensively researched and developed. Recently, for acquisition of large scale scene with high density, simultaneous localisation and mapping (SLAM) technique is preferred because of its simplicity; a single sensor that is moved around freely during 3D scanning. However, to integrate multiple scans, captured data as well as position of each sensor must be highly accurate, making these systems difficult to use in environments not accessible by humans, such as underwater, internal body, or outer space. In this paper, we propose a new, flexible system with multiple line lasers that reconstructs dense and accurate 3D scenes. The advantages of our proposed system are (1) no need of synchronization nor precalibration between lasers and a camera, and (2) the system can reconstruct 3D scenes in extreme conditions, such as underwater. We propose a new self-calibration method leveraging coplanarity and Epipolar constraints is proposed. We also propose a new bundle adjustment (BA) technique that is tailored to the system for a dense integration of multiple line laser scans. Experimental evaluation in both air and underwater environments confirms the advantages of the proposed method.
Genki Nagamatsu, Takaki Ikeda, Takafumi Iwaguchi, Diego Thomas, Jun Takamatsu, Hiroshi Kawasaki
ICPR6
2022 Auto-augmentation with Differentiable Renderer for High-frequency Shape Recovery
abstract
We propose a technique to estimate a high-resolution depth image from a sparse depth image captured by depth camera and a high-resolution shading image obtained by a RGB camera using deep neural network (DNN). In our technique, the network model is pretrained by synthetic images which are generated by rendering high-frequency shapes created by arithmetic model, such as sinusoidal wave of wide variation of parameters. Although the preparation of an appropriate synthetic dataset is critical for such tasks, it is not trivial to find a compact and optimal distribution of shape parameters. In this paper, we propose an auto augmentation technique to optimize hyperparameters for shapes achieving minimum number for training DNN. The proposed augmentation network directly optimizes the hyperparameters of a 3D scene including parameters of procedural shapes and their positions by gradient descent algorithm via a differentiable rendering technique. Unlike previous data augmentation techniques which only have basic image processing methods, such as affine and color transformations, the proposed method can generate optimal training dataset by changing the 3D shape and its position by using a differentiable renderer. In our experiments, we confirmed that our method improved the accuracy of high-resolution depth estimation as well as efficiency of training the network.
Kodai Tokieda, Takafumi Iwaguchi, Hiroshi Kawasaki
ICPR3
2022 Deep Gesture Generation for Social Robots Using Type-Specific Libraries
abstract
Body language such as conversational gesture is a powerful way to ease communication. Conversational gestures do not only make a speech more lively but also contain semantic meaning that helps to stress important information in the discussion. In the field of robotics, giving conversational agents (humanoid robots or virtual avatars) the ability to properly use gestures is critical, yet remain a task of extraordinary difficulty. This is because given only a text as input, there are many possibilities and ambiguities to generate an appropriate gesture. Different to previous works we propose a new method that explicitly takes into account the gesture types to reduce these ambiguities and generate human-like conversational gestures. Key to our proposed system is a new gesture database built on the TED dataset that allows us to map a word to one of three types of gestures: “Imagistic” gestures, which express the content of the speech, “Beat” gestures, which emphasize words, and “No gestures.” We propose a system that first maps the words in the input text to their corresponding gesture type, generate type-specific gestures and combine the generated gestures into one final smooth gesture. In our comparative experiments, the effectiveness of the proposed method was confirmed in user studies for both avatar and humanoid robot.
Hitoshi Teshima, Naoki Wake, Diego Thomas, Yuta Nakashima, Hiroshi Kawasaki, Katsushi Ikeuchi
IROS5
2022 MOTSLAM: MOT-assisted monocular dynamic SLAM using single-view depth estimation
abstract
Visual SLAM systems targeting static scenes have been developed with satisfactory accuracy and robustness. Dynamic 3D object tracking has then become a significant capability in visual SLAM with the requirement of under-standing dynamic surroundings in various scenarios including autonomous driving, augmented and virtual reality. However, performing dynamic SLAM solely with monocular images remains a challenging problem due to the difficulty of asso-ciating dynamic features and estimating their positions. In this paper, we present MOTSLAM, a dynamic visual SLAM system with the monocular configuration that tracks both poses and bounding boxes of dynamic objects. MOTSLAM first performs multiple object tracking (MOT) with associated both 2D and 3D bounding box detection to create initial 3D objects. Then, neural-network-based monocular depth estimation is applied to fetch the depth of dynamic features. Finally, camera poses, object poses, and both static, as well as dynamic map points, are jointly optimized using a novel bundle adjustment. Our experiments on the KITTI dataset demonstrate that our system has reached best performance on both camera ego-motion and object tracking on monocular dynamic SLAM.
Hanwei Zhang 0003, Hideaki Uchiyama, Shintaro Ono, Hiroshi Kawasaki
IROS4
2022 Single-shot dense active stereo with pixel-wise phase estimation based on grid-structure using CNN and correspondence estimation using GCN
abstract
Active stereo systems based on static pattern projection, a.k.a. oneshot scan, have been widely used for measuring dynamic scenes. Many patterns used for oneshot active stereo have grid-structures and grid-wise codes. For such systems, the grid structure is first detected, and graph matching methods are applied to estimate correspondences. However, such graph matching is often vulnerable to graph connection errors caused by grid structure analysis based on image features. Also, dense reconstruction for such systems is an open problem, where pixel-wise correspondence estimation from sparse image features is required. We propose a learning-based method to capture grid structure information and pixel-wise positional information simultaneously. We also propose to represent the grid structure by graphs with augmented connections other than 4-neighbor connections and applying them to a graph convolutional network (GCN). The proposed method can analyze large variety of grid patterns, has auto-calibration capability, can reconstruct dense shapes for fast moving objects.
Ryo Furukawa 0001, Michihiro Mikamo, Ryusuke Sagawa, Hiroshi Kawasaki
WACV4
2022 Orchestrated neuronal migration and cortical folding: A computational and experimental study
abstract
Brain development involves precisely orchestrated genetic, biochemical, and mechanical events. At the cellular level, neuronal proliferation in the innermost zone of the brain followed by migration towards the outermost layer results in a rapid increase in brain surface area, outpacing the volumetric growth of the brain, and forming the highly folded cortex. This work aims to provide mechanistic insights into the process of brain development and cortical folding using a biomechanical model that couples cell division and migration with volumetric growth. Unlike phenomenological growth models, our model tracks the spatio-temporal development of cohorts of neurons born at different times, with each cohort modeled separately as an advection-diffusion process and the total cell density determining the extent of volume growth. We numerically implement our model in Abaqus/Standard (2020) by writing user-defined element (UEL) subroutines. For model calibration, we apply in utero electroporation (IUE) to ferret brains to visualize and track cohorts of neurons born at different stages of embryonic development. Our calibrated simulations of cortical folding align qualitatively with the ferret experiments. We have made our experimental data and finite-element implementation available online to offer other researchers a modeling platform for future study of neurological disorders associated with atypical neurodevelopment and cortical malformations.
Shuolun Wang, Kengo Saito, Hiroshi Kawasaki, Maria A. Holland
PLoS Comput. Biol.3
2021 A Method for Adding Motion-Blur on Arbitrary Objects by Using Auto-Segmentation and Color Compensation Techniques
abstract
When dynamic objects are captured by a camera, motion blur inevitably occurs. Such a blur is sometimes considered as just a noise, however, it sometimes gives an important effect to add dynamism in the scene for photographs or videos. Unlike the similar effects, such as defocus blur, which is now easily controlled even by smartphones, motion blur is still uncontrollable and makes undesired effects on photographs. In this paper, an unified framework to add motion blur on per-object basis is proposed. In the method, multiple frames are captured without motion blur and they are accumulated to create motion blur on target objects. To capture images without motion blur, shutter speed must be short, however, it makes captured images dark, and thus, a sensor gain should be increased to compensate it. Since a sensor gain causes a severe noise on image, we propose a color compensation algorithm based on non-linear filtering technique for solution. Another contribution is that our technique can be used to make HDR images for fast moving objects by using multi-exposure images. In the experiments, effectiveness of the method is confirmed by ablation study using several data sets.
Michihiro Mikamo, Ryo Furukawa 0001, Hiroshi Kawasaki
ICIP3
2021 PoseRN: A 2D Pose Refinement Network For Bias-Free Multi-View 3D Human Pose Estimation
abstract
We propose a new 2D pose refinement network that learns to predict the human bias in the estimated 2D pose. There are biases in 2D pose estimations that are due to differences between annotations of 2D joint locations based on annotators’ perception and those defined by motion capture (MoCap) systems. These biases are crafted into publicly available 2D pose datasets and cannot be removed with existing error reduction approaches. Our proposed pose refinement network allows us to efficiently remove the human bias in the estimated 2D poses and achieve highly accurate multi-view 3D human pose estimation.
Akihiko Sayo, Diego Thomas, Hiroshi Kawasaki, Yuta Nakashima, Katsushi Ikeuchi
ICIP3
2021 High-Frequency Shape Recovery from Shading by CNN and Domain Adaptation
abstract
Importance of structured-light based one-shot scanning technique is increasing because of its simple system configuration and ability of capturing moving objects. One severe limitation of the technique is that it can capture only sparse shape, but not high frequency shapes, because certain area of projection pattern is required to encode spatial information. In this paper, we propose a technique to recover high-frequency shapes by using shading information, which is captured by one-shot RGB-D sensor based on structured light with single camera. Since color image comprises shading information of object surface, high-frequency shapes can be recovered by shape from shading techniques. Although multiple images with different lighting positions are required for shape from shading techniques, we propose a learning based approach to recover shape from a single image. In addition, to overcome the problem of preparing sufficient amount of data for training, we propose a new data augmentation method for high-frequency shapes using synthetic data and domain adaptation. Experimental results are shown to confirm the effectiveness of the proposed method.
Kodai Tokieda, Takafumi Iwaguchi, Hiroshi Kawasaki
ICIP3
2021 Self-calibrated dense 3D sensor using multiple cross line-lasers based on light sectioning method and visual odometry
abstract
Among various 3D capturing systems, since the system with line lasers based on the light sectioning method is simple and accurate, it has widely attracted many developers and used for many purposes. In addition, there is no need to synchronize the camera and the laser and also the configuration of the camera and the lasers is flexible, and thus, the system can be used for extreme conditions, such as underwater. There are two open problems for the system. The first problem is a low density of the 3D shape obtained from a single image, i.e., just several curves. The second problem is the accuracy of line detection in the wild. In this paper, we propose a self-calibration method using visual odometry (VO) to bundle a large number of frames to increase the density to solve the first problem. We also propose a robust line detection algorithm using CNN to solve the second problem. Comparative experiments prove the effectiveness of our proposed method. In addition, the system was tested in the extreme condition for demonstration.
Genki Nagamatsu, Jun Takamatsu, Takafumi Iwaguchi, Diego Thomas, Hiroshi Kawasaki
IROS5
2020 Dense Pixel-Wise Micro-motion Estimation of Object Surface by Using Low Dimensional Embedding of Laser Speckle Pattern
Ryusuke Sagawa, Yusuke Higuchi, Hiroshi Kawasaki, Ryo Furukawa 0001
ACCV (2)3
2020 FreeCam3D: Snapshot Structured Light 3D with Freely-Moving Cameras
Vivek Boominathan, Jacob T. Robinson, Hiroshi Kawasaki, Aswin C. Sankaranarayanan, Ashok Veeraraghavan
ECCV (27)5
2020 3d Imaging For Thermal Cameras Using Structured Light
abstract
Optical 3D sensing technologies are exploited for many applications in autonomous vehicles, manufacturing, and consumer products. However, existing techniques may suffer in certain challenging conditions, where scattering may occur due to particles. While the light in the visible and near IR spectrum is affected by scattering, long-wave IR (LWIR) tends to experience less scattering, especially when the particles are much smaller than the incident radiation. We propose and demonstrate the expansion of structured light scanning approaches into the LWIR spectrum using a thermal camera and black body radiation source. We then validate the results produced against ground truth scans from traditional structured light scanners. Additional means for projecting these scanning patterns are also discussed alongside potential drawbacks and challenges of this technique associated with future adoption.
Jack Erdozain, Kazuto Ichimaru, Tomohiro Maeda, Hiroshi Kawasaki, Ramesh Raskar, Achuta Kadambi
ICIP4
2020 Unsupervised 3D Human Pose Estimation in Multi-view-multi-pose Video
abstract
3D human pose estimation from a single 2D video is an extremely difficult task because computing 3D geometry from 2D images is an ill-posed problem. Recent popular solutions adopt fully-supervised learning strategy, which requires to train a deep network on a large-scale ground truth dataset of 3D poses and 2D images. However, such a large-scale dataset with natural images does not exist, which limits the usability of existing methods. While building a complete 3D dataset is tedious and expensive, abundant 2D in-the-wild data is already publicly available. As a consequence, there is a growing interest in the computer vision community to design efficient techniques that use the unsupervised learning strategy, which does not require any ground truth 3D data. Such methods can be trained with only natural 2D images of humans. In this paper we propose an unsupervised method for estimating 3D human pose in videos. The standard approach for unsupervised learning is to use the Generative Adversarial Network (GAN) framework. To improve the performance of 3D human pose estimation in videos, we propose a new GAN network that enforces body consistency over frames in a video. We evaluate the efficiency of our proposed method on a public 3D human body dataset.
Diego Thomas, Hiroshi Kawasaki
ICPR3
2020 Attention R-CNN for Accident Detection
abstract
This paper addresses accident detection where we not only detect objects with classes, but also recognize their characteristic properties. More specifically, we aim at simultaneously detecting object class bounding boxes on roads and recognizing their status such as safe, dangerous, or crashed. To achieve this goal, we construct a new dataset and propose a baseline method for benchmarking the task of accident detection. We design an accident detection network, called Attention R-CNN, which consists of two streams: one is for object detection with classes and one for characteristic property computation. As an attention mechanism capturing contextual information in the scene, we integrate global contexts exploited from the scene into the stream for object detection. This introduced attention mechanism enables us to recognize object characteristic properties. Extensive experiments on the newly constructed dataset demonstrate the effectiveness of our proposed network. The dataset and source code are publicly available on our project page.
Trung-Nghia Le, Shintaro Ono, Akihiro Sugimoto, Hiroshi Kawasaki
IV4
2020 Toward Interactive Self-Annotation For Video Object Bounding Box: Recurrent Self-Learning And Hierarchical Annotation Based Framework
abstract
Amount and variety of training data drastically affect the performance of CNNs. Thus, annotation methods are becoming more and more critical to collect data efficiently. In this paper, we propose a simple yet efficient Interactive Self-Annotation framework to cut down both time and human labor cost for video object bounding box annotation. Our method is based on recurrent self-supervised learning and consists of two processes: automatic process and interactive process, where the automatic process aims to build a supported detector to speed up the interactive process. In the Automatic Recurrent Annotation, we let an off-the-shelf detector watch unlabeled videos repeatedly to reinforce itself automatically. At each iteration, we utilize the trained model from the previous iteration to generate better pseudo ground-truth bounding boxes than those at the previous iteration, recurrently improving self-supervised training the detector. In the Interactive Recurrent Annotation, we tackle the human-in-the-loop annotation scenario where the detector receives feedback from the human annotator. To this end, we propose a novel Hierarchical Correction module, where the annotated frame-distance binarizedly decreases at each time step, to utilize the strength of CNN for neighbor frames. Experimental results on various video datasets demonstrate the advantages of the proposed framework in generating high-quality annotations while reducing annotation time and human labor costs.
Trung-Nghia Le, Akihiro Sugimoto, Shintaro Ono, Hiroshi Kawasaki
WACV4
2019 Simultaneous Shape Registration and Active Stereo Shape Reconstruction using Modified Bundle Adjustment
abstract
Simultaneous registration and shape fusion using 3D scanners have been proposed for conducting wide-area and dense 3D shape reconstruction. However, because the 3D scanners for such a system must be robust and should provide feedback in real time, only a few devices are available, thereby limiting the application of the technique. In this study, we propose a new wide-area scanning algorithm that only requires an off-the-shelf projector and a camera. In our technique, the devices are not necessarily fixed to each other and the relative positions of the devices as well as the scene shapes can be precisely estimated by bundle adjustment (BA) in case of structured light. To efficiently perform shape registration, a robust and dense shape reconstruction is required, which is currently considered to be an open problem for structured light systems. In this study, we suggest a novel network-based feature detection algorithm as well as shape fusion algorithm for the solution.
Ryo Furukawa 0001, Genki Nagamatsu, Hiroshi Kawasaki
3DV3
2019 Unified Underwater Structure-from-Motion
abstract
This paper shows that accurate underwater 3D shape reconstruction is possible using a single camera, observing a target through a refractive interface. We provide unified reconstruction techniques for a variety of scenarios such as single static camera and moving refractive interface, single moving camera and static refractive interface, and single moving camera and moving refractive interface. In our basic setup, we assume that the refractive interface is planar, and simultaneously estimate the unknown transformations of the planar interface and the camera, and the unknown target shape using bundle adjustment. We also extend it to relax the planarity assumption, which enables us to use waves of the refractive interface for the reconstruction task. Experiments with real data show the superiority of our method to existing methods.
Kazuto Ichimaru, Yuichi Taguchi, Hiroshi Kawasaki
3DV3
2019 Mobile Photometric Stereo with Keypoint-Based SLAM for Dense 3D Reconstruction
abstract
The standard photometric stereo is a technique to densely reconstruct objects' surfaces using light variation under the assumption of a static camera with a moving light source. In this work, we use photometric stereo to reconstruct dense 3D scenes while moving the camera and the light altogether. In such non-static case, camera poses as well as correspondences between pixels of each frame to apply photometric stereo are required. ORB-SLAM is a technique that can be used to estimate camera poses. To retrieve correspondences, our idea is to start from a sparse 3D mesh obtained with ORB SLAM and then densify the mesh by a plane sweep method using a multi-view photometric consistency. By combining ORB-SLAM and photometric stereo, it is possible to reconstruct dense 3D scenes with a off-the-shelf smartphone and its embedded torchlight. Note that SLAM systems usually struggle with textureless object, which is effectively compensated by the photometric stereo in our method. Experiments are conducted to show that our proposed method gives better results than SLAM alone or COLMAP, especially for partially textureless surfaces.
Remy Maxence, Hideaki Uchiyama, Hiroshi Kawasaki, Diego Thomas, Vincent Nozick, Hideo Saito 0001
3DV3
2019 Underwater Stereo Using Refraction-Free Image Synthesized From Light Field Camera
abstract
There is a strong demand on capturing underwater scenes without distortions caused by refraction. Since a light field camera can capture several light rays at each point of an image plane from various directions, if geometrically correct rays are chosen, it is possible to synthesize a refraction-free image. In this paper, we propose a novel technique to efficiently select such rays to synthesize a refraction-free image from an underwater image captured by a light field camera. In addition, we propose a stereo technique to reconstruct 3D shapes using a pair of our refraction-free images, which are central projection. In the experiment, we captured several underwater scenes by two light field cameras, synthesized refraction free images and applied stereo technique to reconstruct 3D shapes. The results are compared with previous techniques which are based on approximation, showing the strength of our method.
Kazuto Ichimaru, Hiroshi Kawasaki
ICIP2
2019 Human Shape Reconstruction with Loose Clothes from Partially Observed Data by Pose Specific Deformation
Akihiko Sayo, Hayato Onizuka, Diego Thomas, Yuta Nakashima, Hiroshi Kawasaki, Katsushi Ikeuchi
PSIVT5
2019 CNN Based Dense Underwater 3D Scene Reconstruction by Transfer Learning Using Bubble Database
abstract
Dense 3D shape acquisition of swimming human or live fish is an important research topic for sports, biological science and so on. For this purpose, active stereo sensor is usually used in the air, however it cannot be applied to the underwater environment because of refraction, strong light attenuation and severe interference of bubbles. Passive stereo is a simple solution for capturing dynamic scenes at underwater environment, however the shape with textureless surfaces or irregular reflections cannot be recovered. Recently, the stereo camera pair with a pattern projector for adding artificial textures on the objects is proposed. However, to use the system for underwater environment, several problems should be compensated, i.e., disturbance by fluctuation and bubbles. Simple solution is to use convolutional neural network for stereo to cancel the effects of bubbles and/or water fluctuation. Since it is not easy to train CNN with small size of database with large variation, we develop a special bubble generation device to efficiently create real bubble database of multiple size and density. In addition, we propose a transfer learning technique for multi-scale CNN to effectively remove bubbles and projected-patterns on the object. Further, we develop a real system and actually captured live swimming human, which has not been done before. Experiments are conducted to show the effectiveness of our method compared with the state of the art techniques.
Kazuto Ichimaru, Ryo Furukawa 0001, Hiroshi Kawasaki
WACV3
2018 Multi-scale CNN Stereo and Pattern Removal Technique for Underwater Active Stereo System
abstract
Demands on capturing dynamic scenes of underwater environments are rapidly growing. Passive stereo is applicable to capture dynamic scenes, however the shape with textureless surfaces or irregular reflections cannot be recovered by the technique. In our system, we add a pattern projector to the stereo camera pair so that artificial textures are augmented on the objects. To use the system at underwater environments, several problems should be compensated, i.e., refraction, disturbance by fluctuation and bubbles. Further, since surface of the objects are interfered by the bubbles, projected patterns, etc., those noises and patterns should be removed from captured images to recover original texture. To solve these problems, we propose three approaches; a depth-dependent calibration, Convolutional Neural Network(CNN)-stereo method and CNN-based texture recovery method. A depth-dependent calibration I sour analysis to find the acceptable depth range for approximation by center projection to find the certain target depth for calibration. In terms of CNN stereo, unlike common CNN based stereo methods which do not consider strong disturbances like refraction or bubbles, we designed a novel CNN architecture for stereo matching using multi-scale information, which is intended to be robust against such disturbances. Finally, we propose a multi-scale method for bubble and a projected-pattern removal method using CNNs to recover original textures. Experimental results are shown to prove the effectiveness of our method compared with the state of the art techniques. Furthermore, reconstruction of a live swimming fish is demonstrated to confirm the feasibility of our techniques.
Kazuto Ichimaru, Ryo Furukawa 0001, Hiroshi Kawasaki
3DV3
2018 A Coded Aperture for Watermark Extraction from Defocused Images
Hiroki Hamasaki, Shingo Takeshita, Kentaro Nakai, Toshiki Sonoda, Hiroshi Kawasaki, Hajime Nagahara, Satoshi Ono
ACCV (6)5
2018 Representing a Partially Observed Non-Rigid 3D Human Using Eigen-Texture and Eigen-Deformation
abstract
Reconstruction of the shape and motion of humans from RGB-D is a challenging problem, receiving much attention in recent years. Recent approaches for full-body reconstruction use a statistic shape model, which is built upon accurate full-body scans of people in skin-tight clothes, to complete invisible parts due to occlusion. Such a statistic model may still be fit to an RGB-D measurement with loose clothes but cannot describe its deformations, such as clothing wrinkles. Observed surfaces may be reconstructed precisely from actual measurements, while we have no cues for unobserved surfaces. For full-body reconstruction with loose clothes, we propose to use lower dimensional embeddings of texture and deformation referred to as eigen-texturing and eigen-deformation, to reproduce views of even unobserved surfaces. Provided a full-body reconstruction from a sequence of partial measurements as 3D meshes, the texture and deformation of each triangle are then embedded using eigen-decomposition. Combined with neural-network-based coefficient regression, our method synthesizes the texture and deformation from arbitrary viewpoints. We evaluate our method using simulated data and visually demonstrate how our method works on real data.
Ryosuke Kimura, Akihiko Sayo, Fabian Lorenzo Dayrit, Yuta Nakashima, Hiroshi Kawasaki, Ambrosio Blanco, Katsushi Ikeuchi
ICPR5
2018 Simultaneous independent information display at multiple depths using multiple projectors and patterns created by epipolar constraint and homography transformation
abstract
Generally, the projected image of a video projector is spatially invariant and cannot project different images at different depths. If independent images can be projected at different depths, there is a great potential for new AR or MR information presentation devices. In the previous research, solution using multiple projectors is presented, however, there are problems in insufficient contrast. In the paper, we propose a new algorithm to achieve high contrast projection, by restricting a projection pattern to only symbolic information such as letters, which allows a large freedom on pattern optimization to present the same information. For efficency, we also introduce epipolar constraint-based pattern optimization algorithm, which divides the original 2D problem into 1D problems. In our demonstration, it will be shown that several words are presented at different depths with static video projectors with high contrast.
Yuto Hirao, Hiroshi Kawasaki
VRST2
2017 Realtime Novel View Synthesis with Eigen-Texture Regression
Yuta Nakashima, Fumio Okura, Norihiko Kawai, Ryosuke Kimura, Hiroshi Kawasaki, Katsushi Ikeuchi, Ambrosio Blanco
BMVC5
2017 Depth Estimation Using Structured Light Flow - Analysis of Projected Pattern Flow on an Object's Surface
abstract
Shape reconstruction techniques using structured light have been widely researched and developed due to their robustness, high precision, and density. Because the techniques are based on decoding a pattern to find correspondences, it implicitly requires that the projected patterns be clearly captured by an image sensor, i.e., to avoid defocus and motion blur of the projected pattern. Although intensive researches have been conducted for solving defocus blur, few researches for motion blur and only solution is to capture with extremely fast shutter speed. In this paper, unlike the previous approaches, we actively utilize motion blur, which we refer to as a light flow, to estimate depth. Analysis reveals that minimum two light flows, which are retrieved from two projected patterns on the object, are required for depth estimation. To retrieve two light flows at the same time, two sets of parallel line patterns are illuminated from two video projectors and the size of motion blur of each line is precisely measured. By analyzing the light flows, i.e. lengths of the blurs, scene depth information is estimated. In the experiments, 3D shapes of fast moving objects, which are inevitably captured with motion blur, are successfully reconstructed by our technique.
Ryo Furukawa 0001, Ryusuke Sagawa, Hiroshi Kawasaki
ICCV3
2017 Temporal Shape Super-Resolution by Intra-frame Motion Encoding Using High-fps Structured Light
abstract
One of the solutions of depth imaging of moving scene is to project a static pattern on the object and use just a single image for reconstruction. However, if the motion of the object is too fast with respect to the exposure time of the image sensor, patterns on the captured image are blurred and reconstruction fails. In this paper, we impose multiple projection patterns into each single captured image to realize temporal super resolution of the depth image sequences. With our method, multiple patterns are projected onto the object with higher fps than possible with a camera. In this case, the observed pattern varies depending on the depth and motion of the object, so we can extract temporal information of the scene from each single image. The decoding process is realized using a learning-based approach where no geometric calibration is needed. Experiments confirm the effectiveness of our method where sequential shapes are reconstructed from a single image. Both quantitative evaluations and comparisons with recent techniques were also conducted.
Yuki Shiba, Satoshi Ono, Ryo Furukawa 0001, Shinsaku Hiura, Hiroshi Kawasaki
ICCV5
2017 Learning-based feature extraction for active 3D scan with reducing color crosstalk of multiple pattern projections
abstract
3D reconstruction methods based on active stereo technique have been widely used for many practical systems. Many of these systems are configured with a single camera and a single projector. Since such systems can only capture one side of the target object, several attempts have been conducted to enlarge the captured area, especially multi-projector systems attract many researchers. For multi-projector based systems, overlap between multiple pattern projections is a serious problem. Even if different color channels are used for each projector, complete separation is not possible because of color crosstalks. Another open problem is decoding errors of the projected patterns, which causes a failure on extracting positional information of the projected pattern form the captured image. Among several reasons for such errors, color crosstalks are crucial because their features are similar to the main signal and difficult to be decomposed. In this paper, we solve these problems by utilizing machine learning techniques where a convolutional neural network is trained to extract low dimensional pattern features for each projector. In addition, it is trained to suppress the color crosstalks from different projectors. Using this new technique, we succeeded in reconstructing 3D shapes from images where multiple patterns are overlapped.
Ryusuke Sagawa, Ryo Furukawa 0001, Akiko Matsumoto, Hiroshi Kawasaki
ICRA4
2017 ReMagicMirror: Action Learning Using Human Reenactment with the Mirror Metaphor
Fabian Lorenzo Dayrit, Ryosuke Kimura, Yuta Nakashima, Ambrosio Blanco, Hiroshi Kawasaki, Katsushi Ikeuchi, Tomokazu Sato, Naokazu Yokoya
MMM (1)5
2017 Auto-calibration Method for Active 3D Endoscope System Using Silhouette of Pattern Projector
Ryo Furukawa 0001, Masahito Naito, Daisuke Miyazaki, Masahi Baba, Shinsaku Hiura, Yoji Sanomura, Shinji Tanaka, Hiroshi Kawasaki
PSIVT8
2017 Calibration Technique for Underwater Active Oneshot Scanning System with Static Pattern Projector and Multiple Cameras
abstract
Underwater 3D shape scanning technique becomes popular because of several rising research topics, such as map making of submarine topography for autonomous underwater vehicle (UAV), shape measurement of live fish, motion capture of swimming human, etc. Structured light systems (SLS) based active 3D scanning systems are widely used in the air and also promising to apply underwater environment. When SLS is used in the air, the stereo correspondences can be efficiently retrieved by epipolar constraint. However, in the underwater environment, the camera and projector are usually set in special housings and refraction occurs at the interfaces between water/glass and glass/air, resulting in invalid conditions for epipolar constraint which severely deteriorates the correspondence search process. In this paper, we propose an efficient technique to calibrate the underwater SLS systems as well as robust 3D shape acquisition technique. In order to avoid the calculation complexity, we approximate the system with central projection model. Although such an approximation produces an inevitable errors in the system, such errors are diminished by a combination of grid based SLS technique and a bundle adjustment algorithm. We tested our method with a real underwater SLS, consisting ofcustom-made laser pattern projector and underwater housings, showing the validity ofour method.
Hiroshi Kawasaki, Hideaki Nakai, Hirohisa Baba, Ryusuke Sagawa, Ryo Furukawa 0001
WACV1
2016 Simultaneous Independent Image Display Technique on Multiple 3D Objects
Takuto Hirukawa, Marco Visentini Scarzanella, Hiroshi Kawasaki, Ryo Furukawa 0001, Shinsaku Hiura
ACCV (4)3
2016 Shape Acquisition and Registration for 3D Endoscope Based on Grid Pattern Projection
Ryo Furukawa 0001, Hiroki Morinaga, Yoji Sanomura, Shinji Tanaka, Shigeto Yoshida, Hiroshi Kawasaki
ECCV (6)6
2016 Lets not stare at smartphones while walking: memorable route recommendation by detecting effective landmarks
abstract
Navigation in unfamiliar cities often requires frequent map checking, which is troublesome for wayfinders. We propose a novel approach for improving real-world navigation by generating short, memorable and intuitive routes. To do so we detect useful landmarks for effective route navigation. This is done by exploiting not only geographic data but also crowd footprints in Social Network Services (SNS) and Location Based Social Networks (LBSN). Specifically, we detect point, area, and line landmarks by using three indicators to measure landmark's utility: visit popularity, direct visibility, and indirect visibility. We then construct an effective route graph based on the extracted landmarks, which facilitates optimal path search. In the experiments, we show that landmark-based routes out-perform the ones created by baseline from the perspectives of the lap time and the number of references necessary to check self-positions for adjusting route directions.
Shoko Wakamiya, Hiroshi Kawasaki, Yukiko Kawai, Adam Jatowt, Eiji Aramaki, Toyokazu Akiyama
UbiComp2
2016 Registration and entire shape acquisition for grid based active one-shot scanning techniques
abstract
One-shot active stereo using structured light is a practical solution for dynamic scene acquisition. Basically, those methods are based on encoding positional information of the pixel into the single projected pattern. A disadvantage of such methods is decreases of the spatial resolution caused by requiring a certain area of the pattern to encode the positional information. Among those methods, grid-based patterns are promising at the point of accuracy and robustness, since triangulation for 3D reconstruction is conducted with light-sectioning method and a line detection is usually a stable image processing. However, no shapes are recovered between the grid lines, and thus, the whole reconstructed shape tends to be sparse. To deal with the problem, integrating multiple shapes that are sequentially captured using registration algorithm such as ICP is one solution. In previous work, we show that naive ICP works poorly for grid-like structured point clouds, and proposed a specialized ICP algorithm for aligning a set of grid-like structured 3D shapes. In this paper, we extend this approach and propose a process for entire shape modeling by capturing objects from all the directions using turn table, and integrating into a single shape using our improved ICP. To achieve this, setting good initial 3D shapes is important. For solution, we interpolation grid shapes to create smooth surface so that common ICP works. Comprehensive experiments are conducted to show the strength of our method compared to common ICP.
Hiroshi Kawasaki, Takuto Hirukawa, Ryo Furukawa 0001
ICPR1
2016 Automatic feature extraction using CNN for robust active one-shot scanning
abstract
Active one-shot scanning techniques have been widely used for various applications. Stereo-based active one-shot scanning embeds a positional information regarding the image plane of a projector onto a projected pattern to retrieve correspondences entirely from a captured image. Many combinations of patterns and decoding algorithms for active one-shot scanning have been proposed. If the capturing environment lacks the assumed conditions, such as the absence of strong external lights, then reconstruction using those methods is degraded, because the pattern decoding fails. In this paper, we propose a general reconstruction algorithm that can be used for any kind of patterns without strict assumptions. The technique is based on an efficient feature extraction function that can drastically reduce redundant information from the raw pixel values of patches of captured images. Shapes are reconstructed by efficiently finding correspondences between a captured image and the pattern using low-dimensional feature vectors. Such a function is created automatically by a convolutional neural network using a large database of pattern images that are efficiently synthesized by using GPU with wide variation of depth and surface orientation. Experimental results show that our technique can be used for several existing patterns without any ad hoc algorithm or information regarding the scene or the sensor.
Ryusuke Sagawa, Yuki Shiba, Takuto Hirukawa, Satoshi Ono, Hiroshi Kawasaki, Ryo Furukawa 0001
ICPR5
2015 Active One-Shot Scan for Wide Depth Range Using a Light Field Projector Based on Coded Aperture
abstract
The central projection model commonly used to model cameras as well as projectors, results in similar advantages and disadvantages in both types of system. Considering the case of active stereo systems using a projector and camera setup, a central projection model creates several problems, among them, narrow depth range and necessity of wide baseline are crucial. In the paper, we solve the problems by introducing a light field projector, which can project a depth-dependent pattern. The light field projector is realized by attaching a coded aperture with a high frequency mask in front of the lens of the video projector, which also projects a high frequency pattern. Because the light field projector cannot be approximated by a thin lens model and a precise calibration method is not established yet, an image-based approach is proposed to apply a stereo technique to the system. Although image-based techniques usually require a large database and often imply heavy computational costs, we propose a hierarchical approach and a feature-based search for solution. In the experiments, it is confirmed that our method can accurately recover the dense shape of curved and textured objects for a wide range of depths from a single captured image.
Hiroshi Kawasaki, Satoshi Ono, Yuuki Horita, Yuki Shiba, Ryo Furukawa 0001, Shinsaku Hiura
ICCV1
2015 Belief-propagation-based robust decoding for two-dimensional barcodes to overcome distortion and occlusion and its extension to multi-view decoding
abstract
This paper proposes a decoding method of 2D code robust against non-uniform, complicated distortion and occlusion. 2D codes printed on paper and cloth cause non-uniform distortions because they are not rigid. The proposed method uses 2D code involving auxiliary lines which allow recognition of distortion and occluded areas, and identify the lines using Belief Propagation (BP). In addition, the proposed method works with more than one camera without complicated processes such as calibration and registration, resulting in enhancing robustness against occlusion. Experimental results show that the proposed method can decode the 2D code with non-uniform, non-smooth distortions and occlusion.
Kohei Kamizuru, Yudai Kawakami, Hiroshi Kawasaki, Satoshi Ono
ICIP3
2015 A Triangle Mesh Reconstruction Method Taking into Account Silhouette Images
Michihiro Mikamo, Yoshinori Oki, Marco Visentini Scarzanella, Hiroshi Kawasaki, Ryo Furukawa 0001, Ryusuke Sagawa
PSIVT4
2015 Underwater Active Oneshot Scan with Static Wave Pattern and Bundle Adjustment
Hiroki Morinaga, Hirohisa Baba, Marco Visentini Scarzanella, Hiroshi Kawasaki, Ryo Furukawa 0001, Ryusuke Sagawa
PSIVT4
2015 Simultaneous Camera, Light Position and Radiant Intensity Distribution Calibration
Marco Visentini Scarzanella, Hiroshi Kawasaki
PSIVT2
2014 Simultaneous Entire Shape Registration of Multiple Depth Images Using Depth Difference and Shape Silhouette
Takuya Ushinohama, Yosuke Sawai, Satoshi Ono, Hiroshi Kawasaki
ACCV (2)4
2014 Simultaneous deblur and super-resolution technique for video sequence captured by hand-held video camera
abstract
Nowadays, video camera is commonly used everywhere and demand of retrieving a single shot from video sequence is increasing. Since resolution of video camera is usually lower than that of digital camera, simply cutting out a frame from a video sequence ends up with low quality. Further, because of the necessity of high fps on video camera, video data inevitably contains motion blur and it leads mis-registration between frames which is critical for multi-frame superresolution. In this paper, we propose a method to restore high-resolution image from a video sequence considering motion blur. Since the frame-rate of a video camera is high, motion of the object in successive frames is small, and thus, stable feature tracking during short sequences is possible even if there is a blur. Thus, we adopt a division/integration approach to realize robust tracking for long sequence. We also propose a simultaneous deblur and super-resolution technique using multiple images based on MAP estimation. Experimental results are shown to prove the strength of our method.
Yuki Matsushita, Hiroshi Kawasaki, Shintaro Ono, Katsushi Ikeuchi
ICIP2
2014 A Two-Dimensional Barcode with Robust Decoding against Distortion and Occlusion for Automatic Recognition of Garbage Bags
abstract
This paper proposes a 2D code and its decoding method robust against non-uniform, complicated distortions, assuming an application to automatic recognition of a plastic garbage bag. Printing a 2D code on a garbage bag is a promising approach for automatic bag recognition from the perspective of information content and cost. However, a 2D code printed on the bag causes non-uniform distortions because the bag is not rigid and does not hold a fixed shape. The proposed 2D code is based on Quick Response (QR) code and has auxiliary lines which allow recognition of distortion and occlusion areas, and the proposed decoding method localizes the lines by reliability calculation and iterated DP matching. Experimental results show that the proposed method in conjunction with the error correction function of QR code could decode the 2D code with non-uniform, non-smooth distortions and occluded areas.
Satoshi Ono, Yudai Kawakami, Hiroshi Kawasaki, Shinsuke Fujita
ICPR3
2014 Interactive 3D Animation Creation and Viewing System based on Motion Graph and Pose Estimation Method
abstract
This paper proposes an interactive 3D animation system specifically aiming efficient control of human motion. However there are various commercial products for creating movies and game contents, those are still difficult to deal with for non-professional users. To ease the creation process and encourage to utilize 3D animation for the general users, e.g., in the field of such as education, medicine and so on, we propose a system using Kinect. The data of skeleton models of human motion estimated by Kinect is processed to generate Motion Graph and finally restructure the data automatically for 3D character models. We also propose an efficient 3D animation viewing system based on touch interface for tablet device, which enables intuitive control of multiple motions of the human activity. To evaluate the effectiveness of the method, we implemented a prototype system and created several 3D animations.
Masayuki Furukawa, Yasuhiro Akagi, Yukiko Kawai, Hiroshi Kawasaki
ACM Multimedia4
2014 Dense 3D Reconstruction from High Frame-Rate Video Using a Static Grid Pattern
abstract
Dense 3D reconstruction of fast moving objects could contribute to various applications such as body structure analysis, accident avoidance, and so on. In this paper, we propose a technique based on a one-shot scanning method, which reconstructs 3D shapes for each frame of a high frame-rate video capturing the scenes projected by a static pattern. To avoid instability of image processing, we restrict the number of colors used in the pattern to less than two. The proposed technique comprises (1) an efficient algorithm to eliminate ambiguity of projected parallel-line patterns by using intersection points, (2) a batch reconstruction algorithm of multiple frames by using spatio-temporal constraints, and (3) an efficient detection method of color-encoded grid pattern based on de Bruijn sequence. In the experiments, the line detection algorithm worked effectively and the dense reconstruction algorithm produces accurate and robust results. We also show the improved results by using temporal constraints. Finally, the dense reconstructions of fast moving objects in a high frame-rate video are presented.
Ryusuke Sagawa, Ryo Furukawa 0001, Hiroshi Kawasaki
IEEE Trans. Pattern Anal. Mach. Intell.3
2013 Robust and Accurate One-Shot 3D Reconstruction by 2C1P System with Wave Grid Pattern
abstract
In this paper, we propose an active 3D reconstruction method with two cameras and one projector (2C1P) system for capturing moving objects. The system reconstructs the shapes from a single frame of each camera by finding the correspondence between the cameras and the projector Based on projecting wave grid pattern. The projected pattern gives the constraint of correspondence between the two cameras in addition to between a projector and a camera. The proposed method finds correspondence by energy minimization on graphs constructed by detecting a grid pattern in camera images. Since the graphs of two cameras are connected as a single graph by using the constraint between cameras, the proposed method simultaneously finds the correspondences for two cameras, which contributes to the robustness of correspondence search. By merging range images created by the correspondence of each camera, we reduce the occluded area compared to the case of one camera. Finally, the proposed method optimizes the shape as three-view stereo to improve the accuracy of shape measurements. In the experiment, we show the effectiveness of using two cameras by making comparison with the case of one camera.
Nozomu Kasuya, Ryusuke Sagawa, Hiroshi Kawasaki, Ryo Furukawa 0001
3DV3
2013 Optimized Aperture for Estimating Depth from Projector's Defocus
abstract
This paper proposes a method for designing a coded aperture, which is installed in a projector for active depth measurement. The aperture design is achieved by genetic algorithm. In this paper, a system for depth measurement using projector and camera system is introduced. The system involves a prototype projector with a coded aperture which is used to reconstruct a shape using Depth from Defocus (DfD) technique. Then, we propose a method to create a coded aperture to achieve the best performance on depth measurement. We also propose an efficient calibration technique on defocus parameters and a cost function for DfD to improve the accuracy and the robustness. Experimental results show that the proposed method can create an aperture pattern using simulation considering noise for fitness calculation. Evaluations are conducted to confirm that the proposed pattern is better than other patterns with our depth measurement algorithm.
Hiroshi Kawasaki, Yuuki Horita, Hitoshi Masuyama, Satoshi Ono, Makoto Kimura, Yasuo Takane
3DV1
2013 Interactive 3D animation system based on touch interface and efficient creation tools
abstract
Recently importance of tablet devices with touch interface increases significantly, because they attract not only mobile phone users, but also people of all generations including small kids to elderly people. One key technology on those devices is 3D graphics; some of them exceed the ability of desktop PCs. Although currently 3D graphics are only used for game applications, it also has a great potential for other purposes. For example, an interactive 3D animation system for tablet devices has been developed recently and attracts many people. With the system, multiple 3D animations are played depending on view direction and user's gesture input. However, the system still has two critical issues: the first one is a complicated process for data creation, and the other is a conflict on several gestures overlapped at the same position. In the paper, the first issue is solved by using a realtime range sensor and the second one is solved by considering the relationship between the 3D object and the input 2D gesture. The effectiveness of the system was proved by making the real system and creating and manipulating the actual 3D animations.
Yasuhiro Akagi, Masayuki Furukawa, Shinya Fukumoto, Yukiko Kawai, Hiroshi Kawasaki
ICME5
2013 Exemplar-Based Hole-Filling Technique for Multiple Dynamic Objects
Matteo Pagliardini, Yasuhiro Akagi, Marcos Slomp, Ryo Furukawa 0001, Ryusuke Sagawa, Hiroshi Kawasaki
PSIVT6
2013 Single colour one-shot scan using modified Penrose tiling pattern
abstract
In this study, the authors propose a new technique to achieve one‐shot scan using single colour and static pattern projector; such a method is ideal for acquisition of moving objects. Since projector–camera systems generally have uncertainties on retrieving correspondences between the captured image and the projected pattern, many solutions have been proposed. Especially for one‐shot scan, which means that only a single image is required for shape reconstruction, positional information of a pixel of the projected pattern should be encoded by spatial and/or colour information. Although colour information is frequently used for encoding, it is severely affected by texture and material of the object and leads unstable reconstruction. In this study, the authors propose a technique to solve the problem by using geometrically unique pattern only with black and white colour that further considers the shape distortion by surface orientation of the shape. The authors technique successfully acquires high‐precision one‐shot scan with an actual system.
Hiroshi Kawasaki, Hitoshi Masuyama, Ryusuke Sagawa, Ryo Furukawa 0001
IET Comput. Vis.1
2012 Structured light with coded aperture for wide range 3D measurement
abstract
In this paper, we propose an active 3D measurement method using structured light system, which can measure wide range of depth. General structured light system requires correspondence between projection patterns and camera observed patterns. Hence, both the projected pattern and the camera image should be in focus on the target. This condition makes a severe limitation on depth range of 3D measurement. Our technique resolves the range limitation by using coded aperture (CA) on projector. It can be understood as a structured light system using Depth from Defocus (DfD) technique, in which defocus of projected light pattern is utilized. By allowing blurry pattern of projection, the measurement range is extended compared to common structured light system. Moreover, CA efficiently improves its accuracy.
Hiroshi Kawasaki, Yuuki Horita, Hiroki Morinaga, Yuuki Matugano, Satoshi Ono, Makoto Kimura, Yasuo Takane
ICIP1
2012 Coded aperture for projector and camera for robust 3D measurement
Yuuki Horita, Yuuki Matugano, Hiroki Morinaga, Hiroshi Kawasaki, Satoshi Ono, Makoto Kimura, Yasuo Takane
ICPR4
2011 Dense one-shot 3D reconstruction by detecting continuous regions with parallel line projection
abstract
3D scanning of moving objects has many applications, for example, marker-less motion capture, analysis on fluid dynamics, object explosion and so on. One of the approach to acquire accurate shape is a projector-camera system, especially the methods that reconstructs a shape by using a single image with static pattern is suitable for capturing fast moving object. In this paper, we propose a method that uses a grid pattern consisting of sets of parallel lines. The pattern is spatially encoded by a periodic color pattern. While informations are sparse in the camera image, the proposed method extracts the dense (pixel-wise) phase informations from the sparse pattern. As the result, continuous regions in the camera images can be extracted by analyzing the phase. Since there remain one DOF for each region, we propose the linear solution to eliminate the DOF by using geometric informations of the devices, i.e. epipolar constraint. In addition, solution space is finite because projected pattern consists of parallel lines with same intervals, the linear equation can be efficiently solved by integer least square method. In this paper, the formulations for both single and multiple projectors are presented. We evaluated the accuracy of correspondences and showed the comparison with respect to the number of projectors by simulation. Finally, the dense 3D reconstruction of moving objects are presented in the experiments.
Ryusuke Sagawa, Hiroshi Kawasaki, Shota Kiyota, Ryo Furukawa 0001
ICCV2
2011 Dynamic Compression of Curve-Based Point Cloud
Ismaël Daribo, Ryo Furukawa 0001, Ryusuke Sagawa, Hiroshi Kawasaki, Shinsaku Hiura, Naoki Asada
PSIVT (2)4
2011 Point cloud compression for grid-pattern-based 3D scanning system
abstract
Recently it is relatively easy to produce digital point sampled 3D geometric models. In sight of the increasing capability of 3D scanning systems to produce models with millions of points, compression efficiency is of paramount importance. In this paper, we propose a novel competition-based predictive method for single-rate compression of 3D models represented as point cloud. In particular we aim at 3D scanning methods based on grid pattern. The proposed method takes advantage of the pattern characteristic made of vertical and horizontal lines, by assuming that the object surface is sampled in curve of points. We then designed and implemented a predictive coder driven by this curve-based point representation. Novel prediction techniques are specifically designed for a curve-based cloud of points, and been competing between them to achieve high quality 3D reconstruction. Experimental results demonstrate the effectiveness of the proposed method.
Ismaël Daribo, Ryo Furukawa 0001, Ryusuke Sagawa, Hiroshi Kawasaki, Shinsaku Hiura, Naoki Asada
VCIP4
2010 Video Deblurring and Super-Resolution Technique for Multiple Moving Objects
Takuma Yamaguchi, Hisato Fukuda, Ryo Furukawa 0001, Hiroshi Kawasaki, Peter F. Sturm
ACCV (4)4
2009 Super-Resolution of Multiple Moving 3D Objects with Pixel-Based Registration
Takuma Yamaguchi, Hiroshi Kawasaki, Ryo Furukawa 0001, Toshihiro Nakayama
ACCV (3)2
2009 Dense 3D reconstruction method using a single pattern for fast moving object
abstract
Dense 3D reconstruction of extremely fast moving objects could contribute to various applications such as body structure analysis and accident avoidance and so on. The actual cases for scanning we assume are, for example, acquiring sequential shape at the moment when an object explodes, or observing fast rotating turbine's blades. In this paper, we propose such a technique based on a one-shot scanning method that reconstructs 3D shape from a single image where dense and simple pattern are projected onto an object. To realize dense 3D reconstruction from a single image, there are several issues to be solved; e.g. instability derived from using multiple colors, and difficulty on detecting dense pattern because of influence of object color and texture compression. This paper describes the solutions of the issues by combining two methods, that is (1) an efficient line detection technique based on de Bruijn sequence and belief propagation, and (2) an extension of shape from intersections of lines method. As a result, a scanning system that can capture an object in fast motion has been actually developed by using a high-speed camera. In the experiments, the proposed method successfully captured the sequence of dense shapes of an exploding balloon, and a breaking ceramic dish at 300–1000 fps.
Ryusuke Sagawa, Yuichi Ota, Yasushi Yagi, Ryo Furukawa 0001, Naoki Asada, Hiroshi Kawasaki
ICCV6
2009 Tour recommendation system based on web information and GIS
abstract
Information recommendation and filtering techniques have been studied intensively. Traditional tour recommendation systems, which can be considered one of information recommendation systems, usually calculate the shortest path in terms of time or distance. Recently, tour recommendation systems for more general purposes have become an important research topic. In this paper we propose an efficient tourist route search system which not only recommends the path simply connecting several tourist spots, but also recommends the path with beautiful scenic sights. We focus on the visibility of scenic sights between one tourist spot and another, which is an important factor for choosing a driving route, but has not been considered in traditional tour recommendation systems. To automatically retrieve tourist spots, we propose a personalized tourist spot recommendation technique using the Web information. Although, for some regions, databases of the famous spots exist and are published, such regions are limited and usually outdated. Our method automatically extracts spots from the Web, thus our system is versatile and up-to-date for large regions. To find a route with attractive scenery, we calculate scores for paths based on the visibility of scenic sights. After generating route candidates using GIS, a 3D virtual space is constructed and the Z-Buffer method is used to decide the visibility of scenic sights for each route candidate. We implemented a prototype and tested the effectiveness of the system.
Yukiko Kawai, Jianwei Zhang 0002, Hiroshi Kawasaki
ICME3
2009 An assurance method for functional system expansion of Tokyo metropolitan railway system
abstract
Autonomous decentralized system (ADS) architecture has been used mainly to realize high assurance control system. However, ADS architecture proposed previously is too simple to apply large-scale complicated system. In this paper, we clarify that large-scale ADS is represented by hierarchical three layers structure and many ADS sub-systems are connected by gateways and agents using railway transport operation control system (ATOS) as an example. ATOS is a system which has been growing and changing continuously. We show that both gateway and agent are effective for system assurance in functional system expansion for large-scale ADS. ATOS: Autonomous Decentralized Transport Operation Control System.
Keiji Kamijyo, Hiroshi Kawasaki, Taichi Arisawa, Keisuke Bekki, Hideki Osumi, Daisuke Yagyu
ISADS2
2009 Laser range scanner based on self-calibration techniques using coplanarities and metric constraints
Ryo Furukawa 0001, Hiroshi Kawasaki
Comput. Vis. Image Underst.2
2009 Shape Reconstruction and Camera Self-Calibration Using Cast Shadows and Scene Geometries
Hiroshi Kawasaki, Ryo Furukawa 0001
Int. J. Comput. Vis.1
2008 Dynamic scene shape reconstruction using a single structured light pattern
abstract
3D acquisition techniques to measure dynamic scenes and deformable objects with little texture are extensively researched for applications like the motion capturing of human facial expression. To allow such measurement, several techniques using structured light have been proposed. These techniques can be largely categorized into two types. The first involves techniques to temporally encode positional information of a projector’s pixels using multiple projected patterns, and the second involves techniques to spatially encode positional information into areas or color spaces. Although the former allows dense reconstruction with a sufficient number of patterns, it has difficulty in scanning objects in rapid motion. The latter technique uses only a single pattern, so this problem can be resolved, however, it often uses complex patterns or color intensities, which are weak to noise, shape distortions, or textures. Thus, it remains an open problem to achieve dense and stable 3D acquisition in real cases. In this paper, we propose a technique to achieve dense shape reconstruction that requires only a single-frame image of a grid pattern. The proposed technique also has the advantage of being robust in terms of image processing.
Hiroshi Kawasaki, Ryo Furukawa 0001, Ryusuke Sagawa, Yasushi Yagi
CVPR1
2008 One-shot range scanner using coplanarity constraints
abstract
Methods for scanning dynamic scenes are important in many applications and many systems using structured light have been proposed. Many of these systems use either multiple patterns projected rapidly or a single pattern. Although the former allows dense reconstruction with a sufficient number of patterns, it has difficulty in capturing objects in rapid motion. The latter technique uses only a single pattern and have no such difficulties, however, they often have stability problems and their result tend to have low resolution. In this paper, we develop a system to achieve dense and accurate 3D measurement from only a single image. The proposed system also has the advantage of being robust in terms of image processing.
Ryo Furukawa 0001, Huynh Quang Huy Viet, Hiroshi Kawasaki, Ryusuke Sagawa, Yasushi Yagi
ICIP3
2008 Efficient meta-information annotation and view-dependent representation system for 3D objects on the Web
abstract
So far, there have been many techniques and applications developed for handling 3D contents efficiently on the Web. However, there is currently no application that successfully makes 3D contents pervasive on the Web. One of the main reasons for this can be considered be the fact that there are only a few pages on the Web where a 3D object is sufficiently annotated with meta-information; meta-information is fundamental for the Web. In this paper, (1) a semi-automatic annotation module which supports users in creating 3D contents with rich meta-information by collecting and filtering meta-data on the Web using 3D information; and (2) a view-dependent information representation module which shows information corresponding to the viewing direction of 3D objects are presented. In those modules, view-dependent information representation module is most important in our system, because such system is highly suitable to show meta-information to users, however, not proposed yet. We implemented the proposed system and conducted experiments which show that it is possible to realize semi-automatic annotation of meta-data on 3D objects and efficiently present view-dependent meta-information to users.
Yukiko Kawai, Shogo Tazawa, Ryo Furukawa 0001, Hiroshi Kawasaki
ICPR4
2008 Real-time image-based rendering system for virtual city based on image compression technique and eigen texture method
abstract
Computer modeling of a large-scale scene such as a city becomes an important topic for computer vision and computer graphics research areas etc. Image-based rendering (IBR) is an effective method for expressing realistic scene, and can construct any arbitrary viewpoint by using the captured real images. However, the large size of the image database in IBR causes serious problems in actual applications, leading to the use of compression techniques. We propose a compression technique based on eigen space combined with a block matching technique to get better result. We also propose a technique to restore the compressed data on Graphic Processing Unit (GPU), allowing us to perform high-speed rendering without raising the load on the CPU.
Ryo Sato, Shintaro Ono, Hiroshi Kawasaki, Katsushi Ikeuchi
ICPR3
2008 The Schemes to Develop Dependable System Using COTS
abstract
This paper provides general design schemes for the dependable system especially long-term projects using commercial off the shelf (COTS) computers based on the concept we established for dependability. As computer technology has shown significantly rapid pace of evolution, it is necessary to take the evolution into consideration when we develop computer systems. Especially in a large-scale project, the evolutions in computer technology during its term are often seen; thus, it is necessary to establish simply applicable schemes as basic concepts to develop dependable system. Basic considerations are presented here by extruding its requirements for dependable systems from wide area railway systems requiring on-time traffic operation and safety under high-density traffic. Many dependable control systems have already been delivered based on autonomous decentralized concept using COTS systems and been showing their sufficient dependability.
Koji Tomita, Kazunori Fujiwara, Hiroshi Kawasaki, Naoki Miwa, Satoru Nagai
PRDC3
2007 Improved Space Carving Method for Merging and Interpolating Multiple Range Images Using Information of Light Sources of Active Stereo
Ryo Furukawa 0001, Tomoya Itano, Akihiko Morisaka, Hiroshi Kawasaki
ACCV (2)4
2007 Shape Reconstruction from Cast Shadows Using Coplanarities and Metric Constraints
Hiroshi Kawasaki, Ryo Furukawa 0001
ACCV (2)1
2007 Geometrical Constraint Based 3D Reconstruction using Implicit Coplanarities
abstract
Coplanarity is a relationship of a set of points that exist on a single plane. Coplanarities can be easily observed in a scene with planer surfaces, and these types of coplanarities have been widely used for 3D reconstructions based on geometrical constraints. Other types of coplanarities that can be observed from images are those observed as cross sections of planes and scenes; for example, points lit by a line laser, or boundary points of a shadow of a straight edge. Although these types of coplanarities have been implicitly used in variations of light sectioning methods, they have not been used in an unified manner with the former types. In this paper, we describe a new 3D reconstruction method based on coplanarities and other geometrical constraints. In particular, we make use of the above two types of coplanarities in an unified manner. This enables us to reconstruct 3D scenes scanned using line lasers or shadows of straight edges observed by a partially-calibrated single camera utilizing geometrical relationships between the planes in the scenes and the planes of line lasers or the planes of shadow boundaries. 1
Ryo Furukawa 0001, Hiroshi Kawasaki
BMVC2
2007 The separation of reflected and transparent layers from real-world image sequence
Thanda Oo, Hiroshi Kawasaki, Yutaka Ohsawa, Katsushi Ikeuchi
Mach. Vis. Appl.2
2006 Dense 3D Reconstruction with an Uncalibrated Active Stereo System
Hiroshi Kawasaki, Yutaka Ohsawa, Ryo Furukawa 0001, Yasuaki Nakamura
ACCV (2)1
2006 Separation of Reflection and Transparency Using Epipolar Plane Image Analysis
Thanda Oo, Hiroshi Kawasaki, Yutaka Ohsawa, Katsushi Ikeuchi
ACCV (1)2
2005 Uncalibrated active stereo for wide and dense 3D data acquisition
abstract
In this paper, we propose an uncalibrated, multi-image 3D reconstruction technique, using coded structured light. Normally, a conventional coded structured light system consists of a camera and a projector and needs precalibration before scanning. Since the camera and the projector have to be fixed after calibration, reconstruction of a wide area of the scene or reducing occlusions are difficult and sometimes impossible. In the proposed method, precalibration can be successfully omitted by applying the uncalibrated stereo technique, thereby multiple scanning while moving the camera or the projector is possible. As the result, users can freely move either the cameras or projectors to scan a wide range of objects.
Hiroshi Kawasaki, Ryo Furukawa 0001
ICIP (3)1
2005 Patch-based BTF synthesis for real-time rendering
abstract
In this paper, we propose a novel synthesis technique for BTFs. A BTF (bidirectional texture function) is a 6D function which can represent appearances of a texture under arbitrary view and lighting conditions. Until now, several approaches of BTF synthesis have been researched. For ordinary textures, patch based methods are promising techniques for texture synthesis. However, it has not been effectively tried to BTFs yet. This is mainly because data size of BTFs is so large and it is not easy to apply the techniques to BTFs. Further, efficient rendering of BTFs is still under research. In this paper, we extend a patch based synthesis technique for BTFs by utilizing compact representation of BTFs. In addition, we propose a novel technique to render the synthesized BTFs efficiently. With our proposed method, we can successfully and effectively render objects with complicated surfaces under arbitrary sizes.
Hiroshi Kawasaki, Kyoung-Dae Seo, Yutaka Ohsawa, Ryo Furukawa 0001
ICIP (1)1
2005 Driving View Simulation Synthesizing Virtual Geometry and Real Images in an Experimental Mixed-Reality Traffic Space
abstract
We propose an efficient and effective image generation system for an experimental mixed-reality traffic space. Our enhanced traffic/driving simulation system represents the view through a hybrid that combines virtual geometry with real images to realize high photo-reality with little human cost. Images for datasets are captured from the real world, and the view for the simulation system is created by synthesizing image datasets - with a conventional driving simulator.
Shintaro Ono, Koichi Ogawara, Masataka Kagesawa, Hiroshi Kawasaki, Masaaki Onuki, Ken Honda, Katsushi Ikeuchi
ISMAR4
2004 Music Compression System Using the GA
Hiroshi Kawasaki, Yasue Mitsukura, Kensuke Mitsukura, Minoru Fukumi, Norio Akamatsu
KES1
2004 Constructing Virtual Cities by Using Panoramic Images
Katsushi Ikeuchi, Masao Sakauchi, Hiroshi Kawasaki, Imari Sato
Int. J. Comput. Vis.3
2001 Light Field Rendering for Large-Scale Scenes
abstract
In this paper, we present an efficient method to synthesize large-scale scenes, such as broad city landscapes. To date, model based approaches have mainly been adopted for this purpose, and some fairly convincing polygon cities have been successfully generated. However, the shapes of real world objects are usually very complicated and it is infeasible to model an entire city realistically. On the other hand, image based approaches have been attempted only recently. Image based methods are effective for realistic rendering, but their huge data sets and restrictions on interactivity pose serious problems for an actual application. Thus, we propose a hybrid method, which uses simple shapes such as planes to model the city, and applies image based techniques to add realism. It can be performed automatically through a simple image capturing process. Further, we also analyze the relationship between error and number of needed images to reduce the data size.
Hiroshi Kawasaki, Katsushi Ikeuchi, Masao Sakauchi
CVPR (2)1
2001 Image-based rendering for mixed reality
abstract
In this paper, we propose an image-based approach to synthesize a novel view image for mixed reality (MR) systems. Theoretically, the image-based method is good for synthesizing realistic images, but it is difficult to achieve interactive handling of the object. As a solution, we propose a new method based on the "surface light field rendering" technique. With this method, we can synthesize the objects with arbitrary deformation and illumination changes. To demonstrate the efficiency of this method, we describe successful experiments that we performed using objects with non-rigid effects (e.g. velvet and Tatami carpet) which are difficult to render correctly by the use of general model-based rendering techniques.
Hiroshi Kawasaki, Hiroyuki Aritaki, Katsushi Ikeuchi, Masao Sakauchi
ICIP (3)1
2000 Spatio-Temporal Analysis of Omni Image
abstract
This paper describes an efficient method to obtain 3D information by using spatio-temporal analysis of omni images for outdoor navigation and map-making in the intelligent transportation system (ITS) application. Two types of omni-directional cameras are employed to make a spatio-temporal volume, which is a sequence of omni images stacked in the spatio-temporal space. For the spatio-temporal analysis of an omni image, we define several different cross sections in such spatio-temporal volumes, and examine characteristics of the traces of image features on the cross sections. We determine that the vertical straight lines in the real world are preserved as straight lines on these cross sections and that the degree of this slope represents the quotient of the velocity of the camera motion and the depth of the object. To acquire 3D information using these characteristics, we propose a hybrid method of the epipolar-plane image (EPI) analysis and the model-based analysis. To demonstrate the effectiveness of this method, we present some experimental results and the ITS applications using an omni-directional video camera to obtain images in outdoor environments.
Hiroshi Kawasaki, Katsushi Ikeuchi, Masao Sakauchi
CVPR1
2000 Arbitrary View Position and Direction Rendering for Large-Scale Scenes
abstract
This paper presents a new method for rendering views, especially those of large-scale scenes, such as broad city landscapes. The main contribution of our method is that we are able to easily render any view from an arbitrary point to an arbitrary direction on the ground in a virtual environment. Our method belongs to the family of work that employs plenoptic functions; however, unlike other works of this type, this particular method allows us to render a novel view from almost any point on the plane at which images are taken. Previous methods, on the other hand, have some restraints concerning their re-constructable area. Thus, when synthesizing a large-scale virtual environment such as a city, our method has a great advantage. One of the applications of our method is a driving simulator in the ITS domain. We can generate any view on any lane on the road from images taken by running along just one lane. Our method, using an omni-directional camera or a measuring device of a similar type, first captures panoramic images by running along a straight line, recording the capturing position of each image. When rendering, the method divides the stored panoramic images into vertical slits, selects some suitable ones based on our theory, and reassembles them for generating an image. The method can make a virtual city with walk-through capabilities. In that virtual city, people can move and look rather freely. In this paper, we describe the basic theory of a new plenoptic function, analyze the applicable areas of the theory and the characteristics of generated images, and demonstrate a complete working system using both indoor and outdoor scenes.
Takuji Takahashi, Hiroshi Kawasaki, Katsushi Ikeuchi, Masao Sakauchi
CVPR2
2000 EPI Analysis of Omni-Camera Image
abstract
The paper describes an efficient method to obtain 3D information from omni-camera images using epipolar-plane image (EPI) analysis. Two types of omni cameras are employed to make a spatio-temporal volume, which is a sequence of omni images stacked in the spatio-temporal space. For the EPI analysis of omni image, we examine different types of cross sections in such spatio-temporal volumes. To conduct the EPI analysis realisticly, we must find out the cross section on which the vertical straight lines in the real world are preserved as straight lines. We define such a cross section as omni EPI. To acquire 3D information using the characteristics of the omni EPI, we propose model based EPI analysis. To demonstrate the effectiveness of this method, we present experimental results using an omni video of outdoor environments.
Hiroshi Kawasaki, Katsushi Ikeuchi, Masao Sakauchi
ICPR1
2000 Expanding Possible View Points of Virtual Environment Using Panoramic Images
abstract
Presents a method for creating a 3D virtual broad city environment with walk-through systems based on image-based rendering (lBR). In that virtual city, people can move rather freely and look at arbitrary views. The strength of our method is that we are able to easily render any view from an arbitrary point to an arbitrary direction on the ground in a virtual environment; previous methods, on the other hand, have strong restraints concerning their reconstructable areas. One of the other applications of our method is a driving simulator in the ITS domain. We can generate any view on any lane on the road from images taken by running along just one lane. Our method first captures panoramic images running along a straight line, indexing the capturing position of each image. The rendering process consists of selecting some suitable slits divided vertically from stored images, and reassembling them to create an image from a novel observation point.
Takuji Takahashi, Hiroshi Kawasaki, Katsushi Ikeuchi, Masao Sakauchi
ICPR2
1999 Automatic modeling of a 3D city map from real-world video
abstract
Mixed reality (MR) systems which integrate the virtual world and the real world have become a major topic in the research area of multimedia. As a practical application of these MR systems, we propose an efficient method for making a 3D map from real-world video data. The proposed method is an automatic organization method focusing on video objects to describe video data in an efficient way, i.e., by collating the real-world video data with map information using DP matching. To demonstrate the reliability of this method, we describe successful experiments that we performed using 3D information obtained from the real-world video data.
Hiroshi Kawasaki, Tomoyuki Yatabe, Katsushi Ikeuchi, Masao Sakauchi
ACM Multimedia (1)1