Masatoshi Okutomi

dblp:17/5585 · DBLP profile ↗
← Back
138ranked-venue papers
10as first author
41since 2021 · last 2026
0000-0001-5787-0742ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 108 · 6 first-author · 29 since 2021Artificial intelligence and machine learning · 92 · 10 first-author · 19 since 2021Systems, architecture and hardware · 10 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Joint 2D-3D Segmentation and Association in Street-Level Imaging
Amir Melnikov, Masayuki Tanaka 0001, Yusuke Monno, Masatoshi Okutomi
ICPR (9)4
2026 Multi-camera Multi-object Tracking Based on Epipolar Distance and Appearance Similarity
Masamune Oka, Masayuki Tanaka 0001, Takashi Shibata 0001, Masatoshi Okutomi
ICPR (4)4
2026 Dashcam-Based 3D Building Generation and Augmentation to Open Geospatial Platform
Masahiko Murakami, Hagad Juan Lorenzo, Eiichiro Yoshioka, Kenichi Yamashita, Kenji Akita, Daisuke Fujioka, Akira Mimura, Osamu Sugahara, Keiichi Matsubara, Seiichi Kataoka, Yusuke Monno, Masatoshi Okutomi
IV13
2026 Reflection removal using recurrent polarization-to-polarization network
abstract
This paper addresses reflection removal, which is the task of separating reflection components from a captured image and deriving the image with only transmission components. Considering that the existence of the reflection changes the polarization state of a scene, some existing methods have exploited polarized images for reflection removal. While these methods apply polarized images as the inputs, they predict the reflection and the transmission directly as non-polarized intensity images. In contrast, we propose a polarization-to-polarization approach that applies polarized images as the inputs and predicts “polarized”reflection and transmission images using two sequential networks to facilitate the separation task by utilizing the interrelated polarization information between the reflection and the transmission. We further adopt a recurrent framework, where the predicted reflection and transmission images are used to iteratively refine each other. To address the lack of a color-polarized image dataset for reflection removal training, we propose a physics-based synthetic dataset generation pipeline designed to produce color-polarized images with reflections. Additionally, to evaluate the generalization capability for real-world scenes, we introduce a new real-world color test dataset captured using a polarization camera. Experimental results on existing grayscale and our color datasets demonstrate that our method outperforms other state-of-the-art approaches.
Wenjiao Bian, Yusuke Monno, Masatoshi Okutomi
Mach. Vis. Appl.3
2025 AP-DPM: A Dual-Path Merging Network Via Adversarial Anatomical Prior Guidance for Wrist Bone Segmentation
abstract
Accurate segmentation of wrist bones from conventional radiographs remains a significant challenge due to severe anatomical overlap and blurred bone boundaries, particularly in patients with rheumatoid arthritis. To address these issues, we propose AP-DPM, a novel Dual-Path Prior-constrained Merging model with adversarial anatomical priors. AP-DPM employs a dual-path architecture to separately predict complete bone masks and overlapping regions, which are subsequently integrated through a residual merging network. To enhance anatomical plausibility, we incorporate an adversarial prior guided by a pre-trained discriminator. Extensive experiments on the publicly available RAM-W600 dataset demonstrate that AP-DPM outperforms state-of-the-art methods across seven quantitative metrics, achieving superior performance particularly in diagnostically critical overlapping regions. Ablation studies further validate that both the dualpath structure for enhanced focusing on overlap regions and the adversarial anatomical prior contribute significantly to performance gains, enhancing local boundary sensitivity and global structural consistency. These results highlight the potential of AP-DPM to improve automated radiographic assessment of rheumatoid arthritis progression. Code is available at https://github.com/YSongxiao/AP-DPM
Songxiao Yang, Haolin Wang 0007, Masayuki Ikebe, Tamotsu Kamishima, Yafei Ou, Masatoshi Okutomi
BIBM6
2025 Multi-Class Smoothed Hinge Loss Function in Pre-Training for Transfer Learning
abstract
Multi-class classification is essential in various machine learning applications, but it often suffers from overfitting when using cross-entropy (CE) loss with the softmax function. A key limitation of CE loss is that it cannot reach zero, even if the model’s output for the target class approaches infinity, leading to models that become overly sensitive to training data. We address these issues by proposing a novel multi-class smoothed hinge loss function that sets a threshold, where the network score over the threshold does not change the loss. Our method is particularly effective in transfer learning and outperforms traditional approaches, achieving outstanding post-transfer accuracy and flatter loss landscapes. Both theoretical and empirical analyses validate the effectiveness of our approach in producing high-quality pre-trained weights for transfer learning.
Wonjik Kim, Masayuki Tanaka 0001, Masatoshi Okutomi, Hirokazu Nosato
ICIP3
2025 Polarization Denoising and Demosaicking: Dataset and Baseline Method
abstract
A division-of-focal-plane (DoFP) polarimeter enables us to acquire images with multiple polarization orientations in one shot and thus it is valuable for many applications using polarimetric information. The image processing pipeline for a DoFP polarimeter entails two crucial tasks: denoising and demosaicking. While polarization demosaicking for a noise-free case has increasingly been studied, the research for the joint task of polarization denoising and demosaicking is scarce due to the lack of a suitable evaluation dataset and a solid baseline method. In this paper, we propose a novel dataset and method for polarization denoising and de-mosaicking. Our dataset contains 40 real-world scenes and three noise-level conditions, consisting of pairs of noisy mosaic inputs and noise-free full images. Our method takes a denoising-then-demosaicking approach based on well-accepted signal processing components to offer a reproducible method. Experimental results demonstrate that our method exhibits higher image reconstruction performance than other alternative methods, offering a solid baseline.
Muhamad Daniel Ariff Bin Abdul Rahman, Yusuke Monno, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP4
2025 TDM: Temporally-Consistent Diffusion Model for All-in-One Real-World Video Restoration
Zihua Liu, Yusuke Monno, Masatoshi Okutomi
MMM (4)4
2025 RAM-W600: A Multi-Task Wrist Dataset and Benchmark for Rheumatoid Arthritis
abstract
Rheumatoid arthritis (RA) is a common autoimmune disease that has been the focus of research in computer-aided diagnosis (CAD) and disease monitoring. In clinical settings, conventional radiography (CR) is widely used for the screening and evaluation of RA due to its low cost and accessibility. The wrist is a critical region for the diagnosis of RA. However, CAD research in this area remains limited, primarily due to the challenges in acquiring high-quality instance-level annotations. (i) The wrist comprises numerous small bones with narrow joint spaces, complex structures, and frequent overlaps, requiring detailed anatomical knowledge for accurate annotation. (ii) Disease progression in RA often leads to osteophyte, bone erosion (BE), and even bony ankylosis, which alter bone morphology and increase annotation difficulty, necessitating expertise in rheumatology.This work presents a multi-task dataset for wrist bone in CR, including two tasks: (i) wrist bone instance segmentation and (ii) Sharp/van der Heijde (SvdH) BE scoring, which is the first public resource for wrist bone instance segmentation. This dataset comprises 1048 wrist conventional radiographs of 388 patients from six medical centers, with pixel-level instance segmentation annotations for 618 images and SvdH BE scores for 800 images. This dataset can potentially support a wide range of research tasks related to RA, including joint space narrowing (JSN) progression quantification, BE detection, bone deformity evaluation, and osteophyte detection. It may also be applied to other wrist-related tasks, such as carpal bone fracture localization.We hope this dataset will significantly lower the barrier to research on wrist RA and accelerate progress in CAD research within the RA-related domain.Benchmark & Code: https://github.com/YSongxiao/RAM-W600Data & Dataset Card: https://huggingface.co/datasets/TokyoTechMagicYang/RAM-W600
Songxiao Yang, Haolin Wang 0007, Tamotsu Kamishima, Masayuki Ikebe, Yafei Ou, Masatoshi Okutomi
NeurIPS8
2024 Disparity Estimation Using a Quad-Pixel Sensor
Zhuofeng Wu 0003, Doehyung Lee, Zihua Liu, Kazunori Yoshizaki, Yusuke Monno, Masatoshi Okutomi
BMVC6
2024 VSRD: Instance-Aware Volumetric Silhouette Rendering for Weakly Supervised 3D Object Detection
abstract
Monocular 3D object detection poses a significant challenge in 3D scene understanding due to its inherently illposed nature in monocular depth estimation. Existing methods heavily rely on supervised learning using abundant 3D labels, typically obtained through expensive and laborintensive annotation on LiDAR point clouds. To tackle this problem, we propose a novel weakly supervised 3D object detection framework named VSRD (Volumetric Silhouette Rendering for Detection) to train 3D object detectors without any 3D supervision but only weak 2D supervision. VSRD consists of multi-view 3D auto-labeling and subsequent training of monocular 3D object detectors using the pseudo labels generated in the auto-labeling stage. In the auto-labeling stage, we represent the surface of each instance as a signed distance field (SDF) and render its silhouette as an instance mask through our proposed instanceaware volumetric silhouette rendering. To directly optimize the 3D bounding boxes through rendering, we decompose the SDF of each instance into the SDF of a cuboid and the residual distance field (RDF) that represents the residual from the cuboid. This mechanism enables us to optimize the 3D bounding boxes in an end-to-end manner by comparing the rendered instance masks with the ground truth instance masks. The optimized 3D bounding boxes serve as effective training data for 3D object detection. We conduct extensive experiments on the KITTI-360 dataset, demonstrating that our method outperforms the existing weakly supervised 3D object detection methods. The code is available at https://github.com/skmhrk1209/VSRD.
Zihua Liu, Hiroki Sakuma, Masatoshi Okutomi
CVPR3
2024 Self-Supervised Spatially Variant PSF Estimation for Aberration-Aware Depth-from-Defocus
abstract
In this paper, we address the task of aberration-aware depth-from- defocus (DfD), which takes account of spatially variant point spread functions (PSFs) of a real camera. To effectively obtain the spatially variant PSFs of a real camera without requiring any ground-truth PSFs, we propose a novel self-supervised learning method that leverages the pair of real sharp and blurred images, which can be easily captured by changing the aperture setting of the camera. In our PSF estimation, we assume rotationally symmetric PSFs and introduce the polar coordinate system to more accurately learn the PSF estimation network. We also handle the focus breathing phenomenon that occurs in real DfD situations. Experimental results on synthetic and real data demonstrate the effectiveness of our method regarding both the PSF estimation and the depth estimation.
Zhuofeng Wu 0003, Yusuke Monno, Masatoshi Okutomi
ICASSP3
2024 Reflection Removal Using Recurrent Polarization-to-Polarization Network
abstract
This paper addresses reflection removal, which is the task of separating reflection components from a captured image and deriving the image with only transmission components. Considering that the existence of the reflection changes the polarization state of a scene, some existing methods have exploited polarized images for reflection removal. While these methods apply polarized images as the inputs, they predict the reflection and the transmission directly as non-polarized intensity images. In contrast, we propose a polarization-to-polarization approach that applies polarized images as the inputs and predicts "polarized" reflection and transmission images using two sequential networks to facilitate the separation task by utilizing the interrelated polarization information between the reflection and the transmission. We further adopt a recurrent framework, where the predicted reflection and transmission images are used to iteratively refine each other. Experimental results on a public dataset demonstrate that our method outperforms other state-of-the-art methods.
Wenjiao Bian, Yusuke Monno, Masatoshi Okutomi
ICASSP3
2024 Object Detection Framework Using Multiple Tone Mappings on High-Dynamic-Range Images
abstract
In practical computer vision applications, such as autonomous driving, the ability to effectively process high-dynamic-range (HDR) scenes is crucial for safe operation. In this paper, we focus on object detection within HDR images. To address this, we propose a simple yet effective framework that employs multiple tone mappings. First, we generate multiple images from an HDR image with varying tone mapping parameters. Then, those images are fed into a high-performance object detector pre-trained with low-dynamic-range (LDR) images. Multiple detection results are merged with non-maximum suppression (NMS). To assess the performance of our method, we have built a validation dataset comprising HDR images captured in outdoor scenes with significant contrast variations. The experimental results using both our dataset and an existing one demonstrate that our method outperforms existing approaches1.1The code and the dataset can be available at https://open-vision.sc.e.titech.ac.jp/research/hdrdet.
Takumi Watanabe, Rei Kawakami, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP4
2024 Few-Shot View Synthesis Based on Geometric and Semantic Consistency
Mizuki Kojima, Rei Kawakami, Masatoshi Okutomi
ICPR (18)3
2024 CFDNet: A Generalizable Foggy Stereo Matching Network with Contrastive Feature Distillation
abstract
Stereo matching under foggy scenes remains a challenging task since the scattering effect degrades the visibility and results in less distinctive features for dense correspondence matching. While some previous learning-based methods integrated a physical scattering function for simultaneous stereo-matching and dehazing, simply removing fog might not aid depth estimation because the fog itself can provide crucial depth cues. In this work, we introduce a framework based on contrastive feature distillation (CFD). This strategy combines feature distillation from merged clean-fog features with contrastive learning, ensuring balanced dependence on fog depth hints and clean matching features. This framework helps to enhance model generalization across both clean and foggy environments. Comprehensive experiments on synthetic and real-world datasets affirm the superior strength and adapt-ability of our method.
Zihua Liu, Masatoshi Okutomi
ICRA3
2024 Global Occlusion-Aware Transformer for Robust Stereo Matching
abstract
Despite the remarkable progress facilitated by learning-based stereo-matching algorithms, the performance in the ill-conditioned regions, such as the occluded regions, remains a bottleneck. Due to the limited receptive field, existing CNN-based methods struggle to handle these ill-conditioned regions effectively. To address this issue, this paper introduces a novel attention-based stereo-matching network called Global Occlusion-Aware Transformer (GOAT) to exploit long-range dependency and occlusion-awareness global context for disparity estimation. In the GOAT architecture, a parallel disparity and occlusion estimation module (PDO) is proposed to estimate the initial disparity map and the occlusion mask using a parallel attention mechanism. To further enhance the disparity estimates in the occluded regions, an occlusion-aware global aggregation module (OGA) is proposed. This module aims to refine the disparity in the occluded regions by leveraging restricted global correlation within the focus scope of the occluded areas. Extensive experiments were conducted on several public benchmark datasets including SceneFlow [15], KITTI 2015 [16], and Middlebury [19]. The results show that proposed GOAT demonstrates outstanding performance among all benchmarks, particularly in the occluded regions.
Zihua Liu, Masatoshi Okutomi
WACV3
2024 Polarimetric PatchMatch Multi-View Stereo
abstract
PatchMatch Multi-View Stereo (PatchMatch MVS) is one of the popular MVS approaches, owing to its balanced accuracy and efficiency. In this paper, we propose Polarimetric PatchMatch multi-view Stereo (PolarPMS), which is the first method exploiting polarization cues to PatchMatch MVS. The key of PatchMatch MVS is to generate depth and normal hypotheses, which form local 3D planes and slanted stereo matching windows, and efficiently search for the best hypothesis based on the consistency among multi-view images. In addition to standard photometric consistency, our PolarPMS evaluates polarimetric consistency to assess the validness of a depth and normal hypothesis, motivated by the physical property that the polarimetric information is related to the object’s surface normal. Experimental results demonstrate that our PolarPMS can improve the accuracy and the completeness of reconstructed 3D models, especially for texture-less surfaces, compared with state-of-the-art PatchMatch MVS methods.
Jinyu Zhao, Jumpei Oishi, Yusuke Monno, Masatoshi Okutomi
WACV4
2024 Dual-Pixel Raindrop Removal
abstract
Removing raindrops in images has been addressed as a significant task for various computer vision applications. In this paper, we propose the first method using a dual-pixel (DP) sensor to better address raindrop removal. Our key observation is that raindrops attached to a glass window yield noticeable disparities in DP's left-half and right-half images, while almost no disparity exists for in-focus backgrounds. Therefore, the DP disparities can be utilized for robust raindrop detection. The DP disparities also bring the advantage that the occluded background regions by raindrops are slightly shifted between the left-half and the right-half images. Therefore, fusing the information from the left-half and the right-half images can lead to more accurate background texture recovery. Based on the above motivation, we propose a DP Raindrop Removal Network (DPRRN) consisting of DP raindrop detection and DP fused raindrop removal. To efficiently generate a large amount of training data, we also propose a novel pipeline to add synthetic raindrops to real-world background DP images. Experimental results on constructed synthetic and real-world datasets demonstrate that our DPRRN outperforms existing state-of-the-art methods, especially showing better robustness to real-world situations.
Yusuke Monno, Masatoshi Okutomi
IEEE Trans. Pattern Anal. Mach. Intell.3
2024 Deep snapshot HDR imaging using multi-exposure color filter array
Yutaro Okamoto, Masayuki Tanaka 0001, Yusuke Monno, Masatoshi Okutomi
Vis. Comput.4
2023 EMR-MSF: Self-Supervised Recurrent Monocular Scene Flow Exploiting Ego-Motion Rigidity
abstract
Self-supervised monocular scene flow estimation, aiming to understand both 3D structures and 3D motions from two temporally consecutive monocular images, has received increasing attention for its simple and economical sensor setup. However, the accuracy of current methods suffers from the bottleneck of less-efficient network architecture and lack of motion rigidity for regularization. In this paper, we propose a superior model named EMR-MSF by borrowing the advantages of network architecture design under the scope of supervised learning. We further impose explicit and robust geometric constraints with an elaborately constructed ego-motion aggregation module where a rigidity soft mask is proposed to filter out dynamic regions for stable ego-motion estimation using static regions. Moreover, we propose a motion consistency loss along with a mask regularization loss to fully exploit static regions. Several efficient training strategies are integrated including a gradient detachment technique and an enhanced view synthesis process for better performance. Our proposed method outperforms the previous self-supervised works by a large margin and catches up to the performance of supervised methods. On the KITTI scene flow benchmark, our approach improves the SF-all metric of the state-of-the-art self-supervised monocular method by 44% and demonstrates superior performance across sub-tasks including depth and visual odometry, amongst other self-supervised single-task or multi-task methods.
Zijie Jiang, Masatoshi Okutomi
ICCV2
2023 AUAAC: Area Under Accuracy-Accuracy Curve for Evaluating Out-of-Distribution Detection
Wonjik Kim, Masayuki Tanaka 0001, Masatoshi Okutomi
PSIVT3
2023 Local Brightness Normalization for Image Classification and Object Detection Robust to Illumination Changes
Yanshuo Lu, Masayuki Tanaka 0001, Rei Kawakami, Masatoshi Okutomi
PSIVT4
2023 Semantic Segmentation of Degraded Images Using Layer-Wise Feature Adjustor
abstract
Semantic segmentation of degraded images is important for practical applications such as autonomous driving and surveillance systems. The degradation level, which represents the strength of degradation, is usually unknown in practice. Therefore, the semantic segmentation algorithm needs to take account of various levels of degradation. In this paper, we propose a convolutional neural network of semantic segmentation which can cope with various levels of degradation. The proposed network is based on the knowledge distillation from a source network trained with only clean images. More concretely, the proposed network is trained to acquire multi-layer features keeping consistency with the source network, while adjusting for various levels of degradation. The effectiveness of the proposed method is confirmed for different types of degradations: JPEG distortion, Gaussian blur and salt&pepper noise. The experimental comparisons validate that the proposed network outperforms existing networks for semantic segmentation of degraded images with various degradation levels.
Kazuki Endo, Masayuki Tanaka 0001, Masatoshi Okutomi
WACV3
2023 Polarimetric Multi-View Inverse Rendering
abstract
A polarization camera has great potential for 3D reconstruction since the angle of polarization (AoP) and the degree of polarization (DoP) of reflected light are related to an object's surface normal. In this paper, we propose a novel 3D reconstruction method called Polarimetric Multi-View Inverse Rendering (Polarimetric MVIR) that effectively exploits geometric, photometric, and polarimetric cues extracted from input multi-view color-polarization images. We first estimate camera poses and an initial 3D model by geometric reconstruction with a standard structure-from-motion and multi-view stereo pipeline. We then refine the initial model by optimizing photometric rendering errors and polarimetric errors using multi-view RGB, AoP, and DoP images, where we propose a novel polarimetric cost function that enables an effective constraint on the estimated surface normal of each vertex, while considering four possible ambiguous azimuth angles revealed from the AoP measurement. The weight for the polarimetric cost is effectively determined based on the DoP measurement, which is regarded as the reliability of polarimetric information. Experimental results using both synthetic and real data demonstrate that our Polarimetric MVIR can reconstruct a detailed 3D shape without assuming a specific surface material and lighting condition.
Jinyu Zhao, Yusuke Monno, Masatoshi Okutomi
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Pro-Cam SSfM: projector-camera system for structure and spectral reflectance from motion
Yusuke Monno, Masatoshi Okutomi
Vis. Comput.3
2022 Robustizing Object Detection Networks Using Augmented Feature Pooling
Takashi Shibata 0001, Masayuki Tanaka 0001, Masatoshi Okutomi
ACCV (5)3
2022 Dual-Pixel Raindrop Removal
Yusuke Monno, Masatoshi Okutomi
BMVC3
2022 Deep Hyperspectral-Depth Reconstruction Using Single Color-Dot Projection
abstract
Depth reconstruction and hyperspectral reflectance reconstruction are two active research topics in computer vision and image processing. Conventionally, these two topics have been studied separately using independent imaging setups and there is no existing method which can acquire depth and spectral reflectance simultaneously in one shot without using special hardware. In this paper, we propose a novel single-shot hyperspectral-depth reconstruction method using an off-the-shelf RGB camera and projector. Our method is based on a single color-dot projection, which simultaneously acts as structured light for depth reconstruction and spatially-varying color illuminations for hyperspectral reflectance reconstruction. To jointly reconstruct the depth and the hyperspectral reflectance from a single color-dot image, we propose a novel end-to-end network architecture that effectively incorporates a geometric color-dot pattern loss and a photometric hyperspectral reflectance loss. Through the experiments, we demonstrate that our hyperspectral-depth reconstruction method outperforms the combination of an existing state-of-the-art single-shot hyperspectral reflectance reconstruction method and depth reconstruction method.
Yusuke Monno, Masatoshi Okutomi
CVPR3
2022 Optimal Noise-Aware Imaging with Switchable Prefilters
abstract
Most consumer digital cameras employ a single-chip image sensor with a color filter array (CFA), where the purpose of an in-camera imaging pipeline is to generate a noise-free and color-corrected standard RGB image from mosaic CFA RAW data. The joint design of camera spectral sensitivity (CSS) and the imaging pipeline has great potential to derive better imaging quality. However, since there is a trade-off between the robustness to noise and the accuracy of color reproduction, one fixed CSS cannot realize optimal imaging in terms of both aspects under various noise levels. Thus, in this paper, we propose noise-aware imaging using camera prefilters for each noise level, where we jointly design the spectral sensitivity of the prefilters, that of CFA, and imaging networks to realize optimal imaging in all noise levels. Experimental results under various noise levels demonstrate that our imaging method using the prefilters outperforms existing methods based on a fixed CSS.
Zilai Gong, Masayuki Tanaka 0001, Yusuke Monno, Masatoshi Okutomi
ICIP4
2022 Two-Step Color-Polarization Demosaicking Network
abstract
Polarization information of light in a scene is valuable for various image processing and computer vision tasks. A division-of-focal-plane polarimeter is a promising approach to capture the polarization images of different orientations in one shot, while it requires color-polarization demosaicking. In this paper, we propose a two-step color-polarization demosaicking network (TCPDNet), which consists of two sub-tasks of color demosaicking and polarization demosaicking. We also introduce a reconstruction loss in the YCbCr color space to improve the performance of TCPDNet. Experimental comparisons demonstrate that TCPDNet outperforms existing methods in terms of the image quality of polarization images and the accuracy of Stokes parameters.
Vy Nguyen, Masayuki Tanaka 0001, Yusuke Monno, Masatoshi Okutomi
ICIP4
2022 Self-Supervised Ego-Motion Estimation Based on Multi-Layer Fusion of RGB and Inferred Depth
abstract
In existing self-supervised depth and ego-motion estimation methods, ego-motion estimation is usually limited to only leveraging RGB information. Recently, several methods have been proposed to further improve the accuracy of self-supervised ego-motion estimation by fusing information from other modalities, e.g., depth, acceleration, and angular velocity. However, they rarely focus on how different fusion strategies affect performance. In this paper, we investigate the effect of different fusion strategies for ego-motion estimation and pro-pose a new framework for self-supervised learning of depth and ego-motion estimation, which performs ego-motion estimation by leveraging RGB and inferred depth information in a Multi-Layer Fusion manner. As a result, we have achieved state-of-the-art performance among learning-based methods on the KITTI odometry benchmark. Detailed studies on the design choices of leveraging inferred depth information and fusion strategies have also been carried out, which clearly demonstrate the advantages of our proposed framework.3
Zijie Jiang, Hajime Taira, Naoyuki Miyashita, Masatoshi Okutomi
ICRA4
2022 Are Realistic Training Data Necessary for Depth-from-Defocus Networks?
abstract
Image-based depth estimation is one of the important tasks in computer vision. Depth-from-defocus (DfD) methods estimate the scene depth from a single or multiple defocused images by exploiting depth-dependent defocus blur cues. Because of the difficulty in obtaining a real-world dataset with ground-truth scene depth, most deep-learning-based DfD methods rely on a synthetic training dataset, where more realistic scene rendering is considered desirable for more accurate depth estimation. In this paper, we consider if realistic 3D objects are really necessary for training DfD networks. To investigate this, we design a very simple and fast synthetic training data generation method for DfD using only two front-parallel texture planes in one scene and compare it with a widely-applied path-tracing method using a common 3D object dataset. Through real-world experiments, we show that the 2-plane method provides comparable and even slightly better performance than the path-tracing method and can be considered as an alternative method for simple and practical DfD network training.
Zhuofeng Wu 0003, Yusuke Monno, Masatoshi Okutomi
IECON3
2022 Digging Into Normal Incorporated Stereo Matching
abstract
Despite the remarkable progress facilitated by learning-based stereo matching algorithms, disparity estimation in low-texture, occluded, and bordered regions still remain bottlenecks that limit the performance. To tackle these challenges, geometric guidance like plane information is necessary as it provides intuitive guidance about disparity consistency and affinity similarity. In this paper, we propose a normal incorporated joint learning that framework consisting of two specific modules named non-local disparity propagation(NDP) and affinity-aware residual learning(ARL). The estimated normal map is first utilized for calculating a non-local affinity matrix as well as a non-local offset to perform spatial propagation at the disparity level. To enhance geometric consistency, especially in low-texture regions, the estimated normal map is then leveraged to calculate a local affinity matrix which provides the residual learning with information about where the correction should refer and thus improve the residual learning efficiency. Extensive experiments on several public datasets including Scene Flow, KITTI 2015, and Middlebury 2014 validate the effectiveness of our proposed method. By the time we finished this work, our approach ranked 1st for stereo matching across foreground pixels on the KITTI 2015 dataset and 3rd on the Scene Flow dataset among all the published works.
Zihua Liu, Songyan Zhang, Zhicheng Wang 0022, Masatoshi Okutomi
ACM Multimedia4
2022 Single Image Deraining Network with Rain Embedding Consistency and Layered LSTM
abstract
Single image deraining is typically addressed as residual learning to predict the rain layer from an input rainy image. For this purpose, an encoder-decoder network draws wide attention, where the encoder is required to encode a high-quality rain embedding which determines the performance of the subsequent decoding stage to reconstruct the rain layer. However, most of existing studies ignore the significance of rain embedding quality, thus leading to limited performance with over/under-deraining. In this paper, with our observation of the high rain layer reconstruction performance by an rain-to-rain autoencoder, we introduce the idea of "Rain Embedding Consistency" by regarding the encoded embedding by the autoencoder as an ideal rain embedding and aim at enhancing the deraining performance by improving the consistency between the ideal rain embedding and the rain embedding derived by the encoder of the deraining network. To achieve this, a Rain Embedding Loss is applied to directly supervise the encoding process, with a Rectified Local Contrast Normalization (RLCN) as the guide that effectively extracts the candidate rain pixels. We also propose Layered LSTM for recurrent deraining and fine-grained encoder feature refinement considering different scales. Qualitative and quantitative experiments demonstrate that our proposed method outperforms previous state-of-the-art methods particularly on a real-world dataset. Our source code is available at http://www.ok.sc.e.titech.ac.jp/res/SIR/.
Yusuke Monno, Masatoshi Okutomi
WACV3
2022 Long-Term Visual Localization Revisited
abstract
Visual localization enables autonomous vehicles to navigate in their surroundings and augmented reality applications to link virtual to real worlds. Practical visual localization approaches need to be robust to a wide variety of viewing conditions, including day-night changes, as well as weather and seasonal variations, while providing highly accurate six degree-of-freedom (6DOF) camera pose estimates. In this paper, we extend three publicly available datasets containing images captured under a wide variety of viewing conditions, but lacking camera pose information, with ground truth pose information, making evaluation of the impact of various factors on 6DOF camera pose estimation accuracy possible. We also discuss the performance of state-of-the-art localization approaches on these datasets. Additionally, we release around half of the poses for all conditions, and keep the remaining half private as a test set, in the hopes that this will stimulate research on long-term visual localization, learned local image features, and related research areas. Our datasets are available at visuallocalization.net, where we are also hosting a benchmarking server for automatic evaluation of results on the test set. The presented state-of-the-art results are to a large degree based on submissions to our server.
Carl Toft, Will Maddern, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Tomás Pajdla, Fredrik Kahl, Torsten Sattler
IEEE Trans. Pattern Anal. Mach. Intell.7
2021 Spectral MVIR: Joint Reconstruction of 3D Shape and Spectral Reflectance
abstract
Reconstructing an object's high-quality 3D shape with inherent spectral reflectance property, beyond typical device-dependent RGB albedos, opens the door to applications requiring a high-fidelity 3D model in terms of both geometry and photometry. In this paper, we propose a novel Multi-View Inverse Rendering (MVIR) method called Spectral MVIR for jointly reconstructing the 3D shape and the spectral reflectance for each point of object surfaces from multi-view images captured using a standard RGB camera and low-cost lighting equipment such as an LED bulb or an LED projector. Our main contributions are twofold: (i) We present a rendering model that considers both geometric and photometric principles in the image formation by explicitly considering camera spectral sensitivity, light's spectral power distribution, and light source positions. (ii) Based on the derived model, we build a cost-optimization MVIR framework for the joint reconstruction of the 3D shape and the per-vertex spectral reflectance while estimating the light source positions and the shadows. Different from most existing spectral-3D acquisition methods, our method does not require expensive special equipment and cumbersome geometric calibration. Experimental results using both synthetic and real-world data demonstrate that our Spectral MVIR can acquire a high-quality 3D model with accurate spectral reflectance property.
Yusuke Monno, Masatoshi Okutomi
ICCP3
2021 Geometric Data Augmentation Based On Feature Map Ensemble
abstract
Deep convolutional networks have become the mainstream in computer vision applications. Although CNNs have been successful in many computer vision tasks, it is not free from drawbacks. The performance of CNN is dramatically degraded by geometric transformation, such as large rotations. In this paper, we propose a novel CNN architecture that can improve the robustness against geometric transformations without modifying the existing backbones of their CNNs. The key is to enclose the existing backbone with a geometric transformation (and the corresponding reverse transformation) and a feature map ensemble. The proposed method can inherit the strengths of existing CNNs that have been presented so far. Furthermore, the proposed method can be employed in combination with state-of-the-art data augmentation algorithms to improve their performance. We demonstrate the effectiveness of the proposed method using standard datasets such as CIFAR, CUB-200, and Mnist-rot-12k.
Takashi Shibata 0001, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP3
2021 InLoc: Indoor Visual Localization with Dense Matching and View Synthesis
abstract
We seek to predict the 6 degree-of-freedom (6DoF) pose of a query photograph with respect to a large indoor 3D map. The contributions of this work are three-fold. First, we develop a new large-scale visual localization method targeted for indoor spaces. The method proceeds along three steps: (i) efficient retrieval of candidate poses that scales to large-scale environments, (ii) pose estimation using dense matching rather than sparse local features to deal with weakly textured indoor scenes, and (iii) pose verification by virtual view synthesis that is robust to significant changes in viewpoint, scene layout, and occlusion. Second, we release a new dataset with reference 6DoF poses for large-scale indoor localization. Query photographs are captured by mobile phones at a different time than the reference 3D map, thus presenting a realistic indoor localization scenario. Third, we demonstrate that our method significantly outperforms current state-of-the-art indoor localization approaches on this new challenging data. Code and data are publicly available.
Hajime Taira, Masatoshi Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomás Pajdla, Akihiko Torii
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Are Large-Scale 3D Models Really Necessary for Accurate Visual Localization?
abstract
Accurate visual localization is a key technology for autonomous navigation. 3D structure-based methods employ 3D models of the scene to estimate the full 6 degree-of-freedom (DOF) pose of a camera very accurately. However, constructing (and extending) large-scale 3D models is still a significant challenge. In contrast, 2D image retrieval-based methods only require a database of geo-tagged images, which is trivial to construct and to maintain. They are often considered inaccurate since they only approximate the positions of the cameras. Yet, the exact camera pose can theoretically be recovered when enough relevant database images are retrieved. In this paper, we demonstrate experimentally that large-scale 3D models are not strictly necessary for accurate visual localization. We create reference poses for a large and challenging urban dataset. Using these poses, we show that combining image-based methods with local reconstructions results in a higher pose accuracy compared to state-of-the-art structure-based methods, albeight at higher run-time costs. We show that some of these run-time costs can be alleviated by exploiting known database image poses. Our results suggest that we might want to reconsider the need for large-scale 3D models in favor of more local models, but also that further research is necessary to accelerate the local reconstruction process.
Akihiko Torii, Hajime Taira, Josef Sivic, Marc Pollefeys, Masatoshi Okutomi, Tomás Pajdla, Torsten Sattler
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 CNN-Based Classification of Degraded Images With Awareness of Degradation Levels
abstract
Image classification needs to consider the existence of image degradations in practice. Although degraded images have various levels of degradation, the degradation levels are usually unknown. This paper proposes a convolutional neural network to classify degraded images by using a restoration network and an ensemble learning. The proposed network can automatically infer ensemble weights by using estimated degradation levels of degraded images and features of restored images, where the degradation levels are estimated internally. The proposed network is mainly discussed with JPEG distortion, while degradations of both Gaussian noise and blurring are also examined. We demonstrate that the proposed network can classify degraded images over various levels of degradation. This paper also reveals how the image-quality of training data for a classification network affects the classification performance of degraded images.
Kazuki Endo, Masayuki Tanaka 0001, Masatoshi Okutomi
IEEE Trans. Circuits Syst. Video Technol.3
2020 Deep Snapshot HDR Imaging Using Multi-exposure Color Filter Array
Takeru Suda, Masayuki Tanaka 0001, Yusuke Monno, Masatoshi Okutomi
ACCV (2)4
2020 Polarimetric Multi-view Inverse Rendering
Jinyu Zhao, Yusuke Monno, Masatoshi Okutomi
ECCV (24)3
2020 Classifying Degraded Images Over Various Levels Of Degradation
abstract
Classification for degraded images having various levels of degradation is very important in practical applications. This paper proposes a convolutional neural network to classify degraded images by using a restoration network and an ensemble learning. The results demonstrate that the proposed network can classify degraded images over various levels of degradation well. This paper also reveals how the image-quality of training data for a classification network affects the classification performance of degraded images.
Kazuki Endo, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP3
2020 Monochrome And Color Polarization Demosaicking Using Edge-Aware Residual Interpolation
abstract
A division-of-focal-plane or microgrid image polarimeter enables us to acquire a set of polarization images in one shot. Since the polarimeter consists of an image sensor equipped with a monochrome or color polarization filter array (MPFA or CPFA), the demosaicking process to interpolate missing pixel values plays a crucial role in obtaining high-quality polarization images. In this paper, we propose a novel MPFA demosaicking method based on edge-aware residual interpolation (EARI) and also extend it to CPFA demosaicking. The key of EARI is a new edge detector for generating an effective guide image used to interpolate the missing pixel values. We also present a newly constructed full color-polarization image dataset captured using a 3-CCD camera and a rotating polarizer. Using the dataset, we experimentally demonstrate that our EARI-based method outperforms existing methods in MPFA and CPFA demosaicking.
Miki Morimatsu, Yusuke Monno, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP4
2020 Human Segmentation with Dynamic LiDAR Data
abstract
Consecutive LiDAR scans compose dynamic 3D sequences, which contain more abundant information than a single frame. Similar to the development history of image and video perception, dynamic 3D sequence perception starts to come into sight after inspiring research on static 3D data perception. This work proposes a spatio-temporal neural network for human segmentation with the dynamic LiDAR point clouds. It takes a sequence of depth images as input. It has a two-branch structure, i.e., the spatial segmentation branch and the temporal velocity estimation branch. The velocity estimation branch is designed to capture motion cues from the input sequence and then propagates them to the other branch. So that the segmentation branch segments humans according to both spatial and temporal features. These two branches are jointly learned on a generated dynamic point cloud dataset for human recognition. Our works fill in the blank of dynamic point cloud perception with the spherical representation of point cloud and achieves high accuracy. The experiments indicate that the introduction of temporal feature benefits the segmentation of dynamic point cloud.
Wonjik Kim, Masayuki Tanaka 0001, Masatoshi Okutomi
ICPR4
2019 Pro-Cam SSfM: Projector-Camera System for Structure and Spectral Reflectance From Motion
abstract
In this paper, we propose a novel projector-camera system for practical and low-cost acquisition of a dense object 3D model with the spectral reflectance property. In our system, we use a standard RGB camera and leverage an off-the-shelf projector as active illumination for both the 3D reconstruction and the spectral reflectance estimation. We first reconstruct the 3D points while estimating the poses of the camera and the projector, which are alternately moved around the object, by combining multi-view structured light and structure-from-motion (SfM) techniques. We then exploit the projector for multispectral imaging and estimate the spectral reflectance of each 3D point based on a novel spectral reflectance estimation model considering the geometric relationship between the reconstructed 3D points and the estimated projector positions. Experimental results on several real objects demonstrate that our system can precisely acquire a dense 3D model with the full spectral reflectance property using off-the-shelf devices.
Yusuke Monno, Hironori Hidaka, Masatoshi Okutomi
ICCV4
2019 Is This the Right Place? Geometric-Semantic Pose Verification for Indoor Visual Localization
abstract
Visual localization in large and complex indoor scenes, dominated by weakly textured rooms and repeating geometric patterns, is a challenging problem with high practical relevance for applications such as Augmented Reality and robotics. To handle the ambiguities arising in this scenario, a common strategy is, first, to generate multiple estimates for the camera pose from which a given query image was taken. The pose with the largest geometric consistency with the query image, e.g., in the form of an inlier count, is then selected in a second stage. While a significant amount of research has concentrated on the first stage, there has been considerably less work on the second stage. In this paper, we thus focus on pose verification. We show that combining different modalities, namely appearance, geometry, and semantics, considerably boosts pose verification and consequently pose accuracy. We develop multiple hand-crafted as well as a trainable approach to join into the geometric-semantic verification and show significant improvements over state-of-the-art on a very challenging indoor dataset.
Hajime Taira, Ignacio Rocco, Jirí Sedlár, Masatoshi Okutomi, Josef Sivic, Tomás Pajdla, Torsten Sattler, Akihiko Torii
ICCV4
2019 Automatic Labeled LiDAR Data Generation based on Precise Human Model
abstract
Following improvements in deep neural networks, state-of-the-art networks have been proposed for human recognition using point clouds captured by LiDAR. However, the performance of these networks strongly depends on the training data. An issue with collecting training data is labeling. Labeling by humans is necessary to obtain the ground truth label; however, labeling requires huge costs. Therefore, we propose an automatic labeled data generation pipeline, for which we can change any parameters or data generation environments. Our approach uses a human model named Dhaiba and a background of Miraikan and consequently generated realistic artificial data. We present 500k + data generated by the proposed pipeline. This paper also describes the specification of the pipeline and data details with evaluations of various approaches.
Wonjik Kim, Masayuki Tanaka 0001, Masatoshi Okutomi, Yoko Sasaki
ICRA3
2018 Benchmarking 6DOF Outdoor Visual Localization in Changing Conditions
abstract
Visual localization enables autonomous vehicles to navigate in their surroundings and augmented reality applications to link virtual to real worlds. Practical visual localization approaches need to be robust to a wide variety of viewing condition, including day-night changes, as well as weather and seasonal variations, while providing highly accurate 6 degree-of-freedom (6DOF) camera pose estimates. In this paper, we introduce the first benchmark datasets specifically designed for analyzing the impact of such factors on visual localization. Using carefully created ground truth poses for query images taken under a wide variety of conditions, we evaluate the impact of various factors on 6DOF camera pose estimation accuracy through extensive experiments with state-of-the-art localization approaches. Based on our results, we draw conclusions about the difficulty of different conditions, showing that long-term localization is far from solved, and propose promising avenues for future work, including sequence-based localization approaches and the need for better local features. Our benchmark is available at visuallocalization.net.
Torsten Sattler, Will Maddern, Carl Toft, Akihiko Torii, Lars Hammarstrand, Erik Stenborg, Daniel Safari, Masatoshi Okutomi, Marc Pollefeys, Josef Sivic, Fredrik Kahl, Tomás Pajdla
CVPR8
2018 InLoc: Indoor Visual Localization With Dense Matching and View Synthesis
abstract
We seek to predict the 6 degree-of-freedom (6DoF) pose of a query photograph with respect to a large indoor 3D map. The contributions of this work are three-fold. First, we develop a new large-scale visual localization method targeted for indoor environments. The method proceeds along three steps: (i) efficient retrieval of candidate poses that ensures scalability to large-scale environments, (ii) pose estimation using dense matching rather than local features to deal with texture less indoor scenes, and (iii) pose verification by virtual view synthesis to cope with significant changes in viewpoint, scene layout, and occluders. Second, we collect a new dataset with reference 6DoF poses for large-scale indoor localization. Query photographs are captured by mobile phones at a different time than the reference 3D map, thus presenting a realistic indoor localization scenario. Third, we demonstrate that our method significantly outperforms current state-of-the-art indoor localization approaches on this new challenging data.
Hajime Taira, Masatoshi Okutomi, Torsten Sattler, Mircea Cimpoi, Marc Pollefeys, Josef Sivic, Tomás Pajdla, Akihiko Torii
CVPR2
2018 Joint Optimization for Compressive Video Sensing and Reconstruction Under Hardware Constraints
Michitaka Yoshida, Akihiko Torii, Masatoshi Okutomi, Kenta Endo, Yukinobu Sugiyama, Rin-Ichiro Taniguchi, Hajime Nagahara
ECCV (10)3
2018 Coupled convolution layer for convolutional neural network
Kazutaka Uchida, Masayuki Tanaka 0001, Masatoshi Okutomi
Neural Networks3
2018 24/7 Place Recognition by View Synthesis
abstract
We address the problem of large-scale visual place recognition for situations where the scene undergoes a major change in appearance, for example, due to illumination (day/night), change of seasons, aging, or structural modifications over time such as buildings being built or destroyed. Such situations represent a major challenge for current large-scale place recognition methods. This work has the following three principal contributions. First, we demonstrate that matching across large changes in the scene appearance becomes much easier when both the query image and the database image depict the scene from approximately the same viewpoint. Second, based on this observation, we develop a new place recognition approach that combines (i) an efficient synthesis of novel views with (ii) a compact indexable image representation. Third, we introduce a new challenging dataset of 1,125 camera-phone query images of Tokyo that contain major changes in illumination (day, sunset, night) as well as structural changes in the scene. We demonstrate that the proposed approach significantly outperforms other large-scale place recognition techniques on this challenging data.
Akihiko Torii, Relja Arandjelovic, Josef Sivic, Masatoshi Okutomi, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.4
2017 Are Large-Scale 3D Models Really Necessary for Accurate Visual Localization?
abstract
Accurate visual localization is a key technology for autonomous navigation. 3D structure-based methods employ 3D models of the scene to estimate the full 6DOF pose of a camera very accurately. However, constructing (and extending) large-scale 3D models is still a significant challenge. In contrast, 2D image retrieval-based methods only require a database of geo-tagged images, which is trivial to construct and to maintain. They are often considered inaccurate since they only approximate the positions of the cameras. Yet, the exact camera pose can theoretically be recovered when enough relevant database images are retrieved. In this paper, we demonstrate experimentally that large-scale 3D models are not strictly necessary for accurate visual localization. We create reference poses for a large and challenging urban dataset. Using these poses, we show that combining image-based methods with local reconstructions results in a pose accuracy similar to the state-of-the-art structure-based methods. Our results suggest that we might want to reconsider the current approach for accurate large-scale localization.
Torsten Sattler, Akihiko Torii, Josef Sivic, Marc Pollefeys, Hajime Taira, Masatoshi Okutomi, Tomás Pajdla
CVPR6
2017 Misalignment-Robust Joint Filter for Cross-Modal Image Pairs
abstract
Although several powerful joint filters for cross-modal image pairs have been proposed, the existing joint filters generate severe artifacts when there are misalignments between a target and a guidance images. Our goal is to generate an artifact-free output image even from the misaligned target and guidance images. We propose a novel misalignment-robust joint filter based on weight-volume-based image composition and joint-filter cost volume. Our proposed method first generates a set of translated guidances. Next, the joint-filter cost volume and a set of filtered images are computed from the target image and the set of the translated guidances. Then, a weight volume is obtained from the joint-filter cost volume while considering a spatial smoothness and a label-sparseness. The final output image is composed by fusing the set of the filtered images with the weight volume for the filtered images. The key is to generate the final output image directly from the set of the filtered images by weighted averaging using the weight volume that is obtained from the joint-filter cost volume. The proposed framework is widely applicable and can involve any kind of joint filter. Experimental results show that the proposed method is effective for various applications including image denosing, image up-sampling, haze removal and depth map interpolation.
Takashi Shibata 0001, Masayuki Tanaka 0001, Masatoshi Okutomi
ICCV3
2017 Tunable color correction between linear and polynomial models for noisy images
abstract
Linear color correction (LCC) and polynomial color correction (PCC) are widely used in a camera imaging pipeline. PCC generally achieves lower colorimetric errors than LCC. However, if an image contains noise, PCC amplifies the noise more severely than LCC. Consequently, there is a trade-off between LCC and PCC in the presence of noise. In this paper, we propose a novel framework for color correction, which we call tunable color correction (TCC). TCC enables us to tune a color correction matrix between linear and polynomial models by a tuning parameter. We also present a way of selecting a suitable parameter value based on the mean squared error calculation model for PCC. Experimental results demonstrate that TCC effectively balances the trade-off and outperforms both LCC and PCC for noisy images.
Ryo Yamakabe, Yusuke Monno, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP4
2016 Gradient-Domain Image Reconstruction Framework with Intensity-Range and Base-Structure Constraints
abstract
This paper presents a novel unified gradient-domain image reconstruction framework with intensity-range constraint and base-structure constraint. The existing method for manipulating base structures and detailed textures are classifiable into two major approaches: i) gradient-domain and ii) layer-decomposition. To generate detail-preserving and artifact-free output images, we combine the benefits of the two approaches into the proposed framework by introducing the intensity-range constraint and the base-structure constraint. To preserve details of the input image, the proposed method takes advantage of reconstructing the output image in the gradient domain, while the output intensity is guaranteed to lie within the specified intensity range, e.g. 0-to-255, by the intensity-range constraint. In addition, the reconstructed image lies close to the base structure by the base-structure constraint, which is effective for restraining artifacts. Experimental results show that the proposed framework is effective for various applications such as tone mapping, seamless image cloning, detail enhancement, and image restoration.
Takashi Shibata 0001, Masayuki Tanaka 0001, Masatoshi Okutomi
CVPR3
2016 Multi-view Inverse Rendering Under Arbitrary Illumination and Albedo
Kichang Kim, Akihiko Torii, Masatoshi Okutomi
ECCV (3)3
2016 Effective color correction pipeline for a noisy image
abstract
Color correction is an essential image processing operation that transforms a camera-dependent RGB color space to a standard color space, e.g., the XYZ or the sRGB color space. The color correction is typically performed by multiplying the camera RGB values by a color correction matrix, which often amplifies image noise. In this paper, we propose an effective color correction pipeline for a noisy image. The proposed pipeline consists of two parts; the color correction and denoising. In the color correction part, we utilize spatially varying color correction (SVCC) that adaptively calculates the color correction matrices for each local image block considering the noise effect. Although the SVCC can effectively suppress the noise amplification, the noise is still included in the color corrected image, where the noise levels spatially vary for each local block. In the denoising part, we propose an effective denoising framework for the color corrected image with spatially varying noise levels. Experimental results demonstrate that the proposed color correction pipeline outperforms existing algorithms for various noise levels.
Kenta Takahashi, Yusuke Monno, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP4
2016 Depth map upsampling by self-guided residual interpolation
abstract
In this paper, we propose a simple and effective depth upsampling technique using self-guided residual interpolation. The original residual interpolation requires guidance information such as high-resolution RGB color image. However, self-guided residual interpolation requires only a single depth map. In the proposed algorithm, a tentative estimation of a high-resolution depth map is first generated from an input low-resolution depth map. Then, re-interpolation is applied to the residual domain, which is defined by differences between the input depth map and the tentative estimate. A precise high-resolution depth map is obtainable by interpolating in the residual domain. Experimental results demonstrate that our algorithm can outperform state-of-the-art depth map upsampling algorithms.
Yosuke Konno, Masayuki Tanaka 0001, Masatoshi Okutomi, Yukiko Yanagawa, Koichi Kinoshita, Masato Kawade
ICPR3
2016 Super high dynamic range video
abstract
High dynamic range (HDR) imaging is highly demanded in computer vision algorithms. An HDR image is composed with several low dynamic range (LDR) images, which usually have some disparities. In many HDR imaging algorithms, the disparities are estimated based on the texture information of the LDR images. However, the texture information is often lost completely if scenes include extremely bright and dark regions simultaneously. Recently, super high dynamic range (SHDR) imaging algorithm has been proposed where the disparities are estimated based on the segment shapes instead of the textures for handling such extreme scenes. In this paper, we extend the SHDR imaging algorithm to SHDR video generation introducing temporal smoothness terms. The temporal smoothness terms improve the temporal stability and the precision of the disparity estimation. Quantitative and qualitative evaluations demonstrate that the proposed algorithm outperforms existing algorithms.
Yuka Ogino, Masayuki Tanaka 0001, Takashi Shibata 0001, Masatoshi Okutomi
ICPR4
2016 Robust feature matching by learning descriptor covariance with viewpoint synthesis
abstract
For images taken from very different viewpoints, we propose a new feature matching algorithm that provides accurate matches while preserving high matchability. Our method first synthesizes images by simulating the viewpoint changes. It then learns variation of local feature descriptors induced by the viewpoint changes. Finally, we robustly match feature descriptors by measuring the similarity using the learned variation. Our method is particularly useful for matching new query images to target image archived in a database. We demonstrate the benefits of the proposed method in terms of accuracy and computational time through experiments using several wide-baseline image datasets.
Hajime Taira, Akihiko Torii, Masatoshi Okutomi
ICPR3
2016 Coupled convolution layer for convolutional neural network
abstract
We introduce a coupled convolution layer comprising two parallel convolutions with mutually constrained weights. Inspired by the human retina mechanism, we constrain the convolution weights such that one set of weights should be the negative of the other to mimic responses of on-center and off-center retinal ganglion cells. Our analysis shows that the retina-like convolution layer, a special case of the coupled convolution layer, can be realized by a normal convolutional layer with a pair of activation functions designated as Biased ON/OFF ReLU. Experimental comparisons demonstrate that the proposed coupled convolution layer performs better without increasing the number of parameters, which reveals two important facts. First, the separation of the positive and negative part into different channels plays an important role. Secondly, constraining weights across convolutions can produce better performance than training weights freely. We evaluate its effect by comparison with ReLU, LReLU, and PReLU using the CIFAR-10, CIFAR-100, and PlanktonSet 1.0 datasets.
Kazutaka Uchida, Masayuki Tanaka 0001, Masatoshi Okutomi
ICPR3
2016 Beyond Color Difference: Residual Interpolation for Color Image Demosaicking
abstract
In this paper, we propose residual interpolation (RI) as an alternative to color difference interpolation, which is a widely accepted technique for color image demosaicking. Our proposed RI performs the interpolation in a residual domain, where the residuals are differences between observed and tentatively estimated pixel values. Our hypothesis for the RI is that if image interpolation is performed in a domain with a smaller Laplacian energy, its accuracy is improved. Based on the hypothesis, we estimate the tentative pixel values to minimize the Laplacian energy of the residuals. We incorporate the RI into the gradient-based threshold free algorithm, which is one of the state-of-the-art Bayer demosaicking algorithms. Experimental results demonstrate that our proposed demosaicking algorithm using the RI surpasses the state-of-the-art algorithms for the Kodak, the IMAX, and the beyond Kodak data sets.
Daisuke Kiku, Yusuke Monno, Masayuki Tanaka 0001, Masatoshi Okutomi
IEEE Trans. Image Process.4
2015 3D Surface Reconstruction from Point-and-Line Cloud
abstract
We present a method for reconstructing 3D surface as triangular meshes from imagery. The surface reconstruction requires 3D point cloud for composing vertices of triangle meshes. A standard approach uses incremental structure from motion (SfM) to obtain camera poses and sparse 3D point cloud that are given based on 2D key-point matching. As the 3D surface directly reconstructed from the sparse 3D point cloud often lack detail of objects, multiple-view stereo (MVS) is commonly used to generate dense 3D point cloud. A known problem with the densification is that MVS generates many small patches even for planar flat objects that degrade the quality of surface model. Using dense 3D point cloud also requires high memory capacity for visualization. In this work, we propose to reconstruct 3D surface using sparse 3D point cloud generated by SfM and 3D line segments (3D line cloud) computed from multiple views since these two elements can complement well for representing man-made structures. The proposed method extends the tetrahedra-carving method as it can use 3D point-and-line cloud under the global optimization framework. We demonstrate that the proposed method can efficiently produce surface models whose quality are at least as good as the baseline method using dense 3D point cloud.
Takayuki Sugiura, Akihiko Torii, Masatoshi Okutomi
3DV3
2015 24/7 place recognition by view synthesis
abstract
We address the problem of large-scale visual place recognition for situations where the scene undergoes a major change in appearance, for example, due to illumination (day/night), change of seasons, aging, or structural modifications over time such as buildings built or destroyed. Such situations represent a major challenge for current large-scale place recognition methods. This work has the following three principal contributions. First, we demonstrate that matching across large changes in the scene appearance becomes much easier when both the query image and the database image depict the scene from approximately the same viewpoint. Second, based on this observation, we develop a new place recognition approach that combines (i) an efficient synthesis of novel views with (ii) a compact indexable image representation. Third, we introduce a new challenging dataset of 1,125 camera-phone query images of Tokyo that contain major changes in illumination (day, sunset, night) as well as structural changes in the scene. We demonstrate that the proposed approach significantly outperforms other large-scale place recognition techniques on this challenging data.
Akihiko Torii, Relja Arandjelovic, Josef Sivic, Masatoshi Okutomi, Tomás Pajdla
CVPR4
2015 Pseudo four-channel image denoising for noisy CFA raw data
abstract
Most demosaicking algorithms only focus on handling noise-free CFA raw data. In practice, the CFA raw data are corrupted by noise, which degrades demosaicking performance. Full-color image quality strongly depends on the performance of the demosaicking. Here, we propose a CFA raw data denoising algorithm. In the proposed algorithm, the CFA raw data is converted to a pseudo four-channel image by rearranging pixels. Then, the four-channel data are transformed based on the principal component analysis (PCA). Existing high-performance gray image denoising algorithm is applied to each transformed image. Finally, the denoised data is rearranged to obtain denoised CFA raw data. We evaluate both the denoised CFA raw data as well as the full-color image reconstructed with the noisy CFA raw data. Experimental comparisons demonstrate that the proposed algorithm outperforms existing state-of-the-art algorithms.
Hiroki Akiyama, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP3
2015 Adaptive residual interpolation for color image demosaicking
abstract
Color image demosaicking is an essential image processing operation for acquiring high-quality color images. Recently, demosaicking algorithms using residual interpolation (RI), which performs the interpolation in a residual domain, have been proposed. An iterative framework has also been introduced into the RI and shown state-of-the-art performance. In this paper, we propose a novel demosaicking algorithm using adaptive residual interpolation (ARI), which adaptively selects a suitable iteration number and combines two different types of RI algorithms at each pixel. Experimental results demonstrate that our demosaicking algorithm can achieve a clear improvement in comparison with existing algorithms.
Yusuke Monno, Daisuke Kiku, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP4
2015 Unified image fusion based on application-adaptive importance measure
abstract
This paper presents a novel unified image fusion framework based on an application-adaptive importance measure. In the proposed method, an important area is selected pixel-by-pixel using the importance measure which is designed for each image type in each application. Then, the fused intensity is generated by a Poisson image editing. The main contribution is to provide a generalized image fusion framework enables us to deal with various different types of images for many applications. Experimental results show that the proposed method is effective for various applications including depth-perceptible image enhancement, temperature-preserving image fusion, optical flow fusion, and haze removal.
Takashi Shibata 0001, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP3
2015 Visual Place Recognition with Repetitive Structures
abstract
Repeated structures such as building facades, fences or road markings often represent a significant challenge for place recognition. Repeated structures are notoriously hard for establishing correspondences using multi-view geometry. They violate the feature independence assumed in the bag-of-visual-words representation which often leads to over-counting evidence and significant degradation of retrieval performance. In this work we show that repeated structures are not a nuisance but, when appropriately represented, they form an important distinguishing feature for many places. We describe a representation of repeated structures suitable for scalable retrieval and geometric verification. The retrieval is based on robust detection of repeated image structures and a suitable modification of weights in the bag-of-visual-word model. We also demonstrate that the explicit detection of repeated patterns is beneficial for robust visual word matching for geometric verification. Place recognition results are shown on datasets of street-level imagery from Pittsburgh and San Francisco demonstrating significant gains in recognition performance compared to the standard bag-of-visual-words baseline as well as the more recently proposed burstiness weighting and Fisher vector encoding.
Akihiko Torii, Josef Sivic, Masatoshi Okutomi, Tomás Pajdla
IEEE Trans. Pattern Anal. Mach. Intell.3
2015 A Practical One-Shot Multispectral Imaging System Using a Single Image Sensor
abstract
Single-sensor imaging using the Bayer color filter array (CFA) and demosaicking is well established for current compact and low-cost color digital cameras. An extension from the CFA to a multispectral filter array (MSFA) enables us to acquire a multispectral image in one shot without increased size or cost. However, multispectral demosaicking for the MSFA has been a challenging problem because of very sparse sampling of each spectral band in the MSFA. In this paper, we propose a high-performance multispectral demosaicking algorithm, and at the same time, a novel MSFA pattern that is suitable for our proposed algorithm. Our key idea is the use of the guided filter to interpolate each spectral band. To generate an effective guide image, in our proposed MSFA pattern, we maintain the sampling density of the G -band as high as the Bayer CFA, and we array each spectral band so that an adaptive kernel can be estimated directly from raw MSFA data. Given these two advantages, we effectively generate the guide image from the most densely sampled G -band using the adaptive kernel. In the experiments, we demonstrate that our proposed algorithm with our proposed MSFA pattern outperforms existing algorithms and provides better color fidelity compared with a conventional color imaging system with the Bayer CFA. We also show some real applications using a multispectral camera prototype we built.
Yusuke Monno, Sunao Kikuchi, Masayuki Tanaka 0001, Masatoshi Okutomi
IEEE Trans. Image Process.4
2014 A General and Simple Method for Camera Pose and Focal Length Determination
abstract
In this paper, we revisit the pose determination problem of a partially calibrated camera with unknown focal length, hereafter referred to as the PnPf problem, by using n(n ≥ 4) 3D-to-2D point correspondences. Our core contribution is to introduce the angle constraint and derive a compact bivariate polynomial equation for each point triplet. Based on this polynomial equation, we propose a truly general method for the PnPf problem, which is suited both to the minimal 4-point based RANSAC application, and also to large scale scenarios with thousands of points, irrespective of the 3D point configuration. In addition, by solving bivariate polynomial systems via the Sylvester resultant, our method is very simple and easy to implement. Its simplicity is especially obvious when one needs to develop a fast solver for the 4-point case on the basis of the characteristic polynomial technique. Experiment results have also demonstrated its superiority in accuracy and efficiency when compared with the existing state-of-the-art solutions.
Yinqiang Zheng, Shigeki Sugimoto, Imari Sato, Masatoshi Okutomi
CVPR4
2014 Signal dependent noise removal from a single image
abstract
State-of-the-art image denoising algorithms usually assume additive white Gaussian noise (AWGN), although they have achieved outstanding performance, modeling and removing real signal dependent noise from a single image still remains a challenging problem. In this paper we propose a segmentation-based image denoising algorithm for signal dependent noise. Incorporating a noise identification algorithm, we integrate these two modules into a full blind, end-to-end denoising algorithm for signal dependent noise. First, we identify the noise level function for a given single noisy image. Then, after initial denoising, segmentation is applied to the pre-filtered image. Assuming the noise level of each segment is constant, we apply AWGN denoising algorithm to each segment. We obtain a final de-noised image by composing the denoised segments. Various experimental results on synthetic and real noisy images show that our algorithm outperforms state-of-the-art denoising algorithms in removing real signal dependent noise.
Xinhao Liu 0002, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP3
2014 Multispectral demosaicking with novel guide image generation and residual interpolation
abstract
A one-shot multispectral imaging system using a multispectral filter array (MSFA) provides a practical solution for compact, low-cost, and real-time multispectral imaging. However, multispectral demosaicking is a challenging problem because each spectral band is significantly undersampled in the MSFA. In this paper, we propose a novel demosaicking algorithm for the MSFA proposed in [1, 2]. Main contributions of this paper are (i) we utilize multispectral correlations for generating a guide image, which is effectively used for interpolation preserving image structures, and (ii) we effectively use residual interpolation (RI) [3] for generating the guide image and interpolating each spectral band. Experimental results demonstrate that our proposed algorithm significantly outperforms existing state-of-the-art algorithms.
Yusuke Monno, Daisuke Kiku, Sunao Kikuchi, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP5
2014 Super-high Dynamic Range Imaging
abstract
We propose a novel high dynamic range (HDR) imaging algorithm for the scenes that contain an extremely wide range of scene radiance. In the HDR imaging, several images are taken under different exposures. Those images usually have displacement from one another due to camera and/or object motions. The challenge of the super HDR imaging is to align those images because any image contains "lost" regions where texture information is completely lost due to overexposure or underexposure. We propose an image alignment algorithm based on similarities of region shapes instead of the similarities of the textures. Experimental comparisons demonstrate that the proposed algorithm outperforms state-of-the-art algorithms.
Takehito Hayami, Masayuki Tanaka 0001, Masatoshi Okutomi, Takashi Shibata 0001, Shuji Senda
ICPR3
2014 A Novel Inference of a Restricted Boltzmann Machine
abstract
A deep neural network (DNN) pre-trained via stacking restricted Boltzmann machines (RBMs) demonstrates high performance. The binary RBM is usually used to construct the DNN. However, a continuous probability of each node is used as real value state, although the state of the binary RBM's node should be represented by a random binary variable. One of main reasons of this abuse is that it works. One of others is to reduce a computational cost. In this paper, we propose a novel inference of the RBM, considering that the input of the RBM is the random binary variable. Straight forward derivation of the proposed inference is intractable. Then, we also propose the closed-form approximation of it. We convince that the proposed inference is more reasonable than a conventional algorithm of the RBM. Experimental comparisons demonstrate that the proposed inference improves the performance of the DNN.
Masayuki Tanaka 0001, Masatoshi Okutomi
ICPR2
2014 Robust ground surface map generation using vehicle-mounted stereo camera
abstract
We propose a robust method for incrementally estimating a regular-grid ground surface map from stereo image sequences captured by nearly front-looking vehicle-mounted stereo cameras. The method simultaneously estimates a camera ego-motion and vertex heights of a regular mesh, which is composed of piecewise triangular patches drawn on a level plane in the ground coordinate system, by minimizing pixel value differences over the ground surface. The method combinationally uses feature-based approach and pixel-based approach for robustly estimating ego-motion parameters. We also show that this combination is beneficial for removing outlier pixels, which mainly represent the edge of the self-shadow area on the ground surface. The validity of the proposed method is demonstrated through experiments using real images.
Kouma Motooka, Shigeki Sugimoto, Masatoshi Okutomi, Takeshi Shima
IROS3
2014 Practical Signal-Dependent Noise Parameter Estimation From a Single Noisy Image
abstract
The additive white Gaussian noise is widely assumed in many image processing algorithms. However, in the real world, the noise from actual cameras is better modeled as signal-dependent noise (SDN). In this paper, we focus on the SDN model and propose an algorithm to automatically estimate its parameters from a single noisy image. The proposed algorithm identifies the noise level function of signal-dependent noise assuming the generalized signal-dependent noise model and is also applicable to the Poisson-Gaussian noise model. The accuracy is achieved by improved estimation of local mean and local noise variance from the selected low-rank patches. We evaluate the proposed algorithm with both synthetic and real noisy images. Experiments demonstrate that the proposed estimation algorithm outperforms the state-of-the-art methods.
Xinhao Liu 0002, Masayuki Tanaka 0001, Masatoshi Okutomi
IEEE Trans. Image Process.3
2013 Visual Place Recognition with Repetitive Structures
abstract
Repeated structures such as building facades, fences or road markings often represent a significant challenge for place recognition. Repeated structures are notoriously hard for establishing correspondences using multi-view geometry. Even more importantly, they violate the feature independence assumed in the bag-of-visual-words representation which often leads to over-counting evidence and significant degradation of retrieval performance. In this work we show that repeated structures are not a nuisance but, when appropriately represented, they form an important distinguishing feature for many places. We describe a representation of repeated structures suitable for scalable retrieval. It is based on robust detection of repeated image structures and a simple modification of weights in the bag-of-visual-word model. Place recognition results are shown on datasets of street-level imagery from Pittsburgh and San Francisco demonstrating significant gains in recognition performance compared to the standard bag-of-visual-words baseline and more recently proposed burstiness weighting.
Akihiko Torii, Josef Sivic, Tomás Pajdla, Masatoshi Okutomi
CVPR4
2013 A Practical Rank-Constrained Eight-Point Algorithm for Fundamental Matrix Estimation
abstract
Due to its simplicity, the eight-point algorithm has been widely used in fundamental matrix estimation. Unfortunately, the rank-2 constraint of a fundamental matrix is enforced via a posterior rank correction step, thus leading to non-optimal solutions to the original problem. To address this drawback, existing algorithms need to solve either a very high order polynomial or a sequence of convex relaxation problems, both of which are computationally ineffective and numerically unstable. In this work, we present a new rank-2 constrained eight-point algorithm, which directly incorporates the rank-2 constraint in the minimization process. To avoid singularities, we propose to solve seven sub problems and retrieve their globally optimal solutions by using tailored polynomial system solvers. Our proposed method is noniterative, computationally efficient and numerically stable. Experiment results have verified its superiority over existing algebraic error based algorithms in terms of accuracy, as well as its advantages when used to initialize geometric error based algorithms.
Yinqiang Zheng, Shigeki Sugimoto, Masatoshi Okutomi
CVPR3
2013 Revisiting the PnP Problem: A Fast, General and Optimal Solution
abstract
In this paper, we revisit the classical perspective-n-point (PnP) problem, and propose the first non-iterative O(n) solution that is fast, generally applicable and globally optimal. Our basic idea is to formulate the PnP problem into a functional minimization problem and retrieve all its stationary points by using the Gr"obner basis technique. The novelty lies in a non-unit quaternion representation to parameterize the rotation and a simple but elegant formulation of the PnP problem into an unconstrained optimization problem. Interestingly, the polynomial system arising from its first-order optimality condition assumes two-fold symmetry, a nice property that can be utilized to improve speed and numerical stability of a Grobner basis solver. Experiment results have demonstrated that, in terms of accuracy, our proposed solution is definitely better than the state-of-the-art O(n) methods, and even comparable with the reprojection error minimization method.
Yinqiang Zheng, Yubin Kuang, Shigeki Sugimoto, Kalle Åström, Masatoshi Okutomi
ICCV5
2013 Residual interpolation for color image demosaicking
abstract
A color difference interpolation technique is widely used for color image demosaicking. In this paper, we propose residual interpolation as an alternative to the color difference interpolation, where the residual is a difference between an observed and a tentatively estimated pixel value. We incorporate the proposed residual interpolation into the gradient based threshold free (GBTF) algorithm, which is one of current state-of-the-art demosaicking algorithms. Experimental results demonstrate that our proposed demosaicking algorithm using the residual interpolation can give state-of-the-art performance for the 30 images of Kodak and IMAX datasets.
Daisuke Kiku, Yusuke Monno, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP4
2013 Estimation of signal dependent noise parameters from a single image
abstract
The additive white Gaussian noise (AWGN) is usually assumed in many image processing algorithms. However, these algorithms cannot effectively deal with the noise from actual cameras which is better modeled as signal dependent noise (SDN). In this paper, we focus on the SDN model and propose an algorithm to accurately estimate its parameters without any assumption of the noise types. The noise parameters are estimated by using the selected weak textured patches from a single noisy image. Experiments on synthetic noisy images are conducted to test the algorithm, which show that our noise parameter estimation outperforms the existing algorithms. And based on our estimation, the performance of image processing applications like Wiener filter can be effectively improved.
Xinhao Liu 0002, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP3
2013 Single-Image Noise Level Estimation for Blind Denoising
abstract
Noise level is an important parameter to many image processing applications. For example, the performance of an image denoising algorithm can be much degraded due to the poor noise level estimation. Most existing denoising algorithms simply assume the noise level is known that largely prevents them from practical use. Moreover, even with the given true noise level, these denoising algorithms still cannot achieve the best performance, especially for scenes with rich texture. In this paper, we propose a patch-based noise level estimation algorithm and suggest that the noise level parameter should be tuned according to the scene complexity. Our approach includes the process of selecting low-rank patches without high frequency components from a single noisy image. The selection is based on the gradients of the patches and their statistics. Then, the noise level is estimated from the selected patches using principal component analysis. Because the true noise level does not always provide the best performance for nonblind denoising algorithms, we further tune the noise level parameter for nonblind denoising. Experiments demonstrate that both the accuracy and stability are superior to the state of the art noise level estimation algorithm for various scenes and noise levels.
Xinhao Liu 0002, Masayuki Tanaka 0001, Masatoshi Okutomi
IEEE Trans. Image Process.3
2012 Stable Two View Reconstruction Using the Six-Point Algorithm
Kazuki Nozawa, Akihiko Torii, Masatoshi Okutomi
ACCV (4)3
2012 Practical low-rank matrix approximation under robust L1-norm
abstract
A great variety of computer vision tasks, such as rigid/nonrigid structure from motion and photometric stereo, can be unified into the problem of approximating a low-rank data matrix in the presence of missing data and outliers. To improve robustness, the L1-norm measurement has long been recommended. Unfortunately, existing methods usually fail to minimize the L1-based nonconvex objective function sufficiently. In this work, we propose to add a convex trace-norm regularization term to improve convergence, without introducing too much heterogenous information. We also customize a scalable first-order optimization algorithm to solve the regularized formulation on the basis of the augmented Lagrange multiplier (ALM) method. Extensive experimental results verify that our regularized formulation is reasonable, and the solving algorithm is very efficient, insensitive to initialization and robust to high percentage of missing data and/or outliers1.
Yinqiang Zheng, Guangcan Liu, Shigeki Sugimoto, Shuicheng Yan, Masatoshi Okutomi
CVPR5
2012 Generalizing Wiberg algorithm for rigid and nonrigid factorizations with missing components and metric constraints
abstract
In spite of intensive endeavor over decades, rigid and nonrigid factorizations under metric constraints, possibly in the presence of missing components, remain to be very challenging. In this work, we try to break the hard nut by generalizing to these problems the Wiberg algorithm, one of the most successful solutions for unconstrained bilinear factorization. To properly handle missing components, we advocate a bilinear factorization formulation with an extra mean vector. In spirit of the Wiberg algorithm, we first propose an efficient and initialization-insensitive algorithm for unconstrained factorization, posterior correction of whose solution offers reasonable initialization for metric upgrade. For factorization with metric constraints, we reformulate it into an unconstrained problem through quaternion parametrization, which merges elegantly into our unconstrained factorization algorithm. Extensive experiment results verify that our proposed methods are fast, accurate and robust to high percentage of missing components.
Yinqiang Zheng, Shigeki Sugimoto, Shuicheng Yan, Masatoshi Okutomi
CVPR4
2012 Noise level estimation using weak textured patches of a single noisy image
abstract
A patch-based noise level estimation algorithm is proposed in this paper, with patches generated from a single noisy image. One can easily estimate the noise level from image patches using principal component analysis (PCA) if the image comprises only weak textured patches. The challenge for patch-based noise level estimation is how to select weak textured patches from a noisy image. As described in this paper, we propose a novel algorithm to select weak textured patches from a single noisy image based on the gradients of the patches and their statistics. Then we estimate the noise level from the selected weak textured patches using PCA. We demonstrate experimentally that the proposed noise level estimation algorithm outperforms the state-of-the-art algorithm.
Xinhao Liu 0002, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP3
2012 Optimal spectral sensitivity functions for a single-camera one-shot multispectral imaging system
abstract
Multispectral imaging is highly demanded for precise color reproduction and for various computer vision applications. Recently, a single-camera one-shot multispectral imaging (SCOS) system that uses a single image sensor equipped with a multispectral filter array (MSFA) has been proposed. In this paper, we develop optimal spectral sensitivity functions (SSFs) for the SCOS system, in which multispectral image quality depends strongly on the performance of multispectral demosaicking. First, we propose a simple optimization algorithm that can incorporate a high-performance multispectral demosaicking algorithm. Then, we experimentally demonstrate that the optimized SSFs by our proposed algorithm improve the performance of spectral reflectance estimation and the accuracy of color reproduction.
Yusuke Monno, Toshihiro Kitao, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP4
2012 Fast global non-rigid registration for mosaic creation
Rafael Henrique Castanheira de Souza, Masatoshi Okutomi, Akihiko Torii
ICPR2
2012 Camera self calibration based on direct image alignment
Shigeki Sugimoto, Masatoshi Okutomi
ICPR2
2012 Augmenting moving planar surfaces interactively with video projection and a color camera
abstract
Traditional applications of augmented reality superimpose generated images onto the real world through goggles or monitors held between objects of interest and the user. To render the augmented surfaces interactive, we may exploit directly existing computer vision techniques. However, when using video projection to alter directly the appearance of surfaces, most vision-based algorithms fail. Even Wear Ur World [5], a recent and otherwise well-received interactive projector-camera system, relies on colored thimbles as markers. As notable exception, Tele-Graffiti [6] was designed for normal visible-light cameras without markers, but still considers the light emitted from the projector as unwanted interference, limiting its application.
Samuel Audet, Masatoshi Okutomi, Masayuki Tanaka 0001
VR2
2011 Deterministically maximizing feasible subsystem for robust model fitting with unit norm constraint
abstract
Many computer vision problems can be accounted for or properly approximated by linearity, and the robust model fitting (parameter estimation) problem in presence of outliers is actually to find the Maximum Feasible Subsystem (MaxFS) of a set of infeasible linear constraints. We propose a deterministic branch and bound method to solve the MaxFS problem with guaranteed global optimality. It can be used in a wide class of computer vision problems, in which the model variables are subject to the unit norm constraint. In contrast to the convex and concave relaxations in existing works, we introduce a piecewise linear relaxation to build very tight under- and over-estimators for square terms by partitioning variable bounds into smaller segments. Based on this novel relaxation technique, our branch and bound method can converge in a few iterations. For homogeneous linear systems, which correspond to some quasi-convex problems based on L∞-L∞-norm, our method is non-iterative and certainly reaches the globally optimal solution at the root node by partitioning each variable range into two segments with equal length. Throughout this work, we rely on the so-called Big-M method, and successfully avoid potential numerical problems by exploiting proper parametrization and problem structure. Experimental results demonstrate the stability and efficiency of our proposed method.
Yinqiang Zheng, Shigeki Sugimoto, Masatoshi Okutomi
CVPR3
2011 A branch and contract algorithm for globally optimal fundamental matrix estimation
abstract
We propose a unified branch and contract method to estimate the fundamental matrix with guaranteed global optimality, by minimizing either the Sampson error or the point to epipolar line distance, and explicitly handling the rank-2 constraint and scale ambiguity. Based on a novel denominator linearization strategy, the fundamental matrix estimation problem can be transformed into an equivalent problem that involves 9 squared univariate, 12 bilinear and 6 trilin-ear terms. We build tight convex and concave relaxations for these nonconvex terms and solve the problem deterministically under the branch and bound framework. For acceleration, a bound contraction mechanism is introduced to reduce the size of the branching region at the root node. Given high-quality correspondences and proper data normalization, our experiments show that the state-of-the-art locally optimal methods generally converge to the globally optimal solution. However, they indeed have the risk of being trapped into local minimum in case of noise. As another important experimental result, we also demonstrate, from the viewpoint of global optimization, that the point to epipolar line distance is slightly inferior to the Sampson error in case of drastically varying object scales across two views.
Yinqiang Zheng, Shigeki Sugimoto, Masatoshi Okutomi
CVPR3
2011 Multispectral demosaicking using adaptive kernel upsampling
abstract
Multispectral demosaicking, which estimates full multispectral images from raw data observed using a single image sensor with a color filter array (CFA), is a challenging task because each spectral component is severely undersampled. In this paper, we propose a novel multispectral demosaicking algorithm. We extend existing upsampling algorithms to adaptive kernel upsampling algorithms using an adaptive kernel as a spatial weight and apply them to multispectral demosaicking. We also propose a new CFA and a direct adaptive kernel estimation from the raw data of the proposed CFA. Experimental results with real multispectral images demonstrate the effectiveness of the proposed algorithm.
Yusuke Monno, Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP3
2011 Real-time step edge estimation using stereo images for biped robot
abstract
A state-of-the-arts biped robot can take foot-steps such that its heels always overhang corner edges while ascending stairs, as humans naturally do. The overhanging footstep is advantageous in terms of relaxation of restrictions on gait planning. However, in a man-made environment without geometry information, the overhanging footstep requires the estimation of the exact step-edge position in real-time. In this paper we propose a real-time method for estimating step edge positions using stereo images. We find a straight edge line, which divides a view area into two regions representing the upper and lower step-able planes at the target edge. The edge line is obtained by minimizing a cost function composed of pixel-value-difference images, which are computed from the two stereo images and the geometry parameters of the planes, estimated by an efficient direct method in high precision. The validity of the proposed method is demonstrated through online experiments using stereo cameras mounted on the body of a biped robot traversing real stairs.
Minami Asatani, Shigeki Sugimoto, Masatoshi Okutomi
IROS3
2011 Real-Time Image Mosaicing Using Non-rigid Registration
Rafael Henrique Castanheira de Souza, Masatoshi Okutomi, Akihiko Torii
PSIVT (1)2
2010 Color Kernel Regression for Robust Direct Upsampling from Raw Data of General Color Filter Array
Masayuki Tanaka 0001, Masatoshi Okutomi
ACCV (3)2
2010 3D Structure Refinement of Nonrigid Surfaces through Efficient Image Alignment
Yinqiang Zheng, Shigeki Sugimoto, Masatoshi Okutomi
ACCV (4)3
2010 Direct image alignment of projector-camera systems with planar surfaces
abstract
Projector-camera systems use computer vision to analyze their surroundings and display feedback directly onto real world objects, as embodied by spatial augmented reality. To be effective, the display must remain aligned even when the target object moves, but the added illumination causes problems for traditional algorithms. Current solutions consider the displayed content as interference and largely depend on channels orthogonal to visible light. They cannot directly align projector images with real world surfaces, even though this may be the actual goal. We propose instead to model the light emitted by projectors and reflected into cameras, and to consider the displayed content as additional information useful for direct alignment. We implemented in software an algorithm that successfully executes on planar surfaces of diffuse reflectance properties at almost two frames per second with subpixel accuracy. Although slow, our work proves the viability of the concept, paving the way for future optimization and generalization.
Samuel Audet, Masatoshi Okutomi, Masayuki Tanaka 0001
CVPR2
2010 Image restoration and disparity estimation from an uncalibrated multi-layered image
abstract
Watching a reflection in a glass window, one can often observe a multi-layered image consisting of a front-surface reflection from the glass and a rear-surface reflection through the glass. That multi-layered image is a composition of dual aspects of the same image, resembling a sound reverberation. As described herein, we propose a method to estimate the original reflection image before the layering. First, we model the multi-layered image generation process; then we derive a restoration filter by assuming that the displacement between the multiple reflection images in the single multi-layered image is known. Second, we propose a method to estimate the original image in conjunction with the displacement estimation. The displacement between the reflection images in a single multi-layered image includes scene depth information. Finally, we demonstrate the effectiveness of our proposed method using both synthetic and real multi-layered images.
Takahiro Yano, Masao Shimizu, Masatoshi Okutomi
CVPR3
2010 Comparison of image alignment on hexagonal and square lattices
abstract
A hexagonal lattice has been researched to improve computer vision and image processing. However, image alignment on the lattice has not been fully discussed yet. In this paper, we perform image alignment on hexagonal and square lattices. Then we compare and evaluate them. We used hexagonal lattices of two sizes. One has the same pixel interval between adjacent pixels as the square lattice, which results in higher resolution than the square one. Another has the same pixel area as the square lattice, which has the same resolution as the square one. The results show that the image alignment of both the hexagonal lattices outperforms that of the square lattice with respect to accuracy and the success rate. Importantly, the results also show that converting an existing large image on the square lattice into the smaller image on the hexagonal ones of both the sizes could improve image alignment.
Tetsuo Shima, Shigeki Sugimoto, Masatoshi Okutomi
ICIP3
2010 Progressive MAP-based Deconvolution with Pixel-Dependent Gaussian Prior
abstract
A deconvolution is a fundamental technique and used in various vision applications. A maximum a posteriori estimation is known as a powerful tool. In this paper, we propose a progressive MAP-based deconvolution algorithm with a pixel dependent Gaussian image prior. In the proposed algorithm, a mean and a variance for each pixel are adaptively estimated. Then, the mean and the variance are progressively updated. We experimentally show that the proposed algorithm is comparable to the state-of-the-art algorithms in the case that the true point spread function (PSF) is used for the deconvolution, and that the proposed algorithm outperforms in the non-true PSF case.
Masayuki Tanaka 0001, Takafumi Kanda, Masatoshi Okutomi
ICPR3
2010 Egomotion estimation using planar and non-planar constraints
abstract
There are two major approaches for estimating camera motion (egomotion) given an image sequence. Each approach has own strengths and weaknesses. One approach is the feature based methods. In this approach the point feature correspondences are taken as the input. Since initially the depths of point features are unknown, the egomotion is estimated by the depth independent epipolar constraints on the point feature correspondences. This approach is robust in practice, but is relatively limited in accuracy since it exploits no structure assumption, such as planarity. The other approach, termed the direct method, has the advantage in its accuracy. In this method, the egomotion is estimated as the parameters of a homography by directly aligning the planar potion of two images. The direct method may be preferable in the cases with known planes that are persistent in the view. The on-board camera system for ground vehicles is a representative example. Despite the potential accuracy, the direct method fails when the plane lacks proper texture. We propose an egomotion estimation method that is based on both the homographic constraint on a planar region, and on the epipolar constraint on generally non-planar regions, so that the both kinds of visual cues contribute to the estimation. We observe that the method improves the egomotion estimation in robustness while retaining the comparable accuracy to the direct method.
Takahiro Azuma, Shigeki Sugimoto, Masatoshi Okutomi
Intelligent Vehicles Symposium3
2010 Panoramic 3D Reconstruction Using Stereo Multi-Perspective Panorama
abstract
In this paper, we present a novel approach to imaging a panoramic (360°) environment and computing its dense depth map. Our approach adopts a multi-baseline stereo strategy using a set of multi-perspective panoramas where large baseline lengths are available. We design two image acquisition rigs for capturing such multi-perspective panoramas. The first one is composed of two parallel stereo cameras. By rotating the rig about a vertical axis, we generate four multi-perspective panoramas by resampling the regular perspective images captured by the stereo cameras. Then a depth map is estimated from the four multi-perspective panoramas and an original perspective image using a multi-baseline matching technique with different types of epipolar constraints. The second one is composed of a single camera and two mirrors. By rotating the rig, we acquire a spatio-temporal volume that is made up of the sequential images captured by the camera. Then we estimate a depth map by extracting trajectories from the spatio-temporal volume by using a multi-baseline stereo technique by considering occlusions. We can consider both rotating rigs as a single rotating camera with a very large field of view (FOV), that offers a large baseline length in depth estimation. In addition, compared with a previous approach using two multi-perspective panoramas from a single rotating camera, our approach can reduce matching errors due to image noise, repeated patterns, and occlusions by multi-baseline stereo techniques. Experimental results using both synthetic and real images show that our approach produces high quality panoramic 3D reconstruction.
Wei Jiang 0009, Shigeki Sugimoto, Masatoshi Okutomi
Int. J. Pattern Recognit. Artif. Intell.3
2009 Single-Camera Multi-baseline Stereo Using Fish-Eye Lens and Mirrors
Masao Shimizu, Masatoshi Okutomi
ACCV (2)3
2009 Disparity Estimation in a Layered Image for Reflection Stereo
Masao Shimizu, Masatoshi Okutomi
ACCV (3)2
2008 Calibration and rectification for reflection stereo
abstract
This paper presents a calibration and rectification method for single-camera range estimation using a single complex image with a transparent plate or a double-sided half-mirror plate, which are collectively called a reflection stereo. The range to an object is obtainable by finding the correspondence on a constraint line in the complex image, which consists of a surface and a rear-surface reflected image in the transparent plate, or also includes a transmitted and internal reflected image through a double-sided half-mirror plate. The range estimation requires extrinsic parameters of the reflection stereo that include the shape and position of the plate and its refraction index. The proposed method assumes that the plate is non-parallel but planar for a local region around the point of interest in the complex image. The method estimates the extrinsic parameters from a set of displacements in the complex images. Experiments using real images demonstrate the effectiveness of the proposed calibration and rectification method along with fine range estimation results.
Masao Shimizu, Masatoshi Okutomi
CVPR2
2008 Super-resolution from image sequence under influence of hot-air optical turbulence
abstract
The appearance of a distant object, when viewed through a telephoto-lens, is often deformed nonuniformly by the influence of hot-air optical turbulence. The deformation is unsteady: an image sequence can include nonuniform movement of the object even if a stationary camera is used for a static object. This study proposes a multi-frame super-resolution reconstruction from such an image sequence. The process consists of the following three stages. In the first stage, an image frame without deformation is estimated from the sequence. However, there is little detailed information about the object. In the second stage, each frame in the sequence is aligned non-rigidly to the estimated image using a non-rigid deformation model. A stable non-rigid registration technique with a B-spline function is also proposed in this study for dealing with a textureless region. In the third stage, a multi-frame super-resolution reconstruction using the non-rigid deformation recovers the detailed information in the frame obtained in the first stage. Experiments using synthetic images demonstrate the accuracy and stability of the proposed non-rigid registration technique. Furthermore, experiments using real sequences underscore the effectiveness of the proposed process.
Masao Shimizu, Shin Yoshimura, Masayuki Tanaka 0001, Masatoshi Okutomi
CVPR4
2008 Locally adaptive learning for translation-variant MRF image priors
abstract
Markov random field (MRF) models are a powerful tool in machine vision applications. However, learning the model parameters is still a challenging problem and a burdensome task. The main contribution of this paper is to propose a locally adaptive learning framework. The proposed learning framework is simple and effective learning framework for translation-variant MRF models. The key idea is to use neighboring patches as a locally adaptive training set. We use multivariate Gaussian MRF models for local image prior models. Although the Gaussian MRF models are too simple for whole natural image priors, the locally adaptive framework enables to express the prior distributions of the every observed image. These locally adaptive learning framework and the multivariate Gaussian translation-variant MRF models simplify the learning procedures. This paper also includes other two contributions; a novel iteration framework by updating the prior information, and a simple and intuitive derivation of the well-known bilateral filter. Experimental results of denoising applications demonstrate that the denoising based on the proposed locally adaptive learning framework outperforms existing high-performance denoising algorithms.
Masayuki Tanaka 0001, Masatoshi Okutomi
CVPR2
2008 Robust and accurate estimation of multiple motions for whole-image super-resolution
abstract
A robust whole-image super-resolution is highly demanded. Whole-image super-resolution requires target-region extraction and motion estimation because an image sequence usually includes multiple targets and a background. A single- motion model is insufficient to represent whole-image motion. In addition, high-accuracy motion estimation is necessary for super-resolution. Simultaneous target-region extraction and high-accuracy motion estimation are challenging problems. We propose a robust and accurate algorithm that can provide extracted single-motion regions and their motion parameters. The proposed algorithm consists of three phases: feature-based motion estimation, single-motion region extraction, and region-based motion estimation. Once we obtain motion parameters of the whole image, we can reconstruct a high-resolution whole image by applying a robust super- resolution. Results of experiments demonstrate that the proposed approach can robustly reconstruct complex multiple- target scenes.
Masayuki Tanaka 0001, Yoichi Yaguchi, Masatoshi Okutomi
ICIP3
2008 Virtual focusing image synthesis for user-specified image region using camera array
abstract
Synthetic Aperture Focusing can produce a virtual image where objects lying on a specified focal plane are focused while object lying off the plane are blurred. by averaging the multiview images warped by planar homographies. In this paper, we propose a method for efficiently estimating a focal plane from multiview stereo images so that a user-specified image region is focused in the virtual image. We estimate the 3D parameters of a focal plane, which corresponds to a planar surface in the image region. by using a fast multiview direct method with image pair selection. We select the ‘best‘ image pair by evaluating both pre-computed condition numbers of Hessian matrices and stereo baseline lengths. Our method can rapidly produce a unique virtual image, referred to as a Virtual Focal Plane (VFP) image, where the image region lying on a non-frontparallel plane in the scene is focused.
Shigeki Sugimoto, Masatoshi Okutomi
ICPR2
2007 Microscopic Surface Shape Estimation of a Transparent Plate Using a Complex Image
Masao Shimizu, Masatoshi Okutomi
ACCV (2)2
2007 Image Correspondence from Motion Subspace Constraint and Epipolar Constraint
Shigeki Sugimoto, Hidekazu Takahashi, Masatoshi Okutomi
ACCV (2)3
2007 Simultaneous Optimization of Structure and Motion in Dynamic Scenes Using Unsynchronized Stereo Cameras
abstract
In this paper, we propose a simultaneous estimation method of structure and motion in dynamic scenes. Usual methods for obtaining structure and motion using stereo cameras require two kinds of operations: stereo correspondence and tracking. Therefore, we must separately determine the correspondence between stereo images and sequential images. This necessity complicates the algorithm and increases the possibility of mismatches because of the object's motion and visibility change in the images. Our proposed method makes two contributions. The first contribution is the method of corresponding all stereo images and sequential images at once. Therefore, we can obtain the structure and motion simultaneously and more accurately. On the other hand, most stereo correspondence algorithms are limited to use under a synchronized status. In a stereo rig using unsynchronized cameras, as are most commercially available cameras, the structure cannot be obtained by stereo correspondence and triangulation because of the unknown time offset between cameras. Therefore, our second contribution is a method of estimating structure, motion, and time offset simultaneously using unsynchronized stereo cameras. This latter task is accomplished by taking advantage of the first contribution scheme. Additionally, our method requires no preprocessing such as motion segmentation for separating identical-motion objects and advance calibration of the time offset. Finally, we present the experimental results using both synthetic and real images.
Akihito Seki, Masatoshi Okutomi
CVPR2
2007 A Direct and Efficient Method for Piecewise-Planar Surface Reconstruction from Stereo Images
abstract
In this paper, we propose a direct method for 3D surface reconstruction from stereo images. We reconstruct a 3D surface by estimating all depths of the vertices of a mesh composed of piecewise triangular patches on the reference (template) image. The analyses described in this paper subsume that the deformation of the mesh between the stereo images is specified by homographies, each of which represents the deformation of a single patch. The homography deforms each patch which has 3 d.o.f under epipolar constraints. We first formulate a fast "direct" method for estimating the three parameters of a 3D plane by incorporating inverse compositional expression into the sum of squared differences (SSD) function of two stereo images. This method is about eight times faster than the conventional method. Then we extend the direct method to the estimation of the vertex depths in the mesh for reconstructing piecewise-planar surfaces. The validity of the proposed method is demonstrated through results of experiments using synthetic and real images.
Shigeki Sugimoto, Masatoshi Okutomi
CVPR2
2007 A footstep-plan-based floor sensing method using stereo images for biped robot control
abstract
In this paper, we propose a floor sensing method using stereo cameras mounted on a biped robot. In the proposed method, we first determine multiple regions of interest (ROI) in a reference image from footstep positions up to several steps, scheduled by a current footstep plan. Then the 3D plane parameters of the floor with respect to each ROI are estimated by a direct method using stereo images. We adopt the fast plane parameter estimation method [5], along with the compensation for the errors of the initial parameters by using the internal state of the robot, for the enhancement of the robustness and efficiency in the optimization process. Additionally, we estimate the shape of the floor including slopes from the set of the estimated plane parameters, and feedback the results for updating the footstep plan. The validity of the proposed method is demonstrated through on-line experiments using stereo cameras mounted on the body of a biped robot traversing a real environment.
Minami Asatani, Shigeki Sugimoto, Masatoshi Okutomi
IROS3
2006 Panoramic 3D Reconstruction Using Rotational Stereo Camera with Simple Epipolar Constraints
abstract
In this paper, we propose a novel method for panoramic 3D scene recovery using rotational stereo cameras with simple epipolar constraints. By rotating two parallel stereo cameras about a vertical axis with a constant velocity, we acquire two sampled spacio-temporal volumes which are made of the sequential images captured by a uniform angular interval. The two spacio-temporal volumes can be resampled into a set of multi-perspective panoramas. We analyze the epipolar geometry among images (panoramas and original images) of two spacio-temporal volumes. The result shows that only three types of simple epipolar constraints (epipolar line is row or column of image) exist in the two spacio-temporal volumes. Then we compute a depth map from four image pairs using a multi-baseline algorithm with the three types of epipolar constraints; that is horizontal, vertical and combination of them. Experimental results using both synthetic and real images show that our approach produces high quality panoramic 3D reconstruction.
Wei Jiang 0009, Masatoshi Okutomi, Shigeki Sugimoto
CVPR (1)2
2006 Super-Resolution using a Multi-Mixture Imaging System
abstract
Pixel mixture accelerates a read-out rate of an image, but it lowers the resolution. Although super-resolution is famous technique to improve the resolution, the super-resolution can not improve the resolution from the pixel mixture images because the pixel mixture images are low-passed to reduce aliasing. Therefore, we propose a novel imaging system that we call "multi-mixture". The proposed imaging system can generate two types of image sequences. This paper demonstrates that the super-resolution using the multi-mixture imaging system greatly improves the image resolution.
Masayuki Tanaka 0001, Masatoshi Okutomi
ICIP2
2006 Ego-motion Estimation by Matching dewarped Road Regions using Stereo Images
abstract
This paper proposes a method for vehicle ego-motion estimation using vehicle-mounted stereo cameras. Estimating ego-motion using cameras requires extraction of static regions from the images. We first use stereo images to estimate regions which correspond to the road plane and can be considered as static areas. Subsequently, we propose a virtual projection plane (VPP) image that is equivalent to the top view of the road scene. A vehicle's ego-motion is obtained by matching the sequential VPP images of the road patterns in the extracted regions. We use a vehicle-motion model and consider a matching method to easily and accurately determine the ego-motion. Finally, we present experimental results obtained using our method
Akihito Seki, Masatoshi Okutomi
ICRA2
2006 Multi-Parameter Simultaneous Estimation on Area-Based Matching
Masao Shimizu, Masatoshi Okutomi
Int. J. Comput. Vis.2
2005 Theoretical Analysis on Reconstruction-Based Super-Resolution for an Arbitrary PSF
abstract
This study presents and proves a condition number theorem for super-resolution (SR). The SR condition number theorem provides the condition number for an arbitrary space-invariant point spread function (PSF) when using an infinite number of low resolution images. A gradient restriction is also derived for maximum likelihood (ML) method. The gradient restriction is presented as an inequality which shows that the power spectrum of the PSF suppresses the spatial frequency component of the gradient of ML cost function. A Box PSF and a Gaussian PSF are analyzed with the SR condition number theorem. Effects of the gradient restriction on super-resolution results are shown using synthetic images.
Masayuki Tanaka 0001, Masatoshi Okutomi
CVPR (2)2
2005 Sub-Pixel Estimation Error Cancellation on Area-Based Matching
Masao Shimizu, Masatoshi Okutomi
Int. J. Comput. Vis.2
2004 Direct Super-Resolution and Registration Using Raw CFA Images
Tomomasa Gotoh, Masatoshi Okutomi
CVPR (2)2
2002 Camera Calibration with Precise Extraction of Feature Points using Projective Transformation
abstract
In this paper, we propose a method for precise camera calibration which conducts feature point extraction and camera parameter estimation iteratively. Many of conventional researches on camera calibration have focused on how to calculate camera parameters using data obtained from input images, that is, location of feature points. However, these input images suffer from distortions caused by perspective and lens imperfections. In our proposed method, at the beginning of the procedure, projective transformation matrices between image planes and a calibration target and lens distortion parameters are approximately estimated. These parameters are used in order to reduce the influence of distortions in input images. After the removal of distortions, feature points in processed images are localized precisely and used to update/calculate the projective transformation matrices, the lens distortion parameters and intrinsic camera parameters. These procedures are iterated until they converge and this iteration results in precise estimation of the parameters. The effectiveness of the proposed method has been recognized through experiments using synthesized data and real images.
Katsuyuki Nakano, Masatoshi Okutomi, Yuichi Hasegawa
ICRA2
2002 Robust Estimation of Planar Regions for Visual Navigation using Sequential Stereo Images
abstract
In this paper, we propose a robust method to estimate planar regions using sequential stereo images for visual navigation of an autonomous vehicle. The proposed method estimates 2D projective transformations dynamically, which represent planes in space, for both stereo images and sequential images. This can be done robustly by utilizing sequential information, i.e. previous estimation of both the projective transformations and the corresponding planar region. In addition, a method for preventing misdetection due to textureless areas is proposed. The experimental results, using sequential stereo images taken from a moving vehicle, have shown that the proposed method can work robustly even in the conditions of undulation of the road and rolling and pitching of the vehicle.
Masatoshi Okutomi, Katsuyuki Nakano, Junichi Maruyama, Tomoaki Hara
ICRA1
2002 A Simple Stereo Algorithm to Recover Precise Object Boundaries and Smooth Surfaces
Masatoshi Okutomi, Yasuhiro Katayama, Setsuko Oka
Int. J. Comput. Vis.1
2001 A Simple Stereo Algorithm to Recover Precise Object Boundaries and Smooth Surfaces
abstract
In area-based stereo matching, there is a problem called "boundary overreach", i.e. the recovered object boundary turns out to be wrongly located away from the real one. This is especially harmful to segmenting objects using depth information. A few approaches have been proposed to solve this problem. However, these techniques tend to degrade on smooth surfaces. That is, there seems to be a trade-off problem between recovering precise object edges and obtaining smooth surfaces. In this paper, we propose a new simple method to solve this problem. Using multiple stereo pairs and multiple windowing, our method detects the region where the boundary overreach is likely to occur (let us call it "BO region") and adopts appropriate methods for the BO and non-BO regions. Although the proposed method is quite simple, the experimental results have shown that it is very effective at recovering both sharp object edges at their correct locations and smooth object surfaces.
Masatoshi Okutomi, Yasuhiro Katayama, Setsuko Oka
CVPR (2)1
2001 Precise Sub-Pixel Estimation on Area-Based Matching
Masao Shimizu, Masatoshi Okutomi
ICCV2
2000 Shape Recovery of Rotating Object Using Weighted Voting of Spatio-Temporal Image
abstract
We propose a method to recover the 3D shape of an object rotating on a turnable. A spacio-temporal image is made of the sequential images taken by a single camera. Then, the trajectories which correspond to the 3D points on the surface are extracted in the spacio-temporal image by using "weighted voting" of all intensity values on a constrained surface. Since the method consequently utilize intensity information as it is, the method can recover dense 3D positions compared with the one using feature extraction and tracking. Also, it can recover concave shapes unlike the one using silhouettes of the object. The experimental results with real images show the effectiveness of our method.
Masatoshi Okutomi, Shigeki Sugimoto
ICPR1
1998 Extraction of road region using stereo images
abstract
Extracting the road region in an observed image is an important technique for visual navigation of an autonomous vehicle. In this paper, we propose a road extraction method using stereo images. The method does not rely on the existence of any specific road painting or texture. Instead, it supposes that a road (or a passable part) can be approximated by a plane. Then, a homography matrix which represents a geometric relation between the road plane and the stereo images can be computed from the stereo images. And the road region can be extracted in the observed image by transforming one image by the homography matrix and simple matching. In this method, neither a predetermined geometric relation between the cameras and the road nor a strong camera calibration are necessary. Experimental results with real scenes have shown the effectiveness of the proposed method.
Masatoshi Okutomi, Suguru Noguchi
ICPR1
1994 A Stereo Matching Algorithm with an Adaptive Window: Theory and Experiment
abstract
A central problem in stereo matching by computing correlation or sum of squared differences (SSD) lies in selecting an appropriate window size. The window size must be large enough to include enough intensity variation for reliable matching, but small enough to avoid the effects of projective distortion. If the window is too small and does not cover enough intensity variation, it gives a poor disparity estimate, because the signal (intensity variation) to noise ratio is low. If, on the other hand, the window is too large and covers a region in which the depth of scene points (i.e., disparity) varies, then the position of maximum correlation or minimum SSD may not represent correct matching due to different projective distortions in the left and right images. For this reason, a window size must be selected adaptively depending on local variations of intensity and disparity. The authors present a method to select an appropriate window by evaluating the local variation of the intensity and the disparity. The authors employ a statistical model of the disparity distribution within the window. This modeling enables the authors to assess how disparity variation, as well as intensity variation, within a window affects the uncertainty of disparity estimate at the center point of the window. As a result, the authors devise a method which searches for a window that produces the estimate of disparity with the least uncertainty for each pixel of an image: the method controls not only the size but also the shape (rectangle) of the window. The authors have embedded this adaptive-window method in an iterative stereo matching algorithm: starting with an initial estimate of the disparity map, the algorithm iteratively updates the disparity estimate for each point by choosing the size and shape of a window till it converges. The stereo matching algorithm has been tested on both synthetic and real images, and the quality of the disparity maps obtained demonstrates the effectiveness of the adaptive window method.>
Takeo Kanade, Masatoshi Okutomi
IEEE Trans. Pattern Anal. Mach. Intell.2
1993 A Multiple-Baseline Stereo
abstract
A stereo matching method that uses multiple stereo pairs with various baselines generated by a lateral displacement of a camera to obtain precise distance estimates without suffering from ambiguity is presented. Matching is performed simply by computing the sum of squared-difference (SSD) values. The SSD functions for individual stereo pairs are represented with respect to the inverse distance and are then added to produce the sum of SSDs. This resulting function is called the SSSD-in-inverse-distance. It is shown that the SSSD-in-inverse-distance function exhibits a unique and clear minimum at the correct matching position, even when the underlying intensity patterns of the scene include ambiguities or repetitive patterns. The authors first define a stereo algorithm based on the SSSD-in-inverse-distance and present a mathematical analysis to show how the algorithm can remove ambiguity and increase precision. Experimental results with real stereo images are presented to demonstrate the effectiveness of the algorithm.>
Masatoshi Okutomi, Takeo Kanade
IEEE Trans. Pattern Anal. Mach. Intell.1
1992 Color stereo matching and its application to 3-D measurement of optic nerve head
abstract
Presents a color stereo matching method. The authors first show the effect of using color information in stereo matching mathematically by using sum-of-squared-differences (SSD) criterion and then experimentally by using synthesized images. They then present a color stereo matching algorithm for a medical application. The algorithm takes into account both the intensity and the disparity variations within the matching window, and estimates the disparity at subpixel resolution exploiting color stereo images. The authors have applied the algorithm to three-dimensional measurement of optic nerve heads using stereo fundus images. The experimental results shows that the proposed stereo algorithm together with various means of displaying the results could give useful information for diagnosing and monitoring glaucoma.>
Masatoshi Okutomi, Osamu Yoshizaki, Goji Tomita
ICPR (1)1
1992 A locally adaptive window for signal matching
Masatoshi Okutomi, Takeo Kanade
Int. J. Comput. Vis.1
1991 A multiple-baseline stereo
abstract
A stereo matching method is presented which uses multiple stereo pairs with various baselines to obtain precise depth estimates without suffering from ambiguity. The stereo matching method uses multiple stereo pairs with different baselines generated by a lateral displacement of a camera. Matching is performed by computing the sum of squared-difference (SSD) values. The SSD functions for individual stereo pairs are represented with respect to the inverse depth (rather than the disparity, as is usually done), and then are simply added to produce the sum of SSDs. This resulting function is called the SSSD-in-inverse-depth. The authors define a stereo algorithm, based on the SSSD-in-inverse-depth and then present a mathematical analysis to show how the algorithm can remove ambiguity and increase precision. Experimental results for stereo images are presented to demonstrate the effectiveness of the algorithm.>
Masatoshi Okutomi, Takeo Kanade
CVPR1
1990 A locally adaptive window for signal matching
abstract
The authors presents a signal matching algorithm that can select an appropriate window size adaptively so as to obtain both precise and stable estimation of correspondences. A statistical model is presented for disparity variation within a window, and it is used to establish a link between the window size and the uncertainty of the computed disparity. This makes it possible to choose the window size that minimizes uncertainty in the disparity computed at each point. A theory is presented for the model and the resultant algorithm, together with analytical and experimental results that demonstrate their effectiveness.>
Masatoshi Okutomi, Takeo Kanade
ICCV1