Keita Takahashi 0001

dblp:46/3729-1 · DBLP profile ↗
← Back
58ranked-venue papers
21as first author
9since 2021 · last 2024
0000-0001-9429-5273ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 57 · 21 first-author · 9 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2024 Time-Efficient Light-Field Acquisition Using Coded Aperture and Events
abstract
We propose a computational imaging method for time-efficient light-field acquisition that combines a coded aperture with an event-based camera. Differentfrom the conventional coded-aperture imaging method, our method applies a sequence of coding patterns during a single exposure for an image frame. The parallax information, which is related to the differences in coding patterns, is recorded as events. The image frame and events, all of which are measured in a single exposure, are jointly used to computationally reconstruct a light field. We also designed an algorithm pipeline for our method that is end-to-end trainable on the basis of deep optics and compatible with real camera hardware. We experimentally showed that our method can achieve more accurate reconstruction than several other imaging methods with a single exposure. We also developed a hardware prototype with the potential to complete the measurement on the camera within 22 msec and demonstrated that light fields from real 3-D scenes can be obtained with convincing visual quality. Our software and supplementary video are available from our project website1:
Shuji Habuchi, Keita Takahashi 0001, Chihiro Tsutake, Toshiaki Fujii, Hajime Nagahara
CVPR2
2024 Mono+Sub: Compressing Light Field as Monocular Image and Subsidiary Data
abstract
A light field is usually represented as a set of multi-view images captured from a two-dimensional (2-D) array of viewpoints and requires a large amount of data compared with a standard 2-D image. We propose a 2-D compatible light-field compression method for encoding a light field as a 2-D monocular image and subsidiary data. In terms of the image quality, we prioritize the central image (regarded as the 2-D monocular image) over the other images in the light field, because the light field is considered an extension of the 2-D monocular image. To this end, we encode and decode the monocular image using a standard image codec and introduce a learned encoder and decoder pair for the subsidiary data. Experimental results indicate that our method achieved promising rate-distortion performance, especially for extremely low bit-rate ranges. Even though our method requires only a small amount of subsidiary data compared with those for the monocular image, the entire light field can be reconstructed with reasonable visual quality.
Ryosuke Imazu, Chihiro Tsutake, Keita Takahashi 0001, Toshiaki Fujii
VCIP3
2024 Warm-start NeRF: Accelerating Per-scene Training of NeRF-based Light-Field Representation
abstract
A light field is represented as a set of multi-view images captured from a dense 2-D array of viewpoints. To treat a light field as being continuous, we represent it as a neural radiance field (NeRF), which is a learned representation of a 3-D scene. NeRFs are renowned for their ability to reconstruct a target 3-D scene with compelling visual quality, but they are slow to train. A solution for this problem is to use a tiny neural network and trainable volumetric features as the scene representation, which is considered the baseline of our research. For further acceleration, we propose a method for warm-starting the per-scene training by setting good initial values for the trainable parameters. To this end, we introduce another encoder network to obtain the initial volumetric features from the target light field. Starting with the appropriate initial values, our method can achieve better rendering quality with fewer training iterations than the baseline.
Takuto Nishio, Chihiro Tsutake, Keita Takahashi 0001, Toshiaki Fujii
VCIP3
2022 Acquiring a Dynamic Light Field through a Single-Shot Coded Image
abstract
We propose a method for compressively acquiring a dynamic light field (a 5-D volume) through a single-shot coded image (a 2-D measurement). We designed an imaging model that synchronously applies aperture coding and pixel-wise exposure coding within a single exposure time. This coding scheme enables us to effectively embed the original information into a single observed image. The observed image is then fed to a convolutional neural network (CNN) for light-field reconstruction, which is jointly trained with the camera-side coding patterns. We also developed a hardware prototype to capture a real 3-D scene moving over time. We succeeded in acquiring a dynamic light field with 5x5 viewpoints over 4 temporal sub-frames (100 views in total)from a single observed image. Repeating capture and reconstruction processes over time, we can acquire a dynamic light field at 4x the frame rate of the camera. To our knowledge, our method is the first to achieve a finer temporal resolution than the camera itself in compressive light-field acquisition. Our software is available from our project webpage.11https://www.fujii.nuee.nagoya-u.ac.jp/Research/CompCam2
Ryoya Mizuno, Keita Takahashi 0001, Michitaka Yoshida, Chihiro Tsutake, Toshiaki Fujii, Hajime Nagahara
CVPR2
2022 Unsupervised disparity estimation from light field using plug-and-play weighted warping loss
abstract
We investigated disparity estimation from a light field using a convolutional neural network (CNN). Most of the methods implemented a supervised learning framework, where the predicted disparity map was compared directly to the corresponding ground-truth disparity map in the training stage. However, light field data accompanied with ground-truth disparity maps were insufficient and rarely available for real-world scenes. The lack of training data resulted in limited generality of the methods trained with them. To tackle this problem, we took a simple Figure-and-play approach to remake a supervised method into an unsupervised (self-supervised) one. We replaced the loss function of the original method with one that does not depend on the ground-truth disparity maps. More specifically, our loss function is designed to indirectly evaluate the accuracy of the disparity map by using warping errors among the input light field views. We designed pixel-wise weights to properly evaluate the warping errors in the presence of occlusions, and an edge loss to encourage edge alignment between the image and the disparity map. As a result of this unsupervised learning framework, our method can use more abundant training datasets (even those without ground-truth disparity maps) than the original supervised method. Our method was evaluated on computer-generated scenes (4D Light Field Benchmark) and real-world scenes captured by Lytro Illum cameras. Our method achieved the state-of-the-art performance as an unsupervised method on the benchmark. We also demonstrated that our method can estimate disparity maps more accurately than the original supervised method for various real-world scenes.
Taisei Iwatsuki, Keita Takahashi 0001, Toshiaki Fujii
Signal Process. Image Commun.2
2022 Denoising multi-view images by soft thresholding: A short-time DFT approach
abstract
Short-time discrete Fourier transform (ST-DFT) is known as a promising technique for image and video denoising. The seminal work by Saito and Komatsu hypothesized that natural video sequences can be represented by sparse ST-DFT coefficients and noisy video sequences can be denoised on the basis of statistical modeling and shrinkage of the ST-DFT coefficients. Motivated by their theory, we develop an application of ST-DFT for denoising multi-view images. We first show that multi-view images have sparse ST-DFT coefficients as well and then propose a new statistical model, which we call the multi-block Laplacian model, based on the block-wise sparsity of ST-DFT coefficients. We finally utilize this model to carry out denoising by solving a convex optimization problem, referred to as the least absolute shrinkage and selection operator. A closed-form solution can be computed by soft thresholding, and the optimal threshold value is derived by minimizing the error function in the ST-DFT domain. We demonstrate through experiments the effectiveness of our denoising method compared with several previous denoising techniques. Our method implemented in Python language is available from https://github.com/ctsutake/mviden.
Keigo Tomita, Chihiro Tsutake, Keita Takahashi 0001, Toshiaki Fujii
Signal Process. Image Commun.3
2021 An Efficient Image Compression Method Based On Neural Network: An Overfitting Approach
abstract
Over the past decade, nonlinear image compression techniques based on neural networks have been rapidly developed to achieve more efficient storage and transmission of images compared with conventional linear techniques. A typical nonlinear technique is implemented as a neural network trained on a vast set of images, and the latent representation of a target image is transmitted. In contrast to the previous nonlinear techniques, we propose a new image compression method in which a neural network model is trained exclusively on a single target image, rather than a set of images. Such an overfitting strategy enables us to embed fine image features in not only the latent representation but also the network parameters, which helps reduce the reconstruction error against the target image. The effectiveness of our method is validated through a comparison with conventional image compression techniques in terms of a rate-distortion criterion.
Yu Mikami, Chihiro Tsutake, Keita Takahashi 0001, Toshiaki Fujii
ICIP3
2021 Factorized Modulation For Singleshot Lightfield Acquisition
abstract
A light field (LF), which is represented as a set of dense multiview images, has been utilized in various 3-D applications. To make LF acquisition more efficient, researchers have investigated compressive sensing methods by incorporating modulation or coding functions into the camera. In this work, we investigate a challenging case of compressive LF acquisition in which an entire LF should be reconstructed from only a single coded image. To achieve this goal, we propose a new modulation scheme called factorized modulation that can approximate arbitrary 4-D modulation patterns in a factorized manner. Our method can be hardware-implemented by combining the architectures for coded aperture and pixel-wise coded exposure imaging. The modulation pattern is jointly optimized with a CNN-based reconstruction algorithm. Our method is validated through extensive evaluations against other modulation schemes.
Kohei Tateishi, Kohei Sakai, Chihiro Tsutake, Keita Takahashi 0001, Toshiaki Fujii
ICIP4
2021 An Efficient Compression Method For Sign Information Of DCT Coefficients Via Sign Retrieval
abstract
Compression of the sign information of discrete cosine transform coefficients is an intractable problem in image compression schemes due to the equiprobable occurrence of the sign bits. To overcome this difficulty, we propose an efficient compression method for such sign information based on phase retrieval, which is a classical signal restoration problem attempting to find the phase information of discrete Fourier transform coefficients from their magnitudes. In our compression strategy, the sign bits of all the AC components in the cosine domain are excluded from a bitstream at the encoder and are complemented at the decoder by solving a sign recovery problem, which we call sign retrieval. The experimental results demonstrate that the proposed method outperforms previous techniques for sign compression in terms of a rate-distortion criterion. Our method implemented in Python language is available from https://github.com/ctsutake/sr.
Chihiro Tsutake, Keita Takahashi 0001, Toshiaki Fujii
ICIP2
2020 Acquiring Dynamic Light Fields Through Coded Aperture Camera
Kohei Sakai, Keita Takahashi 0001, Toshiaki Fujii, Hajime Nagahara
ECCV (19)2
2019 A 3-D Display Pipeline from Coded-Aperture Camera to Tensor Light-Field Display Through CNN
abstract
We propose an efficient pipeline from input to output for a tensor light-field display. Conventionally, a dense light field (i.e., tens of images taken with narrow viewpoint intervals) is required as an input in such displays. However, obtaining dense light fields is a challenging task for real scenes. To make the acquisition process more efficient, we adopted a coded-aperture camera as an input device, which is suitable for acquiring dense light fields in a compressive manner. Moreover, we modeled the entire process from acquisition to display using a convolutional neural network. As a result of training the network on a massive light field data, we can reproduce the whole light field on the display from only a few images taken with the camera. Both simulative and real experiments were conducted to show the effectiveness of our method.
Keita Maruyama, Yasutaka Inagaki, Keita Takahashi 0001, Toshiaki Fujii, Hajime Nagahara
ICIP3
2019 LF-TSP: Traveling salesman problem for HEVC-based light-field coding
abstract
We studied a coding scheme where light field (LF) images (dense multi-view images) are regarded as a sequence of temporal video frames and encoded with video codecs such as High Efficiency Video Coding (HEVC). An important issue with this scheme is how to determine the frame order of the LF images. We propose a method to find the optimum frame order through a formulation of the traveling salesman problem (TSP). Under the assumption that video codecs are more effective with temporally smooth videos, our method, named LF-TSP, defines frame-to-frame distances for each image pairs in an LF, and attempted to find the shortest route that visits all frames. Experiments showed that our method achieved an overall better rate-distortion performance than several previous methods.
Kota Imaeda, Kohei Isechi, Keita Takahashi 0001, Toshiaki Fujii, Yukihiro Bandoh, Takehito Miyazawa, Seishi Takamura, Atsushi Shimizu
VCIP3
2018 Learning to Capture Light Fields Through a Coded Aperture Camera
Yasutaka Inagaki, Yuto Kobayashi, Keita Takahashi 0001, Toshiaki Fujii, Hajime Nagahara
ECCV (7)3
2018 How Should we Handle 4D Light Fields with CNNS?
abstract
We investigated how we should handle high dimensional light fields (LFs) with convolutional neural networks (CNNs). An LF is a 4-D signal representation of light rays, and it is interpreted as a set of dense multi-view images. As an important building block of various light field applications, we focused on signal restoration problems for LFs, and we adopted CNN s as the solver for them because of its striking performance on the conventional 2-D images. In applying CNN s, the high dimensionality of LFs should be carefully addressed. Instead of treating the full 4- D signal as it is, we followed a divide and conquer strategy. Specifically, we cascade two or three CNNs, each of which works only on 2-D subspace of the full 4-D LFs. Combining CNN s that work on different subspaces, we can eventually handle the full 4- D structure. Moreover, considering different properties of those subspaces, we experimentally explored the best combination in which different subspaces are cascaded. Although our experiments are currently limited to a denoising problem, the lessons found from our results will benefit the prospective research on the full4-D LF processing.
Shu Fujita, Keita Takahashi 0001, Toshiaki Fujii
ICIP2
2018 Fast and Robust Disparity Estimation for Noisy Light Fields
abstract
Depth (disparity) estimation from a light field (a set of dense multiview images) has attracted much research interest recently. This paper is focused on how to handle noisy light field for disparity estimation' because if left as it is the noise deteriorates the accuracy of estimated disparity maps. Several researchers have worked on this problem, e.g. by introducing disparity cues that are robust to noise. However, it is not easy to break the trade-off between the accuracy and computational speed. To tackle this trade-off, we have integrated a fast denoising scheme in a fast disparity estimation framework that works in the epipolar plane image (EPI) domain. Specifically, we found that a simple 1-D slanted filter is very effective for reducing noise while preserving the underlying structure in an EPI. Experimental results show that our method can achieve good accuracy with much less computational time compared to some state-of-the-art methods.
Gou Houben, Shu Fujita, Keita Takahashi 0001, Toshiaki Fujii
ICIP3
2018 Scalable Light Field Coding Using Weighted Binary Images
abstract
We propose an efficient coding scheme for a dense light field, i.e., a set of multi-viewpoint images taken with very small viewpoint intervals. The key idea behind our proposal is that a light field is represented only using weighted binary images, where several binary images and corresponding weight values are to be chosen to optimally approximate the light field. The coding scheme derived from this idea is completely different from those of modern image/video coding standards. However, we found that our scheme can achieve comparable coding efficiency (rate-distortion performance) to that of modern highly-sophisticated video codecs. Moreover, the decoding process of our scheme is extremely simple, which will lead to a faster and less power-hungry decoder than those of the modern codecs. Furthermore, our scheme can be made scalable, where the accuracy of the decoded light field is improved in a progressive manner as we use more encoded information. Thanks to the divide-and-conquer strategy adopted for the scalable coding, we can also drastically reduce the computational complexity of the encoding process.
Koji Komatsu, Keita Takahashi 0001, Toshiaki Fujii
ICIP2
2018 From Focal Stack to Tensor Light-Field Display
abstract
We propose a method of using a focal stack, i.e., a set of differently focused images, as the input for a novel light field display called a "tensor display." Although this display consists of only a few light attenuating layers located in front of a backlight, it can be viewed from many directions (angles) simultaneously without the resolution of each viewing direction being sacrificed. Conventionally, a transmittance pattern is calculated for each layer from a light field, namely, a set of dense multi-view images (typically dozens) that are to be observed from different directions. However, preparing such a massive amount of images is often cumbersome for real objects. We developed a method that does not require a complete light field as the input; instead, a focal stack composed of only a few differently focused images is directly transformed into layer patterns. Our method greatly reduces the cost of acquiring data while also maintaining the quality of the output light field. We validated the method with experiments using synthetic light field datasets and a focal stack acquired by an ordinary camera.
Keita Takahashi 0001, Yuto Kobayashi, Toshiaki Fujii
IEEE Trans. Image Process.1
2017 From focal stacks to tensor display: A method for light field visualization without multi-view images
abstract
A new type of light field display called a tensor display was investigated. Although this display consists of only a few light attenuating layers located in front of a backlight, many views can be emitted in different directions simultaneously without sacrificing the resolution of each view. The transmittance pattern of each layer is calculated from a light field, namely, a set of dense multi-view images (typically dozens) that are to be observed from different directions. However, preparing such images is often cumbersome for real objects. We propose a method that does not require multi-view images as the input; instead, a focal stack composed of only a few differently focused images is directly transformed into the layer patterns. Our method greatly reduces the data acquisition cost while also maintaining the quality of the output light field. We validated the method with experiments using synthetic light field datasets and a focal stack acquired by an ordinary camera.
Yuto Kobayashi, Keita Takahashi 0001, Toshiaki Fujii
ICASSP2
2017 Good group sparsity prior for light field interpolation
abstract
A light field, which is equivalent to a dense set of multi-view images, has various applications such as depth estimation and 3D display. One of the essential problems is light field interpolation, which is obtaining sufficiently dense views from sparser views. The accuracy of interpolation will be enhance by exploiting an inherent property of a light field. Specifically, an epipolar plane image (EPI), which is a 2D subset of the 4D light field, consists of many lines. This structure induces a sparse representation in the frequency domain, where most of the energy resides on a line passing through the origin. On the basis of this observation, we propose a group sparsity prior suitable for light fields to fully exploit their line structure for interpolation. Our experimental results show that the proposed method can achieve better quality than a state-of-the-art shearlet-based method.
Keita Takahashi 0001, Shu Fujita, Toshiaki Fujii
ICIP1
2017 PCA-coded aperture for light field photography
abstract
A light field, which is often understood as a set of dense multi-view images, has been utilized in various 2D/3D applications. Efficient light field acquisition using a coded aperture camera is the target problem considered in this paper. Specifically, the entire light field, which consists of many images, should be reconstructed from only a few images that are captured through different aperture patterns. In previous work, this problem has often been discussed from the context of compressed sensing (CS). In contrast, we formulated this problem from the perspective of principal component analysis (PCA) to derive optimal non-negative aperture patterns and a straight-forward reconstruction algorithm. Even though it is based on a conventional technique, our method has proven to be more accurate and much faster than a state-of-the-art CS-based method.
Yusuke Yagi, Keita Takahashi 0001, Toshiaki Fujii, Toshiki Sonoda, Hajime Nagahara
ICIP2
2016 Disparity estimation from light fields using sheared EPI analysis
abstract
Structure tensor analysis on epipolar plane images (EPIs) is a successful approach to estimate disparity from a light field, i.e. a dense set of multi-view images. However, the disparity range allowable for the light field is limited, because the estimation becomes less accurate as the range of disparities become larger. To overcome this limitation, we propose a new method called sheared EPI analysis, where EPIs are sheared before the structure tensor analysis. The results of analysis obtained with different shear values are integrated into a final disparity map. As verified by extensive evaluations on 12 datasets with large disparity ranges, our method is comparably accurate to and much faster than a multi-view stereo method.
Keita Takahashi 0001, Toshiaki Fujii
ICIP2
2015 Joint directional-positional multiplexing for light field acquisition by Kronecker compressed sensing
abstract
In this paper, we propose a joint and unified framework to compressively capture a light field in the consideration of both directional and positional multiplexing based on Kronecker compressed sensing (KCS). First of all, both of the 2D angular and 2D spatial correlations of the light field can be fully utilized in the compressive acquisition, and the multiplexing is more flexible and balanced during the acquisition. Secondly, other types of light field acquisition can be unified into our proposed framework. In the experiment, it is shown that more balanced allocation between directional and positional multiplexing achieves better reconstruction quality of light field given the same number of total acquisitions. Furthermore, the experimental result also illustrates that the proposed method can capture a light field with full resolution and achieve better reconstruction quality than other previous methods.
Keita Takahashi 0001, Mehrdad Panahpour Tehrani, Toshiaki Fujii
ICASSP2
2015 Reconstruction of compressively sampled light fields using a weighted 4D-DCT basis
abstract
The coded aperture/mask technique enables us to capture light field data in a compressive way through a single camera. A pixel value recorded by such a camera is a summation of the light rays that pass though different positions on the coded aperture/mask. The target light field can be reconstructed from the recorded pixel values by using prior information of the light field signal. As prior information, a dictionary (light field atoms), which was learned from training datasets, was used in the current state of the art. Meanwhile, it was reported that general bases such as DCT were not suitable to efficiently represent prior information. In this work, however, we demonstrate that a 4D-DCT basis works surprisingly better if it is combined with a weighting scheme in which the amplitude difference in DCT coefficients is considered. Simulation results using 18 light field datasets are reported to show the superior performance of the weighted 4D-DCT basis to the learned dictionary.
Yusuke Miyagi, Keita Takahashi 0001, Mehrdad Panahpour Tehrani, Toshiaki Fujii
ICIP2
2015 Super-resolution image synthesis using the physical pixel arrangementofalight field camera
abstract
We propose a method for super-resolution image synthesis that accurately handles the physical pixel arrangement of a light field (plenoptic) camera. We use a Lytro camera to obtain 4D light field data (a set of multi-viewpoint images) through a micro-lens array. The light field data are multiplexed on a single image sensor, and thus, the data is first de-multiplexed into a set of multi-viewpoint (sub-aperture) images. However, the de-multiplexing process usually involves interpolation of the original data such as demosaicing for a color filter array and pixel resampling for the non-square micro-lens arrangement. During this interpolation, some information is added or lost to/from the original data. In contrast, our method can preserve the originally captured data as they are, and directly use them for the super-resolution image synthesis, where the super-resolved image and the corresponding depth map are alternatively refined. We experimentally demonstrate that our method can achieve higher image quality than that with a standard Light Field Toolbox.
Kazuki Ohashi, Keita Takahashi 0001, Mehrdad Panahpour Tehrani, Toshiaki Fujii
ICIP2
2015 Rank analysis of a light field for dual-layer 3D displays
abstract
In this paper, a new type of 3D display, called a layered light-field display, was investigated. By using only a few light-attenuating layers located in front of a backlight, this display can present many views in different directions simultaneously without sacrificing the resolutions of each view. The essential factor for efficient layer-based representation, which has not been deeply analyzed in previous works, is redundancy- namely, a low rank structure-of the light-field data. Accordingly, to reveal the origin of the redundancy, a generative model, in which a textured surface located at a certain depth generates a light field, was formulated and evaluated in this paper. Our theoretical analysis shows that the redundancy depends on not only the texture complexity but also the depth of the surface from the light attenuating layers. The theoretical model was validated through experimental simulation of the display.
Keita Takahashi 0001, Toyohiro Saito, Mehrdad Panahpour Tehrani, Toshiaki Fujii
ICIP1
2015 Data format and view synthesis for free-viewpoint video streaming of super multiview video
abstract
Our goal is to propose and evaluate a new streaming data format that can be adapted to the limited bandwidth and capable of free-viewpoint video streaming using super multi-view video plus depth (MVD). Additionally, we proposed a view synthesis method for our data format. Given a requested free-viewpoint, we use the two closest views and corresponding depth maps to perform free-viewpoint video synthesis. The new data format consists of all views and corresponding depth maps in a lowered resolution, and the two closest views to the requested viewpoint in the high resolution. When the requested viewpoint changes, the two closest viewpoints will change, but one or both views are transmitted only in the low resolution during periods of large round-trip delay time. Therefore, the resolution compensation is required before view synthesis. Experimental results show that our proposed framework achieves view synthesis quality close to view synthesis using high resolution multi-view video plus depth.
Takaaki Emori, Mehrdad Panahpour Tehrani, Keita Takahashi 0001, Toshiaki Fujii
PCS3
2015 View synthesis using superpixel based inpainting capable of occlusion handling and hole filling
abstract
The existing virtual view synthesis methods generate the images with many artifacts that are annoying, especially for forward virtual viewpoint, and virtual viewpoint generated by reference views with large baseline, due to occlusions and the limited sampling density. In this paper, we propose a new view synthesis method, robust to the above-mentioned problem, consist of three steps, using stereo contents. Firstly, view plus depth data of each viewpoint is 3D warped to the virtual viewpoint. We determine which neighboring pixels should be connected or kept isolated. Polygons enclosed by the connected pixels, i.e. superpixel, are interpolated. Secondly, we blend those warped images by comparing each pixel's depth value to obtain the virtual view, in which non-occlusion holes have already been interpolated by the process in the first step. Thirdly, the remaining holes are filled by inpainting. Our experimental results and comparisons show that the proposed view synthesis method allows smoother view reconstruction, while holes due to occlusion and 3D warping are filled with less artifacts.
Tomoyuki Tezuka, Mehrdad Panahpour Tehrani, Kazuyoshi Suzuki, Keita Takahashi 0001, Toshiaki Fujii
PCS4
2014 Least MSE Regression for View Synthesis
abstract
View synthesis is the process of combining given multi-view images to generate an image from a new viewpoint. Assuming that each pixel of the new view is obtained as the weighted sum of the corresponding pixels from the input views, we focus on the problem of how to optimize the weight for each of the input views. Our weighting method is called least mean squared error (MSE) regression because it is formulated as a regression problem in which second order statistics among the viewpoints are exploited to minimize the MSE of the resulting image. More specifically, the affinity across the viewpoints is represented as a covariance and approximated using a linear model whose parameters are adapted for each dataset. By using the approximated covariance, the optimal weights can be successfully estimated. As a result, the weights derived using our method are data dependent and significantly differ from those obtained using current empirical methods such as distance penalty. Our method is still effective if the given correspondence is not completely accurate due to noise. We report on experimental results using several multi-view datasets to validate our theory and method.
Keita Takahashi 0001, Toshiaki Fujii
3DV1
2014 Synthesis Error COmpeNsateD Multiview Video plus Depth for representation of multiview video
abstract
SECOND-MVD (Synthesis Error COmpeNsateD Multiview Video plus Depth) is an alternative 3D format that we introduce for representation of multiview video. In this data format, images at some viewpoint remain original, and the others are converted to a novel format. Residual based representation, such as layered depth video and free-viewpoint TV data unit were proposed. We propose hybrid image that not only consists of residual but also remainder pixels. Generation and reconstruction process of a hybrid image uses virtual image synthesized by the images that remained original in SECOND-MVD. In this paper, we investigate the compression performance of SECOND-MVD using hybrid image. Experiments demonstrate reduction in bit rate using hybrid image against residual image in SECOND-MVD framework.
Mehrdad Panahpour Tehrani, Akio Ishikawa, Makoto Okui, Naomi Inoue, Keita Takahashi 0001, Toshiaki Fujii
ICASSP5
2013 Unified environment-adaptive control of accompanying robots using artificial potential field
Kazushi Nakazawa, Keita Takahashi 0001, Masahide Kaneko
HRI2
2013 View interpolation sensitive to pixel positions
abstract
Theoretical analysis and an optimization scheme to solve the basic problem with view interpolation are presented. Two error factors, i.e., disparity error and pixel interpolation error, are considered. I found that the method of view interpolation should change according to which error factor is dominant. When disparity information is very accurate, the method of view interpolation should be especially sensitive to pixel positions, which has rarely been noticed in previous studies.
Keita Takahashi 0001
ICIP1
2012 Image Segmentation using Dual Distribution Matching
abstract
We propose an image segmentation method that divides an image into foreground and background regions when the approximate color distributions for these regions are given.Our approach was inspired by global consistency measures that directly evaluate the similarity between a given distribution and the distribution of the resulting segmentation, which were recently proposed in order to overcome the limitations of traditional pixelwise (local) consistency measures.The main feature of our proposal is that it uses two (foreground and background) input distributions, which increases the robustness compared to previous studies.To achieve this, we formulated a new mathematical model that describes the consistencies between the two input distributions and the segmentation, in which weighting parameters for the two distribution matching terms are set to be approximately proportional to the size of the foreground and background areas.We call this dual distribution matching (DDM).We also derived an optimization method that uses graph cuts.Experimental results that show the effectiveness of our method and comparisons between local and global consistency measures are presented.
Tatsunori Taniai, Viet Quoc Pham, Keita Takahashi 0001, Takeshi Naemura
BMVC3
2012 Theoretical analysis on interframe predictive coding with subpixel displacement accuracy - An exhaustive approach
abstract
A theoretical analysis on interframe predictive coding is presented in which special attention is paid on two issues that are practically important. First, the displacements between the target and reference images are usually given with limited accuracy due to quantization. Second, when the displacements are provided in subpixel accuracies, interpolation between the pixels is necessary to produce an estimation image. Our formulation takes an exhaustive approach in which we faithfully follow the process of subpixel interpolations with limited displacement accuracies. Our analysis can be used to determine how accurate the displacement vectors should be and which subpixel interpolation method should be used to optimize the rate-distortion performance.
Keita Takahashi 0001, Masahide Kaneko
VCIP1
2012 Theoretical Analysis of View Interpolation With Inaccurate Depth Information
abstract
A problem of view interpolation from a pair of rectified stereo images with inaccurate depth information is addressed. Errors in geometric information greatly affect the quality of the resulting images since inaccurate geometry causes miscorrespondences between the input images. A new theory for quantitatively analyzing the effect of depth errors and providing a principled optimization scheme based on the mean-squared error metric is proposed. The theory clarifies that, if the probabilistic distribution of the depth errors is given, an optimized view-interpolation scheme that outperforms conventional linear interpolation can be derived. It also reveals that, under specific conditions, linear interpolation is acceptable as an approximation of the optimized-interpolation scheme. Furthermore, band limitation combined with linear interpolation is also analyzed, leading to an optimal cutoff frequency, which achieves better results than the antialias scheme proposed in previous studies. Experimental results using real scenes are also presented to confirm this theory.
Keita Takahashi 0001
IEEE Trans. Image Process.1
2011 Foreground-background segmentation using iterated distribution matching
abstract
This paper addresses the problem of image segmentation with a reference distribution. Recent studies have shown that segmentation with global consistency measures outperforms conventional techniques based on pixel-wise measures. However, such global approaches require a precise distribution to obtain the correct extraction. To overcome this strict assumption, we propose a new approach in which the given reference distribution plays a guiding role in inferring the latent distribution and its consistent region. The inference is based on an assumption that the latent distribution resembles the distribution of the consistent region but is distinct from the distribution of the complement region. We state the problem as the minimization of an energy function consisting of global similarities based on the Bhattacharyya distance and then implement a novel iterated distribution matching process for jointly optimizing distribution and segmentation. We evaluate the proposed algorithm on the GrabCut dataset, and demonstrate the advantages of using our approach with various segmentation problems, including interactive segmentation, background subtraction, and co-segmentation.
Viet Quoc Pham, Keita Takahashi 0001, Takeshi Naemura
CVPR2
2011 Super-resolution plane sweeping for free-viewpoint image synthesis
abstract
Free-viewpoint image synthesis (FVIS) refers to the process of generating novel viewpoint images from a set of multi-view images. Most of the conventional FVIS methods were based on image blending, so that they are subject to a fundamental limitation in resolution: the output resolution is lower than or at most equal to that of the input images. A reasonable approach to overcome this limitation is to replace image blending with reconstruction-based super-resolution. Following this idea, we propose a new FVIS method named as super-resolution plane sweeping by extending general plane sweeping methods. We also propose an adaptive weighting scheme to make super-resolution reconstruction operate only on the pixels where it improve the quality. Experimental results with real images are presented to show the effectiveness of our method.
Keita Takahashi 0001, Masato Ishii, Takeshi Naemura
ICIP1
2011 Rate-distortion analysis of super-resolution image/video decoding
abstract
An image/video communication scenario with super-resolution (SR) decoding, where the decoded images are upsampled using SR reconstruction at the receiver side, is analyzed from a theoretical perspective. To formulate the rate-distortion performance for such cases, we propose a new numerical model that combines a frequency-domain SR model and a rate-distortion theory for lossy image compression. We considered several factors that affect the reconstruction quality, and revealed that SR decoding performs better in low bitrates. We also conducted real-image simulations and confirmed that both the numerical analysis and real-image simulations exhibit quite similar tendencies, which supports the effectiveness of our numerical model.
Keita Takahashi 0001, Takeshi Naemura, Masayuki Tanaka 0001
ICIP1
2011 Theoretical Analysis of Multi-view Camera Arrangement and Light-Field Super-Resolution
Ryo Nakashima, Keita Takahashi 0001, Takeshi Naemura
PSIVT (1)2
2011 Super-Resolved Free-Viewpoint Image Synthesis Using Semi-global Depth Estimation and Depth-Reliability-Based Regularization
Keita Takahashi 0001, Takeshi Naemura
PSIVT (1)1
2010 Theory of Optimal View Interpolation with Depth Inaccuracy
Keita Takahashi 0001
ECCV (4)1
2010 Performance analysis on multi-view coding with depth map distortion
abstract
Multi-view video plus depth (MVD) representations have been studied for 3-D TV and free-viewpoint video applications. In this paper, I quantitatively analyze the effect of depth map distortion on the performance of prediction-based multi-view coding. A new metric of this effect is proposed based on the rate-distortion theory, and successfully applied to both real image data and a theoretical model of view-synthesis prediction. This study reveals the types of images that are more/less sensitive to depth map distortion, which will be useful for designing/optimizing MVD coding schemes.
Keita Takahashi 0001
ICIP1
2010 Bounding-Box Based Segmentation with Single Min-cut Using Distant Pixel Similarity
abstract
This paper addresses the problem of interactive image segmentation with a user-supplied object bounding box. The underlying problem is the classification of pixels into foreground and background, where only background information is provided with sample pixels. Many approaches treat appearance models as an unknown variable and optimize the segmentation and appearance alternatively, in an expectation maximization manner. In this paper, we describe a novel approach to this problem: the objective function is expressed purely in terms of the unknown segmentation and can be optimized using only one minimum cut calculation. We aim to optimize the trade-off of making the foreground layer as large as possible while keeping the similarity between the foreground and background layers as small as possible. This similarity is formulated using the similarities of distant pixel pairs. We evaluated our algorithm on the GrabCut dataset and demonstrated that high-quality segmentations were attained at a fast calculation speed.
Viet Quoc Pham, Keita Takahashi 0001, Takeshi Naemura
ICPR2
2010 Direction-adaptive hierarchical decomposition for image coding
abstract
A new model of decomposing an image hierarchically into direction-adaptive subbands using pixel-wise direction estimation is presented. For each decomposing operation, an input image is divided into two parts: a base image subsampled from the input image and subband components. The subband components consist of residuals of estimating the pixels skipped through the subsampling, which ensures the invertibility of the decomposition. The estimation is performed in a direction-adaptive way, whose optimal direction is determined by a L1 norm criterion for each pixel, aiming to achieve good energy compaction that is suitable for image coding. Furthermore, since the L1 norms are obtained from the base image alone, we do not need to retain the directional information explicitly, which is another advantage of our model. Experimental results show that the proposed model can achieve lower entropy than conventional Haar or D5/3 discrete wavelet transform in case of lossless coding.
Tomokazu Murakami, Keita Takahashi 0001, Takeshi Naemura
PCS2
2009 Real-Time Video Matting Based on Bilayer Segmentation
Viet Quoc Pham, Keita Takahashi 0001, Takeshi Naemura
ACCV (2)2
2009 Joint rendering and segmentation of free-viewpoint images
abstract
This paper presents a method that jointly performs synthesis of free-viewpoint images and object segmentation from that viewpoint. This method works efficiently and online by sharing a certain calculation process among rendering and segmentation steps. Since the segmentation is performed for arbitrary viewpoints directly, the extracted object can be superimposed onto another 3-D scene with geometric consistency. Experimental results using a 25-camera array show the effectiveness of our method.
Masato Ishii, Keita Takahashi 0001, Takeshi Naemura
ICIP2
2009 Live Video Segmentation in Dynamic Backgrounds Using Thermal Vision
Viet Quoc Pham, Keita Takahashi 0001, Takeshi Naemura
PSIVT2
2009 TransCAIP: A Live 3D TV System Using a Camera Array and an Integral Photography Display with Interactive Control of Viewing Parameters
abstract
The system described in this paper provides a real-time 3D visual experience by using an array of 64 video cameras and an integral photography display with 60 viewing directions. The live 3D scene in front of the camera array is reproduced by the full-color, full-parallax autostereoscopic display with interactive control of viewing parameters. The main technical challenge is fast and flexible conversion of the data from the 64 multicamera images to the integral photography format. Based on image-based rendering techniques, our conversion method first renders 60 novel images corresponding to the viewing directions of the display, and then arranges the rendered pixels to produce an integral photography image. For real-time processing on a single PC, all the conversion processes are implemented on a GPU with GPGPU techniques. The conversion method also allows a user to interactively control viewing parameters of the displayed image for reproducing the dynamic 3D scene with desirable parameters. This control is performed as a software process, without reconfiguring the hardware system, by changing the rendering parameters such as the convergence point of the rendering cameras and the interval between the viewpoints of the rendering cameras.
Yuichi Taguchi, Takafumi Koike, Keita Takahashi 0001, Takeshi Naemura
IEEE Trans. Vis. Comput. Graph.3
2008 Rate-distortion performance of multi-view image coding with subsampling of viewpoints
abstract
In this paper, we consider the rate-distortion performance of a multi- view image set with subsampling of the viewpoints. As a basic analysis, we compare two scenarios: (i) all images would be coded and transmitted in even quality, and (ii) a half of the images would be discarded by subsampling at the sender side, and the remaining half would be coded and transmitted. In the second scenario, the discarded images would be reconstructed at the receiver side using some view interpolation technology. We first introduce a theoretical model describing the rate-distortion performance of the scenarios above. Then, we present numerical simulations and experiments, showing that which scenario yields better performance depends on the bitrate, the properties of the image set, and the accuracy of the view interpolation.
Masato Ishii, Keita Takahashi 0001, Takeshi Naemura
ICIP2
2008 Theoretical model and optimal prefilter for view interpolation
abstract
This paper discusses a view interpolation problem where a new view would be synthesized at the midst of two parallel views. We first introduce a new theoretical model that describes the relation between accuracy of disparities and the quality of the interpolated view. Then, we derive the optimal prefilter that minimizes the power of errors on the interpolated view. Our approach is inspired by the rate-distortion theory for video coding, and can be also regarded as an attempt to reconsider the plenoptic sampling theory. Finally, we present numerical simulations and experimental results to validate our theory.
Keita Takahashi 0001, Takeshi Naemura
ICIP1
2008 Foreground segmentation with single reference frame using iterative likelihood estimation and graph-cut
abstract
This paper introduces a new foreground segmentation method. In contrast to most of the related works, our method uses only two image frames, a target frame to process, and a single reference frame. Our method first conducts simple thresholding like background subtraction, but then applies an iteration scheme we propose to estimate the pixel-wise likelihood of belonging to the foreground/background from the frame-to-frame difference. Finally, a further refinement considering edges is applied using graph-cut optimization. Experimental results show the effectiveness of our method, especially in that it keeps good performance over a wide range of the threshold value. That consistent performance will become an important step toward fully-automatic segmentation.
Keita Takahashi 0001, Taketoshi Mori
ICME1
2007 How does Subsampling of Multi-View Images Affect the Rate-Distortion Performance?
abstract
This paper introduces a new theoretical model on the rate-distortion performance in transmitting multi-view image data. To reduce the data amount, a practical solution is just decreasing the number of images by subsampling. The questions we focus on are (i) how much bit-rate can be reduced, and (ii) how much additional distortion would be caused, by the subsampling of images. In our theoretical model, the rate-distortion theory and the plenoptic sampling theory are combined to consider the relation between the sampling condition (the cameras' interval) and the compression efficiency. Numerical simulations which show the theoretical lower bounds for (i) and (ii) are presented with discussions.
Keita Takahashi 0001, Takeshi Naemura
ICIP (1)1
2006 A Theory of Aliasing Separation for Light Field Data
abstract
A light field means a 4-D function which characterizes the flow of light rays from a target scene, and used for image-based rendering. This paper presents a novel theoretical framework which considers the aliasing problem in dealing with discrete light field data. We introduce a new scheme called the aliasing separation to isolate the additional aliasing component caused by subsampling, and give a new perspective for how to optimize the reconstruction filter to interpolate light field data without aliasing artifacts. This optimization is closely related with depth estimation. Both the focus measure for light field rendering and the multiple baseline stereo can be derived from our theoretical framework.
Keita Takahashi 0001, Takeshi Naemura
ICIP1
2006 Layered light-field rendering with focus measurement
Keita Takahashi 0001, Takeshi Naemura
Signal Process. Image Commun.1
2005 Spatial domain analysis on the focus measurement for light field rendering
abstract
Light field rendering (LFR) is an image-based rendering method for synthesizing free-viewpoint images from a set of multi-view images. In LFR, no/little knowledge of geometry is required, since scene objects are assumed to be at a constant depth. This assumption causes focus-like effects on synthetic images, in which only the objects near to the assumed depth are synthesized clearly and sharply. In order to generate all-in-focus synthetic images, the authors proposed a focus measurement method which uses several kinds of reconstruction filters. Using this method, we can detect in-focus parts on multiple images that are synthesized with different assumed depths, and integrate them into one final image. Though this algorithm was shown to be effective for some real scenes, the optimal combination of reconstruction filters has not been discussed yet. In this paper, we introduce a detailed spatial domain analysis of this focus measurement method, and compare three combinations by the theoretical analysis and an experiment.
Keita Takahashi 0001, Takeshi Naemura
ICIP (3)1
2005 Unstructured light field rendering using on-the-fly focus measurement
abstract
This paper introduces a novel image-based rendering method which uses inputs from unstructured cameras and synthesizes free-viewpoint images of high quality. Our method uses a set of depth layers in order to deal with scenes with large depth ranges. To each pixel on the synthesized image, the optimal depth layer is assigned automatically based on the on-the-fly focus measurement algorithm that we propose. We implemented this method efficiently on a PC and achieved nearly interactive frame-rates.
Keita Takahashi 0001, Takeshi Naemura
ICME1
2004 A focus measure for light field rendering
abstract
Light field rendering is a fundamental method for synthesizing free-viewpoint images from a set of multiviewpoint images. In the simplest case, the scene structure is approximated by a simple plane: a focal plane. This approximation leads to focus-like effects on synthetic images where the focused depth is determined by the focal plane. A serious problem is that the range of the focused depth is too small in most practical cases. In this paper, we propose a focus measure that is specialized for synthetic images by light field rendering. When a set of differently-focused images is generated at a given viewpoint, the proposed focus measure enables us to obtain a depth map and an all in-focus image. Our approach has some remarkable differences from other related techniques, such as depth-from-stereo and depth-from-focus methods. Experimental results show that the proposed method effectively enhances PSNR of the final synthetic images.
Keita Takahashi 0001, Akira Kubota, Takeshi Naemura
ICIP1
2004 Focus Measurement on Programmable Graphics Hardware for All in-Focus Rendering from Light Fields
Kaoru Sugita, Keita Takahashi 0001, Takeshi Naemura, Hiroshi Harashima
VR2
2003 Depth of field in light field rendering
abstract
This paper focuses on the sampling problem in light field rendering (LFR) that is a fundamental approach to image based rendering. Quality of LFR depends on a light ray database generated from pre-acquired images, since image synthesis is a process of gathering appropriate light ray data from the database. For improving the quality, interpolation of light ray data is effective. It is based on an assumption that scene objects are placed on a plane called "focal plane". According to the depth of the focal plane (distance between cameras and focal plane), focus-like effect would appear on the synthesized images. In this paper, we formulate the depth of field In light field rendering to address the range of depth where scene objects are rendered in focus. In our theory, the plenoptic sampling theory is generalized to compare with some other related works. Our representation is intuitive and useful for quantitative analysis of LFR.
Keita Takahashi 0001, Takeshi Naemura, Hiroshi Harashima
ICIP (1)1