EDBT 2026 Demo / reviewers in the wild / expert
Toshiaki Fujii
dblp:18/5373
· DBLP profile ↗
70ranked-venue papers
3as first author
13since 2021 · last 2024
0000-0002-3440-5132ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 3 since 2021Computer networks · 4Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Synchronization for VLC Using Orthogonally Aligned Rolling Shutter Cameras and LED ArrayabstractThis research addresses challenges in visible light communication (VLC) systems that use two orthogonally aligned rolling shutters (RS) image sensors as receivers and an LED array as a transmitter. The study focuses on overcoming burst signal loss caused by unsensed periods between frames in RS image sensors. To improve VLC performance, we propose and compare three different schemes. The first is a conventional approach that superimposes a Barker code synchronization signal on the transmission signal using pulse width modulation (PWM). The second introduces a new method of spatial synchronization by dedicating specific LEDs in an array for synchronization signals. The third scheme extends this spatial approach by incorporating a 4 -level PWM for information transmission to increase data rates. The study aims to evaluate and compare the error rate characteristics of these three schemes, assessing their effectiveness in mitigating burst errors and enhancing overall VLC system performance. This research has potential applications, including vehicle communication, Internet of Things devices, and smartphones. Ayumu Otsuka, Takaya Yamazato, Hiraku Okada, Toshiaki Fujii, Koji Kamakura, Masayuki Kinoshita, Shintaro Arai, Tomohiro Yendo, Shan Lu 0003 |
APCC | 4 |
| 2024 | Time-Efficient Light-Field Acquisition Using Coded Aperture and EventsabstractWe propose a computational imaging method for time-efficient light-field acquisition that combines a coded aperture with an event-based camera. Differentfrom the conventional coded-aperture imaging method, our method applies a sequence of coding patterns during a single exposure for an image frame. The parallax information, which is related to the differences in coding patterns, is recorded as events. The image frame and events, all of which are measured in a single exposure, are jointly used to computationally reconstruct a light field. We also designed an algorithm pipeline for our method that is end-to-end trainable on the basis of deep optics and compatible with real camera hardware. We experimentally showed that our method can achieve more accurate reconstruction than several other imaging methods with a single exposure. We also developed a hardware prototype with the potential to complete the measurement on the camera within 22 msec and demonstrated that light fields from real 3-D scenes can be obtained with convincing visual quality. Our software and supplementary video are available from our project website1: Shuji Habuchi, Keita Takahashi 0001, Chihiro Tsutake, Toshiaki Fujii, Hajime Nagahara |
CVPR | 4 |
| 2024 | Mono+Sub: Compressing Light Field as Monocular Image and Subsidiary DataabstractA light field is usually represented as a set of multi-view images captured from a two-dimensional (2-D) array of viewpoints and requires a large amount of data compared with a standard 2-D image. We propose a 2-D compatible light-field compression method for encoding a light field as a 2-D monocular image and subsidiary data. In terms of the image quality, we prioritize the central image (regarded as the 2-D monocular image) over the other images in the light field, because the light field is considered an extension of the 2-D monocular image. To this end, we encode and decode the monocular image using a standard image codec and introduce a learned encoder and decoder pair for the subsidiary data. Experimental results indicate that our method achieved promising rate-distortion performance, especially for extremely low bit-rate ranges. Even though our method requires only a small amount of subsidiary data compared with those for the monocular image, the entire light field can be reconstructed with reasonable visual quality. Ryosuke Imazu, Chihiro Tsutake, Keita Takahashi 0001, Toshiaki Fujii |
VCIP | 4 |
| 2024 | Advancements in Lenslet Video Coding: Insights from MPEG LVCabstractBeing a general representation format of the dense light field, lenslet video, where each frame consists of a 2D grid of micro-images, shows high potential for applications in immersive media such as glasses-free 3D displays and virtual reality. However, its distinct spatial-temporal-angular distribution places a significant challenge on conventional video coding. In July 2021, the Moving Picture Experts Group (MPEG) established an Ad-Hoc group, Lenslet Video Coding (LVC), to explore use cases, efficient compression methods, testing sequences, conversion tools, and coding architectures towards a new compression standard. This paper provides an overview of recent progress in the LVC Ad Hoc group and presents the compression efficiency of state-of-the-art codec agnostic coding tools to encourage contributions in the future. Mehrdad Teratani, Byeungwoo Jeon, Toshiaki Fujii, Ruibo Zhao, Eline Soetens |
VCIP | 4 |
| 2024 | Warm-start NeRF: Accelerating Per-scene Training of NeRF-based Light-Field RepresentationabstractA light field is represented as a set of multi-view images captured from a dense 2-D array of viewpoints. To treat a light field as being continuous, we represent it as a neural radiance field (NeRF), which is a learned representation of a 3-D scene. NeRFs are renowned for their ability to reconstruct a target 3-D scene with compelling visual quality, but they are slow to train. A solution for this problem is to use a tiny neural network and trainable volumetric features as the scene representation, which is considered the baseline of our research. For further acceleration, we propose a method for warm-starting the per-scene training by setting good initial values for the trainable parameters. To this end, we introduce another encoder network to obtain the initial volumetric features from the target light field. Starting with the appropriate initial values, our method can achieve better rendering quality with fewer training iterations than the baseline. Takuto Nishio, Chihiro Tsutake, Keita Takahashi 0001, Toshiaki Fujii |
VCIP | 4 |
| 2024 | Editorial
Caroline Conti, Atanas P. Gotchev, Robert Bregovic, Donald G. Dansereau, Cristian Perra, Toshiaki Fujii |
Signal Process. Image Commun. | 6 |
| 2022 | Acquiring a Dynamic Light Field through a Single-Shot Coded ImageabstractWe propose a method for compressively acquiring a dynamic light field (a 5-D volume) through a single-shot coded image (a 2-D measurement). We designed an imaging model that synchronously applies aperture coding and pixel-wise exposure coding within a single exposure time. This coding scheme enables us to effectively embed the original information into a single observed image. The observed image is then fed to a convolutional neural network (CNN) for light-field reconstruction, which is jointly trained with the camera-side coding patterns. We also developed a hardware prototype to capture a real 3-D scene moving over time. We succeeded in acquiring a dynamic light field with 5x5 viewpoints over 4 temporal sub-frames (100 views in total)from a single observed image. Repeating capture and reconstruction processes over time, we can acquire a dynamic light field at 4x the frame rate of the camera. To our knowledge, our method is the first to achieve a finer temporal resolution than the camera itself in compressive light-field acquisition. Our software is available from our project webpage.11https://www.fujii.nuee.nagoya-u.ac.jp/Research/CompCam2 Ryoya Mizuno, Keita Takahashi 0001, Michitaka Yoshida, Chihiro Tsutake, Toshiaki Fujii, Hajime Nagahara |
CVPR | 5 |
| 2022 | Unsupervised disparity estimation from light field using plug-and-play weighted warping lossabstractWe investigated disparity estimation from a light field using a convolutional neural network (CNN). Most of the methods implemented a supervised learning framework, where the predicted disparity map was compared directly to the corresponding ground-truth disparity map in the training stage. However, light field data accompanied with ground-truth disparity maps were insufficient and rarely available for real-world scenes. The lack of training data resulted in limited generality of the methods trained with them. To tackle this problem, we took a simple Figure-and-play approach to remake a supervised method into an unsupervised (self-supervised) one. We replaced the loss function of the original method with one that does not depend on the ground-truth disparity maps. More specifically, our loss function is designed to indirectly evaluate the accuracy of the disparity map by using warping errors among the input light field views. We designed pixel-wise weights to properly evaluate the warping errors in the presence of occlusions, and an edge loss to encourage edge alignment between the image and the disparity map. As a result of this unsupervised learning framework, our method can use more abundant training datasets (even those without ground-truth disparity maps) than the original supervised method. Our method was evaluated on computer-generated scenes (4D Light Field Benchmark) and real-world scenes captured by Lytro Illum cameras. Our method achieved the state-of-the-art performance as an unsupervised method on the benchmark. We also demonstrated that our method can estimate disparity maps more accurately than the original supervised method for various real-world scenes. Taisei Iwatsuki, Keita Takahashi 0001, Toshiaki Fujii |
Signal Process. Image Commun. | 3 |
| 2022 | Denoising multi-view images by soft thresholding: A short-time DFT approachabstractShort-time discrete Fourier transform (ST-DFT) is known as a promising technique for image and video denoising. The seminal work by Saito and Komatsu hypothesized that natural video sequences can be represented by sparse ST-DFT coefficients and noisy video sequences can be denoised on the basis of statistical modeling and shrinkage of the ST-DFT coefficients. Motivated by their theory, we develop an application of ST-DFT for denoising multi-view images. We first show that multi-view images have sparse ST-DFT coefficients as well and then propose a new statistical model, which we call the multi-block Laplacian model, based on the block-wise sparsity of ST-DFT coefficients. We finally utilize this model to carry out denoising by solving a convex optimization problem, referred to as the least absolute shrinkage and selection operator. A closed-form solution can be computed by soft thresholding, and the optimal threshold value is derived by minimizing the error function in the ST-DFT domain. We demonstrate through experiments the effectiveness of our denoising method compared with several previous denoising techniques. Our method implemented in Python language is available from https://github.com/ctsutake/mviden. Keigo Tomita, Chihiro Tsutake, Keita Takahashi 0001, Toshiaki Fujii |
Signal Process. Image Commun. | 4 |
| 2021 | An Efficient Image Compression Method Based On Neural Network: An Overfitting ApproachabstractOver the past decade, nonlinear image compression techniques based on neural networks have been rapidly developed to achieve more efficient storage and transmission of images compared with conventional linear techniques. A typical nonlinear technique is implemented as a neural network trained on a vast set of images, and the latent representation of a target image is transmitted. In contrast to the previous nonlinear techniques, we propose a new image compression method in which a neural network model is trained exclusively on a single target image, rather than a set of images. Such an overfitting strategy enables us to embed fine image features in not only the latent representation but also the network parameters, which helps reduce the reconstruction error against the target image. The effectiveness of our method is validated through a comparison with conventional image compression techniques in terms of a rate-distortion criterion. Yu Mikami, Chihiro Tsutake, Keita Takahashi 0001, Toshiaki Fujii |
ICIP | 4 |
| 2021 | Factorized Modulation For Singleshot Lightfield AcquisitionabstractA light field (LF), which is represented as a set of dense multiview images, has been utilized in various 3-D applications. To make LF acquisition more efficient, researchers have investigated compressive sensing methods by incorporating modulation or coding functions into the camera. In this work, we investigate a challenging case of compressive LF acquisition in which an entire LF should be reconstructed from only a single coded image. To achieve this goal, we propose a new modulation scheme called factorized modulation that can approximate arbitrary 4-D modulation patterns in a factorized manner. Our method can be hardware-implemented by combining the architectures for coded aperture and pixel-wise coded exposure imaging. The modulation pattern is jointly optimized with a CNN-based reconstruction algorithm. Our method is validated through extensive evaluations against other modulation schemes. Kohei Tateishi, Kohei Sakai, Chihiro Tsutake, Keita Takahashi 0001, Toshiaki Fujii |
ICIP | 5 |
| 2021 | An Efficient Compression Method For Sign Information Of DCT Coefficients Via Sign RetrievalabstractCompression of the sign information of discrete cosine transform coefficients is an intractable problem in image compression schemes due to the equiprobable occurrence of the sign bits. To overcome this difficulty, we propose an efficient compression method for such sign information based on phase retrieval, which is a classical signal restoration problem attempting to find the phase information of discrete Fourier transform coefficients from their magnitudes. In our compression strategy, the sign bits of all the AC components in the cosine domain are excluded from a bitstream at the encoder and are complemented at the decoder by solving a sign recovery problem, which we call sign retrieval. The experimental results demonstrate that the proposed method outperforms previous techniques for sign compression in terms of a rate-distortion criterion. Our method implemented in Python language is available from https://github.com/ctsutake/sr. Chihiro Tsutake, Keita Takahashi 0001, Toshiaki Fujii |
ICIP | 3 |
| 2021 | Vehicle Distance Measurement based on Visible Light Communication Using Stereo CamerasabstractVisible light communication based intelligent transportation systems (ITS-VLC) show great potential for future urban mobility. This study presents a performance evaluation of range estimation between vehicles and infrastructures in an ITS-VLC system. In the proposed ITS-VLC system, it is easy to simultaneously conduct communication and ranging using stereo cameras. However, the stereo camera calibration becomes a problem during simultaneous communication and ranging due to vehicle vibration. Using the data from LED transmitters and stereo cameras, it can obtain multiple measurements of distance. The monocular-stereo fusion algorithm is applied to visible light ranging in the proposed scheme using particle swarm optimization. We employed real data from the field trial experiment and achieved a ranging accuracy of 60 ± 1.0 m. Ruiyi Huang, Takaya Yamazato, Masayuki Kinoshita, Hiraku Okada, Koji Kamakura, Shintaro Arai, Tomohiro Yendo, Toshiaki Fujii |
IV | 8 |
| 2020 | BER Measurement for Transmission Pattern Design of ITS Image Sensor Communication Using DMD ProjectorabstractThis paper presents an image sensor communication (ISC) using a digital micromirror device (DMD) projector as a transmitter. In particular, we focus on ISC for intelligent transport systems (ITSs) because DMD projectors are expected to be used in road traffic, such as vehicle headlights, street lights, and traffic signs. The DMD projector controls light patterns by switching tilt of micromirrors at high speed. Compared to the conventional LED array transmitter, DMD projector can design and alter the shape of the light pattern and data rate of transmission more easily. In the proposed system, the data rate can be easily increased by increasing the number of multiplexes. However, as the number of multiplexes is increased, the number of received pixels on the image sensor is decreased, and thus the performance of symbol detection deteriorates. Therefore, the purpose of this paper is to examine the relation between the number of received pixels per cell and communication performance for transmission pattern design. Hence, we experimentally clarify this relation using our prototype system via a DMD projector and a high-speed camera. Tomoya Arisue, Takaya Yamazato, Hiraku Okada, Toshiaki Fujii, Masayuki Kinoshita, Koji Kamakura, Shintaro Arai, Tomohiro Yendo |
CCNC | 4 |
| 2020 | Acquiring Dynamic Light Fields Through Coded Aperture Camera
Kohei Sakai, Keita Takahashi 0001, Toshiaki Fujii, Hajime Nagahara |
ECCV (19) | 3 |
| 2019 | Stereo Ranging Method Using LED Transmitter for Visible Light CommunicationabstractThis paper focuses on high-speed stereo camera receivers for visible light communication (VLC) in intelligent transport systems (ITSs), contrary to conventional studies where a single (high-speed) camera is used. An advantage of high-speed stereo cameras as VLC receivers is that both front-facing ranging and wireless communication can be achieved in one device. In this paper, utilizing features of VLC, we consider stereo ranging with high accuracy. Moreover, it is known that an LED transmitter for VLC has specific features, high-luminance and high- speed blinking. This paper proposes stereo ranging by applying LED transmitter detection (utilizing these features) to disparity estimation. Additionally, we introduce subpixel estimation using equiangular line fitting to obtain more accurate disparity. Finally, we experimentally demonstrate that the proposed stereo ranging achieves an accuracy of 60 m ± 0.1 m. Masayuki Kinoshita, Koji Kamakura, Takaya Yamazato, Hiraku Okada, Toshiaki Fujii, Shintaro Arai, Tomohiro Yendo |
GLOBECOM | 5 |
| 2019 | A 3-D Display Pipeline from Coded-Aperture Camera to Tensor Light-Field Display Through CNNabstractWe propose an efficient pipeline from input to output for a tensor light-field display. Conventionally, a dense light field (i.e., tens of images taken with narrow viewpoint intervals) is required as an input in such displays. However, obtaining dense light fields is a challenging task for real scenes. To make the acquisition process more efficient, we adopted a coded-aperture camera as an input device, which is suitable for acquiring dense light fields in a compressive manner. Moreover, we modeled the entire process from acquisition to display using a convolutional neural network. As a result of training the network on a massive light field data, we can reproduce the whole light field on the display from only a few images taken with the camera. Both simulative and real experiments were conducted to show the effectiveness of our method. Keita Maruyama, Yasutaka Inagaki, Keita Takahashi 0001, Toshiaki Fujii, Hajime Nagahara |
ICIP | 4 |
| 2019 | LF-TSP: Traveling salesman problem for HEVC-based light-field codingabstractWe studied a coding scheme where light field (LF) images (dense multi-view images) are regarded as a sequence of temporal video frames and encoded with video codecs such as High Efficiency Video Coding (HEVC). An important issue with this scheme is how to determine the frame order of the LF images. We propose a method to find the optimum frame order through a formulation of the traveling salesman problem (TSP). Under the assumption that video codecs are more effective with temporally smooth videos, our method, named LF-TSP, defines frame-to-frame distances for each image pairs in an LF, and attempted to find the shortest route that visits all frames. Experiments showed that our method achieved an overall better rate-distortion performance than several previous methods. Kota Imaeda, Kohei Isechi, Keita Takahashi 0001, Toshiaki Fujii, Yukihiro Bandoh, Takehito Miyazawa, Seishi Takamura, Atsushi Shimizu |
VCIP | 4 |
| 2018 | Learning to Capture Light Fields Through a Coded Aperture Camera
Yasutaka Inagaki, Yuto Kobayashi, Keita Takahashi 0001, Toshiaki Fujii, Hajime Nagahara |
ECCV (7) | 4 |
| 2018 | A Comparison of Reception Methods for Visible Light Communication Using High-Speed Stereo CamerasabstractIn-vehicle cameras can be used as receivers for visible light communication in addition to their primary functions: for viewing assistance and object recognition. Most studies with this design perform tests via a single high-speed camera that is used as the receiver, but this study employs a receiver that uses a high- speed stereo camera. The reception method developed in this paper effectively utilizes left and right stereo images to improve communication performance. Although the selective diversity reception method proposed in previous work improved communication performance, it cannot deal with cases wherein both the left and right cameras introduce errors. To further improve the communication performance of this system, maximal ratio combining (MRC) diversity is applied to the high-speed stereo camera in this study. Communication performance is then evaluated comparing to the selective diversity reception and conventional single-camera reception by experimental tests conducted under static and driving conditions. Masayuki Kinoshita, Takaya Yamazato, Hiraku Okada, Toshiaki Fujii, Shintaro Arai, Tomohiro Yendo, Koji Kamakura |
GLOBECOM | 4 |
| 2018 | How Should we Handle 4D Light Fields with CNNS?abstractWe investigated how we should handle high dimensional light fields (LFs) with convolutional neural networks (CNNs). An LF is a 4-D signal representation of light rays, and it is interpreted as a set of dense multi-view images. As an important building block of various light field applications, we focused on signal restoration problems for LFs, and we adopted CNN s as the solver for them because of its striking performance on the conventional 2-D images. In applying CNN s, the high dimensionality of LFs should be carefully addressed. Instead of treating the full 4- D signal as it is, we followed a divide and conquer strategy. Specifically, we cascade two or three CNNs, each of which works only on 2-D subspace of the full 4-D LFs. Combining CNN s that work on different subspaces, we can eventually handle the full 4- D structure. Moreover, considering different properties of those subspaces, we experimentally explored the best combination in which different subspaces are cascaded. Although our experiments are currently limited to a denoising problem, the lessons found from our results will benefit the prospective research on the full4-D LF processing. Shu Fujita, Keita Takahashi 0001, Toshiaki Fujii |
ICIP | 3 |
| 2018 | Fast and Robust Disparity Estimation for Noisy Light FieldsabstractDepth (disparity) estimation from a light field (a set of dense multiview images) has attracted much research interest recently. This paper is focused on how to handle noisy light field for disparity estimation' because if left as it is the noise deteriorates the accuracy of estimated disparity maps. Several researchers have worked on this problem, e.g. by introducing disparity cues that are robust to noise. However, it is not easy to break the trade-off between the accuracy and computational speed. To tackle this trade-off, we have integrated a fast denoising scheme in a fast disparity estimation framework that works in the epipolar plane image (EPI) domain. Specifically, we found that a simple 1-D slanted filter is very effective for reducing noise while preserving the underlying structure in an EPI. Experimental results show that our method can achieve good accuracy with much less computational time compared to some state-of-the-art methods. Gou Houben, Shu Fujita, Keita Takahashi 0001, Toshiaki Fujii |
ICIP | 4 |
| 2018 | Scalable Light Field Coding Using Weighted Binary ImagesabstractWe propose an efficient coding scheme for a dense light field, i.e., a set of multi-viewpoint images taken with very small viewpoint intervals. The key idea behind our proposal is that a light field is represented only using weighted binary images, where several binary images and corresponding weight values are to be chosen to optimally approximate the light field. The coding scheme derived from this idea is completely different from those of modern image/video coding standards. However, we found that our scheme can achieve comparable coding efficiency (rate-distortion performance) to that of modern highly-sophisticated video codecs. Moreover, the decoding process of our scheme is extremely simple, which will lead to a faster and less power-hungry decoder than those of the modern codecs. Furthermore, our scheme can be made scalable, where the accuracy of the decoded light field is improved in a progressive manner as we use more encoded information. Thanks to the divide-and-conquer strategy adopted for the scalable coding, we can also drastically reduce the computational complexity of the encoding process. Koji Komatsu, Keita Takahashi 0001, Toshiaki Fujii |
ICIP | 3 |
| 2018 | From Focal Stack to Tensor Light-Field DisplayabstractWe propose a method of using a focal stack, i.e., a set of differently focused images, as the input for a novel light field display called a "tensor display." Although this display consists of only a few light attenuating layers located in front of a backlight, it can be viewed from many directions (angles) simultaneously without the resolution of each viewing direction being sacrificed. Conventionally, a transmittance pattern is calculated for each layer from a light field, namely, a set of dense multi-view images (typically dozens) that are to be observed from different directions. However, preparing such a massive amount of images is often cumbersome for real objects. We developed a method that does not require a complete light field as the input; instead, a focal stack composed of only a few differently focused images is directly transformed into layer patterns. Our method greatly reduces the cost of acquiring data while also maintaining the quality of the output light field. We validated the method with experiments using synthetic light field datasets and a focal stack acquired by an ordinary camera. Keita Takahashi 0001, Yuto Kobayashi, Toshiaki Fujii |
IEEE Trans. Image Process. | 3 |
| 2017 | From focal stacks to tensor display: A method for light field visualization without multi-view imagesabstractA new type of light field display called a tensor display was investigated. Although this display consists of only a few light attenuating layers located in front of a backlight, many views can be emitted in different directions simultaneously without sacrificing the resolution of each view. The transmittance pattern of each layer is calculated from a light field, namely, a set of dense multi-view images (typically dozens) that are to be observed from different directions. However, preparing such images is often cumbersome for real objects. We propose a method that does not require multi-view images as the input; instead, a focal stack composed of only a few differently focused images is directly transformed into the layer patterns. Our method greatly reduces the data acquisition cost while also maintaining the quality of the output light field. We validated the method with experiments using synthetic light field datasets and a focal stack acquired by an ordinary camera. Yuto Kobayashi, Keita Takahashi 0001, Toshiaki Fujii |
ICASSP | 3 |
| 2017 | Good group sparsity prior for light field interpolationabstractA light field, which is equivalent to a dense set of multi-view images, has various applications such as depth estimation and 3D display. One of the essential problems is light field interpolation, which is obtaining sufficiently dense views from sparser views. The accuracy of interpolation will be enhance by exploiting an inherent property of a light field. Specifically, an epipolar plane image (EPI), which is a 2D subset of the 4D light field, consists of many lines. This structure induces a sparse representation in the frequency domain, where most of the energy resides on a line passing through the origin. On the basis of this observation, we propose a group sparsity prior suitable for light fields to fully exploit their line structure for interpolation. Our experimental results show that the proposed method can achieve better quality than a state-of-the-art shearlet-based method. Keita Takahashi 0001, Shu Fujita, Toshiaki Fujii |
ICIP | 3 |
| 2017 | PCA-coded aperture for light field photographyabstractA light field, which is often understood as a set of dense multi-view images, has been utilized in various 2D/3D applications. Efficient light field acquisition using a coded aperture camera is the target problem considered in this paper. Specifically, the entire light field, which consists of many images, should be reconstructed from only a few images that are captured through different aperture patterns. In previous work, this problem has often been discussed from the context of compressed sensing (CS). In contrast, we formulated this problem from the perspective of principal component analysis (PCA) to derive optimal non-negative aperture patterns and a straight-forward reconstruction algorithm. Even though it is based on a conventional technique, our method has proven to be more accurate and much faster than a state-of-the-art CS-based method. Yusuke Yagi, Keita Takahashi 0001, Toshiaki Fujii, Toshiki Sonoda, Hajime Nagahara |
ICIP | 3 |
| 2016 | Disparity estimation from light fields using sheared EPI analysisabstractStructure tensor analysis on epipolar plane images (EPIs) is a successful approach to estimate disparity from a light field, i.e. a dense set of multi-view images. However, the disparity range allowable for the light field is limited, because the estimation becomes less accurate as the range of disparities become larger. To overcome this limitation, we propose a new method called sheared EPI analysis, where EPIs are sheared before the structure tensor analysis. The results of analysis obtained with different shear values are integrated into a final disparity map. As verified by extensive evaluations on 12 datasets with large disparity ranges, our method is comparably accurate to and much faster than a multi-view stereo method. Keita Takahashi 0001, Toshiaki Fujii |
ICIP | 3 |
| 2015 | Joint directional-positional multiplexing for light field acquisition by Kronecker compressed sensingabstractIn this paper, we propose a joint and unified framework to compressively capture a light field in the consideration of both directional and positional multiplexing based on Kronecker compressed sensing (KCS). First of all, both of the 2D angular and 2D spatial correlations of the light field can be fully utilized in the compressive acquisition, and the multiplexing is more flexible and balanced during the acquisition. Secondly, other types of light field acquisition can be unified into our proposed framework. In the experiment, it is shown that more balanced allocation between directional and positional multiplexing achieves better reconstruction quality of light field given the same number of total acquisitions. Furthermore, the experimental result also illustrates that the proposed method can capture a light field with full resolution and achieve better reconstruction quality than other previous methods. Keita Takahashi 0001, Mehrdad Panahpour Tehrani, Toshiaki Fujii |
ICASSP | 4 |
| 2015 | Reconstruction of compressively sampled light fields using a weighted 4D-DCT basisabstractThe coded aperture/mask technique enables us to capture light field data in a compressive way through a single camera. A pixel value recorded by such a camera is a summation of the light rays that pass though different positions on the coded aperture/mask. The target light field can be reconstructed from the recorded pixel values by using prior information of the light field signal. As prior information, a dictionary (light field atoms), which was learned from training datasets, was used in the current state of the art. Meanwhile, it was reported that general bases such as DCT were not suitable to efficiently represent prior information. In this work, however, we demonstrate that a 4D-DCT basis works surprisingly better if it is combined with a weighting scheme in which the amplitude difference in DCT coefficients is considered. Simulation results using 18 light field datasets are reported to show the superior performance of the weighted 4D-DCT basis to the learned dictionary. Yusuke Miyagi, Keita Takahashi 0001, Mehrdad Panahpour Tehrani, Toshiaki Fujii |
ICIP | 4 |
| 2015 | Super-resolution image synthesis using the physical pixel arrangementofalight field cameraabstractWe propose a method for super-resolution image synthesis that accurately handles the physical pixel arrangement of a light field (plenoptic) camera. We use a Lytro camera to obtain 4D light field data (a set of multi-viewpoint images) through a micro-lens array. The light field data are multiplexed on a single image sensor, and thus, the data is first de-multiplexed into a set of multi-viewpoint (sub-aperture) images. However, the de-multiplexing process usually involves interpolation of the original data such as demosaicing for a color filter array and pixel resampling for the non-square micro-lens arrangement. During this interpolation, some information is added or lost to/from the original data. In contrast, our method can preserve the originally captured data as they are, and directly use them for the super-resolution image synthesis, where the super-resolved image and the corresponding depth map are alternatively refined. We experimentally demonstrate that our method can achieve higher image quality than that with a standard Light Field Toolbox. Kazuki Ohashi, Keita Takahashi 0001, Mehrdad Panahpour Tehrani, Toshiaki Fujii |
ICIP | 4 |
| 2015 | Rank analysis of a light field for dual-layer 3D displaysabstractIn this paper, a new type of 3D display, called a layered light-field display, was investigated. By using only a few light-attenuating layers located in front of a backlight, this display can present many views in different directions simultaneously without sacrificing the resolutions of each view. The essential factor for efficient layer-based representation, which has not been deeply analyzed in previous works, is redundancy- namely, a low rank structure-of the light-field data. Accordingly, to reveal the origin of the redundancy, a generative model, in which a textured surface located at a certain depth generates a light field, was formulated and evaluated in this paper. Our theoretical analysis shows that the redundancy depends on not only the texture complexity but also the depth of the surface from the light attenuating layers. The theoretical model was validated through experimental simulation of the display. Keita Takahashi 0001, Toyohiro Saito, Mehrdad Panahpour Tehrani, Toshiaki Fujii |
ICIP | 4 |
| 2015 | Data format and view synthesis for free-viewpoint video streaming of super multiview videoabstractOur goal is to propose and evaluate a new streaming data format that can be adapted to the limited bandwidth and capable of free-viewpoint video streaming using super multi-view video plus depth (MVD). Additionally, we proposed a view synthesis method for our data format. Given a requested free-viewpoint, we use the two closest views and corresponding depth maps to perform free-viewpoint video synthesis. The new data format consists of all views and corresponding depth maps in a lowered resolution, and the two closest views to the requested viewpoint in the high resolution. When the requested viewpoint changes, the two closest viewpoints will change, but one or both views are transmitted only in the low resolution during periods of large round-trip delay time. Therefore, the resolution compensation is required before view synthesis. Experimental results show that our proposed framework achieves view synthesis quality close to view synthesis using high resolution multi-view video plus depth. Takaaki Emori, Mehrdad Panahpour Tehrani, Keita Takahashi 0001, Toshiaki Fujii |
PCS | 4 |
| 2015 | View synthesis using superpixel based inpainting capable of occlusion handling and hole fillingabstractThe existing virtual view synthesis methods generate the images with many artifacts that are annoying, especially for forward virtual viewpoint, and virtual viewpoint generated by reference views with large baseline, due to occlusions and the limited sampling density. In this paper, we propose a new view synthesis method, robust to the above-mentioned problem, consist of three steps, using stereo contents. Firstly, view plus depth data of each viewpoint is 3D warped to the virtual viewpoint. We determine which neighboring pixels should be connected or kept isolated. Polygons enclosed by the connected pixels, i.e. superpixel, are interpolated. Secondly, we blend those warped images by comparing each pixel's depth value to obtain the virtual view, in which non-occlusion holes have already been interpolated by the process in the first step. Thirdly, the remaining holes are filled by inpainting. Our experimental results and comparisons show that the proposed view synthesis method allows smoother view reconstruction, while holes due to occlusion and 3D warping are filled with less artifacts. Tomoyuki Tezuka, Mehrdad Panahpour Tehrani, Kazuyoshi Suzuki, Keita Takahashi 0001, Toshiaki Fujii |
PCS | 5 |
| 2015 | Vehicle Motion and Pixel Illumination Modeling for Image Sensor Based Visible Light CommunicationabstractChannel modeling is critical for the design and performance evaluation of visible light communication (VLC). Although a considerable amount of research has focused on indoor VLC systems using single-element photodiodes, there remains a need for channel modeling of VLC systems for outdoor mobile environments. In this paper, we describe and provide results for modeling image sensor based VLC for automotive applications. In particular, we examine the channel model for mobile movements in the image plane as well as channel decay according to the distance between the transmitter and the receiver. Optical flow measurements were conducted for three VLC situations for automotive use: infrastructure to vehicle VLC (I2V-VLC); vehicle to infrastructure VLC (V2I-VLC); and vehicle to vehicle VLC (V2V-VLC). We describe vehicle motion by optical flow with subpixel accuracy using phase-only correlation (POC) analysis and show that a single-pinhole camera model successfully describes these three VLC cases. In addition, the luminance of the central pixel from the projected LED area versus the distance between the LED and the camera was measured. Our key findings are twofold. First, a single-pinhole camera model can be applied to vehicle motion modeling of a I2V-VLC, V2I-VLC, and V2V-VLC. Second, the DC gain at a pixel remains constant as long as the projected image of the transmitter LED occupies several pixels. In other words, if we choose a pixel with highest luminance among the projected image of transmitter LED, the value remains constant, and the signal-to-noise ratio does not change according to the distance. Takaya Yamazato, Masayuki Kinoshita, Shintaro Arai, Eisho Souke, Tomohiro Yendo, Toshiaki Fujii, Koji Kamakura, Hiraku Okada |
IEEE J. Sel. Areas Commun. | 6 |
| 2014 | Least MSE Regression for View SynthesisabstractView synthesis is the process of combining given multi-view images to generate an image from a new viewpoint. Assuming that each pixel of the new view is obtained as the weighted sum of the corresponding pixels from the input views, we focus on the problem of how to optimize the weight for each of the input views. Our weighting method is called least mean squared error (MSE) regression because it is formulated as a regression problem in which second order statistics among the viewpoints are exploited to minimize the MSE of the resulting image. More specifically, the affinity across the viewpoints is represented as a covariance and approximated using a linear model whose parameters are adapted for each dataset. By using the approximated covariance, the optimal weights can be successfully estimated. As a result, the weights derived using our method are data dependent and significantly differ from those obtained using current empirical methods such as distance penalty. Our method is still effective if the given correspondence is not completely accurate due to noise. We report on experimental results using several multi-view datasets to validate our theory and method. Keita Takahashi 0001, Toshiaki Fujii |
3DV | 2 |
| 2014 | Multiple LED arrays acquisition for image-sensor-based I2V-VLC using block matchingabstractThe present paper proposes a novel multiple-LED-arrays acquisition for an infrastructure-to-vehicle visible light communication (I2V-VLC) using LED arrays (transmitter) and an in-vehicle high-speed image sensor (receiver). In order to achieve a robust detection of LED arrays, we employ the block matching algorithm, which is a way of finding a corresponding position between two successive frames. The proposed method divides a captured image into a number of small domains (blocks) and determines if the LED array is present or absent using the block matching. We perform I2V-VLC experiments with multiple-LED arrays and evaluate the acquisition capability of the proposed method. Shintaro Arai, Yasutaka Shiraki, Takaya Yamazato, Hiraku Okada, Toshiaki Fujii, Tomohiro Yendo |
CCNC | 5 |
| 2014 | Synthesis Error COmpeNsateD Multiview Video plus Depth for representation of multiview videoabstractSECOND-MVD (Synthesis Error COmpeNsateD Multiview Video plus Depth) is an alternative 3D format that we introduce for representation of multiview video. In this data format, images at some viewpoint remain original, and the others are converted to a novel format. Residual based representation, such as layered depth video and free-viewpoint TV data unit were proposed. We propose hybrid image that not only consists of residual but also remainder pixels. Generation and reconstruction process of a hybrid image uses virtual image synthesized by the images that remained original in SECOND-MVD. In this paper, we investigate the compression performance of SECOND-MVD using hybrid image. Experiments demonstrate reduction in bit rate using hybrid image against residual image in SECOND-MVD framework. Mehrdad Panahpour Tehrani, Akio Ishikawa, Makoto Okui, Naomi Inoue, Keita Takahashi 0001, Toshiaki Fujii |
ICASSP | 6 |
| 2012 | Multi-view video contents viewing system by synchronized multi-view streaming architectureabstractWe developed a novel networked video viewing system for multi-view video contents with video streaming technology. This work's contribution is that our developed system confirmed the validity of simultaneous multiple video streaming architecture that incorporates real-time channel switching and a target-oriented viewing interface. In addition, we newly introduced an inter-channel bandwidth management scheme to achieve cost-effective multi-view streaming. Takafumi Marutani, Kenji Mase, Toshiaki Fujii, Tetsuya Kawamoto |
ACM Multimedia | 3 |
| 2012 | FTV format using global view and depth mapabstractWe propose a novel 3D space representation for multi-view video, using epipolar plane depth images (EPDI). Multi-view video plus depth (MVD) is used as common data format for FTV (Free-viewpoint TV), which enables synthesizing virtual view images. Due to the large amount of data and complexity of the multi-view video coding (MVC), compression of MVD is a challenging issue. We address this problem and propose a new representation that is constructed from MVD using ray-space. MVD is converted into image and depth ray-spaces. The proposed representation is obtained by converting each of ray-spaces into a global depth map and a global view using EPDI. Experiments demonstrate the analysis of this representation. Takashi Ishibashi, Mehrdad Panahpour Tehrani, Toshiaki Fujii, Masayuki Tanimoto |
PCS | 3 |
| 2012 | FTV for 3-D Spatial CommunicationabstractFree-viewpoint TV (FTV) is cutting the frontier of audiovisual communications. FTV is an innovative media that enables us to view 3-D space by freely changing our viewpoints. It also allows us to listen at any listening point in the 3-D space. Since FTV transmits all audiovisual information of the 3-D space, it can reconstruct an audiovisual replica of the 3-D space anywhere and anytime over distance and time. For video, FTV captures a part of rays in 3-D space by using many cameras, and the other rays that are not captured are obtained by interpolating the captured rays. We constructed real-time FTV systems including the complete chain of operation from image capture to display. We also carried out FTV on a laptop computer and a mobile player. For audio, two kinds of free listening-point systems are demonstrated. MPEG regarded FTV as the most challenging 3-D media and has been conducting its international standardization activities. The first phase of FTV was multiview video coding (MVC) and the second phase of FTV is 3-D video (3DV). MVC enables the efficient coding of multiple camera views and was completed in 2009. MVC has been adopted by Blu-ray 3-D. 3DV is a standard that targets serving a variety of 3-D displays and its call for proposals was issued in March 2011. Masayuki Tanimoto, Mehrdad Panahpour Tehrani, Toshiaki Fujii, Tomohiro Yendo |
Proc. IEEE | 3 |
| 2011 | Erasure coding for road-to-vehicle visible light communication systemsabstractIn this paper, we focus on a road-to-vehicle visible light communication (VLC) system using LED traffic lights. In this system, an LED traffic light consists of a 2-dimensional LED array (2D LED array), and cars are equipped with high-speed 2-dimensional cameras (2D image sensors). An important issue of this system is frame loss. Sometimes, 2D image sensor in a car fails to get a frame. So as to mitigate the influence of frame loss, we propose to apply erasure coding to the road-to-vehicle VLC system. In the proposed system, a data sequence is encoded by LDPC code whose length is much longer than the size of 2D LED array. In addition, code synchronization is required for the proposed system. We also propose a code synchronization scheme, which makes use of error detection capability of LDPC code. We evaluate the performance of our proposed system, and show that it can recover frame loss for high SNR when less than 8 frames among total 18 frames are lost. In addition, we show that our proposed code synchronization scheme can ignore its errors. Hiraku Okada, Takuya Ishizaki, Takaya Yamazato, Tomohiro Yendo, Toshiaki Fujii |
CCNC | 5 |
| 2011 | All-around ray-reproducing 3DTVabstractWe have developed a 3DTV system that captures and shows 3D images that covers 360-degree of viewing zone horizontally. 3D images are captured and reproduced as gathering of dense light rays. This system consists of a cylinder-shaped 3D display that allows viewers to see 3D images from 360-degree, a ray capturing unit, a realtime image correction unit, and data transferring system. The capturing unit acquires multiview images from all horizontal directions around an object with narrow view interval, which the display needs as light ray data. The image correction unit can corrects rotation and distortion that caused by optics of ray capturing unit in real time. Tomohiro Yendo, Toshiaki Fujii, Mehrdad Panahpour Tehrani, Masayuki Tanimoto |
ICME | 2 |
| 2011 | Traffic sign detection in dual-focal active camera systemabstractTraffic sign recognition systems can be applied to assist drivers and improve the traffic safety. High resolution image of traffic sign can improve the recognition result, especially for some complex traffic signs. In this paper, a dual-focal active camera system is proposed to obtain a high resolution image of traffic sign. In the system an active telephoto camera is equipped as an assistant of a wide angle camera. To make the proposed system correctly capture a high resolution image of traffic sign, the shape, color features and the relationship between continuous frames are used together in the traffic sign detection. The experiment results demonstrate that the proposed method makes the dual-focal active camera system effectively work when the system is installed on an automobile and moves on the road. Yanlei Gu, Tomohiro Yendo, Mehrdad Panahpour Tehrani, Toshiaki Fujii, Masayuki Tanimoto |
Intelligent Vehicles Symposium | 4 |
| 2010 | Probabilistic reliability based view synthesis for FTVabstractView synthesis using depth maps is an important application in 3D image processing. In this paper, a novel method is proposed for the plausible view synthesis of Free-viewpoint TV (FTV), using two input images and their depth maps. The depth estimation based on stereo matching is known to be error-prone, leading to noticeable artifacts in the synthesized new views. To produce high-quality view synthesis, we introduce a probabilistic framework which constrains the reliability of each pixel of new view by Maximizing Likelihood (ML). The spatial adaptive reliability is provided by incorporating Gamma hyper-prior and the synthesis error approximation. Furthermore, we generate the virtual view by solving a Maximum a Posterior (MAP) problem using graph cuts. We compare the proposed method with other depth based view synthesis approaches on MPEG test sequences. The results show the outperformance of our method both at subjective artifacts reduction and objective PSNR improvement. Tomohiro Yendo, Mehrdad Panahpour Tehrani, Toshiaki Fujii, Masayuki Tanimoto |
ICIP | 4 |
| 2010 | A new vision system for traffic sign recognitionabstractIn this paper, we introduce a new vision system for traffic sign recognition. The new vision system is a hybrid camera system, including a wide angle camera, a narrow angle camera and control part. The traffic sign candidates are detected in the image of the wide angle camera, based on shape and local feature. The control part of the vision system is adjusted to the correct positions, which makes the narrow angle camera of this system to get a high resolution image of each detected candidate. This high resolution image is used for traffic sign classification. In the new vision system, the traffic sign detection and classification are processed in different resolution images separately. This advantage of the system makes traffic sign recognition at long distance possible. Yanlei Gu, Tomohiro Yendo, Mehrdad Panahpour Tehrani, Toshiaki Fujii, Masayuki Tanimoto |
Intelligent Vehicles Symposium | 4 |
| 2010 | High-speed-camera image processing based LED traffic light detection for road-to-vehicle visible light communicationabstractAs one of ITS technique, a new visible light road-to-vehicle communication system at intersections is proposed. In this system, the communication between a vehicle and an LED traffic light is conducted using an LED traffic light as a transmitter, and an on-vehicle high-speed camera as a receiver. The LEDs in the transmitter emit light in high frequency and those emitting LEDs are captured by the high-speed camera for making communication. Here, the luminance value of LEDs in the transmitter should be captured in consecutive frames to achieve effective communication. For this purpose, first the transmitter should be found, then it should be tracked in consecutive frames by processing the images from the high-speed camera. In this paper, we propose new effective algorithms for finding and tracking the transmitter, which result in a increased communication speed, compared to the previous methods. Experiments using appropriate images showed the effectiveness of the proposals. Chinthaka Premachandra, Tomohiro Yendo, Mehrdad Panahpour Tehrani, Takaya Yamazato, Hiraku Okada, Toshiaki Fujii, Masayuki Tanimoto |
Intelligent Vehicles Symposium | 6 |
| 2010 | Free-viewpoint image generation using different focal length camera arrayabstractThe availability of multi-view images of a scene makes new and exciting applications possible, including Free-Viewpoint TV (FTV). FTV allows us to change viewpoint freely in a 3D world, where the virtual viewpoint images are synthesized by Image-Based Rendering (IBR). In this paper, we introduce a FTV depth estimation method for forward virtual viewpoints. Moreover, we introduce a view generation method by using a zoom camera in our camera setup to improve virtual viewpoint-ts' image quality. Simulation results confirm reduced error during depth estimation using our proposed method in comparison with conventional stereo matching scheme. We have demonstrated the improvement in image resolution of virtually moved forward camera using a zoom camera setup. Kengo Ando, Norishige Fukushima, Tomohiro Yendo, Mehrdad Panahpour Tehrani, Toshiaki Fujii, Masayuki Tanimoto |
PCS | 5 |
| 2010 | 3D space representation using epipolar plane depth imageabstractWe propose a novel 3D space representation for multi-view video, using epipolar plane depth images (EPDI). Multi-view video plus depth (MVD) is used as common data format for FTV(Free-viewpoint TV), which enables synthesizing virtual view images. Due to large amount of data and complexity of the multi-view video coding (MVC), compression of MVD is a challenging issue. We address this problem and propose a new representation that is constructed from MVD using rayspace. MVD is converted into image and depth ray-spaces. The proposed representation is obtained by converting each of ray-spaces into a global depth map and a texture map using EPDI. Experiments demonstrate the analysis of this representation, and its efficiency. Takashi Ishibashi, Tomohiro Yendo, Mehrdad Panahpour Tehrani, Toshiaki Fujii, Masayuki Tanimoto |
PCS | 4 |
| 2010 | Parallel processing method for realtime FTVabstractIn this paper, we propose a parallel processing method to generate free viewpoint image in realtime. It is impossible to arrange the cameras in a high density realistically though it is necessary to capture images of the scene from innumerable cameras to express the free viewpoint image. Therefore, it is necessary to interpolate the image of arbitrary viewpoint from limited captured images. However, this process has the relation of the trade-off between the image quality and the computing time. In proposed method, it aimed to generate the high-quality free viewpoint image in realtime by applying the parallel processing to time-consuming interpolation part. Kazuma Suzuki, Norishige Fukushima, Tomohiro Yendo, Mehrdad Panahpour Tehrani, Toshiaki Fujii, Masayuki Tanimoto |
PCS | 5 |
| 2010 | Color based depth up-sampling for depth compressionabstract3D scene information can be represented in several ways. In applications based on a (N-)view plus (N-)depth representation, both view and depth data is compressed. In this paper we present a depth compression method containing an depth up-sample filter which uses the color view as prior. Our method of depth down-/up-sampling is able to maintain clear object boundaries in the reconstructed depth maps. Our experimental results show that the proposed depth re-sampling filter, used in combination with a standard state-of-the art video encoder, can increase both the coding efficiency and rendering quality. Meindert Onno Wildeboer, Tomohiro Yendo, Mehrdad Panahpour Tehrani, Toshiaki Fujii, Masayuki Tanimoto |
PCS | 4 |
| 2010 | Reducing bitrates of compressed video with enhanced view synthesis for FTVabstractView synthesis using depth maps is a well-known technique for exploiting the redundancy between multi-view videos. In this paper, we deal with the bitrates of view synthesis at the decoder side of FTV that would use compressed depth maps and views. Both inherent depth estimation error and coding distortion would degrade synthesis quality. The focus is to reduce bitrates required for generating the high-quality virtual view. We employ a reliable view synthesis method which is compared with standard MPEG view synthesis software. The experimental results show that the bitrates required for synthesizing high-quality virtual view could be reduced by utilizing our enhanced view synthesis technique to improve the PSNR at medium bitrates. Meindert Onno Wildeboer, Tomohiro Yendo, Mehrdad Panahpour Tehrani, Toshiaki Fujii, Masayuki Tanimoto |
PCS | 5 |
| 2010 | A semi-automatic multi-view depth estimation methodabstractIn this paper, we propose a semi-automatic depth estimation algorithm whereby the user defines object depth boundaries and disparity initialization. Automatic depth estimation methods generally have difficulty to obtain good depth results around object edges and in areas with low texture. The goal of our method is to improve the depth in these areas and reduce view synthesis artifacts in Depth Image Based Rendering. Good view synthesis quality is very important in applications such as 3DTV and Free-viewpoint Television (FTV). In our proposed method, initial disparity values for smooth areas can be input through a so-called manual disparity map, and depth boundaries are defined by a manually created edge map which can be supplied for one or multiple frames. For evaluation we used MPEG multi-view videos and we demonstrate our algorithm can significantly improve the depth maps and reduce view synthesis artifacts. Meindert Onno Wildeboer, Norishige Fukushima, Tomohiro Yendo, Mehrdad Panahpour Tehrani, Toshiaki Fujii, Masayuki Tanimoto |
VCIP | 5 |
| 2010 | Improved Decoding Methods of Visible Light Communication System for ITS Using LED Array and High-Speed CameraabstractIn this paper, we consider visible light communication systems using LED array as a transmitter and high-speed camera as a receiver for Intelligent Transport System (ITS). Previously, we have proposed the hierarchical coding scheme which allocates data to spatial frequency components of the image depending on the priority. This scheme is possible to receive information of the high-priority even if communication distance is long. However, we need to distinguish multi-valued data from the received image by using a hierarchical coding. In this paper, we propose two improved decoding methods, and demonstrate to distinguish multi-valued data more correctly in the experiment. Toru Nagura, Takaya Yamazato, Masaaki Katayama, Tomohiro Yendo, Toshiaki Fujii, Hiraku Okada |
VTC Spring | 5 |
| 2010 | Artifact reduction using reliability reasoning for image generation of FTV
Tomohiro Yendo, Mehrdad Panahpour Tehrani, Toshiaki Fujii, Masayuki Tanimoto |
J. Vis. Commun. Image Represent. | 4 |
| 2010 | The Seelinder: Cylindrical 3D display viewable from 360 degrees
Tomohiro Yendo, Toshiaki Fujii, Masayuki Tanimoto, Mehrdad Panahpour Tehrani |
J. Vis. Commun. Image Represent. | 2 |
| 2009 | View generation with 3D warping using depth information for FTV
Yuji Mori, Norishige Fukushima, Tomohiro Yendo, Toshiaki Fujii, Masayuki Tanimoto |
Signal Process. Image Commun. | 4 |
| 2008 | 3DAV integrated system featuring arbitrary listening-point and viewpoint generationabstractIn this paper, we propose two novel methods for arbitrary listening-point generation for 3D audio-video (3DAV) integration in a large-scale multipoint cameras and microphones system with abilities to process, and display information of any recorded 3D scene in realtime. With this system, users are able to control their own viewpoint/listening-point position, freely. Arbitrary listening-point can be generated by either (i) ray-space representation of sound wave field (i.e. source sound independent) for multi frequency layers, or (ii) acoustic transfer function estimation (i.e. source sound dependent) and blind separation of sources of sounds. Arbitrary viewpoint generation is based on ray-space method, which is enhanced by using multipass dynamic programming for geometry compensation. Integration is done by either (i) ray-space representation of sound wave and image together, or (ii) integrating each camera video signal and acoustic transfer function of the same location as integrated 3DAV data. The prototype system of integrated audio-visual viewer achieves both good image and sound qualities with 15 frames/second. Mehrdad Panahpour Tehrani, Kenta Niwa, Norishige Fukushima, Yasushi Hirano, Toshiaki Fujii, Masayuki Tanimoto, Kazuya Takeda, Kenji Mase, Akio Ishikawa, Shigeyuki Sakazawa, Atsushi Koike |
MMSP | 5 |
| 2007 | Experimental on Hierarchical Transmission Scheme for Visible Light Communication using LED Traffic Light and High-Speed CameraabstractLEDs are expected as lighting sources for next generation, and data transmission system using LEDs attract attention. In this paper, we present hierarchical coding scheme using LED traffic lights and high-speed camera for intelligent transport systems (ITS) application. Further, if each of LEDs in traffic lights is individually modulated, parallel data transmissions are possible using a camera as a reception device. Such parallel LED-camera channel can be modeled as spatial low-pass filtered channel of which the cut-off frequency varies according to the distance. To overcome, we propose hierarchical coding scheme based on 2D fast Haar wavelet transform. As results, the proposed hierarchical transmission schemes outperform the conventional on-off keying and the reception of high priority data is guaranteed even LED-camera distance is further. Shintaro Arai, Shohei Mase, Takaya Yamazato, Tomohiro Endo, Toshiaki Fujii, Masayuki Tanimoto, Kiyosumi Kidono, Yoshikatsu Kimura, Yoshiki Ninomiya |
VTC Fall | 5 |
| 2007 | Multiview Video Coding Using View Interpolation and Color CorrectionabstractNeighboring views must be highly correlated in multiview video systems. We should therefore use various neighboring views to efficiently compress videos. There are many approaches to doing this. However, most of these treat pictures of other views in the same way as they treat pictures of the current view, i.e., pictures of other views are used as reference pictures (inter-view prediction). We introduce two approaches to improving compression efficiency in this paper. The first is by synthesizing pictures at a given time and a given position by using view interpolation and using them as reference pictures (view-interpolation prediction). In other words, we tried to compensate for geometry to obtain precise predictions. The second approach is to correct the luminance and chrominance of other views by using lookup tables to compensate for photoelectric variations in individual cameras. We implemented these ideas in H.264/AVC with inter-view prediction and confirmed that they worked well. The experimental results revealed that these ideas can reduce the number of generated bits by approximately 15% without loss of PSNR. Kenji Yamamoto, Masaki Kitahara, Hideaki Kimata, Tomohiro Yendo, Toshiaki Fujii, Masayuki Tanimoto, Shinya Shimizu, Kazuto Kamikura, Yoshiyuki Yashima |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2006 | Arbitrary Listening-Point Generation Using Sub-Band Representation of Sound Wave Ray-SpaceabstractThis paper proposes arbitrary listening-point generation of sound without source localization, and a theory based on the ray-space representation of light rays using sub-band signal processing. An array of beam-formed dynamic microphone-arrays (MAs), are set and each MA generates a multi frequency layer sound-image (SImage) by scanning the viewing range of a camera. Each layer is captured by active microphones in MA for a given frequency range which makes the correct interval to capture that frequency layer, and a sound wave ray-space. Arbitrary listening-point generation of each layer is done by geometry compensation of corresponding images in the location of each MA or SImages or their combination. Each layer sound of an SIamge is generated by averaging the sound wave in each pixel or group of pixel. The listening-point is generated after combining each layer sound in frequency domain and inverse transformation from frequency to time domain Mehrdad Panahpour Tehrani, Yasushi Hirano, Toshiaki Fujii, Shoji Kajita, Kazuya Takeda, Kenji Mase |
ICASSP (5) | 3 |
| 2006 | Multipoint Measuring System for Video and Sound - 100-camera and microphone systemabstractWe developed a novel multipoint measurement system capable of acquiring video and sound at more than 100 points in a "synchronized" manner. In this paper, we first describe the specification of the system and how the system works in detail. Then we report some experimental results that confirm the performance of the system. We also describe test data set we provided for MPEG (moving picture experts group) multi-viewpoint video coding activities. Using this system, we are planning to conduct projects to measure humans and their activities, collect a large volume of real-world data of video and sound, and release them to the public Toshiaki Fujii, Kensaku Mori, Kazuya Takeda, Kenji Mase, Masayuki Tanimoto, Yasuhito Suenaga |
ICME | 1 |
| 2006 | Multi-View Video Coding using View Interpolation and Reference Picture SelectionabstractWe propose a new multi-view video coding method using adaptive selection of motion/disparity compensation based on H.264/AVC. One of the key points of the proposed method is the use of view interpolation as a tool for disparity compensation by assigning reference picture indices to interpolated images. Experimental results show that significant gains can be obtained compared to the conventional approach that was often used Masaki Kitahara, Hideaki Kimata, Shinya Shimizu, Kazuto Kamikura, Yoshiyuki Yashima, Kenji Yamamoto, Tomohiro Yendo, Toshiaki Fujii, Masayuki Tanimoto |
ICME | 8 |
| 2005 | The sound wave ray-spaceabstractThis paper addresses the problem of 3D sound representation without sound source localization and proposes a theory based on the ray-space representation of light rays, which is independent of object's specifications. An array of beam-formed microphone-arrays (MAs), are set and each MA generates a sound-image (SImage) by scanning the viewing range of a camera in the same location. SImage has the same size of an image and contains of blocks of sound wave with duration of one image-frame. Captured SImages with the array of MAs generate the sound wave ray-space. To make a dense SImage ray-space, we propose to use the geometry compensation of corresponding images in the location of each MA. By a dense sound ray-space, any virtual SImage, which corresponds to an arbitrary listening-point, can be generated. The listening-point sound is generated by averaging the sound wave in each pixel or group of pixel of the virtual SImage. Mehrdad Panahpour Tehrani, Yasushi Hirano, Toshiaki Fujii, Shoji Kajita, Kazuya Takeda, Masayuki Tanimoto, Kenji Mase |
ICME | 3 |
| 2004 | Distributed source coding of multiview imagesabstractConsidering nodes energy and channel bandwidth limitations in a multiview images network, avoiding inter-node communication in the coding scheme is necessary and makes the communication efficient, in comparison with conventional coding methods with internode communication. In our system, we consider a multiview images network as an array of nodes on a line with the same distance to each other. Each node in this system includes a camera to capture with a limited processing and communication abilities. We propose a multiview images coding without requiring inter-node communication, to gain the advantage of correlation at decoder side for two network configurations, but the main concept of the proposed coding for both network configurations is the same. The two network configurations are distinguished based on presence of parent node in each cluster, which sends full information to central node (i.e. joint decoder). The other nodes called children nodes and send their partial information to central node. In this paper, a coding scheme is proposed for multiview images network with parent node based on the number of parent nodes (i.e. “one parent node cluster” and “two parent node cluster”) in each cluster. Not that, because we are involved with multiview images, the decoding procedure is searching for correspondence between partially sent multiview images data. In coding scheme with parent node, if each parent node fails the decoding task of children nodes will be failed or the error caused by corresponding search during the decoding of children nodes is increased if network uses other parent nodes to decode the cluster without parent node. To make the network robust in case of node failure, we avoid parent nodes, and developed a network configuration where all nodes are children nodes (i.e. "without parent node cluster"). For such a network, we also propose a coding algorithm for multiview images that the joint decoder at central node is able to decode the partial information of each children node with side information of other children node. So, in case of losing any viewpoint, joint decoder is still able to decode other received partial information of children nodes. Finally, we have compared coding scheme of network with and without inter-node communication for different cluster configuration mentioned herein, and “all parent node cluster” configuration, considering communication rate in the network, symmetry in communication and decoding quality of children node, and their robustness in case of nodes failure. Mehrdad Panahpour Tehrani, Toshiaki Fujii, Masayuki Tanimoto |
VCIP | 2 |
| 2000 | Fast calculation of IFS parameters for fractal image coding
Masaki Harada, Tadahiko Kimoto, Toshiaki Fujii, Masayuki Tanimoto |
VCIP | 3 |
| 2000 | A new flexible acquisition system of ray-space data for arbitrary objectsabstractConventional ray-space acquisition systems require very precise mechanisms to control the small movement of cameras or objects. Most of them adopt camera with a gantry or a turntable. Although they are good for acquiring the ray-space of small objects, they are not suitable for ray-space acquisition of very large structures, such as a building, tower, etc. This paper proposes a new ray-space acquisition system which consists of a camera and a 3-D position and orientation sensor. It is not only a compact and easy-to-handle system, but is also free from limitations of size or shape, in principle. It can obtain any ray-space data as far as the camera is located within the coverage of the 3-D sensor. This paper describes our system and its specifications. Experimental results are also presented. Toshiaki Fujii, Tadahiko Kimoto, Masayuki Tanimoto |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1997 | Automated Testing System for Switching Systems with Multiple ModulesabstractEfficient testing of switching systems is essential to provide new services quickly and to reduce development costs. This paper presents an automated testing system for switching systems with multiple modules. This system allocates hardware resources, which include test equipment and inter-module connector ports, automatically constructs a test environment for each test item, and executes test scripts. It can reduce the complexity of scheduling hardware and setting up the test environment, reduce the testing cycle time, and improve hardware resource utilization. We also show the effects of this system through our experiences developing the NTT's NS8000 Series switching systems. Kumiko Ono, Keisuke Ohmori, Toshiaki Fujii, Takeshi Asakura |
ICC (1) | 3 |
| 1996 | A new fractal image coding scheme employing blocks of variable shapesabstractIn the fractal coding schemes proposed so far based on iterated function systems, an image is represented by a self-affine set of square blocks. We propose a scheme of fractal image coding with blocks of variable shapes. In this scheme, the range blocks are determined by the so-called splitting-and-merging method. Also, we define two kinds of range blocks, shade blocks and non-shade blocks, according to the variance of the brightness level inside blocks. By this method, range blocks both of various size and of various shapes are extracted from an input image. The larger range blocks result in the a greater reduction of the total number of blocks and, hence, the amount of encoded bits. Because of the specified order used in merging of primitive blocks, only a few more bits are added for each range block to encode its shape. The results of computer simulation with a test gray-scale image show that compared with the conventional scheme using square range blocks, the proposed coding scheme reduces about 0.2 bits/pel to achieve almost the same reproduced quality. Masayuki Tanimoto, Hiroshi Ohyama, Tadahiko Kimoto, Sakae Katsuyama, Toshiaki Fujii |
ICIP (1) | 5 |
| 1994 | 3-D image coding based on affine transformabstractThis paper is concerned with the data compression and interpolation of multi-view images. We propose the affine-based disparity compensation based on a geometric relationship. We first investigate the geometric relationship between the point in object space and its projection onto a view image. Then, we propose the disparity compensation based on the affine transform, which utilize the geometric constraints between view images. In this scheme, multi-view images are compressed into the structure and texture of the triangular patches. This scheme not only compresses the multi-view image but also synthesize the view images from any viewpoints in the viewing zone, because the geometric relationship is taken into account. Finally, we report an experiment, where 19 view images were used as the original multi-view image and the amount of data was reduced to 1/19 with an SNR of 34 dB.> Toshiaki Fujii, Hiroshi Harashima |
ICASSP (5) | 1 |