Akira Kubota

dblp:29/55 · DBLP profile ↗
← Back
46ranked-venue papers
19as first author
7since 2021 · last 2025
0000-0001-5323-2401ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 45 · 18 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2025 Diversity-Aware Active Learning for Object Detection Utilizing Time-of-Day Metadata
abstract
We investigate batch active learning within a pool-based model for object detection in driving videos (ODDV). This paper proposes a novel diversity sampling method that uses time-of-day metadata to reduce skew in the data distribution. Active learning allows a model to actively select images for labeling, improving performance with fewer labeled samples. In a pool-based setting, the model selects informative images from a large unlabeled dataset and updates the training iteratively over multiple rounds. In driving video object detection, metadata such as time-of-day (day, night, dawn/dusk), weather, and scene context are often automatically assigned, which adds no annotation burden since they are not used as labels. Despite this, most active learning methods do not directly incorporate metadata into the acquisition strategy. Our method uses an acquisition function that linearly combines three scores: (i) a distance-based score (S1) that favors images different from the labeled set, (ii) a non-redundancy score (S2) based on distance to reduce duplicates within each round, and (iii) a rarity score (S3) that promotes sampling of time-of-day rare conditions. This approach selects informative images from the unlabeled pool. In experiments with the BDD100K dataset, starting with 350 images and adding 200 images per round up to 1,750, our method—using Faster R-CNN—achieves a final AP50 of 0.335. This exceeds the performance of Random (0.296), Core-set (0.316), and Entropy (0.315), and is comparable to the S1 + S2 method (0.333). Our approach maintains a daytime selection ratio of 40.3% while increasing nighttime and dawn/dusk samples. It achieves the highest AP50 on night images (0.281) and performs similarly to S1 + S2 on dawn/dusk images. Overall, these results show that incorporating time-of-day metadata into sampling strategies reduces dataset bias without degrading overall detection accuracy.
Fumiya Higashide, Akira Kubota
ISM2
2025 Fewer-Shot Self-Supervised Image Recoloring for Deutan Deficiency Based on Laplacian Pyramid
abstract
Color vision deficiency (CVD) reduces the ability to distinguish certain colors, leading to difficulties in daily tasks. Conventional methods improve color discrimination but often cause unnatural color changes for individuals without CVD. This paper presents a self-supervised recoloring method for Deutan deficiency that adapts to different CVD levels. The method replaces the low-frequency component of the Laplacian pyramid with a recolored version generated by a self-supervised module while preserving high-frequency details. Experiments with 500 training images demonstrated that the proposed method enhances color discrimination for individuals with CVD while maintaining naturalness for those without CVD.
Onhi Kato, Akira Kubota
ISM2
2025 LPConv: Laplacian Pyramid Convolutions for Parameter-Efficient Receptive Field Expansion
Naoki Nishiya, Akira Kubota
ISM2
2025 Facial Similarity-Guided Fine-Tuning for Hand Shape Correction in AI-Generated Human Images
abstract
In this study, we propose a method to correct inaccurate hand shapes and finger counts in AI-generated human images. AI-based image generation enables easy creation of illustrations or designs from text or image inputs, but accurately rendering hands remains challenging due to their complex structure; overlapping fingers are often not rendered correctly. Existing approaches detect and mask hand regions for regeneration or overlay depth maps similar to detected hand shapes. However, these methods rely heavily on hand detection accuracy and the quality of the depth map dataset, which can result in hands that differ significantly from the original image. To address these issues, we introduce an alternative approach that applies fine-tuning from a specific step in the generation process. The optimal start step is selected based on facial similarity in images throughout the generation process. This method leverages the observation that noise removal in AI generation progresses from low to high frequencies, allowing the coarse image structure to form initially, with fine-tuning subsequently refining details. Faces, as salient high-frequency features, serve as indicators for determining the optimal step. Unlike methods that only modify masked regions, our approach corrects the entire image, avoiding boundary artifacts. Additionally, it allows adjustment of the balance between hand correction and overall impression by choosing the start step. Experimental results reveal two key findings. First, applying fine-tuning with Textual Inversion at a later start step improves hand correction while maintaining similarity to the original image, demonstrating that adjusting the start step can balance hand accuracy and overall visual impression. Second, by comparing FaceNet feature vector distances at each step, we identified the step with the greatest difference as the optimal start step, effectively estimating the step that best preserves the original image's impression.
Yuki Ryu, Akira Kubota
ISM2
2025 Syntax-Aware Transformer for Sentiment Analysis of Japanese SNS Text
abstract
Japanese SNS sentiment analysis has broad applications in marketing, public opinion surveys, and mental health support. However, because posts are often short, colloquial, and elliptical, they tend to be misclassified. To address this issue, we propose three methods that incorporate syntactic information into a Japanese BERT classifier. Using the GiNZA dependency parser, we extract four types of syntactic features—POS tags, dependency labels, parent position, and tree depth—and encode them via: (1) an embedding-based method that adds syntactic embeddings, (2) a syntactic positional encoding method that applies sinusoidal encodings to positional information, and (3) a hybrid method that combines these two methods. Experiments on the WRIME ver. 2 dataset show that the embedding variant utilizing POS tags and parent position achieves the best performance in 5-class polarity classification, while the hybrid method combining all syntactic features yields the best results in 8-class emotion classification. To the best of our knowledge, this is the first systematic study demonstrating the effectiveness of syntactic integration for Japanese SNS sentiment analysis and identifying optimal combinations of syntactic features.
Sotaro Shiozawa, Akira Kubota
ISM2
2025 Lightweight High-Accuracy Tomato Detection and Classification by Efficientnet-Enhanced YOLOv8
abstract
This study focuses on real-time, high-accuracy tomato detection. In recent years, image recognition technology has been increasingly applied to smart agriculture to support precision cultivation and reduce labor burdens. As labor shortages intensify, accurately estimating harvest timing and fruit quality has become critically important. Deep learning-based object detection models, particularly the YOLO (You Only Look Once) series, have attracted significant attention for real-time and high-accuracy applications. However, these models still face difficulties in detecting tomatoes in complex environments, as fruits resemble each other and the surrounding foliage. Detecting small or occluded tomatoes is especially challenging, since conventional methods often fail to capture sufficient features from limited regions. To overcome these challenges, we incorporate MBConv and Fused-MBConv modules from EfficientNetV2 into the backbone network of YOLOv8. By carefully adjusting the number of layers and the expansion rate at each stage, the proposed method aims to achieve a good balance between computational efficiency and detection confidence. Additionally, we design a new structure within the detection head that better captures the spatial features of small objects. Experiments on a re-annotated tomato dataset show that the proposed method improves the mean Average Precision (mAP) at IoU = 0.5 by 4.4% compared to standard YOLOv8, while maintaining real-time inference speed and reducing parameter count. The improved model also reduces false detections and missed tomatoes caused by occlusion, resulting in more reliable performance. These results demonstrate that combining EfficientNetV2 modules with structural enhancements offers a promising direction for developing lightweight, high-accuracy detection systems tailored for practical smart agriculture applications.
Hayato Tsukada, Akira Kubota
ISM2
2022 Evaluation of Pseudo-Haptics system feedbacking muscle activity
abstract
Differences in perceptions between virtual reality (VR) and reality prevent immersion in VR. To improve immersion in VR, many methods have adopted haptic feedback in VR using pseudo-haptics. However, these methods have little evaluated the effect of force feedback on pseudo-haptics that reflect the user’s state. This paper proposes and evaluates the pseudo-haptics system that manipulates the control/display (C/D) ratio between reality and VR using muscle activity measured. We conducted a user study under three conditions: the C/D ratio is constant, large, or small, depending on the muscle activity. Our results indicated that pseudo-haptics were effective for small C/D ratio settings during low myoelectric intensity.
Yoshihito Tanaka, Akira Kubota
VRST2
2018 Audio Feature Extraction Based on Sub-Band Signal Correlations for Music Genre Classification
abstract
We present novel low-level audio features that are based on correlations between sub-band audio signals decomposed by undecimated wavelet transform. Under the assumption that SVM is used for classifier learning, the experimental results on GTZAN dataset showed that the proposed method demonstrated the best accuracy of 81.5%, outperforming the conventional methods.
Takuya Kobayashi, Akira Kubota, Yusuke Suzuki
ISM2
2017 Fast camera self-calibration for synthesizing Free Viewpoint soccer Video
abstract
Recently, non-fixed camera-based free viewpoint sports video synthesis has become very popular. Camera calibration is an indispensable step in free viewpoint video synthesis, and the calibration has to be done frame by frame for a non-fixed camera. Thus, calibration speed is of great significance in real-time application. In this paper, a fast self-calibration method for a non-fixed camera is proposed to estimate the homography matrix between a camera image and a soccer field model. As far as we know, it is the first time to propose constructing feature vectors by analyzing crossing points of field lines in both camera image and field model. Therefore, different from previous methods that evaluate all the possible homography matrices and select the best one, our proposed method only evaluates a small number of homography matrices based on the matching result of the constructed feature vectors. Experimental results show that the proposed method is much faster than other methods with only a slight loss of calibration accuracy that is negligible in final synthesized videos.
Akira Kubota, Kaoru Kawakita, Keisuke Nonaka, Hiroshi Sankoh, Sei Naito
ICASSP2
2016 View synthesis from defocused stereo images with coded aperture
abstract
In this paper, we propose a linear view interpolation method using defocused stereo images with coded aperture. The proposed method, without requiring depth estimation, reconstructs intermediate view by iteratively minimizing the squared error of stereo images model that is a linear combination of depth textures. Using the defocused images with Zhou's aperture code makes view interpolation higher quality.
Yuta Narukiyo, Akira Kubota
PCS2
2016 Adding Search Queries to Picture Lifelogs for Memory Retrieval
abstract
A picture lifelog is a type of lifelog that consists of pictures, mainly taken by the user. Recently, users have been able to easily create picture lifelogs because many portable devices such as smart phones have a camera. When a user sees a picture in their picture lifelog, it is sometimes difficult to recall the events related to the picture. Therefore, we proposed to combine search queries on a picture lifelog in order to support memory retrieval. Search queries are input into a web search engine to satisfy a user's need for information. Recently, because of the prevalence of smart phones, the opportunity to input search queries has increased to anytime and anywhere. Search queries are stored in a cloud user database such as Google search history. In addition, those search queries imply what the user was thinking at the time. We investigated whether search queries enable a user to recall their thoughts regarding picture lifelogs. Thus, we conducted an experiment to ascertain whether search queries reminded a user of past events. As a result, we reveal that displaying a picture with search queries performed around the time it was taken tends to improve users' memories better than its time, location, or emails sent during that time.
Akira Kubota, Tomu Tominaga, Yoshinori Hijikata, Nobuchika Sakata
WI1
2014 Linear view/image restoration for dense light fields
abstract
Scene refocusing and free viewpoint image reconstruction from a dense light field acquired by a massive camera array or a lens array are actively studied. However, it is difficult to directly integrate acquired information into a light field without defects. For example, actually, a massive camera array hardly works without failure of several cameras, and part of multi-view images are missing. On the other hand, because a lens array consists of very small lenses gathering only limited rays, acquired images themselves are often noisy. This paper proposes simple linear filters achieving view/image restoration to obtain a dense light field robustly in suppressing such defects without block matching used by the conventional approaches of view synthesis and denoising. Experimental results show that the proposed filters can be flexibly designed for various defects based on the relation between a 4D light field and the corresponding 3D multi-focus images.
Kazuya Kodama, Akira Kubota
ICIP2
2013 Efficient Reconstruction of All-in-Focus Images Through Shifted Pinholes From Multi-Focus Images for Dense Light Field Synthesis and Rendering
abstract
Scene refocusing beyond extended depth of field for users to observe objects effectively is aimed by researchers in computational photography, microscopic imaging, and so on. Ordinary all-in-focus image reconstruction from a sequence of multi-focus images achieves extended depth of field, where reconstructed images would be captured through a pinhole in the center on the lens. In this paper, we propose a novel method for reconstructing all-in-focus images through shifted pinholes on the lens based on 3D frequency analysis of multi-focus images. Such shifted pinhole images are obtained by a linear combination of multi-focus images with scene-independent 2D filters in the frequency domain. The proposed method enables us to efficiently synthesize dense 4D light field on the lens plane for image-based rendering, especially, robust scene refocusing with arbitrary bokeh. Our novel method using simple linear filters achieves not only reconstruction of all-in-focus images even for shifted pinholes more robustly than the conventional methods depending on scene/focus estimation, but also scene refocusing without suffering from limitation of resolution in comparison with recent approaches using special devices such as lens arrays in computational photography.
Kazuya Kodama, Akira Kubota
IEEE Trans. Image Process.2
2012 Quality-optimized encoding of JPEG images using transform domain sparsification
abstract
To account for the unique characteristics and limitations of the human visual system (HVS) when perceiving images, a variety of perceptual quality metrics have been proposed in the literature. Tailoring rate-distortion (RD) optimization for each metric is cumbersome and time-consuming. In this paper, we propose a general RD-optimization strategy called “transform domain bounding box” (BB) that can easily adapt to different quality metrics for JPEG-like block-based encoding of images. First, we define an objective function that is a weighted sum of the l0-norm of the transform coefficients (a proxy for rate) and distortion from the transform domain representation. Next, for a given distortion target τ, we define a don't care region (DCR) that specifies a search region of representations with distortion ≤τ. We then show that the sparsest transform domain representation (lowest encoding rate) inside a BB that tightly contains the DCR can be constructed efficiently. Varying τ to induce different DCRs and corresponding BBs results in a set of constructed sparse representations of different sparsity counts, and the one that optimally trades off rate and distortion can be easily identified as solution to our objective. We show that our proposed BB strategy can be easily re-targeted for three common quality metrics: MSE, MSE-HVS-M and SSIM. Experimental results show that our BB strategy outperformed unoptimized JPEG compression by up to 1dB in PSNR when distortion metric is MSE, up to 2dB when metric is MSE-HVS-M, and up to 0.005 when metric is SSIM.
Junichi Ishida, Gene Cheung, Akira Kubota, Antonio Ortega
MMSP3
2011 Transform domain sparsification of depth maps using iterative quadratic programming
abstract
Compression of depth maps is important for “texture plus depth” format of multiview images, which enables synthesis of novel intermediate views via depth-image-based rendering (DIBR) at decoder. Previous depth map coding schemes exploit unique depth data characteristics to compactly and faithfully reproduce the original signal. In contrast, since depth map is only a means to the end of view synthesis and not itself viewed, in this paper we explicitly manipulate depth values, without causing severe synthesized view distortion, in order to maximize representation sparsity in the transform domain for compression gain - we call this process transform domain spar-sification (TDS). Specifically, for each pixel in the depth map, we first define a quadratic penalty function, with minimum at ground truth depth value, based on synthesized view's distortion sensitivity to the pixel's depth value during DIBR. We then define an objective for a depth signal in a block as a weighted sum of: i) signal's sparsity in the transform domain, and ii) per-pixel synthesized view distortion penalties for the chosen signal. Given that sparsity (l0-norm) is non-convex and difficult to optimize, we replace the l0-norm in the objective with a computationally inexpensive weighted l2-norm; the optimization is then an unconstrained quadratic program, solvable via a set of linear equations. For the weighted l2-norm to promote sparsity, we solve the optimization iteratively, where at each iteration weights are readjusted to mimic sparsity-promoting lτ-norm, 0 ≤ τ ≤ 1. Using JPEG as an example transform codec, we show that our TDS approach gained up to 1.7dB in rate-distortion performance for the interpolated view over compression of unaltered depth maps.
Gene Cheung, Junichi Ishida, Akira Kubota, Antonio Ortega
ICIP3
2011 Depth map coding using graph based transform and transform domain sparsification
abstract
Depth map compression is important for compact “texture-plus-depth” representation of a 3D scene, where texture and depth maps captured from multiple camera viewpoints are coded into the same format. Having received such format, the decoder can synthesize any novel intermediate view using texture and depth maps of two neighboring captured views via depth-image-based rendering (DIBR). In this paper, we combine two previously proposed depth map compression techniques that promote sparsity in the transform domain for coding gain-graph-based transform (GBT) and transform domain sparsification (TDS) - together under one unified optimization framework. The key to combining GBT and TDS is to adaptively select the simplest transform per block that leads to a sparse representation. For blocks without detected prominent edges, the synthesized view's distortion sensitivity to depth map errors is low, and TDS can effectively identify a sparse depth signal in fixed DCT domain within a large search space of good signals with small synthesized view distortion. For blocks with detected prominent edges, the synthesized view's distortion sensitivity to depth map errors is high, and the search space of good depth signals for TDS to find sparse representations in DCT domain is small. In this case, GBT is first performed on a graph defining all detected edges, so that filtering across edges is avoided, resulting in a sparsity count ρ in GBT. We then incrementally add the most important edge to an initial no-edge graph, each time performing TDS in the resulting GBT domain, until the same sparsity count ρ is achieved. Experimentation on two sets of multiview images showed gain of up to 0.7dB in PSNR in synthesized view quality compared to previous techniques that employ either GBT or TDS alone.
Gene Cheung, Woo-Shik Kim, Antonio Ortega, Junichi Ishida, Akira Kubota
MMSP5
2010 Robust reconstruction of arbitrarily deformed bokeh from ordinary multiple differently focused images
abstract
This paper deals with a method of generating seriously deformed bokeh on reconstructed images from ordinary multiple differently focused images including just simple bokeh such as Gaussian blurs. We previously proposed scene re-focusing with various iris shapes by applying a three-dimensional filter to the multi-focus images. However, actually the proposed method implicitly assumed that the feature of the iris can be expressed mathematically and it has some symmetry like a horizontally open iris. In this paper, at first, the captured multi-focus images are robustly decomposed into components, each of which goes through its own corresponding pin-hole on the lens, by using dimension reduction and a two-dimensional filter. Then, based on the appropriate composition of the components, reconstruction of arbitrarily deformed bokeh introduced by any user-defined iris is achieved. By some experiments, we show that our novel method can generate even seriously deformed bokeh that does not have simple symmetry.
Kazuya Kodama, Ippeita Izawa, Akira Kubota
ICIP3
2010 Sparse representation of depth maps for efficient transform coding
abstract
Compression of depth maps is important for “image plus depth” representation of multiview images, which enables synthesis of novel intermediate views via depth-image-based rendering (DIBR) at decoder. Previous depth map coding schemes exploit unique depth characteristics to compactly and faithfully reproduce the original signal. In contrast, given that depth maps are not directly viewed but are only used for view synthesis, in this paper we manipulate depth values themselves, without causing severe synthesized view distortion, in order to maximize sparsity in the transform domain for compression gain. We formulate the sparsity maximization problem as an l0-norm optimization. Given l0-norm optimization is hard in general, we first find a sparse representation by iteratively solving a weighted l1minimization via linear programming (LP). We then design a heuristic to push resulting LP solution away from constraint boundaries to avoid quantization errors. Using JPEG as an example transform codec, we show that our approach gained up to 2.5 dB in rate-distortion performance for the interpolated view.
Gene Cheung, Akira Kubota, Antonio Ortega
PCS2
2010 A new hybrid parallel intra coding method based on interpolative prediction
abstract
The hybrid coding method combining the predictive coding with the orthogonal transformation and the quantization is mainly used recently. This paper proposes a new hybrid parallel Intra Coding based on interpolative prediction which uses correlations between neighboring pixels, including non-causal pixels. In order to get high prediction performance, the optimal quantizing scheme, which is used to cancel the error that expands when decoding, is used. Furthermore, a new type of block shape, which enables parallel coding, is proposed to simplify the processing of interpolative prediction. The result of comparison between proposed method and intra coding method in H.264 shows that the PSNR of proposed technique achieves 1 dB to 4 dB improvement in Luminance, especially for image with more details.
Akira Kubota, Yoshinori Hatori
PCS2
2010 A high efficiency coding framework for multiple image compression of circular camera array
abstract
Many existing multi-view video coding techniques remove inter-viewpoint redundancy by applying disparity compensation in conventional video coding frameworks, e.g., H264/MPEG4. However, conventional methodology works ineffectively as they ignore the special features of the inter-view-point disparity. This paper proposes a framework using virtual plane (VP) for multi-view image compression, such that we can largely reduce the disparity compensation cost. Based on this VP predictor, we design a poxel (probabilistic voxelized volume) framework, which integrates the information of the cameras in different view-points in the polar axis to obtain a more effective compression performance. In addition, considering the replay convenience of the multi-view video at the receiving side, we reform overhead information in polar axis at the sending side in advance.
Dongming Xue, Akira Kubota, Yoshinori Hatori
PCS2
2009 View Interpolation using defocused stereo images: A space-invariant filtering approach
abstract
This paper presents a novel view interpolation method for reconstructing an intermediate all-in-focus image from defocused stereo images of a scene consisting of two depths. In the presented method, the novel view can be reconstructed as a sum of the filtered stereo images. The reconstruction filters are derived as stable and linear space-invariant ones; thus the presented method does not require estimation of feature correspondence. Experimental results on real images showed that view interpolation was possible with sufficient quality.
Akira Kubota, Tomohito Hamanaka, Yoshinori Hatori
ICIP1
2007 Simple and Fast All-in-Focus Image Reconstruction Based on Three-Dimensional/Two-Dimensional Transform and Filtering
abstract
This paper deals with all-in-focus image reconstruction by merging multiple differently focused images. We previously proposed a method of generating an all-in-focus image from multi-focus imaging sequences based on a 3-D filtering. In this paper, we first combine the sequence into a 2-D image. Just by applying a 2-D filter to the image, we realize fast reconstruction of all-in-focus images robustly. In order to reduce the cost of image acquisition, the optimal number of multiple differently focused images is also discussed. In addition, we introduce a simple estimation method utilizing the 2-D filter for the parameter of 3-D blurs. We show experimental results of fast all-in-focus image reconstruction by using synthetic and real images.
Kazuya Kodama, Hiroshi Mo, Akira Kubota
ICASSP (1)3
2007 A Variational Recovery Method for Virtual View Synthesis
abstract
This paper presents a novel method based on image recovery scheme for virtual view synthesis. First, using multiple hypothetical depths, we generate multiple candidate images for the desired virtual view. The generated images suffer from blending artifacts (seen like blur) due to pixel mis-correspondence. From these blurry images, we recover an image without artifacts (i.e., an all infocus image) by minimizing an energy functional of unknown textures at all the hypothetical depths. The desired image is finally reconstructed as the sum of all the estimated textures. Simulation result shows that texture color value exist over all the hypothetical depths (i.e., depth is not uniquely identified for every pixel) nevertheless the desired image can be reconstructed with adequate quality.
Akira Kubota, Takahiro Saito
ICIP (4)1
2007 Image-Based Refocusing by 3D Filtering
Akira Kubota, Kazuya Kodama, Yoshinori Hatori
PSIVT1
2007 View interpolation by inverse filtering: generating the center view using multiview images of circular camera array
abstract
This paper deals with a view interpolation problem using multiple images captured with a circular camera array. An inverse filtering method for reconstructing a virtual image at the center of the camera array is proposed. First, we generate a candidate image by adding all correspondence pixel values based on multiple layers assumed in a scene. Second, we model a linear relationship between the desired virtual image and the candidate image, and then derive the inverse filter to reconstruct the virtual image. Since we can derive the inverse filter independent of the scene structure, the proposed method requires no depth estimation. Simulation results using synthetic images show that increasing the number of cameras and layers improves the quality of the virtual image.
Akira Kubota, Kazuya Kodama, Yoshinori Hatori
VCIP1
2007 Reconstructing Dense Light Field From Array of Multifocus Images for Novel View Synthesis
abstract
This paper presents a novel method for synthesizing a novel view from two sets of differently focused images taken by an aperture camera array for a scene consisting of two approximately constant depths. The proposed method consists of two steps. The first step is a view interpolation to reconstruct an all-in-focus dense light field of the scene. The second step is to synthesize a novel view by a light-field rendering technique from the reconstructed dense light field. The view interpolation in the first step can be achieved simply by linear filters that are designed to shift different object regions separately, without region segmentation. The proposed method can effectively create a dense array of pin-hole cameras (i.e., all-in-focus images), so that the novel view can be synthesized with better quality.
Akira Kubota, Kiyoharu Aizawa, Tsuhan Chen
IEEE Trans. Image Process.1
2006 Free Viewpoint, Iris and Focus Image Generation by Using a Three-Dimensional Filtering Based on Frequency Analysis of Blurs
abstract
This paper describes a method of image generation based on transformation integrating certain sequences of multiple differently focused images. First, we assume that a scene is defocused by a geometrical blurring model. Then we combine spatial frequencies of the scene and the sequence with a 3-D convolution filter that expresses how the scene is defocused on the sequence. The filter can be represented with a linear combination of ray-sets through each point of the lens. Based on the relation, in the 3-D frequency domain we extract each ray-set from the filter as certain frequency components and merge them to reconstruct various filters that can generate images with different viewpoints and blurs.
Kazuya Kodama, Hiroshi Mo, Akira Kubota
ICASSP (2)3
2006 Free Iris Scene Re-Focusing Based on a Three-Dimensional Filtering of Multiple Differently Focused Images
abstract
This paper describes a method of scene re-focusing with various iris shapes by integrating a sequence of multiple differently focused images. First, we introduce a formula that combines the sequence and spatial information of a scene with a convolution of a 3-D blur. The blur expresses how the scene is defocused in the sequence. Based on the formula, in the 3-D frequency domain we can design filters merging certain frequency components of the sequence to generate images that would be acquired with various iris shapes. Some experiments of scene re-focusing using synthetic and real images indicate that we can not only arbitrarily suppress blurs of the sequence but also generate images with asymmetrical blurs like motion blurs.
Kazuya Kodama, Hiroshi Mo, Akira Kubota
ICIP3
2006 Deconvolution Method for View Interpolation Using Multiple Images of Circular Camera Array
abstract
This paper deals with a view interpolation problem using multiple images captured with a circular camera array. A novel deconvolution method for reconstructing a virtual image at the center of the camera array is presented and discussed in a framework of image restoration in the frequency domain. The reconstruction filter does not depend on the scene structure; therefore the presented method requires no depth estimation. The simulation result using synthetic images shows that increasing the number of the cameras improves the quality of the virtual image.
Akira Kubota, Kazuya Kodama, Yoshinori Hatori
ICIP1
2006 Special issue on multi-view image processing and its application in image-based rendering
Akira Kubota, Tsuhan Chen
Signal Process. Image Commun.1
2005 Direct filtering method for image based rendering
abstract
Image based rendering (IBR) is basically a light ray resampling method for generating a novel image without aliasing artifacts from a given set of sampled light rays. To avoid aliasing artifacts, conventional methods require an estimate of the scene geometry. In this paper, we present a novel IBR method by linear filtering without estimating the scene geometry, for the simplified case when generating a virtual image at the center of 2/spl times/2 sparse camera array for a two depth-layers scene. The reconstruction filter used in the proposed method can derive by integrating all the process in our previously proposed method into a one-shot process.
Akira Kubota, Kiyoharu Aizawa, Tsuhan Chen
ICIP (3)1
2005 Reconstructing arbitrarily focused images from two differently focused images using linear filters
abstract
We present a novel filtering method for reconstructing an all-in-focus image or an arbitrarily focused image from two images that are focused differently. The method can arbitrarily manipulate the degree of blur of the objects using linear filters without segmentation. The filters are uniquely determined from a linear imaging model in the Fourier domain. An effective and accurate blur estimation method is developed. The simulation results show that the accuracy and computational time of the proposed method are improved compared with the previous iterative method and that the effects of blur estimation error on the quality of the reconstructed image are very small. The method performs well for real images acquired without visible artifacts.
Akira Kubota, Kiyoharu Aizawa
IEEE Trans. Image Process.1
2004 Virtual view synthesis through linear processing without geometry
abstract
This paper presents a new approach for virtual view synthesis that does not require any information of scene geometry. Our approach first generates multiple virtual views at the same position based on multiple depths by the conventional view interpolation method. The interpolated views suffer from blurring and ghosting artifacts due to the pixel mis-correspondence. Secondly, the multiple views are integrated into a novel view where all regions are focused. This integration problem can be formulated as the problem of solving a set of linear equations that relates the multiple views. To solve this set of equations, two methods using projection onto convex sets (POCS) and inverse filtering are presented that effectively integrate the focused regions in each view into a novel view. Experimental results using real images show the validity of our methods.
Akira Kubota, Kiyoharu Aizawa, Tsuhan Chen
ICIP1
2004 A focus measure for light field rendering
abstract
Light field rendering is a fundamental method for synthesizing free-viewpoint images from a set of multiviewpoint images. In the simplest case, the scene structure is approximated by a simple plane: a focal plane. This approximation leads to focus-like effects on synthetic images where the focused depth is determined by the focal plane. A serious problem is that the range of the focused depth is too small in most practical cases. In this paper, we propose a focus measure that is specialized for synthetic images by light field rendering. When a set of differently-focused images is generated at a given viewpoint, the proposed focus measure enables us to obtain a depth map and an all in-focus image. Our approach has some remarkable differences from other related techniques, such as depth-from-stereo and depth-from-focus methods. Experimental results show that the proposed method effectively enhances PSNR of the final synthetic images.
Keita Takahashi 0001, Akira Kubota, Takeshi Naemura
ICIP2
2004 Reconstructing dense light field from a multi-focus images array
abstract
The work presents a novel method for synthesizing a novel view from two sets of differently focused images taken by a sparse camera array for a scene of two approximately constant depths. The proposed method consists of two steps. The first step is a view interpolation to reconstruct an all-focused dense light field of the scene. The second step is to synthesize a novel view by a light-field rendering technique from the reconstructed dense light field. The view interpolation can be achieved simply by linear filters that are designed to convert defocus effects to parallax effects without estimating the depth map of the scene. The proposed method can effectively create a dense array of pin-hole cameras (i.e., all-focused images), so that the final novel view is better than traditional method using a sparse array of cameras. Experimental results on real images from four aligned cameras are shown.
Akira Kubota, Kiyoharu Aizawa, Tsuhan Chen
ICME1
2003 A novel image-based rendering method by linear filtering of multiple focused images acquired by a camera array
abstract
In this paper, we present a novel approach to image-based rendering (IBR) for generating an arbitrary view image with arbitrary focus for a scene consisting two approximately constant depths. The presented method differs from the conventional IBRs using multiple view images in that we acquire two differently focused images from each camera position and render parallax and focus effects on objects at different depth simply by linear filtering of the acquired images without segmentation. Experimental results on the real images acquired with 4 cameras located in parallel are presented.
Akira Kubota, Kiyoharu Aizawa
ICIP (3)1
2003 Object-based approach to image-based rendering with linear filters using defocus information
Akira Kubota, Kiyoharu Aizawa
VCIP1
2002 Arbitrary view and focus image generation: rendering object-based shifting and focussing effect by linear filtering
abstract
This paper presents a novel method to render shifting (parallax) and focussing effects on an object in a scene for generating a virtual view image with arbitrary focus. Two differently focused images of the same scene-near-focused image and far-focused image-are used as input under the assumption that a scene has near and far objects. The proposed method can freely handle the shifting and the focussing effect on each object according to the view point and the focus depth of the virtual camera only by linear filtering of the input images without any segmentation or modeling of the objects. Experimental results using real images are shown to test the performance of the method.
Akira Kubota, Kiyoharu Aizawa
ICIP (1)1
2001 The results in the clinical trial of CAD system for lung cancer using helical CT images
abstract
We have developed a computer assisted automatic detection system for lung cancer that detects tumor candidates at an early stage from helical CT images. In July 1997, we started the comparative field trial using our system prospectively. Chest CT images obtained by helical CT scanner have drawn great interest in the detection of suspicious regions. However, mass screening based on helical CT images leads to a considerable number of images to be diagnosed. We expect that our system can reduce the time complexity and increase diagnostic confidence. We describe the results of nodule detection for definite diagnosis. We show the prospective results and the retrospective results. These results show that the system can detect lung cancer candidates at an early stage successfully and can be applied to a mass screening.
Yoshiki Kawata, Hironobu Ohmatsu, Ryutaro Kakinuma, Noriyuki Moriyama, Mitsuru Kubo, Noboru Niki, Kenji Eguchi, Masahiro Kaneko, Akira Kubota
ICIP (1)9
2001 A new approach to depth range detection by producing depth-dependent blurring effect
abstract
This paper presents a new method for detecting the depth range of a scene from two differently focused images. This method first generates three images with differently emphasized blur using linear filtering of the two acquired images for more depth-sensitive estimation. In this point, this method is different from conventional approaches. The discrete Fourier transform of the ratio between the generated images is then introduced as a criterion of the depth range. Finally, through the thresholding process of the criterion according to its correspondence to the depth on the basis of the imaging model, the depth range is estimated in five depth steps. Experiments for synthesized and real images are performed to test the proposed method.
Akira Kubota, Kiyoharu Aizawa
ICIP (3)1
2001 All-focused image generation and 3D modeling of microscopic images of insects
abstract
We discuss a method which generates an all-focused image from a large number of microscopic images of insects. First, we describe our previously proposed select-and-merge method for all-focused image acquisition. We can get good results by using this method for two differently focused images. However, this method can not give a good result when applied to a large number of images. We propose a method which compares each images with its preceding and subsequent images using the estimation method of focused regions. We compare the results of reconstruction using the novel method and using our previously proposed method. Finally, we generate a 3D image using the intensity of the all-focused image and depths of the in-focus regions, both acquired by our proposed method.
Y. Tsubaki, Akira Kubota, Kiyoharu Aizawa
ICIP (2)2
2000 Inverse Filters for Reconstruction of Arbitrarily Focused Images from two Differently Focused Images
abstract
This paper describes a novel filtering method to reconstruct an arbitrarily focused image from two differently focused images. Based on the assumption that image scene has two layers-foreground and background-, two differently focused images are used as inputs, one of which is focused on the foreground and the other is focused on the background. The linear equation that holds between these images and the desired image, which is derived from their imaging models, can be formulated as an image restoration problem. This paper shows that the solution of this problem exists as an inverse filter and the desired image can be reconstructed only by the linear filters. As a result, fast reconstruction with high accuracy can be achieved. Experiments using real images are shown.
Akira Kubota, Kiyoharu Aizawa
ICIP1
2000 Inverse filters for generation of arbitrarily focused images
Akira Kubota, Kiyoharu Aizawa
VCIP1
2000 Producing object-based special effects by fusing multiple differently focused images
abstract
We propose a novel approach for producing special visual effects by fusing multiple differently focused images. This method differs from conventional image fusion techniques because it enables us to arbitrarily generate object-based visual effects such as blurring, enhancement, and shifting. Notably, the method does not need any segmentation. Using a linear imaging model, it directly generates the desired image from multiple differently focused images.
Kiyoharu Aizawa, Kazuya Kodama, Akira Kubota
IEEE Trans. Circuits Syst. Video Technol.3
1999 Producing Object-Based Special Visual Effects by Integrating Multiple Differently Focused Images: Implicit 3D Approach to Image Content Manipulation
abstract
We propose a novel approach to image content manipulation. It enables us to arbitrarily manipulate an object in a scene by linear processings such as blurring, enhancement and shift etc. Notably, the method does not need any segmentation nor 3D modeling. Making use of multiple differently focused images and a linear imaging model, it directly generates the desired image from the original images. A special camera is developed which can acquire three differently focused image sequences, in order to extend the proposed method to image sequence processing.
Kiyoharu Aizawa, Kazuya Kodama, Akira Kubota
ICIP (2)3
1999 Registration and Blur Estimation Methods for Multiple Differently Focused Images
abstract
In this paper, we propose a registration method between multiple differently focused images using the hierarchical block matching technique in which displacement, scale and rotation are taken into account. Local deformation due to lens distortion is further corrected by local matching. We also propose an efficient estimation method of blur parameters of defocused regions in these focused images. Simulation results showed that the proposed methods achieve high accuracy. In experiments using real images captured by hand-held camera, an all focused image with good quality was able to be automatically generated using the corrected images and the estimated parameters.
Akira Kubota, Kazuya Kodama, Kiyoharu Aizawa
ICIP (2)1