VLDB 2026 Research / reviewers in the wild / expert
Masaki Kitahara
dblp:06/1672
· DBLP profile ↗
15ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0009-4260-1174ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pre-capture Privacy via Adaptive Single-Pixel ImagingabstractAs cameras become ubiquitous in our living environment, invasion of privacy is becoming a significant concern. A common approach to privacy preservation is to remove personally identifiable information from a captured image, but there is a risk of the original image being leaked. In this paper, we propose a pre-capture privacy-aware imaging method that captures images from which the details of a pre-specified anonymized target have been eliminated. The proposed method applies a single-pixel imaging frame-work in which we introduce a feedback mechanism called an aperture pattern generator (APG). The introduced APG adaptively outputs the next aperture pattern to avoid sampling the anonymized target by using already acquired data as a clue. Furthermore, the anonymized target can be set to any object without changing hardware. Except for the removed detailed features of the anonymized target, the captured images are of comparable quality to those captured by a general camera and can be used for various computer vision applications. We target faces and license plates and experimentally show that the proposed method can capture clear images in which detailed features of the anonymized target are eliminated, achieving both privacy and utility. Yoko Sogabe, Shiori Sugimoto, Ayumi Matsumoto, Masaki Kitahara |
WACV | 4 |
| 2024 | Sparse Regularization Based on Reverse Ordered Weighted L1-Norm and Its Application to Edge-Preserving SmoothingabstractSparse regularization is being applied to solve indeterminate inverse problems. However, current regularization is unable to manage sparsity and small perturbations at the same time, and does not perform well enough for some applications. In this study, we propose reversed ordered weighted L1-norm regularization (ROWL) that can tolerate small perturbations while well-handling sparsity. Since ROWL can make proximity mapping easy to compute, it is possible to construct an algorithm to find a suboptimal solution to the inverse problem using the proximity splitting method. Using ROWL for image edge-preserving smoothing, allows us to control both edge sharpness and gradation smoothness. Takayuki Sasaki, Yukihiro Bandoh, Masaki Kitahara |
ICASSP | 3 |
| 2024 | Pose-Invariant Learning for Efficient Person Identification from Hyperspectral Hand ImagesabstractWhile person identification from multi/hyperspectral images of hands has advantages such as contactless and high flexibility in capturing images, it remains a difficult task because individual characteristics are not as clear as those of a face or fingerprints. The state-of-the-art method uses a 3D CNN classifier to capture detailed spectral information. However, this is computationally expensive and is negatively affected by undesired spectral variations caused by changes in hand pose. We propose a new method to address these problems. The key technical components of the proposed method are in the introduction of adversarial learning to learn pose-invariant features and in usage of the separable convolutions to decouple the operations in the channel and spatial directions to improve efficiency. Furthermore, these technical components are integrated into a unified supervised contrastive learning framework, which is suitable for person identification. Experimental results demonstrate that our method achieves higher accuracy than the existing method while significantly reducing computational complexity. Keigo Kunikata, Amane Kashino, Yota Yamamoto, Yukinobu Taniguchi, Yoko Sogabe, Ayumi Matsumoto, Masaki Kitahara, Go Irie |
ICIP | 7 |
| 2024 | Deep Counterfactual Representation Learning for Visual Recognition Against Weather CorruptionsabstractDeep learning has been widely studied for processing and understanding multimedia data, and it does help improve performance. Recent research has shown that deep models are vulnerable to images containing adverse weather corruptions, leading to a safety risk for numerous safety-critical systems (e.g., autonomous driving systems). There are two problems with the current situation. First, collecting data under different weather scenarios is highly difficult in practice. Second, the performance degrades significantly when the training and test data are from different distributions, as exemplified by the weather corrupted test data. As a result, it is challenging to train a model without access to the images containing variations of various weather conditions, and it is difficult to make trained model generalized to unknown data under different weather conditions. In this paper, we introduce aCounterfactual Representation Learning(CRL) method to address these problems. Without access to training data including weather condition variations, our CRL makes the model resistant to unseen test data that has been corrupted by weather condition variations. Our basic idea is inspired by the perspective of counterfactual regularization. We build a causal model that introduces a counterfactual variable to eliminate the unobserved characteristics brought about by weather conditions. In particular, such a counterfactual variable is approximated by randomly shuffled features, echoing the previous empirical observation that the shuffling technique can perturb the shape details while preserving the local textures. We use information theoretic representation learning to encourage the neural networks to learn more powerful and robust features, which consist of two components. We conduct experiments on five benchmark datasets, namely, CIFAR-100-C, ImageNet-C, KITTI-C, BDD100 k, and CityScapes-C, all of which contain weather corruption. The results of our experiments show that our proposed method can not only be a plug-and-play technique but also work nicely for both object recognition and detection. Hong Liu 0009, Yongqing Sun, Yukihiro Bandoh, Masaki Kitahara, Shin'ichi Satoh 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Complexity Reduction of Graph Signal Denoising Based on Fast Graph Fourier TransformabstractDenoising is one of the most fundamental and important problems in signal processing, and graph signal denoising methods have been actively studied. Several graph signal denoising methods based on mathematical programming require solving linear equations involving Laplacian matrix, which creates problem with computational accuracy and running time. This study proposes a fast and accurate solution of linear equations for denoising based on the fast graph Fourier transform method. Moreover, the proposed method can perform denoising not only on graphs for which the fast graph Fourier transform can be performed, but also on a wide class of graphs with more relaxed conditions, without loss of accuracy. Experiments demonstrate the efficiency of the proposed method and confirm that denoising can be performed up to 167.3 times faster without loss of accuracy. Takayuki Sasaki, Yukihiro Bandoh, Masaki Kitahara |
ICIP | 3 |
| 2016 | Motion vector prediction methods considering prediction continuity in HEVCabstractIn video coding standards such as H.265/HEVC and H.264/AVC using motion compensation, motion vector prediction (MVP) refers to motion vectors predicted from the neighboring encoded or decoded blocks. However, if the reference block is encoded in intra prediction mode, the motion vector cannot be acquired because the block has no motion information. This decreases motion vector prediction efficiency in AMVP (Advanced Motion Vector Prediction) mode and motion compensation efficiency in Merge mode. To address this problem, this paper describes a new motion vector prediction method we propose that considers prediction continuity. The method improves coding efficiency by setting a MVP candidate list to a block encoded in intra prediction mode, which the blocks that follow it refer when MVP is performed. Experiment results show that the method provides up to 1.0% rate-distortion improvement over HM16.6. Shinobu Kudo, Masaki Kitahara, Atsushi Shimizu |
PCS | 2 |
| 2015 | A highly parallel motion estimation method based on temporal motion vector prediction for a many-core platformabstractIn hybrid video coding such as H.264/AVC and H.265/HEVC using motion compensation, most coding processes are mainly used for motion estimation. Recently, highly parallel processing devices such as graphics processing units (GPUs) or many-core processors have been utilized to accelerate motion estimation. Although a straightforward way to parallelize motion estimation is block-based parallelization within a frame, motion information of the neighboring block is not available so coding efficiency loss is inevitable. A method using motion vectors of coded frames has been proposed to tackle this problem; however, it causes a decrease in the precision of motion vector predictions in a hierarchical reference structure. This paper proposes a temporal motion vector prediction-based, highly parallel motion estimation method that is applicable to a hierarchical reference structure. It utilizes motion vectors of non-encoded frames referring to the same reference frame of an encoding frame and a block size decision and a concatenation of motion vectors to improve the estimation accuracy. Experiments show that the proposed method achieves up to 9.2% rate-distortion improvement over the conventional method with the similar encoding speed improvement over HM. Shinobu Kudo, Masaki Kitahara, Atsushi Shimizu |
PCS | 2 |
| 2007 | Progressive Coding of Surface Light Fields for Efficient Image Based Rendering
Masaki Kitahara, Hideaki Kimata, Shinya Shimizu, Kazuto Kamikura, Yoshiyuki Yashima |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | View Scalable Multiview Video Coding Using 3-D Warping With Depth MapabstractMultiview video coding demands high compression rates as well as view scalability, which enables the video to be displayed on a multitude of different terminals. In order to achieve view scalability, it is necessary to limit the inter-view prediction structure. In this paper, we propose a new multiview video coding scheme that can improve the compression efficiency under such a limited inter-view prediction structure. All views are divided into two groups in the proposed scheme: base view and enhancement views. The proposed scheme first estimates a view-dependent geometry of the base view. It then uses a video encoder to encode the video of base view. The view-dependent geometry is also encoded by the video encoder. The scheme then generates prediction images of enhancement views from the decoded video and the view-dependent geometry by using image-based rendering techniques, and it makes residual signals for each enhancement view. Finally, it encodes residual signals by the conventional video encoder as if they were regular video signals. We implement one encoder that employs this scheme by using a depth map as the view-dependent geometry and 3-D warping as the view generation method. In order to increase the coding efficiency, we adopt the following three modifications: (1) object-based interpolation on 3-D warping; (2) depth estimation with consideration of rate-distortion costs; and (3) quarter-pel accuracy depth representation. Experiments show that the proposed scheme offers about 30% higher compression efficiency than the conventional scheme, even though one depth map video is added to the original multiview video. Shinya Shimizu, Masaki Kitahara, Hideaki Kimata, Kazuto Kamikura, Yoshiyuki Yashima |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Multiview Video Coding Using View Interpolation and Color CorrectionabstractNeighboring views must be highly correlated in multiview video systems. We should therefore use various neighboring views to efficiently compress videos. There are many approaches to doing this. However, most of these treat pictures of other views in the same way as they treat pictures of the current view, i.e., pictures of other views are used as reference pictures (inter-view prediction). We introduce two approaches to improving compression efficiency in this paper. The first is by synthesizing pictures at a given time and a given position by using view interpolation and using them as reference pictures (view-interpolation prediction). In other words, we tried to compensate for geometry to obtain precise predictions. The second approach is to correct the luminance and chrominance of other views by using lookup tables to compensate for photoelectric variations in individual cameras. We implemented these ideas in H.264/AVC with inter-view prediction and confirmed that they worked well. The experimental results revealed that these ideas can reduce the number of generated bits by approximately 15% without loss of PSNR. Kenji Yamamoto, Masaki Kitahara, Hideaki Kimata, Tomohiro Yendo, Toshiaki Fujii, Masayuki Tanimoto, Shinya Shimizu, Kazuto Kamikura, Yoshiyuki Yashima |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2006 | Multi-View Video Coding using View Interpolation and Reference Picture SelectionabstractWe propose a new multi-view video coding method using adaptive selection of motion/disparity compensation based on H.264/AVC. One of the key points of the proposed method is the use of view interpolation as a tool for disparity compensation by assigning reference picture indices to interpolated images. Experimental results show that significant gains can be obtained compared to the conventional approach that was often used Masaki Kitahara, Hideaki Kimata, Shinya Shimizu, Kazuto Kamikura, Yoshiyuki Yashima, Kenji Yamamoto, Tomohiro Yendo, Toshiaki Fujii, Masayuki Tanimoto |
ICME | 1 |
| 2004 | Hierarchical reference picture selection method for temporal scalability beyond H.264abstractTemporal scalability is effective to adapt the bitstream adaptation for various capabilities of the video terminals and the delivery network, e.g. the processing speed of video terminals and the transmission rate, respectively. This technique uses the reference picture selection method on both the current layer and lower layers. The conventional scalability selects only from the last previous picture of the current layer and that of the first lower layer. This paper proposes a novel prediction scheme, the HRPS (hierarchical reference picture selection) method in which the reference picture is selected from more previous pictures in more layers, in order to improve coding efficiency, keeping the temporal scalability functionality. The proposed method is developed with a modification of the H.264. This paper demonstrates the effectiveness of the HRPS compared with the conventional temporal scalable methods. Hideaki Kimata, Masaki Kitahara, Kazuto Kamikura, Yoshiyuki Yashima |
ICME | 2 |
| 2003 | 3D motion vector coding with block base adaptive interpolation filter on H.264abstractFractional pel motion compensation generally improves coding efficiency due to more precise motion accuracy and low path filtering effect in generating an image at fractional pel positions. In H.264, quarter pel motion compensation is applied, where the image at half pel position is generated by a 6 tap Wiener filter. And the adaptive interpolation filter technique, which adaptively changes filter characteristics for half pel positions has been proposed. That technique also changes the image at quarter pel positions, so it can be exploited to extend motion accuracy to be more precise. In this paper, a 3D motion vector coding (3DMVC) technique with block base adaptive interpolation filter (BAIF) is proposed. This paper also demonstrates the proposed method ensures filter data is successfully integrated into motion vector coding and outperforms the normal H.264. Hideaki Kimata, Masaki Kitahara, Yoshiyuki Yashima |
ICASSP (3) | 2 |
| 2003 | 3D motion vector coding with block base adaptive interpolation filter on H.264abstractFractional pel motion compensation generally improves coding efficiency due to more precise motion accuracy and low path filtering effect in generating image at fractional pel positions. In H.264, quarter pel motion compensation is applied, where image at half pel position is generated by 6 tap Wiener filter. And the adaptive interpolation filter technique, which adaptively changes filter characteristics for half pel positions have been proposed. That technique also changes image at quarter pel positions, so it can be exploited to extent motion accuracy to be more precise. In this paper, 3D motion vector coding (3DMVC) technique with block base adaptive interpolation filter (BAIF) is proposed. This paper also demonstrates the proposed method achieves filter data is successfully integrated into motion vector coding and outperforms the normal H.264. Hideaki Kimata, Masaki Kitahara, Yoshiyuki Yashima |
ICME | 2 |
| 2003 | Recursively weighting pixel domain intra prediction on H.264
Hideaki Kimata, Masaki Kitahara, Yoshiyuki Yashima |
VCIP | 2 |