Kohei Matsuzaki

dblp:181/4520 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0001-9386-2192ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 TS-PCI: Point Cloud Frame Interpolation with Time-Aware Point Cloud Sampling and Self-Supervised Learning Strategy
abstract
Recent point cloud frame interpolation methods predict an interpolated frame through the merging of two intermediate frames constructed by scene flow estimation. However, these methods introduce errors due to generation since they adopt a generative approach to merge the frames, degrading frame interpolation performance. In this paper, we propose a point cloud frame interpolation method with time-aware point cloud sampling and a self-supervised learning strategy, termed TS-PCI. The proposed method introduces a time-aware learning-based point cloud sampling model to merge the two frames into a single frame in a non-generative approach. The proposed method also introduces an attention-based geometry refinement model to improve the geometric quality of the sampled point clouds. Furthermore, the proposed method adopts a self-supervised strategy that dynamically creates ground truth labels for point cloud sampling, allowing the models to be trained in an end-to-end manner. Experimental results on three large-scale datasets show that the proposed method achieves superior performance compared to state-of-the-art methods.
Kohei Matsuzaki, Keisuke Nonaka
WACV1
2025 Point Cloud Color Upsampling with Attention-Based Coarse Colorization and Refinement
abstract
Point cloud color upsampling is an important and less explored research topic. State-of-the-art methods colorize points based on the colors of neighboring points and geometric distances. However, these methods often suffer from blurring and noise at color boundaries since object textures can have large color variations even between geometrically neighboring positions. In this paper, we propose a point cloud color upsampling method with attention weights for neighboring points. The proposed method first performs coarse colorization with the colors of low-resolution points neighboring the high-resolution points and predicted weights. Then, it refines the colors by predicting offsets for high-resolution points with aggregate features obtained from the low-resolution points. Both quantitative and qualitative experimental results on datasets acquired in real-world environments demonstrate that the proposed method achieves significantly superior color upsampling performance compared to state-of-the-art methods.
Kohei Matsuzaki, Keisuke Nonaka
WACV1
2025 MDLPCC: Misalignment-aware dynamic LiDAR point cloud compression
abstract
LiDAR point cloud plays an important role in various real-world areas. It is usually generated as sequences by LiDAR on moving vehicles. Regarding the large data size of LiDAR point clouds, Dynamic Point Cloud Compression (DPCC) methods are developed to reduce transmission and storage data costs. However, most existing DPCC methods neglect the intrinsic misalignment in LiDAR point cloud sequences, limiting the rate–distortion (RD) performance. This paper proposes a Misalignment-aware Dynamic LiDAR Point Cloud Compression method (MDLPCC), which alleviates the misalignment problem in both macroscope and microscope. MDLPCC exploits a global transformation (GlobTrans) method to eliminate the macroscopic misalignment problem, which is the obvious gap between two continuous point cloud frames. MDLPCC also uses a spatial–temporal mixed structure to alleviate the microscopic misalignment, which still exists in the detailed parts of two point clouds after GlobTrans. The experiments on our MDLPCC show superior performance over existing point cloud compression methods.
Ao Luo, Linxin Song, Keisuke Nonaka, Jinming Liu 0001, Kyohei Unno, Kohei Matsuzaki, Heming Sun, Jiro Katto
J. Vis. Commun. Image Represent.6
2023 Point Cloud Sampling Preserving Local Geometry for Surface Reconstruction
Kohei Matsuzaki, Keisuke Nonaka
BMVC1
2023 Rate-Distortion Optimized Variable-Node-size Trisoup for Point Cloud Coding
abstract
Triangle soup (Trisoup) is being studied as a new coding tool for Geometry-based Point Cloud Compression (G-PCC) stan-dardized in the Moving Picture Experts Group (MPEG). Outside of MPEG, a variable-node-size extension of Trisoup is studied to increase the flexibility of G-PCC. A primary advantage of variable node size is to achieve better coding performance by selecting appropriate node size according to local geometric complexity and required bits. However, the node size is not optimized in terms of bit rate and distortion in the conventional extension. To maximize the coding performances of the variable-node-size method, we propose a new cost function considering both bit rates and distortions. The experimental results show that the proposed method provides -1.5 % coding performance improvement in point-to-point PSNR versus bit rate against the conventional extension.
Kyohei Unno, Kohei Matsuzaki, Satoshi Komorita, Kei Kawamura
ICASSP2
2022 Relative Viewpoint Estimation Based on Structured 3d Representation Alignment
abstract
Relative viewpoint estimation is a fundamental problem in various image processing applications. Traditional estimation approaches can fail if sufficient appearance overlap is not observed between two images. Recent advances in 3D representation learning from images have made it possible to exploit the underlying 3D structure. In this paper, we propose a relative viewpoint estimation method using an end-to-end trainable network that learns structured 3D representations. In the proposed method, an independent coordinate system is set for each image in order to construct a structured 3D representation. This makes it possible to estimate the relative viewpoint by aligning those representations through coordinate transformations. Experimental results on the ShapeNet, Pix3D, and Thingi10K datasets demonstrated that the proposed method achieves accurate estimation even if there is not sufficient observable appearance overlap between the images.
Kohei Matsuzaki, Kei Kawamura
ICASSP1
2019 Representation Learning via Parallel Subset Reconstruction for 3D Point Cloud Generation
abstract
Three-dimensional (3D) point cloud processing has attracted a great deal of attention in computer vision, robotics, and the machine learning community because of significant progress in deep neural networks on 3D data. Another trend in the community is learning of generative models based on generative adversarial networks. In this paper, we propose a framework for 3D point cloud generation based on a combination of auto-encoders and generative adversarial networks. The framework first trains auto-encoders to learn latent representations, and then trains generative adversarial networks in the learned latent space. We focus on improving the training method for auto-encoders in order to generate 3D point clouds with higher fidelity and coverage. We add parallel sub-decoders that reconstruct subsets of the input point cloud. In order to construct these subsets, we introduce a point sampling algorithm that imposes a method to sample spatially localized point sets. These local subsets are utilized to measure local reconstruction losses. We train auto-encoders to learn an effective latent representation for both global and local shape reconstruction based on the multi-task learning approach. Furthermore, we add global and local adversarial losses to generate more plausible point clouds. Quantitative and qualitative evaluations demonstrate that the proposed method outperforms state-of-the-art method on the task of 3D point cloud generation.
Kohei Matsuzaki, Kazuyuki Tasaka
IROS1
2018 A Compact Map Representation for Large-scale Environments and Localization Method based on Similarity Measure
abstract
Targeted at autonomous driving, compact map representation for localization is needed to deal with limited disk space or communication bandwidth. State-of-the-art compact map representations distinguish partial-objects such as lane markings from whole-object data. However, they require that objects are correctly detected in both map generation and localization. Therefore, the localization may fail if either of these conditions is violated due to occlusion, poor road texture, or other complications. In this paper, we propose a novel map generation and localization method that uses whole-object data without object detection. The novel map so produced has compactness comparable to state-of-the-art methods involving object detection because it is described as blocks of quantized representation in a voxel grid. In the localization step, we formulate localization as a similarity search between the map data and differently sampled sensor data. The proposed localization method performs the search at a sub-voxel level while alleviating voxel discretization error. This enables accurate localization even with low-resolution voxels. Experiments show the proposed method achieves sufficient localization accuracy and real-time processing with a map having a data size of 32 kB per km. The storage efficiency is at least 161 times greater than state-of-the-art maps covering the same area.
Kohei Matsuzaki, Hiromasa Yanagihara
Intelligent Vehicles Symposium1
2016 Geometric verification using semi-2D constraints for 3D object retrieval
abstract
Geometric verification with epipolar geometry often results in a high score for an incorrect image pair due to ambiguity in its geometric constraints. The ambiguity is caused by a high degree of freedom in the epipolar geometry and a weak constraint from the fitting between a point and a line. In order to mitigate the ambiguity, we propose to filter geometrically inconsistent components, namely correspondences, a sample, a model, and inliers in a RANSAC-based geometric verification. For the filtering, we introduce novel semi-2D constraints whose geometric constraint is weaker than full-2D constraint, but stronger than pure-epipolar constraint. Additionally, an advantage of the proposed approach is that it requires only an image pair instead of neither additional information nor prior learning. Experiments on the public dataset containing 3D object images show that the proposed approach improves the true positive rate when the false positive rate is low, and greatly reduces computational time for the geometric verification of both a correct image pair and an incorrect image pair.
Kohei Matsuzaki, Yusuke Uchida, Shigeyuki Sakazawa, Shin'ichi Satoh 0001
ICPR1