VLDB 2026 Research / reviewers in the wild / expert
Ryuichi Tanida
dblp:74/11276
· DBLP profile ↗
10ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-5379-3150ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Training-Free Photo-Realistic Point Cloud Rendering via Geometry-Aware Densification and Multi-view Refinement
Shogo Sato, Kazuhiko Murasaki, Ryuichi Tanida |
ICPR (2) | 3 |
| 2026 | IPCD: Intrinsic Point-Cloud DecompositionabstractPoint clouds are widely used in various fields, including augmented reality (AR) and robotics, where relighting and texture editing are crucial for realistic visualization. Achieving these tasks requires accurately separating albedo from shade. However, performing this separation on point clouds presents two key challenges: (1) the non-grid structure of point clouds makes conventional image-based decomposition models ineffective, and (2) point-cloud models designed for other tasks do not explicitly consider global-light direction, resulting in inaccurate shade. In this paper, we introduce Intrinsic Point-Cloud Decomposition (IPCD), which extends image decomposition to the direct decomposition of colored point clouds into albedo and shade. To overcome challenge (1), we propose IPCD-Net that extends image-based model with point-wise feature aggregation for non-grid data processing. For challenge (2), we introduce Projection-based Luminance Distribution (PLD) with a hierarchical feature refinement, capturing global-light ques via multi-view projection. For comprehensive evaluation, we create a synthetic outdoor-scene dataset. Experimental results demonstrate that IPCD-Net reduces cast shadows in albedo and enhances color accuracy in shade. Furthermore, we showcase its applications in texture editing, relighting, and point-cloud registration under varying illumination. Finally, we verify the real-world applicability of IPCD-Net. Shogo Sato, Takuhiro Kaneko, Shoichiro Takeda, Tomoyasu Shimada, Kazuhiko Murasaki, Taiga Yoshida, Ryuichi Tanida, Akisato Kimura |
WACV | 7 |
| 2025 | Objective, Absolute and Hue-Aware Metrics for Intrinsic Image Decomposition on Real-World Scenes: A Proof of ConceptabstractIntrinsic image decomposition (IID) is the task of separating an image into albedo and shade. In real-world scenes, it is difficult to quantitatively assess IID quality due to the unavailability of ground truth. The existing method provides the relative reflection intensities based on human-judged annotations. However, these annotations have challenges in subjectivity, relative evaluation, and hue non-assessment. To address these, we propose a concept of quantitative evaluation with a calculated albedo from a hyperspectral imaging and light detection and ranging (LiDAR) intensity. Additionally, we introduce an optional albedo densification approach based on spectral similarity. This paper conducted a concept verification in a laboratory environment, and suggested the feasibility of an objective, absolute, and hue-aware assessment.1 Shogo Sato, Masaru Tsuchida, Mariko Yamaguchi, Takuhiro Kaneko, Kazuhiko Murasaki, Taiga Yoshida, Ryuichi Tanida |
ICIP | 7 |
| 2025 | Unsupervised Single-Image Intrinsic Image Decomposition with LiDAR Intensity Enhanced TrainingabstractUnsupervised intrinsic image decomposition (IID) is the task of separating a natural image into albedo and shade without ground truth during training. Although a recent model employing light detection and ranging (LiDAR) intensity demonstrated impressive performance, the necessity of LiDAR intensity during inference restricts its practicality. To expand the usage scenario while maintaining the IID quality achieved by using both an image and its corresponding LiDAR intensity, we propose a novel approach that utilizes an image without LiDAR intensity during inference while utilizing both an image and LiDAR intensity during training. Specifically, our proposed model processes an image and LiDAR intensity individually using distinct encoder paths during training, but utilizes only an imageencoder path during inference. Additionally, we introduce an albedo-alignment loss aligning the gray-scale albedo from an image to that from its corresponding LiDAR intensity. LiDAR intensity is not affected by illumination effects including cast shadows, thus albedo-alignment loss transfers the illumination-invariant property of LiDAR intensity to the image-encoder path. Furthermore, we also propose image-LiDAR conversion (ILC) paths that mutually translates the style of an image and LiDAR intensity. IID models translate an image into albedo and shade styles while keeping the image contents, thus it is important to separate the image into contents and style. Trained with pairs of an image and its corresponding LiDAR intensity which share contents but differ in style, the mutual translation in ILC paths improve the accuracy of the separation. Consequently, our model achieves comparable IID quality to the existing model with LiDAR intensity, while utilizing only an image without LiDAR intensity during inference. Shogo Sato, Takuhiro Kaneko, Kazuhiko Murasaki, Taiga Yoshida, Ryuichi Tanida, Akisato Kimura |
WACV | 5 |
| 2024 | Memory-Efficient Point Cloud Registration via Overlapping Region SamplingabstractRecent advances in deep learning have improved 3D point cloud registration but increased graphics processing unit (GPU) memory usage, often requiring preliminary sampling that reduces accuracy. We propose an overlapping region sampling method to reduce memory usage while maintaining accuracy. Our approach estimates the overlapping region and intensively samples from it, using a k-nearest-neighbor (kNN) based point compression mechanism with multi layer perceptron (MLP) and transformer architectures. Evaluations on 3DMatch and 3DLoMatch datasets show our method outperforms other sampling methods in registration recall, especially at lower GPU memory levels. For 3DMatch, we achieve 94% recall with 33% reduced memory usage, with greater advantages in 3DLoMatch. Our method enables efficient large-scale point cloud registration in resource-constrained environments, maintaining high accuracy while significantly reducing memory requirements. Tomoyasu Shimada, Kazuhiko Murasaki, Shogo Sato, Toshihiko Nishimura, Taiga Yoshida, Ryuichi Tanida |
VCIP | 6 |
| 2023 | Distorted image classification using neural activation pattern matching lossabstractIn image classification, a deep neural network (DNN) that is trained on undistorted images constitutes an effective decision boundary. Unfortunately, this boundary does not support distorted images, such as noisy or blurry ones, leading to accuracy drop-off. As a simple approach for classifying distorted images as well as undistorted ones, previous methods have optimized the trained DNN again on both kinds of images. However, in these methods, the decision boundary may become overly complicated during optimization because there is no regularization of the decision boundary. Consequently, this decision boundary limits efficient optimization. In this paper, we study a simple yet effective decision boundary for distorted image classification through the use of a novel loss, called a "neural activation pattern matching (NAPM) loss". The NAPM loss is based on recent findings that the decision boundary is a piecewise linear function, where each linear segment is constructed from a neural activation pattern in the DNN when an image is fed to it. The NAPM loss extracts the neural activation patterns when the distorted image and its undistorted version are fed to the DNN and then matches them with each other via the sigmoid cross-entropy. Therefore, it constrains the DNN to classify the distorted image and its undistorted version by the same linear segment. As a result, our loss accelerates efficient optimization by preventing the decision boundary from becoming overly complicated. Our experiments demonstrate that our loss increases the accuracy of the previous methods in all conditions evaluated. Shoichiro Takeda, Ryuichi Tanida, Yukihiro Bandoh, Hayaru Shouno |
Neural Networks | 3 |
| 2022 | Deep Feature Compression Using Spatio-Temporal Arrangement Toward Collaborative Intelligent WorldabstractCollaborative Intelligence is a new paradigm that splits a deep neural network (DNN) into an edge and cloud for deploying a DNN-based image recognition application. In this paradigm, deep features, which are the outputs of the edge DNN, are compressed and transmitted to the cloud DNN. Because the deep features have a number of responses that are similar to each other, for efficient compression, previous methods spatially arrange and compress the deep features as an image to utilize the similarity as a spatial correlation. However, if the deep features are arranged in not only spatial but also temporal directions like those in a video, it may be possible to compress them more efficiently by increasing a temporal correlation. To explore this possibility, we propose a “spatio-temporal arrangement”. This method spatially arranges the deep features as images and temporally arranges them as a video with a novel ordering search algorithm. Our method effectively increases the spatial and temporal correlations hidden in the deep features and achieves high compression efficiency compared with the previous methods. Experimental results demonstrate the compression efficiency of our method is better than that of the previous methods (1.50% to 4.98% on BD-Rate evaluation in a lossy setting). Our analysis shows that our method effectively increases the correlation when the input is an image with rich edges and textures. Shoichiro Takeda, Motohiro Takagi, Ryuichi Tanida, Hideaki Kimata, Hayaru Shouno |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Knowledge Transferred Fine-Tuning for Anti-Aliased Convolutional Neural Network in Data-Limited SituationabstractAnti-aliased convolutional neural networks (CNNs) introduce blur filters to intermediate representations in CNNs to achieve high accuracy. A promising way to build a new antialiased CNN is to fine-tune a pre-trained CNN, which can easily be found online, with blur filters. However, blur filters drastically degrade the pre-trained representation, so the fine-tuning needs to rebuild the representation by using massive training data. Therefore, if the training data is limited, the fine-tuning cannot work well because it induces overfitting to the limited training data. To tackle this problem, this paper proposes “knowledge transferred fine-tuning”. On the basis of the idea of knowledge transfer, our method transfers the knowledge from intermediate representations in the pre-trained CNN to the anti-aliased CNN while fine-tuning. We transfer only essential knowledge using a pixel-level loss that transfers detailed knowledge and a global-level loss that transfers coarse knowledge. Experimental results demonstrate that our method significantly outperforms the simple fine-tuning method. Shoichiro Takeda, Ryuichi Tanida, Hideaki Kimata, Hayaru Shouno |
ICIP | 3 |
| 2020 | Deep Feature Compression With Spatio-Temporal Arranging for Collaborative IntelligenceabstractCollaborative Intelligence is a new paradigm that splits a deep neural network (DNN) into an edge DNN and a cloud DNN. In this paradigm, deep features, which are the outputs of the edge DNN, are compressed and transmitted to the cloud DNN. Since the deep features have a few responses that are similar to each other, previous studies have proposed compressing them as an image with a spatial arrangement to utilize spatial correlation between the deep features. However, this method may not sufficiently consider the similarity because only the spatial correlation is utilized. In this work, we propose a “spatio-temporal arranging” that considers the similarity of deep features in both the spatial and temporal directions. This method arranges the deep features spatiotemporally and compresses them as a video to utilize the spatio-temporal correlation. We perform spatial arranging as images to increase the spatial correlation and temporal arranging as a video with a novel ordering search to increase the temporal correlation. Experimental results demonstrate that our method performs better than previous methods. Motohiro Takagi, Shoichiro Takeda, Ryuichi Tanida, Hideaki Kimata |
ICIP | 4 |
| 2019 | GAN-based Image Compression Using Mutual Information Maximizing RegularizationabstractRecently, image compression systems based on convolutional neural networks that use flexible nonlinear analysis and synthesis transformations have been developed to improve the restoration accuracy of decoded images. A method using a framework called a generative adversarial network [1] has been reported as one of the methods aiming to improve the subjective image quality [2][3]. It optimizes the distribution of restored images to be close to that of natural images; thus it suppresses visual artifacts such as blurring, ringing, and blocking. However, since methods of this type are optimized to focus on whether the restored image is subjectively natural or not, components that are not correlated with the original image are mixed in the coding features obtained from the encoder. Thus, even though the appearance looks natural, it may be subjectively seen as a different object from the original image or the impression may be changed. In this paper, we describe a method we have developed to maximize mutual information between the coding features and the restored images. This method, which we call "regularization", makes it possible to develop image compression systems that suppress appearance differences with subjective naturalness. Shinobu Kudo, Shota Orihashi, Ryuichi Tanida, Atsushi Shimizu |
PCS | 3 |