Kazuto Nakashima

dblp:156/3026 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-6773-7811ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Systems, architecture and hardware · 5 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Fast LiDAR Data Generation with Rectified Flows
abstract
Building LiDAR generative models holds promise as powerful data priors for restoration, scene manipulation, and scalable simulation in autonomous mobile robots. In recent years, approaches using diffusion models have emerged, significantly improving training stability and generation quality. Despite their success, diffusion models require numerous iterations of running neural networks to generate high-quality samples, making the increasing computational cost a potential barrier for robotics applications. To address this challenge, this paper presents R2Flow, a fast and high-fidelity generative model for LiDAR data. Our method is based on rectified flows that learn straight trajectories, simulating data generation with significantly fewer sampling steps compared to diffusion models. We also propose an efficient Transformer-based model architecture for processing the image representation of LiDAR range and reflectance measurements. Our experiments on unconditional LiDAR data generation using the KITTI-360 dataset demonstrate the effectiveness of our approach in terms of both efficiency and quality.
Kazuto Nakashima, Tomoya Miyawaki, Yumi Iwashita, Ryo Kurazume
ICRA1
2025 Enhancing the Quality of 3D Lunar Maps Using JAXA's Kaguya Imagery
abstract
As global efforts to explore the Moon intensify, the need for high-quality 3D lunar maps becomes increasingly critical—particularly for long-distance missions such as NASA’s Endurance mission concept, in which a rover aims to traverse 2,000 km across the South Pole–Aitken basin. Kaguya TC (Terrain Camera) images, though globally available at 10 m/pixel, suffer from altitude inaccuracies caused by stereo matching errors and JPEG-based compression artifacts. This paper presents a method to improve the quality of 3D maps generated from Kaguya TC images, focusing on mitigating the effects of compression-induced noise in disparity maps. We analyze the compression behavior of Kaguya TC imagery, and identify systematic disparity noise patterns, especially in darker regions. In this paper, we propose an approach to enhance 3D map quality by reducing residual noise in disparity images derived from compressed images. Our experimental results show that the proposed approach effectively reduces elevation noise, enhancing the safety and reliability of terrain data for future lunar missions.
Yumi Iwashita, Haakon Moe, Adnan Ansar, Georgios Georgakis, Adrian Stoica, Kazuto Nakashima, Ryo Kurazume, Jim Tørresen
SMC7
2024 LiDAR Data Synthesis with Denoising Diffusion Probabilistic Models
abstract
Generative modeling of 3D LiDAR data is an emerging task with promising applications for autonomous mobile robots, such as scalable simulation, scene manipulation, and sparse-to-dense completion of LiDAR point clouds. While existing approaches have demonstrated the feasibility of image-based LiDAR data generation using deep generative models, they still struggle with fidelity and training stability. In this work, we present R2DM, a novel generative model for LiDAR data that can generate diverse and high-fidelity 3D scene point clouds based on the image representation of range and reflectance intensity. Our method is built upon denoising diffusion probabilistic models (DDPMs), which have shown impressive results among generative model frameworks in recent years. To effectively train DDPMs in the LiDAR domain, we first conduct an in-depth analysis of data representation, loss functions, and spatial inductive biases. Leveraging our R2DM model, we also introduce a flexible LiDAR completion pipeline based on the powerful capabilities of DDPMs. We demonstrate that our method surpasses existing methods in generating tasks on the KITTI-360 and KITTI-Raw datasets, as well as in the completion task on the KITTI-360 dataset. Our project page can be found at https://kazuto1011.github.io/r2dm.
Kazuto Nakashima, Ryo Kurazume
ICRA1
2024 Fast LiDAR Upsampling using Conditional Diffusion Models
abstract
The search for refining 3D LiDAR data has attracted growing interest motivated by recent techniques such as supervised learning or generative model-based methods. Existing approaches have shown the possibilities for using diffusion models to generate refined LiDAR data with high fidelity, although the performance and speed of such methods have been limited. These limitations make it difficult to execute in real-time, causing the approaches to struggle in real-world tasks such as autonomous navigation and human-robot interaction. In this work, we introduce a novel approach based on conditional diffusion models for fast and high-quality sparse-to-dense upsampling of 3D scene point clouds through an image representation. Our method employs denoising diffusion probabilistic models trained with conditional inpainting masks, which have been shown to give high performance on image completion tasks. We introduce a series of experiments, including multiple datasets, sampling steps, and conditional masks. This paper illustrates that our method outperforms the baselines in sampling speed and quality on upsampling tasks using the KITTI-360 dataset. Furthermore, we illustrate the generalization ability of our approach by simultaneously training on real-world and synthetic datasets, introducing variance in quality and environments.
Sander Elias Magnussen Helgesen, Kazuto Nakashima, Jim Tørresen, Ryo Kurazume
RO-MAN2
2024 Development of Dementia Care Training System Using AR and Large Language Model
abstract
We have developed HEARTS, a dementia care training system using augmented reality based on Humanitude. Humanitude is a multimodal comprehensive care technique for dementia, and has attracted attention as a method to reduce the burden on both caregivers and patients. However, the HEARTS developed so far could not evaluate “speaking” skills based on the content of conversations among “seeing,” “touching,” and “speaking,” all of which are fundamental skills in Humanitude. Therefore, we attempted a new quantitative evaluation of trainees' “speaking” skills by estimating the emotional value of conversational content using GPT-4 named HEARTS 5. A survey of caregivers was conducted using the developed system and was well received by the participants. We also developed a HEARTS 5 conversational version based on HEARTS 5, with the addition of GPT-4 conversation generation and Azure Text-to-Speech.
Tomoya Miyawaki, Yuki Nishiura, Ryouta Fukuda, Kazuto Nakashima, Ryo Kurazume
SMC4
2023 Generative Range Imaging for Learning Scene Priors of 3D LiDAR Data
abstract
3D LiDAR sensors are indispensable for the robust vision of autonomous mobile robots. However, deploying LiDAR-based perception algorithms often fails due to a domain gap from the training environment, such as inconsistent angular resolution and missing properties. Existing studies have tackled the issue by learning inter-domain mapping, while the transferability is constrained by the training configuration and the training is susceptible to peculiar lossy noises called ray-drop. To address the issue, this paper proposes a generative model of LiDAR range images applicable to the data-level domain transfer. Motivated by the fact that LiDAR measurement is based on point-by-point range imaging, we train an implicit image representation-based generative adversarial networks along with a differentiable ray-drop effect. We demonstrate the fidelity and diversity of our model in comparison with the point-based and image-based state-of-the-art generative models. We also showcase upsampling and restoration applications. Furthermore, we introduce a Sim2Real application for LiDAR semantic segmentation. We demonstrate that our method is effective as a realistic ray-drop simulator and outperforms state-of-the-art methods.
Kazuto Nakashima, Yumi Iwashita, Ryo Kurazume
WACV1
2022 Understanding Humanitude Care for Sit-to-stand Motion by Wearable Sensors
abstract
Assisting patients with dementia is a significant social issue. Currently, to assist patients with dementia, a multimodal care technique called Humanitude is gaining popularity. In Humanitude, the patients are assisted through various techniques to stand up independently by utilizing their motor functions as much as possible. Humanitude care techniques encourage caregivers to increase the area of contact with patients during the sit-to-stand motion. However, Humanitude care techniques are not accurately performed by novice caregivers. Therefore, in this study, a smock-type wearable sensor was developed to measure the proximity between caregivers and care recipients during sit-to-stand motion assistance. A measurement experiment was conducted to evaluate the proximity differences between Humanitude care and simulated novice care. In addition, the effects of different care techniques on the center of mass (CoM) trajectory and muscle activity of the care recipients were investigated. The results showed that the caregivers tend to bring their top and middle trunk closer in Humanitude care compared with novice simulated care. Furthermore, it was observed that the CoM trajectory and muscle activity under Humanitude care were similar to those observed when the care recipient stands up independently. These results validate the effectiveness of Humanitude care and provide useful information for teaching techniques in Humanitude.
Qi An 0001, Akito Tanaka, Kazuto Nakashima, Hidenobu Sumioka, Masahiro Shiomi, Ryo Kurazume
SMC3
2021 Learning to Drop Points for LiDAR Scan Synthesis
abstract
3D laser scanning by LiDAR sensors plays an important role for mobile robots to understand their surroundings. Nevertheless, not all systems have high resolution and accuracy due to hardware limitations, weather conditions, and so on. Generative modeling of LiDAR data as scene priors is one of the promising solutions to compensate for unreliable or incomplete observations. In this paper, we propose a novel generative model for learning LiDAR data based on generative adversarial networks. As in the related studies, we process LiDAR data as a compact yet lossless representation, a cylindrical depth map. However, despite the smoothness of real-world objects, many points on the depth map are dropped out through the laser measurement, which causes learning difficulty on generative models. To circumvent this issue, we introduce measurement uncertainty into the generation process, which allows the model to learn a disentangled representation of the underlying shape and the dropout noises from a collection of real LiDAR data. To simulate the lossy measurement, we adopt a differentiable sampling framework to drop points based on the learned uncertainty. We demonstrate the effectiveness of our method on synthesis and reconstruction tasks using two datasets. We further showcase potential applications by restoring LiDAR data with various types of corruption.
Kazuto Nakashima, Ryo Kurazume
IROS1
2018 Fourth-Person Captioning: Describing Daily Events by Uni-supervised and Tri-regularized Training
abstract
We aim to develop a supporting system which enhances the ability of human's short-term visual memory in an intelligent space where the human and a service robot coexist. Particularly, this paper focuses on how we can interpret and record diverse and complex life events on behalf of humans, from a multi-perspective viewpoint. We propose a novel method named "fourth-person captioning", which generates natural language descriptions by summarizing visual contexts complementarily from three types of cameras corresponding the first-, second-, and third-person viewpoint. We first extend the latest image captioning technique and design a new model to generate a sequence of words given the multiple images. Then we provide an effective training strategy that needs only annotations supervising images from a single viewpoint in a general caption dataset and unsupervised triplet instances in the intelligent space. As the three types of cameras, we select a wearable camera on the human, a robot-mounted camera, and an embedded camera, which can be defined as the first-, second-, and third-person viewpoint, respectively. We hope our work will accelerate a cross-modal interaction bridging the human's egocentric cognition and multi-perspective intelligence.
Kazuto Nakashima, Yumi Iwashita, Akihiro Kawamura, Ryo Kurazume
SMC1
2017 Feasibility study of IoRT platform "Big Sensor Box"
abstract
This paper proposes new software and hardware platforms named ROS-TMS and Big Sensor Box, respectively, for an informationally structured environment. We started the development of a management system for an informationally structured environment named Town Management System (TMS) in the Robot Town Project in 2005. Since then we have been continuing our efforts to improve performance and to enhance TMS functions. Recently, we launched a new version of TMS named ROS-TMS, which resolves some critical problems in TMS by adopting the Robot Operating System (ROS) and utilizing the high scalability and numerous resources of ROS. In this paper, we first discuss the structure of a software platform for the informationally structured environment and describe in detail our latest system, ROS-TMS version 4.0. Next, we introduce a hardware platform for the informationally structured environment named Big Sensor Box, in which a variety of sensors are embedded and service robots are operated according to the structured information under the management of ROS-TMS. Robot service experiments including a fetch-and-give task and autonomous control of a wheelchair robot are also conducted in Big Sensor Box.
Ryo Kurazume, YoonSeok Pyo, Kazuto Nakashima, Akihiro Kawamura, Tokuo Tsuji
ICRA3
2017 Previewed reality: Near-future perception system
abstract
This paper presents a near-future perception system named “Previewed Reality”. The system consists of an informationally structured environment (ISE), an immersive VR display, a stereo camera, an optical tracking system, and a dynamic simulator. In an ISE, a number of sensors are embedded, and information such as the position of furniture, objects, humans, and robots, is sensed and stored in a database. The position and orientation of the immersive VR display are also tracked by an optical tracking system. Therefore, we can forecast the next possible events using a dynamic simulator and synthesize virtual images of what users will see in the near future from their own viewpoint. The synthesized images, overlaid on a real scene by using augmented reality technology, are presented to the user. The proposed system can allow a human and a robot to coexist more safely by showing possible hazardous situations to the human intuitively in advance.
Yuta Horikawa, Asuka Egashira, Kazuto Nakashima, Akihiro Kawamura, Ryo Kurazume
IROS3