Ken Sakurada

dblp:88/10003 · DBLP profile ↗
← Back
25ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0003-3386-1547ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 first-author · 5 since 2021Systems, architecture and hardware · 8 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2025 Low-Latency Privacy-Aware Robot Behavior guided by Automatically Generated Text Datasets
abstract
Humans typically avert their gaze when faced with situations involving another person’s privacy, and humanoid robots should exhibit similar behaviors. Various approaches exist for privacy recognition, including an image privacy recognition model and a Large Vision-Language Model (LVLM). The former relies on datasets of labeled images, which raise ethical concerns, while the latter requires more time to recognize images accurately, making real-time responses difficult. To this end, we propose a method of automatically constructing the LLM Privacy Text Dataset (LPT Dataset), a privacy-related text dataset with privacy indicators, and a method of recognizing whether observing a scene violates privacy without ethically sensitive training images. In constructing the LPT Dataset, which consists of both private and public scenes, we use an LLM to define privacy indicators and generate texts scored for each indicator. Our model recognizes whether a given image is private or public by retrieving texts with privacy scores similar to the image in a multi-modal feature space. In our experiments, we evaluated the performance of our model on three image privacy datasets and a realistic experiment with a humanoid robot in terms of accuracy and responsibility. The experiments show that our approach identifies the private image as accurately as the highly tuned LVLM without delay.
Yuta Irisawa, Tomoaki Yamazaki, Seiya Ito, Shuhei Kurita, Ryota Akasaka, Masaki Onishi, Kouzou Ohara, Ken Sakurada
IROS8
2024 G2fR: Frequency Regularization in Grid-Based Feature Encoding Neural Radiance Fields
Shuxiang Xie, Ken Sakurada, Ryoichi Ishikawa, Masaki Onishi, Takeshi Oishi
ECCV (22)3
2024 Wavefront Neural Radiance Fields for Multi-depth Reconstruction
Tsubasa Nakamura, Ken Sakurada, Gaku Nakano
ICPR (18)2
2024 Implicit Neural Fusion of RGB and Far-Infrared 3D Imagery for Invisible Scenes
abstract
Optical sensors, such as the Far Infrared (FIR) sensor, have demonstrated advantages over traditional imaging. For example, 3D reconstruction in the FIR field captures the heat distribution of a scene that is invisible to RGB, aiding various applications like gas leak detection. However, less texture information and challenges in acquiring FIR frames hinder the reconstruction process. Given that implicit neural representations (INRs) can integrate geometric information across different sensors, we propose Implicit Neural Fusion (INF) of RGB and FIR for 3D reconstruction of invisible scenes in the FIR field. Our method first obtains a neural density field of objects from RGB frames. Then, with the trained object density field, a separate neural density field of gases is optimized using limited view inputs of FIR frames. Our method not only demonstrates outstanding reconstruction quality in the FIR field through extensive experiments but also can isolate the geometric information of the invisible, offering a new dimension of scene understanding.
Xiangjie Li, Shuxiang Xie, Ken Sakurada, Ryusuke Sagawa, Takeshi Oishi
IROS3
2024 SDFT: Structural Discrete Fourier Transform for Place Recognition and Traversability Analysis
abstract
The ability to associate the current location with previously visited places is an essential aspect of autonomous ground robots. Unstructured environments such as planetary surfaces pose a significant challenge for robots because their terrain is less distinctive. Meanwhile, traversability must be analyzed simultaneously for safe navigation. In the past, place recognition research has rarely considered traversability analysis despite its significance. This is because the structural information of terrains becomes quickly implicit during the encoding process. This paper provides a method that explicitly addresses both problems: place recognition and traversability analysis. It proposes a discrete Fourier transform (DFT) to represent the frequency components embedded in ground curvature, which underlies both concepts. Our place recognition function demonstrates excellent performance in extensive experiments using challenging planetary & urban datasets while estimating traversability that other approaches find difficult to handle.
Ayumi Umemura, Ken Sakurada, Masaki Onishi, Kazuya Yoshida
IROS2
2023 Hierarchical Neural Memory Network for Low Latency Event Processing
abstract
This paper proposes a low latency neural network architecture for event-based dense prediction tasks. Conventional architectures encode entire scene contents at a fixed rate regardless of their temporal characteristics. Instead, the proposed network encodes contents at a proper temporal scale depending on its movement speed. We achieve this by constructing temporal hierarchy using stacked latent memories that operate at different rates. Given low latency event steams, the multi-level memories gradually extract dynamic to static scene contents by propagating information from the fast to the slow memory modules. The architecture not only reduces the redundancy of conventional architectures but also exploits long-term dependencies. Further-more, an attention-based event representation efficiently encodes sparse event streams into the memory cells. We conduct extensive evaluations on three event-based dense prediction tasks, where the proposed approach outperforms the existing methods on accuracy and latency, while demonstrating effective event and image fusion capabilities. The code is available at https://hamarh.github.io/hmnet/.
Ryuhei Hamaguchi, Yasutaka Furukawa, Masaki Onishi, Ken Sakurada
CVPR4
2023 INF: Implicit Neural Fusion for LiDAR and Camera
abstract
Sensor fusion has become a popular topic in robotics. However, conventional fusion methods encounter many difficulties, such as data representation differences, sensor variations, and extrinsic calibration. For example, the calibration methods used for LiDAR-camera fusion often require manual operation and auxiliary calibration targets. Implicit neural representations (INRs) have been developed for 3D scenes, and the volume density distribution involved in an INR unifies the scene information obtained by different types of sensors. Therefore, we propose implicit neural fusion (INF) for LiDAR and camera. INF first trains a neural density field of the target scene using LiDAR frames. Then, a separate neural color field is trained using camera images and the trained neural density field. Along with the training process, INF both estimates LiDAR poses and optimizes extrinsic parameters. Our experiments demonstrate the high accuracy and stable performance of the proposed method.
Shuxiang Xie, Ryoichi Ishikawa, Ken Sakurada, Masaki Onishi, Takeshi Oishi
IROS4
2022 LB-NERF: Light Bending Neural Radiance Fields for Transparent Medium
abstract
Neural radiance fields (NeRFs) have been proposed as methods of novel view synthesis and have been used to address various problems because of its versatility. NeRF can represent colors and densities in 3D space using neural rendering assuming a straight light path. However, a medium with a different refractive index in the scene, such as a transparent medium, causes light refraction and breaks the assumption of the straight path of light. Therefore, the NeRFs cannot be learned consistently across multi-view images. To solve this problem, this study proposes a method to learn consistent radiance fields across multiple viewpoints by introducing the light refraction effect as an offset from the straight line originating from the camera center. The experimental results quantitatively and qualitatively verified that our method can interpolate viewpoints better than the conventional NeRF method when considering the refraction of transparent objects.
Taku Fujitomi, Ken Sakurada, Ryuhei Hamaguchi, Hidehiko Shishido, Masaki Onishi, Yoshinari Kameda
ICIP2
2022 Fast Structural Representation and Structure-aware Loop Closing for Visual SLAM
abstract
Perceptual Aliasing is one of the main problems in simultaneous localization and mapping (SLAM). Wrong associations between different places may lead to failure of the whole map. Research on structure information is rarely investigated among existing solutions to this problem. In cases of visual SLAM without sensors, such as LiDAR or Inertial Measurement Unit (IMU), structure information can rarely be obtained due to the sparsity of 3D points, which also makes structure analysis complex. This study provides a spherical harmonics (SH) based fast structural representation (SH-FS) in visual SLAM using sparse point clouds, which extracts the structure information from sparse points into single vector. SH-FS was applied in conventional feature-based loop closing process. Furthermore, a structure-aware loop closing method in visual SLAM was proposed to improve the robustness of SLAM systems. Moreover, our methods show a favorable performance in extensive experiments on different large-scale real world datasets.
Shuxiang Xie, Ryoichi Ishikawa, Ken Sakurada, Masaki Onishi, Takeshi Oishi
IROS3
2021 Heterogeneous Grid Convolution for Adaptive, Efficient, and Controllable Computation
abstract
This paper proposes a novel heterogeneous grid convolution that builds a graph-based image representation by exploiting heterogeneity in the image content, enabling adaptive, efficient, and controllable computations in a convolutional architecture. More concretely, the approach builds a data-adaptive graph structure from a convolutional layer by a differentiable clustering method, pools features to the graph, performs a novel direction-aware graph convolution, and unpool features back to the convolutional layer. By using the developed module, the paper proposes heterogeneous grid convolutional networks, highly efficient yet strong extension of existing architectures. We have evaluated the proposed approach on four image understanding tasks, semantic segmentation, object localization, road extraction, and salient object detection. The proposed method is effective on three of the four tasks. Especially, the method outperforms a strong baseline with more than 90% reduction in floating-point operations for semantic segmentation, and achieves the state-of-the-art result for road extraction. We will share our code, model, and data.
Ryuhei Hamaguchi, Yasutaka Furukawa, Masaki Onishi, Ken Sakurada
CVPR4
2020 Privacy Preserving Visual SLAM
Mikiya Shibuya, Shinya Sumikura, Ken Sakurada
ECCV (22)3
2020 Benchmarking Cameras for Open VSLAM Indoors
abstract
In this paper we benchmark different types of cameras and evaluate their performance in terms of reliable localization reliability and precision in Visual Simultaneous Localization and Mapping (vSLAM). Such benchmarking is merely found for visual odometry, but never for vSLAM. Existing studies usually compare several algorithms for a given camera. The evaluation methodology we propose is applied to the recent OpenVSLAM framework. The latter is versatile enough to natively deal with perspective, fisheye, 360 cameras in a monocular or stereoscopic setup, an in RGB or RGB-D modalities. Results in various sequences containing light variation and scenery modifications in the scene assess quantitatively the maximum localization rate for 360 vision. In the contrary, RGB-D vision shows the lowest localization rate, but highest precision when localization is possible. Stereo-fisheye trades-off with localization rates and precision between 360 vision and RGB-D vision. The dataset with ground truth will be made available in open access to allow evaluating other/future vSLAM algorithms with respect to these camera types.
Kevin Chappellet, Guillaume Caron, Fumio Kanehiro, Ken Sakurada, Abderrahmane Kheddar
ICPR4
2020 Weakly Supervised Silhouette-based Semantic Scene Change Detection
abstract
This paper presents a novel semantic scene change detection scheme with only weak supervision. A straightforward approach for this task is to train a semantic change detection network directly from a large-scale dataset in an end-to-end manner. However, a specific dataset for this task, which is usually labor-intensive and time-consuming, becomes indispensable. To avoid this problem, we propose to train this kind of network from existing datasets by dividing this task into change detection and semantic extraction. On the other hand, the difference in camera viewpoints, for example, images of the same scene captured from a vehicle-mounted camera at different time points, usually brings a challenge to the change detection task. To address this challenge, we propose a new siamese network structure with the introduction of correlation layer. In addition, we create a publicly available dataset for semantic change detection to evaluate the proposed method. The experimental results verified both the robustness to viewpoint difference in change detection task and the effectiveness for semantic change detection of the proposed networks. Our code and dataset are available at https://github.com/xdspacelab/sscdnet.
Ken Sakurada, Mikiya Shibuya, Weimin Wang 0007
ICRA1
2020 GAN-Based SAR-to-Optical Image Translation with Region Information
abstract
In this paper, we propose a SAR-to-optical image translation method based on conditional generative adversarial networks (cGANs). Though cGANs have achieved great success in image translation, some problems remain in SAR-to-optical image translation. One of the problems is the colorization error owing to the lack of color information in SAR data. Since the colors of optical images are varied, while SAR images have no color information, the generator network is confused and fail to generate correctly colorized optical images. To prevent it, we introduce a region information to the image translation network. Specifically, the feature vector from the pre-trained classification network is fed to the generator and discriminator network. Experimental results with SEN1-2 dataset show the advantage of our proposed method over the baseline method that does not use any additional information.
Kento Doi, Ken Sakurada, Masaki Onishi, Akira Iwasaki
IGARSS2
2020 Self-supervised Simultaneous Alignment and Change Detection
abstract
This study proposes a self-supervised method for detecting scene changes from an image pair. For mobile cameras such as drive recorders, to alleviate the camera viewpoints' difference, image alignment and change detection must be optimized simultaneously because they depend on each other. Moreover, lighting condition makes the scene change detection more difficult because it widely varies in images taken at different times. To solve these challenges, we propose a self-supervised simultaneous alignment and change detection net-work (SACD-Net). The proposed network is robust specifically in differences of camera viewpoints and lighting conditions to simultaneously estimate warping parameters and multi-scale change probability maps while change regions are not taken into account of calculation of the feature consistency and semantic losses. Based on comparative analysis between our self-supervised and the previous supervised models as well as ablation study of the losses of SACD-Net, the results show the effectiveness of the proposed method using a synthetic dataset and our new real dataset.
Yukuko Furukawa, Kumiko Suzuki, Ryuhei Hamaguchi, Masaki Onishi, Ken Sakurada
IROS5
2019 Rare Event Detection Using Disentangled Representation Learning
abstract
This paper presents a novel method for rare event detection from an image pair with class-imbalanced datasets. A straightforward approach for event detection tasks is to train a detection network from a large-scale dataset in an end-to-end manner. However, in many applications such as building change detection on satellite images, few positive samples are available for the training. Moreover, an image pair of scenes contains many trivial events, such as in illumination changes or background motions. These many trivial events and the class imbalance problem lead to false alarms for rare event detection. In order to overcome these difficulties, we propose a novel method to learn disentangled representations from only low-cost negative samples. The proposed method disentangles the different aspects in a pair of observations: variant and invariant factors that represent trivial events and image contents, respectively. The effectiveness of the proposed approach is verified by the quantitative evaluations on four change detection datasets, and the qualitative analysis shows that the proposed method can acquire the representations that disentangle rare events from trivial ones.
Ryuhei Hamaguchi, Ken Sakurada, Ryosuke Nakamura
CVPR2
2019 OpenVSLAM: A Versatile Visual SLAM Framework
abstract
In this paper, we introduce OpenVSLAM, a visual SLAM framework with high usability and extensibility. Visual SLAM systems are essential for AR devices, autonomous control of robots and drones, etc. However, conventional open-source visual SLAM frameworks are not appropriately designed as libraries called from third-party programs. To overcome this situation, we have developed a novel visual SLAM framework. This software is designed to be easily used and extended. It incorporates several useful features and functions for research and development. OpenVSLAM is released at https://github.com/xdspacelab/openvslam under the 2-clause BSD license.
Shinya Sumikura, Mikiya Shibuya, Ken Sakurada
ACM Multimedia3
2018 Scale Estimation of Monocular SfM for a Multi-modal Stereo Camera
Shinya Sumikura, Ken Sakurada, Nobuo Kawaguchi, Ryosuke Nakamura
ACCV (3)2
2018 Image Translation Between Sar and Optical Imagery with Generative Adversarial Nets
abstract
In this paper, we propose a method for the translation from Synthetic Aperture Radar (SAR) to optical images using conditional Generative Adversarial Networks (cGANs). Satellite images have been widely utilized for various purposes, such as natural environment monitoring (pollution, forest or rivers), transportation improvement and prompt emergency response to disasters. However, the obscurity caused by clouds leads to unstable monitoring of the ground situation while using the optical camera. Images captured by a longer wavelength are introduced to reduce the effects of clouds. In particular, SAR images are known to be nearly unaffected by clouds and are often used for stably observing the ground situation. On the other hand, SAR images have lower spatial resolution and visibility than optical images. Therefore, we propose a deep neural network that generates optical images from SAR images. Finally, we confirm the feasibility of the proposed network on a dataset consisting of optical images and the corresponding SAR images.
Kenji Enomoto, Ken Sakurada, Nobuo Kawaguchi, Masashi Matsuoka, Ryosuke Nakamura
IGARSS2
2017 Temporal city modeling using street level imagery
Ken Sakurada, Daiki Tetsuka, Takayuki Okatani
Comput. Vis. Image Underst.1
2016 Hybrid macro-micro visual analysis for city-scale state estimation
abstract
We address the task of estimating large-scale land surface conditions using overhead aerial (macro-level) images and street view (micro-level) images. These two types of images are captured from orthogonal viewpoints and have different resolutions, thus conveying very different types of information that can be used in a complementary way. Moreover, their integration is necessary to enable an accurate understanding of changes in natural phenomena over massive city-scale landscapes. The key technical challenge is devising a method to integrate these two disparate types of image data in an effective manner, to leverage the wide coverage capabilities of macro-level images and detailed resolution of micro-level images. The strategy proposed in this work uses macro-level imaging to learn the extent to which the land condition corresponds between land regions that share similar visual characteristics (e.g., mountains, streets, buildings, rivers), whereas micro-level images are used to acquire high resolution statistics of land conditions (e.g., the amount of debris on the ground). By combining macro- and micro-level information about regional correspondences and surface conditions, our proposed method is capable of generating detailed estimates of land surface conditions over an entire city.
Ken Sakurada, Takayuki Okatani, Kris Makoto Kitani
Comput. Vis. Image Underst.1
2015 Change Detection from a Street Image Pair using CNN Features and Superpixel Segmentation
Ken Sakurada, Takayuki Okatani
BMVC1
2014 Massive City-Scale Surface Condition Analysis Using Ground and Aerial Imagery
Ken Sakurada, Takayuki Okatani, Kris Makoto Kitani
ACCV (1)1
2013 Detecting Changes in 3D Structure of a Scene from Multi-view Images Captured by a Vehicle-Mounted Camera
abstract
This paper proposes a method for detecting temporal changes of the three-dimensional structure of an outdoor scene from its multi-view images captured at two separate times. For the images, we consider those captured by a camera mounted on a vehicle running in a city street. The method estimates scene structures probabilistically, not deterministically, and based on their estimates, it evaluates the probability of structural changes in the scene, where the inputs are the similarity of the local image patches among the multi-view images. The aim of the probabilistic treatment is to maximize the accuracy of change detection, behind which there is our conjecture that although it is difficult to estimate the scene structures deterministically, it should be easier to detect their changes. The proposed method is compared with the methods that use multi-view stereo (MVS) to reconstruct the scene structures of the two time points and then differentiate them to detect changes. The experimental results show that the proposed method outperforms such MVS-based methods.
Ken Sakurada, Takayuki Okatani, Koichiro Deguchi
CVPR1
2010 Development of motion model and position correction method using terrain information for tracked vehicles with sub-tracks
abstract
Gyro-based odometry is an easy-to-use localization method for tracked vehicles because it uses only internal sensors. However, on account of track-terrain slippage and transformation caused by changes in sub-track angles, gyro-based odometry for tracked vehicles with sub-tracks experiences difficulties in estimating the exact location of the vehicles. In order to solve this problem, we propose an estimation method with 6 degrees of freedom (DOF) for determining the position and pose of the tracked vehicles using terrain information. (In this study, position refers to the robot's position and pose.) In the proposed method, position are estimated using a particle filter. The subsequent position of each particle are predicted using a motion model that separately considers each contact point of the vehicle with the ground. In addition, each particle is analyzed using terrain and gravity information. Experimental results demonstrate the effectiveness of this method.
Ken Sakurada, Eijiro Takeuchi, Kazunori Ohno, Satoshi Tadokoro
IROS1