Naoya Chiba

dblp:202/5745 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0003-3332-4426ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HybridSphere: Enhancing Hybrid Meetings with Avatar-Based VR Environments
abstract
With recent advances in information and communication technologies, Hybrid meetings, where local attendees are physically present and remote participants join virtually, have become increasingly common. However, remote participants often experience reduced contextual awareness and a sense of isolation. To address these issues, we propose HybridSphere, a hybrid meeting system that reconstructs a shared virtual reality environment for remote participants. The system employs a 360-degree camera and pose estimation to generate real-time avatar representations of local attendees to allow remote users with head-mounted displays to experience the meeting as if all participants are in the same virtual space. We conducted a user study comparing HybridSphere to a baseline condition in which remote participants viewed an unmodified 360-degree video. Although the avatars were rated lower in perceived trustworthiness and likability due to limited visual fidelity, participants appreciated the seated-6DoF function. These results suggest that immersive and spatially flexible VR representations can enhance remote engagement in hybrid meetings.
Koji Momota, Shizuka Shirai, Masato Kobayashi 0001, Naoya Chiba, Photchara Ratsamee, Kiyoshi Kiyokawa, Yuuki Uranishi
IEEE Trans. Vis. Comput. Graph.4
2025 Evaluating Data Quality and Preprocessing Methods to Enhance Skeleton-Based Action Recognition in Retail Environments
abstract
The accuracy of skeleton-based action recognition models can be significantly improved using data processing techniques, particularly in complicated environments such as retail stores where people’s activities are dynamic and prone to noise and occlusions. This paper investigates how various preprocessing techniques, including smoothing, data compensation, and noise augmentation, influence the accuracy of two action recognition models, STGCN and STGCN++. We evaluated these methods using different data sources: key points of YOLO, HRNet, and motion capture (MoCap). Our findings demonstrate that higher quality data, such as projected 2D MoCap, lead to significantly better model performance, with STGCN++ achieving a mean accuracy of $91.26 \%$, compared to $81.06 \%$ with YOLO key points. We also show that applying preprocessing methods to noisier data, such as YOLO, can reduce performance degradation, improving their robustness for real-world applications. These results highlight the importance of preprocessing in implementing robust action recognition systems that can operate in dynamic, real-world settings and open the way for further improvements in the field.
Samer Yousef, Chengjun Han, Naoya Chiba, Koichi Hashimoto
FG3
2025 Focused Blind Switching Manipulation Based on Constrained and Regional Touch States of Multi-Fingered Hand Using Deep Learning
abstract
To achieve a desired grasping posture (including object position and orientation), multi-finger motions need to be conducted according to the the current touch state. Specifically, when subtle changes happen during correcting the object state, not only proprioception but also tactile information from the entire hand can be beneficial. However, switching motions with high-DOFs of multiple fingers and abundant tactile information is still challenging. In this study, we propose a loss function with constraints of touch states and an attention mechanism for focusing on important modalities depending on the touch states. The policy model is AE-LSTM which consists of Autoencoder (AE) which compresses abundant tactile information and Long Short-Term Memory (LSTM) which switches the motion depending on the touch states. Motion for cap-opening was chosen as a target task which consists of sub tasks of sliding an object and opening its cap. As a result, the proposed method achieved the best success rates with a variety of objects for real time cap-opening manipulation. Furthermore, we could confirm that the proposed model acquired the features of each subtask and attention on specific modalities.
Satoshi Funabashi, Atsumu Hiramoto, Naoya Chiba, Alexander Schmitz, Shardul Kulkarni, Tetsuya Ogata
ICRA3
2025 MRHaD: Mixed Reality-based Hand-Drawn Map Editing Interface for Mobile Robot Navigation
abstract
Mobile robot navigation systems are increasingly relied upon in dynamic and complex environments, yet they often struggle with map inaccuracies and the resulting inefficient path planning. This paper presents MRHaD, a Mixed Reality-based Hand-drawn Map Editing Interface that enables intuitive, real-time map modifications through natural hand gestures. By integrating the MR head-mounted display with the robotic navigation system, operators can directly create hand-drawn restricted zones (HRZ), thereby bridging the gap between 2D map representations and the real-world environment. Comparative experiments against conventional 2D editing methods demonstrate that MRHaD significantly improves editing efficiency, map accuracy, and overall usability, contributing to safer and more efficient mobile robot operations. The proposed approach provides a robust technical foundation for advancing human-robot collaboration and establishing innovative interaction models that enhance the hybrid future of robotics and human society. For additional material, please check: https://mertcookimg.github.io/mrhad/
Takumi Taki, Masato Kobayashi 0001, Eduardo Iglesius, Naoya Chiba, Shizuka Shirai, Yuuki Uranishi
RO-MAN4
2024 Learning 3D Point Cloud Registration as a Single Optimization Problem
Rintaro Yanagi, Atsushi Hashimoto 0001, Naoya Chiba, Shusaku Sone, Yoshitaka Ushiku
ACCV (9)3
2024 Crystalformer: Infinitely Connected Attention for Periodic Structure Encoding
abstract
Predicting physical properties of materials from their crystal structures is a fundamental problem in materials science. In peripheral areas such as the prediction of molecular properties, fully connected attention networks have been shown to be successful. However, unlike these finite atom arrangements, crystal structures are infinitely repeating, periodic arrangements of atoms, whose fully connected attention results in *infinitely connected attention*. In this work, we show that this infinitely connected attention can lead to a computationally tractable formulation, interpreted as *neural potential summation*, that performs infinite interatomic potential summations in a deeply learned feature space. We then propose a simple yet effective Transformer-based encoder architecture for crystal structures called *Crystalformer*. Compared to an existing Transformer-based model, the proposed model requires only 29.4% of the number of parameters, with minimal modifications to the original Transformer architecture. Despite the architectural simplicity, the proposed method outperforms state-of-the-art methods for various property regression tasks on the Materials Project and JARVIS-DFT datasets.
Tatsunori Taniai, Ryo Igarashi 0002, Naoya Chiba, Kotaro Saito, Yoshitaka Ushiku, Kanta Ono
ICLR4
2024 NeuralLabeling: A versatile toolset for labeling vision datasets using Neural Radiance Fields
abstract
We present NeuralLabeling, a labeling approach and toolset for annotating 3D scenes using either bounding boxes or meshes and generating segmentation masks, affordance maps, 2D bounding boxes, 3D bounding boxes, 6DOF object poses, depth maps, and object meshes. NeuralLabeling uses Neural Radiance Fields (NeRF) as a renderer, allowing labeling to be performed using 3D spatial tools while incorporating geometric clues such as occlusions, relying only on images captured from multiple viewpoints as input. To demonstrate the applicability of NeuralLabeling to a practical problem in robotics, we added ground truth depth maps to 30000 frames of transparent object RGB and noisy depth maps of glasses placed in a dishwasher captured using an RGBD sensor, yielding the Dishwasher30k dataset. We show that training a simple deep neural network with supervision using the annotated depth maps yields a higher reconstruction performance than training with the previously applied weakly supervised approach. We also show how instance segmentation and depth completion datasets generated using NeuralLabeling can be incorporated into a robot application for grasping transparent objects placed in a dishwasher with an accuracy of 83.3%, compared to 16.3% without depth completion. Supplementary URI: https://florise.github.io/neural_labeling_web/.
Floris Erich, Naoya Chiba, Abdullah Mustafa, Yusuke Yoshiyasu, Noriaki Ando, Ryo Hanai, Yukiyasu Domae
IROS2
2023 Multi-Timestep-Ahead Prediction with Mixture of Experts for Embodied Question Answering
Kanata Suzuki, Yuya Kamiwano, Naoya Chiba, Hiroki Mori, Tetsuya Ogata
ICANN (6)3
2023 Reference-based Dense Pose Estimation via Partial 3D Point Cloud Matching
abstract
Interacting with real-world objects is one of the fundamental tasks in multimedia. Despite its importance, existing object pose estimation targets only rigid objects. This demonstration proposes a novel application for non-rigid object pose estimation. Inspired by human dense pose estimation, we represent a pose of a non-rigid object as an indexed point cloud, where each index corresponds to that in a template. The correspondence is identified by a machine-learning-based 3D point cloud matching. Finding correspondence to the template point cloud enables a dense pose estimation with no object-specific learning processes. In the demonstration, we visualize the correspondence of points in observed depth images and the template. We also provide a demonstration of template point cloud reconstruction. Through these systems, onsite visitors can test our system with objects brought by themselves and have an experience with a state-of-the-art 3D point cloud matching method as well as this novel task.
Rintaro Yanagi, Atsushi Hashimoto 0001, Naoya Chiba, Yoshitaka Ushiku
ACM Multimedia3
2022 Point Cloud Pre-training with Natural 3D Structures
abstract
The construction of 3D point cloud datasets requires a great deal of human effort. Therefore, constructing a large-scale 3D point clouds dataset is difficult. In order to rem-edy this issue, we propose a newly developed point cloud fractal database (PC-FractalDB), which is a novel family of formula-driven supervised learning inspired by fractal geometry encountered in natural 3D structures. Our re-search is based on the hypothesis that we could learn rep-resentations from more real-world 3D patterns than con-ventional 3D datasets by learning fractal geometry. We show how the PC-FractalDB facilitates solving several re-cent dataset-related problems in 3D scene understanding, such as 3D model collection and labor-intensive annotation. The experimental section shows how we achieved the performance rate of up to 61.9% and 59.0% for the Scan-NetV2 and SUN RGB-D datasets, respectively, over the current highest scores obtained with the PointContrast, con-trastive scene contexts (CSC), and RandomRooms. More-over, the PC-FractalDB pre-trained model is especially ef-fective in training with limited data. For example, in 10% of training data on ScanNetV2, the PC-FractalDB pre-trained VoteNet performs at 38.3%, which is +14.8% higher accu-racy than CSC. Of particular note, we found that the pro-posed method achieves the highest results for 3D object de-tection pre-training in limited point cloud data.11Dataset release: https://ryosuke-yamada.github.io/PointCloud-FractalDataBase/
Ryosuke Yamada, Hirokatsu Kataoka, Naoya Chiba, Yukiyasu Domae, Tetsuya Ogata
CVPR3
2018 Sparse Estimation of Light Transport Matrix under Saturated Condition
Naoya Chiba, Koichi Hashimoto
BMVC1
2018 Ultra-Fast Multi-Scale Shape Estimation of Light Transport Matrix for Complex Light Reflection Objects
abstract
3D measurement of target objects characterized by specular reflection or subsurface scatterings cannot be measured by traditional 3D measurement methods because these targets have multiple light paths that make it difficult to determine the unique surface. We define these objects as complex light reflection objects. In this case, 3D measurement methods based on Light Transport (LT) Matrix estimation may be a solution to measure these complex light reflection objects, because LT Matrix captures every light path, and we can identify all 3D points on the target shape by using LT Matrix. However, these methods either provide low resolution results, or they are too slow for use in robot vision in practice. In this paper, we suppress the computational cost of LT Matrix estimation by dividing LT Matrix estimation into multi-scale. The proposed method reduces the number of candidate combinations between camera pixels and projector pixels greatly by using the information given by low resolution observations. The proposed algorithm allows high resolution measurement of the LT Matrix very efficiently. Furthermore, careful implementation of our method by using a sparse matrix representation achieves memory efficiency. We evaluated our method by measuring 3D points for a 256 × 256 resolution projector and camera system, which is an LT matrix 4096 times larger than that developed in our previous study [1] and 100 times faster than our naïve implementation of [2].
Naoya Chiba, Koichi Hashimoto
ICRA1