Yoshinobu Hagiwara

dblp:29/10948 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
6since 2021 · last 2025
0009-0006-1208-3159ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 since 2021Systems, architecture and hardware · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2025 Take That for Me: Multimodal Exophora Resolution with Interactive Questioning for Ambiguous Out-of-View Instructions
abstract
Daily life support robots must interpret ambiguous verbal instructions involving demonstratives such as "Bring me that cup," even when objects or users are out of the robot’s view. Existing approaches to exophora resolution primarily rely on visual data and thus fail in real-world scenarios where the object or user is not visible. We propose Multimodal Interactive Exophora resolution with user Localization (MIEL), which is a multimodal exophora resolution framework leveraging sound source localization (SSL), semantic mapping, visual-language models (VLMs), and interactive questioning with GPT-4o. Our approach first constructs a semantic map of the environment and estimates candidate objects from a linguistic query with the user’s skeletal data. SSL is utilized to orient the robot toward users who are initially outside its visual field, enabling accurate identification of user gestures and pointing directions. When ambiguities remain, the robot proactively interacts with the user, employing GPT-4o to formulate clarifying questions. Experiments in a real-world environment showed results that were approximately 1.3 times better when the user was visible to the robot and 2.0 times better when the user was not visible to the robot, compared to the methods without SSL and interactive questioning. The project website is https://emergentsystemlabstudent.github.io/MIEL/.
Akira Oyama, Shoichi Hasegawa, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi
RO-MAN4
2024 Object Instance Retrieval in Assistive Robotics: Leveraging Fine-Tuned SimSiam with Multi-View Images Based on 3D Semantic Map
abstract
Robots that assist humans in their daily lives should be able to locate specific instances of objects in an environment that match a user’s desired objects. This task is known as instance-specific image goal navigation (InstanceImageNav), which requires a model that can distinguish different instances of an object within the same class. A significant challenge in robotics is that when a robot observes the same object from various 3D viewpoints, its appearance may differ significantly, making it difficult to recognize and locate accurately. In this paper, we introduce a method called SimView, which leverages multi-view images based on a 3D semantic map of an environment and self-supervised learning using SimSiam to train an instance-identification model on-site. The effectiveness of our approach was validated using a photorealistic simulator, Habitat Matterport 3D, created by scanning actual home environments. Our results demonstrate a 1.7-fold improvement in task accuracy compared with contrastive language-image pre-training (CLIP), a pre-trained multimodal contrastive learning method for object searching. This improvement highlights the benefits of our proposed fine-tuning method in enhancing the performance of assistive robots in InstanceImageNav tasks. The project website is https://emergentsystemlabstudent.github.io/MultiViewRetrieve/.
Taichi Sakaguchi, Akira Taniguchi, Yoshinobu Hagiwara, Lotfi El Hafi, Shoichi Hasegawa, Tadahiro Taniguchi
IROS3
2024 Pointing Frame Estimation With Audio-Visual Time Series Data for Daily Life Service Robots
abstract
Daily life support robots in the home environment interpret the user's pointing and understand the instructions, thereby increasing the number of instructions accomplished. This study aims to improve the estimation performance of pointing frames by using speech information when a person gives pointing or verbal instructions to the robot. The estimation of the pointing frame, which represents the moment when the user points, can help the user understand the instructions. Therefore, we perform pointing frame estimation using a time-series model, utilizing the user's speech, images, and speech-recognized text observed by the robot. In our experiments, we set up realistic communication conditions, such as speech containing everyday conversation, non-upright posture, actions other than pointing, and reference objects outside the robot's field of view. The results showed that adding speech information improved the estimation performance, especially the Transformer model with Mel-Spectrogram as a feature. This study will lead to be applied to object localization and action planning in 3D environments by robots in the future. The project website is https://emergentsystemlabstudent.github.io/PointingImgEst/.
Hikaru Nakagawa, Shoichi Hasegawa, Yoshinobu Hagiwara, Akira Taniguchi, Tadahiro Taniguchi
SMC3
2023 Exophora Resolution of Linguistic Instructions with a Demonstrative based on Real-World Multimodal Information
abstract
To enable a robot to provide support in a home environment through human-robot interaction, exophora resolution is crucial for accurately identifying the target of ambiguous linguistic instructions, which may include a demonstrative, such as “Take that one”. Unlike endophora resolution, which involves predicting the corresponding word from given sentences, exophora resolution necessitates comprehensive utilization of external real-world information to identify and disambiguate the target from the on-site environment. This study aims to resolve ambiguity in language instructions containing a demonstrative through exophora resolution, utilizing real-world multimodal information. The robot accomplishes this by using three types of information: 1) object categories, 2) demonstratives, and 3) pointing, as well as knowledge about objects obtained from the robot’s pre-exploration of the environment. We evaluated the accuracy of object identification under multiple conditions by identifying a user-indicated object in a field that mimics a home environment. Our results demonstrate that our proposed method of exophora resolution using multimodal information can identify the target with two to three times higher accuracy than baseline methods in cases where information is missing.
Akira Oyama, Shoichi Hasegawa, Hikaru Nakagawa, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi
RO-MAN5
2023 Active Semantic Mapping for Household Robots: Rapid Indoor Adaptation and Reduced User Burden
abstract
Active semantic mapping is essential for service robots to quickly capture both the map of an environment and its spatial meaning, while also minimizing the burden on users during robot operation and data collection. SpCoSLAM, a method of semantic mapping with place categorization and simultaneous localization and mapping (SLAM), is well suited to environmental adaptation, as it is not limited to predefined labels. However, SpCoSLAM presents two issues that increase the burden on users: 1) users struggle to efficiently determine a destination for the robot's quick adaptation, and 2) providing instructions to the robot becomes repetitive and cumbersome. To address these challenges, we propose Active-SpCoSLAM, which enables the robot to actively explore uncharted areas and employs CLIP for image captioning to provide a flexible vocabulary that replaces human instructions. The robot determines its actions by calculating information gain integrated from both semantics and SLAM uncertainties. We conducted experiments in a simulated environment, comparing the proposed method to other methods in terms of efficiency and applicability to object discovery tasks. Additionally, we tested the proposed method, which combines user instruction and CLIP, in a real environment. Our results demonstrated that the robot explored its environment with approximately five fewer iterations and 11 minutes faster compared to the case of random exploration. Moreover, our method achieved a higher success rate in object discovery tasks during earlier stages of learning compared to other methods. In conclusion, the proposed method rapidly covers an environment while gathering useful data for object discovery tasks, thus reducing the burden on users and enhancing the robot's adaptability. The project website is https://tomochika-ishikawa.github.io/Active-SpCoSLAM/.
Tomochika Ishikawa, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi
SMC3
2022 Multimodal Object Categorization with Reduced User Load through Human-Robot Interaction in Mixed Reality
abstract
Enabling robots to learn from interactions with users is essential to perform service tasks. However, as a robot categorizes objects from multimodal information obtained by its sensors during interactive onsite teaching, the inferred names of unknown objects do not always match the human user's expectation, especially when the robot is introduced to new environments. Confirming the learning results through natural speech interaction with the robot often puts an additional burden on the user who can only listen to the robot to validate the results. Therefore, we propose a human-robot interface to reduce the burden on the user by visualizing the inferred results in mixed reality (MR). In particular, we evaluate the proposed interface on the system usability scale (SUS) and the NASA task load index (NASA-TLX) with three experimental object categorization scenarios based on multimodal latent Dirichlet allocation (MLDA) in which the robot: 1) does not share the inferred results with the user at all, 2) shares the inferred results through speech interaction with the user (baseline), and 3) shares the inferred results with the user through an MR interface (proposed). We show that providing feedback through an MR interface significantly reduces the temporal, physical, and mental burden on the human user compared to speech interaction with the robot.
Hitoshi Nakamura, Lotfi El Hafi, Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi
IROS4
2020 SpCoMapGAN: Spatial Concept Formation-based Semantic Mapping with Generative Adversarial Networks
abstract
In semantic mapping, which connects semantic information to an environment map, it is a challenging task for robots to deal with both local and global information of environments. In addition, it is important to estimate semantic information of unobserved areas from already acquired partial observations in a newly visited environment. On the other hand, previous studies on spatial concept formation enabled a robot to relate multiple words to places from bottom-up observations even when the vocabulary was not provided beforehand. However, the robot could not transfer global information related to the room arrangement between semantic maps from other environments. In this paper, we propose SpCoMapGAN, which generates the semantic map in a newly visited environment by training an inference model using previously estimated semantic maps. SpCoMapGAN uses generative adversarial networks (GANs) to transfer semantic information based on room arrangements to a newly visited environment. Our proposed method assigns semantics to the map of an unknown environment using the prior distribution of the map trained in known environments and the multimodal observations made in the unknown environment. We experimentally show in simulation that SpCoMapGAN can use global information for estimating the semantic map and is superior to previous methods. Finally, we also demonstrate in a real environment that SpCoMapGAN can accurately 1) deal with local information, and 2) acquire the semantic information of real places.
Yuki Katsumata, Akira Taniguchi, Lotfi El Hafi, Yoshinobu Hagiwara, Tadahiro Taniguchi
IROS4
2017 Online spatial concept and lexical acquisition with simultaneous localization and mapping
abstract
In this paper, we propose an online learning algorithm based on a Rao-Blackwellized particle filter for spatial concept acquisition and mapping. We have proposed a nonparametric Bayesian spatial concept acquisition model (SpCoA). We propose a novel method (SpCoSLAM) integrating SpCoA and FastSLAM in the theoretical framework of the Bayesian generative model. The proposed method can simultaneously learn place categories and lexicons while incrementally generating an environmental map. Furthermore, the proposed method has scene image features and a language model added to SpCoA. In the experiments, we tested online learning of spatial concepts and environmental maps in a novel environment of which the robot did not have a map. Then, we evaluated the results of online learning of spatial concepts and lexical acquisition. The experimental results demonstrated that the robot was able to more accurately learn the relationships between words and the place in the environmental map incrementally by using the proposed method.
Akira Taniguchi, Yoshinobu Hagiwara, Tadahiro Taniguchi, Tetsunari Inamura
IROS2
2015 Statistical localization exploiting convolutional neural network for an autonomous vehicle
abstract
In this paper, we propose a self-localization method that exploits object recognition results by using convolutional neural networks (CNNs) for autonomous vehicles. Monte-Carlo localization (MCL) is one of the most popular localization methods that use odometry and distance sensor data for determining vehicle position. Some errors are often observed in the localization tasks and MCL often suffers from global positional errors. A global positional error means that particles representing a vehicle's position are distributed in the form of a multimodal distribution, i.e., the distribution has several peaks. To overcome this problem, we propose a method in which an autonomous vehicle employs object recognition results, obtained using CNNs, as the measurement data with a Bag-of-Features representation in an integrative manner. The semantic information found in the recognition results obtained using the CNN reduces the global errors in localization. The experimental results show that the proposed method can converge the distribution of the vehicle positions and particle orientations and reduce the global positional errors.
Satoshi Ishibushi, Akira Taniguchi, Toshiaki Takano, Yoshinobu Hagiwara, Tadahiro Taniguchi
IECON4
2014 A new dimension for RoboCup @home: human-robot interaction between virtual and real worlds
abstract
This work proposes a new approach to realize embodied and multimodal HRI between virtual robot and real world human for HRI challenges in RoboCup @Home.
Jeffrey Too Chuan Tan, Tetsunari Inamura, Yoshinobu Hagiwara, Komei Sugiura, Takayuki Nagai
HRI3