Akishige Yuguchi

dblp:202/0234 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0002-2985-1162ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 What Should Autonomous Robots Verbalize and What Should They Not?
Daichi Yoshihara, Akishige Yuguchi, Seiya Kawano, Takamasa Iio, Koichiro Yoshino
MMM (5)2
2025 Multi-step or Direct: A Proactive Home-Assistant System Based on Commonsense Reasoning
abstract
There is a growing expectation for the realization of proactive home-assistant robots that can assist users in their daily lives. It is essential to develop a framework that closely observes the user’s surrounding context, selectively extracts relevant information, and infers the user’s needs to proactively propose appropriate assistance. In this study, we first extend the Do-I-Demand dataset to define expected proactive assistance actions in domestic situations, where users make ambiguous utterances. These behaviors were defined based on common patterns of support that a majority of users would expect from a robot. We subsequently constructed a framework that infers users’ expected assistance actions from ambiguous utterances through commonsense reasoning. We explored two approaches: (1) multi-step reasoning using COMET as a commonsense reasoning engine, and (2) direct reasoning using large language models. Our experimental results suggest that both the multi-step and direct reasoning methods can successfully derive necessary assistance actions even when dealing with ambiguous user utterances.
Konosuke Yamasaki, Shohei Tanaka, Akishige Yuguchi, Seiya Kawano, Koichiro Yoshino
SIGDIAL3
2024 A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions
abstract
Situated conversations, which refer to visual information as visual question answering (VQA), often contain ambiguities caused by reliance on directive information. This problem is exacerbated because some languages, such as Japanese, often omit subjective or objective terms. Such ambiguities in questions are often clarified by the contexts in conversational situations, such as joint attention with a user or user gaze information. In this study, we propose the Gaze-grounded VQA dataset (GazeVQA) that clarifies ambiguous questions using gaze information by focusing on a clarification process complemented by gaze information. We also propose a method that utilizes gaze target estimation results to improve the accuracy of GazeVQA tasks. Our experimental results showed that the proposed method improved the performance in some cases of a VQA system on GazeVQA and identified some typical problems of GazeVQA tasks that need to be improved.
Shun Inadumi, Seiya Kawano, Akishige Yuguchi, Yasutomo Kawanishi, Koichiro Yoshino
LREC/COLING3
2024 J-CRe3: A Japanese Conversation Dataset for Real-world Reference Resolution
abstract
Understanding expressions that refer to the physical world is crucial for such human-assisting systems in the real world, as robots that must perform actions that are expected by users. In real-world reference resolution, a system must ground the verbal information that appears in user interactions to the visual information observed in egocentric views. To this end, we propose a multimodal reference resolution task and construct a Japanese Conversation dataset for Real-world Reference Resolution (J-CRe3). Our dataset contains egocentric video and dialogue audio of real-world conversations between two people acting as a master and an assistant robot at home. The dataset is annotated with crossmodal tags between phrases in the utterances and the object bounding boxes in the video frames. These tags include indirect reference relations, such as predicate-argument structures and bridging references as well as direct reference relations. We also constructed an experimental model and clarified the challenges in multimodal reference resolution tasks.
Nobuhiro Ueda, Hideko Habe, Akishige Yuguchi, Seiya Kawano, Yasutomo Kawanishi, Sadao Kurohashi, Koichiro Yoshino
LREC/COLING3
2024 Prompt Design Using Past Dialogue Summarization for LLMs to Generate the Current Appropriate Dialogue
Yuya Okadome, Akishige Yuguchi, Ryota Fukui, Yoshio Matsumoto
ICANN (9)2
2024 Evaluation of Preference on Context-Aware Utterances based on Personality Traits using a Conversational Android Robot System
abstract
With recent technological improvements, the performance of conversational robots such as context-aware dialogue is expected to be dramatically enhanced. While context-aware utterance is considered one of the crucial functions of conversational robots, only a few studies investigate the relationship between context awareness and user preference. As is the case with human social relationships, human-robot interactions are under the effect of user’s attributes such as personality, and thus considering these factors of users is essential to continue interacting with robots. In this paper, we evaluate the preference for context awareness based on personality trait effects in an open-domain dialogue using a conversational android robot system. Two types of utterances, context-aware and context-free, are implemented in the robot system, and the participants talk with and evaluate our conversational system. The preference for two systems and the personality traits of each participant are collected. We analyze the relationship between preferences and personalities. The results suggest that the participants’ preference for whether an utterance is context-aware or context-free is distinguished by some personality traits.
Ryota Fukui, Akishige Yuguchi, Yoshio Matsumoto, Yuya Okadome
RO-MAN2
2024 An Immersive Mirror-Reversal Interface for Teleoperated Bring-me Task using a Mobile Manipulator
abstract
One of the major tasks of assistive robots for future domestic applications is to bring objects requested by the user, which is called a “Bring-me Task.” Many of those previous studies including the Bring-me Task focus on the technologies to perform the task autonomously. However, object recognition, grasping, and navigation on autonomous robots are still difficult and not yet practical in a real-world unstructured environment. As an alternative approach, Bring-me Tasks by teleoperated robots have also been studied, but such systems require remote operators to perform the task for the user. In this paper, we propose a novel approach to the Teleoperated Bring-me Task in which a user teleoperates a robot by him/herself and receives an object from the robot. The user wears an HMD showing a mirror-reversal image from the viewpoint of the robot. The results from the experiments on the Teleoperated Bring-me Task suggest that the proposed interface with the mirror-reversal image improved the performance of receiving the object by hand compared with the original image. Furthermore, it was also confirmed that our interface improved the operability of the robot. (A video introducing the proposed method can be viewed11https://youtu.be/Qwrj8tIK_lI?si=CgjGGOC4IZzgO-JC.)
Keiichi Ono, Akishige Yuguchi, Yoshio Matsumoto
SMC2
2024 Capturing Contact Surfaces by a Frustrated Total Internal Reflection System Using a Curved Plate for Comparison of the Beginning of Touching Motions by Humanitude Experts and Novices
abstract
Analyzing time-series changes of the contact surface by touching is important to elucidate the touching skills of Humanitude as one of the pervasive multimodal comprehensive care methodologies. For the analysis, there is a frustrated total internal reflection (FTIR) method to capture contact surfaces on a transparent flat plate by a camera. However, this conventional flat plate is far from the actual surfaces of care receivers because the surfaces of humans consist of curved shapes. In this paper, we propose an FTIR sensing system using a transparent curved plate to capture more ideal contact states with a surface shape more similar to the human body. We collect the contact surface data of the beginning of touching motions by Humanitude experts and novices using the FTIR sensing system with the curved and flat plates. Then, we compare the data by the experts and novices in terms of time-series contact areas and the quantitative indices and subjective evaluation and discuss the analysis results. Through these experiments, we confirm that the proposed system has the potential for the novices to perform more correctly the beginning of Humanitude's touching motions.
Akishige Yuguchi, Mayuki Toyoda, Sung-Gwi Cho, Atsushi Nakazawa, Jun Takamatsu, Koichiro Yoshino, Tsukasa Ogasawara
SMC1
2023 Operative Action Captioning for Estimating System Actions
abstract
Human-assistive systems, such as robots, need to correctly understand the surrounding situation based on obser-vations and output the required support actions for humans. Language is one of the important channels to communicate with humans, and robots are required to have the ability to express their understanding and action-planning results. In this study, we propose a new task of operative action captioning that estimates and verbalizes the actions to be taken by the system in a human-assisting domain. We constructed a system that outputs a verbal description of a possible operative action that changes the current state to the given target state. We collected a dataset consisting of two images as observations, which express the current state and the state changed by actions and a caption that describes the actions that change the current state to the target state, by crowdsourcing in daily life situations. Then we constructed a system that estimates an operative action by a caption. Since the operative action's caption is expected to contain some state-changing actions, we use scene graph prediction as an auxiliary task because the events written in the scene graphs correspond to the state changes. Experimental results showed that our system successfully described the operative actions that should be conducted between the current and target states. The auxiliary tasks that predict the scene graphs improved the quality of the estimation results.
Taiki Nakamura, Seiya Kawano, Akishige Yuguchi, Yasutomo Kawanishi, Koichiro Yoshino
ICRA3
2022 Butsukusa: A Conversational Mobile Robot Describing Its Own Observations and Internal States
abstract
This paper presents an autonomous conversational mobile robot Butsukusa that can describe its own observations and internal states during patrolling tasks. The proposed robot can observe the surrounding environment using the recognition module for objects, humans, environment, localization, and speech and then move autonomously around an indoor living space. Interaction skills via language are required for the robot to perform in such human-centered spaces. To investigate a better communication protocol with users, we evaluate various language generation patterns based on different observations and interaction patterns. The evaluation results indicate that the importance of describing the robot's observation results and internal states, as well as the necessity of an appropriate description, depends on the situation.
Akishige Yuguchi, Seiya Kawano, Koichiro Yoshino, Carlos Toshinori Ishi, Yasutomo Kawanishi, Yutaka Nakamura, Takashi Minato, Yasuki Saito, Michihiko Minoh
HRI1
2022 Multimodal Persuasive Dialogue Corpus using Teleoperated Android
Seiya Kawano, Muteki Arioka, Akishige Yuguchi, Kenta Yamamoto, Koji Inoue, Tatsuya Kawahara, Satoshi Nakamura 0001, Koichiro Yoshino
INTERSPEECH3
2019 Evaluating Imitation of Human Eye Contact and Blinking Behavior Using an Android for Human-like Communication
abstract
The appearance of android robots is very similar to that of human beings. From their appearance, we expect that androids might provide us with high-level communication. The imitation of human behavior gives us the feeling of natural behavior even if we do not know what drives high-level communication. In this paper, we evaluate the imitation of human eye behavior by an android. We consider that the android imitates human eye behavior while explaining some research topic and a person acts as a listener. Then, we construct a method to imitate the eye behavior obtained from eye trackers. For the evaluation, we asked seventeen male subjects for their subjective evaluation and compared the imitation with an android that controlled eye-contact duration and eyeblinks by editing the imitation or programming rule-based behavior. From the results, we found out that 1) the rule-based behaviors kept human-likeness, 2) 3-second eye contact obtained better scores regardless of the imitation-based or rule-based eye behavior, and 3) the subjects might regard the longer eyeblinks as voluntary eyeblinks, with the intention to break eye contacts.
Tetsuya Sano, Akishige Yuguchi, Gustavo Alfonso Garcia Ricardez, Jun Takamatsu, Atsushi Nakazawa, Tsukasa Ogasawara
RO-MAN2
2019 Real-Time Gazed Object Identification with a Variable Point of View Using a Mobile Service Robot
abstract
As sensing and image recognition technologies advance, the environments where service robots operate expand into human-centered environments. Since the roles of service robots depend on the user situations, it is important for the robots to understand human intentions. Gaze information, such as gazed objects (i. e., the objects humans are looking at) can help to understand the users' intentions. In this paper, we propose a real-time gazed object identification method from RGBD images captured by a camera mounted on a mobile service robot. First, we search for the candidate gazed objects using state-of-the-art, real-time object detection. Second, we estimate the human face direction using facial landmarks extracted by a real-time face detection tool. Then, by searching for an object along the estimated face direction, we identify the gazed object. If the gazed object identification fails even though a user is looking at an object, i. e., has a fixed gaze direction, the robot can determine whether the object is inside or outside the robot's view based on the face direction, and, then, change its point of view to improve the identification. Finally, through multiple evaluation experiments with the mobile service robot Pepper, we verified the effectiveness of the proposed identification and the improvement of the identification accuracy by changing the robot's point of view.
Akishige Yuguchi, Tomoaki Inoue, Gustavo Alfonso Garcia Ricardez, Ming Ding 0002, Jun Takamatsu, Tsukasa Ogasawara
RO-MAN1
2018 Human-like Subconscious Behaviors for an Android when Telling a Lie
abstract
Since the appearance of androids is very similar to the appearance of humans, androids are expected to perform human-like interactions. Among these interactions, telling a lie is especially interesting since such an interaction is not always malicious and is employed due to social politeness and to protect self-esteem of the person who tells a lie. In this paper, we develop and evaluate an android that performs humanlike subconscious behaviors when telling a lie in a simple game. For this purpose, we design a game to evaluate the effectiveness of such behaviors. Based on a literature review, we devise a strategy to implement such behaviors on the android. By observing the game being played between two subjects, we analyze and implement the behaviors obtained. The experimental results indicated that the android exhibiting the human-like subconscious behaviors gave users an impression of higher affinity, activity, and honesty. Finally, we concluded that playing games with the android was more amusing for subjects than simple interaction.
Masahiro Iwamoto, Akishige Yuguchi, Masahiro Yoshikawa, Gustavo Alfonso Garcia Ricardez, Jun Takamatsu, Tsukasa Ogasawara
RO-MAN2