VLDB 2026 Research / reviewers in the wild / expert
Kohei Okuoka
dblp:231/5171
· DBLP profile ↗
14ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0002-5569-3356ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 13 · 1 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conversational Robot System for Travel Memoir GenerationabstractReflecting on memories has a positive effect on mental health. For robots that interact with older adults, interactions that look back on memories are also an important application area. Despite the importance of such reminiscence, no study has simultaneously addressed methods to both facilitate robots conversing about users’ memories and generate content from those memories. In this article, we propose TRAVOT, a system that explores the events behind travel photos through conversation and generates travel memoirs with such photos. TRAVOT uses a large language model (LLM) for flexible information collection. Moreover, it can deepen the conversation topic to obtain a more profound story from the user via a topic control mechanism with Meta-LLM. This not only elicits information that cannot be obtained from the photos based on a prepared list of questions but also allows for deepening the discussion by generating additional questions. In addition, it can eliminate redundant questions that result from naive use of an LLM by applying matching judgment with certain required questions. We conducted a user experiment to evaluate TRAVOT’s effectiveness, and we found that the participants could recall more interesting and unusual things that happened during their trips when they had conversations with TRAVOT. The users could also reminisce about their trips when they read the travel memoirs generated by TRAVOT. In addition, TRAVOT increased the amount of information contained in the conversations and travel memoirs. Kaon Shimoyama, Kohei Okuoka, Mitsuhiko Kimoto, Michita Imai |
ACM Trans. Hum. Robot Interact. | 2 |
| 2025 | LLM-Based Evaluation of Utterances with Implicature Understanding: A Preliminary StudyabstractAlthough Large Language Models (LLMs) have recently shown remarkable performance in many language comprehension tasks, they struggle to perform adequately in communicative contexts involving implicature. In our previous study, we proposed LLM-based agents by integrating LLMs with cognitive models. In three dialogue scenarios, these agents generated appropriate utterances as if inferring the speaker’s intentions (i.e., implicature). Further investigation of the agents’ performances requires an examination of their utterances in a large number of scenarios. In addition, it is also important to consistently evaluate the agents’ utterances. Thus, this study proposes a method in which LLMs evaluate agents’ generated utterances in the same way as human evaluators. Using our pilot prompt, we demonstrated that the evaluations of LLMs and human evaluators were similar. Ayu Iida, Kohei Okuoka, Takashi Omori, Ryoichi Nakashima, Masahiko Osawa |
HAI | 2 |
| 2025 | Semantic Babbling: Interactive Baby Robot System Using Large Language ModelsabstractTherapeutic robots are used in facilities for older people, but many of them can only provide mechanical responses to users. To improve the quality of interactions, it is crucial to address emotional consistency. Thus, we propose Semantic Babbling, a system that enables interaction with users by utilizing the baby robot's “inner voice”. By generating inner voices and selecting babbling based on sentiment, Semantic Babbling aims to mimic infant-like responses. A crowd-sourcing survey demonstrated the significance of emotional consistency and the display of inner voice in Semantic Babbling, revealing that it improves the impression of baby robots. Serina Miyake, Ryuki Matsuoka, Kohei Okuoka, Takuto Akiyoshi, Hidenobu Sumioka, Masahiro Shiomi, Michita Imai |
HRI | 3 |
| 2025 | RelBot: Building a Balanced Relationship From Conversation ContentabstractTo construct a conversational system for robots that considers the interpersonal relationships among three parties, this paper proposes a system called RelBot. The system uses a large language model (LLM) to estimate the current interpersonal relationship and ideal balanced interpersonal relationship from the content of a three-way conversation between a user and two robots; then, it generates statements for the robots to establish the ideal balanced relationship according to the user's desire and Heider's balance theory. The challenging points here are to investigate whether an LLM can recognize the current relationships between a user and each robot and between two robots from the content of their conversation, and whether it can estimate the relationships that the user wants to achieve. Moreover, RelBot provides a new way to adjust the three-way relationship to the ideal balanced one. We conducted two evaluation experiments. The results indicated that participants could intentionally change the relationships to the desired ones, while RelBot accurately recognized both the current relationship and the participant's desired relationship. Furthermore, the final relationship in the free conversation with two robots could be balanced through the effectiveness of RelBot's generated statements. Yoshinari Onodera, Hikaru Matsuzaki, Ayano Kawara, Kohei Okuoka, Mitsuhiko Kimoto, Michita Imai |
HRI | 4 |
| 2025 | RoDiL: Giving Route Directions with Landmarks by Robots
Kanta Tachikawa, Shota Akahori, Kohei Okuoka, Mitsuhiko Kimoto, Michita Imai |
ICAART (1) | 3 |
| 2024 | Integrating Large Language Model and Mental Model of Others: Studies on Dialogue Communication Based on ImplicatureabstractDespite the significant development of Large Language Models (LLMs), they struggle with “dialogue communication based on implicature,” which humans handle easily. In this study, we aim to improve the performance of LLMs in this type of dialogue tasks by integrating LLMs with a cognitive model of dialogue. The cognitive model of dialogue comprises beliefs, desires, and intentions, which is defined as the Mental Model of Others (MMO) for predicting or estimating other’s mental states and behaviors. We propose two integration methods: the LLM Embedded in Cognitive Model (LEC) and the Cognitive Model Embedded in LLM (CEL). We examined the performances of our proposed method using the dialogue task, showing that the LEC can respond appropriately in the dialogues with implicatures, which cannot be achieved by conventional LLMs. Ayu Iida, Kohei Okuoka, Satoko Fukuda, Takashi Omori, Ryoichi Nakashima, Masahiko Osawa |
HAI | 2 |
| 2022 | Advantage Mapping: Learning Operation Mapping for User-Preferred Manipulation by Extracting Scenes with Advantage FunctionabstractWhen a user manipulates a system, a user input through an interface, or an operation, is converted to the user’s intended action according to the mapping that links operations and actions, which we call “operation mapping”. Although many operation mappings are created by designers assuming how a typical user would operate the system, the optimal operation mapping may vary from user to user. The designer cannot prepare in advance all possible operation mappings. One approach to solve this problem involves autonomous learning of an operation mapping during the operation. However, existing methods require manual preparation of scenes for learning mappings. We propose advantage mapping, which enables the efficient learning of operation mappings. Working from the idea that scenes in which the user’s desired action is predictable are useful for learning operation mappings, advantage mapping extracts scenes according to the magnitude of entropy in the output of the action value function acquired from reinforcement learning. In our experiment, the user’s ideal operation mapping was more accurately obtained from the scenes selected by advantage mapping than from learning through actual play. Rintaro Hasegawa, Yosuke Fukuchi, Kohei Okuoka, Michita Imai |
HAI | 3 |
| 2022 | VISTURE: A System for Video-Based Gesture and Speech Generation by RobotsabstractThis paper proposes VISTURE, a system for generating a robot’s gesture and speech by using video as input. VISTURE assumes a situation in which a robot conveys what it saw with a camera to a person who was absent. The value of this paper is that we have performed a case study to investigate the expressions that Japanese people use to describe video scenes, and used the results to build VISTURE. In particular, we found classification of expressions depicting the video scenes throughout the case study: Foreground information that is the relevant event of the scene and Background one that is not the main point of the description giving the entire scene. Foreground and Background are referred in combination. VISTURE employs the classification to generate human-like expressions. Moreover, we designed the method to determine Foreground and Background, and it can generate multiple combinations of expressions. We investigated the people’s impression of a robot performing the gestures and speech generated by VISTURE to evaluate the quality of those gestures and speech. The results showed that the robot was perceived as more likable and capable when it performed gestures. Kaon Shimoyama, Kohei Okuoka, Mitsuhiko Kimoto, Michita Imai |
HAI | 2 |
| 2022 | $Q$-Mapping: Learning User-Preferred Operation Mappings With Operation-Action Value FunctionabstractUser interfaces have been designed to fit typical users and their usage styles as assumed by designers. However, it is impossible to cover all the possible use cases. To address this problem, we propose$Q$-Mapping, which is a method for user interfaces to acquire the operation mapping, or mapping from user operations to their effects.$Q$-Mapping has an advantage over previous techniques in that it can acquire operation mapping interactively. The core idea of$Q$-Mapping is that what a user selects as an ideal action has a tendency to be the same as the action that has the highest$Q$-value. On the basis of this concept, we defined the operation-action value function, which can be calculated from the value that a user expects to gain when a particular mapping is given in that state and is updated each time an operation occurs. We conducted a simulation experiment and a user study to investigate the$Q$-Mapping performance and the effects of the acquisition of interactive operation mapping. The simulation results showed that the changeability of operation mapping could be controlled by a coefficient called the balancing parameter. As for the user study, we found that$Q$-Mapping with a balancing parameter that decays with time was able to acquire operation mapping that was easy for users to understand. These results demonstrate the importance of balancing consistency and adaptability in the interactive acquisition of operation mapping. Riki Satogata, Mitsuhiko Kimoto, Yosuke Fukuchi, Kohei Okuoka, Michita Imai |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2021 | Examining the Factors that Make Co-Watching with Agents EffectiveabstractThe spread of video streaming services has increased the opportunities for video watching, and research on co-watching with agents is also being conducted. However, few studies have been conducted and not insufficient knowledge has been obtained. In this study, we will conduct an investigation on the relationship among the background of the co-watchers, the video evaluation, and the agent’s impression. In the experiment, we asked the participants to co-watch a basketball game with an agent who responded appropriately then, and after watching, they answered a questionnaire. As a result, there was a moderate correlation between the Likeability and the evaluation of the video watching. There was also a correlation between the evaluation of the video watching and the empathy for the agent. Masaki Abe, Kohei Okuoka, Masahiko Osawa |
HAI | 2 |
| 2020 | PredGaze: A Incongruity Prediction Model for User's Gaze MovementabstractWith digital signage and communication robots, digital agents have gradually become popular and will become more popular. It is important to make humans notice the intentions of agents throughout the interaction between them. This paper is focused on the gaze behavior of an agent and the phenomenon that if the gaze behavior of an agent is different from human expectations, human will have a incongruity and feel the existence of the agent's intention behind the behavioral changes instinctively. We propose PredGaze, a model of estimating this incongruity which humans have according to the shift in gaze behavior from the human's expectations. In particular, PredGaze uses the variance in the agent behavior model to express how well humans sense the behavioral tendency of the agent. We expect that this variance will improve the estimation of the incongruity. PredGaze uses three variables to estimate the internal state of how much a human senses the agent's intention: error, confidence, and incongruity. To evaluate the effectiveness of PredGaze with these three variables, we conducted an experiment to investigate the effects of the timing of gaze behavior change and incongruity. The experimental results indicated that there were significant differences in the subjective scores of the naturalness of agents and incongruity with agents according to the difference in the timing of the agent's change in its gaze behavior. Yohei Otsuka, Shohei Akita, Kohei Okuoka, Mitsuhiko Kimoto, Michita Imai |
RO-MAN | 3 |
| 2019 | Notification Timing of Agent with Vection and Character for Semi-Automatic Wheelchair OperationabstractAutomatic driving systems not only for cars, but also for wheelchairs are being developed. Improving the operability and safety of electric wheelchairs is an important issue. For example, if a driving system changes the speed of the vehicle without a driver's operation, it makes the driver uneasy. We developed a system to address this uneasiness. Our system, called the MIZUSAKI system, notifies drivers of a change in the speed gain, which is controlled by the system, before the change. This system uses 4 in-screen effects, which are intended to be seen in the peripheral vision and do not inhibit drivers' attention to driving. In developing this system, we have always considered the notification timing to be a key factor. We tested this system to find the best notification timing and found that if the anticipatory timing is 3 seconds before the speed gain changes or later. Kouichi Enami, Kohei Okuoka, Shohei Akita, Michita Imai |
HAI | 2 |
| 2018 | Semi-Autonomous Telepresence Robot for Adaptively Switching Operation Using Inhibition and Disinhibition MechanismabstractIn research on semi-autonomous telepresence robots, a problem in which remote operators become frustrated with autonomous operations that do not match their intention has been reported. However, in previous research, a general-purpose method for automatically switching between remote and autonomous operations has not been proposed. In this paper, through the use of a general purpose arbitration model, called the accumulator based arbitration model (ABAM), we propose an adaptive switching architecture for remote and autonomous operations, named "One Minder." We incorporated One Minder into a semi-autonomous telepresence system autonomizing contingent behaviors, and conducted experiments to verify its utility using a robot implementing the proposed architecture. As the experiment results indicate, it was shown that One Minder can adaptively switch between remote and autonomous operations without manual switching. In addition, One Minder was also shown to reduce the operational load and frustration given to a remote operator by allowing the arbitration to properly output an autonomous operation. Kohei Okuoka, Yusuke Takimoto, Masahiko Osawa, Michita Imai |
HAI | 1 |
| 2018 | Adaptive Semi-autonomous Agents via Episodic ControlabstractShared autonomy is a situation in which agents adapt to users based on feedbacks to jointly accomplish tasks. In the case of semiautonomous agents such as mobile robots operational load can be reduced by adaptively automating their behavior. Machine learning is one of the method for adapting teleoperated agents to users as control tendency depends on users and environments. It is difficult to regard user inputs as supervisions because user controls are not always available during descisionmaking processes for agents that continually make decisions (such as navigation robots). Takuma Seno, Kohei Okuoka, Masahiko Osawa, Michita Imai |
HAI | 2 |