VLDB 2026 Research / reviewers in the wild / expert
Mitsuhiko Kimoto
dblp:171/5771
· DBLP profile ↗
20ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0001-8441-8815ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 17 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Conversational Robot System for Travel Memoir GenerationabstractReflecting on memories has a positive effect on mental health. For robots that interact with older adults, interactions that look back on memories are also an important application area. Despite the importance of such reminiscence, no study has simultaneously addressed methods to both facilitate robots conversing about users’ memories and generate content from those memories. In this article, we propose TRAVOT, a system that explores the events behind travel photos through conversation and generates travel memoirs with such photos. TRAVOT uses a large language model (LLM) for flexible information collection. Moreover, it can deepen the conversation topic to obtain a more profound story from the user via a topic control mechanism with Meta-LLM. This not only elicits information that cannot be obtained from the photos based on a prepared list of questions but also allows for deepening the discussion by generating additional questions. In addition, it can eliminate redundant questions that result from naive use of an LLM by applying matching judgment with certain required questions. We conducted a user experiment to evaluate TRAVOT’s effectiveness, and we found that the participants could recall more interesting and unusual things that happened during their trips when they had conversations with TRAVOT. The users could also reminisce about their trips when they read the travel memoirs generated by TRAVOT. In addition, TRAVOT increased the amount of information contained in the conversations and travel memoirs. Kaon Shimoyama, Kohei Okuoka, Mitsuhiko Kimoto, Michita Imai |
ACM Trans. Hum. Robot Interact. | 3 |
| 2025 | RelBot: Building a Balanced Relationship From Conversation ContentabstractTo construct a conversational system for robots that considers the interpersonal relationships among three parties, this paper proposes a system called RelBot. The system uses a large language model (LLM) to estimate the current interpersonal relationship and ideal balanced interpersonal relationship from the content of a three-way conversation between a user and two robots; then, it generates statements for the robots to establish the ideal balanced relationship according to the user's desire and Heider's balance theory. The challenging points here are to investigate whether an LLM can recognize the current relationships between a user and each robot and between two robots from the content of their conversation, and whether it can estimate the relationships that the user wants to achieve. Moreover, RelBot provides a new way to adjust the three-way relationship to the ideal balanced one. We conducted two evaluation experiments. The results indicated that participants could intentionally change the relationships to the desired ones, while RelBot accurately recognized both the current relationship and the participant's desired relationship. Furthermore, the final relationship in the free conversation with two robots could be balanced through the effectiveness of RelBot's generated statements. Yoshinari Onodera, Hikaru Matsuzaki, Ayano Kawara, Kohei Okuoka, Mitsuhiko Kimoto, Michita Imai |
HRI | 5 |
| 2025 | RoDiL: Giving Route Directions with Landmarks by Robots
Kanta Tachikawa, Shota Akahori, Kohei Okuoka, Mitsuhiko Kimoto, Michita Imai |
ICAART (1) | 4 |
| 2024 | What Kinds of Facial Self-Touches Strengthen Expressed Emotions?abstractAs virtual spaces continue to attract more and more attention, the development of virtual agents is also gaining momentum. Agents capable of expressing emotions are being developed, and various methods exist through which they express such feelings. Although much research has focused on the implementation of human actions in agents, no research has yet implemented facial self-touch for agents used in emotional expressions, even though such transmission of human emotions is a critical tool for them. In this study, we developed a system that enriches agents' emotional expressions by implementing self-touches in virtual agents. From an evaluation of our system, we found that it enhances the emotions of virtual agents. The naturalness of their emotional expressions varied depending on the emotion. Junya Kawano, Mitsuhiko Kimoto, Takamasa Iio, Masahiro Shiomi |
RO-MAN | 2 |
| 2024 | Two is Better than One: Cultural Differences in the Number of Apologizing Robots in the U.S. and JapanabstractApology behavior design is becoming important for social robots that work in daily environments because of their widespread use. Robotics researchers have reported that multiple robots can effectively achieve more acceptance and trust in apology situations. Unfortunately, such effects have only been confirmed in a single country: Japan. Some studies investigated cultural differences in apologies between Japan and other cultures and reported how the former influences the perceived function and meaning of apologies. Therefore, we conducted a web-based survey to investigate whether using multiple robots in apology situations is effective in another country. We compared such perceived feelings as forgiveness and trust toward a robot’s apologies between Japan and the U.S. by using the visual stimuli of one and two robots. The experiment results showed that U.S. people felt that multiple robot apologies are more acceptable than apologies from just one robot, similar to results with Japanese participants. Perceived trust did show a different phenomenon between the two countries. Masahiro Shiomi, Taichi Hirayama, Mitsuhiko Kimoto, Takamasa Iio, Katsunori Shimohara |
RO-MAN | 3 |
| 2023 | Conversational Context-sensitive Ad Generation with a Few Core-QueriesabstractWhen people are talking together in front of digital signage, advertisements that are aware of the context of the dialogue will work the most effectively. However, it has been challenging for computer systems to retrieve the appropriate advertisement from among the many options presented in large databases. Our proposed system, the Conversational Context-sensitive Advertisement generator (CoCoA), is the first attempt to apply masked word prediction to web information retrieval that takes into account the dialogue context. The novelty of CoCoA is that advertisers simply need to prepare a few abstract phrases, called Core-Queries, and then CoCoA automatically generates a context-sensitive expression as a complete search query by utilizing a masked word prediction technique that adds a word related to the dialogue context to one of the prepared Core-Queries. This automatic generation frees the advertisers from having to come up with context-sensitive phrases to attract users’ attention. Another unique point is that the modified Core-Query offers users speaking in front of the CoCoA system a list of context-sensitive advertisements. CoCoA was evaluated by crowd workers regarding the context-sensitivity of the generated search queries against the dialogue text of multiple domains prepared in advance. The results indicated that CoCoA could present more contextual and practical advertisements than other web-retrieval systems. Moreover, CoCoA acquired a higher evaluation in a particular conversation that included many travel topics to which the Core-Queries were designated, implying that it succeeded in adapting the Core-Queries for the specific ongoing context better than the compared method without any effort on the part of the advertisers. In addition, case studies with users and advertisers revealed that the context-sensitive advertisements generated by CoCoA also had an effect on the content of the ongoing dialogue. Specifically, since pairs unfamiliar with each other more frequently referred to the advertisement CoCoA displayed, the advertisements had an effect on the topics about which the pairs spoke. Moreover, participants of an advertiser role recognized that some of the search queries generated by CoCoA fit the context of a conversation and that CoCoA improved the effect of the advertisement. In particular, they learned how to design of designing a good Core-Query at ease by observing the users’ response to the advertisements retrieved with the generated search queries. Ryoichi Shibata, Shoya Matsumori, Yosuke Fukuchi, Tomoyuki Maekawa, Mitsuhiko Kimoto, Michita Imai |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2022 | What is the speed boundary between patting and slapping by a robot?: Investigating perceptions toward robot-robot interactionabstractSuch multiple agent interactions as robot-robot interactions are one effective approach to indirectly provide information, unlike direct interactions between people and agents. Several studies reported the effectiveness of this approach and developed conversational mechanisms for robot-robot interaction, although their physical interaction design for robot-robot interaction remains relatively less focused. This study concentrated on the effects of speed in touching behaviors between robots, because people may perceive the relationships shared between robots differently due to various motion speeds even though the motion is identical. Therefore, we conducted two web-survey experiments to investigate the perceptions of a robot's touching motion and how the relationship between two robots was perceived when a robot touched another at different speeds. The first experiment showed two peak speeds where people perceived a robot's touch as patting (friendly touch) or slapping (aggressive touch). The second experiment results showed similar peak speeds where people perceived the relationships between robots as positive or negative to patting or slapping behaviors. We believe that knowledge about the relationships between motion speeds and perceived friendliness between robots will contribute to design physical interaction between robots. Taichi Hirayama, Yuka Okada, Mitsuhiko Kimoto, Takamasa Iio, Katsunori Shimohara, Masahiro Shiomi |
HAI | 3 |
| 2022 | VISTURE: A System for Video-Based Gesture and Speech Generation by RobotsabstractThis paper proposes VISTURE, a system for generating a robot’s gesture and speech by using video as input. VISTURE assumes a situation in which a robot conveys what it saw with a camera to a person who was absent. The value of this paper is that we have performed a case study to investigate the expressions that Japanese people use to describe video scenes, and used the results to build VISTURE. In particular, we found classification of expressions depicting the video scenes throughout the case study: Foreground information that is the relevant event of the scene and Background one that is not the main point of the description giving the entire scene. Foreground and Background are referred in combination. VISTURE employs the classification to generate human-like expressions. Moreover, we designed the method to determine Foreground and Background, and it can generate multiple combinations of expressions. We investigated the people’s impression of a robot performing the gestures and speech generated by VISTURE to evaluate the quality of those gestures and speech. The results showed that the robot was perceived as more likable and capable when it performed gestures. Kaon Shimoyama, Kohei Okuoka, Mitsuhiko Kimoto, Michita Imai |
HAI | 3 |
| 2022 | Utilizing Core-Query for Context-Sensitive Ad Generation Based on DialogueabstractIn this work, we present a system that sequentially generates advertisements within the context of a dialogue. Advertisements tailored to the user have long been displayed on the digital signage in stores, on web pages, and on smartphone applications. Advertisements will work more effectively if they are aware of the context of the dialogue between the users. Creating an advertising sentence as a query and searching the web by using that query is one way to present a variety of advertisements, but there is currently no method to create an appropriate search query for the search in accordance with the dialogue context. Therefore, we developed a method called the Conversational Context-sensitive Advertisement generator (CoCoA). The novelty of CoCoA is that advertisers simply need to prepare a few abstract phrases, called Core-Queries, and then CoCoA dynamically transforms the Core-Queries into complete search queries in accordance with the dialogue context. Here, “transforms” means to add words related to the context in the dialogue to the prepared Core-Queries. The transformation is enabled by a masked word prediction technique that predicts a word that is hidden in a sentence. Our attempt is the first to apply masked word prediction to a web information retrieval framework that takes into account the dialogue context. We asked users to evaluate the search query presented by CoCoA against the dialogue text of multiple domains prepared in advance and found that CoCoA could present more contextual and effective advertisements than Google Suggest or a method without the query transformation. In addition, we found that CoCoA generated high-quality advertisements that advertisers had not expected when they created the Core-Queries. Ryoichi Shibata, Shoya Matsumori, Yosuke Fukuchi, Tomoyuki Maekawa, Mitsuhiko Kimoto, Michita Imai |
IUI | 5 |
| 2022 | $Q$-Mapping: Learning User-Preferred Operation Mappings With Operation-Action Value FunctionabstractUser interfaces have been designed to fit typical users and their usage styles as assumed by designers. However, it is impossible to cover all the possible use cases. To address this problem, we propose$Q$-Mapping, which is a method for user interfaces to acquire the operation mapping, or mapping from user operations to their effects.$Q$-Mapping has an advantage over previous techniques in that it can acquire operation mapping interactively. The core idea of$Q$-Mapping is that what a user selects as an ideal action has a tendency to be the same as the action that has the highest$Q$-value. On the basis of this concept, we defined the operation-action value function, which can be calculated from the value that a user expects to gain when a particular mapping is given in that state and is updated each time an operation occurs. We conducted a simulation experiment and a user study to investigate the$Q$-Mapping performance and the effects of the acquisition of interactive operation mapping. The simulation results showed that the changeability of operation mapping could be controlled by a coefficient called the balancing parameter. As for the user study, we found that$Q$-Mapping with a balancing parameter that decays with time was able to acquire operation mapping that was easy for users to understand. These results demonstrate the importance of balancing consistency and adaptability in the interactive acquisition of operation mapping. Riki Satogata, Mitsuhiko Kimoto, Yosuke Fukuchi, Kohei Okuoka, Michita Imai |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2020 | Can a Robot's Touches Express the Feeling of Kawaii toward an Object?abstractKawaii, a Japanese word that means "cute," is an essential design concept in consumer and pop culture in Japan. In this study, we focused on a situation where a social robot describes an object during an information-providing task, which is commonly required for social robots in daily environments. Since past studies reported kawaii feelings are associated with a motivation to approach a target, our robot expressed feelings of kawaii to objects touch behaviors. We also focused on whether touch behaviors that emphasize style increase the feeling of kawaii of the touched object, following a phenomenon where people strongly touch a target when they overwhelmingly feel positive emotion: cute aggression. Our experimental results showed the effectiveness of touch behaviors to express the feelings of kawaii from the robot toward objects and to increase the participants' feelings of kawaii toward the object. We identified fewer effects from the participants to the robot. The emphasized motion style did not show any significant effects for the kawaii feelings. Yuka Okada, Mitsuhiko Kimoto, Takamasa Iio, Katsunori Shimohara, Hiroshi Nittono, Masahiro Shiomi |
IROS | 2 |
| 2020 | Effects of Social Touch from an Agent in Virtual Space: Comparing Visual Stimuli and Virtual-Tactile StimuliabstractAlthough social touch in physical space has been scrutinized for its positive effects on people, the effects of social touch in virtual space has been neglected. In virtual space, two types of social touch can be designed, a touch with only visual stimuli and a touch with both visual and tactile stimuli. This paper investigates the effects of the two types of agent's social touch on users in the context where the agent praises their performance of a task in virtual space. Based on past studies of social touch, we hypothesized that the two types of agent's social touch in virtual space would improve the user's task motivation, task performance, and the agent's likability. We experimentally tested our hypotheses by comparing those variables among no-touch, visual-touch, and visual-tactile touch groups. Since our results showed no significant differences among these groups, our hypotheses were not supported. However, a post-hoc analysis by gender suggests that the agent's social touch with both visual and tactile stimuli while praising male users increased their task motivation in virtual space. This result suggests that the effects of an agent's social touch in virtual space may be different due to genders, and shows a possibility of positive effects for the touch behavior design of the agent in virtual space. Kana Higashino, Mitsuhiko Kimoto, Takamasa Iio, Katsunori Shimohara, Masahiro Shiomi |
RO-MAN | 2 |
| 2020 | Effects of Touch Behaviors and Whispering Voices in Robot-Robot Interaction for Information Providing TasksabstractUsing multiple robots for information providing tasks is one effective approach to attract people by showing conversational interaction between robots. In this study, we investigate the effectiveness of showing such intimate interaction between robots as touching because it has positive effects on human-robot interactions, even though it has received less attention than robot-robot interaction. We prepared touch behaviors and whispering voices between robots and conducted experiments where robots provided information to participants. Our experiment results showed the effectiveness of both touch behaviors and whispering voices between robots in information providing tasks. On the other hands, we also found that whispering voices decreased their likeability, although touch behavior did not have such negative effects. Yuka Okada, Riko Taniguchi, Akihiro Tatsumi, Masashi Okubo, Mitsuhiko Kimoto, Takamasa Iio, Katsunori Shimohara, Masahiro Shiomi |
RO-MAN | 5 |
| 2020 | PredGaze: A Incongruity Prediction Model for User's Gaze MovementabstractWith digital signage and communication robots, digital agents have gradually become popular and will become more popular. It is important to make humans notice the intentions of agents throughout the interaction between them. This paper is focused on the gaze behavior of an agent and the phenomenon that if the gaze behavior of an agent is different from human expectations, human will have a incongruity and feel the existence of the agent's intention behind the behavioral changes instinctively. We propose PredGaze, a model of estimating this incongruity which humans have according to the shift in gaze behavior from the human's expectations. In particular, PredGaze uses the variance in the agent behavior model to express how well humans sense the behavioral tendency of the agent. We expect that this variance will improve the estimation of the incongruity. PredGaze uses three variables to estimate the internal state of how much a human senses the agent's intention: error, confidence, and incongruity. To evaluate the effectiveness of PredGaze with these three variables, we conducted an experiment to investigate the effects of the timing of gaze behavior change and incongruity. The experimental results indicated that there were significant differences in the subjective scores of the naturalness of agents and incongruity with agents according to the difference in the timing of the agent's change in its gaze behavior. Yohei Otsuka, Shohei Akita, Kohei Okuoka, Mitsuhiko Kimoto, Michita Imai |
RO-MAN | 4 |
| 2019 | Preliminary Investigation of Pre-Touch Reaction Distances toward Virtual AgentsabstractThis study addresses the pre-touch reaction distance effects in human-agent touch interaction in a VR environment. Past studies on human-agent interaction focused on post-touch situations, i.e., pre-touch situations received less attention, and such pre-touch situation was only investigated with a physical agent, i.e., robot. In this study, we conducted a preliminary data collection to investigate the minimum comfortable distance to virtual agent's touch by using VR application. For this purpose, we prepared two kinds of virtual agents (female and male) to investigate gender effects in touch settings. We analyzed the collected data to investigate about people's perceptions towards a touch from virtual agents, and the results showed similar phenomenon with a physical agent. Aoba Sato, Mitsuhiko Kimoto, Takamasa Iio, Katsunori Shimohara, Masahiro Shiomi |
HAI | 2 |
| 2018 | Calibrating Depth Sensors for Pedestrian Tracking Using a Robot as a Movable and Localized LandmarkabstractUsing multiple depth sensors enables us to accurately track pedestrians in real environments. Accurate pedestrian positions are essential for building an effective human-centered cyber world provided by location-based services. In particular, ceiling-mounted depth sensors can robustly track people in such environments. However, one important problem for this approach is the accurate calibration of the absolute sensor positions. This problem remains unsolved due to limited range and sensor distortions from a distance. Manual calibration is complicated and time-consuming, and the existing calibration method still has several limitations since it used a pedestrian as a movable landmark. Instead of a human landmark, we propose a method that uses a mobile robot as a movable and localized landmark to calibrate each sensor. We compared the calibration performance of the proposed and existing methods and showed that the former achieved more accurate calibration for both the absolute sensor and tracked pedestrian positions. Our proposed method with a mobile robot not only increased the accuracy of the calibration processes but also decreased human efforts. Mitsuhiko Kimoto, Masahiro Shiomi, Takamasa Iio, Katsunori Shimohara, Norihiro Hagita |
SMC | 1 |
| 2017 | Effects of a Listener Robot with Children in StorytellingabstractThis paper investigates the effects of a listener robot that joins a storytelling situation for children as a side-participant. For this purpose, we develop a storytelling-robot system that consists of both reader and listener robots. Our semi-autonomous system involves a human operator who makes correct responses and provides easily understandable answers to the children's questions during storytelling. We develop a gaze model for the natural and autonomous gaze behaviors of a reader robot by considering multiple listeners and a storytelling object (a display that shows images). We conducted an experiment with 16 children to investigate whether they preferred storytelling with/without the listener robot and the changes of speech activities during the storytelling. Children preferred storytelling with the listener robot to storytelling without it. Their speech activities decreased when the listener robot was involved in the storytelling. Yumiko Tamura, Mitsuhiko Kimoto, Masahiro Shiomi, Takamasa Iio, Katsunori Shimohara, Norihiro Hagita |
HAI | 2 |
| 2016 | Communication Cues in a Human-Robot Touch InteractionabstractHaptic interaction is a key capability for social robots that closely interact with people in daily environments. Such human communication cues as gaze behaviors make haptic interaction look natural. Since the purpose of this study is to increase human-robot touch interaction, we conducted an experiment with 20 participants who interacted with a robot with different combinations of gaze behaviors and touch styles. The experimental results showed that both gaze behaviors and touch styles influence the changes in the perceived feelings of touch interaction with a robot. Takahiro Hirano, Masahiro Shiomi, Takamasa Iio, Mitsuhiko Kimoto, Takuya Nagashio, Ivan Tanev, Katsunori Shimohara, Norihiro Hagita |
HAI | 4 |
| 2016 | Alignment Approach Comparison between Implicit and Explicit Suggestions in Object Reference ConversationsabstractThe recognition of an indicated object by an interacting person is an essential function for a robot that acts in daily environments. To improve recognition accuracy, clarifying the goal of the indicating behaviors is needed. For this purpose, we experimentally compared two kinds of interaction strategies: a robot that explicitly provides instructions to people about how to refer to objects or a robot that implicitly aligns with the people's indicating behaviors. Even though our results showed that participants evaluated the implicit approach to be more natural than the explicit approach, the recognition performances of the two approaches were not significantly different. Mitsuhiko Kimoto, Takamasa Iio, Masahiro Shiomi, Ivan Tanev, Katsunori Shimohara, Norihiro Hagita |
HAI | 1 |
| 2015 | Improvement of object reference recognition through human robot alignmentabstractThis paper reports an interactive approach to improve the recognition performance by robots of objects indicated by humans during human-robot interaction. We developed an approach based on two findings in conversations where a human refers to an object, which is confirmed by a robot. First, humans tend to use the same words or gestures as the robot in a phenomenon called alignment. Second, humans tend to decrease the amount of information in their references when the robot uses excess information in its confirmations: in other words, alignment inhibition. These findings lead to the following design; a robot should use enough information without being excessive to identify objects to improve recognition accuracy because humans will eventually use similar information to refer to those objects by alignment. If humans more frequently use the same information to identify objects, the robot can more easily recognize those being indicated by humans. To verify our design, we developed a robotic system to recognize the objects to which humans referred and conducted a control experiment that had 2 × 3 conditions; one factor was the robot's confirmation way and another was the arrangement of the objects. The first factor had two levels to identify objects: enough information and excess information. The second factor had three levels: congestion, two groups, and a sparse set. We measured the recognition accuracy of the objects humans referred to and the amount of information in their references. The success rate of the recognition and information amount was higher in the adequate information condition than in the excess condition in a particular situation. The results suggested the possibility that our proposed interactive approach improved recognition performance. Mitsuhiko Kimoto, Takamasa Iio, Masahiro Shiomi, Ivan Tanev, Katsunori Shimohara, Norihiro Hagita |
RO-MAN | 1 |