VLDB 2026 Research / reviewers in the wild / expert
Harumi Kobayashi
dblp:42/7940
· DBLP profile ↗
26ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0002-6076-6097ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 16 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Is "Two" Just Two? Human Expectations of How Robots Interpret Number WordsabstractThis study examined peoples’ expectations about robots’ thinking about numbers in terms of "exact" interpretation and "at least" interpretation. Japanese adult participants were shown three boxes (e.g., a box with one orange, a covered box with no visible contents, and a box with three oranges) and were asked how they or the robot would respond to instructions such as “Give me the box with two oranges.” In the number condition, the word “two” was interpreted as “exactly two,” and when no visible box matched this quantity, participants tended to select the covered box. However, when predicting the robot’s choice, participants showed a tendency to avoid selecting the covered box compared to when making choices themselves. This discrepancy did not occur in the scalar condition. The study suggested that people’s expectations about number words are different between humans and robots. Zixin He, Harumi Kobayashi, Tetsuya Yasuda |
HAI | 2 |
| 2025 | Do People Choose an Incorrect Option if Other Humans or Robots are Looking at it?abstractGaze direction is a powerful nonverbal cue that conveys attention, intention, and social relevance. It is unclear to what extent people are influenced not only by human gaze but also by robotic gaze. This study examined how agent type (Human, Robot, or Visor Robot which wore a visor) and gaze direction influence human choices in a task where participants judged which of two grids had higher black area density, with correct answers not easily discernible. Participants observed four agents whose gaze favored either the correct option (3vs1 Correct), incorrect option (3vs1 Incorrect), or neither (2vs2). Results showed participants were more likely to choose incorrectly when most agents looked at the wrong option. Under difficult conditions, robot gaze had stronger influence than human gaze, whereas in easier conditions, the opposite was true. This finding offers implications for designing effective gaze-based interactions in HAI. Yousei Morimoto, Harumi Kobayashi, Tetsuya Yasuda |
HAI | 2 |
| 2024 | Co-speech gestures complement motion state information expressed by verbs
Hina Kimura, Tetsuya Yasuda, Harumi Kobayashi |
CogSci | 3 |
| 2024 | Understanding demonstratives produced by agentsabstractIt is essential to investigate how humans interpret a communicative robot’s “utterances.” This study examined how Japanese participants interpreted an agent’s use of demonstratives kono (this) and ano (that) when the agent was either human or robot, since demonstratives are important words for joint attention. In the experiment, participants observed an agent (robot or human agent) uttering an expression using demonstratives with or without eye contact. Then participants were asked to respond the intended referent. Participants’ eye gaze was also recorded. The results were that, in the use of demonstratives, newly appearing referents were assumed to be referred. Shorter reaction times were observed when eye contact was initially established compared to when there was no eye contact, and when the robot agent spoke rather than the human agent. It suggested participants paid more attention to human agent’s referential intention. Interpretation of demonstratives seems to be influenced by agent types. Hina Kimura, Harumi Kobayashi, Tetsuya Yasuda |
HAI | 2 |
| 2024 | People's expectation on the use of pragmatic inferences in "robot utterances"abstractPeople’s utterances often mean information that is not stated in the utterance itself. For example, if you invited your friends Naomi and Ken for party and another good friend said, “Naomi will come,” people may assume that Ken will not come. The reason is that if Ken will also come, your friend would have said so, thus, not saying may mean that he will not come. This kind of inference is called pragmatic inference and assumed to be important for human communication. Then do we assume communicative robots use such pragmatic inferences? This study examined whether people think robots’ utterances include certain types of pragmatic inferences when people interacted with robots via virtual reality who appeared to be autonomous/non-autonomous and produced considerate words/no considerate words, in addition to human avatar. The result was that when the implicature task was more ambiguous, participants thought the human avatar with considerate words, the autonomous robot with considerate words, and the non-autonomous robot without considerate words would use similar implicature (“ad hoc implicature”). However, participants’ expectation for the autonomous robot without considerate words was more varied. This difference was not found when the task was more ad-hoc implicature inclined. The study suggests that people may assume different types of implicature for robots according to the robot types. Chisato Nishihata, Harumi Kobayashi, Tetsuya Yasuda |
HAI | 2 |
| 2023 | Spontaneous co-speech gestures with prompt phrases reflect linguistic structures
Hina Kimura, Tetsuya Yasuda, Harumi Kobayashi |
CogSci | 3 |
| 2023 | Human-like "agents" or "tools"?: Exploring the implicature-of-quantity in HAIabstractA quantitative implicature task was used to investigate whether humans believe that robots can make human-like inferences. As quantifier terms, such as “little” and “much,” do not specify an exact amount, the actual quantity must be inferred from various pieces of information. The participants (N = 24) had to encounter either 100% perfectly-controllable-robots or randomly-controllable-robots. The participants then determined the exact amount of energy to charge to each vehicle depending on the initial-state of the vehicle and the agent’s request. The results indicated that the participants’ estimates of the amounts for randomly-controlled-robots were similar to those for human agents. However, participants with experience in perfectly-controlled-robots estimated a lower amount when the request was labeled as “little,” indicating their estimation was based on literal interpretation. People seemed to predict that perfectly-controlled-robots would not use human-like inference, implying that perfectly-controlled-robots can be assumed to be just “tools” not human-like “agents.” Chisato Nishihata, Harumi Kobayashi, Tetsuya Yasuda |
HAI | 2 |
| 2021 | The Use of Co-Speech Gestures in Conveying Japanese Phrases with Verbs
Yuki Handa, Tetsuya Yasuda, Harumi Kobayashi |
CogSci | 3 |
| 2020 | Gesture and pause can facilitate chunking syntactic information in ambiguous phrases
Harumi Kobayashi |
CogSci | 1 |
| 2020 | Understanding scalar implicature without scale markers SOME and ALL in Japanese preschoolers and adults
Tetsuya Yasuda, Harumi Kobayashi |
CogSci | 2 |
| 2019 | Do people use gestures differently to disambiguate the meanings of Japanese compounds?
Kei Kashiwadate, Tetsuya Yasuda, Harumi Kobayashi |
CogSci | 3 |
| 2019 | Demonstrative "This" and Hand Pointing Can Promote Socio-Centric Interpretations About Invisible Objects
Tetsuya Yasuda, Kei Kashiwadate, Harumi Kobayashi |
CogSci | 3 |
| 2017 | Enforced pointing gesture can indicate invisible objects behind a wall
Hajime Takahashi, Tetsuya Yasuda, Harumi Kobayashi |
CogSci | 3 |
| 2014 | Addressers' gesture changes according to addressees' interpretation of communicative intentionabstractTo construct mind-reading robots, it is important to correctly identify the communicative intention of humans. In human environments, there are usually an enormous number of objects, and each object has many different parts. Therefore, referring to an object or an object part in current communication is not an easy task. We use both linguistic and nonlinguistic information to specify such referential intentions, but exactly how nonlinguistic information is used has not been examined in detail. In this study, video data from eight pairs of adult addressers and addressees who communicated about whole-object or object-part labels were analyzed. The addressers' utterances were transcribed, and nouns and demonstrative words were extracted. Hand motions were divided into five categories: showing, pointing, stroking, functional action, and other. Gaze data obtained using eye movement recordings were also coded using the following categories: gaze on object, face, or body of the experimenter and other. Gaze on object was further divided into two categories; critical and noncritical object parts. The results showed that participants used different patterns of hand motions and gaze to refer to whole-object and object-part labels. When they taught whole-object labels, they showed the object and looked at the addressee's face. When they taught object-part labels, they pointed at, stroked, and looked at the object part. Based on the results, we propose the concept of action contrast in human communication of referential intention. We propose that the need for specification and the use of actions in a contrasting manner are related. Nonlinguistic cues, such as pointing and showing, and timing of utterances are important sources of information that specify human referential intentions in the environment. Action contrast is an important source of information specifying addressers' referential intentions, which can be utilized to construct mind-reading robots. Tetsuya Yasuda, Harumi Kobayashi |
RO-MAN | 2 |
| 2013 | KANSEI Differences between the Child and the Caregiver in Room Arrangement Examined by the Room Arrangement Workshop for Child Room Spaces
Yoko Katsumata, Hideaki Kitazume, Tetsuya Yasuda, Harumi Kobayashi |
CogSci | 4 |
| 2012 | Roles of Adult's Gestures and Eye Gaze in Whole or Object Part Presenting
Tetsuya Yasuda, Harumi Kobayashi |
CogSci | 2 |
| 2012 | Changes of action ontology in conversation among collaborators using virtual spaceabstractAnalyzer of conversation is effective to estimate a person's intention and to predict the person's actions, and such function of the analyzer can be utilized for watch-over system. It is, however, difficult to make such perfect analyzer with a simple computer system. Hence in this study, we focused on change in utilization of corpus by limiting the situation of human interaction. We investigated construction of common grounds and sharing ontology based on action words using utterances occurred in the skill acquisition of virtual collaborative conveyer task. Specifically, we constructed corpus and analyzed conversation between three adults that occurred while completing 10 task trials. We extracted action words from the conversation and classified them into three categories, Move, Stop, and Turn. Action words were defined as the word directly related to transportation task. The results were that the number of action word types decreased in all action word categories and in all experimental groups, showing that the ontology of action concepts converged into a few words in each category. The results also showed that the construction of common grounds among the participants seems to contribute to the convergence of action ontology. Harumi Kobayashi, Tetsuya Yasuda, Hiroshi Igarashi |
RO-MAN | 1 |
| 2012 | Activity recognition for children using self-organizing mapabstractBased on a new concept of Kinder-GuaRdian which aims to realize a support system for kindergarten's staffs and parents by watching over children's activity and their growing process, activity recognition (AR) for children is treated in this paper. In case of children, smaller amount of sensors is better since easy fitting of the sensor device to child's body is strongly preferable. In this study, using one acceleration sensor attached to the upper arm, AR in cases of adults and children were evaluated using several orthodox classifiers and self-organizing map (SOM). Through the several types of validation, effectiveness, issues, and benefits of their methods are discussed. Yasue Mitsukura, Hiroshi Igarashi, Harumi Kobayashi, Fumio Harashima |
RO-MAN | 4 |
| 2012 | Use of demonstrative words between speaker and hearer in virtual spaceabstractTo realize interactive robots that watch for humans, it is necessary for robots to keep distance adequately from humans. The robot must keep away from the person to monitor for avoiding unnecessary interference, while the robot must approach to the person for active communication or supportive actions. Thus, in addition to that distance between the robot and the person must be properly controlled, the human intention of either “I act, without your help” or “You act, help me” must be specified. Demonstrative words such as “this” and “that” may be used as one way to control this distance information and understand human desire for support. We investigated the relationship between distance and demonstrative words when a speaker utters demonstrative words to a hearer in cooperative carrying task using virtual space. We used 3 corpus of speech to extract demonstrative words, then we calculated the distance between the speaker's and the hearer's position to the object to be moved. Results indicated 1) Demonstrative “Kore (This)” was used when the speaker's position was near to the objects. 2) Demonstrative “Sore (That-proximal)” was used when the hearer's position was near to the objects. 3) Demonstrative “Are (That-distal)” was not observed. It was also suggested that the aspect of distance (proximal/distal) was not the key factor of the use of Japanese demonstratives in virtual space. “Kore” was often used for objects that the speaker oneself wanted to act on, while “Sore” was often used for objects that the speaker wanted the other person (who was close to the object) act on. Whether the speaker himself wanted to act or he wanted the hearer act was the key factor of the use of these demonstratives Tetsuya Yasuda, Harumi Kobayashi, Hiroshi Igarashi |
RO-MAN | 2 |
| 2011 | Relation between skill acquisition and task specific human speech in collaborative workabstractTo accomplish the objective to make human-collaborative robots, we need to clarify how humans actually interact each other when they do collaborative work. In this study, we transcribed all utterances produced while participants' completing human-human collaborative conveyer task, and computed and categorized all morphemes (minimal unit of language meaning) using 4 categories based on the morpheme's role in the task. The role categories were Robot Action, User Action, Modifier, Object. We analyzed the utterances produced by 4 groups, 3 participants in each group. Results were that frequency of each category per minute decreased over ten trials. However, the variety of words in each category tended to show an inverted U-shaped pattern. Based on these results, we proposed three stages of language skill acquisition in a collaborative work. Shuichi Nakata, Harumi Kobayashi, Tetsuya Yasuda, Masafumi Kumata, Hiroshi Igarashi |
RO-MAN | 2 |
| 2010 | Changes of utterances in the skill acquisition of collaborative conveyer taskabstractImportance of developing human adaptive robots or systems is increasing these days. To accomplish it, we have to clarify features of human-human communication. In this study, we analyzed human speech while completing computerized collaborate task to clarify how humans speak in collaborative work. We extracted questioning speeches from conversations using CLAN, and classified them into four categories according to their content. We investigated whether the number and ratio of each type of questions changed over ten trials. The results indicate that leader role person is sensitive to task planning Shuichi Nakata, Harumi Kobayashi, Hiroshi Igarashi |
HRI | 2 |
| 2010 | Social cooperation analysis for multi-user cooperative tasksabstractThis paper addresses evaluation methods for social cooperation characteristics in cooperative tasks. Using these we found that the ratio of four typical indexes relate to task performance. Most of conventional human-machine systems assumed to assist only a single operator. In such cases, the system should only pay attention to his/her intention or operation characteristics. If there were multiple operators they also included social behaviours affected from others. Such behaviours could not be regarded by the one-on-one assist systems. Since intelligent assist systems are expected to apply in human society, the system requires work among humans. A challenge of our research is to quantify the human social characteristics, especially when relating to cooperation in the task involving multiple operators. In this paper, human social characteristics are quantified in order to evaluate cooperative performance. For the evaluations, ratio indexes by four kinds of behaviours are focused on. One is observation behaviour for other operator's work, which is typical social behaviour in a cooperative task. Second is the ratio of approaching to other robots for physical supporting. Third, total activity by summation of robot motion distance. Then, an index of visibility of robots each other. Finally, correlation between expansion of the index distribution and a task completion time is analyzed. Furthermore advanced assist using the evaluation is discussed. Hiroshi Igarashi, Harumi Kobayashi, Fumio Harashima |
RO-MAN | 3 |
| 2010 | Question utterances in a collaborative conveyer task in virtual spaceabstractA group of participants performed a collaborative conveyer task that needs to transfer objects to designated positions by vehicles in a virtual space. All utterances were transcribed and questioning utterances were specified based on rising intonation patterns and content of the speech. Then we classified these utterances into four types: planning (ask whole transporting strategy in task), partner's (asking the movement and status of the partners), self (asking own movement and status), and other, and analyzed the changes of the frequency of these question types during 10 trials. We also constructed a collaboration levels and evaluated collaboration status for each completion of a box conveying. This collaboration level was determined by the number of people who collaborated to move one box during one trial. We investigated the relationship between strategy changes, collaboration levels, and question utterances. The results were that in a group that showed a strategy change, collaboration level was generally higher and more varied than the collaboration level observed in the group that showed no strategy change. In addition, in a group that showed a strategy change, question utterances about whole plan decreased but question utterances about the self status increased. However, in a group who showed no strategy change, all types of question utterances decreased. It was found that when a strategy change occurred, people tend to make more questions about self status to confirm the details of movement to complete the task. Collaboration level changes can be a good index to infer occurrences of strategy changes. Masafumi Kumata, Shuichi Nakata, Hiroshi Igarashi, Harumi Kobayashi |
RO-MAN | 5 |
| 2010 | Influence of 10 seconds' interval in pragmatic interpretationabstractThis study examined interpreting word meanings and movement of line-of-regard of participants in a joint attention experiment. In addition to immediately giving an object label to a child, we tested an effect of 10 seconds' interval on children and adults using a joint attention experiment. Results were that adults used both pragmatic and eye gaze cues and interpreted word meanings appropriately. However, 4-year-old children tended to use only a pragmatic cue in the similar task ignoring eye gaze when a label was given after 10 seconds elapsed. 2-year-old children used only eye gaze. The study suggested that with an immature interactive system a child tends to rely only one or a few non-linguistic cues in interaction. Implication of this study is that robots must be able to use various non-linguistic cues with an integrated manner in human-robot interactions. Tetsuya Yasuda, Harumi Kobayashi |
RO-MAN | 2 |
| 2009 | Role sharing analysis on multi-operator cooperative workabstractThis paper addresses a quantification method for role sharing in cooperative tasks. By the method, we found that the ratio of three typical indexes relate to task performance. Most of conventional human-machine systems assumed to assist single operator. In such case, the systems should only pay attention his/her operation characteristics. If there were multiple operators, they also include altruistic behaviours during the work. Such behaviours could not be regarded by such one-on-one assist system. Since intelligent assist systems are expected to apply in human society, the systems are required work among humans. A challenge of our research is to quantify such humans' altruistic behaviours, especially what relating cooperation in the task involving multiple participants. In this paper, role sharing characteristics are quantified in order to evaluate the cooperative performance. For this quantification, ratio indexes of three kinds of behaviours is focused on. One is observation behaviours for other operator's work, which is typical altruistic behaviour in a cooperative task. Second is ratio of egoistic behaviour as active conveyance. Then, total activity by motion distance of the robot could be indexes of the role sharing. Finally, correlation between these ratio of indexes and a task performance is analyzed. Furthermore an advanced assist using the evaluation is discussed. Hiroshi Igarashi, Harumi Kobayashi, Fumio Harashima |
RO-MAN | 3 |
| 2009 | Relations between eye gaze and cognitive factors in a transportation task using remote controlabstractWe investigated eye movement in a transportation task using remote control. Participants controlled two sets of joysticks to move a toy excavator and a toy truck monitoring the task process through video displays. They repeated the task ten times. At an early stage of skill acquisition, the participants tended to look at both task related objects (e.g., material to transport and the shovel) and the course (e.g., ground and wall), but did not look much at obstacles (trees). At a late stage of skill acquisition, the participants tended to look at the course, other than related objects or obstacles. However, fixations to related objects did not decrease in ldquolook aheadrdquo fixations or fixations prior to manipulation of the joystick. It can be said that as humans acquire a skill, the positions of the related objects are learned so they did not actually look at these objects when they manipulate joystick. We came to a temporal conclusion that the participants might allocate ldquocovert attentionrdquo or attention without gaze to these related objects and obstacles for better performance. Harumi Kobayashi, Tetsuya Yasuda, Shuichi Nakata |
RO-MAN | 1 |