Tetsuya Yasuda

dblp:42/1411 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0002-2181-7816ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 6 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 11 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Is "Two" Just Two? Human Expectations of How Robots Interpret Number Words
abstract
This study examined peoples’ expectations about robots’ thinking about numbers in terms of "exact" interpretation and "at least" interpretation. Japanese adult participants were shown three boxes (e.g., a box with one orange, a covered box with no visible contents, and a box with three oranges) and were asked how they or the robot would respond to instructions such as “Give me the box with two oranges.” In the number condition, the word “two” was interpreted as “exactly two,” and when no visible box matched this quantity, participants tended to select the covered box. However, when predicting the robot’s choice, participants showed a tendency to avoid selecting the covered box compared to when making choices themselves. This discrepancy did not occur in the scalar condition. The study suggested that people’s expectations about number words are different between humans and robots.
Zixin He, Harumi Kobayashi, Tetsuya Yasuda
HAI3
2025 Do People Choose an Incorrect Option if Other Humans or Robots are Looking at it?
abstract
Gaze direction is a powerful nonverbal cue that conveys attention, intention, and social relevance. It is unclear to what extent people are influenced not only by human gaze but also by robotic gaze. This study examined how agent type (Human, Robot, or Visor Robot which wore a visor) and gaze direction influence human choices in a task where participants judged which of two grids had higher black area density, with correct answers not easily discernible. Participants observed four agents whose gaze favored either the correct option (3vs1 Correct), incorrect option (3vs1 Incorrect), or neither (2vs2). Results showed participants were more likely to choose incorrectly when most agents looked at the wrong option. Under difficult conditions, robot gaze had stronger influence than human gaze, whereas in easier conditions, the opposite was true. This finding offers implications for designing effective gaze-based interactions in HAI.
Yousei Morimoto, Harumi Kobayashi, Tetsuya Yasuda
HAI3
2024 Co-speech gestures complement motion state information expressed by verbs
Hina Kimura, Tetsuya Yasuda, Harumi Kobayashi
CogSci2
2024 Understanding demonstratives produced by agents
abstract
It is essential to investigate how humans interpret a communicative robot’s “utterances.” This study examined how Japanese participants interpreted an agent’s use of demonstratives kono (this) and ano (that) when the agent was either human or robot, since demonstratives are important words for joint attention. In the experiment, participants observed an agent (robot or human agent) uttering an expression using demonstratives with or without eye contact. Then participants were asked to respond the intended referent. Participants’ eye gaze was also recorded. The results were that, in the use of demonstratives, newly appearing referents were assumed to be referred. Shorter reaction times were observed when eye contact was initially established compared to when there was no eye contact, and when the robot agent spoke rather than the human agent. It suggested participants paid more attention to human agent’s referential intention. Interpretation of demonstratives seems to be influenced by agent types.
Hina Kimura, Harumi Kobayashi, Tetsuya Yasuda
HAI3
2024 People's expectation on the use of pragmatic inferences in "robot utterances"
abstract
People’s utterances often mean information that is not stated in the utterance itself. For example, if you invited your friends Naomi and Ken for party and another good friend said, “Naomi will come,” people may assume that Ken will not come. The reason is that if Ken will also come, your friend would have said so, thus, not saying may mean that he will not come. This kind of inference is called pragmatic inference and assumed to be important for human communication. Then do we assume communicative robots use such pragmatic inferences? This study examined whether people think robots’ utterances include certain types of pragmatic inferences when people interacted with robots via virtual reality who appeared to be autonomous/non-autonomous and produced considerate words/no considerate words, in addition to human avatar. The result was that when the implicature task was more ambiguous, participants thought the human avatar with considerate words, the autonomous robot with considerate words, and the non-autonomous robot without considerate words would use similar implicature (“ad hoc implicature”). However, participants’ expectation for the autonomous robot without considerate words was more varied. This difference was not found when the task was more ad-hoc implicature inclined. The study suggests that people may assume different types of implicature for robots according to the robot types.
Chisato Nishihata, Harumi Kobayashi, Tetsuya Yasuda
HAI3
2023 Spontaneous co-speech gestures with prompt phrases reflect linguistic structures
Hina Kimura, Tetsuya Yasuda, Harumi Kobayashi
CogSci2
2023 Human-like "agents" or "tools"?: Exploring the implicature-of-quantity in HAI
abstract
A quantitative implicature task was used to investigate whether humans believe that robots can make human-like inferences. As quantifier terms, such as “little” and “much,” do not specify an exact amount, the actual quantity must be inferred from various pieces of information. The participants (N = 24) had to encounter either 100% perfectly-controllable-robots or randomly-controllable-robots. The participants then determined the exact amount of energy to charge to each vehicle depending on the initial-state of the vehicle and the agent’s request. The results indicated that the participants’ estimates of the amounts for randomly-controlled-robots were similar to those for human agents. However, participants with experience in perfectly-controlled-robots estimated a lower amount when the request was labeled as “little,” indicating their estimation was based on literal interpretation. People seemed to predict that perfectly-controlled-robots would not use human-like inference, implying that perfectly-controlled-robots can be assumed to be just “tools” not human-like “agents.”
Chisato Nishihata, Harumi Kobayashi, Tetsuya Yasuda
HAI3
2021 The Use of Co-Speech Gestures in Conveying Japanese Phrases with Verbs
Yuki Handa, Tetsuya Yasuda, Harumi Kobayashi
CogSci2
2020 Understanding scalar implicature without scale markers SOME and ALL in Japanese preschoolers and adults
Tetsuya Yasuda, Harumi Kobayashi
CogSci1
2019 Do people use gestures differently to disambiguate the meanings of Japanese compounds?
Kei Kashiwadate, Tetsuya Yasuda, Harumi Kobayashi
CogSci2
2019 Demonstrative "This" and Hand Pointing Can Promote Socio-Centric Interpretations About Invisible Objects
Tetsuya Yasuda, Kei Kashiwadate, Harumi Kobayashi
CogSci1
2017 Enforced pointing gesture can indicate invisible objects behind a wall
Hajime Takahashi, Tetsuya Yasuda, Harumi Kobayashi
CogSci2
2014 Addressers' gesture changes according to addressees' interpretation of communicative intention
abstract
To construct mind-reading robots, it is important to correctly identify the communicative intention of humans. In human environments, there are usually an enormous number of objects, and each object has many different parts. Therefore, referring to an object or an object part in current communication is not an easy task. We use both linguistic and nonlinguistic information to specify such referential intentions, but exactly how nonlinguistic information is used has not been examined in detail. In this study, video data from eight pairs of adult addressers and addressees who communicated about whole-object or object-part labels were analyzed. The addressers' utterances were transcribed, and nouns and demonstrative words were extracted. Hand motions were divided into five categories: showing, pointing, stroking, functional action, and other. Gaze data obtained using eye movement recordings were also coded using the following categories: gaze on object, face, or body of the experimenter and other. Gaze on object was further divided into two categories; critical and noncritical object parts. The results showed that participants used different patterns of hand motions and gaze to refer to whole-object and object-part labels. When they taught whole-object labels, they showed the object and looked at the addressee's face. When they taught object-part labels, they pointed at, stroked, and looked at the object part. Based on the results, we propose the concept of action contrast in human communication of referential intention. We propose that the need for specification and the use of actions in a contrasting manner are related. Nonlinguistic cues, such as pointing and showing, and timing of utterances are important sources of information that specify human referential intentions in the environment. Action contrast is an important source of information specifying addressers' referential intentions, which can be utilized to construct mind-reading robots.
Tetsuya Yasuda, Harumi Kobayashi
RO-MAN1
2013 KANSEI Differences between the Child and the Caregiver in Room Arrangement Examined by the Room Arrangement Workshop for Child Room Spaces
Yoko Katsumata, Hideaki Kitazume, Tetsuya Yasuda, Harumi Kobayashi
CogSci3
2012 Roles of Adult's Gestures and Eye Gaze in Whole or Object Part Presenting
Tetsuya Yasuda, Harumi Kobayashi
CogSci1
2012 Changes of action ontology in conversation among collaborators using virtual space
abstract
Analyzer of conversation is effective to estimate a person's intention and to predict the person's actions, and such function of the analyzer can be utilized for watch-over system. It is, however, difficult to make such perfect analyzer with a simple computer system. Hence in this study, we focused on change in utilization of corpus by limiting the situation of human interaction. We investigated construction of common grounds and sharing ontology based on action words using utterances occurred in the skill acquisition of virtual collaborative conveyer task. Specifically, we constructed corpus and analyzed conversation between three adults that occurred while completing 10 task trials. We extracted action words from the conversation and classified them into three categories, Move, Stop, and Turn. Action words were defined as the word directly related to transportation task. The results were that the number of action word types decreased in all action word categories and in all experimental groups, showing that the ontology of action concepts converged into a few words in each category. The results also showed that the construction of common grounds among the participants seems to contribute to the convergence of action ontology.
Harumi Kobayashi, Tetsuya Yasuda, Hiroshi Igarashi
RO-MAN2
2012 Use of demonstrative words between speaker and hearer in virtual space
abstract
To realize interactive robots that watch for humans, it is necessary for robots to keep distance adequately from humans. The robot must keep away from the person to monitor for avoiding unnecessary interference, while the robot must approach to the person for active communication or supportive actions. Thus, in addition to that distance between the robot and the person must be properly controlled, the human intention of either “I act, without your help” or “You act, help me” must be specified. Demonstrative words such as “this” and “that” may be used as one way to control this distance information and understand human desire for support. We investigated the relationship between distance and demonstrative words when a speaker utters demonstrative words to a hearer in cooperative carrying task using virtual space. We used 3 corpus of speech to extract demonstrative words, then we calculated the distance between the speaker's and the hearer's position to the object to be moved. Results indicated 1) Demonstrative “Kore (This)” was used when the speaker's position was near to the objects. 2) Demonstrative “Sore (That-proximal)” was used when the hearer's position was near to the objects. 3) Demonstrative “Are (That-distal)” was not observed. It was also suggested that the aspect of distance (proximal/distal) was not the key factor of the use of Japanese demonstratives in virtual space. “Kore” was often used for objects that the speaker oneself wanted to act on, while “Sore” was often used for objects that the speaker wanted the other person (who was close to the object) act on. Whether the speaker himself wanted to act or he wanted the hearer act was the key factor of the use of these demonstratives
Tetsuya Yasuda, Harumi Kobayashi, Hiroshi Igarashi
RO-MAN1
2011 Relation between skill acquisition and task specific human speech in collaborative work
abstract
To accomplish the objective to make human-collaborative robots, we need to clarify how humans actually interact each other when they do collaborative work. In this study, we transcribed all utterances produced while participants' completing human-human collaborative conveyer task, and computed and categorized all morphemes (minimal unit of language meaning) using 4 categories based on the morpheme's role in the task. The role categories were Robot Action, User Action, Modifier, Object. We analyzed the utterances produced by 4 groups, 3 participants in each group. Results were that frequency of each category per minute decreased over ten trials. However, the variety of words in each category tended to show an inverted U-shaped pattern. Based on these results, we proposed three stages of language skill acquisition in a collaborative work.
Shuichi Nakata, Harumi Kobayashi, Tetsuya Yasuda, Masafumi Kumata, Hiroshi Igarashi
RO-MAN3
2010 Influence of 10 seconds' interval in pragmatic interpretation
abstract
This study examined interpreting word meanings and movement of line-of-regard of participants in a joint attention experiment. In addition to immediately giving an object label to a child, we tested an effect of 10 seconds' interval on children and adults using a joint attention experiment. Results were that adults used both pragmatic and eye gaze cues and interpreted word meanings appropriately. However, 4-year-old children tended to use only a pragmatic cue in the similar task ignoring eye gaze when a label was given after 10 seconds elapsed. 2-year-old children used only eye gaze. The study suggested that with an immature interactive system a child tends to rely only one or a few non-linguistic cues in interaction. Implication of this study is that robots must be able to use various non-linguistic cues with an integrated manner in human-robot interactions.
Tetsuya Yasuda, Harumi Kobayashi
RO-MAN1
2009 Relations between eye gaze and cognitive factors in a transportation task using remote control
abstract
We investigated eye movement in a transportation task using remote control. Participants controlled two sets of joysticks to move a toy excavator and a toy truck monitoring the task process through video displays. They repeated the task ten times. At an early stage of skill acquisition, the participants tended to look at both task related objects (e.g., material to transport and the shovel) and the course (e.g., ground and wall), but did not look much at obstacles (trees). At a late stage of skill acquisition, the participants tended to look at the course, other than related objects or obstacles. However, fixations to related objects did not decrease in ldquolook aheadrdquo fixations or fixations prior to manipulation of the joystick. It can be said that as humans acquire a skill, the positions of the related objects are learned so they did not actually look at these objects when they manipulate joystick. We came to a temporal conclusion that the participants might allocate ldquocovert attentionrdquo or attention without gaze to these related objects and obstacles for better performance.
Harumi Kobayashi, Tetsuya Yasuda, Shuichi Nakata
RO-MAN2
1998 A gabor filter-based method for recognizing handwritten numerals
Yoshihiko Hamamoto, Shunji Uchimura, Masanori Watanabe, Tetsuya Yasuda, Yoshihiro Mitani, Shingo Tomita
Pattern Recognit.4
1997 Normalization Techniques of Handwritten Numerals for Gabor Filters
abstract
When recognizing handwritten characters, normalization of characters has been usually done prior to feature extraction. Some authors point out that improvement of normalization increases the recognition rate. Thus, normalization techniques must be carefully examined. We propose a Gabor filter based handwritten character recognition system. Our main interest is to select the best normalization technique for Gabor filters. We conduct comparative experiments and show that position, slant, and thickness normalization is best, with regard to minimizing the error.
Masanori Watanabe, Yoshihiko Hamamoto, Tetsuya Yasuda, Shingo Tomita
ICDAR3
1996 Recognition of handwritten numerals using Gabor features
abstract
We study a Gabor filter-based feature extraction method for handwritten numeral character recognition. The performance of the Gabor filter-based method is demonstrated on the ETL-1 database. Experimental results suggest that the Gabor filter-based method should be considered in recognition of handwritten numeric characters.
Yoshihiko Hamamoto, Shunji Uchimura, Masanori Watanabe, Tetsuya Yasuda, Shingo Tomita
ICPR4