VLDB 2026 Research / reviewers in the wild / expert
Shiro Kumano
dblp:94/53
· DBLP profile ↗
44ranked-venue papers
11as first author
14since 2021 · last 2025
0000-0002-1231-5566ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 24 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Disentangling Perceptual Ambiguity in Multifunctional Nonverbal Behaviors in Conversations via Tensor Spectrum Decomposition
Issa Tamura, Momoka Tajima, Shiro Kumano, Kazuhiro Otsuka |
ICMI | 3 |
| 2025 | Synergistic Functional Spectrum Analysis: A Framework for Exploring the Multifunctional Interplay Among Multimodal Nonverbal Behaviours in ConversationsabstractA novel framework named thesynergistic functional spectrum analysis(sFSA) is proposed to explore the multifunctional interplay among multimodal nonverbal behaviours in human conversations. This study aims to reveal how multimodal nonverbal behaviours cooperatively perform communicative functions in conversations. To capture the intrinsic nature of nonverbal expressions, functional multiplicity, and interpretational ambiguity, e.g., a single head nod could imply listening, agreeing, or both, a novel concept named thefunctional spectrum, which is defined as the distribution of perceptual intensities of multiple functions by multiple observers, is introduced in the sFSA. Based on this concept, this paper presentsfunctional spectrum corpora, which target 44 facial expression and 32 head movement functions. Then, spectrum decomposition is conducted to reduce the multimodal functional spectrum to asynergetic functional spectrumin a lower dimensionfunctional spacethat is spanned byfunctional basisvectors representing primary and distinctive functionalities across multiple modalities. To that end, we propose a semiorthogonal nonnegative matrix factorization (SO-NMF) method, which assumes the additivity of multiple functions and aims to balance the distinctiveness and expressiveness of the factorization. The results confirm that some primary functional bases can be identified, which can be interpreted as the listener’s backchannel, thinking, and affirmative response functions, and the speaker’s thinking and addressing functions, and their positive emotion functions. In addition, regression models based on convolutional neural networks (CNNs) are presented to estimate thesynergistic functional spectrumfrom the head poses and facial action units measured from conversation data. The results of these analyses and experiments confirm the potential of the sFSA and may lead to future extensions. Mai Imamura, Ayane Tashiro, Shiro Kumano, Kazuhiro Otsuka |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Editorial
Rachael Jack, Desmond C. Ong, Khiet Truong, Gale M. Lucas, Shiro Kumano |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | VAD Emotion Control in Visual Art Captioning via Disentangled Multimodal RepresentationabstractArt evokes distinct affective responses, leading to the generation of emotional verbal expressions in response to visual stimuli like visual art images. While previous research has addressed controlling linguistic impressions in terms of basic emotion categories, less attention has been given to their continuous, dimensional nature. This paper aims to continuously modulate emotions across valence, arousal, and dominance (VAD) dimensions by employing multimodal representation learning (MMRL) and crossmodal inference. We utilize a multimodal variational autoencoder for MMRL, encoding visual and textual stimuli into a continuous joint vector representation. We then auxiliarily tune this representation to explicitly include the VAD dimensions in a disentangled manner. This enables nuanced emotional control in image captioning, where visual inputs are converted into vectors and then into text. Trained on the ArtEmis dataset, which includes emotion-evoking captions for visual art, as well as additional datasets annotated with VAD scores for either text or visual art, our model demonstrates that manipulating VAD intensities in the vector representation results in continuous caption variations, reflecting the intended emotional changes. Ryo Ueda, Hiromi Narimatsu, Yusuke Miyao, Shiro Kumano |
ACII | 4 |
| 2024 | Relationship between emotional linkages and perceived emotion during a joint task
Aiko Murata, Shiro Kumano |
CogSci | 2 |
| 2024 | Exploring Interlocutor Gaze Interactions in Conversations based on Functional Spectrum AnalysisabstractA novel framework named a gaze interactional functional spectrum analysis (GI-FSA) is proposed to explore the functional aspects of gaze interactions among interlocutors in conversations. It aims to reveal the primary and distinctive interactional functionalities that emerge via the gaze behaviors of the speaker, and the listener whom the speaker looks at. To capture the intrinsic nature of gaze functions, such as multiple functionalities and ambiguity, this study introduces a novel representation called a gaze functional spectrum representing the distribution of perceptual intensity of multiple gaze functions and presents a gaze functional spectrum corpus that targets 43 gaze functions covering various speech-related, listening-related and other functions. Then, semiorthogonal nonnegative matrix factorization (SO-NMF) is employed to decompose the concatenated speaker-listener functional spectra into a interactional functional spectrum in a lower-dimensional functional space spanned with functional bases, each of which represents a distinct aspect of interactional functionalities. Targeting four female conversations, the GI-FSA revealed interpretable functional bases such as addressing-listening and joint positive emotion. In addition, this paper proposes convolutional neural networks (CNNs) that can recognize the binary level of the interactional functional spectrum from observable multimodal nonverbal behaviors, including head pose, utterance status, eyeball direction and facial expressions. These experimental findings validate the potential of the GI-FSA as a promising framework for analyzing gaze interactions among interlocutors, and understanding communication dynamics. Ayane Tashiro, Mai Imamura, Shiro Kumano, Kazuhiro Otsuka |
ICMI | 3 |
| 2024 | Guest Editorial Best of ACII 2021abstractThe 9TH AAAC Conference on Affective Computing and Intelligent Interaction 2021 was held in a virtual format in the fall of 2021. It was technically co-sponsored by the IEEE Computer Society and featured the recent work on Affective Computing. The six best papers from this conference were selected by the technical program chairs. They were invited to submit their extended version to be considered for this special section at the IEEE Transactions on Affective Computing. Each submission was reviewed by at least three expert reviewers and was evaluated in terms of overall contribution and the adequacy of the additional content to warrant a new article. This special section features five accepted submissions whose major contributions are summarized below. Mohammad Soleymani 0001, Shiro Kumano, Emily Mower Provost, Nadia Bianchi-Berthouze, Akane Sano, Kenji Suzuki 0002 |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Emotion-Controllable Impression Utterance Generation for Visual ArtabstractThe degree of subjectivity and the type of emotions when people express their impressions of an object depend on various factors, including their affective states, psychological traits, goals, and social norms. However, most previous efforts in generating people’s impressions of objects have focused on extreme cases, i.e., fully objective or fully emotional, as represented by the MS-COCO and ArtEmis datasets, respectively. We propose an emotional impression generation method that continuously controls the degree of subjectivity versus objectivity and the type of emotion in a unified manner. In the framework of ConCap, which allows controlling text style by modifying an auxiliary input called prefix, we propose to use two types of prefixes to jointly control the degree of subjectivity and the type of emotion: subjective-style prefix and categorical-emotional-style prefix. The subjective-style prefixes are holistic and attempt to describe the entire emotion felt by a group of people, although different individuals may feel different emotions. The categorical-emotional-style prefix is more individual-oriented and tries to focus on a single or few specific emotions. An experiment using both objectively and subjectively descriptive datasets shows qualitatively and quantitatively that the proposed method provides good gradual control of impression expressions of visual art in terms of both subjectivity and objectivity as well as emotion categories. Ryo Ueda, Hiromi Narimatsu, Yusuke Miyao, Shiro Kumano |
ACII | 4 |
| 2023 | Collision Probability Matching Loss for Disentangling Epistemic Uncertainty from Aleatoric UncertaintyabstractTwo important aspects of machine learning, uncertainty and calibration, have previously been studied separately. The first aspect involves knowing whether inaccuracy is due to the epistemic uncertainty of the model, which is theoretically reducible, or to the aleatoric uncertainty in the data per se, which thus becomes the upper bound of model performance. As for the second aspect, numerous calibration methods have been proposed to correct predictive probabilities to better reflect the true probabilities of being correct. In this paper, we aim to obtain the squared error of predictive distribution from the true distribution as epistemic uncertainty. Our formulation, based on second-order Rényi entropy, integrates the two problems into a unified framework and obtains the epistemic (un)certainty as the difference between the aleatoric and predictive (un)certainties. As an auxiliary loss to ordinary losses, such as cross-entropy loss, the proposed collision probability matching (CPM) loss matches the cross-collision probability between the true and predictive distributions to the collision probability of the predictive distribution, where these probabilities correspond to accuracy and confidence, respectively. Unlike previous Shannon-entropy-based uncertainty methods, the proposed method makes the aleatoric uncertainty directly measurable as test-retest reliability, which is a summary statistic of the true distribution frequently used in scientific research on humans. We provide mathematical proof and strong experimental evidence for our formulation using both a real dataset consisting of real human ratings toward emotional faces and simulation. Hiromi Narimatsu, Mayuko Ozawa, Shiro Kumano |
AISTATS | 3 |
| 2022 | Consistent Smile Intensity Estimation from Wearable Optical SensorsabstractSmiling plays a crucial role in human communication. It is the most frequent expression shown in daily life. Smile analysis usually employs computer vision-based methods that use data sets annotated by experts. However, cameras have space constraints in most realistic scenarios due to occlusions. Wearable electromyography is a promising alternative; however, issue of user comfort is a barrier to long-term use. Other wearable-based methods can detect smiles, but they lack consistency because they use subjective criteria without expert annotation. We investigate a wearable-based method that uses optical sensors for consistent smile intensity estimation while reducing manual annotation cost. First, we use a state-of-art computer vision method (OpenFace) to train a regression model to estimate smile intensity from sensor data. Then, we compare the estimation result to that of OpenFace. We also compared their results to human annotation. The results show that the wearable method has a higher matching coefficient (r=0.67) with human annotated smile intensity than OpenFace (r=0.56). Also, when the sensor data and OpenFace output were fused, the multimodal method produced estimates closer to human annotation (r=0.74). Finally, we investigate how the synchrony of smile dynamics among subjects and their average smile intensity are correlated to assess the potential of wearable smile intensity estimation. Katsutoshi Masai, Monica Perusquía-Hernández, Maki Sugimoto, Shiro Kumano, Toshitaka Kimura |
ACII | 4 |
| 2022 | Cross-Linguistic Study on Affective Impression and Language for Visual Art Using Neural SpeakerabstractVisual art is one of the ideal targets for affective computing because viewing artworks is a regular experience for many people, and it elicits various affective appraisals, cognitions, and reactions in the viewer. To gain a detailed understanding of the interplay between visual content, its emotional impact, and linguistic explanations of this impact, visual art datasets have been proposed. One example is the ArtEmis dataset, which contains emotion categorization and linguistic expressions by crowds of people in reaction to numerous paintings. Linguistic expressions are influenced both by culture in how to appraise art and by language in how to verbalize one's impression. However, cultural and linguistic differences in this domain have not been fully explored. Therefore, we collected a new dataset (ArtEmis-JP) consisting of 16 k emotion labels and utterances in Japan, one of the countries most frequently compared with Western countries in psychology, while carefully following the procedures of the original ArtEmis study conducted with English speakers. In this paper, we report the commonalities and differences between the original ArtEmis dataset and our Japanese dataset using basic statistical comparisons and performance comparisons for a state-of-the-art neural speaker in utterance generation. Going beyond the original study, we newly examined the impact of expertise on emotional categorization and linguistic expressions by examining the differences between experts and non-experts. Hiromi Narimatsu, Ryo Ueda, Shiro Kumano |
ACII | 3 |
| 2022 | Real-time Auditory Feedback System for Bow-tilt Correction while Aiming in ArcheryabstractIn archery, archers aim statically at a target then release an arrow; thus, bow stability during aiming is important for high performance. The stability of the bow is controlled as a part of posture by using the skeletal muscles of the trunk. Previous studies have proposed methods for measuring postural stability, but the extent to which these methods are effective in training is largely unexplored. Therefore, we introduce a real-time auditory feedback system to correct the bow stability of archers during their daily training. Our system provides audible feedback to archers when their bow is tilted from the desired vertical or upright position while aiming. The bow's tilt is measured with a small, lightweight accelerometer attached to the bow, and the system emits tones depending on the tilt angle to alert the archer to correct the bow's tilt. We conducted an experiment to verify the proposed system, and the results indicate a statistically significant 38% reduction in tilt from nine experienced university archers. This effect persisted even after the auditory feedback was turned off. Our system also helped 67% of the participants notice the deviation of their perceived tilt of the bow from the measured tilt. This experiment demonstrated the persistent effect of bow-tilt correction and the usefulness of the proposed auditory feedback system. Takayuki Ogasawara, Hanako Fukamachi, Kenryu Aoyagi, Shiro Kumano, Hiroyoshi Togo, Koichiro Oka |
BIBE | 4 |
| 2021 | Deep Explanatory Polytomous Item-Response Model for Predicting Idiosyncratic Affective RatingsabstractTowards explainable affective computing (XAC), researchers have invested considerable effort into post hoc approaches and reverse engineering to seek explanations for deep learning models. However, alternative, intrinsic approaches that aim to build inherently interpretable models by restricting their complexity are yet to be widely explored. In this study, we integrate an explanatory polytomous item response model that provides a well-established psychological interpretation for ordinal scales with deep neural networks to realize high prediction performance and good result interpretability. We conducted an experiment on a growing task (i.e., predicting the idiosyncratic perception of emotional faces of an individual); as expected theoretically, the topmost parameters of our model demonstrated strong correlations with those of the corresponding ordinal item response model: r = 0.928 to 1.00. Our proposed intrinsic approach can used as a complementary framework for post-hoc methods in XAC to coach and support human social interactions. Tsukasa Ishigaki, Shiro Kumano |
ACII | 3 |
| 2021 | Archery Skill Assessment Using an Acceleration SensorabstractA key skill in archery is the ability to suppress postural tremor while aiming at a target. Providing feedback during daily archery practice is a potentially effective way of suppressing tremor. However, postural tremor is subtle and difficult to measure using vision-based techniques. This article proposes a feedback method that uses a bow equipped with a small, lightweight acceleration sensor. First, we automatically detect an archer's shooting execution cycle, including the aiming, release, and follow-through phases, by using binary classification, and then, we quantify postural tremor during aiming. Then, from the quantified postural tremor, we regress the expected total score that the archer would obtain in a series of shots during a real game. We performed an experiment with 11 members of a university archery club and achieved 1) a precision of 0.72 and recall of 0.80 in shooting detection and 2) an absolute correlation coefficient of 0.74 in score prediction with leave-one-subject-out cross-validation. Takayuki Ogasawara, Hanako Fukamachi, Kenryu Aoyagi, Shiro Kumano, Hiroyoshi Togo, Koichiro Oka |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2020 | Interpersonal physiological linkage is related to excitement during a joint task
Aiko Murata, Shiro Kumano, Junji Watanabe |
CogSci | 2 |
| 2019 | Multitask Item Response Models for Response Bias Removal from Affective RatingsabstractResponse style (RS) is a tendency to choose specific categories regardless of content, e.g, extreme or midpoint categories. It degrades the validity of the analysis of subjective ratings such as correlation and variance-based analyses. However, the computational removal of RS has received little attention from the affective computing community. RS removal techniques have been proposed in areas such as marketing research. However, most of these techniques do not exploit the content-independence of RS; i.e. it should be observed consistently in various tasks, such as affective judgment tasks and standard psychological questionnaires. Therefore, this paper proposes a multitask RS removal method. An individual's responses in multiple tasks are modeled using task-independent RS parameters, and task-dependent parameters, including the item and respondent's characteristic parameters based on item response models (IRM). Through Bayesian modeling, we observed that: i) the proposed model outperformed traditional IRMs in terms of predictive accuracy; ii) our multitask framework estimated RS with higher precision than previous single-task-based RS removal methods; iii) our model replicated Japanese midpoint RS, which has been demonstrated repeatedly in previous cross-cultural studies; and iv) RS-removed predictive ratings showed higher inter-rater agreement than those including RS in valence/arousal judgment tasks. Shiro Kumano, Keishi Nomura |
ACII | 1 |
| 2019 | The Invisible Potential of Facial Electromyography: A Comparison of EMG and Computer Vision when Distinguishing Posed from Spontaneous SmilesabstractPositive experiences are a success metric in product and service design. Quantifying smiles is a method of assessing them continuously. Smiles are usually a cue of positive affect, but they can also be fabricated voluntarily. Automatic detection is a promising complement to human perception in terms of identifying the differences between smile types. Computer vision (CV) and facial distal electromyography (EMG) have been proven successful in this task. This is the first study to use a wearable EMG that does not obstruct the face to compare the performance of CV and EMG measurements in the task of distinguishing between posed and spontaneous smiles. The results showed that EMG has the advantage of being able to identify covert behavior not available through vision. Moreover, CV appears to be able to identify visible dynamic features that human judges cannot account for. This sheds light on the role of non-observable behavior in distinguishing affect-related smiles from polite positive affect displays. Monica Perusquía-Hernández, Saho Ayabe-Kanamura, Kenji Suzuki 0002, Shiro Kumano |
CHI | 4 |
| 2019 | Bayesian Item Response Model with Condition-specific Parameters for Evaluating the Differential Effects of Perspective-taking on Emotional Sharing
Keishi Nomura, Aiko Murata, Yuko Yotsumoto, Shiro Kumano |
CogSci | 4 |
| 2018 | Analyzing Gaze Behavior and Dialogue Act during Turn-taking for Estimating Empathy Skill LevelabstractWe explored the gaze behavior towards the end of utterances and dialogue act (DA), i.e., verbal-behavior information indicating the intension of an utterance, during turn-keeping/changing to estimate empathy skill levels in multiparty discussions. This is the first attempt to explore the relationship between such a combination. First, we collected data on Davis' Interpersonal Reactivity Index (which measures empathy skill level), utterances that include the DA categories of Provision, Self-disclosure, Empathy, Turn-yielding, and Others, and gaze behavior from participants in four-person discussions. The results of analysis indicate that the gaze behavior accompanying utterances that include these DA categories during turn-keeping/changing differs in accordance with people's empathy skill levels. The most noteworthy result was that speakers with low empathy skill levels tend to avoid making eye contact with the listener when the DA category is Self-disclosure during turn-keeping. However, they tend to maintain eye contact when the DA category is Empathy. A listener who has a high empathy skill level often looks away from the speaker during turn-changing when the DA category of a speaker's utterance is Provision or Empathy. There was also no difference in gaze behavior between empathy skill levels when the DA category of the speaker's utterance was turn-yielding. From these findings, we constructed and evaluated models for estimating empathy skill level using gaze behavior and DA information. The evaluation results indicate that using both gaze behavior and DA during turn-keeping/changing is effective for estimating an individual's empathy skill level in multi-party discussions. Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Ryuichiro Higashinaka, Junji Tomita |
ICMI | 3 |
| 2017 | Computational model of idiosyncratic perception of others' emotionsabstractThis paper deals with computational modelling for predicting the idiosyncratic perception of others' emotions, namely how individual external observers will score the emotional states of others interacting with each other. We separately model the observer effect (or individual differences of observers), and the conversational-scene effect or the video-clip effect (how interlocutors are interacting), based on Bayes' theorem with the assumption of their conditional independence. The observer term describes the observer's cognitive tendency, including bias, in a probabilistic form, and does not include any clip information. In contrast, the clip term describes how a target clip is recognized by an unspecified observer. The perceived emotion is predicted to be the state that maximizes the conditional probability given the observer and target clip. An experiment with 100 observers and 97 clips demonstrated, in a leave-one-out cross-validation scenario, that 1) there is in fact no statistically and practically significant interaction between observer and clip, and 2) our Bayesian modelling achieves a 97 percent accuracy as a reference of test-retest reliability. Furthermore, when combined with existing observer and clip models that can handle unknown observers and clips, our model yielded an accuracy of around 50 percent in a more challenging leave-one-subject-and-clip-out cross-validation scenario. Shiro Kumano, Ryo Ishii, Kazuhiro Otsuka |
ACII | 1 |
| 2017 | Comparing empathy perceived by interlocutors in multiparty conversation and external observersabstractThis paper investigates the basic characteristics of perceived empathy in Breithaupt's three-person model to consider a way of realizing its automatic prediction or empathy reading machines. More specifically, we report the extent to which interlocutors differ from external observers in perceiving the empathy aroused during group conversation. We also evaluate the accuracies of various frequently used models, including majority voting and multiple regression, in predicting an interlocutor's ratings from those of other interlocutors and/or those of the observers. Defining empathy as the emotional congruence between pairs of interlocutors, we studied a four-person conversation, in which previously unacquainted people held a decision-making discussion. We used a 5-point Likert scale when collecting self-reports of empathy from the interlocutors, and reports from a total of forty external observers (ten for each interlocutor) who adopted a target interlocutor's perspective. We obtained three indications. First, when no empathy ratings are available from the target interlocutor for model training, it is beneficial to ask observers to take the target interlocutor's perspective. Second, when target interlocutors' self-reports are available, it is advantageous to instruct observers not to take the target interlocutor's perspective. Third, in both scenarios, it is useful to ask interlocutors to rate the pairs excluding themselves. These findings provide some insights into good rating procedure as regards studying perceived empathy. Shiro Kumano, Ryo Ishii, Kazuhiro Otsuka |
ACII | 1 |
| 2017 | Prediction of Next-Utterance Timing using Head Movement in Multi-Party MeetingsabstractTo build a conversational interface wherein an agent system can smoothly communicate with multiple persons, it is imperative to know how the timing of speaking is decided. In this research, we explore the head movements of participants as an easy-to-measure nonverbal behavior to predict the nest-utterance timing, i.e., the interval between the end of the current speaker's utterance and the start of the next speaker's utterance, in turn-changing in multi-party meetings. First, we collected data on participants' six degree-of-freedom head movements and utterances in four-person meetings. The results of the analysis revealed that the amount of head movements of current speaker, next speaker, and listeners have a positive correlation with the utterance interval. Moreover, the degree of synchrony of the head position and posture between the current speaker and next speaker is negatively correlated with the utterance interval. On the basis of these findings, we used their head movements and the synchrony of their head movements as feature values and devised several prediction models. A model using all features performed the best and was able to predict the next-utterance timing well. Therefore, this research revealed that the participants' head movement is useful for predicting the next-utterance timing in turn-changing in multi-party meetings. Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka |
HAI | 2 |
| 2017 | Analyzing gaze behavior during turn-taking for estimating empathy skill levelabstractTechniques that use nonverbal behaviors to estimate communication skill in discussions have been receiving a lot of attention in recent research. In this study, we explored the gaze behavior towards the end of an utterance during turn-keeping/changing to estimate empathy skills in multiparty discussions. First, we collected data on Davis' Interpersonal Reactivity Index (IRI) (which measures empathy skill), utterances, and gaze behavior from participants in four-person discussions. The results of the analysis showed that the gaze behavior during turn-keeping/changing differs in accordance with people's empathy skill levels. The most noteworthy result is that the amount of a person's empathy skill is inversely proportional to the frequency of eye contact with the conversational partner during turn-keeping/changing. Specifically, if the current speaker has a high skill level, she often does not look at listener during turn-keeping and turn-changing. Moreover, when a person with a high skill level is the next speaker, she does not look at the speaker during turn-changing. In contrast, people who have a low skill level often continue to make eye contact with speakers and listeners. On the basis of these findings, we constructed and evaluated four models for estimating empathy skill levels. The evaluation results showed that the average absolute error of estimation is only 0.22 for the gaze transition pattern (GTP) model. This model uses the occurrence probability of GTPs when the person is a speaker and listener during turn-keeping and speaker, next-speaker, and listener during turn-changing. It outperformed the models that used the amount of utterances and duration of gazes. This suggests that the GTP during turn-keeping and turn-changing is effective for estimating an individual's empathy skills in multi-party discussions. Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka |
ICMI | 2 |
| 2017 | Collective First-Person Vision for Automatic Gaze Analysis in Multiparty ConversationsabstractThis paper targets smallto medium-sized-group face-to-face conversations where each person wears a dual-view camera, consisting of inwardand outward-looking cameras, and presents an almost fully automatic but accurate ofline gaze analysis framework that does not require users to perform any calibration steps. Our collective first-person vision framework, where captured audio-visual signals are gathered and processed in a centralized system, jointly undertakes the fundamental functions required for group gaze analysis, including speaker detection, face tracking, and gaze tracking. Of particular note is our self-calibration of gaze trackers by exploiting a general conversation rule, namely that listeners are likely to look at the speaker. From the rough conversational prior knowledge, our system visualizes fine-grained participants' gaze behavior as a gazee-centered heat map, which quantitatively reveals what parts of the gazee's body the participant looked at and for how long while the gazer was speaking or listening. An experiment using conversations amounting to a total of 140 min, each lasting an average of 8.7 min and engaged in by 37 participants in groups of three to six, achieves a mean absolute error of 2.8° in gaze tracking. A statistical test reveals neither a group size effect nor a conversation type effect. Our method achieves F-scores of over 0.89 and 0.87 in gazee and eye contact recognition, respectively, in comparison with human annotation. Shiro Kumano, Kazuhiro Otsuka, Ryo Ishii, Junji Yamato |
IEEE Trans. Multim. | 1 |
| 2016 | Analyzing mouth-opening transition pattern for predicting next speaker in multi-party meetingsabstractTechniques that use nonverbal behaviors to predict turn-changing situations—e.g., predicting who will speak next and when, in multi-party meetings—have been receiving a lot of attention in recent research. In this research, we explored the transition pattern of the degree of mouth opening (MOTP) towards the end of an utterance to predict the next speaker in multiparty meetings. First, we collected data on utterances and on the degree of mouth opening (closed, slightly open, and wide open) from participants in four-person meetings. The representative results of the analysis of the MOTPs showed that the speaker often continues to open the mouth slightly in turn-keeping and starts to close the mouth from opening it slightly or continues to open the mouth largely in turn-changing. The next speaker often starts to open the mouth slightly from closing it in turn-changing. On the basis of these findings, we constructed next-speaker prediction models using the MOTPs. In addition, as a multimodal fusion, we constructed models using the MOTPs and gaze information, which is known to be one of the most useful types of information for the next-speaker prediction. The evaluation of the models suggests that the speaker's and listeners' MOTPs are effective for predicting the next speaker in multi-party meetings. It also suggests the multimodal fusion using the MOTP and gaze information is more useful for the prediction than using one or the other. Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka |
ICMI | 2 |
| 2016 | Prediction of Who Will Be the Next Speaker and When Using Gaze Behavior in Multiparty MeetingsabstractIn multiparty meetings, participants need to predict the end of the speaker’s utterance and who will start speaking next, as well as consider a strategy for good timing to speak next. Gaze behavior plays an important role in smooth turn-changing. This article proposes a prediction model that features three processing steps to predict (I) whether turn-changing or turn-keeping will occur, (II) who will be the next speaker in turn-changing, and (III) the timing of the start of the next speaker’s utterance. For the feature values of the model, we focused on gaze transition patterns and the timing structure of eye contact between a speaker and a listener near the end of the speaker’s utterance. Gaze transition patterns provide information about the order in which gaze behavior changes. The timing structure of eye contact is defined as who looks at whom and who looks away first, the speaker or listener, when eye contact between the speaker and a listener occurs. We collected corpus data of multiparty meetings, using the data to demonstrate relationships between gaze transition patterns and timing structure and situations (I), (II), and (III). The results of our analyses indicate that the gaze transition pattern of the speaker and listener and the timing structure of eye contact have a strong association with turn-changing, the next speaker in turn-changing, and the start time of the next utterance. On the basis of the results, we constructed prediction models using the gaze transition patterns and timing structure. The gaze transition patterns were found to be useful in predicting turn-changing, the next speaker in turn-changing, and the start time of the next utterance. Contrary to expectations, we did not find that the timing structure is useful for predicting the next speaker and the start time. This study opens up new possibilities for predicting the next speaker and the timing of the next utterance using gaze transition patterns in multiparty meetings. Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2016 | Using Respiration to Predict Who Will Speak Next and When in Multiparty MeetingsabstractTechniques that use nonverbal behaviors to predict turn-changing situations—such as, in multiparty meetings, who the next speaker will be and when the next utterance will occur—have been receiving a lot of attention in recent research. To build a model for predicting these behaviors we conducted a research study to determine whether respiration could be effectively used as a basis for the prediction. Results of analyses of utterance and respiration data collected from participants in multiparty meetings reveal that the speaker takes a breath more quickly and deeply after the end of an utterance in turn-keeping than in turn-changing. They also indicate that the listener who will be the next speaker takes a bigger breath more quickly and deeply in turn-changing than the other listeners. On the basis of these results, we constructed and evaluated models for predicting the next speaker and the time of the next utterance in multiparty meetings. The results of the evaluation suggest that the characteristics of the speaker's inhalation right after an utterance unit—the points in time at which the inhalation starts and ends after the end of the utterance unit and the amplitude, slope, and duration of the inhalation phase—are effective for predicting the next speaker in multiparty meetings. They further suggest that the characteristics of listeners' inhalation—the points in time at which the inhalation starts and ends after the end of the utterance unit and the minimum and maximum inspiration, amplitude, and slope of the inhalation phase—are effective for predicting the next speaker. The start time and end time of the next speaker's inhalation are also useful for predicting the time of the next utterance in turn-changing. Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato |
ACM Trans. Interact. Intell. Syst. | 3 |
| 2015 | Predicting next speaker based on head movement in multi-party meetingsabstractWe proposed a model for predicting the next speaker in multi-party meetings by focusing on the participants' head movements measured by using a six degrees-of-freedom head tracker. Results of an analysis of head movements collected from multi-party meetings revealed differences in the amounts, amplitude, and frequency of movement of the head position and rotation of the speaker near the end of an utterance in turn-keeping and turn-taking. The results also revealed the differences in the amounts of movement, amplitude, and frequency of head position movement and rotation between the listeners in turn-keeping, turn-taking, and the next speaker in turn-taking. We then built a next speaker prediction model that features two processing steps to predict whether turn-taking or turn-keeping will occur and who the next speaker will be in turn-taking. The evaluation results for the model suggest that the speaker's and listeners' head movements contribute to predicting the next speaker. Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka |
ICASSP | 2 |
| 2015 | Multimodal Fusion using Respiration and Gaze for Predicting Next Speaker in Multi-Party MeetingsabstractTechniques that use nonverbal behaviors to predict turn-taking situations, such as who will be the next speaker and the next utterance timing in multi-party meetings are receiving a lot of attention recently. It has long been known that gaze is a physical behavior that plays an important role in transferring the speaking turn between humans. Recently, a line of research has focused on the relationship between turn-taking and respiration, a biological signal that conveys information about the intention or preliminary action to start to speak. It has been demonstrated that respiration and gaze behavior separately have the potential to allow predicting the next speaker and the next utterance timing in multi-party meetings. As a multimodal fusion to create models for predicting the next speaker in multi-party meetings, we integrated respiration and gaze behavior, which were extracted from different modalities and are completely different in quality, and implemented a model uses information about them to predict the next speaker at the end of an utterance. The model has a two-step processing. The first is to predict whether turn-keeping or turn-taking happens; the second is to predict the next speaker in turn-taking. We constructed prediction models with either respiration or gaze behavior and with both respiration and gaze behaviors as features and compared their performance. The results suggest that the model with both respiration and gaze behaviors performs better than the one using only respiration or gaze behavior. It is revealed that multimodal fusion using respiration and gaze behavior is effective for predicting the next speaker in multi-party meetings. It was found that gaze behavior is more useful for predicting turn-keeping/turn-taking than respiration and that respiration is more useful for predicting the next speaker in turn-taking. Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka |
ICMI | 2 |
| 2015 | Analyzing Interpersonal Empathy via Collective ImpressionsabstractThis paper presents a research framework for understanding the empathy that arises between people while they are conversing. By focusing on the process by which empathy is perceived by other people, this paper aims to develop a computational model that automatically infers perceived empathy from participant behavior. To describe such perceived empathy objectively, we introduce the idea of using the collective impressions of external observers. In particular, we focus on the fact that the perception of other's empathy varies from person to person, and take the standpoint that this individual difference itself is an essential attribute of human communication for building, for example, successful human relationships and consensus. This paper describes a probabilistic model of the process that we built based on the Bayesian network, and that relates the empathy perceived by observers to how the gaze and facial expressions of participants co-occur between a pair. In this model, the probability distribution represents the diversity of observers' impression, which reflects the individual differences in the schema when perceiving others' empathy from their behaviors, and the ambiguity of the behaviors. Comprehensive experiments demonstrate that the inferred distributions are similar to those made by observers. Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Masafumi Matsuda, Junji Yamato |
IEEE Trans. Affect. Comput. | 1 |
| 2015 | In the Mood for Vlog: Multimodal Inference in Conversational Social VideoabstractThe prevalent “share what's on your mind” paradigm of social media can be examined from the perspective of mood: short-term affective states revealed by the shared data. This view takes on new relevance given the emergence of conversational social video as a popular genre among viewers looking for entertainment and among video contributors as a channel for debate, expertise sharing, and artistic expression. From the perspective of human behavior understanding, in conversational social video both verbal and nonverbal information is conveyed by speakers and decoded by viewers. We present a systematic study of classification and ranking of mood impressions in social video, using vlogs from YouTube. Our approach considers eleven natural mood categories labeled through crowdsourcing by external observers on a diverse set of conversational vlogs. We extract a comprehensive number of nonverbal and verbal behavioral cues from the audio and video channels to characterize the mood of vloggers. Then we implement and validate vlog classification and vlog ranking tasks using supervised learning methods. Following a reliability and correlation analysis of the mood impression data, our study demonstrates that, while the problem is challenging, several mood categories can be inferred with promising performance. Furthermore, multimodal features perform consistently better than single-channel features. Finally, we show that addressing mood as a ranking problem is a promising practical direction for several of the mood categories studied. Dairazalia Sanchez-Cortes, Shiro Kumano, Kazuhiro Otsuka, Daniel Gatica-Perez |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2014 | Analysis and modeling of next speaking start timing based on gaze behavior in multi-party meetingsabstractTo realize a conversational interface where an agent system can smoothly communicate with multiple persons, it is imperative to know how the start timing of speaking is decided. In this research, we demonstrate a relationship between gaze transition patterns and the start timing of next speaking against the end of the last speaking in multi-party meetings. Then, we construct a prediction model for the start timing using gaze transition patterns near the end of an utterance. An analysis of data collected from natural multi-party meetings reveals a strong relationship between gaze transition patterns of the speaker, next speaker, and listener and the start timing of the next speaker. On the basis of the results, we used gaze transition patterns of the speaker, next speaker, and listener and mutual gaze as variables, and devised several prediction models. A model using all features performed the best and was able to predict the start timing well. Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato |
ICASSP | 3 |
| 2014 | Analysis of Respiration for Prediction of "Who Will Be Next Speaker and When?" in Multi-Party MeetingsabstractTo build a model for predicting the next speaker and the start time of the next utterance in multi-party meetings, we performed a fundamental study of how respiration could be effective for the prediction model. The results of the analysis reveal that a speaker inhales more rapidly and quickly right after the end of a unit of utterance in turn-keeping. The next speaker takes a bigger breath toward speaking in turn-changing than listeners who will not become the next speaker. Based on the results of the analysis, we constructed the prediction models to evaluate how effective the parameters are. The results of the evaluation suggest that the speaker's inhalation right after a unit of utterance, such as the start time from the end of the unit of utterance and the slope and duration of the inhalation phase, is effective for predicting whether turn-keeping or turn-changing happen about 350 ms before the start time of the next utterance on average and that listener's inhalation before the next utterance, such as the maximal inspiration and amplitude of the inhalation phase, is effective for predicting the next speaker in turn-changing about 900 ms before the start time of the next utterance on average. Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato |
ICMI | 3 |
| 2013 | Using a Probabilistic Topic Model to Link Observers' Perception Tendency to PersonalityabstractTargeting multiparty conversations, the present study aims to elucidate how an observer will tend to perceive others' emotional states, develops a computational model that realizes the automatic inferencing of the observer's perception tendency. This paper proposes a probabilistic model that automatically discovers the correlation between perception tendency, gender, and personality traits of a target observer. Perception tendency, a probability distribution, explains how likely the observer is to perceive a certain state/level of a target emotion. Personality traits are measured by a variety of questionnaires. The proposed model links these three factors via a latent variable and explains observer's characteristics as a mixture of prototypical characters. An experiment is conducted with fifty observers. They watch 97 short conversation videos and give their impressions about the empathy between each interacting pair. The results demonstrate that the proposed method can find a reasonable framework that underlies the factors: e.g. 1) people who have high scores in Davis's empathy measures show empathy-biased response tendency, and 2) people who have strong sense of consideration for others tend to show an extreme response tendency, and such people are likely to be females. The proposed method shows promise in estimating an observer's perception tendency from his/her gender and personality traits, even when the target perception tendency is quite different from the average perception tendency among observers. Shiro Kumano, Kazuhiro Otsuka, Masafumi Matsuda, Ryo Ishii, Junji Yamato |
ACII | 1 |
| 2013 | Predicting next speaker and timing from gaze transition patterns in multi-party meetingsabstractIn multi-party meetings, participants need to predict the end of the speaker's utterance and who will start speaking next, and to consider a strategy for good timing to speak next. Gaze behavior plays an important role for smooth turn-taking. This paper proposes a mathematical prediction model that features three processing steps to predict (I) whether turn-taking or turn-keeping will occur, (II) who will be the next speaker in turn-taking, and (III) the timing of the start of the next speaker's utterance. For the feature quantity of the model, we focused on gaze transition patterns near the end of utterance. We collected corpus data of multi party meetings and analyzed how the frequencies of appearance of gaze transition patterns differs depending on situations of (I), (II), and (III). On the basis of the analysis, we construct a probabilistic mathematical model that uses the frequencies of appearance of all participants' gaze transition patterns. The results of an evaluation of the model show the proposed models succeed with high precision compared to ones that do not take gaze transition patterns into account. Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Masafumi Matsuda, Junji Yamato |
ICMI | 3 |
| 2013 | MM+Space: n x 4 degree-of-freedom kinetic display for recreating multiparty conversation spacesabstractA novel system, called MM+Space, is presented for recreating multiparty face-to-face conversation scenes in the real world. It aims to display and playback pre-recorded conversations as if the people were talking in front of the viewer(s). This system consists of multiple projectors and transparent screens, which display the life-size faces of people. The key idea is the physical augmentation of human head motions, i.e. the screen pose is dynamically controlled to emulate the head motions, for boosting the viewers' perception of nonverbal behaviors and interactions. In particular, MM+Space newly introduces 2-Degree-of-Freedom (DoF) translations, in forward-backward and right-left directions, in addition to 2-DoF head rotations (nodding and shaking), which were proposed in our former MM-Space system. The full 4-DoF kinetic display is expected to enhance the expressibility of head and body motions, and to create more realistic representation of interacting people. Experiments showed that the proposed system with 4-DoF motions outperformed the rotation-only system in the increased perception of people's presence and in expressing their postures. In addition, it was reported that the proposed system allowed the viewers to experience rich emotional expressibility, immersion in conversations, and potential behavioral/emotional contagion. Kazuhiro Otsuka, Shiro Kumano, Ryo Ishii, Maja Zbogar, Junji Yamato |
ICMI | 2 |
| 2013 | Inferring mood in ubiquitous conversational videoabstractConversational social video is becoming a worldwide trend. Video communication allows a more natural interaction, when aiming to share personal news, ideas, and opinions, by transmitting both verbal content and nonverbal behavior. However, the automatic analysis of natural mood is challenging, since it is displayed in parallel via voice, face, and body. This paper presents an automatic approach to infer 11 natural mood categories in conversational social video using single and multimodal nonverbal cues extracted from video blogs (vlogs) from YouTube. The mood labels used in our work were collected via crowdsourcing. Our approach is promising for several of the studied mood categories. Our study demonstrates that although multimodal features perform better than single channel features, not always all the available channels are needed to accurately discriminate mood in videos. Dairazalia Sanchez-Cortes, Joan-Isaac Biel, Shiro Kumano, Junji Yamato, Kazuhiro Otsuka, Daniel Gatica-Perez |
MUM | 3 |
| 2011 | Analyzing empathetic interactions based on the probabilistic modeling of the co-occurrence patterns of facial expressions in group meetingsabstractThis paper presents a novel research framework for the estimation of emotional interactions produced between meeting participants. The types of emotional interaction targeted in this paper are empathy, antipathy, and unconcern. We define here emotional interaction as a brief contiguous event wherein a pair exchange emotional messages via verbal and non-verbal behaviors. As the key behaviors, we focus on facial expression and gaze, because their combination realizes the rapid and directed transmission of a large number of emotional messages. We assume that there is a strong link between the emotional interaction and the participants' facial expressions that occur simultaneously with the type of the emotional interactions. Based on this assumption, we build a probabilistic model that represents a hierarchical structure involving the emotional interactions, facial expressions and other behaviors including utterance and gaze direction. Using this model, the type of emotional interaction is estimated from interpersonal gaze directions, facial expressions, and utterances. Our estimation is based on the Bayesian approach, and uses the Markov chain Monte Carlo method to approximate joint posterior probability distributions of the emotional interaction and model parameters present within the observed data. An experiment on four-party conversations demonstrates the promising effectiveness of the proposed method. Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Junji Yamato |
FG | 1 |
| 2011 | A system for reconstructing multiparty conversation field based on augmented head motion by dynamic projectionabstractA novel system is presented for reconstructing, in the real world, multiparty face-to-face conversation scenes; it uses dynamics projection to augment human head motion. This system aims to display and playback pre-recorded conversations to the viewers as if the remote people were taking in front of them. This system consists of multiple projectors and transparent screens. Each screen separately displays the life-size face of one meeting participant, and are spatially arranged to recreate the actual scene. The main feature of this system is dynamics projection, screen pose is dynamically controlled to emulate the head motions of the participants, especially rotation around the vertical axis, that are typical of shifts in visual attention, i.e. turning gaze from one to another. This recreation of head motion by physical screen motion, in addition to image motion, aims to more clearly express the interactions involving visual attention among the participants. The minimal design, frameless-projector-screen, with augmented head motion is expected to create a feeling that the remote participants are actually present in the same room. This demo presents our initial system and discusses its potential impact on future visual communications. Kazuhiro Otsuka, Kamil Sebastian Mucha, Shiro Kumano, Dan Mikami, Masafumi Matsuda, Junji Yamato |
ACM Multimedia | 3 |
| 2011 | Early facial expression recognition with high-frame rate 3D sensingabstractThis work investigates a new challenging problem: how to exactly recognize facial expression as early as possible, while most works generally focus on improving the recognition rate of facial expression recognition. The features of facial expressions in their early stage are unfortunately very sensitive to noise due to their low intensity. So, we propose a novel wavelet spectral subtraction method to spatio-temporally refine the subtle facial expression features. Moreover, in order to achieve early facial expression recognition, we newly introduce an early AdaBoost algorithm for facial expression recognition problem. Experiments using our database established by using a high-frame rate 3D sensing showed that the proposed method has a promising performance on early facial expression recognition. Lumei Su, Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Junji Yamato, Yoichi Sato 0001 |
SMC | 2 |
| 2009 | Recognizing communicative facial expressions for discovering interpersonal emotions in group meetingsabstractThis paper proposes a novel facial expression recognizer and describes its application to group meeting analysis. Our goal is to automatically discover the interpersonal emotions that evolve over time in meetings, e.g. how each person feels about the others, or who affectively influences the others the most. As the emotion cue, we focus on facial expression, more specifically smile, and aim to recognize ``who is smiling at whom, when, and how often'', since frequently smiling carries affective messages that are strongly directed to the person being looked at; this point of view is our novelty. To detect such communicative smiles, we propose a new algorithm that jointly estimates facial pose and expression in the framework of the particle filter. The main feature is its automatic selection of interest points that can robustly capture small changes in expression even in the presence of large head rotations. Based on the recognized facial expressions and their directions to others, which are indicated by the estimated head poses, we visualize interpersonal smile events as a graph structure, we call it the interpersonal emotional network; it is intended to indicate the emotional relationships among meeting participants. A four-person meeting captured by an omnidirectional video system is used to confirm the effectiveness of the proposed method and the potential of our approach for deep understanding of human relationships developed through communications. Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Junji Yamato |
ICMI | 1 |
| 2009 | Pose-Invariant Facial Expression Recognition Using Variable-Intensity Templates
Shiro Kumano, Kazuhiro Otsuka, Junji Yamato, Eisaku Maeda, Yoichi Sato 0001 |
Int. J. Comput. Vis. | 1 |
| 2008 | Combining Stochastic and Deterministic Search for Pose-Invariant Facial Expression RecognitionabstractWe propose a novel method for pose-invariant facial expression recognition from monocular video sequences that combines stochastic and determinis-tic search processes. We use the simple face model called variable-intensity template, which can be prepared with very little time and effort. We tackle the two issues found in previous work on the variable-intensity template: low accuracy in head pose estimation, and assumption violations due to external intensity changes such as illumination change. We mitigate these issues by introducing the deterministic approach into the stochastic approach imple-mented as a particle filter. Our experiment demonstrates significant improve-ments in recognition performance for horizontal and vertical head orienta-tions in the range of ±40 degrees and ±20 degrees, respectively, from the frontal view. 1 Shiro Kumano, Kazuhiro Otsuka, Junji Yamato, Eisaku Maeda, Yoichi Sato 0001 |
BMVC | 1 |
| 2007 | Pose-Invariant Facial Expression Recognition Using Variable-Intensity Templates
Shiro Kumano, Kazuhiro Otsuka, Junji Yamato, Eisaku Maeda, Yoichi Sato 0001 |
ACCV (1) | 1 |