Kazuhiro Otsuka

dblp:56/5734 · DBLP profile ↗
← Back
55ranked-venue papers
13as first author
6since 2021 · last 2025
0000-0003-4352-3955ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 29 · 7 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 7 first-authorArtificial intelligence and machine learning · 16 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Disentangling Perceptual Ambiguity in Multifunctional Nonverbal Behaviors in Conversations via Tensor Spectrum Decomposition
Issa Tamura, Momoka Tajima, Shiro Kumano, Kazuhiro Otsuka
ICMI4
2025 Synergistic Functional Spectrum Analysis: A Framework for Exploring the Multifunctional Interplay Among Multimodal Nonverbal Behaviours in Conversations
abstract
A novel framework named thesynergistic functional spectrum analysis(sFSA) is proposed to explore the multifunctional interplay among multimodal nonverbal behaviours in human conversations. This study aims to reveal how multimodal nonverbal behaviours cooperatively perform communicative functions in conversations. To capture the intrinsic nature of nonverbal expressions, functional multiplicity, and interpretational ambiguity, e.g., a single head nod could imply listening, agreeing, or both, a novel concept named thefunctional spectrum, which is defined as the distribution of perceptual intensities of multiple functions by multiple observers, is introduced in the sFSA. Based on this concept, this paper presentsfunctional spectrum corpora, which target 44 facial expression and 32 head movement functions. Then, spectrum decomposition is conducted to reduce the multimodal functional spectrum to asynergetic functional spectrumin a lower dimensionfunctional spacethat is spanned byfunctional basisvectors representing primary and distinctive functionalities across multiple modalities. To that end, we propose a semiorthogonal nonnegative matrix factorization (SO-NMF) method, which assumes the additivity of multiple functions and aims to balance the distinctiveness and expressiveness of the factorization. The results confirm that some primary functional bases can be identified, which can be interpreted as the listener’s backchannel, thinking, and affirmative response functions, and the speaker’s thinking and addressing functions, and their positive emotion functions. In addition, regression models based on convolutional neural networks (CNNs) are presented to estimate thesynergistic functional spectrumfrom the head poses and facial action units measured from conversation data. The results of these analyses and experiments confirm the potential of the sFSA and may lead to future extensions.
Mai Imamura, Ayane Tashiro, Shiro Kumano, Kazuhiro Otsuka
IEEE Trans. Affect. Comput.4
2024 Exploring Interlocutor Gaze Interactions in Conversations based on Functional Spectrum Analysis
abstract
A novel framework named a gaze interactional functional spectrum analysis (GI-FSA) is proposed to explore the functional aspects of gaze interactions among interlocutors in conversations. It aims to reveal the primary and distinctive interactional functionalities that emerge via the gaze behaviors of the speaker, and the listener whom the speaker looks at. To capture the intrinsic nature of gaze functions, such as multiple functionalities and ambiguity, this study introduces a novel representation called a gaze functional spectrum representing the distribution of perceptual intensity of multiple gaze functions and presents a gaze functional spectrum corpus that targets 43 gaze functions covering various speech-related, listening-related and other functions. Then, semiorthogonal nonnegative matrix factorization (SO-NMF) is employed to decompose the concatenated speaker-listener functional spectra into a interactional functional spectrum in a lower-dimensional functional space spanned with functional bases, each of which represents a distinct aspect of interactional functionalities. Targeting four female conversations, the GI-FSA revealed interpretable functional bases such as addressing-listening and joint positive emotion. In addition, this paper proposes convolutional neural networks (CNNs) that can recognize the binary level of the interactional functional spectrum from observable multimodal nonverbal behaviors, including head pose, utterance status, eyeball direction and facial expressions. These experimental findings validate the potential of the GI-FSA as a promising framework for analyzing gaze interactions among interlocutors, and understanding communication dynamics.
Ayane Tashiro, Mai Imamura, Shiro Kumano, Kazuhiro Otsuka
ICMI4
2021 Deep Transfer Learning for Recognizing Functional Interactions via Head Movements in Multiparty Conversations
abstract
Head movements play various functions in multiparty conversations. To date, convolutional neural networks (CNNs) have been proposed to recognize the functions of individual interlocutors’ head movements. This paper extends the concept of head-movement functions to the interaction functions between speaker and listener, which are performed through their head movements, e.g., a listener’s back-channel nodding in response to a speaker’s rhythmic movements. Then, we propose transfer strategies to build deep neural networks (DNNs) to recognize these interaction functions by reusing pretrained CNNs for individual head-movement functions. One of the proposed strategies uses CNNs as the feature extractor and identifies the interaction function with another classifier using the extracted features. Compared with the baseline model that employs the logical product of the output of two individual CNNs, the transferred DNNs outperform the baseline model in four out of five interaction functions. For example, the F-measure is improved by 13.9 points for the interaction of a listener’s positive emotion in response to a speaker’s rhythmic movements. These results confirm the potential of the proposed transfer strategies for recognizing interaction functions based on head movements.
Takashi Mori, Kazuhiro Otsuka
ICMI2
2021 Prediction of Interlocutors' Subjective Impressions Based on Functional Head-Movement Features in Group Meetings
abstract
A novel model is proposed to predict interlocutors’ subjective impressions during group meetings based on their head movements. The goal is to explore the potential of the communicative functions performed through head movements. To this end, this paper focuses on ten frequent functions, such as speaker emphasis and listener back-channel responses, which are detected using convolutional neural networks (CNNs) from head-pose sequences. Regarding the prediction target, this study employs four items of subjective impressions—atmosphere, enjoyment, willingness, and concentration—which are scored at every two-minute interval by the interlocutors themselves in four-party meetings with 17 groups. From the detected head-movement functions, this paper newly defines the functional features, including function occurrence rates and composition ratios, in addition to the kinetic features representing head-movement activity. Using these features as the input, random forest regressors predict the impression scores. Compared with the baseline model using only kinetic features, the prediction model using both kinetic and functional features improved prediction performance. The percentage of moderate or higher correlations (≥ 0.5) between reported scores and predictions increased from 25% to 41% in all groups and all items. These results suggest that head-movement functions could be effective cues for predicting interlocutors’ subjective impressions.
Shumpei Otsuchi, Yoko Ishii, Momoko Nakatani, Kazuhiro Otsuka
ICMI4
2021 Inflation-Deflation Networks for Recognizing Head-Movement Functions in Face-to-Face Conversations
abstract
Head movements have various functions in face-to-face conversations. Recently, convolutional neural networks (CNNs) have been proposed to recognize the communicative functions performed by the head movements from the time series of interlocutors’ head pose angles during multiparty conversations. However, there is room for improvement in the recognition performance. To boost the CNNs’ performance, this paper proposes a feature Inflation-Deflation module (I/DeF module) as an additional module attached ahead of the CNNs’ input layer to facilitate the feature learning of the head-movement dynamics. The I/DeF module consists of repeated inflation and deflation processes. The inflation process upscales and extrapolates the windowed input time series by a transposed convolution. The deflation process compresses the inflated data and recovers its original data length. Targeting the ten frequent head-movement functions, the experiments showed that CNNs with the I/DeF module (I/DeF-CNNs) outperformed the previous CNNs in all function categories up to 4.5 points in F-measure. We also integrated the I/DeF module into VGG and ResNet. Comparison to these methods showed that I/DeF-CNNs surpassed the other models for 8 out of 10 functions. These results confirmed the effectiveness of the I/DeF module and its potential for advancing nonverbal behavior recognition.
Kazuaki Takeda 0001, Kazuhiro Otsuka
ICMI2
2018 Analyzing Gaze Behavior and Dialogue Act during Turn-taking for Estimating Empathy Skill Level
abstract
We explored the gaze behavior towards the end of utterances and dialogue act (DA), i.e., verbal-behavior information indicating the intension of an utterance, during turn-keeping/changing to estimate empathy skill levels in multiparty discussions. This is the first attempt to explore the relationship between such a combination. First, we collected data on Davis' Interpersonal Reactivity Index (which measures empathy skill level), utterances that include the DA categories of Provision, Self-disclosure, Empathy, Turn-yielding, and Others, and gaze behavior from participants in four-person discussions. The results of analysis indicate that the gaze behavior accompanying utterances that include these DA categories during turn-keeping/changing differs in accordance with people's empathy skill levels. The most noteworthy result was that speakers with low empathy skill levels tend to avoid making eye contact with the listener when the DA category is Self-disclosure during turn-keeping. However, they tend to maintain eye contact when the DA category is Empathy. A listener who has a high empathy skill level often looks away from the speaker during turn-changing when the DA category of a speaker's utterance is Provision or Empathy. There was also no difference in gaze behavior between empathy skill levels when the DA category of the speaker's utterance was turn-yielding. From these findings, we constructed and evaluated models for estimating empathy skill level using gaze behavior and DA information. The evaluation results indicate that using both gaze behavior and DA during turn-keeping/changing is effective for estimating an individual's empathy skill level in multi-party discussions.
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Ryuichiro Higashinaka, Junji Tomita
ICMI2
2018 Estimating Visual Focus of Attention in Multiparty Meetings using Deep Convolutional Neural Networks
abstract
Convolutional neural networks (CNNs) are employed to estimate the visual focus of attention (VFoA), also called gaze direction , in multiparty face-to-face meetings on the basis of multimodal nonverbal behaviors including head pose, direction of the eyeball, and presence/absence of utterance. To reveal the potential of CNNs, we focus on aspects of multimodal and multiparty fusion including individual/group models, early/late fusion, and robustness when using inputs from image-based trackers. In contrast to the individual model that separately targets each person specific to one's seat, the group model aims to jointly estimate the gaze directions of all participants. Experiments confirmed that the group model outperformed the individual model especially in predicting listeners' VFoA when the inputs did not include eyeball directions. This result indicates that the group CNN model can implicitly learn underlying conversation structures, e.g., the listeners' gazes converge on the speaker. When the eyeball direction feature is available, both models outperformed the Bayes models used for comparison. In this case, the individual model was superior to the group model, particularly in estimating the speaker's VFoA. Moreover, it was revealed that in group models, two-stage late fusion, which integrates an individual features first, and multiparty features second, outperformed other structures. Furthermore, our experiment confirmed that image-based tracking can provide a comparable level of performance to that of sensor-based measurements. Overall, the results suggest that the CNN is a promising approach for VFoA estimation.
Kazuhiro Otsuka, Keisuke Kasuga, Martina Köhler
ICMI1
2018 Vlogging Over Time: Longitudinal Impressions and Behavior in YouTube
abstract
YouTube vlogging, as a popular genre of ubiquitous social video, engages people in entertainment, civic, and social activities. Although several aspects of vlogging have been studied in media studies and multimedia analysis, the longitudinal angle of vlogging regarding recognition of personal state and trait impressions from behavior has not been yet analyzed. We present a study using behavioral data of vloggers who posted vlogs on YouTube for a period between three and six years. We use online crowdsourcing to collect a rich set of 21 impression variables for each video, including perceived personality, mood, skills, and expertise. Acoustic and motion features are extracted to characterize basic nonverbal behavior. The analysis shows that only a couple of perceived variables, including perceived expertise and perceived quality of audio and video, display weak temporal patterns. Furthermore, we show that the use of longitudinal data helps to improve the automatic inference of impressions for several of the impression variables.
Daniel Gatica-Perez, Dairazalia Sanchez-Cortes, Trinh Minh Tri Do, Dinesh Babu Jayagopi, Kazuhiro Otsuka
MUM5
2018 Behavioral Analysis of Kinetic Telepresence for Small Symmetric Group-to-Group Meetings
abstract
Nonverbal behavior analysis revealed the effect of MMSpace, a kinetic telepresence developed for social telepresence, on small symmetric group-to-group conversations. MMSpace consists of kinetic avatars, equipped with flat projection screen panels as faces, that can change their pose and position automatically to mirror the remote user's head motions. The advantage is the realistic kinetic expression of human head movements, which form gestures like nodding and indicate the focus of visual attention, through the use of four degree-of-freedom low-latency precision actuators. Another feature is the support of eye contact among remote participants, which is made possible by the avatar's kinetic pose changes and by adaptive camera selection for orienting the user's face toward the remote addressee. Its limitation is its room-scale infrastructure and restricted participant positions. Targeting a symmetric 2 × 2 setting, participants' nonverbal behaviors, including gaze directions and head gestures, were compared among three conditions, MMSpace with/without physical motions and face-to-face settings. There was a significant difference between the conditions in terms of the duration of glance/mutual glances, total gaze transition time, amount of head gesturing, and co-occurrences of head gestures in the remote participants. The results indicate that the avatar's physical motion can elicit longer (mutual) glances with a shorter total transition time and more (co-)occurrences of head gestures, and it makes MMSpace -based conversations closer, in terms of these nonverbal statistics, to face-to-face ones compared with those of a static version of MMSpace without physical motion.
Kazuhiro Otsuka
IEEE Trans. Multim.1
2017 Computational model of idiosyncratic perception of others' emotions
abstract
This paper deals with computational modelling for predicting the idiosyncratic perception of others' emotions, namely how individual external observers will score the emotional states of others interacting with each other. We separately model the observer effect (or individual differences of observers), and the conversational-scene effect or the video-clip effect (how interlocutors are interacting), based on Bayes' theorem with the assumption of their conditional independence. The observer term describes the observer's cognitive tendency, including bias, in a probabilistic form, and does not include any clip information. In contrast, the clip term describes how a target clip is recognized by an unspecified observer. The perceived emotion is predicted to be the state that maximizes the conditional probability given the observer and target clip. An experiment with 100 observers and 97 clips demonstrated, in a leave-one-out cross-validation scenario, that 1) there is in fact no statistically and practically significant interaction between observer and clip, and 2) our Bayesian modelling achieves a 97 percent accuracy as a reference of test-retest reliability. Furthermore, when combined with existing observer and clip models that can handle unknown observers and clips, our model yielded an accuracy of around 50 percent in a more challenging leave-one-subject-and-clip-out cross-validation scenario.
Shiro Kumano, Ryo Ishii, Kazuhiro Otsuka
ACII3
2017 Comparing empathy perceived by interlocutors in multiparty conversation and external observers
abstract
This paper investigates the basic characteristics of perceived empathy in Breithaupt's three-person model to consider a way of realizing its automatic prediction or empathy reading machines. More specifically, we report the extent to which interlocutors differ from external observers in perceiving the empathy aroused during group conversation. We also evaluate the accuracies of various frequently used models, including majority voting and multiple regression, in predicting an interlocutor's ratings from those of other interlocutors and/or those of the observers. Defining empathy as the emotional congruence between pairs of interlocutors, we studied a four-person conversation, in which previously unacquainted people held a decision-making discussion. We used a 5-point Likert scale when collecting self-reports of empathy from the interlocutors, and reports from a total of forty external observers (ten for each interlocutor) who adopted a target interlocutor's perspective. We obtained three indications. First, when no empathy ratings are available from the target interlocutor for model training, it is beneficial to ask observers to take the target interlocutor's perspective. Second, when target interlocutors' self-reports are available, it is advantageous to instruct observers not to take the target interlocutor's perspective. Third, in both scenarios, it is useful to ask interlocutors to rate the pairs excluding themselves. These findings provide some insights into good rating procedure as regards studying perceived empathy.
Shiro Kumano, Ryo Ishii, Kazuhiro Otsuka
ACII3
2017 Recognizing Words from Gestures: Discovering Gesture Descriptors Associated with Spoken Utterances
abstract
This study investigates a new challenge: modelling the relationship between gestures and spoken words during continuous hand motions, speech and language data. The problem setting and modelling is defined as “gesture association” (associating gestures with words). We present a framework to associate spoken words with hand motion data observed from a optical motion-capture system. The framework identifies pairs of hand motions (gesture) and uttered words as training samples by autonomously aligning the speech and the co-occurring gestures. Using the samples, a supervised learning approach is undertaken to learn a model to discriminate between gestures that co-occur with a spoken word (w) and gestures when the word w is not being spoken. To detect gestures and associate them with spoken words, we extract (1) a trajectory signal feature set, (2) features of various gesture phases from hand motion data and (3) primitive gesture patterns learned by Sift Invariant Sparse Coding (SISC). Then Hidden Markov Model (HMM) and Support Vector Machine (SVM) classifiers were trained with these three-feature sets. In experiments, the classification accuracy achieved for 80 words was above 60 % (the maximum accuracy achieved was 71 %) using the proposed framework. In particular, SVMs trained with features (2) and (3) outperform an HMM trained with feature (1) (prepared as a baseline model). The results show that the gesture phase and primitive patterns trained by SISC are effective gesture features for recognizing the words accompanying gestures.
Shogo Okada, Kazuhiro Otsuka
FG2
2017 Prediction of Next-Utterance Timing using Head Movement in Multi-Party Meetings
abstract
To build a conversational interface wherein an agent system can smoothly communicate with multiple persons, it is imperative to know how the timing of speaking is decided. In this research, we explore the head movements of participants as an easy-to-measure nonverbal behavior to predict the nest-utterance timing, i.e., the interval between the end of the current speaker's utterance and the start of the next speaker's utterance, in turn-changing in multi-party meetings. First, we collected data on participants' six degree-of-freedom head movements and utterances in four-person meetings. The results of the analysis revealed that the amount of head movements of current speaker, next speaker, and listeners have a positive correlation with the utterance interval. Moreover, the degree of synchrony of the head position and posture between the current speaker and next speaker is negatively correlated with the utterance interval. On the basis of these findings, we used their head movements and the synchrony of their head movements as feature values and devised several prediction models. A model using all features performed the best and was able to predict the next-utterance timing well. Therefore, this research revealed that the participants' head movement is useful for predicting the next-utterance timing in turn-changing in multi-party meetings.
Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka
HAI3
2017 Analyzing gaze behavior during turn-taking for estimating empathy skill level
abstract
Techniques that use nonverbal behaviors to estimate communication skill in discussions have been receiving a lot of attention in recent research. In this study, we explored the gaze behavior towards the end of an utterance during turn-keeping/changing to estimate empathy skills in multiparty discussions. First, we collected data on Davis' Interpersonal Reactivity Index (IRI) (which measures empathy skill), utterances, and gaze behavior from participants in four-person discussions. The results of the analysis showed that the gaze behavior during turn-keeping/changing differs in accordance with people's empathy skill levels. The most noteworthy result is that the amount of a person's empathy skill is inversely proportional to the frequency of eye contact with the conversational partner during turn-keeping/changing. Specifically, if the current speaker has a high skill level, she often does not look at listener during turn-keeping and turn-changing. Moreover, when a person with a high skill level is the next speaker, she does not look at the speaker during turn-changing. In contrast, people who have a low skill level often continue to make eye contact with speakers and listeners. On the basis of these findings, we constructed and evaluated four models for estimating empathy skill levels. The evaluation results showed that the average absolute error of estimation is only 0.22 for the gaze transition pattern (GTP) model. This model uses the occurrence probability of GTPs when the person is a speaker and listener during turn-keeping and speaker, next-speaker, and listener during turn-changing. It outperformed the models that used the amount of utterances and duration of gazes. This suggests that the GTP during turn-keeping and turn-changing is effective for estimating an individual's empathy skills in multi-party discussions.
Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka
ICMI3
2017 Collective First-Person Vision for Automatic Gaze Analysis in Multiparty Conversations
abstract
This paper targets smallto medium-sized-group face-to-face conversations where each person wears a dual-view camera, consisting of inwardand outward-looking cameras, and presents an almost fully automatic but accurate ofline gaze analysis framework that does not require users to perform any calibration steps. Our collective first-person vision framework, where captured audio-visual signals are gathered and processed in a centralized system, jointly undertakes the fundamental functions required for group gaze analysis, including speaker detection, face tracking, and gaze tracking. Of particular note is our self-calibration of gaze trackers by exploiting a general conversation rule, namely that listeners are likely to look at the speaker. From the rough conversational prior knowledge, our system visualizes fine-grained participants' gaze behavior as a gazee-centered heat map, which quantitatively reveals what parts of the gazee's body the participant looked at and for how long while the gazer was speaking or listening. An experiment using conversations amounting to a total of 140 min, each lasting an average of 8.7 min and engaged in by 37 participants in groups of three to six, achieves a mean absolute error of 2.8° in gaze tracking. A statistical test reveals neither a group size effect nor a conversation type effect. Our method achieves F-scores of over 0.89 and 0.87 in gazee and eye contact recognition, respectively, in comparison with human annotation.
Shiro Kumano, Kazuhiro Otsuka, Ryo Ishii, Junji Yamato
IEEE Trans. Multim.2
2016 Analyzing mouth-opening transition pattern for predicting next speaker in multi-party meetings
abstract
Techniques that use nonverbal behaviors to predict turn-changing situations—e.g., predicting who will speak next and when, in multi-party meetings—have been receiving a lot of attention in recent research. In this research, we explored the transition pattern of the degree of mouth opening (MOTP) towards the end of an utterance to predict the next speaker in multiparty meetings. First, we collected data on utterances and on the degree of mouth opening (closed, slightly open, and wide open) from participants in four-person meetings. The representative results of the analysis of the MOTPs showed that the speaker often continues to open the mouth slightly in turn-keeping and starts to close the mouth from opening it slightly or continues to open the mouth largely in turn-changing. The next speaker often starts to open the mouth slightly from closing it in turn-changing. On the basis of these findings, we constructed next-speaker prediction models using the MOTPs. In addition, as a multimodal fusion, we constructed models using the MOTPs and gaze information, which is known to be one of the most useful types of information for the next-speaker prediction. The evaluation of the models suggests that the speaker's and listeners' MOTPs are effective for predicting the next speaker in multi-party meetings. It also suggests the multimodal fusion using the MOTP and gaze information is more useful for the prediction than using one or the other.
Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka
ICMI3
2016 MMSpace: Kinetically-augmented telepresence for small group-to-group conversations
abstract
A novel research prototype, called MMSpace, was developed for realistic social telepresence in small group-to-group conversations. MMSpace consists of kinetic display avatars, which can change pose and position by automatically mirroring the remote user's head motions. To fully explore its potential beyond previous alternatives, MMSpace has the following novel features. First, it targets symmetric group-to-group telepresence. Second, the kinetic avatars of MMSpace can produce highly accurate, low latency, and silent physical motions, by using 4-Degree-of-Freedom (DoF) direct-drive actuators, and they can express a wide range of natural human behaviors like head gestures and changing attitudes, as well as indicating the focus of attention. Third, MMSpace supports eye contact between every pair of participants, by integrating i) directional visual attention cues indicated by avatar's kinetic pose change, ii) line-of-sight alignment among the positions of persons, avatars and cameras, and iii) attention-based camera switching, which allows an avatar to always show its owner's face looking directly toward the person that the avatar's owner is looking at. The prototype targets the 2 × 2 setting, and subjective evaluations based on group discussions indicate that the kinetic display avatar is superior to static displays in various aspects including gaze-awareness, eye-contact, perception of other nonverbal behaviors, mutual understanding, and sense of telepresence.
Kazuhiro Otsuka
VR1
2016 Prediction of Who Will Be the Next Speaker and When Using Gaze Behavior in Multiparty Meetings
abstract
In multiparty meetings, participants need to predict the end of the speaker’s utterance and who will start speaking next, as well as consider a strategy for good timing to speak next. Gaze behavior plays an important role in smooth turn-changing. This article proposes a prediction model that features three processing steps to predict (I) whether turn-changing or turn-keeping will occur, (II) who will be the next speaker in turn-changing, and (III) the timing of the start of the next speaker’s utterance. For the feature values of the model, we focused on gaze transition patterns and the timing structure of eye contact between a speaker and a listener near the end of the speaker’s utterance. Gaze transition patterns provide information about the order in which gaze behavior changes. The timing structure of eye contact is defined as who looks at whom and who looks away first, the speaker or listener, when eye contact between the speaker and a listener occurs. We collected corpus data of multiparty meetings, using the data to demonstrate relationships between gaze transition patterns and timing structure and situations (I), (II), and (III). The results of our analyses indicate that the gaze transition pattern of the speaker and listener and the timing structure of eye contact have a strong association with turn-changing, the next speaker in turn-changing, and the start time of the next utterance. On the basis of the results, we constructed prediction models using the gaze transition patterns and timing structure. The gaze transition patterns were found to be useful in predicting turn-changing, the next speaker in turn-changing, and the start time of the next utterance. Contrary to expectations, we did not find that the timing structure is useful for predicting the next speaker and the start time. This study opens up new possibilities for predicting the next speaker and the timing of the next utterance using gaze transition patterns in multiparty meetings.
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato
ACM Trans. Interact. Intell. Syst.2
2016 Using Respiration to Predict Who Will Speak Next and When in Multiparty Meetings
abstract
Techniques that use nonverbal behaviors to predict turn-changing situations—such as, in multiparty meetings, who the next speaker will be and when the next utterance will occur—have been receiving a lot of attention in recent research. To build a model for predicting these behaviors we conducted a research study to determine whether respiration could be effectively used as a basis for the prediction. Results of analyses of utterance and respiration data collected from participants in multiparty meetings reveal that the speaker takes a breath more quickly and deeply after the end of an utterance in turn-keeping than in turn-changing. They also indicate that the listener who will be the next speaker takes a bigger breath more quickly and deeply in turn-changing than the other listeners. On the basis of these results, we constructed and evaluated models for predicting the next speaker and the time of the next utterance in multiparty meetings. The results of the evaluation suggest that the characteristics of the speaker's inhalation right after an utterance unit—the points in time at which the inhalation starts and ends after the end of the utterance unit and the amplitude, slope, and duration of the inhalation phase—are effective for predicting the next speaker in multiparty meetings. They further suggest that the characteristics of listeners' inhalation—the points in time at which the inhalation starts and ends after the end of the utterance unit and the minimum and maximum inspiration, amplitude, and slope of the inhalation phase—are effective for predicting the next speaker. The start time and end time of the next speaker's inhalation are also useful for predicting the time of the next utterance in turn-changing.
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato
ACM Trans. Interact. Intell. Syst.2
2015 Predicting next speaker based on head movement in multi-party meetings
abstract
We proposed a model for predicting the next speaker in multi-party meetings by focusing on the participants' head movements measured by using a six degrees-of-freedom head tracker. Results of an analysis of head movements collected from multi-party meetings revealed differences in the amounts, amplitude, and frequency of movement of the head position and rotation of the speaker near the end of an utterance in turn-keeping and turn-taking. The results also revealed the differences in the amounts of movement, amplitude, and frequency of head position movement and rotation between the listeners in turn-keeping, turn-taking, and the next speaker in turn-taking. We then built a next speaker prediction model that features two processing steps to predict whether turn-taking or turn-keeping will occur and who the next speaker will be in turn-taking. The evaluation results for the model suggest that the speaker's and listeners' head movements contribute to predicting the next speaker.
Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka
ICASSP3
2015 Multimodal Fusion using Respiration and Gaze for Predicting Next Speaker in Multi-Party Meetings
abstract
Techniques that use nonverbal behaviors to predict turn-taking situations, such as who will be the next speaker and the next utterance timing in multi-party meetings are receiving a lot of attention recently. It has long been known that gaze is a physical behavior that plays an important role in transferring the speaking turn between humans. Recently, a line of research has focused on the relationship between turn-taking and respiration, a biological signal that conveys information about the intention or preliminary action to start to speak. It has been demonstrated that respiration and gaze behavior separately have the potential to allow predicting the next speaker and the next utterance timing in multi-party meetings. As a multimodal fusion to create models for predicting the next speaker in multi-party meetings, we integrated respiration and gaze behavior, which were extracted from different modalities and are completely different in quality, and implemented a model uses information about them to predict the next speaker at the end of an utterance. The model has a two-step processing. The first is to predict whether turn-keeping or turn-taking happens; the second is to predict the next speaker in turn-taking. We constructed prediction models with either respiration or gaze behavior and with both respiration and gaze behaviors as features and compared their performance. The results suggest that the model with both respiration and gaze behaviors performs better than the one using only respiration or gaze behavior. It is revealed that multimodal fusion using respiration and gaze behavior is effective for predicting the next speaker in multi-party meetings. It was found that gaze behavior is more useful for predicting turn-keeping/turn-taking than respiration and that respiration is more useful for predicting the next speaker in turn-taking.
Ryo Ishii, Shiro Kumano, Kazuhiro Otsuka
ICMI3
2015 Design and Evaluation of Mirror Interface MIOSS to Overlay Remote 3D Spaces
Ryo Ishii, Shiro Ozawa, Akira Kojima, Kazuhiro Otsuka, Yuki Hayashi, Yukiko I. Nakano
INTERACT (4)4
2015 Analyzing Interpersonal Empathy via Collective Impressions
abstract
This paper presents a research framework for understanding the empathy that arises between people while they are conversing. By focusing on the process by which empathy is perceived by other people, this paper aims to develop a computational model that automatically infers perceived empathy from participant behavior. To describe such perceived empathy objectively, we introduce the idea of using the collective impressions of external observers. In particular, we focus on the fact that the perception of other's empathy varies from person to person, and take the standpoint that this individual difference itself is an essential attribute of human communication for building, for example, successful human relationships and consensus. This paper describes a probabilistic model of the process that we built based on the Bayesian network, and that relates the empathy perceived by observers to how the gaze and facial expressions of participants co-occur between a pair. In this model, the probability distribution represents the diversity of observers' impression, which reflects the individual differences in the schema when perceiving others' empathy from their behaviors, and the ambiguity of the behaviors. Comprehensive experiments demonstrate that the inferred distributions are similar to those made by observers.
Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Masafumi Matsuda, Junji Yamato
IEEE Trans. Affect. Comput.2
2015 In the Mood for Vlog: Multimodal Inference in Conversational Social Video
abstract
The prevalent “share what's on your mind” paradigm of social media can be examined from the perspective of mood: short-term affective states revealed by the shared data. This view takes on new relevance given the emergence of conversational social video as a popular genre among viewers looking for entertainment and among video contributors as a channel for debate, expertise sharing, and artistic expression. From the perspective of human behavior understanding, in conversational social video both verbal and nonverbal information is conveyed by speakers and decoded by viewers. We present a systematic study of classification and ranking of mood impressions in social video, using vlogs from YouTube. Our approach considers eleven natural mood categories labeled through crowdsourcing by external observers on a diverse set of conversational vlogs. We extract a comprehensive number of nonverbal and verbal behavioral cues from the audio and video channels to characterize the mood of vloggers. Then we implement and validate vlog classification and vlog ranking tasks using supervised learning methods. Following a reliability and correlation analysis of the mood impression data, our study demonstrates that, while the problem is challenging, several mood categories can be inferred with promising performance. Furthermore, multimodal features perform consistently better than single-channel features. Finally, we show that addressing mood as a ranking problem is a promising practical direction for several of the mood categories studied.
Dairazalia Sanchez-Cortes, Shiro Kumano, Kazuhiro Otsuka, Daniel Gatica-Perez
ACM Trans. Interact. Intell. Syst.3
2014 Analysis and modeling of next speaking start timing based on gaze behavior in multi-party meetings
abstract
To realize a conversational interface where an agent system can smoothly communicate with multiple persons, it is imperative to know how the start timing of speaking is decided. In this research, we demonstrate a relationship between gaze transition patterns and the start timing of next speaking against the end of the last speaking in multi-party meetings. Then, we construct a prediction model for the start timing using gaze transition patterns near the end of an utterance. An analysis of data collected from natural multi-party meetings reveals a strong relationship between gaze transition patterns of the speaker, next speaker, and listener and the start timing of the next speaker. On the basis of the results, we used gaze transition patterns of the speaker, next speaker, and listener and mutual gaze as variables, and devised several prediction models. A model using all features performed the best and was able to predict the start timing well.
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato
ICASSP2
2014 Analysis of Respiration for Prediction of "Who Will Be Next Speaker and When?" in Multi-Party Meetings
abstract
To build a model for predicting the next speaker and the start time of the next utterance in multi-party meetings, we performed a fundamental study of how respiration could be effective for the prediction model. The results of the analysis reveal that a speaker inhales more rapidly and quickly right after the end of a unit of utterance in turn-keeping. The next speaker takes a bigger breath toward speaking in turn-changing than listeners who will not become the next speaker. Based on the results of the analysis, we constructed the prediction models to evaluate how effective the parameters are. The results of the evaluation suggest that the speaker's inhalation right after a unit of utterance, such as the start time from the end of the unit of utterance and the slope and duration of the inhalation phase, is effective for predicting whether turn-keeping or turn-changing happen about 350 ms before the start time of the next utterance on average and that listener's inhalation before the next utterance, such as the maximal inspiration and amplitude of the inhalation phase, is effective for predicting the next speaker in turn-changing about 900 ms before the start time of the next utterance on average.
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato
ICMI2
2013 Using a Probabilistic Topic Model to Link Observers' Perception Tendency to Personality
abstract
Targeting multiparty conversations, the present study aims to elucidate how an observer will tend to perceive others' emotional states, develops a computational model that realizes the automatic inferencing of the observer's perception tendency. This paper proposes a probabilistic model that automatically discovers the correlation between perception tendency, gender, and personality traits of a target observer. Perception tendency, a probability distribution, explains how likely the observer is to perceive a certain state/level of a target emotion. Personality traits are measured by a variety of questionnaires. The proposed model links these three factors via a latent variable and explains observer's characteristics as a mixture of prototypical characters. An experiment is conducted with fifty observers. They watch 97 short conversation videos and give their impressions about the empathy between each interacting pair. The results demonstrate that the proposed method can find a reasonable framework that underlies the factors: e.g. 1) people who have high scores in Davis's empathy measures show empathy-biased response tendency, and 2) people who have strong sense of consideration for others tend to show an extreme response tendency, and such people are likely to be females. The proposed method shows promise in estimating an observer's perception tendency from his/her gender and personality traits, even when the target perception tendency is quite different from the average perception tendency among observers.
Shiro Kumano, Kazuhiro Otsuka, Masafumi Matsuda, Ryo Ishii, Junji Yamato
ACII2
2013 Predicting next speaker and timing from gaze transition patterns in multi-party meetings
abstract
In multi-party meetings, participants need to predict the end of the speaker's utterance and who will start speaking next, and to consider a strategy for good timing to speak next. Gaze behavior plays an important role for smooth turn-taking. This paper proposes a mathematical prediction model that features three processing steps to predict (I) whether turn-taking or turn-keeping will occur, (II) who will be the next speaker in turn-taking, and (III) the timing of the start of the next speaker's utterance. For the feature quantity of the model, we focused on gaze transition patterns near the end of utterance. We collected corpus data of multi party meetings and analyzed how the frequencies of appearance of gaze transition patterns differs depending on situations of (I), (II), and (III). On the basis of the analysis, we construct a probabilistic mathematical model that uses the frequencies of appearance of all participants' gaze transition patterns. The results of an evaluation of the model show the proposed models succeed with high precision compared to ones that do not take gaze transition patterns into account.
Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Masafumi Matsuda, Junji Yamato
ICMI2
2013 MM+Space: n x 4 degree-of-freedom kinetic display for recreating multiparty conversation spaces
abstract
A novel system, called MM+Space, is presented for recreating multiparty face-to-face conversation scenes in the real world. It aims to display and playback pre-recorded conversations as if the people were talking in front of the viewer(s). This system consists of multiple projectors and transparent screens, which display the life-size faces of people. The key idea is the physical augmentation of human head motions, i.e. the screen pose is dynamically controlled to emulate the head motions, for boosting the viewers' perception of nonverbal behaviors and interactions. In particular, MM+Space newly introduces 2-Degree-of-Freedom (DoF) translations, in forward-backward and right-left directions, in addition to 2-DoF head rotations (nodding and shaking), which were proposed in our former MM-Space system. The full 4-DoF kinetic display is expected to enhance the expressibility of head and body motions, and to create more realistic representation of interacting people. Experiments showed that the proposed system with 4-DoF motions outperformed the rotation-only system in the increased perception of people's presence and in expressing their postures. In addition, it was reported that the proposed system allowed the viewers to experience rich emotional expressibility, immersion in conversations, and potential behavioral/emotional contagion.
Kazuhiro Otsuka, Shiro Kumano, Ryo Ishii, Maja Zbogar, Junji Yamato
ICMI1
2013 Inferring mood in ubiquitous conversational video
abstract
Conversational social video is becoming a worldwide trend. Video communication allows a more natural interaction, when aiming to share personal news, ideas, and opinions, by transmitting both verbal content and nonverbal behavior. However, the automatic analysis of natural mood is challenging, since it is displayed in parallel via voice, face, and body. This paper presents an automatic approach to infer 11 natural mood categories in conversational social video using single and multimodal nonverbal cues extracted from video blogs (vlogs) from YouTube. The mood labels used in our work were collected via crowdsourcing. Our approach is promising for several of the studied mood categories. Our study demonstrates that although multimodal features perform better than single channel features, not always all the available channels are needed to accurately discriminate mood in videos.
Dairazalia Sanchez-Cortes, Joan-Isaac Biel, Shiro Kumano, Junji Yamato, Kazuhiro Otsuka, Daniel Gatica-Perez
MUM5
2012 Linking speaking and looking behavior patterns with group composition, perception, and performance
abstract
This paper addresses the task of mining typical behavioral patterns from small group face-to-face interactions and linking them to social-psychological group variables. Towards this goal, we define group speaking and looking cues by aggregating automatically extracted cues at the individual and dyadic levels. Then, we define a bag of nonverbal patterns (Bag-of-NVPs) to discretize the group cues. The topics learnt using the Latent Dirichlet Allocation (LDA) topic model are then interpreted by studying the correlations with group variables such as group composition, group interpersonal perception, and group performance. Our results show that both group behavior cues and topics have significant correlations with (and predictive information for) all the above variables. For our study, we use interactions with unacquainted members i.e. newly formed groups.
Dinesh Babu Jayagopi, Dairazalia Sanchez-Cortes, Kazuhiro Otsuka, Junji Yamato, Daniel Gatica-Perez
ICMI3
2012 Low-Latency Real-Time Meeting Recognition and Understanding Using Distant Microphones and Omni-Directional Camera
abstract
This paper presents our real-time meeting analyzer for monitoring conversations in an ongoing group meeting. The goal of the system is to recognize automatically “who is speaking what” in an online manner for meeting assistance. Our system continuously captures the utterances and face poses of each speaker using a microphone array and an omni-directional camera positioned at the center of the meeting table. Through a series of advanced audio processing operations, an overlapping speech signal is enhanced and the components are separated into individual speaker's channels. Then the utterances are sequentially transcribed by our speech recognizer with low latency. In parallel with speech recognition, the activity of each participant (e.g., speaking, laughing, watching someone) and the circumstances of the meeting (e.g., topic, activeness, casualness) are detected and displayed on a browser together with the transcripts. In this paper, we describe our techniques and our attempt to achieve the low-latency monitoring of meetings, and we show our experimental results for real-time meeting transcription.
Takaaki Hori, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto, Shinji Watanabe 0001, Takanobu Oba, Atsunori Ogawa, Kazuhiro Otsuka, Dan Mikami, Keisuke Kinoshita, Tomohiro Nakatani, Atsushi Nakamura, Junji Yamato
IEEE Trans. Speech Audio Process.8
2011 Analyzing empathetic interactions based on the probabilistic modeling of the co-occurrence patterns of facial expressions in group meetings
abstract
This paper presents a novel research framework for the estimation of emotional interactions produced between meeting participants. The types of emotional interaction targeted in this paper are empathy, antipathy, and unconcern. We define here emotional interaction as a brief contiguous event wherein a pair exchange emotional messages via verbal and non-verbal behaviors. As the key behaviors, we focus on facial expression and gaze, because their combination realizes the rapid and directed transmission of a large number of emotional messages. We assume that there is a strong link between the emotional interaction and the participants' facial expressions that occur simultaneously with the type of the emotional interactions. Based on this assumption, we build a probabilistic model that represents a hierarchical structure involving the emotional interactions, facial expressions and other behaviors including utterance and gaze direction. Using this model, the type of emotional interaction is estimated from interpersonal gaze directions, facial expressions, and utterances. Our estimation is based on the Bayesian approach, and uses the Markov chain Monte Carlo method to approximate joint posterior probability distributions of the emotional interaction and model parameters present within the observed data. An experiment on four-party conversations demonstrates the promising effectiveness of the proposed method.
Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Junji Yamato
FG2
2011 A system for reconstructing multiparty conversation field based on augmented head motion by dynamic projection
abstract
A novel system is presented for reconstructing, in the real world, multiparty face-to-face conversation scenes; it uses dynamics projection to augment human head motion. This system aims to display and playback pre-recorded conversations to the viewers as if the remote people were taking in front of them. This system consists of multiple projectors and transparent screens. Each screen separately displays the life-size face of one meeting participant, and are spatially arranged to recreate the actual scene. The main feature of this system is dynamics projection, screen pose is dynamically controlled to emulate the head motions of the participants, especially rotation around the vertical axis, that are typical of shifts in visual attention, i.e. turning gaze from one to another. This recreation of head motion by physical screen motion, in addition to image motion, aims to more clearly express the interactions involving visual attention among the participants. The minimal design, frameless-projector-screen, with augmented head motion is expected to create a feeling that the remote participants are actually present in the same room. This demo presents our initial system and discusses its potential impact on future visual communications.
Kazuhiro Otsuka, Kamil Sebastian Mucha, Shiro Kumano, Dan Mikami, Masafumi Matsuda, Junji Yamato
ACM Multimedia1
2011 Early facial expression recognition with high-frame rate 3D sensing
abstract
This work investigates a new challenging problem: how to exactly recognize facial expression as early as possible, while most works generally focus on improving the recognition rate of facial expression recognition. The features of facial expressions in their early stage are unfortunately very sensitive to noise due to their low intensity. So, we propose a novel wavelet spectral subtraction method to spatio-temporally refine the subtle facial expression features. Moreover, in order to achieve early facial expression recognition, we newly introduce an early AdaBoost algorithm for facial expression recognition problem. Experiments using our database established by using a high-frame rate 3D sensing showed that the proposed method has a promising performance on early facial expression recognition.
Lumei Su, Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Junji Yamato, Yoichi Sato 0001
SMC3
2010 Memory-Based Particle Filter for Tracking Objects with Large Variation in Pose and Appearance
Dan Mikami, Kazuhiro Otsuka, Junji Yamato
ECCV (3)2
2010 Real-time meeting recognition and understanding using distant microphones and omni-directional camera
abstract
This paper presents our newly developed real-time meeting analyzer for monitoring conversations in an ongoing group meeting. The goal of the system is to automatically recognize “who is speaking what” in an online manner for meeting assistance. Our system continuously captures the utterances and the face pose of each speaker using a distant microphone array and an omni-directional camera at the center of the meeting table. Through a series of advanced audio processing operations, an overlapping speech signal is enhanced and the components are separated into individual speaker's channels. Then the utterances are sequentially transcribed by our speech recognizer with low latency. In parallel with speech recognition, the activity of each participant (e.g. speaking, laughing, watching someone) and the situation of the meeting (e.g. topic, activeness, casualness) are detected and displayed on a browser together with the transcripts. In this paper, we describe our techniques and our attempt to achieve the low-latency monitoring of meetings, and we show our experimental results for real-time meeting transcription.
Takaaki Hori, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto, Shinji Watanabe 0001, Takanobu Oba, Atsunori Ogawa, Kazuhiro Otsuka, Dan Mikami, Keisuke Kinoshita, Tomohiro Nakatani, Atsushi Nakamura, Junji Yamato
SLT8
2009 Memory-based Particle Filter for face pose tracking robust under complex dynamics
abstract
A novel particle filter, the memory-based particle filter (M-PF), is proposed that can visually track moving objects that have complex dynamics. We aim to realize robustness against abrupt object movements and quick recovery from tracking failure caused by factors such as occlusions. To that end, we eliminate the Markov assumption from the previous particle filtering framework and predict the prior distribution of the target state from the long-term dynamics. More concretely, M-PF stores the past history of the estimated target states, and employs a random sampling from the history to generate prior distribution; it represents a novel PF formulation.Our method can handle nonlinear, time-variant, and non-Markov dynamics, which is not possible within existing PF frameworks. Accurate prior prediction based on proper dynamics model is especially effective for recovering lost tracks, because it can provide possible target states, which can drastically change since the track was lost. We target the face pose of seated humans in this paper. Quantitative evaluations with magnetic sensors confirm improved accuracy in face pose estimation and successful recovery from tracking loss. The proposed M-PF suggests a new paradigm for modeling systems with complex dynamics and so offers a various visual tracking applications.
Dan Mikami, Kazuhiro Otsuka, Junji Yamato
CVPR2
2009 A speaker diarization method based on the probabilistic fusion of audio-visual location information
abstract
This paper proposes a speaker diarization method for determining ""who spoke when"" in multi-party conversations, based on the probabilistic fusion of audio and visual location information. The audio and visual information is obtained from a compact system designed to analyze round table multi-party conversations. The system consists of two cameras and a triangular microphone array with three microphones, and can cover a spherical region. Speaker locations are estimated from audio and visual observations in terms of azimuths from this recording system. Unlike conventional speech diarization methods, our proposed method estimates the probability of the presence of multiple simultaneous speakers in a physical space with a small microphone setup instead of using a cascade consisting of speech activity detection, direction of arrival estimation, acoustic feature extraction, and information criteria based speaker segmentation. To estimate the speaker presence more correctly, the speech presence probabilities in a physical space are integrated with the probabilities estimated from participants' face locations obtained with a robust particle filtering based face tracker with two cameras equipped with fisheye lenses. The locations in a physical space with highly integrated probabilities are then classified into a certain number of speaker classes by using on-line classification to realize speaker diarization. The probability calculations and speaker classifications are conducted on-line, making it unnecessary to observe all the conversation data. An experiment using real casual conversations, which include more overlaps and short speech segments than formal meetings, showed the advantages of the proposed method.
Kentaro Ishizuka, Shoko Araki, Kazuhiro Otsuka, Tomohiro Nakatani, Masakiyo Fujimoto
ICMI3
2009 Recognizing communicative facial expressions for discovering interpersonal emotions in group meetings
abstract
This paper proposes a novel facial expression recognizer and describes its application to group meeting analysis. Our goal is to automatically discover the interpersonal emotions that evolve over time in meetings, e.g. how each person feels about the others, or who affectively influences the others the most. As the emotion cue, we focus on facial expression, more specifically smile, and aim to recognize ``who is smiling at whom, when, and how often'', since frequently smiling carries affective messages that are strongly directed to the person being looked at; this point of view is our novelty. To detect such communicative smiles, we propose a new algorithm that jointly estimates facial pose and expression in the framework of the particle filter. The main feature is its automatic selection of interest points that can robustly capture small changes in expression even in the presence of large head rotations. Based on the recognized facial expressions and their directions to others, which are indicated by the estimated head poses, we visualize interpersonal smile events as a graph structure, we call it the interpersonal emotional network; it is intended to indicate the emotional relationships among meeting participants. A four-person meeting captured by an omnidirectional video system is used to confirm the effectiveness of the proposed method and the potential of our approach for deep understanding of human relationships developed through communications.
Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Junji Yamato
ICMI2
2009 Realtime meeting analysis and 3D meeting viewer based on omnidirectional multimodal sensors
abstract
This demo presents a realtime system for analyzing group meetings. Targeting round-table meetings, this system employs an omnidirectional camera-microphone system. The goal of this system is to automatically discover "who is talking to whom and when". To that purpose, the face pose/position of meeting participants are tracked on panorama images acquired from fisheye-based omnidirectional cameras. From audio signals obtained with microphone array, speaker diarization, i.e. the estimation of "who is speaking and when", is carried out. The visual focus of attention, i.e. "who is looking at whom", is esimated from the result of face tracking. The results are displayed based on a 3D visualization scheme. The advantage of our system is its realtimeness. We will demonstrate the portable version of the system consisting of two laptop PCs. In addition, we will showcase our meeting playback viewer with man-machine interfaces that allow users to freely control space and time of meeting scenes. With this viewer, users can also experince 3D positional sound effect linked with 3D viewpoint, using enhanced audio tracks for each participant.
Kazuhiro Otsuka, Shoko Araki, Dan Mikami, Kentaro Ishizuka, Masakiyo Fujimoto, Junji Yamato
ICMI1
2009 Pose-Invariant Facial Expression Recognition Using Variable-Intensity Templates
Shiro Kumano, Kazuhiro Otsuka, Junji Yamato, Eisaku Maeda, Yoichi Sato 0001
Int. J. Comput. Vis.2
2008 Combining Stochastic and Deterministic Search for Pose-Invariant Facial Expression Recognition
abstract
We propose a novel method for pose-invariant facial expression recognition from monocular video sequences that combines stochastic and determinis-tic search processes. We use the simple face model called variable-intensity template, which can be prepared with very little time and effort. We tackle the two issues found in previous work on the variable-intensity template: low accuracy in head pose estimation, and assumption violations due to external intensity changes such as illumination change. We mitigate these issues by introducing the deterministic approach into the stochastic approach imple-mented as a particle filter. Our experiment demonstrates significant improve-ments in recognition performance for horizontal and vertical head orienta-tions in the range of ±40 degrees and ±20 degrees, respectively, from the frontal view. 1
Shiro Kumano, Kazuhiro Otsuka, Junji Yamato, Eisaku Maeda, Yoichi Sato 0001
BMVC2
2008 Simultaneous and fast 3D tracking of multiple faces in video by GPU-based stream processing
abstract
In this work, we implement a real-time visual tracker that targets the position and 3D pose of objects in video sequences, specifically faces. Using stream processors for performing the computations as well as efficient sparse-template-based particle filtering allows us to achieve real-time processing even when tracking multiple objects simultaneously in high- resolution video frames. Stream processing is a relatively new computing paradigm that permits the expression and execution of data-parallel algorithms with great efficiency and minimum effort. Using a GPU (graphics processing unit, a consumer-grade stream processor) and the NVIDIA CUDAtrade technology, we can achieve real-time performance even when tracking multiple objects in high-quality videos.
Oscar Mateo Lozano, Kazuhiro Otsuka
ICASSP2
2008 A realtime multimodal system for analyzing group meetings by combining face pose tracking and speaker diarization
abstract
This paper presents a realtime system for analyzing group meetings that uses a novel omnidirectional camera-microphone system. The goal is to automatically discover the visual focus of attention (VFOA), i.e. "who is looking at whom", in addition to speaker diarization, i.e. "who is speaking and when". First, a novel tabletop sensing device for round-table meetings is presented; it consists of two cameras with two fisheye lenses and a triangular microphone array. Second, from high-resolution omnidirectional images captured with the cameras, the position and pose of people's faces are estimated by STCTracker (Sparse Template Condensation Tracker); it realizes realtime robust tracking of multiple faces by utilizing GPUs (Graphics Processing Units). The face position/pose data output by the face tracker is used to estimate the focus of attention in the group. Using the microphone array, robust speaker diarization is carried out by a VAD (Voice Activity Detection) and a DOA (Direction of Arrival) estimation followed by sound source clustering. This paper also presents new 3-D visualization schemes for meeting scenes and the results of an analysis. Using two PCs, one for vision and one for audio processing, the system runs at about 20 frames per second for 5-person meetings.
Kazuhiro Otsuka, Shoko Araki, Kentaro Ishizuka, Masakiyo Fujimoto, Martin Heinrich, Junji Yamato
ICMI1
2007 Pose-Invariant Facial Expression Recognition Using Variable-Intensity Templates
Shiro Kumano, Kazuhiro Otsuka, Junji Yamato, Eisaku Maeda, Yoichi Sato 0001
ACCV (1)2
2007 Automatic inference of cross-modal nonverbal interactions in multiparty conversations: "who responds to whom, when, and how?" from gaze, head gestures, and utterances
abstract
A novel probabilistic framework is proposed for analyzing cross-modal nonverbal interactions in multiparty face-to-face conversations. The goal is to determine "who responds to whom, when, and how" from multimodal cues including gaze, head gestures, and utterances. We formulate this problem as the probabilistic inference of the causal relationship among participants' behaviors involving head gestures and utterances. To solve this problem, this paper proposes a hierarchical probabilistic model; the structures of interactions are probabilistically determined from high-level conversation regimes (such as monologue or dialogue) and gaze directions. Based on the model, the interaction structures, gaze, and conversation regimes, are simultaneously inferred from observed head motion and utterances, using a Markov chain Monte Carlo method. The head gestures, including nodding, shaking and tilt, are recognized with a novel Wavelet-based technique from magnetic sensor signals. The utterances are detected using data captured by lapel microphones. Experiments on four-person conversations confirm the effectiveness of the framework in discovering interactions such as question-and-answer and addressing behavior followed by back-channel responses.
Kazuhiro Otsuka, Hiroshi Sawada, Junji Yamato
ICMI1
2006 Conversation Scene Analysis with Dynamic Bayesian Network Basedon Visual Head Tracking
abstract
A novel method based on a probabilistic model for conversation scene analysis is proposed that can infer conversation structure from video sequences of face-to-face communication. Conversation structure represents the type of conversation such as monologue or dialogue, and can indicate who is talking/listening to whom. This study assumes that the gaze directions of participants provide cues for discerning the conversation structure, and can be identified from head directions. For measuring head directions, the proposed method newly employs a visual head tracker based on sparse-template condensation. The conversation model is built on a dynamic Bayesian network and is used to estimate the conversation structure and gaze directions from observed head directions and utterances. Visual tracking is conventionally thought to be less reliable than contact sensors, but experiments confirm that the proposed method achieves almost comparable performance in estimating gaze directions and conversation structure to a conventional sensor-based method
Kazuhiro Otsuka, Junji Yamato, Yoshinao Takemae, Hiroshi Murase
ICME1
2005 Effects of Automatic Video Editing System Using Stereo-Based Head Tracking for Archiving Meetings
abstract
This paper presents an automatic video editing system based on head tracking for archiving meetings. Systems that archive meetings are attracting considerable interest. Conventional systems use a fixed-viewpoint camera and simple camera selection based on participants' utterances. However, conventional systems fail to adequately convey who is talking to whom and nonverbal information about participants etc. We focus on the participants' head orientation since this information is useful in detecting the speaker and who the speaker is talking to. In order to automatically estimate each participant's head orientation, our system combines several modules to realize stereo-based head tracking. The system selects the shot of the participant that most participants are looking at, based on majority decision. Experiments on presenting videos to viewers confirm the effectiveness of our system in several 3-participant conversations
Yoshinao Takemae, Kazuhiro Otsuka, Junji Yamato
ICME2
2005 A probabilistic inference of multiparty-conversation structure based on Markov-switching models of gaze patterns, head directions, and utterances
abstract
A novel probabilistic framework is proposed for inferring the structure of conversation in face-to-face multiparty communication, based on gaze patterns, head directions and the presence/absence of utterances. As the structure of conversation, this study focuses on the combination of participants and their participation roles. First, we assess the gaze patterns that frequently appear in conversations, and define typical types of conversation structure, called conversational regime, and hypothesize that the regime represents the high-level process that governs how people interact during conversations. Next, assuming that the regime changes over time exhibit Markov properties, we propose a probabilistic conversation model based on Markov-switching; the regime controls the dynamics of utterances and gaze patterns, which stochastically yield measurable head-direction changes. Furthermore, a Gibbs sampler is used to realize the Bayesian estimation of regime, gaze pattern, and model parameters from observed head directions and utterances. Experiments on four-person conversations confirm the effectiveness of the framework in identifying conversation structures.
Kazuhiro Otsuka, Yoshinao Takemae, Junji Yamato
ICMI1
2004 Multiview Occlusion Analysis for Tracking Densely Populated Objects Based on 2-D Visual Angles
Kazuhiro Otsuka, Naoki Mukawa
CVPR (1)1
2003 Video cut editing rule based on participants' gaze in multiparty conversation
abstract
This paper proposes a video cut editing rule based on participants' gaze for extracting and conveying the flow of conversation in multiparty conversation. Systems that record meetings and those that support teleconferences are attracting considerable interest. Conventional systems use a fixed-viewpoint camera and simple camera selection based on participants' utterances. However, conventional systems fail to convey a sufficient amount of nonverbal information about the participants and the flow of conversation. We focus on participants' gaze since it is a good indicator of the participants' intent and emotion, conversational attention etc. We propose a video cut editing rule based on the convergence of participants' gaze direction. We conduct an experiment to evaluate the effectiveness of the proposed method. The results indicate that the proposed method can successfully convey who is taking to whom, which is a key indicator of the flow of conversation.
Yoshinao Takemae, Kazuhiro Otsuka, Naoki Mukawa
ACM Multimedia2
1998 Feature extraction of temporal texture based on spatiotemporal motion trajectory
abstract
A framework and method are proposed to extract local features of a certain kind of naturally occurring, non-rigid motion pattern, referred to as temporal texture. To catch both the spatial and temporal features of this complex pattern, we focus on the surfaces of motion trajectories in spatiotemporal space derived from multiple frames of an image sequence, and represent the surfaces as a set of tangent planes of the surfaces. From the distribution of the tangent planes in local regions in time and space, spatial and temporal texture features are computed The features considered here include spatial arrangement of dominant contours, uniformity of velocity components, and trajectory run length. Experimental results show that the newly defined features have the capability of quantifying the features of complex motion patterns such as weather radar images.
Kazuhiro Otsuka, Tsutomu Horikoshi, Masaharu Fujii
ICPR1
1997 Image velocity estimation from trajectory surface in spatiotemporal space
abstract
A new framework and method, based on image motion trajectories in spatiotemporal space (x-y-t space), are proposed to estimate image velocity from an image sequence. We focus on the surfaces of the trajectories in the x-y-t space formed by the edges and contours of moving objects and obtain image velocity from the orientation of the intersection line formed by tangent planes on the trajectories. The proposed method includes two Hough transforms to detect the most dominant orientation in all possible intersection lines and reliably produces the dominant translational image velocity semi-locally. Also, the confidence measure of estimates is defined to decide the optimal size of patch that suppresses the aperture problem. Experimental results from several synthetic and real image sequences are presented to verify the effectiveness of the method and to confirm its robustness against noise and occlusion.
Kazuhiro Otsuka, Tsutomu Horikoshi
CVPR1