EDBT 2026 Demo / reviewers in the wild / expert
Junji Yamato
dblp:85/5476
· DBLP profile ↗
45ranked-venue papers
4as first author
1since 2021 · last 2022
0000-0003-1634-7984ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 1 first-authorArtificial intelligence and machine learning · 17 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 17 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 since 2021Computer networks · 3 · 2 first-authorSystems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Video understanding and tracking · 37% Face, body and person analysis · 37% Speech recognition and synthesis · 25% | |
| Human-computer interaction and pervasive computing
3 papers |
Immersive interaction · 44% Wearable and physiological sensing · 31% Human-robot interaction · 19% |
Topics — the 11 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Face, body and person analysis
gaze analysis |
0.3 | 1 | 2017 | Collective First-Person Vision for Automatic Gaze Analysis in Multiparty Conversations · IEEE Trans. Multim. 2017 |
Computer vision › Video understanding and tracking
object tracking |
0.2 | 2 | 2010 | Memory-Based Particle Filter for Tracking Objects with Large Variation in Pose and Appearance · ECCV (3) 2010 Memory-based Particle Filter for face pose tracking robust under complex dynamics · CVPR 2009 |
Natural language and speech › Speech recognition and synthesis
speaker diarization |
0.1 | 1 | 2012 | Low-Latency Real-Time Meeting Recognition and Understanding Using Distant Microphones and Omni-Directional Camera · IEEE Trans. Speech Audio Process. 2012 |
Immersive interaction
telepresence |
0.1 | 1 | 2011 | A system for reconstructing multiparty conversation field based on augmented head motion by dynamic projection · ACM Multimedia 2011 |
Computer vision › Video understanding and tracking › object tracking › adaptive tracking
appearance-adaptive tracking |
0.1 | 1 | 2010 | Memory-Based Particle Filter for Tracking Objects with Large Variation in Pose and Appearance · ECCV (3) 2010 |
Computer vision › Video understanding and tracking › object tracking › probabilistic tracking
particle filter tracking |
0.1 | 1 | 2009 | Memory-based Particle Filter for face pose tracking robust under complex dynamics · CVPR 2009 |
Computer vision › Face, body and person analysis
head pose estimation |
0.0 | 1 | 2012 | Low-Latency Real-Time Meeting Recognition and Understanding Using Distant Microphones and Omni-Directional Camera · IEEE Trans. Speech Audio Process. 2012 |
Usability and user experience research
decision-making |
0.0 | 1 | 2005 | Differences in effect of robot and screen agent recommendations on human decision-making · Int. J. Hum. Comput. Stud. 2005 |
Computer vision › Video understanding and tracking
action recognition |
0.0 | 1 | 1992 | Recognizing human action in time-sequential images using hidden Markov model · CVPR 1992 |
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
hidden markov model-based recognition |
0.0 | 1 | 1992 | Recognizing human action in time-sequential images using hidden Markov model · CVPR 1992 |
Internet architecture and protocols
network synchronization |
0.0 | 2 | 1974 | Dynamic Behavior of a Synchronization Control System for an Integrated Telephone Network · IEEE Trans. Commun. 1974 Stability of a Synchronization Control System for an Integrated Telephone Network · IEEE Trans. Commun. 1974 |
Methods — techniques the papers use, named apart from their topics
speaker detection · 0.6self-calibration · 0.6overlapping speech separation · 0.1omnidirectional camera · 0.1microphone array processing · 0.1particle filter · 0.1memory-based appearance model · 0.1random sampling from history · 0.1memory-based particle filter · 0.1feature-based bottom-up approach · 0.0stability analysis · 0.0sampled-data control · 0.0impulse response analysis · 0.0computer simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Preliminary Investigation of Collision Risk Assessment with Vision for Selecting Targets Paid Attention to by Mobile RobotabstractVision plays an important role in motion planning for mobile robots which coexist with humans. Because a method predicting a pedestrian path with a camera has a trade-off relationship between the calculation speed and accuracy, such a path prediction method is not good at instantaneously detecting multiple people at a distance. In this study, we thus present a method with visual recognition and prediction of transition of human action states to assess the risk of collision for selecting the avoidance target. The proposed system calculates the risk assessment score based on recognition of human body direction, human walking patterns with an object, and face orientation as well as prediction of transition of human action states. First, we investigated the validation of each recognition model, and we confirmed that the proposed system can recognize and predict human actions with high accuracy ahead of 3 m. Then, we compared the risk assessment score with video interviews to ask a human whom a mobile robot should pay attention to, and we found that the proposed system could capture the features of human states that people pay attention to when avoiding collision with other people from vision. Masaaki Hayashi, Tamon Miyake, Mitsuhiro Kamezaki, Junji Yamato, Kyosuke Saito, Taro Hamada, Eriko Sakurai, Shigeki Sugano, Jun Ohya |
RO-MAN | 4 |
| 2020 | Adversarial Knowledge Distillation for a Compact GeneratorabstractIn this paper, we propose memory-efficient Generative Adversarial Nets (GANs) in line with knowledge distillation. Most existing GANs have a shortcoming in terms of the number of model parameters and low processing speed. Here, to tackle the problem, we propose Adversarial Knowledge Distillation for Generative models (AKDG) for highly efficient GANs, in terms of unconditional generation. Using AKDG, model size and processing speed are substantively reduced. Through an adversarial training exercise with a distillation discriminator, a student generator successfully mimics a teacher generator in fewer model layers and fewer parameters and at a higher processing speed. Moreover, our AKDG is network architecture-agnostic. A Comparison of AKDG-applied models to vanilla models suggests that it achieves closer scores to a teacher generator and more efficient performance than a baseline method with respect to Inception Score (IS) and Frechet Inception Distance (FID). In CIFAR-10 experiments, improving IS/FID 1.17pt/55.19pt and in LSUN bedroom experiments, improving FID 71.1pt in comparison to the conventional distillation method for GANs. Our project page is https://maguro27.github.io/AKDG/. Hideki Tsunashima, Hirokatsu Kataoka, Junji Yamato, Qiu Chen, Shigeo Morishima |
ICPR | 3 |
| 2018 | Improving Dialogue Continuity using Inter-Robot InteractionabstractResearch on conversational dialogue systems has attracted attention for achieving dialogue systems that can build social relationships with users. Although such systems are required for continuing long conversations with users to build relationships, they sometimes make sentences that are not related to the dialogue context, causing the dialogue to easily break down. The cause of the difficulty is simple. Unlike task-oriented dialogues, unexpected user utterances whose meaning the system cannot capture are frequently spoken in conversational dialogues, since the dialogue domain is much less restricted and the dialogue goal is less obvious. In this paper, we propose a novel strategy for dialogue systems through which two robots coordinate to create long conversations by avoiding dialogue breakdowns. If one of the two robots accepts user utterances with backchannel, the responsibility of responding to the user utterances is resolved and the other robot can change the dialogue topic, which decreases the risk of generating discontinuous utterances. Even if a dialogue nearly breaks down, such multiple robots can present predefined natural interactions among themselves to repair it. Our experiments show that the inter-robot interaction effectively improves the establishment of conversational dialogue that helps users continue the dialogue. Hiroaki Sugiyama, Toyomi Meguro, Yuichiro Yoshikawa, Junji Yamato |
RO-MAN | 4 |
| 2017 | Collective First-Person Vision for Automatic Gaze Analysis in Multiparty ConversationsabstractThis paper targets smallto medium-sized-group face-to-face conversations where each person wears a dual-view camera, consisting of inwardand outward-looking cameras, and presents an almost fully automatic but accurate ofline gaze analysis framework that does not require users to perform any calibration steps. Our collective first-person vision framework, where captured audio-visual signals are gathered and processed in a centralized system, jointly undertakes the fundamental functions required for group gaze analysis, including speaker detection, face tracking, and gaze tracking. Of particular note is our self-calibration of gaze trackers by exploiting a general conversation rule, namely that listeners are likely to look at the speaker. From the rough conversational prior knowledge, our system visualizes fine-grained participants' gaze behavior as a gazee-centered heat map, which quantitatively reveals what parts of the gazee's body the participant looked at and for how long while the gazer was speaking or listening. An experiment using conversations amounting to a total of 140 min, each lasting an average of 8.7 min and engaged in by 37 participants in groups of three to six, achieves a mean absolute error of 2.8° in gaze tracking. A statistical test reveals neither a group size effect nor a conversation type effect. Our method achieves F-scores of over 0.89 and 0.87 in gazee and eye contact recognition, respectively, in comparison with human annotation. Shiro Kumano, Kazuhiro Otsuka, Ryo Ishii, Junji Yamato |
IEEE Trans. Multim. | 4 |
| 2016 | Prediction of Who Will Be the Next Speaker and When Using Gaze Behavior in Multiparty MeetingsabstractIn multiparty meetings, participants need to predict the end of the speaker’s utterance and who will start speaking next, as well as consider a strategy for good timing to speak next. Gaze behavior plays an important role in smooth turn-changing. This article proposes a prediction model that features three processing steps to predict (I) whether turn-changing or turn-keeping will occur, (II) who will be the next speaker in turn-changing, and (III) the timing of the start of the next speaker’s utterance. For the feature values of the model, we focused on gaze transition patterns and the timing structure of eye contact between a speaker and a listener near the end of the speaker’s utterance. Gaze transition patterns provide information about the order in which gaze behavior changes. The timing structure of eye contact is defined as who looks at whom and who looks away first, the speaker or listener, when eye contact between the speaker and a listener occurs. We collected corpus data of multiparty meetings, using the data to demonstrate relationships between gaze transition patterns and timing structure and situations (I), (II), and (III). The results of our analyses indicate that the gaze transition pattern of the speaker and listener and the timing structure of eye contact have a strong association with turn-changing, the next speaker in turn-changing, and the start time of the next utterance. On the basis of the results, we constructed prediction models using the gaze transition patterns and timing structure. The gaze transition patterns were found to be useful in predicting turn-changing, the next speaker in turn-changing, and the start time of the next utterance. Contrary to expectations, we did not find that the timing structure is useful for predicting the next speaker and the start time. This study opens up new possibilities for predicting the next speaker and the timing of the next utterance using gaze transition patterns in multiparty meetings. Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato |
ACM Trans. Interact. Intell. Syst. | 4 |
| 2016 | Using Respiration to Predict Who Will Speak Next and When in Multiparty MeetingsabstractTechniques that use nonverbal behaviors to predict turn-changing situations—such as, in multiparty meetings, who the next speaker will be and when the next utterance will occur—have been receiving a lot of attention in recent research. To build a model for predicting these behaviors we conducted a research study to determine whether respiration could be effectively used as a basis for the prediction. Results of analyses of utterance and respiration data collected from participants in multiparty meetings reveal that the speaker takes a breath more quickly and deeply after the end of an utterance in turn-keeping than in turn-changing. They also indicate that the listener who will be the next speaker takes a bigger breath more quickly and deeply in turn-changing than the other listeners. On the basis of these results, we constructed and evaluated models for predicting the next speaker and the time of the next utterance in multiparty meetings. The results of the evaluation suggest that the characteristics of the speaker's inhalation right after an utterance unit—the points in time at which the inhalation starts and ends after the end of the utterance unit and the amplitude, slope, and duration of the inhalation phase—are effective for predicting the next speaker in multiparty meetings. They further suggest that the characteristics of listeners' inhalation—the points in time at which the inhalation starts and ends after the end of the utterance unit and the minimum and maximum inspiration, amplitude, and slope of the inhalation phase—are effective for predicting the next speaker. The start time and end time of the next speaker's inhalation are also useful for predicting the time of the next utterance in turn-changing. Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato |
ACM Trans. Interact. Intell. Syst. | 4 |
| 2015 | A conversational robot with vocal and bodily fillers for recovering from awkward silence at turn-takingsabstractWhen there is a lull in conversation, many people feel awkward and make sounds, such as “ummm” (a vocal filler), or stroke their chins or other parts of their bodies (a bodily filler). These fillers suggest the intent to recommence and continue the conversation. Thus, the purpose of conversation is to facilitate the sharing of beneficial information as well as a comfortable moment with an interlocutor. In this manner, a robot intended to help foster a comfortable or relaxing conversational atmosphere through interaction with a user (e.g., therapeutic robots) needs to be designed to convey such a cooperative demeanor to its human interlocutor. In this study, we analyze the effects of a robot's vocal and bodily fillers during awkward silences between turns in conversations with humans. The results of our study show that people feel awkward during silences, even when in conversation with a robot, and that the robot's conversational fillers help mitigate awkwardness and express a cooperative attitude in verbal interactions with people. Our experiments also revealed that subjects who were less socially adept reported feeling that their robot interlocutor was more sincere than its human counterpart. Naoki Ohshima, Keita Kimijima, Junji Yamato, Naoki Mukawa |
RO-MAN | 3 |
| 2015 | Analyzing Interpersonal Empathy via Collective ImpressionsabstractThis paper presents a research framework for understanding the empathy that arises between people while they are conversing. By focusing on the process by which empathy is perceived by other people, this paper aims to develop a computational model that automatically infers perceived empathy from participant behavior. To describe such perceived empathy objectively, we introduce the idea of using the collective impressions of external observers. In particular, we focus on the fact that the perception of other's empathy varies from person to person, and take the standpoint that this individual difference itself is an essential attribute of human communication for building, for example, successful human relationships and consensus. This paper describes a probabilistic model of the process that we built based on the Bayesian network, and that relates the empathy perceived by observers to how the gaze and facial expressions of participants co-occur between a pair. In this model, the probability distribution represents the diversity of observers' impression, which reflects the individual differences in the schema when perceiving others' empathy from their behaviors, and the ambiguity of the behaviors. Comprehensive experiments demonstrate that the inferred distributions are similar to those made by observers. Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Masafumi Matsuda, Junji Yamato |
IEEE Trans. Affect. Comput. | 5 |
| 2014 | Analysis and modeling of next speaking start timing based on gaze behavior in multi-party meetingsabstractTo realize a conversational interface where an agent system can smoothly communicate with multiple persons, it is imperative to know how the start timing of speaking is decided. In this research, we demonstrate a relationship between gaze transition patterns and the start timing of next speaking against the end of the last speaking in multi-party meetings. Then, we construct a prediction model for the start timing using gaze transition patterns near the end of an utterance. An analysis of data collected from natural multi-party meetings reveals a strong relationship between gaze transition patterns of the speaker, next speaker, and listener and the start timing of the next speaker. On the basis of the results, we used gaze transition patterns of the speaker, next speaker, and listener and mutual gaze as variables, and devised several prediction models. A model using all features performed the best and was able to predict the start timing well. Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato |
ICASSP | 4 |
| 2014 | Experimental Evaluation of Chromostereopsis with Varying Center Wavelength and FWHM of Spectral Power Distribution
Masaru Tsuchida, Kunio Kashino, Junji Yamato |
ICISP | 3 |
| 2014 | Analysis of Respiration for Prediction of "Who Will Be Next Speaker and When?" in Multi-Party MeetingsabstractTo build a model for predicting the next speaker and the start time of the next utterance in multi-party meetings, we performed a fundamental study of how respiration could be effective for the prediction model. The results of the analysis reveal that a speaker inhales more rapidly and quickly right after the end of a unit of utterance in turn-keeping. The next speaker takes a bigger breath toward speaking in turn-changing than listeners who will not become the next speaker. Based on the results of the analysis, we constructed the prediction models to evaluate how effective the parameters are. The results of the evaluation suggest that the speaker's inhalation right after a unit of utterance, such as the start time from the end of the unit of utterance and the slope and duration of the inhalation phase, is effective for predicting whether turn-keeping or turn-changing happen about 350 ms before the start time of the next utterance on average and that listener's inhalation before the next utterance, such as the maximal inspiration and amplitude of the inhalation phase, is effective for predicting the next speaker in turn-changing about 900 ms before the start time of the next utterance on average. Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Junji Yamato |
ICMI | 4 |
| 2013 | Using a Probabilistic Topic Model to Link Observers' Perception Tendency to PersonalityabstractTargeting multiparty conversations, the present study aims to elucidate how an observer will tend to perceive others' emotional states, develops a computational model that realizes the automatic inferencing of the observer's perception tendency. This paper proposes a probabilistic model that automatically discovers the correlation between perception tendency, gender, and personality traits of a target observer. Perception tendency, a probability distribution, explains how likely the observer is to perceive a certain state/level of a target emotion. Personality traits are measured by a variety of questionnaires. The proposed model links these three factors via a latent variable and explains observer's characteristics as a mixture of prototypical characters. An experiment is conducted with fifty observers. They watch 97 short conversation videos and give their impressions about the empathy between each interacting pair. The results demonstrate that the proposed method can find a reasonable framework that underlies the factors: e.g. 1) people who have high scores in Davis's empathy measures show empathy-biased response tendency, and 2) people who have strong sense of consideration for others tend to show an extreme response tendency, and such people are likely to be females. The proposed method shows promise in estimating an observer's perception tendency from his/her gender and personality traits, even when the target perception tendency is quite different from the average perception tendency among observers. Shiro Kumano, Kazuhiro Otsuka, Masafumi Matsuda, Ryo Ishii, Junji Yamato |
ACII | 5 |
| 2013 | Predicting next speaker and timing from gaze transition patterns in multi-party meetingsabstractIn multi-party meetings, participants need to predict the end of the speaker's utterance and who will start speaking next, and to consider a strategy for good timing to speak next. Gaze behavior plays an important role for smooth turn-taking. This paper proposes a mathematical prediction model that features three processing steps to predict (I) whether turn-taking or turn-keeping will occur, (II) who will be the next speaker in turn-taking, and (III) the timing of the start of the next speaker's utterance. For the feature quantity of the model, we focused on gaze transition patterns near the end of utterance. We collected corpus data of multi party meetings and analyzed how the frequencies of appearance of gaze transition patterns differs depending on situations of (I), (II), and (III). On the basis of the analysis, we construct a probabilistic mathematical model that uses the frequencies of appearance of all participants' gaze transition patterns. The results of an evaluation of the model show the proposed models succeed with high precision compared to ones that do not take gaze transition patterns into account. Ryo Ishii, Kazuhiro Otsuka, Shiro Kumano, Masafumi Matsuda, Junji Yamato |
ICMI | 5 |
| 2013 | MM+Space: n x 4 degree-of-freedom kinetic display for recreating multiparty conversation spacesabstractA novel system, called MM+Space, is presented for recreating multiparty face-to-face conversation scenes in the real world. It aims to display and playback pre-recorded conversations as if the people were talking in front of the viewer(s). This system consists of multiple projectors and transparent screens, which display the life-size faces of people. The key idea is the physical augmentation of human head motions, i.e. the screen pose is dynamically controlled to emulate the head motions, for boosting the viewers' perception of nonverbal behaviors and interactions. In particular, MM+Space newly introduces 2-Degree-of-Freedom (DoF) translations, in forward-backward and right-left directions, in addition to 2-DoF head rotations (nodding and shaking), which were proposed in our former MM-Space system. The full 4-DoF kinetic display is expected to enhance the expressibility of head and body motions, and to create more realistic representation of interacting people. Experiments showed that the proposed system with 4-DoF motions outperformed the rotation-only system in the increased perception of people's presence and in expressing their postures. In addition, it was reported that the proposed system allowed the viewers to experience rich emotional expressibility, immersion in conversations, and potential behavioral/emotional contagion. Kazuhiro Otsuka, Shiro Kumano, Ryo Ishii, Maja Zbogar, Junji Yamato |
ICMI | 5 |
| 2013 | Inferring mood in ubiquitous conversational videoabstractConversational social video is becoming a worldwide trend. Video communication allows a more natural interaction, when aiming to share personal news, ideas, and opinions, by transmitting both verbal content and nonverbal behavior. However, the automatic analysis of natural mood is challenging, since it is displayed in parallel via voice, face, and body. This paper presents an automatic approach to infer 11 natural mood categories in conversational social video using single and multimodal nonverbal cues extracted from video blogs (vlogs) from YouTube. The mood labels used in our work were collected via crowdsourcing. Our approach is promising for several of the studied mood categories. Our study demonstrates that although multimodal features perform better than single channel features, not always all the available channels are needed to accurately discriminate mood in videos. Dairazalia Sanchez-Cortes, Joan-Isaac Biel, Shiro Kumano, Junji Yamato, Kazuhiro Otsuka, Daniel Gatica-Perez |
MUM | 4 |
| 2012 | Linking speaking and looking behavior patterns with group composition, perception, and performanceabstractThis paper addresses the task of mining typical behavioral patterns from small group face-to-face interactions and linking them to social-psychological group variables. Towards this goal, we define group speaking and looking cues by aggregating automatically extracted cues at the individual and dyadic levels. Then, we define a bag of nonverbal patterns (Bag-of-NVPs) to discretize the group cues. The topics learnt using the Latent Dirichlet Allocation (LDA) topic model are then interpreted by studying the correlations with group variables such as group composition, group interpersonal perception, and group performance. Our results show that both group behavior cues and topics have significant correlations with (and predictive information for) all the above variables. For our study, we use interactions with unacquainted members i.e. newly formed groups. Dinesh Babu Jayagopi, Dairazalia Sanchez-Cortes, Kazuhiro Otsuka, Junji Yamato, Daniel Gatica-Perez |
ICMI | 4 |
| 2012 | Low-Latency Real-Time Meeting Recognition and Understanding Using Distant Microphones and Omni-Directional CameraabstractThis paper presents our real-time meeting analyzer for monitoring conversations in an ongoing group meeting. The goal of the system is to recognize automatically “who is speaking what” in an online manner for meeting assistance. Our system continuously captures the utterances and face poses of each speaker using a microphone array and an omni-directional camera positioned at the center of the meeting table. Through a series of advanced audio processing operations, an overlapping speech signal is enhanced and the components are separated into individual speaker's channels. Then the utterances are sequentially transcribed by our speech recognizer with low latency. In parallel with speech recognition, the activity of each participant (e.g., speaking, laughing, watching someone) and the circumstances of the meeting (e.g., topic, activeness, casualness) are detected and displayed on a browser together with the transcripts. In this paper, we describe our techniques and our attempt to achieve the low-latency monitoring of meetings, and we show our experimental results for real-time meeting transcription. Takaaki Hori, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto, Shinji Watanabe 0001, Takanobu Oba, Atsunori Ogawa, Kazuhiro Otsuka, Dan Mikami, Keisuke Kinoshita, Tomohiro Nakatani, Atsushi Nakamura, Junji Yamato |
IEEE Trans. Speech Audio Process. | 13 |
| 2011 | Analyzing empathetic interactions based on the probabilistic modeling of the co-occurrence patterns of facial expressions in group meetingsabstractThis paper presents a novel research framework for the estimation of emotional interactions produced between meeting participants. The types of emotional interaction targeted in this paper are empathy, antipathy, and unconcern. We define here emotional interaction as a brief contiguous event wherein a pair exchange emotional messages via verbal and non-verbal behaviors. As the key behaviors, we focus on facial expression and gaze, because their combination realizes the rapid and directed transmission of a large number of emotional messages. We assume that there is a strong link between the emotional interaction and the participants' facial expressions that occur simultaneously with the type of the emotional interactions. Based on this assumption, we build a probabilistic model that represents a hierarchical structure involving the emotional interactions, facial expressions and other behaviors including utterance and gaze direction. Using this model, the type of emotional interaction is estimated from interpersonal gaze directions, facial expressions, and utterances. Our estimation is based on the Bayesian approach, and uses the Markov chain Monte Carlo method to approximate joint posterior probability distributions of the emotional interaction and model parameters present within the observed data. An experiment on four-party conversations demonstrates the promising effectiveness of the proposed method. Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Junji Yamato |
FG | 4 |
| 2011 | A system for reconstructing multiparty conversation field based on augmented head motion by dynamic projectionabstractA novel system is presented for reconstructing, in the real world, multiparty face-to-face conversation scenes; it uses dynamics projection to augment human head motion. This system aims to display and playback pre-recorded conversations to the viewers as if the remote people were taking in front of them. This system consists of multiple projectors and transparent screens. Each screen separately displays the life-size face of one meeting participant, and are spatially arranged to recreate the actual scene. The main feature of this system is dynamics projection, screen pose is dynamically controlled to emulate the head motions of the participants, especially rotation around the vertical axis, that are typical of shifts in visual attention, i.e. turning gaze from one to another. This recreation of head motion by physical screen motion, in addition to image motion, aims to more clearly express the interactions involving visual attention among the participants. The minimal design, frameless-projector-screen, with augmented head motion is expected to create a feeling that the remote participants are actually present in the same room. This demo presents our initial system and discusses its potential impact on future visual communications. Kazuhiro Otsuka, Kamil Sebastian Mucha, Shiro Kumano, Dan Mikami, Masafumi Matsuda, Junji Yamato |
ACM Multimedia | 6 |
| 2011 | Early facial expression recognition with high-frame rate 3D sensingabstractThis work investigates a new challenging problem: how to exactly recognize facial expression as early as possible, while most works generally focus on improving the recognition rate of facial expression recognition. The features of facial expressions in their early stage are unfortunately very sensitive to noise due to their low intensity. So, we propose a novel wavelet spectral subtraction method to spatio-temporally refine the subtle facial expression features. Moreover, in order to achieve early facial expression recognition, we newly introduce an early AdaBoost algorithm for facial expression recognition problem. Experiments using our database established by using a high-frame rate 3D sensing showed that the proposed method has a promising performance on early facial expression recognition. Lumei Su, Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Junji Yamato, Yoichi Sato 0001 |
SMC | 5 |
| 2010 | Memory-Based Particle Filter for Tracking Objects with Large Variation in Pose and Appearance
Dan Mikami, Kazuhiro Otsuka, Junji Yamato |
ECCV (3) | 3 |
| 2010 | Real-time meeting recognition and understanding using distant microphones and omni-directional cameraabstractThis paper presents our newly developed real-time meeting analyzer for monitoring conversations in an ongoing group meeting. The goal of the system is to automatically recognize “who is speaking what” in an online manner for meeting assistance. Our system continuously captures the utterances and the face pose of each speaker using a distant microphone array and an omni-directional camera at the center of the meeting table. Through a series of advanced audio processing operations, an overlapping speech signal is enhanced and the components are separated into individual speaker's channels. Then the utterances are sequentially transcribed by our speech recognizer with low latency. In parallel with speech recognition, the activity of each participant (e.g. speaking, laughing, watching someone) and the situation of the meeting (e.g. topic, activeness, casualness) are detected and displayed on a browser together with the transcripts. In this paper, we describe our techniques and our attempt to achieve the low-latency monitoring of meetings, and we show our experimental results for real-time meeting transcription. Takaaki Hori, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto, Shinji Watanabe 0001, Takanobu Oba, Atsunori Ogawa, Kazuhiro Otsuka, Dan Mikami, Keisuke Kinoshita, Tomohiro Nakatani, Atsushi Nakamura, Junji Yamato |
SLT | 13 |
| 2009 | Memory-based Particle Filter for face pose tracking robust under complex dynamicsabstractA novel particle filter, the memory-based particle filter (M-PF), is proposed that can visually track moving objects that have complex dynamics. We aim to realize robustness against abrupt object movements and quick recovery from tracking failure caused by factors such as occlusions. To that end, we eliminate the Markov assumption from the previous particle filtering framework and predict the prior distribution of the target state from the long-term dynamics. More concretely, M-PF stores the past history of the estimated target states, and employs a random sampling from the history to generate prior distribution; it represents a novel PF formulation.Our method can handle nonlinear, time-variant, and non-Markov dynamics, which is not possible within existing PF frameworks. Accurate prior prediction based on proper dynamics model is especially effective for recovering lost tracks, because it can provide possible target states, which can drastically change since the track was lost. We target the face pose of seated humans in this paper. Quantitative evaluations with magnetic sensors confirm improved accuracy in face pose estimation and successful recovery from tracking loss. The proposed M-PF suggests a new paradigm for modeling systems with complex dynamics and so offers a various visual tracking applications. Dan Mikami, Kazuhiro Otsuka, Junji Yamato |
CVPR | 3 |
| 2009 | Saliency-based video segmentation with graph cuts and sequentially updated priorsabstractThis paper proposes a new method for achieving precise video segmentation without any supervision or interaction. The main contributions of this report include 1) the introduction of fully automatic segmentation based on the maximum a posteriori (MAP) estimation of the Markov random field (MRF) with graph cuts and saliency-driven priors and 2) the updating of priors and feature likelihoods by integrating the previous segmentation results and the currently estimated saliency-based visual attention. Test results indicate that our new method precisely extracts probable regions from videos without any supervised interactions. Ken Fukuchi, Kouji Miyazato, Akisato Kimura, Shigeru Takagi, Junji Yamato |
ICME | 5 |
| 2009 | Real-time estimation of human visual attention with dynamic Bayesian network and MCMC-based particle filterabstractRecent studies in signal detection theory suggest that the human responses to the stimuli on a visual display are nondeterministic. People may attend to different locations on the same visual input at the same time. Constructing a stochastic model of human visual attention would be promising to tackle the above problem. This paper proposes a new method to achieve a quick and precise estimation of human visual attention based on our previous stochastic model with a dynamic Bayesian network. A particle filter with Markov chain Monte-Carlo (MCMC) sampling make it possible to achieve a quick and precise estimation through stream processing. Experimental results indicate that the proposed method can estimate human visual attention in real time and more precisely than previous methods. Kouji Miyazato, Akisato Kimura, Shigeru Takagi, Junji Yamato |
ICME | 4 |
| 2009 | Recognizing communicative facial expressions for discovering interpersonal emotions in group meetingsabstractThis paper proposes a novel facial expression recognizer and describes its application to group meeting analysis. Our goal is to automatically discover the interpersonal emotions that evolve over time in meetings, e.g. how each person feels about the others, or who affectively influences the others the most. As the emotion cue, we focus on facial expression, more specifically smile, and aim to recognize ``who is smiling at whom, when, and how often'', since frequently smiling carries affective messages that are strongly directed to the person being looked at; this point of view is our novelty. To detect such communicative smiles, we propose a new algorithm that jointly estimates facial pose and expression in the framework of the particle filter. The main feature is its automatic selection of interest points that can robustly capture small changes in expression even in the presence of large head rotations. Based on the recognized facial expressions and their directions to others, which are indicated by the estimated head poses, we visualize interpersonal smile events as a graph structure, we call it the interpersonal emotional network; it is intended to indicate the emotional relationships among meeting participants. A four-person meeting captured by an omnidirectional video system is used to confirm the effectiveness of the proposed method and the potential of our approach for deep understanding of human relationships developed through communications. Shiro Kumano, Kazuhiro Otsuka, Dan Mikami, Junji Yamato |
ICMI | 4 |
| 2009 | Realtime meeting analysis and 3D meeting viewer based on omnidirectional multimodal sensorsabstractThis demo presents a realtime system for analyzing group meetings. Targeting round-table meetings, this system employs an omnidirectional camera-microphone system. The goal of this system is to automatically discover "who is talking to whom and when". To that purpose, the face pose/position of meeting participants are tracked on panorama images acquired from fisheye-based omnidirectional cameras. From audio signals obtained with microphone array, speaker diarization, i.e. the estimation of "who is speaking and when", is carried out. The visual focus of attention, i.e. "who is looking at whom", is esimated from the result of face tracking. The results are displayed based on a 3D visualization scheme. The advantage of our system is its realtimeness. We will demonstrate the portable version of the system consisting of two laptop PCs. In addition, we will showcase our meeting playback viewer with man-machine interfaces that allow users to freely control space and time of meeting scenes. With this viewer, users can also experince 3D positional sound effect linked with 3D viewpoint, using enhanced audio tracks for each participant. Kazuhiro Otsuka, Shoko Araki, Dan Mikami, Kentaro Ishizuka, Masakiyo Fujimoto, Junji Yamato |
ICMI | 6 |
| 2009 | Pose-Invariant Facial Expression Recognition Using Variable-Intensity Templates
Shiro Kumano, Kazuhiro Otsuka, Junji Yamato, Eisaku Maeda, Yoichi Sato 0001 |
Int. J. Comput. Vis. | 3 |
| 2008 | Combining Stochastic and Deterministic Search for Pose-Invariant Facial Expression RecognitionabstractWe propose a novel method for pose-invariant facial expression recognition from monocular video sequences that combines stochastic and determinis-tic search processes. We use the simple face model called variable-intensity template, which can be prepared with very little time and effort. We tackle the two issues found in previous work on the variable-intensity template: low accuracy in head pose estimation, and assumption violations due to external intensity changes such as illumination change. We mitigate these issues by introducing the deterministic approach into the stochastic approach imple-mented as a particle filter. Our experiment demonstrates significant improve-ments in recognition performance for horizontal and vertical head orienta-tions in the range of ±40 degrees and ±20 degrees, respectively, from the frontal view. 1 Shiro Kumano, Kazuhiro Otsuka, Junji Yamato, Eisaku Maeda, Yoichi Sato 0001 |
BMVC | 3 |
| 2008 | t-Room: Next Generation Video Communication SystemabstractIn this paper, we present t-Room, the next generation video communication system we are developing. Our approach is to build rooms with identical layouts, including walls of display panels on which users and physical or virtual objects are all shown at life-size. In this way, the user space enclosed by t-Room's surrounding displays can be shared as a common space at any other site. In other words, the enclosed spaces overlap each other. This configuration effectively provides symmetric reproduction of the audio-visual information surrounding local and remote users and objects. The feeling provided by t-Room is different from that by conventional videoconferencing systems, since there is no spatial barrier separating users such as the video screen of a conventional videoconferencing system. Furthermore, t-Room benefits in every way from Next Generation Network (NGN) technology: QoS, service productivity, and security. We view t-Room as a future form of telephone service. Keiji Hirata 0001, Yasunori Harada, Toshihiro Takada, Shigemi Aoyagi, Yoshinari Shirai, Naomi Yamashita, Katsuhiko Kaji, Junji Yamato, Kenji Nakazawa |
GLOBECOM | 8 |
| 2008 | A stochastic model of selective visual attention with a dynamic Bayesian networkabstractRecent studies in signal detection theory suggest that the human responses to the stimuli on a visual display are nondeterministic. People may attend to different locations on the same visual input at the same time. To predict the likelihood of where humans typically focus on a video scene, we propose a new stochastic model of visual attention by introducing a dynamic Bayesian network. Our model simulates and combines the visual saliency response and the cognitive state of a person to estimate the most probable attended regions. Experimental results have demonstrated that our model performs significantly better in predicting human visual attention compared to the previous deterministic model. Derek Pang, Akisato Kimura, Tatsuto Takeuchi, Junji Yamato, Kunio Kashino |
ICME | 4 |
| 2008 | A realtime multimodal system for analyzing group meetings by combining face pose tracking and speaker diarizationabstractThis paper presents a realtime system for analyzing group meetings that uses a novel omnidirectional camera-microphone system. The goal is to automatically discover the visual focus of attention (VFOA), i.e. "who is looking at whom", in addition to speaker diarization, i.e. "who is speaking and when". First, a novel tabletop sensing device for round-table meetings is presented; it consists of two cameras with two fisheye lenses and a triangular microphone array. Second, from high-resolution omnidirectional images captured with the cameras, the position and pose of people's faces are estimated by STCTracker (Sparse Template Condensation Tracker); it realizes realtime robust tracking of multiple faces by utilizing GPUs (Graphics Processing Units). The face position/pose data output by the face tracker is used to estimate the focus of attention in the group. Using the microphone array, robust speaker diarization is carried out by a VAD (Voice Activity Detection) and a DOA (Direction of Arrival) estimation followed by sound source clustering. This paper also presents new 3-D visualization schemes for meeting scenes and the results of an analysis. Using two PCs, one for vision and one for audio processing, the system runs at about 20 frames per second for 5-person meetings. Kazuhiro Otsuka, Shoko Araki, Kentaro Ishizuka, Masakiyo Fujimoto, Martin Heinrich, Junji Yamato |
ICMI | 6 |
| 2008 | Dynamic Markov random fields for stochastic modeling of visual attentionabstractThis report proposes a new stochastic model of visual attention to predict the likelihood of where humans typically focus on a video scene. The proposed model is composed of a dynamic Bayesian network that simulates and combines a person’s visual saliency response and eye movement patterns to estimate the most probable regions of attention. Dynamic Markov random field (MRF) models are newly introduced to include spatiotemporal relationships of visual saliency responses. Experimental results have revealed that the propose model outperforms the previous deterministic model and the stochastic model without dynamic MRF in predicting human visual attention. Akisato Kimura, Derek Pang, Tatsuto Takeuchi, Junji Yamato, Kunio Kashino |
ICPR | 4 |
| 2007 | Pose-Invariant Facial Expression Recognition Using Variable-Intensity Templates
Shiro Kumano, Kazuhiro Otsuka, Junji Yamato, Eisaku Maeda, Yoichi Sato 0001 |
ACCV (1) | 3 |
| 2007 | Automatic inference of cross-modal nonverbal interactions in multiparty conversations: "who responds to whom, when, and how?" from gaze, head gestures, and utterancesabstractA novel probabilistic framework is proposed for analyzing cross-modal nonverbal interactions in multiparty face-to-face conversations. The goal is to determine "who responds to whom, when, and how" from multimodal cues including gaze, head gestures, and utterances. We formulate this problem as the probabilistic inference of the causal relationship among participants' behaviors involving head gestures and utterances. To solve this problem, this paper proposes a hierarchical probabilistic model; the structures of interactions are probabilistically determined from high-level conversation regimes (such as monologue or dialogue) and gaze directions. Based on the model, the interaction structures, gaze, and conversation regimes, are simultaneously inferred from observed head motion and utterances, using a Markov chain Monte Carlo method. The head gestures, including nodding, shaking and tilt, are recognized with a novel Wavelet-based technique from magnetic sensor signals. The utterances are detected using data captured by lapel microphones. Experiments on four-person conversations confirm the effectiveness of the framework in discovering interactions such as question-and-answer and addressing behavior followed by back-channel responses. Kazuhiro Otsuka, Hiroshi Sawada, Junji Yamato |
ICMI | 3 |
| 2006 | Poster Image Matching by Color Scheme and Layout InformationabstractIn this paper, we demonstrate a novel poster image matching system for wireless multimedia applications. We propose a method that incorporates both color and layout information of the poster image to achieve a robust performance in poster image matching. We apply both color compensation and background separation to extract a poster from an image effectively. Based on our experiment, we show that even under the effects of lighting, image rotation, scaling, and occlusion, our system can still maintain high recall and precision. We also show that our system can recognize the correct image from a database which contains several poster images with similar features. Finally, the promising performance of the poster image matching encourages us to further enrich the information retrieval for wireless environment Cheng-Yao Chen, Takayuki Kurozumi, Junji Yamato |
ICME | 3 |
| 2006 | Conversation Scene Analysis with Dynamic Bayesian Network Basedon Visual Head TrackingabstractA novel method based on a probabilistic model for conversation scene analysis is proposed that can infer conversation structure from video sequences of face-to-face communication. Conversation structure represents the type of conversation such as monologue or dialogue, and can indicate who is talking/listening to whom. This study assumes that the gaze directions of participants provide cues for discerning the conversation structure, and can be identified from head directions. For measuring head directions, the proposed method newly employs a visual head tracker based on sparse-template condensation. The conversation model is built on a dynamic Bayesian network and is used to estimate the conversation structure and gaze directions from observed head directions and utterances. Visual tracking is conventionally thought to be less reliable than contact sensors, but experiments confirm that the proposed method achieves almost comparable performance in estimating gaze directions and conversation structure to a conventional sensor-based method Kazuhiro Otsuka, Junji Yamato, Yoshinao Takemae, Hiroshi Murase |
ICME | 2 |
| 2005 | Effects of Automatic Video Editing System Using Stereo-Based Head Tracking for Archiving MeetingsabstractThis paper presents an automatic video editing system based on head tracking for archiving meetings. Systems that archive meetings are attracting considerable interest. Conventional systems use a fixed-viewpoint camera and simple camera selection based on participants' utterances. However, conventional systems fail to adequately convey who is talking to whom and nonverbal information about participants etc. We focus on the participants' head orientation since this information is useful in detecting the speaker and who the speaker is talking to. In order to automatically estimate each participant's head orientation, our system combines several modules to realize stereo-based head tracking. The system selects the shot of the participant that most participants are looking at, based on majority decision. Experiments on presenting videos to viewers confirm the effectiveness of our system in several 3-participant conversations Yoshinao Takemae, Kazuhiro Otsuka, Junji Yamato |
ICME | 3 |
| 2005 | A probabilistic inference of multiparty-conversation structure based on Markov-switching models of gaze patterns, head directions, and utterancesabstractA novel probabilistic framework is proposed for inferring the structure of conversation in face-to-face multiparty communication, based on gaze patterns, head directions and the presence/absence of utterances. As the structure of conversation, this study focuses on the combination of participants and their participation roles. First, we assess the gaze patterns that frequently appear in conversations, and define typical types of conversation structure, called conversational regime, and hypothesize that the regime represents the high-level process that governs how people interact during conversations. Next, assuming that the regime changes over time exhibit Markov properties, we propose a probabilistic conversation model based on Markov-switching; the regime controls the dynamics of utterances and gaze patterns, which stochastically yield measurable head-direction changes. Furthermore, a Gibbs sampler is used to realize the Bayesian estimation of regime, gaze pattern, and model parameters from observed head directions and utterances. Experiments on four-person conversations confirm the effectiveness of the framework in identifying conversation structures. Kazuhiro Otsuka, Yoshinao Takemae, Junji Yamato |
ICMI | 3 |
| 2005 | Differences in effect of robot and screen agent recommendations on human decision-making
Kazuhiko Shinozawa, Futoshi Naya, Junji Yamato, Kiyoshi Kogure |
Int. J. Hum. Comput. Stud. | 3 |
| 2004 | Effect of robot's tracking users on human decision makingabstractThis study clarified that a robot tracking a user can influence user decision-making. Two types of robot were used and the robots tracked or did not track the user. A laboratory experiment (n = 118) tested differences in selection ratios corresponding to a color name recommended by the robots. Subjects selected the recommended color name significantly more often when the robot tracked subjects' faces than in the no-recommendation situation, however, when the robot did not track subjects' faces with recommendation, there was no significant difference. The physiological data also showed that the tracking condition generated higher levels of arousal than the non-tracking condition. This suggests that tracking the user's face influences user decision-making. Kazuhiko Shinozawa, Futoshi Naya, Kiyoshi Kogure, Junji Yamato |
IROS | 4 |
| 2001 | Evaluation of Communication with Robot and Agent: Are Robots Better Social Actors than Agents?
Junji Yamato, Kazuhiko Shinozawa, Futoshi Naya, Kiyoshi Kogure |
INTERACT | 1 |
| 1992 | Recognizing human action in time-sequential images using hidden Markov modelabstractA human action recognition method based on a hidden Markov model (HMM) is proposed. It is a feature-based bottom-up approach that is characterized by its learning capability and time-scale invariability. To apply HMMs, one set of time-sequential images is transformed into an image feature vector sequence, and the sequence is converted into a symbol sequence by vector quantization. In learning human action categories, the parameters of the HMMs, one per category, are optimized so as to best describe the training sequences from the category. To recognize an observed sequence, the HMM which best matches the sequence is chosen. Experimental results for real time-sequential images of sports scenes show recognition rates higher than 90%. The recognition rate is improved by increasing the number of people used to generate the training data, indicating the possibility of establishing a person-independent action recognizer.> Junji Yamato, Jun Ohya, Kenichiro Ishii |
CVPR | 1 |
| 1974 | Stability of a Synchronization Control System for an Integrated Telephone NetworkabstractThis paper is a sequel to previous papers which considered the characteristics of a proposed synchronization control system for an integrated telephone network. In those papers, delay of the paths of each junction, one for transmission and the other for control signal, were assumed to have the same value. Though the analysis produced exact expressions for the static and dynamic characteristics of the control system, this assumption is not likely to be found realistic, in many cases. The present work extends that condition into a case where the above-mentioned delays of the paths are not equal. The conclusion is reached that the stability limit of the network is chiefly defined by the path that has the shorter delay. The stability of a synchronization control system in sampled-data control mode is also discussed. It was clarified that the reduction of the sampling rate results in higher stability of the network. Junji Yamato |
IEEE Trans. Commun. | 1 |
| 1974 | Dynamic Behavior of a Synchronization Control System for an Integrated Telephone NetworkabstractAn integrated telephone network requires synchronization of digital signal streams arriving at each switching system with its switching points. A new synchronization control system for this purpose was already proposed by one of the authors and the final clock rate and stability criterion were discussed in the preceding material. Continuing with this discussion, impulse responses of the proposed control system applied to mesh networks are mathematically analyzed under the condition that all of the clock pulse generators have respective clock rate derivations from their nominal value. The conclusion was arrived at that the more stations included in a network, the shorter the time elapsed before the frame pulse phase disturbances are settled down to a steady state, although this tendency is rather gradual. The dynamic response of the synchronization control system applied to mesh networks, a ring network, a chain network, and an assumed worldwide network are computer simulated in linear and three-value clock rate control modes and the characteristic of the control system is clarified. Junji Yamato, Seiichi Nakajima, Kyuta Saito |
IEEE Trans. Commun. | 1 |