Dinesh Babu Jayagopi

dblp:39/1552 · also Dinesh Babu J. · DBLP profile ↗
← Back
38ranked-venue papers
11as first author
11since 2021 · last 2025
0000-0003-0080-452XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 4 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 15 · 6 first-author · 2 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Efficient Label Refinement for Face Parsing Under Extreme Poses Using 3D Gaussian Splatting
abstract
Accurate face parsing under extreme viewing angles remains a significant challenge due to limited labeled data in such poses. Manual annotation is costly and often impractical at scale. We propose a novel label refinement pipeline that leverages 3D Gaussian Splatting (3DGS) to generate accurate segmentation masks from noisy multiview predictions. By jointly fitting two 3DGS models, one to RGB images and one to their initial segmentation maps, our method enforces multiview consistency through shared geometry, enabling the synthesis of pose-diverse training data with only minimal post-processing. Fine-tuning a face parsing model on this refined dataset significantly improves accuracy on challenging head poses, while maintaining strong performance on standard views. Extensive experiments, including human evaluations, demonstrate that our approach achieves superior results compared to state-of-the-art methods, despite requiring no ground-truth 3D annotations and using only a small set of initial images. Our method offers a scalable and effective solution for improving face parsing robustness in real-world settings.
Ankit Gahlawat, Dinesh Babu Jayagopi
VCIP3
2024 Content-Based Objective Evaluation of Artificially Generated Sign Language Videos
abstract
Sign language is vital for communication within the deaf and hard-of-hearing community. Avatar-based methods and deep learning techniques like Generative Adversarial Networks have shown promise in generating sign language video content. One of the challenges in sign language generation is the evaluation of the generated video content. One possible solution is to subjectively evaluate using human raters. This is time-consuming and costly. The other possible solution is objective evaluation. In the literature, video quality metrics such as PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index) and skeleton-based measures such as MSE have been proposed. A limitation of these approaches is that they do not provide information about the generated video content. In this paper, we propose a novel phonology-based approach that evaluates the generated video along different channels, namely, hand movement and handshape, which convey the linguistic information in sign language. More precisely, in this approach an objective score is obtained by extracting sequences of hand movement sub-units and handshape sub-units class conditional probabilities (posterior features) from the source and generated videos and comparing them using dynamic time warping. Our experimental studies demonstrate that the proposed objective scoring method yields a better correlation to subjective human ratings than PSNR, SSIM, and MSE-based metrics.
Neha Tarigopula, Preyas Garg, Skanda Muralidhar, Sandrine Tornay, Dinesh Babu Jayagopi, Mathew Magimai-Doss
ICASSP5
2024 Modeling essay grading with pre-trained BERT features
Annapurna Sharma, Dinesh Babu Jayagopi
Appl. Intell.2
2024 On the potential of supporting autonomy in online video interview training platforms
abstract
Rising unemployment has led to many discouraged job seekers. While the impact of job seekers’ motivation on interview performance is acknowledged in previous research, little attention has been given to understanding the effect of training on interview motivation and performance. We present InterviewApp, an online interview training tool aiming to support interview motivation through autonomy, relatedness and competence needs derived from Self-Determination Theory and, in turn, performance. Through a four-month study (N=135), we assess its effectiveness in supporting job seekers’ interview motivation and performance. Our results demonstrate the role of autonomy in mediating the effect of training on performance. We found that the intervention significantly affected the job seekers’ perceived autonomy. Furthermore, engagement with the recording and feedback features of the tool positively impacted performance. Overall, job seekers found InterviewApp helpful for online interview training and valued the provided expert feedback. These findings have implications for the design of online interview training tools and for behaviour change interventions to support employment.
Pooja S. B. Rao, Laetitia A. Renier, Marc-Olivier Boldi, Marianne Schmid Mast, Dinesh Babu Jayagopi, Mauro Cherubini
Int. J. Hum. Comput. Stud.5
2024 Bag of states: a non-sequential approach to video-based engagement measurement
Ali Abedi 0003, Chinchu Thomas, Dinesh Babu Jayagopi, Shehroz S. Khan
Multim. Syst.3
2023 Full-page handwriting recognition and automated essay scoring for in-the-wild essays
Annapurna Sharma, Rohit Katlaa, Gurleen Kaur, Dinesh Babu Jayagopi
Multim. Tools Appl.4
2022 Detecting A Child's Stimming Behaviours for Autism Spectrum Disorder Diagnosis using Rgbpose-Slowfast Network
abstract
Autism Spectrum Disoder (ASD) is a neurodevelopmental disorder characterized by (a) persistent deficits in social communication and interaction, and (b) presence of restrictive, repetitive patterns of behaviours, interests or activities. The stereotyped repetitive behaviours are also referred to as stimming behaviours. We propose a deep learning based approach to automatically predict a child’s stimming behaviours from videos recorded in unconstrained conditions. The child’s region in the video is tracked and its skeletal joints are derived using the pose estimator. The heatmap representation of skeletal joints and the raw video signals are used as inputs to the two pathways of the RGBPose-SlowFast deep network to model stimming behaviours. The proposed model is evaluated using the publicly available Self-Stimulatory Behaviour Dataset (SSBD) of stimming behaviours. The generalization ability of the model is validated using the Autism dataset containing child’s motor actions. Our experiments demonstrate state-of-the-art results on both datasets.
Jeba Berlin, Deepak Pandian, Shyam Sundar Rajagopalan, Dinesh Babu Jayagopi
ICIP4
2022 Review of realistic behavior and appearance generation in embodied conversational agents: A comparison between traditional and modern approaches
abstract
With recent technological advancements, many firms that formerly relied on traditional face-to-face conversations have switched to an online conversation mode. This has not only helped businesses to increase their revenues, but it has also enabled users and customers to access world-class services. Due to scalability concerns with these platforms, many of the online conversations have lost the personal touch associated with a face-to-face based communication. An embodied conversational agent (ECA) can address this void by creating suitable behavior as well as the required voice output via a realistic-looking avatar. However, in order to scale up to a large customer base, the behavior and appearance of agents must be adequately modeled. Traditionally, rule-based methods were used to generate the animations associated with the avatars, but because of their limitations modern approaches use deep learning models to create end-to-end systems. We present various conventional and current methodologies for behavior and appearance modeling in our work. We will discuss similarities between these systems as well as their limitations. We believe that our work will be useful in developing a hybrid system that uses both traditional and modern approach to handle these challenges in creating modern embodied conversational agents.
Kumar Shubham, Dinesh Babu Jayagopi
ICMI3
2021 Learning a Deep Reinforcement Learning Policy Over the Latent Space of a Pre-trained GAN for Semantic Age Manipulation
abstract
Learning a disentangled representation of the latent space has become one of the most fundamental problems studied in computer vision. Recently, many Generative Adversarial Networks (GANs) have shown promising results in generating high fidelity images. However, studies to understand the semantic layout of the latent space of pre-trained models are still limited. Several works train conditional GANs to generate faces with required semantic attributes. Unfortunately, in these attempts, the generated output is often not as photo-realistic as the unconditional state-of-the-art models. Besides, they also require large computational resources and specific datasets to generate high fidelity images. In our work, we have formulated a Markov Decision Process (MDP) over the latent space of a pre-trained GAN model to learn a conditional policy for semantic manipulation along specific attributes under defined identity bounds. Further, we have defined a semantic age manipulation scheme using a locally linear approximation over the latent space. Results show that our learned policy samples high fidelity images with required age alterations, while preserving the identity of the person.
Kumar Shubham, Gopalakrishnan Venkatesh, Reijul Sachdev, Akshi, Dinesh Babu Jayagopi, G. Srinivasaraghavan 0001
IJCNN5
2021 Towards efficient unconstrained handwriting recognition using Dilated Temporal Convolution Network
Annapurna Sharma, Dinesh Babu Jayagopi
Expert Syst. Appl.2
2021 Approaches for Multilingual Phone Recognition in Code-switched and Non-code-switched Scenarios Using Indian Languages
abstract
In this study, we evaluate and compare two different approaches for multilingual phone recognition in code-switched and non-code-switched scenarios. First approach is a front-end Language Identification (LID)-switched to a monolingual phone recognizer (LID-Mono), trained individually on each of the languages present in multilingual dataset. In the second approach, a common multilingual phone-set derived from the International Phonetic Alphabet (IPA) transcription of the multilingual dataset is used to develop a Multilingual Phone Recognition System (Multi-PRS). The bilingual code-switching experiments are conducted using Kannada and Urdu languages. In the first approach, LID is performed using the state-of-the-art i-vectors. Both monolingual and multilingual phone recognition systems are trained using Deep Neural Networks. The performance of LID-Mono and Multi-PRS approaches are compared and analysed in detail. It is found that the performance of Multi-PRS approach is superior compared to more conventional LID-Mono approach in both code-switched and non-code-switched scenarios. For code-switched speech, the effect of length of segments (that are used to perform LID) on the performance of LID-Mono system is studied by varying the window size from 500 ms to 5.0 s, and full utterance. The LID-Mono approach heavily depends on the accuracy of the LID system and the LID errors cannot be recovered. But, the Multi-PRS system by virtue of not having to do a front-end LID switching and designed based on the common multilingual phone-set derived from several languages, is not constrained by the accuracy of the LID system, and hence performs effectively on code-switched and non-code-switched speech, offering low Phone Error Rates than the LID-Mono system.
K. Manjunath, Srinivasa Raghavan K. M., K. Sreenivasa Rao, Dinesh Babu Jayagopi, V. Ramasubramanian 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2020 Using Deep 3D Features and an LSTM Based Sequence Model for Automatic Pain Detection in the Wild
abstract
Automatic pain detection is an important problem in diagnostic and therapeutic applications. In this paper, we aim to develop a computational framework to automatically detect pain in videos in the wild. The videos in the wild vary with respect to gender, age, ethnicity and even other qualitative attributes like upbringing. Previous systems focused on methodologies confined to one particular dataset that is hard to generalize for the population in the wild, or based on invasive methods that collect data using many physiological sensors and induced stressors. We propose a method to automatically detect pain in videos using state-of-the-art expression recognition system along with deep learning. We curated a dataset of 194 videos in the wild with pain and non-pain. We used a sliding window strategy to obtain a fixed-length input sample for the LSTM (Long Short Term Memory) network. We then carefully concatenate the network output of every segment to generate a video-level output. The proposed end-to-end framework can predict binary classification label (pain/non-pain) at video level. Our method achieves promising results on the dataset we collected.
Sowmya Rasipuram, Bukka Nikhil Sai, Dinesh Babu Jayagopi, Anutosh Maitra
FG3
2020 Conventional and Non-conventional Job Interviewing Methods: A Comparative Study in Two Countries
abstract
With recent advancements in technology, new platforms have come up to substitute face-to-face interviews. Of particular interest are asynchronous video interviewing (AVI) platforms, where candidates talk to a screen with questions, and virtual agent based interviewing platforms, where a human-like avatar interviews candidates. These anytime-anywhere interviewing systems scale up the overall reach of the interviewing process for firms, though they may not provide the best experience for the candidates. An important research question is how the candidates perceive such platforms and its impact on their performance and behavior. Also, is there an advantage of one setting vs. another i.e., Avatar vs. Platform? Finally, would such differences be consistent across cultures? In this paper, we present the results of a comparative study conducted in three different interview settings (i.e., Face-to-face, Avatar, and Platform), as well as two different cultural contexts (i.e., India and Switzerland), and analyze the differences in self-rated, others-rated performance, and automatic audiovisual behavioral cues.
Kumar Shubham, Emmanuelle Patricia Kleinlogel, Anaïs Butera, Marianne Schmid Mast, Dinesh Babu Jayagopi
ICMI5
2020 Automatic multimodal assessment of soft skills in social interactions: a review
Sowmya Rasipuram, Dinesh Babu Jayagopi
Multim. Tools Appl.2
2018 Unsupervised Speaker Cue Usage Detection in Public Speaking Videos
Dinesh Babu Jayagopi
BMVC2
2018 Automated Grading of Handwritten Essays
abstract
Automatic grading of handwritten essays is vital in evaluating the performance of students in educational settings, particularly in situations where language experts are rare. We build a system capable of taking the input as handwritten essays in image format and outputs the grading on the scale of 0-5; 0 being the worst and 5 being the best. The overall system integrates Optical Handwriting Recognition (OHR) and Automated Essay Scoring (AES)/grading. The handwritten essay is transcribed using a network composed of Multi-Dimensional Long Short Term Memory (MDLSTM) and convolution layers. The loss function is Connectionist Temporal Classification (CTC). The AES model is a 2-layer artificial neural network with a feature set based on pretrained GloVe word vectors. The results of grading of essays are compared for transcriptions of essays received from OHR system and transcriptions of essays done manually. The mutual agreement between the two shows a Quadratic Weighted Kappa score of 0.88. The results indicate that though the current OHR systems have transcription errors but as a whole can perform well for an application like AES.
Annapurna Sharma, Dinesh Babu Jayagopi
ICFHR2
2018 Predicting Engagement Intensity in the Wild Using Temporal Convolutional Network
abstract
Engagement is the holy grail of learning whether it is in a classroom setting or an online learning platform. Studies have shown that engagement of the student while learning can benefit students as well as the teacher if the engagement level of the student is known. It is difficult to keep track of the engagement of each student in a face-to-face learning happening in a large classroom. It is even more difficult in an online learning platform where, the user is accessing the material at different instances. Automatic analysis of the engagement of students can help to better understand the state of the student in a classroom setting as well as online learning platforms and is more scalable. In this paper we propose a framework that uses Temporal Convolutional Network (TCN) to understand the intensity of engagement of students attending video material from Massive Open Online Courses (MOOCs). The input to the TCN network is the statistical features computed on 10 second segments of the video from the gaze, head pose and action unit intensities available in OpenFace library. The ability of the TCN architecture to capture long term dependencies gives it the ability to outperform other sequential models like LSTMs. On the given test set in the EmotiW 2018 sub challenge-"Engagement in the Wild", the proposed approach with Dilated-TCN achieved an average mean square error of 0.079.
Chinchu Thomas, Nitin Nair, Dinesh Babu Jayagopi
ICMI3
2018 Indian Languages ASR: A Multilingual Phone Recognition Framework with IPA Based Common Phone-set, Predicted Articulatory Features and Feature fusion
K. Manjunath, K. Sreenivasa Rao, Dinesh Babu Jayagopi, V. Ramasubramanian 0001
INTERSPEECH3
2018 Vlogging Over Time: Longitudinal Impressions and Behavior in YouTube
abstract
YouTube vlogging, as a popular genre of ubiquitous social video, engages people in entertainment, civic, and social activities. Although several aspects of vlogging have been studied in media studies and multimedia analysis, the longitudinal angle of vlogging regarding recognition of personal state and trait impressions from behavior has not been yet analyzed. We present a study using behavioral data of vloggers who posted vlogs on YouTube for a period between three and six years. We use online crowdsourcing to collect a rich set of 21 impression variables for each video, including perceived personality, mood, skills, and expertise. Acoustic and motion features are extracted to characterize basic nonverbal behavior. The analysis shows that only a couple of perceived variables, including perceived expertise and perceived quality of audio and video, display weak temporal patterns. Furthermore, we show that the use of longitudinal data helps to improve the automatic inference of impressions for several of the impression variables.
Daniel Gatica-Perez, Dairazalia Sanchez-Cortes, Trinh Minh Tri Do, Dinesh Babu Jayagopi, Kazuhiro Otsuka
MUM4
2018 Automatic assessment of communication skill in interview-based interactions
Sowmya Rasipuram, Dinesh Babu Jayagopi
Multim. Tools Appl.2
2017 Automatic assessment of communication skill in non-conventional interview settings: a comparative study
abstract
Effective communication is an important social skill that facilitates us to interpret and connect with people around us and is of utmost importance in employment based interviews. This paper presents a methodical study and automatic measurement of communication skill of candidates in different modes of behavioural interviews. It demonstrates a comparative analysis of non-conventional methods of employment interviews namely 1) Interface-based asynchronous video interviews and 2) Written interviews (including a short essay). In order to achieve this, we have collected a dataset of 100 structured interviews from participants. These interviews are evaluated independently by two human expert annotators on rubrics specific to each of the settings. We, then propose a predictive model using automatically extracted multimodal features like audio, visual and lexical, applying classical machine learning algorithms. Our best model performs with an accuracy of 75% for a binary classification task in all the three contexts. We also study the differences between the expert perception and the automatic prediction across the settings.
Pooja S. B. Rao, Sowmya Rasipuram, Rahul Das, Dinesh Babu Jayagopi
ICMI4
2016 Asynchronous video interviews vs. face-to-face interviews for communication skill measurement: a systematic study
abstract
Communication skill is an important social variable in em- ployment interviews. As recent trends point to, increasingly asynchronous or interface-based video interviews are becom- ing popular. Also getting increasing interest is automatic hiring analysis, of which automatic communication skill pre- diction is one such task. In this context, a research gap that exists and which our paper addresses is â€oeAre there any differences in perception of communication skill and the accuracy of automatic prediction of say classes of communicators (e.g. those below average) when we compare interface-based and face-to-face interviewsâ€❝. To this end, we have collected a set of 106 interview videos from graduate students in both the settings i.e., interface-based and face-to-face. We observe that perception of behavior of participants in interface-based (when no person is involved) vs. face-to-face (when inter- viewer is involved) according to the external naive observers is slightly different. In this paper, we present an automatic system to predict the communication skill of a person in interface-based and face-to-face interviews by automatically extracting several low level features based on audio, visual and lexical behavior of the participants and using Machine Learning algorithms like Linear Regression, Support Vec- tor Machine (SVM) and Logistic Regression. We also make an extensive study of the verbal behavior of the participant when the spoken response is obtained from manual tran- scriptions and Automatic Speech Recognition (ASR) tool. Our best automatic prediction results achieve an accuracy of 80% in interface-based and 83% in face-to-face setting.
Sowmya Rasipuram, Pooja S. B. Rao, Dinesh Babu Jayagopi
ICMI3
2015 A robust lane detection and departure warning system
abstract
In this work, we have developed a robust lane detection and departure warning technique. Our system is based on single camera sensor. For lane detection a modified Inverse Perspective Mapping using only a few extrinsic camera parameters and illuminant Invariant techniques is used. Lane markings are represented using a combination of 2nd and 4th order steerable filters, robust to shadowing. Effect of shadowing and extra sun light are removed using Lab color space, and illuminant invariant representation. Lanes are assumed to be cubic curves and fitted using robust RANSAC. This method can reliably detect lanes of the road and its boundary. This method has been experimented in Indian road conditions under different challenging situations and the result obtained were very good. For lane departure angle an optical flow based method were used.
Mrinal Haloi, Dinesh Babu Jayagopi
Intelligent Vehicles Symposium2
2013 Given that, should i respond?: contextual addressee estimation in multi-party human-robot interactions
Dinesh Babu Jayagopi, Jean-Marc Odobez
HRI1
2013 The vernissage corpus: a conversational human-robot-interaction dataset
Dinesh Babu Jayagopi, Samira Sheikhi, David Klotz, Johannes Wienke, Jean-Marc Odobez, Sebastian Wrede 0001, Vasil Khalidov, Laurent Nyugen, Britta Wrede, Daniel Gatica-Perez
HRI1
2012 Linking speaking and looking behavior patterns with group composition, perception, and performance
abstract
This paper addresses the task of mining typical behavioral patterns from small group face-to-face interactions and linking them to social-psychological group variables. Towards this goal, we define group speaking and looking cues by aggregating automatically extracted cues at the individual and dyadic levels. Then, we define a bag of nonverbal patterns (Bag-of-NVPs) to discretize the group cues. The topics learnt using the Latent Dirichlet Allocation (LDA) topic model are then interpreted by studying the correlations with group variables such as group composition, group interpersonal perception, and group performance. Our results show that both group behavior cues and topics have significant correlations with (and predictive information for) all the above variables. For our study, we use interactions with unacquainted members i.e. newly formed groups.
Dinesh Babu Jayagopi, Dairazalia Sanchez-Cortes, Kazuhiro Otsuka, Junji Yamato, Daniel Gatica-Perez
ICMI1
2012 Modeling dominance effects on nonverbal behaviors using granger causality
abstract
In this paper we modeled the effects that dominant people might induce on the nonverbal behavior (speech energy and body motion) of the other meeting participants using Granger causality technique. Our initial hypothesis that more dominant people have generalized higher influence was not validated when using the DOME-AMI corpus as data source. However, from the correlational analysis some interesting patterns emerged: contradicting our initial hypothesis dominant individuals are not accounting for the majority of the causal flow in a social interaction. Moreover, they seem to have more intense causal effects as their causal density was significantly higher. Finally dominant individuals tend to respond to the causal effects more often with complementarity than with mimicry.
Kyriaki Kalimeri, Bruno Lepri, Oya Aran, Dinesh Babu Jayagopi, Daniel Gatica-Perez, Fabio Pianesi
ICMI4
2012 Privacy-sensitive recognition of group conversational context with sociometers
Dinesh Babu Jayagopi, Taemie Jung Kim, Alex Pentland, Daniel Gatica-Perez
Multim. Syst.1
2010 Recognizing conversational context in group interaction using privacy-sensitive mobile sensors
abstract
The availability of mobile sociometric sensors allows Computer-Supported Cooperative Work (CSCW) designers the possibility to enhance online meeting support through automatic recognition of conversational context. This paper addresses the task of discriminating one conversational context against another, specifically brainstorming from decision-making interactions using easily computable nonverbal behavioral cues. We hypothesize that the difference in the dynamics between brainstorming and decision-making discussions is significant and measurable using speech activity based nonverbal cues. We employ a set of nonverbal cues to characterize the entire group by the aggregation (both temporal and person-wise) of their nonverbal behavior. Our results on a dataset collected using privacy-sensitive sociometric badges show that the floor-occupation patterns in a brain-storming interaction are different from a decision-making interaction and we can obtain a discrimination accuracy as high as 87.5%.
Dinesh Babu Jayagopi, Taemie Jung Kim, Alex Pentland, Daniel Gatica-Perez
MUM1
2010 Mining Group Nonverbal Conversational Patterns Using Probabilistic Topic Models
abstract
The automatic discovery of group conversational behavior is a relevant problem in social computing. In this paper, we present an approach to address this problem by defining a novel group descriptor called bag of group-nonverbal-patterns (NVPs) defined on brief observations of group interaction, and by using principled probabilistic topic models to discover topics. The proposed bag of group NVPs allows fusion of individual cues and facilitates the eventual comparison of groups of varying sizes. The use of topic models helps to cluster group interactions and to quantify how different they are from each other in a formal probabilistic sense. Results of behavioral topics discovered on the Augmented Multi-Party Interaction (AMI) meeting corpus are shown to be meaningful using human annotation with multiple observers. Our method facilitates “group behavior-based” retrieval of group conversational segments without the need of any previous labeling.
Dinesh Babu Jayagopi, Daniel Gatica-Perez
IEEE Trans. Multim.1
2009 Characterizing conversational group dynamics using nonverbal behaviour
abstract
This paper addresses the novel problem of characterizing conversational group dynamics. It is well documented in social psychology that depending on the objectives a group, the dynamics are different. For example, a competitive meeting has a different objective from that of a collaborative meeting. We propose a method to characterize group dynamics based on the joint description of a group members' aggregated acoustical nonverbal behaviour to classify two meeting datasets (one being cooperative-type and the other being competitive-type). We use 4.5 hours of real behavioural multi-party data and show that our methodology can achieve a classification rate of upto 100%.
Dinesh Babu Jayagopi, Bogdan Raducanu, Daniel Gatica-Perez
ICME1
2009 Discovering group nonverbal conversational patterns with topics
abstract
This paper addresses the problem of discovering conversational group dynamics from nonverbal cues extracted from thin-slices of interaction. We first propose and analyze a novel thin-slice interaction descriptor - a bag of group nonverbal patterns - which robustly captures the turn-taking behavior of the members of a group while integrating its leader's position. We then rely on probabilistic topic modeling of the interaction descriptors which, in a fully unsupervised way, is able to discover group interaction patterns that resemble prototypical leadership styles proposed in social psychology. Our method, validated on the Augmented Multi-Party Interaction (AMI) meeting corpus, facilitates the retrieval of group conversational segments where semantically meaningful group behaviours emerge, without the need of any previous labeling.
Dinesh Babu Jayagopi, Daniel Gatica-Perez
ICMI1
2009 Modeling Dominance in Group Conversations Using Nonverbal Activity Cues
abstract
Dominance - a behavioral expression of power - is a fundamental mechanism of social interaction, expressed and perceived in conversations through spoken words and audiovisual nonverbal cues. The automatic modeling of dominance patterns from sensor data represents a relevant problem in social computing. In this paper, we present a systematic study on dominance modeling in group meetings from fully automatic nonverbal activity cues, in a multi-camera, multi-microphone setting. We investigate efficient audio and visual activity cues for the characterization of dominant behavior, analyzing single and joint modalities. Unsupervised and supervised approaches for dominance modeling are also investigated. Activity cues and models are objectively evaluated on a set of dominance-related classification tasks, derived from an analysis of the variability of human judgment of perceived dominance in group discussions. Our investigation highlights the power of relatively simple yet efficient approaches and the challenges of audiovisual integration. This constitutes the most detailed study on automatic dominance modeling in meetings to date.
Dinesh Babu Jayagopi, Hayley Hung, Chuohao Yeo, Daniel Gatica-Perez
IEEE Trans. Speech Audio Process.1
2008 Investigating automatic dominance estimation in groups from visual attention and speaking activity
abstract
We study the automation of the visual dominance ratio (VDR); a classic measure of displayed dominance in social psychology literature, which combines both gaze and speaking activity cues. The VDR is modified to estimate dominance in multi-party group discussions where natural verbal exchanges are possible and other visual targets such as a table and slide screen are present. Our findings suggest that fully automated versions of these measures can estimate effectively the most dominant person in a meeting and can match the dominance estimation performance when manual labels of visual attention are used.
Hayley Hung, Dinesh Babu Jayagopi, Sileye O. Ba, Jean-Marc Odobez, Daniel Gatica-Perez
ICMI2
2008 Predicting two facets of social verticality in meetings from five-minute time slices and nonverbal cues
abstract
This paper addresses the automatic estimation of two aspects of social verticality (status and dominance) in small-group meetings using nonverbal cues. The correlation of nonverbal behavior with these social constructs have been extensively documented in social psychology, but their value for computational models is, in many cases, still unknown. We present a systematic study of automatically extracted cues - including vocalic, visual activity, and visual attention cues - and investigate their relative effectiveness to predict both the most-dominant person and the high-status project manager from relative short observations. We use five hours of task-oriented meeting data with natural behavior for our experiments. Our work suggests that, although dominance and role-based status are related concepts, they are not equivalent and are thus not equally explained by the same nonverbal cues. Furthermore, the best cues can correctly predict the person with highest dominance or role-based status with an accuracy of 70% approximately.
Dinesh Babu Jayagopi, Sileye O. Ba, Jean-Marc Odobez, Daniel Gatica-Perez
ICMI1
2008 Predicting the dominant clique in meetings through fusion of nonverbal cues
abstract
This paper addresses the problem of automatically predicting the dominant clique (i.e., the set of K-dominant people) in face-to-face small group meetings recorded by multiple audio and video sensors. For this goal, we present a framework that integrates automatically extracted nonverbal cues and dominance prediction models. Easily computable audio and visual activity cues are automatically extracted from cameras and microphones. Such nonverbal cues, correlated to human display and perception of dominance, are well documented in the social psychology literature. The effectiveness of the cues were systematically investigated as single cues as well as in unimodal and multimodal combinations using unsupervised and supervised learning approaches for dominant clique estimation. Our framework was evaluated on a five-hour public corpus of teamwork meetings with third-party manual annotation of perceived dominance. Our best approaches can exactly predict the dominant clique with 80.8% accuracy in four-person meetings in which multiple human annotators agree on their judgments of perceived dominance.
Dinesh Babu Jayagopi, Hayley Hung, Chuohao Yeo, Daniel Gatica-Perez
ACM Multimedia1
2007 Using audio and video features to classify the most dominant person in a group meeting
abstract
The automated extraction of semantically meaningful information from multi-modal data is becoming increasingly necessary due to the escalation of captured data for archival. A novel area of multi-modal data labelling, which has received relatively little attention, is the automatic estimation of the most dominant person in a group meeting. In this paper, we provide a framework for detecting dominance in group meetings using different audio and video cues. We show that by using a simple model for dominance estimation we can obtain promising results.
Hayley Hung, Dinesh Babu Jayagopi, Chuohao Yeo, Gerald Friedland, Sileye O. Ba, Jean-Marc Odobez, Kannan Ramchandran, Nikki Mirghafori, Daniel Gatica-Perez
ACM Multimedia2
2007 Generalized adaptive IFIR filter bank structures
K. Rajgopal, Dinesh Babu Jayagopi, S. Venkataraman
Signal Process.2