Hayley Hung

dblp:13/4646 · DBLP profile ↗
← Back
69ranked-venue papers
17as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 11 first-author · 2 since 2021Artificial intelligence and machine learning · 25 · 2 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 20 · 5 first-author · 6 since 2021Computer networks · 2 · 1 first-author
YearPublicationVenuePosition
2026 Reflecti-Mate: A Conversational Agent for Adaptive Decision-Making Support Through System 1 and System 2 Thinking
abstract
Making high-stakes personal decisions involves cognitive, emotional, and intuitive processes, and individuals differ in how they allocate attention across these modes. Integration of these processes has shown to benefit decision making. Yet, most current decision-support systems focus primarily on supporting cognitive aspects, rather than adapting to the individual’s thinking profile to support integration of different types of thoughts. In this study, we investigate an agent designed to encourage integration by adapting to the individual user’s thought patterns. We explore its effects on participants’ perceptions of the agent and their reflective behavior, in comparison with unaided pre-reflection and a baseline agent. In a between-subjects study (N = 128), our agent, which fostered broad and elaborated thinking, enabled more personalized reflective trajectories, elicited more integrative reflective language, and was perceived as providing stronger support for holistic reflection. In contrast, the baseline agent produced homogenized profiles dominated by cognitive language across participants.
Morita Tarvirdians, Senthil Chandrasegaran, Hayley Hung, Catholijn M. Jonker, Catharine Oertel
UMAP3
2025 Technologies Supporting Self-Reflection on Social Interactions: A Systematic Review
abstract
As intelligent technology and applications have become an integral part of nearly all aspects of people's daily lives, many intelligent systems have been designed to help people navigate the complex space of social interactions. One prominent strategy for such intelligent support is providing meaningful Ad Hoc Interventions (ADI), e.g., through timely notifications. An alternative is Technology-Supported Reflection (TSR), e.g., by offering information about activities in one's past for personal insights. In contrast to straight-up interventions, the aim of the latter strategy is not to directly augment human skills but instead support learning and personal growth over time. However, while TSR has seen widespread interest in applications in some areas, such as physical fitness and mental health, its use for improving human social interactions has not yet been systematically explored. Concretely, it is currently unclear 1) what forms of self-reflection systems intend to support, 2) how their different technological components (e.g., data collection, information integration) are involved in providing support, and 3) what common limitations and design challenges they face. In this article, we present the results of a systematic literature review focusing on these questions to provide a structured foundation for targeted research. Concretely, we identified and analysed a collection of 23 relevant papers, each describing a system deploying TSR to support humans with elements of social interactions.We constructed a framework with a set of features to comprehensively describe and analyze the systems that support self-reflection, including their application domains, how they fit into the existing design framework, how they facilitate learning through reflection, how adaptive they are to individual users, and how they were evaluated. Finally, we propose a direction for designing systems that support individual's social interactions through self-reflection in an adaptive manner.
Chenxu Hao, Tiffany Matej Hrkalovic, Daniel Balliet, Hayley Hung, Bernd Dudzik
IUI4
2025 PARSEL: A Multimodal Dataset for Modeling Decision-Making Processes Involved in Selecting Partners for Joint Tasks
abstract
How people evaluate, select, and engage with others in cooperative settings significantly impacts their well-being, happiness, and success. However, navigating these processes is complex. Equipping systems with the ability to recognize, interpret, and even engage during such socio-cognitive processes can increase their potential to support humans in these socio-cognitive processes and be more successful in adjusting to the social environment they are embedded in (e.g., understanding human preferences and attitudes), leading to better quality interactions and decision-making for future partners. Yet, the developments of such systems depend on available datasets. However, based on our knowledge, no dataset exists that can be used to model partner selection for joint tasks. To support research focused on creating such intelligent systems, we introduce the PARSEL dataset – a comprehensive corpus of dyadic interactions designed for computational modeling of PARtner SELection processes and collaborative behavior. In total, 297 participants took part in the datasets. The dataset contains measurements of partner selection decisions over three different stages, as well as factors that may influence partner selection in the context of (online) social interactions. It includes audiovisual recordings that offer fine-grained behavioral cues used during these interactions, self-reported traits, and reported perceptions of person-, situation- and team-specific phenomena. By providing this resource, we aim to foster advancements in computational methods that can effectively model and augment socio-cognitive processes, contributing to socially aware intelligent systems and enhanced human-system interactions.
Tiffany Matej Hrkalovic, Bernd Dudzik, Daniel Balliet, Hayley Hung
IEEE Trans. Affect. Comput.4
2024 How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
Ailin Liu, Pepijn Vunderink, José Vargas Quiros, Chirag Raman, Hayley Hung
INTERSPEECH5
2024 Impact of Annotation Modality on Label Quality and Model Performance in the Automatic Assessment of Laughter In-the-Wild
abstract
Although laughter is known to be a multimodal signal, it is primarily annotated from audio. It is unclear how laughter labels may differ when annotated from modalities like video, which capture body movements and are relevant in in-the-wild studies. In this work we ask whether annotations of laughter are congruent across modalities, and compare the effect that labeling modality has on machine learning model performance. We compare annotations and models for laughter detection, intensity estimation, and segmentation, using a challenging in-the-wild conversational dataset with a variety of camera angles, noise conditions and voices. Our study with 48 annotators revealed evidence for incongruity in the perception of laughter and its intensity between modalities, mainly due to lower recall in the video condition. Our machine learning experiments compared the performance of modern unimodal and multi-modal models for different combinations of input modalities, training, and testing label modalities. In addition to the same input modalities rated by annotators (audio and video), we trained models with body acceleration inputs, robust to cross-contamination, occlusion and perspective differences. Our results show that performance of models with body movement inputs does not suffer when trained with video-acquired labels, despite their lower inter-rater agreement.
José Vargas Quiros, Laura Cabrera Quiros, Catharine Oertel, Hayley Hung
IEEE Trans. Affect. Comput.4
2023 Why Did This Model Forecast This Future? Information-Theoretic Saliency for Counterfactual Explanations of Probabilistic Regression Models
abstract
We propose a post hoc saliency-based explanation framework for counterfactual reasoning in probabilistic multivariate time-series forecasting (regression) settings. Building upon Miller's framework of explanations derived from research in multiple social science disciplines, we establish a conceptual link between counterfactual reasoning and saliency-based explanation techniques. To address the lack of a principled notion of saliency, we leverage a unifying definition of information-theoretic saliency grounded in preattentive human visual cognition and extend it to forecasting settings. Specifically, we obtain a closed-form expression for commonly used density functions to identify which observed timesteps appear salient to an underlying model in making its probabilistic forecasts. We empirically validate our framework in a principled manner using synthetic data to establish ground-truth saliency that is unavailable for real-world data. Finally, using real-world data and forecasting models, we demonstrate how our framework can assist domain experts in forming new data-driven hypotheses about the causal relationships between features in the wild.
Chirag Raman, Alec Nonnemaker, Amelia Villegas-Morcillo, Hayley Hung, Marco Loog
NeurIPS4
2023 Collecting Mementos: A Multimodal Dataset for Context-Sensitive Modeling of Affect and Memory Processing in Responses to Videos
abstract
In this article we introduceMementos: the first multimodal corpus for computational modeling of affect and memory processing in response to video content. It was collected online via crowdsourcing and captures 1995 individual responses collected from 297 unique viewers responding to 42 different segments of music videos. Apart from webcam recordings of their upper-body behavior (totaling 2012 minutes) and self-reports of their emotional experience, it contains detailed descriptions of the occurrence and content of 989 personal memories triggered by the video content. Finally, the dataset includes self-report measures related to individual differences in participants’ background and situation (Demographics,Personality, andMood), thereby facilitating the exploration of important contextual factors in research using the dataset. We describe 1) the construction and contents of the corpus itself, 2) analyse thevalidityof its content by investigating biases and consistency with existing research on affect and memory processing, 3) review previously published work that demonstrates theusefulnessof the multimodal data in the corpus for research on automated detection and prediction tasks, and 4) provide suggestions for how the dataset can be used in future research on modelingVideo-Induced Emotions,Memory-Associated Affect, andMemory Evocation.
Bernd Dudzik, Hayley Hung, Mark A. Neerincx, Joost Broekens
IEEE Trans. Affect. Comput.2
2023 Capturing Interaction Quality in Long Duration (Simulated) Space Missions With Wearables
abstract
Space exploration is evolving with the recent increase in interest and investment. For the success of planned long-duration crewed missions, good interpersonal interactions between crew members are crucial. In this study, we evaluate the use of wearables for detection and estimation of the quality of each social interaction participants have throughout a long mission rather than aggregate measures of interactions. Our proposed method utilizes Temporal Convolutional Networks(TCNs) for extracting individual representations from acceleration and audio streams and learnable pooling layers(NetVLAD) to aggregate these representations into fixed-size representations. Use of NetVLAD layers provides an intelligent alternative to simple aggregation for handling variable-sized interactions and interactions with missing data. We evaluate our method on a 4-month simulated space mission where 5 participants wore Sociometric Badges and provided reports on their interactions in terms of effectiveness, frustration, and satisfaction. Our method provides an average ROC-AUC score of 0.64. Since we are not aware of any comparable baselines, we compare our method to hand-crafted features formerly utilized for cohesion estimation in similar scenarios and show it significantly outperforms them. We also present ablation studies where we replace the components in our approach with well-known alternatives and show that they provide better performance than their respective counterparts.
Ekin Gedik, Jeffrey Olenick, Chu-Hsiang Chang, Steve W. J. Kozlowski, Hayley Hung
IEEE Trans. Affect. Comput.5
2023 Individual and Joint Body Movement Assessed by Wearable Sensing as a Predictor of Attraction in Speed Dates
abstract
Interpersonal attraction is known to motivate behavioral responses in the person experiencing this subjective phenomenon. Such responses may involve the imitation of behavior, as in mirroring or mimicry of postures or gestures, which have been found to be associated with the desire to be liked by an interlocutor. Speed dating provides a unique opportunity for the study of such behavioral manifestations of interpersonal attraction through the elimination of barriers to initiating communication, while maintaining significant ecological validity. In this paper we investigate the relationship between body movement, measured via accelerometer sensors, and self-reports or ratings of attraction and affiliation in a dataset of 399 speed dates between 72 subjects. Through machine learning experiments, we found that both features derived from a single individual's body movement and features designed to measure aspects of synchrony and convergence of the couple's body movement signals were predictive of different attraction ratings. Our statistical analysis revealed that the overall increase or decrease in an individual's body movement throughout an interaction is a potential indicator of friendly intentions, possibly related to the desire to affiliate.
José Vargas Quiros, Öykü Kapcak, Hayley Hung, Laura Cabrera Quiros
IEEE Trans. Affect. Comput.3
2023 Perceived Conversation Quality in Spontaneous Interactions
abstract
The quality of daily spontaneous conversations is of importance towards both our well-being as well as the development of interactive social agents. Prior research directly studying the quality of social conversations has operationalized it in narrow terms, associating greater quality to less small talk. Other works taking a broader perspective of interaction experience have indirectly studied quality through one of the several overlapping constructs such as rapport or engagement, in isolation. In this work we bridge this gap by proposing a holistic conceptualization of conversation quality, building upon the collaborative attributes of cooperative conversation floors. Taking a multilevel perspective of conversation, we develop and validate two instruments for perceived conversation quality (PCQ) at the individual and group levels. Specifically, we motivate capturing external raters’ gestalt impressions of participant experiences from thin slices of behavior, and collect annotations of PCQ on the publicly available MatchNMingle dataset of in-the-wild mingling conversations. Finally, we present an analysis of behavioral features that are predictive of PCQ. We find that for the conversations in MatchNMingle, raters tend to associate smaller group sizes, equitable speaking turns with fewer interruptions, and time taken for synchronous bodily coordination with higher PCQ.
Chirag Raman, Navin Raj Prabhu, Hayley Hung
IEEE Trans. Affect. Comput.3
2022 Exploring the Detection of Spontaneous Recollections during Video-viewing In-the-Wild using Facial Behavior Analysis
abstract
Intelligent systems might benefit from automatically detecting when a stimulus has triggered a user’s recollection of personal memories, e.g., to identify that a piece of media content holds personal significance for them. While computational research has demonstrated the potential to identify related states based on facial behavior (e.g., mind-wandering), the automatic detection of spontaneous recollections specifically has not been investigated this far. Motivated by this, we present machine learning experiments exploring the feasibility of detecting whether a video clip has triggered personal memories in a viewer based on the analysis of their Head Rotation, Head Position, Eye Gaze, and Facial Expressions. Concretely, we introduce an approach for automatic detection and evaluate its potential for predictions using in-the-wild webcam recordings. Overall, our findings demonstrate the capacity for above chance detections in both settings, with substantially better performance for the video-independent variant. Beyond this, we investigate the role of person-specific recollection biases for predictions of our video-independent models and the importance of specific modalities of facial behavior. Finally, we discuss the implications of our findings for detecting recollections and user-modeling in adaptive systems.
Bernd Dudzik, Hayley Hung
ICMI2
2022 Conversation Group Detection With Spatio-Temporal Context
abstract
In this work, we propose an approach for detecting conversation groups in social scenarios like cocktail parties and networking events, from overhead camera recordings. We posit the detection of conversation groups as a learning problem that could benefit from leveraging the spatial context of the surroundings, and the inherent temporal context in interpersonal dynamics which is reflected in the temporal dynamics in human behavior signals, an aspect that has not been addressed in recent prior works. This motivates our approach which consists of a dynamic LSTM-based deep learning model that predicts continuous pairwise affinity values indicating how likely two people are in the same conversation group. These affinity values are also continuous in time, since relationships and group membership do not occur instantaneously, even though the ground truths of group membership are binary. Using the predicted affinity values, we apply a graph clustering method based on Dominant Set extraction to identify the conversation groups. We benchmark the proposed method against established methods on multiple social interaction datasets. Our results showed that the proposed method improves group detection performance in data that has more temporal granularity in conversation group labels. Additionally, we provide an analysis in the predicted affinity values in relation to the conversation group detection. Finally, we demonstrate the usability of the predicted affinity values in a forecasting framework to predict group membership for a given forecast horizon.
Stephanie Tan, David M. J. Tax, Hayley Hung
ICMI3
2022 ConfLab: A Data Collection Concept, Dataset, and Benchmark for Machine Analysis of Free-Standing Social Interactions in the Wild
abstract
Recording the dynamics of unscripted human interactions in the wild is challenging due to the delicate trade-offs between several factors: participant privacy, ecological validity, data fidelity, and logistical overheads. To address these, following a 'datasets for the community by the community' ethos, we propose the Conference Living Lab (ConfLab): a new concept for multimodal multisensor data collection of in-the-wild free-standing social conversations. For the first instantiation of ConfLab described here, we organized a real-life professional networking event at a major international conference. Involving 48 conference attendees, the dataset captures a diverse mix of status, acquaintance, and networking motivations. Our capture setup improves upon the data fidelity of prior in-the-wild datasets while retaining privacy sensitivity: 8 videos (1920x1080, 60 fps) from a non-invasive overhead view, and custom wearable sensors with onboard recording of body motion (full 9-axis IMU), privacy-preserving low-frequency audio (1250 Hz), and Bluetooth-based proximity. Additionally, we developed custom solutions for distributed hardware synchronization at acquisition, and time-efficient continuous annotation of body keypoints and actions at high sampling rates. Our benchmarks showcase some of the open research tasks related to in-the-wild privacy-preserving social data analysis: keypoints detection from overhead camera views, skeleton-based no-audio speaker detection, and F-formation detection.
Chirag Raman, José Vargas Quiros, Stephanie Tan, Ashraful Islam, Ekin Gedik, Hayley Hung
NeurIPS6
2022 Multimodal Self-Assessed Personality Estimation During Crowded Mingle Scenarios Using Wearables Devices and Cameras
abstract
This paper focuses on the automatic classification of self-assessed personality traits from the HEXACO inventory during crowded mingle scenarios. These scenarios provide rich study cases for social behavior analysis but are also challenging to analyze automatically as people in them interact dynamically and freely in anin-the-wildface-to-face setting. To do so, we leverage the use of wearable sensors recording acceleration and proximity, and video from overhead cameras. We use 3 different behavioral modality types (movement, speech and proximity) coming from 2 sensors (wearable and camera). Unlike other works, we extract an individual’s speaking status from a single body worn triaxial accelerometer instead of audio, which scales easily to large populations. Additionally, we study the effect of different combinations of modality types on the personality estimation, and how this relates to the nature of each trait. We also include an analysis of feature complementarity and an evaluation of feature importance for the classification, showing that combining complementary modality types further improves the classification performance. We estimate the self-assessed personality traits both using a binary classification (community’s standard) and as a regression over the trait scores. Finally, we analyze the impact of the accuracy of the speech detection on the overall performance of the personality estimation.
Laura Cabrera Quiros, Ekin Gedik, Hayley Hung
IEEE Trans. Affect. Comput.3
2021 Insights on Group and Team Dynamics
abstract
We are organizing again the workshop on Interdisciplinary Insights into Group and Team Dynamics which is a joint effort between researchers in the the ICMI and INGRoup (Interdisciplinary Network for Group Research) communities. This workshop aims to provide a common destination for researchers to exchange ideas and collaborate. We have found in previous years that instigating interdisciplinary collaborations can be hard. The aim of this workshop is to sustain a joint community to foster continued cross-disciplinary exchange and mutual understanding.
Joseph A. Allen, Hayley Hung, Joann Keyton, Gabriel Murray, Catharine Oertel, Giovanna Varni
ICMI2
2021 Recognizing Perceived Interdependence in Face-to-Face Negotiations through Multimodal Analysis of Nonverbal Behavior
abstract
Enabling computer-based applications to display intelligent behavior in complex social settings requires them to relate to important aspects of how humans experience and understand such situations. One crucial driver of peoples’ social behavior during an interaction is the interdependence they perceive, i.e., how the outcome of an interaction is determined by their own and others’ actions. According to psychological studies, both the nonverbal behavior displayed by Motivated by this, we present a series of experiments to automatically recognize interdependence perceptions in dyadic face-to-face negotiations using these sources. Concretely, our approach draws on a combination of features describing individuals’ Facial, Upper Body, and Vocal Behavior with state-of-the-art algorithms for multivariate time series classification. Our findings demonstrate that differences in some types of interdependence perceptions can be detected through the automatic analysis of nonverbal behaviors. We discuss implications for developing socially intelligent systems and opportunities for future research.
Bernd Dudzik, Simon Columbus, Tiffany Matej Hrkalovic, Daniel Balliet, Hayley Hung
ICMI5
2021 Social Signals and Multimedia: Past, Present, Future
abstract
The rising popularity of Artificial Intelligence (AI) has brought considerable public interest as well faster and more direct transfer of research ideas into practice. One of the aspects of AI that still trails behind considerably is the role of machines in interpreting, enhancing, modeling, generating, and influencing social behavior. Such behavior is captured as social signals, usually by sensors recording multiple modalities, making it classic multimedia data. Such behavior can also be generated by an AI system when interacting with humans. Using AI techniques in combination with multimedia data can be used to pursue multiple goals, two of which are high-lighted here. First, supporting people during social interactions and helping them to fulfil their social needs either actively or passively.Second, improving our understanding of how people collaborate, build relationships, and process self identity. Despite the rise of fields such as Social Signal Processing, a similar panel organised at ACM Multimedia 2014, and an area on social and emotional signal sat the ACM MM since 2014, we argue that we have yet to truly fulfil the potential of the combining social signals and multimedia. This panel asks where we have come far enough and what remaining challenges there are in light of recent global events.
Hayley Hung, Cathal Gurrin, Martha A. Larson, Hatice Gunes, Fabien Ringeval, Elisabeth André, Louis-Philippe Morency
ACM Multimedia1
2021 Towards Analyzing and Predicting the Experience of Live Performances with Wearable Sensing
abstract
We present an approach to interpret the response of audiences to live performances by processing mobile sensor data. We apply our method on three different datasets obtained from three live performances, where each audience member wore a single tri-axial accelerometer and proximity sensor embedded inside a smart sensor pack. Using these sensor data, we developed a novel approach to predict audience members’ self-reported experience of the performances in terms of enjoyment, immersion, willingness to recommend the event to others, and change in mood. The proposed method uses an unsupervised method to identify informative intervals of the event, using the linkage of the audience members’ bodily movements, and uses data from these intervals only to estimate the audience members’ experience. We also analyze how the relative location of members of the audience can affect their experience and present an automatic way of recovering neighborhood information based on proximity sensors. We further show that the linkage of the audience members’ bodily movements is informative of memorable moments which were later reported by the audience.
Ekin Gedik, Laura Cabrera Quiros, Claudio Martella, Gwenn Englebienne, Hayley Hung
IEEE Trans. Affect. Comput.5
2021 The MatchNMingle Dataset: A Novel Multi-Sensor Resource for the Analysis of Social Interactions and Group Dynamics In-the-Wild During Free-Standing Conversations and Speed Dates
abstract
We present MatchNMingle, a novel multimodal/multisensor dataset for the analysis of free-standing conversational groups and speed-dates in-the-wild. MatchNMingle leverages the use of wearable devices and overhead cameras to record social interactions of 92 people during real-life speed-dates, followed by a cocktail party. To our knowledge, MatchNMingle has the largest number of participants, longest recording time and largest set of manual annotations for social actions available in this context in a real-life scenario. It consists of 2 hours of data from wearable acceleration, binary proximity, video, audio, personality surveys, frontal pictures and speed-date responses. Participants' positions and group formations were manually annotated; as were social actions (eg. speaking, hand gesture) for 30 minutes at 20 FPS making it the first dataset to incorporate the annotation of such cues in this context. We present an empirical analysis of the performance of crowdsourcing workers against trained annotators in simple and complex annotation tasks, founding that although efficient for simple tasks, using crowdsourcing workers for more complex tasks like social action annotation led to additional overhead and poor inter-annotator agreement compared to trained annotators (differences up to 0.4 in Fleiss' Kappa coefficients). We also provide example experiments of how MatchNMingle can be used.
Laura Cabrera Quiros, Andrew M. Demetriou, Ekin Gedik, Leander van der Meij, Hayley Hung
IEEE Trans. Affect. Comput.5
2021 On Social Involvement in Mingling Scenarios: Detecting Associates of F-Formations in Still Images
abstract
In this paper, we carry out an extensive study of social involvement in free standing conversing groups (the so-called F-formations) from static images. By introducing a novel feature representation, we show that the standard features which have been used to represent full membership in an F-formation cannot be applied to the detection of so-called associates of F-formations due to their sparser nature. We also enrich state-of-the-art F-formation modelling by learning a frustum of attention that accounts for the spatial context. That is, F-formation configurations vary with respect to the arrangement of furniture and the non-uniform crowdedness in the space during mingling scenarios. Moroever, the majority of prior works have considered the labelling of conversing groups as an objective task, requiring only a single annotator. However, we show that by embracing the subjectivity of social involvement, we not only generate a richer model of the social interactions in a scene but can use the detected associates to improve initial estimates of the full members of an F-formation. We carry out extensive experimental validation of our proposed approach by collecting a novel set of multi-annotator labels of involvement on two publicly available datasets; The Idiap Poster Data and SALSA data set. Moreover, we show that parameters learned from the Idiap Poster Data can be transferred to the SALSA data, showing the power of our proposed representation in generalising over new unseen data from a different environment.
Lu Zhang 0018, Hayley Hung
IEEE Trans. Affect. Comput.2
2020 Exploring Personal Memories and Video Content as Context for Facial Behavior in Predictions of Video-Induced Emotions
abstract
Empirical evidence suggests that the emotional meaning of facial behavior in isolation is often ambiguous in real-world conditions. While humans complement interpretations of others' faces with additional reasoning about context, automated approaches rarely display such context-sensitivity. Empirical findings indicate that the personal memories triggered by videos are crucial for predicting viewers' emotional response to such videos ?- in some cases, even more so than the video's audiovisual content. In this article, we explore the benefits of personal memories as context for facial behavior analysis. We conduct a series of multimodal machine learning experiments combining the automatic analysis of video-viewers' faces with that of two types of context information for affective predictions: \beginenumerate* [label=(\arabic*)] \item self-reported free-text descriptions of triggered memories and \item a video's audiovisual content \endenumerate*. Our results demonstrate that both sources of context provide models with information about variation in viewers' affective responses that complement facial analysis and each other.
Bernd Dudzik, Joost Broekens, Mark A. Neerincx, Hayley Hung
ICMI4
2020 Workshop on Interdisciplinary Insights into Group and Team Dynamics
abstract
There has been gathering momentum over the last 10 years in the study of group behavior in multimodal multiparty interactions. While many works in the computer science community focus on the analysis of individual or dyadic interactions, we believe that the study of groups adds an additional layer of complexity with respect to how humans cooperate and what outcomes can be achieved in these settings. Moreover, the development of technologies that can help to interpret and enhance group behaviours dynamically is still an emerging field. Social theories that accompany the study of groups dynamics are in their infancy and there is a need for more interdisciplinary dialogue between computer scientists and social scientists on this topic. This workshop has been organised to facilitate those discussions and strengthen the bonds between these overlapping research communities
Hayley Hung, Gabriel Murray, Giovanna Varni, Nale Lehmann-Willenbrock, Fabiola H. Gerpott, Catharine Oertel
ICMI1
2020 ConfFlow: A Tool to Encourage New Diverse Collaborations
abstract
ConfFlow is an interactive web application that allows conference participants to inspect other attendees through a visualized similarity space. The construction of the similarity space is done in a similar manner to the well-known Toronto Paper Matching System (TPMS) and based on the publicly available former publications of the attendees, obtained by crawling through the Web. ConfFlow aims to help attendees initiate new connections and collaborations with participants that have similar and/or complementary research interests. It has multiple functionalities that allow users to customize their experience and identify the perfect connection for their next collaboration.
Ekin Gedik, Hayley Hung
ACM Multimedia2
2020 A Modular Approach for Synchronized Wireless Multimodal Multisensor Data Acquisition in Highly Dynamic Social Settings
abstract
Existing data acquisition literature for human behavior research provides wired solutions, mainly for controlled laboratory setups. In uncontrolled free-standing conversation settings, where participants are free to walk around, these solutions are unsuitable. While wireless solutions are employed in the broadcasting industry, they can be prohibitively expensive. In this work, we propose a modular and cost-effective wireless approach for synchronized multisensor data acquisition of social human behavior. Our core idea involves a cost-accuracy trade-off by using Network Time Protocol (NTP) as a source reference for all sensors. While commonly used as a reference in ubiquitous computing, NTP is widely considered to be insufficiently accurate as a reference for video applications, where Precision Time Protocol (PTP) or Global Positioning System (GPS) based references are preferred. We argue and show, however, that the latency introduced by using NTP as a source reference is adequate for human behavior research, and the subsequent cost and modularity benefits are a desirable trade-off for applications in this domain. We also describe one instantiation of the approach deployed in a real-world experiment to demonstrate the practicality of our setup in-the-wild.
Chirag Raman, Stephanie Tan, Hayley Hung
ACM Multimedia3
2020 Investigating the Influence of Personal Memories on Video-Induced Emotions
abstract
This paper contributes to the automatic estimation of the subjective emotional experience that audio-visual media content induces in individual viewers, e.g. to support affect-based recommendations. Making accurate predictions of these responses is a challenging task because of their highly person-dependent and situation-specific nature. Findings from psychology indicate that an important driver for the emotional impact of media is the triggering of personal memories in observers. However, existing research on automated predictions focuses on the isolated analysis of audiovisual content, ignoring such contextual influences. In a series of empirical investigations, we (1) quantify the impact of associated personal memories on viewers' emotional responses to music videos in-the-wild and (2) assess the potential value of information about triggered memories for personalizing automatic predictions in this setting. Our findings indicate that the occurrence of memories intensifies emotional responses to videos. Moreover, information about viewers' memory response explains more variation in video-induced emotions than either the identity of videos or relevant viewer-characteristics (e.g. personality or mood). We discuss the implications of these results for existing approaches to automated predictions and describe ways for progress towards developing memory-sensitive alternatives.
Bernd Dudzik, Hayley Hung, Mark A. Neerincx, Joost Broekens
UMAP2
2020 Facial feedback for reinforcement learning: a case study and offline analysis using the TAMER framework
abstract
Abstract Interactive reinforcement learning provides a way for agents to learn to solve tasks from evaluative feedback provided by a human user. Previous research showed that humans give copious feedback early in training but very sparsely thereafter. In this article, we investigate the potential of agent learning from trainers’ facial expressions via interpreting them as evaluative feedback. To do so, we implemented TAMER which is a popular interactive reinforcement learning method in a reinforcement-learning benchmark problem—Infinite Mario, and conducted the first large-scale study of TAMER involving 561 participants. With designed CNN–RNN model, our analysis shows that telling trainers to use facial expressions and competition can improve the accuracies for estimating positive and negative feedback using facial expressions. In addition, our results with a simulation experiment show that learning solely from predicted feedback based on facial expressions is possible and using strong/effective prediction models or a regression method, facial responses would significantly improve the performance of agents. Furthermore, our experiment supports previous studies demonstrating the importance of bi-directional feedback and competitive elements in the training interface.
Guangliang Li, Hamdi Dibeklioglu, Shimon Whiteson, Hayley Hung
Auton. Agents Multi Agent Syst.4
2020 Gestures In-The-Wild: Detecting Conversational Hand Gestures in Crowded Scenes Using a Multimodal Fusion of Bags of Video Trajectories and Body Worn Acceleration
abstract
This paper addresses the detection of hand gestures during free-standing conversations in crowded mingle scenarios. Unlike the scenarios of the previous works in gesture detection and recognition, crowded mingle scenes have additional challenges such as cross-contamination between subjects, strong occlusions, and nonstationary backgrounds. This makes them more complex to analyze using computer vision techniques alone. We propose a multimodal approach using video and wearable acceleration data recorded via smart badges hung around the neck. In the video modality, we propose to treat noisy dense trajectories as bags-of-trajectories. For a given bag, we can have good trajectories corresponding to the subject, and bad trajectories due for instance to cross-contamination. However, we hypothesize that for a given class, it should be possible to learn trajectories that are discriminative while ignoring noisy trajectories. We do this by exploiting multiple instance learning via embedded instance selection as our multiple instance learning approach. This technique also allows us to identify which instances contribute more to the classification. By fusing the decisions of the classifiers from the video and wearable acceleration modalities, we show improvements over the unimodal approaches with an AUC of 0.69. We also present a static analysis and a dynamic analysis to assess the impact of noisy data on the fused detection results, showing that the moments of high occlusion in the video are compensated by the information from the wearables. Finally, we applied our method to detect speaking status, leveraging the close relationship found in the literature between hand gestures and speech.
Laura Cabrera Quiros, David M. J. Tax, Hayley Hung
IEEE Trans. Multim.3
2019 Context in Human Emotion Perception for Automatic Affect Detection: A Survey of Audiovisual Databases
abstract
An important aspect of human emotion perception is the use of contextual information to understand others' feelings even in situations where their behavior is not very expressive or has an emotionally ambiguous meaning. For technology to successfully detect affect, it must mimic this human ability when analyzing audiovisual input. Databases upon which machine learning algorithms are trained should capture the context of social interactions as well as the behavior expressed in them. However, there is a lack of consensus about what constitutes relevant context in such databases. In this article, we make two contributions towards overcoming this challenge: (a) we identify two principal sources of context for emotion perceptions based on psychological theory, and (b) we provide an overview of how each of these has been considered in published databases covering social interactions. Our results show that a similar set of contextual features are present across the reviewed databases. Between all the different databases researchers seem to have taken into account a set of contextual features reflecting the sources of context seen in psychological theory. However, within individual databases, these features are not yet systematically varied. This is problematic because it prevents them from being used directly as resources for the modeling of context-sensitive affect detection. Based on our findings, we suggest improvements for the future development of affective databases.
Bernd Dudzik, Michel-Pierre Jansen, Franziska Burger, Frank Kaptein, Joost Broekens, Dirk Heylen, Hayley Hung, Mark A. Neerincx, Khiet P. Truong
ACII7
2019 PANEL: Challenges for Multimedia/Multimodal Research in the Next Decade
abstract
The multimedia and multi-modal community is witnessing an explosive transformation in the recent years with major societal impact. With the unprecedented deployment of multimedia devices and systems, multimedia research is critical to our abilities and prospects in advancing state-of-the-art technologies and solving real-world challenges facing the society and the nation. To respond to these challenges and further advance the frontiers of the field of multimedia, this panel will discuss the challenges and visions that may guide future research in the next ten years.
Shih-Fu Chang, Louis-Philippe Morency, Alex Hauptmann 0001, Alberto Del Bimbo, Cathal Gurrin, Hayley Hung, Heng Ji 0001, Alan F. Smeaton
ACM Multimedia6
2019 Multimodal Data Collection for Social Interaction Analysis In-the-Wild
abstract
The benefits of exploiting multi-modality in the analysis of human-human social behaviour has been demonstrated widely in the community. An important aspect of this problem is the collection of data-sets that provide a rich and realistic representation of how people actually socialize with each other in real life. These subtle coordination patterns are influenced by individual beliefs, goals, and, desires related to what an individual stands to lose or gain in the activities they perform in their every day life. These conditions cannot be easily replicated in a lab setting and require a radical re-thinking of both how and what to collect. This tutorial provides a guide on how to create such multi-modal multi-sensor data sets when holistically considering the entire experimental design and data collection process.
Hayley Hung, Chirag Raman, Ekin Gedik, Stephanie Tan, José Vargas Quiros
ACM Multimedia1
2019 A Hierarchical Approach for Associating Body-Worn Sensors to Video Regions in Crowded Mingling Scenarios
abstract
We address the complex problem of associating several wearable devices with the spatio-temporal region of their wearers in video during crowded mingling events using only acceleration and proximity. This is a particularly important first step for multisensor behavior analysis using video and wearable technologies, where the privacy of the participants must be maintained. Most state-of-the-art works using these two modalities perform their association manually, which becomes practically unfeasible as the number of people in the scene increases. We proposed an automatic association method based on a hierarchical linear assignment optimization, which exploits the spatial context of the scene. Moreover, we present extensive experiments on matching from 2 to more than 69 acceleration and video streams, showing significant improvements over a random baseline in a real-world crowded mingling scenario. We also show the effectiveness of our method for incomplete or missing streams (up to a certain limit) and analyze the tradeoff between length of the streams and number of participants. Finally, we provide an analysis of failure cases, showing that deep understanding of the social actions within the context of the event is necessary to further improve performance on this intriguing task.
Laura Cabrera Quiros, Hayley Hung
IEEE Trans. Multim.2
2018 Group Interaction Frontiers in Technology
abstract
Analysis of group interaction and team dynamics is an important topic in a wide variety of fields, owing to the amount of time that individuals typically spend in small groups for both professional and personal purposes, and given how crucial group cohesion and productivity are to the success of businesses and other organizations. This fact is attested by the rapid growth of fields such as People Analytics and Human Resource Analytics, which in turn have grown out of many decades of research in social psychology, organizational behaviour, computing, and network science, amongst other fields. The goal of this workshop is to bring together researchers from diverse fields related to group interaction, team dynamics, people analytics, multi-modal speech and language processing, social psychology, and organizational behaviour.
Gabriel Murray, Hayley Hung, Joann Keyton, Catherine Lai, Nale Lehmann-Willenbrock, Catharine Oertel
ICMI2
2018 The I in Team: Mining Personal Social Interaction Routine with Topic Models from Long-Term Team Data
abstract
Social interaction plays a key role in assessing teamwork and collaboration. It becomes particularly critical in team performance when coupled with isolated, confined, and extreme conditions such as undersea missions. This work investigates how social interactions of individual members in a small team evolve during the course of a long duration mission. We propose to use a topic model to mine individual social interaction patterns and examine how the dynamics of these patterns have an effect on self-assessment of mood and team cohesion. Specifically, we analyzed data from a 6-person crew wearing Sociometric badges over a 4-month mission. Our results show that our method can extract the latent structure of social contexts without supervision. We demonstrate how the extracted patterns based on probabilistic models can provide insights on common behaviors at various temporal resolutions and exhibit links with self-report affective states and team cohesion.
Jeffrey Olenick, Chu-Hsiang Chang, Steve W. J. Kozlowski, Hayley Hung
IUI5
2018 Social interaction for efficient agent learning from human reward
abstract
Learning from rewards generated by a human trainer observing an agent in action has been proven to be a powerful method for teaching autonomous agents to perform challenging tasks, especially for those non-technical users. Since the efficacy of this approach depends critically on the reward the trainer provides, we consider how the interaction between the trainer and the agent should be designed so as to increase the efficiency of the training process. This article investigates the influence of the agent’s socio-competitive feedback on the human trainer’s training behavior and the agent’s learning. The results of our user study with 85 participants suggest that the agent’s passive socio-competitive feedback—showing performance and score of agents trained by trainers in a leaderboard—substantially increases the engagement of the participants in the game task and improves the agents’ performance, even though the participants do not directly play the game but instead train the agent to do so. Moreover, making this feedback active—sending the trainer her agent’s performance relative to others—further induces more participants to train agents longer and improves the agent’s learning. Our further analysis shows that agents trained by trainers affected by both the passive and active social feedback could obtain a higher performance under a score mechanism that could be optimized from the trainer’s perspective and the agent’s additional active social feedback can keep participants to further train agents to learn policies that can obtain a higher performance under such a score mechanism.
Guangliang Li, Shimon Whiteson, W. Bradley Knox, Hayley Hung
Auton. Agents Multi Agent Syst.4
2017 Estimating verbal expressions of task and social cohesion in meetings by quantifying paralinguistic mimicry
abstract
In this paper we propose a novel method of estimating verbal expressions of task and social cohesion by quantifying the dynamic alignment of nonverbal behaviors in speech. As team cohesion has been linked to team effectiveness and productivity, automatically estimating team cohesion can be a useful tool for assessing meeting quality and broader team functioning. In total, more than 20 hours of business meetings (3-8 people) were recorded and annotated for behavioral indicators of group cohesion, distinguishing between social and task cohesion. We hypothesized that behaviors commonly referred to as mimicry can be indicative of verbal expressions of social and task cohesion. Where most prior work targets mimicry of dyads, we investigated the effectiveness of quantifying group-level phenomena. A dynamic approach was adopted in which both the cohesion expressions and the paralinguistic mimicry were quantified on small time windows. By extracting features solely related to the alignment of paralinguistic speech behavior, we found that 2-minute high and low social cohesive regions could be classified with a 0.71 Area under the ROC curve, performing on par with the state-of-the-art where turn-taking features were used. Estimating task cohesion was more challenging, obtaining an accuracy of 0.64 AUC, outperforming the state-of-the-art. Our results suggest that our proposed methodology is successful in quantifying group-level paralinguistic mimicry. As both the state-of-the-art turn-taking features and mimicry features performed worse on estimating task cohesion, we conclude that social cohesion is more openly expressed by nonverbal vocal behavior than task cohesion.
Marjolein C. Nanninga, Nale Lehmann-Willenbrock, Zoltán Szlávik, Hayley Hung
ICMI5
2017 Personalised models for speech detection from body movements using transductive parameter transfer
abstract
We investigate the task of detecting speakers in crowded environments using a single body worn triaxial accelerometer. Detection of such behaviour is very challenging to model as people’s body movements during speech vary greatly. Similar to previous studies, by assuming that body movements are indicative of speech, we show experimentally, on a real-world dataset of 3 h including 18 people, that transductive parameter transfer learning (Zen et al. in Proceedings of the 16th international conference on multimodal interaction. ACM, 2014 ) can better model individual differences in speaking behaviour, significantly improving on the state-of-the-art performance. We also discuss the challenges introduced by the in-the-wild nature of our dataset and experimentally show how they affect detection performance. We strengthen the need for an adaptive approach by comparing the speech detection problem to a more traditional activity (i.e. walking). We provide an analysis of the transfer by considering different source sets which provides a deeper investigation of the nature of both speech and body movements, in the context of transfer learning.
Ekin Gedik, Hayley Hung
Pers. Ubiquitous Comput.2
2016 Beyond F-Formations: Determining Social Involvement in Free Standing Conversing Groups from Static Images
abstract
In this paper, we present the first attempt to analyse differing levels of social involvement in free standing conversing groups (or the so-called F-formations) from static images. In addition, we enrich state-of-the-art F-formation modelling by learning a frustum of attention that accounts for the spatial context. That is, F-formation configurations vary with respect to the arrangement of furniture and the non-uniform crowdedness in the space during mingling scenarios. The majority of prior works have considered the labelling of conversing group as an objective task, requiring only a single annotator. However, we show that by embracing the subjectivity of social involvement, we not only generate a richer model of the social interactions in a scene but also significantly improve F-formation detection. We carry out extensive experimental validation of our proposed approach by collecting a novel set of multi-annotator labels of involvement on the publicly available Idiap Poster Data, the only multi-annotator labelled database of free standing conversing groups that is currently available.
Lu Zhang 0018, Hayley Hung
CVPR2
2016 Estimating self-assessed personality from body movements and proximity in crowded mingling scenarios
abstract
This paper focuses on the automatic classification of self-assessed personality traits from the HEXACO inventory during crowded mingle scenarios. We exploit acceleration and proximity data from a wearable device hung around the neck. Unlike most state-of-the-art studies, addressing personality estimation during mingle scenarios provides a challenging social context as people interact dynamically and freely in a face-to-face setting. While many former studies use audio to extract speech-related features, we present a novel method of extracting an individual’s speaking status from a single body worn triaxial accelerometer which scales easily to large populations. Moreover, by fusing both speech and movement energy related cues from just acceleration, our experimental results show improvements on the estimation of Humility over features extracted from a single behavioral modality. We validated our method on 71 participants where we obtained an accuracy of 69% for Honesty, Conscientiousness and Openness to Experience. To our knowledge, this is the largest validation of personality estimation carried out in such a social context with simple wearable sensors.
Laura Cabrera Quiros, Ekin Gedik, Hayley Hung
ICMI3
2016 Who is where?: Matching People in Video to Wearable Acceleration During Crowded Mingling Events
abstract
We address the challenging problem of associating acceleration data from a wearable sensor with the corresponding spatio-temporal region of a person in video during crowded mingling scenarios. This is an important first step for multi-sensor behavior analysis using these two modalities. Clearly, as the numbers of people in a scene increases, there is also a need to robustly and automatically associate a region of the video with each person's device. We propose a hierarchical association approach which exploits the spatial context of the scene, outperforming the state-of-the-art approaches significantly. Moreover, we present experiments on matching from 3 to more than 130 acceleration and video streams which, to our knowledge, is significantly larger than prior works where only up to 5 device streams are associated.
Laura Cabrera Quiros, Hayley Hung
ACM Multimedia2
2016 Using informative behavior to increase engagement while learning from human reward
abstract
In this work, we address a relatively unexplored aspect of designing agents that learn from human reward. We investigate how an agent’s non-task behavior can affect a human trainer’s training and agent learning. We use the TAMER framework, which facilitates the training of agents by human-generated reward signals, i.e., judgements of the quality of the agent’s actions, as the foundation for our investigation. Then, starting from the premise that the interaction between the agent and the trainer should be bi-directional, we propose two new training interfaces to increase a human trainer’s active involvement in the training process and thereby improve the agent’s task performance. One provides information on the agent’s uncertainty which is a metric calculated as data coverage, the other on its performance. Our results from a 51-subject user study show that these interfaces can induce the trainers to train longer and give more feedback. The agent’s performance, however, increases only in response to the addition of performance-oriented information, not by sharing uncertainty levels. These results suggest that the organizational maxim about human behavior, “you get what you measure”—i.e., sharing metrics with people causes them to focus on optimizing those metrics while de-emphasizing other objectives—also applies to the training of agents. Using principle component analysis, we show how trainers in the two conditions train agents differently. In addition, by simulating the influence of the agent’s uncertainty–informative behavior on a human’s training behavior, we show that trainers could be distracted by the agent sharing its uncertainty levels about its actions, giving poor feedback for the sake of reducing the agent’s uncertainty without improving the agent’s performance.
Guangliang Li, Shimon Whiteson, W. Bradley Knox, Hayley Hung
Auton. Agents Multi Agent Syst.4
2016 Detecting conversational groups in images and sequences: A robust game-theoretic approach
Sebastiano Vascon, Eyasu Zemene Mequanint, Marco Cristani, Hayley Hung, Marcello Pelillo, Vittorio Murino
Comput. Vis. Image Underst.4
2016 Is automatic facial expression recognition of emotions coming to a dead end? The rise of the new kids on the block
abstract
Hatice Gunes’ work is partially supported the EPSRC under its IDEAS Factory Sandpits call on Digital Personhood (Grant Ref: EP/L00416X/1). Hayley Hung was partially supported by the Dutch national program COMMIT, by the European Commission under contract number FP7-ICT-600877 (SPENCER), and is affiliated with the Delft Data Science consortium.
Hatice Gunes, Hayley Hung
Image Vis. Comput.2
2015 Tutorial on Emotional and Social Signals for Multimedia Research
abstract
No abstract available.
Hayley Hung, Hatice Gunes
ACM Multimedia1
2015 How Was It?: Exploiting Smartphone Sensing to Measure Implicit Audience Responses to Live Performances
abstract
In this paper, we present an approach to understand the response of an audience to a live dance performance by the processing of mobile sensor data. We argue that exploiting sensing capabilities already available in smart phones enables a potentially large scale measurement of an audience's implicit response to a performance. In this work, we leverage both tri-axial accelerometers, worn by ordinary members of the public during a dance performance, to predict responses to a number of survey answers, comprising enjoyment, immersion, willingness to recommend the event to others, and change in mood. We also analyse how behaviour as a result of seeing a dance performance might be reflected in a people's subsequent social behaviour using proximity and acceleration sensing. To our knowledge, this is the first work where pervasive mobile sensing has been used to investigate spontaneous responses to predict the affective evaluation of a live performance. Using a single body worn accelerometer to monitor a set of audience members, we were able to predict whether they enjoyed the event with a balanced classification accuracy of 90\%. The collective coordination of the audience's bodily movements also highlighted memorable moments that were reported later by the audience. The effective use of body movements to measure affective responses in such a setting is particularly surprising given that traditionally, physiological signals such as skin conductance or brain-based signals are the more commonly accepted methods to measure implicit affective response. Our experiments open interesting new directions for research on both automated techniques and applications for the implicit tagging of real world events via spontaneous and implicit audience responses during as well as after a performance.
Claudio Martella, Ekin Gedik, Laura Cabrera Quiros, Gwenn Englebienne, Hayley Hung
ACM Multimedia5
2015 Brief Introduction to the Special Issue on Behavior Understanding for Arts and Entertainment
abstract
This editorial introduction describes the aims and scope of the special issue of the ACM Transactions on Interactive Intelligent Systems on Behavior Understanding for Arts and Entertainment, which is being published in issues 2 and 3 of volume 5 of the journal. Here we offer a brief introduction to the use of behavior analysis for interactive systems that involve creativity in either the creator or the consumer of a work of art. We then characterize each of the five articles included in this first part of the special issue, which span a wide range of applications.
Albert Ali Salah, Hayley Hung, Oya Aran, Hatice Gunes, Matthew Turk 0001
ACM Trans. Interact. Intell. Syst.2
2015 Behavior Understanding for Arts and Entertainment
abstract
This editorial introduction complements the shorter introduction to the first part of the two-part special issue on Behavior Understanding for Arts and Entertainment. It offers a more expansive discussion of the use of behavior analysis for interactive systems that involve creativity, either for the producer or the consumer of such a system. We first summarise the two articles that appear in this second part of the special issue. We then discuss general questions and challenges in this domain that were suggested by the entire set of seven articles of the special issue and by the comments of the reviewers of these articles.
Albert Ali Salah, Hayley Hung, Oya Aran, Hatice Gunes, Matthew Turk 0001
ACM Trans. Interact. Intell. Syst.2
2015 Introduction to: Special Issue on Extended Best Papers from ACM Multimedia 2014
abstract
No abstract available.
Hayley Hung, George Toderici
ACM Trans. Multim. Comput. Commun. Appl.1
2014 A Game-Theoretic Probabilistic Approach for Detecting Conversational Groups
Sebastiano Vascon, Eyasu Zemene Mequanint, Marco Cristani, Hayley Hung, Marcello Pelillo, Vittorio Murino
ACCV (5)4
2014 Detecting conversing groups with a single worn accelerometer
abstract
In this paper we propose the novel task of detecting groups of conversing people using only a single body-worn accelerometer per person. Our approach estimates each individual's social actions and uses the co-ordination of these social actions between pairs to identify group membership. The aim of such an approach is to be deployed in dense crowded environments. Our work differs significantly from previous approaches, which have tended to rely on audio and/or proximity sensing, often in much less crowded scenarios, for estimating whether people are talking together or who is speaking. Ultimately, we are interested in detecting who is speaking, who is conversing with whom, and from that, to infer socially relevant information about the interaction such as whether people are enjoying themselves, or the quality of their relationship in these extremely dense crowded scenarios. Striving towards this long-term goal, this paper presents a systematic study to understand how to detect groups of people who are conversing together in this setting, where we achieve a $64%$ classification accuracy using a fully automated system.
Hayley Hung, Gwenn Englebienne, Laura Cabrera Quiros
ICMI1
2013 Classifying social actions with a single accelerometer
abstract
In this paper, we estimate different types of social actions from a single body-worn accelerometer in a crowded social setting. Accelerometers have many advantages in such settings: they are impervious to environmental noise, unobtrusive, cheap, low-powered, and their readings are specific to a single person. Our experiments show that they are surprisingly informative of different types of social actions. The social actions we address in this paper are whether a person is speaking, laughing, gesturing, drinking, or stepping. To our knowledge, this is the first work to carry out experiments on estimating social actions from conversational behavior using only a wearable accelerometer. The ability to estimate such actions using just the acceleration opens up the potential for analyzing more about social aspects of people's interactions without explicitly recording what they are saying.
Hayley Hung, Gwenn Englebienne, Jeroen Kools
UbiComp1
2013 Fourth international workshop on human behavior understanding (HBU 2013)
abstract
With advances in pattern recognition and multimedia computing, it became possible to analyze human behavior via multimodal sensors, at different time-scales and at different levels of interaction and interpretation. This ability opens up enormous possibilities for multimedia and multimodal interaction, with a potential of endowing the computers with a capacity to attribute meaning to users' attitudes, preferences, personality, social relationships, etc., as well as to understand what people are doing, the activities they have been engaged in, their routines and lifestyles. This workshop gathers researchers dealing with the problem of modeling human behavior under its multiple facets with particular attention to interactions in arts, creativity, entertainment and edutainment.
Albert Ali Salah, Hayley Hung, Oya Aran, Hatice Gunes
ACM Multimedia2
2011 Detecting F-formations as dominant sets
abstract
The first step towards analysing social interactive behaviour in crowded environments is to identify who is interacting with whom. This paper presents a new method for detecting focused encounters or F-formations in a crowded, real-life social environment. An F-formation is a specific instance of a group of people who are congregated together with the intent of conversing and exchanging information with each other. We propose a new method of estimating F-formations using a graph clustering algorithm by formulating the problem in terms of identifying dominant sets. A dominant set is a form of maximal clique which occurs in edge weighted graphs. As well as using the proximity between people, body orientation information is used; we propose a socially motivated estimate of focus orientation (SMEFO), which is calculated with location information only. Our experiments show significant improvements in performance over the existing modularity cut algorithm and indicates the effectiveness of using a local social context for detecting F-formations.
Hayley Hung, Ben J. A. Kröse
ICMI1
2011 Move, and i will tell you who you are: detecting deceptive roles in low-quality data
abstract
Motion, like speech, provides information about one's emotional state. This work introduces an automated non-verbal audio-visual approach for detecting deceptive roles in multi-party conversations using low resolution video. We show how using simple features extracted from motion and speech improves over speech-only for the detection of deceptive roles. Our results show that deceptive players were recognised with significantly higher precision when video features were used. We improve the classification performance with 22.6% compared to our baseline.
Nimrod Raiman, Hayley Hung, Gwenn Englebienne
ICMI2
2011 Estimating Dominance in Multi-Party Meetings Using Speaker Diarization
abstract
With the increase in cheap commercially available sensors, recording meetings is becoming an increasingly practical option. With this trend comes the need to summarize the recorded data in semantically meaningful ways. Here, we investigate the task of automatically measuring dominance in small group meetings when only a single audio source is available. Past research has found that speaking length as a single feature, provides a very good estimate of dominance. For these tasks we use speaker segmentations generated by our automated faster than real-time speaker diarization algorithm, where the number of speakers is not known beforehand. From user-annotated data, we analyze how the inherent variability of the annotations affects the performance of our dominance estimation method. We primarily focus on examining of how the performance of the speaker diarization and our dominance tasks vary under different experimental conditions and computationally efficient strategies, and how this would impact on a practical implementation of such a system. Despite the use of a state-of-the-art speaker diarization algorithm, speaker segments can be noisy. On conducting experiments on almost 5 hours of audio-visual meeting data, our results show that the dominance estimation is robust to increasing diarization noise.
Hayley Hung, Gerald Friedland, Daniel Gatica-Perez
IEEE Trans. Speech Audio Process.1
2010 Are you Awerewolf? Detecting deceptive roles and outcomes in a conversational role-playing game
abstract
This paper addresses the task of automatically detecting outcomes of social interaction patterns, using non-verbal audio cues in competitive role-playing games (RPGs). For our experiments, we introduce a new data set which features 3 hours of audio-visual recordings of the popular “Are you a Werewolf?” RPG. Two problems are approached in this paper: Detecting lying or suspicious behavior using non-verbal audio cues in a social context and predicting participants' decisions in a game-day by analyzing speaker turns. Our best classifier exhibits a performance improvement of 87% over the baseline for detecting deceptive roles. Also, we show that speaker turn based features can be used to determine the outcomes in the initial stages of the game, when the group is large.
Gokul Chittaranjan, Hayley Hung
ICASSP2
2010 Speech/Non-Speech Detection in Meetings from Automatically Extracted low Resolution Visual Features
Hayley Hung, Sileye O. Ba
ICASSP1
2010 The idiap wolf corpus: exploring group behaviour in a competitive role-playing game
abstract
In this paper we present the Idiap Wolf Database. This is a audio-visual corpus containing natural conversational data of volunteers who took part in a competitive role-playing game. Four groups of 8-12 people were recorded. In total, just over 7 hours of interactive conversational data was collected. The data has been annotated in terms of the roles and outcomes of the game. There are 371 examples of different roles played over 50 games. Recordings were made with headset microphones, an 8-microphone array, and 3 video cameras and are fully synchronised. The novelty of this data is that some players have deceptive roles and the participants do not know what roles other people play.
Hayley Hung, Gokul Chittaranjan
ACM Multimedia1
2010 Encounter (resonances)
abstract
This work is about the remediation of one of Mark Rothko's Seagram murals through the composition of several online sources and additional digital rendering. Based on reproductions of Rothko's "Red on Maroon" found on the Internet, and using computer graphics compositing associated with moiré and specular lighting effects, "Encounter (Resonances)" offers a new approach to the presentation of a piece of work that allows a viewer to perceive some of its very subtle nuances. The work echoes Rothko's mixed media layered painting technique by using reproductions of various color palettes and resolutions as metaphors for the layers of paint in his original works. While each of these copies may instantly remind us of the original work, the graphical rendering of "Encounter (Resonances)" combines them at three levels of representation (global shape, micro and macro structure), in an effort to encourage a level of prolonged engagement and gradual discovery in the artwork.
Hayley Hung, Christian Jacquemin
ACM Multimedia1
2010 Estimating Cohesion in Small Groups Using Audio-Visual Nonverbal Behavior
abstract
Cohesiveness in teams is an essential part of ensuring the smooth running of task-oriented groups. Research in social psychology and management has shown that good cohesion in groups can be correlated with team effectiveness or productivity, so automatically estimating group cohesion for team training can be a useful tool. This paper addresses the problem of analyzing group behavior within the context of cohesion. Four hours of audio-visual group meeting data were used for collecting annotations on the cohesiveness of four-participant teams. We propose a series of audio and video features, which are inspired by findings in the social sciences literature. Our study is validated on a set of 61 2-min meeting segments which showed high agreement amongst human annotators when asked to identify meetings that have high or low cohesion.
Hayley Hung, Daniel Gatica-Perez
IEEE Trans. Multim.1
2010 Dialocalization: Acoustic speaker diarization and visual localization as joint optimization problem
abstract
The following article presents a novel audio-visual approach for unsupervised speaker localization in both time and space and systematically analyzes its unique properties. Using recordings from a single, low-resolution room overview camera and a single far-field microphone, a state-of-the-art audio-only speaker diarization system (speaker localization in time) is extended so that both acoustic and visual models are estimated as part of a joint unsupervised optimization problem. The speaker diarization system first automatically determines the speech regions and estimates “who spoke when,” then, in a second step, the visual models are used to infer the location of the speakers in the video. We call this process “dialocalization.” The experiments were performed on real-world meetings using 4.5 hours of the publicly available AMI meeting corpus. The proposed system is able to exploit audio-visual integration to not only improve the accuracy of a state-of-the-art (audio-only) speaker diarization, but also adds visual speaker localization at little incremental engineering and computation costs. The combined algorithm has different properties, such as increased robustness, that cannot be observed in algorithms based on single modalities. The article describes the algorithm, presents benchmarking results, explains its properties, and systematically discusses the contributions of each modality.
Gerald Friedland, Chuohao Yeo, Hayley Hung
ACM Trans. Multim. Comput. Commun. Appl.3
2009 Multi-modal speaker diarization of real-world meetings using compressed-domain video features
abstract
Speaker diarization is originally defined as the task of determining ldquowho spoke whenrdquo given an audio track and no other prior knowledge of any kind. The following article shows a multi-modal approach where we improve a state-of-the-art speaker diarization system by combining standard acoustic features (MFCCs) with compressed domain video features. The approach is evaluated on over 4.5 hours of the publicly available AMI meetings dataset which contains challenges such as people standing up and walking out of the room. We show a consistent improvement of about 34% relative in speaker error rate (21% DER) compared to a state-of-the-art audio-only baseline.
Gerald Friedland, Hayley Hung, Chuohao Yeo
ICASSP2
2009 Visual activity context for focus of attention estimation in dynamic meetings
abstract
We address the problem of recognizing, in dynamic meetings in which people do not remain seated all the time, the visual focus of attention (VFOA) of seated people from their head pose and contextual activity cues. We propose a model that comprises the VFOA of a meeting participant as the hidden state, and his head pose as the observation. To account for the presence of moving visual targets due to the dynamic nature of the meeting, the locations of the visual targets are used as an input variables to the head pose observation model. Contextual information is introduced in the VFOA dynamics through a slide activity variable and speaking or visual activity variables that relate people's focus to the meeting activity context. The main novelty of this paper is the introduction of visual activity context for FOA recognition to account for the correlation between a person's focus and the other people's gestures, hand and body motions. We evaluate our model on a large dataset of 5 hours. Our results show that, for VFOA estimation in meetings, visual activity contextual information can be as effective as speaking context.
Sileye O. Ba, Hayley Hung, Jean-Marc Odobez
ICME2
2009 Visual speaker localization aided by acoustic models
abstract
The following paper presents a novel audio-visual approach for unsupervised speaker locationing. Using recordings from a single, low-resolution room overview camera and a single far-field microphone, a state-of-the art audio-only speaker localization system (traditionally called speaker diarization) is extended so that both acoustic and visual models are estimated as part of a joint unsupervised optimization problem. The speaker diarization system first automatically determines the number of speakers and estimates "who spoke when", then, in a second step, the visual models are used to infer the location of the speakers in the video. The experiments were performed on real-world meetings using 4.5 hours of the publicly available AMI meeting corpus. The proposed system is able to exploit audio-visual integration to not only improve the accuracy of a state-of-the-art (audio-only) speaker diarization, but also adds visual speaker locationing at little incremental engineering and computation costs.
Gerald Friedland, Chuohao Yeo, Hayley Hung
ACM Multimedia3
2009 Modeling Dominance in Group Conversations Using Nonverbal Activity Cues
abstract
Dominance - a behavioral expression of power - is a fundamental mechanism of social interaction, expressed and perceived in conversations through spoken words and audiovisual nonverbal cues. The automatic modeling of dominance patterns from sensor data represents a relevant problem in social computing. In this paper, we present a systematic study on dominance modeling in group meetings from fully automatic nonverbal activity cues, in a multi-camera, multi-microphone setting. We investigate efficient audio and visual activity cues for the characterization of dominant behavior, analyzing single and joint modalities. Unsupervised and supervised approaches for dominance modeling are also investigated. Activity cues and models are objectively evaluated on a set of dominance-related classification tasks, derived from an analysis of the variability of human judgment of perceived dominance in group discussions. Our investigation highlights the power of relatively simple yet efficient approaches and the challenges of audiovisual integration. This constitutes the most detailed study on automatic dominance modeling in meetings to date.
Dinesh Babu Jayagopi, Hayley Hung, Chuohao Yeo, Daniel Gatica-Perez
IEEE Trans. Speech Audio Process.2
2008 Identifying dominant people in meetings from audio-visual sensors
abstract
This paper provides an overview of the area of automated dominance estimation in group meetings. We describe research in social psychology and use this to explain the motivations behind suggested automated systems. With the growth in availability of conversational data captured in meeting rooms, it is possible to investigate how multi-sensor data allows us to characterize non-verbal behaviors that contribute towards dominance. We use an overview of our own work to address the challenges and opportunities in this area of research.
Hayley Hung, Daniel Gatica-Perez
FG1
2008 Estimating the dominant person in multi-party conversations using speaker diarization strategies
abstract
In this paper, we apply speaker diarization strategies from a single source to the task of estimating the dominant person in a group meeting. Previous work has shown that speaking length is strongly correlated with perceived dominance. Here we investigate this in more depth by considering two dominance tasks where there is full and majority agreement amongst ground-truth annotators. In addition, we investigate how 24 different speed-up and algorithmic strategies, and source types lead to interesting outcomes when applied to dominance estimation. We obtained the best performance of 77% using our slowest scheme and a single distant microphone (SDM). Within the top 3 out of 24 performing experiments in both dominance tasks, we show that we can use the furthest SDM, with no prior knowledge of the number of speakers and the fastest diarization scheme, which performs 1.3 times faster than real-time.
Hayley Hung, Gerald Friedland, Daniel Gatica-Perez
ICASSP1
2008 Investigating automatic dominance estimation in groups from visual attention and speaking activity
abstract
We study the automation of the visual dominance ratio (VDR); a classic measure of displayed dominance in social psychology literature, which combines both gaze and speaking activity cues. The VDR is modified to estimate dominance in multi-party group discussions where natural verbal exchanges are possible and other visual targets such as a table and slide screen are present. Our findings suggest that fully automated versions of these measures can estimate effectively the most dominant person in a meeting and can match the dominance estimation performance when manual labels of visual attention are used.
Hayley Hung, Dinesh Babu Jayagopi, Sileye O. Ba, Jean-Marc Odobez, Daniel Gatica-Perez
ICMI1
2008 Predicting the dominant clique in meetings through fusion of nonverbal cues
abstract
This paper addresses the problem of automatically predicting the dominant clique (i.e., the set of K-dominant people) in face-to-face small group meetings recorded by multiple audio and video sensors. For this goal, we present a framework that integrates automatically extracted nonverbal cues and dominance prediction models. Easily computable audio and visual activity cues are automatically extracted from cameras and microphones. Such nonverbal cues, correlated to human display and perception of dominance, are well documented in the social psychology literature. The effectiveness of the cues were systematically investigated as single cues as well as in unimodal and multimodal combinations using unsupervised and supervised learning approaches for dominant clique estimation. Our framework was evaluated on a five-hour public corpus of teamwork meetings with third-party manual annotation of perceived dominance. Our best approaches can exactly predict the dominant clique with 80.8% accuracy in four-person meetings in which multiple human annotators agree on their judgments of perceived dominance.
Dinesh Babu Jayagopi, Hayley Hung, Chuohao Yeo, Daniel Gatica-Perez
ACM Multimedia2
2007 Using audio and video features to classify the most dominant person in a group meeting
abstract
The automated extraction of semantically meaningful information from multi-modal data is becoming increasingly necessary due to the escalation of captured data for archival. A novel area of multi-modal data labelling, which has received relatively little attention, is the automatic estimation of the most dominant person in a group meeting. In this paper, we provide a framework for detecting dominance in group meetings using different audio and video cues. We show that by using a simple model for dominance estimation we can obtain promising results.
Hayley Hung, Dinesh Babu Jayagopi, Chuohao Yeo, Gerald Friedland, Sileye O. Ba, Jean-Marc Odobez, Kannan Ramchandran, Nikki Mirghafori, Daniel Gatica-Perez
ACM Multimedia1