Hatice Gunes

dblp:38/1743 · DBLP profile ↗
← Back
132ranked-venue papers
14as first author
63since 2021 · last 2026
0000-0003-2407-3012ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 76 · 6 first-author · 43 since 2021Human-computer interaction and ubiquitous computing · 53 · 7 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 3 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 3 first-author · 15 since 2021Computer networks · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Social Robots as Active Safeguards for Children's Welfare: Community Stakeholder Insights and Design Recommendations
abstract
Integrating robots into children's spaces demands careful attention to the likelihood that children may disclose information indicating that their welfare is at risk. To consider the active involvement of social robots in safeguarding children, we interviewed 22 community stakeholders who work with children professionally in education, social services, healthcare or legal services in the UK and the USA. All stakeholders saw value in having robots create playful, safe spaces that can help alleviate the anxiety and reduce the emotional burden of the safeguarding process. However, they worried about a robot's ability to handle disclosures, interpret context and not overburden welfare services. We present hypothetical roles that robots could adopt within the safeguarding pipeline and map them to design considerations that emphasise the need for relational understanding, structured thresholds, and transparent data collection. Together, these considerations provide directions for ethical, child-centred robot technologies in safeguarding youth.
Nida Itrat Abbasi, Leigh Levinson, Selma Sabanovic, Hatice Gunes
HRI4
2026 Reimagining Social Robots as Recommender Systems: Foundations, Framework, and Applications
abstract
Personalization in social robots refers to the ability of the robot to meet the needs and/or preferences of an individual user. Existing approaches typically rely on large language models (LLMs) to generate context-aware responses based on user metadata and historical interactions or on adaptive methods such as reinforcement learning (RL) to learn from users’ immediate reactions in real time. However, these approaches fall short of comprehensively capturing user preferences–including long-term, short-term, and fine-grained aspects–, and of using them to rank and select actions, proactively personalize interactions, and ensure ethically responsible adaptations. To address the limitations, we propose drawing on recommender systems (RSs), which specialize in modeling user preferences and providing personalized recommendations. To ensure the integration of RS techniques is well-grounded and seamless throughout the social robot pipeline, we (i) align the paradigms underlying social robots and RSs, (ii) identify key techniques that can enhance personalization in social robots, and (iii) design them as modular, plug-and-play components. This work not only establishes a framework for integrating RS techniques into social robots but also opens a pathway for deep collaboration between the RS and HRI communities, accelerating innovation in both fields.
Jin Huang 0010, Fethiye Irmak Dogan, Hatice Gunes
HRI3
2026 Social Robotics for Disabled Students: An Empirical Investigation of Embodiment, Roles, and Interaction
abstract
Institutional and social barriers in higher education often prevent students with disabilities from effectively accessing support, including lengthy procedures, insufficient information, and high social-emotional demands. This study empirically explores how disabled students perceive robot-based support, comparing two interaction roles, one information based (signposting) and one disclosure based (sounding board), and two embodiment types (physical robot/disembodied voice agent). Participants assessed these systems across five dimensions: perceived understanding, social energy demands, information access/clarity, task difficulty, and data privacy concerns. The main findings of the study reveal that the physical robot was perceived as more understanding than the voice-only agent, with embodiment significantly shaping perceptions of sociability, animacy, and privacy. We also analyse differences between disability types. These results provide critical insights into the potential of social robots to mitigate accessibility barriers in higher education, while highlighting ethical, social and technical challenges.
Alva Markelius, Fethiye Irmak Dogan, Julie Bailey, Guy Laban, Jenny L. Gibson, Hatice Gunes
HRI6
2026 Graph in Graph Neural Network
abstract
Abstract Existing Graph Neural Networks (GNNs) are limited to process graphs each of whose vertices is represented by a vector or a single value, limited their representing capability to describe complex objects. In this paper, we propose a novel GNN (called Graph in Graph Neural (GIG) Network) which can process graph-style data (called GIG sample) whose vertices are further represented by graphs. Given a set of graphs or a data sample whose components can be represented by a set of graphs (called multi-graph data sample), our GIG network starts with a GIG sample generation (GSG) module which encodes the input as a GIG sample , where each GIG vertex includes a graph. Then, a set of GIG hidden layers are stacked, with each consisting of: (1) a GIG vertex-level updating (GVU) module that individually updates the graph in every GIG vertex based on its internal information; and (2) a global-level GIG sample updating (GGU) module that updates graphs in all GIG vertices based on their relationships, making the updated GIG vertices become global context-aware. This way, both internal cues within the graph contained in each GIG vertex and the relationships among GIG vertices could be utilized for down-stream tasks. Experimental results demonstrate that our GIG network generalizes well for not only various generic graph analysis tasks but also real-world multi-graph data analysis (e.g., human skeleton video-based action recognition), which achieved the new state-of-the-art results on 15 out of 16 evaluated datasets. Our code is publicly available at https://github.com/wangjs96/Graph-in-Graph-Neural-Network .
Jiongshu Wang, Jing Yang 0038, Jiankang Deng, Hatice Gunes, Siyang Song
Int. J. Comput. Vis.4
2026 Past, Present, and Future: A Survey of the Evolution of Affective Robotics for Well-Being
abstract
Recent research in affective robots has recognized their potential in supporting human well-being. Due to rapidly developing affective and artificial intelligence technologies, this field of research has undergone explosive expansion and advancement in recent years. In order to develop a deeper understanding of recent advancements, we present a systematic review of the past 10 years of research in affective robotics for wellbeing. In this review, we identify the domains of well-being that have been studied, the methods used to investigate affective robots for well-being, and how these have evolved over time. We also examine the evolution of the multifaceted research topic from three lenses: technical, design, and ethical. Finally, we discuss future opportunities for research based on the gaps we have identified in our review – proposing pathways to take affective robotics from the past and present to the future. The results of our review are of interest to human-robot interaction and affective computing researchers, as well as clinicians and well-being professionals who may wish to examine and incorporate affective robotics in their practices.
Micol Spitale, Minja Axelsson, Sooyeon Jeong, Paige Tuttosi, Caitlin A. Stamatis, Guy Laban, Angelica Lim, Hatice Gunes
IEEE Trans. Affect. Comput.8
2026 AsyReC: A Multimodal Graph-Based Framework for Spatio-Temporal Asymmetric Dyadic Relationship Classification
abstract
Dyadic social relationships, which refer to relationships between two individuals who know each other through repeated interactions (or not), are shaped by shared spatial and temporal experiences. Current computational methods for modeling these relationships face three major challenges: (1) the failure to model asymmetric relationships, e.g., one individual may perceive the other as afriendwhile the other perceives them as anacquaintance, (2) the disruption of continuous interactions by discrete frame sampling, which segments the temporal continuity of interaction in real-world scenarios, and (3) the limitation to consider periodic behavioral cues, such as rhythmic vocalizations or recurrent gestures, which are crucial for inferring the evolution of dyadic relationships. To address these challenges, we propose AsyReC, a multimodal graph-based framework for asymmetric dyadic relationship classification, with three core innovations: (i) a triplet graph neural network with node-edge dual attention that dynamically weights multimodal cues to capture interaction asymmetries (addressing challenge 1); (ii) a clip-level relationship learning architecture that preserves temporal continuity, enabling fine-grained modeling of real-world interaction dynamics (addressing challenge 2); and (iii) a periodic temporal encoder that projects time indices onto sine/cosine waveforms to model recurrent behavioral patterns (addressing challenge 3). Extensive experiments on two public datasets demonstrate state-of-the-art performance, while ablation studies validate the critical role of asymmetric interaction modelling and periodic temporal encoding in improving the robustness of dyadic relationship classification in real-world scenarios. Our code is publicly available at: https://github.com/tw-repository/AsyReC.
Wang Tang, Fethiye Irmak Dogan, Linbo Qing, Hatice Gunes
IEEE Trans. Circuits Syst. Video Technol.4
2025 PerReactor: Offline Personalised Multiple Appropriate Facial Reaction Generation
abstract
In dyadic human-human interactions, individuals may express multiple different facial reactions in response to the same/similar behaviours expressed by their conversational partners depending on their personalised behaviour patterns. As a result, frequently-employed reconstruction loss-based strategies lead the training of previous automatic facial reaction generation (FRG) models to not only suffer from the 'one-to-many mapping' problem, but also fail to comprehensively consider the quality of the generated facial reactions. Besides, none of them considered such personalised behaviour patterns in generating facial reactions. In this paper, we propose the first adversarial FRG model training strategy which jointly learns appropriateness and realism discriminators to provide comprehensive task-specific supervision for training the target facial reaction generators, and reformulates the 'one-to-many (facial reactions) mapping' training problem as a 'one-to-one (distribution) mapping' training task, i.e., the FRG model is trained to output a distribution representing multiple appropriate/plausible facial reaction from each input human behaviour. In addition, our approach also serves as the first offline FRG approach that considers personalised behaviour patterns in generating of target individuals' facial reactions. Experiments show that our PerReactor not only largely outperformed all existing offline solutions for generating more appropriate, diverse and realistic facial reactions, but also is the first approach that can effectively generate personalised appropriate facial reactions.
Hengde Zhu, Xiangyu Kong 0001, Weicheng Xie 0001, Xilin He, Lu Liu 0001, LinLin Shen, Wei Zhang 0243, Hatice Gunes, Siyang Song
AAAI9
2025 Gender Fairness of Machine Learning Algorithms for Pain Detection
abstract
Automated pain detection through machine learning (ML) and deep learning (DL) algorithms holds significant potential in healthcare, particularly for patients unable to self-report pain levels. However, the accuracy and fairness of these algorithms across different demographic groups (e.g., gender) remain under-researched. This paper investigates the gender fairness of ML and DL models trained on the UNBC-McMaster Shoulder Pain Expression Archive Database, evaluating the performance of various models in detecting pain based solely on the visual modality of participants’ facial expressions. We compare traditional ML algorithms, Linear Support Vector Machine (L SVM) and Radial Basis Function SVM (RBF SVM), with DL methods, Convolutional Neural Network (CNN) and Vision Transformer (ViT), using a range of performance and fairness metrics. While ViT achieved the highest accuracy and a selection of fairness metrics, all models exhibited gender-based biases. These findings highlight the persistent trade-off between accuracy and fairness, emphasising the need for fairness-aware techniques to mitigate biases in automated healthcare systems.
Yuting Shang, Jiaee Cheong, Yang Liu 0182, Hatice Gunes
FG5
2025 GRACE: Generating Socially Appropriate Robot Actions Leveraging LLMs and Human Explanations
abstract
When operating in human environments, robots need to handle complex tasks while both adhering to social norms and accommodating individual preferences. For instance, based on common sense knowledge, a household robot can pre-dict that it should avoid vacuuming during a social gathering, but it may still be uncertain whether it should vacuum before or after having guests. In such cases, integrating common-sense knowledge with human preferences, often conveyed through human explanations, is fundamental yet a challenge for existing systems. In this paper, we introduce GRACE, a novel approach addressing this while generating socially appropriate robot actions. GRACE leverages common sense knowledge from LLMs, and it integrates this knowledge with human explanations through a generative network. The bidirectional structure of GRACE enables robots to refine and enhance LLM predictions by utilizing human explanations and makes robots capable of generating such explanations for human-specified actions. Our evaluations show that integrating human explanations boosts GRACE's performance, where it outperforms several baselines and provides sensible explanations.
Fethiye Irmak Dogan, Umut Ozyurt, Gizem Cinar, Hatice Gunes
ICRA4
2025 ERR@HRI 2.0 Challenge: Multimodal Detection of Errors and Failures in Human-Robot Conversations
abstract
The integration of large language models (LLMs) into conversational robots has made human-robot conversations more dynamic. Yet, LLM-powered conversational robots remain prone to errors, e.g., misunderstanding user intent, prematurely interrupting users, or failing to respond altogether. Detecting and addressing these failures is critical for preventing conversational breakdowns, avoiding task disruptions, and sustaining user trust. To tackle this problem, the ERR@HRI 2.0 Challenge provides a multimodal dataset of LLM-powered conversational robot failures during human-robot conversations and encourages researchers to benchmark machine learning models designed to detect robot failures. The dataset includes 16 hours of dyadic human-robot interactions, incorporating facial, speech, and head movement features. Each interaction is annotated with the presence or absence of robot errors from the system perspective, and perceived user intention to correct for a mismatch between robot behavior and user expectation. Participants are invited to form teams and develop machine learning models that detect these failures using multimodal data. Submissions will be evaluated using various performance metrics, including detection accuracy and false positive rate. This challenge represents another key step toward improving failure detection in human-robot interaction through social signal analysis.
Shiye Cao, Maia Stiber, Amama Mahmood, Maria Teresa Parreira, Wendy Ju, Micol Spitale, Hatice Gunes, Chien-Ming Huang 0001
ACM Multimedia7
2025 REACT 2025: the Third Multiple Appropriate Facial Reaction Generation Challenge
abstract
In dyadic interactions, a broad spectrum of human facial reactions might be appropriate for responding to each human speaker behaviour. Following the successful organisation of the REACT 2023 and REACT 2024 challenges, we are proposing the REACT 2025 challenge encouraging the development and benchmarking of Machine Learning (ML) models that can be used to generate multiple appropriate, diverse, realistic and synchronised human-style facial reactions expressed by human listeners in response to an input stimulus (i.e., audio-visual behaviours expressed by their corresponding speakers). As a key of the challenge, we provide challenge participants with the first natural and large-scale multi-modal Multiple Appropriate Facial Reaction Generation (MAFRG) dataset (called MARS) recording 136 human-human dyadic interactions containing a total of 2856 interaction sessions covering five different topics. In addition, this paper also presents the challenge guidelines and the performance of our baselines on the two proposed sub-challenges: Offline MAFRG and Online MAFRG, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2025
Siyang Song, Micol Spitale, Xiangyu Kong 0001, Hengde Zhu, Cristina Palmero, Germán Barquero, Sergio Escalera, Michel F. Valstar, Mohamed Daoudi, Tobias Baur 0001, Fabien Ringeval, Andrew Howes 0001, Elisabeth André, Hatice Gunes
ACM Multimedia15
2025 Some Optimizers are More Equal: Understanding the Role of Optimizers in Group Fairness
abstract
We study whether and how the choice of optimization algorithm can impact group fairness in deep neural networks. Through stochastic differential equation analysis of optimization dynamics in an analytically tractable setup, we demonstrate that the choice of optimization algorithm indeed influences fairness outcomes, particularly under severe imbalance. Furthermore, we show that when comparing two categories of optimizers, adaptive methods and stochastic methods, RMSProp (from the adaptive category) has a higher likelihood of converging to fairer minima than SGD (from the stochastic category). Building on this insight, we derive two new theoretical guarantees showing that, under appropriate conditions, RMSProp exhibits fairer parameter updates and improved fairness in a single optimization step compared to SGD. We then validate these findings through extensive experiments on three publicly available datasets, namely CelebA, FairFace, and MS-COCO, across different tasks as facial expression recognition, gender classification, and multi-label classification, using various backbones. Considering multiple fairness definitions including equalized odds, equal opportunity, and demographic parity, adaptive optimizers like RMSProp and Adam consistently outperform SGD in terms of group fairness, while maintaining comparable predictive accuracy. Our results highlight the role of adaptive updates as a crucial yet overlooked mechanism for promoting fair outcomes. We release the source code at: https://github.com/Mkolahdoozi/Some-Optimizers-Are-More-Equal.
Mojtaba Kolahdouzi, Hatice Gunes, Ali Etemad
NeurIPS2
2025 Robot-Led Vision Language Model Wellbeing Assessment of Children
abstract
This study presents a novel robot-led approach to assessing children’s mental wellbeing using a Vision Language Model (VLM). Inspired by the Child Apperception Test (CAT), the social robot NAO presented children with pictorial stimuli to elicit their verbal narratives of the images, which were then evaluated by a VLM in accordance with CAT assessment guidelines. The VLM’s assessments were systematically compared to those provided by a trained psychologist. The results reveal that while the VLM demonstrates moderate reliability in identifying cases with no wellbeing concerns, its ability to accurately classify assessments with wellbeing concerns remains limited. Moreover, although the model’s performance was generally consistent when prompted with varying demographic factors such as age and gender, a significantly higher false positive rate was observed for girls, indicating potential sensitivity to gender attribute. These findings highlight both the promise and the challenges of integrating VLMs into robot-led assessments of children’s wellbeing.
Nida Itrat Abbasi, Fethiye Irmak Dogan, Guy Laban, Joanna Anderson, Tamsin Ford, Peter B. Jones, Hatice Gunes
RO-MAN7
2025 Self-Disclosure Themes and Semantics Across Human, Robotic, and Disembodied Conversational Partners
abstract
As social robots and other artificial agents become more conversationally capable, it is important to understand whether the content and meaning of self-disclosure towards these agents changes depending on the agent’s embodiment. In this study, we analysed conversational data from three controlled experiments in which participants self-disclosed to a human, a humanoid social robot, and a disembodied conversational agent. Using sentence embeddings and clustering, we identified themes in participants’ disclosures, which were then labelled and explained by a large language model. We subsequently assessed whether these themes and the underlying semantic structure of the disclosures varied by agent embodiment. Our findings reveal strong consistency: thematic distributions did not significantly differ across embodiments, and semantic similarity analyses showed that disclosures were expressed in highly comparable ways. These results suggest that while embodiment may influence human behaviour in human–robot and human–agent interactions, people tend to maintain a consistent thematic focus and semantic structure in their disclosures, whether speaking to humans or artificial interlocutors.
Sophie Chiang, Guy Laban, Emily S. Cross, Hatice Gunes
RO-MAN4
2025 What People Share With a Robot When Feeling Lonely and Stressed and How It Helps Over Time
abstract
Loneliness and stress are prevalent among young adults and are linked to significant psychological and health-related consequences. Social robots may offer a promising avenue for emotional support, especially when considering the ongoing advancements in conversational AI. This study investigates how repeated interactions with a social robot influence feelings of loneliness and perceived stress, and how such feelings are reflected in the themes of user disclosures towards the robot. Participants engaged in a five-session robot-led intervention, where a LLM-powered QTrobot facilitated structured conversations designed to support cognitive reappraisal. Results from linear mixed-effects models show significant reductions in both loneliness and perceived stress over time. Additionally, semantic clustering of 560 user disclosures towards the robot revealed six distinct conversational themes. Results from Kruskal-Wallis H-test demonstrate that participants reporting higher loneliness and stress, more frequently engaged in socially focused disclosures, such as friendship and connection, whereas lower distress was associated with introspective and goal-oriented themes (e.g., academic ambitions). By exploring both how the intervention affects well-being, as well as how well-being shapes the content of robot-directed conversations, we aim to capture the dynamic nature of emotional support in human–robot interaction.
Guy Laban, Sophie Chiang, Hatice Gunes
RO-MAN3
2025 Exploring Causality for HRI: A Case Study on Robotic Mental Well-being Coaching
abstract
One of the primary goals of Human-Robot Interaction (HRI) research is to develop robots that can interpret human behavior and adapt their responses accordingly. Adaptive learning models, such as continual and reinforcement learning, play a crucial role in improving robots’ ability to interact effectively in real-world settings. However, these models face significant challenges due to the limited availability of real-world data, particularly in sensitive domains like healthcare and well-being. To address these challenges, causality provides a structured framework for understanding and modeling the underlying relationships between actions, events, and outcomes. By moving beyond mere pattern recognition, causality enables robots to make more explainable and generalizable decisions. This paper presents an exploratory causality-based analysis through a case study of an adaptive robotic coach delivering positive psychology exercises over four weeks in a workplace setting. The robotic coach autonomously adapts to multimodal human behaviors, such as facial valence and speech duration. By conducting both macro- and micro-level causal analyses, this study aims to gain deeper insights into how adaptability can enhance well-being during interactions. Ultimately, this research seeks to advance our understanding of how causality can help overcome challenges in HRI, particularly in real-world applications.
Micol Spitale, Srikar Babu, Serhan Cakmak, Jiaee Cheong, Hatice Gunes
RO-MAN5
2025 Two-Stage Temporal Modelling Framework for Video-Based Depression Recognition Using Graph Representation
abstract
Video-based automatic depression analysis provides a fast, objective and repeatable self-assessment solution, which has been widely developed in recent years. While depression cues may be reflected by human facial behaviours of various temporal scales, most existing approaches either focused on modelling depression from short-term or video-level facial behaviours. In this sense, we propose a two-stage framework that models depression severity from multi-scale short-term and video-level facial behaviours. The short-term depressive behaviour modelling stage first deep learns depression-related facial behavioural features from multiple short temporal scales, where a Depression Feature Enhancement (DFE) module is proposed to enhance the depression-related cues for all temporal scales and remove non-depression related noise. Two novel graph encoding strategies are proposed in the video-level depressive behavior modeling stage, i.e., Sequential Graph Representation (SEG) and Spectral Graph Representation (SPG), to re-encode all short-term features of the target video into a video-level graph representation, summarizing depression-related multi-scale video-level temporal information. As a result, the produced graph representations predict depression severity using both short-term and long-term facial behaviour patterns. The experimental results on AVEC 2013, AVEC 2014 and AVEC 2019 datasets show that the proposed DFE module constantly enhanced the depression severity estimation performance for various CNN models while the SPG is superior than other video-level modelling methods. More importantly, the result achieved for the proposed two-stage framework shows its promising and solid performance compared to widely-used one-stage modelling approaches.
Hatice Gunes, Keerthy Kusumam, Michel F. Valstar, Siyang Song
IEEE Trans. Affect. Comput.2
2025 A Longitudinal Study of Child Wellbeing Assessment via Online Interactions with a Social Robot
abstract
Socially Assistive Robots are studied in different child–robot interaction settings. However, logistical constraints limit accessibility, particularly affecting timely support for mental wellbeing. In this work, we have investigated whether online interactions with a robot can be used for the assessment of mental wellbeing in children. The children (N = 40, 20 girls and 20 boys; 8–13 years) interacted with the Nao robot (30–45 mins) over three sessions, at least a week apart. Audio-visual recordings were collected throughout the sessions that concluded with the children answering user perception questionnaires pertaining to their anxiety toward the robot, and the robot’s abilities. We divided the participants into three wellbeing clusters (low, med, and high tertiles) using their responses to the Short Moods and Feelings Questionnaire (SMFQ) and further analyzed how their wellbeing and their perceptions of the robot changed over the wellbeing tertiles, across sessions and across participants’ gender. Our primary findings suggest that (I) online-mediated interactions with robots can be effective in assessing children’s mental wellbeing over time, and (II) children’s overall perception of the robot either improved or remained consistent across time. Supplementary exploratory analyses have also revealed that the gender of the children affected their wellbeing assessments with interactions effectively distinguishing between varying levels of wellbeing for both boys and girls for the first session and only for boys during the second session. The analyses have also revealed that girls have a higher opinion of the robot as a confidante as compared with boys. Findings from this work affirm the potential of using online-mediated interactions with robots for the assessment of the mental wellbeing of children.
Nida Itrat Abbasi, Guy Laban, Tamsin Ford, Peter B. Jones, Hatice Gunes
ACM Trans. Hum. Robot Interact.5
2025 Participant Perceptions of a Robotic Coach Conducting Positive Psychology Exercises: A Qualitative Analysis
abstract
This article presents a qualitative analysis of participants’ perceptions of a robotic coach conducting Positive Psychology exercises, providing insights for the future design of robotic coaches. Participants \((n=20)\) took part in a single-session (avg. \(31\pm 10\) minutes) Human–Robot Interaction study in a laboratory setting. We created the design of the robotic coach, and its affective adaptation , based on user-centred design research and collaboration with a professional coach. We transcribed post-study participant interviews and conducted a Thematic Analysis. We discuss the results of that analysis, presenting aspects participants found particularly helpful (e.g., the robot asked the correct questions and helped them think of new positive things in their life), and what should be improved (e.g., the robot’s utterance content should be more responsive). We found that participants had no clear preference for affective adaptation or no affective adaptation, which may be due to both positive and negative user perceptions being heightened in the case of adaptation. Based on our qualitative analysis, we highlight insights for the future design of robotic coaches, and areas for future investigation (e.g., examining how participants with different personality traits, or participants experiencing isolation, could benefit from an interaction with a robotic coach).
Minja Axelsson, Nikhil Churamani, Atahan Caldir, Hatice Gunes
ACM Trans. Hum. Robot Interact.4
2025 VITA: A Multi-Modal LLM-Based System for Longitudinal, Autonomous and Adaptive Robotic Mental Well-Being Coaching
abstract
Recently, several works have explored if and how robotic coaches can promote and maintain mental well-being in different settings. However, findings from these studies revealed that these robotic coaches are not ready to be used and deployed in real-world settings due to several limitations that span from technological challenges to coaching success. To overcome these challenges, this article presents VITA, a novel multi-modal LLM-based system that allows robotic coaches to autonomously adapt to the coachee’s multi-modal behaviours (facial valence and speech duration) and deliver coaching exercises in order to promote mental well-being in adults. We identified five objectives that correspond to the challenges in the recent literature, and we show how the VITA system addresses these via experimental validations that include one in-lab pilot study ( N = 4) that enabled us to test different robotic coach configurations (pre-scripted, generic and adaptive models) and inform its design for using it in the real world, and one real-world study ( N = 17) conducted in a workplace over 4 weeks. Our results show that: (i) coachees perceived the VITA adaptive and generic configurations more positively than the pre-scripted one, and they felt understood and heard by the adaptive robotic coach, (ii) the VITA adaptive robotic coach kept learning successfully by personalising to each coachee over time and did not detect any interaction ruptures during the coaching and (iii) coachees had significant mental well-being improvements via the VITA-based robotic coach practice. The code for the VITA system is openly available via https://github.com/Cambridge-AFAR/VITA-system .
Micol Spitale, Minja Axelsson, Hatice Gunes
ACM Trans. Hum. Robot Interact.3
2025 ReactFace: Online Multiple Appropriate Facial Reaction Generation in Dyadic Interactions
abstract
In dyadic interaction, predicting the listener's facial reactions is challenging as different reactions could be appropriate in response to the same speaker's behaviour. Previous approaches predominantly treated this task as an interpolation or fitting problem, emphasizing deterministic outcomes but ignoring the diversity and uncertainty of human facial reactions. Furthermore, these methods often failed to model short-range and long-range dependencies within the interaction context, leading to issues in the synchrony and appropriateness of the generated facial reactions. To address these limitations, this paper reformulates the task as an extrapolation or prediction problem, and proposes an novel framework (called ReactFace) to generate multiple different but appropriate facial reactions from a speaker behaviour rather than merely replicating the corresponding listener facial behaviours. Our ReactFace generates multiple different but appropriate photo-realistic human facial reactions by: (i) learning an appropriate facial reaction distribution representing multiple different but appropriate facial reactions; and (ii) synchronizing the generated facial reactions with the speaker verbal and non-verbal behaviours at each time stamp, resulting in realistic 2D facial reaction sequences. Experimental results demonstrate the effectiveness of our approach in generating multiple diverse, synchronized, and appropriate facial reactions from each speaker's behaviour. The quality of the generated facial reactions is intimately tied to the speaker's speech and facial expressions, achieved through our novel speaker-listener interaction modules.
Siyang Song, Weicheng Xie 0001, Micol Spitale, ZongYuan Ge, LinLin Shen, Hatice Gunes
IEEE Trans. Vis. Comput. Graph.7
2024 Expert Insights on Robots for Safeguarding Children: How (not) and Why (not)?
abstract
To investigate a robot’s role in children’s welfare and safety, we conducted interviews with 8 Subject Matter Experts and Professionals (SMEs) across the disciplines of robotics, child technology, psychology, and psychiatry disciplines. Through qualitative analysis, we synthesize the challenges of safeguarding, compounding limitations, and potential solutions for involving robots in safeguarding as broadly defined in SME interviews. While they agree robots should not be responsible for making judgement calls, the experts also identified the ways robots can be a valuable addition to the safeguarding team. However, more conversations spanning disciplines need to occur to inform policy and legal frameworks that will better establish a robot’s role in intimate spaces in children’s lives. While this line of inquiry is specific to robots in safeguarding, many of the themes reflect the nuances of finding appropriate places for child-robot interactions in the context of children’s welfare.
Leigh Levinson, Nida Itrat Abbasi, Selma Sabanovic, Hatice Gunes
IDC4
2024 Multi-modal Human Behaviour Graph Representation Learning for Automatic Depression Assessment
abstract
Automatic depression assessment (ADA) often relies on crucial cues embedded in human verbal and non-verbal behaviors, which exists in video, audio, and text modalities. Although these modalities often show in time-series forms, current research offers limited exploration of the complex intra-modal temporal dynamics inherent to each modality, failing to extract the depression-related cues in a global view. While many methodologies attempt to exploit the multifaceted information encoded across modalities via decision-level or feature-level fusion techniques, they often fall short in effectively representing pairwise inter-modal relationships, which is the key to utilize the distinct complementary relationship between each modality pair. This paper presents a novel graph-based multimodal fusion approach, which can model intra-modal and inter-modal dynamics conveniently using a graph representation. It adopts undirected edges to link not only temporally continuous, pre-extracted features of each modality, but also temporally aligned features across each pair of modalities. This ensures the seamless propagation of global information across temporal dimensions and helps capture the pairwise inter-modal dynamics. We conduct experiments on the E-DAIC dataset to prove our approach's effectiveness, with an RMSE of 4.80 and a CCC value of 0.563, which rival the top-performing method. We also experiment on the AFAR-BSFP dataset to show the generality of our approach. Our code will be made publicly available.
Haotian Shen, Siyang Song, Hatice Gunes
FG3
2024 REACT 2024: the Second Multiple Appropriate Facial Reaction Generation Challenge
abstract
In dyadic interactions, humans communicate their intentions and state of mind using verbal and non-verbal cues, where multiple different facial reactions might be appropriate in response to a specific speaker behaviour. Then, how to develop a machine learning (ML) model that can automatically generate multiple appropriate, diverse, realistic and synchronised human facial reactions from an previously unseen speaker behaviour is a challenging task. Following the successful organisation of the first REACT challenge (REACT 2023), this edition of the challenge (REACT 2024) employs a subset used by the previous challenge, which contains segmented 30-secs dyadic interaction clips originally recorded as part of the NOXI and RECOLA datasets, encouraging participants to develop and benchmark Machine Learning (ML) models that can generate multiple appropriate facial reactions (including facial image sequences and their attributes) given an input conversational partner's stimulus under various dyadic video conference scenarios. This paper presents: (i) the guidelines of the REACT 2024 challenge; (ii) the dataset utilized in the challenge; and (iii) the performance of the baseline systems on the two proposed sub-challenges: Offline Multiple Appropriate Facial Reaction Generation and Online Multiple Appropriate Facial Reaction Generation, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2024.
Siyang Song, Micol Spitale, Cristina Palmero, Germán Barquero, Hengde Zhu, Sergio Escalera, Michel F. Valstar, Tobias Baur 0001, Fabien Ringeval, Elisabeth André, Hatice Gunes
FG12
2024 LEXI: Large Language Models Experimentation Interface
abstract
The recent developments in Large Language Models (LLMs) mark a significant moment in the research and development of social interactions with artificial agents. These agents are widely deployed in a variety of settings, with potential impact on users. However, the study of social interactions with agents powered by LLMs is still emerging, limited by access to the technology and to data, the absence of standardised interfaces, and challenges to establishing controlled experimental setups using the currently available platforms. To address these gaps, we developed LEXI, LLMs Experimentation Interface, an open-source tool for deploying artificial agents powered by LLMs in social interaction behavioural experiments. Using a graphical interface, LEXI allows researchers to build agents and deploy them in experimental setups along with forms for collecting self-reported data while collecting interaction logs. The outcomes of usability testing indicate LEXI’s broad utility, high usability, and minimal mental workload requirement, with benefits observed across disciplines. A proof-of-concept study exploring the tool’s efficacy in evaluating social human–agent interactions was conducted, resulting in high-quality data. A comparison of empathetic versus neutral agents indicated that people perceive empathetic agents as more social, and write longer and more positive messages towards them.
Guy Laban, Tomer Laban, Hatice Gunes
HAI3
2024 "Oh, Sorry, I Think I Interrupted You": Designing Repair Strategies for Robotic Longitudinal Well-being Coaching
abstract
Robotic well-being coaches have been shown to successfully promote people's mental well-being. To provide successful coaching, a robotic coach should have the capability to repair the mistakes it makes. Past investigations of robot mistakes are limited to game or task-based, one-off and in-lab studies. This paper presents a 4-phase design process to design repair strategies for robotic longitudinal well-being coaching with the involvement of real-world stakeholders: 1) designing repair strategies with a professional well-being coach; 2) a longitudinal study with the involvement of experienced users (i.e., who had already interacted with a robotic coach) to investigate the repair strategies defined in (1); 3) a design workshop with users from the study in (2) to gather their perspectives on the robotic coach's repair strategies; 4) discussing the results obtained in (2) and (3) with the mental well-being professional to reflect on how to design repair strategies for robotic coaching. Our results show that users have different expectations for a robotic coach than a human coach, which influences how repair strategies should be designed. We show that different repair strategies (e.g., apologizing, explaining, or repairing empathically) are appropriate in different scenarios, and that preferences for repair strategies change during longitudinal interactions with the robotic coach.
Minja Axelsson, Micol Spitale, Hatice Gunes
HRI3
2024 MERG: Multi-Dimensional Edge Representation Generation Layer for Graph Neural Networks
abstract
Edges are essential in describing relationships among nodes. While existing graphs frequently use a single-value edge to describe association between each pair of node vectors, crucial relationships may be disregarded if they are not linearly correlated, which may limit graph analysis performance. Although some recent Graph Neural Networks (GNNs) can process graphs containing multi-dimensional edge features, they cannot convert single-value edge graphs to multi-dimensional edge graphs during propagation. This paper proposes a generic Multi-dimensional Edge Representation Generation (MERG) layer that can be inserted into any GNNs for heterogeneous graph analysis. It assigns multi-dimensional edge features for the input single-value edge graph, describing multiple task-specific and global context-aware relationship cues between each connected node pair. Results on eight graph benchmark datasets demonstrate that inserting the MERG layer into widely-used GNNs (e.g., GatedGCN and GAT) leads to major performance improvements, resulting in state-of-the-art (SOTA) results on seven out of eight evaluated datasets. Our code is publicly available at1.
YuXin Song 0001, Aaron S. Jackson, Xi Jia, Weicheng Xie 0001, LinLin Shen, Hatice Gunes, Siyang Song
ICASSP7
2024 ERR@HRI 2024 Challenge: Multimodal Detection of Errors and Failures in Human-Robot Interactions
abstract
Despite the recent advancements in robotics and machine learning (ML), the deployment of autonomous robots in our everyday lives is still an open challenge. This is due to multiple reasons among which are their frequent mistakes, such as interrupting people or having delayed responses, as well as their limited ability to understand human speech, i.e., failure in tasks like transcribing speech to text. These mistakes may disrupt interactions and negatively influence human perception of these robots. To address this problem, robots need to have the ability to detect human-robot interaction (HRI) failures. The ERR@HRI 2024 challenge tackles this by offering a benchmark multimodal dataset of robot failures during human-robot interactions, encouraging researchers to develop and benchmark multimodal machine learning models to detect these failures. We created a dataset featuring multimodal non-verbal interaction data, including facial, speech, and pose features from video clips of interactions with a robotic coach, annotated with labels indicating the presence or absence of robot mistakes, user awkwardness, and interaction ruptures, allowing for the training and evaluation of predictive models. Challenge participants have been invited to submit their multimodal ML models for detection of robot errors, to be evaluated against various performance metrics such as accuracy, precision, recall, F1 score, with and without a margin of error reflecting the time-sensitivity of these metrics. The results of this challenge will help the research field in better understanding the robot failures in human-robot interactions and designing autonomous robots that can mitigate their own errors after successfully detecting them.
Micol Spitale, Maria Teresa Parreira, Maia Stiber, Minja Axelsson, Neval Kara, Garima Kankariya, Chien-Ming Huang 0001, Malte F. Jung, Wendy Ju, Hatice Gunes
ICMI10
2024 FairReFuse: Referee-Guided Fusion for Multi-Modal Causal Fairness in Depression Detection
Jiaee Cheong, Sinan Kalkan, Hatice Gunes
IJCAI3
2024 PerFRDiff: Personalised Weight Editing for Multiple Appropriate Facial Reaction Generation
abstract
Human facial reactions play crucial roles in dyadic human-human interactions, where individuals (i.e., listeners) with varying cognitive process styles may display different but appropriate facial reactions in response to an identical behaviour expressed by their conversational partners. While several existing facial reaction generation approaches are capable of generating multiple appropriate facial reactions (AFRs) in response to each given human behaviour, they fail to take human's personalised cognitive process in AFRs generation. In this paper, we propose the first online personalised multiple appropriate facial reaction generation (MAFRG) approach which learns a unique personalised cognitive style from the target human listener's previous facial behaviours and represents it as a set of network weight shifts. These personalised weight shifts are then applied to edit the weights of a pre-trained generic MAFRG model, allowing the obtained personalised model to naturally mimic the target human listener's cognitive process in its reasoning for multiple AFRs generations. Experimental results show that our approach not only largely outperformed all existing approaches in generating more appropriate and diverse generic AFRs, but also serves as the first reliable personalised MAFRG solution. Our code is made available at https://github.com/xk0720/PerFRDiff.
Hengde Zhu, Xiangyu Kong 0001, Weicheng Xie 0001, LinLin Shen, Lu Liu 0001, Hatice Gunes, Siyang Song
ACM Multimedia7
2024 Robotising Psychometrics: Validating Wellbeing Assessment Tools in Child-Robot Interactions
abstract
The interdisciplinary nature of Child-Robot Interaction (CRI) fosters incorporating measures and methodologies from many established domains. However, when employing CRI approaches to sensitive avenues of health and wellbeing, caution is critical in adapting metrics to retain their safety standards and ensure accurate utilisation. We conducted a secondary analysis to previous empirical work, investigating the reliability and construct validity of established psychological questionnaires such as the Short Moods and Feelings Questionnaire (SMFQ) and three subscales (generalised anxiety, panic and low mood) of the Revised Child Anxiety and Depression Scale (RCADS) within a CRI setting for the assessment of mental wellbeing. Through confirmatory principal component analysis, we have observed that these measures are reliable and valid in the context of CRI. Furthermore, our analysis revealed that scales communicated by a robot demonstrated a better fit than when self-reported, underscoring the efficiency and effectiveness of robot-mediated psychological assessments in these settings. Nevertheless, we have also observed variations in item contributions to the main factor, suggesting potential areas of examination and revision (e.g., relating to physiological changes, inactivity and cognitive demands) when used in CRI. Our findings highlight the importance of verifying the reliability and validity of standardised metrics and assessment tools when employed in CRI settings, thus, aiming to avoid any misinterpretations and misrepresentations.
Nida Itrat Abbasi, Guy Laban, Tamsin Ford, Peter B. Jones, Hatice Gunes
RO-MAN5
2024 Appropriateness of LLM-equipped Robotic Well-being Coach Language in the Workplace: A Qualitative Evaluation
abstract
Robotic coaches have been recently investigated to promote mental well-being in various contexts such as workplaces and homes. With the widespread use of Large Language Models (LLMs), HRI researchers are called to consider language appropriateness when using such generated language for robotic mental well-being coaches in the real world. Therefore, this paper presents the first work that investigated the language appropriateness of robot mental well-being coach in the workplace. To this end, we conducted an empirical study that involved 17 employees who interacted over 4 weeks with a robotic mental well-being coach equipped with LLM-based capabilities. After the study, we individually interviewed them and we conducted a focus group of 1.5 hours with 11 of them. The focus group consisted of: i) an ice-breaking activity, ii) evaluation of robotic coach language appropriateness in various scenarios, and iii) listing shoulds and shouldn’ts for designing appropriate robotic coach language for mental well-being. From our qualitative evaluation, we found that a language-appropriate robotic coach should (1) ask deep questions which explore feelings of the coachees, rather than superficial questions, (2) express and show emotional and empathic understanding of the context, and (3) not make any assumptions without clarifying with follow-up questions to avoid bias and stereotyping. These results can inform the design of language-appropriate robotic coach to promote mental well-being in real-world contexts.
Micol Spitale, Minja Axelsson, Hatice Gunes
RO-MAN3
2024 HRI Wasn't Built In a Day: A Call To Action For Responsible HRI Research
abstract
In recent years, the awareness of the academy around responsible research has notably increased. For instance, with advances in machine learning and artificial intelligence, recent efforts have been made to promote ethical, fair, and inclusive AI and robotics. To better understand if and to what extent HRI is incentivizing researchers to engage in responsible research, we conducted an exploratory review of the publishing guidelines for the most popular HRI conference venues. We identified 18 conferences which published at least 7 HRI papers in 2022. From these, we discuss four themes relevant to conducting responsible HRI research in line with the Responsible Research and Innovation framework: ethical and human participant considerations, transparency and reproducibility, accessibility and inclusion, and plagiarism and LLM use. We identify several gaps and room for improvement within HRI regarding responsible research. Finally, we establish a call to action to provoke conversations among HRI researchers about the importance of conducting responsible research within emerging fields like HRI.
Micol Spitale, Rebecca Stower, Maria Teresa Parreira, Elmira Yadollahi, Iolanda Leite, Hatice Gunes
RO-MAN6
2024 Loss Relaxation Strategy for Noisy Facial Video-based Automatic Depression Recognition
abstract
Automatic depression analysis has been widely investigated on face videos that have been carefully collected and annotated in lab conditions. However, videos collected under real-world conditions may suffer from various types of noise due to challenging data acquisition conditions and lack of annotators. Although deep learning (DL) models frequently show excellent depression analysis performances on datasets collected in controlled lab conditions, such noise may degrade their generalization abilities for real-world depression analysis tasks. In this article, we uncovered that noisy facial data and annotations consistently change the distribution of training losses for facial depression DL models; i.e., noisy data–label pairs cause larger loss values compared to clean data–label pairs. Since different loss functions could be applied depending on the employed model and task, we propose a generic loss function relaxation strategy that can jointly reduce the negative impact of various noisy data and annotation problems occurring in both classification and regression loss functions for face video-based depression analysis, where the parameters of the proposed strategy can be automatically adapted during depression model training. The experimental results on 25 different artificially created noisy depression conditions (i.e., five noise types with five different noise levels) show that our loss relaxation strategy can clearly enhance both classification and regression loss functions, enabling the generation of superior face video-based depression analysis models under almost all noisy conditions. Our approach is robust to its main variable settings and can adaptively and automatically obtain its parameters during training.
Siyang Song, Tugba Tümer, Changzeng Fu, Michel F. Valstar, Hatice Gunes
ACM Trans. Comput. Heal.6
2024 Uncertainty as a Fairness Measure
abstract
Unfair predictions of machine learning (ML) models impede their broad acceptance in real-world settings. Tackling this arduous challenge first necessitates defining what it means for an ML model to be fair. This has been addressed by the ML community with various measures of fairness that depend on the prediction outcomes of the ML models, either at the group-level or the individual-level. These fairness measures are limited in that they utilize point predictions, neglecting their variances, or uncertainties, making them susceptible to noise, missingness and shifts in data. In this paper, we first show that a ML model may appear to be fair with existing point-based fairness measures but biased against a demographic group in terms of prediction uncertainties. Then, we introduce new fairness measures based on different types of uncertainties, namely, aleatoric uncertainty and epistemic uncertainty. We demonstrate on many datasets that (i) our uncertaintybased measures are complementary to existing measures of fairness, and (ii) they provide more insights about the underlying issues leading to bias.
Selim Kuzucu, Jiaee Cheong, Hatice Gunes, Sinan Kalkan
J. Artif. Intell. Res.3
2024 An Open-Source Benchmark of Deep Learning Models for Audio-Visual Apparent and Self-Reported Personality Recognition
abstract
Personality determines a wide variety of human daily and working behaviours, and is crucial for understanding human internal and external states. In recent years, a large number of automatic personality computing approaches have been developed to predict either the apparent personality or self-reported personality of the subject based on non-verbal audio-visual behaviours. However, the majority of them suffer from complex and dataset-specific pre-processing steps and model training tricks. In the absence of a standardized benchmark with consistent experimental settings, it is not only impossible to fairly compare the real performances of these personality computing models but also makes them difficult to be reproduced. In this paper, we present the first reproducible audio-visual benchmarking framework to provide a fair and consistent evaluation of eight existing personality computing models (e.g., audio, visual and audio-visual) and seven standard deep learning models on both self-reported and apparent personality recognition tasks. Building upon a set of benchmarked models, we also investigate the impact of two previously-used long-term modelling strategies for summarising short-term/frame-level predictions on personality computing results. We conduct a comprehensive investigation into all the benchmarked models to demonstrate their capabilities in modelling personality traits on two publicly available datasets, audio-visual apparent personality (ChaLearn First Impression) and self-reported personality (UDIVA) datasets. The experimental results conclude: (i) apparent personality traits, inferred from facial behaviours by most benchmarked deep learning models, show more reliability than self-reported ones; (ii) visual models frequently achieved superior performances than audio models on personality recognition; (iii) non-verbal behaviours contribute differently in predicting different personality traits; and (iv) our reproduced personality computing models generally achieved worse performances than their original reported results. We make all the code and settings of this personality computing benchmark publicly available athttps://github.com/liaorongfan/DeepPersonality.
Rongfan Liao, Siyang Song, Hatice Gunes
IEEE Trans. Affect. Comput.3
2024 Automatic Context-Aware Inference of Engagement in HMI: A Survey
abstract
Engagement is the process by which participants establish, maintain, and end their perceived connection. Automatic engagement inference is one of the tasks required to develop successful human-centered HMI applications. Engagement is a multi-faceted multimodal construct requiring high accuracy in interpretating contextual, verbal and non-verbal cues, making the development of an intelligent automated engagement inference system challenging. Existing surveys concentrate on specific application settings, and a comprehensive survey covering the different engagement facets, definition and inference across various contexts is lacking. Moreover, despite the importance of context-aware modeling, the literature lacks a systematic context-aware overview on the topic. This paper presents a comprehensive survey on previous work in engagement for HMI, entailing interdisciplinary definition, engagement components, publicly available datasets, ground truth assessment, and commonly used features and methods, serving as a guide for the development of future HMI interfaces with reliable context-aware engagement inference capability. An in-depth review across embodied and disembodied interaction modes, and an emphasis on the interaction context of which engagement is studied sets apart this survey from existing ones. Our findings suggest four important directions for future research: (1) context-aware computational modeling, (2) temporal dynamics, (3) personalised computing, and (4) bias and fairness of engagement inference systems.
Hanan Salam, Oya Çeliktutan, Hatice Gunes, Mohamed Chetouani
IEEE Trans. Affect. Comput.3
2024 Robots as Mental Well-being Coaches: Design and Ethical Recommendations
abstract
The last decade has shown a growing interest in robots as well-being coaches. However, insightful guidelines for the design of robots as coaches to promote mental well-being have not yet been proposed. This article details design and ethical recommendations based on a qualitative analysis drawing on a grounded theory approach, which was conducted with a three-step iterative design process which included user-centered design studies involving robotic well-being coaches, namely: (1) a user-centred design study conducted with 11 participants consisting of both prospective users who had participated in a Brief Solution-Focused Practice study with a human coach, as well as coaches of different disciplines, (2) semi-structured individual interview data gathered from 20 participants attending a Positive Psychology intervention study with the robotic well-being coach Pepper, and (3) a user-centred design study conducted with 3 participants of the Positive Psychology study as well as 2 relevant well-being coaches. After conducting a thematic analysis and a qualitative analysis, we collated the data gathered into convergent and divergent themes, and we distilled from those results a set of design guidelines and ethical considerations. Our findings can inform researchers and roboticists on the key aspects to take into account when designing robotic mental well-being coaches.
Minja Axelsson, Micol Spitale, Hatice Gunes
ACM Trans. Hum. Robot Interact.3
2023 "It's not Fair!" - Fairness for a Small Dataset of Multi-modal Dyadic Mental Well-being Coaching
abstract
In recent years, the affective computing research community has put ethics at the centre of its research agenda. However, many of the currently available datasets for affective computing are ‘small’, making bias and debias analysis challenging. This paper presents the first work to explore bias analysis and mitigation of a small temporal multi-modal dataset for mental well-being by adopting different data augmentation techniques. This proof-of-concept work’s contributions include: i) introducing a novel small temporal multi-modal dataset of dyadic interactions during mental well-being coaching; ii) providing multi-modal and feature importance analyses evaluated via modelling performance and fairness metrics across both high and low-level features; and iii) proposing a simple and effective data augmentation strategy (MixFeat) to debias the small dataset presented in this paper. We conduct extensive experiments and analyses to compare our proposed method against other baseline data augmentation method across various uni-modal and multi-modal setups. Our results indicate that, regardless of the dimensionality of the dataset at hand, the inclusion of a bias analysis section in the conference papers is viable. This paper is therefore a call to the community to include a bias analysis section in ACII conference submissions, similar to the ablation studies conducted in papers submitted to major machine learning conferences.
Jiaee Cheong, Micol Spitale, Hatice Gunes
ACII3
2023 Latent Generative Replay for Resource-Efficient Continual Learning of Facial Expressions
abstract
Real-world Facial Expression Recognition (FER) systems require models to constantly learn and adapt with novel data. Traditional Machine Learning (ML) approaches struggle to adapt to such dynamics as models need to be re-trained from scratch with a combination of both old and new data. Replay-based Continual Learning (CL) provides a solution to this problem, either by storing previously seen data samples in memory, sampling and interleaving them with novel data (rehearsal) or by using a generative model to simulate pseudo-samples to replay past knowledge (pseudo-rehearsal). Yet, the high memory footprint of rehearsal and the high computational cost of pseudo-rehearsal limit the real-world application of such methods, especially on resource-constrained devices. To address this, we propose Latent Generative Replay (LGR) for pseudo-rehearsal of low-dimensional latent features to mitigate forgetting in a resource-efficient manner. We adapt popular CL strategies to use LGR instead of generating pseudo-samples, resulting in performance upgrades when evaluated on the CK+, RAF-DB and AffectNet FER benchmarks where LGR significantly reduces the memory and resource consumption of replay-based CL without compromising model performance.
Samuil Stoychev, Nikhil Churamani, Hatice Gunes
FG3
2023 Robotic Mental Well-being Coaches for the Workplace: An In-the-Wild Study on Form
abstract
The World Health Organization recommends that employers take action to protect and promote mental well-being at work. However, the extent to which these recommended practices can be implemented in the workplace is limited by the lack of resources and personnel availability. Robots have been shown to have great potential for promoting mental well-being, and the gradual adoption of such assistive technology may allow employers to overcome the aforementioned resource barriers. This paper presents the first study that investigates the deployment and use of two different forms of robotic well-being coaches in the workplace in collaboration with a tech company whose employees (26 coachees) interacted with either a QTrobot (QT ) or a Misty robot (M). We endowed the robots with a coaching personality to deliver positive psychology exercises over four weeks (one exercise per week). Our results show that the robot form significantly impacts coachees' perceptions of the robotic coach in the workplace. Coachees perceived the robotic coach in M more positively than in QT (both in terms of behaviour appropriateness and perceived personality), and they felt more connection with the robotic coach in M. Our study provides valuable insights for robotic well-being coach design and deployment, and contributes to the vision of taking robotic coaches into the real world.
Micol Spitale, Minja Axelsson, Hatice Gunes
HRI3
2023 Towards Gender Fairness for Mental Health Prediction
abstract
Mental health is becoming an increasingly prominent health challenge. Despite a plethora of studies analysing and mitigating bias for a variety of tasks such as face recognition and credit scoring, research on machine learning (ML) fairness for mental health has been sparse to date. In this work, we focus on gender bias in mental health and make the following contributions. First, we examine whether bias exists in existing mental health datasets and algorithms. Our experiments were conducted using Depresjon, Psykose and D-Vlog. We identify that both data and algorithmic bias exist. Second, we analyse strategies that can be deployed at the pre-processing, in-processing and post-processing stages to mitigate for bias and evaluate their effectiveness. Third, we investigate factors that impact the efficacy of existing bias mitigation strategies and outline recommendations to achieve greater gender fairness for mental health. Upon obtaining counter-intuitive results on D-Vlog dataset, we undertake further experiments and analyses, and provide practical suggestions to avoid hampering bias mitigation efforts in ML for mental health.
Jiaee Cheong, Selim Kuzucu, Sinan Kalkan, Hatice Gunes
IJCAI4
2023 REACT2023: The First Multiple Appropriate Facial Reaction Generation Challenge
abstract
The Multiple Appropriate Facial Reaction Generation Challenge (REACT2023) is the first competition event focused on evaluating multimedia processing and machine learning techniques for generating human-appropriate facial reactions in various dyadic interaction scenarios, with all participants competing strictly under the same conditions. The goal of the challenge is to provide the first benchmark test set for multi-modal information processing and to foster collaboration among the audio, visual, and audio-visual behaviour analysis and behaviour generation (a.k.a generative AI) communities, to compare the relative merits of the approaches to automatic appropriate facial reaction generation under different spontaneous dyadic interaction conditions. This paper presents: (i) the novelties, contributions and guidelines of the REACT2023 challenge; (ii) the dataset utilized in the challenge; and (iii) the performance of the baseline systems on the two proposed sub-challenges: Offline Multiple Appropriate Facial Reaction Generation and Online Multiple Appropriate Facial Reaction Generation, respectively. The challenge baseline code is publicly available at https://github.com/reactmultimodalchallenge/baseline_react2023.
Siyang Song, Micol Spitale, Germán Barquero, Cristina Palmero, Sergio Escalera, Michel F. Valstar, Tobias Baur 0001, Fabien Ringeval, Elisabeth André, Hatice Gunes
ACM Multimedia11
2023 Humanoid Robots for Wellbeing Assessment in Children: How Does Anxiety towards the Robot Affect Perceptions of Robot Role, Behaviour and Capabilities?
abstract
With the introduction of socially assistive robots in many avenues of children’s lives, it is becoming increasingly vital to understand how children’s perceptions of the robot affect their evaluation and interaction. The main objective of this work is to investigate how children’s anxiety towards robots has influenced their perceptions of their interaction with a Nao robot. We collected data from 37 children (8 - 13 years old) who interacted, for about 30-45 minutes, with the robot which delivered initial pleasantries and four different tasks to help assess their mental wellbeing in a lab setting. We collected audio-visual recordings of the interaction. At the end of the session, we asked children to answer three self-report questionnaires to evaluate: the robot’s role as a confidante, the anxiety towards the robot, and the children’s perception of the robot’s behaviour and capabilities. Based on their responses to the robot’s anxiety questionnaire, children were divided into two categories: “low anxiety” (anxiety scoremedian anxiety score). Our results show that i) most children (89.2%) irrespective of their wellbeing, experience some degree of anxiety towards the robot, ii) children’s anxiety has influenced their willingness to participate in the initial pleasantries conducted by the robot, and iii) children’s anxiety has also affected their evaluations of the robot as a confidante and their perceptions of the robot’s behaviour and capabilities. Findings from this work have significant implications for designing effective and successful robot-led initiatives for assessing mental wellbeing in children, by taking into account their mindsets and dispositions.
Nida Itrat Abbasi, Micol Spitale, Joanna Anderson, Tamsin Ford, Peter B. Jones, Hatice Gunes
RO-MAN6
2023 Federated Continual Learning for Socially Aware Robotics
abstract
From learning assistance to companionship, socially aware and socially assistive robotics is promising for enhancing many aspects of daily life. However, socially aware robots face many challenges preventing their widespread public adoption. Two such major challenges are (1) lack in behavior adaptation to new environments, contexts and users, and (2) insufficient capability for privacy protection. The commonly employed centralized learning paradigm, whereby training data is gathered and centralized in a single location (i.e., machine/ server) and the centralized entity trains and hosts the model, contributes to these limitations by preventing online learning of new experiences and requiring storage of privacy-sensitive data. In this work, we propose a decentralized learning paradigm that aims to improve the personalization capability of social robots while also paving the way towards privacy preservation. First, we present a new framework by capitalising on two machine learning approaches, Federated Learning and Continual Learning, to capture interaction dynamics distributed physically across robots and temporally across repeated robot encounters. Second, we introduce four criteria (adaptation quality, adaptation time, knowledge sharing, and model overhead) that should be balanced within our decentralized robot learning framework. Third, we develop a new algorithm – Elastic Transfer – that leverages importance-based regularization to preserve relevant parameters across robots and interactions with multiple humans (users). We show that decentralized learning is a viable alternative to centralized learning in a proof-of-concept Socially-Aware Navigation domain, and demonstrate the efficacy of Elastic Transfer across our proposed evaluation criteria.
Luke Guerdan, Hatice Gunes
RO-MAN2
2023 Affective Computing for Human-Robot Interaction Research: Four Critical Lessons for the Hitchhiker
abstract
Social Robotics and Human-Robot Interaction (HRI) research relies on different Affective Computing (AC) solutions for sensing, perceiving and understanding human affective behaviour during interactions. This may include utilising off-the-shelf affect perception models that are pre-trained on popular affect recognition benchmarks and directly applied to situated interactions. However, the conditions in situated human-robot interactions differ significantly from the training data and settings of these models. Thus, there is a need to deepen our understanding of how AC solutions can be best leveraged, customised and applied for situated HRI. This paper, while critiquing the existing practices, presents four critical lessons to be noted by the hitchhiker when applying AC for HRI research. These lessons conclude that: (i) The six basic emotions categories are not always relevant in situated interactions, (ii) Affect recognition accuracy (%) improvement as the sole goal is inappropriate for situated interactions, (iii) Affect recognition may not generalise across contexts, and (iv) Affect recognition alone is insufficient for adaptation and personalisation. By describing the background and the context for each lesson, and demonstrating how these lessons have been compiled from the various studies of the authors, this paper aims to enable the hitchhiker to successfully leverage AC solutions for advancing HRI research.
Hatice Gunes, Nikhil Churamani
RO-MAN1
2023 Longitudinal Evolution of Coachees' Behavioural Responses to Interaction Ruptures in Robotic Positive Psychology Coaching
abstract
Robotic mental well-being coaches could be used to help people maintain their well-being, and improve access to mental healthcare. In coaching, the alliance between the coach and coachee is important for the success of the practice. However, this alliance might be negatively affected by interaction ruptures (e.g., the robot making mistakes and the user feeling awkward) that still commonly occur in human-robot interactions. Therefore, robotic coaches should be able to recognize ruptures occurring during their interactions with human users to guarantee the success of the well-being practice. To this aim, we analyse coachee behavioural responses to interaction ruptures during a robotic positive psychology coaching practice and how these behavioural cues evolve over time. We focus our analysis on a dataset we collected in a previous work, where 26 participants interacted with either a QTrobot or a Misty II robot at their workplace over 4 weeks. We undertake a longitudinal analysis of coachees’ multimodal nonverbal cues (i.e., facial expressions, vocal acoustic features, and body pose features) to investigate the contribution of individual modalities for detecting interaction ruptures. Our results show that coachees: i) displayed facial cues of rupture (e.g, laughing at the robot) and suspicion more in the first week than in the last week; ii) talked more and were less silent in the last week than in the previous weeks; and iii) exhibited a higher number of hand-over-face gestures (a cue for self-disclosure) in the last week than in the previous weeks. Our findings aim to inform the development of AI models for multi-modal detection of interaction ruptures which can be used to improve the effectiveness and the success of robotic well-being coaching.
Micol Spitale, Minja Axelsson, Neval Kara, Hatice Gunes
RO-MAN4
2023 Domain-Incremental Continual Learning for Mitigating Bias in Facial Expression and Action Unit Recognition
abstract
As Facial Expression Recognition (FER) systems become integrated into our daily lives, these systems need to prioritise makingfairdecisions instead of only aiming at higher individual accuracy scores. From surveillance systems, to monitoring the mental and emotional health of individuals, these systems need to balance theaccuracy versus fairnesstrade-off to make decisions that do not unjustly discriminate against specific under-represented demographic groups. Identifyingbiasas a critical problem in facial analysis systems, different methods have been proposed that aim to mitigate bias both at data and algorithmic levels. In this work, we propose the novel use of Continual Learning (CL), in particular, using Domain-Incremental Learning (Domain-IL) settings, as a potent bias mitigation method to enhance thefairnessof Facial Expression Recognition (FER) systems. We compare different non-CL-based and CL-based methods for theirperformanceandfairness scoreson expression recognition and Action Unit (AU) detection tasks using two popular benchmarks, the RAF-DB and BP4D datasets, respectively. Our experimental results show that CL-based methods, on average, outperform other popular bias mitigation techniques on bothaccuracyandfairnessmetrics.
Nikhil Churamani, Özgür Kara, Hatice Gunes
IEEE Trans. Affect. Comput.3
2023 Learning Person-Specific Cognition From Facial Reactions for Automatic Personality Recognition
abstract
This article proposes to recognise the true (self-reported) personality traits from the target subject's cognition simulated from facial reactions. This approach builds on the following two findings in cognitive science: (i) human cognition partially determines expressed behaviour and is directly linked to true personality traits; and (ii) in dyadic interactions, individuals’ nonverbal behaviours are influenced by their conversational partner's behaviours. In this context, we hypothesise that during a dyadic interaction, a target subject's facial reactions are driven by two main factors: their internal (person-specific) cognitive process, and the externalised nonverbal behaviours of their conversational partner. Consequently, we propose to represent the target subject's (defined as the listener) person-specific cognition in the form of a person-specific CNN architecture that has unique architectural parameters and depth, which takes audio-visual non-verbal cues displayed by the conversational partner (defined as the speaker) as input, and is able to reproduce the target subject's facial reactions. Each person-specific CNN is explored by the Neural Architecture Search (NAS) and a novel adaptive loss function, which is then represented as a graph representation for recognising the target subject's true personality. Experimental results not only show that the produced graph representations are well associated with target subjects’ personality traits in both human-human and human-machine interaction scenarios, and outperform the existing approaches with significant advantages, but also demonstrate that the proposed novel strategies help in learning more reliable personality representations.
Siyang Song, Zilong Shao, Shashank Jaiswal, LinLin Shen, Michel F. Valstar, Hatice Gunes
IEEE Trans. Affect. Comput.6
2022 Lifelong Learning and Personalization in Long-Term Human-Robot Interaction (LEAP-HRI)
abstract
While most research in Human-Robot Interaction (HRI) studies one-off or short-term interactions in constrained laboratory settings, a growing body of research focuses on breaking through these boundaries and studying long-term interactions that arise through deployments of robots “in the wild”. Under these conditions, robots need to incrementally learn new concepts or abilities (i.e., “lifelong learning”) to adapt their behaviors within new situations and personalize their interactions with users to maintain their interest and engagement. The second edition of the “Lifelong Learning and Personalization in Long-Term Human-Robot Interaction (LEAP-HRI)” workshop aims to address the developments and challenges in these areas and create a medium for researchers to share their work in progress, present preliminary results, learn from the experience of invited researchers and discuss relevant topics. The workshop focuses on studies on lifelong learning and adaptivity to users, context, environment, and tasks in long-term interactions in a variety of fields such as education, rehabilitation, elderly care, collaborative tasks, service, and companion robots.
Bahar Irfan, Aditi Ramachandran, Samuel Spaulding, German Ignacio Parisi, Hatice Gunes
HRI5
2022 Learning Socially Appropriate Robo-waiter Behaviours through Real-time User Feedback
abstract
Current Humanoid Service Robot (HSR) behaviours mainly rely on static models that cannot adapt dynamically to meet individual customer attitudes and preferences. In this work, we focus on empowering HSRs with adaptive feedback mechanisms driven by either implicit reward, by estimating facial affect, or explicit reward, by incorporating verbal responses of the human ‘customer’. To achieve this, we first create a custom dataset, annotated using crowd-sourced labels, to learn appropri-ate approach (positioning and movement) behaviours for a Robo-waiter. This dataset is used to pre-train a Reinforcement Learning (RL) agent to learn behaviours deemed socially appropriate for the robo-waiter. This model is later extended to include separate implicit and explicit reward mechanisms to allow for interactive learning and adaptation from user social feedback. We present a within-subjects Human-Robot Interaction (HRI) study with 21 participants implementing interactions between the robo-waiter and human customers implementing the above-mentioned model variations. Our results show that both explicit and implicit adaptation mechanisms enabled the adaptive robo-waiter to be rated as more enjoyable and sociable, and its positioning relative to the participants as more appropriate compared to using the pre-trained model or a randomised control implementation.
Emily McQuillin, Nikhil Churamani, Hatice Gunes
HRI3
2022 Statistical, Spectral and Graph Representations for Video-Based Facial Expression Recognition in Children
abstract
Child facial expression recognition is a relatively less investigated area within affective computing. Children’s facial expressions differ significantly from adults; thus, it is necessary to develop emotion recognition frameworks that are more objective, descriptive and specific to this target user group. In this paper we propose the first approach that (i) constructs video-level heterogeneous graph representation for facial expression recognition in children, and (ii) predicts children’s facial expressions using the automatically detected Action Units (AUs). To this aim, we construct three separate length-independent representations, namely, statistical, spectral and graph at video-level for detailed multi-level facial behaviour decoding (AU activation status, AU temporal dynamics and spatio-temporal AU activation patterns, respectively). Our experimental results on the LIRIS Children Spontaneous Facial Expression Video Database demonstrate that combining these three feature representations provides the highest accuracy for expression recognition in children.
Nida Itrat Abbasi, Siyang Song, Hatice Gunes
ICASSP3
2022 Dataset Bias in Deception Detection
abstract
With the advances in Machine Learning, lie detection technology gained significant attention. In recent years, several multi-modal techniques achieved as high as 99% accuracy results using the Real-life Trial dataset with only 121 data points. This led to considerable media hype and research interest in lie detection with machine learning. In this paper, we analyze the effect of dataset bias in deception detection. More specifically, we train a classifier to predict the sex of the identity appearing in the video. On a test data point, we use the sex predictor to predict sex which we use as a proxy for predicting deception, predicting lie for females and truth for males. This lie predictor simulates a classifier that uses nothing but dataset bias. Nevertheless, we find that the performance of this biased classifier is comparable to those of state-of-the-art papers. More specifically, when using IDT features, our biased classifier achieves 64.6% and 59.3% AUC while a classifier trained normally on truth/lie labels achieves 57.4% accuracy and 69.3% AUC. We perform similar experiments on the Bag-of-Lies dataset and show that it too is biased with respect to sex. In addition, we apply the state-ofthe-art techniques on an unbiased dataset and show that their performance is no better than chance. Our experiments strongly suggest that the results of recent deception detection techniques can be explained by the bias inherent in the datasets.
Ara Mambreyan, Elena Punskaya, Hatice Gunes
ICPR3
2022 Learning Multi-dimensional Edge Feature-based AU Relation Graph for Facial Action Unit Recognition
abstract
The activations of Facial Action Units (AUs) mutually influence one another. While the relationship between a pair of AUs can be complex and unique, existing approaches fail to specifically and explicitly represent such cues for each pair of AUs in each facial display. This paper proposes an AU relationship modelling approach that deep learns a unique graph to explicitly describe the relationship between each pair of AUs of the target facial display. Our approach first encodes each AU's activation status and its association with other AUs into a node feature. Then, it learns a pair of multi-dimensional edge features to describe multiple task-specific relationship cues between each pair of AUs. During both node and edge feature learning, our approach also considers the influence of the unique facial display on AUs' relationship by taking the full face representation as an input. Experimental results on BP4D and DISFA datasets show that both node and edge feature learning modules provide large performance improvements for CNN and transformer-based backbones, with our best systems achieving the state-of-the-art AU recognition results. Our approach not only has a strong capability in modelling relationship cues for AU recognition but also can be easily incorporated into various backbones. Our PyTorch code is made available at https://github.com/CVI-SZU/ME-GraphAU.
Siyang Song, Weicheng Xie 0001, LinLin Shen, Hatice Gunes
IJCAI5
2022 Can Robots Help in the Evaluation of Mental Wellbeing in Children? An Empirical Study
abstract
Socially Assistive Robots (SARs) show promise in helping children during therapeutic and clinical interventions. However, using SARs for the evaluation of mental wellbeing of children has not yet been explored. Thus, this paper presents an empirical study with 28 children 8-13 years old interacting with a Nao robot in a 45-minute session where the robot administered (robotised) the Short Mood and Feelings Questionnaire (SMFQ) and the Revised Child Anxiety and Depression Scale (RCADS). Prior to the experimental session, we also evaluated children’s wellbeing using established standardised approaches via online RCADS questionnaires filled by the children (self-report) and their parents (parent-report). We clustered the participants into three groups (lower, medium, and higher tertile) based on their SMFQ scores. Further, we analysed the questionnaire responses across the three clusters and across the different modes of administration (self-report, parent-report, and robotised). Our results show that the robotised evaluation seems to be the most suitable mode in identifying wellbeing related anomalies in children across the three clusters of participants as compared with the self-report and the parent-report modes. Further, children with decreasing levels of wellbeing (lower, medium and higher tertiles) exhibit different response patterns: children of higher tertile are more negative in their responses to the robot while the ones of lower tertile are more positive in their responses to the robot. Findings from this work show that SARs can be a promising tool to potentially evaluate mental wellbeing related concerns in children.
Nida Itrat Abbasi, Micol Spitale, Joanna Anderson, Tamsin Ford, Peter B. Jones, Hatice Gunes
RO-MAN6
2021 Multi-dimensional Affect in Poetry (POCA) Dataset: Acquisition, Annotation and Baseline Results
abstract
Detecting emotions and affect in text has received enormous attention in recent years, and yet majority of the works in this area reduce the nuanced emotional responses into ‘positive’, ‘negative’ and ‘neutral’. In this paper, we introduce a novel multi-dimensional affect in poetry (POCA) dataset for sentiment analysis annotated using the Geneva Emotion Wheel (GEW), to capture and analyse the multi-dimensional affect evoked in listeners. The POCA dataset is based on poems and their corresponding recitals from an online poetry database where recitals are curated by the website, and performed by the poet or an approved artist. The POCA dataset contains 330 poems (text and audio), from the English language, each of which is annotated across 20 different emotion classes, by 5 listeners. A subset of the dataset (50 poems) have also been annotated by an inlab study by 3 listeners each while their Electrodermal activity (EDA) was being recorded. As a proof of concept, we (i) introduce representative problem formulations to be addressed by machine learning approaches using the POCA dataset, from single emotion recognition (e.g., does this poem evoke joy?) to continuous affect prediction (e.g., what level arousal and valence does this poem evoke?), (ii) provide baseline results for text-based affect recognition using several classification and regression models, and (iii) provide baseline results for EDA-based affect prediction. Our results show that (i) for text-based affect recognition, classical approaches can provide as accurate results as their fine-tuned neural network counterparts, and (ii) in EDA-based affect prediction, in general there is a strong relation between the EDA signals and the self-reported valence and arousal quadrants, while predictions are better for arousal than valence.
Akbir Khan, Jack Hopkins, Hatice Gunes
ACII3
2021 AULA-Caps: Lifecycle-Aware Capsule Networks for Spatio-Temporal Analysis of Facial Actions
abstract
Most state-of-the-art approaches for Facial Action Unit (AU) detection rely on evaluating static frames, encoding a snapshot of heightened facial activity. In real-world interactions, however, facial expressions are more subtle and evolve over time requiring AU detection models to learn spatial as well as temporal information. In this work, we focus on both spatial and spatio-temporal features encoding the temporal evolution of facial AU activation. We propose the Action Unit Lifecycle-Aware Capsule Network (AULA-Caps) for AU detection using both frame and sequence-level features. While, at the frame-level, the capsule layers of AULA-Caps learn spatial feature primitives to determine AU activations, at the sequence-level, it learns temporal dependencies between contiguous frames by focusing on relevant spatio-temporal segments in the sequence. The learnt feature capsules are routed together such that the model learns to selectively focus on spatial or spatio-temporal information depending upon the AU lifecycle. The proposed model is evaluated on popular benchmarks, namely BP4D and GFT datasets, obtaining state-of-the-art results for both.
Nikhil Churamani, Sinan Kalkan, Hatice Gunes
FG3
2021 Social Signals and Multimedia: Past, Present, Future
abstract
The rising popularity of Artificial Intelligence (AI) has brought considerable public interest as well faster and more direct transfer of research ideas into practice. One of the aspects of AI that still trails behind considerably is the role of machines in interpreting, enhancing, modeling, generating, and influencing social behavior. Such behavior is captured as social signals, usually by sensors recording multiple modalities, making it classic multimedia data. Such behavior can also be generated by an AI system when interacting with humans. Using AI techniques in combination with multimedia data can be used to pursue multiple goals, two of which are high-lighted here. First, supporting people during social interactions and helping them to fulfil their social needs either actively or passively.Second, improving our understanding of how people collaborate, build relationships, and process self identity. Despite the rise of fields such as Social Signal Processing, a similar panel organised at ACM Multimedia 2014, and an area on social and emotional signal sat the ACM MM since 2014, we argue that we have yet to truly fulfil the potential of the combining social signals and multimedia. This panel asks where we have come far enough and what remaining challenges there are in light of recent global events.
Hayley Hung, Cathal Gurrin, Martha A. Larson, Hatice Gunes, Fabien Ringeval, Elisabeth André, Louis-Philippe Morency
ACM Multimedia4
2021 Personality Recognition by Modelling Person-specific Cognitive Processes using Graph Representation
abstract
Recent research shows that in dyadic and group interactions individuals' nonverbal behaviours are influenced by the behaviours of their conversational partner(s). Therefore, in this work we hypothesise that during a dyadic interaction, the target subject's facial reactions are driven by two main factors: (i) their internal (person-specific) cognition, and (ii) the externalised nonverbal behaviours of their conversational partner. Subsequently, our novel proposition is to simulate and represent the target subject's (i.e., the listener) cognitive process in the form of a person-specific CNN architecture whose input is the audio-visual non-verbal cues displayed by the conversational partner (i.e., the speaker), and the output is the target subject's (i.e., the listener) facial reactions. We then undertake a search for the optimal CNN architecture whose results are used to create a person-specific graph representation for recognising the target subject's personality. The graph representation, fortified with a novel end-to-end edge feature learning strategy, helps with retaining both the unique parameters of the person-specific CNN and the geometrical relationship between its layers. Consequently, the proposed approach is the first work that aims to recognize the true (self-reported) personality of a target subject (i.e., the listener) from the learned simulation of their cognitive process (i.e., parameters of the person-specific CNN). The experimental results show that the CNN architectures are well associated with target subjects' personality traits and the proposed approach clearly outperforms multiple existing approaches that predict personality directly from non-verbal behaviours. In light of these findings, this work opens up a new avenue of research for predicting and recognizing socio-emotional phenomena (personality, affect, engagement etc.) from simulations of person-specific cognitive processes.
Zilong Shao, Siyang Song, Shashank Jaiswal, LinLin Shen, Michel F. Valstar, Hatice Gunes
ACM Multimedia6
2021 Participatory Design of a Robotic Mental Well-being Coach
abstract
Recent research is emerging in the field of Social Robotics where robots have the potential to serve as tools to improve human well-being. However, research exploring the expectations and perceptions of prospective users of such robots, and the professionals who currently deliver these interventions, is limited. In this paper, we present qualitative analysis of discussions with prospective users and experienced coaches regarding the design of robot well-being coaches. We invited participants interested in well-being practices to take-part in a Participatory Design (PD) study, consisting of individual interviews and a focus group discussion (NP= 8). Discussions focused on ideating how a robot could function as a mental well-being coach, based on their experiences with well-being practices. Data triangulation was employed by interviewing three professional coaches as additional sources of information. This resulted in a rich set of data, which we transcribed and analysed using Thematic Analysis (TA). The developed themes regarding robot features, form, behaviours, robot-led well-being practices, and the advantages and disadvantages these could provide, were compiled and are discussed in detail. We present this data together with tabulated quotes from the participants and coaches, to pave the way towards designing robot coaches that can provide supportive interventions to improve the mental health and well-being of their users.
Minja Axelsson, Indu P. Bodala, Hatice Gunes
RO-MAN3
2021 Teleoperated Robot Coaching for Mindfulness Training: A Longitudinal Study
abstract
Social robots are becoming incorporated in daily human lives, assisting in the promotion of the physical and mental wellbeing of individuals. To investigate the design and use of social robots for delivering mindfulness training, we develop a teleoperation framework that enables an experienced Human Coach (HC) to conduct mindfulness training sessions virtually, by replicating their upper-body and head movements onto the Pepper robot, in real-time. Pepper’s vision is mapped onto a Head-Mounted Display (HMD) worn by the HC and a bidirectional audio pipeline is set up, enabling the HC to communicate with the participants through the robot. To evaluate the participants’ perceptions of the teleoperated Robot Coach (RC), we study the interactions between a group of participants and the RC over 5 weeks and compare these with another group of participants interacting directly with the HC. Growth modelling analysis of this longitudinal data shows that the HC ratings are consistently greater than 4 (on a scale of 1 5) for all aspects while an increase is witnessed in the RC ratings over the weeks, for the Robot Motion and Conversation dimensions. Mindfulness training delivered by both types of coaching evokes positive responses from the participants across all the sessions, with the HC rated significantly higher than the RC on Animacy, Likeability and Perceived Intelligence. Participants’ personality traits such as Conscientiousness and Neuroticism are found to influence their perception of the RC. These findings enable an understanding of the differences between the perceptions of HC and RC delivering mindfulness training, and provide insights towards the development of robot coaches for improving the psychological wellbeing of individuals.
Indu P. Bodala, Nikhil Churamani, Hatice Gunes
RO-MAN3
2021 Special Issue on Automated Perception of Human Affect from Longitudinal Behavioral Data
abstract
The papers in this special section are aimed at contributions from computational neuroscience and psychology, artificial intelligence, machine learning, and affective computing, challenging and expanding current research on interpretation and estimation of human affective behavior from longitudinal data, i.e., single or multiple modalities captured over extended periods of time allowing efficient representation of behavior and inference in terms of affect and other socio-cognitive dimensions.
Pablo V. A. Barros, Stefan Wermter, Ognjen Rudovic, Hatice Gunes
IEEE Trans. Affect. Comput.4
2021 Audio-Driven Robot Upper-Body Motion Synthesis
abstract
Body language is an important aspect of human communication, which an effective human-robot interaction interface should mimic well. Human beings exchange information and convey their thoughts and feelings through gaze, facial expressions, body language, and tone of voice along with spoken words, and infer 65% of the meaning of the communicated messages from these nonverbal cues. Modern robotic platforms are, however, limited in their ability to automatically generate behaviors that align with their speech. In this article, we develop a neural-network-based system that takes audio from a user as an input and generates upper-body gestures, including head, hand, and torso movements of the user on a humanoid robot, namely, Softbank Robotics' Pepper. Our system was evaluated quantitatively as well as qualitatively using Web surveys when driven by natural speech and synthetic speech. We compare the impact of generic and person-specific neural-network models on the quality of synthesized movements. We further investigate the relationships between quantitative and qualitative evaluations and examine how the speaker's personality traits affect the synthesized movements.
Jan Ondras, Oya Çeliktutan, Paul Bremner, Hatice Gunes
IEEE Trans. Cybern.4
2020 CLIFER: Continual Learning with Imagination for Facial Expression Recognition
abstract
Current Facial Expression Recognition (FER) approaches tend to be insensitive to individual differences in expression and interaction contexts. They are unable to adapt to the dynamics of real-world environments where data is only available incrementally, acquired by the system during interactions. In this paper, we propose a novel continual learning framework with imagination for FER (CLIFER) that (i) implements imagination to simulate expression data for particular subjects and integrates it with (ii) a complementary learning-based dual-memory (episodic and semantic) model, to augment person-specific learning. The framework is evaluated on its ability to remember previously seen classes as well as on generalising to yet unseen classes, resulting in high F1-scores for multiple FER datasets: RAVDESS (episodic: F1=0.98 ± 0.01, semantic: F1=0.75 ± 0.01), MMI (episodic: F1=0.75 ± 0.07, semantic: F1=0.46 ± 0.04) and BAUM-I (episodic: F1=0.87 ± 0.05, semantic: F1=0.51 ± 0.04).
Nikhil Churamani, Hatice Gunes
FG2
2020 Facial Electromyography-based Adaptive Virtual Reality Gaming for Cognitive Training
abstract
Cognitive training has shown promising results for delivering improvements in human cognition related to attention, problem solving, reading comprehension and information retrieval. However, two frequently cited problems in cognitive training literature are a lack of user engagement with the training programme, and a failure of developed skills to generalise to daily life. This paper introduces a new cognitive training (CT) paradigm designed to address these two limitations by combining the benefits of gamification, virtual reality (VR), and affective adaptation in the development of an engaging, ecologically valid, CT task. Additionally, it incorporates facial electromyography (EMG) as a means of determining user affect while engaged in the CT task. This information is then utilised to dynamically adjust the game's difficulty in real-time as users play, with the aim of leading them into a state of flow. Affect recognition rates of 64.1% and 76.2%, for valence and arousal respectively, were achieved by classifying a DWT-Haar approximation of the input signal using kNN. The affect-aware VR cognitive training intervention was then evaluated with a control group of older adults. The results obtained substantiate the notion that adaptation techniques can lead to greater feelings of competence and a more appropriate challenge of the user's skills.
Lorcan Reidy, Dennis Chan, Charles Nduka, Hatice Gunes
ICMI4
2020 Continual Learning for Affective Robotics: Why, What and How?
abstract
Creating and sustaining closed-loop dynamic and social interactions with humans require robots to continually adapt towards their users' behaviours, their affective states and moods while keeping them engaged in the task they are performing. Analysing, understanding and appropriately responding to human nonverbal behaviour and affective states are the central objectives of affective robotics research. Conventional machine learning approaches do not scale well to the dynamic nature of such real-world interactions as they require samples from stationary data distributions. The real-world is not stationary, it changes continuously. In such contexts, the training data and learning objectives may also change rapidly. Continual Learning (CL), by design, is able to address this very problem by learning incrementally. In this paper, we argue that CL is an essential paradigm for creating fully adaptive affective robots (why). To support this argument, we first provide an introduction to CL approaches and what they can offer for various dynamic (interactive) situations (what). We then formulate guidelines for the affective robotics community on how to utilise CL for perception and behaviour learning with adaptation (how). For each case, we reformulate the problem as a CL problem and outline a corresponding CL-based solution. We conclude the paper by highlighting the potential challenges to be faced and by providing specific recommendations on how to utilise CL for affective robotics.
Nikhil Churamani, Sinan Kalkan, Hatice Gunes
RO-MAN3
2020 Investigating Taste-liking with a Humanoid Robot Facilitator
abstract
Tasting is an essential activity in our daily lives. Implementing social robots in the food and drink service industry requires the social robots to be able to understand customers' nonverbal behaviours, including taste-liking. Little is known about whether people alter their behavioural responses related to taste-liking when interacting with a humanoid social robot. We conducted the first beverage tasting study where the facilitator is a human versus a humanoid social robot with priming versus non-priming instruction styles. We found that the facilitator type and facilitation style had no significant influence on cognitive taste-liking. However, in robot facilitator scenarios, people were more willing to follow the instruction and felt more comfortable when facilitated with priming. Our study provides new empirical findings and design implications for using humanoid social robots in the hospitality industry.
Zhuoni Jie, Hatice Gunes
RO-MAN2
2019 Your Fellows Matter: Affect Analysis across Subjects in Group Videos
abstract
Automatic affect analysis has become a well established research area in the last two decades. Recent works have started moving from individual to group scenarios. However, little attention has been paid to investigating how individuals in a group influence the affective states of each other. In this paper, we propose a novel framework for cross-subjects affect analysis in group videos. Specifically, we analyze the correlation of the affect among group members and investigate the automatic recognition of the affect of one subject using the behaviours expressed by another subject in the same group. A set of experiments are conducted using a recently collected database aimed at affect analysis in group settings. Our results show that (1) people in the same group do share more information in terms of behaviours and emotions than people in different groups; and (2) the affect of one subject in a group can be better predicted using the expressive behaviours of another subject within the same group than using that of a subject from a different group. This work is of great importance for affect recognition in group settings: when the information of one subject is unavailable due to occlusion, head/body poses etc., we can predict his/her affect by employing the expressive behaviours of the other subject(s).
Wenxuan Mou, Hatice Gunes, Ioannis Patras
FG2
2019 Registration-free Face-SSD: Single shot analysis of smiles, facial attributes, and affect in the wild
Youngkyoon Jang, Hatice Gunes, Ioannis Patras
Comput. Vis. Image Underst.2
2019 A deep generic to specific recognition model for group membership analysis using non-verbal cues
abstract
Automatic understanding and analysis of groups has attracted increasing attention in the vision and multimedia communities in recent years. However, little attention has been paid to the automatic analysis of the non-verbal behaviors and how this can be utilized for analysis of group membership, i.e., recognizing which group each individual is part of. This paper presents a novel Support Vector Machine (SVM) based Deep Specific Recognition Model (DeepSRM) that is learned based on a generic recognition model . The generic recognition model refers to the model trained with data across different conditions, i.e., when people are watching movies of different types. Although the generic recognition model can provide a baseline for the recognition model trained for each specific condition, the different behaviors people exhibit in different conditions limit the recognition performance of the generic model. Therefore, the specific recognition model is proposed for each condition separately and built on top of the generic recognition model . A number of experiments are conducted using a database aiming to study group analysis while each group (i.e., four participants together) were watching a number of long movie segments. Our experimental results show that the proposed deep specific recognition model (44%) outperforms the generic recognition model (26%). The recognition of group membership also indicates that the non-verbal behaviors of individuals within a group share commonalities.
Wenxuan Mou, Christos Tzelepis, Vasileios Mezaris, Hatice Gunes, Ioannis Patras
Image Vis. Comput.4
2019 Multimodal Human-Human-Robot Interactions (MHHRI) Dataset for Studying Personality and Engagement
abstract
In this paper we introduce a novel dataset, the Multimodal Human-Human-Robot-Interactions (MHHRI) dataset, with the aim of studying personality simultaneously in human-human interactions (HHI) and human-robot interactions (HRI) and its relationship with engagement. Multimodal data was collected during a controlled interaction study where dyadic interactions between two human participants and triadic interactions between two human participants and a robot took place with interactants asking a set of personal questions to each other. Interactions were recorded using two static and two dynamic cameras as well as two biosensors, and meta-data was collected by having participants fill in two types of questionnaires, for assessing their own personality traits and their perceived engagement with their partners (self labels) and for assessing personality traits of the other participants partaking in the study (acquaintance labels). As a proof of concept, we present baseline results for personality and engagement classification. Our results show that (i) trends in personality classification performance remain the same with respect to the self and the acquaintance labels across the HHI and HRI settings; (ii) for extroversion, the acquaintance labels yield better results as compared to the self labels; (iii) in general, multi-modality yields better performance for the classification of personality traits.
Oya Çeliktutan, Efstratios Skordos, Hatice Gunes
IEEE Trans. Affect. Comput.3
2019 Alone versus In-a-group: A Multi-modal Framework for Automatic Affect Recognition
abstract
Recognition and analysis of human affect has been researched extensively within the field of computer science in the past two decades. However, most of the past research in automatic analysis of human affect has focused on the recognition of affect displayed by people in individual settings and little attention has been paid to the analysis of the affect expressed in group settings. In this article, we first analyze the affect expressed by each individual in terms of arousal and valence dimensions in both individual and group videos and then propose methods to recognize the contextual information, i.e., whether a person is alone or in-a-group by analyzing their face and body behavioral cues. For affect analysis, we first devise affect recognition models separately in individual and group videos and then introduce a cross-condition affect recognition model that is trained by combining the two different types of data. We conduct a set of experiments on two datasets that contain both individual and group videos. Our experiments show that (1) the proposed Volume Quantized Local Zernike Moments Fisher Vector outperforms other unimodal features in affect analysis; (2) the temporal learning model, Long-Short Term Memory Networks, works better than the static learning model, Support Vector Machine; (3) decision fusion helps to improve affect recognition, indicating that body behaviors carry emotional information that is complementary rather than redundant to the emotion content in facial behaviors; and (4) it is possible to predict the context, i.e., whether a person is alone or in-a-group, using their non-verbal behavioral cues.
Wenxuan Mou, Hatice Gunes, Ioannis Patras
ACM Trans. Multim. Comput. Commun. Appl.2
2018 Detecting Deception and Suspicion in Dyadic Game Interactions
abstract
In this paper we focus on detection of deception and suspicion from electrodermal activity (EDA) measured on left and right wrists during a dyadic game interaction. We aim to answer three research questions: (i) Is it possible to reliably distinguish deception from truth based on EDA measurements during a dyadic game interaction? (ii) Is it possible to reliably distinguish the state of suspicion from trust based on EDA measurements during a card game? (iii) What is the relative importance of EDA measured on left and right wrists? To answer our research questions we conducted a study in which 20 participants were playing the game Cheat in pairs with one EDA sensor placed on each of their wrists. Our experimental results show that EDA measures from left and right wrists provide more information for suspicion detection than for deception detection and that the person-dependent detection is more reliable than the person-independent detection. In particular, classifying the EDA signal with Support Vector Machine (SVM) yields accuracies of 52% and 57% for person-independent prediction of deception and suspicion respectively, and 63% and 76% for person-dependent prediction of deception and suspicion respectively. Also, we found that: (i) the optimal interval of informative EDA signal for deception detection is about 1 s while it is around 3.5 s for suspicion detection; (ii) the EDA signal relevant for deception/suspicion detection can be captured after around 3.0 seconds after a stimulus occurrence regardless of the stimulus type (deception/truthfulness/suspicion/trust); and that (iii) features extracted from EDA from both wrists are important for classification of both deception and suspicion. To the best of our knowledge, this is the first work that uses EDA data to automatically detect both deception and suspicion in a dyadic game interaction setting.
Jan Ondras, Hatice Gunes
ICMI2
2017 Effects of valence and arousal on working memory performance in virtual reality gaming
abstract
The role of affective states in cognitive performance has long been an area of interest in cognitive science. Recent research in game-based cognitive training suggest that cognitive games should incorporate real-time adaptive mechanisms. These adaptive mechanisms would change the game's difficulty according to the player's performance in order to provide appropriate challenges and thus, achieve a real cognitive improvement. However, these mechanisms currently ignore the effects of valence and arousal on the player's cognitive skills. In this paper we investigate how working memory (WM) performance is affected when playing a VR game, and the effects of valence and arousal in this context. To this aim, a custom video game was created for Desktop and VR. Three difficulty levels were designed to evoke different levels of arousal while maintaining the same memory load for each difficulty level. We found an improvement in WM performance when playing in VR compared to Desktop. This effect was particularly pronounced in those with a low WM capacity. Significantly higher levels of valence and arousal were self-reported when playing in VR. We explore the impact that reported affective states could have in the player's WM performance. We suggest that high levels of arousal and positive valence can lead players to a flow state [1] that may have a positive impact on the player's WM performance.
Daniel Gábana Arellano, Laurissa N. Tokarchuk, Emily Hannon, Hatice Gunes
ACII4
2017 Computational analysis of valence and arousal in virtual reality gaming using lower arm electromyograms
abstract
Progress in the affective computing field has led to the creation of affect-aware games that aim to adapt to the emotions experienced by the players. In this paper we focus on affect recognition in virtual reality (VR) gaming, a problem that to the best of our knowledge has not yet been sufficiently explored. We aim to answer two research questions: (i) Is it possible to reliably capture and recognize the affective state of a person based on EMG sensors placed on their lower arms, while they interact with the virtual environment? and (ii) Is EMG signal from one arm sufficient for detecting affect? We conducted a study in which 8 people were playing a set of VR games with two EMG sensors placed on their arms. We analysed the EMG signals and extracted a number of features to infer the affective states of the players. Our experimental results show that the EMG measures from left and right arms provide sufficient information to detect emotions experienced by a player of a VR game. Our results also show that classifying a DWT-dbl signal with Support Vector Machine (SVM) yields F1=0.91 for predicting low/high arousal and F1=0.85 for predicting positive/negative valence when using just the left-arm EMG signal. To the best of our knowledge, this is the first work that uses EMG data from arm movements as a single source of affective information and addresses affect recognition in VR gaming.
Ilia Shumailov, Hatice Gunes
ACII2
2017 Generic to Specific Recognition Models for Membership Analysis in Group Videos
abstract
Automatic understanding and analysis of groups has attracted increasing attention in the vision and multimedia communities in recent years. However, little attention has been paid to the automatic analysis of group membership - i.e., recognizing which group the individual in question is part of. This paper presents a novel two-phase Support Vector Machine (SVM) based specific recognition model that is learned using an optimized generic recognition model. We conduct a set of experiments using a database collected to study group analysis from multimodal cues while each group (i.e., four participants together) were watching a number of long movie segments. Our experimental results show that the proposed specific recognition model (52%) outperforms the generic recognition model trained across all different videos (35%) and the independent recognition model trained directly on each specific video (33%) using linear SVM.
Wenxuan Mou, Christos Tzelepis, Vasileios Mezaris, Hatice Gunes, Ioannis Patras
FG4
2017 Automatic replication of teleoperator head movements and facial expressions on a humanoid robot
abstract
Robotic telepresence aims to create a physical presence for a remotely located human (teleoperator) by reproducing their verbal and nonverbal behaviours (e.g. speech, gestures, facial expressions) on a robotic platform. In this work, we propose a novel teleoperation system that combines the replication of facial expressions of emotions (neutral, disgust, happiness, and surprise) and head movements on the fly on the humanoid robot Nao. Robots' expression of emotions is constrained by their physical and behavioural capabilities. As the Nao robot has a static face, we use the LEDs located around its eyes to reproduce the teleoperator expressions of emotions. Using a web camera, we computationally detect the facial action units and measure the head pose of the operator. The emotion to be replicated is inferred from the detected action units by a neural network. Simultaneously, the measured head motion is smoothed and bounded to the robot's physical limits by applying a constrained-state Kalman filter. In order to evaluate the proposed system, we conducted a user study by asking 28 participants to use the replication system by displaying facial expressions and head movements while being recorded by a web camera. Subsequently, 18 external observers viewed the recorded clips via an online survey and assessed the quality of the robot's replication of the participants' behaviours. Our results show that the proposed teleoperation system can successfully communicate emotions and head movements, resulting in a high agreement among the external observers (ICC_E = 0.91, ICC_HP = 0.72).
Jan Ondras, Oya Çeliktutan, Evangelos Sariyanidi, Hatice Gunes
RO-MAN4
2017 Automatic Prediction of Impressions in Time and across Varying Context: Personality, Attractiveness and Likeability
abstract
In this paper, we propose a novel multimodal framework for automatically predicting the impressions of extroversion, agreeableness, conscientiousness, neuroticism , openness, attractiveness and likeability continuously in time and across varying situational contexts. Differently from the existing works, we obtain visual-only and audio-only annotations continuously in time for the same set of subjects, for the first time in the literature, and compare them to their audio-visual annotations. We propose a time-continuous prediction approach that learns the temporal relationships rather than treating each time instant separately. Our experiments show that the best prediction results are obtained when regression models are learned from audio-visual annotations and visual cues, and from audio-visual annotations and visual cues combined with audio cues at the decision level. Continuously generated annotations have the potential to provide insight into better understanding which impressions can be formed and predicted more dynamically, varying with situational context, and which ones appear to be more static and stable over time.
Oya Çeliktutan, Hatice Gunes
IEEE Trans. Affect. Comput.2
2017 Robust Registration of Dynamic Facial Sequences
abstract
Accurate face registration is a key step for several image analysis applications. However, existing registration methods are prone to temporal drift errors or jitter among consecutive frames. In this paper, we propose an iterative rigid registration framework that estimates the misalignment with trained regressors. The input of the regressors is a robust motion representation that encodes the motion between a misaligned frame and the reference frame(s), and enables reliable performance under non-uniform illumination variations. Drift errors are reduced when the motion representation is computed from multiple reference frames. Furthermore, we use the L2norm of the representation as a cue for performing coarse-to-fine registration efficiently. Importantly, the framework can identify registration failures and correct them. Experiments show that the proposed approach achieves significantly higher registration accuracy than the state-of-the-art techniques in challenging sequences.
Evangelos Sariyanidi, Hatice Gunes, Andrea Cavallaro
IEEE Trans. Image Process.2
2017 Biologically Inspired Motion Encoding for Robust Global Motion Estimation
abstract
The growing use of cameras embedded in autonomous robotic platforms and worn by people is increasing the importance of accurate global motion estimation (GME). However, existing GME methods may degrade considerably under illumination variations. In this paper, we address this problem by proposing a biologically-inspired GME method that achieves high estimation accuracy in the presence of illumination variations. We mimic the early layers of the human visual cortex with the spatio-temporal Gabor motion energy by adopting the pioneering model of Adelson and Bergen and we provide the closed-form expressions that enable the study and adaptation of this model to different application needs. Moreover, we propose a normalisation scheme for motion energy to tackle temporal illumination variations. Finally, we provide an overall GME scheme which, to the best of our knowledge, achieves the highest accuracy on the Pose, Illumination, and Expression (PIE) database.
Evangelos Sariyanidi, Hatice Gunes, Andrea Cavallaro
IEEE Trans. Image Process.2
2017 Learning Bases of Activity for Facial Expression Recognition
abstract
The extraction of descriptive features from sequences of faces is a fundamental problem in facial expression analysis. Facial expressions are represented by psychologists as a combination of elementary movements known as action units: each movement is localised and its intensity is specified with a score that is small when the movement is subtle and large when the movement is pronounced. Inspired by this approach, we propose a novel data-driven feature extraction framework that represents facial expression variations as a linear combination of localised basis functions, whose coefficients are proportional to movement intensity. We show that the linear basis functions required by this framework can be obtained by training a sparse linear model with Gabor phase shifts computed from facial videos. The proposed framework addresses generalisation issues that are not addressed by existing learnt representations, and achieves, with the same learning parameters, state-of-the-art results in recognising both posed expressions and spontaneous micro-expressions. This performance is confirmed even when the data used to train the model differ from test data in terms of the intensity of facial movements and frame rate.
Evangelos Sariyanidi, Hatice Gunes, Andrea Cavallaro
IEEE Trans. Image Process.2
2016 Personality Perception of Robot Avatar Tele-operators
abstract
Nowadays a significant part of human-human interaction takes place over distance. Tele-operated robot avatars, in which an operator's behaviours are portrayed by a robot proxy, have the potential to improve distance interaction, e.g., improving social presence and trust. However, having communication mediated by a robot changes the perception of the operator's appearance and behaviour, which have been shown to be used alongside vocal cues in judging personality. In this paper we present a study that investigates how robot mediation affects the way the personality of the operator is perceived. More specifically, we aim to investigate if judges of personality can be consistent in assessing personality traits, can agree with one another, can agree with operators' self-assessed personality, and shift their perceptions to incorporate characteristics associated with the robot's appearance. Our experiments show that (i) judges utilise robot appearance cues along with operator vocal cues to make their judgements, (ii) operators' arm gestures reproduced on the robot aid personality judgements, and (iii) how personality cues are perceived and evaluated through speech, gesture and robot appearance is highly operator-dependent. We discuss the implications of these results for both tele-operated and autonomous robots that aim to portray personality.
Paul Bremner, Oya Çeliktutan, Hatice Gunes
HRI3
2016 Alone versus In-a-group: A Comparative Analysis of Facial Affect Recognition
abstract
Automatic affect analysis and understanding has become a well established research area in the last two decades. Recent works have started moving from individual to group scenarios. However, little attention has been paid to comparing the affect expressed in individual and group settings. This paper presents a framework to investigate the differences in affect recognition models along arousal and valence dimensions in individual and group settings. We analyse how a model trained on data collected from an individual setting performs on test data collected from a group setting, and vice versa. A third model combining data from both individual and group settings is also investigated. A set of experiments is conducted to predict the affective states along both arousal and valence dimensions on two newly collected databases that contain sixteen participants watching affective movie stimuli in individual and group settings, respectively. The experimental results show that (1) the affect model trained with group data performs better on individual test data than the model trained with individual data tested on group data, indicating that facial behaviours expressed in a group setting capture more variation than in an individual setting; and (2) the combined model does not show better performance than the affect model trained with a specific type of data (i.e., individual or group), but proves a good compromise. These results indicate that in settings where multiple affect models trained with different types of data are not available, using the affect model trained with group data is a viable solution.
Wenxuan Mou, Hatice Gunes, Ioannis Patras
ACM Multimedia2
2016 Is automatic facial expression recognition of emotions coming to a dead end? The rise of the new kids on the block
abstract
Hatice Gunes’ work is partially supported the EPSRC under its IDEAS Factory Sandpits call on Digital Personhood (Grant Ref: EP/L00416X/1). Hayley Hung was partially supported by the Dutch national program COMMIT, by the European Commission under contract number FP7-ICT-600877 (SPENCER), and is affiliated with the Delft Data Science consortium.
Hatice Gunes, Hayley Hung
Image Vis. Comput.1
2015 Building autonomous sensitive artificial listeners (Extended abstract)
abstract
This paper describes a substantial effort to build a real-time interactive multimodal dialogue system with a focus on emotional and non-verbal interaction capabilities. The work is motivated by the aim to provide technology with competences in perceiving and producing the emotional and non-verbal behaviours required to sustain a conversational dialogue. We present the Sensitive Artificial Listener (SAL) scenario as a setting which seems particularly suited for the study of emotional and non-verbal behaviour, since it requires only very limited verbal understanding on the part of the machine. This scenario allows us to concentrate on non-verbal capabilities without having to address at the same time the challenges of spoken language understanding, task modeling etc. We first summarise three prototype versions of the SAL scenario, in which the behaviour of the Sensitive Artificial Listener characters was determined by a human operator. These prototypes served the purpose of verifying the effectiveness of the SAL scenario and allowed us to collect data required for building system components for analysing and synthesising the respective behaviours. We then describe the fully autonomous integrated real-time system we created, which combines incremental analysis of user behaviour, dialogue management, and synthesis of speaker and listener behaviour of a SAL character displayed as a virtual agent. We discuss principles that should underlie the evaluation of SAL-type systems. Since the system is designed for modularity and reuse, and since it is publicly available, the SAL system has potential as a joint research tool in the affective computing research community.
Marc Schröder 0001, Elisabetta Bevacqua, Roddy Cowie, Florian Eyben, Hatice Gunes, Dirk Heylen, Mark ter Maat, Gary McKeown, Sathish Pammi, Maja Pantic, Catherine Pelachaud, Björn W. Schuller, Etienne de Sevin, Michel F. Valstar, Martin Wöllmer
ACII5
2015 Face Alignment Assisted by Head Pose Estimation
abstract
In this paper we propose a supervised initialization scheme for cascaded face alignment based on explicit head pose estimation. We first investigate the failure cases of most state of the art face alignment approaches and observe that these failures often share one common global property, i.e. the head pose variation is usually large. Inspired by this, we propose a deep convolutional network model for reliable and accurate head pose estimation. Instead of using a mean face shape, or randomly selected shapes for cascaded face alignment initialisation, we propose two schemes for generating initialisation: the first one relies on projecting a mean 3D face shape (represented by 3D facial landmarks) onto 2D image under the estimated head pose; the second one searches nearest neighbour shapes from the training set according to head pose distance. By doing so, the initialisation gets closer to the actual shape, which enhances the possibility of convergence and in turn improves the face alignment performance. We demonstrate the proposed method on the benchmark 300W dataset and show very competitive performance in both head pose estimation and face alignment.
Heng Yang 0001, Wenxuan Mou, Ioannis Patras, Hatice Gunes, Peter Robinson 0001
BMVC5
2015 The LuminUs: Providing Musicians with Visual Feedback on the Gaze and Body Motion of Their Co-performers
Evan Morgan, Hatice Gunes, Nick Bryan-Kinns
INTERACT (2)2
2015 Tutorial on Emotional and Social Signals for Multimedia Research
abstract
No abstract available.
Hayley Hung, Hatice Gunes
ACM Multimedia2
2015 Computational analysis of human-robot interactions through first-person vision: Personality and interaction experience
abstract
In this paper, we analyse interactions with Nao, a small humanoid robot, from the viewpoint of human participants through an ego-centric camera placed on their forehead. We focus on human participants' and robot's personalities and their impact on the human-robot interactions. We automatically extract nonverbal cues (e.g., head movement) from first-person perspective and explore the relationship of nonverbal cues with participants' self-reported personality and their interaction experience. We generate two types of behaviours for the robot (i.e., extroverted vs. introverted) and examine how robot's personality and behaviour affect the findings. Significant correlations are obtained between the extroversion and agreeable-ness traits of the participants and the perceived enjoyment with the extroverted robot. Plausible relationships are also found between the measures of interaction experience and personality and the first-person vision features. We then use computational models to automatically predict the participants' personality traits from these features. Promising results are achieved for the traits of agreeableness, conscientiousness and extroversion.
Oya Çeliktutan, Hatice Gunes
RO-MAN2
2015 Using affective and behavioural sensors to explore aspects of collaborative music making
Evan Morgan, Hatice Gunes, Nick Bryan-Kinns
Int. J. Hum. Comput. Stud.2
2015 Automatic Analysis of Facial Affect: A Survey of Registration, Representation, and Recognition
abstract
Automatic affect analysis has attracted great interest in various contexts including the recognition of action units and basic or non-basic emotions. In spite of major efforts, there are several open questions on what the important cues to interpret facial expressions are and how to encode them. In this paper, we review the progress across a range of affect recognition applications to shed light on these fundamental questions. We analyse the state-of-the-art solutions by decomposing their pipelines into fundamental components, namely face registration, representation, dimensionality reduction and recognition. We discuss the role of these components and highlight the models and new trends that are followed in their design. Moreover, we provide a comprehensive analysis of facial representations by uncovering their advantages and limitations; we elaborate on the type of information they encode and discuss how they deal with the key challenges of illumination variations, registration errors, head-pose variations, occlusions, and identity bias. This survey allows us to identify open issues and to define future directions for designing real-world affect recognition systems.
Evangelos Sariyanidi, Hatice Gunes, Andrea Cavallaro
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Brief Introduction to the Special Issue on Behavior Understanding for Arts and Entertainment
abstract
This editorial introduction describes the aims and scope of the special issue of the ACM Transactions on Interactive Intelligent Systems on Behavior Understanding for Arts and Entertainment, which is being published in issues 2 and 3 of volume 5 of the journal. Here we offer a brief introduction to the use of behavior analysis for interactive systems that involve creativity in either the creator or the consumer of a work of art. We then characterize each of the five articles included in this first part of the special issue, which span a wide range of applications.
Albert Ali Salah, Hayley Hung, Oya Aran, Hatice Gunes, Matthew Turk 0001
ACM Trans. Interact. Intell. Syst.4
2015 Behavior Understanding for Arts and Entertainment
abstract
This editorial introduction complements the shorter introduction to the first part of the two-part special issue on Behavior Understanding for Arts and Entertainment. It offers a more expansive discussion of the use of behavior analysis for interactive systems that involve creativity, either for the producer or the consumer of such a system. We first summarise the two articles that appear in this second part of the special issue. We then discuss general questions and challenges in this domain that were suggested by the entire set of seven articles of the special issue and by the comments of the reviewers of these articles.
Albert Ali Salah, Hayley Hung, Oya Aran, Hatice Gunes, Matthew Turk 0001
ACM Trans. Interact. Intell. Syst.4
2014 Probabilistic Subpixel Temporal Registration for Facial Expression Analysis
Evangelos Sariyanidi, Hatice Gunes, Andrea Cavallaro
ACCV (4)2
2014 Continuous prediction of perceived traits and social dimensions in space and time
abstract
Developing automatic personality predictors requires generating reliable annotations, i.e., ground truth. To date, researchers have relied on the overall ratings provided for a whole video sequence, either obtained by self-assessment or provided by external observers. In this paper, we propose a novel personality assessment approach, where we ask external observers to continuously provide ratings along multiple dimensions ranging from 0 to 100 along time, and we generate continuous annotations in space and time. In addition to the widely used Big Five personality dimensions, we introduce three more dimensions that have the potential to gauge the reliability of the perceived social and trait judgements in the context of varying situational interactions between a human subject and virtual characters. Our results demonstrate the viability of the proposed approach and the plausible relationship between the extracted features and perceived trait and social dimensions. Annotations obtained continuously in time and in trait-social dimensional space showed that a number of dimensions appear to be more static and stable over time while other dimensions appear to be more dynamic.
Oya Çeliktutan, Hatice Gunes
ICIP2
2014 Automatic analysis of facial attractiveness from video
abstract
There has been a growing interest in the computer science field for automatic analysis and recognition of facial beauty and attractiveness. Most of the proposed studies attempt to model and predict facial attractiveness using a single static facial image. While a static image provides limited information about facial attractiveness, using a video clip that contains information about the motion and the dynamic behaviour of the face provides a richer understanding and valuable insights into analysing facial attractiveness. With this motivation, we propose to use dynamic features obtained from video clips along with static features obtained from static frames for automatic analysis of facial attractiveness. Support Vector Machine (SVM) and Random Forest (RF) are utilised to create and train models of attractiveness using the features extracted. Experimental results show that combining static and dynamic features improve performance over using either of these feature sets alone, and SVM provides the best prediction performance.
Sacide Kalayci, Hazim Kemal Ekenel, Hatice Gunes
ICIP3
2014 MAPTRAITS 2014 - The First Audio/Visual Mapping Personality Traits Challenge - An Introduction: Perceived Personality and Social Dimensions
abstract
The Audio/Visual Mapping Personality Challenge and Workshop (MAPTRAITS) is a competition event that is organised to facilitate the development of signal processing and machine learning techniques for the automatic analysis of personality traits and social dimensions. MAPTRAITS includes two sub-challenges, the continuous space-time sub-challenge and the quantised space-time sub-challenge. The continuous sub-challenge evaluated how systems predict the variation of perceived personality traits and social dimensions in time, whereas the quantised challenge evaluated the ability of systems to predict the overall perceived traits and dimensions in shorter video clips. To analyse the effect of audio and visual modalities on personality perception, we compared systems under three different settings: visual-only, audio-only and audio-visual. With MAPTRAITS we aimed at improving the knowledge on the automatic analysis of personality traits and social dimensions by producing a benchmarking protocol and encouraging the participation of various research groups from different backgrounds.
Oya Çeliktutan, Florian Eyben, Evangelos Sariyanidi, Hatice Gunes, Björn W. Schuller
ICMI4
2014 Automatic Prediction of Perceived Traits Using Visual Cues under Varied Situational Context
abstract
Automatic assessment of human personality traits is a non-trivial problem, especially when perception is marked over a fairly short duration of time. In this study, thin slices of behavioral data are analyzed. Perceived physical and behavioral traits are assessed by external observers (raters). Along with the big-five personality trait model, four new traits are introduced and assessed in this work. The relationship between various traits is investigated to obtain a better understanding of observer perception and assessment. Perception change is also considered when participants interact with several virtual characters each with a distinct emotional style. Encapsulating these observations and analysis, an automated system is proposed by firstly computing low level visual features. Using these features a separate model is trained for each trait and performance is evaluated. Further, a weighted model based on rater credibility is proposed to address observer biases. Experimental results indicate that a weighted model show major improvement for automatic prediction of perceived physical and behavioral traits.
Jyoti Joshi, Hatice Gunes, Roland Göcke
ICPR2
2013 Measuring Affect for the Study and Enhancement of Co-present Creative Collaboration
abstract
Affective computing research has tended to focus on the recognition of emotional states in individuals, with the intention of enhancing human-computer interaction. In this paper we advocate the need for a shift of attention towards emotional communication between people. To contextualise our views we discuss the ways in which rapid technological advances have impacted society and human psychology over the last decade. By outlining our doctoral research topic, we then highlight how affective computing based research could help us understand and enhance co-present human-human interactions. We are especially interested in studying situations where the interaction is directed towards collaborative creativity, as there is little existing work in this area and we see great potential for real-world applications to stem from our research.
Evan Morgan, Hatice Gunes, Nick Bryan-Kinns
ACII2
2013 Local Zernike Moment Representation for Facial Affect Recognition
abstract
Local representations became popular for facial affect recognition as they efficiently capture the image discontinuities, which play an important role for interpreting facial actions. We propose to use Local Zernike Moments (ZMs) [4] due to their useful and compact description of the image discontinuities and texture. Their main advantage in comparison to well-established alternatives such as Local Binary Patterns (LBPs) [5], is their flexibility in terms of the size and level of detail of the local description. We introduce a local ZM-based representation which involves a non-linear encoding layer (quantisation). The functionality of this layer is mapping similar facial configurations together and increasing compactness. We demonstrate the use of the local ZM-based representation for posed and naturalistic affect recognition on standard datasets, and show its superiority to alternative approaches for both tasks. Contemporary representations are often designed as frameworks consisting of three layers [2]: (Local) feature extraction, non-linear encoding and pooling. Non-linear encoding aims at enhancing the relevance of local features by increasing their robustness against image noise. Pooling describes small spatial neighbourhoods as single entities, ignoring the precise location of the encoded features, and increasing the tolerance against small geometric inconsistencies. In what follows, we describe the proposed local ZM-based representation scheme in terms of this threelayered framework. Feature Extraction – Local Zernike Moments: The computation of (complex) ZMs can be considered equivalent to representing an image in an alternative space. As shown in Figure 1-a, an image is decomposed onto a set of basis matrices (ZM bases), which are useful for describing the variation at different directions and scales. ZM bases are orthogonal, therefore there is no overlap in the information conveyed by each feature (ZM coefficient). ZMs are usually computed for the entire image, however in this case, ZMs cannot capture the local variation due to ZM bases lacking localisation [3]. In contrary, when computed around local neighbourhoods across the image, they become an efficient tool for describing the image discontinuities which are essential to interpreting facial activity. Non-linear Encoding – Quantisation: We perform quantisation via converting local features into binary values. Such coarse quantisation increases compactness and allows us to code each local block only with a single integer. Figure 1-b illustrates the process of obtaining the Quantised Local ZM (QLZM) image. Firstly, local ZM coefficients are computed across the input image (LZM layer) — each image in the LZM layer (LZM image) contains the features that are extracted through a particular ZM basis. Next, each LZM image is converted into a binary image by quantising each pixel via the signum(·) function. Finally, the QLZM image is obtained by combining all of the binary images. Specifically, each pixel in a particular location of the QLZM image is an integer (QLZM integer), computed by concatenating all of the binary values in the corresponding location of all binary images. The QLZM image is similar to an LBP-transformed image, in the sense that it contains integers of a limited range. Yet, the physical meaning of the information encoded by each integer is quite different. LBP integers describe a circular block by considering only the values along the border, neglecting the pixels that remain inside the block. Therefore, the efficient operation scale of LBPs is usually limited to 3-5 pixels [1, 5]. QLZM integers, on the other hand, describe blocks as a whole, and provide flexibility in terms of operation scale without major loss of information. Pooling – Histograms: Our representation scheme pools encoded features over local histograms. Figure 1-c illustrates the overall pipeline of the proposed representation scheme. Firstly, the QLZM image is computed through the process that is illustrated in detail in Figure 1-b. Next, . . . . . . ... = ZM Coefficients (local features) ZM Bases
Evangelos Sariyanidi, Hatice Gunes, Muhittin Gökmen, Andrea Cavallaro
BMVC2
2013 Fourth international workshop on human behavior understanding (HBU 2013)
abstract
With advances in pattern recognition and multimedia computing, it became possible to analyze human behavior via multimodal sensors, at different time-scales and at different levels of interaction and interpretation. This ability opens up enormous possibilities for multimedia and multimodal interaction, with a potential of endowing the computers with a capacity to attribute meaning to users' attitudes, preferences, personality, social relationships, etc., as well as to understand what people are doing, the activities they have been engaged in, their routines and lifestyles. This workshop gathers researchers dealing with the problem of modeling human behavior under its multiple facets with particular attention to interactions in arts, creativity, entertainment and edutainment.
Albert Ali Salah, Hayley Hung, Oya Aran, Hatice Gunes
ACM Multimedia4
2013 Introduction To The Special Issue On Affect Analysis In Continuous Input
Hatice Gunes, Björn W. Schuller
Image Vis. Comput.1
2013 Categorical and dimensional affect analysis in continuous input: Current trends and future directions
Hatice Gunes, Björn W. Schuller
Image Vis. Comput.1
2012 Dimensional and continuous analysis of emotions for multimedia applications: a tutorial overview
abstract
No abstract available.
Hatice Gunes, Björn W. Schuller
ACM Multimedia1
2012 Output-associative RVM regression for dimensional and continuous emotion prediction
Mihalis A. Nicolaou, Hatice Gunes, Maja Pantic
Image Vis. Comput.2
2012 Building Autonomous Sensitive Artificial Listeners
abstract
This paper describes a substantial effort to build a real-time interactive multimodal dialogue system with a focus on emotional and nonverbal interaction capabilities. The work is motivated by the aim to provide technology with competences in perceiving and producing the emotional and nonverbal behaviors required to sustain a conversational dialogue. We present the Sensitive Artificial Listener (SAL) scenario as a setting which seems particularly suited for the study of emotional and nonverbal behavior since it requires only very limited verbal understanding on the part of the machine. This scenario allows us to concentrate on nonverbal capabilities without having to address at the same time the challenges of spoken language understanding, task modeling, etc. We first report on three prototype versions of the SAL scenario in which the behavior of the Sensitive Artificial Listener characters was determined by a human operator. These prototypes served the purpose of verifying the effectiveness of the SAL scenario and allowed us to collect data required for building system components for analyzing and synthesizing the respective behaviors. We then describe the fully autonomous integrated real-time system we created, which combines incremental analysis of user behavior, dialogue management, and synthesis of speaker and listener behavior of a SAL character displayed as a virtual agent. We discuss principles that should underlie the evaluation of SAL-type systems. Since the system is designed for modularity and reuse and since it is publicly available, the SAL system has potential as a joint research tool in the affective computing research community.
Marc Schröder 0001, Elisabetta Bevacqua, Roddy Cowie, Florian Eyben, Hatice Gunes, Dirk Heylen, Mark ter Maat, Gary McKeown, Sathish Pammi, Maja Pantic, Catherine Pelachaud, Björn W. Schuller, Etienne de Sevin, Michel F. Valstar, Martin Wöllmer
IEEE Trans. Affect. Comput.5
2011 String-based audiovisual fusion of behavioural events for the assessment of dimensional affect
abstract
The automatic assessment of affect is mostly based on feature-level approaches, such as distances between facial points or prosodic and spectral information when it comes to audiovisual analysis. However, it is known and intuitive that behavioural events such as smiles, head shakes or laughter and sighs also bear highly relevant information regarding a subject's affective display. Accordingly, we propose a novel string-based prediction approach to fuse such events and to predict human affect in a continuous dimensional space. Extensive analysis and evaluation has been conducted using the newly released SEMAINE database of human-to-agent communication. For a thorough understanding of the obtained results, we provide additional benchmarks by more conventional feature-level modelling, and compare these and the string-based approach to fusion of signal-based features and string-based events. Our experimental results show that the proposed string-based approach is the best performing approach for automatic prediction of Valence and Expectation dimensions, and improves prediction performance for the other dimensions when combined with at least acoustic signal-based features.
Florian Eyben, Martin Wöllmer, Michel F. Valstar, Hatice Gunes, Björn W. Schuller, Maja Pantic
FG4
2011 Emotion representation, analysis and synthesis in continuous space: A survey
abstract
Despite major advances within the affective computing research field, modelling, analysing, interpreting and responding to naturalistic human affective behaviour still remains as a challenge for automated systems as emotions are complex constructs with fuzzy boundaries and with substantial individual variations in expression and experience. Thus, a small number of discrete categories (e.g., happiness and sadness) may not reflect the subtlety and complexity of the affective states conveyed by such rich sources of information. Therefore, affective and behavioural computing researchers have recently invested increased effort in exploring how to best model, analyse, interpret and respond to the subtlety, complexity and continuity (represented along a continuum e.g., from -1 to +1) of affective behaviour in terms of latent dimensions (e.g., arousal, power and valence) and appraisals. Accordingly, this paper aims to present the current state of the art and the new challenges in automatic, dimensional and continuous analysis and synthesis of human emotional behaviour in an interdisciplinary perspective.
Hatice Gunes, Björn W. Schuller, Maja Pantic, Roddy Cowie
FG1
2011 Output-associative RVM regression for dimensional and continuous emotion prediction
abstract
Many problems in machine learning and computer vision consist of predicting multi-dimensional output vectors given a specific set of input features. In many of these problems, there exist inherent temporal and spacial dependencies between the output vectors, as well as repeating output patterns and input-output associations, that can provide more robust and accurate predictors when modelled properly. With this intrinsic motivation, we propose a novel Output-Associative Relevance Vector Machine (OA-RVM) regression framework that augments the traditional RVM regression by being able to learn non-linear input and output dependencies. Instead of depending solely on the input patterns, OA-RVM models output structure and covariances within a predefined temporal window, thus capturing past, current and future context. As a result, output patterns manifested in the training data are captured within a formal probabilistic framework, and subsequently used during inference. As a proof of concept, we target the highly challenging problem of dimensional and continuous prediction of emotions from naturalistic facial expressions. We demonstrate the advantages of the proposed OA-RVM regression by performing both subject-dependent and subject-independent experiments using the SAL database. The experimental results show that OA-RVM regression outperforms the traditional RVM and SVM regression approaches in prediction accuracy, generating more robust and accurate models.
Mihalis A. Nicolaou, Hatice Gunes, Maja Pantic
FG2
2011 Come and have an emotional workout with sensitive artificial listeners!
abstract
This demonstration aims to showcase the recently completed SEMAINE system. The SEMAINE system is a publicly available, fully autonomous Sensitive Artificial Listeners (SAL) system that consists of virtual dialog partners based on audiovisual analysis and synthesis (see http://semaine.opendfki.de/wiki). The system runs in real-time, and combines incremental analysis of user behavior, dialog management, and synthesis of speaker and listener behavior of a SAL character, displayed as a virtual agent. The SAL characters intend to engage the user in a conversation by paying attention to the user's emotions and nonverbal expressions. The characters have their own emotionally defined personality. During an interaction, the characters attempt to create an emotional workout for the user by drawing her/him towards their dominant emotion, through a combination of verbal and nonverbal expressions.
Marc Schröder 0001, Sathish Pammi, Hatice Gunes, Maja Pantic, Michel F. Valstar, Roddy Cowie, Gary McKeown, Dirk Heylen, Mark ter Maat, Florian Eyben, Björn W. Schuller, Martin Wöllmer, Elisabetta Bevacqua, Catherine Pelachaud, Etienne de Sevin
FG3
2011 A multi-layer hybrid framework for dimensional emotion classification
abstract
This paper investigates dimensional emotion prediction and classification from naturalistic facial expressions. Similarly to many pattern recognition problems, dimensional emotion classification requires generating multi-dimensional outputs. To date, classification for valence and arousal dimensions has been done separately, assuming that they are independent. However, various psychological findings suggest that these dimensions are correlated. We therefore propose a novel, multi-layer hybrid framework for emotion classification that is able to model inter-dimensional correlations. Firstly, we derive a novel geometric feature set based on the (a)symmetric spatio-temporal characteristics of facial expressions. Subsequently, we use the proposed feature set to train a multi-layer hybrid framework composed of a tem- poral regression layer for predicting emotion dimensions, a graphical model layer for modeling valence-arousal correlations, and a final classification and fusion layer exploiting informative statistics extracted from the lower layers. This framework (i) introduces the Auto-Regressive Coupled HMM (ACHMM), a graphical model specifically tailored to accommodate not only inter-dimensional correlations but also to exploit the internal dynamics of the actual observations, and (ii) replaces the commonly used Maximum Likelihood principle with a more robust final classification and fusion layer. Subject-independent experimental validation, performed on a naturalistic set of facial expressions, demonstrates the effectiveness of the derived feature set, and the robustness and flexibility of the proposed framework.
Mihalis A. Nicolaou, Hatice Gunes, Maja Pantic
ACM Multimedia2
2011 MLiT: mixtures of Gaussians under linear transformations
Ahmed Fawzi Otoom, Hatice Gunes, Óscar Pérez, Massimo Piccardi
Pattern Anal. Appl.2
2011 Continuous Prediction of Spontaneous Affect from Multiple Cues and Modalities in Valence-Arousal Space
abstract
Past research in analysis of human affect has focused on recognition of prototypic expressions of six basic emotions based on posed data acquired in laboratory settings. Recently, there has been a shift toward subtle, continuous, and context-specific interpretations of affective displays recorded in naturalistic and real-world settings, and toward multimodal analysis and recognition of human affect. Converging with this shift, this paper presents, to the best of our knowledge, the first approach in the literature that: 1) fuses facial expression, shoulder gesture, and audio cues for dimensional and continuous prediction of emotions in valence and arousal space, 2) compares the performance of two state-of-the-art machine learning techniques applied to the target problem, the bidirectional Long Short-Term Memory neural networks (BLSTM-NNs), and Support Vector Machines for Regression (SVR), and 3) proposes an output-associative fusion framework that incorporates correlations and covariances between the emotion dimensions. Evaluation of the proposed approach has been done using the spontaneous SAL data from four subjects and subject-dependent leave-one-sequence-out cross validation. The experimental results obtained show that: 1) on average, BLSTM-NNs outperform SVR due to their ability to learn past and future context, 2) the proposed output-associative fusion framework outperforms feature-level and model-level fusion by modeling and learning correlations and patterns between the valence and arousal dimensions, and 3) the proposed system is well able to reproduce the valence and arousal ground truth obtained from human coders.
Mihalis A. Nicolaou, Hatice Gunes, Maja Pantic
IEEE Trans. Affect. Comput.2
2010 Audio-Visual Classification and Fusion of Spontaneous Affective Data in Likelihood Space
abstract
This paper focuses on audio-visual (using facial expression, shoulder and audio cues) classification of spontaneous affect, utilising generative models for classification (i) in terms of Maximum Likelihood Classification with the assumption that the generative model structure in the classifier is correct, and (ii) Likelihood Space Classification with the assumption that the generative model structure in the classifier may be incorrect, and therefore, the classification performance can be improved by projecting the results of generative classifiers onto likelihood space, and then using discriminative classifiers. Experiments are conducted by utilising Hidden Markov Models for single cue classification, and 2 and 3-chain coupled Hidden Markov Models for fusing multiple cues and modalities. For discriminative classification, we utilise Support Vector Machines. Results show that Likelihood Space Classification improves the performance (91.76%) of Maximum Likelihood Classification (79.1%). Thereafter, we introduce the concept of fusion in the likelihood space, which is shown to outperform the typically used model-level fusion, attaining a classification accuracy of 94.01% and further improving all previous results.
Mihalis A. Nicolaou, Hatice Gunes, Maja Pantic
ICPR2
2010 Dimensional Emotion Prediction from Spontaneous Head Gestures for Interaction with Sensitive Artificial Listeners
Hatice Gunes, Maja Pantic
IVA1
2009 Mixtures of Normalized Linear Projections
Ahmed Fawzi Otoom, Óscar Pérez, Hatice Gunes, Massimo Piccardi
ACIVS3
2009 Head detection for video surveillance based on categorical hair and skin colour models
abstract
We propose a new robust head detection algorithm that is capable of handling significantly different conditions in terms of viewpoint, tilt angle, scale and resolution. To this aim, we built a new model for the head based on appearance distributions and shape constraints. We construct a categorical model for hair and skin, separately, and train the models for four categories of hair (brown, red, blond and black) and three categories of skin representing the different illumination conditions (bright, standard and dark). The shape constraint fits an elliptical model to the candidate region and compares its parameters with priors based on human anatomy. The experimental results validate the usability of the proposed algorithm in various video surveillance and multimedia applications.
Zui Zhang, Hatice Gunes, Massimo Piccardi
ICIP2
2009 Static vs. dynamic modeling of human nonverbal behavior from multiple cues and modalities
abstract
Human nonverbal behavior recognition from multiple cues and modalities has attracted a lot of interest in recent years. Despite the interest, many research questions, including the type of feature representation, choice of static vs. dynamic classification schemes, the number and type of cues or modalities to use, and the optimal way of fusing these, remain open research questions. This paper compares frame-based vs window-based feature representation and employs static vs. dynamic classification schemes for two distinct problems in the field of automatic human nonverbal behavior analysis: multicue discrimination between posed and spontaneous smiles from facial expressions, head and shoulder movements, and audio-visual discrimination between laughter and speech. Single cue and single modality results are compared to multicue and multimodal results by employing Neural Networks, Hidden Markov Models (HMMs), and 2- and 3-chain coupled HMMs. Subject independent experimental evaluation shows that: 1) both for static and dynamic classification, fusing data coming from multiple cues and modalities proves useful to the overall task of recognition, 2) the type of feature representation appears to have a direct impact on the classification performance, and 3) static classification is comparable to dynamic classification both for multicue discrimination between posed and spontaneous smiles, and audio-visual discrimination between laughter and speech.
Stavros Petridis, Hatice Gunes, Sebastian Kaltwang, Maja Pantic
ICMI2
2009 Automatic Temporal Segment Detection and Affect Recognition From Face and Body Display
abstract
Psychologists have long explored mechanisms with which humans recognize other humans' affective states from modalities, such as voice and face display. This exploration has led to the identification of the main mechanisms, including the important role played in the recognition process by the modalities' dynamics. Constrained by the human physiology, the temporal evolution of a modality appears to be well approximated by a sequence of temporal segments called onset, apex, and offset. Stemming from these findings, computer scientists, over the past 15 years, have proposed various methodologies to automate the recognition process. We note, however, two main limitations to date. The first is that much of the past research has focused on affect recognition from single modalities. The second is that even the few multimodal systems have not paid sufficient attention to the modalities' dynamics: The automatic determination of their temporal segments, their synchronization to the purpose of modality fusion, and their role in affect recognition are yet to be adequately explored. To address this issue, this paper focuses on affective face and body display, proposes a method to automatically detect their temporal segments or phases, explores whether the detection of the temporal phases can effectively support recognition of affective states, and recognizes affective states based on phase synchronization/alignment. The experimental results obtained show the following: 1) affective face and body displays are simultaneous but not strictly synchronous; 2) explicit detection of the temporal phases can improve the accuracy of affect recognition; 3) recognition from fused face and body modalities performs better than that from the face or the body modality alone; and 4) synchronized feature-level fusion achieves better performance than decision-level fusion.
Hatice Gunes, Massimo Piccardi
IEEE Trans. Syst. Man Cybern. Part B1
2008 Tracking People in Crowds by a Part Matching Approach
abstract
The major difficulty in human tracking is the problem raised by challenging occlusions where the target person is repeatedly and extensively occluded by either the background or another moving object. These types of occlusions may cause significant changes in the person¿s shape, appearance or motion, thus making the data association problem extremely difficult to solve. Unlike most of the existing methods for human tracking that handle occlusions by data association of the complete human body, in this paper we propose a method that tracks people under challenging spatial occlusions based on body part tracking. The human model we propose consists of five body parts with six degrees of freedom and each part is represented by a rich set of features. The tracking is solved using a layered data association approach, direct comparison between features (feature layer) and subsequently matching between parts of the same bodies (part layer) lead to a final decision for the global match (global layer). Experimental results have confirmed the effectiveness of the proposed method.
Zui Zhang, Hatice Gunes, Massimo Piccardi
AVSS2
2008 Feature extraction techniques for abandoned object classification in video surveillance
abstract
We address the problem of abandoned object classification in video surveillance. Our aim is to determine (i) which feature extraction technique proves more useful for accurate object classification in a video surveillance context (scale invariant image transform (SIFT) keypoints vs. geometric primitive features), and (ii) how the resulting features affect classification accuracy and false positive rates for different classification schemes used. Objects are classified into four different categories: bag (s), person (s), trolley (s), and group (s) of people. Our experimental results show that the highest recognition accuracy and the lowest false alarm rate are achieved by building a classifier based on our proposed set of statistics of geometric primitives' features. Moreover, classification performance based on this set of features proves to be more invariant across different learning algorithms.
Ahmed Fawzi Otoom, Hatice Gunes, Massimo Piccardi
ICIP2
2008 An accurate algorithm for head detection based on XYZ and HSV hair and skin color models
abstract
Head detection in images and videos plays an important role in a wide range of computer vision and multimedia applications. In this paper, we propose a new head detection algorithm that is capable of handling significantly variable conditions in terms of viewpoint (i.e. frontal, profile, back view, from −180 degrees to +180 degrees), tilt angle (i.e. from horizontal to aerial), scale and resolution. To this aim, we built a new model for the head based on appearance distributions and shape constraints. The appearance distribution models the colors of hair and skin by sets of Gaussian mixtures in the XYZ and HSV color spaces. The shape constraint fits an elliptical model to the candidate region and compares its parameters with priors based on the human anatomy. This presents a pixel-level measurement of accuracy for the proposed algorithm both prior and after applying the spatial constraints referenced by the elliptical model. The excellent accuracy at both levels confirms the accuracy of the appearance model and the appropriateness of the spatial and topological process.
Zui Zhang, Hatice Gunes, Massimo Piccardi
ICIP2
2008 Maximum-likelihood dimensionality reduction in gaussian mixture models with an application to object classification
abstract
Accurate classification of objects of interest for video surveillance is difficult due to occlusions, deformations and variable views/illumination. The adopted feature sets tend to overcome these issues by including many and complementary features; however, their large dimensionality poses an intrinsic challenge to the classification task. In this paper, we present a novel technique providing maximum-likelihood dimensionality reduction in Gaussian mixture models for classification. The technique, called hereafter mixture of maximum-likelihood normalized projections (mixture of ML-NP), was used in this work to classify a 44-dimensional data set into 4 classes (bag, trolley, single person, group of people). The accuracy achieved on an independent test set is 98% vs. 80% of the runner-up (MultiBoost/AdaBoost).
Massimo Piccardi, Hatice Gunes, Ahmed Fawzi Otoom
ICPR2
2008 Comparative performance analysis of feature sets for abandoned object classification
abstract
Accurate classification of abandoned objects is crucial in video surveillance systems. In this paper, we experiment with different validation techniques (hold-out and 10-fold cross validation), with the aim of determining which feature set proves more useful for accurate object classification in a video surveillance context (scale invariant image transform (SIFT) keypoints vs. geometric primitive features). Moreover, we show how the resulting features affect classification performance across different classifiers. We also further analyze the best performing classifier in order to have better understanding of its classification results. Objects are classified into four different categories: bag (s), person (s), trolley (s), and group (s) of people. Our experimental results show that the highest recognition accuracy and the lowest false alarm rate are achieved by building a classifier based on our proposed set of statistics of geometric primitives' features. This set of features maximizes inter-class separation and simplifies the classification process. Classification based on this set of features thus outperforms the second best approach based on SIFT keypoint histograms by providing on average 22% higher recognition accuracy and 7% lower false alarm rate.
Ahmed Fawzi Otoom, Hatice Gunes, Massimo Piccardi
SMC2
2007 How to distinguish posed from spontaneous smiles using geometric features
abstract
Automatic distinction between posed and spontaneous expressions is an unsolved problem. Previously cognitive sciences' studies indicated that the automatic separation of posed from spontaneous expressions is possible using the face modality alone. However, little is known about the information contained in head and shoulder motion. In this work, we propose to (i) distinguish between posed and spontaneous smiles by fusing the head, face, and shoulder modalities, (ii) investigate which modalities carry important information and how the information of the modalities relate to each other, and (iii) to which extent the temporal dynamics of these signals attribute to solving the problem. We use a cylindrical head tracker to track the head movements and two particle filtering techniques to track the facial and shoulder movements. Classification is performed by kernel methods combined with ensemble learning techniques. We investigated two aspects of multimodal fusion: the level of abstraction (i.e., early, mid-level, and late fusion) and the fusion rule used (i.e., sum, product and weight criteria). Experimental results from 100 videos displaying posed smiles and 102 videos displaying spontaneous smiles are presented. Best results were obtained with late fusion of all modalities when 94.0% of the videos were classified correctly.
Michel F. Valstar, Hatice Gunes, Maja Pantic
ICMI2
2007 Bi-modal emotion recognition from expressive face and body gestures
Hatice Gunes, Massimo Piccardi
J. Netw. Comput. Appl.1
2006 Creating and Annotating Affect Databases from Face and Body Display: A Contemporary Survey
abstract
Databases containing representative samples of human multi-modal expressive behavior are needed for the development of affect recognition systems. However, at present publicly-available databases exist mainly for single expressive modalities such as facial expressions, static and dynamic hand postures, and dynamic hand gestures. Only recently, a first bimodal affect database consisting of expressive face and upper-body display has been released. To foster development of affect recognition systems, this paper presents a comprehensive survey of the current state-of-the art in affect database creation from face and body display and elicits the requirements of an ideal multi-modal affect database.
Hatice Gunes, Massimo Piccardi
SMC1
2006 Assessing facial beauty through proportion analysis by image processing and supervised learning
Hatice Gunes, Massimo Piccardi
Int. J. Hum. Comput. Stud.1
2005 Fusing Face and Body Display for Bi-modal Emotion Recognition: Single Frame Analysis and Multi-frame Post Integration
Hatice Gunes, Massimo Piccardi
ACII1
2005 Affect recognition from face and body: early fusion vs. late fusion
abstract
This paper presents an approach to automatic visual emotion recognition from two modalities: face and body. Firstly, individual classifiers are trained from individual modalities. Secondly, we fuse facial expression and affective body gesture information first at a feature-level, in which the data from both modalities are combined before classification, and later at a decision-level, in which we integrate the outputs of the monomodal systems by the use of suitable criteria. We then evaluate these two fusion approaches, in terms of performance over monomodal emotion recognition based on facial expression modality only. In the experiments performed the emotion classification using the two modalities achieved a better recognition accuracy outperforming the classification using the individual facial modality. Moreover, fusion at the feature-level proved better recognition than fusion at the decision-level.
Hatice Gunes, Massimo Piccardi
SMC1
2004 Video object encoder using selective local-space support vector machines
abstract
Recently, support vector machine (SVM) has been shown to be a good classifier; however, its large computational requirement prohibited its use in real time video processing applications. In this paper, a model is proposed that enables use of SVM in video applications. The proposed model allows selected image scales (of interest) to be encoded and classified more accurately by complex classifier such as SVM, whilst other image scales of less significance to be encoded and classified by simpler encoder/classifier. Experiment with video object encoding shows that the performance of the proposed model is comparable with other models, however with reduced computational requirements.
Po Hsiang Tsai, Seun Jan, Hatice Gunes
MMSP3
2004 Automated classification of female facial beauty by image analysis and supervised learning
abstract
The fact that perception of facial beauty may be a universal concept has long been debated amongst psychologists and anthropologists. In this paper, we performed experiments to evaluate the extent of beauty universality by asking a number of diverse human referees to grade a same collection of female facial images. Results obtained show that the different individuals gave similar votes, thus well supporting the concept of beauty universality. We then trained an automated classifier using the human votes as the ground truth and used it to classify an independent test set of facial images. The high accuracy achieved proves that this classifier can be used as a general, automated tool for objective classification of female facial beauty. Potential applications exist in the entertainment industry and plastic surgery.
Hatice Gunes, Massimo Piccardi, Tony Jan
VCIP1