EDBT 2026 Demo / reviewers in the wild / expert
Rosalind W. Picard
dblp:p/RWPicard
· DBLP profile ↗
166ranked-venue papers
27as first author
22since 2021 · last 2026
0000-0002-5661-0022ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 74 · 10 first-author · 12 since 2021Artificial intelligence and machine learning · 58 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 10 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1Computer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Designing an Affective Mobile Probe to Measure Smile Dynamics in DepressionabstractDepression is a complex disorder for which there is growing interest in identifying objective behavioral markers that measure precise symptoms, such as anhedonia and blunted emotional reactivity. This study explores the feasibility of using smile and smirk expression dynamics, captured through our novel stimulus-based mobile affective probe, as candidate digital biomarkers of depression severity within a large-scale mobile health intervention trial, BeWell. Data from 684 BeWell participants (2,702 observations) are analyzed longitudinally for 16 weeks, comparing their PHQ-8 survey scores with their facial responses to short videos intended to elicit smiles. Mixed-effects models reveal that higher maximum Duchenne smile intensity in reaction to liked stimuli is associated with lower depression scores over time at both within- and between-person levels. We additionally share insights from our tool, including ease of use, perceptions of the stimulus, and technical challenges, which offer considerations for the future development of stimulus-based affect probes in real-world settings. Nelson Hidalgo Julia, Robert Lewis 0001, Craig Ferguson, Joshua Angulo Lopez, Hahrin Jung, Rosalind W. Picard, Simon Goldberg, Raquel Tatar, Wendy Lau, Caroline Swords, Christine D. Wilson-Mendenhall, Gabriela Valdivia, Molly Schaefer, Richard Davidson |
CHI | 6 |
| 2025 | Identifying Vocal and Facial Biomarkers of Depression in Large-Scale Remote Recordings: A Multimodal Study Using Mixed-Effects Modeling
Nelson Hidalgo Julia, Robert Lewis 0001, Craig Ferguson, Simon Goldberg, Wendy Lau, Caroline Swords, Gabriela Valdivia, Christine D. Wilson-Mendenhall, Raquel Tatar, Rosalind W. Picard, Richard Davidson |
INTERSPEECH | 10 |
| 2025 | Towards the Objective Characterisation of Major Depressive Disorder Using Speech Data from a 12-week Observational Study with Daily Measurements
Robert Lewis 0001, Szymon Fedor, Nelson Hidalgo Julia, Joshua Curtiss, Jiyeon Kim, Noah Jones, David Mischoulon, Thomas F. Quatieri, Nicholas Cummins, Paola Pedrelli, Rosalind W. Picard |
INTERSPEECH | 11 |
| 2025 | Cultivating a Supportive Sphere: Designing Technology to Increase Social Support for Foster-Involved YouthabstractApproximately 400,000 youth in the US are living in foster care due to experiences with abuse or neglect at home[17]. For multiple reasons, these youth often don't receive adequate social support from those around them. Despite technology's potential, very little work has explored how these tools can provide more support to foster-involved youth. To begin to fill this gap, we worked with current and former foster-involved youth to develop the first digital tool that aims to increase social support for this population, creating a novel system in which users complete reflective check-ins in an online community setting. We then conducted a pilot study with 15 current and former foster-involved youth, comparing the effect of using the app for two weeks to two weeks of no intervention. We collected qualitative and quantitative data, which demonstrated that this type of interface can provide youth with types of social support that are often not provided by foster care services and other digital interventions. The paper details the motivation behind the app, the trauma-informed design process, and insights gained from this initial evaluation study. Finally, the paper concludes with recommendations for designing digital tools that effectively provide social support to foster-involved youth. Ila Krishna Kumar, Craig Ferguson, Jiayi Wu 0009, Rosalind W. Picard |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2025 | Connecting through Comics: Design and Evaluation of Cube, an Arts-Based Digital Platform for Trauma-Impacted YouthabstractThis paper explores the design, development and evaluation of a digital platform that aims to assist young people who have experienced trauma in understanding and expressing their emotions and fostering social connections. Integrating principles from expressive arts and narrative-based therapies, we collaborate with lived experts to iteratively design a novel, user-centered digital tool for young people to create and share comics that represent their experiences. Specifically, we conduct a series of nine workshops with N=54 trauma-impacted youth and young adults to test and refine our tool, beginning with three workshops using low-fidelity prototypes, followed by six workshops with Cube, a web version of the tool. A qualitative analysis of workshop feedback and empathic relations analysis of artifacts provides valuable insights into the usability and potential impact of the tool, as well as the specific needs of young people who have experienced trauma. Our findings suggest that the integration of expressive and narrative therapy principles into Cube can offer a unique avenue for trauma-impacted young people to process their experiences, more easily communicate their emotions, and connect with supportive communities. We end by presenting implications for the design of social technologies that aim to support the emotional well-being and social integration of youth and young adults who have faced trauma. Ila Krishna Kumar, Jocelyn Shen, Craig Ferguson, Rosalind W. Picard |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2025 | Analyzing the Visual Road Scene for Driver Stress EstimationabstractThis paper studies the contribution of the visual road scene to estimate the driver-reported stress levels. Our research leverages on previous work showing that environmental factors, such as traffic congestion, weather conditions, and driving context, impact driver’s stress. Each of the models we evaluated is trained and tested with the publicly available AffectiveROAD dataset to estimate three categories of driver-reported stress level. We test three types of modelling approaches: (i) single-frame baselines (Random Forest, SVM, and Convolutional Neural Networks); (ii) Temporal Segment Networks (TSN) and two variants of it, which use learned weights (TSN-w) and LSTM (TSN-LSTM) as consensus functions; and (iii) video classification Transformers. Our experiments reveal that the TSN-w, TSN-LSTM, and Transformer models achieve statistically equivalent performances, all significantly outperforming the other models. Particularly noteworthy is TSN-w, which attains the highest performance observed with an average accuracy of 0.77. We further provide an explainability analysis using Class Activation Mapping and image semantic segmentation to identify the elements of the road scene that contribute the most to high levels of stress. Our results demonstrate that the visible road scene offers significant contextual information for estimating driver-reported stress levels, with potential implications for the design of safer urban road environments. Cristina Bustos, Albert Solé-Ribalta, Neska El Haouij, Javier Borge-Holthoefer, Àgata Lapedriza, Rosalind W. Picard |
IEEE Trans. Affect. Comput. | 6 |
| 2024 | A HeARTfelt Robot: Social Robot-Driven Deep Emotional Art Reflection with ChildrenabstractSocial-emotional learning (SEL) skills are essential for children to develop to provide a foundation for future relational and academic success. Using art as a medium for creation or as a topic to provoke conversation is a well-known method of SEL learning. Similarly, social robots have been used to teach SEL competencies like empathy, but the combination of art and social robotics has been minimally explored. In this paper, we present a novel child-robot interaction designed to foster empathy and promote SEL competencies via a conversation about art scaffolded by a social robot. Participants (N=11, age range: 7-11) conversed with a social robot about emotional and neutral art. Analysis of video and speech data demonstrated that this interaction design successfully engaged children in the practice of SEL skills, like emotion recognition and self-awareness, and greater rates of empathetic reasoning were observed when children engaged with the robot about emotional art. This study demonstrated that art-based reflection with a social robot, particularly on emotional art, can foster empathy in children, and interactions with a social robot help alleviate discomfort when sharing deep or vulnerable emotions. Isabella Pu, Golda Nguyen, Lama Alsultan, Rosalind W. Picard, Cynthia Breazeal, Sharifa Alghowinem |
RO-MAN | 4 |
| 2024 | Guest Editorial: Ethics in Affective ComputingabstractStunning advances in machine learning are heralding a new era in sensing, interpreting, simulating and stimulating human emotion. In the human sciences, research is increasingly highlighting the explanatory power of emotions, feelings, and other affective processes to predict how we think and behave. This is beginning to translate into an explosion of applications that can improve human wellbeing including methods to reduce stress and improve emotion regulation skills, techniques to support healthier social media use, pain monitoring in neonates, and decision-support tools that recognize emotional bias. Jonathan Gratch, Gretchen Greene, Rosalind W. Picard, Lachlan Urquhart, Michel F. Valstar |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Assessing Affective Engagement with Narratives of Invisible DisabilityabstractNarratives about invisible disabilities are poorly represented in public discourse and often go undisclosed [1], leading to false assumptions, discrimination, and stigma [2] against those who experience these conditions. To address these issues, recent studies have suggested that disclosure of first-person narratives of invisible disabilities should be increased [3]. To understand the mechanisms affecting recipients of such narratives, the present study evaluates how social media users (N = 124) engage affectively with this content in a digitally mediated narrative-form intervention designed to reduce harmful assumptions against persons who experience invisible disabilities. Results of this study indicate that such an intervention may prove effective at reducing harmful assumptions on the basis of visual cues, and in line with past research, finds that affect may play an important role in assumption-making processes [4]. Findings from this study may be used to inform novel digital interventions capable of counteracting harmful assumptions that drive prejudicial behaviors against a wide range of populations and communities. Daniel T. Kessler, David Y. J. Kim, Grace S. Ahn, Neska El Haouij, Rosalind W. Picard |
ACII | 5 |
| 2023 | A Robotic Companion for Psychological Well-being: A Long-term Investigation of Companionship and Therapeutic AllianceabstractSocial support plays a crucial role in managing and enhancing one's mental health and well-being. In order to explore the role of a robot's companion-like behavior on its therapeutic interventions, we conducted an eight-week-long deployment study with seventy participants to compare the impact of (1) acontrol robot with only assistant-like skills, (2) acoach-like robot with additional instructive positive psychology interventions, and (3) acompanion-like robot that delivered the same interventions in a peer-like and supportive manner. The companion-like robot was shown to be the most effective in building a positive therapeutic alliance with people, enhancing participants' well-being and readiness for change. Our work offers valuable insights into how companion AI agents could further enhance the efficacy of the mental health interventions by strengthening their therapeutic alliance with people for long-term mental health support. Sooyeon Jeong, Laura Aymerich-Franch, Sharifa Alghowinem, Rosalind W. Picard, Cynthia Breazeal, Hae Won Park 0001 |
HRI | 4 |
| 2023 | MultiPar-T: Multiparty-Transformer for Capturing Contingent Behaviors in Group ConversationsabstractAs we move closer to real-world social AI systems, AI agents must be able to deal with multiparty (group) conversations. Recognizing and interpreting multiparty behaviors is challenging, as the system must recognize individual behavioral cues, deal with the complexity of multiple streams of data from multiple people, and recognize the subtle contingent social exchanges that take place amongst group members. To tackle this challenge, we propose the Multiparty-Transformer (Multipar- T), a transformer model for multiparty behavior modeling. The core component of our proposed approach is Crossperson Attention, which is specifically designed to detect contingent behavior between pairs of people. We verify the effectiveness of Multipar-T on a publicly available video-based group engagement detection benchmark, where it outperforms state-of-the-art approaches in average F-1 scores by 5.2% and individual class F-1 scores by up to 10.0%. Through qualitative analysis, we show that our Crossperson Attention module is able to discover contingent behaviors. Dong Won Lee 0007, Yubin Kim 0002, Rosalind W. Picard, Cynthia Breazeal, Hae Won Park 0001 |
IJCAI | 3 |
| 2023 | Deploying a robotic positive psychology coach to improve college students' psychological well-beingabstractDespite the increase in awareness and support for mental health, college students' mental health is reported to decline every year in many countries. Several interactive technologies for mental health have been proposed and are aiming to make therapeutic service more accessible, but most of them only provide one-way passive contents for their users, such as psycho-education, health monitoring, and clinical assessment. We present a robotic coach that not only delivers interactive positive psychology interventions but also provides other useful skills to build rapport with college students. Results from our on-campus housing deployment feasibility study showed that the robotic intervention showed significant association with increases in students' psychological well-being, mood, and motivation to change. We further found that students' personality traits were associated with the intervention outcomes as well as their working alliance with the robot and their satisfaction with the interventions. Also, students' working alliance with the robot was shown to be associated with their pre-to-post change in motivation for better well-being. Analyses on students' behavioral cues showed that several verbal and nonverbal behaviors were associated with the change in self-reported intervention outcomes. The qualitative analyses on the post-study interview suggest that the robotic coach's companionship made a positive impression on students, but also revealed areas for improvement in the design of the robotic coach. Results from our feasibility study give insight into how learning users' traits and recognizing behavioral cues can help an AI agent provide personalized intervention experiences for better mental health outcomes. Sooyeon Jeong, Laura Aymerich-Franch, Kika Arias, Sharifa Alghowinem, Àgata Lapedriza, Rosalind W. Picard, Hae Won Park 0001, Cynthia Breazeal |
User Model. User Adapt. Interact. | 6 |
| 2022 | Computational Empathy Counteracts the Negative Effects of Anger on Creative Problem SolvingabstractHow does empathy influence creative problem solving? We introduce a computational empathy intervention based on context-specific affective mimicry and perspective taking by a virtual agent appearing in the form of a well-dressed polar bear. In an online experiment with 1,006 participants randomly assigned to an emotion elicitation intervention (with a control elicitation condition and anger elicitation condition) and a computational empathy intervention (with control virtual agent and an empathic virtual agent), we examine how anger and empathy influence participants' performance in solving a word game based on Wordle. We find participants who are assigned to the anger elicitation condition perform significantly worse on multiple performance metrics than participants assigned to the control condition. However, we find the empathic virtual agent counteracts the drop in performance induced by the anger condition such that participants assigned to both the empathic virtual agent and the anger condition perform no differently than participants in the control elicitation condition and significantly better than participants assigned to the control virtual agent and the anger elicitation condition. While empathy reduces the negative effects of anger, we do not find evidence that the empathic virtual agent influences performance of participants who are assigned to the control elicitation condition. By introducing a framework for computational empathy interventions and conducting a two-by-two factorial design randomized experiment, we provide rigorous, empirical evidence that computational empathy can counteract the negative effects of anger on creative problem solving. Matthew Groh, Craig Ferguson, Robert Lewis 0001, Rosalind W. Picard |
ACII | 4 |
| 2022 | Affective Ratings of Nonverbal Vocalizations Produced by Minimally-Speaking Individuals: What Do Naive Listeners Perceive?abstractIndividuals who produce few spoken words (“minimally-speaking” individuals) often convey rich affective and communicative information through nonverbal vocalizations, such as grunts, yells, babbles, and monosyllabic expressions. Yet, little data exists on the affective content of the vocal expressions of this population. Here, we present 78,624 arousal and valence ratings of nonverbal vocalizations from the online ReCANVo (Real-World Communicative and Affective Nonverbal Vocalizations) database. This dataset contains over 7,000 vocalizations that have been labeled with their expressive functions (delight, frustration, etc.) from eight minimally-speaking individuals. Our results suggest that raters who have no knowledge of the context or meaning of a nonverbal vocalization are still able to detect arousal and valence differences between different types of vocalizations based on Likert-scale ratings. Moreover, these ratings are consistent with hypothesized arousal and valence rankings for the different vocalization types. Raters are also able to detect arousal and valence differences between different vocalization types within individual speakers. To our knowledge, this is the first large-scale analysis of affective content within nonverbal vocalizations from minimally verbal individuals. These results complement affective computing research of nonverbal vocalizations that occur within typical verbal speech (e.g., grunts, sighs) and serve as a foundation for further understanding of how humans perceive emotions in sounds. Kristina T. Johnson, Amanda O'Brien, Ayelet M. Kershenbaum, Jaya Narain, Simon Radhakrishnan, Thomas F. Quatieri, Rosalind W. Picard |
ACII | 7 |
| 2022 | DISSECT: Disentangled Simultaneous Explanations via Concept Traversals
Asma Ghandeharioun, Been Kim, Chun-Liang Li, Brendan Jou, Brian Eoff, Rosalind W. Picard |
ICLR | 6 |
| 2022 | Modeling Real-World Affective and Communicative Nonverbal Vocalizations From Minimally Speaking IndividualsabstractNonverbal vocalizations from non- and minimally speaking individuals who speak fewer than 20 words (mv* individuals) convey important communicative and affective information. While nonverbal vocalizations that occur amidst typical speech and infant vocalizations have been studied extensively in the literature, there is limited prior work on vocalizations by mv* individuals. Our work is among the first studies of the communicative and affective information expressed in nonverbal vocalizations by mv* children and adults. We collected labeled vocalizations in real-world settings with eight mv* communicators, with communicative and affective labels provided in-the-moment by a close family member. Using evaluation strategies suitable for messy, real-world data, we show that nonverbal vocalizations can be classified by function (with 4- and 5-way classifications) with F1 scores above chance for all participants. We analyze labeling and data collection practices for each participating family, and discuss the classification results in the context of our novel real-world data collection protocol. The presented work includes results from the largest classification experiments with nonverbal vocalizations from mv* communicators to date. Jaya Narain, Kristina T. Johnson, Thomas F. Quatieri, Rosalind W. Picard, Pattie Maes |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | Automatic Recognition Methods Supporting Pain Assessment: A SurveyabstractPain is a complex phenomenon, involving sensory and emotional experience, that is often poorly understood, especially in infants, anesthetized patients, and others who cannot speak. Technology supporting pain assessment has the potential to help reduce suffering; however, advances are needed before it can be adopted clinically. This survey paper assesses the state of the art and provides guidance for researchers to help make such advances. First, we overview pain’s biological mechanisms, physiological and behavioral responses, emotional components, as well as assessment methods commonly used in the clinic. Next, we discuss the challenges hampering the development and validation of pain recognition technology, and we survey existing datasets together with evaluation methods. We then present an overview of all automated pain recognition publications indexed in the Web of Science as well as from the proceedings of the major conferences on biomedical informatics and artificial intelligence, to provide understanding of the current advances that have been made. We highlight progress in both non-contact and contact-based approaches, tools using face, voice, physiology, and multi-modal information, the importance of context, and discuss challenges that exist, including identification of ground truth. Finally, we identify underexplored areas such as chronic pain and connections to treatments, and describe promising opportunities for continued advances. Philipp Werner, Daniel Lopez Martinez, Steffen Walter 0001, Ayoub Al-Hamadi, Sascha Gruss, Rosalind W. Picard |
IEEE Trans. Affect. Comput. | 6 |
| 2022 | Holistic Affect Recognition Using PaNDA: Paralinguistic Non-Metric Dimensional AnalysisabstractHumans perceive emotion from each other using a holistic perspective, accounting for diverse personal, non-emotional variables, such as age and personality, that shape expression. In contrast, today’s algorithms are mainly designed to recognize emotion in isolation, and are usually demonstrated only within one relatively narrow database. In this article, we propose a multi-task learning approach to jointly learn the recognition of affective states from speech along with various speaker attributes. A problem with multi-task learning is that sometimes inductive transfer can negatively impact performance. To mitigate negative transfer, we introduce the Paralinguistic Non-metric Dimensional Analysis (PaNDA) method that systematically measures task relatedness and also enables visualizing the topology of affective phenomena as a whole. In addition, we present a generic framework that conflates the concepts of single-task and multi-task learning. Using this framework, we construct two models that demonstrate holistic affect recognition: one treats all tasks as equally related, whereas the other one incorporates the task correlations between a main task and its supporting tasks obtained from PaNDA. Both models employ a multi-task deep neural network, in which separate output layers are used to predict discrete and continuous attributes, while hidden layers are shared across different tasks. On average across 18 classification and regression tasks, the weighted multi-task learning with PaNDA significantly improves performance compared to single-task and unweighted multi-task learning. Yue Zhang 0014, Felix Weninger, Björn W. Schuller, Rosalind W. Picard |
IEEE Trans. Affect. Comput. | 4 |
| 2021 | Predicting Driver Self-Reported Stress by Analyzing the Road SceneabstractSeveral studies have shown the relevance of biosignals in driver stress recognition. In this work, we examine something important that has been less frequently explored: We develop methods to test if the visual driving scene can be used to estimate a drivers’ subjective stress levels. For this purpose, we use the AffectiveROAD video recordings and their corresponding stress labels, a continuous human-driver-provided stress metric. We use the common class discretization for stress, dividing its continuous values into three classes: low, medium, and high. We design and evaluate three computer vision modeling approaches to classify the driver’s stress levels: (1) object presence features, where features are computed using automatic scene segmentation; (2) end-to-end image classification; and (3) end-to-end video classification. All three approaches show promising results, suggesting that it is possible to approximate the drivers’ subjective stress from the information found in the visual scene. We observe that the video classification, which processes the temporal information integrated with the visual information, obtains the highest accuracy of 0.72, compared to a random baseline accuracy of 0.33 when tested on a set of nine drivers. Cristina Bustos, Neska El Haouij, Albert Solé-Ribalta, Javier Borge-Holthoefer, Àgata Lapedriza, Rosalind W. Picard |
ACII | 6 |
| 2021 | Guidelines for Assessing and Minimizing Risks of Emotion Recognition ApplicationsabstractSociety has witnessed a rapid increase in the adoption of commercial uses of emotion recognition. Tools that were traditionally used by domain experts are now being used by individuals who are often unaware of the technology’s limitations and may use them in potentially harmful settings. The change in scale and agency, paired with gaps in regulation, urge the research community to rethink how we design, position, implement and ultimately deploy emotion recognition to anticipate and minimize potential risks. To help understand the current ecosystem of applied emotion recognition, this work provides an overview of some of the most frequent commercial applications and identifies some of the potential sources of harm. Informed by these, we then propose 12 guidelines for systematically assessing and reducing the risks presented by emotion recognition applications. These guidelines can help identify potential misuses and inform future deployments of emotion recognition. Javier Hernandez, Josh Lovejoy, Daniel McDuff, Jina Suh, Tim O'Brien, Arathi Sethumadhavan, Gretchen Greene, Rosalind W. Picard, Mary Czerwinski |
ACII | 8 |
| 2021 | The Guardians: Designing a Game for Long-term Engagement with Mental Health TherapyabstractThis work introduces The Guardians: Unite the Realms, a novel free-to-play and publicly released mobile game that encourages the adoption of healthy real-world behaviours in exchange for rewards that enrich the gaming experience. We describe the game, its grounding in a mental health therapy known as behavioural activation, and how we designed it to keep players engaged over time. Instead of using traditional digital health gamification techniques such as badges or leaderboards, The Guardians creates a motivational pull by embedding the therapy into a complete mobile game. In-game items earned via the therapy have an immediate purpose in the game and, thus, they are considered intrinsically valuable by players. Analysis of game interaction data from 7,782 real-world users suggests 15-day and 30-day retention rates of 10.0% and 6.6%, respectively, which is more than double the average retention levels of most digital mental health interventions. Furthermore, players reported completion of a healthy real-world task on 69.0% of days played (37,574 completed tasks in 54,461 total days). We also report interaction metrics with game features and the effectiveness of the players' chosen real-world activities. Craig Ferguson, Robert Lewis 0001, Chelsey Wilks, Rosalind W. Picard |
CoG | 4 |
| 2021 | Beyond the Words: Analysis and Detection of Self-Disclosure Behavior during Robot Positive Psychology InteractionabstractSelf-disclosure is an important part of mental health treatment process. As interactive technologies are becoming more widely available, many AI agents for mental health prompt their users to self-disclose as part of the intervention activities. However, most existing works focus on linguistic features to classify self-disclosure behavior, and do not utilize other multi-modal behavioral cues. We present analyses of people's non-verbal cues (vocal acoustic features, head orientation and body gestures/movements) exhibited during self-disclosure tasks based on the human-robot interaction data collected in our previous work. Results from the classification experiments suggest that prosody, head pose, and body postures can be independently used to detect self-disclosure behavior with high accuracy (up to 81%). Moreover, positive emotions, high engagement, self-soothing and positive attitudes behavioral cues were found to be positively correlated to self-disclosure. Insights from our work can help build a self-disclosure detection model that can be used in real time during multi-modal interactions between humans and AI agents. Sharifa Alghowinem, Sooyeon Jeong, Kika Arias, Rosalind W. Picard, Cynthia Breazeal, Hae Won Park 0001 |
FG | 4 |
| 2020 | Hierarchical Reinforcement Learning for Open-Domain DialogabstractOpen-domain dialog generation is a challenging problem; maximum likelihood training can lead to repetitive outputs, models have difficulty tracking long-term conversational goals, and training on standard movie or online datasets may lead to the generation of inappropriate, biased, or offensive text. Reinforcement Learning (RL) is a powerful framework that could potentially address these issues, for example by allowing a dialog model to optimize for reducing toxicity and repetitiveness. However, previous approaches which apply RL to open-domain dialog generation do so at the word level, making it difficult for the model to learn proper credit assignment for long-term conversational rewards. In this paper, we propose a novel approach to hierarchical reinforcement learning (HRL), VHRL, which uses policy gradients to tune the utterance-level embedding of a variational sequence model. This hierarchical approach provides greater flexibility for learning long-term, conversational rewards. We use self-play and RL to optimize for a set of human-centered conversation metrics, and show that our approach provides significant improvements – in terms of both human evaluation and automatic metrics – over state-of-the-art dialog models, including Transformers. Abdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Rosalind W. Picard |
AAAI | 5 |
| 2020 | Human-centric dialog training via offline reinforcement learningabstractNatasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, Rosalind Picard. Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020. Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun, Craig Ferguson, Àgata Lapedriza, Noah Jones, Shixiang Gu, Rosalind W. Picard |
EMNLP (1) | 8 |
| 2020 | Dyadic Speech-based Affect Recognition using DAMI-P2C Parent-child Multimodal Interaction DatasetabstractAutomatic speech-based affect recognition of individuals in dyadic conversation is a challenging task, in part because of its heavy reliance on manual pre-processing. Traditional approaches frequently require hand-crafted speech features and segmentation of speaker turns. In this work, we design end-to-end deep learning methods to recognize each person's affective expression in an audio stream with two speakers, automatically discovering features and time regions relevant to the target speaker's affect. We integrate a local attention mechanism into the end-to-end architecture and compare the performance of three attention implementations - one mean pooling and two weighted pooling methods. Our results show that the proposed weighted-pooling attention solutions are able to learn to focus on the regions containing target speaker's affective information and successfully extract the individual's valence and arousal intensity. Here we introduce and use a "dyadic affect in multimodal interaction - parent to child" (DAMI-P2C) dataset collected in a study of 34 families, where a parent and a child (3-7 years old) engage in reading storybooks together. In contrast to existing public datasets for affect recognition, each instance for both speakers in the DAMI-P2C dataset is annotated for the perceived affect by three labelers. To encourage more research on the challenging task of multi-speaker affect sensing, we make the annotated DAMI-P2C dataset publicly available, including acoustic features of the dyads' raw audios, affect annotations, and a diverse set of developmental, social, and demographic profiles of each dyad. Huili Chen, Yue Zhang 0014, Felix Weninger, Rosalind W. Picard, Cynthia Breazeal, Hae Won Park 0001 |
ICMI | 4 |
| 2020 | Personalized Modeling of Real-World Vocalizations from Nonverbal IndividualsabstractNonverbal vocalizations contain important affective and communicative information, especially for those who do not use traditional speech, including individuals who have autism and are non- or minimally verbal (nv/mv). Although these vocalizations are often understood by those who know them well, they can be challenging to understand for the community-at-large. This work presents (1) a methodology for collecting spontaneous vocalizations from nv/mv individuals in natural environments, with no researcher present, and personalized in-the-moment labels from a family member; (2) speaker-dependent classification of these real-world sounds for three nv/mv individuals; and (3) an interactive application to translate the nonverbal vocalizations in real time. Using support-vector machine and random forest models, we achieved speaker-dependent unweighted average recalls (UARs) of 0.75, 0.53, and 0.79 for the three individuals, respectively, with each model discriminating between 5 nonverbal vocalization classes. We also present first results for real-time binary classification of positive- and negative-affect nonverbal vocalizations, trained using a commercial wearable microphone and tested in real time using a smartphone. This work informs personalized machine learning methods for non-traditional communicators and advances real-world interactive augmentative technology for an underserved population. Jaya Narain, Kristina T. Johnson, Craig Ferguson, Amanda O'Brien, Tanya Talkar, Yue Zhang 0014, Peter Wofford, Thomas F. Quatieri, Rosalind W. Picard, Pattie Maes |
ICMI | 9 |
| 2020 | Going with our Guts: Potentials of Wearable Electrogastrography (EGG) for Affect DetectionabstractA hard challenge for wearable systems is to measure differences in emotional valence, i.e. positive and negative affect via physiology. However, the stomach or gastric signal is an unexplored modality that could offer new affective information. We created a wearable device and software to record gastric signals, known as electrogastrography (EGG). An in-laboratory study was conducted to compare EGG with electrodermal activity (EDA) in 33 individuals viewing affective stimuli. We found that negative stimuli attenuate EGG's indicators of parasympathetic activation, or "rest and digest" activity. We compare EGG to the remaining physiological signals and describe implications for affect detection. Further, we introduce how wearable EGG may support future applications in areas as diverse as reducing nausea in virtual reality and helping treat emotion-related eating disorders. Angela Vujic, Stephanie Tong, Rosalind W. Picard, Pattie Maes |
ICMI | 3 |
| 2020 | A Robotic Positive Psychology Coach to Improve College Students' WellbeingabstractA significant number of college students suffer from mental health issues that impact their physical, social, and occupational outcomes. Various scalable technologies have been proposed in order to mitigate the negative impact of mental health disorders. However, the evaluation for these technologies, if done at all, often reports mixed results on improving users' mental health. We need to better understand the factors that align a user's attributes and needs with technology-based interventions for positive outcomes. In psychotherapy theory, therapeutic alliance and rapport between a therapist and a client is regarded as the basis for therapeutic success. In prior works, social robots have shown the potential to build rapport and a working alliance with users in various settings. In this work, we explore the use of a social robot coach to deliver positive psychology interventions to college students living in on-campus dormitories. We recruited 35 college students to participate in our study and deployed a social robot coach in their room. The robot delivered daily positive psychology sessions among other useful skills like delivering the weather forecast, scheduling reminders, etc. We found a statistically significant improvement in participants' psychological wellbeing, mood, and readiness to change behavior for improved wellbeing after they completed the study. Furthermore, students' personality traits were found to have a significant association with intervention efficacy. Analysis of the post-study interview revealed students' appreciation of the robot's companionship and their concerns for privacy. Sooyeon Jeong, Sharifa Alghowinem, Laura Aymerich-Franch, Kika Arias, Àgata Lapedriza, Rosalind W. Picard, Hae Won Park 0001, Cynthia Breazeal |
RO-MAN | 6 |
| 2020 | Studying Personalized Just-in-time Auditory Breathing Guides and Potential Safety Implications during Simulated DrivingabstractDriving can occupy a considerable part of our daily lives and is often associated with high levels of stress. Motivated by the effectiveness of controlled breathing, this work studies the potential use of breathing interventions while driving to help manage stress. In particular, we implemented and evaluated a closed-loop system that monitored the breathing rate of drivers in real-time and delivered either a conscious or an unconscious personalized acoustic breathing guide whenever needed. In a study with 24 participants, we observed that conscious interventions more effectively reduced the breathing rate but also increased the number of driving mistakes. We observed that prior driving experience as well as personality are significantly associated with the effect of the interventions, which highlights the importance of considering user profiles for in-car stress management interventions. Sebastian Zepf, Neska El Haouij, Jinmo Lee, Asma Ghandeharioun, Javier Hernandez, Rosalind W. Picard |
UMAP | 6 |
| 2020 | Toward Assessing and Recommending Combinations of Behaviors for Improving Health and Well-BeingabstractMultiple behaviors typically work together to influence health, making it hard to understand how one behavior might compensate for another. Rich multi-modal datasets from mobile sensors and advances in machine learning are today enabling new kinds of associations to be made between combinations of behaviors objectively assessed from daily life and self-reported levels of stress, mood, and health. In this article, we present a framework to (1) map multi-modal messy data collected in the “wild” to meaningful feature representations of health-related behaviors, (2) uncover latent patterns comprising combinations of behaviors that best predict health and well-being, and (3) use these learned patterns to make evidence-based recommendations that may improve health and well-being. We show how to use supervised latent Dirichlet allocation to model the observed behaviors, and we apply variational inference to uncover the latent patterns. Implementing and evaluating the model on 5,397 days of data from a group of 244 college students, we find that these latent patterns are indeed predictive of daily self-reported levels of stressed-calm, sad-happy, and sick-healthy states. We investigate the patterns of modifiable behaviors present on different days and uncover several ways in which they relate to stress, mood, and health. This work contributes a new method using objective data analysis to help advance understanding of how combinations of modifiable human behaviors may promote human health and well-being. Ehimwenma Nosakhare, Rosalind W. Picard |
ACM Trans. Comput. Heal. | 2 |
| 2020 | Personalized Multitask Learning for Predicting Tomorrow's Mood, Stress, and HealthabstractWhile accurately predicting mood and wellbeing could have a number of important clinical benefits, traditional machine learning (ML) methods frequently yield low performance in this domain. We posit that this is because a one-size-fits-all machine learning model is inherently ill-suited to predicting outcomes like mood and stress, which vary greatly due to individual differences. Therefore, we employ Multitask Learning (MTL) techniques to train personalized ML models which are customized to the needs of each individual, but still leverage data from across the population. Three formulations of MTL are compared: i) MTL deep neural networks, which share several hidden layers but have final layers unique to each task; ii) Multi-task Multi-Kernel learning, which feeds information across tasks through kernel weights on feature types; and iii) a Hierarchical Bayesian model in which tasks share a common Dirichlet Process prior. We offer the code for this work in open source. These techniques are investigated in the context of predicting future mood, stress, and health using data collected from surveys, wearable sensors, smartphone logs, and the weather. Empirical results demonstrate that using MTL to account for individual differences provides large performance improvements over traditional machine learning methods and provides personalized, actionable insights. Sara Taylor, Natasha Jaques, Ehimwenma Nosakhare, Akane Sano, Rosalind W. Picard |
IEEE Trans. Affect. Comput. | 5 |
| 2019 | Tweet Moodifier: Towards giving emotional awareness to Twitter usersabstractEmotional contagion in online social networks has been of great interest over the past years. Previous studies have focused mainly on finding evidence of affect contagion in homophilic atmospheres. However, these studies have overlooked users' awareness of the sentiments they share and consume online. In this paper, we present an experiment with Twitter users that aims to help them better understand which emotions they experience on this social network. We introduce Tweet Moodifier (T-Moodifier), a Google Chrome extension that enables Twitter users to filter and make explicit (through colored visual marks) the emotional content in their News Feed. We compare behavioral changes between 55 participants and 5089 of their public “friends.” The comparison period spans from two weeks before installing T-Moodifier to one week thereafter. The results suggest that the use of T-Moodifier might help Twitter users increase their emotional awareness: T-Moodifier users who had access to emotional statistics about their posts produced a significantly higher percentage of neutral content. This behavioral change suggests that people could behave differently while using real-time mechanisms that increase their affect reflection. Also, post-experience, those who completed both pre- and post-surveys could assert more confidently the main emotions they shared and perceived on Twitter. This shows T-Moodifier's potential to effectively make users reflect on their News Feed. Belén Saldías-Fuentes, Rosalind W. Picard |
ACII | 2 |
| 2019 | Engineering Music to Slow Breathing and Invite Relaxed PhysiologyabstractWe engineered an interactive music system that influences a user's breathing rate to induce a relaxation response. This system generates ambient music containing periodic shifts in loudness that are determined by the user's own breathing patterns. We evaluated the efficacy of this music intervention for participants who were engaged in an attention-demanding task, and thus explicitly not focusing on their breathing or on listening to the music. We measured breathing patterns in addition to multiple peripheral and cortical indicators of physiological arousal while users experienced three different interaction designs: (1) a “Fixed Tempo” amplitude modulation rate at six beats per minute; (2) a “Personalized Tempo” modulation rate fixed at 75% of each individual's breathing rate baseline, and (3) a “Personalized Envelope” design in which the amplitude modulation matches each individual's breathing pattern in real-time. Our results revealed that each interactive music design slowed down breathing rates, with the “Personalized Tempo” design having the largest effect, one that was more significant than the non-personalized design. The physiological arousal indicators (electrodermal activity, heart rate, and slow cortical potentials measured in EEG) showed concomitant reductions, suggesting that slowing users' breathing rates shifted them towards a more calmed state. These results suggest that interactive music incorporating biometric data may have greater effects on physiology than traditional recorded music. Grace Leslie, Asma Ghandeharioun, Diane Y. Zhou, Rosalind W. Picard |
ACII | 4 |
| 2019 | Unintentional affective priming during labeling may bias labelsabstractOnline platforms displaying long streams of examples are often employed to gather labels from both experts and crowd workers. While previous work in crowdsourcing focused on objective tasks and estimating error parameters of annotators, collecting labels in a subjective setting (e.g. emotion recognition) is more complicated due to different interpretations of examples. These interpretations could be influenced by many factors such as annotator mood and previously seen examples. In this work, we examine two hypotheses of order-dependent biases in sequential labeling tasks: negatively auto-correlated sequential decision making and positively auto-correlated affective priming. Using controlled generation of facial expressions, we find that i) annotators achieve higher agreement when presented examples in the same sequential order, ii) the valence label of the current image positively correlates with the previous labels given. While we also observe a positive correlation between labels and the number of preceding positive and negative images seen, this correlation is highly dependent on example ordering. Our findings demonstrate that randomized examples given to annotators may produce systematic bias in labels. Future data collection should present examples in orderings which mitigate such bias. Judy Hanwen Shen, Àgata Lapedriza, Rosalind W. Picard |
ACII | 3 |
| 2019 | Multi-modal Active Learning From Human Data: A Deep Reinforcement Learning ApproachabstractHuman behavior expression and experience are inherently multimodal, and characterized by vast individual and contextual heterogeneity. To achieve meaningful human-computer and human-robot interactions, multi-modal models of the user’s states (e.g., engagement) are therefore needed. Most of the existing works that try to build classifiers for the user’s states assume that the data to train the models are fully labeled. Nevertheless, data labeling is costly and tedious, and also prone to subjective interpretations by the human coders. This is even more pronounced when the data are multi-modal (e.g., some users are more expressive with their facial expressions, some with their voice). Thus, building models that can accurately estimate the user’s states during an interaction is challenging. To tackle this, we propose a novel multi-modal active learning (AL) approach that uses the notion of deep reinforcement learning (RL) to find an optimal policy for active selection of the user’s data, needed to train the target (modality-specific) models. We investigate different strategies for multi-modal data fusion, and show that the proposed model-level fusion coupled with RL outperforms the feature-level and modality-specific models, and the naïve AL strategies such as random sampling, and the standard heuristics such as uncertainty sampling. We show the benefits of this approach on the task of engagement estimation from real-world child-robot interactions during an autism therapy. Importantly, we show that the proposed multi-modal AL approach can be used to efficiently personalize the engagement classifiers to the target user using a small amount of actively selected user’s data. Ognjen Rudovic, Meiru Zhang, Björn W. Schuller, Rosalind W. Picard |
ICMI | 4 |
| 2019 | Probabilistic Latent Variable Modeling for Assessing Behavioral Influences on Well-BeingabstractHealth research has an increasing focus on promoting well-being and positive mental health, to prevent disease and to more effectively treat disorders. The availability of rich multi-modal datasets and advances in machine learning methods are now enabling data science research to begin to objectively assess well-being. However, most existing studies focus on detecting the current state or predicting the future state of well-being using stand-alone health behaviors. There is a need for methods that can handle a complex combination of health behaviors, as arise in real-world data. Ehimwenma Nosakhare, Rosalind W. Picard |
KDD | 2 |
| 2019 | Approximating Interactive Human Evaluation with Self-Play for Open-Domain Dialog SystemsabstractBuilding an open-domain conversational agent is a challenging problem. Current evaluation methods, mostly post-hoc judgments of static conversation, do not capture conversation quality in a realistic interactive context. In this paper, we investigate interactive human evaluation and provide evidence for its necessity; we then introduce a novel, model-agnostic, and dataset-agnostic method to approximate it. In particular, we propose a self-play scenario where the dialog system talks to itself and we calculate a combination of proxies such as sentiment and semantic coherence on the conversation trajectory. We show that this metric is capable of capturing the human-rated quality of a dialog model better than any automated metric known to-date, achieving a significant Pearson correlation (r>.7, p<.05). To investigate the strengths of this novel metric and interactive evaluation in comparison to state-of-the-art metrics and human evaluation of static conversations, we perform extended experiments with a set of models, including several that make novel improvements to recent hierarchical dialog generation architectures through sentiment and semantic knowledge distillation on the utterance level. Finally, we open-source the interactive evaluation platform we built and the dataset we collected to allow researchers to efficiently deploy and evaluate dialog models. Asma Ghandeharioun, Judy Hanwen Shen, Natasha Jaques, Craig Ferguson, Noah Jones, Àgata Lapedriza, Rosalind W. Picard |
NeurIPS | 7 |
| 2019 | Wearable Motion-Based Heart Rate at Rest: A Workplace EvaluationabstractThis paper studies the feasibility of using low-cost motion sensors to provide opportunistic heart rate assessments from ballistocardiographic signals during restful periods of daily life. Three wearable devices were used to capture peripheral motions at specific body locations (head, wrist, and trouser pocket) of 15 participants during five regular workdays each. Three methods were implemented to extract heart rate from motion data and their performance was compared to those obtained with an FDA-cleared device. With a total of 1358 h of naturalistic sensor data, our results show that providing accurate heart rate estimations from peripheral motion signals is possible during relatively "still" moments. In our real-life workplace study, the head-mounted device yielded the most frequent assessments (22.98% of the time under 5 beats per minute of error) followed by the smartphone in the pocket (5.02%) and the wrist-worn device (3.48%). Most importantly, accurate assessments were automatically detected by using a custom threshold based on the device jerk. Due to the pervasiveness and low cost of wearable motion sensors, this paper demonstrates the feasibility of providing opportunistic large-scale low-cost samples of resting heart rate. Javier Hernandez, Daniel McDuff, Karen S. Quigley, Pattie Maes, Rosalind W. Picard |
IEEE J. Biomed. Health Informatics | 5 |
| 2019 | Multimodal Ambulatory Sleep Detection Using LSTM Recurrent Neural NetworksabstractUnobtrusive and accurate ambulatory methods are needed to monitor long-term sleep patterns for improving health. Previously developed ambulatory sleep detection methods rely either in whole or in part on self-reported diary data as ground truth, which is a problem, since people often do not fill them out accurately. This paper presents an algorithm that uses multimodal data from smart-phones and wearable technologies to detect sleep/wake state and sleep onset/offset using a type of recurrent neural network with long-short-term memory (LSTM) cells for synthesizing temporal information. We collected 5580 days of multimodal data from 186 participants and compared the new method for sleep/wake classification and sleep onset/offset detection to, first, nontemporal machine learning methods and, second, a state-of-the-art actigraphy software. The new LSTM method achieved a sleep/wake classification accuracy of 96.5%, and sleep onset/offset detection F1 scores of 0.86 and 0.84, respectively, with mean absolute errors of 5.0 and 5.5 min, respectively, when compared with sleep/wake state and sleep onset/offset assessed using actigraphy and sleep diaries. The LSTM results were statistically superior to those from non-temporal machine learning algorithms and the actigraphy software. We show good generalization of the new algorithm by comparing participant-dependent and participant-independent models, and we show how to make the model nearly realtime with slightly reduced performance. Akane Sano, Weixuan 'Vincent' Chen, Daniel Lopez Martinez, Sara Taylor, Rosalind W. Picard |
IEEE J. Biomed. Health Informatics | 5 |
| 2018 | A Trip to the Moon: Personalized Animated Movies for Self-reflectionabstractSelf-tracking physiological and psychological data poses the challenge of presentation and interpretation. Insightful narratives for self-tracking data can motivate the user towards constructive self-reflection. One powerful form of narrative that engages audience across various culture and age groups is animated movies. We collected a week of self-reported mood and behavior data from each user and created in Unity a personalized animation based on their data. We evaluated the impact of their video in a randomized control trial with a non-personalized animated video as control. We found that personalized videos tend to be more emotionally engaging, encouraging greater and lengthier writing that indicated self-reflection about moods and behaviors, compared to non-personalized control videos. Fengjiao Peng, Veronica LaBelle, Emily Christen Yue, Rosalind W. Picard |
CHI | 4 |
| 2018 | Multi-task multiple kernel machines for personalized pain recognition from functional near-infrared spectroscopy brain signalsabstractCurrently there is no validated objective measure of pain. Recent neuroimaging studies have explored the feasibility of using functional near-infrared spectroscopy (fNIRS) to measure alterations in brain function in evoked and ongoing pain. In this study, we applied multi-task machine learning methods to derive a practical algorithm for pain detection derived from fNIRS signals in healthy volunteers exposed to a painful stimulus. Especially, we employed multi-task multiple kernel learning to account for the inter-subject variability in pain response. Our results support the use of fNIRS and machine learning techniques in developing objective pain detection, and also highlight the importance of adopting personalized analysis in the process. Daniel Lopez Martinez, Sarah C. Steele, Arielle J. Lee, David Borsook, Rosalind W. Picard |
ICPR | 6 |
| 2018 | CultureNet: A Deep Learning Approach for Engagement Intensity Estimation from Face Images of Children with AutismabstractMany children on autism spectrum have atypical behavioral expressions of engagement compared to their neu-rotypical peers. In this paper, we investigate the performance of deep learning models in the task of automated engagement estimation from face images of children with autism. Specifically, we use the video data of 30 children with different cultural backgrounds (Asia vs. Europe) recorded during a single session of a robot-assisted autism therapy. We perform a thorough evaluation of the proposed deep architectures for the target task, including within- and across-culture evaluations, as well as when using the child-independent and child-dependent settings. We also introduce a novel deep learning model, named CultureNet, which efficiently leverages the multi-cultural data when performing the adaptation of the proposed deep architecture to the target culture and child. We show that due to the highly heterogeneous nature of the image data of children with autism, the child-independent models lead to overall poor estimation of target engagement levels. On the other hand, when a small amount of data of target children is used to enhance the model learning, the estimation performance on the held-out data from those children increases significantly. This is the first time that the effects of individual and cultural differences in children with autism have empirically been studied in the context of deep learning performed directly from face images. Ognjen Rudovic, Yuria Utsumi, Jaeryoung Lee, Javier Hernandez, Eduardo Castelló Ferrer, Björn W. Schuller, Rosalind W. Picard |
IROS | 7 |
| 2017 | GIFGIF+: Collecting emotional animated GIFs with clustered multi-task learningabstractAnimated GIFs are widely used on the Internet to express emotions, but their automatic analysis is largely unexplored. Existing GIF datasets with emotion labels are too small for training contemporary machine learning models, so we propose a semi-automatic method to collect emotional animated GIFs from the Internet with the least amount of human labor. The method trains weak emotion recognizers on labeled data, and uses them to sort a large quantity of unlabeled GIFs. We found that by exploiting the clustered structure of emotions, the number of GIFs a labeler needs to check can be greatly reduced. Using the proposed method, a dataset called GIFGIF+ with 23,544 GIFs over 17 emotions was created, which provides a promising platform for affective computing research. Weixuan 'Vincent' Chen, Ognjen Rudovic, Rosalind W. Picard |
ACII | 3 |
| 2017 | Objective assessment of depressive symptoms with machine learning and wearable sensors dataabstractDepression is the major cause of years lived in disability world-wide; however, its diagnosis and tracking methods still rely mainly on assessing self-reported depressive symptoms, methods that originated more than fifty years ago. These methods, which usually involve filling out surveys or engaging in face-to-face interviews, provide limited accuracy and reliability and are costly to track and scale. In this paper, we develop and test the efficacy of machine learning techniques applied to objective data captured passively and continuously from E4 wearable wristbands and from sensors in an Android phone for predicting the Hamilton Depression Rating Scale (HDRS). Input data include electrodermal activity (EDA), sleep behavior, motion, phone-based communication, location changes, and phone usage patterns. We introduce our feature generation and transformation process, imputing missing clinical scores from self-reported measures, and predicting depression severity from continuous sensor measurements. While HDRS ranges between 0 and 52, we were able to impute it with 2.8 RMSE and predict it with 4.5 RMSE which are low relative errors. Analyzing the features and their relation to depressive symptoms, we found that poor mental health was accompanied by more irregular sleep, less motion, fewer incoming messages, less variability in location patterns, and higher asymmetry of EDA between the right and the left wrists. Asma Ghandeharioun, Szymon Fedor, Lisa Sangermano, Dawn Ionescu, Jonathan Alpert, Chelsea Dale, David A. Sontag, Rosalind W. Picard |
ACII | 8 |
| 2017 | Stress measurement from tongue color imagingabstractA growing number of studies show links between changes in tongue appearance and human health conditions. This paper studies tongue color changes in the context of stress to explore the feasibility of providing a novel and non-invasive stress measurement method. In a laboratory study, 24 participants were asked to perform a calm and a stressful math task and to take a photo of their tongue right after each of the tasks. We observed subtle but consistent color differences between calm and stress tasks for up to 75% of the participants, which was consistent with both self-report and physiological metrics of stress. Moreover, we observed significant correlations of up to 0.72 between certain tongue colors and long-term stress assessed with the 10-item Perceived Stress Scale questionnaire. We discuss the potential implications of this work and highlight some lines of future research. Javier Hernandez, Craig Ferguson, Akane Sano, Weixuan 'Vincent' Chen, Weihui Li, Albert S. Yeung, Rosalind W. Picard |
ACII | 7 |
| 2017 | Multimodal autoencoder: A deep learning approach to filling in missing sensor data and enabling better mood predictionabstractTo accomplish forecasting of mood in real-world situations, affective computing systems need to collect and learn from multimodal data collected over weeks or months of daily use. Such systems are likely to encounter frequent data loss, e.g. when a phone loses location access, or when a sensor is recharging. Lost data can handicap classifiers trained with all modalities present in the data. This paper describes a new technique for handling missing multimodal data using a specialized denoising autoencoder: the Multimodal Autoencoder (MMAE). Empirical results from over 200 participants and 5500 days of data demonstrate that the MMAE is able to predict the feature values from multiple missing modalities more accurately than reconstruction methods such as principal components analysis (PCA). We discuss several practical benefits of the MMAE's encoding and show that it can provide robust mood prediction even when up to three quarters of the data sources are lost. Natasha Jaques, Sara Taylor, Akane Sano, Rosalind W. Picard |
ACII | 4 |
| 2017 | SPRING: Customizable, Motivation-Driven Technology for Children with Autism or Neurodevelopmental DifferencesabstractCurrent research to understand and enhance the development of children with neurological differences, including Autism Spectrum Disorder (ASD), is often severely limited by small sample sizes of human-gathered data in artificially structured learning environments. SPRING: Smart Platform for Research, Intervention, and Neurodevelopmental Growth is a new hardware and software system designed to 1) automate quantitative data acquisition, 2) optimize learning progressions through customized, motivating stimuli, and 3) encourage social, cognitive, and motor development in a personalized, child-led play environment. SPRING can also be paired with sensors to probe the physiological underpinnings of motivation, engagement, and cognition. Here, we present the design principles and methodology for SPRING, as well as two heterogeneous case studies. The first case highlights enhanced attention and accelerated skill development using SPRING, while the second pairs SPRING data with electrodermal activity measurements to identify a possible physiological signature of engagement and challenge in learning. Kristina T. Johnson, Rosalind W. Picard |
IDC | 2 |
| 2017 | Eliminating Physiological Information from Facial VideosabstractVital signs, cognitive load, and stress can be remotely measured from human faces using video-capturing devices under ambient light, which raises both wide applications and privacy issues. To avoid immoral use of this technology, there is a need for methods to eliminate physiological information from facial videos without affecting their visual appearance. To meet the need, we develop a novel algorithm based on motion component magnification that inputs a video and outputs its replica with physiological signals removed. Facial video data has been collected from 18 participants in a study to assess the performance of our algorithm in thwarting heart rate measurement based on remote photoplethysmography. Our results show that the mean absolute error of heart rate measurement averaged among participants was increased from 0.254 beats per minute to above 17 beats per minute without causing visible artifact. This is the first demonstration of an algorithm that can achieve this kind of functionality. Weixuan 'Vincent' Chen, Rosalind W. Picard |
FG | 2 |
| 2016 | COGCAM: Contact-free Measurement of Cognitive Stress During Computer Tasks with a Digital CameraabstractContact-free camera-based measurement of cognitive stress opens up new possibilities for human-computer interaction with applications in remote learning, stress monitoring, and optimization of workload for user experience. The autonomic nervous system controls the inter-beat intervals of the heart and breathing patterns, and these signals change under cognitive stress. We built a participant-independent cognitive stress recognition model based on photoplethysmographic signals measured remotely at a distance of 3 meters. We tested the model on naturalistic responses from 10 individuals completing randomized-order computer-based tasks (ball control and card sorting). The system successfully detected increased stress during the tasks, which were consistent with self-report measures. Changes in heart rate variability were more discriminative indicators of cognitive stress than were heart rate and breathing rate. Daniel McDuff, Javier Hernandez, Sarah Gontarek, Rosalind W. Picard |
CHI | 4 |
| 2016 | Emotion technology, wearables, and surprisesabstractCould we help people have healthier lives and better experiences if computers could measure and help communicate our emotion? Years ago, my students at MIT and I began to design, build, and test both wearable and other sensors for recognizing emotion. We designed studies, gathered data, and developed signal processing and machine learning techniques to see what could be reliably extracted. In this talk I will highlight several of the most surprising findings during this adventure. These include new insights about the "true smile of happiness," discovering that regular cameras (and your smartphone, even in your handbag) can compute some of your biosignals, finding electrical signals on the wrist that give insight into deep brain activity, and learning surprising implications of wearable sensing for autism, anxiety, depression, sleep-memory consolidation, epilepsy, and more. Rosalind W. Picard |
UbiComp | 1 |
| 2016 | Predicting Perceived Emotions in Animated GIFs with 3D Convolutional Neural NetworksabstractAnimated GIFs are widely used on the Internet to express emotions, but their automatic analysis is largely unexplored before. To help with the search and recommendation of GIFs, we aim to predict their emotions perceived by humans based on their contents. Since previous solutions to this problem only utilize image-based features and lose all the motion information, we propose to use 3D convolutional neural networks (CNNs) to extract spatiotemporal features from GIFs. We evaluate our methodology on a crowd-sourcing platform called GIFGIF with more than 6000 animated GIFs, and achieve a better accuracy then any previous approach in predicting crowd-sourced intensity scores of 17 emotions. It is also found that our trained model can be used to distinguish and cluster emotions in terms of valence and risk perception. Weixuan 'Vincent' Chen, Rosalind W. Picard |
ISM | 2 |
| 2016 | Personality, Attitudes, and Bonding in Conversations
Natasha Jaques, Yoo Lim Kim, Rosalind W. Picard |
IVA | 3 |
| 2016 | Understanding and Predicting Bonding in Conversations Using Thin Slices of Facial Expressions and Body Language
Natasha Jaques, Daniel McDuff, Yoo Lim Kim, Rosalind W. Picard |
IVA | 4 |
| 2016 | Wearable ESM: differences in the experience sampling method across wearable devicesabstractThe Experience Sampling Method is widely used for collecting self-report responses from people in natural settings. While most traditional approaches rely on using a phone to trigger prompts and record information, wearable devices now offer new opportunities that may improve this method. This research quantitatively and qualitatively studies the experience sampling process on head-worn and wrist-worn wearable devices, and compares them to the traditional "smartphone in the pocket." To enable this work, we designed and implemented a custom application to provide similar prompts across the three types of devices and evaluated it with 15 individuals for five days (75 days total), in the context of real-life stress measurement. We found significant differences in response times across devices, and captured tradeoffs in interaction types, screen size, and device familiarity that can affect both users' experience and the reports made by users. Javier Hernandez, Daniel McDuff, Christian Infante, Pattie Maes, Karen S. Quigley, Rosalind W. Picard |
MobileHCI | 6 |
| 2015 | Exploring temporal patterns in classifying frustrated and delighted smiles (Extended abstract)abstractWe created two experimental situations to elicit two affective states: frustration and delight. In the first experiment, participants were asked to recall situations while expressing either frustration or delight. The second experiment tried to elicit these states naturally with a frustrating experience and a delightful video. There were two significant differences between the acted and natural occurrences of the expressions. First, the acted instances were much easier for the computer to classify. Second, in 90 percent of the acted cases, participants did not smile when frustrated. In 90 percent of the natural cases, participants smiled during the frustrating interaction, despite self-reporting significant frustration with the experience. As a follow up study, we develop an automated system to distinguish between naturally occurring spontaneous smiles under frustrating and delightful stimuli by exploring their temporal patterns, given video of both. We extracted local and global features related to human smile dynamics. Next, we evaluated and compared two variants of Support Vector Machines (SVM), Hidden Markov Models (HMM), and Hidden-state Conditional Random Fields (HCRF) for binary classification. While human classification of the smile videos under frustrating stimuli was below chance, a dynamic SVM classifier obtained an accuracy of 92 percent in distinguishing smiles under frustrating and delighted stimuli. Mohammed E. Hoque 0001, Daniel McDuff, Rosalind W. Picard |
ACII | 3 |
| 2015 | Predicting students' happiness from physiology, phone, mobility, and behavioral dataabstractIn order to model students' happiness, we apply machine learning methods to data collected from undergrad students monitored over the course of one month each. The data collected include physiological signals, location, smartphone logs, and survey responses to behavioral questions. Each day, participants reported their wellbeing on measures including stress, health, and happiness. Because of the relationship between happiness and depression, modeling happiness may help us to detect individuals who are at risk of depression and guide interventions to help them. We are also interested in how behavioral factors (such as sleep and social activity) affect happiness positively and negatively. A variety of machine learning and feature selection techniques are compared, including Gaussian Mixture Models and ensemble classification. We achieve 70% classification accuracy of self-reported happiness on held-out test data. Natasha Jaques, Sara Taylor, Asaph Azaria, Asma Ghandeharioun, Akane Sano, Rosalind W. Picard |
ACII | 6 |
| 2015 | Crowdsourcing facial responses to online videos: Extended abstractabstractTraditional observational research methods required an experimenter's presence in order to record videos of participants, and limited the scalability of data collection to typically less than a few hundred people in a single location. In order to make a significant leap forward in affective expression data collection and the insights based on it, our work has created and validated a novel framework for collecting and analyzing facial responses over the Internet. The first experiment using this framework enabled 3,268 trackable face videos to be collected and analyzed in under two months. Each participant viewed one or more commercials while their facial response was recorded and analyzed. Our data showed significantly different intensity and dynamics patterns of smile responses between subgroups who reported liking the commercials versus those who did not. Since this framework appeared in 2011, we have collected over three million videos of facial responses in over 75 countries using this same methodology, enabling facial analytics to become significantly more accurate and validated across five continents. Many new insights have been discovered based on crowd-sourced facial data, enabling Internet-based measurement of facial responses to become reliable and proven. We are now able to provide large-scale evidence for gender, cultural and age differences in behaviors. Today such methods are used as part of standard practice in industry for copy-testing advertisements and are increasingly used for online media evaluations, distance learning, and mobile applications. Daniel McDuff, Rana El Kaliouby, Rosalind W. Picard |
ACII | 3 |
| 2015 | BioInsights: Extracting personal data from "Still" wearable motion sensorsabstractDuring recent years a large variety of wearable devices have become commercially available. As these devices are in close contact with the body, they have the potential to capture sensitive and unexpected personal data even when the wearer is not moving. This work demonstrates that wearable motion sensors such as accelerometers and gyroscopes embedded in head-mounted and wrist-worn wearable devices can be used to identify the wearer (among 12 participants) and his/her body posture (among 3 positions) from only 10 seconds of “still” motion data. Instead of focusing on large and apparent motions such as steps or gait, the proposed methods amplify and analyze very subtle body motions associated with the beating of the heart. Our findings have the potential to increase the value of pervasive wearable motion sensors but also raise important privacy concerns that need to be considered. Javier Hernandez, Daniel McDuff, Rosalind W. Picard |
BSN | 3 |
| 2015 | Recognizing academic performance, sleep quality, stress level, and mental health using personality traits, wearable sensors and mobile phonesabstractWhat can wearable sensors and usage of smart phones tell us about academic performance, self-reported sleep quality, stress and mental health condition? To answer this question, we collected extensive subjective and objective data using mobile phones, surveys, and wearable sensors worn day and night from 66 participants, for 30 days each, totaling 1,980 days of data. We analyzed daily and monthly behavioral and physiological patterns and identified factors that affect academic performance (GPA), Pittsburg Sleep Quality Index (PSQI) score, perceived stress scale (PSS), and mental health composite score (MCS) from SF-12, using these month-long data. We also examined how accurately the collected data classified the participants into groups of high/low GPA, good/poor sleep quality, high/low self-reported stress, high/low MCS using feature selection and machine learning techniques. We found associations among PSQI, PSS, MCS, and GPA and personality types. Classification accuracies using the objective data from wearable sensors and mobile phones ranged from 67-92%. Akane Sano, Andrew J. K. Phillips, Amy Z. Yu, Andrew W. McHill, Sara Taylor, Natasha Jaques, Charles A. Czeisler, Elizabeth B. Klerman, Rosalind W. Picard |
BSN | 9 |
| 2015 | Common Sense Reasoning for Detection, Prevention, and Mitigation of Cyberbullying (Extended Abstract)
Karthik Dinakar, Rosalind W. Picard, Henry Lieberman |
IJCAI | 2 |
| 2015 | Mixed-Initiative Real-Time Topic Modeling & Visualization for Crisis CounselingabstractText-based counseling and support systems have seen an increasing proliferation in the past decade. We present Fathom, a natural language interface to help crisis counselors on Crisis Text Line, a new 911-like crisis hotline that takes calls via text messaging rather than voice. Text messaging opens up the opportunity for software to read the messages as well as people, and to provide assistance for human counselors who give clients emotional and practical support. Crisis counseling is a tough job that requires dealing with emotionally stressed people in possibly life-critical situations, under time constraints. Fathom is a system that provides topic modeling of calls and graphical visualization of topic distributions, updated in real time. We develop a mixed-initiative paradigm to train coherent topic and word distributions and use them to power real-time visualizations aimed at reducing counselor cognitive overload. We believe Fathom to be the first real-time computational framework to assist in crisis counseling. Karthik Dinakar, Jackie Chen, Henry Lieberman, Rosalind W. Picard, Robert Filbin |
IUI | 4 |
| 2015 | Recognizing Stress, Engagement, and Positive EmotionabstractAn intelligent interaction should not typically call attention to emotion. However, it almost always involves emotion: For example, it should engage, not inflict undesirable stress and frustration, and perhaps elicit positive emotions such as joy or delight. How would the system sense or recognize if it was succeeding in these elements of intelligent interaction? This keynote talk will address some ways that our work at the MIT Media Lab has advanced solutions for recognizing user emotion during everyday experiences. Rosalind W. Picard |
IUI | 1 |
| 2015 | Predicting Ad Liking and Purchase Intent: Large-Scale Analysis of Facial Responses to AdsabstractBillions of online video ads are viewed every month. We present a large-scale analysis of facial responses to video content measured over the Internet and their relationship to marketing effectiveness. We collected over 12,000 facial responses from 1,223 people to 170 ads from a range of markets and product categories. The facial responses were automatically coded frame-by-frame. Collection and coding of these 3.7 million frames would not have been feasible with traditional research methods. We show that detected expressions are sparse but that aggregate responses reveal rich emotion trajectories. By modeling the relationship between the facial responses and ad effectiveness, we show that ad liking can be predicted accurately (ROC AUC = 0.85) from webcam facial responses. Furthermore, the prediction of a change in purchase intent is possible (ROC AUC = 0.78). Ad liking is shown by eliciting expressions, particularly positive expressions. Driving purchase intent is more complex than just making viewers smile: peak positive responses that are immediately preceded by a brand appearance are more likely to be effective. The results presented here demonstrate a reliable and generalizable system for predicting ad effectiveness automatically from facial responses without a need to elicit self-report responses from the viewers. In addition we can gain insight into the structure of effective ads. Daniel McDuff, Rana El Kaliouby, Jeffrey F. Cohn, Rosalind W. Picard |
IEEE Trans. Affect. Comput. | 4 |
| 2015 | Guest Editorial Sensor Informatics and Quantified SelfabstractThe articles in this special issue focus on new technologies and applications for medical services that incorporate wearable sensors, signal processing, machine learning, and data mining techniques. The ability to collect large sets of human data comfortably 24/7, are advancing new ways to learn about human well being. Measurements that used to be confined to short-term sampling in a lab or medical facility are now able to be conducted continuously, while at home, work, sleep, or play. Studies are no longer limited to a focus on disease progression or to the effect of therapeutic measures provided in clinical settings—instead, it is becoming possible to quantify healthy activities and behavior, and capture how these slowly change as illness develops or progresses. The quantified self movement, where people can monitor their own health and fitness-related data, is closely linked to the emergence of new methods in biometric sensing. Rosalind W. Picard, Gary Wolf |
IEEE J. Biomed. Health Informatics | 1 |
| 2014 | Modeling Subjective Experience-Based Learning under Uncertainty and FramesabstractIn this paper we computationally examine how subjective experience may help or harm the decision maker's learning under uncertain outcomes, frames and their interactions. To model subjective experience, we propose the "experienced-utility function" based on a prospect theory (PT)-based parameterized subjective value function. Our analysis and simulations of two-armed bandit tasks present that the task domain (underlying outcome distributions) and framing (reference point selection) influence experienced utilities and in turn, the "subjective discriminability" of choices under uncertainty. Experiments demonstrate that subjective discriminability improves on objective discriminability by the use of the experienced-utility function with appropriate framing for a given task domain, and that bigger subjective discriminability leads to more optimal decisions in learning under uncertainty. Hyungil Ahn, Rosalind W. Picard |
AAAI | 2 |
| 2014 | Using electrodermal activity to recognize ease of engagement in children during social interactionsabstractThe recent emergence of comfortable wearable sensors has focused almost entirely on monitoring physical activity, ignoring opportunities to monitor more subtle phenomena, such as the quality of social interactions. We argue that it is compelling to address whether physiological sensors can shed light on quality of social interactive behavior. This work leverages the use of a wearable electrodermal activity (EDA) sensor to recognize ease of engagement of children during a social interaction with an adult. In particular, we monitored 51 child-adult dyads in a semi-structured play interaction and used Support Vector Machines to automatically identify children who had been rated by the adult as more or less difficult to engage. We report on the classification value of several features extracted from the child's EDA responses, as well as several other features capturing the physiological synchrony between the child and the adult. Javier Hernandez, Ivan Riobo, Agata Rozga, Gregory D. Abowd, Rosalind W. Picard |
UbiComp | 5 |
| 2014 | Affective media and wearables: surprising findingsabstractOver a decade ago, I suggested that computers will need the skills of emotional intelligence in order to interact with regular people in ways that they perceive as intelligent. Our lab embarked on this journey of 'affective computing' with a focus on first enabling computers to better understand and communicate human emotion. Our main tools have been wearable sensors (several which we created), video, and audio, coupled with signal processing, machine learning and pattern analysis of multimodal human data. Along the way we encountered several surprises. This talk will highlight some of the challenges we have faced, some accomplishments, and the most surprising and rewarding findings. Our findings reveal the power of the human emotion system not only in intelligence, in social interaction, and in everyday media consumption, but also in autism, epilepsy, and sleep memory formation. Rosalind W. Picard |
ACM Multimedia | 1 |
| 2014 | Automatic measurement of ad preferences from facial responses gathered over the Internet
Daniel McDuff, Rana El Kaliouby, Thibaud Senechal, David Demirdjian, Rosalind W. Picard |
Image Vis. Comput. | 5 |
| 2014 | Measuring Affective-Cognitive Experience and Predicting Market SuccessabstractWe present a new affective-behavioral-cognitive (ABC) framework to measure the usual cognitive self-report information and behavioral information, together with affective information while a customer makes repeated selections in a random-outcome two-option decision task to obtain their preferred product. The affective information consists of human-labeled facial expression valence taken from two contexts: one where the facial valence is associated with affective wanting, and the other with affective liking. The new “affective wanting” measure is made by setting up a condition where the person shows desire to receive one of two products, and we measure if the face looks satisfied or disappointed when each of the products arrives. The “affective liking” measure captures facial expressions after sampling a product. The ABC framework is tested in a real-world beverage taste experiment, comparing two similar products that actually went to market, where we know the market outcomes. We find that the affective measure provides significant improvement over the cognitive measure, increasing the discriminability between the two similar products, making it easier to tell which is most preferred using a small number of people. We also find that the new facial valence “affective wanting” measure provides a significant boost in discrimination and accuracy. Hyung-Il Ahn, Rosalind W. Picard |
IEEE Trans. Affect. Comput. | 2 |
| 2014 | FEEL: A System for Frequent Event and Electrodermal Activity LabelingabstractThe wide availability of low-cost wearable biophysiological sensors enables us to measure how the environment and our experiences impact our physiology. This creates a challenge: in order to interpret the longitudinal data, we require the matching contextual information as well. Collecting continuous biophysiological data makes it unfeasible to rely solely on our memory for contextual information. In this paper, we first present an architecture and implementation of a system for the acquisition, processing, and visualization of biophysiological signals and contextual information. Next, we present the results of a user study: users wore electrodermal activity wrist sensors that measured their autonomic arousal. These users uploaded the sensor data at the end of each day. At first, they annotated their events at the end of each day; then, after a two-day break, they annotated the data from two days earlier. One group of users had access to both the signal and the contextual information collected by the mobile phone and the other group could only access the biophysiological signal. At the end of the study, the users filled in a system usability scale and user experience surveys. Our results show that the system enables the users to annotate biophysiological signals at a greater effectiveness than the current state of the art while also providing very good usability. Yadid Ayzenberg, Rosalind W. Picard |
IEEE J. Biomed. Health Informatics | 2 |
| 2013 | Automated Coach to Practice ConversationsabstractWe present a real-time system including a 3D character that can converse, capture, analyze and interpret subtle and multidimensional human nonverbal behaviors for possible applications such as job interviews, public speaking, or even automated speech therapy. The system works in a personal computer and senses nonverbal data from video (i.e., facial expressions) and audio (i.e., speech recognition and prosody analysis) using a standard web cam. We contextualized the development and evaluation of our system as a training scenario for job interviews. Using user-centered design and iterations, we determine how the nonverbal data could be presented to the user in an intuitive and educational manner. We tested efficacy of the system in context of job interviews with 90 MIT undergraduate students. Our results suggest that the participants who used our system to improve their interview skills were perceived to be better candidates by human judges. Participants reported that the most useful feature was being given feedback on their speaking rate, and overall they reported strong agreement that would consider using this system again for self-reflection. Mohammed E. Hoque 0001, Rosalind W. Picard |
ACII | 2 |
| 2013 | Measuring Voter's Candidate Preference Based on Affective Responses to Election DebatesabstractIn this paper we present the first analysis of facial responses to electoral debates measured automatically over the Internet. We show that significantly different responses can be detected from viewers with different political preferences and that similar expressions at significant moments can have very different meanings depending on the actions that appear subsequently. We used an Internet based framework to collect 611 naturalistic and spontaneous facial responses to five video clips from the 3rd presidential debate during the 2012 American presidential election campaign. Using this framework we were able to collect over 60% of these video responses (374 videos) within one day of the live debate and over 80% within three days. No participants were compensated for taking the survey. We present and evaluate a method for predicting independent voter preference based on automatically measured facial responses and self-reported preferences from the viewers. We predict voter preference with an average accuracy of over 73% (AUC 0.779). Daniel McDuff, Rana El Kaliouby, Evan Kodra, Rosalind W. Picard |
ACII | 4 |
| 2013 | Stress Recognition Using Wearable Sensors and Mobile PhonesabstractIn this study, we aim to find physiological or behavioral markers for stress. We collected 5 days of data for 18 participants: a wrist sensor (accelerometer and skin conductance), mobile phone usage (call, short message service, location and screen on/off) and surveys (stress, mood, sleep, tiredness, general health, alcohol or caffeinated beverage intake and electronics usage). We applied correlation analysis to find statistically significant features associated with stress and used machine learning to classify whether the participants were stressed or not. In comparison to a baseline 87.5% accuracy using the surveys, our results showed over 75% accuracy in a binary classification using screen on, mobility, call or activity level information (some showed higher accuracy than the baseline). The correlation analysis showed that the higher-reported stress level was related to activity level, SMS and screen on/off patterns. Akane Sano, Rosalind W. Picard |
ACII | 2 |
| 2013 | Recognition of sleep dependent memory consolidation with multi-modal sensor dataabstractThis paper presents the possibility of recognizing sleep dependent memory consolidation using multi-modal sensor data. We collected visual discrimination task (VDT) performance before and after sleep at laboratory, hospital and home for N=24 participants while recording EEG (electroencepharogram), EDA (electrodermal activity) and ACC (accelerometer) or actigraphy data during sleep. We extracted features and applied machine learning techniques (discriminant analysis, support vector machine and k-nearest neighbor) from the sleep data to classify whether the participants showed improvement in the memory task. Our results showed 60–70% accuracy in a binary classification of task performance using EDA or EDA+ACC features, which provided an improvement over the more traditional use of sleep stages (the percentages of slow wave sleep (SWS) in the 1stquarter and rapid eye movement (REM) in the 4th quarter of the night) to predict VDT improvement. Akane Sano, Rosalind W. Picard |
BSN | 2 |
| 2013 | MACH: my automated conversation coachabstractMACH--My Automated Conversation coacH--is a novel system that provides ubiquitous access to social skills training. The system includes a virtual agent that reads facial expressions, speech, and prosody and responds with verbal and nonverbal behaviors in real time. This paper presents an application of MACH in the context of training for job interviews. During the training, MACH asks interview questions, automatically mimics certain behavior issued by the user, and exhibit appropriate nonverbal behaviors. Following the interaction, MACH provides visual feedback on the user's performance. The development of this application draws on data from 28 interview sessions, involving employment-seeking students and career counselors. The effectiveness of MACH was assessed through a weeklong trial with 90 MIT undergraduates. Students who interacted with MACH were rated by human experts to have improved in overall interview performance, while the ratings of students in control groups did not improve. Post-experiment interviews indicate that participants found the interview experience informative about their behaviors and expressed interest in using MACH in the future. Mohammed E. Hoque 0001, Matthieu Courgeon, Jean-Claude Martin, Bilge Mutlu, Rosalind W. Picard |
UbiComp | 5 |
| 2013 | Surprising discoveries from emotion sensors
Rosalind W. Picard |
SEKE | 1 |
| 2012 | Mood meter: counting smiles in the wildabstractIn this study, we created and evaluated a computer vision based system that automatically encouraged, recognized and counted smiles on a college campus. During a ten-week installation, passersby were able to interact with the system at four public locations. The aggregated data was displayed in real time in various intuitive and interactive formats on a public website. We found privacy to be one of the main design constraints, and transparency to be the best strategy to gain participants' acceptance. In a survey (with 300 responses), participants reported that the system made them smile more than they expected, and it made them and others around them feel momentarily better. Quantitative analysis of the interactions revealed periodic patterns (e.g., more smiles during the weekends) and strong correlation with campus events (e.g., fewer smiles during exams, most smiles the day after graduation), reflecting the emotional responses of a large community. Javier Hernandez, Mohammed E. Hoque 0001, Will Drevo, Rosalind W. Picard |
UbiComp | 4 |
| 2012 | Multimodal annotation tool for challenging behaviors in people with Autism spectrum disordersabstractIndividuals diagnosed with Autism Spectrum Disorders (ASD) often have challenging behaviors (CB's), such as self-injury or emotional outbursts, which can negatively impact the quality of life of themselves and those around them. Recent advances in mobile and ubiquitous technologies provide an opportunity to efficiently and accurately capture important information preceding and associated with these CB's. The ability to obtain this type of data will help with both intervention and behavioral phenotyping efforts. Through collaboration with behavioral scientists and therapists, we identified relevant design requirements and created an easy-to-use mobile application for collecting, labeling, and sharing in-situ behavior data in individuals diagnosed with ASD. Furthermore, we have released the application to the community as an open-source project so it can be validated and extended by other researchers. Akane Sano, Javier Hernandez, Jean Deprey, Micah Eckhardt, Matthew S. Goodwin, Rosalind W. Picard |
UbiComp | 6 |
| 2012 | You Too?! Mixed-Initiative LDA Story Matching to Help Teens in Distress
Karthik Dinakar, Birago Jones, Henry Lieberman, Rosalind W. Picard, Carolyn P. Rosé, Matthew Thoman, Roi Reichart |
ICWSM | 4 |
| 2012 | Exploring Temporal Patterns in Classifying Frustrated and Delighted SmilesabstractWe create two experimental situations to elicit two affective states: frustration, and delight. In the first experiment, participants were asked to recall situations while expressing either delight or frustration, while the second experiment tried to elicit these states naturally through a frustrating experience and through a delightful video. There were two significant differences in the nature of the acted versus natural occurrences of expressions. First, the acted instances were much easier for the computer to classify. Second, in 90 percent of the acted cases, participants did not smile when frustrated, whereas in 90 percent of the natural cases, participants smiled during the frustrating interaction, despite self-reporting significant frustration with the experience. As a follow up study, we develop an automated system to distinguish between naturally occurring spontaneous smiles under frustrating and delightful stimuli by exploring their temporal patterns given video of both. We extracted local and global features related to human smile dynamics. Next, we evaluated and compared two variants of Support Vector Machine (SVM), Hidden Markov Models (HMM), and Hidden-state Conditional Random Fields (HCRF) for binary classification. While human classification of the smile videos under frustrating stimuli was below chance, an accuracy of 92 percent distinguishing smiles under frustrating and delighted stimuli was obtained using a dynamic SVM classifier. Mohammed E. Hoque 0001, Daniel McDuff, Rosalind W. Picard |
IEEE Trans. Affect. Comput. | 3 |
| 2012 | Crowdsourcing Facial Responses to Online VideosabstractWe present results validating a novel framework for collecting and analyzing facial responses to media content over the Internet. This system allowed 3,268 trackable face videos to be collected and analyzed in under two months. We characterize the data and present analysis of the smile responses of viewers to three commercials. We compare statistics from this corpus to those from the Cohn-Kanade+ (CK+) and MMI databases and show that distributions of position, scale, pose, movement, and luminance of the facial region are significantly different from those represented in these traditionally used datasets. Next, we analyze the intensity and dynamics of smile responses, and show that there are significantly different facial responses from subgroups who report liking the commercials compared to those that report not liking the commercials. Similarly, we unveil significant differences between groups who were previously familiar with a commercial and those that were not and propose a link to virality. Finally, we present relationships between head movement and facial behavior that were observed within the data. The framework, data collected, and analysis demonstrate an ecologically valid method for unobtrusive evaluation of facial responses to media content that is robust to challenging real-world conditions and requires no explicit recruitment or compensation of participants. Daniel McDuff, Rana El Kaliouby, Rosalind W. Picard |
IEEE Trans. Affect. Comput. | 3 |
| 2012 | Common Sense Reasoning for Detection, Prevention, and Mitigation of CyberbullyingabstractCyberbullying (harassment on social networks) is widely recognized as a serious social problem, especially for adolescents. It is as much a threat to the viability of online social networks for youth today as spam once was to email in the early days of the Internet. Current work to tackle this problem has involved social and psychological studies on its prevalence as well as its negative effects on adolescents. While true solutions rest on teaching youth to have healthy personal relationships, few have considered innovative design of social network software as a tool for mitigating this problem. Mitigating cyberbullying involves two key components: robust techniques for effective detection and reflective user interfaces that encourage users to reflect upon their behavior and their choices. Spam filters have been successful by applying statistical approaches like Bayesian networks and hidden Markov models. They can, like Google’s GMail, aggregate human spam judgments because spam is sent nearly identically to many people. Bullying is more personalized, varied, and contextual. In this work, we present an approach for bullying detection based on state-of-the-art natural language processing and a common sense knowledge base, which permits recognition over a broad spectrum of topics in everyday life. We analyze a more narrow range of particular subject matter associated with bullying (e.g. appearance, intelligence, racial and ethnic slurs, social acceptance, and rejection), and construct BullySpace , a common sense knowledge base that encodes particular knowledge about bullying situations. We then perform joint reasoning with common sense knowledge about a wide range of everyday life topics. We analyze messages using our novel AnalogySpace common sense reasoning technique. We also take into account social network analysis and other factors. We evaluate the model on real-world instances that have been reported by users on Formspring, a social networking website that is popular with teenagers. On the intervention side, we explore a set of reflective user-interaction paradigms with the goal of promoting empathy among social network participants. We propose an “air traffic control”-like dashboard, which alerts moderators to large-scale outbreaks that appear to be escalating or spreading and helps them prioritize the current deluge of user complaints. For potential victims, we provide educational material that informs them about how to cope with the situation, and connects them with emotional support from others. A user evaluation shows that in-context, targeted, and dynamic help during cyberbullying situations fosters end-user reflection that promotes better coping strategies. Karthik Dinakar, Birago Jones, Catherine Havasi, Henry Lieberman, Rosalind W. Picard |
ACM Trans. Interact. Intell. Syst. | 5 |
| 2011 | Call Center Stress Recognition with Person-Specific Models
Javier Hernandez, Robert R. Morris, Rosalind W. Picard |
ACII (1) | 3 |
| 2011 | Machine Learning for Affective Computing
Mohammed E. Hoque 0001, Daniel McDuff, Louis-Philippe Morency, Rosalind W. Picard |
ACII (2) | 4 |
| 2011 | Are You Friendly or Just Polite? - Analysis of Smiles in Spontaneous Face-to-Face Interactions
Mohammed E. Hoque 0001, Louis-Philippe Morency, Rosalind W. Picard |
ACII (1) | 3 |
| 2011 | Measuring Affect in the Wild
Rosalind W. Picard |
ACII (1) | 1 |
| 2011 | Real-time inference of mental states from facial expressions and upper body gesturesabstractWe present a real-time system for detecting facial action units and inferring emotional states from head and shoulder gestures and facial expressions. The dynamic system uses three levels of inference on progressively longer time scales. Firstly, facial action units and head orientation are identified from 22 feature points and Gabor filters. Secondly, Hidden Markov Models are used to classify sequences of actions into head and shoulder gestures. Finally, a multi level Dynamic Bayesian Network is used to model the unfolding emotional state based on probabilities of different gestures. The most probable state over a given video clip is chosen as the label for that clip. The average F1 score for 12 action units (AUs 1, 2, 4, 6, 7, 10, 12, 15, 17, 18, 25, 26), labelled on a frame by frame basis, was 0.461. The average classification rate for five emotional states (anger, fear, joy, relief, sadness) was 0.440. Sadness had the greatest rate, 0.64, anger the smallest, 0.11. Tadas Baltrusaitis, Daniel McDuff, Ntombikayise Banda, Marwa Mahmoud, Rana El Kaliouby, Peter Robinson 0001, Rosalind W. Picard |
FG | 7 |
| 2011 | Acted vs. natural frustration and delight: Many people smile in natural frustrationabstractThis work is part of research to build a system to combine facial and prosodic information to recognize commonly occurring user states such as delight and frustration. We create two experimental situations to elicit two emotional states: the first involves recalling situations while expressing either delight or frustration; the second experiment tries to elicit these states directly through a frustrating experience and through a delightful video. We find two significant differences in the nature of the acted vs. natural occurrences of expressions. First, the acted ones are much easier for the computer to recognize. Second, in 90% of the acted cases, participants did not smile when frustrated, whereas in 90% of the natural cases, participants smiled during the frustrating interaction, despite self-reporting significant frustration with the experience. This paper begins to explore the differences in the patterns of smiling that are seen under natural frustration and delight conditions, to see if there might be something measurably different about the smiles in these two cases, which could ultimately improve the performance of classifiers applied to natural expressions. Mohammed E. Hoque 0001, Rosalind W. Picard |
FG | 2 |
| 2011 | Acume: A new visualization tool for understanding facial expression and gesture dataabstractFacial and head actions contain significant affective information. To date, these actions have mostly been studied in isolation because the space of naturalistic combinations is vast. Interactive visualization tools could enable new explorations of dynamically changing combinations of actions as people interact with natural stimuli. This paper describes a new open-source tool that enables navigation of and interaction with dynamic face and gesture data across large groups of people, making it easy to see when multiple facial actions co-occur, and how these patterns compare and cluster across groups of participants. We share two case studies that demonstrate how the tool allows researchers to quickly view an entire corpus of data for single or multiple participants, stimuli and actions. Acume yielded patterns of actions across participants and across stimuli, and helped give insight into how our automated facial analysis methods could be better designed. The results of these case studies are used to demonstrate the efficacy of the tool. The open-source code is designed to directly address the needs of the face and gesture research community, while also being extensible and flexible for accommodating other kinds of behavioral data. Source code, application and documentation are available at http://affect.media.mit.edu/acume. Daniel McDuff, Rana El Kaliouby, Karim Kassam, Rosalind W. Picard |
FG | 4 |
| 2011 | Crowdsourced data collection of facial responsesabstractIn the past, collecting data to train facial expression and affect recognition systems has been time consuming and often led to data that do not include spontaneous expressions. We present the first crowdsourced data collection of dynamic, natural and spontaneous facial responses as viewers watch media online. This system allowed a corpus of 3,268 videos to be collected in under two months. Daniel McDuff, Rana El Kaliouby, Rosalind W. Picard |
ICMI | 3 |
| 2011 | Recognizing affect from speech prosody using hierarchical graphical models
Raul Fernandez, Rosalind W. Picard |
Speech Commun. | 2 |
| 2010 | Broadening accessibility through special interests: a new approach for software customizationabstractIndividuals diagnosed with autism spectrum disorder (ASD) often fixate on narrow, restricted interests. These interests can be highly motivating, but they can also create attentional myopia, preventing individuals from pursuing a broad range of activities. Interestingly, researchers have found that preferred interests can be used to help individuals with ASD branch out and participate in educational, therapeutic, or social situations they might otherwise shun. When interventions are modified, such that an individual's interest is properly represented, task adherence and performance can increase. While this strategy has seen success in the research literature, it is difficult to implement on a large scale and therefore has not been widely adopted. This paper describes a software approach designed to solve this problem. The approach facilitates customization, allowing users to easily embed images of almost any special interest into computer-based interventions. Specifically, we describe an algorithm that will: (1) retrieve any image from the Google image database; (2) strip it of its background; and (3) embed it seamlessly into Flash-based computer programs. To evaluate our algorithm, we employed it in a naturalistic setting with eleven individuals (nine diagnosed with ASD and two diagnosed with other developmental disorders). We also tested its ability to retrieve and process examples of preferred interests previously reported in the ASD literature. The results indicate that our method was an easy and efficient way for users to customize our software programs. While we believe this model is uniquely suited for individuals with ASD, we also foresee this approach being useful for anyone that might like a quick and simple way to personalize software programs. Robert R. Morris, Connor R. Kirschbaum, Rosalind W. Picard |
ASSETS | 3 |
| 2010 | Mobile emotional intelligenceabstractThis keynote will highlight work from the Affective Computing group at the MIT Media Laboratory bringing new skills of emotional intelligence to mobile technology. Here are some of the things I plan to discuss and/or to demonstrate live: (1) New comfortable wearable technology for measuring physiological components of emotion, for example, I will show measurement of sympathetic nervous system activity using a new wrist-worn stretchy washable sensor and cardiovascular changes using regular earbuds plugged into your mobile phone; (2) Technology that reads facial expressions using a mobile device, which we are using to help people on the autism spectrum better learn about nonverbal communication; (3) Collection of affective information from mobile devices without being annoying, even when interrupting people over a dozen times a day; (4) Applications of these technologies to help people with sensory disorders, stress disorders, sleep disorders, and substance abuse. As time permits, I will also discuss a growing vision of how these new technologies can enable more research "by the people for the people. Rosalind W. Picard |
MobiSys | 1 |
| 2010 | Technology for Changing Feelings
Rosalind W. Picard |
PERSUASIVE | 1 |
| 2010 | iCalm: wearable sensor and network architecture for wirelessly communicating and logging autonomic activityabstractWidespread use of affective sensing in healthcare applications has been limited due to several practical factors, such as lack of comfortable wearable sensors, lack of wireless standards, and lack of low-power affordable hardware. In this paper, we present a new low-cost, low-power wireless sensor platform implemented using the IEEE 802.15.4 wireless standard, and describe the design of compact wearable sensors for long-term measurement of electrodermal activity, temperature, motor activity, and photoplethysmography. We also illustrate the use of this new technology for continuous long-term monitoring of autonomic nervous system and motion data from active infants, children, and adults. We describe several new applications enabled by this system, discuss two specific wearable designs for the wrist and foot, and present sample data. Rich Fletcher, Kelly Dobson, Matthew S. Goodwin, Hoda Eydgahi, Oliver Wilder-Smith, David Fernholz, Yuta Kuboyama, Elliott Bruce Hedman, Ming-Zher Poh, Rosalind W. Picard |
IEEE Trans. Inf. Technol. Biomed. | 10 |
| 2010 | Motion-tolerant magnetic earring sensor and wireless earpiece for wearable photoplethysmographyabstractThis paper addresses the design considerations and critical evaluation of a novel embodiment for wearable photoplethysmography (PPG) comprising a magnetic earring sensor and wireless earpiece. The miniaturized sensor can be worn comfortably on the earlobe and contains an embedded accelerometer to provide motion reference for adaptive noise cancellation. The compact wireless earpiece provides analog signal conditioning and acts as a data-forwarding device via a radio frequency transceiver. Using Bland-Altman and correlation analysis, we evaluated the performance of the proposed system against an FDA-approved ECG measurement device during daily activities. The mean +/- standard deviation (SD) of the differences between heart rate measurements from the proposed device and ECG (expressed as percentage of the average between the two techniques) along with the 95% limits of agreement (LOA = +/-1.96 SD) was 0.62% +/- 4.51% (LOA = -8.23% and 9.46%), -0.49% +/- 8.65% (-17.39% and 16.42%), and -0.32% +/- 10.63% (-21.15% and 20.52%) during standing, walking, and running, respectively. Linear regression indicated a high correlation between the two measurements across the three evaluated conditions (r = 0.97, 0.82, and 0.76, respectively with p < 0.001). The new earring PPG system provides a platform for comfortable, robust, unobtrusive, and discreet monitoring of cardiovascular function. Ming-Zher Poh, Nicholas C. Swenson, Rosalind W. Picard |
IEEE Trans. Inf. Technol. Biomed. | 3 |
| 2009 | Robots with emotional intelligenceabstractThis keynote talk will illustrate a basic set of skills of emotional intelligence, how they are important for robots and agents that interact with people, and how our research at MIT addresses part of the problem of giving robots such skills. One of the most important skills is the ability to perceive and understand expressions of emotion, which I will highlight by demonstrating our latest technologies developed to read joint facial-head movements in real-time and associate these with complex affective-cognitive states, and technologies to read paralinguistic vocal cues from speech. The latter have been made open-source and are available for free. I will also show some non-traditional ways robots might sense and learn about human emotion, and ways they can respond to what they sense that can help or hurt people. I will discuss social and ethical issues these technologies raise. Finally, I will present some new possibilities for robots to both learn from people and help teach skills of emotional intelligence to people, especially to those with nonverbal learning impairments who may want to learn these skills, including many people with diagnoses of autism spectrum disorders such as Aspergers Syndrome. Rosalind W. Picard |
HRI | 1 |
| 2009 | Exploring speech therapy games with children on the autism spectrumabstractIndividuals on the autism spectrum often have difficulties producing intelligible speech with either high or low speech rate, and atypical pitch and/or amplitude affect.In this study, we present a novel intervention towards customizing speech enabled games to help them produce intelligible speech.In this approach, we clinically and computationally identify the areas of speech production difficulties of our participants.We provide an interactive and customized interface for the participants to meaningfully manipulate the prosodic aspects of their speech.Over the course of 12 months, we have conducted several pilots to set up the experimental design, developed a suite of games and audio processing algorithms for prosodic analysis of speech.Preliminary results demonstrate our intervention being engaging and effective for our participants. Mohammed E. Hoque 0001, Joseph K. Lane, Rana El Kaliouby, Matthew S. Goodwin, Rosalind W. Picard |
INTERSPEECH | 5 |
| 2009 | When Human Coders (and Machines) Disagree on the Meaning of Facial Affect in Spontaneous Videos
Mohammed E. Hoque 0001, Rana El Kaliouby, Rosalind W. Picard |
IVA | 3 |
| 2008 | Technology for just-in-time in-situ learning of facial affect for persons diagnosed with an autism spectrum disorderabstractMany first-hand accounts from individuals diagnosed with autism spectrum disorders (ASD) highlight the challenges inherent in processing high-speed, complex, and unpredictable social information such as facial expressions in real-time. In this paper, we describe a new technology aimed at helping people capture, analyze, and reflect on a set of social-emotional signals communicated by facial and head movements in live social interaction that occurs with their everyday social companions. We describe our development of a new combination of hardware using a miniature camera connected to an ultramobile PC together with custom software developed to track, capture, interpret, and intuitively present various interpretations of the facial-head movements (e.g., presenting that there is a high probability the person looks "confused"). This paper describes this new technology together with the results of a series of pilot studies conducted with adolescents diagnosed with ASD who used the technology in their peer-group setting and contributed to its development via their feedback. Miriam Madsen, Rana El Kaliouby, Matthew S. Goodwin, Rosalind W. Picard |
ASSETS | 4 |
| 2007 | Stoop to Conquer: Posture and Affect Interact to Influence Computer Users' Persistence
Hyungil Ahn, Alea Teeters, Andrew J. Wang, Cynthia Breazeal, Rosalind W. Picard |
ACII | 5 |
| 2007 | Experiments with a robotic computer: body, affect and cognition interactionsabstractWe present RoCo, the first robotic computer designed with the ability to move its monitor in subtly expressive ways that respond to and encourage its user's own postural movement. We use RoCo in a novel user study to explore whether a computer's "posture" can in fluence its use''s subsequent posture, and if the interaction of the user's body state with their affective state during a task leads to improved task measures such as persistence in problem solving. We believe this is possible in light of new theories that link physical posture and its in uence on affect and cognition. Initial results with 71 subjects support the hypothesis that RoCo's posture not only manipulates the user's posture, but also is associated with hypothesized posture-affect interactions. Specifically, we found effects on increased persistence on a subsequent cognitive task, and effects on perceived level of comfort. Cynthia Breazeal, Andrew J. Wang, Rosalind W. Picard |
HRI | 3 |
| 2007 | Automatic prediction of frustration
Ashish Kapoor, Winslow Burleson, Rosalind W. Picard |
Int. J. Hum. Comput. Stud. | 3 |
| 2007 | Relative subjective count and assessment of interruptive technologies applied to mobile monitoring of stress
Rosalind W. Picard, Karen K. Liu |
Int. J. Hum. Comput. Stud. | 1 |
| 2006 | Building an Affective Learning Companion
Rosalind W. Picard |
Intelligent Tutoring Systems | 1 |
| 2006 | Special issue on dialog systems for health communication
Timothy W. Bickmore, Toni Giorgino, Nancy L. Green, Rosalind W. Picard |
J. Biomed. Informatics | 4 |
| 2005 | Affective-Cognitive Learning and Decision Making: A Motivational Reward Framework for Affective Agents
Hyungil Ahn, Rosalind W. Picard |
ACII | 2 |
| 2005 | The HandWave Bluetooth Skin Conductance Sensor
Marc Strauss, Carson Reynolds, Stephen Hughes, Kyoung Park, Gary McDarby, Rosalind W. Picard |
ACII | 6 |
| 2005 | Classical and novel discriminant features for affect recognition from speechabstractThis paper investigates the performance and relevance of a set of acoustic features for the task of automatic recognition of affect from speech using machine learning techniques. Eighty seven novel and classical features related to loudness, intonation, and voice quality, are examined. Using feature selection, the results yield a performance level of 49.4% recognition rate (compared to a human performance rate of 60.4% and a chance level of 20%), while the relevance results show that the more exploratory and novel subset of these features outrank the more classical features in the recognition task. In the active research area of recognition of affect from speech it is of particular interest to obtain acoustic features that provide results closer to those of human recognition abilities. While many now “classic” features have been proposed in the literature, their performance has still fallen short of human recognition, suggesting the need to continue a search for novel features and methods. This paper briefly highlights results from an extensive investigation developing new features, and comparing them side-by-side with classical ones using machine learning techniques. (See [1] for many details omitted in this paper.) Algorithms and features associated with modeling loudness, intonation, and voice quality are highlighted in § 2, 3 and 4 respectively, and results of the experiments in §5 with some concluding remarks in § 6. Raul Fernandez, Rosalind W. Picard |
INTERSPEECH | 2 |
| 2005 | Multimodal affect recognition in learning environmentsabstractWe propose a multi-sensor affect recognition system and evaluate it on the challenging task of classifying interest (or disinterest) in children trying to solve an educational puzzle on the computer. The multimodal sensory information from facial expressions and postural shifts of the learner is combined with information about the learner's activity on the computer. We propose a unified approach, based on a mixture of Gaussian Processes, for achieving sensor fusion under the problematic conditions of missing channels and noisy labels. This approach generates separate class labels corresponding to each individual modality. The final classification is based upon a hidden random variable, which probabilistically combines the sensors. The multimodal Gaussian Process approach achieves accuracy of over 86%, significantly outperforming classification using the individual modalities, and several other combination schemes. Ashish Kapoor, Rosalind W. Picard |
ACM Multimedia | 2 |
| 2005 | Hyperparameter and Kernel Learning for Graph Based Semi-Supervised ClassificationabstractThere have been many graph-based approaches for semi-supervised clas- sification. One problem is that of hyperparameter learning: performance depends greatly on the hyperparameters of the similarity graph, trans- formation of the graph Laplacian and the noise model. We present a Bayesian framework for learning hyperparameters for graph-based semi- supervised classification. Given some labeled data, which can contain inaccurate labels, we pose the semi-supervised classification as an in- ference problem over the unknown labels. Expectation Propagation is used for approximate inference and the mean of the posterior is used for classification. The hyperparameters are learned using EM for evidence maximization. We also show that the posterior mean can be written in terms of the kernel matrix, providing a Bayesian classifier to classify new points. Tests on synthetic and real datasets show cases where there are significant improvements in performance over the existing approaches. Ashish Kapoor, Yuan Qi 0001, Hyungil Ahn, Rosalind W. Picard |
NIPS | 4 |
| 2005 | Detecting stress during real-world driving tasks using physiological sensorsabstractThis paper presents methods for collecting and analyzing physiological data during real-world driving tasks to determine a driver's relative stress level. Electrocardiogram, electromyogram, skin conductance, and respiration were recorded continuously while drivers followed a set route through open roads in the greater Boston area. Data from 24 drives of at least 50-min duration were collected for analysis. The data were analyzed in two ways. Analysis I used features from 5-min intervals of data during the rest, highway, and city driving conditions to distinguish three levels of driver stress with an accuracy of over 97% across multiple drivers and driving days. Analysis II compared continuous features, calculated at 1-s intervals throughout the entire drive, with a metric of observable stressors created by independent coders from videotapes. The results show that for most drivers studied, skin conductivity and heart rate metrics are most closely correlated with driver stress level. These findings indicate that physiological signals can provide a metric of driver stress in future cars capable of physiological monitoring. Such a metric could be used to help manage noncritical in-vehicle information systems and could also provide a continuous measure of how different road and traffic conditions affect drivers. Jennifer A. Healey, Rosalind W. Picard |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2005 | Establishing and maintaining long-term human-computer relationshipsabstractThis research investigates the meaning of “human-computer relationship” and presents techniques for constructing, maintaining, and evaluating such relationships, based on research in social psychology, sociolinguistics, communication and other social sciences. Contexts in which relationships are particularly important are described, together with specific benefits (like trust) and task outcomes (like improved learning) known to be associated with relationship quality. We especially consider the problem of designing for long-term interaction, and define relational agents as computational artifacts designed to establish and maintain long-term social-emotional relationships with their users. We construct the first such agent, and evaluate it in a controlled experiment with 101 users who were asked to interact daily with an exercise adoption system for a month. Compared to an equivalent task-oriented agent without any deliberate social-emotional or relationship-building skills, the relational agent was respected more, liked more, and trusted more, even after four weeks of interaction. Additionally, users expressed a significantly greater desire to continue working with the relational agent after the termination of the study. We conclude by discussing future directions for this research together with ethical and other ramifications of this work for HCI designers. Timothy W. Bickmore, Rosalind W. Picard |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2004 | Predictive automatic relevance determination by expectation propagationabstractIn many real-world classification problems the input contains a large number of potentially ir-relevant features. This paper proposes a new Bayesian framework for determining the rele-vance of input features. This approach extends one of the most successful Bayesian methods for feature selection and sparse learning, known as Automatic Relevance Determination (ARD). ARD finds the relevance of features by optimiz-ing the model marginal likelihood, also known as the evidence. We show that this can lead to over-fitting. To address this problem, we propose Pre-dictive ARD based on estimating the predictive performance of the classifier. While the actual leave-one-out predictive performance is generally very costly to compute, the expectation propaga-tion (EP) algorithm proposed by Minka provides an estimate of this predictive performance as a side-effect of its iterations. We exploit this in our algorithm to do feature selection, and to select data points in a sparse Bayesian kernel classifier. Moreover, we provide two other improvements to previous algorithms, by replacing Laplace’s approximation with the generally more accurate EP, and by incorporating the fast optimization algorithm proposed by Faul and Tipping. Our experiments show that our method based on the EP estimate of predictive performance is more accurate on test data than relevance determina-tion by optimizing the evidence. Yuan Qi 0001, Tom Minka, Rosalind W. Picard, Zoubin Ghahramani |
ICML | 3 |
| 2004 | Workshop on Social and Emotional Intelligence in Learning Environments
Claude Frasson, Kaska Porayska-Pomsta, Cristina Conati, Guy Gouardères, W. Lewis Johnson, Helen Pain, Elisabeth André, Timothy W. Bickmore, Paul Brna, Isabel Fernández de Castro, Stefano A. Cerri, Cleide Jane Costa, James C. Lester, Christine L. Lisetti, Stacy Marsella, Jack Mostow, Roger Nkambou, Magalie Ochs, Ana Paiva 0001, Fábio Paraguaçu, Natalie K. Person, Rosalind W. Picard, Candace L. Sidner, Angel de Vicente |
Intelligent Tutoring Systems | 22 |
| 2003 | Exertion interfaces: sports over a distance for social bonding and funabstractAn Exertion Interface is an interface that deliberately requires intense physical effort. Exertion Interfaces have applications in "Sports over a Distance", potentially capitalizing on the power of traditional physical sports in supporting social bonding. We designed, developed, and evaluated an Exertion Interface that allows people who are miles apart to play a physically exhausting ball game together. Players interact through a life-size video-conference screen using a regular soccer ball as an input device. The Exertion Interface users said that they got to know the other player better, had more fun, became better friends, and were happier with the transmitted audio and video quality, in comparison to those who played the same game using a non-exertion keyboard interface. These results suggest that an Exertion Interface, as compared to a traditional interface, offers increased opportunities for connecting people socially, especially when they have never met before. Florian 'Floyd' Mueller, Stefan Agamanolis, Rosalind W. Picard |
CHI | 3 |
| 2003 | Affective computing: challenges
Rosalind W. Picard |
Int. J. Hum. Comput. Stud. | 1 |
| 2003 | Modeling drivers' speech under stress
Raul Fernandez, Rosalind W. Picard |
Speech Commun. | 2 |
| 2002 | Bayesian spectrum estimation of unevenly sampled nonstationary dataabstractSpectral estimation methods typically assume stationarity and uniform spacing between samples of data. The non-stationarity of real data is usually accommodated by windowing methods, while the lack of uniformly-spaced samples is typically addressed by methods that “fill in” the data in some way. This paper presents a new approach to both of these problems: We use a non-stationary Kalman filter within a Bayesian framework to jointly estimate all spectral coefficients instantaneously. The new method works regardless of how the signal samples are spaced. We illustrate the method on several data sets, showing that it provides more accurate estimation than the Lomb-Scargle method and several classical spectral estimation methods. Yuan Qi 0001, Tom Minka, Rosalind W. Picard |
ICASSP | 3 |
| 2002 | Experimentally Augmenting an Intelligent Tutoring System with Human-Supplied Capabilities: Adding Human-Provided Emotional Scaffolding to an Automated Reading Tutor that ListensabstractWe present the first statistically reliable empirical evidence from a controlled study for the effect of human-provided emotional scaffolding on student persistence in an intelligent tutoring system. We describe an experiment that added human-provided emotional scaffolding to an automated Reading Tutor that listens, and discuss the methodology we developed to conduct this experiment. Each student participated in one (experimental) session with emotional scaffolding, and in one (control) session without emotional scaffolding, counterbalanced by order of session. Each session was divided into several portions. After each portion of the session was completed, the Reading Tutor gave the student a choice: continue, or quit. We measured persistence as the number of portions the student completed. Human-provided emotional scaffolding added to the automated Reading Tutor resulted in increased student persistence, compared to the Reading Tutor alone. Increased persistence means increased time on task, which ought lead to improved learning. If these results for reading turn out to hold for other domains too, the implication for intelligent tutoring systems is that they should respond with not just cognitive support-but emotional scaffolding as well. Furthermore, the general technique of adding human-supplied capabilities to an existing intelligent tutoring system should prove useful for studying other ITSs too. Gregory Aist, Barry Kort, Rob Reilly, Jack Mostow, Rosalind W. Picard |
ICMI | 5 |
| 2002 | Adding Human-Provided Emotional Scaffolding to an Automated Reading Tutor That Listens Increases Student Persistence
Gregory Aist, Barry Kort, Rob Reilly, Jack Mostow, Rosalind W. Picard |
Intelligent Tutoring Systems | 5 |
| 2001 | An Affective Model of Interplay between Emotions and Learning: Reengineering Educational Pedagogy - Building a Learning CompanionabstractThere is an interplay, between emotions and learning, but this interaction is far more complex than previous theories have articulated. The article proffers a novel model by which to: 1). regard the interplay of emotions upon learning for, 2). the larger practical aim of crafting computer-based models that will recognize a learner's affective state and respond appropriately to it, so that learning will proceed at an optimal pace. Barry Kort, Rob Reilly, Rosalind W. Picard |
ICALT | 3 |
| 2001 | Affective and Wearable Interfaces: Sensing and Responding to Human Emotion
Rosalind W. Picard |
INTERACT | 1 |
| 2001 | This computer responds to user frustration: Theory, design, and resultsabstractUse of technology often has unpleasant side effects, which may include strong, negative emotional states that arise during interaction with computers. Frustration, confusion, anger, anxiety and similar emotional states can affect not only the interaction itself, but also productivity, learning, social relationships, and overall well-being. This paper suggests a new solution to this problem: designing human–computer interaction systems to actively support users in their ability to manage and recover from negative emotional states. An interactive affect–support agent was designed and built to test the proposed solution in a situation where users were feeling frustration. The agent, which used only text and buttons in a graphical user interface for its interaction, demonstrated components of active listening, empathy, and sympathy in an effort to support users in their ability to recover from frustration. The agent's effectiveness was evaluated against two control conditions, which were also text-based interactions: (1) users’ emotions were ignored, and (2) users were able to report problems and ‘vent’ their feelings and concerns to the computer. Behavioral results showed that users chose to continue to interact with the system that had caused their frustration significantly longer after interacting with the affect–support agent, in comparison with the two controls. These results support the prediction that the computer can undo some of the negative feelings it causes by helping a user manage his or her emotional state. Jonathan Klein, Youngme Moon, Rosalind W. Picard |
Interact. Comput. | 3 |
| 2001 | Computers that recognise and respond to user emotion: theoretical and practical implicationsabstractPrototypes of interactive computer systems have been built that can begin to detect and label aspects of human emotional expression, and that respond to users experiencing frustration and other negative emotions with emotionally supportive interactions, demonstrating components of human skills such as active listening, empathy, and sympathy. These working systems support the prediction that a computer can begin to undo some of the negative feelings it causes by helping a user manage his or her emotional state. This paper clarifies the philosophy of this new approach to human–computer interaction: deliberately recognising and responding to an individual user's emotions in ways, that help users meet their needs. We define user needs in a broader perspective than has been hitherto discussed in the HCI community, to include emotional and social needs, and examine technology's emerging capability to address and support such needs. We raise and discuss potential concerns and objections regarding this technology, and describe several opportunities for future work. Rosalind W. Picard, Jonathan Klein |
Interact. Comput. | 1 |
| 2001 | Frustrating the user on purpose: a step toward building an affective computerabstractUsing a deliberately slow computer–game-interface to induce a state of hypothesised frustration in users, we collected physiological, video and behavioural data, and developed a strategy for coupling these data with real-world events. The effectiveness of our strategy was tested in a study with thirty six subjects, where the system was shown to reliably synchronise and gather data for affect analysis. A pattern-recognition strategy known as Hidden Markov Models was applied to each subject's physiological signals of skin conductivity and blood volume pressure in an effort to see if regimes of likely frustration could be automatically discriminated from regimes when frustration was much less likely. This pattern-recognition approach performed significantly better than random guessing at classifying the two regimes. Mouse-clicking behaviour was also synchronised to frustration-eliciting events and analysed, revealing four distinct patterns of clicking responses. We provide recommendations and guidelines for using physiology as a dependent measure for HCI experiments, especially when considering human emotions in the HCI equation. Jocelyn Scheirer, Raul Fernandez, Jonathan Klein, Rosalind W. Picard |
Interact. Comput. | 4 |
| 2001 | Tools for Browsing a TV Situation Comedy Based on Content Specific Attributes
Joshua S. Wachman, Rosalind W. Picard |
Multim. Tools Appl. | 2 |
| 2001 | Toward Machine Emotional Intelligence: Analysis of Affective Physiological StateabstractThe ability to recognize emotion is one of the hallmarks of emotional intelligence, an aspect of human intelligence that has been argued to be even more important than mathematical and verbal intelligences. This paper proposes that machine intelligence needs to include emotional intelligence and demonstrates results toward this goal: developing a machine's ability to recognize the human affective state given four physiological signals. We describe difficult issues unique to obtaining reliable affective data and collect a large set of data from a subject trying to elicit and experience each of eight emotional states, daily, over multiple weeks. This paper presents and compares multiple algorithms for feature-based recognition of emotional state from this data. We analyze four physiological signals that exhibit problematic day-to-day variations: The features of different emotions on the same day tend to cluster more tightly than do the features of the same emotion on different days. To handle the daily variations, we propose new features and algorithms and compare their performance. We find that the technique of seeding a Fisher Projection with the results of sequential floating forward search improves the performance of the Fisher Projection and provides the highest recognition rates reported to date for classification of affect from physiology: 81 percent recognition accuracy on eight classes of emotion, including neutral. Rosalind W. Picard, E. Vyzas, Jennifer A. Healey |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | SmartCar: Detecting Driver StressabstractSmart physiological sensors embedded in an automobile afford a novel opportunity to capture naturally occurring episodes of driver stress. In a series of ten ninety minute drives on public roads and highways, ECG, EMG, respiration and skin conductance sensors were used to measure the autonomic nervous system activation. The signals were digitized in real time and stored on the SmartCar's Pentium class computer. Each drive followed a pre-specified route through fifteen different events, from which four stress level categories were created according to the results of the subjects self report questionnaires. In total, 545 one minute segments were classified. A linear discriminant function was used to rank each feature individually based on the recognition performance, and a sequential forward floating selection algorithm was used to find an optimal set of features for recognizing patterns of driver stress. Using multiple features improved performance significantly over the best single feature performance. Jennifer A. Healey, Rosalind W. Picard |
ICPR | 2 |
| 1999 | A spectral 2-D Wold decomposition algorithm for homogeneous random fieldsabstractThe theory of the 2-D Wold decomposition of homogeneous random fields is effective in image and video analysis, synthesis, and modeling. However, a robust and computationally efficient decomposition algorithm is needed for use of the theory in practical applications. This paper presents a spectral 3-D Wold decomposition algorithm for homogeneous and near homogeneous random fields. The algorithm relies on the intrinsic fundamental-harmonic relationship among Fourier spectral peaks to identify harmonic frequencies, and uses a Hough transformation to detect spectral evanescent components. A local variance based procedure is developed to determine the spectral peak support. Compared to the two other existing methods for Wold decompositions, global thresholding and maximum-likelihood parameter estimation, this algorithm is more robust and flexible for the large variety of natural images, as well as computationally more efficient than the maximum-likelihood method. Fang Liu 0027, Rosalind W. Picard |
ICASSP | 2 |
| 1998 | Signal processing for recognition of human frustrationabstractIn this work, inspired by the application of human-machine interaction and the potential use that human-computer interfaces can make of knowledge regarding the affective state of a user, we investigate the problem of sensing and recognizing typical affective experiences that arise when people communicate with computers. In particular, we address the problem of detecting "frustration" in human computer interfaces. By first sensing human biophysiological correlates of internal affective states, we proceed to stochastically model the biological time series with hidden Markov models to obtain user-dependent recognition systems that learn affective patterns from a set of training data. Labeling criteria to classify the data are discussed, and generalization of the results to a set of unobserved data is evaluated. Significant recognition results (greater than random) are reported for 21 of 24 subjects. Raul Fernandez, Rosalind W. Picard |
ICASSP | 2 |
| 1998 | Digital processing of affective signalsabstractAffective signal processing algorithms were developed to allow a digital computer to recognize the affective state of a user who is intentionally expressing that state. This paper describes the method used for collecting the training data, the feature extraction algorithms used and the results of pattern recognition using a Fisher linear discriminant and the leave one out test method. Four physiological signals, skin conductivity, blood volume pressure, respiration and an electromyogram (EMG) on the masseter muscle were analyzed. It was found that anger was well differentiated from peaceful emotions (90%-100%), that high and low arousal states were distinguished (80%-88%), but positive and negative valence states were difficult to distinguish (50%-82%). Subsets of three emotion states could be well separated (75%-87%) and characteristic patterns for single emotions were found. Jennifer A. Healey, Rosalind W. Picard |
ICASSP | 2 |
| 1998 | Finding Periodicity in Space and TimeabstractAn algorithm for simultaneous detection, segmentation, and characterization of spatiotemporal periodicity is presented. The use of periodicity templates is proposed to localize and characterize temporal activities. The templates not only indicate the presence and location of a periodic event, but also give an accurate quantitative periodicity measure. Hence, they can be used as a new means of periodicity representation. The proposed algorithm can also be considered as a "periodicity filter", a low-level model of periodicity perception. The algorithm is computationally simple, and shown to be more robust than optical flow based techniques in the presence of noise. A variety of real-world examples are used to demonstrate the performance of the algorithm. Fang Liu 0027, Rosalind W. Picard |
ICCV | 2 |
| 1998 | Panel on Affect and Emotion in the User InterfaceabstractArticle Free Access Share on Panel on affect and emotion in the user interface Authors: Barbara Hayes-Roth Stanford University, Gates Computer Science Building, Stanford, CA Stanford University, Gates Computer Science Building, Stanford, CAView Profile , Gene Ball Microsoft Research, One Microsoft Way, Redmond, WA Microsoft Research, One Microsoft Way, Redmond, WAView Profile , Christine Lisetti Stanford University, Department of Computer Science and Department of Psychology, Stanford, CA Stanford University, Department of Computer Science and Department of Psychology, Stanford, CAView Profile , Rosalind W. Picard MIT Media Laboratory, E15-392, 20 Ames Street, Cambridge, MA MIT Media Laboratory, E15-392, 20 Ames Street, Cambridge, MAView Profile , Andrew Stern PF. Magic, 501 2nd St., Suite 400, San Francisco, CA PF. Magic, 501 2nd St., Suite 400, San Francisco, CAView Profile Authors Info & Claims IUI '98: Proceedings of the 3rd international conference on Intelligent user interfacesJanuary 1998 Pages 91–94https://doi.org/10.1145/268389.268406Online:01 January 1998Publication History 16citation1,051DownloadsMetricsTotal Citations16Total Downloads1,051Last 12 Months29Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Barbara Hayes-Roth, Gene Ball, Christine L. Lisetti, Rosalind W. Picard, Andrew Stern |
IUI | 4 |
| 1998 | Human-computer couplingabstract"Man-machine coupling," effecting the transfer of a human's thoughts, movements, and feelings to a computer, remains a challenge today. Alas, the 1960's reigning belief in the use of control theory for directly coupling us to machines and removing the need for "teams of programmers" has not provided the solution, although control theory has been of great use in the development of new technologies. Instead, teams of programmers are still necessary for enabling machines to interpret the signals that carry human communication. Such teams appear to be growing, not shrinking, in the coming decade. Rosalind W. Picard |
Proc. IEEE | 1 |
| 1997 | Interactive learning with a "society of models"
Tom Minka, Rosalind W. Picard |
Pattern Recognit. | 2 |
| 1997 | Affective WearablesabstractAn "affective wearable" is a wearable system equipped with sensors and tools which enables recognition of its wearer's affective patterns. Affective patterns include expressions of emotion such as a joyful smile, an angry gesture, a strained voice or a change in autonomic nervous system activity such as accelerated head rate or increasing skin conductivity. This paper describes new applications of affective wearables, and presents a prototype which gathers physiological signals and their annotations from its wearer. Results of preliminary experiments of its performance are reported for a user wearing four different sensors and engaging in several natural activities. Rosalind W. Picard, Jennifer A. Healey |
Pers. Ubiquitous Comput. | 1 |
| 1997 | Video orbits of the projective group a simple approach to featureless estimation of parametersabstractWe present direct featureless methods for estimating the eight parameters of an "exact" projective (homographic) coordinate transformation to register pairs of images, together with the application of seamlessly combining a plurality of images of the same scene, resulting in a single image (or new image sequence) of greater resolution or spatial extent. The approach is "exact" for two cases of static scenes: (1) images taken from the same location of an arbitrary three-dimensional (3-D) scene, with a camera that is free to pan, tilt, rotate about its optical axis, and zoom, or (2) images of a flat scene taken from arbitrary locations. The featureless projective approach generalizes interframe camera motion estimation methods that have previously used a camera model (which lacks the degrees of freedom to "exactly" characterize such phenomena as camera pan and tilt) and/or which have relied upon finding points of correspondence between the image frames. The featureless projective approach, which operates directly on the image pixels, is shown to be superior in accuracy and the ability to enhance the resolution. The proposed methods work well on image data collected from both good-quality and poor-quality video under a wide variety of conditions (sunny, cloudy, day, night). These new fully automatic methods are also shown to be robust to deviations from the assumptions of static scene and no parallax. Steve Mann 0001, Rosalind W. Picard |
IEEE Trans. Image Process. | 2 |
| 1997 | Cluster-based probability model and its application to image and texture processingabstractWe develop, analyze, and apply a specific form of mixture modeling for density estimation within the context of image and texture processing. The technique captures much of the higher order, nonlinear statistical relationships present among vector elements by combining aspects of kernel estimation and cluster analysis. Experimental results are presented in the following applications: image restoration, image and texture compression, and texture classification. Kris Popat, Rosalind W. Picard |
IEEE Trans. Image Process. | 2 |
| 1996 | Interactive Learning with a "Society of Models" abstractDigital library access is driven by features, but the relevance of a feature for a query is not always obvious. This paper describes an approach for integrating a large number of context-dependent features into a semi-automated tool. Instead of requiring universal similarity measures or manual selection of relevant features, the approach provides a learning algorithm for selecting and combining groupings of the data, where groupings can be induced by highly specialized features. The selection process is guided by positive and negative examples from the user. The inherent combinatorics of using multiple features is reduced by a multistage grouping generation, weighting, and collection process. The stages closest to the user are trained fastest and slowly propagate their adaptations back to earlier stages. The weighting stage adapts the collection stage's search space across uses, so that, in later interactions, good groupings are found given few examples from the user. Tom Minka, Rosalind W. Picard |
CVPR | 2 |
| 1996 | Modeling user subjectivity in image librariesabstractIn addition to the problem of which image analysis models to use in digital libraries, e.g. wavelet, Wold, color histograms, is the problem of how to combine these models with their different strengths. Most present systems place the burden of combination on the user, e.g. the user specifies 50% texture features, 20% color features, etc. This is a problem since most users do not know how to best pick the settings for the given data and search problem. The paper addresses this problem, describing research in progress for a system that: (1) automatically infers which combination of models best represents the data of interest to the user; and (2) learns continuously during interaction with each user. In particular, these two components-inference and learning-provide a solution that adapts to the subjective and hard to predict behaviors frequently seen when people query or browse image libraries. Rosalind W. Picard, Tom Minka, Martin Szummer |
ICIP (2) | 1 |
| 1996 | Temporal texture modelingabstractTemporal textures are textures with motion. Examples include wavy water, rising steam and fire. We model image sequences of temporal textures using the spatio-temporal autoregressive model (STAR). This model expresses each pixel as a linear combination of surrounding pixels lagged both in space and in time. The model provides a base for both recognition and synthesis. We show how the least squares method can accurately estimate model parameters for large, causal neighborhoods with more than 1000 parameters. Synthesis results show that the model can adequately capture the spatial and temporal characteristics of many temporal textures. A 95% recognition rate is achieved for a 135 element database with 15 texture classes. Martin Szummer, Rosalind W. Picard |
ICIP (3) | 2 |
| 1996 | Photobook: Content-based manipulation of image databases
Alex Pentland, Rosalind W. Picard, Stan Sclaroff |
Int. J. Comput. Vis. | 2 |
| 1996 | Periodicity, Directionality, and Randomness: Wold Features for Image Modeling and RetrievalabstractOne of the fundamental challenges in pattern recognition is choosing a set of features appropriate to a class of problems. In applications such as database retrieval, it is important that image features used in pattern comparison provide good measures of image perceptual similarities. We present an image model with a new set of features that address the challenge of perceptual similarity. The model is based on the 2D Wold decomposition of homogeneous random fields. The three resulting mutually orthogonal subfields have perceptual properties which can be described as "periodicity," "directionality," and "randomness," approximating what are indicated to be the three most important dimensions of human texture perception. The method presented improves upon earlier Wold-based models in its tolerance to a variety of local inhomogeneities which arise in natural textures and its invariance under image transformation such as rotation. An image retrieval algorithm based on the new texture model is presented. Different types of image features are aggregated for similarity comparison by using a Bayesian probabilistic approach. The, effectiveness of the Wold model at retrieving perceptually similar natural textures is demonstrated in comparison to that of two other well-known pattern recognition methods. The Wold model appears to offer a perceptually more satisfying measure of pattern similarity while exceeding the performance of these other methods by traditional pattern recognition criteria. Examples of natural scene Wold texture modeling are also presented. Fang Liu 0027, Rosalind W. Picard |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1996 | Introduction to the Special Section on Digital Libraries: Representation and Retrieval
Rosalind W. Picard, Alex Pentland |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1996 | M-lattice: from morphogenesis to image processingabstractThe paper is based on reaction-diffusion, a nonlinear mechanism first proposed by Turing in 1952 to account for morphogenesis, the formation of shape and pattern in nature. One of the key limitations of reaction-diffusion systems is that they are generally unbounded, making them awkward for digital image processing. In this paper we introduce the "M-lattice", a system that preserves the pattern-formation properties of reaction-diffusion and is bounded. On the theoretical front, we establish how the M-lattice is closely related to the analog Hopfield network and the cellular neural network, but has more flexibility in how its variables interact. Like many "neurally inspired" systems, the bounded M-lattice also enables computer or analog VLSI implementations to simulate a variety of partial and ordinary differential equations. On the practical front, we demonstrate two novel applications of reaction-diffusion formulated as the new M-lattice. These are adaptive filtering, applied to the restoration and enhancement of fingerprint images, and nonlinear programming, applied to image halftoning in both "faithful" and "special effects" styles. Alex Sherstinsky, Rosalind W. Picard |
IEEE Trans. Image Process. | 2 |
| 1996 | On the efficiency of the orthogonal least squares training method for radial basis function networksabstractThe efficiency of the orthogonal least squares (OLS) method for training approximation networks is examined using the criterion of energy compaction. We show that the selection of basis vectors produced by the procedure is not the most compact when the approximation is performed using a nonorthogonal basis. Hence, the algorithm does not produce the smallest possible networks for a given approximation error. Specific examples are given using the Gaussian radial basis functions type of approximation networks. Alex Sherstinsky, Rosalind W. Picard |
IEEE Trans. Neural Networks | 2 |
| 1995 | Digital Libraries: Meeting Place for Low-Level And High-Level Vision
Rosalind W. Picard |
ACCV | 1 |
| 1995 | Light-years from Lena: video and image libraries of the futureabstractThe average consumer with a personal computer will soon have access to the world's collections of digital video and images. However, the theory and tools that facilitate browsing, querying, retrieval, and manipulation of imagery are still in their infancy. For example, people would like to access content in movies, e.g. "fast forward to where they bicycle through the sky". This new application area reveals an abundance of unsolved scientific problems for image processing. An overview is provided of the key technical challenges that the image processing community should embrace. The scope of the paper is restricted to image processing, with particular focus on the problems of representation and analysis of image content, and frame-to-frame motion. Rosalind W. Picard |
ICIP | 1 |
| 1995 | Color halftoning with M-latticeabstractThis paper presents new results in the halftoning of color images. The first automatic generation of the Wall Street Journal type halftones of gray-scale images was done using the M-lattice system and reported in Sherstinsky and Picard [1994]. The M-lattice system was derived from the reaction-diffusion model, first proposed by Turing [1952] in order to explain mammal coat patterns. The M-lattice is a non-linear dynamical system that is well-suited for a variety of applications formulated as constrained non-linear optimization. In particular, it can perform image processing operations that emphasize oriented patterns. The present study uses this property to extend the special-effects halftoning of gray-scale images to that of color images. The duality metric comes from the directionality information extracted by steerable filters from the gray-scale version of the original color image. The binary requirement is stated as an explicit constraint, and all three (red, green, and blue) halftone components are synthesized simultaneously by the M-lattice. Alex Sherstinsky, Rosalind W. Picard |
ICIP | 2 |
| 1995 | Vision Texture for Annotation
Rosalind W. Picard, Tom Minka |
Multim. Syst. | 1 |
| 1994 | A new Wold ordering for image similarityabstractThe problem of measuring perceptual similarity between images is addressed using a new image model based on the Weld decomposition. The model permits separate treatment of image components which correspond approximately to periodicity, directionality, and randomness. The authors compare its performance in an image search application to two other methods-one based on shift-invariant principle components and one based on a multiscale simultaneous autoregressive model. When textured images are ordered by distances between their Wold components, the results appear to be much closer to the human perception of similarity. The authors discuss how decoupling the three components can increase flexibility for measuring image similarity and can save computation, permitting the "quickest" matches when the features are the most perceptually "salient".> Rosalind W. Picard, Fang Liu 0027 |
ICASSP (5) | 1 |
| 1994 | Cluster-based probability model applied to image restoration and compressionabstractThe performance of a statistical signal processing system is determined in large part by the accuracy of the probabilistic model it employs. Accurate modeling often requires working in several dimensions, but doing so can introduce dimensionality-related difficulties. A previously introduced model circumvents some of these difficulties while maintaining accuracy sufficient to account for much of the high-order, nonlinear statistical interdependence of samples. Properties of this model are reviewed, and its power demonstrated by application to image restoration and compression. Also described is a vector quantization (VQ) scheme which employs the model in entropy coding a Z/sup N/-lattice. The scheme has the advantage over standard VQ of bounding maximum instantaneous errors.> Kris Popat, Rosalind W. Picard |
ICASSP (5) | 2 |
| 1994 | M-lattice: a novel non-linear dynamical system and its application to halftoningabstractThis paper presents a novel non-linear dynamical system called the "M-lattice system". This system is rooted in the reaction-diffusion model, first proposed be Turing in 1952 to explain the formation of animal patterns such as zebra stripes and leopard spots. The M-lattice system is closely related to the analog Hopfield network and the cellular neural network, but has more flexibility in how its variables interact. In particular, the model is well-suited to a variety of applications formulated as constrained nonlinear optimization. The present study demonstrates the use of this model for two different image halftoning examples. The first example synthesizes a halftone of Einstein in the "hand-drawn" style of the Wall Street Journal portraits; it illustrates how a more flexible quality metric can be used when the binary requirement is stated as an explicit constraint. The second example synthesizes halftones free of correlated artifacts; it illustrates the noise-shaping capability of the M-lattice system.> Alex Sherstinsky, Rosalind W. Picard |
ICASSP (2) | 2 |
| 1994 | Virtual Bellows: Constructing High Quality Stills from VideoabstractCameras with bellows give photographers flexibility for controlling perspective, but once the picture is taken, its perspective is set. We introduce 'virtual bellows' to provide control over perspective after a picture has been taken. Virtual bellows can be used to align images taken from different viewpoints, an important initial step in applications such as creating a high-resolution still image from video. We show how the virtual bellows, which implements the projective group, is an exact model fit to both pan and tilt. Specifically, we identify two important classes of image sequences accommodated by the virtual bellows. Examples of constructing high-quality stills are shown for the two cases: multiple frames taken of a flat object, and multiple frames taken from a fixed point.> Steve Mann 0001, Rosalind W. Picard |
ICIP (1) | 2 |
| 1994 | Exaggerated Consensus in Lossless Image CompressionabstractGood probabilistic models are needed in data compression and many other applications. A good model must exploit contextual information, which requires high-order conditioning. As the number of conditioning variables increases, direct estimation of the distribution becomes exponentially more difficult. To circumvent this, we consider a means of adaptively combining several low-order conditional probability distributions into a single higher-order estimate, based on their degree of agreement. Though the technique is broadly applicable, image compression is singled out as a testing ground of its abilities. Good performance is demonstrated by experimental results.> Kris Popat, Rosalind W. Picard |
ICIP (3) | 2 |
| 1994 | Orientation-sensitive Image Processing with M-Lattice - A Novel Nonlinear Dynamical SystemabstractResearchers in image processing have long recognized the importance of modeling the human observer. Although a full human vision model remains elusive, orientation detection, one of the key components of human vision, can be directly incorporated into a variety of image processing algorithms. Orientation detection also provides cues that allow an algorithm to adapt to inhomogeneities in images. The authors show how the M-lattice system, a new non-linear dynamical system, can easily incorporate orientation sensitivity for two different types of problems. First, simultaneous adaptive filtering and non-linear restoration is illustrated for fingerprint enhancement. Second, constrained non-linear optimization is illustrated for halftoning in a "hand-drawn" style.> Alex Sherstinsky, Rosalind W. Picard |
ICIP (3) | 2 |
| 1994 | Texture orientation for sorting photos "at a glance"abstractInvestigates a measure of "dominant perceived orientation" that has been developed to match the output of a human study involving 40 subjects. The results of this measure are compared with humans analyzing seven "teaser" images to test its effectiveness for finding perceptually dominant orientations. The use of low-level orientation is then applied to a "quick search" problem important in image database applications. Since both pigeons and humans are able to perform coarse classification of certain kinds of scenes, e.g., city from country, without taking time or brain-power to solve the image understanding problem, the authors conjecture that the collective behavior of low-level textural features such as orientation may be doing most of the work. The authors demonstrate a simple test of global multiscale orientation for quickly searching a database of vacation photos for likely "city/suburb" shots. The orientation features achieve agreement with human classification in 91 out of 98 of the scenes. Monika M. Gorkani, Rosalind W. Picard |
ICPR (1) | 2 |
| 1994 | Periodicity, directionality, and randomness: Wold features for perceptual pattern recognitionabstractOne of the fundamental challenges in pattern recognition is choosing a set of features appropriate to a class of problems. In applications such as image retrieval, if is important that features used by the system in pattern comparison provide good measures of "perceptual similarity". The authors present a new set of features and an image model based on the three mutually orthogonal components produced by the 2-D Wold decomposition of random fields. These components have visual properties which approximate the three most important perceptual dimensions of human texture perception. The method presented here is different from the existing Wold-based models in that it tolerates certain local inhomogeneities which arise in natural textures and reduces computation for comparison of patterns subjected to transformations such as rotation. An image retrieval algorithm based on the new texture model is presented. The effectiveness of the new Wold features for retrieving perceptually similar natural textures is demonstrated by comparing it to that of other well-known pattern recognition methods. The Wold model appears to offer a perceptually more satisfying measure of pattern similarity. Fang Liu 0027, Rosalind W. Picard |
ICPR (2) | 2 |
| 1994 | Restoration and enhancement of fingerprint images using M-lattice-a novel nonlinear dynamical systemabstractDevelops a method for the simultaneous restoration and halftoning of fingerprints using the "M-lattice", a new nonlinear dynamical system. This system is rooted in the reaction-diffusion model, first proposed by Turing to explain morphogenesis (the formation of patterns in nature). But in contrast with the general reaction-diffusion, the state variables of the M-lattice are guaranteed to be bounded. The M-lattice system is closely related to the analog Hopfield network and the cellular neural network, but has more flexibility in how its variables interact. These properties make it better suited than reaction-diffusion for several new engineering applications. The proposed method for enhancing fingerprints explores the ability of the M-lattice to form oriented spatial patterns (like reaction-diffusion), while producing binary outputs (like feedback neural networks). The fingerprints synthesized by the M-lattice retain and emphasize more of the relevant detail than do those obtained by adaptive thresholding, a common halftoning method employed in traditional fingerprint classification systems. Alex Sherstinsky, Rosalind W. Picard |
ICPR (2) | 2 |
| 1994 | Gibbs Random Fields, Cooccurrences, and Texture ModelingabstractGibbs random field (GRF) models and features from cooccurrence matrices are typically considered as separate but useful tools for texture discrimination. The authors show an explicit relationship between cooccurrences and a large class of GRF's. This result comes from a new framework based on a set-theoretic concept called the "aura set" and on measures of this set, "aura measures." This framework is also shown to be useful for relating different texture analysis tools. The authors show how the aura set can be constructed with morphological dilation, how its measure yields cooccurrences, and how it can be applied to characterizing the behavior of the Gibbs model for texture. In particular, they show how the aura measure generalizes, to any number of gray levels and neighborhood order, some properties previously known for just the binary, nearest-neighbor GRF. Finally, the authors illustrate how these properties can guide one's intuition about the types of GRF patterns which are most likely to form.> Ibrahim M. Elfadel, Rosalind W. Picard |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1993 | Real-time recognition with the entire Brodatz texture databaseabstractThe Brodatz Album has become the de facto standard for evaluating texture algorithms, with hundreds of studies having been applied to small sets of its images. The authors compare two powerful recognition algorithms, principal components analysis and multiscale autoregressive models, by evaluating them on a 999-image database derived from the entire Brodatz Album. The variety of homogeneous and nonhomogeneous images studied is thus nearly an order of magnitude larger than has been compared before, giving one snapshot of the state of the art in real-time texture recognition.> Rosalind W. Picard, Tanweer Kabir, Fang Liu 0027 |
CVPR | 1 |
| 1993 | Finding similar patterns in large image databases
Rosalind W. Picard, Tanweer Kabir |
ICASSP (5) | 1 |
| 1993 | Novel cluster-based probability model for texture synthesis, classification, and compressionabstractWe present a new probabilistic modeling technique for high-dimensional vector sources, and consider its application to the problems of texture synthesis, classification, and compression. Our model combines kernel estimation with clustering, to obtain a semiparametric probability mass function estimate which summarizes -- rather than contains -- the training data. Because the model is cluster based, it is inferable from a limited set of training data, despite the model's high dimensionality. Moreover, its functional form allows recursive implementation that avoids exponential growth in required memory as the number of dimensions increases. Experimental results are presented for each of the three applications considered. Kris Popat, Rosalind W. Picard |
VCIP | 2 |
| 1992 | Gibbs random fields: temperature and parameter analysisabstractGibbs random field (GRF) models work well for synthesizing complex natural-looking image data with a small number of parameters; however, estimation methods for these parameters have a lot of problems. The analysis problem is addressed in a new way by examining the role of the temperature parameter of the Gibbs distribution. Studies of the model energy with respect to the temperature are used to indicate pattern equilibrium and regions of different behaviour, analogous to the existence of distinct phases in a physical system. The results on equilibrium and regions of different phases are offered as explanations for some of the peculiar behaviour of current estimation algorithms.> Rosalind W. Picard |
ICASSP | 1 |
| 1991 | Markov/Gibbs texture modeling: aura matrices and temperature effectsabstractAn 'aura' framework is used to rewrite the nonlinear energy function of a homogeneous anisotropic Markov/Gibbs random field (MRF) as a linear sum of aura measures. The formulation relates MRFs to co-occurrence matrices. It also provides a physical interpretation of MRF textures in terms of the mixing and separation of gray-level sets, and in terms of boundary maximization and minimization. Within this framework, the authors introduce the use of temperature for texture modeling and show how the parameters of the MRF can be interpreted as temperature annealing rates. In particular, they show evidence for a transition temperature, above which all patterns generated will be visually similar, and below which a pattern evolves down to its ground state. Results which characterize the ground state patterns are described.> Rosalind W. Picard, Ibrahim M. Elfadel, Alex Pentland |
CVPR | 1 |