VLDB 2026 Research / reviewers in the wild / expert
Daniel McDuff
dblp:63/9606 · also Daniel J. McDuff
· DBLP profile ↗
94ranked-venue papers
25as first author
41since 2021 · last 2025
0000-0001-7313-0082ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 14 first-author · 20 since 2021Human-computer interaction and ubiquitous computing · 39 · 9 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Substance over Style: Evaluating Proactive Conversational Coaching AgentsabstractWhile NLP research has made strides in conversational tasks, many approaches focus on single-turn responses with well-defined objectives or evaluation criteria. In contrast, coaching presents unique challenges with initially undefined goals that evolve through multi-turn interactions, subjective evaluation criteria, mixed-initiative dialogue. In this work, we describe and implement five multi-turn coaching agents that exhibit distinct conversational styles, and evaluate them through a user study, collecting first-person feedback on 155 conversations. We find that users highly value core functionality, and that stylistic components in absence of core components are viewed negatively. By comparing user feedback with third-person evaluations from health experts and an LM, we reveal significant misalignment across evaluation approaches. Our findings provide insights into design and evaluation of conversational coaching agents and contribute toward improving human-centered NLP applications. Vidya Srinivas, Xuhai Xu, Xin Liu 0034, Kumar Ayush, Isaac R. Galatzer-Levy, Shwetak N. Patel, Daniel McDuff, Tim Althoff |
ACL (1) | 7 |
| 2025 | Non-Contact Health Monitoring During Daily Personal Care RoutinesabstractRemote photoplethysmography (rPPG) enables noncontact, continuous monitoring of physiological signals and offers a practical alternative to traditional health sensing methods. Although rPPG is promising for daily health monitoring, its application in long-term personal care scenarios-such as mirrorfacing routines in high-altitude environments-remains challenging due to ambient lighting variations, frequent occlusions from hand movements, and dynamic facial postures. To address these challenges, we present the Long-term Altitude Daily Health (LADH) dataset, the first long-term rPPG dataset containing 240 synchronized RGB and infrared (IR) facial videos from 21 participants across five common personal care scenarios, along with ground-truth PPG, respiration, and blood oxygen signals. Our experiments demonstrate that combining RGB and IR video inputs improves the accuracy and robustness of non-contact physiological monitoring, achieving a mean absolute error (MAE) of 4.99 BPM in heart rate estimation. Furthermore, we find that multi-task learning enhances performance across multiple physiological indicators simultaneously. Dataset and code are open at https://github.com/McJackTang/FusionVitals. Xulin Ma, Jiankai Tang, Zhang Jiang, Songqin Cheng, Yuanchun Shi, Xin Liu 0034, Daniel McDuff, Yuntao Wang 0001 |
BSN | 8 |
| 2025 | Promoting Prosociality via Micro-acts of Joy: A Large-Scale Well-Being Intervention StudyabstractProsociality has been well-documented to positively impact mental, social, and physical well-being.However, existing studies of interventions for promoting prosociality have limitations such as Hitesh Goel, Yoobin Park, Jin Liou, Darwin A. Guevarra, Peggy Callahan, Jolene Smith, Bingsheng Yao, Dakuo Wang, Xin Liu 0034, Daniel McDuff, Noémie Elhadad, Emiliana Simon-Thomas, Elissa Epel, Xuhai Xu |
CHI | 10 |
| 2025 | Scaling Wearable Foundation ModelsabstractWearable sensors have become ubiquitous thanks to a variety of health tracking features. The resulting continuous and longitudinal measurements from everyday life generate large volumes of data. However, making sense of these observations for scientific and actionable insights is non-trivial. Inspired by the empirical success of generative modeling, where large neural networks learn powerful representations from vast amounts of text, image, video, or audio data, we investigate the scaling properties of wearable sensor foundation models across compute, data, and model size. Using a dataset of up to 40 million hours of in-situ heart rate, heart rate variability, accelerometer, electrodermal activity, skin temperature, and altimeter per-minute data from over 165,000 people, we create LSM, a multimodal foundation model built on the largest wearable-signals dataset with the most extensive range of sensor modalities to date. Our results establish the scaling laws of LSM for tasks such as imputation, interpolation and extrapolation across both time and sensor modalities. Moreover, we highlight how LSM enables sample-efficient downstream learning for tasks including exercise and activity recognition. Girish Narayanswamy, Xin Liu 0034, Kumar Ayush, Yuzhe Yang 0003, Xuhai Xu, Shun Liao, Jake Garrison, Shyam A. Tailor, Jacob E. Sunshine, Yun Liu 0013, Tim Althoff, Shri Narayanan, Pushmeet Kohli, Jiening Zhan, Mark Malhotra, Shwetak N. Patel, Samy Abdel-Ghaffar, Daniel McDuff |
ICLR | 18 |
| 2025 | VocalAgent: Large Language Models for Vocal Health Diagnostics with Safety-Aware EvaluationabstractVocal health plays a crucial role in peoples' lives, significantly impacting their communicative abilities and interactions. However, despite the global prevalence of voice disorders, many lack access to convenient diagnosis and treatment. This paper introduces VocalAgent, an audio large language model (LLM) to address these challenges through vocal health diagnosis. We leverage Qwen-Audio-Chat fine-tuned on three datasets collected in-situ from hospital patients, and present a multifaceted evaluation framework encompassing a safety assessment to mitigate diagnostic biases, cross-lingual performance analysis, and modality ablation studies. VocalAgent demonstrates superior accuracy on voice disorder classification compared to state-of-the-art baselines. Its LLM-based method offers a scalable solution for broader adoption of health diagnostics, while underscoring the importance of ethical and technical validation. Yubin Kim 0002, Taehan Kim, Wonjune Kang, Eugene Park, Joonsik Yoon, Xin Liu 0034, Daniel McDuff, Hyeonhoon Lee, Cynthia Breazeal, Hae Won Park 0001 |
INTERSPEECH | 8 |
| 2025 | RADAR: Benchmarking Language Models on Imperfect Tabular DataabstractLanguage models (LMs) are increasingly being deployed to perform autonomous data analyses. However, their data awareness—the ability to recognize, reason over, and appropriately handle data artifacts such as missing values, outliers, and logical inconsistencies—remains underexplored. These artifacts are especially common in real-world tabular data and, if mishandled, can significantly compromise the validity of analytical conclusions. To address this gap, we present RADAR, a benchmark for systematically evaluating data-aware reasoning on tabular data. We develop a framework to simulate data artifacts via programmatic perturbations to enable targeted evaluation of model behavior. RADAR comprises 2,980 table-query pairs, grounded in real-world data spanning 9 domains and 5 data artifact types. In addition to evaluating artifact handling, RADAR systematically varies table size to study how reasoning performance holds when increasing table size. Our evaluation reveals that, despite decent performance on tables without data artifacts, frontier models degrade significantly when data artifacts are introduced, exposing critical gaps in their capacity for robust, data-aware analysis. Designed to be flexible and extensible, RADAR supports diverse perturbation types and controllable table sizes, offering a valuable resource for advancing tabular reasoning. Ken Gu, Zhihan Zhang 0002, Kate Lin, Yuwei Zhang 0001, Akshay Paruchuri, Hong Yu 0001, Mehran Kazemi, Kumar Ayush, A. Ali Heydari, Maxwell A. Xu, Yun Liu 0013, Ming-Zher Poh, Yuzhe Yang 0003, Mark Malhotra, Shwetak N. Patel, Hamid Palangi, Xuhai Xu, Daniel McDuff, Tim Althoff, Xin Liu 0034 |
NeurIPS | 18 |
| 2025 | SensorLM: Learning the Language of Wearable SensorsabstractWe present SensorLM, a family of sensor-language foundation models that enable wearable sensor data understanding with natural language. Despite its pervasive nature, aligning and interpreting sensor data with language remains challenging due to the lack of paired, richly annotated sensor-text descriptions in uncurated, real-world wearable data. We introduce a hierarchical caption generation pipeline designed to capture statistical, structural, and semantic information from sensor data. This approach enabled the curation of the largest sensor-language dataset to date, comprising over 59.7 million hours of data from more than 103,000 people. Furthermore, SensorLM extends prominent multimodal pretraining architectures (e.g., CLIP, CoCa) and recovers them as specific variants within a generic architecture. Extensive experiments on real-world tasks in human activity analysis and healthcare verify the superior performance of SensorLM over state-of-the-art in zero-shot recognition, few-shot learning, and cross-modal retrieval. SensorLM also demonstrates intriguing capabilities including scaling behaviors, label efficiency, sensor captioning, and zero-shot generalization to unseen tasks. Code is available at https://github.com/Google-Health/consumer-health-research/tree/main/sensorlm. Yuwei Zhang 0001, Kumar Ayush, Siyuan Qiao, A. Ali Heydari, Girish Narayanswamy, Maxwell A. Xu, Ahmed Metwally 0002, Jinhua Xu, Jake Garrison, Xuhai Xu, Tim Althoff, Yun Liu 0013, Pushmeet Kohli, Jiening Zhan, Mark Malhotra, Shwetak N. Patel, Cecilia Mascolo, Xin Liu 0034, Daniel McDuff, Yuzhe Yang 0003 |
NeurIPS | 19 |
| 2025 | Triple Peak Day: Work Rhythms of Software Developers in Hybrid WorkabstractThe future of work is rapidly changing, with remote and hybrid settings blurring the boundaries between professional and personal life. To understand how work rhythms vary across different work settings, we conducted a month-long study of 65 software developers, collecting anonymized computer activity data as well as daily ratings for perceived stress, productivity, and work setting. In addition to confirming the double-peak pattern of activity at 10:00 am and 2:00 pm observed in prior research, we observed a significant third peak around 9:00 pm. This third peak was associated with higher perceived productivity during remote days but increased stress during onsite and hybrid days, highlighting a nuanced interplay between work demands and work settings. Additionally, we found strong correlations between computer activity, productivity, and stress, including an inverted U-shaped relationship where productivity peaked at around six hours of computer activity before declining on more active days. These findings provide new insights into evolving work rhythms and highlight the impact of different work settings on productivity and stress. Javier Hernandez, Vedant Das Swain, Jina Suh, Daniel McDuff, Judith Amores, Gonzalo A. Ramos, Kael Rowan, Brian Houck, Shamsi T. Iqbal, Mary Czerwinski |
IEEE Trans. Software Eng. | 4 |
| 2024 | What Are the Odds? Language Models Are Capable of Probabilistic ReasoningabstractAkshay Paruchuri, Jake Garrison, Shun Liao, John B Hernandez, Jacob Sunshine, Tim Althoff, Xin Liu, Daniel McDuff. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Akshay Paruchuri, Jake Garrison, Shun Liao, John Hernandez, Jacob E. Sunshine, Tim Althoff, Xin Liu 0034, Daniel McDuff |
EMNLP | 8 |
| 2024 | Position: Standardization of Behavioral Use Clauses is Necessary for the Adoption of Responsible Licensing of AIabstractGrowing concerns over negligent or malicious uses of AI have increased the appetite for tools that help manage the risks of the technology. In 2018, licenses with behaviorial-use clauses (commonly referred to as Responsible AI Licenses) were proposed to give developers a framework for releasing AI assets while specifying their users to mitigate negative applications. As of the end of 2023, on the order of 40,000 software and model repositories have adopted responsible AI licenses licenses. Notable models licensed with behavioral use clauses include BLOOM (language) and LLaMA2 (language), Stable Diffusion (image), and GRID (robotics). This paper explores why and how these licenses have been adopted, and why and how they have been adapted to fit particular use cases. We use a mixed-methods methodology of qualitative interviews, clustering of license clauses, and quantitative analysis of license adoption. Based on this evidence we take the position that responsible AI licenses need standardization to avoid confusing users or diluting their impact. At the same time, customization of behavioral restrictions is also appropriate in some contexts (e.g., medical domains). We advocate for “standardized customization” that can meet users’ needs and can be supported via tooling. Daniel McDuff, Tim Korjakow, Scott Cambo, Jesse Josua Benjamin, Jenny Lee, Yacine Jernite, Carlos Muñoz Ferrandis, Aaron Gokaslan, Alek Tarkowski, Joseph Lindley, A. Feder Cooper, Danish Contractor |
ICML | 1 |
| 2024 | MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingabstractFoundation models are becoming valuable tools in medicine. Yet despite their promise, the best way to leverage Large Language Models (LLMs) in complex medical tasks remains an open question. We introduce a novel multi-agent framework, named **M**edical **D**ecision-making **Agents** (**MDAgents**) that helps to address this gap by automatically assigning a collaboration structure to a team of LLMs. The assigned solo or group collaboration structure is tailored to the medical task at hand, a simple emulation inspired by the way real-world medical decision-making processes are adapted to tasks of different complexities. We evaluate our framework and baseline methods using state-of-the-art LLMs across a suite of real-world medical knowledge and clinical diagnosis benchmarks, including a comparison of
LLMs’ medical complexity classification against human physicians. MDAgents achieved the **best performance in seven out of ten** benchmarks on tasks requiring an understanding of medical knowledge and multi-modal reasoning, showing a significant **improvement of up to 4.2\%** ($p$ < 0.05) compared to previous methods' best performances. Ablation studies reveal that MDAgents effectively determines medical complexity to optimize for efficiency and accuracy across diverse medical tasks. Notably, the combination of moderator review and external medical knowledge in group collaboration resulted in an average accuracy **improvement of 11.8\%**. Our code can be found at https://github.com/mitmedialab/MDAgents. Yubin Kim 0002, Chanwoo Park, Hyewon Jeong, Yik Siu Chan, Xuhai Xu, Daniel McDuff, Hyeonhoon Lee, Marzyeh Ghassemi, Cynthia Breazeal, Hae Won Park 0001 |
NeurIPS | 6 |
| 2024 | BigSmall: Efficient Multi-Task Learning for Disparate Spatial and Temporal Physiological MeasurementsabstractUnderstanding of human visual perception has historically inspired the design of computer vision architectures. As an example, perception occurs at different scales both spatially and temporally, suggesting that the extraction of salient visual information may be made more effective by attending to specific features at varying scales. Visual changes in the body, due to physiological processes, also occur at varying scales and with modality-specific characteristic properties. Inspired by this, we present BigSmall, an efficient architecture for physiological and behavioral measurement. We present the first joint camera-based facial action, cardiac, and pulmonary measurement model. We propose a multi-branch network with wrapping temporal shift modules that yields efficiency gains and accuracy on par with task-optimized methods. We observe that fusing low-level features leads to suboptimal performance, but that fusing high level features enables efficiency gains with negligible losses in accuracy. We experimentally validate that BigSmall significantly reduces computational cost while achieving comparable results on multiple physiological measurement tasks simultaneously with a unified model. Girish Narayanswamy, Yuzhe Yang 0003, Chengqian Ma, Xin Liu 0034, Daniel McDuff, Shwetak N. Patel |
WACV | 6 |
| 2024 | Motion Matters: Neural Motion Transfer for Better Camera Physiological MeasurementabstractMachine learning models for camera-based physiological measurement can have weak generalization due to a lack of representative training data. Body motion is one of the most significant sources of noise when attempting to recover the subtle cardiac pulse from a video. We explore motion transfer as a form of data augmentation to introduce motion variation while preserving physiological changes of interest. We adapt a neural video synthesis approach to augment videos for the task of remote photoplethysmography (rPPG) and study the effects of motion augmentation with respect to 1) the magnitude and 2) the type of motion. After training on motion-augmented versions of publicly available datasets, we demonstrate a 47% improvement over existing inter-dataset results using various state-of-the-art methods on the PURE dataset. We also present inter-dataset results on five benchmark datasets to show improvements of up to 79% using TS-CAN, a neural rPPG estimation method. Our findings illustrate the usefulness of motion transfer as a data augmentation technique for improving the generalization of models for camera-based physiological sensing. We release our code for using motion transfer as a data augmentation technique on three publicly available datasets, UBFC-rPPG, PURE, and SCAMPS, and models pre-trained on motion-augmented data here: https://motion-matters.github.io/ Akshay Paruchuri, Xin Liu 0034, Yulu Pan, Shwetak N. Patel, Daniel McDuff, Roni Sengupta |
WACV | 5 |
| 2023 | Do You Even Need Sensors?: Synthetic Biomusic as an Empathic TechnologyabstractPrevious research suggests that biomusic, a type of biosignal sharing, is effective at promoting empathy and closeness among individuals. However, it is unclear whether these effects are due to the information it encodes or other emotional aspects of its resulting music. To explore this question, we developed a Generative Adversarial Network (GAN) to create synthetic biomusic that approximates real biomusic, and employed deception to evaluate its effects on 24 pairs of participants engaged in real-time emotional disclosure. Users reported that both real and synthetic biomusic provided the same amount of information about their conversational partner as observing body language, facial expressions, or vocal tone. Further, both conditions increased users’ ratings of closeness and empathy with each other compared to listening to no music. However, we found no statistically significant differences between the two biomusic conditions across any of our metrics. We discuss the implications of these results for the design of future biomusic systems. Daway Chou-Ren, Mike Winters, Javier Hernandez, Daniel McDuff, Jina Suh, Vanessa Rodriguez, Gonzalo A. Ramos, Mary Czerwinski |
ACII | 4 |
| 2023 | SimPer: Simple Self-Supervised Learning of Periodic Targets
Yuzhe Yang 0003, Xin Liu 0034, Silviu Borac, Dina Katabi, Ming-Zher Poh, Daniel McDuff |
ICLR | 7 |
| 2023 | rPPG-Toolbox: Deep Remote PPG ToolboxabstractCamera-based physiological measurement is a fast growing field of computer vision. Remote photoplethysmography (rPPG) utilizes imaging devices (e.g., cameras) to measure the peripheral blood volume pulse (BVP) via photoplethysmography, and enables cardiac measurement via webcams and smartphones. However, the task is non-trivial with important pre-processing, modeling and post-processing steps required to obtain state-of-the-art results. Replication of results and benchmarking of new models is critical for scientific progress; however, as with many other applications of deep learning, reliable codebases are not easy to find or use. We present a comprehensive toolbox, rPPG-Toolbox, unsupervised and supervised rPPG models with support for public benchmark datasets, data augmentation and systematic evaluation: https://github.com/ubicomplab/rPPG-Toolbox. Xin Liu 0034, Girish Narayanswamy, Akshay Paruchuri, Jiankai Tang, Roni Sengupta, Shwetak N. Patel, Yuntao Wang 0001, Daniel McDuff |
NeurIPS | 10 |
| 2023 | EfficientPhys: Enabling Simple, Fast and Accurate Camera-Based Cardiac MeasurementabstractCamera-based physiological measurement is a growing field with neural models providing state-of-the-art performance. Prior research has explored various "end-to-end" architectures; however these methods still require several preprocessing steps and are not able to run directly on mobile and edge devices. The operations are often non-trivial to implement, making replication and deployment difficult and can even have a higher computational budget than the "core" network itself. In this paper, we propose two novel and efficient neural models for camera-based physiological measurement called EfficientPhys that remove the need for face detection, segmentation, normalization, color space transformation or any other preprocessing steps. Using an input of raw video frames, our models achieve strong accuracy on three public datasets. We show that this is the case whether using a transformer or convolutional backbone. We further evaluate the latency of the proposed networks and show that our most lightweight network also achieves a 33% improvement in efficiency. Xin Liu 0034, Brian L. Hill, Ziheng Jiang, Shwetak N. Patel, Daniel McDuff |
WACV | 5 |
| 2023 | Editorial: Special Issue on Unobtrusive Physiological Measurement Methods for Affective ApplicationsabstractIn The formative years of Affective Computing [1], from the late 1990s and into the early 2000s, a significant fraction of research attention was focused on the development of methods forunobtrusive physiological measurement. It quickly became obvious that wiring people with electrodes and strapping cumbersome hardware to their bodies was not only restricting the types of experiments that could be performed but also was not conducive to unbiased observations. For instance, subjects with fingers wrapped with electrodermal activity (EDA) and photoplethysmography (PPG) sensors could hardly type, drive or sleep comfortably. Hence, there was a need for more elegant and scalable physiological measurement methods [2]. Ioannis Pavlidis, Theodora Chaspari, Daniel McDuff |
IEEE Trans. Affect. Comput. | 3 |
| 2022 | DOC2PPT: Automatic Presentation Slides Generation from Scientific DocumentsabstractCreating presentation materials requires complex multimodal reasoning skills to summarize key concepts and arrange them in a logical and visually pleasing manner. Can machines learn to emulate this laborious process? We present a novel task and approach for document-to-slide generation. Solving this involves document summarization, image and text retrieval, slide structure and layout prediction to arrange key elements in a form suitable for presentation. We propose a hierarchical sequence-to-sequence approach to tackle our task in an end-to-end manner. Our approach exploits the inherent structures within documents and slides and incorporates paraphrasing and layout prediction modules to generate slides. To help accelerate research in this domain, we release a dataset about 6K paired documents and slide decks used in our experiments. We show that our approach outperforms strong baselines and produces slides with rich content and aligned imagery. Tsu-Jui Fu, William Yang Wang, Daniel McDuff, Yale Song |
AAAI | 3 |
| 2022 | DeepFN: Towards Generalizable Facial Action Unit Recognition with Deep Face NormalizationabstractDeployment of facial action unit recognition models has been impeded due to their limited generalization to unseen people and demographics. This work conducts an in-depth generalization analysis across several sources of variance: individuals (40 subjects), genders (male and female), skin types (darker and lighter), and databases (BP4D and DISFA). To help suppress the variance in data, we propose using self-supervised denoising autoencoders to transfer facial expressions of different people onto a common facial template which is then used to train and evaluate each of the models. We show that person-independent models yielded significantly lower performance (55% average F1 and accuracy across 40 subjects) than person-dependent models (60.3 %), leading to a generalization gap of 5.3%. However, normalizing the data with the proposed method significantly increased the performance of person-independent models (59.6%). Similarly, the proposed method was able to significantly reduce the generalization gap when considering gender (2.4%), skin type (5.3%), and dataset (9.4%). These findings represent an important step towards the creation of more generalizable facial action unit recognition systems. Javier Hernandez, Daniel McDuff, Ognjen Rudovic, Alberto Fung, Mary Czerwinski |
ACII | 2 |
| 2022 | Advancing the Understanding and Measurement of Workplace Stress in Remote Information Workers from Passive Sensors and Behavioral DataabstractWorkplace stress has been increasing in recent decades and has worsened by the unique demands imposed by COVID-19 and the new remote/hybrid work settings. High-stress working conditions can be detrimental to the health and wellness of workers and can lead to significant business costs in terms of productivity loss and medical expenses. An essential step toward managing stress involves finding comfortable ways to sense workers and recognizing stress as soon as it happens. This work explores the potential value of using pervasive sensors such as keyboards, webcams, and behavioral data such as calendar and e-mail activity to passively assess individual stress levels of work in real-life. In particular, we collected a large corpus of such data from 46 remote information workers over one month and asked them to self-report their stress levels and other relevant factors several times a day. Analysis of the data demonstrates that passive sensors can effectively detect both triggers and manifestations of workplace stress and that having access to prior data of the worker is critical for developing well-performing stress recognition models. Furthermore, we provide qualitative feedback capturing workers' preferences in workplace stress monitoring. Mehrab Bin Morshed, Javier Hernandez, Daniel McDuff, Jina Suh, Esther Howe, Kael Rowan, Marah Ihab Abdin, Gonzalo A. Ramos, Tracy Tran, Mary Czerwinski |
ACII | 3 |
| 2022 | Design of Digital Workplace Stress-Reduction Intervention Systems: Effects of Intervention Type and TimingabstractWorkplace stress-reduction interventions have produced mixed results due to engagement and adherence barriers. Leveraging technology to integrate such interventions into the workday may address these barriers and help mitigate the mental, physical, and monetary effects of workplace stress. To inform the design of a workplace stress-reduction intervention system, we conducted a four-week longitudinal study with 86 participants, examining the effects of intervention type and timing on usage, stress reduction impact, and user preferences. We compared three intervention types and two delivery timing conditions: Pre-scheduled (PS) by users and Just-in-time (JIT) prompted by the system-identified user stress-levels. We found JIT participants completed significantly more interventions than PS participants, but post-intervention and study-long stress reduction was not significantly different between conditions. Participants rated low-effort interventions highest, but high-effort interventions reduced the most stress. Participants felt JIT provided accountability but desired partial agency over timing. We present type and timing implications. Esther Howe, Jina Suh, Mehrab Bin Morshed, Daniel McDuff, Kael Rowan, Javier Hernandez, Marah Ihab Abdin, Gonzalo A. Ramos, Tracy Tran, Mary Czerwinski |
CHI | 4 |
| 2022 | "I Didn't Know I Looked Angry": Characterizing Observed Emotion and Reported Affect at WorkabstractWith the growing prevalence of affective computing applications, Automatic Emotion Recognition (AER) technologies have garnered attention in both research and industry settings. Initially limited to speech-based applications, AER technologies now include analysis of facial landmarks to provide predicted probabilities of a common subset of emotions (e.g., anger, happiness) for faces observed in an image or video frame. In this paper, we study the relationship between AER outputs and self-reports of affect employed by prior work, in the context of information work at a technology company. We compare the continuous observed emotion output from an AER tool to discrete reported affect obtained via a one-day combined tool-use and diary study (N = 15). We provide empirical evidence showing that these signals do not completely align, and find that using additional workplace context only improves alignment up to 58.6%. These results suggest affect must be studied in the context it is being expressed, and observed emotion signal should not replace internal reported affect for affective computing applications. Harmanpreet Kaur, Daniel McDuff, Alex C. Williams, Jaime Teevan, Shamsi T. Iqbal |
CHI | 2 |
| 2022 | COMPASS: Contrastive Multimodal Pretraining for Autonomous SystemsabstractLearning representations that generalize across tasks and domains is challenging yet necessary for autonomous systems. Although task-driven approaches are appealing, de-signing models specific to each application can be difficult in the face of limited data, especially when dealing with highly variable multimodal input spaces arising from different tasks in different environments. We introduce the first general-purpose pretraining pipeline, COntrastive Multimodal Pretraining for AutonomouS Systems (COMPASS), to overcome the limitations of task-specific models and existing pretraining approaches. COMPASS constructs a multimodal graph by considering the essential information for autonomous systems and the proper-ties of different modalities. Through this graph, multimodal signals are connected and mapped into two factorized spatio-temporal latent spaces: a “motion pattern space” and a “current state space.” By learning from multimodal correspondences in each latent space, COMPASS creates state representations that models necessary information such as temporal dynamics, geometry, and semantics. We pretrain COMPASS on a large-scale multimodal simulation dataset TartanAir [1] and evaluate it on drone navigation, vehicle racing, and visual odometry tasks. The experiments indicate that COMPASS can tackle all three scenarios and can also generalize to unseen environments and real-world data.11Our code implementation can be found at https://github.com/microsoft/COMPASS Sai Vemprala, Jayesh K. Gupta, Yale Song, Daniel McDuff, Ashish Kapoor |
IROS | 6 |
| 2022 | SCAMPS: Synthetics for Camera Measurement of Physiological SignalsabstractThe use of cameras and computational algorithms for noninvasive, low-cost and scalable measurement of physiological (e.g., cardiac and pulmonary) vital signs is very attractive. However, diverse data representing a range of environments, body motions, illumination conditions and physiological states is laborious, time consuming and expensive to obtain. Synthetic data have proven a valuable tool in several areas of machine learning, yet are not widely available for camera measurement of physiological states. Synthetic data offer "perfect" labels (e.g., without noise and with precise synchronization), labels that may not be possible to obtain otherwise (e.g., precise pixel level segmentation maps) and provide a high degree of control over variation and diversity in the dataset. We present SCAMPS, a dataset of synthetics containing 2,800 videos (1.68M frames) with aligned cardiac and respiratory signals and facial action intensities. The RGB frames are provided alongside segmentation maps and precise descriptive statistics about the underlying waveforms, including inter-beat interval, heart rate variability, and pulse arrival time. Finally, we present baseline results training on these synthetic data and testing on real-world datasets to illustrate generalizability. Daniel McDuff, Miah Wander, Xin Liu 0034, Brian L. Hill, Javier Hernandez, Jonathan Lester, Tadas Baltrusaitis |
NeurIPS | 1 |
| 2022 | Spectral Synthesis for Geostationary Satellite-to-Satellite TranslationabstractEarth-observing satellites carrying multispectral sensors are widely used to monitor the physical and biological states of the atmosphere, land, and oceans. These satellites have different vantage points above the Earth and different spectral imaging bands resulting in inconsistent imagery from one to another. This presents challenges in building downstream applications. What if we could generate synthetic bands for existing satellites from the union of all domains? We tackle the problem of generating synthetic spectral imagery for multispectral sensors as an unsupervised image-to-image translation problem modeled with a variational autoencoder (VAE) and generative adversarial network (GAN) architecture. Our approach introduces a novel shared spectral reconstruction loss to constrain the high-dimensional feature space of multispectral images. Simulated experiments performed by dropping one or more spectral bands show that cross-domain reconstruction outperforms measurements obtained from a second vantage point. Our proposed approach enables the synchronization of multispectral data and provides a basis for more homogeneous remote sensing datasets. Thomas Vandal, Daniel McDuff, Weile Wang, Kate Duffy, Andrew R. Michaelis, Ramakrishna R. Nemani |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Guidelines for Assessing and Minimizing Risks of Emotion Recognition ApplicationsabstractSociety has witnessed a rapid increase in the adoption of commercial uses of emotion recognition. Tools that were traditionally used by domain experts are now being used by individuals who are often unaware of the technology’s limitations and may use them in potentially harmful settings. The change in scale and agency, paired with gaps in regulation, urge the research community to rethink how we design, position, implement and ultimately deploy emotion recognition to anticipate and minimize potential risks. To help understand the current ecosystem of applied emotion recognition, this work provides an overview of some of the most frequent commercial applications and identifies some of the potential sources of harm. Informed by these, we then propose 12 guidelines for systematically assessing and reducing the risks presented by emotion recognition applications. These guidelines can help identify potential misuses and inform future deployments of emotion recognition. Javier Hernandez, Josh Lovejoy, Daniel McDuff, Jina Suh, Tim O'Brien, Arathi Sethumadhavan, Gretchen Greene, Rosalind W. Picard, Mary Czerwinski |
ACII | 3 |
| 2021 | Exploring the Effects of Virtual Agents' Smiles on Human-Agent Interaction: A Mixed-Methods StudyabstractArtificial agents’ smiling behaviour is likely to influence their likeability and the quality of user experience. While studies of human interaction highlight the importance of smile dynamics, this feature is often lacking in artificial agents, presenting a design opportunity. We developed a virtual motivational therapist with four smiling behaviours, varying in terms of quality and dynamism . We video-recorded experimental sessions with participants who posed as patients in a therapy session. The data were analysed combining a mix of quantitative and qualitative methods, focusing on participants’ own facial expressions during the interaction. Results suggest that the condition driven using data from a real therapist, where smiles are dynamic and occur at specific moments, is the most effective. We further discuss the particular importance of smile as a multipurpose emotional display in human-machine interaction. Ilaria Torre 0002, Sylvaine Tuncer, Daniel McDuff, Mary Czerwinski |
ACII | 3 |
| 2021 | Understanding Conversational and Expressive Style in a Multimodal Embodied Conversational AgentabstractEmbodied conversational agents have changed the ways we can interact with machines. However, these systems often do not meet users’ expectations. A limitation is that the agents are monotonic in behavior and do not adapt to an interlocutor. We present SIVA (a Socially Intelligent Virtual Agent), an expressive, embodied conversational agent that can recognize human behavior during open-ended conversations and automatically align its responses to the conversational and expressive style of the other party. SIVA leverages multimodal inputs to produce rich and perceptually valid responses (lip syncing and facial expressions) during the conversation. We conducted a user study (N=30) in which participants rated SIVA as being more empathetic and believable than the control (agent without style matching). Based on almost 10 hours of interaction, participants who preferred interpersonal involvement evaluated SIVA as significantly more animate than the participants who valued consideration and independence. Deepali Aneja, Jessie Hoegen, Daniel McDuff, Mary Czerwinski |
CHI | 3 |
| 2021 | AffectiveSpotlight: Facilitating the Communication of Affective Responses from Audience Members during Online PresentationsabstractThe ability to monitor audience reactions is critical when delivering presentations. However, current videoconferencing platforms offer limited solutions to support this. This work leverages recent advances in affect sensing to capture and facilitate communication of relevant audience signals. Using an exploratory survey (N=175), we assessed the most relevant audience responses such as confusion, engagement, and head-nods. We then implemented AffectiveSpotlight, a Microsoft Teams bot that analyzes facial responses and head gestures of audience members and dynamically spotlights the most expressive ones. In a within-subjects study with 14 groups (N=117), we observed that the system made presenters significantly more aware of their audience, speak for a longer period of time, and self-assess the quality of their talk more similarly to the audience members, compared to two control conditions (randomly-selected spotlight and default platform UI). We provide design recommendations for future affective interfaces for online presentations based on feedback from the study. Prasanth Murali, Javier Hernandez, Daniel McDuff, Kael Rowan, Jina Suh, Mary Czerwinski |
CHI | 3 |
| 2021 | MeetingCoach: An Intelligent Dashboard for Supporting Effective & Inclusive MeetingsabstractVideo-conferencing is essential for many companies, but its limitations in conveying social cues can lead to ineffective meetings. We present MeetingCoach, an intelligent post-meeting feedback dashboard that summarizes contextual and behavioral meeting information. Through an exploratory survey (N=120), we identified important signals (e.g., turn taking, sentiment) and used these insights to create a wireframe dashboard. The design was evaluated with in situ participants (N=16) who helped identify the components they would prefer in a post-meeting dashboard. After recording video-conferencing meetings of eight teams over four weeks, we developed an AI system to quantify the meeting features and created personalized dashboards for each participant. Through interviews and surveys (N=23), we found that reviewing the dashboard helped improve attendees’ awareness of meeting dynamics, with implications for improved effectiveness and inclusivity. Based on our findings, we provide suggestions for future feedback system designs of video-conferencing meetings. Samiha Samrose, Daniel McDuff, Robert Sim, Jina Suh, Kael Rowan, Javier Hernandez, Sean Rintel, Kevin Moynihan, Mary Czerwinski |
CHI | 2 |
| 2021 | The Benefit of Distraction: Denoising Camera-Based Physiological Measurements using Inverse AttentionabstractAttention networks perform well on diverse computer vision tasks. The core idea is that the signal of interest is stronger in some pixels ("foreground"), and by selectively focusing computation on these pixels, networks can extract subtle information buried in noise and other sources of corruption. Our paper is based on one key observation: in many real-world applications, many sources of corruption, such as illumination and motion, are often shared between the "foreground" and the "background" pixels. Can we utilize this to our advantage? We propose the utility of inverse attention networks, which focus on extracting information about these shared sources of corruption. We show that this helps to effectively suppress shared covariates and amplify signal information, resulting in improved performance. We illustrate this on the task of camera-based physiological measurement where the signal of interest is weak and global illumination variations and motion act as significant shared sources of corruption. We perform experiments on three datasets and show that our approach of inverse attention produces state-of-the-art results, increasing the signal-to-noise ratio by up to 5.8 dB, reducing heart rate and breathing rate estimation errors by as much as 30 %, recovering subtle waveform dynamics, and generalizing from RGB to NIR videos without retraining. Ewa Magdalena Nowara, Daniel McDuff, Ashok Veeraraghavan |
ICCV | 2 |
| 2021 | Active Contrastive Learning of Audio-Visual Video Representations
Zhaoyang Zeng, Daniel McDuff, Yale Song |
ICLR | 3 |
| 2021 | Modeling Affect-based Intrinsic Rewards for Exploration and LearningabstractPositive affect has been linked to increased interest, curiosity and satisfaction in human learning. In reinforcement learning, extrinsic rewards are often sparse and difficult to define, intrinsically motivated learning can help address these challenges. We argue that positive affect is an important intrinsic reward that effectively helps drive exploration that is useful in gathering experiences. We present a novel approach leveraging a task-independent reward function trained on spontaneous smile behavior that reflects the intrinsic reward of positive affect. To evaluate our approach we trained several downstream computer vision tasks on data collected with our policy and several baseline methods. We show that the policy based on our affective rewards successfully increases the duration of episodes, the area explored and reduces collisions. The impact is the increased speed of learning for several downstream computer vision tasks. Dean Zadok, Daniel McDuff, Ashish Kapoor |
ICRA | 2 |
| 2021 | Contrastive Learning of Global and Local Video RepresentationsabstractContrastive learning has delivered impressive results for various tasks in the self-supervised regime. However, existing approaches optimize for learning representations specific to downstream scenarios, i.e., global representations suitable for tasks such as classification or local representations for tasks such as detection and localization. While they produce satisfactory results in the intended downstream scenarios, they often fail to generalize to tasks that they were not originally designed for. In this work, we propose to learn video representations that generalize to both the tasks which require global semantic information (e.g., classification) and the tasks that require local fine-grained spatio-temporal information (e.g., localization). We achieve this by optimizing two contrastive objectives that together encourage our model to learn global-local visual information given audio signals. We show that the two objectives mutually improve the generalizability of the learned global-local representations, significantly outperforming their disjointly learned counterparts. We demonstrate our approach on various tasks including action/sound classification, lipreading, deepfake detection, event and sound localization. Zhaoyang Zeng, Daniel McDuff, Yale Song |
NeurIPS | 3 |
| 2021 | Do Affective Cues Validate Behavioural Metrics for Search?abstractTraces of searcher behaviour, such as query reformulation or clicks, are commonly used to evaluate a running search engine. The underlying expectation is that these behaviours are proxies for something more important, such as relevance, utility, or satisfaction. Affective computing technology gives us the tools to help confirm some of these expectations, by examining visceral expressive responses during search sessions. However, work to date has only studied small populations in laboratory settings and with a limited number of contrived search tasks. In this study, we analysed longitudinal, in-situ, search behaviours of 152 information workers, over the course of several weeks while simultaneously tracking their facial expressions. Results from over 20,000 search sessions and 45,000 queries allow us to observe that indeed affective expressions are consistent with, and complementary to, existing "click-based'' metrics. On a query-level, searches that result in a short dwell time are associated with a decrease in smiles (expressions of "happiness'') and that if a query is reformulated the results of the reformulation are associated with an increase in smiling---suggesting a positive outcome as people converge on the information they need. On a session-level, sessions that feature reformulations are more commonly associated with fewer smiles and more furrowed brows (expressions of "anger/frustration''). Similarly, sessions with short-dwell clicks are also associated with fewer smiles. These data provide an insight into visceral aspects of search experience and present a new dimension for evaluating engine performance. Daniel McDuff, Paul Thomas 0001, Nick Craswell, Kael Rowan, Mary Czerwinski |
SIGIR | 1 |
| 2021 | Faded Smiles? A Largescale Observational Study of Smiling from Adolescence to Old AgeabstractA relatively large body of work exists examining sex differences in expressiveness; however, there remains little research of differences in expressiveness associated with aging. Observational studies of facial expressivity across ages are limited in part due to the poor scalability of traditional research methods. We collected over 17,000 videos of natural facial behavior using the Internet and performed a large observational study of smiling responses of people ages 18 to 70 years. Using automated facial coding we quantified the presence of smiles as people watched a set of controlled mundane online content. The likelihood of smiles and the duration of smiles increased with age. We attribute this to greater expression of positive emotion in older people. Women smiled more than men over all and gender differences increased significantly with age. We question whether results may be influenced by the effect of age on the accuracy of the automated smile detection; however, validation on a large set of human coded videos shows that the observed effects were not due to smile detection performance. Daniel McDuff, Stephanie Glass |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | Longitudinal Observational Evidence of the Impact of Emotion Regulation Strategies on Affective ExpressionabstractThe ability to regulate our emotions plays an important role in our psychological and physical health. Regulating emotions influences how and when emotions are expressed. We performed a large scale, longitudinal observational study to investigate the effect of emotion regulation ability on expressed affect. We found that expression of negative affect increased throughout the day. For people who suppress emotion this increase is slower that for those who do not. For those with stronger cognitive reappraisal abilities, though not significant, there was a trend for higher positive affect and negative affect increased significantly less steeply, suggesting that they might experience more positive and less negative affect. These results reflect some of the first results based on large scale, continuous tracking of behavioral expression of emotion longitudinally. Our results demonstrate the need to carefully consider the time of day and emotion regulation ability, in addition to gender and age, when attempting to automatically infer affective states for facial behavior. Daniel McDuff, Eunice Jun, Kael Rowan, Mary Czerwinski |
IEEE Trans. Affect. Comput. | 1 |
| 2021 | Guest Editorial: Camera-Based Monitoring for Pervasive Healthcare InformaticsabstractThe papers in this special section focus on camera-based monitoring for pervasive healthcare informatics. Measuring physiological signals from the human face and body using video cameras is an emerging research topic that has grown rapidly in the last decade. Remote cameras (in both visible and infrared wavelengths) can be used to measure vital signs from a human body based on skin optics or body movements thereby avoiding mechanical contact with the skin. Camera-based health monitoring will bring a rich set of compelling healthcare applications that directly improve upon contact-based monitoring solutions and impact people’s care experience and quality of life in various scenarios, such as in hospital care units, sleep/senior centers, assisted-living homes, telemedicine and e-health, baby/elderly care at home, fitness and sports, driver monitoring in automotive applications, cardiac/ respiratory gating for MRI/CT, AR/VR based therapy and clinical training, e Wenjin Wang 0002, Steffen Leonhardt, Lionel Tarassenko, Caifeng Shan, Daniel McDuff |
IEEE J. Biomed. Health Informatics | 5 |
| 2021 | DeepMag: Source-Specific Change Magnification Using Gradient AscentabstractMany important physical phenomena involve subtle signals that are difficult to observe with the unaided eye, yet visualizing them can be very informative. Current motion magnification techniques can reveal these small temporal variations in video, but require precise prior knowledge about the target signal, and cannot deal with interference motions at a similar frequency. We present DeepMag, an end-to-end deep neural video-processing framework based on gradient ascent that enables automated magnification of subtle color and motion signals from a specific source, even in the presence of large motions of various velocities. The advantages of DeepMag are highlighted via the task of video-based physiological visualization. Through systematic quantitative and qualitative evaluation of the approach on videos with different levels of head motion, we compare the magnification of pulse and respiration to existing state-of-the-art methods. Our method produces magnified videos with substantially fewer artifacts and blurring whilst magnifying the physiological changes by a similar degree. Weixuan 'Vincent' Chen, Daniel McDuff |
ACM Trans. Graph. | 2 |
| 2021 | Theories of Conversation for Conversational IRabstractConversational information retrieval is a relatively new and fast-developing research area, but conversation itself has been well studied for decades. Researchers have analysed linguistic phenomena such as structure and semantics but also paralinguistic features such as tone, body language, and even the physiological states of interlocutors. We tend to treat computers as social agents—especially if they have some humanlike features in their design—and so work from human-to-human conversation is highly relevant to how we think about the design of human-to-computer applications. In this article, we summarise some salient past work, focusing on social norms; structures; and affect, prosody, and style. We examine social communication theories briefly as a review to see what we have learned about how humans interact with each other and how that might pertain to agents and robots. We also discuss some implications for research and design of conversational IR systems. Paul Thomas 0001, Mary Czerwinski, Daniel McDuff, Nick Craswell |
ACM Trans. Inf. Syst. | 3 |
| 2020 | Lessons Learned in Designing AI for Autistic AdultsabstractThrough an iterative design process using Wizard of Oz (WOz) prototypes, we designed a video calling application for people with Autism Spectrum Disorder. Our Video Calling for Autism prototype provided an Expressiveness Mirror that gave feedback to autistic people on how their facial expressions might be interpreted by their neurotypical conversation partners. This feedback was in the form of emojis representing six emotions and a bar indicating the amount of overall expressiveness demonstrated by the user. However, when we built a working prototype and conducted a user study with autistic participants, their negative feedback caused us to reconsider how our design process led to a prototype that they did not find useful. We reflect on the design challenges around developing AI technology for an autistic user population, how Wizard of Oz prototypes can be overly optimistic in representing AI-driven prototypes, how autistic research participants can respond differently to user experience prototypes of varying fidelity, and how designing for people with diverse abilities needs to include that population in the development process. Andrew Begel, John C. Tang, Sean Andrist, Michael Barnett 0001, Tony Carbary, Piali Choudhury, Edward Cutrell, Alberto Fung, Sasa Junuzovic, Daniel McDuff, Kael Rowan, Shibashankar Sahoo, Jennifer Frances Waldern, Jessica Wolk, Annuska Z. Perkins |
ASSETS | 10 |
| 2020 | Optimizing for Happiness and Productivity: Modeling Opportune Moments for Transitions and Breaks at WorkabstractInformation workers perform jobs that demand constant multitasking, leading to context switches, productivity loss, stress, and unhappiness. Systems that can mediate task transitions and breaks have the potential to keep people both productive and happy. We explore a crucial initial step for this goal: finding opportune moments to recommend transitions and breaks without disrupting people during focused states. Using affect, workstation activity, and task data from a three-week field study (N=25), we build models to predict whether a person should continue their task, transition to a new task, or take a break. The R-squared values of our models are as high as 0.7, with only 15% error cases. We ask users to evaluate the timing of recommendations provided by a recommender that relies on these models. Our study shows that users find our transition and break recommendations to be well-timed, rating them as 86% and 77% accurate, respectively. We conclude with a discussion of the implications for intelligent systems that seek to guide task transitions and manage interruptions at work. Harmanpreet Kaur, Alex C. Williams, Daniel McDuff, Mary Czerwinski, Jaime Teevan, Shamsi T. Iqbal |
CHI | 3 |
| 2020 | Spatio-Temporal Attention and Magnification for Classification of Parkinson's Disease from Videos Collected via the InternetabstractWe present an automated framework for detecting Parkinson's disease (PD) from videos collected through a scalable online platform. We analyzed 1380 videos of age-matched participants performing four standard motor tasks from the MDS-UPDRS. Our proposed framework leverages multiple deep neural networks to temporally and spatially segment the videos as well as magnify relevant motions. Frequency domain representations of the resulting data are then classified using supervised learning. Overall, the proposed framework achieves an accuracy of 82.5% when discriminating between those with PD and those without, and 61.8% when discriminating between those with PD with treatment, with PD without treatment, and those without PD. These results increased up to 91.8% and 73.5%, respectively, when combining the predictions of multiple models. To understand the contributions of each part of our framework we perform systematic ablation studies. We also compare between motion features based on pixel, phase-based and deep learning-based representations. This work demonstrates the possibility of identifying PD cues in challenging real-life settings with inexpensive webcams. Mohammad Rafayet Ali, Javier Hernandez, Earl Ray Dorsey, Mohammed E. Hoque 0001, Daniel McDuff |
FG | 5 |
| 2020 | Multi-Reference Neural TTS Stylization with Adversarial Cycle ConsistencyabstractCurrent multi-reference style transfer models for Text-to-Speech (TTS) perform sub-optimally on disjoints datasets, where one dataset contains only a single style class for one of the style dimensions.These models generally fail to produce style transfer for the dimension that is underrepresented in the dataset.In this paper, we propose an adversarial cycle consistency training scheme with paired and unpaired triplets to ensure the use of information from all style dimensions.During training, we incorporate unpaired triplets with randomly selected reference audio samples and encourage the synthesized speech to preserve the appropriate styles using adversarial cycle consistency.We use this method to transfer emotion from a dataset containing four emotions to a dataset with only a single emotion.This results in a 78% improvement in style transfer (based on emotion classification) with minimal reduction in fidelity and naturalness.In subjective evaluations our method was consistently rated as closer to the reference style than the baseline.Synthesized speech samples are available at: https://sites.google. Matt Whitehill, Daniel McDuff, Yale Song |
INTERSPEECH | 3 |
| 2020 | Design and evaluation of intelligent agent prototypes for assistance with focus and productivity at workabstractCurrent research on building intelligent agents for aiding with productivity and focus in the workplace is quite limited, despite the ubiquity of information workers across the globe. In our work, we present a productivity agent which helps users schedule and block out time on their calendar to focus on important tasks, monitor and intervene with distractions, and reflect on their daily mood and goals in a single, standalone application. We created two different prototype versions of our agent: a text-based (TB) agent with a similar UI to a standard chatbot, and a more emotionally expressive virtual agent (VA) that employs a video avatar and the ability to detect and respond appropriately to users' emotions. We evaluated these two agent prototypes against an existing product (control) condition through a three-week, within subjects study design with 40 participants, across different work roles in a large organization. We found that participants scheduled 134% more time with the TB prototype, and 110% more time with the VA prototype for focused tasks compared to the control condition. Users reported that they felt more satisfied and productive with the VA agent. However, The perception of anthropomorphism in the VA was polarized, with several participants suggesting that the human appearance was unnecessary. We discuss important insights from our work for the future design of conversational agents for productivity, wellbeing, and focus in the workplace. Ted Grover, Kael Rowan, Jina Suh, Daniel McDuff, Mary Czerwinski |
IUI | 4 |
| 2020 | Conversational Error Analysis in Human-Agent InteractionabstractConversational Agents (CAs) present many opportunities for changing how we interact with information and computer systems in a more natural, accessible way. Building on research in machine learning and HCI, it is now possible to design and test multi-turn CA that is capable of extended interactions. However, there are many ways in which these CAs can "fail" and fall short of human expectations. We systematically analyzed how five different types of conversational errors impacted perceptions of an embodied CA. Not all errors negatively impacted perceptions of the agent. Repetitions by the agent and clarifications by the human significantly decreased the perceived intelligence and anthropomorphism of the agent. Turn-taking errors significantly decreased the likability of the agent. However, coherence errors significantly positively increased likability, and these errors were also associated with positive valence via facial expressions, suggesting that the users found them amusing. We believe this work is the first to identify that turn-taking, repetition, clarification, and coherence errors directly affect users' acceptance of an embodied CA, and are worth taking note by designers of such systems during dialog configuration. We release the Agent Conversational Error (ACE) dataset, a set of transcripts and error annotations of human-agent conversations. The dataset can be found at the GITHUB link: https://github.com/deepalianeja/ACE-dataset Deepali Aneja, Daniel McDuff, Mary Czerwinski |
IVA | 2 |
| 2020 | Multi-Task Temporal Shift Attention Networks for On-Device Contactless Vitals MeasurementabstractTelehealth and remote health monitoring have become increasingly important during the SARS-CoV-2 pandemic and it is widely expected that this will have a lasting impact on healthcare practices. These tools can help reduce the risk of exposing patients and medical staff to infection, make healthcare services more accessible, and allow providers to see more patients. However, objective measurement of vital signs is challenging without direct contact with a patient. We present a video-based and on-device optical cardiopulmonary vital sign measurement approach. It leverages a novel multi-task temporal shift convolutional attention network (MTTS-CAN) and enables real-time cardiovascular and respiratory measurements on mobile platforms. We evaluate our system on an Advanced RISC Machine (ARM) CPU and achieve state-of-the-art accuracy while running at over 150 frames per second which enables real-time applications. Systematic experimentation on large benchmark datasets reveals that our approach leads to substantial (20%-50%) reductions in error and generalizes well across datasets. Xin Liu 0034, Josh Fromm, Shwetak N. Patel, Daniel McDuff |
NeurIPS | 4 |
| 2020 | Expressions of Style in Information Seeking Conversation with an AgentabstractPast work in information-seeking conversation has demonstrated that people exhibit different conversational styles---for example, in word choice or prosody---that differences in style lead to poorer conversations, and that partners actively align their styles over time. One might assume that this would also be true for conversations with an artificial agent such as Cortana, Siri, or Alexa; and that agents should therefore track and mimic a user's style. We examine this hypothesis with reference to a lab study, where 24 participants carried out relatively long information-seeking tasks with an embodied conversational agent. The agent combined topical language models with a conversational dialogue engine, style recognition and alignment modules. We see that "style'' can be measured in human-to-agent conversation, although it looks somewhat different to style in human-to-human conversation and does not correlate with self-reported preferences. There is evidence that people align their style to the agent, and that conversations run more smoothly if the agent detects, and aligns to, the human's style as well. Paul Thomas 0001, Daniel McDuff, Mary Czerwinski, Nick Craswell |
SIGIR | 2 |
| 2019 | EMMA: An Emotion-Aware Wellbeing ChatbotabstractThe delivery of mental health interventions via ubiquitous devices has shown much promise. A conversational chatbot is a promising oracle for delivering appropriate just-in-time interventions. However, designing emotionally-aware agents, specially in this context, is under-explored. Furthermore, the feasibility of automating the delivery of just-in-time mHealth interventions via such an agent has not been fully studied. In this paper, we present the design and evaluation of EMMA (EMotion-Aware mHealth Agent) through a two-week long human-subject experiment with N=39 participants. EMMA provides emotionally appropriate micro-activities in an empathetic manner. We show that the system can be extended to detect a user's mood purely from smartphone sensor data. Our results show that our personalized machine learning model was perceived as likable via self-reports of emotion from users. Finally, we provide a set of guidelines for the design of emotion-aware bots for mHealth. Asma Ghandeharioun, Daniel McDuff, Mary Czerwinski, Kael Rowan |
ACII | 2 |
| 2019 | Towards Understanding Emotional Intelligence for Behavior Change ChatbotsabstractA natural conversational interface that allows longitudinal symptom tracking would be extremely valuable in health/wellness applications. However, the task of designing emotionally-aware agents for behavior change is still poorly understood. In this paper, we present the design and evaluation of an emotion-aware chatbot that conducts experience sampling in an empathetic manner. We evaluate it through a human-subject experiment with N=39 participants over the course of a week. Our results show that extraverts preferred the emotion-aware chatbot significantly more than introverts. Also, participants reported a higher percentage of positive mood reports when interacting with the empathetic bot. Finally, we provide guidelines for the design of emotion-aware chatbots for potential use in mHealth contexts. Asma Ghandeharioun, Daniel McDuff, Mary Czerwinski, Kael Rowan |
ACII | 2 |
| 2019 | A Conversational Agent in Support of Productivity and Wellbeing at WorkabstractConversational agents have the potential to support users in many tasks. However, support for productivity and well-being in the workplace has received little attention. We present the first design of a conversational system that supports information workers with multiple work-related goals, informed by a survey of the current and potential use of conversational agents in the workplace. The goals of this research include the evaluation of using an agent for scheduling and prioritizing tasks, switching tasks, providing break reminders, dealing with social media distractions and for end of the day reflection on tasks accomplished. We deployed a chat-based intelligent agent, named Amber, in a field study with 24 information workers over the course of 6 days. We present our preliminary findings from the field study and discuss implications for the design of future workplace conversational agents. Everlyne Kimani, Kael Rowan, Daniel McDuff, Mary Czerwinski, Gloria Mark |
ACII | 3 |
| 2019 | Democratizing Psychological Insights from Analysis of Nonverbal BehaviorabstractThe affective computing community has invested heavily in building automated tools for the analysis of facial behavior and the expression of emotion. These tools present a valuable, but largely untapped, opportunity for social scientists to perform observational analyses of nonverbal behavior at very large scale. Various tech companies are collecting huge corpora of images and videos from around the world that could be used to study important scientific questions. However, privacy restrictions and intellectual property concerns render these data inaccessible to most academics. Unfortunately, this limits the potential for scientific advancement and leads to the consolidation of data and opportunity into the hands of a few powerful institutions. In this paper, we ask whether similar psychological insights can be gained by analyzing smaller, public datasets that are more within reach for academic researchers. As a proof-of-concept for this idea, we gather, analyze, and release a corpus of public images and metadata and use it to replicate recent psychological findings about smiling, gender, and culture. In so doing, we provide evidence that psychological insights can indeed by democratized through the automated analysis of nonverbal behavior. Daniel McDuff, Jeffrey M. Girard |
ACII | 1 |
| 2019 | Unpaired Image-to-Speech Synthesis With Multimodal Information BottleneckabstractDeep generative models have led to significant advances in cross-modal generation such as text-to-image synthesis. Training these models typically requires paired data with direct correspondence between modalities. We introduce the novel problem of translating instances from one modality to another without paired data by leveraging an intermediate modality shared by the two other modalities. To demonstrate this, we take the problem of translating images to speech. In this case, one could leverage disjoint datasets with one shared modality, e.g., image-text pairs and text-speech pairs, with text as the shared modality. We call this problem “skip-modal generation” because the shared modality is skipped during the generation process. We propose a multimodal information bottleneck approach that learns the correspondence between modalities from unpaired data (image and speech) by leveraging the shared modality (text). We address fundamental challenges of skip-modal generation: 1) learning multimodal representations using a single model, 2) bridging the domain gap between two unrelated datasets, and 3) learning the correspondence between modalities from unpaired data. We show qualitative results on image-to-speech synthesis; this is the first time such results have been reported in the literature. We also show that our approach improves performance on traditional cross-modal generation, suggesting that it improves data efficiency in solving individual tasks. Daniel McDuff, Yale Song |
ICCV | 2 |
| 2019 | Neural TTS Stylization with Adversarial and Collaborative Games
Daniel McDuff, Yale Song |
ICLR (Poster) | 2 |
| 2019 | Visceral Machines: Risk-Aversion in Reinforcement Learning with Intrinsic Physiological Rewards
Daniel McDuff, Ashish Kapoor |
ICLR (Poster) | 1 |
| 2019 | A High-Fidelity Open Embodied Avatar with Lip Syncing and Expression CapabilitiesabstractEmbodied avatars as virtual agents have many applications and provide benefits over disembodied agents, allowing nonverbal social and interactional cues to be leveraged, in a similar manner to how humans interact with each other. We present an open embodied avatar built upon the Unreal Engine that can be controlled via a simple python programming interface. The avatar has lip syncing (phoneme control), head gesture and facial expression (using either facial action units or cardinal emotion categories) capabilities. We release code and models to illustrate how the avatar can be controlled like a puppet or used to create a simple conversational agent using public application programming interfaces (APIs). GITHUB link: https://github.com/danmcduff/AvatarSim Deepali Aneja, Daniel McDuff, Shital Shah |
ICMI | 2 |
| 2019 | An End-to-End Conversational Style Matching AgentabstractWe present an end-to-end voice-based conversational agent that is able to engage in naturalistic multi-turn dialogue and align with the interlocutor's conversational style. The system uses a series of deep neural network components for speech recognition, dialogue generation, prosodic analysis and speech synthesis to generate language and prosodic expression with qualities that match those of the user. We conducted a user study (N=30) in which participants talked with the agent for 15 to 20 minutes, resulting in over 8 hours of natural interaction data. Users with high consideration conversational styles reported the agent to be more trustworthy when it matched their conversational style. Whereas, users with high involvement conversational styles were indifferent. Finally, we provide design guidelines for multi-turn dialogue interactions using conversational style adaptation. Jessie Hoegen, Deepali Aneja, Daniel McDuff, Mary Czerwinski |
IVA | 3 |
| 2019 | Characterizing Bias in Classifiers using Generative ModelsabstractModels that are learned from real-world data are often biased because the data used to train them is biased. This can propagate systemic human biases that exist and ultimately lead to inequitable treatment of people, especially minorities. To characterize bias in learned classifiers, existing approaches rely on human oracles labeling real-world examples to identify the "blind spots" of the classifiers; these are ultimately limited due to the human labor required and the finite nature of existing image examples. We propose a simulation-based approach for interrogating classifiers using generative adversarial models in a systematic manner. We incorporate a progressive conditional generative model for synthesizing photo-realistic facial images and Bayesian Optimization for an efficient interrogation of independent facial image classification systems. We show how this approach can be used to efficiently characterize racial and gender biases in commercial systems. Daniel McDuff, Yale Song, Ashish Kapoor |
NeurIPS | 1 |
| 2019 | Circadian Rhythms and Physiological Synchrony: Evidence of the Impact of Diversity on Small Group CreativityabstractCircadian rhythms determine daily sleep cycles, mood, and cognition. Depending on an individual's circadian preference, or chronotype (i.e.,"early birds" and "night owls"), the rhythms shift earlier or later in the day. Early birds experience circadian arousal peaks earlier in the morning than night owls. Prior work has shown that individuals are more effective at analytic tasks during their peak arousal times but are more creative during their off-peak times. We investigate if these findings hold true for small groups. We find that time of day and a group's majority chronotype impact performance on analytic and creative tasks. Physiological synchrony among group members positively predicts group satisfaction. Specifically, homogeneous groups perform worse on all tasks regardless of time of day, but they achieve greater physiological synchrony and feel more satisfied as a group. Based on these findings, we present and advocate for a temporal dimension of group diversity. Eunice Jun, Daniel McDuff, Mary Czerwinski |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2019 | Accessible Video Calling: Enabling Nonvisual Perception of Visual Conversation CuesabstractNonvisually Accessible Video Calling (NAVC) is a prototype that detects visual conversation cues in a video call and uses audio cues to convey them to a user who is blind or low-vision. NAVC uses audio cues inspired by movie soundtracks to convey Attention, Agreement, Disagreement, Happiness, Thinking, and Surprise. When designing NAVC, we partnered with people who are blind or low-vision through a user-centered design process that included need-finding interviews and design reviews. To evaluate NAVC, we conducted a user study with 16 participants. The study provided feedback on the NAVC prototype and showed that the participants could easily discern some cues, like Attention and Agreement, but had trouble distinguishing others. The accuracy of the prototype in detecting conversation cues emerged as a key concern, especially in avoiding false positives and in detecting negative emotions, which tend to be masked in social conversations. This research identified challenges and design opportunities in using AI models to enable accessible video calling. Lei Shi 0020, Brianna J. Tomlinson, John C. Tang, Edward Cutrell, Daniel McDuff, Gina Venolia, Paul Johns, Kael Rowan |
Proc. ACM Hum. Comput. Interact. | 5 |
| 2019 | Managing Stress: The Needs of Autistic Adults in Video CallingabstractVideo calling (VC) aims to create multi-modal, collaborative environments that are "just like being there." However, we found that autistic individuals, who exhibit atypical social and cognitive processing, may not share this goal. We interviewed autistic adults about their perceptions of VC compared to other computer- mediated communications (CMC) and face-to-face interactions. We developed a neurodiversity-sensitive model of CMC that describes how stressors such as sensory sensitivities, cognitive load, and anxiety, contribute to their preferences for CMC channels. We learned that they apply significant effort to construct coping strategies to support their sensory, cognitive, and social needs. These strategies include moderating their sensory inputs, creating mental models of conversation partners, and attempting to mask their autism by adopting neurotypical behaviors. Without effective strategies, interviewees experience more stress, have less capacity to interpret verbal and non-verbal cues, and feel less empowered to participate. Our findings reveal critical needs for autistic users. We suggest design opportunities to support their ability to comfortably use VC, and in doing so, point the way towards making VC more comfortable for all. Annuska Z. Perkins, Andrew Begel, Jennifer Frances Waldern, John C. Tang, Michael Barnett 0001, Edward Cutrell, Daniel McDuff, Sean Andrist, Meredith Ringel Morris |
Proc. ACM Hum. Comput. Interact. | 7 |
| 2019 | AM-FED+: An Extended Dataset of Naturalistic Facial Expressions Collected in Everyday SettingsabstractPublic datasets have played a significant role in advancing the state-of-the-art in automated facial coding. Many of these datasets contain posed expressions and/or videos recorded in controlled lab conditions with little variation in lighting or head pose. As such, the data do not reflect the conditions observed in many real-world applications. We present AM-FED+ an extended dataset of naturalistic facial response videos collected in everyday settings. The dataset contains 1,044 videos of which 545 videos (263,705 frames or 21,859 seconds) have been comprehensively manually coded for facial action units. These videos act as a challenging benchmark for automated facial coding systems. All the videos contain gender labels and a large subset (77 percent) contain age and country information. Subject self-reported liking and familiarity with the stimuli are also included. We provide automated facial landmark detection locations for the videos. Finally, baseline action unit classification results are presented for the coded videos. The dataset is available to download online: https://www.affectiva.com/facial-expression-dataset/. Daniel McDuff, May Amr, Rana El Kaliouby |
IEEE Trans. Affect. Comput. | 1 |
| 2019 | Wearable Motion-Based Heart Rate at Rest: A Workplace EvaluationabstractThis paper studies the feasibility of using low-cost motion sensors to provide opportunistic heart rate assessments from ballistocardiographic signals during restful periods of daily life. Three wearable devices were used to capture peripheral motions at specific body locations (head, wrist, and trouser pocket) of 15 participants during five regular workdays each. Three methods were implemented to extract heart rate from motion data and their performance was compared to those obtained with an FDA-cleared device. With a total of 1358 h of naturalistic sensor data, our results show that providing accurate heart rate estimations from peripheral motion signals is possible during relatively "still" moments. In our real-life workplace study, the head-mounted device yielded the most frequent assessments (22.98% of the time under 5 beats per minute of error) followed by the smartphone in the pocket (5.02%) and the wrist-worn device (3.48%). Most importantly, accurate assessments were automatically detected by using a custom threshold based on the device jerk. Due to the pervasiveness and low cost of wearable motion sensors, this paper demonstrates the feasibility of providing opportunistic large-scale low-cost samples of resting heart rate. Javier Hernandez, Daniel McDuff, Karen S. Quigley, Pattie Maes, Rosalind W. Picard |
IEEE J. Biomed. Health Informatics | 2 |
| 2018 | Emotional Dialogue Generation using Image-Grounded Language ModelsabstractComputer-based conversational agents are becoming ubiquitous. However, for these systems to be engaging and valuable to the user, they must be able to express emotion, in addition to providing informative responses. Humans rely on much more than language during conversations; visual information is key to providing context. We present the first example of an image-grounded conversational agent using visual sentiment, facial expression and scene features. We show that key qualities of the generated dialogue can be manipulated by the features used for training the agent. We evaluate our model on a large and very challenging real-world dataset of conversations from social media (Twitter). The image-grounding leads to significantly more informative, emotional and specific responses, and the exact qualities can be tuned depending on the image features used. Furthermore, our model improves the objective quality of dialogue responses when evaluated on standard natural language metrics. Bernd Huber, Daniel McDuff, Chris Brockett, Michel Galley, William B. Dolan |
CHI | 2 |
| 2018 | Style and Alignment in Information-Seeking ConversationabstractAnalysis of casual chit-chat indicates that differences in conversational style---the way things are said---can significantly impact a participants» impressions of the conversation and of each other. However, prior work has not systematically analyzed how important style is in task-oriented, information-seeking exchanges of the sort we might have with a conversational search agent. We examine recordings from the MISC data set, where pairs of "users" and "intermediaries" collaborate on information-seeking tasks, and look for indications of style which can be computed at scale. We find that stylistic markers identified by Tannen in casual chat do exist in information-seeking dialogue, and that participants can be arranged along a single stylistic dimension: "considerate" to "involved". This labelling for style needs no manual intervention. Furthermore, we find that there is no clear best style; but that differences in style, previously thought to impede communication, are only a problem for shorter tasks. This result is likely due to alignment of conversational style over the course of an interaction. Paul Thomas 0001, Mary Czerwinski, Daniel McDuff, Nick Craswell, Gloria Mark |
CHIIR | 3 |
| 2018 | DeepPhys: Video-Based Physiological Measurement Using Convolutional Attention Networks
Weixuan 'Vincent' Chen, Daniel McDuff |
ECCV (2) | 2 |
| 2018 | Facial Expression Grounded Conversational Dialogue GenerationabstractWe present a novel conversational language model that is grounded with information about facial expressions. To our knowledge this is the first in-depth examination of grounding natural language models with facial cues. We train a neural language model that uses automatically detected facial action unit intensity information in images alongside text to generate conversational dialogue. We evaluate our model on a large and very challenging unconstrained real-world dataset from social media (Twitter), featuring 450,000 conversations with associated facial expressions. Systematic linguistic and crowdsourced analyses reveal the properties of our models: The facial expression grounding strengthens the sentiment of the resulting dialogue such that it is consistent with the valence of the facial expressions. Furthermore, the automatically generated conversational responses are rated as equivalent to the human gold-standard responses on relevance and emotion dimensions. Bernd Huber, Daniel McDuff |
FG | 2 |
| 2017 | Smiling from adolescence to old age: A large observational studyabstractThe following topics are dealt with: emotion recognition; learning (artificial intelligence); psychology; feature extraction; face recognition; behavioural sciences computing; human computer interaction; neural nets; speech recognition; medical signal processing. Daniel McDuff |
ACII | 1 |
| 2017 | Historical Heterogeneity Predicts Smiling: Evidence from Large-Scale Observational AnalysesabstractFacial behavior is a valuable source of information about an individual's feelings and intentions. However, many factors combine to influence and moderate facial behavior including personality, gender, context, and culture. Due to the high cost of traditional observational methods, the relationship between culture and facial behavior is not well-understood. In the current study, we explored the sociocultural factors that influence facial behavior using large-scale observational analyses. We developed and implemented an algorithm to automatically analyze the smiling of 866,726 participants across 31 different countries. We found that participants smiled more when from a country that is higher in individualism, has a lower population density, and has a long history of immigration diversity (i.e., historical heterogeneity). Our findings provide the first evidence that historical heterogeneity predicts actual smiling behavior. Furthermore, they converge with previous findings using selfreport methods. Taken together, these findings support the theory that historical heterogeneity explains, and may even contribute to the development of, permissive cultural display rules that encourage the open expression of emotion. Jeffrey M. Girard, Daniel McDuff |
FG | 2 |
| 2017 | The Impact of Video Compression on Remote Cardiac Pulse Measurement Using Imaging PhotoplethysmographyabstractRemote physiological measurement has great potential in healthcare and affective computing applications. Imaging photoplethysmography (iPPG) leverages digital cameras to recover the blood volume pulse from the human body. While the impact of video parameters such as resolution and frame rate on iPPG accuracy have been studied, there has not been a systematic analysis of video compression algorithms. We compared a set of commonly used video compression algorithms (x264 and x265) and varied the Constant Rate Factor (CRF) to measure pulse rate recovery for a range of bit rates (file sizes) and video qualities. We found that compression, even at a low CRF, degrades the blood volume pulse (BVP) signal-tonoise ratio considerably. However, the bit rate of a video can be substantially decreased (by a factor of over 1000) without destroying the BVP signal entirely. We found an approximately linear relationship between bit rate and BVP signal-to-noise ratio up to a CRF of 36. A faster decrease in SNR was observed for videos of the task involving larger head motions and the x265 algorithm appeared to work more effectively in these cases. Daniel McDuff, Ethan B. Blackford, Justin Estepp |
FG | 1 |
| 2017 | Large-scale Affective Content Analysis: Combining Media Content Features and Facial ReactionsabstractWe present a novel multimodal fusion model for affective content analysis, combining visual, audio and deep visual-sentiment descriptors from the media content with automated facial action measurements from naturalistic responses to the media. We collected a dataset of 48,867 facial responses to 384 media clips and extracted a rich feature set from the facial responses and media content. The stimulus videos were validated to be informative, inspiring, persuasive, sentimental or amusing. By combining the features, we were able to obtain a classification accuracy of 63% (weighted F1-score: 0.62) for a five-class task. This was a significant improvement over using the media content features alone. By analyzing the feature sets independently, we found that states of informed and persuaded were difficult to differentiate from facial responses alone due to the presence of similar sets of action units in each state (AU 2 occurring frequently in both cases). Facial actions were beneficial in differentiating between amused and informed states whereas media content features alone performed less well due to similarities in the visual and audio make up of the content. We highlight examples of content and reactions from each class. This is the first affective content analysis based on reactions of 10,000s of people. Daniel McDuff, Mohammad Soleymani 0001 |
FG | 1 |
| 2017 | Multimodal analysis of vocal collaborative search: a public corpus and resultsabstractIntelligent agents have the potential to help with many tasks. Information seeking and voice-enabled search assistants are becoming very common. However, there remain questions as to the extent by which these agents should sense and respond to emotional signals. We designed a set of information seeking tasks and recruited participants to complete them using a human intermediary. In total we collected data from 22 pairs of individuals, each completing five search tasks. The participants could communicate only using voice, over a VoIP service. Using automated methods we extracted facial action, voice prosody and linguistic features from the audio-visual recordings. We analyzed the characteristics of these interactions that correlated with successful communication and understanding between the pairs. We found that those who were expressive in channels that were missing from the communication channel (e.g., facial actions and gaze) were rated as communicating poorly, being less helpful and understanding. Having a way of reinstating nonverbal cues into these interactions would improve the experience, even when the tasks are purely information seeking exercises. The dataset used for this analysis contains over 15 hours of video, audio and transcripts and reported ratings. It is publicly available for researchers at: http://aka.ms/MISCv1. Daniel McDuff, Paul Thomas 0001, Mary Czerwinski, Nick Craswell |
ICMI | 1 |
| 2017 | Pulse and vital sign measurement in mixed reality using a HoloLensabstractCardiography, quantitative measurement of the functioning of the heart, traditionally requires customized obtrusive contact sensors. Using new methods photoplethysmography and ballistocardiography signals can be captured using ubiquitous sensors, such as webcams and accelerometers. However, these signals are not visible to the unaided eye. We present Cardiolens - a mixed reality system that enables real-time, hands-free measurement and visualization of blood flow and vital signs from multiple people. The system combines a front-facing webcam, imaging ballistocardiography, and remote imaging photoplethysmography methods for recovering pulse signals. A heads up display allows users to view their own heart rate whenever they are wearing the device and the heart rate and heart rate variability of another person simply by looking at them. Cardiolens provides the wearer with a new way to understand physiological signals and has applications in human-computer interaction and in the study of social psychology. Daniel McDuff, Christophe Hurter, Mar González-Franco |
VRST | 1 |
| 2017 | Applications of Automated Facial Coding in Media MeasurementabstractFacial coding has become a common tool in media measurement, with large companies (e.g., Unilever) using it to test all of their new video ad content. Facial reactions capture the in-the-moment response of an individual and these data complement self-report measures. Two advancements in affective computing have made measurement possible at scale: 1) computer vision algorithms are used to automatically code sign and message judgments based on facial muscle movements, 2) video data are collected by recording responses in everyday environments via the viewer's own webcam over the Internet. We present results of online facial coding studies of video ads, movie trailers, political content, and long-form TV shows. We explain how these data can be used in market research. Despite the ability to measure facial behavior in a scalable and quantifiable way, the interpretation of these data is still challenging without baselines and comparative measures. Over the past four years we have collected and coded over two million responses to everyday media content. Our huge dataset allows us to calculate reliable normative distributions of responses across different media types. We present these data and argue that this provides a context within which to interpret facial responses more accurately. Daniel McDuff, Rana El Kaliouby |
IEEE Trans. Affect. Comput. | 1 |
| 2016 | COGCAM: Contact-free Measurement of Cognitive Stress During Computer Tasks with a Digital CameraabstractContact-free camera-based measurement of cognitive stress opens up new possibilities for human-computer interaction with applications in remote learning, stress monitoring, and optimization of workload for user experience. The autonomic nervous system controls the inter-beat intervals of the heart and breathing patterns, and these signals change under cognitive stress. We built a participant-independent cognitive stress recognition model based on photoplethysmographic signals measured remotely at a distance of 3 meters. We tested the model on naturalistic responses from 10 individuals completing randomized-order computer-based tasks (ball control and card sorting). The system successfully detected increased stress during the tasks, which were consistent with self-report measures. Changes in heart rate variability were more discriminative indicators of cognitive stress than were heart rate and breathing rate. Daniel McDuff, Javier Hernandez, Sarah Gontarek, Rosalind W. Picard |
CHI | 1 |
| 2016 | Discovering facial expressions for states of amused, persuaded, informed, sentimental and inspiredabstractFacial expressions play a significant role in everyday interactions. A majority of the research on facial expressions of emotion has focused on a small set of "basic" states. However, in real-life the expression of emotions is highly context dependent and prototypic expressions of "basic" emotions may not always be present. In this paper we attempt to discover expressions associated with alternate states of informed, inspired, persuaded, sentimental and amused based on a very large dataset of observed facial responses. We used a curated set of 395 everyday videos that were found to reliably elicit the states and recorded 49,869 facial responses as viewers watched the videos in their homes. Using automated facial coding we quantified the presence of 18 facial actions in each of the 23.4 million frames. Lip corner pulls, lip sucks and inner brow raises were prominent in sentimental responses. Outer brow raises and eye widening were prominent in persuaded and informed responses. More brow furrowing distinguished informed from persuaded responses potentially indicating higher cognition. Daniel McDuff |
ICMI | 1 |
| 2016 | Driver Frustration Detection from Audio and Video in the Wild
Irman Abdic, Lex Fridman 0001, Daniel McDuff, Erik Marchi, Bryan Reimer, Björn W. Schuller |
IJCAI | 3 |
| 2016 | Understanding and Predicting Bonding in Conversations Using Thin Slices of Facial Expressions and Body Language
Natasha Jaques, Daniel McDuff, Yoo Lim Kim, Rosalind W. Picard |
IVA | 2 |
| 2016 | Wearable ESM: differences in the experience sampling method across wearable devicesabstractThe Experience Sampling Method is widely used for collecting self-report responses from people in natural settings. While most traditional approaches rely on using a phone to trigger prompts and record information, wearable devices now offer new opportunities that may improve this method. This research quantitatively and qualitatively studies the experience sampling process on head-worn and wrist-worn wearable devices, and compares them to the traditional "smartphone in the pocket." To enable this work, we designed and implemented a custom application to provide similar prompts across the three types of devices and evaluated it with 15 individuals for five days (75 days total), in the context of real-life stress measurement. We found significant differences in response times across devices, and captured tradeoffs in interaction types, screen size, and device familiarity that can affect both users' experience and the reports made by users. Javier Hernandez, Daniel McDuff, Christian Infante, Pattie Maes, Karen S. Quigley, Rosalind W. Picard |
MobileHCI | 2 |
| 2015 | Exploring temporal patterns in classifying frustrated and delighted smiles (Extended abstract)abstractWe created two experimental situations to elicit two affective states: frustration and delight. In the first experiment, participants were asked to recall situations while expressing either frustration or delight. The second experiment tried to elicit these states naturally with a frustrating experience and a delightful video. There were two significant differences between the acted and natural occurrences of the expressions. First, the acted instances were much easier for the computer to classify. Second, in 90 percent of the acted cases, participants did not smile when frustrated. In 90 percent of the natural cases, participants smiled during the frustrating interaction, despite self-reporting significant frustration with the experience. As a follow up study, we develop an automated system to distinguish between naturally occurring spontaneous smiles under frustrating and delightful stimuli by exploring their temporal patterns, given video of both. We extracted local and global features related to human smile dynamics. Next, we evaluated and compared two variants of Support Vector Machines (SVM), Hidden Markov Models (HMM), and Hidden-state Conditional Random Fields (HCRF) for binary classification. While human classification of the smile videos under frustrating stimuli was below chance, a dynamic SVM classifier obtained an accuracy of 92 percent in distinguishing smiles under frustrating and delighted stimuli. Mohammed E. Hoque 0001, Daniel McDuff, Rosalind W. Picard |
ACII | 2 |
| 2015 | Crowdsourcing facial responses to online videos: Extended abstractabstractTraditional observational research methods required an experimenter's presence in order to record videos of participants, and limited the scalability of data collection to typically less than a few hundred people in a single location. In order to make a significant leap forward in affective expression data collection and the insights based on it, our work has created and validated a novel framework for collecting and analyzing facial responses over the Internet. The first experiment using this framework enabled 3,268 trackable face videos to be collected and analyzed in under two months. Each participant viewed one or more commercials while their facial response was recorded and analyzed. Our data showed significantly different intensity and dynamics patterns of smile responses between subgroups who reported liking the commercials versus those who did not. Since this framework appeared in 2011, we have collected over three million videos of facial responses in over 75 countries using this same methodology, enabling facial analytics to become significantly more accurate and validated across five continents. Many new insights have been discovered based on crowd-sourced facial data, enabling Internet-based measurement of facial responses to become reliable and proven. We are now able to provide large-scale evidence for gender, cultural and age differences in behaviors. Today such methods are used as part of standard practice in industry for copy-testing advertisements and are increasingly used for online media evaluations, distance learning, and mobile applications. Daniel McDuff, Rana El Kaliouby, Rosalind W. Picard |
ACII | 1 |
| 2015 | BioInsights: Extracting personal data from "Still" wearable motion sensorsabstractDuring recent years a large variety of wearable devices have become commercially available. As these devices are in close contact with the body, they have the potential to capture sensitive and unexpected personal data even when the wearer is not moving. This work demonstrates that wearable motion sensors such as accelerometers and gyroscopes embedded in head-mounted and wrist-worn wearable devices can be used to identify the wearer (among 12 participants) and his/her body posture (among 3 positions) from only 10 seconds of “still” motion data. Instead of focusing on large and apparent motions such as steps or gait, the proposed methods amplify and analyze very subtle body motions associated with the beating of the heart. Our findings have the potential to increase the value of pervasive wearable motion sensors but also raise important privacy concerns that need to be considered. Javier Hernandez, Daniel McDuff, Rosalind W. Picard |
BSN | 2 |
| 2015 | Predicting Ad Liking and Purchase Intent: Large-Scale Analysis of Facial Responses to AdsabstractBillions of online video ads are viewed every month. We present a large-scale analysis of facial responses to video content measured over the Internet and their relationship to marketing effectiveness. We collected over 12,000 facial responses from 1,223 people to 170 ads from a range of markets and product categories. The facial responses were automatically coded frame-by-frame. Collection and coding of these 3.7 million frames would not have been feasible with traditional research methods. We show that detected expressions are sparse but that aggregate responses reveal rich emotion trajectories. By modeling the relationship between the facial responses and ad effectiveness, we show that ad liking can be predicted accurately (ROC AUC = 0.85) from webcam facial responses. Furthermore, the prediction of a change in purchase intent is possible (ROC AUC = 0.78). Ad liking is shown by eliciting expressions, particularly positive expressions. Driving purchase intent is more complex than just making viewers smile: peak positive responses that are immediately preceded by a brand appearance are more likely to be effective. The results presented here demonstrate a reliable and generalizable system for predicting ad effectiveness automatically from facial responses without a need to elicit self-report responses from the viewers. In addition we can gain insight into the structure of effective ads. Daniel McDuff, Rana El Kaliouby, Jeffrey F. Cohn, Rosalind W. Picard |
IEEE Trans. Affect. Comput. | 1 |
| 2014 | Automatic measurement of ad preferences from facial responses gathered over the Internet
Daniel McDuff, Rana El Kaliouby, Thibaud Senechal, David Demirdjian, Rosalind W. Picard |
Image Vis. Comput. | 1 |
| 2013 | Measuring Voter's Candidate Preference Based on Affective Responses to Election DebatesabstractIn this paper we present the first analysis of facial responses to electoral debates measured automatically over the Internet. We show that significantly different responses can be detected from viewers with different political preferences and that similar expressions at significant moments can have very different meanings depending on the actions that appear subsequently. We used an Internet based framework to collect 611 naturalistic and spontaneous facial responses to five video clips from the 3rd presidential debate during the 2012 American presidential election campaign. Using this framework we were able to collect over 60% of these video responses (374 videos) within one day of the live debate and over 80% within three days. No participants were compensated for taking the survey. We present and evaluate a method for predicting independent voter preference based on automatically measured facial responses and self-reported preferences from the viewers. We predict voter preference with an average accuracy of over 73% (AUC 0.779). Daniel McDuff, Rana El Kaliouby, Evan Kodra, Rosalind W. Picard |
ACII | 1 |
| 2012 | AffectAura: an intelligent system for emotional memoryabstractWe present AffectAura, an emotional prosthetic that allows users to reflect on their emotional states over long periods of time. We designed a multimodal sensor set-up for continuous logging of audio, visual, physiological and contextual data, a classification scheme for predicting user affective state and an interface for user reflection. The system continuously predicts a user's valence, arousal and engage-ment, and correlates this with information on events, communications and data interactions. We evaluate the interface through a user study consisting of six users and over 240 hours of data, and demonstrate the utility of such a reflection tool. We show that users could reason forward and backward in time about their emotional experiences using the interface, and found this useful. Daniel McDuff, Amy K. Karlson, Ashish Kapoor, Asta Roseway, Mary Czerwinski |
CHI | 1 |
| 2012 | Towards sensing the influence of visual narratives on human affectabstractIn this paper, we explore a multimodal approach to sensing affective state during exposure to visual narratives. Using four different modalities, consisting of visual facial behaviors, thermal imaging, heart rate measurements, and verbal descriptions, we show that we can effectively predict changes in human affect. Our experiments show that these modalities complement each other, and illustrate the role played by each of the four modalities in detecting human affect. Mihai Burzo, Daniel McDuff, Rada Mihalcea, Louis-Philippe Morency, Alexis Narvaez, Verónica Pérez-Rosas |
ICMI | 2 |
| 2012 | Exploring Temporal Patterns in Classifying Frustrated and Delighted SmilesabstractWe create two experimental situations to elicit two affective states: frustration, and delight. In the first experiment, participants were asked to recall situations while expressing either delight or frustration, while the second experiment tried to elicit these states naturally through a frustrating experience and through a delightful video. There were two significant differences in the nature of the acted versus natural occurrences of expressions. First, the acted instances were much easier for the computer to classify. Second, in 90 percent of the acted cases, participants did not smile when frustrated, whereas in 90 percent of the natural cases, participants smiled during the frustrating interaction, despite self-reporting significant frustration with the experience. As a follow up study, we develop an automated system to distinguish between naturally occurring spontaneous smiles under frustrating and delightful stimuli by exploring their temporal patterns given video of both. We extracted local and global features related to human smile dynamics. Next, we evaluated and compared two variants of Support Vector Machine (SVM), Hidden Markov Models (HMM), and Hidden-state Conditional Random Fields (HCRF) for binary classification. While human classification of the smile videos under frustrating stimuli was below chance, an accuracy of 92 percent distinguishing smiles under frustrating and delighted stimuli was obtained using a dynamic SVM classifier. Mohammed E. Hoque 0001, Daniel McDuff, Rosalind W. Picard |
IEEE Trans. Affect. Comput. | 2 |
| 2012 | Crowdsourcing Facial Responses to Online VideosabstractWe present results validating a novel framework for collecting and analyzing facial responses to media content over the Internet. This system allowed 3,268 trackable face videos to be collected and analyzed in under two months. We characterize the data and present analysis of the smile responses of viewers to three commercials. We compare statistics from this corpus to those from the Cohn-Kanade+ (CK+) and MMI databases and show that distributions of position, scale, pose, movement, and luminance of the facial region are significantly different from those represented in these traditionally used datasets. Next, we analyze the intensity and dynamics of smile responses, and show that there are significantly different facial responses from subgroups who report liking the commercials compared to those that report not liking the commercials. Similarly, we unveil significant differences between groups who were previously familiar with a commercial and those that were not and propose a link to virality. Finally, we present relationships between head movement and facial behavior that were observed within the data. The framework, data collected, and analysis demonstrate an ecologically valid method for unobtrusive evaluation of facial responses to media content that is robust to challenging real-world conditions and requires no explicit recruitment or compensation of participants. Daniel McDuff, Rana El Kaliouby, Rosalind W. Picard |
IEEE Trans. Affect. Comput. | 1 |
| 2011 | Machine Learning for Affective Computing
Mohammed E. Hoque 0001, Daniel McDuff, Louis-Philippe Morency, Rosalind W. Picard |
ACII (2) | 2 |
| 2011 | Real-time inference of mental states from facial expressions and upper body gesturesabstractWe present a real-time system for detecting facial action units and inferring emotional states from head and shoulder gestures and facial expressions. The dynamic system uses three levels of inference on progressively longer time scales. Firstly, facial action units and head orientation are identified from 22 feature points and Gabor filters. Secondly, Hidden Markov Models are used to classify sequences of actions into head and shoulder gestures. Finally, a multi level Dynamic Bayesian Network is used to model the unfolding emotional state based on probabilities of different gestures. The most probable state over a given video clip is chosen as the label for that clip. The average F1 score for 12 action units (AUs 1, 2, 4, 6, 7, 10, 12, 15, 17, 18, 25, 26), labelled on a frame by frame basis, was 0.461. The average classification rate for five emotional states (anger, fear, joy, relief, sadness) was 0.440. Sadness had the greatest rate, 0.64, anger the smallest, 0.11. Tadas Baltrusaitis, Daniel McDuff, Ntombikayise Banda, Marwa Mahmoud, Rana El Kaliouby, Peter Robinson 0001, Rosalind W. Picard |
FG | 2 |
| 2011 | Acume: A new visualization tool for understanding facial expression and gesture dataabstractFacial and head actions contain significant affective information. To date, these actions have mostly been studied in isolation because the space of naturalistic combinations is vast. Interactive visualization tools could enable new explorations of dynamically changing combinations of actions as people interact with natural stimuli. This paper describes a new open-source tool that enables navigation of and interaction with dynamic face and gesture data across large groups of people, making it easy to see when multiple facial actions co-occur, and how these patterns compare and cluster across groups of participants. We share two case studies that demonstrate how the tool allows researchers to quickly view an entire corpus of data for single or multiple participants, stimuli and actions. Acume yielded patterns of actions across participants and across stimuli, and helped give insight into how our automated facial analysis methods could be better designed. The results of these case studies are used to demonstrate the efficacy of the tool. The open-source code is designed to directly address the needs of the face and gesture research community, while also being extensible and flexible for accommodating other kinds of behavioral data. Source code, application and documentation are available at http://affect.media.mit.edu/acume. Daniel McDuff, Rana El Kaliouby, Karim Kassam, Rosalind W. Picard |
FG | 1 |
| 2011 | Crowdsourced data collection of facial responsesabstractIn the past, collecting data to train facial expression and affect recognition systems has been time consuming and often led to data that do not include spontaneous expressions. We present the first crowdsourced data collection of dynamic, natural and spontaneous facial responses as viewers watch media online. This system allowed a corpus of 3,268 videos to be collected in under two months. Daniel McDuff, Rana El Kaliouby, Rosalind W. Picard |
ICMI | 1 |