Sidney K. D'Mello

dblp:42/1682 · also Sidney D'Mello · DBLP profile ↗
← Back
196ranked-venue papers
24as first author
61since 2021 · last 2026
0000-0003-0347-2807ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 130 · 15 first-author · 43 since 2021Applied, interdisciplinary, general and emerging computing · 124 · 12 first-author · 37 since 2021Artificial intelligence and machine learning · 29 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 A Matter of Perspective: Contrasting User and Subject-Matter Experts' Sensemaking of LLM Feedback on Instructional Discourse
Chelsea Brown, Chelsea Chandler, Sandra Sawaya, Sidney K. D'Mello
AIED (5)4
2026 AI Partners that Support Productive Uncertainty Within "Jigsaw" Activities During Small Group Collaborative Learning in Classrooms
Monlin Ko, Chelsea Chandler, Sierra Rose, Brooklyn Cline, Emily Watts, Jason G. Reitman, Peter W. Foltz, Sidney K. D'Mello
AIED (3)8
2026 Scaling Teacher Professional Learning Through Automated Feedback: An Evaluation in K-12 Classrooms
abstract
Automated feedback provides a scalable approach to supporting teacher reflection and teaching practice. However, there is limited evidence on how these technologies impact multiple dimensions of teacher and student outcomes in authentic school settings. This mixed-methods study evaluated the implementation and short-term impacts of a 4–-5 week pilot providing K–-12 teachers with professional learning and automated feedback on teaching practices via the TeachFX application. We examined TeachFX’s impact on teacher well-being, teaching practice, classroom talk, and student engagement. Findings indicate that automated feedback is associated with reduced teacher burnout, improved self-efficacy, and increased perceptions of student engagement. Impacts on teaching practices were mixed, demonstrating both increases and decreases throughout the intervention period and varied by context. Results also suggested a shift from teacher discourse towards more student discourse over the course of the intervention. These findings suggest the potential of automated feedback to support teacher well-being and inform instructional changes, while highlighting areas for further research on sustaining and broadening practice improvements.
Jessica Vitale, Alyssa Van Camp, Mercedes Ortiz, Sidney K. D'Mello
LAK5
2026 Testing an AI-Enhanced Coached-Tutor Professional Learning Model for Scaling High-Dosage Tutoring
Robert Moulder, Sandra Sawaya, Sidney K. D'Mello
L@S3
2025 Efficacy of a Computer Tutor that Models Expert Human Tutors
Andrew Olney, Sidney K. D'Mello, Natalie K. Person, Whitney L. Cade, Patrick Hays, Claire W. Dempsey, Blair Lehman, Betsy Williams Sanders, Arthur C. Graesser
AIED (5)2
2025 Sense-Making with an AI-Enhanced Coaching Tool: A Think-Aloud Study
Sandra Sawaya, Sidney K. D'Mello
AIED (3)2
2025 Improving Tutor Discourse Practices via AI-Enhanced Coaching: A Piecewise Latent Growth Curve Modeling Approach
Sandra Sawaya, Jennifer Jacobs 0002, Robert G. Moulder, Chelsea Chandler, Brent Milne, Tom Fischaber, Sidney K. D'Mello
AIED (4)7
2025 "It feels like we're not meeting the criteria": Examining and Mitigating the Cascading Effects of Bias in Automatic Speech Recognition in Spoken Language Interfaces
abstract
Researchers have demonstrated that Automatic Speech Recognition (ASR) systems perform differently across demographic groups (i.e. show bias), yet their downstream impact on spoken language interfaces remains unexplored. We examined this question in the context of a real-world AI-powered interface that provides tutors with feedback on the quality of their discourse. We found that the Whisper ASR had lower accuracy for Black vs. white tutors, likely due to differences in acoustic patterns of speech. The downstream automated discourse classifiers of tutor talk were correspondingly less accurate for Black tutors when presented with ASR input. As a result, although Black tutors demonstrated higher-quality discourse on human transcripts, this trend was not evident on ASR transcripts. We experimented with methods to reduce ASR bias, finding that fine-tuning the ASR on Black speech reduced, but did not eliminate, ASR bias and its downstream effects. We discuss implications for AI-based spoken language interfaces aimed at providing unbiased assessments to improve performance outcomes.
Kelechi Ezema, Chelsea Chandler, Rosy Southwell, Niranjan Cholendiran, Sidney K. D'Mello
CHI5
2025 Improving the Generalizability of Models of Collaborative Discourse
Chelsea Chandler, Rohit Raju, Jason G. Reitman, William R. Penuel, Monlin Ko, Jeffrey Bush 0001, Quentin Biddy, Sidney K. D'Mello
EDM8
2025 Interactive Workshop: Multimodal, Multiparty Learning Analytics (MMLA)
Peter W. Foltz, Gautam Biswas, Sidney K. D'Mello
EDM3
2025 From Discourse to Dynamics: Understanding Team Interactions Through Temporally Sensitive NLP
Seehee Park, Danielle Shariff, Mohammad Amin Samadi, Nia Nixon, Sidney K. D'Mello
EDM5
2025 The Relationship between Collaborative Problem-Solving Skills and Group-to-Individual Learning Transfer in a Game-based Learning Environment
abstract
Collaborative problem solving (CPS) is viewed as an essential 21st century skill for the modern workforce. Accordingly, researchers have been investigating how to conceptualize, assess, and develop pedagogical approaches to improve CPS. These efforts require theoretically-grounded and empirically-validated frameworks of CPS which have been emerging over the past decade with various levels of validity data. The present paper focuses on validating the generalized competency model (GCM) of CPS with respect to predicting individual learning outcomes following CPS among triads. The GCM consists of three main facets–constructing shared knowledge, negotiation/coordination, and maintaining team function–mapped to behavioral indicators (i.e., observable evidence). It hypothesizes that scores on all three facets should positively predict CPS outcomes, including group-to-individual learning transfer. We tested this hypothesis in a study where 249 students who comprised 83 triads engaged in collaborative gameplay with the Physics Playground game environment remotely via videoconferencing. We found that the only CPS facet predicting individual physics learning was maintaining team function, after accounting for pretest scores, students’ perceptions of team collaboration, and their perceived physics self-efficacy. This facet was also the only significant predictor of individual learning regardless of how facet scores were computed (i.e., reverse coding of negative indicators, separating the sums of positive and negative indicators, and no reverse coding of negative indicators). Implications for the GCM and other CPS frameworks are discussed.
Chen Sun 0011, Valerie J. Shute, Angela Stewart, Sidney K. D'Mello
LAK4
2025 Understanding Collaborative Learning Processes and Outcomes Through Student Discourse Dynamics
Seehee Park, Nia Nixon, Sidney K. D'Mello, Danielle Shariff, Jaeyoon Choi
LAK3
2024 Aligning Tutor Discourse Supporting Rigorous Thinking with Tutee Content Mastery for Predicting Math Achievement
Mark Abdelshiheed, Jennifer Jacobs 0002, Sidney K. D'Mello
AIED (2)3
2024 Co-design Partners as Transformative Learners: Imagining Ideal Technology for Schools by Centering Speculative Relationships
abstract
Emergent technologies like artificial intelligence have been proposed to address issues of inequity in schools, yet tend to ossify the status quo because they address needs within an already inequitable system. In this paper, we draw from speculative participatory approaches across HCI and the learning sciences, and present a novel approach to co-design that forefronts supporting historically minoritized youth in developing transformative agency to change their schools based on their valued hopes, practices, and concerns. We argue that when co-design spaces forefront relational development, expansive technological objects emerge as a byproduct. We present a case study of expansive dreaming with U.S. historically minoritized students about the use of artificial intelligence to support classroom collaboration. Methodologically, we demonstrate how physically visiting spaces of collective agency serves as a powerful perceptual bridge to imagining joyful, equitable possibilities for schooling. Our approach yields new visions for schooling and new metaphors for artificial intelligence.
Michael Alan Chang, Richmond Y. Wong, Thomas Breideband, Thomas M. Philip, Ashieda McKoy, Arturo Cortez, Sidney K. D'Mello
CHI7
2024 Prompting as Panacea? A Case Study of In-Context Learning Performance for Qualitative Coding of Classroom Dialog
Ananya Ganesh, Chelsea Chandler, Sidney K. D'Mello, Martha Palmer, Katharina Kann
EDM3
2024 Characterizing Learners' Complex Attentional States During Online Multimedia Learning Using Eye-tracking, Egocentric Camera, Webcam, and Retrospective recalls
abstract
As online learning becomes increasingly ubiquitous, a key challenge is maintaining learners' sustained attention. Using eye-tracking, together with observing and interviewing learners, we can characterize both 1) whether they are looking at their learning materials, and 2) whether they are thinking about them. Critically, eye-tracking only speaks to the first distinction, not the second. To overcome this limitation, we supplemented eye-tracking with an egocentric camera, a webcam, a retrospective recall, and mind-wandering probes to capture a 2×2 matrix of attentional/cognitive states. We then categorized N=101 learners' attentional/cognitive states while they completed a multimedia physics module. This meets two goals: 1) allowing basic research to understand the relationship between attentional/cognitive states and behavioral outcomes; and 2) facilitating applied research by generating rich ground truth for future use in training machine learning to categorize this 2×2 set of attentional states, for which eye-tracking is necessary, but not sufficient.
Prasanth Chandran, Jeremy Munsell, Brian Howatt, Brayden Wallace, Lindsey Wilson, Sidney K. D'Mello, Minh Hoai, N. Sanjay Rebello, Lester C. Loschky
ETRA7
2024 Automatic Speech Recognition Tuned for Child Speech in the Classroom
abstract
K-12 school classrooms have proven to be a challenging environment for Automatic Speech Recognition (ASR) systems, both due to background noise and conversation, and differences in linguistic and acoustic properties from adult speech, on which the majority of ASR systems are trained and evaluated. We report on experiments to improve ASR for child speech in the classroom by training and fine-tuning transformer models on public corpora of adult and child speech augmented with classroom background noise. By tuning OpenAI’s Whisper model we achieve a 38% relative reduction in word error rate (WER) to 9.2% on the public MyST dataset of child speech – the lowest yet reported – and a 7% relative reduction to reach 54% WER on a more challenging classroom speech dataset (ISAT). We also introduce a novel beam hypothesis rescoring method that incorporates a speed-aware term to capture prior knowledge of human speaking rates, as well as a Large Language Model, to select among hypotheses. We demonstrate the effectiveness of this technique on both publicly-available datasets and a classroom speech dataset.
Rosy Southwell, Wayne H. Ward, Viet Anh Trinh, Charis Clevenger, Clay Clevenger, Emily Watts, Jason G. Reitman, Sidney K. D'Mello, Jacob Whitehill
ICASSP8
2024 Putting the "Brain" Back in the Eye-Mind Link: Aligning Eye Movements and Brain Activations During Naturalistic Reading
abstract
Eye movements have long been used to reflect ongoing cognitive processing to develop explanatory and predictive models of mental states and processes. This relationship, deemed the eye-mind link, contains underlying assumptions of the mental processes occurring in the brain, which have rarely been explicitly investigated. We propose a multimodal approach to investigate alignment of eye movements and brain activations (eye-brain alignment) and how it might be predicted by unfolding cognitive processes. We applied this method to a dataset of 76 participants who read long, connected texts while their eye movements and the hemodynamic responses in their brains were tracked using functional near-infrared spectroscopy (fNIRS). We found that reliable eye-brain alignment signals varied based on the participants’ cognitive state during reading. Implications for multimodal modeling of cognitive processes are discussed.
Megan Caruso, Rosy Southwell, Leanne M. Hirshfield, Sidney K. D'Mello
ICMI4
2024 Human-tutor Coaching Technology (HTCT): Automated Discourse Analytics in a Coached Tutoring Model
abstract
High-dosage tutoring has become an effective strategy for bolstering K-12 academic performance and combating education declines accelerated by the COVID-19 pandemic. To achieve high-dosage tutoring at scale, tutoring programs often rely on paraprofessional tutors—recruited tutors with college degrees who lack formal training in education—however, these tutors may require consistent and targeted feedback from instructional coaches for improvement. Accordingly, we developed a human-tutor coaching technology (HTCT) system to automatically extract discourse analytics pertaining to accountable talk moves (or academically productive talk) from tutoring sessions and provide feedback visualizations to coaches to aid their coaching sessions with tutors. We deployed HTCT in a user study using a virtual tutoring platform with 11 real coaches, 40 tutors, and their students to investigate coaches’ usage patterns with HTCT, perceptions of its utility, and changes in tutors’ talk. Overall, we found that coaches had positive perceptions of the system. We also observed an increase in accountable talk from tutors whose coaches used HTCT compared to tutors whose coaches did not. We discuss implications for AI-based applications which offer coaches a promising way to provide personalized, automated, and data-driven feedback to scale high-dosage tutoring.
Brandon M. Booth, Jennifer Jacobs 0002, Jeffrey Bush 0001, Brent Milne, Tom Fischaber, Sidney K. D'Mello
LAK6
2024 Computational Modeling of Collaborative Discourse to Enable Feedback and Reflection in Middle School Classrooms
abstract
Collaboration analytics has the potential to empower teachers and students with valuable insights to facilitate more meaningful and engaging collaborative learning experiences. Towards this end, we developed computational models of student speech during small group work, identifying instances of uplifting behavior related to three Community Agreements: community building, moving thinking forward, and being respectful. Pre-trained RoBERTa language models were fine-tuned and evaluated on human annotated data (N = 9,607 student utterances from 100 unique 5-minute classroom recordings). The models achieved moderate accuracies (AUROCs between 0.67-0.84) and were robust to speech recognition errors. Preliminary generalizability studies indicated that the models generalized well to two other domains (transfer ratios between 0.46-0.85; with 1.0 indicating perfect transfer). We also developed four approaches to provide qualitative feedback in the form of noticings (i.e., specific exemplars) of positive instances of the Community Agreements, finding moderate alignment with human ratings. This research contributes to the computational modeling of the relationship dimension of collaboration from noisy classroom data, selection of positive examples for qualitative feedback, and towards the empowerment of teachers to support diverse learners during collaborative learning.
Chelsea Chandler, Thomas Breideband, Jason G. Reitman, Marissa Chitwood, Jeffrey Bush 0001, Amanda Howard, Sarah Leonhart, Peter W. Foltz, William R. Penuel, Sidney K. D'Mello
LAK10
2024 Not a Team but Learning as One: The Impact of Consistent Attendance on Discourse Diversification in Math Group Modeling
abstract
This work investigates relationships between consistent attendance —attendance rates in a group that maintains the same tutor and students across the school year— and learning in small group tutoring sessions. We analyzed data from two large urban districts consisting of 206 9th-grade student groups (3 − 6 students per group) for a total of 803 students and 75 tutors. The students attended small group tutorials approximately every other day during the school year and completed a pre and post-assessment of math skills at the start and end of the year, respectively. First, we found that the attendance rates of the group predicted individual assessment scores better than the individual attendance rates of students comprising that group. Second, we found that groups with high consistent attendance had more frequent and diverse tutor and student talk centering around rich mathematical discussions. Whereas we emphasize that changing tutors or groups might be necessary, our findings suggest that consistently attending tutorial sessions as a group with the same tutor might lead the group to implicitly learn as a team despite not being one.
Mark Abdelshiheed, Jennifer Jacobs 0002, Sidney K. D'Mello
UMAP3
2024 Improving collaborative problem-solving skills via automated feedback and scaffolding: a quasi-experimental study with CPSCoach 2.0
Sidney K. D'Mello, Nicholas D. Duran, Amanda Michaels, Angela Stewart
User Model. User Adapt. Interact.1
2023 A Multi-theoretic Analysis of Collaborative Discourse: A Step Towards AI-Facilitated Student Collaborations
Jason G. Reitman, Charis Clevenger, Quinton Beck-White, Amanda Howard, Sierra Rose, Jacob Elick, Julianna Harris, Peter W. Foltz, Sidney K. D'Mello
AIED9
2023 Getting the Wiggles Out: Movement Between Tasks Predicts Future Mind Wandering During Learning Activities
Rosy Southwell, Candace E. Peacock, Sidney K. D'Mello
AIED3
2023 CPSCoach: The Design and Implementation of Intelligent Collaborative Problem Solving Feedback
Angela Stewart, Arjun Ramesh Rao, Amanda Michaels, Chen Sun 0011, Nicholas D. Duran, Valerie J. Shute, Sidney K. D'Mello
AIED7
2023 Do Associations Between Mind Wandering and Learning from Complex Texts Vary by Assessment Depth and Time?
abstract
We examined associations between mind wandering – where attention shifts from the task at hand to task-unrelated thoughts – and learning outcomes. Our data consisted of 177 students who self-reported mind wandering while reading five long, connected texts on scientific research methods and completed learning assessments targeting multiple depths of processing (rote, inference, integration) at different timescales (during and after reading each text, after reading all texts, and after a week-long delay). We found that mind wandering negatively predicted measures of factual, text-based (explicit) information and global integration of information across multiple parts of the text, but not measures requiring a local inference on a single sentence. Further, mind wandering only predicted comprehension measures assessed during the reading session and not after a week-long delay. Our findings provide important nuances to the established negative link between mind wandering and learning outcomes, which has predominantly focused on rote comprehension assessed during the learning session itself. Implications for interventions to address mind wandering during learning are discussed.
Megan Caruso, Sidney K. D'Mello
LAK2
2023 Recurrence Quantification Analysis of Eye Gaze Dynamics During Team Collaboration
abstract
Shared visual attention between team members facilitates collaborative problem solving (CPS), but little is known about how team-level eye gaze dynamics influence the quality and successfulness of CPS. To better understand the role of shared visual attention during CPS, we collected eye gaze data from 279 individuals solving computer-based physics puzzles while in teams of three. We converted eye gaze into discrete screen locations and quantified team-level gaze dynamics using recurrence quantification analysis (RQA). Specifically, we used a centroid-based auto-RQA approach, a pairwise team member cross-RQAs approach, and a multi-dimensional RQA approach to quantify team-level eye gaze dynamics from the eye gaze data of team members. We find that teams differing in composition based on prior task knowledge, gender, and race show few differences in team-level eye gaze dynamics. We also find that RQA metrics of team-level eye gaze dynamics were predictive of task success (all ps < .001). However, the same metrics showed different patterns of feature importance depending on predictive model and RQA type, suggesting some redundancy in task-relevant information. These findings signify that team-level eye gaze dynamics play an important role in CPS and that different forms of RQA pick up on unique aspects of shared attention between team-members.
Robert G. Moulder, Brandon M. Booth, Angelina Abitino, Sidney K. D'Mello
LAK4
2023 A Comparative Analysis of Automatic Speech Recognition Errors in Small Group Classroom Discourse
abstract
In collaborative learning environments, effective intelligent learning systems need to accurately analyze and understand the collaborative discourse between learners (i.e., group modeling) to provide adaptive support. We investigate how automatic speech recognition (ASR) errors influence discourse models of small group collaboration in noisy real-world classrooms. Our dataset consisted of 30 students recorded by consumer off-the-shelf microphones (Yeti Blue) while engaging in dyadic- and triadic- collaborative learning in a multi-day STEM curriculum unit. We found that two state-of-the-art ASR systems (Google Speech and OpenAI Whisper) yielded very high word error rates (0.822, 0.847) but very different profiles of error with Google being more conservative, rejecting 38% of utterances instead of 12% for Whisper. Next, we examined how these ASR errors influenced down-stream small group modeling based on pre-trained large language models for three tasks: Abstract Meaning Representation parsing (AMRParsing), on-task/off-task detection (OnTask), and Accountable Productive Talk prediction (TalkMove). As expected, models trained on clean human transcripts yielded degraded performance on all three tasks, measured by the transfer ratio (TR). However, the TR of the specific sentence-level AMRParsing task (.39 - .62) was much lower than that of the abstract discourse-level OnTask (.63- .94) and TalkMove tasks (.64-.72). Furthermore, different training strategies that incorporated ASR transcripts alone or as augmentations of human transcripts increased accuracy for the discourse-level tasks (OnTask and TalkMove) but not AMRParsing. Simulation experiments suggested that the models were tolerant of missing utterances in the dialog context, and that jointly improving ASR accuracy on important word classes (e.g., verbs and nouns) can improve performance across all tasks. Overall, our results provide insights into how different types of NLP-based tasks might be tolerant of ASR errors under extremely noisy conditions and provide suggestions for how to improve accuracy in small group modeling settings for a more equitable, engaging, and adaptive collaborative learning environment.
Jie Cao 0010, Ananya Ganesh, Jon Z. Cai, Rosy Southwell, Margaret Perkoff, Michael Regan, Katharina Kann, James H. Martin, Martha Palmer, Sidney K. D'Mello
UMAP10
2023 'Location, Location, Location': An Exploration of Different Workplace Contexts in Remote Teamwork during the COVID-19 Pandemic
abstract
Much emphasis has been placed on how the affordances and layouts of an office setting can influence co-worker interactions and perceived team outcomes. Little is known, however, whether perceptions of teamwork and team conflict are affected when the location of work changes from the office to the home. To address this gap, we present findings from a ten-week,in situ study of 91 information workers from 27 US-based teams. We compare three distinct work locations---private and shared workspaces at home as well at the office---and explore how each location may impact individual perceptions of teamwork. While there was no significant association with participants' perceptions of teamwork, results revealed associations of work location with team conflict: participants who worked in a private room at home reported significantly lower team conflict compared to those working in the office. No difference was found for the office and the shared workspace. We further found that the influence of work location on team conflict interacted with job decision latitude and the level of task interdependence among co-workers. We discuss practical implications for full-time work from home (WFH) on teams. Our study adds an important environmental dimension to the literature on remote teaming, which in turn may help organizations as they consider, prepare, or implement more permanent WFH and/or hybrid work policies in the future.
Thomas Breideband, Robert G. Moulder, Gonzalo J. Martínez, Megan Caruso, Gloria Mark, Aaron Striegel, Sidney K. D'Mello
Proc. ACM Hum. Comput. Interact.7
2023 Engagement Detection and Its Applications in Learning: A Tutorial and Selective Review
abstract
Engagement is critical to satisfaction and performance in a number of domains but is challenging to measure and sustain. Thus, there is considerable interest in developing affective computing technologies to automatically measure and enhance engagement, especially in the wild and at scale. This article provides an accessible introduction to affective computing research on engagement detection and enhancement using educational applications as an application domain. We begin with defining engagement as a multicomponential construct (i.e., a conceptual entity) situated within a context and bounded by time and review how the past six years of research has conceptualized it. Next, we examine traditional and affective computing methods for measuring engagement and discuss their relative strengths and limitations. Then, we move to a review of proactive and reactive approaches to enhancing engagement toward improving the learning experience and outcomes. We underscore key concerns in engagement measurement and enhancement, especially in digitally enhanced learning contexts, and conclude with several open questions and promising opportunities for future work.
Brandon M. Booth, Nigel Bosch, Sidney K. D'Mello
Proc. IEEE3
2023 Multimodal Engagement Analysis From Facial Videos in the Classroom
abstract
Student engagement is a key component of learning and teaching, resulting in a plethora of automated methods to measure it. Whereas most of the literature explores student engagement analysis using computer-based learning often in the lab, we focus on using classroom instruction in authentic learning environments. We collected audiovisual recordings of secondary school classes over a one and a half month period, acquired continuous engagement labeling per student (N=15) in repeated sessions, and explored computer vision methods to classify engagement from facial videos. We learned deep embeddings for attentional and affective features by training Attention-Net for head pose estimation and Affect-Net for facial expression recognition using previously-collected large-scale datasets. We used these representations to train engagement classifiers on our data, in individual and multiple channel settings, considering temporal dependencies. The best performing engagement classifiers achieved student-independent AUCs of .620 and .720 for grades 8 and 12, respectively, with attention-based features outperforming affective features. Score-level fusion either improved the engagement classifiers or was on par with the best performing modality. We also investigated the effect of personalization and found that only 60 seconds of person-specific data, selected by margin uncertainty of the base classifier, yielded an average AUC improvement of .084.
Ömer Sümer, Patricia Goldberg, Sidney K. D'Mello, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci
IEEE Trans. Affect. Comput.3
2023 Gaze-based predictive models of deep reading comprehension
Rosy Southwell, Caitlin Mills 0001, Megan Caruso, Sidney K. D'Mello
User Model. User Adapt. Interact.4
2022 Eye to Eye: Gaze Patterns Predict Remote Collaborative Problem Solving Behaviors in Triads
Angelina Abitino, Samuel L. Pugh, Candace E. Peacock, Sidney K. D'Mello
AIED (1)4
2022 Going Deep and Far: Gaze-based Models Predict Multiple Depths of Comprehension During and One Week Following Reading
Megan Caruso, Candace E. Peacock, Rosy Southwell, Guojing Zhou, Sidney K. D'Mello
EDM5
2022 Challenges and Feasibility of Automatic Speech Recognition for Modeling Student Collaborative Discourse in Classrooms
Rosy Southwell, Samuel L. Pugh, Margaret Perkoff, Charis Clevenger, Jeffrey Bush 0001, Rachel Lieber, Wayne H. Ward, Peter W. Foltz, Sidney K. D'Mello
EDM9
2022 Investigating Temporal Dynamics Underlying Successful Collaborative Problem Solving Behaviors with Multilevel Vector Autoregression
Guojing Zhou, Robert Moulder, Chen Sun 0011, Sidney K. D'Mello
EDM4
2022 Evaluating Calibration-free Webcam-based Eye Tracking for Gaze-based User Modeling
abstract
Eye tracking has been a research tool for decades, providing insights into interactions, usability, and, more recently, gaze-enabled interfaces. Recent work has utilized consumer-grade and webcam-based eye tracking, but is limited by the need to repeatedly calibrate the tracker, which becomes cumbersome for use outside the lab. To address this limitation, we developed an unsupervised algorithm that maps gaze vectors from a webcam to fixation features used for user modeling, bypassing the need for screen-based gaze coordinates, which require a calibration process. We evaluated our approach using three datasets (N=377) encompassing different UIs (computerized reading, an Intelligent Tutoring System), environments (laboratory or the classroom), and a traditional gaze tracker used for comparison. Our research shows that webcam-based gaze features correlate moderately with eye-tracker-based features and can model user engagement and comprehension as accurately as the latter. We discuss applications for research and gaze-enabled user interfaces for long-term use in the wild.
Stephen Hutt, Sidney K. D'Mello
ICMI2
2022 Assessing Multimodal Dynamics in Multi-Party Collaborative Interactions with Multi-Level Vector Autoregression
abstract
Multi-level vector autoregression (mlVAR) is a recently developed dynamic network model for assessing multimodal temporal data streams derived from multiple users over time. Importantly, mlVAR facilitates investigations into highly complex collaborative interactions within a unified framework. In order to demonstrate the utility of mlVAR for understanding the temporal dynamics of multimodal multi-party (MMP) interactions, we apply it to 9 signals measured from 201 users (67 triads) who engaged in a 15-minute collaborative problem solving task. Measured signals reflect participants’ affective states (positive valence and negative valence), physiological states (skin conductance and heart rate), attention (gaze fixation duration and gaze dispersion), nonverbal communication (head acceleration and facial expressiveness), and verbal communication (speech rate). Using node-level metrics of in-strength, out-strength, and synchrony, we show that mlVAR is capable of teasing apart complex role-based dynamics (controller, primary contributor, or secondary contributor) between participants. Our findings also provide evidence for a complex feedback system between individuals where internal states (i.e., skin conductance) are influenced by external signals of shared attention and communication (i.e., gaze and speech).
Robert G. Moulder, Nicholas D. Duran, Sidney K. D'Mello
ICMI3
2022 "Beautiful work, you're rock stars!": Teacher Analytics to Uncover Discourse that Supports or Undermines Student Motivation, Identity, and Belonging in Classrooms
abstract
From carefully crafted messages to flippant remarks, warm expressions to unfriendly grunts, teachers’ behaviors set the tone, expectations, and attitudes of the classroom. Thus, it is prudent to identify the ways in which teachers foster motivation, positive identity, and a strong sense of belonging through inclusive messaging and other interactions. We leveraged a new coding of teacher supportive discourse in 156 video clips from 73 6th to 8th grade math teachers from the archival Measures of Effective Teaching (MET) project. We trained Random Forest classifiers using verbal (words used) and paraverbal (acoustic-prosodic cues, e.g., speech rate) features to detect seven features of teacher discourse (e.g., public admonishment, autonomy supportive messages) from transcripts and audio, respectively. While both modalities performed over chance guessing, the specific language content was more predictive than paraverbal cues (mean correlation = .546 vs. .276); combining the two yielded no improvement. We examined the most predictive cues in order to gain a deeper understanding of the underlying messages in teacher talk. We discuss implications of our work for teacher analytics tools that aim to provide educators and researchers with insight into supportive discourse.
Nicholas C. Hunkins, Sean Kelly, Sidney K. D'Mello
LAK3
2022 A novel video recommendation system for algebra: An effectiveness evaluation study
abstract
This study presents a novel video recommendation system for an algebra virtual learning environment (VLE) that leverages ideas and methods from engagement measurement, item response theory, and reinforcement learning. Following Vygotsky's Zone of Proximal Development (ZPD) theory, but considering low affect and high affect students separately, we developed a system of five categories of video recommendations: 1) Watch new video; 2) Review current topic video with a new tutor; 3) Review segment of current video with current tutor; 4) Review segment of current video with a new tutor; 5) Watch next video in curriculum sequence. The category of recommendation was determined by student scores on a quiz and a sensor-free engagement detection model. New video recommendations (i.e., category 1) were selected based on a novel reinforcement learning algorithm that takes input from an item response theory model. The recommendation system was evaluated in a large field experiment, both before and after school closures due to the COVID-19 pandemic. The results show evidence of effectiveness of the video recommendation algorithm during the period of normal school operations, but the effect disappears after school closures. Implications for teacher orchestration of technology for normal classroom use and periods of school closure are discussed.
Walter L. Leite, Samrat Roy, Nilanjana Chakraborty, George Michailidis, Anne Corinne Huggins-Manley, Sidney K. D'Mello, Mohamad Kazem Shirani Faradonbeh, Emily Jensen, Huan Kuang, Zeyuan Jing
LAK6
2022 Do Speech-Based Collaboration Analytics Generalize Across Task Contexts?
abstract
We investigated the generalizability of language-based analytics models across two collaborative problem solving (CPS) tasks: an educational physics game and a block programming challenge. We analyzed a dataset of 95 triads (N=285) who used videoconferencing to collaborate on both tasks for an hour. We trained supervised natural language processing classifiers on automatic speech recognition transcripts to predict the human-coded CPS facets (skills) of constructing shared knowledge, negotiation / coordination, and maintaining team function. We tested three methods for representing collaborative discourse: (1) deep transfer learning (using BERT), (2) n-grams (counts of words/phrases), and (3) word categories (using the Linguistic Inquiry Word Count [LIWC] dictionary). We found that the BERT and LIWC methods generalized across tasks with only a small degradation in performance (Transfer Ratio of .93 with 1 indicating perfect transfer), while the n-grams had limited generalizability (Transfer Ratio of .86), suggesting overfitting to task-specific language. We discuss the implications of our findings for deploying language-based collaboration analytics in authentic educational environments.
Samuel L. Pugh, Arjun Ramesh Rao, Angela Stewart, Sidney K. D'Mello
LAK4
2022 Heterogeneity of Treatment Effects of a Video Recommendation System for Algebra
abstract
Previous research has shown that providing video recommendations to students in virtual learning environments implemented at scale positively affects student achievement. However, it is also critical to evaluate whether the treatment effects are heterogeneous, and whether they depend on contextual variables such as disadvantaged student status and characteristics of the school settings. The current study extends the evaluation of a novel video recommendation system by performing an exploratory search for sources of heterogeneity of treatment effects. This study's design is a multi-site randomized controlled trial with an assignment at the student level across three large and diverse school districts in the southeast United States. The study occurred in Spring 2021, when some students were in regular classrooms and others in online classrooms. The results of the current study replicate positive effects found in a previous field experiment that occurred in Spring 2020, at the onset of the COVID-19 pandemic. Then, causal forests were used to investigate the heterogeneity of treatment effects. This study contributes to the literature on content sequencing systems and recommendation systems by showing how these systems can disproportionally benefit the groups of students who had higher levels of previous algebra ability, followed more recommendations, learned remotely, were Hispanic, and received free or reduced-price lunch, which has implications for the fairness of implementation of educational technology solutions.
Walter L. Leite, Huan Kuang, Zuchao Shen, Nilanjana Chakraborty, George Michailidis, Sidney K. D'Mello, Wanli Xing 0001
L@S6
2022 What do Students' Interactions with Online Lecture Videos Reveal about their Learning?
abstract
Video viewing is an important component of online learning, yet little is known about what information about learning outcomes can be derived from students’ video control actions. We investigate the extent to which information on student learning is contained in their video-watching clickstreams (e.g. pausing, playing) immediately after watching a video. We analyzed data from 10,492 students who used an online learning platform for their Algebra 1 course. Our experiments encode students’ video-control clickstreams into sequences in several ways (e.g. aggregate actions, shuffle actions, and merge action types), and train Long Short-term Memory (LSTM) neural network models to predict after-video quiz scores (N = 32,482) from the sequences in a student-independent fashion. The results suggest that the action sequences contain a limited amount of information about student learning (r = 0.108 between model-predicted- and actual- quiz scores), with most of the information in simple counts of actions (r = 0.081) rather than the temporal ordering of actions. Combining information from video action sequences and traditional knowledge estimates from item-response theory (IRT) outperformed (r = 0.224) either approach independently. Implications for student modeling and adaptive learning support for viewing lecture videos are discussed.
Guojing Zhou, Tetsumichi Umada, Sidney K. D'Mello
UMAP3
2022 Flexibility Versus Routineness in Multimodal Health Indicators: A Sensor-based Longitudinal in Situ Study of Information Workers
abstract
Although some research highlights the benefits of behavioral routines for individual functioning, other research indicates that routines can reflect an individual's inflexibility and lower well-being. Given conflicting accounts on the benefits of routine, research is needed to examine how routineness versus flexibility in health-related behaviors correspond to personality traits, health, and occupational outcomes. We adopt a nonlinear dynamical systems approach to understanding routine using automatically sensed health-related behaviors collected from 483 information workers over a roughly two-month period. We utilized multidimensional recurrence quantification analysis to derive a measure of health regularity (routineness) from measures of daily step count, sleep duration, and heart rate variability (which relates to stress). Participants also completed measures of personality, health, and job performance at the start of the study and for two months via Ecological Momentary Assessments. Greater regularity was associated with higher neuroticism, lower agreeableness, and greater interpersonal and organizational deviance. Importantly, these results were independent of overall levels of each health indicator in addition to demographics. It is often believed that routine is desirable, but the results suggest that associations with routineness are more nuanced, and wearable sensors can provide insights into beneficial health behaviors.
Mary Jean Amon, Stephen M. Mattingly, Aaron Necaise, Gloria Mark, Nitesh V. Chawla, Anind K. Dey, Sidney K. D'Mello
ACM Trans. Comput. Heal.7
2022 Sleep Patterns and Sleep Alignment in Remote Teams during COVID-19
abstract
Working remotely from home during the COVID-19 pandemic has resulted in significant shifts and disruptions in the personal and work lives of millions of information workers and their teams. We examined how sleep patterns---an important component of mental and physical health---relates to teamwork. We used wearable sensing and daily questionnaires to examine sleep patterns, affect, and perceptions of teamwork in 71 information workers from 22 teams over a ten-week period. Participants reported delays in sleep onset and offset as well as longer sleep duration during the pandemic. A similar shift was found in work schedules, though total work hours did not change significantly. Surprisingly, we found that more sleep was negatively related to positive affect, perceptions of teamwork, and perceptions of team productivity. However, a greater misalignment in the sleep patterns of members in a team predicted positive affect and teamwork after accounting for individual differences in sleep preferences. A follow-up analysis of exit interviews with participants revealed team-working conventions and collaborative mindsets as prominent themes that might help explain some of the ways that misalignment in sleep can affect teamwork. We discuss implications of sleep and sleep misalignment in work-from-home contexts with an eye towards leveraging sleep data to facilitate remote teamwork.
Thomas Breideband, Gonzalo J. Martínez, Poorna Talkad Sukumar, Megan Caruso, Sidney K. D'Mello, Aaron Striegel, Gloria Mark
Proc. ACM Hum. Comput. Interact.5
2022 Home-Life and Work Rhythm Diversity in Distributed Teamwork: A Study with Information Workers during the COVID-19 Pandemic
abstract
During the COVID-19 pandemic, millions of previously co-located information workers had to work from home, a trend expected to become much more commonplace in the future. We interviewed 53 information workers from 17 U.S. teams to understand how this unique extended work-from-home setting influenced teamwork and how they adapted to it. Using a grounded theory approach, we discovered that extended remote work highlighted diversity in team members' home-lives and daily work rhythms. Whereas these types of diversity played only marginal roles for teams in the co-located office, they had a more tangible impact in the work-from-home setting, from coordination delays and interruptions to conflicts related to workload fairness, miscommunication, and trust. Importantly, workers reported that their teams adapted to these challenges by setting explicit norms and standards for online communication and asynchronous collaboration and by promoting general social and situational awareness. We discuss computer-supported designs to help teams manage these latent diversities in an extended remote teamwork setting.
Thomas Breideband, Poorna Talkad Sukumar, Gloria Mark, Megan Caruso, Sidney K. D'Mello, Aaron Striegel
Proc. ACM Hum. Comput. Interact.5
2022 Feasibility of Longitudinal Eye-Gaze Tracking in the Workplace
abstract
Eye movements provide a window into cognitive processes, but much of the research harnessing this data has been confined to the laboratory. We address whether eye gaze can be passively, reliably, and privately recorded in real-world environments across extended timeframes using commercial-off-the-shelf (COTS) sensors. We recorded eye gaze data from a COTS tracker embedded in participants (N=20) work environments at pseudorandom intervals across a two-week period. We found that valid samples were recorded approximately 30% of the time despite calibrating the eye tracker only once and without placing any other restrictions on participants. The number of valid samples decreased over days with the degree of decrease dependent on contextual variables (i.e., frequency of video conferencing) and individual difference attributes (e.g., sleep quality and multitasking ability). Participants reported that sensors did not change or impact their work. Our findings suggest the potential for the collection of eye-gaze in authentic environments.
Stephen Hutt, Angela Stewart, Julie M. Gregg, Stephen M. Mattingly, Sidney K. D'Mello
Proc. ACM Hum. Comput. Interact.5
2022 Toward Robust Stress Prediction in the Age of Wearables: Modeling Perceived Stress in a Longitudinal Study With Information Workers
abstract
Given the widespread adverse outcomes of stress – exacerbated by the current pandemic – wearable sensing provides unique opportunities for automated stress tracking to inform well-being interventions. However, its success in the wild and at scale depends on the robustness and validity of automated stress inference, which is limited in current systems. In this work, we enumerate the properties of robustness and validity necessary for achieving viable automated stress inference using wearable sensors, and we underscore present challenges to constructing and evaluating these systems. Using these criteria as guiding principles, we present automated stress inference results from a large (N=606)in situlongitudinal wearable and contextual sensing study of information workers. Using a multimodal approach encompassing a wearable sensor, relative location tracking, smartphone usage, and environmental sensing, we trained regression models to predict daily self-reported perceived stress in a participant-independent fashion. Our models significantly outperformed baseline variants with shuffled stress scores and were consistent with small-to-moderate effects. Our findings highlight the performance disparity between robust and valid approaches to automated perceived stress inference and current approaches and suggest that further performance gains might require additional sensing modalities and enhanced contextual awareness than existing approaches.
Brandon M. Booth, Hana Vrzakova, Stephen M. Mattingly, Gonzalo J. Martínez, Louis Faust, Sidney K. D'Mello
IEEE Trans. Affect. Comput.6
2022 Can Computers Outperform Humans in Detecting User Zone-Outs? Implications for Intelligent Interfaces
abstract
The ability to identify whether a user is “zoning out” (mind wandering) from video has many HCI (e.g., distance learning, high-stakes vigilance tasks). However, it remains unknown how well humans can perform this task, how they compare to automatic computerized approaches, and how a fusion of the two might improve accuracy. We analyzed videos of users’ faces and upper bodies recorded 10s prior to self-reported mind wandering (i.e., ground truth) while they engaged in a computerized reading task. We found that a state-of-the-art machine learning model had comparable accuracy to aggregated judgments of nine untrained human observers (area under receiver operating characteristic curve [AUC] = .598 versus .589). A fusion of the two (AUC = .644) outperformed each, presumably because each focused on complementary cues. Furthermore, adding more humans beyond 3–4 observers yielded diminishing returns. We discuss implications of human–computer fusion as a means to improve accuracy in complex tasks.
Nigel Bosch, Sidney K. D'Mello
ACM Trans. Comput. Hum. Interact.2
2021 Annotating Student Engagement Across Grades 1-12: Associations with Demographics and Expressivity
Nese Alyüz, Sinem Aslan, Sidney K. D'Mello, Lama Nachman, Asli Arslan Esme
AIED (1)3
2021 Breaking out of the Lab: Mitigating Mind Wandering with Gaze-Based Attention-Aware Technology in Classrooms
abstract
We designed and tested an attention-aware learning technology (AALT) that detects and responds to mind wandering (MW), a shift in attention from task-related to task-unrelated thoughts, that is negatively associated with learning. We leveraged an existing gaze-based mind wandering detector that uses commercial off the shelf eye tracking to inform real-time interventions during learning with an Intelligent Tutoring System in real-world classrooms. The intervention strategies, co-designed with students and teachers, consisted of using student names, reiterating content, and asking questions, with the aim to reengage wandering minds and improve learning. After several rounds of iterative refinement, we tested our AALT in two classroom studies with 287 high-school students. We found that interventions successfully reoriented attention, and compared to two control conditions, reduced mind wandering, and improved retention (measured via a delayed assessment) for students with low prior-knowledge who occasionally (but not excessively) mind wandered. We discuss implications for developing gaze-based AALTs for real-world contexts.
Stephen Hutt, Kristina Krasich, James R. Brockmole, Sidney K. D'Mello
CHI4
2021 Multi-Level Linguistic Alignment in a Virtual Collaborative Problem-Solving Task
Nicholas D. Duran, Amie Paige, Sidney K. D'Mello
CogSci3
2021 Improving Automated Teacher Discourse Classification via Automated Reliability Modeling
Nicholas C. Hunkins, Emily Jensen, Sidney K. D'Mello
EDM3
2021 Say What? Automatic Modeling of Collaborative Problem Solving Skills from Student Speech in the Wild
Samuel L. Pugh, Shree Krishna Subburaj, Arjun Ramesh Rao, Angela Stewart, Jessica Andrews-Todd, Sidney K. D'Mello
EDM6
2021 Bias and Fairness in Multimodal Machine Learning: A Case Study of Automated Video Interviews
abstract
We introduce the psychometric concepts of bias and fairness in a multimodal machine learning context assessing individuals’ hireability from prerecorded video interviews. We collected interviews from 733 participants and hireability ratings from a panel of trained annotators in a simulated hiring study, and then trained interpretable machine learning models on verbal, paraverbal, and visual features extracted from the videos to investigate unimodal versus multimodal bias and fairness. Our results demonstrate that, in the absence of any bias mitigation strategy, combining multiple modalities only marginally improves prediction accuracy at the cost of increasing bias and reducing fairness compared to the least biased and most fair unimodal predictor set (verbal). We further show that gender-norming predictors only reduces gender predictability for paraverbal and visual modalities, while removing gender-biased features can achieve gender blindness, minimal bias, and fairness (for all modalities except for visual) at the cost of some prediction accuracy. Overall, the reduced-feature approach using predictors from all modalities achieved the best balance between accuracy, bias, and fairness, with the verbal modality alone performing almost as well. Our analysis highlights how optimizing model prediction accuracy in isolation and in a multimodal context may cause bias, disparate impact, and potential social harm, while a more holistic optimization approach based on accuracy, bias, and fairness can avoid these pitfalls.
Brandon M. Booth, Louis Hickman, Shree Krishna Subburaj, Louis Tay, Sang Eun Woo, Sidney K. D'Mello
ICMI6
2021 A Deep Transfer Learning Approach to Modeling Teacher Discourse in the Classroom
abstract
Teachers, like everyone else, need objective reliable feedback in order to improve their effectiveness. However, developing a system for automated teacher feedback entails many decisions regarding data collection procedures, automated analysis, and presentation of feedback for reflection. We address the latter two questions by comparing two different machine learning approaches to automatically model seven features of teacher discourse (e.g., use of questions, elaborated evaluations). We compared a traditional open-vocabulary approach using n-grams and Random Forest classifiers with a state-of-the-art deep transfer learning approach for natural language processing (BERT). We found a tradeoff between data quantity and accuracy, where deep models had an advantage on larger datasets, but not for smaller datasets, particularly for variables with low incidence rates. We also compared the models based on the level of feedback granularity: utterance-level (e.g., whether an utterance is a question or a statement), class session-level proportions by averaging across utterances (e.g., question incidence score of 48%), and session-level ordinal feedback based on pre-determined thresholds (e.g., question asking score is medium [vs. low or high]) and found that BERT generally provided more accurate feedback at all levels of granularity. Thus, BERT appears to be the most viable approach to providing automatic feedback on teacher discourse provided there is sufficient data to fine tune the model.
Emily Jensen, Samuel L. Pugh, Sidney K. D'Mello
LAK3
2021 What You Do Predicts How You Do: Prospectively Modeling Student Quiz Performance Using Activity Features in an Online Learning Environment
abstract
Students using online learning environments need to effectively self-regulate their learning. However, with an absence of teacher-provided structure, students often resort to less effective, passive learning strategies versus constructive ones. We consider the potential benefits of interventions that promote retrieval practice – retrieving learned content from memory – which is an effective strategy for learning and retention. The goal is to nudge students towards completing short, formative quizzes when they are likely to succeed on those assessments. Towards this goal, we developed a machine-learning model using data from 32,685 students who used an online mathematics platform over an entire school year to prospectively predict scores on three-item assessments (N = 210,020) from interaction patterns up to 9 minutes before the assessment as well as Item Response Theory (IRT) estimates of student ability and quiz difficulty. These models achieved a student-independent correlation of 0.55 between predicted and actual scores on the assessments and outperformed IRT-only predictions (r = 0.34). Model performance was largely independent of the length of the analyzed window preceding a quiz. We discuss potential for future applications of the models to trigger dynamic interventions that aim to encourage students to engage with formative assessments rather than more passive learning strategies.
Emily Jensen, Tetsumichi Umada, Nicholas C. Hunkins, Stephen Hutt, Anne Corinne Huggins-Manley, Sidney K. D'Mello
LAK6
2021 Eye-Mind reader: an intelligent reading interface that promotes long-term comprehension by detecting and responding to mind wandering
abstract
We zone out roughly 20-40% of the time during reading – a rate that is concerning given the negative relationship between mind-wandering and comprehension. We tested if Eye-Mind Reader – an intelligent interface that targeted mind-wandering as it occurred – could mitigate its negative impact on reading comprehension. When an eye-gaze-based classifier indicated that a reader was mind-wandering, those in a MW-Intervention condition were asked to self-explain the concept they were reading about. If the self-explanation quality was deemed subpar by an automated scoring mechanism, readers were asked to re-read parts of the text in order to correct their comprehension deficits and improve their self-explanation. Each participant in the MW-Intervention condition was paired with a Yoked-Control counterpart who received the exact same interventions regardless of whether they were mind-wandering. Results indicate that re-reading improved self-explanation quality for the MW-Intervention group, but not the control group. The two conditions performed equally well on textbase (i.e. fact-based) and inference-level comprehension questions immediately after reading. However, after a week-long delay, the MW-Intervention condition significantly outperformed the yoked-control condition on both comprehension assessments (ds = .352 and .307). Our findings suggest that real-time interventions during critical periods of mind-wandering can promote long-term retention and comprehension.
Caitlin Mills 0001, Julie M. Gregg, Robert Bixler, Sidney K. D'Mello
Hum. Comput. Interact.4
2021 Automatic Detection of Mind Wandering from Video in the Lab and in the Classroom
abstract
We report two studies that used facial features to automatically detect mind wandering, a ubiquitous phenomenon whereby attention drifts from the current task to unrelated thoughts. In a laboratory study, university students$(N = 152)$read a scientific text, whereas in a classroom study high school students$(N = 135)$learned biology from an intelligent tutoring system. Mind wandering was measured using validated self-report methods. In the lab, we recorded face videos and analyzed these at six levels of granularity: (1) upper-body movement; (2) head pose; (3) facial textures; (4) facial action units (AUs); (5) co-occurring AUs; and (6) temporal dynamics of AUs. Due to privacy constraints, videos were not recorded in the classroom. Instead, we extracted head pose, AUs, and AU co-occurrences in real-time. Machine learning models, consisting of support vector machines (SVM) and deep neural networks, achieved$F_{1}$scores of .478 and .414 (25.4 and 20.9 percent above-chance improvements, both with SVMs) for detecting mind wandering in the lab and classroom, respectively. The lab-based detectors achieved 8.4 percent improvement over the previous state-of-the-art; no comparison is available for classroom detectors. We discuss how the detectors can integrate into intelligent interfaces to increase engagement and learning by responding to wandering minds.
Nigel Bosch, Sidney K. D'Mello
IEEE Trans. Affect. Comput.2
2021 Multimodal modeling of collaborative problem-solving facets in triads
Angela Stewart, Zachary A. Keirn, Sidney K. D'Mello
User Model. User Adapt. Interact.3
2020 MBead: Semi-supervised Multilabel Behaviour Anomaly Detection on Multivariate Temporal Sensory Data
abstract
Human abnormal physical and psychological behaviors, such as high level of stress, may result in negative impacts on work and life, if not handled efficiently. However, the continuous collection of behavioral data from questionnaires is not feasible, as is often the case for the natural downside of survey data gathering. Thanks to the proliferation of mobile sensors, it brings compelling opportunities for us to more deeply analyze human behavior. In this work, we ask the question of detecting anomalies in human physical and psychological behaviors from multivariate temporal data from multi-modal sensors. In the past decades, many efforts have been made in developing anomaly detection methods, but there remain several challenges in this specific domain problem: 1) data contains missing values at random positions. 2) data from multiple sensors is of multi-resolution and multivariate. 3) human behaviors are correlated to each other, thus it poses a multi-label problem. 4) the available labeled instances are limited, which requires the semi-supervised learning setting. 5) the frequency of anomaly occurrence is much smaller than that of normal instances, leading to imbalance problems. We propose a novel framework MBead to resolve these concerns. MBead consists of three key components: reweighted autoencoder to capture the dependency across temporal domain and multiple modalities, relevance learning module to learn the pairwise relations among labeled instances, and temporal prediction module to detect the anomalies while trained in semi-supervised settings. Extensive experiments show our MBead outperforms seven state-of-art baselines on three tasks of behavior anomaly detection: stress, affect, and work performance.
Suwen Lin, Louis Faust, Sidney K. D'Mello, Gonzalo J. Martínez, Nitesh V. Chawla
IEEE BigData3
2020 Toward Automated Feedback on Teacher Discourse to Enhance Teacher Learning
abstract
Like anyone, teachers need feedback to improve. Due to the high cost of human classroom observation, teachers receive infrequent feedback which is often more focused on evaluating performance than on improving practice. To address this critical barrier to teacher learning, we aim to provide teachers with detailed and actionable automated feedback. Towards this end, we developed an approach that enables teachers to easily record high-quality audio from their classes. Using this approach, teachers recorded 142 classroom sessions, of which 127 (89%) were usable. Next, we used speech recognition and machine learning to develop teacher-generalizable computer-scored estimates of key dimensions of teacher discourse. We found that automated models were moderately accurate when compared to human coders and that speech recognition errors did not influence performance. We conclude that authentic teacher discourse can be recorded and analyzed for automatic feedback. Our next step is to incorporate the automatic models into an interactive visualization tool that will provide teachers with objective feedback on the quality of their discourse.
Emily Jensen, Meghan Dale, Patrick J. Donnelly, Cathlyn Stone, Sean Kelly, Amanda Godley, Sidney K. D'Mello
CHI7
2020 Beyond Team Makeup: Diversity in Teams Predicts Valued Outcomes in Computer-Mediated Collaborations
abstract
In an increasingly globalized and service-oriented economy, people need to engage in computer-mediated collaborative problem solving (CPS) with diverse teams. However, teams routinely fail to live up to expectations, showcasing the need for technologies that help develop effective collaboration skills. We take a step in this direction by investigating how different dimensions of team diversity (demographic, personality, attitudes towards teamwork, prior domain experience) predict objective (e.g. effective solutions) and subjective (e.g. positive perceptions) collaborative outcomes. We collected data from 96 triads who engaged in a 30-minute CPS task via videoconferencing. We found that demographic diversity and differing attitudes towards teamwork predicted impressions of positive engagement, while personality diversity predicted learning outcomes. Importantly, these relationships were maintained after accounting for team makeup. None of the diversity measures predicted task performance. We discuss how our findings can be incorporated into technologies that aim to help diverse teams develop CPS skills.
Angela Stewart, Mary Jean Amon, Nicholas D. Duran, Sidney K. D'Mello
CHI4
2020 Multimodal, Multiparty Modeling of Collaborative Problem Solving Performance
abstract
Modeling team phenomena from multiparty interactions inherently requires combining signals from multiple teammates, often by weighting strategies. Here, we explored the hypothesis that strategic weighting signals from individual teammates would outperform an equal weighting baseline. Accordingly, we explored role-, trait-, and behavior-based weighting of behavioral signals across team members. We analyzed data from 101 triads engaged in computer-mediated collaborative problem solving (CPS) in an educational physics game. We investigated the accuracy of machine-learned models trained on facial expressions, acoustic-prosodics, eye gaze, and task context information, computed one-minute prior to the end of a game level, at predicting success at solving that level. AUROCs for unimodal models that equally weighted features from the three teammates ranged from .54 to .67, whereas a combination of gaze, face, and task context features, achieved an AUROC of .73. The various multiparty weighting strategies did not outperform an equal-weighting baseline. However, our best nonverbal model (AUROC = .73) outperformed a language-based model (AUROC = .67), and there were some advantages to combining the two (AUROC = .75). Finally, models aimed at prospectively predicting performance on a minute-by-minute basis from the start of the level achieved a lower, but still above-chance, AUROC of .60. We discuss implications for multiparty modeling of team performance and other team constructs.
Shree Krishna Subburaj, Angela Stewart, Arjun Ramesh Rao, Sidney K. D'Mello
ICMI4
2020 Focused or stuck together: multimodal patterns reveal triads' performance in collaborative problem solving
abstract
Collaborative problem solving (CPS) in virtual environments is an increasingly important context of 21st century learning. However, our understanding of this complex and dynamic phenomenon is still limited. Here, we examine unimodal primitives (activity on the screen, speech, and body movements), and their multimodal combinations during remote CPS. We analyze two datasets where 116 triads collaboratively engaged in a challenging visual programming task using video conferencing software. We investigate how UI-interactions, behavioral primitives, and multimodal patterns were associated with teams' subjective and objective performance outcomes. We found that idling with limited speech (i.e., silence or backchannel feedback only) and without movement was negatively correlated with task performance and with participants' subjective perceptions of the collaboration. However, being silent and focused during solution execution was positively correlated with task performance. Results illustrate that in some cases, multimodal patterns improved the predictions and improved explanatory power over the unimodal primitives. We discuss how the findings can inform the design of real-time interventions for remote CPS.
Hana Vrzakova, Mary Jean Amon, Angela Stewart, Nicholas D. Duran, Sidney K. D'Mello
LAK5
2020 Painometry: wearable and objective quantification system for acute postoperative pain
abstract
Over 50 million people undergo surgeries each year in the United States, with over 70% of them filling opioid prescriptions within one week of the surgery. Due to the highly addictive nature of these opiates, a post-surgical window is a crucial time for pain management to ensure accurate prescription of opioids. Drug prescription nowadays relies primarily on self-reported pain levels to determine the frequency and dosage of pain drug. Patient pain self-reports are, however, influenced by subjective pain tolerance, memories of past painful episodes, current context, and the patient's integrity in reporting their pain level. Therefore, objective measures of pain are needed to better inform pain management.
Hoang Truong 0002, Nam Bui, Zohreh Raghebi, Marta Ceko, Nhat Pham, Phuc Nguyen 0002, Anh Nguyen 0001, Katrina Siegfried, Evan Stene, Taylor Tvrdy, Logan Weinman, Thomas H. Payne, Devin Burke, Thang N. Dinh, Sidney K. D'Mello, Farnoush Banaei Kashani, Tor D. Wager, Pavel Goldstein, Tam Vu 0001
MobiSys16
2020 Looking for a Deal?: Visual Social Attention during Negotiations via Mixed Media Videoconferencing
abstract
Whereas social visual attention has been examined in computer-mediated (e.g., shared screen) or video-mediated (e.g., FaceTime) interaction, it has yet to be studied in mixed-media interfaces that combine video of the conversant along with other UI elements. We analyzed eye gaze of 37 dyads (74 participants) who were tasked with negotiating the price of a new car (as a buyer and seller) using mixed-media video conferencing under competitive or cooperative negotiation instructions (experimental manipulation). We used multidimensional recurrence quantification analysis to extract spatio-temporal patterns corresponding to mutual gaze (individuals look at each other), joint attention (individuals focus on the same elements of the interface), and gaze aversion (an individual looks at their partner, who is looking elsewhere). Our results indicated that joint attention predicted the sum of points attained by the buyer and seller (i.e., the joint score). In contrast, gaze aversion was associated with faster time to complete the negotiation, but with a lower joint score. Unexpectedly, mutual gaze was highly infrequent and unrelated to the negotiation outcomes and none of the gaze patterns predicted subjective perceptions of the negotiation. There were also no effects of gender composition or negotiation condition on the gaze patterns or negotiation outcomes. Our results suggest that social visual attention may operate differently in mixed-media collaborative interfaces than in face-to-face interaction. As mixed-media collaborative interfaces gain prominence, our work can be leveraged to inform the design of gaze-sensitive user interfaces that support remote negotiations among other tasks.
Hana Vrzakova, Mary Jean Amon, McKenzie Rees, Myrthe Faber, Sidney K. D'Mello
Proc. ACM Hum. Comput. Interact.5
2019 Imputing Missing Social Media Data Stream in Multisensor Studies of Human Behavior
abstract
The ubiquitous use of social media enables researchers to obtain self-recorded longitudinal data of individuals in real-time. Because this data can be collected in an inexpensive and unobtrusive way at scale, social media has been adopted as a “passive sensor” to study human behavior. However, such research is impacted by the lack of homogeneity in the use of social media, and the engineering challenges in obtaining such data. This paper proposes a statistical framework to leverage the potential of social media in sensing studies of human behavior, while navigating the challenges associated with its sparsity. Our framework is situated in a large-scale in-situ study concerning the passive assessment of psychological constructs of 757 information workers wherein of four sensing streams was deployed - bluetooth beacons, wearable, smartphone, and social media. Our framework includes principled feature transformation and machine learning models that predict latent social media features from the other passive sensors. We demonstrate the efficacy of this imputation framework via a high correlation of 0.78 between actual and imputed social media features. With the imputed features we test and validate predictions on psychological constructs like personality traits and affect. We find that adding the social media data streams, in their imputed form, improves the prediction of these measures. We discuss how our framework can be valuable in multimodal sensing studies that aim to gather comprehensive signals about an individual's state or situation.
Koustuv Saha, Raghu Mulukutla, Kari Nies, Pablo Robles-Granda, Anusha Sirigiri, Dong Whi Yoo, Pino G. Audia, Andrew T. Campbell, Nitesh V. Chawla, Sidney K. D'Mello, Anind K. Dey, Manikanta D. Reddy, Kaifeng Jiang, Gloria Mark, Edward Moskal, Aaron Striegel, Munmun De Choudhury, Vedant Das Swain, Julie M. Gregg, Ted Grover, Suwen Lin, Gonzalo J. Martínez, Stephen M. Mattingly, Shayan Mirjafari
ACII10
2019 Reducing Mind-Wandering During Vicarious Learning from an Intelligent Tutoring System
Caitlin Mills 0001, Nigel Bosch, Kristina Krasich, Sidney K. D'Mello
AIED (1)4
2019 Investigating the Impact of a Real-time, Multimodal Student Engagement Analytics Technology in Authentic Classrooms
abstract
We developed a real-time, multimodal Student Engagement Analytics Technology so that teachers can provide just-in-time personalized support to students who risk disengagement. To investigate the impact of the technology, we ran an exploratory semester-long study with a teacher in two classrooms. We used a multi-method approach consisting of a quasi-experimental design to evaluate the impact of the technology and a case study design to understand the environmental and social factors surrounding the classroom setting. The results show that the technology had a significant impact on the teacher's classroom practices (i.e., increased scaffolding to the students) and student engagement (i.e., less boredom). These results suggest that the technology has the potential to support teachers' role of being a coach in technology-mediated learning environments.
Sinem Aslan, Nese Alyüz, Cagri Tanriover, Sinem Emine Mete, Eda Okur, Sidney K. D'Mello, Asli Arslan Esme
CHI6
2019 Time to Scale: Generalizable Affect Detection for Tens of Thousands of Students across An Entire School Year
abstract
We developed generalizable affect detectors using 133,966 instances of 18 affective states collected from 69,174 students who interacted with an online math learning platform called Algebra Nation over the entire school year. To enable scalability and generalizability, we used generic interaction features (e.g., viewing a video, taking a quiz), which do not require specialized sensors and are domain- and (to a certain extent) system-independent. We experimented with standard classifiers, recurrent neural networks, and genetically evolved neural networks for affect modeling. Prediction accuracies, quantified with Spearman's rho, were modest and ranged from .08 (for surprise) to .34 (for happiness) with a mean of .25. Our model trained on Algebra students generalized to a different set of Geometry students (n = 28,458) on the same platform. We discuss implications for scaling up affect detection for affect-sensitive online learning environments which aim to improve engagement and learning by detecting and responding to student affect.
Stephen Hutt, Joseph F. Grafsgaard, Sidney K. D'Mello
CHI3
2019 Dynamics of Visual Attention in Multiparty Collaborative Problem Solving using Multidimensional Recurrence Quantification Analysis
abstract
Multiparty collaborative problem solving - an increasingly important context in the 21st century workforce - suffers from a degradation of social and behavioral signals when attempted remotely, resulting in suboptimal outcomes. We investigate teams' multidimensional patterns of visual attention during a collaborative problem-solving task with an eye for leveraging insights to improve collaborative interfaces. Fifty-seven novices (forming 19 triads) engaged in a challenging programming task (Minecraft Hour of Code) using videoconferencing software with screen sharing. To discover patterns of individual-level gaze-UI coupling(coordination of a teammate's attention with respect to changes in the user interface) and team-level gaze-UI regularity (dynamics of teams' collective attention in context with changes in the user interface), we applied cross- and multidimensional recurrence quantification analyses, respectively. Individuals' eye gaze was significantly coupled with the ongoing screen activity whereas teams displayed significant patterns of gaze regularity, suggesting repetitive patterns in teams' attention. These measures predicted expert-coded collaborative processes of constructing shared knowledge and negotiation and coordination (but not maintaining team function) and correlated with task score (r = .425). They also predicted individually assessed subjective perceptions of team performance and the collaboration process, but not individual's learning or team's task scores. We discuss implications of our findings for the design of intelligent collaborative interfaces.
Hana Vrzakova, Mary Jean Amon, Angela Stewart, Sidney K. D'Mello
CHI4
2019 Evaluating Fairness and Generalizability in Models Predicting On-Time Graduation from College Applications
Stephen Hutt, Margo Gardner, Angela Lee Duckworth, Sidney K. D'Mello
EDM4
2019 Generalizability of Sensor-Free Affect Detection Models in a Longitudinal Dataset of Tens of Thousands of Students
Emily Jensen, Stephen Hutt, Sidney K. D'Mello
EDM3
2019 Utterance-level Modeling of Indicators of Engaging Classroom Discourse
Cathlyn Stone, Patrick J. Donnelly, Meghan Dale, Sarah Capello, Sean Kelly, Amanda Godley, Sidney K. D'Mello
EDM7
2019 Modeling Team-level Multimodal Dynamics during Multiparty Collaboration
abstract
We adopt a multimodal approach to investigating team interactions in the context of remote collaborative problem solving (CPS). Our goal is to understand multimodal patterns that emerge and their relation with collaborative outcomes. We measured speech rate, body movement, and galvanic skin response from 101 triads (303 participants) who used video conferencing software to collaboratively solve challenging levels in an educational physics game. We use multi-dimensional recurrence quantification analysis (MdRQA) to quantify patterns of team-level regularity, or repeated patterns of activity in these three modalities. We found that teams exhibit significant regularity above chance baselines. Regularity was unaffected by task factors. but had a quadratic relationship with session time in that it initially increased but then decreased as the session progressed. Importantly, teams that produce more varied behavioral patterns (irregularity) reported higher emotional valence and performed better on a subset of the problem solving tasks. Regularity did not predict arousal or subjective perceptions of the collaboration. We discuss implications of our findings for the design of systems that aim to improve collaborative outcomes by monitoring the ongoing collaboration and intervening accordingly.
Lucca Eloy, Angela Stewart, Mary Jean Amon, Caroline Reinhardt, Amanda Michaels, Chen Sun 0011, Valerie J. Shute, Nicholas D. Duran, Sidney K. D'Mello
ICMI9
2019 Language as Thought: Using Natural Language Processing to Model Noncognitive Traits that Predict College Success
abstract
It is widely acknowledged that the language we use reflects numerous psychological constructs, including our thoughts, feelings, and desires. Can the so called "noncognitive" traits with known links to success, such as growth mindset, leadership ability, and intrinsic motivation, be similarly revealed through language? We investigated this question by analyzing students' 150-word open-ended descriptions of their own extracurricular activities or work experiences included in their college applications. We used the Common Application-National Student Clearinghouse data set, a six-year longitudinal dataset that includes college application data and graduation outcomes for 278,201 U.S. high-school students. We first developed a coding scheme from a stratified sample of 4,000 essays and used it to code seven traits: growth mindset, perseverance, goal orientation, leadership, psychological connection (intrinsic motivation), self-transcendent (prosocial) purpose, and team orientation, along with earned accolades. Then, we used standard classifiers with bag-of-n-grams as features and deep learning techniques (recurrent neural networks) with word embeddings to automate the coding. The models demonstrated convergent validity with the human coding with AUCs ranging from .770 to .925 and correlations ranging from .418 to .734. There was also evidence of discriminant validity in the pattern of inter-correlations (rs between -.206 to .306) for both human- and model-coded traits. Finally, the models demonstrated incremental predictive validity in predicting six-year graduation outcomes net of sociodemographics, intelligence, academic achievement, and institutional graduation rates. We conclude that language provides a lens into noncognitive traits important for college success, which can be captured with automated methods.
Cathlyn Stone, Abigail Quirk, Margo Gardner, Stephen Hutt, Angela Lee Duckworth, Sidney K. D'Mello
LAK6
2019 I Say, You Say, We Say: Using Spoken Language to Model Socio-Cognitive Processes during Computer-Supported Collaborative Problem Solving
abstract
Collaborative problem solving (CPS) is a crucial 21st century skill; however, current technologies fall short of effectively supporting CPS processes, especially for remote, computer-enabled interactions. In order to develop next-generation computer-supported collaborative systems that enhance CPS processes and outcomes by monitoring and responding to the unfolding collaboration, we investigate automated detection of three critical CPS process ? construction of shared knowledge, negotiation/coordination, and maintaining team function ? derived from a validated CPS framework. Our data consists of 32 triads who were tasked with collaboratively solving a challenging visual computer programming task for 20 minutes using commercial videoconferencing software. We used automatic speech recognition to generate transcripts of 11,163 utterances, which trained humans coded for evidence of the above three CPS processes using a set of behavioral indicators. We aimed to automate the trained human-raters' codes in a team-independent fashion (current study) in order to provide automatic real-time or offline feedback (future work). We used Random Forest classifiers trained on the words themselves (bag of n-grams) or with word categories (e.g., emotions, thinking styles, social constructs) from the Linguistic Inquiry Word Count (LIWC) tool. Despite imperfect automatic speech recognition, the n-gram models achieved AUROC (area under the receiver operating characteristic curve) scores of .85, .77, and .77 for construction of shared knowledge, negotiation/coordination, and maintaining team function, respectively; these reflect 70%, 54%, and 54% improvements over chance. The LIWC-category models achieved similar scores of .82, .74, and .73 (64%, 48%, and 46% improvement over chance). Further, the LIWC model-derived scores predicted CPS outcomes more similar to human codes, demonstrating predictive validity. We discuss embedding our models in collaborative interfaces for assessment and dynamic intervention aimed at improving CPS outcomes.
Angela Stewart, Hana Vrzakova, Chen Sun 0011, Jade Yonehiro, Cathlyn Stone, Nicholas D. Duran, Valerie J. Shute, Sidney K. D'Mello
Proc. ACM Hum. Comput. Interact.8
2019 Automated gaze-based mind wandering detection during computerized learning in classrooms
Stephen Hutt, Kristina Krasich, Caitlin Mills 0001, Nigel Bosch, Shelby White, James R. Brockmole, Sidney K. D'Mello
User Model. User Adapt. Interact.7
2018 "Mind" TS: Testing a Brief Mindfulness Intervention with an Intelligent Tutoring System
Kristina Krasich, Stephen Hutt, Caitlin Mills 0001, Catherine A. Spann, James R. Brockmole, Sidney K. D'Mello
AIED (2)6
2018 Connecting the Dots Towards Collaborative AIED: Linking Group Makeup to Process to Learning
Angela Stewart, Sidney K. D'Mello
AIED (1)2
2018 Mind wandering during conversations affects subjective but not objective outcomes
Myrthe Faber, McKenzie Rees, Sidney K. D'Mello
CogSci3
2018 Predicting Reading Comprehension From Eye Gaze
Julie M. Gregg, Sidney K. D'Mello
CogSci2
2018 An Open Vocabulary Approach for Estimating Teacher Use of Authentic Questions in Classroom Discourse
Connor Cook, Andrew Olney, Sean Kelly, Sidney K. D'Mello
EDM4
2018 Generative Multimodal Models of Nonverbal Synchrony in Close Relationships
abstract
Positive interpersonal relationships require shared understanding along with a sense of rapport. A key facet of rapport is mirroring and convergence of facial expression and body language, known as nonverbal synchrony. We examined nonverbal synchrony in a study of 29 heterosexual romantic couples, in which audio, video, and bracelet accelerometer were recorded during three conversations. We extracted facial expression, body movement, and acoustic-prosodic features to train neural network models that predicted the nonverbal behaviors of one partner from those of the other. Recurrent models (LSTMs) outperformed feed-forward neural networks and other chance baselines. The models learned behaviors encompassing facial responses, speech-related facial movements, and head movement. However, they did not capture fleeting or periodic behaviors, such as nodding, head turning, and hand gestures. Notably, a preliminary analysis of clinical measures showed greater association with our model outputs than correlation of raw signals. We discuss potential uses of these generative models as a research tool to complement current analytical methods along with real-world applications (e.g., as a tool in therapy).
Joseph F. Grafsgaard, Nicholas D. Duran, Ashley K. Randall, Chun Tao, Sidney K. D'Mello
FG5
2018 Multimodal Modeling of Coordination and Coregulation Patterns in Speech Rate during Triadic Collaborative Problem Solving
abstract
We model coordination and coregulation patterns in 33 triads engaged in collaboratively solving a challenging computer programming task for approximately 20 minutes. Our goal is to prospectively model speech rate (words/sec) - an important signal of turn taking and active participation - of one teammate (A or B or C) from time lagged nonverbal signals (speech rate and acoustic-prosodic features) of the other two (i.e., A + B → C; A + C → B; B + C → A) and task-related context features. We trained feed-forward neural networks (FFNNs) and long short-term memory recurrent neural networks (LSTMs) using group-level nested cross-validation. LSTMs outperformed FFNNs and a chance baseline and could predict speech rate up to 6s into the future. A multimodal combination of speech rate, acoustic-prosodic, and task context features outperformed unimodal and bimodal signals. The extent to which the models could predict an individual's speech rate was positively related to that individual's scores on a subsequent posttest, suggesting a link between coordination/coregulation and collaborative learning outcomes. We discuss applications of the models for real-time systems that monitor the collaborative process and intervene to promote positive collaborative outcomes.
Angela Stewart, Zachary A. Keirn, Sidney K. D'Mello
ICMI3
2018 Prospectively predicting 4-year college graduation from student applications
abstract
We leverage a unique national dataset of 41,359 college applications to prospectively predict 4-year bachelor's graduation in a generalizable manner. Our features include sociodemographics, institutional graduation rates, academic achievement, standardized test scores, engagement in extracurricular activities, work experiences, and ratings by teachers and high-school guidance counselors. A random forest classifier successfully predicted 4-year graduation for 71.4% of the students (base rate = 44%) using all 166 of the aforementioned features and a split-half validation method. A stochastic hill-climbing feature selection procedure effectively maintained the same classification accuracy, but with a minimal set of 37 features, consisting of an approximately equal representation of sociodemographics, cognitive, and noncognitive factors. We advocate against using these results for admissions decisions, instead contemplating how they might be used to provide parents and educators with actionable information to guide students towards college success.
Stephen Hutt, Margo Gardner, Donald Kamentz, Angela Lee Duckworth, Sidney K. D'Mello
LAK5
2017 Face Forward: Detecting Mind Wandering from Video During Narrative Film Comprehension
Angela Stewart, Nigel Bosch, Huili Chen, Patrick J. Donnelly, Sidney K. D'Mello
AIED5
2017 What's on your wandering mind? The content of mind wandering during text- and film comprehension
Myrthe Faber, Sidney K. D'Mello
CogSci2
2017 Zone out no more: Mitigating mind wandering during computerized reading
Sidney K. D'Mello, Caitlin Mills 0001, Robert Bixler, Nigel Bosch
EDM1
2017 Gaze-based Detection of Mind Wandering during Lecture Viewing
Stephen Hutt, Jessica Hardey, Robert Bixler, Angela Stewart, Evan F. Risko, Sidney K. D'Mello
EDM6
2017 Tracking Online Reading of College Students
Andrew Olney, Eric Hosman, Arthur C. Graesser, Sidney K. D'Mello
EDM4
2017 Assessing the Dialogic Properties of Classroom Discourse: Proportion Models for Imbalanced Classes
Andrew Olney, Borhan Samei, Patrick J. Donnelly, Sidney K. D'Mello
EDM4
2017 Generalizability of Face-Based Mind Wandering Detection Across Task Contexts
Angela Stewart, Nigel Bosch, Sidney K. D'Mello
EDM3
2017 Words matter: automatic detection of teacher questions in live classroom discourse using linguistics, acoustics, and context
abstract
We investigate automatic detection of teacher questions from audio recordings collected in live classrooms with the goal of providing automated feedback to teachers. Using a dataset of audio recordings from 11 teachers across 37 class sessions, we automatically segment the audio into individual teacher utterances and code each as containing a question or not. We train supervised machine learning models to detect the human-coded questions using high-level linguistic features extracted from automatic speech recognition (ASR) transcripts, acoustic and prosodic features from the audio recordings, as well as context features, such as timing and turn-taking dynamics. Models are trained and validated independently of the teacher to ensure generalization to new teachers. We are able to distinguish questions and non-questions with a weighted F1 score of 0.69. A comparison of the three feature sets indicates that a model using linguistic features outperforms those using acoustic-prosodic and context features for question detection, but the combination of features yields a 5% improvement in overall accuracy compared to linguistic features alone. We discuss applications for pedagogical research, teacher formative assessment, and teacher professional development.
Patrick J. Donnelly, Nathaniel Blanchard, Andrew Olney, Sean Kelly, Martin Nystrand, Sidney K. D'Mello
LAK6
2017 Put your thinking cap on: detecting cognitive load using EEG during learning
abstract
Current learning technologies have no direct way to assess students' mental effort: are they in deep thought, struggling to overcome an impasse, or are they zoned out? To address this challenge, we propose the use of EEG-based cognitive load detectors during learning. Despite its potential, EEG has not yet been utilized as a way to optimize instructional strategies. We take an initial step towards this goal by assessing how experimentally manipulated (easy and difficult) sections of an intelligent tutoring system (ITS) influenced EEG-based estimates of students' cognitive load. We found a main effect of task difficulty on EEG-based cognitive load estimates, which were also correlated with learning performance. Our results show that EEG can be a viable source of data to model learners' mental states across a 90-minute session.
Caitlin Mills 0001, Igor Fridman, Walid Soussou, Disha Waghray, Andrew Olney, Sidney K. D'Mello
LAK6
2017 "Out of the Fr-Eye-ing Pan": Towards Gaze-Based Models of Attention during Learning with Technology in the Classroom
abstract
Attention is critical to learning. Hence, advanced learning technologies should benefit from mechanisms to monitor and respond to learners' attentional states. We study the feasibility of integrating commercial off-the-shelf (COTS) eye trackers to monitor attention during interactions with a learning technology called GuruTutor. We tested our implementation on 135 students in a noisy computer-enabled high school classroom and were able to collect a median 95% valid eye gaze data in 85% of the sessions where gaze data was successfully recorded. Machine learning methods were employed to develop automated detectors of mind wandering (MW) -- a phenomenon involving a shift in attention from task-related to task-unrelated thoughts that is negatively correlated with performance. Our student-independent, gaze-based models could detect MW with an accuracy (F1 of MW = 0.59) significantly greater than chance (F1 of MW = 0.24). Predicted rates of mind wandering were negatively related to posttest performance, providing evidence for the predictive validity of the detector. We discuss next steps towards developing gaze-based, attention-aware, learning technologies that can be deployed in noisy, real-world environments.
Stephen Hutt, Caitlin Mills 0001, Nigel Bosch, Kristina Krasich, James R. Brockmole, Sidney K. D'Mello
UMAP6
2017 ETGraph: A graph-based approach for visual analytics of eye-tracking data
Chaoli Wang 0001, Robert Bixler, Sidney K. D'Mello
Comput. Graph.4
2017 Automated Detection of Engagement Using Video-Based Estimation of Facial Expressions and Heart Rate
abstract
We explored how computer vision techniques can be used to detect engagement while students (N = 22) completed a structured writing activity (draft-feedback-review) similar to activities encountered in educational settings. Students provided engagement annotations both concurrently during the writing activity and retrospectively from videos of their faces after the activity. We used computer vision techniques to extract three sets of features from videos, heart rate, Animation Units (from Microsoft Kinect Face Tracker), and local binary patterns in three orthogonal planes (LBP-TOP). These features were used in supervised learning for detection of concurrent and retrospective self-reported engagement. Area under the ROC Curve (AUC) was used to evaluate classifier accuracy using leave-several-students-out cross validation. We achieved an AUC = .758 for concurrent annotations and AUC = .733 for retrospective annotations. The Kinect Face Tracker features produced the best results among the individual channels, but the overall best results were found using a fusion of channels.
Hamed Monkaresi, Nigel Bosch, Rafael A. Calvo, Sidney K. D'Mello
IEEE Trans. Affect. Comput.4
2016 Mind Wandering during Film Comprehension: The Role of Prior Knowledge and Situational Interest
Sidney K. D'Mello, Kristopher Kopp, Caitlin Mills 0001
CogSci1
2016 The effect of disfluency on mind wandering during text comprehension
Myrthe Faber, Caitlin Mills 0001, Kristopher Kopp, Sidney K. D'Mello
CogSci4
2016 Semi-Automatic Detection of Teacher Questions from Human-Transcripts of Audio in Live Classrooms
Nathaniel Blanchard, Patrick J. Donnelly, Andrew Olney, Borhan Samei, Sean Kelly, Xiaoyi Sun, Brooke Ward, Martin Nystrand, Sidney K. D'Mello
EDM9
2016 Student Emotion, Co-occurrence, and Dropout in a MOOC Context
John Z. Dillon, Nigel Bosch, Malolan Chetlur, Nirandika Wanigasekara, G. Alex Ambrose, Bikram Sengupta, Sidney K. D'Mello
EDM7
2016 The Eyes Have It: Gaze-based Detection of Mind Wandering during Learning with an Intelligent Tutoring System
Stephen Hutt, Caitlin Mills 0001, Shelby White, Patrick J. Donnelly, Sidney K. D'Mello
EDM5
2016 Automatic Gaze-Based Detection of Mind Wandering during Film Viewing
Caitlin Mills 0001, Robert Bixler, Sidney K. D'Mello
EDM4
2016 Multi-sensor modeling of teacher instructional segments in live classrooms
abstract
We investigate multi-sensor modeling of teachers’ instructional segments (e.g., lecture, group work) from audio recordings collected in 56 classes from eight teachers across five middle schools. Our approach fuses two sensors: a unidirectional microphone for teacher audio and a pressure zone microphone for general classroom audio. We segment and analyze the audio streams with respect to discourse timing, linguistic, and paralinguistic features. We train supervised classifiers to identify the five instructional segments that collectively comprised a majority of the data, achieving teacher-independent F1 scores ranging from 0.49 to 0.60. With respect to individual segments, the individual sensor models and the fused model were on par for Question & Answer and Procedures & Directions segments. For Supervised Seatwork, Small Group Work, and Lecture segments, the classroom model outperformed both the teacher and fusion models. Across all segments, a multi-sensor approach led to an average 8% improvement over the state of the art approach that only analyzed teacher audio. We discuss implications of our findings for the emerging field of multimodal learning analytics.
Patrick J. Donnelly, Nathaniel Blanchard, Borhan Samei, Andrew Olney, Xiaoyi Sun, Brooke Ward, Sean Kelly, Martin Nystrand, Sidney K. D'Mello
ICMI9
2016 Detecting Student Emotions in Computer-Enabled Classrooms
Nigel Bosch, Sidney K. D'Mello, Ryan Baker 0001, Jaclyn Ocumpaugh, Valerie J. Shute, Matthew Ventura, Weinan Zhao
IJCAI2
2016 Investigating boredom and engagement during writing using multiple sources of information: the essay, the writer, and keystrokes
abstract
Writing training systems have been developed to provide students with instruction and deliberate practice on their writing. Although generally successful in providing accurate scores, a common criticism of these systems is their lack of personalization and adaptive instruction. In particular, these systems tend to place the strongest emphasis on delivering accurate scores, and therefore, tend to overlook additional indices that may contribute to students' success, such as their affective states during writing practice. This study takes an initial step toward addressing this gap by building a predictive model of students' affect using information that can potentially be collected by computer systems. We used individual difference measures, text indices, and keystroke analyses to predict engagement and boredom in 132 writing sessions. The results suggest that these three categories of indices were successful in modeling students' affective states during writing. Taken together, indices related to students' academic abilities, text properties, and keystroke logs were able classify high and low engagement and boredom in writing sessions with accuracies between 76.5% and 77.3%. These results suggest that information readily available in writing training systems can inform affect detectors and ultimately improve student models within intelligent tutoring systems.
Laura K. Allen, Caitlin Mills 0001, Matthew E. Jacovina, Scott A. Crossley, Sidney K. D'Mello, Danielle S. McNamara
LAK5
2016 Student affect during learning with a MOOC
abstract
This paper presents affect data collected from periodic emotion detection surveys throughout an introductory Statistics MOOC called "I Heart Stats." This is the first MOOC, to our knowledge, to capture valuable student affect data through self-reported surveys. To collect student affect, we used two self-reporting methods: (1) The Self-Assessment Manikin and (2) A discrete emotion list. We found that the most common reported MOOC emotion was Hope followed by Enjoyment and Contentment. There were substantial shifts in affective states over the course, notably with Anxiety and Pride. The most valuable result of our study is a preliminary description of the methods for collecting self-reported student affect at scale in a MOOC setting.
John Z. Dillon, G. Alex Ambrose, Nirandika Wanigasekara, Malolan Chetlur, Bikram Sengupta, Sidney K. D'Mello
LAK7
2016 Identifying Teacher Questions Using Automatic Speech Recognition in Classrooms
abstract
Nathaniel Blanchard, Patrick Donnelly, Andrew M. Olney, Borhan Samei, Brooke Ward, Xiaoyi Sun, Sean Kelly, Martin Nystrand, Sidney K. D’Mello. Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2016.
Nathaniel Blanchard, Patrick J. Donnelly, Andrew Olney, Borhan Samei, Brooke Ward, Xiaoyi Sun, Sean Kelly, Martin Nystrand, Sidney K. D'Mello
SIGDIAL Conference9
2016 Automatic Teacher Modeling from Live Classroom Audio
abstract
We investigate automatic analysis of teachers' instructional strategies from audio recordings collected in live classrooms. We collected a data set of teacher audio and human-coded instructional activities (e.g., lecture, question and answer, group work) in 76 middle school literature, language arts, and civics classes from eleven teachers across six schools. We automatically segment teacher audio to analyze speech vs. rest patterns, generate automatic transcripts of the teachers' speech to extract natural language features, and compute low-level acoustic features. We train supervised machine learning models to identify occurrences of five key instructional segments (Question & Answer, Procedures and Directions, Supervised Seatwork, Small Group Work, and Lecture) that collectively comprise 76% of the data. Models are validated independently of teacher in order to increase generalizability to new teachers from the same sample. We were able to identify the five instructional segments above chance levels with F1 scores ranging from 0.64 to 0.78. We discuss key findings in the context of teacher modeling for formative assessment and professional development.
Patrick J. Donnelly, Nathaniel Blanchard, Borhan Samei, Andrew Olney, Xiaoyi Sun, Brooke Ward, Sean Kelly, Martin Nystrand, Sidney K. D'Mello
UMAP9
2016 Where's Your Mind At?: Video-Based Mind Wandering Detection During Film Viewing
abstract
Mind wandering (MW) is a ubiquitous phenomenon in which attention involuntarily shifts from task-related processing to task-unrelated thoughts. This study reports preliminary results of a video-based MW detector during film viewing. We collected training data in a study where participants self-reported when they caught themselves MW over the course of watching a 32.5 minute commercial film. We trained classification models on automatically extracted facial features and bodily movement and were able to detect MW with an F1 of .30. The model was successful in reproducing the MW distribution obtained from the self-reports
Angela Stewart, Nigel Bosch, Huili Chen, Patrick J. Donnelly, Sidney K. D'Mello
UMAP5
2016 On the Influence of an Iterative Affect Annotation Approach on Inter-Observer and Self-Observer Reliability
abstract
Affect detection systems require reliable methods to annotate affective data. Typically, two or more observers independently annotate audio-visual affective data. This approach results in inter-observer reliabilities that can be categorized as fair (Cohen's kappas of approximately .40). In an alternative iterative approach, observers independently annotate small amounts of data, discuss their annotations, and annotate a different sample of data. After a pre-determined reliability threshold is reached, the observers independently annotate the remainder of the data. The effectiveness of the iterative approach was tested in an annotation study where pairs of observers annotated affective video data in nine annotate-discuss iterations. Self-annotations were previously collected on the same data. Mixed effects linear regression models indicated that inter-observer agreement increased (unstandardized coefficient B = .031) across iterations, with agreement in the final iteration reflecting a 64 percent improvement over the first iteration. Follow-up analyses indicated that the improvement was nonlinear in that most of the improvement occurred after the first three iterations (B = .043), after which agreement plateaued (B ≈ 0). There was no notable complementary improvement (B ≈ 0) in self-observer agreement, which was considerably lower than observer-observer agreement. Strengths, limitations, and applications of the iterative affective annotation approach are discussed.
Sidney K. D'Mello
IEEE Trans. Affect. Comput.1
2016 Using Video to Automatically Detect Learner Affect in Computer-Enabled Classrooms
abstract
Affect detection is a key component in intelligent educational interfaces that respond to students’ affective states. We use computer vision and machine-learning techniques to detect students’ affect from facial expressions (primary channel) and gross body movements (secondary channel) during interactions with an educational physics game. We collected data in the real-world environment of a school computer lab with up to 30 students simultaneously playing the game while moving around, gesturing, and talking to each other. The results were cross-validated at the student level to ensure generalization to new students. Classification accuracies, quantified as area under the receiver operating characteristic curve (AUC), were above chance (AUC of 0.5) for all the affective states observed, namely, boredom (AUC = .610), confusion (AUC = .649), delight (AUC = .867), engagement (AUC = .679), frustration (AUC = .631), and for off-task behavior (AUC = .816). Furthermore, the detectors showed temporal generalizability in that there was less than a 2% decrease in accuracy when tested on data collected from different times of the day and from different days. There was also some evidence of generalizability across ethnicity (as perceived by human coders) and gender, although with a higher degree of variability attributable to differences in affect base rates across subpopulations. In summary, our results demonstrate the feasibility of generalizable video-based detectors of naturalistic affect in a real-world setting, suggesting that the time is ripe for affect-sensitive interventions in educational games and other intelligent interfaces.
Nigel Bosch, Sidney K. D'Mello, Jaclyn Ocumpaugh, Ryan Baker 0001, Valerie J. Shute
ACM Trans. Interact. Intell. Syst.2
2016 Automatic gaze-based user-independent detection of mind wandering during computerized reading
Robert Bixler, Sidney K. D'Mello
User Model. User Adapt. Interact.2
2015 A Study of Automatic Speech Recognition in Noisy Classroom Environments for Automated Dialog Analysis
Nathaniel Blanchard, Michael Brady 0003, Andrew Olney, Marci Glaus, Xiaoyi Sun, Martin Nystrand, Borhan Samei, Sean Kelly, Sidney K. D'Mello
AIED9
2015 Temporal Generalizability of Face-Based Affect Detection in Noisy Classroom Environments
Nigel Bosch, Sidney K. D'Mello, Ryan Baker 0001, Jaclyn Ocumpaugh, Valerie J. Shute
AIED2
2015 Mind Wandering During Learning with an Intelligent Tutoring System
Caitlin Mills 0001, Sidney K. D'Mello, Nigel Bosch, Andrew Olney
AIED2
2015 Classifying Q&A from Teachers' Speech: Moving Toward an Automated System of Dialogic Analysis
Nathaniel Blanchard, Sidney K. D'Mello, Andrew Olney, Martin Nystrand
EDM2
2015 Video-Based Affect Detection in Noninteractive Learning Environments
Nigel Bosch, Sidney K. D'Mello
EDM3
2015 Breaking Off Engagement: Readers' Cognitive Decoupling as a Function of Reader and Text Characteristics
Patricia Goedecke, Daqi Dong, Genghu Shi, Evan F. Risko, Andrew Olney, Sidney K. D'Mello, Arthur C. Graesser
EDM7
2015 A Comparison of Face-based and Interaction-based Affect Detectors in Physics Playground
Shiming Kai, Luc Paquette, Ryan Baker 0001, Nigel Bosch, Sidney K. D'Mello, Jaclyn Ocumpaugh, Valerie J. Shute, Matthew Ventura
EDM5
2015 Toward a Real-time (Day) Dreamcatcher: Detecting Mind Wandering Episodes During Online Reading
Caitlin Mills 0001, Sidney K. D'Mello
EDM2
2015 Modeling Classroom Discourse: Do Models of Predicting Dialogic Instruction Properties Generalize across Populations?
Borhan Samei, Andrew Olney, Sean Kelly, Martin Nystrand, Sidney K. D'Mello, Nathaniel Blanchard, Arthur C. Graesser
EDM5
2015 Automatic Detection of Mind Wandering During Reading Using Gaze and Physiology
abstract
Mind wandering (MW) entails an involuntary shift in attention from task-related thoughts to task-unrelated thoughts, and has been shown to have detrimental effects on performance in a number of contexts. This paper proposes an automated multimodal detector of MW using eye gaze and physiology (skin conductance and skin temperature) and aspects of the context (e.g., time on task, task difficulty). Data in the form of eye gaze and physiological signals were collected as 178 participants read four instructional texts from a computer interface. Participants periodically provided self-reports of MW in response to pseudorandom auditory probes during reading. Supervised machine learning models trained on features extracted from participants' gaze fixations, physiological signals, and contextual cues were used to detect pages where participants provided positive responses of MW to the auditory probes. Two methods of combining gaze and physiology features were explored. Feature level fusion entailed building a single model by combining feature vectors from individual modalities. Decision level fusion entailed building individual models for each modality and adjudicating amongst individual decisions. Feature level fusion resulted in an 11% improvement in classification accuracy over the best unimodal model, but there was no comparable improvement for decision level fusion. This was reflected by a small improvement in both precision and recall. An analysis of the features indicated that MW was associated with fewer and longer fixations and saccades, and a higher more deterministic skin temperature. Possible applications of the detector are discussed.
Robert Bixler, Nathaniel Blanchard, Luke Garrison, Sidney K. D'Mello
ICMI4
2015 Accuracy vs. Availability Heuristic in Multimodal Affect Detection in the Wild
abstract
This paper discusses multimodal affect detection from a fusion of facial expressions and interaction features derived from students' interactions with an educational game in the noisy real-world context of a computer-enabled classroom. Log data of students' interactions with the game and face videos from 133 students were recorded in a computer-enabled classroom over a two day period. Human observers live annotated learning-centered affective states such as engagement, confusion, and frustration. The face-only detectors were more accurate than interaction-only detectors. Multimodal affect detectors did not show any substantial improvement in accuracy over the face-only detectors. However, the face-only detectors were only applicable to 65% of the cases due to face registration errors caused by excessive movement, occlusion, poor lighting, and other factors. Multimodal fusion techniques were able to improve the applicability of detectors to 98% of cases without sacrificing classification accuracy. Balancing the accuracy vs. applicability tradeoff appears to be an important feature of multimodal affect detection.
Nigel Bosch, Huili Chen, Sidney K. D'Mello, Ryan Baker 0001, Valerie J. Shute
ICMI3
2015 Multimodal Capture of Teacher-Student Interactions for Automated Dialogic Analysis in Live Classrooms
abstract
We focus on data collection designs for the automated analysis of teacher-student interactions in live classrooms with the goal of identifying instructional activities (e.g., lecturing, discussion) and assessing the quality of dialogic instruction (e.g., analysis of questions). Our designs were motivated by multiple technical requirements and constraints. Most importantly, teachers could be individually micfied but their audio needed to be of excellent quality for automatic speech recognition (ASR) and spoken utterance segmentation. Individual students could not be micfied but classroom audio quality only needed to be sufficient to detect student spoken utterances. Visual information could only be recorded if students could not be identified. Design 1 used an omnidirectional laptop microphone to record both teacher and classroom audio and was quickly deemed unsuitable. In Designs 2 and 3, teachers wore a wireless Samson AirLine 77 vocal headset system, which is a unidirectional microphone with a cardioid pickup pattern. In Design 2, classroom audio was recorded with dual first- generation Microsoft Kinects placed at the front corners of the class. Design 3 used a Crown PZM-30D pressure zone microphone mounted on the blackboard to record classroom audio. Designs 2 and 3 were tested by recording audio in 38 live middle school classrooms from six U.S. schools while trained human coders simultaneously performed live coding of classroom discourse. Qualitative and quantitative analyses revealed that Design 3 was suitable for three of our core tasks: (1) ASR on teacher speech (word recognition rate of 66% and word overlap rate of 69% using Google Speech ASR engine); (2) teacher utterance segmentation (F-measure of 97%); and (3) student utterance segmentation (F-measure of 66%). Ideas to incorporate video and skeletal tracking with dual second-generation Kinects to produce Design 4 are discussed.
Sidney K. D'Mello, Andrew Olney, Nathaniel Blanchard, Borhan Samei, Xiaoyi Sun, Brooke Ward, Sean Kelly
ICMI1
2015 Automatic Detection of Learning-Centered Affective States in the Wild
abstract
Affect detection is a key component in developing intelligent educational interfaces that are capable of responding to the affective needs of students. In this paper, computer vision and machine learning techniques were used to detect students' affect as they used an educational game designed to teach fundamental principles of Newtonian physics. Data were collected in the real-world environment of a school computer lab, which provides unique challenges for detection of affect from facial expressions (primary channel) and gross body movements (secondary channel) - up to thirty students at a time participated in the class, moving around, gesturing, and talking to each other. Results were cross validated at the student level to ensure generalization to new students. Classification was successful at levels above chance for off-task behavior (area under receiver operating characteristic curve or (AUC = .816) and each affective state including boredom (AUC =.610), confusion (.649), delight (.867), engagement (.679), and frustration (.631) as well as a five-way overall classification of affect (.655), despite the noisy nature of the data. Implications and prospects for affect-sensitive interfaces for educational software in classroom environments are discussed.
Nigel Bosch, Sidney K. D'Mello, Ryan Baker 0001, Jaclyn Ocumpaugh, Valerie J. Shute, Matthew Ventura, Weinan Zhao
IUI2
2015 Automatic Gaze-Based Detection of Mind Wandering with Metacognitive Awareness
Robert Bixler, Sidney K. D'Mello
UMAP2
2015 Introduction to the "Best of ACII 2013" Special Section
abstract
The papers in this special section were presented at the 6th Biannual International Conference on Affective Computing and Intelligent Interaction (ACII 2013), held from September 2-5, 2013, in Geneva, Switzerland. These papers reflect the diversity of affective computing research, ranging from physiological models for affect detection, methodological issues, computational models of affect, and human perceptions of affective virtual agents.
Sidney K. D'Mello, Maja Pantic, Anton Nijholt
IEEE Trans. Affect. Comput.1
2014 Gesturing May Not Always Make Learning Last
Caroline Hornburg, Nicole M. McNeil, Sidney K. D'Mello, Susan Wagner Cook
CogSci3
2014 Automatic assessment of student reading comprehension from short summaries
Lisa Mintz, Dan Stefanescu, Sidney K. D'Mello, Arthur C. Graesser
EDM4
2014 Domain Independent Assessment of Dialogic Properties of Classroom Discourse
Borhan Samei, Andrew Olney, Sean Kelly, Martin Nystrand, Sidney K. D'Mello, Nathaniel Blanchard, Xiaoyi Sun, Marci Glaus, Arthur C. Graesser
EDM5
2014 Improving automated source code summarization via an eye-tracking study of programmers
abstract
Source Code Summarization is an emerging technology for automatically generating brief descriptions of code. Current summarization techniques work by selecting a subset of the statements and keywords from the code, and then including information from those statements and keywords in the summary. The quality of the summary depends heavily on the process of selecting the subset: a high-quality selection would contain the same statements and keywords that a programmer would choose. Unfortunately, little evidence exists about the statements and keywords that programmers view as important when they summarize source code. In this paper, we present an eye-tracking study of 10 professional Java programmers in which the programmers read Java methods and wrote English summaries of those methods. We apply the findings to build a novel summarization tool. Then, we evaluate this tool and provide evidence to support the development of source code summarization systems.
Paige Rodeghero, Collin McMillan, Paul W. McBurney, Nigel Bosch, Sidney K. D'Mello
ICSE5
2014 Automated Physiological-Based Detection of Mind Wandering during Learning
Nathaniel Blanchard, Robert Bixler, Tera Joyce, Sidney K. D'Mello
Intelligent Tutoring Systems4
2014 It's Written on Your Face: Detecting Affective States from Facial Expressions while Learning Computer Programming
Nigel Bosch, Sidney K. D'Mello
Intelligent Tutoring Systems3
2014 It Takes Two: Momentary Co-occurrence of Affective States during Computerized Learning
Nigel Bosch, Sidney K. D'Mello
Intelligent Tutoring Systems2
2014 Identifying Learning Conditions that Minimize Mind Wandering by Modeling Individual Attributes
Kristopher Kopp, Robert Bixler, Sidney K. D'Mello
Intelligent Tutoring Systems3
2014 To Quit or Not to Quit: Predicting Future Behavioral Disengagement from Reading Patterns
Caitlin Mills 0001, Nigel Bosch, Arthur C. Graesser, Sidney K. D'Mello
Intelligent Tutoring Systems4
2014 Toward Fully Automated Person-Independent Detection of Mind Wandering
Robert Bixler, Sidney K. D'Mello
UMAP2
2013 Towards Automated Detection and Regulation of Affective States During Academic Writing
Robert Bixler, Sidney K. D'Mello
AIED2
2013 Programming with Your Heart on Your Sleeve: Analyzing the Affective States of Computer Programming Students
Nigel Bosch, Sidney K. D'Mello
AIED2
2013 What Emotions Do Novices Experience during Their First Computer Programming Learning Session?
Nigel Bosch, Sidney K. D'Mello, Caitlin Mills 0001
AIED2
2013 Who Benefits from Confusion Induction during Learning? An Individual Differences Cluster Analysis
Blair Lehman, Sidney K. D'Mello, Arthur C. Graesser
AIED2
2013 Sorry, I Must Have Zoned Out: Tracking Mind Wandering Episodes in an Interactive Learning Environment
Caitlin Mills 0001, Sidney K. D'Mello
AIED2
2013 What Makes Learning Fun? Exploring the Influence of Choice and Difficulty on Mind Wandering and Engagement during Learning
Caitlin Mills 0001, Sidney K. D'Mello, Blair Lehman, Nigel Bosch, Amber Chauncey Strain, Arthur C. Graesser
AIED2
2013 Automatic Gaze-Based Detection of Mind Wandering during Reading
Sidney K. D'Mello, Jonathan Cobian, Matthew Hunter
EDM1
2013 Reading into the Text: Investigating the Influence of Text Complexity on Cognitive Engagement
Benjamin Vega, Blair Lehman, Arthur C. Graesser, Sidney K. D'Mello
EDM5
2013 Affect Detection and Classification from the Non-stationary Physiological Data
abstract
Affect detection from physiological signals has received a great deal of attention recently. One arising challenge is that physiological measures are expected to exhibit considerable variations or non-stationarities over multiple days/sessions recordings. These variations pose challenges to effectively classify affective sates from future physiological data. The present study collects affective physiological data (electrocardiogram (ECG), electromyogram (EMG), skin conductivity (SC), and respiration (RSP)) from four participants over five sessions each. The study provides insights on how diagnostic physiological features of affect change over time. We compare the classification performance of two feature sets, pooled features (obtained from pooled day data) and day-specific features using an up datable classifier ensemble algorithm. The study also provides an analysis on the performance of individual physiological channels for affect detection. Our results show that using pooled feature set for affect detection is more accurate than using day-specific features. The corrugator and zygomatic facial EMGs were more reliable measures for detecting valence than arousal compared to ECG, RSP and SC over the span of multi-session recordings. It is also found that corrugator EMG features and a fusion of features from all physiological channels have the highest affect detection accuracy for both valence and arousal.
Omar AlZoubi, Davide Fossati, Sidney K. D'Mello, Rafael A. Calvo
ICMLA (1)3
2013 Detecting boredom and engagement during writing with keystroke analysis, task appraisals, and stable traits
abstract
It is hypothesized that the ability for a system to automatically detect and respond to users' affective states can greatly enhance the human-computer interaction experience. Although there are currently many options for affect detection, keystroke analysis offers several attractive advantages to traditional methods. In this paper, we consider the possibility of automatically discriminating between natural occurrences of boredom, engagement, and neutral by analyzing keystrokes, task appraisals, and stable traits of 44 individuals engaged in a writing task. The analyses explored several different arrangements of the data: using downsampled and/or standardized data; distinguishing between three different affect states or groups of two; and using keystroke/timing features in isolation or coupled with stable traits and/or task appraisals. The results indicated that the use of raw data and the feature set that combined keystroke/timing features with task appraisals and stable traits, yielded accuracies that were 11% to 38% above random guessing and generalized to new individuals. Applications of our affect detector for intelligent interfaces that provide engagement support during writing are discussed.
Robert Bixler, Sidney K. D'Mello
IUI2
2013 Unimodal and Multimodal Human Perceptionof Naturalistic Non-Basic Affective Statesduring Human-Computer Interactions
abstract
The present study investigated unimodal and multimodal emotion perception by humans, with an eye for applying the findings towards automated affect detection. The focus was on assessing the reliability by which untrained human observers could detect naturalistic expressions of non-basic affective states (boredom, engagement/flow, confusion, frustration, and neutral) from previously recorded videos of learners interacting with a computer tutor. The experiment manipulated three modalities to produce seven conditions: face, speech, context, face+speech, face+context, speech+context, face+speech+context. Agreement between two observers (OO) and between an observer and a learner (LO) were computed and analyzed with mixed-effects logistic regression models. The results indicated that agreement was generally low (kappas ranged from .030 to .183), but, with one exception, was greater than chance. Comparisons of overall agreement (across affective states) between the unimodal and multimodal conditions supported redundancy effects between modalities, but there were superadditive, additive, redundant, and inhibitory effects when affective states were individually considered. There was both convergence and divergence of patterns in the OO and LO data sets; however, LO models yielded lower agreement but higher multimodal effects compared to OO models. Implications of the findings for automated affect detection are discussed.
Sidney K. D'Mello, Nia Nixon, Arthur C. Graesser
IEEE Trans. Affect. Comput.1
2012 Consistent but modest: a meta-analysis on unimodal and multimodal affect detection accuracies from 30 studies
abstract
The recent influx of multimodal affect classifiers raises the important question of whether these classifiers yield accuracy rates that exceed their unimodal counterparts. This question was addressed by performing a meta-analysis on 30 published studies that reported both multimodal and unimodal affect detection accuracies. The results indicated that multimodal accuracies were consistently better than unimodal accuracies and yielded an average 8.12% improvement over the best unimodal classifiers. However, performance improvements were three times lower when classifiers were trained on natural or seminatural data (4.39% improvement) compared to acted data (12.1% improvement). Importantly, performance of the best unimodal classifier explained an impressive 80.6% (cross-validated) of the variance in multimodal accuracy. The results also indicated that multimodal accuracies were substantially higher than accuracies of the second-best unimodal classifiers (an average improvement of 29.4%) irrespective of the naturalness of the training data. Theoretical and applied implications of the findings are discussed.
Sidney K. D'Mello, Jacqueline Kory Westlund
ICMI1
2012 How Do They Do It? Investigating Dialogue Moves within Dialogue Modes in Expert Human Tutoring
Blair Lehman, Sidney K. D'Mello, Whitney L. Cade, Natalie K. Person
ITS2
2012 Interventions to Regulate Confusion during Learning
Blair Lehman, Sidney K. D'Mello, Arthur C. Graesser
ITS2
2012 Automatic Evaluation of Learner Self-Explanations and Erroneous Responses for Dialogue-Based ITSs
Blair Lehman, Caitlin Mills 0001, Sidney K. D'Mello, Arthur C. Graesser
ITS3
2012 Emotions during Writing on Topics That Align or Misalign with Personal Beliefs
Caitlin Mills 0001, Sidney K. D'Mello
ITS2
2012 Guru: A Computer Tutor That Models Expert Human Tutors
Andrew Olney, Sidney K. D'Mello, Natalie K. Person, Whitney L. Cade, Patrick Hays, Claire Williams, Blair Lehman, Arthur C. Graesser
ITS2
2012 Exploring Relationships between Learners' Affective States, Metacognitive Processes, and Learning Outcomes
Amber Chauncey Strain, Roger Azevedo, Sidney K. D'Mello
ITS3
2012 How Do Learners Regulate Their Emotions?
Amber Chauncey Strain, Sidney K. D'Mello, Melissa R. Gross
ITS2
2012 Gaze tutor: A gaze-reactive intelligent tutoring system
Sidney K. D'Mello, Andrew Olney, Claire Williams, Patrick Hays
Int. J. Hum. Comput. Stud.1
2012 Detecting Naturalistic Expressions of Nonbasic Affect Using Physiological Signals
abstract
Signals from peripheral physiology (e.g., ECG, EMG, and GSR) in conjunction with machine learning techniques can be used for the automatic detection of affective states. The affect detector can be user-independent, where it is expected to generalize to novel users, or user-dependent, where it is tailored to a specific user. Previous studies have reported some success in detecting affect from physiological signals, but much of the work has focused on induced affect or acted expressions instead of contextually constrained spontaneous expressions of affect. This study addresses these issues by developing and evaluating user-independent and user-dependent physiology-based detectors of nonbasic affective states (e.g., boredom, confusion, curiosity) that were trained and validated on naturalistic data collected during interactions between 27 students and AutoTutor, an intelligent tutoring system with conversational dialogues. There is also no consensus on which techniques (i.e., feature selection or classification methods) work best for this type of data. Therefore, this study also evaluates the efficacy of affect detection using a host of feature selection and classification techniques on three physiological signals (ECG, EMG, and GSR) and their combinations. Two feature selection methods and nine classifiers were applied to the problem of recognizing eight affective states (boredom, confusion, curiosity, delight, flow/-engagement, surprise, and neutral). The results indicated that the user-independent modeling approach was not feasible; however, a mean kappa score of 0.25 was obtained for user-dependent models that discriminated among the most frequent emotions. The results also indicated that k-nearest neighbor and Linear Bayes Normal Classifier (LBNC) classifiers yielded the best affect detection rates. Single channel ECG, EMG, and GSR and three-channel multimodal models were generally more diagnostic than two--channel models.
Omar AlZoubi, Sidney K. D'Mello, Rafael A. Calvo
IEEE Trans. Affect. Comput.2
2012 AutoTutor and affective autotutor: Learning by talking with cognitively and emotionally intelligent computers that talk back
abstract
We present AutoTutor and Affective AutoTutor as examples of innovative 21 st century interactive intelligent systems that promote learning and engagement. AutoTutor is an intelligent tutoring system that helps students compose explanations of difficult concepts in Newtonian physics and enhances computer literacy and critical thinking by interacting with them in natural language with adaptive dialog moves similar to those of human tutors. AutoTutor constructs a cognitive model of students' knowledge levels by analyzing the text of their typed or spoken responses to its questions. The model is used to dynamically tailor the interaction toward individual students' zones of proximal development. Affective AutoTutor takes the individualized instruction and human-like interactivity to a new level by automatically detecting and responding to students' emotional states in addition to their cognitive states. Over 20 controlled experiments comparing AutoTutor with ecological and experimental controls such reading a textbook have consistently yielded learning improvements of approximately one letter grade after brief 30--60-minute interactions. Furthermore, Affective AutoTutor shows even more dramatic improvements in learning than the original AutoTutor system, particularly for struggling students with low domain knowledge. In addition to providing a detailed description of the implementation and evaluation of AutoTutor and Affective AutoTutor, we also discuss new and exciting technologies motivated by AutoTutor such as AutoTutor-Lite, Operation ARIES, GuruTutor, DeepTutor, MetaTutor, and AutoMentor. We conclude this article with our vision for future work on interactive and engaging intelligent tutoring systems.
Sidney K. D'Mello, Arthur C. Graesser
ACM Trans. Interact. Intell. Syst.1
2011 Affective Modeling from Multichannel Physiology: Analysis of Day Differences
Omar AlZoubi, M. Sazzad Hussain, Sidney K. D'Mello, Rafael A. Calvo
ACII (1)3
2011 Does Topic Matter? Topic Influences on Linguistic and Rubric-Based Evaluation of Writing
Nia Nixon, Sidney K. D'Mello, Caitlin Mills 0001, Arthur C. Graesser
AIED2
2011 Affect Detection from Multichannel Physiology during Learning Sessions with AutoTutor
M. Sazzad Hussain, Omar AlZoubi, Rafael A. Calvo, Sidney K. D'Mello
AIED4
2011 Inducing and Tracking Confusion with Contradictions during Critical Thinking and Scientific Reasoning
Blair Lehman, Sidney K. D'Mello, Amber Chauncey Strain, Melissa R. Gross, Allyson Dobbins, Patricia S. Wallace, Keith K. Millis, Arthur C. Graesser
AIED2
2011 Emotion Regulation during Learning
Amber Chauncey Strain, Sidney K. D'Mello
AIED2
2011 Training Emotion Regulation Strategies During Computerized Learning: A Method for Improving Learner Self-Regulation
Amber Chauncey Strain, Sidney K. D'Mello, Arthur C. Graesser
AIED2
2011 Dynamical Emotions: Bodily Dynamics of Affect during Problem Solving
Sidney K. D'Mello
CogSci1
2011 Strategy Shifting in a Procedural-Motor Drawing Task
Brent Morgan, Sidney K. D'Mello, Jenna Fielding, Karl Fike, Andrea Tamplin, Gabriel Radvansky, James Arnett, Robert G. Abbott, Arthur C. Graesser
CogSci2
2010 Mining Bodily Patterns of Affective Experience during Learning
Sidney K. D'Mello, Arthur C. Graesser
EDM1
2010 Collaborative Lecturing by Human and Computer Tutors
Sidney K. D'Mello, Patrick Hays, Claire Williams, Whitney L. Cade, Andrew Olney
Intelligent Tutoring Systems (2)1
2010 A Time for Emoting: When Affect-Sensitivity Is and Isn't Effective at Promoting Deep Learning
Sidney K. D'Mello, Blair Lehman, Jeremiah Sullins, Rosaire Daigle, Rebekah Combs, Kimberly Vogt, Lydia Perkins, Arthur C. Graesser
Intelligent Tutoring Systems (1)1
2010 The Intricate Dance between Cognition and Emotion during Expert Tutoring
Blair Lehman, Sidney K. D'Mello, Natalie K. Person
Intelligent Tutoring Systems (2)2
2010 A DIY Pressure Sensitive Chair for Intelligent Tutoring Systems
Andrew Olney, Sidney K. D'Mello
Intelligent Tutoring Systems (2)2
2010 The Impact of System Feedback on Learners' Affective and Physiological States
Payam Aghaei Pour, M. Sazzad Hussain, Omar AlZoubi, Sidney K. D'Mello, Rafael A. Calvo
Intelligent Tutoring Systems (1)4
2010 Predicting Student Knowledge Level from Domain-Independent Function and Content Words
Claire Williams, Sidney K. D'Mello
Intelligent Tutoring Systems (2)2
2010 Toward Spoken Human-Computer Tutorial Dialogues
abstract
Oral discourse is the primary form of human–human communication, hence, computer interfaces that communicate via unstructured spoken dialogues will presumably provide a more efficient, meaningful, and naturalistic interaction experience. Within the context of learning environments, there are theoretical positions supporting a speech facilitation hypothesis that predicts that spoken tutorial dialogues will increase learning more than typed dialogues. We evaluated this hypothesis in an experiment where 24 participants learned computer literacy via a spoken and a typed conversation with AutoTutor, an intelligent tutoring system with conversational dialogues. The results indicated that (a) enhanced content coverage was achieved in the spoken condition; (b) learning gains for both modalities were on par and greater than a no-instruction control; (c) although speech recognition errors were unrelated to learning gains, they were linked to participants' evaluations of the tutor; (d) participants adjusted their conversational styles when speaking compared to typing; (e) semantic and statistical natural language understanding approaches to comprehending learners' responses were more resilient to speech recognition errors than syntactic and symbolic-based approaches; and (f) simulated speech recognition errors had differential impacts on the fidelity of different semantic algorithms. We discuss the impact of our findings on the speech facilitation hypothesis and on human–computer interfaces that support spoken dialogues.
Sidney K. D'Mello, Arthur C. Graesser, Brandon G. King
Hum. Comput. Interact.1
2010 Better to be frustrated than bored: The incidence, persistence, and impact of learners' cognitive-affective states during interactions with three different computer-based learning environments
Ryan Baker 0001, Sidney K. D'Mello, Ma. Mercedes T. Rodrigo, Arthur C. Graesser
Int. J. Hum. Comput. Stud.2
2010 Affect Detection: An Interdisciplinary Review of Models, Methods, and Their Applications
abstract
This survey describes recent progress in the field of Affective Computing (AC), with a focus on affect detection. Although many AC researchers have traditionally attempted to remain agnostic to the different emotion theories proposed by psychologists, the affective technologies being developed are rife with theoretical assumptions that impact their effectiveness. Hence, an informed and integrated examination of emotion theories from multiple areas will need to become part of computing practice if truly effective real-world systems are to be achieved. This survey discusses theoretical perspectives that view emotions as expressions, embodiments, outcomes of cognitive appraisal, social constructs, products of neural circuitry, and psychological interpretations of basic feelings. It provides meta-analyses on existing reviews of affect detection systems that focus on traditional affect detection modalities like physiology, face, and voice, and also reviews emerging research on more novel channels such as text, body language, and complex multimodal systems. This survey explicitly explores the multidisciplinary foundation that underlies all AC applications by describing how AC researchers have incorporated psychological theories of emotion and how these theories affect research questions, methods, results, and their interpretations. In this way, models and methods can be compared, and emerging insights from various disciplines can be more expertly integrated.
Rafael A. Calvo, Sidney K. D'Mello
IEEE Trans. Affect. Comput.2
2010 Multimodal semi-automated affect detection from conversational cues, gross body language, and facial features
Sidney K. D'Mello, Arthur C. Graesser
User Model. User Adapt. Interact.1
2009 Cohesion Relationships in Tutorial Dialogue as Predictors of Affective States
abstract
We explored the possibility of predicting learners' affective states (boredom, flow/engagement, confusion, and frustration) by monitoring variations in the cohesiveness of tutorial dialogues during interactions with AutoTutor, an intelligent tutoring system with conversational dialogues. Multiple measures of cohesion (e.g., pronouns, connectives, semantic overlap, causal cohesion, coreference) were automatically computed using the Coh-Metrix facility for analyzing discourse and language characteristics of text. Cohesion measures in multiple regression models predicted the proportional occurrence of each affective state, yielding medium to large effect sizes. The incidence of negations, pronoun referential cohesion, causal cohesion, and co-reference cohesion were the most diagnostic predictors of the affective states. We discuss the generalizability of our findings to other domains and tutoring systems, as well as the possibility of constructing real-time, cohesion-based affect detectors.
Sidney K. D'Mello, Nia Nixon, Arthur C. Graesser
AIED1
2009 Antecedent-Consequent Relationships and Cyclical Patterns between Affective States and Problem Solving Outcomes
abstract
We explored the complex interplay between students' affective states and problem solving outcomes. We conducted a study where 41 students solved 28 analytical reasoning problems from the Law School Admission Test. Participants viewed videos of their interaction history and judged their emotions at theoretically relevant points in the problem solving session (after new problem is displayed, in the midst of problem solving, after feedback is received). We explore excitatory and inhibitory relationships between the affective states and problem solving outcomes (i.e. success or failure, and associated positive or negative feedback). We isolate affective states that are consequences of outcomes and associated feedback as well as affective states that are antecedents of positive or negative outcomes. Follow-up analyses focused on cyclical patterns that incorporate complex relationships between the affective states and problem solving outcomes. Implications of our results for affect-sensitive artificial learning environments are discussed.
Sidney K. D'Mello, Natalie K. Person, Blair Lehman
AIED1
2009 The Relationship Between Modality and Metacognition While Interacting with AutoTutor
abstract
In this paper we explored the relationship between metacognitive statements and learning gains with students' typed and spoken interactions with an intelligent tutoring system, called AutoTutor. Analyses revealed that students who entered their contributions via speech showed a significantly higher proportion of metacognitive statements (e.g., I'm not following, I understand). There was a significant negative correlation between metacognitive statements and posttest scores on both typed and spoken interactions. Students with low prior knowledge expressed more metacognitive statements than did students with high prior knowledge. Therefore, metacognitive expressions reflect the learners' knowledge deficits as opposed to improved knowledge monitoring from greater subject matter knowledge.
Jeremiah Sullins, Moongee Jeon, Sidney K. D'Mello, Arthur C. Graesser
AIED3
2008 Dialogue Modes in Expert Tutoring
Whitney L. Cade, Jessica L. Copeland, Natalie K. Person, Sidney K. D'Mello
Intelligent Tutoring Systems4
2008 Self Versus Teacher Judgments of Learner Emotions During a Tutoring Session with AutoTutor
Sidney K. D'Mello, Roger Taylor, Kelly Davidson, Arthur C. Graesser
Intelligent Tutoring Systems1
2008 What Are You Feeling? Investigating Student Affective States During Expert Human Tutoring Sessions
Blair Lehman, Melanie Matthews, Sidney K. D'Mello, Natalie K. Person
Intelligent Tutoring Systems3
2008 Comparing Learners' Affect While Using an Intelligent Tutoring System and a Simulation Problem Solving Game
Ma. Mercedes T. Rodrigo, Ryan Baker 0001, Sidney K. D'Mello, Ma. Celeste T. Gonzalez, Maria Carminda V. Lagud, Sheryl Ann L. Lim, Alexis F. Macapanpan, Sheila A. M. S. Pascua, Jerry Q. Santillano, Jessica O. Sugay, Sinath Tep, Norma J. B. Viehland
Intelligent Tutoring Systems3
2008 The Dynamics of Self-regulatory Processes within Self-and Externally Regulated Learning Episodes During Complex Science Learning with Hypermedia
Amy M. Witherspoon, Roger Azevedo, Sidney K. D'Mello
Intelligent Tutoring Systems3
2008 Automatic detection of learner's affect from conversational cues
Sidney K. D'Mello, Scotty D. Craig, Amy M. Witherspoon, Bethany McDaniel, Arthur C. Graesser
User Model. User Adapt. Interact.1
2007 Modeling and Scaffolding Affective Experiences to Impact Learning
Sidney K. D'Mello, Scotty D. Craig, Rana El Kaliouby, Madeline Alsmeyer, Genaro Rebolledo-Mendez
AIED1
2007 Mind and Body: Dialogue and Posture for Affect Detection in Learning Environments
Sidney K. D'Mello, Arthur C. Graesser
AIED1
2007 Emotions and Learning with AutoTutor
Arthur C. Graesser, Patrick Chipman, Brandon G. King, Bethany McDaniel, Sidney K. D'Mello
AIED5
2006 Affect Detection from Human-Computer Dialogue with an Intelligent Tutoring System
Sidney K. D'Mello, Arthur C. Graesser
IVA1
2006 MIKI: A Speech Enabled Intelligent Kiosk
Lee McCauley, Sidney K. D'Mello
IVA2