Jacob Whitehill

dblp:17/2833 · DBLP profile ↗
← Back
49ranked-venue papers
15as first author
15since 2021 · last 2026
0000-0002-5851-312XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 11 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 6 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Survey of end-to-end multi-speaker automatic speech recognition for monaural audio
abstract
Monaural multi-speaker automatic speech recognition (ASR) remains challenging due to data scarcity and the intrinsic difficulty of recognizing and attributing words to individual speakers, particularly in overlapping speech. Recent advances have driven the shift from cascade systems to end-to-end (E2E) architectures, which reduce error propagation and better exploit the synergy between speech content and speaker identity. Despite rapid progress in E2E multi-speaker ASR, the field lacks a comprehensive review of recent developments. This survey provides a systematic taxonomy of E2E neural approaches for multi-speaker ASR, highlighting recent advances and comparative analysis. Specifically, we analyze: (1) architectural paradigms (single-input-multiple-output (SIMO) vs. single-input-single-output (SISO)) for pre-segmented audio, analyzing their distinct characteristics and trade-offs; (2) recent architectural and algorithmic improvements based on these two paradigms, including multi-modal inputs; (3) extensions to long-form speech, including segmentation strategy and speaker-consistent hypothesis stitching. Further, we (4) evaluate and compare methods across standard benchmarks. We conclude with a discussion of open challenges and future research directions towards building robust and scalable multi-speaker ASR.
Xinlu He, Jacob Whitehill
Comput. Speech Lang.2
2025 Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?
abstract
Decoder-only discrete-token language models have recently achieved significant success in automatic speech recognition. However, systematic analyses of how different modalities impact performance in specific scenarios remain limited. In this paper, we investigate the effects of multiple modalities on speech recognition accuracy on both synthetic and real-world datasets. Our experiments suggest that: (1) Integrating more modalities can increase accuracy but the benefit depends on the amount of auditory noise. We also show for the first time the benefit of combining audio, image context, and lip information in one speech recognition model. (2) Images as a supplementary modality for speech recognition provide their greatest benefit at moderate audio noise levels; moreover, they exhibit a different trend compared to inherently synchronized modalities like lip movements. (3) Performance improves on both synthetic and real-world datasets when the most relevant visual information is filtered as a preprocessing step.
Yiwen Guan, Viet Anh Trinh, Vivek Voleti, Jacob Whitehill
ICME4
2024 Tracking Classroom Movement Patterns with Person Re-ID
Xinlu He, Viet Anh Trinh, Andrew A. McReynolds, Jacob Whitehill
EDM5
2024 Speaker Diarization in the Classroom: How Much Does Each Student Speak in Group Discussions?
Shiran Dudy, Xinlu He, Rosy Southwell, Jacob Whitehill
EDM6
2024 Automatic Speech Recognition Tuned for Child Speech in the Classroom
abstract
K-12 school classrooms have proven to be a challenging environment for Automatic Speech Recognition (ASR) systems, both due to background noise and conversation, and differences in linguistic and acoustic properties from adult speech, on which the majority of ASR systems are trained and evaluated. We report on experiments to improve ASR for child speech in the classroom by training and fine-tuning transformer models on public corpora of adult and child speech augmented with classroom background noise. By tuning OpenAI’s Whisper model we achieve a 38% relative reduction in word error rate (WER) to 9.2% on the public MyST dataset of child speech – the lowest yet reported – and a 7% relative reduction to reach 54% WER on a more challenging classroom speech dataset (ISAT). We also introduce a novel beam hypothesis rescoring method that incorporates a speed-aware term to capture prior knowledge of human speaking rates, as well as a Large Language Model, to select among hypotheses. We demonstrate the effectiveness of this technique on both publicly-available datasets and a classroom speech dataset.
Rosy Southwell, Wayne H. Ward, Viet Anh Trinh, Charis Clevenger, Clay Clevenger, Emily Watts, Jason G. Reitman, Sidney K. D'Mello, Jacob Whitehill
ICASSP9
2023 In Search of Negative Moments: Multi-Modal Analysis of Teacher Negativity in Classroom Observation Videos
Zilin Dai, Andrew A. McReynolds, Jacob Whitehill
EDM3
2023 Compositional clustering: Applications to multi-label object recognition and speaker identification
Xinlu He, Jacob Whitehill
Pattern Recognit.3
2023 Toward Automated Classroom Observation: Multimodal Machine Learning to Estimate CLASS Positive Climate and Negative Climate
abstract
In this article we present a multi-modal machine learning-based system, which we call ACORN, to analyze videos of school classrooms for the Positive Climate (PC) and Negative Climate (NC) dimensions of the CLASS [1] observation protocol that is widely used in educational research. ACORN uses convolutional neural networks to analyze spectral audio features, the faces of teachers and students, and the pixels of each image frame, and then integrates this information over time using Temporal Convolutional Networks. The audiovisual ACORN's PC and NC predictions have Pearson correlations of 0.55 and 0.63 with ground-truth scores provided by expert CLASS coders on the UVA Toddler dataset (cross-validation on$n=300$15-min video segments), and a purely auditory ACORN predicts PC and NC with correlations of 0.36 and 0.41 on the MET dataset (test set of$n=2000$videos segments). These numbers are similar to inter-coder reliability of human coders. Finally, using Graph Convolutional Networks we make early strides (AUC=0.70) toward predicting the specific moments (45-90sec clips) when the PC is particularly weak/strong. Our findings inform the design of automatic classroom observation and also more general video activity recognition and summary recognition systems.
Anand Ramakrishnan, Brian Zylich, Erin Ottmar, Jennifer LoCasale-Crouch, Jacob Whitehill
IEEE Trans. Affect. Comput.5
2022 How to Give Imperfect Automated Guidance to Learners: A Case-Study in Workplace Learning
Jacob Whitehill, Amitai Erfanian
AIED (1)1
2021 Affective Teacher Tools: Affective Class Report Card and Dashboard
Ankit Gupta 0016, Neeraj Menon, William Lee 0002, William Rebelsky, Danielle Allessio, Tom Murray 0001, Beverly P. Woolf, Jacob Whitehill, Ivon Arroyo
AIED (1)8
2021 Text Representations of Math Tutorial Videos forClustering, Retrieval, and Learning Gain Prediction
Pichayut Liamthong, Jacob Whitehill
EDM2
2021 Leveraging Affect Transfer Learning for Behavior Prediction in an Intelligent Tutoring System
abstract
In this work, we propose a video-based transfer learning approach for predicting problem outcomes of students working with an intelligent tutoring system (ITS). By analyzing a student's face and gestures, our method predicts the outcome of a student answering a problem in an ITS from a video feed. Our work is motivated by the reasoning that the ability to predict such outcomes enables tutoring systems to adjust interventions, such as hints and encouragement, and to ultimately yield improved student learning. We collected a large labeled dataset of student interactions with an intelligent online math tutor consisting of 68 sessions, where 54 individual students solved 2,749 problems. We will release this dataset publicly upon publication of this paper. It will be available at https://www.cs.bu.edu/faculty/betke/research/learning/. Working with this dataset, our transfer-learning challenge was to design a representation in the source domain of pictures obtained “in the wild” for the task of facial expression analysis, and transferring this learned representation to the task of human behavior prediction in the domain of webcam videos of students in a classroom environment. We developed a novel facial affect representation and a user-personalized training scheme that unlocks the potential of this representation. We designed several variants of a recurrent neural network that models the temporal structure of video sequences of students solving math problems. Our final model, named ATL-BP for Affect Transfer Learning for Behavior Prediction, achieves a relative increase in mean F -score of 50 % over the state-of-the-art method on this new dataset.
Nataniel Ruiz, Hao Yu 0014, Danielle Allessio, Mona Jalal, Ajjen Joshi, Tom Murray 0001, John J. Magee, Jacob Whitehill, Vitaly Ablavsky, Ivon Arroyo, Beverly P. Woolf, Stan Sclaroff, Margrit Betke
FG8
2021 Compositional Embedding Models for Speaker Identification and Diarization with Simultaneous Speech From 2+ Speakers
abstract
We propose a new method for speaker diarization that can handle overlapping speech with 2+ people. Our method is based on compositional embeddings [1]: Like standard speaker embedding methods such as x-vector [2], compositional embedding models contain a function f that separates speech from different speakers. In addition, they include a composition function g to compute set-union operations in the embedding space so as to infer the set of speakers within the input audio. In an experiment on multi-person speaker identification using synthesized LibriSpeech data, the proposed method outperforms traditional embedding methods that are only trained to separate single speakers (not speaker sets). In a speaker diarization experiment on the AMI Headset Mix corpus, we achieve state-of-the-art accuracy (DER=22.93%), slightly better than the previous best result (23.82% from [3]).
Jacob Whitehill
ICASSP2
2021 Learning to Work in a Materials Recovery Facility: Can Humans and Machines Learn from Each Other?
abstract
Workplace learning often requires workers to learn new perceptual and motor skills. The future of work will increasingly feature human users who cooperate with machines, both to learn the tasks and to perform them. In this paper, we examine workplace learning in Materials Recovery Facilities (MRFs), i.e., recycling plants, where workers separate waste items on conveyer belts before they are formed into bales and reprocessed. Using a simulated MRF, we explored the benefit of machine learning assistants (MLAs) that help workers, and help train them, to sort objects efficiently by providing automated perceptual guidance. In a randomized experiment (n = 140), we found: (1) A low-accuracy MLA is worse than no MLA at all, both in terms of task performance and learning. (2) A perfect MLA led to the best task performance, but was no better in helping users to learn than having no MLA at all. (3) Users tend to follow the MLA’s judgments too often, even when they were incorrect. Finally, (4) we devised a novel learning analytics algorithm to assess the worker’s accuracy, with the goal of obtaining additional training labels that can be used for fine-tuning the machine. A simulation study illustrates how even noisy labels can increase the machine’s accuracy.
Harrison Kyriacou, Anand Ramakrishnan, Jacob Whitehill
LAK3
2021 Compositional Embeddings for Multi-Label One-Shot Learning
abstract
We present a compositional embedding framework that infers not just a single class per input image, but a set of classes, in the setting of one-shot learning. Specifically, we propose and evaluate several novel models consisting of (1) an embedding function f trained jointly with a "composition" function g that computes set union operations between the classes encoded in two embedding vectors; and (2) embedding f trained jointly with a "query" function h that computes whether the classes encoded in one embedding subsume the classes encoded in another embedding. In contrast to prior work, these models must both perceive the classes associated with the input examples and encode the relationships between different class label sets, and they are trained using only weak one-shot supervision consisting of the label-set relationships among training examples. Experiments on the OmniGlot, Open Images, and COCO datasets show that the proposed compositional embedding models outperform existing embedding methods. Our compositional embedding models have applications to multi-label object recognition for both one-shot and supervised learning.
Michael C. Mozer, Jacob Whitehill
WACV3
2020 Toward Better Speaker Embeddings: Automated Collection of Speech Samples From Unknown Distinct Speakers
abstract
The accuracy of speaker verification and diarization models depends on the quality of the speaker embeddings used to separate audio samples from different speakers. With the goal of training better embedding models, we devise an automatic pipeline for large-scale collection of speech samples from unique speakers that is significantly more automated than previous approaches. With this pipeline, we collect and publish the BookTubeSpeech dataset, containing 8,450 YouTube videos (7.74 min per video on average) that each contains a single unique speaker. Using this dataset combined with VoxCeleb2, we show a substantial improvement in the quality of embeddings when tested on LibriSpeech compared to a model trained on only VoxCeleb2.
Minh Pham 0005, Jacob Whitehill
ICASSP3
2020 Noise-Robust Key-Phrase Detectors for Automated Classroom Feedback
abstract
With the goal of giving teachers automated feedback about their classrooms, we investigate how to train automatic speech detectors of key phrases such as good job, thank you, please, and you’re welcome. This kind of language conveys support and respect from teacher to student and is one of the behavioral markers used in the established CLASS [1] classroom observation protocol. School classrooms are noisy and contain overlapping speech, presenting a highly challenging environment for automatic speech recognition (ASR), even for state-of-the-art approaches. We train deep neural networks using hierarchical multitask learning (MTL) on a modest-sized but highly-tailored dataset of classroom speech. Compared to 2 state-of-the-art ASR systems for general-purpose speech recognition (Google [2] and Deep-Speech [3]), our system delivers a substantially improved recall rate (50.4% versus 20.5%) while matching their precision (30%). Moreover, our system’s predictions correlate with several dimensions of the CLASS.
Brian Zylich, Jacob Whitehill
ICASSP2
2020 How Does Label Noise Affect the Quality of Speaker Embeddings?
Minh Pham 0005, Jacob Whitehill
INTERSPEECH3
2019 How Does Knowledge of the AUC Constrain the Set of Possible Ground-Truth Labelings?
abstract
Recent work on privacy-preserving machine learning has considered how datamining competitions such as Kaggle could potentially be “hacked”, either intentionally or inadvertently, by using information from an oracle that reports a classifier’s accuracy on the test set (Blum and Hardt 2015; Hardt and Ullman 2014; Zheng 2015; Whitehill 2016). For binary classification tasks in particular, one of the most common accuracy metrics is the Area Under the ROC Curve (AUC), and in this paper we explore the mathematical structure of how the AUC is computed from an n-vector of real-valued “guesses” with respect to the ground-truth labels. Under the assumption of perfect knowledge of the test set AUC c=p/q, we show how knowing c constrains the set W of possible ground-truth labelings, and we derive an algorithm both to compute the exact number of such labelings and to enumerate efficiently over them. We also provide empirical evidence that, surprisingly, the number of compatible labelings can actually decrease as n grows, until a test set-dependent threshold is reached. Finally, we show how W can be efficiently whittled down, through pairs of oracle queries, to infer all the groundtruth test labels with complete certainty.
Jacob Whitehill
AAAI1
2019 How Should Online Teachers of English as a Foreign Language (EFL) Write Feedback to Students?
Cecilia Aguerrebere, Monica Bulger, Cristobal Cobo, Sofía García, Gabriela Kaplan, Jacob Whitehill
EDM6
2019 Measuring students' thermal comfort and its impact on learning
Han Jiang 0009, Matthew Iandoli, Steven Van Dessel, Shichao Liu 0006, Jacob Whitehill
EDM5
2019 Do Learners Know What's Good for Them? Crowdsourcing Subjective Ratings of OERs to Predict Learning Gains
Jacob Whitehill, Cecilia Aguerrebere, Benjamin Hylak
EDM1
2019 Affect-driven Learning Outcomes Prediction in Intelligent Tutoring Systems
abstract
Equipping an Intelligent Tutoring System (ITS) with the ability to interpret affective signals from students could potentially improve the learning experience of students by enabling the tutor to monitor the students' progress and provide timely interventions as well as present appropriate affective reactions via a virtual tutor. Most ITSs equipped with affect modeling capabilities attempt to predict the emotional state of users. However, the focus in this work is instead on trying to directly predict the learning outcomes of students from a stream of video capturing the students faces as they work on a set of math problems. Using facial features extracted from a video stream, we train classifiers to directly predict the success or failure of a student's attempt to answer a question while the student has just begun to work on the problem. In this work, we first introduce a novel dataset of student interactions with MathSpring, a popular ITS. We provide an exploratory analysis of the different problem outcome classes using typical facial action unit activations. We develop baseline models to predict the problem outcome labels of students solving math problems and discuss how early problem outcome labels can be forecasted and utilized to provide possible interventions.
Ajjen Joshi, Danielle Allessio, John J. Magee, Jacob Whitehill, Ivon Arroyo, Beverly P. Woolf, Stan Sclaroff, Margrit Betke
FG4
2019 Toward Automated Classroom Observation: Predicting Positive and Negative Climate
abstract
We devised and evaluated a multi-modal machine learning-based system to analyze videos of school classrooms for "positive climate" and "negative climate", which are two dimensions of the Classroom Assessment Scoring System (CLASS) [1]. School classrooms are highly cluttered audiovisual scenes containing many overlapping faces and voices. Due to the difficulty of labeling them (reliable coding requires weeks of training) and their sensitive nature (students and teachers may be in stressful or potentially embarrassing situations), CLASS- labeled classroom video datasets are scarce, and their labels are sparse (just a few labels per 15-minute video dip). Thus, the overarching challenge was how to harness modern deep perceptual architectures despite the paucity of labeled data. Through training low-level CNN-based facial attribute detectors (facial expression & adult/child) as well as a direct audio-to- climate regressor, and by integrating low-level information over time using a Bi-LSTM, we constructed automated detectors of positive and negative classroom climate with accuracy (10- fold cross-validation Pearson correlation on 241 CLASS-labeled videos) of 0.40 and 0.51, respectively. These numbers are superior to what we obtained using shallower architectures. This work represents the first automated system designed to detect specific dimensions of the CLASS.
Anand Ramakrishnan, Erin Ottmar, Jennifer LoCasale-Crouch, Jacob Whitehill
FG4
2019 Automatic Classifiers as Scientific Instruments: One Step Further Away from Ground-Truth
abstract
Automatic machine learning-based detectors of various psychological and social phenomena (e.g., emotion, stress, engagement) have great potential to advance basic science. However, when a detector d is trained to approximate an existing measurement tool (e.g., a questionnaire, observation protocol), then care must be taken when interpreting measurements collected using d since they are one step further removed from the under- lying construct. We examine how the accuracy of d, as quantified by the correlation q of d’s out- puts with the ground-truth construct U, impacts the estimated correlation between U (e.g., stress) and some other phenomenon V (e.g., academic performance). In particular: (1) We show that if the true correlation between U and V is r, then the expected sample correlation, over all vectors T n whose correlation with U is q, is qr. (2) We derive a formula for the probability that the sample correlation (over n subjects) using d is positive given that the true correlation is negative (and vice-versa); this probability can be substantial (around 20 - 30%) for values of n and q that have been used in recent affective computing studies. (3) With the goal to reduce the variance of correlations estimated by an automatic detector, we show that training multiple neural networks d(1) , . . . , d(m) using different training architectures and hyperparameters for the same detection task provides only limited “coverage” of T^n.
Jacob Whitehill, Anand Ramakrishnan
ICML1
2019 Deep Learning with Domain Randomization for Optimal Filtering
abstract
Filtering is the process of recovering a signal, x(t), from noisy measurements z(t). One common filter is the Kalman Filter, which is proven to be the conditional minimum variance estimator of x(t) when the measurements are a Gaussian random processes. However, in practice (a) the measurements are not necessarily Gaussian and (b) an estimation of the measurement covariance is problematic and often tuned using cross-validation and domain knowledge. In order to address the Kalman Filter's suboptimal performance in situations where non-Gaussian noise is present, we train a deep autoencoder to learn a mapping of the noisy measurements to a "learned" noise process covariance R(t) and measurements z(t), which are passed to a Kalman Filter. The Kalman Filter's output is then mapped to the original measurement space via the decoder portion of the autoencoder, where this output and original measurements are compared. Training this autoencoder-Kalman Filter (AEKF) via domain randomization on simulated noisy sensor responses, we show the AEKF's estimate of x(t), in the majority of cases, achieves a lower MSE than both a standard Kalman Filter and a Long Short-Term Memory (LSTM) recurrent deep neural network. Most significantly, the AEKF outperforms all methods in the case of simulated Cauchy noise.
Matthew L. Weiss, Randy C. Paffenroth, Jacob Whitehill, Joshua R. Uzarski
ICMLA3
2018 Estimating the Treatment Effect of New Device Deployment on Uruguayan Students' Online Learning Activity
Cecilia Aguerrebere, Cristobal Cobo, Jacob Whitehill
EDM3
2018 Who are they looking at? Automatic Eye Gaze Following for Classroom Observation Video Analysis
Arkar Min Aung, Anand Ramakrishnan, Jacob Whitehill
EDM3
2018 Harnessing Label Uncertainty to Improve Modeling: An Application to Student Engagement Recognition
abstract
Automatic facial expression recognition systems are usually trained from target labels that model each example as belonging unambiguously to a single class (e.g., "non-engaged", "very engaged", etc.). However, in some settings, ground-truth labels can be more aptly modeled as probability distributions (e.g., [0.1, 0.1, 0.5, 0.3] over 4 engagement categories) that capture the uncertainty that can arise during the annotation process. In this paper, we explore how harnessing the full probability distribution of each label ("soft labels"), rather than just a scalar summary statistic ("hard labels", e.g., majority class or mean), can yield better recognition accuracy when training automated detectors. Our results on a face image dataset (10698 faces over 20 subjects) labeled for perceived student engagement suggest that training on soft labels can deliver engagement detectors that fit the data stat. sig. more accurately (lower cross-entropy for classification, higher Pearson correlation for regression) than when training on hard labels. Moreover, we explore possible reasons for this effect and provide evidence that it is due to implicit regularization that the soft labels enact on the trained engagement detector. This effect is similar to, but empirically seems stronger than, the "label smoothing" approach proposed by Szegedy, et al. [1].
Arkar Min Aung, Jacob Whitehill
FG2
2018 Predicting When Teachers Look at Their Students in 1-on-1 Tutoring Sessions
abstract
We propose and evaluate a neural network archi- tecture for predicting when human teachers shift their eye-gaze to look at their students during 1-on-1 math tutoring sessions. Such models may be useful when developing affect-sensitive intelligent tutoring systems (ITS) because they can function as an attention model that informs the ITS when the student's face, body posture, and other visual cues are most important to observe. Our approach combines both feed-forward (FF) and recurrent (LSTM) components for predicting gaze shifts based on the history of tutoring actions (e.g., request assistance from the teacher, pose a new problem to the student, give a hint, etc.), as well as the teacher's prior gaze events. Despite the challenging nature of the task - we are asking the network to predict whether or not the teacher will shift her/his eye gaze during the next one- second time interval - the network achieves an AUC (averaged over 2 teachers) of 0.75. In addition, we identify some of the factors that the human teachers in our study used when making gaze decisions and show evidence that the two teachers' gaze patterns share common characteristics.
Han Jiang 0009, Karmen Dykstra, Jacob Whitehill
FG3
2018 Permutation-Invariant Consensus over Crowdsourced Labels
abstract
This paper introduces a novel crowdsourcing consensus model and inference algorithm — which we call PICA (Permutation-Invariant Crowdsourcing Aggregation) — that is designed to recover the ground-truth labels of a dataset while being invariant to the class permutations enacted by the different annotators. This is particularly useful for settings in which annotators may have systematic confusions about the meanings of different classes, as well as clustering problems (e.g., dense pixel-wise image segmentation) in which the names/numbers assigned to each cluster have no inherent meaning.The PICA model is constructed by endowing each annotator with a doubly-stochastic matrix (DSM), which models the probabilities that an annotator will perceive one class and transcribe it into another. We conduct simulations and experiments to show the advantage of PICA compared to two baselines (Majority Vote, and an "unpermutation" heuristic) for three different clustering/labeling tasks. We also explore the conditions under which PICA provides better inference accuracy compared to a simpler but related model based on right-stochastic matrices. Finally, we show that PICA can be used to crowdsource responses for dense image segmentation tasks, and provide a proof-of-concept that aggregating responses in this way could improve the accuracy of this labor-intensive task.
Michael Giancola, Randy C. Paffenroth, Jacob Whitehill
HCOMP3
2017 Getting to Know English Language Learners in MOOCs: Their Motivations, Behaviors, and Outcomes
abstract
Massive Open Online Courses (MOOCs) promise to engage a global audience and emphasize the democratic achievement of free, university-level education. While such open access enables participation, it is unclear how learners who are not fluent in English (ELLs) engage with MOOC content. After all, the language of MOOCs is English. In order to improve accessibility for ELLs in digital learning environments, we must first have a clear understanding of the educational landscape: who are the non-native English speakers enrolled in MOOCs? Where are they located geographically? What are their current online learning behaviors, motivations and outcomes? In this paper we start answering some of these questions by analyzing data from 100 HarvardX courses, using self-report and log data. Preliminary analysis show evidence that ELLs are motivated by more utilitarian goals compared to non-ELLs.
Selen Türkay, Hadas Eidelman, Yigal Rosen, Daniel T. Seaton, Glenn Lopez, Jacob Whitehill
L@S6
2017 MOOC Dropout Prediction: How to Measure Accuracy?
abstract
In order to obtain reliable accuracy estimates for automatic MOOC dropout predictors, it is important to train and test them in a manner consistent with how they will be used in practice. Yet most prior research on MOOC dropout prediction has measured test accuracy on the same course used for training, which can lead to overly optimistic accuracy estimates. In order to understand better how accuracy is affected by the training+testing regime, we compared the accuracy of a standard dropout prediction architecture (clickstream features + logistic regression) across 4 different training paradigms. Results suggest that (1) training and testing on the same course ("post-hoc") can significantly overestimate accuracy. Moreover, (2) training dropout classifiers using proxy labels based on students' persistence -- which are available before a MOOC finishes -- is surprisingly competitive with post-hoc training (87.33% v.~90.20% AUC averaged over 8 weeks of 40 HarvardX MOOCs) and can support real-time MOOC interventions.
Jacob Whitehill, Kiran Mohan, Daniel T. Seaton, Yigal Rosen, Dustin Tingley
L@S1
2017 A Crowdsourcing Approach to Collecting Tutorial Videos - Toward Personalized Learning-at-Scale
abstract
We investigated the feasibility of crowdsourcing full- fledged tutorial videos from ordinary people on the Web on how to solve math problems related to logarithms. This kind of approach (a form of learnersourcing [9, 11]) to efficiently collecting tutorial videos and other learning resources could be useful for realizing personalized learning-at-scale, whereby students receive specific learning resources -- drawn from a large and diverse set -- that are tailored to their individual and time-varying needs. Results of our study, in which we collected 399 videos from 66 unique "teachers" on Mechanical Turk, suggest that (1) approximately 100 videos -- over 80% of which are mathematically fully correct -- can be crowdsourced per week for $5/video; (2) the average learning gains (posttest minus pretest score) associated with watching the videos was stat. sig. higher than for a control video (0.105 versus 0.045); and (3) the average learning gains (0.1416) from watching the best tested crowdsourced videos was comparable to the learning gains (0.1506) from watching a popular Khan Academy video on logarithms.
Jacob Whitehill, Margo I. Seltzer
L@S1
2016 Exploiting an Oracle That Reports AUC Scores in Machine Learning Contests
abstract
In machine learning contests such as the ImageNet Large Scale Visual Recognition Challenge and the KDD Cup, contestants can submit candidate solutions and receive from an oracle (typically the organizers of the competition) the accuracy of their guesses compared to the ground-truth labels. One of the most commonly used accuracy metrics for binary classification tasks is the Area Under the Receiver Operating Characteristics Curve (AUC). In this paper we provide proofs-of-concept of how knowledge of the AUC of a set of guesses can be used, in two different kinds of attacks, to improve the accuracy of those guesses. On the other hand, we also demonstrate the intractability of one kind of AUC exploit by proving that the number of possible binary labelings of n examples for which a candidate solution obtains a AUC score of c grows exponentially in n, for every c in (0,1).
Jacob Whitehill
AAAI1
2015 Beyond Prediction: Towards Automatic Intervention in MOOC Student Stop-out
Jacob Whitehill, Joseph Jay Williams, Glenn Lopez, Cody A. Coleman, Justin Reich
EDM1
2014 The Faces of Engagement: Automatic Recognition of Student Engagementfrom Facial Expressions
abstract
Student engagement is a key concept in contemporary education, where it is valued as a goal in its own right. In this paper we explore approaches for automatic recognition of engagement from students' facial expressions. We studied whether human observers can reliably judge engagement from the face; analyzed the signals observers use to make these judgments; and automated the process using machine learning. We found that human observers reliably agree when discriminating low versus high degrees of engagement (Cohen's κ = 0.96). When fine discrimination is required (four distinct levels) the reliability decreases, but is still quite high ( κ = 0.56). Furthermore, we found that engagement labels of 10-second video clips can be reliably predicted from the average labels of their constituent frames (Pearson r=0.85), suggesting that static expressions contain the bulk of the information used by observers. We used machine learning to develop automatic engagement detectors and found that for binary classification (e.g., high engagement versus low engagement), automated engagement detectors perform with comparable accuracy to humans. Finally, we show that both human and automatic engagement judgments correlate with task performance. In our experiment, student post-test performance was predicted with comparable accuracy from engagement labels ( r=0.47) as from pre-test scores ( r=0.44).
Jacob Whitehill, Zewelanji Serpell, Yi-Ching Lin, Aysha Foster, Javier R. Movellan
IEEE Trans. Affect. Comput.1
2012 Discriminately decreasing discriminability with learned image filters
abstract
In machine learning and computer vision, input signals are often filtered to increase data discriminability. For example, preprocessing face images with Gabor band-pass filters is known to improve performance in expression recognition tasks [1]. Sometimes, however, one may wish to purposely decrease discriminability of one classification task (a “distractor” task), while simultaneously preserving information relevant to another task (the target task): For example, due to privacy concerns, it may be important to mask the identity of persons contained in face images before submitting them to a crowdsourcing site (e.g., Mechanical Turk) when labeling them for certain facial attributes. Suppressing discriminability in distractor tasks may also be needed to improve inter-dataset generalization: training datasets may sometimes contain spurious correlations between a target attribute (e.g., facial expression) and a distractor attribute (e.g., gender). We might improve generalization to new datasets by suppressing the signal related to the distractor task in the training dataset. This can be seen as a special form of supervised regularization. In this paper we present an approach to automatically learning preprocessing filters that suppress discriminability in distractor tasks while preserving it in target tasks. We present promising results in simulated image classification problems and in a realistic expression recognition problem.
Jacob Whitehill, Javier R. Movellan
CVPR1
2012 Multilayer Architectures for Facial Action Unit Recognition
abstract
In expression recognition and many other computer vision applications, the recognition performance is greatly improved by adding a layer of nonlinear texture filters between the raw input pixels and the classifier. The function of this layer is typically known as feature extraction. Popular filter types for this layer are Gabor energy filters (GEFs) and local binary patterns (LBPs). Recent work [1] suggests that adding a second layer of nonlinear filters on top of the first layer may be beneficial. However, it is unclear what is the best architecture of layers and selection of filters. In this paper, we present a thorough empirical analysis of the performance of single-layer and dual-layer texture-based approaches for action unit recognition. For the single hidden layer case, GEFs perform consistently better than LBPs, which may be due to their robustness to jitter and illumination noise as well as to their ability to encode texture at multiple resolutions. For dual-layer case, we confirm that, while small, the benefit of adding this second layer is reliable and consistent across data sets. Interestingly for this second layer, LBPs appear to perform better than GEFs.
Tingfan Wu, Nicholas J. Butko, Paul Ruvolo, Jacob Whitehill, Marian Stewart Bartlett, Javier R. Movellan
IEEE Trans. Syst. Man Cybern. Part B4
2011 The motion in emotion - A CERT based approach to the FERA emotion challenge
abstract
This paper assesses the performance of measures of facial expression dynamics derived from the Computer Expression Recognition Toolbox (CERT) for classifying emotions in the Facial Expression Recognition and Analysis (FERA) Challenge. The CERT system automatically estimates facial action intensity and head position using learned appearance-based models on single frames of video. CERT outputs were used to derive a representation of the intensity and motion in each video, consisting of the extremes of displacement, velocity and acceleration. Using this representation, emotion detectors were trained on the FERA training examples. Experiments on the released portion of the FERA dataset are presented, as well as results on the blind test. No consideration of subject identity was taken into account in the blind test. The F1 scores were well above the baseline criterion for success.
Gwen Littlewort, Jacob Whitehill, Tingfan Wu, Nicholas J. Butko, Paul Ruvolo, Javier R. Movellan, Marian Stewart Bartlett
FG2
2011 The computer expression recognition toolbox (CERT)
abstract
We present the Computer Expression Recognition Toolbox (CERT), a software tool for fully automatic real-time facial expression recognition, and officially release it for free academic use. CERT can automatically code the intensity of 19 different facial actions from the Facial Action Unit Coding System (FACS) and 6 different prototypical facial expressions. It also estimates the locations of 10 facial features as well as the 3-D orientation (yaw, pitch, roll) of the head. On a database of posed facial expressions, Extended Cohn-Kanade (CK+[1]), CERT achieves an average recognition performance (probability of correctness on a two-alternative forced choice (2AFC) task between one positive and one negative example) of 90.1% when analyzing facial actions. On a spontaneous facial expression dataset, CERT achieves an accuracy of nearly 80%. In a standard dual core laptop, CERT can process 320 × 240 video images in real time at approximately 10 frames per second.
Gwen Littlewort, Jacob Whitehill, Tingfan Wu, Ian R. Fasel, Mark G. Frank, Javier R. Movellan, Marian Stewart Bartlett
FG2
2011 Action unit recognition transfer across datasets
abstract
We explore how CERT, a computer expression recognition toolbox trained on a large dataset of spontaneous facial expressions (FFD07), generalizes to a new, previously unseen dataset (FERA). The experiment was unique in that the authors had no access to the test labels, which were guarded as part of the FERA challenge. We show that without any training or special adaptation to the new database, CERT performs better than a baseline method trained exclusively on that database. Best results are achieved by retraining CERT with a combination of old and new data. We also found that the FERA dataset may be too small and idiosyncratic to generalize to other datasets. Training on FERA alone produced good results on FERA but very poor results on FFD07. We reflect on the importance of challenges like this for the future of the field, and discuss suggestions for standardization of future challenges.
Tingfan Wu, Nicholas J. Butko, Paul Ruvolo, Jacob Whitehill, Marian Stewart Bartlett, Javier R. Movellan
FG4
2010 Monocular head pose estimation using generalized adaptive view-based appearance model
Louis-Philippe Morency, Jacob Whitehill, Javier R. Movellan
Image Vis. Comput.2
2009 Whose Vote Should Count More: Optimal Integration of Labels from Labelers of Unknown Expertise
abstract
Modern machine learning-based approaches to computer vision require very large databases of labeled images. Some contemporary vision systems already require on the order of millions of images for training (e.g., Omron face detector). While the collection of these large databases is becoming a bottleneck, new Internet-based services that allow labelers from around the world to be easily hired and managed provide a promising solution. However, using these services to label large databases brings with it new theoretical and practical challenges: (1) The labelers may have wide ranging levels of expertise which are unknown a priori, and in some cases may be adversarial; (2) images may vary in their level of difficulty; and (3) multiple labels for the same image must be combined to provide an estimate of the actual label of the image. Probabilistic approaches provide a principled way to approach these problems. In this paper we present a probabilistic model and use it to simultaneously infer the label of each image, the expertise of each labeler, and the difficulty of each image. On both simulated and real data, we demonstrate that the model outperforms the commonly used ``Majority Vote heuristic for inferring image labels, and is robust to both adversarial and noisy labelers.
Jacob Whitehill, Paul Ruvolo, Tingfan Wu, Jacob Bergsma, Javier R. Movellan
NIPS1
2009 Toward Practical Smile Detection
abstract
Machine learning approaches have produced some of the highest reported performances for facial expression recognition. However, to date, nearly all automatic facial expression recognition research has focused on optimizing performance on a few databases that were collected under controlled lighting conditions on a relatively small number of subjects. This paper explores whether current machine learning methods can be used to develop an expression recognition system that operates reliably in more realistic conditions. We explore the necessary characteristics of the training data set, image registration, feature representation, and machine learning algorithms. A new database, GENKI, is presented which contains pictures, photographed by the subjects themselves, from thousands of different people in many different real-world imaging conditions. Results suggest that human-level expression recognition accuracy in real-life illumination conditions is achievable with machine learning technology. However, the data sets currently used in the automatic expression recognition literature to evaluate progress may be overly constrained and could potentially lead research into locally optimal algorithmic solutions.
Jacob Whitehill, Gwen Littlewort, Ian R. Fasel, Marian Stewart Bartlett, Javier R. Movellan
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Generalized adaptive view-based appearance model: Integrated framework for monocular head pose estimation
abstract
Accurately estimating the person's head position and orientation is an important task for a wide range of applications such as driver awareness and human-robot interaction. Over the past two decades, many approaches have been suggested to solve this problem, each with its own advantages and disadvantages. In this paper, we present a probabilistic framework called generalized adaptive viewbased appearance model (GAVAM) which integrates the advantages from three of these approaches: (1) the automatic initialization and stability of static head pose estimation, (2) the relative precision and user-independence of differential registration, and (3) the robustness and bounded drift of keyframe tracking. In our experiments, we show how the GAVAM model can be used to estimate head position and orientation in real-time using a simple monocular camera. Our experiments on two previously published datasets show that the GAVAM framework can accurately track for a long period of time (>2 minutes) with an average accuracy of 3.5deg and 0.75 in with an inertial sensor and a 3D magnetic sensor.
Louis-Philippe Morency, Jacob Whitehill, Javier R. Movellan
FG2
2008 Personalized facial attractiveness prediction
abstract
We present a fully automatic approach to learning the personal facial attractiveness preferences of individual users directly from example images. The target application is computer assisted search of partners in online dating services. The proposed approach is based on the use of epsiv-SVMs to learn a regression function that maps low level image features onto attractiveness ratings. We present empirical results based on a dataset of images collected from a large online dating site. Our system achieved correlations of up to 0.45 (Pearson correlation) on the attractiveness predictions for individual users. We show evidence that the approach learned not just a universal sense of attraction shared by multiple users, but capitalized on the preferences of individual subjects. Our results are promising and could already be used to facilitate the personalized search of partners in online dating.
Jacob Whitehill, Javier R. Movellan
FG1
2008 A discriminative approach to frame-by-frame head pose tracking
abstract
We present a discriminative approach to frame-by-frame head pose tracking that is robust to a wide range of illuminations and facial appearances and that is inherently immune to accuracy drift. Most previous research on head pose tracking has been validated on test datasets spanning only a small (< 20) subjects under controlled illumination conditions on continuous video sequences. In contrast, the system presented in this paper was both trained and tested on a much larger database, GENKI, spanning tens of thousands of different subjects, illuminations, and geographical locations from images on the Web. Our pose estimator achieves accuracy of 5.82deg, 5.65deg, and 2.96deg root-mean-square (RMS) error for yaw, pitch, and roll, respectively. A set of 4000 images from this dataset, labeled for pose, was collected and released for use by the research community.
Jacob Whitehill, Javier R. Movellan
FG1
2008 Measuring the Perceived Difficulty of a Lecture Using Automatic Facial Expression Recognition
Jacob Whitehill, Marian Stewart Bartlett, Javier R. Movellan
Intelligent Tutoring Systems1