Nathaniel Blanchard

dblp:145/7729 · also Nathan Blanchard · DBLP profile ↗
← Back
31ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-2653-0873ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 16 · 4 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 15 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Ordered Network Analysis of Epistemic Emotions During Collaborative Problem Solving
Sifatul Anindho, Videep Venkatesha, Jaclyn Ocumpaugh, Nathaniel Blanchard
AIED4
2026 Using LLMs to Annotate Pedagogical Moves: You Know What I Mean?
Videep Venkatesha, Sifatul Anindho, Ethan Seefried, Nathaniel Blanchard
AIED4
2025 The Impact of Background Speech on Interruption Detection in Collaborative Groups
Mariah Bradford, Nikhil Krishnaswamy, Nathaniel Blanchard
AIED (3)3
2025 HuTCH: Human Teachable Concept Highlighter for Post-hoc Visual Explanations
Erfan Mirhaji, Nikhil Krishnaswamy, Jill Zarestky, Lisa Mason, Sarath Sreedharan, Nathaniel Blanchard
AIED (5)6
2025 A Linguistic Analysis of Spontaneous Thoughts: Investigating Experiences of Deja Vu, Unexpected Thoughts, and Involuntary Autobiographical Memories
Videep Venkatesha, Mary Cati Poulos, Christopher Steadman, Caitlin Mills 0001, Anne M. Cleary, Nathaniel Blanchard
CogSci6
2025 Clean Training, Clear Skies: Virtual Reality Training for Expert Smoke Opacity Certification
abstract
Current smoke opacity certification methods are environmentally harmful and logistically inefficient, requiring real emissions, travel, and time off work. In this study, we present a Virtual Reality (VR) training and assessment system that simulates black and white smoke emissions in an immersive, controlled environment, designed to mirror real-world training and testing standards. Participants in the VR group completed two training sessions and a standardized test adapted from an official certification exam. Their performance was compared to a control group that received traditional lecture-based training and paper-based testing. Certification required a test score below 37 for both smoke colors. The VR group achieved mean scores of 33.44 (black smoke) and 57.75 (white smoke), outperforming the control group, which scored 46.00 and 72.00 respectively. The VR group significantly outperformed the control on black smoke$(p<0.05)$achieving passing scores on nearly all tested opacities ($<3$points (15%) deviation per reading). Subjectively, VR participants reported feeling more confident, better prepared, and found the application enjoyable. These results suggest VR is an effective, scalable alternative to traditional training, enhancing performance, increasing user satisfaction, and reducing environmental impact.
Ethan Seefried, Videep Venkatesha, Nathaniel Blanchard, Mohammed Safayet Arefin
ISMAR3
2025 The Choice to Use Automation: Improvements from Evidence Accumulation
abstract
This study examined discretionary automation – the choice to engage optional, automated support. In a dynamic decision-making task, after manual and aided trials, participants chose whether to engage a machine learning decision aid. The aid improved performance but participants frequently disused automation and lagged in compliance. Experiment 2 forced greater evidence accumulation, improving compliance, automation use, and performance. Experiment 3 required an initial judgment coupled with subsequent evidence accumulation. This reduced automation use but retained higher levels of compliance and performance. Across experiments, trust in automation and self-confidence influenced use decisions, but task difficulty had little impact. Implications for human-automation systems are discussed.
Colleen E. Patton, Benjamin A. Clegg, Blake C. Davis, Turgay Caglar, Caspian Siebert, Nathaniel Blanchard
Int. J. Hum. Comput. Interact.6
2024 Common Ground Tracking in Multimodal Dialogue
abstract
Within Dialogue Modeling research in AI and NLP, considerable attention has been spent on “dialogue state tracking” (DST), which is the ability to update the representations of the speaker’s needs at each turn in the dialogue by taking into account the past dialogue moves and history. Less studied but just as important to dialogue modeling, however, is “common ground tracking” (CGT), which identifies the shared belief space held by all of the participants in a task-oriented dialogue: the task-relevant propositions all participants accept as true. In this paper we present a method for automatically identifying the current set of shared beliefs and ”questions under discussion” (QUDs) of a group with a shared goal. We annotate a dataset of multimodal interactions in a shared physical space with speech transcriptions, prosodic features, gestures, actions, and facets of collaboration, and operationalize these features for use in a deep neural model to predict moves toward construction of common ground. Model outputs cascade into a set of formal closure rules derived from situated evidence and belief axioms and update operations. We empirically assess the contribution of each feature type toward successful construction of common ground relative to ground truth, establishing a benchmark in this novel, challenging task.
Ibrahim Khebour, Kenneth Lai, Mariah Bradford, Yifan Zhu 0014, Richard Brutti, Christopher Tam, Jingxuan Tu, Benjamin Ibarra, Nathaniel Blanchard, Nikhil Krishnaswamy, James Pustejovsky
LREC/COLING9
2024 Multimodal Cross-Document Event Coreference Resolution Using Linear Semantic Transfer and Mixed-Modality Ensembles
abstract
Event coreference resolution (ECR) is the task of determining whether distinct mentions of events within a multi-document corpus are actually linked to the same underlying occurrence. Images of the events can help facilitate resolution when language is ambiguous. Here, we propose a multimodal cross-document event coreference resolution method that integrates visual and textual cues with a simple linear map between vision and language models. As existing ECR benchmark datasets rarely provide images for all event mentions, we augment the popular ECB+ dataset with event-centric images scraped from the internet and generated using image diffusion models. We establish three methods that incorporate images and text for coreference: 1) a standard fused model with finetuning, 2) a novel linear mapping method without finetuning and 3) an ensembling approach based on splitting mention pairs by semantic and discourse-level difficulty. We evaluate on 2 datasets: the augmented ECB+, and AIDA Phase 1. Our ensemble systems using cross-modal linear mapping establish an upper limit (91.9 CoNLL F1) on ECB+ ECR performance given the preprocessing assumptions used, and establish a novel baseline on AIDA Phase 1. Our results demonstrate the utility of multimodal information in ECR for certain challenging coreference problems, and highlight a need for more multimodal resources in the coreference resolution space.
Abhijnan Nath, Huma Jamil, Shafiuddin Rehan Ahmed, George Arthur Baker, Rahul Ghosh, James H. Martin, Nathaniel Blanchard, Nikhil Krishnaswamy
LREC/COLING7
2024 Propositional Extraction from Natural Speech in Small Group Collaborative Tasks
Videep Venkatesha, Abhijnan Nath, Ibrahim Khebour, Avyakta Chelle, Mariah Bradford, Jingxuan Tu, James Pustejovsky, Nathaniel Blanchard, Nikhil Krishnaswamy
EDM8
2023 Automatic Detection of Collaborative States in Small Groups Using Multimodal Features
Mariah Bradford, Ibrahim Khebour, Nathaniel Blanchard, Nikhil Krishnaswamy
AIED3
2023 Automatically detecting task-unrelated thoughts during conversations using keystroke analysis
Vishal Kiran Kuvar, Nathaniel Blanchard, Alexander Colby, Laura K. Allen, Caitlin Mills 0001
User Model. User Adapt. Interact.2
2022 Dual Graphs of Polyhedral Decompositions for the Detection of Adversarial Attacks
abstract
Previous work has shown that a neural network with the rectified linear unit (ReLU) activation function leads to a convex polyhedral decomposition of the input space. These decompositions can be represented by a dual graph with vertices corresponding to polyhedra and edges corresponding to polyhedra sharing a facet, which is a subgraph of a Hamming graph. This paper illustrates how one can utilize the dual graph to detect and analyze adversarial attacks in the context of digital images. When an image passes through a network containing ReLU nodes, the firing or non-firing at a node can be encoded as a bit (1 for ReLU activation, 0 for ReLU non-activation). The sequence of all bit activations identifies the image with a bit vector, which identifies it with a polyhedron in the decomposition and, in turn, identifies it with a vertex in the dual graph. We identify ReLU bits that are discriminators between non-adversarial and adversarial images and examine how well collections of these discriminators can ensemble vote to build an adversarial image detector. Specifically, we examine the similarities and differences of ReLU bit vectors for adversarial images, and their non-adversarial counterparts, using a pre-trained ResNet-50 architecture. While this paper focuses on adversarial digital images, ResNet-50 architecture, and the ReLU activation function, our methods extend to other network architectures, activation functions, and types of datasets.
Huma Jamil, Christina M. Cole, Nathaniel Blanchard, Emily J. King, Michael Kirby, Chris Peterson 0001
IEEE Big Data4
2022 A deep dive into microphone hardware for recording collaborative group work
Mariah Bradford, Paige Hansen, J. Ross Beveridge, Nikhil Krishnaswamy, Nathaniel Blanchard
EDM5
2022 The VoxWorld Platform for Multimodal Embodied Agents
abstract
We present a five-year retrospective on the development of the VoxWorld platform, first introduced as a multimodal platform for modeling motion language, that has evolved into a platform for rapidly building and deploying embodied agents with contextual and situational awareness, capable of interacting with humans in multiple modalities, and exploring their environments. In particular, we discuss the evolution from the theoretical underpinnings of the VoxML modeling language to a platform that accommodates both neural and symbolic inputs to build agents capable of multimodal interaction and hybrid reasoning. We focus on three distinct agent implementations and the functionality needed to accommodate all of them: Diana, a virtual collaborative agent; Kirby, a mobile robot; and BabyBAW, an agent who self-guides its own exploration of the world.
Nikhil Krishnaswamy, William Pickard, Brittany Cates, Nathaniel Blanchard, James Pustejovsky
LREC4
2021 A Pose Proposal and Refinement Network for Better 6D Object Pose Estimation
abstract
In this paper, we present a novel, end-to-end 6D object pose estimation method that operates on RGB inputs. Our approach is composed of 2 main components: the first component classifies the objects in the input image and proposes an initial 6D pose estimate through a multi-task, CNN-based encoder/multi-decoder module. The second component, a refinement module, includes a renderer and a multi-attentional pose refinement network, which iteratively refines the estimated poses by utilizing both appearance features and flow vectors. Our refiner takes advantage of the hybrid representation of the initial pose estimates to predict the relative errors with respect to the target poses. It is further augmented by a spatial multi-attention block that emphasizes objects' discriminative feature parts. Experiments on three benchmarks for 6D pose estimation show that our proposed pipeline outperforms state-of-the-art RGB-based methods with competitive runtime performance.
Ameni Trabelsi, Mohamed Chaabane, Nathaniel Blanchard, J. Ross Beveridge
WACV3
2020 Looking Ahead: Anticipating Pedestrians Crossing with Future Frames Prediction
abstract
In this paper, we present an end-to-end future-prediction model that focuses on pedestrian safety. Specifically, our model uses previous video frames, recorded from the perspective of the vehicle, to predict if a pedestrian will cross in front of the vehicle. The long term goal of this work is to design a fully autonomous system that acts and reacts as a defensive human driver would - predicting future events and reacting to mitigate risk. We focus on pedestrian-vehicle interactions because of the high risk of harm to the pedestrian if their actions are miss-predicted. Our end-to-end model consists of two stages: the first stage is an encoder/decoder network that learns to predict future video frames. The second stage is a deep spatiotemporal network that utilizes the predicted frames of the first stage to predict the pedestrian's future action. Our system achieves state-of-the-art accuracy on the Joint Attention for Autonomous Driving (JAAD) dataset on both future frames prediction, with a pixel-wise prediction l1error of 1.12, and pedestrian behavior prediction with an average precision of 86.7.
Mohamed Chaabane, Ameni Trabelsi, Nathaniel Blanchard, J. Ross Beveridge
WACV3
2019 A Neurobiological Evaluation Metric for Neural Network Model Search
abstract
Neuroscience theory posits that the brain's visual system coarsely identifies broad object categories via neural activation patterns, with similar objects producing similar neural responses. Artificial neural networks also have internal activation behavior in response to stimuli. We hypothesize that networks exhibiting brain-like activation behavior will demonstrate brain-like characteristics, e.g., stronger generalization capabilities. In this paper we introduce a human-model similarity (HMS) metric, which quantifies the similarity of human fMRI and network activation behavior. To calculate HMS, representational dissimilarity matrices (RDMs) are created as abstractions of activation behavior, measured by the correlations of activations to stimulus pairs. HMS is then the correlation between the fMRI RDM and the neural network RDM across all stimulus pairs. We test the metric on unsupervised predictive coding networks, which specifically model visual perception, and assess the metric for statistical significance over a large range of hyperparameters. Our experiments show that networks with increased human-model similarity are correlated with better performance on two computer vision tasks: next frame prediction and object matching accuracy. Further, HMS identifies networks with high performance on both tasks. An unexpected secondary finding is that the metric can be employed during training as an early-stopping mechanism.
Nathaniel Blanchard, Jeffery Kinnison, Brandon RichardWebster, Pouya Bashivan, Walter J. Scheirer
CVPR1
2019 "Keep Me In, Coach!": A Computer Vision Perspective on Assessing ACL Injury Risk in Female Athletes
abstract
We present and share a foundational dataset of multi-angle video recordings of scripted athletic movements to enable the development of computer vision research applications that evaluate and identify lower-body injury risk. The focus of the dataset is female athletes, who are at a substantially increased risk of anterior cruciate ligament (ACL) injury and are therefore a top priority for sports science. In our study, varsity and club sport athletes perform two assessment movements (the countermovement jump and the drop jump). These jump tasks are used ubiquitously in sports medicine research to characterize athleticism and to identify risk factors that indicate ACL injury propensity. The novelty of the dataset centers on (i) the type of movement data (purposeful, evaluative movements that need to be tracked with a high degree of precision), (ii) our generalized collection method that can be replicated with ease by non-experts, and (iii) the amount of data collected (we collected data from 55 division one (D1) female athletes performing 3-5 iterations of each jumps, for a total of 480 jumps). Data from each camera was manually aligned and a fully automated pipeline was built to extract knee information from athletes. Ideally, any athlete or researcher will be able to easily replicate our setup and assemble a compatible and complementary dataset to propel the development and assessment of injury propensity models.
Nathaniel Blanchard, Kyle Skinner, Aden Kemp, Walter J. Scheirer, Patrick J. Flynn
WACV1
2017 Words matter: automatic detection of teacher questions in live classroom discourse using linguistics, acoustics, and context
abstract
We investigate automatic detection of teacher questions from audio recordings collected in live classrooms with the goal of providing automated feedback to teachers. Using a dataset of audio recordings from 11 teachers across 37 class sessions, we automatically segment the audio into individual teacher utterances and code each as containing a question or not. We train supervised machine learning models to detect the human-coded questions using high-level linguistic features extracted from automatic speech recognition (ASR) transcripts, acoustic and prosodic features from the audio recordings, as well as context features, such as timing and turn-taking dynamics. Models are trained and validated independently of the teacher to ensure generalization to new teachers. We are able to distinguish questions and non-questions with a weighted F1 score of 0.69. A comparison of the three feature sets indicates that a model using linguistic features outperforms those using acoustic-prosodic and context features for question detection, but the combination of features yields a 5% improvement in overall accuracy compared to linguistic features alone. We discuss applications for pedagogical research, teacher formative assessment, and teacher professional development.
Patrick J. Donnelly, Nathaniel Blanchard, Andrew Olney, Sean Kelly, Martin Nystrand, Sidney K. D'Mello
LAK2
2016 Semi-Automatic Detection of Teacher Questions from Human-Transcripts of Audio in Live Classrooms
Nathaniel Blanchard, Patrick J. Donnelly, Andrew Olney, Borhan Samei, Sean Kelly, Xiaoyi Sun, Brooke Ward, Martin Nystrand, Sidney K. D'Mello
EDM1
2016 Multi-sensor modeling of teacher instructional segments in live classrooms
abstract
We investigate multi-sensor modeling of teachers’ instructional segments (e.g., lecture, group work) from audio recordings collected in 56 classes from eight teachers across five middle schools. Our approach fuses two sensors: a unidirectional microphone for teacher audio and a pressure zone microphone for general classroom audio. We segment and analyze the audio streams with respect to discourse timing, linguistic, and paralinguistic features. We train supervised classifiers to identify the five instructional segments that collectively comprised a majority of the data, achieving teacher-independent F1 scores ranging from 0.49 to 0.60. With respect to individual segments, the individual sensor models and the fused model were on par for Question & Answer and Procedures & Directions segments. For Supervised Seatwork, Small Group Work, and Lecture segments, the classroom model outperformed both the teacher and fusion models. Across all segments, a multi-sensor approach led to an average 8% improvement over the state of the art approach that only analyzed teacher audio. We discuss implications of our findings for the emerging field of multimodal learning analytics.
Patrick J. Donnelly, Nathaniel Blanchard, Borhan Samei, Andrew Olney, Xiaoyi Sun, Brooke Ward, Sean Kelly, Martin Nystrand, Sidney K. D'Mello
ICMI2
2016 Identifying Teacher Questions Using Automatic Speech Recognition in Classrooms
abstract
Nathaniel Blanchard, Patrick Donnelly, Andrew M. Olney, Borhan Samei, Brooke Ward, Xiaoyi Sun, Sean Kelly, Martin Nystrand, Sidney K. D’Mello. Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2016.
Nathaniel Blanchard, Patrick J. Donnelly, Andrew Olney, Borhan Samei, Brooke Ward, Xiaoyi Sun, Sean Kelly, Martin Nystrand, Sidney K. D'Mello
SIGDIAL Conference1
2016 Automatic Teacher Modeling from Live Classroom Audio
abstract
We investigate automatic analysis of teachers' instructional strategies from audio recordings collected in live classrooms. We collected a data set of teacher audio and human-coded instructional activities (e.g., lecture, question and answer, group work) in 76 middle school literature, language arts, and civics classes from eleven teachers across six schools. We automatically segment teacher audio to analyze speech vs. rest patterns, generate automatic transcripts of the teachers' speech to extract natural language features, and compute low-level acoustic features. We train supervised machine learning models to identify occurrences of five key instructional segments (Question & Answer, Procedures and Directions, Supervised Seatwork, Small Group Work, and Lecture) that collectively comprise 76% of the data. Models are validated independently of teacher in order to increase generalizability to new teachers from the same sample. We were able to identify the five instructional segments above chance levels with F1 scores ranging from 0.64 to 0.78. We discuss key findings in the context of teacher modeling for formative assessment and professional development.
Patrick J. Donnelly, Nathaniel Blanchard, Borhan Samei, Andrew Olney, Xiaoyi Sun, Brooke Ward, Sean Kelly, Martin Nystrand, Sidney K. D'Mello
UMAP2
2015 A Study of Automatic Speech Recognition in Noisy Classroom Environments for Automated Dialog Analysis
Nathaniel Blanchard, Michael Brady 0003, Andrew Olney, Marci Glaus, Xiaoyi Sun, Martin Nystrand, Borhan Samei, Sean Kelly, Sidney K. D'Mello
AIED1
2015 Classifying Q&A from Teachers' Speech: Moving Toward an Automated System of Dialogic Analysis
Nathaniel Blanchard, Sidney K. D'Mello, Andrew Olney, Martin Nystrand
EDM1
2015 Modeling Classroom Discourse: Do Models of Predicting Dialogic Instruction Properties Generalize across Populations?
Borhan Samei, Andrew Olney, Sean Kelly, Martin Nystrand, Sidney K. D'Mello, Nathaniel Blanchard, Arthur C. Graesser
EDM6
2015 Automatic Detection of Mind Wandering During Reading Using Gaze and Physiology
abstract
Mind wandering (MW) entails an involuntary shift in attention from task-related thoughts to task-unrelated thoughts, and has been shown to have detrimental effects on performance in a number of contexts. This paper proposes an automated multimodal detector of MW using eye gaze and physiology (skin conductance and skin temperature) and aspects of the context (e.g., time on task, task difficulty). Data in the form of eye gaze and physiological signals were collected as 178 participants read four instructional texts from a computer interface. Participants periodically provided self-reports of MW in response to pseudorandom auditory probes during reading. Supervised machine learning models trained on features extracted from participants' gaze fixations, physiological signals, and contextual cues were used to detect pages where participants provided positive responses of MW to the auditory probes. Two methods of combining gaze and physiology features were explored. Feature level fusion entailed building a single model by combining feature vectors from individual modalities. Decision level fusion entailed building individual models for each modality and adjudicating amongst individual decisions. Feature level fusion resulted in an 11% improvement in classification accuracy over the best unimodal model, but there was no comparable improvement for decision level fusion. This was reflected by a small improvement in both precision and recall. An analysis of the features indicated that MW was associated with fewer and longer fixations and saccades, and a higher more deterministic skin temperature. Possible applications of the detector are discussed.
Robert Bixler, Nathaniel Blanchard, Luke Garrison, Sidney K. D'Mello
ICMI2
2015 Multimodal Capture of Teacher-Student Interactions for Automated Dialogic Analysis in Live Classrooms
abstract
We focus on data collection designs for the automated analysis of teacher-student interactions in live classrooms with the goal of identifying instructional activities (e.g., lecturing, discussion) and assessing the quality of dialogic instruction (e.g., analysis of questions). Our designs were motivated by multiple technical requirements and constraints. Most importantly, teachers could be individually micfied but their audio needed to be of excellent quality for automatic speech recognition (ASR) and spoken utterance segmentation. Individual students could not be micfied but classroom audio quality only needed to be sufficient to detect student spoken utterances. Visual information could only be recorded if students could not be identified. Design 1 used an omnidirectional laptop microphone to record both teacher and classroom audio and was quickly deemed unsuitable. In Designs 2 and 3, teachers wore a wireless Samson AirLine 77 vocal headset system, which is a unidirectional microphone with a cardioid pickup pattern. In Design 2, classroom audio was recorded with dual first- generation Microsoft Kinects placed at the front corners of the class. Design 3 used a Crown PZM-30D pressure zone microphone mounted on the blackboard to record classroom audio. Designs 2 and 3 were tested by recording audio in 38 live middle school classrooms from six U.S. schools while trained human coders simultaneously performed live coding of classroom discourse. Qualitative and quantitative analyses revealed that Design 3 was suitable for three of our core tasks: (1) ASR on teacher speech (word recognition rate of 66% and word overlap rate of 69% using Google Speech ASR engine); (2) teacher utterance segmentation (F-measure of 97%); and (3) student utterance segmentation (F-measure of 66%). Ideas to incorporate video and skeletal tracking with dual second-generation Kinects to produce Design 4 are discussed.
Sidney K. D'Mello, Andrew Olney, Nathaniel Blanchard, Borhan Samei, Xiaoyi Sun, Brooke Ward, Sean Kelly
ICMI3
2014 Domain Independent Assessment of Dialogic Properties of Classroom Discourse
Borhan Samei, Andrew Olney, Sean Kelly, Martin Nystrand, Sidney K. D'Mello, Nathaniel Blanchard, Xiaoyi Sun, Marci Glaus, Arthur C. Graesser
EDM6
2014 Automated Physiological-Based Detection of Mind Wandering during Learning
Nathaniel Blanchard, Robert Bixler, Tera Joyce, Sidney K. D'Mello
Intelligent Tutoring Systems1