VLDB 2026 Research / reviewers in the wild / expert
Kirk Vanacore
dblp:278/7409 · also Kirk P. Vanacore
· DBLP profile ↗
24ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0003-0673-5721ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 24 · 7 first-author · 23 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 8 · 3 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Does the TalkMoves Codebook Generalize to One-on-One Tutoring and Multimodal Interaction?
Corina Luca Focsan, Marie Cynthia Abijuru Kamikazi, Tamisha Thompson, Jennifer St. John, Kirk Vanacore, Danielle R. Thomas, Kenneth R. Koedinger, René F. Kizilcec |
AIED (5) | 5 |
| 2026 | Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education
Danielle R. Thomas, Conrad Borchers, Kirk Vanacore, Kenneth R. Koedinger, René F. Kizilcec |
AIED (6) | 3 |
| 2026 | AI Annotation Orchestration: Evaluating LLM Verifiers to Improve the Quality of LLM Annotations in Learning AnalyticsabstractLarge Language Models (LLMs) are increasingly used to annotate learning interactions, yet concerns about reliability limit their utility. We test whether verification-oriented orchestration-prompting models to check their own labels (self-verification) or audit one another (cross-verification)-improves qualitative coding of tutoring discourse. Using transcripts from 30 one-to-one math sessions, we compare three production LLMs (GPT, Claude, Gemini) under three conditions: unverified annotation, self-verification, and cross-verification across all orchestration configurations. Outputs are benchmarked against a blinded, disagreement-focused human adjudication using Cohen's kappa. Overall, orchestration yields a 58 percent improvement in kappa. Self-verification nearly doubles agreement relative to unverified baselines, with the largest gains for challenging tutor moves. Cross-verification achieves a 37 percent improvement on average, with pair- and construct-dependent effects: some verifier-annotator pairs exceed self-verification, while others reduce alignment, reflecting differences in verifier strictness. We contribute: (1) a flexible orchestration framework instantiating control, self-, and cross-verification; (2) an empirical comparison across frontier LLMs on authentic tutoring data with blinded human "gold" labels; and (3) a concise notation, verifier(annotator) (e.g., Gemini(GPT) or Claude(Claude)), to standardize reporting and make directional effects explicit for replication. Results position verification as a principled design lever for reliable, scalable LLM-assisted annotation in Learning Analytics. Bakhtawar Ahtisham, Kirk Vanacore, Jinsook Lee, Zhuqian Zhou, Doug Pietrzak, René F. Kizilcec |
LAK | 2 |
| 2026 | Sentiment Gaps Between AI and Human Tutors: A Work-in-Progress InvestigationabstractAs generative AI tutors become increasingly common in educational settings, understanding how their communication patterns compare to human tutors is critical. This pilot study applies sentiment analysis to tutoring transcripts from a free online tutoring platform, comparing sessions with an AI tutor (N=787) to matched human tutor sessions. We find that AI tutors display significantly higher positive sentiment (M = 0.624 vs. M = 0.199, d = 2.13). Furthermore, students respond differently in terms of sentiment to human and AI tutors. While human tutor sessions show converging sentiment trajectories (54%), AI sessions more often diverge (54%). Critically, large sentiment gaps in AI sessions predicted negative student sentiment change, whereas human sessions showed no such relationship. We outline ongoing work, including BERT-based sentiment analysis, NLP-based question pattern analysis, and qualitative examination of sentiment trajectories to deepen understanding of these patterns and inform AI tutor design. Sarah Shaw, Aly Murray, Kirk Vanacore |
L@S | 3 |
| 2026 | How Well Do Large Language Models Recognize Instructional Moves? Establishing Baselines for Foundation Models in Educational DiscourseabstractLarge language models (LLMs) are increasingly used in educational contexts, yet their ability to interpret authentic instructional discourse out-of-the-box remains unclear. We benchmark six state-of-the-art LLMs on classifying instructional moves in K-12 mathematics classroom transcripts annotated by expert educators (κ>0.90). We evaluated four prompting strategies, including zero-shot, one-shot, and few-shot prompts derived from the human coding manual. Zero-shot prompting achieved fair-to-moderate agreement (κ = 0.38–0.48, F1 = 0.45–0.53). Providing comprehensive examples improved performance for some models (e.g., κ = 0.48 to 0.58 for Claude 4.5 Opus; κ = 0.38 to 0.57 for Gemini 2.5 Pro), but gains were uneven and precision remained limited (best precision = 0.56, recall = 0.75). Errors were concentrated in constructs that require inference about instructor intent; for example, models confused Press for Reasoning with Press for Accuracy (42%–53% false-positive rates). Overall, our analysis found that LLMs demonstrate meaningful but limited capacity to identify aspects of instructional discourse, providing a baseline for educational discourse benchmarking and for designing more reliable annotation workflows. This work also points to a potential weakness in LLMs' ability to interpret key nuances of educational instruction. Kirk Vanacore, René F. Kizilcec |
L@S | 1 |
| 2026 | From Tutor Moves to Tutoring States: Modeling the Timing and Sequencing of Pedagogical Strategies for Student EngagementabstractUnderstanding how tutoring unfolds requires capturing not only what tutors and students say, but also the multi-turn pedagogical strategies that structure their conversation. In this study, we analyze 77 online math tutoring sessions (about 40 hours of tutoring) using a combination of LLM-assisted annotation and statistical modeling. Building on utterance-level annotations with a taxonomy of tutor moves, we identify three recurrent latent states using a Hidden Markov Model: Inquiry Elicitation (dominated by prompting and probing student reasoning), Direct Instruction (centered on explanation and scaffolding), and Affective Support (characterized by praise and socioemotional support). Sequential pattern mining revealed that these states exhibit distinct instructional motifs—for example, Inquiry Elicitation involves repeated prompting, while Direct Instruction features sustained explanatory scaffolding. These dynamics were correlated with different levels of student engagement. Sessions starting with Inquiry Elicitation and ending with Affective Support elicited more student talk, while sustaining Direct Instruction was associated with a higher likelihood that students met tutoring dosage recommendations by returning for additional sessions. Together, these results suggest that the impact of tutoring may lie not just in which strategies tutors employ, but in how those strategies are ordered and coordinated over time—patterns that reveal the deeper conversational architecture shaping student engagement. More broadly, this work demonstrates how integrating generative AI annotation with advanced statistical modeling can scale analyses of tutoring dialogue and identify conversational practices that may prompt sustained engagement. Kirk Vanacore, Jinsook Lee, Bakhtawar Ahtisham, Sarah Shaw, Justin Reich, René F. Kizilcec |
L@S | 1 |
| 2025 | How Much Mastery is Enough Mastery? The Relationship between Mastery in a Lesson and the Performance on the Subsequent Lesson
Jiayi Zhang 0004, Kirk Vanacore, Ryan Baker 0001, Nabil Ch, Caitlin Mills 0001, Owen Henkel |
EDM | 2 |
| 2025 | CausalEDM: Linking Innovations in Instructional Design and the Complex Behaviors that Underlie Learning Processes and Outcomes
Kirk Vanacore, Anthony Botelho, Avery Harrison Closser, Adam Sales, Neil T. Heffernan |
EDM | 1 |
| 2025 | The Half-Life of Epistemic Emotions: How Motivation Influences Affective Chronometry
Andres Felipe Zambrano, Jaclyn Ocumpaugh, Ryan Baker 0001, Kirk Vanacore, Jordan Esiason, Jessica Vandenberg |
EDM | 4 |
| 2025 | Understanding MOOC Stopout Patterns: Course and Assessment-Level InsightsabstractThis study investigates stopout patterns in MOOCs to understand course and assessment-level factors that influence student stopout behavior. We expanded previous work on stopout by assessing the exponential decay of assessment-level stopout rates across courses. Results confirm a disproportionate stopout rate on the first graded assessment. We then evaluated which course and assessment level features were associated with stopout on the first assessment. Findings suggest that a higher number of questions and estimated time commitment in the early assessments and more assessments in a course may be associated with a higher proportion of early stopout behavior. Yunlang Dai, Kirk Vanacore, Ryan Baker 0001, Stefan Slater |
L@S | 2 |
| 2025 | Do MOOC Conversations Matter? Investigating the Role of Social Presence and Course-Relevant Discussion in Career AdvancementabstractWhile MOOCs have been widely studied in terms of student engagement and academic performance, the extent to which engagement within MOOCs predict career advancement remains underexplored. Building on prior work, this study investigates how participation in discussion forums, specifically social presence and the use of course-relevant keywords, affects career advancement. Using GPT-assisted content analysis of forum posts, we assess how these engagement factors relate to both achievement during the course and post-course career advancement. Our findings indicate that social presence and use of course-relevant keywords has a positive relationship with course achievement during the MOOC. However, no significant relationship was found between career advancement and either social presence or course-related keywords in discussion forums. These findings suggest that while active engagement in MOOC discussion forums enhances academic achievement, it might not directly translate into career advancement, highlighting a possible disconnect between learning participation in MOOCs and professional outcomes. Shruti Mehta, Namrata Srivastava, Xiner Liu, Kirk Vanacore, Ryan Baker 0001 |
L@S | 4 |
| 2025 | Scaling Effective AI-Generated Explanations for Middle School Mathematics in Online Learning Platforms
Eamon Worden, Kirk Vanacore, Aaron Haim, Neil T. Heffernan |
L@S | 2 |
| 2024 | Causal Inference in Educational Data Mining
Anthony Botelho, Avery Harrison Closser, Adam Sales, Neil T. Heffernan, Kirk Vanacore |
EDM | 5 |
| 2024 | LOOL: Towards Personalization with Flexible \& Robust Estimation of Heterogeneous Treatment Effects
Duy M. Pham, Kirk Vanacore, Adam Sales, Johann Gagnon-Bartsch |
EDM | 2 |
| 2024 | Problem-Solving Behavior and EdTech Effectiveness: A Model for Exploratory Causal Analysis
Adam Sales, Kirk Vanacore, Hyeon-Ah Kang, Tiffany A. Whittaker |
EDM | 2 |
| 2024 | Multiple Choice vs. Fill-In Problems: The Trade-off Between Scalability and LearningabstractLearning experience designers consistently balance the trade-off between open and close-ended activities. The growth and scalability of Computer Based Learning Platforms (CBLPs) have only magnified the importance of these design trade-offs. CBLPs often utilize close-ended activities (i.e. Multiple-Choice Questions [MCQs]) due to feasibility constraints associated with the use of open-ended activities. MCQs offer certain affordances, such as immediate grading and the use of distractors, setting them apart from open-ended activities. Our current study examines the effectiveness of Fill-In problems as an alternative to MCQs for middle school mathematics. We report on a randomized study conducted from 2017 to 2022, with a total of 6,768 students from middle schools across the US. We observe that, on average, Fill-In problems lead to better post-test performance than MCQs; albeit deeper explorations indicate differences between the two design paradigms to be more nuanced. We find evidence that students with higher math knowledge benefit more from Fill-In problems than those with lower math knowledge. Ashish Gurung, Kirk Vanacore, Andrew A. McReynolds, Korinn S. Ostrow, Eamon Worden, Adam Sales, Neil T. Heffernan |
LAK | 2 |
| 2024 | The Effect of Assistance on Gamers: Assessing The Impact of On-Demand Hints & Feedback Availability on Learning for Students Who Game the SystemabstractGaming the system, characterized by attempting to progress through a learning activity without engaging in essential learning behaviors, remains a persistent problem in computer-based learning platforms. This paper examines a simple intervention to mitigate the harmful effects of gaming the system by evaluating the impact of immediate feedback on students prone to gaming the system. Using a randomized controlled trial comparing two conditions - one with immediate hints and feedback and another with delayed access to such resources - this study employs a Fully Latent Principal Stratification model to determine whether students inclined to game the system would benefit more from the delayed hints and feedback. The results suggest differential effects on learning, indicating that students prone to gaming the system may benefit from restricted or delayed access to on-demand support. However, removing immediate hints and feedback did not fully alleviate the learning disadvantage associated with gaming the system. Additionally, this paper highlights the utility of combining detection methods and causal models to comprehend and effectively respond to students’ behaviors. Overall, these findings contribute to our understanding of effective intervention design that addresses gaming the system behaviors, consequently enhancing learning outcomes in computer-based learning platforms. Kirk Vanacore, Ashish Gurung, Adam Sales, Neil T. Heffernan |
LAK | 1 |
| 2023 | Effective Evaluation of Online Learning Interventions with Surrogate Measures
Ethan Prihar, Kirk Vanacore, Adam Sales, Neil T. Heffernan |
EDM | 2 |
| 2023 | Identification, Exploration, and Remediation: Can Teachers Predict Common Wrong Answers?abstractPrior work analyzing tutoring sessions provided evidence that highly effective tutors, through their interaction with students and their experience, can perceptively recognize incorrect processes or “bugs” when students incorrectly answer problems. Researchers have studied these tutoring interactions examining instructional approaches to address incorrect processes and observed that the format of the feedback can influence learning outcomes. In this work, we recognize the incorrect answers caused by these buggy processes as Common Wrong Answers (CWAs). We examine the ability of teachers and instructional designers to identify CWAs proactively. As teachers and instructional designers deeply understand the common approaches and mistakes students make when solving mathematical problems, we examine the feasibility of proactively identifying CWAs and generating Common Wrong Answer Feedback (CWAFs) as a formative feedback intervention for addressing student learning needs. As such, we analyze CWAFs in three sets of analyses. We first report on the accuracy of the CWAs predicted by the teachers and instructional designers on the problems across two activities. We then measure the effectiveness of the CWAFs using an intent-to-treat analysis. Finally, we explore the existence of personalization effects of the CWAFs for the students working on the two mathematics activities. Ashish Gurung, Sami Baral, Kirk Vanacore, Andrew A. McReynolds, Hilary Kreisberg, Anthony Botelho, Stacy T. Shaw, Neil T. Heffernan |
LAK | 3 |
| 2023 | Impact of Non-Cognitive Interventions on Student Learning Behaviors and Outcomes: An analysis of seven large-scale experimental inventionsabstractAs evidence grows supporting the importance of non-cognitive factors in learning, computer-assisted learning platforms increasingly incorporate non-academic interventions to influence student learning and learning related-behaviors. Non-cognitive interventions often attempt to influence students’ mindset, motivation, or metacognitive reflection to impact learning behaviors and outcomes. In the current paper, we analyze data from five experiments, involving seven treatment conditions embedded in mastery-based learning activities hosted on a computer-assisted learning platform focused on middle school mathematics. Each treatment condition embodied a specific non-cognitive theoretical perspective. Over seven school years, 20,472 students participated in the experiments. We estimated the effects of each treatment condition on students’ response time, hint usage, likelihood of mastering knowledge components, learning efficiency, and post-tests performance. Our analyses reveal a mix of both positive and negative treatment effects on student learning behaviors and performance. Few interventions impacted learning as assessed by the post-tests. These findings highlight the difficulty in positively influencing student learning behaviors and outcomes using non-cognitive interventions. Kirk Vanacore, Ashish Gurung, Andrew A. McReynolds, Allison S. Liu, Stacy T. Shaw, Neil T. Heffernan |
LAK | 1 |
| 2023 | How Common are Common Wrong Answers? Crowdsourcing Remediation at ScaleabstractSolving mathematical problems is cognitively complex, involving strategy formulation, solution development, and the application of learned concepts. However, gaps in students' knowledge or weakly grasped concepts can lead to errors. Teachers play a crucial role in predicting and addressing these difficulties, which directly influence learning outcomes. However, preemptively identifying misconceptions leading to errors can be challenging. This study leverages historical data to assist teachers in recognizing common errors and addressing gaps in knowledge through feedback. We present a longitudinal analysis of incorrect answers from the 2015-2020 academic years on two curricula, Illustrative Math and EngageNY, for grades 6, 7, and 8. We find consistent errors across 5 years despite varying student and teacher populations. Based on these Common Wrong Answers (CWAs), we designed a crowdsourcing platform for teachers to provide Common Wrong Answer Feedback (CWAF). This paper reports on an in vivo randomized study testing the effectiveness of CWAFs in two scenarios: next-problem-correctness within-skill and next-problem-correctness within-assignment, regardless of the skill. We find that receiving CWAF leads to a significant increase in correctness for consecutive problems within-skill. However, the effect was not significant for all consecutive problems within-assignment, irrespective of the associated skill. This paper investigates the potential of scalable approaches in identifying Common Wrong Answers (CWAs) and how the use of crowdsourced CWAFs can enhance student learning through remediation. Ashish Gurung, Sami Baral, Morgan P. Lee, Adam Sales, Aaron Haim, Kirk Vanacore, Andrew A. McReynolds, Hilary Kreisberg, Cristina Heffernan, Neil T. Heffernan |
L@S | 6 |
| 2023 | Benefit of Gamification for Persistent Learners: Propensity to Replay Problems Moderates Algebra-Game EffectivenessabstractComputer-assisted learning platforms (CALPS) increasingly include gamified elements to improve student outcomes by enhancing their engagement with content. Although evidence exists that gamified programs increase engagement and learning outcomes, there is little causal research on what programmatic mechanisms drive the effect between engagement and learning. In the following paper, we explore this relationship through a method of causal moderation known as fully latent principal stratification. Using data from a large-scale randomized control trial assessing gamified and traditional CALP systems' effects on algebraic knowledge, we estimate the impact of using the gamified CALP on students who engage with one of its key gamification elements---replaying a problem after a suboptimal attempt. The gamified CALP asks students to manipulate algebraic expressions from start to goal states and provides feedback based on the efficiency of these manipulations, allowing students to replay the problems when their efficiency can be improved. We find that the effect of gamification is greater for students with a higher propensity to replay problems. This finding suggests that gamification elements that provide students with opportunities to retry problems are driving the game's efficacy and provide evidence for a scalable mechanism of gamification that can improve students' learning. Kirk Vanacore, Adam Sales, Allison S. Liu, Erin Ottmar |
L@S | 1 |
| 2021 | Longitudinal Clusters of Online Educator Portal Access: Connecting Educator Behavior to Student OutcomesabstractThe rising prevalence of blended learning programs has provided educators with an abundance of information about students' specific educational needs through educator portals. Full implementation of blended learning models requires educators to utilize these data to inform their teaching practices, yet most research on blended learning programs focuses solely on student engagement with the digital learning environment. In this paper, we utilize a longitudinal clustering method to identify patterns of educator portal usage and examine the associations between these clusters and student program outcomes. The clusters of educators varied in intensity and consistency of educator portal access across a school year and were associated with significant differences in student usage and progress in the program. The analyses allowed us to identify preferable educator usage patterns based upon their associated students’ program outcomes, which provides novel information about the potential impact of educator engagement on overall implementation fidelity of blended learning programs. Kirk Vanacore, Kevin Dieter, Lisa Hurwitz, Jamie Studwell |
LAK | 1 |
| 2020 | Differential Responses to Personalized Learning Recommendations Revealed by Event-Related Analysis
Kevin Dieter, Jamie Studwell, Kirk Vanacore |
EDM | 3 |