EDBT 2026 Demo / reviewers in the wild / expert
Danielle R. Thomas
dblp:340/6741 · also Danielle R. Chine, Danielle Thomas
· DBLP profile ↗
19ranked-venue papers
12as first author
19since 2021 · last 2026
0000-0001-8196-3252ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 12 first-author · 18 since 2021Human-computer interaction and ubiquitous computing · 13 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Does the TalkMoves Codebook Generalize to One-on-One Tutoring and Multimodal Interaction?
Corina Luca Focsan, Marie Cynthia Abijuru Kamikazi, Tamisha Thompson, Jennifer St. John, Kirk Vanacore, Danielle R. Thomas, Kenneth R. Koedinger, René F. Kizilcec |
AIED (5) | 6 |
| 2026 | Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education
Danielle R. Thomas, Conrad Borchers, Kirk Vanacore, Kenneth R. Koedinger, René F. Kizilcec |
AIED (6) | 1 |
| 2026 | Brief but Impactful: How Human Tutoring Interactions Shape Engagement in Online LearningabstractLearning analytics can guide human tutors to efficiently address motivational barriers to learning that AI systems struggle to support. Students become more engaged when they receive human attention. However, what occurs during short interventions, and when are they most effective? We align student–tutor dialogue transcripts with MATHia tutoring system log data to study brief human-tutor interactions on Zoom drawn from 2,075 hours of 191 middle school students’ classroom math practice. Mixed-effect models reveal that engagement, measured as successful solution steps per minute, is higher during a human-tutor visit and remains elevated afterward. Visit length exhibits diminishing returns: engagement rises during and shortly after visits, irrespective of visit length. Timing also matters: later visits yield larger immediate lifts than earlier ones, though an early visit remains important to counteract engagement decline. We create analytics that identify which tutor-student dialogues raise engagement the most. Qualitative analysis reveals that interactions with concrete, stepwise scaffolding with explicit work organization elevate engagement most strongly. We discuss implications for resource-constrained tutoring, prioritizing several brief, well-timed check-ins by a human tutor while ensuring at least one early contact. Our analytics can guide the prioritization of students for support and surface effective tutor moves in real-time. Conrad Borchers, Ashish Gurung, Qinyi Liu, Danielle R. Thomas, Mohammad Khalil, Kenneth R. Koedinger |
LAK | 4 |
| 2026 | Coasting Through Class: Learning Opportunity Loss from Practice Avoidance During Individual SeatworkabstractMeasures of disengagement provide insights into unproductive use of learning opportunities. Although measures of active disengagement, such as gaming the system and mind-wandering, are well studied, loss of practice time due to outright task avoidance remains relatively understudied. The current study addresses this gap by extending existing within-task measures (idle time) with two new session-level measures (delayed start and early stop) to capture loss of practice time due to task avoidance. We characterize the combined lost time as coasted time and the associated behavior as coasting behavior. Using ASSISTments logs (N = 1,425), we find that students dedicate only 40% of available classwork time to math practice and coast through the remaining 60%. Of the coasted time, 36% resulted from delayed starts, 2% from mid-practice idling, and 62% from stopping early. Delayed start and early stop showed moderate temporal stability (G = 0.73 and 0.71, respectively), suggesting that coasting is a consistent behavioral pattern. Even after excluding early stops attributable to assignment completion (i.e., early stop = 0), coasted time remained substantial at 32%. While we observe significant differences in coasting by gender and IEP status, we do not observe them by other demographic factors or school locale. Critically, students who continued working beyond the first assignment completion (''extra effort'') performed significantly better on standardized tests. For research, coasting offers a new lens on opportunity loss by combining session-level disengagement with within-task disengagement. For practitioners, our results highlight the need for platform affordances that support sustained engagement and more productive use of available practice time. Ashish Gurung, Jordan Gutterman, Danielle R. Thomas, Mingyu Feng, Vincent Aleven, Kenneth R. Koedinger |
L@S | 3 |
| 2025 | Human Tutoring Improves the Impact of AI Tutor Use on Learning Outcomes
Ashish Gurung, Jionghao Lin, Jordan Gutterman, Danielle R. Thomas, Alex Houk, Shivang Gupta, Emma Brunskill, Lee G. Branstetter, Vincent Aleven, Kenneth R. Koedinger |
AIED (4) | 4 |
| 2025 | Improving Open-Response Assessment with LearnLM
Danielle R. Thomas, Conrad Borchers, Shambhavi Bhushan, Sanjit Kakarla, Alex Houk, Ralph Abboud, Shivang Gupta, Erin Gatz, Kenneth R. Koedinger |
AIED (5) | 1 |
| 2025 | Leveraging LLMs to Assess Tutor Moves in Real-Life Dialogues: A Feasibility Study
Danielle R. Thomas, Conrad Borchers, Jionghao Lin, Sanjit Kakarla, Shambhavi Bhushan, Erin Gatz, Shivang Gupta, Ralph Abboud, Kenneth R. Koedinger |
EC-TEL (2) | 1 |
| 2025 | Does Multiple Choice Have a Future in the Age of Generative AI? A Posttest-only RCTabstractThe role of multiple-choice questions (MCQs) as effective learning tools has been debated in past research. While MCQs are widely used due to their ease in grading, open response questions are increasingly used for instruction, given advances in large language models (LLMs) for automated grading. This study evaluates MCQs effectiveness relative to open-response questions, both individually and in combination, on learning. These activities are embedded within six tutor lessons on advocacy. Using a posttest-only randomized control design, we compare the performance of 234 tutors (790 lesson completions) across three conditions: MCQ only, open response only, and a combination of both. We find no significant learning differences across conditions at posttest, but tutors in the MCQ condition took significantly less time to complete instruction. These findings suggest that MCQs are as effective, and more efficient, than open response tasks for learning when practice time is limited. To further enhance efficiency, we autograded open responses using GPT-4o and GPT-4-turbo. GPT models demonstrate proficiency for purposes of low-stakes assessment, though further research is needed for broader use. This study contributes a dataset of lesson log data, human annotation rubrics, and LLM prompts to promote transparency and reproducibility. Danielle R. Thomas, Conrad Borchers, Sanjit Kakarla, Jionghao Lin, Shambhavi Bhushan, Boyuan Guo, Erin Gatz, Kenneth R. Koedinger |
LAK | 1 |
| 2025 | Do Tutors Learn from Equity Training and Can Generative AI Assess It?abstractEquity is a core concern of learning analytics. However, applications that teach and assess equity skills, particularly at scale are lacking, often due to barriers in evaluating language. Advances in generative AI via large language models (LLMs) are being used in a wide range of applications, with this present work assessing its use in the equity domain. We evaluate tutor performance within an online lesson on enhancing tutors' skills when responding to students in potentially inequitable situations. We apply a mixed-method approach to analyze the performance of 81 undergraduate remote tutors. We find marginally significant learning gains with increases in tutors' self-reported confidence in their knowledge in responding to middle school students experiencing possible inequities from pretest to posttest. Both GPT-4o and GPT-4-turbo demonstrate proficiency in assessing tutors ability to predict and explain the best approach. Balancing performance, efficiency, and cost, we determine that few-shot learning using GPT-4o is the preferred model. This work makes available a dataset of lesson log data, tutor responses, rubrics for human annotation, and generative AI prompts. Future work involves leveling the difficulty among scenarios and enhancing LLM prompts for large-scale grading and assessment. Danielle R. Thomas, Conrad Borchers, Sanjit Kakarla, Jionghao Lin, Shambhavi Bhushan, Boyuan Guo, Erin Gatz, Kenneth R. Koedinger |
LAK | 1 |
| 2025 | Advancing the Science of Teaching with Tutoring Data: A Collaborative Workshop with the National Tutoring ObservatoryabstractL@S ’25, Palermo, Italy Danielle R. Thomas, Dorottya Demszky, Kenneth R. Koedinger, Josh Marland, Doug Pietrzak, Justin Reich, Rachel Slama, Amalia Christina Toutziaridi, René F. Kizilcec |
L@S | 1 |
| 2024 | The Neglected 15%: Positive Effects of Hybrid Human-AI Tutoring Among Students with Disabilities
Danielle R. Thomas, Erin Gatz, Shivang Gupta, Vincent Aleven, Kenneth R. Koedinger |
AIED (1) | 1 |
| 2024 | How Can I Improve? Using GPT to Highlight the Desired and Undesired Parts of Open-ended Responses
Jionghao Lin, Eason Chen, Zifei FeiFei Han, Ashish Gurung, Danielle R. Thomas, Ngoc Dang Nguyen, Kenneth R. Koedinger |
EDM | 5 |
| 2024 | Improving Student Learning with Hybrid Human-AI Tutoring: A Three-Study Quasi-Experimental InvestigationabstractArtificial intelligence (AI) applications to support human tutoring have potential to significantly improve learning outcomes, but engagement issues persist, especially among students from low-income backgrounds. We introduce an AI-assisted tutoring model that combines human and AI tutoring and hypothesize this synergy will have positive impacts on learning processes. To investigate this hypothesis, we conduct a three-study quasi-experiment across three urban and low-income middle schools: 1) 125 students in a Pennsylvania school; 2) 385 students (50% Latinx) in a California school, and 3) 75 students (100% Black) in a Pennsylvania charter school, all implementing analogous tutoring models. We compare learning analytics of students engaged in human-AI tutoring compared to students using math software only. We find human-AI tutoring has positive effects, particularly in student’s proficiency and usage, with evidence suggesting lower achieving students may benefit more compared to higher achieving students. We illustrate the use of quasi-experimental methods adapted to the particulars of different schools and data-availability contexts so as to achieve the rapid data-driven iteration needed to guide an inspired creation into effective innovation. Future work focuses on improving the tutor dashboard and optimizing tutor-student ratios, while maintaining annual costs per student of approximately $700 annually. Danielle R. Thomas, Jionghao Lin, Erin Gatz, Ashish Gurung, Shivang Gupta, Kole Norberg, Stephen Fancsali, Vincent Aleven, Lee G. Branstetter, Emma Brunskill, Kenneth R. Koedinger |
LAK | 1 |
| 2024 | Learning and AI Evaluation of Tutors Responding to Students Engaging in Negative Self-TalkabstractAddressing negative self-talk by students, such as responding to a student when saying, "I am dumb"or "I can't do this"can be difficult for even the most experienced tutor. Despite potential tutor learning from scenario-based lessons on this topic, human-graded assessment remains time-consuming. Leveraging generative AI for evaluating textual responses in online training presents a scalable solution. Research suggests a tutor validates student's feelings when they speak negatively of themselves, e.g., by a tutor responding, "I understand how you feel"or "I recognize this is difficult."This ongoing work assesses the performance of 60 undergraduate tutors within an online lesson on enhancing tutors' abilities to respond to students engaging in negative self-talk. We find statistically significant tutor learning gains from pretest to posttest. Additionally, we describe a method of using generative AI for assessing tutors' responses to predict the best approach and subsequently explain the rationale behind it. Using the large language model GPT-4, we find high absolute performance when evaluating tutor responses involving predicting (F1 = 0.85) and explaining (F1 = 0.83) the best approach. Minor improvements are needed to the lesson itself. A future goal of this work is to fully develop automated systems of assessing tutor learning attending to barriers to students' motivation and doing so at scale. Danielle R. Thomas, Jionghao Lin, Shambhavi Bhushan, Ralph Abboud, Erin Gatz, Shivang Gupta, Kenneth R. Koedinger |
L@S | 1 |
| 2023 | When the Tutor Becomes the Student: Design and Evaluation of Efficient Scenario-based Lessons for TutorsabstractTutoring is among the most impactful educational influences on student achievement, with perhaps the greatest promise of combating student learning loss. Due to its high impact, organizations are rapidly developing tutoring programs and discovering a common problem- a shortage of qualified, experienced tutors. This mixed methods investigation focuses on the impact of short (∼15 min.), online lessons in which tutors participate in situational judgment tests based on everyday tutoring scenarios. We developed three lessons on strategies for supporting student self-efficacy and motivation and tested them with 80 tutors from a national, online tutoring organization. Using a mixed-effects logistic regression model, we found a statistically significant learning effect indicating tutors performed about 20% higher post-instruction than pre-instruction (β = 0.811, p < 0.01). Tutors scored ∼30% better on selected compared to constructed responses at posttest with evidence that tutors are learning from selected-response questions alone. Learning analytics and qualitative feedback suggest future design modifications for larger scale deployment, such as creating more authentically challenging selected-response options, capturing common misconceptions using learnersourced data, and varying modalities of scenario delivery with the aim of maintaining learning gains while reducing time and effort for tutor participants and trainers. Danielle R. Thomas, Shivang Gupta, Adetunji Adeniran, Elizabeth A. McLaughlin, Kenneth R. Koedinger |
LAK | 1 |
| 2023 | Using latent variable models to make gaming-the-system detection robust to context variationsabstractGaming the system, a behavior in which learners exploit a system's properties to make progress while avoiding learning, has frequently been shown to be associated with lower learning. However, when we applied a previously validated gaming detector across conditions in experiments with an algebra tutor, the detected gaming was not associated with reduced learning, challenging its validity in our study context. Our exploratory data analysis suggested that varying contextual factors across and within conditions contributed to this lack of association. We present a new approach, latent variable-based gaming detection (LV-GD), that controls for contextual factors and more robustly estimates student-level latent gaming tendencies. In LV-GD, a student is estimated as having a high gaming tendency if the student is detected to game more than the expected level of the population given the context. LV-GD applies a statistical model on top of an existing action-level gaming detector developed based on a typical human labeling process, without additional labeling effort. Across three datasets, we find that LV-GD consistently outperformed the original detector in validity measured by association between gaming and learning as well as reliability. LV-GD also afforded high practical utility: it more accurately revealed intervention effects on gaming, revealed a correlation between gaming and perceived competence in math and helped understand productive detected gaming behaviors. Our approach is not only useful for others wanting a cost-effective way to adapt a gaming detector to their context but is also generally applicable in creating robust behavioral measures. Yun Huang 0002, Steven Dang, J. Elizabeth Richey, Pallavi Chhabra, Danielle R. Thomas, Michael W. Asher, Nikki G. Lobczowski, Elizabeth A. McLaughlin, Judith M. Harackiewicz, Vincent Aleven, Kenneth R. Koedinger |
User Model. User Adapt. Interact. | 5 |
| 2022 | Educational Equity Through Combined Human-AI Personalization: A Propensity Matching Evaluation
Danielle R. Thomas, Cassandra Brentley, Carmen Thomas-Browne, J. Elizabeth Richey, Abdulmenaf Gul, Paulo Carvalho 0004, Lee G. Branstetter, Kenneth R. Koedinger |
AIED (1) | 1 |
| 2022 | Item Response Theory-Based Gaming Detection
Yun Huang 0002, Steven Dang, J. Elizabeth Richey, Michael W. Asher, Nikki G. Lobczowski, Danielle R. Thomas, Elizabeth A. McLaughlin, Judith M. Harackiewicz, Vincent Aleven, Kenneth R. Koedinger |
EDM | 6 |
| 2022 | Development of Scenario-based Mentor Lessons: An Iterative Design Process for Training at ScaleabstractIn this demonstration, we showcase the recent advancement of scenario-based tutor training with a focus to scale by applying the learn-by-doing approach to teaching strategies to provide socio-motivational support. These short (~15 min.) self-paced lessons use the predict-observe-explain inquiry method to develop mentor capacity in bolstering student motivation (i.e., fostering growth mindset). These custom training modules are being created to provide supplemental mentor support within the Personalized Learning2 system, an app which combines human tutoring and student math software to improve mentoring efficiency by connecting mentors to personalized resources, such as scenario-based mentor lessons, based on individual needs. Enhancing mentor training will aid in better quality mentoring at low cost. Mentor training is most effective when scenario-based practice provides trainees with response-specific feedback. To achieve feedback at scale, we illustrate an iterative design effort toward creating selected-response tasks that maintain some of the authenticity benefits of constructed-response. These scenario-based mentor lessons will be used by national level mentoring organizations as part of our efforts to scale. Danielle R. Thomas, Pallavi Chhabra, Adetunji Adeniran, Shivang Gupta, Kenneth R. Koedinger |
L@S | 1 |