Conrad Borchers

dblp:315/3884 · DBLP profile ↗
← Back
58ranked-venue papers
25as first author
58since 2021 · last 2026
0000-0003-3437-8979ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 58 · 25 first-author · 58 since 2021Human-computer interaction and ubiquitous computing · 39 · 17 first-author · 39 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 8 · 3 first-author · 8 since 2021
YearPublicationVenuePosition
2026 Understanding Teacher Revisions of Large Language Model-Generated Feedback
Conrad Borchers, Luiz A. L. Rodrigues, Newarney Torrezão da Costa, Cleon Xavier, Rafael Ferreira Leite de Mello
AIED1
2026 Multimodal Analytics of Cybersecurity Crisis Preparation Exercises: What Predicts Success?
Conrad Borchers, Valdemar Svábenský, Sandesh K. Kafle, Kevin K. Tang, Jan Vykopal
AIED (3)1
2026 Representation Learning to Study Temporal Dynamics in Tutorial Scaffolding
Conrad Borchers, Jiayi Zhang 0004, Ashish Gurung
AIED (3)1
2026 A Benchmark for Gender Bias in Large Language Model Feedback on Student Essays
Yishan Du, Conrad Borchers, Mutlu Cukurova
AIED (6)2
2026 Physiological and Semantic Patterns in Medical Teams Using an Intelligent Tutoring System
Xiaoshan Huang, Conrad Borchers, Jiayi Zhang 0004, Susanne P. Lajoie
AIED (5)2
2026 Evaluating a Data-Driven Redesign Process for Intelligent Tutoring Systems
Qianru Lyu, Conrad Borchers, Meng Xia 0002, Karen Xiao, Paulo Carvalho 0004, Kenneth R. Koedinger, Vincent Aleven
AIED (3)2
2026 Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education
Danielle R. Thomas, Conrad Borchers, Kirk Vanacore, Kenneth R. Koedinger, René F. Kizilcec
AIED (6)2
2026 Toward Trait-Aware Learning Analytics
abstract
Learning analytics (LA) draws from the learning sciences to interpret learner behavior and inform system design. Yet, past personalization remains largely at the content or performance level (during learner-system interactions), overlooking relatively stable individual differences such as personality (unfolding over long-term learning trajectories such as college degrees). The latter could bring underappreciated benefits to the design, implementation, and impact of LA. In this position paper, we conduct an ad hoc literature review and argue for an expanded framing of LA that centers on learner traits as key to both interpreting and designing close-the-loop experiments in LA. We show that personality traits are relevant to LA's central outcomes (e.g., engagement and achievement) and conducive to action, as their established ties to human-computer interaction (HCI) inform how systems time, frame, and personalize support. Drawing inspiration from HCI, where psychometrics inform personalization strategies, we propose that LA can evolve by treating traits not only as predictive features but as design resources and moderators of analytics efficacy. In line with past position papers published at LAK, we present a research agenda grounded in the LA cycle and discuss methodological and ethical challenges.
Conrad Borchers, Hannah Deininger, Zachary A. Pardos
LAK1
2026 Brief but Impactful: How Human Tutoring Interactions Shape Engagement in Online Learning
abstract
Learning analytics can guide human tutors to efficiently address motivational barriers to learning that AI systems struggle to support. Students become more engaged when they receive human attention. However, what occurs during short interventions, and when are they most effective? We align student–tutor dialogue transcripts with MATHia tutoring system log data to study brief human-tutor interactions on Zoom drawn from 2,075 hours of 191 middle school students’ classroom math practice. Mixed-effect models reveal that engagement, measured as successful solution steps per minute, is higher during a human-tutor visit and remains elevated afterward. Visit length exhibits diminishing returns: engagement rises during and shortly after visits, irrespective of visit length. Timing also matters: later visits yield larger immediate lifts than earlier ones, though an early visit remains important to counteract engagement decline. We create analytics that identify which tutor-student dialogues raise engagement the most. Qualitative analysis reveals that interactions with concrete, stepwise scaffolding with explicit work organization elevate engagement most strongly. We discuss implications for resource-constrained tutoring, prioritizing several brief, well-timed check-ins by a human tutor while ensuring at least one early contact. Our analytics can guide the prioritization of students for support and surface effective tutor moves in real-time.
Conrad Borchers, Ashish Gurung, Qinyi Liu, Danielle R. Thomas, Mohammad Khalil, Kenneth R. Koedinger
LAK1
2026 Disentangling Learning from Judgment: Representation Learning for Open Response Analytics
abstract
Open-ended responses are central to learning, yet automated scoring often conflates what students wrote with how teachers grade. We present an analytics-first framework that separates content signals from rater tendencies, making judgments visible and auditable via analytics. Using de-identified ASSISTments mathematics responses, we model teacher histories as dynamic priors and represent text with sentence embeddings. We apply centroid normalization and response–problem embedding differences, and explicitly model teacher effects with priors to reduce problem- and teacher-related confounds. Temporally-validated linear models quantify the contributions of each signal, and model disagreements surface observations for qualitative inspection. Results show that teacher priors heavily influence grade predictions; the strongest results arise when priors are combined with content embeddings (AUC ≈ 0.815), while content-only models remain above chance but substantially weaker (AUC ≈ 0.626). Adjusting for rater effects sharpens the selection of features derived from content representations, retaining more informative embedding dimensions and revealing cases where semantic evidence supports understanding as opposed to surface-level differences in how students respond. The contribution presents a practical pipeline that transforms embeddings from mere features into learning analytics for reflection, enabling teachers and researchers to examine where grading practices align (or conflict) with evidence of student reasoning and learning.
Conrad Borchers, Manit Patel, Seiyon M. Lee, Anthony Botelho
LAK1
2026 Sticky Help, Bounded Effects: Session-by-Session Analytics of Teacher Interventions in K-12 Classrooms
abstract
Teachers’ in-the-moment support is a limited resource in technology-supported classrooms, and teachers must decide whom to help and when during ongoing student work. However, less is known about how students’ prior help history (whether they were helped earlier) and their engagement states (e.g., idle, struggle) shape teachers’ decisions, and whether observed learning benefits associated with teacher help extend beyond the current class session. To address these questions, we first conducted interviews with nine K–12 mathematics teachers to identify candidate decision factors for teacher help. We then analyzed 1.4 million student–system interactions from 339 students across 14 classes in the MATHia intelligent tutoring system by linking teacher-logged help events with fine-grained engagement states. Mixed-effects models show that students who received help earlier were more likely to receive additional help later, even after accounting for current engagement state. Cross-lagged panel analyses further show that teacher help recurred across sessions, whereas idle behavior did not receive sustained attention over time. Finally, help coincided with immediate learning within sessions, but did not predict skill acquisition in later sessions, as estimated by additive factor modeling. These findings suggest that teacher help is “sticky” in that it recurs for previously supported students, while its measurable learning benefits in our data are largely session-bound. We discuss implications for designing real-time analytics that track attention coverage and highlight under-visited students to support a more equitable and effective allocation of teacher attention.
Qiao Jin 0002, Conrad Borchers, Ashish Gurung, Sean Jackson, Sameeksha Agarwal, Yichen Andy Yu, Pragati Maheshwary, Vincent Aleven
LAK2
2026 Measuring the Impact of Student Gaming Behaviors on Learner Modeling
abstract
The expansion of large-scale online education platforms has yielded vast amounts of student interaction data for knowledge tracing (KT). KT models estimate students’ concept mastery from interaction data, but the models’ performance is sensitive to input data quality. Gaming behaviors, such as excessive hint use, may misrepresent students’ knowledge and undermine model reliability. However, systematic investigations of how different types of gaming behaviors affect KT remain scarce, and existing studies rely on costly manual analysis that does not capture behavioral diversity. In this study, we conceptualize gaming behaviors as a form of data poisoning, defined as the deliberate submission of incorrect or misleading interaction data to corrupt a model’s learning process. We design Data Poisoning Attacks (DPA) to simulate diverse gaming patterns and systematically evaluate their impact on KT model performance. Moreover, drawing on advances in DPA detection, we explore unsupervised approaches to enhance the generalizability of gaming behavior detection. We find that KT models performance tend to decrease especially for random guess behaviors. Our findings provide insights into the vulnerabilities of KT models and highlight the potential of adversarial methods for improving the robustness of learning analytics systems.
Qinyi Liu, Lin Li 0039, Valdemar Svábenský, Conrad Borchers, Mohammad Khalil
LAK4
2026 Fifteen Years of Learning Analytics Research: Topics, Trends, and Challenges
abstract
The learning analytics (LA) community has recently reached two important milestones: celebrating the 15th LAK conference and updating the 2011 definition of LA to reflect the 15 years of changes in the discipline. However, despite LA’s growth, little is known about how research topics, funding, and collaboration, as well as the relationships among them, have developed within the community over time. This study addressed this gap by analyzing all 936 full and short papers published at LAK over a 15-year period using unsupervised machine learning, natural language processing, and network analytics. The analysis revealed a stable core of prolific authors alongside high turnover of newcomers, systematic links between funding sources and research directions, and six enduring topical centers that remain globally shared but vary in prominence across countries. These six topical centers, which encompass LA research, are: self-regulated learning, dashboards and theory, social learning, automated feedback, multimodal analytics, and outcome prediction. Our findings highlight key challenges for the future: widening participation, reducing dependency on a narrow set of funders, and ensuring that emerging research trajectories remain responsive to educational practice and societal needs.
Valdemar Svábenský, Conrad Borchers, Elvin Fortuna, Elizabeth B. Cloude, Dragan Gasevic
LAK2
2026 Disagreement as Data: Reasoning Trace Analytics in Multi-Agent Systems
abstract
Learning analytics researchers often analyze qualitative student data such as coded annotations or interview transcripts to understand learning processes. With the rise of generative AI, fully automated and human–AI workflows have emerged as promising methods for analysis. However, methodological standards to guide such workflows remain limited. In this study, we propose that reasoning traces generated by large language model (LLM) agents, especially within multi-agent systems, constitute a novel and rich form of process data to enhance interpretive practices in qualitative coding. We apply cosine similarity to LLM reasoning traces to systematically detect, quantify, and interpret disagreements among agents, reframing disagreement as a meaningful analytic signal. Analyzing nearly 10,000 instances of agent pairs coding human tutoring dialog segments, we show that LLM agents’ semantic reasoning similarity robustly differentiates consensus from disagreement and correlates with human coding reliability. Qualitative analysis guided by this metric reveals nuanced instructional sub-functions within codes and opportunities for conceptual codebook refinement. By integrating quantitative similarity metrics with qualitative review, our method bears potential to improve and accelerate the process of establishing inter-rater reliability during coding by surfacing interpretive ambiguity, especially when LLMs collaborate with humans. We discuss how reasoning-trace disagreements represent a valuable new class of analytic signals advancing methodological rigor and interpretive depth in educational research.
Elham Tajik, Conrad Borchers, Bahar Shahrokhian, Sebastian Simon, Ali Keramati, Sonika Pal, Sreecharan Sankaranarayanan
LAK2
2026 Using Large Language Models to Detect Socially Shared Regulation of Collaborative Learning
abstract
The field of learning analytics has made notable strides in automating the detection of complex learning processes in multimodal data. However, most advancements have focused on individualized problem-solving instead of collaborative, open-ended problem-solving, which may offer both affordances (richer data) and challenges (low cohesion) to behavioral prediction. Here, we extend predictive models to automatically detect socially shared regulation of learning (SSRL) behaviors in collaborative computational modeling environments using embedding-based approaches. We leverage large language models (LLMs) as summarization tools to generate task-aware representations of student dialogue aligned with system logs. These summaries, combined with text-only embeddings, context-enriched embeddings, and log-derived features, were used to train predictive models. Results show that text-only embeddings often achieve stronger performance in detecting SSRL behaviors related to enactment or group dynamics (e.g., off-task behavior or requesting assistance). In contrast, contextual and multimodal features provide complementary benefits for constructs such as planning and reflection. Overall, our findings highlight the promise of embedding-based models for extending learning analytics by enabling scalable detection of SSRL behaviors, ultimately supporting real-time feedback and adaptive scaffolding in collaborative learning environments that teachers value.
Jiayi Zhang 0004, Conrad Borchers, Clayton Cohn, Namrata Srivastava, Caitlin Snyder, T. S. Ashwin, Naveeduddin Mohammed, Haley Noh, Gautam Biswas
LAK2
2026 Understanding Gaming the System by Analyzing Self-Regulated Learning in Think-Aloud Protocols
abstract
In digital learning systems, gaming the system refers to occasions when students attempt to succeed in an educational task by systematically taking advantage of system features rather than engaging meaningfully with the content. Often viewed as a form of behavioral disengagement, gaming the system is negatively associated with short- and long-term learning outcomes. However, little research has explored this phenomenon beyond its behavioral representation, leaving questions such as whether students are cognitively disengaged or whether they engage in different self-regulated learning (SRL) strategies when gaming largely unanswered. This study employs a mixed-methods approach to examine students’ cognitive engagement and SRL processes during gaming versus non-gaming periods, using utterance length and SRL codes inferred from think-aloud protocols collected while students interacted with an intelligent tutoring system for chemistry. We found that gaming does not simply reflect a lack of cognitive effort; during gaming, students often produced longer utterances, were more likely to engage in processing information and realizing errors, but less likely to engage in planning, and exhibited reactive rather than proactive self-regulatory strategies. These findings provide empirical evidence supporting the interpretation that gaming may represent a maladaptive form of SRL. With this understanding, future work can address gaming and its negative impacts by designing systems that target maladaptive self-regulation to promote better learning.
Jiayi Zhang 0004, Conrad Borchers, Canwen Wang, Leah Teffera, Bruce M. McLaren, Ryan Baker 0001
LAK2
2026 Push and Pull in Community College Cross-Enrollment: Remoteness, Articulation, and Student Mobility
abstract
Cross-enrollment across institutions can expand access to courses and support student progression. Still, little is known about how geographic constraints and institutional policies jointly shape cross-enrollment within community college (CC) systems. We adopt a push--pull framework: geographic remoteness constrains feasible cross-institution mobility, while credit mobility may attract enrollment expressed as articulation (CC-to-university: credit toward a four-year partner) and course equivalencies (CC-to-CC: equivalencies across the system). Using de-identified administrative records from a 12-institution community college system (100,547 students; 1,290,311 course enrollments), we quantify outgoing and incoming cross-enrollment and relate these patterns to institutional remoteness and credit mobility. We find that less remote colleges exhibit higher outgoing and incoming cross-enrollment than more remote colleges. Further, cross-enrolled students are more likely to take articulated courses, and institutions with higher equivalency ratios receive higher incoming cross-enrollment (8.62% vs. 6.70%). This association was slightly stronger at more remote colleges. This study demonstrates how analysis of complex college systems can surface factors shaping student mobility and inform the design of cross-enrollment and articulation policies in CC systems.
Conrad Borchers, Robin Schmucker, Zachary A. Pardos
L@S1
2026 Who Decides in AI-Mediated Learning? The Agency Allocation Framework
abstract
As AI-mediated learning systems increasingly shape how learners plan, make decisions, and progress through education, learner agency is becoming both more consequential and harder to conceptualize at scale. Existing research often treats agency as a proxy for engagement and self-regulation, leaving unclear who actually holds decision-making authority in large-scale, automated learning environments. This paper reframes learner agency as the allocation of decision authority across learners, educators, institutions, and AI systems. We introduce the Agency Allocation Framework (AAF) for analyzing how decisions are distributed, how choices are architected, what evidence supports them, and over what time horizons their consequences unfold. Drawing on a focused review of Learning at Scale literature and an illustrative tutoring-system example, we identify four recurring challenges for studying learner agency at scale: (1) conceptual ambiguity, (2) reliance on behavioral proxies, (3) trade-offs between efficiency and learner control, and (4) the redistribution of agency through AI-mediated systems. Rather than advocating more or less automation, the AAF supports systematic analysis of when AI scaffolds learners' capacity to act and when it substitutes for it. By making decision authority explicit, the framework provides researchers and designers with analytic tools for studying, comparing, and evaluating agency-preserving learning systems in increasingly automated educational contexts.
Conrad Borchers, Olga Viberg, René F. Kizilcec
L@S1
2026 Understanding Student Effort Using Response-Time Propensities During Problem Solving
Conrad Borchers, Lijin Zhang, Tomohiro Nagashima, Benjamin W. Domingue
L@S1
2026 Student, Course Design, or Context? Studying the Determinants of Academic Procrastination at Scale
Jinwon Kim, Qiujie Li, Conrad Borchers, Zilu Jiang, Di Xu 0005
L@S3
2026 The Digital Divide in Generative AI: Evidence from Large Language Model Use in College Admissions Essays
abstract
Large language models (LLMs) have become popular writing tools among students and may expand access to high-quality feedback for students with less access to traditional writing support. At the same time, LLMs may standardize student voice or invite overreliance. This study examines how adoption of LLM-assisted writing varies across socioeconomic groups and how it relates to outcomes in a high-stakes context: U.S. college admissions. We analyze a de-identified longitudinal dataset of applications to a selective university from 2020 to 2024 (N = 81,663). Estimating LLM use using a distribution-based detector trained on synthetic and historical essays, we tracked how student writing changed as LLM use proliferated, how adoption differed by socioeconomic status (SES), and whether potential benefits translated equitably into admissions outcomes. Using fee-waiver status as a proxy for SES, we observe post-2023 convergence in surface-level linguistic features, with the largest changes among lower SES and rejected applicants. Estimated LLM use rose sharply in 2024 across all groups, with disproportionately larger increases among lower SES applicants, consistent with the hypothesis that LLMs substitute for scarce writing support. However, increased estimated LLM use was associated with larger declines in predicted admission probability for lower SES applicants than for higher SES applicants, even after controlling for academic credentials and stylometric features. These findings raise concerns about equity and the validity of essay-based evaluation in an era of AI-assisted writing and provide the first large-scale longitudinal evidence linking LLM adoption, linguistic change, and evaluative outcomes in college admissions.
Jinsook Lee, Conrad Borchers, A. J. Alvero, Thorsten Joachims, René F. Kizilcec
L@S2
2025 Engagement and Learning Benefits of Goal Setting with Rewards in Human-AI Tutoring
Conrad Borchers, Alex Houk, Vincent Aleven, Kenneth R. Koedinger
AIED (4)1
2025 Involving Parents in Tutoring Systems to Increase Content Confidence: A Design Probe Study
Conrad Borchers, Ha Tien Nguyen, Paulo Carvalho 0004, Kenneth R. Koedinger, Vincent Aleven
AIED (6)1
2025 Student Perceptions of Adaptive Goal Setting Recommendations: A Design Prototyping Study
Conrad Borchers, Cindy Peng, Qianru Lyu, Paulo Carvalho 0004, Kenneth R. Koedinger, Vincent Aleven
AIED (5)1
2025 Can Large Language Models Match Tutoring System Adaptivity? A Benchmarking Study
Conrad Borchers, Tianze Shou
AIED (2)1
2025 Comparing a Human's and a Multi-Agent System's Thematic Analysis: Assessing Qualitative Coding Consistency
Sebastian Simon, Sreecharan Sankaranarayanan, Elham Tajik, Conrad Borchers, Bahar Shahrokhian, Francesco Balzan, Sebastian Strauss, Sree Aurovindh Viswanathan, Amine Hatun Atas, Mia Carapina, Berkan Celik
AIED (3)4
2025 Improving Open-Response Assessment with LearnLM
Danielle R. Thomas, Conrad Borchers, Shambhavi Bhushan, Sanjit Kakarla, Alex Houk, Ralph Abboud, Shivang Gupta, Erin Gatz, Kenneth R. Koedinger
AIED (5)2
2025 Designing the Course Load Analytics Platform
Conrad Borchers, Shreya K. Sheel, Anirudh Pai, Sher Shah, Zachary A. Pardos
EC-TEL (2)1
2025 How Expertise Levels Shape Preferences and Reflection Needs: Towards AI Reflection Systems for Teacher Empowerment
Ann-Christin Falhs, Conrad Borchers, Vanessa Echeverría, Kexin Bella Yang, Nikol Rummel, Vincent Aleven
EC-TEL (2)2
2025 Error Classification in Stoichiometry Tutoring Systems with Different Levels of Scaffolding: Comparing Rule-Based Classification and Machine Learning
Hendrik Fleischer, Conrad Borchers, Sascha Schanze, Vincent Aleven
EC-TEL (2)2
2025 Leveraging LLMs to Assess Tutor Moves in Real-Life Dialogues: A Feasibility Study
Danielle R. Thomas, Conrad Borchers, Jionghao Lin, Sanjit Kakarla, Shambhavi Bhushan, Erin Gatz, Shivang Gupta, Ralph Abboud, Kenneth R. Koedinger
EC-TEL (2)2
2025 Toward Sufficient Statistical Power in Algorithmic Bias Assessment: A Test for ABROCA
Conrad Borchers
EDM1
2025 Does Student Learning Rate Depend on Feedback Type and Prior Knowledge?
Hendrik Fleischer, Arne Noglik, Conrad Borchers, Sascha Schanze
EDM3
2025 Starting Seatwork Earlier as a Valid Measure of Student Engagement
Ashish Gurung, Jionghao Lin, Zhongtian Huang, Conrad Borchers, Ryan Baker 0001, Vincent Aleven, Kenneth R. Koedinger
EDM4
2025 Who to Help? A Time-Slice Analysis of K-12 Teachers' Decisions in Classes with AI-Supported Tutoring
Qiao Jin 0002, Conrad Borchers, Stephen Fancsali, Vincent Aleven
EDM2
2025 ABROCA Distributions For Algorithmic Bias Assessment: Considerations Around Interpretation
abstract
Algorithmic bias continues to be a key concern of learning analytics.We study the statistical properties of the Absolute Between-ROC Area (ABROCA) metric.This fairness measure quantifies grouplevel differences in classifier performance through the absolute difference in ROC curves.ABROCA is particularly useful for detecting nuanced performance differences even when overall Area Under the ROC Curve (AUC) values are similar.We sample ABROCA under various conditions, including varying AUC differences and class distributions.We find that ABROCA distributions exhibit high skewness dependent on sample sizes, AUC differences, and class imbalance.When assessing whether a classifier is biased, this skewness inflates ABROCA values by chance, even when data is drawn (by simulation) from populations with equivalent ROC curves.These findings suggest that ABROCA requires careful interpretation given its distributional properties, especially when used to assess the degree of bias and when classes are imbalanced.
Conrad Borchers, Ryan Baker 0001
LAK1
2025 How Learner Control and Explainable Learning Analytics About Skill Mastery Shape Student Desires to Finish and Avoid Loss in Tutored Practice
abstract
Personalized problem selection enhances student practice in tutoring systems.Prior research has focused on transparent problem selection that supports learner control but rarely engages learners in selecting practice materials.We explored how different levels of control (i.e., full AI control, shared control, and full learner control), combined with showing learning analytics on skill mastery and visual what-if explanations, can support students in practice contexts requiring high degrees of self-regulation, such as homework.Semistructured interviews with six middle school students revealed three key insights: (1) participants highly valued learner control for an enhanced learning experience and better self-regulation, especially because most wanted to avoid losses in skill mastery;(2) only seeing their skill mastery estimates often made participants base problem selection on their weaknesses; and (3) what-if explanations stimulated participants to focus more on their strengths and improve skills until they were mastered.These findings show how explainable learning analytics could shape students' selection strategies when they have control over what to practice.They suggest promising avenues for helping students learn to regulate their effort, motivation, and goals during practice with tutoring systems.
Conrad Borchers, Jeroen Ooge, Cindy Peng, Vincent Aleven
LAK1
2025 Does the Doer Effect Generalize To Non-WEIRD Populations? Toward Analytics in Radio and Phone-Based Learning
abstract
The Doer Effect states that completing more active learning activities, like practice questions, is more strongly related to positive learning outcomes than passive learning activities, like reading, watching, or listening to course materials. Although broad, most evidence has emerged from practice with tutoring systems in Western, Industrialized, Rich, Educated, and Democratic (WEIRD) populations in North America and Europe. Does the Doer Effect generalize beyond WEIRD populations, where learners may practice in remote locales through different technologies? Through learning analytics, we provide evidence from N = 234 Ugandan students answering multiple-choice questions via phones and listening to lectures via community radio. Our findings support the hypothesis that active learning is more associated with learning outcomes than passive learning. We find this relationship is weaker for learners with higher prior educational attainment. Our findings motivate further study of the Doer Effect in diverse populations. We offer considerations for future research in designing and evaluating contextually relevant active and passive learning opportunities including leveraging familiar technology, increasing the number of practice opportunities, and aligning multiple data sources.
Darren Butler, Conrad Borchers, Michael W. Asher, Yongmin Lee, Sonya Karnataki, Sameeksha Dangi, Samyukta Athreya, John C. Stamper, Amy Ogan, Paulo Carvalho 0004
LAK2
2025 Evaluating the Impact of Data Augmentation on Predictive Model Performance
abstract
In supervised machine learning (SML) research, large training datasets are essential for valid results. However, obtaining primary data in learning analytics (LA) is challenging. Data augmentation can address this by expanding and diversifying data, though its use in LA remains underexplored. This paper systematically compares data augmentation techniques and their impact on prediction performance in a typical LA task: prediction of academic outcomes. Augmentation is demonstrated on four SML models, which we successfully replicated from a previous LAK study based on AUC values. Among 21 augmentation techniques, SMOTE-ENN sampling performed the best, improving the average AUC by 0.01 and approximately halving the training time compared to the baseline models. In addition, we compared 99 combinations of chaining 21 techniques, and found minor, although statistically significant, improvements across models when adding noise to SMOTE-ENN (+0.014). Notably, some augmentation techniques significantly lowered predictive performance or increased performance fluctuation related to random chance. This paper's contribution is twofold. Primarily, our empirical findings show that sampling techniques provide the most statistically reliable performance improvements for LA applications of SML, and are computationally more efficient than deep generation methods with complex hyperparameter settings. Second, the LA community may benefit from validating a recent study through independent replication.
Valdemar Svábenský, Conrad Borchers, Elizabeth B. Cloude, Atsushi Shimada 0001
LAK2
2025 Does Multiple Choice Have a Future in the Age of Generative AI? A Posttest-only RCT
abstract
The role of multiple-choice questions (MCQs) as effective learning tools has been debated in past research. While MCQs are widely used due to their ease in grading, open response questions are increasingly used for instruction, given advances in large language models (LLMs) for automated grading. This study evaluates MCQs effectiveness relative to open-response questions, both individually and in combination, on learning. These activities are embedded within six tutor lessons on advocacy. Using a posttest-only randomized control design, we compare the performance of 234 tutors (790 lesson completions) across three conditions: MCQ only, open response only, and a combination of both. We find no significant learning differences across conditions at posttest, but tutors in the MCQ condition took significantly less time to complete instruction. These findings suggest that MCQs are as effective, and more efficient, than open response tasks for learning when practice time is limited. To further enhance efficiency, we autograded open responses using GPT-4o and GPT-4-turbo. GPT models demonstrate proficiency for purposes of low-stakes assessment, though further research is needed for broader use. This study contributes a dataset of lesson log data, human annotation rubrics, and LLM prompts to promote transparency and reproducibility.
Danielle R. Thomas, Conrad Borchers, Sanjit Kakarla, Jionghao Lin, Shambhavi Bhushan, Boyuan Guo, Erin Gatz, Kenneth R. Koedinger
LAK2
2025 Do Tutors Learn from Equity Training and Can Generative AI Assess It?
abstract
Equity is a core concern of learning analytics. However, applications that teach and assess equity skills, particularly at scale are lacking, often due to barriers in evaluating language. Advances in generative AI via large language models (LLMs) are being used in a wide range of applications, with this present work assessing its use in the equity domain. We evaluate tutor performance within an online lesson on enhancing tutors' skills when responding to students in potentially inequitable situations. We apply a mixed-method approach to analyze the performance of 81 undergraduate remote tutors. We find marginally significant learning gains with increases in tutors' self-reported confidence in their knowledge in responding to middle school students experiencing possible inequities from pretest to posttest. Both GPT-4o and GPT-4-turbo demonstrate proficiency in assessing tutors ability to predict and explain the best approach. Balancing performance, efficiency, and cost, we determine that few-shot learning using GPT-4o is the preferred model. This work makes available a dataset of lesson log data, tutor responses, rubrics for human annotation, and generative AI prompts. Future work involves leveling the difficulty among scenarios and enhancing LLM prompts for large-scale grading and assessment.
Danielle R. Thomas, Conrad Borchers, Sanjit Kakarla, Jionghao Lin, Shambhavi Bhushan, Boyuan Guo, Erin Gatz, Kenneth R. Koedinger
LAK2
2025 Combining Large Language Models with Tutoring System Intelligence: A Case Study in Caregiver Homework Support
abstract
Caregivers (i.e., parents and members of a child's caring community) are underappreciated stakeholders in learning analytics. Although caregiver involvement can enhance student academic outcomes, many obstacles hinder involvement, most notably knowledge gaps with respect to modern school curricula. An emerging topic of interest in learning analytics is hybrid tutoring, which includes instructional and motivational support. Caregivers assert similar roles in homework, yet it is unknown how learning analytics can support them. Our past work with caregivers suggested that conversational support is a promising method of providing caregivers with the guidance needed to effectively support student learning. We developed a system that provides instructional support to caregivers through conversational recommendations generated by a Large Language Model (LLM). Addressing known instructional limitations of LLMs, we use instructional intelligence from tutoring systems while conducting prompt engineering experiments with the open-source Llama 3 LLM. This LLM generated message recommendations for caregivers supporting their child's math practice via chat. Few-shot prompting and combining real-time problem-solving context from tutoring systems with examples of tutoring practices yielded desirable message recommendations. These recommendations were evaluated with ten middle school caregivers, who valued recommendations facilitating content-level support and student metacognition through self-explanation. We contribute insights into how tutoring systems can best be merged with LLMs to support hybrid tutoring settings through conversational assistance, facilitating effective caregiver involvement in tutoring systems.
Devika Venugopalan, Ziwen Yan, Conrad Borchers, Jionghao Lin, Vincent Aleven
LAK3
2025 VTutor for High-Impact Tutoring at Scale: Managing Engagement and Real-Time Multi-Screen Monitoring with P2P Connections
Eason Chen, Aprille J. Xi, Chenyu Lin, Conrad Borchers, Shivang Gupta, Jionghao Lin, Kenneth R. Koedinger
L@S5
2025 Demo of VTutor for High-Impact Tutoring at Scale: A Real-Time Multi-Screen Tutor Support System with P2P Connections
abstract
published_or_final_version
Eason Chen, Aprille Xi, Chenyu Lin, Conrad Borchers, Shivang Gupta, Jionghao Lin, Kenneth R. Koedinger
L@S5
2024 Leveraging Multimodal Classroom Data for Teacher Reflection: Teachers' Preferences, Practices, and Privacy Considerations
Kexin Bella Yang, Conrad Borchers, Ann-Christin Falhs, Vanessa Echeverría, Shamya Karumbaiah, Nikol Rummel, Vincent Aleven
EC-TEL (1)2
2024 Using Large Language Models to Detect Self-Regulated Learning in Think-Aloud Protocols
Jiayi Zhang 0004, Conrad Borchers, Vincent Aleven, Ryan Baker 0001
EDM2
2024 Are You an Early Dropper or Late Shopper? Mining Enrollment Transaction Data to Study Procrastination in Higher Education
Conrad Borchers, Yinuo Xu, Zachary A. Pardos
EDM1
2024 Combining Dialog Acts and Skill Modeling: What Chat Interactions Enhance Learning Rates During AI-Supported Peer Tutoring?
Conrad Borchers, Jionghao Lin, Nikol Rummel, Kenneth R. Koedinger, Vincent Aleven
EDM1
2024 Using Think-Aloud Data to Understand Relations between Self-Regulation Cycle Characteristics and Student Performance in Intelligent Tutoring Systems
abstract
Numerous studies demonstrate the importance of self-regulation during learning by problem-solving. Recent work in learning analytics has largely examined students’ use of SRL concerning overall learning gains. Limited research has related SRL to in-the-moment performance differences among learners. The present study investigates SRL behaviors in relationship to learners’ moment-by-moment performance while working with intelligent tutoring systems for stoichiometry chemistry. We demonstrate the feasibility of labeling SRL behaviors based on AI-generated think-aloud transcripts, identifying the presence or absence of four SRL categories (processing information, planning, enacting, and realizing errors) in each utterance. Using the SRL codes, we conducted regression analyses to examine how the use of SRL in terms of presence, frequency, cyclical characteristics, and recency relate to student performance on subsequent steps in multi-step problems. A model considering students’ SRL cycle characteristics outperformed a model only using in-the-moment SRL assessment. In line with theoretical predictions, students’ actions during earlier, process-heavy stages of SRL cycles exhibited lower moment-by-moment correctness during problem-solving than later SRL cycle stages. We discuss system re-design opportunities to add SRL support during stages of processing and paths forward for using machine learning to speed research depending on the assessment of SRL based on transcription of think-aloud data.
Conrad Borchers, Jiayi Zhang 0004, Ryan Baker 0001, Vincent Aleven
LAK1
2024 Revealing Networks: Understanding Effective Teacher Practices in AI-Supported Classrooms using Transmodal Ordered Network Analysis
abstract
Learning analytics research increasingly studies classroom learning with AI-based systems through rich contextual data from outside these systems, especially student-teacher interactions. One key challenge in leveraging such data is generating meaningful insights into effective teacher practices. Quantitative ethnography bears the potential to close this gap by combining multimodal data streams into networks of co-occurring behavior that drive insight into favorable learning conditions. The present study uses transmodal ordered network analysis to understand effective teacher practices in relationship to traditional metrics of in-system learning in a mathematics classroom working with AI tutors. Incorporating teacher practices captured by position tracking and human observation codes into modeling significantly improved the inference of how efficiently students improved in the AI tutor beyond a model with tutor log data features only. Comparing teacher practices by student learning rates, we find that students with low learning rates exhibited more hint use after monitoring. However, after an extended visit, students with low learning rates showed learning behavior similar to their high learning rate peers, achieving repeated correct attempts in the tutor. Observation notes suggest conceptual and procedural support differences can help explain visit effectiveness. Taken together, offering early conceptual support to students with low learning rates could make classroom practice with AI tutors more effective. This study advances the scientific understanding of effective teacher practice in classrooms learning with AI tutors and methodologies to make such practices visible.
Conrad Borchers, Yeyu Wang, Shamya Karumbaiah, Muhammad Ashiq, David Williamson Shaffer, Vincent Aleven
LAK1
2024 Gaining Insights into Group-Level Course Difficulty via Differential Course Functioning
abstract
Curriculum Analytics (CA) studies curriculum structure and student data to ensure the quality of educational programs. One desirable property of courses within curricula is that they are not unexpectedly more difficult for students of different backgrounds. While prior work points to likely variations in course difficulty across student groups, robust methodologies for capturing such variations are scarce, and existing approaches do not adequately decouple course-specific difficulty from students' general performance levels. The present study introduces Differential Course Functioning (DCF) as an Item Response Theory (IRT)-based CA methodology. DCF controls for student performance levels and examines whether significant differences exist in how distinct student groups succeed in a given course. Leveraging data from over 20,000 students at a large public university, we demonstrate DCF's ability to detect inequities in undergraduate course difficulty across student groups described by grade achievement. We compare major pairs with high co-enrollment and transfer students to their non-transfer peers. For the former, our findings suggest a link between DCF effect sizes and the alignment of course content to student home department motivating interventions targeted towards improving course preparedness. For the latter, results suggest minor variations in course-specific difficulty between transfer and non-transfer students. While this is desirable, it also suggests that interventions targeted toward mitigating grade achievement gaps in transfer students should encompass comprehensive support beyond enhancing preparedness for individual courses. By providing more nuanced and equitable assessments of academic performance and difficulties experienced by diverse student populations, DCF could support policymakers, course articulation officers, and student advisors.
Frederik Baucks, Robin Schmucker, Conrad Borchers, Zachary A. Pardos, Laurenz Wiskott
L@S3
2023 A Spatiotemporal Analysis of Teacher Practices in Supporting Student Learning and Engagement in an AI-Enabled Classroom
Shamya Karumbaiah, Conrad Borchers, Tianze Shou, Ann-Christin Falhs, Pinyang Liu, Tomohiro Nagashima, Nikol Rummel, Vincent Aleven
AIED2
2023 What Makes Problem-Solving Practice Effective? Comparing Paper and AI Tutoring
abstract
Abstract In numerous studies, intelligent tutoring systems (ITSs) have proven effective in helping students learn mathematics. Prior work posits that their effectiveness derives from efficiently providing eventually-correct practice opportunities. Yet, there is little empirical evidence on how learning processes with ITSs compare to other forms of instruction. The current study compares problem-solving with an ITS versus solving the same problems on paper. We analyze the learning process and pre-post gain data from N = 97 middle school students practicing linear graphs in three curricular units. We find that (i) working with the ITS, students had more than twice the number of eventually-correct practice opportunities than when working on paper and (ii) omission errors on paper were associated with lower learning gains. Yet, contrary to our hypothesis, tutor practice did not yield greater learning gains, with tutor and paper comparing differently across curricular units. These findings align with tutoring allowing students to grapple with challenging steps through tutor assistance but not with eventually-correct opportunities driving learning gains. Gaming-the-system, lack of transfer to an unfamiliar test format, potentially ineffective tutor design, and learning affordances of paper can help explain this gap. This study provides first-of-its-kind quantitative evidence that ITSs yield more learning opportunities than equivalent paper-and-pencil practice and reveals that the relation between opportunities and learning gains emerges only when the instruction is effective.
Conrad Borchers, Paulo Carvalho 0004, Meng Xia 0002, Pinyang Liu, Kenneth R. Koedinger, Vincent Aleven
EC-TEL1
2023 Timing Matters: Inferring Educational Twitter Community Switching from Membership Characteristics
Conrad Borchers, Lennart Klein, Hayden Johnson, Christian Fischer 0007
EDM1
2023 Optimizing Parameters for Accurate Position Data Mining in Diverse Classrooms Layouts
Tianze Shou, Conrad Borchers, Shamya Karumbaiah, Vincent Aleven
EDM2
2023 Insights into undergraduate pathways using course load analytics
abstract
Course load analytics (CLA) inferred from LMS and enrollment features can offer a more accurate representation of course workload to students than credit hours and potentially aid in their course selection decisions. In this study, we produce and evaluate the first machine-learned predictions of student course load ratings and generalize our model to the full 10,000 course catalog of a large public university. We then retrospectively analyze longitudinal differences in the semester load of student course selections throughout their degree. CLA by semester shows that a student’s first semester at the university is among their highest load semesters, as opposed to a credit hour-based analysis, which would indicate it is among their lowest. Investigating what role predicted course load may play in program retention, we find that students who maintain a semester load that is low as measured by credit hours but high as measured by CLA are more likely to leave their program of study. This discrepancy in course load is particularly pertinent in STEM and associated with high prerequisite courses. Our findings have implications for academic advising, institutional handling of the freshman experience, and student-facing analytics to help students better plan, anticipate, and prepare for their selected courses.
Conrad Borchers, Zachary A. Pardos
LAK1
2021 To Scale or Not to Scale: Comparing Popular Sentiment Analysis Dictionaries on Educational Twitter Data
Conrad Borchers, Joshua Rosenberg 0001, Benjamin Gibbons, Macy Alana Burchfield, Christian Fischer 0007
EDM1
2021 Are Violations of Student Privacy "Quick and Easy"? Investigating the Privacy of Students' Images and Names in the Context of K-12 Educational Institution's Posts on Facebook
Macy Burchfield, Joshua Rosenberg 0001, Conrad Borchers, Tayla Thomas, Benjamin Gibbons, Christian Fischer 0007
EDM3