EDBT 2026 Demo / reviewers in the wild / expert
Paulo Carvalho 0004
dblp:176/2876 · also Paulo F. Carvalho
· DBLP profile ↗
51ranked-venue papers
14as first author
28since 2021 · last 2026
0000-0002-0449-3733ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 47 · 14 first-author · 24 since 2021Artificial intelligence and machine learning · 30 · 13 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 17 · 14 since 2021Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating a Data-Driven Redesign Process for Intelligent Tutoring Systems
Qianru Lyu, Conrad Borchers, Meng Xia 0002, Karen Xiao, Paulo Carvalho 0004, Kenneth R. Koedinger, Vincent Aleven |
AIED (3) | 5 |
| 2026 | Will They Try Again? A Large-Scale RCT on Scaffolds that Support Persistence in an Intelligent Tutoring SystemabstractPersistence after failure is critical for learning—but when students make mistakes in intelligent tutoring systems, they often choose not to try again. How can digital platforms encourage students to persist at these moments? We conducted a randomized controlled trial in an intelligent tutoring system for math and science, involving 164,532 students (Grades 8-12) who completed 17 million practice problems. We tested two scalable interventions: a brief persuasive prompt encouraging students to try again, and a visual default nudge that highlighted the retry option. Both interventions increased persistence after failure, and when combined, their effects were additive—suggesting they operate through distinct psychological mechanisms. The nudge had a much larger immediate effect, but the prompt showed proportionally greater spillover to untreated problems. These findings advance theories of persuasive design, demonstrating that implicit, interface-level nudges and explicit motivational prompts can be combined to avoid redundancy while amplifying impact. Michael W. Asher, Yumou Wei, Adam Daniel Reynolds, Amy Ogan, Paulo Carvalho 0004 |
CHI | 5 |
| 2026 | Inclusive Mobile Learning: How Technology-Enabled Language Choice Supports Multilingual StudentsabstractMost learners worldwide are multilingual, yet implementing multilingual education remains challenging in practice. EdTech offers an opportunity to bridge this gap and expand access for linguistically diverse learners. We conducted a quasi-experiment in Uganda with 2,931 participants enrolled in a non-formal radio- and mobile-based engineering course, where learners self-selected instruction in Leb Lango (a local language), English, or a Hybrid option combining both languages. The Leb Lango version of the course was used disproportionately by learners from rural areas, those with less formal education, and those with lower prior knowledge, broadening participation among disadvantaged learners. Moreover, the availability of Leb Lango instruction was associated with higher active participation, even among learners who registered for English instruction. Although Leb Lango learners began with lower performance, they demonstrated faster learning gains and achieved comparable final examination outcomes to English and Hybrid learners. These results suggest that providing local language options to learners is an effective way to make EdTech more accessible. Phenyo Phemelo Moletsane, Michael W. Asher, Christine Kwon, Paulo Carvalho 0004, Amy Ogan |
CHI | 4 |
| 2026 | AI Knows Best? The Paradox of Expertise, AI-Reliance, and Performance in Educational Tutoring Decision-Making TasksabstractWe present an empirical study examining how experienced tutors (experts) and non-tutors (novices) evaluate the correctness of tutor praise responses under different AI-assisted decision-support interfaces and explanation styles. We examine human-AI reliance patterns by decomposing interaction errors into over-reliance (accepting incorrect AI suggestions) and under-reliance (rejecting correct AI suggestions), together with time cost as a process-level indicator. Across conditions, human-AI collaboration improved accuracy compared to humans working alone, but consistently underperformed an AI-only baseline, indicating that human judgment introduced additional errors even when assisted by a highly accurate model. Novices benefited more from AI support since they tend to follow AI suggestions, whereas experts frequently overrode correct AI advice, resulting in lower overall performance, revealing a paradox of expertise in educational decision-making. We further compare two explanation modalities: textual reasoning and inline highlighting. Textual reasoning reduced under-reliance when the AI was correct but increased over-reliance when the AI was wrong, while inline highlighting exerted minimal influence on either behavior. Notably, neither explanation modality improved accuracy, and both increased time costs. As a contribution to learning analytics, we demonstrate how reliance patterns (over-reliance and under-reliance) and time cost function as process-level indicators that reveal how users integrate, or fail to integrate, AI recommendations. Our findings underscore the need for adaptive, trust-calibrated explanation strategies in tutor-facing decision support systems that balance accuracy, efficiency, and accountability in human-AI collaboration. Eason Chen, Jeffrey Li, Scarlett Huang, Jionghao Lin, Paulo Carvalho 0004, Kenneth R. Koedinger |
LAK | 6 |
| 2026 | Active Learning Beyond Borders: PEOE Enhancement of Explanatory Understanding in Japanese Undergraduates
Yugo Hayashi, Shigen Shimojyo, Paulo Carvalho 0004, Kenneth R. Koedinger |
LAK | 3 |
| 2026 | Generate-Then-Validate: A Novel Question Generation Approach Using Small Language ModelsabstractWe explore the use of small language models (SLMs) for automatic question generation as a complement to the prevalent use of their large counterparts in learning analytics research. We present a novel question generation pipeline that leverages both the text generation and the probabilistic reasoning abilities of SLMs to generate high-quality questions. Adopting a “generate-then-validate” strategy, our pipeline first performs expansive generation to create an abundance of candidate questions and refine them through selective validation based on novel probabilistic reasoning. We conducted two evaluation studies, one with seven human experts and the other with a large language model (LLM), to assess the quality of the generated questions. Most judges (humans or LLMs) agreed that the generated questions had clear answers and generally aligned well with the intended learning objectives. Our findings suggest that an SLM can effectively generate high-quality questions when guided by a well-designed pipeline that leverages its strengths. Yumou Wei, John C. Stamper, Paulo Carvalho 0004 |
LAK | 3 |
| 2026 | Benefit or Bottleneck? Assessing the Impact of Structured Reflection on Learning from AI-Driven Explanatory FeedbackabstractAs AI tutors become increasingly capable of delivering rich, personalized feedback at scale, a key challenge remains: novice learners often struggle to process detailed explanations on their own. Structured reflection, grounded in decades of self-explanation research, is a theoretically compelling solution. By helping learners parse feedback and prompting them to actively interpret it, reflection activities are designed to reduce cognitive overload and deepen understanding. But does adding reflection to already rich AI-generated feedback actually help, or does it simply add friction? We tested this in a randomized experiment comparing Python practice with AI-generated, personalized feedback to ''reflective practice,'' which paired identical feedback with structured self-explanation prompts. Contrary to our predictions, reflection never improved performance on any measure. Instead, it proved to be a temporal bottleneck: it doubled time spent on feedback and reduced practice volume by 40%, without making each learning opportunity more effective. Learners who cycled through more practice-and-feedback iterations outperformed reflective learners at the end of the session and maintained a small, nonsignificant advantage on transfer. Notably, reflection did not provide the scaffolding benefit we predicted for novices—and when individual differences did emerge, they favored higher-volume practice for more knowledgeable learners. Both practice conditions also substantially outperformed a high-quality video baseline (d = 0.66–0.93), replicating benefits of active practice with feedback over passive instruction. These findings provide initial evidence that when feedback is already elaborated and personalized, self-explanation activities may add little value, particularly when their time-related costs are considered. As AI-generated feedback reaches learners at scale, these findings underscore the necessity of empirically validating pedagogical scaffolds—even those with strong theoretical support—before deploying them broadly. Michael W. Asher, Gillian Gold, Paulo Carvalho 0004 |
L@S | 3 |
| 2026 | Can Multilingual Environments Promote Scalable EdTech? Evidence from a Randomized Controlled Trial
Phenyo Phemelo Moletsane, Christine Kwon, John C. Stamper, Amy Ogan, Paulo Carvalho 0004 |
L@S | 5 |
| 2026 | Investigating the Efficacy of Mastery-Based Tests in Fostering Effective Self-Regulated Learning Behaviors in CS1 CoursesabstractGiven the cumulative nature of computer science, success in introductory computing (CS1) courses requires students to not only learn the material but also develop effective self-regulated learning (SRL) habits. While theories of SRL emphasize planning, performance, and self-reflection as essential phases of effective learning, there is limited evidence on how to help learners put these phases into practice. In this context, Mastery-Based Tests (MBT), which allow students to retake assessments after receiving feedback, have shown promise for improving learning outcomes. However, prior work in computer science is largely observational and does not directly test MBT's impact on SRL behaviors. This paper presents a pilot study (N = 6) exploring this relationship in CS1. Using a between-subjects design, we observed that learners who first completed an MBT achieved higher post-test scores, demonstrated higher metacognitive accuracy, and self-reported more productive SRL behaviors. These patterns suggest that MBTs warrant further investigation as a viable scaffold for fostering self-regulation in CS1. Joyce Gill, Michael W. Asher, Paulo Carvalho 0004 |
SIGCSE (2) | 3 |
| 2025 | Involving Parents in Tutoring Systems to Increase Content Confidence: A Design Probe Study
Conrad Borchers, Ha Tien Nguyen, Paulo Carvalho 0004, Kenneth R. Koedinger, Vincent Aleven |
AIED (6) | 3 |
| 2025 | Student Perceptions of Adaptive Goal Setting Recommendations: A Design Prototyping Study
Conrad Borchers, Cindy Peng, Qianru Lyu, Paulo Carvalho 0004, Kenneth R. Koedinger, Vincent Aleven |
AIED (5) | 4 |
| 2025 | Identifying Effective Praise in Tutoring: Large Language Models with Transparent Explanations
Eason Chen, Jeffrey Li, Scarlett Huang, Jionghao Lin, Paulo Carvalho 0004, Kenneth R. Koedinger |
AIED (6) | 6 |
| 2025 | Adaptive Spaced Retrieval Practice in Algebra I: A Classroom-Based Study
Paulo Carvalho 0004 |
CogSci | 2 |
| 2025 | Optimizing Learning Efficiency: Balancing Spacing and Repetition Under Time Constraints
Veronica X. Yan, Faria Sana, Paulo Carvalho 0004 |
CogSci | 4 |
| 2025 | What's going on? Surprising difficulties in complex relational rule discovery
Julia J. Conti, Kenneth R. Koedinger, Paulo Carvalho 0004 |
CogSci | 3 |
| 2025 | To Honor or Dishonor Student Choices? The Impact of Self-Regulation on Instructional Methods and Learning Outcomes
Gillian Gold, Michael W. Asher, Paulo Carvalho 0004 |
CogSci | 3 |
| 2025 | 9th Educational Data Mining in Computer Science Education (CSEDM) Workshop
Bita Akram, Yang Shi 0004, Peter Brusilovsky, Thomas W. Price, Kenneth R. Koedinger, Paulo Carvalho 0004, Shan Zhang 0003, Andrew S. Lan, Juho Leinonen 0001 |
EDM | 6 |
| 2025 | KCluster: An LLM-based Clustering Approach to Knowledge Component Discovery
Yumou Wei, Paulo Carvalho 0004, John C. Stamper |
EDM | 2 |
| 2025 | Does the Doer Effect Generalize To Non-WEIRD Populations? Toward Analytics in Radio and Phone-Based LearningabstractThe Doer Effect states that completing more active learning activities, like practice questions, is more strongly related to positive learning outcomes than passive learning activities, like reading, watching, or listening to course materials. Although broad, most evidence has emerged from practice with tutoring systems in Western, Industrialized, Rich, Educated, and Democratic (WEIRD) populations in North America and Europe. Does the Doer Effect generalize beyond WEIRD populations, where learners may practice in remote locales through different technologies? Through learning analytics, we provide evidence from N = 234 Ugandan students answering multiple-choice questions via phones and listening to lectures via community radio. Our findings support the hypothesis that active learning is more associated with learning outcomes than passive learning. We find this relationship is weaker for learners with higher prior educational attainment. Our findings motivate further study of the Doer Effect in diverse populations. We offer considerations for future research in designing and evaluating contextually relevant active and passive learning opportunities including leveraging familiar technology, increasing the number of practice opportunities, and aligning multiple data sources. Darren Butler, Conrad Borchers, Michael W. Asher, Yongmin Lee, Sonya Karnataki, Sameeksha Dangi, Samyukta Athreya, John C. Stamper, Amy Ogan, Paulo Carvalho 0004 |
LAK | 10 |
| 2025 | Validating a New Approach for Measuring Student Engagement in Remote, Low-Infrastructure Learning EnvironmentsabstractExpanding access to education in rural African communities remains difficult, largely due to limited internet connectivity. Mobile learning courses delivered via radio and offline mobile phones offer a promising, scalable solution. However, it is challenging to track student engagement in these environments due to the absence of tools that monitor students' interactions with the radio. In this study, we investigate the potential of ''Prize Codes'' -- codes read aloud during broadcasts that students enter via text message -- to serve as a real-time measure of student engagement with mobile-learning broadcasts. Using data from a 2024 implementation of Yiya AirScience, a mobile engineering course in Uganda, we evaluate the validity of Prize Codes as an engagement metric. Specifically, we test whether Prize Code measures (1) demonstrate reliability, with students who enter correct codes in one lesson being more likely to do so in subsequent lessons; (2) demonstrate convergent validity with existing measures of engagement; and (3) demonstrate predictive validity, predicting learning outcomes in the course. Our findings suggest that Prize Codes are a reliable and valid measure of engagement. Prize-Code accuracy demonstrates strong internal consistency (alpha = .97) and moderate test-retest reliability (ICC = .44). The measure aligns closely with synchronous participation (87% agreement, Cohen's kappa = .50), indicating it captures similar engagement patterns. Importantly, students who consistently enter correct Prize Codes perform significantly better on assessments, with Prize Code engagement predicting final exam scores above and beyond other engagement metrics. After establishing the measure's validity, we use it to (1) characterize patterns of engagement with Yiya broadcasts, (2) investigate early engagement with the broadcasts as a predictor of course persistence, and (3) replicate findings about the benefits of learning by doing. This study suggests that Prize Codes can be a feasible, scalable approach for tracking real-time engagement in resource-limited mobile learning settings at scale. Michael W. Asher, Christine Kwon, John C. Stamper, Amy Ogan, Paulo Carvalho 0004 |
L@S | 5 |
| 2024 | Students Can Learn More Efficiently When Lectures Are Replaced with Practice Opportunities and Feedback
Michael W. Asher, Faria Sana, Kenneth R. Koedinger, Paulo Carvalho 0004 |
CogSci | 4 |
| 2024 | Beyond Accuracy: Embracing Meaningful Parameters in Educational Data Mining
Napol Rachatasumrit, Paulo Carvalho 0004, Kenneth R. Koedinger |
EDM | 2 |
| 2024 | Content Matters: A Computational Investigation into the Effectiveness of Retrieval Practice and Worked Examples (Extended Abstract)
Napol Rachatasumrit, Paulo Carvalho 0004, Sophie Li, Kenneth R. Koedinger |
IJCAI | 2 |
| 2024 | Beyond Repetition: The Role of Varied Questioning and Feedback in Knowledge GeneralizationabstractThis study examines the effects of question type and feedback on learning outcomes in a hybrid graduate-level course. By analyzing data from 32 students over 30,198 interactions, we assess the efficacy of unique versus repeated questions and the impact of feedback on student learning. The findings reveal students demonstrate significantly better knowledge generalization when encountering unique questions compared to repeated ones, even though they perform better with repeated opportunities. Moreover, we find that the timing of explanatory feedback is a more robust predictor of learning outcomes than the practice opportunities themselves. These insights suggest that educational practices and technological platforms should prioritize a variety of questions to enhance the learning process. The study also highlights the critical role of feedback; opportunities preceding feedback are less effective in enhancing learning. Gautam Yadav, Paulo Carvalho 0004, Elizabeth A. McLaughlin, Kenneth R. Koedinger |
L@S | 2 |
| 2023 | Content Matters: A Computational Investigation into the Effectiveness of Retrieval Practice and Worked Examples
Napol Rachatasumrit, Paulo Carvalho 0004, Sophie Li, Kenneth R. Koedinger |
AIED | 2 |
| 2023 | What Makes Problem-Solving Practice Effective? Comparing Paper and AI TutoringabstractAbstract In numerous studies, intelligent tutoring systems (ITSs) have proven effective in helping students learn mathematics. Prior work posits that their effectiveness derives from efficiently providing eventually-correct practice opportunities. Yet, there is little empirical evidence on how learning processes with ITSs compare to other forms of instruction. The current study compares problem-solving with an ITS versus solving the same problems on paper. We analyze the learning process and pre-post gain data from N = 97 middle school students practicing linear graphs in three curricular units. We find that (i) working with the ITS, students had more than twice the number of eventually-correct practice opportunities than when working on paper and (ii) omission errors on paper were associated with lower learning gains. Yet, contrary to our hypothesis, tutor practice did not yield greater learning gains, with tutor and paper comparing differently across curricular units. These findings align with tutoring allowing students to grapple with challenging steps through tutor assistance but not with eventually-correct opportunities driving learning gains. Gaming-the-system, lack of transfer to an unfamiliar test format, potentially ineffective tutor design, and learning affordances of paper can help explain this gap. This study provides first-of-its-kind quantitative evidence that ITSs yield more learning opportunities than equivalent paper-and-pencil practice and reveals that the relation between opportunities and learning gains emerges only when the instruction is effective. Conrad Borchers, Paulo Carvalho 0004, Meng Xia 0002, Pinyang Liu, Kenneth R. Koedinger, Vincent Aleven |
EC-TEL | 2 |
| 2022 | Educational Equity Through Combined Human-AI Personalization: A Propensity Matching Evaluation
Danielle R. Thomas, Cassandra Brentley, Carmen Thomas-Browne, J. Elizabeth Richey, Abdulmenaf Gul, Paulo Carvalho 0004, Lee G. Branstetter, Kenneth R. Koedinger |
AIED (1) | 6 |
| 2022 | Learning depends on knowledge: The benefits of retrieval practice vary for facts and skills
Paulo Carvalho 0004, Napol Rachatasumrit, Kenneth R. Koedinger |
CogSci | 1 |
| 2020 | Comprehensive Views of Math Learners: A Case for Modeling and Supporting Non-math Factors in Adaptive Math Software
J. Elizabeth Richey, Nikki G. Lobczowski, Paulo Carvalho 0004, Kenneth R. Koedinger |
AIED (1) | 3 |
| 2020 | The attentional demands of learning by doing: A developmental study
Karrie E. Godwin, Paulo Carvalho 0004 |
CogSci | 2 |
| 2019 | Square it up!: How to model step duration when predicting student performanceabstractIn this paper, we explore how we can model students' response times to predict student performance in Intelligent Tutoring Systems. Related research suggests that response time can provide information with respect to correctness. However, time is not consistently used when modeling students' performance. Here, we build on previous work that indicated that the relationship between response time and student performance is non-linear. Based on this concept, we compare three models: a standard Additive Factors Analysis Model (AFM), an AFM model enhanced with a linear step duration parameter and an AFM model enhanced with a quadratic, step duration parameter. The results of this comparison show that the AFM model that is enhanced with the quadratic step duration parameter outperforms the other models over four different datasets and for most of the metrics we used to evaluate the models in cross validation and prediction. Irene-Angelica Chounta, Paulo Carvalho 0004 |
LAK | 2 |
| 2019 | Comprehension Factor Analysis: Modeling student's reading behaviour: Accounting for reading practice in predicting students' learning in MOOCsabstractMassive Open Online Courses (MOOCs) often incorporate lecture-based learning along with lecture notes, textbooks, and videos to students. Moreover, MOOCs also incorporate practice activities and quizzes. Student learning in MOOCs can be tracked and improved using state-of-the-art student modeling. Currently, this means employing conventional student models that are constructed around Intelligent Tutoring Systems (ITS). Traditional ITS systems only utilize students performance interactions (quiz, problem-solving or practice activities). Therefore, text interactions are entirely ignored while modeling students performance in MOOCs using these cognitive models. In this work, we propose a Comprehension Factor Analysis model (CFM) for online courses, which integrates student reading interactions in student models to track and predict learning outcomes. Our model evaluation shows that CFM outperforms state-of-the-art models in predicting students' performance in a MOOC. These models can help better student-wise adaptation in the context of MOOCs. Khushboo Thaker, Paulo Carvalho 0004, Kenneth R. Koedinger |
LAK | 2 |
| 2018 | Understanding the Dynamics of Learning: The Case for Studying Interactions
Paulo Carvalho 0004 |
CogSci | 1 |
| 2018 | Not all Active Learning is Equal: Predicting and Explaining Improves Transfer Relative to Answering Practice Questions
Paulo Carvalho 0004, Kody Manke, Kenneth R. Koedinger |
CogSci | 1 |
| 2018 | Week-long practice matching 2D objects by shape improves 3D shape bias and accelerates children vocabulary growth
Paulo Carvalho 0004, Linda B. Smith |
CogSci | 1 |
| 2018 | Analyzing the relative learning benefits of completing required activities and optional readings in online courses
Paulo Carvalho 0004, Benjamin Motz 0002, Kenneth R. Koedinger |
EDM | 1 |
| 2017 | The most efficient sequence of study depends on the type of test
Paulo Carvalho 0004, Robert L. Goldstone |
CogSci | 1 |
| 2017 | Is there an explicit learning bias? Students beliefs, behaviors and learning outcomes
Paulo Carvalho 0004, Elizabeth A. McLaughlin, Kenneth R. Koedinger |
CogSci | 1 |
| 2016 | The sequence of study changes what is encoded during category learning
Paulo Carvalho 0004, Robert L. Goldstone |
CogSci | 1 |
| 2016 | Which Statistic Matters? Effects of Category Size and Distribution on Statistical Category Learning
Chi-hsin Chen, Paulo Carvalho 0004, Chen Yu 0001 |
CogSci | 2 |
| 2016 | Variability in category learning: The Effect of Context Change and Item Variation on Knowledge Generalization
Dustin Finch, Paulo Carvalho 0004, Robert L. Goldstone |
CogSci | 2 |
| 2015 | Effectiveness of Learner-Regulated Study Sequence: An in-vivo study in Introductory Psychology course
Paulo Carvalho 0004, David W. Braithwaite, Josh de Leeuw, Benjamin Motz 0002, Robert L. Goldstone |
CogSci | 1 |
| 2015 | Can You Repeat That? The Effect of Item Repetition on Interleaved and Blocked Study
Abigail Kost, Paulo Carvalho 0004, Robert L. Goldstone |
CogSci | 2 |
| 2014 | Effects of interleaved and blocked study in a 24 hour delayed transfer test
Paulo Carvalho 0004, Robert L. Goldstone |
CogSci | 1 |
| 2014 | Similarity-based Ordering of Instances for Efficient Concept Learning
Erik Weitnauer, Paulo Carvalho 0004, Robert L. Goldstone, Helge J. Ritter |
CogSci | 2 |
| 2013 | How to present exemplars of several categories? Interleave during active learning and block during passive learning
Paulo Carvalho 0004, Robert L. Goldstone |
CogSci | 1 |
| 2013 | An eyetracking study of children's relational thinking: The role of labels and sustained attention
Paulo Carvalho 0004, Catarina Vales, Caitlin M. Fausey, Linda B. Smith |
CogSci | 1 |
| 2013 | Grouping by Similarity Helps Concept Learning
Erik Weitnauer, Paulo Carvalho 0004, Robert L. Goldstone, Helge J. Ritter |
CogSci | 2 |
| 2012 | Category structure modulates interleaving and blocking advantage in inductive category acquisition
Paulo Carvalho 0004, Robert L. Goldstone |
CogSci | 1 |
| 2012 | Going to Extremes: The influence of unsupervised categories on the mental caricaturization of faces and asymmetries in perceptual discrimination
Andrew Hendrickson, Paulo Carvalho 0004, Robert L. Goldstone |
CogSci | 2 |
| 2011 | Sequential similarity and comparison effects in category learning
Paulo Carvalho 0004, Robert L. Goldstone |
CogSci | 1 |