EDBT 2026 Demo / reviewers in the wild / expert
Brent N. Reeves
dblp:36/317
· DBLP profile ↗
17ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0001-5781-1136ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 17 · 2 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Validated Scale Measuring Student Self-Efficacy for Programming with Generative AIabstractThe rise of generative artificial intelligence (GenAI) has sparked a rapid change in computing curricula and teaching approaches. GenAI coding tools can accurately complete assignments, answer test questions, and perform other tasks traditionally associated with learning programming, especially at the introductory level. Because GenAI is still so new, researchers investigating student usage of GenAI have used informal rubrics and questionnaires. To advance, the field needs validated instruments that measure student perception and use of GenAI. This paper presents the development and initial validation of an instrument to measure self-efficacy while using GenAI to learn programming. Self-efficacy is an important construct in education research because it robustly correlates with student success, across disciplines and ages, including undergraduate computing education. Computing education researchers have presented several validated self-efficacy instruments, most recently by Steinhorst et al. in 2020. Critically, this instrument was created before the rise of GenAI’s popularity in 2022. To complement this instrument, we created a GenAI scale similar in style to the Steinhorst self-efficacy instrument, consisting originally of 11 items and revised to 5 items. We report two important findings in this paper. First, we found strong support for the validity of the existing Steinhorst instrument in a new context, specifically an introductory programming course that fully integrates GenAI. Second, the new GenAI scale shows strong internal reliability, discriminant validity with items in the Steinhorst subscales, and criterion validity with students’ GenAI usage patterns. Based on statistical analysis and cognitive probing interviews, we argue for the validity of the five-item scale to measure students’ GenAI self-efficacy in the context of programming. James Prather, Lauren E. Margulieux, Yekaterina Kharitonova, Yonggao Yang, Brent N. Reeves, Paul Denny 0001, Jamie Gorson Benario, Ernest D. V. Holmes, Erin M. Spaulding, Gweneth Barbre, Musa Blake, Juho Leinonen 0001 |
ICER (1) | 5 |
| 2026 | Scaffolding Autocomplete: Improving Guidance for Learners using Generative Code SuggestionsabstractModern programming tools use generative AI (GenAI) to suggest code to the user as they type, interrupting their problem-solving behavior and undermining the development of their programming critical thinking skills. In this paper, we present a scaffolded programming exercise designed to support student differentiation between good and bad GenAI code suggestions based on negative expertise–that identifying why an answer is wrong is part of developing conceptual knowledge. We compare a version of the tool that showed one suggestion (correct or not), to a version that showed three suggestions (one of which was correct). We present results on performance and error rates as well as qualitative findings centered on Pintrich and DeGroot’s theory of self-regulation. Students reported that the single suggestion version better aligned with industry tools and presented a lower cognitive load. Students also reported that the multiple suggestion version caused them to slow down and think critically about the line under consideration, the overall purpose of the code, and the benefits of planning. James Prather, Stephen MacNeil, Andrew Luxton-Reilly, Lauren E. Margulieux, Brent N. Reeves, Paul Denny 0001, Juho Leinonen 0001, John Homer, Rahad Arman Nabid, Rachel Louise Rossetti |
ICER (1) | 5 |
| 2025 | Fostering Responsible AI Use Through Negative Expertise: A Contextualized Autocompletion QuizabstractPublisher Copyright: © 2025 Copyright held by the owner/author(s). Stephen MacNeil, James Prather, Rahad Arman Nabid, Sebastian Gutierrez, Silas Carvalho, Saimon Shrestha, Paul Denny 0001, Brent N. Reeves, Juho Leinonen 0001, Rachel Louise Rossetti |
ITiCSE (1) | 8 |
| 2024 | "Backseat Gaming" A Study of Co-Regulated Learning within a Collegiate Male Esports CommunityabstractPrevious work demonstrated that esports players often leverage insights from other players and communities to learn and improve. However, little research examined social learning in esports, over time, in granular detail. Understanding the role of others in the esports learning process has implications for the design of computational support systems that can help esports players learn and make the games more accessible. Therefore, we perform an exploration of this topic using Co-Regulated Learning as a theoretical lens. In doing so, we hope to enrich existing knowledge on social learning in esports, provide insights for the future development of computational support, and a road-map for future work. Through an interview study of an esports community consisting of 14, college-aged, male players, we uncovered 10 themes regarding how Co-Regulated learning occurs within their teams. Based on these, we discuss three main takeaways and their implications for future research and development. Erica Kleinman, Reza Habibi, Garrett B. Powell, Brent N. Reeves, James Prather, Magy Seif El-Nasr |
CHI | 4 |
| 2024 | The Widening Gap: The Benefits and Harms of Generative AI for Novice ProgrammersabstractNovice programmers often struggle through programming problem solving due to a lack of metacognitive awareness and strategies. Previous research has shown that novices can encounter multiple metacognitive difficulties while programming, such as forming incorrect conceptual models of the problem or having a false sense of progress after testing their solution. Novices are typically unaware of how these difficulties are hindering their progress. Meanwhile, many novices are now programming with generative AI (GenAI), which can provide complete solutions to most introductory programming problems, code suggestions, hints for next steps when stuck, and explain cryptic error messages. Its impact on novice metacognition has only started to be explored. Here we replicate a previous study that examined novice programming problem solving behavior and extend it by incorporating GenAI tools. Through 21 lab sessions consisting of participant observation, interview, and eye tracking, we explore how novices are coding with GenAI tools. Although 20 of 21 students completed the assigned programming problem, our findings show an unfortunate divide in the use of GenAI tools between students who did and did not struggle. Some students who did not struggle were able to use GenAI to accelerate, creating code they already intended to make, and were able to ignore unhelpful or incorrect inline code suggestions. But for students who struggled, our findings indicate that previously known metacognitive difficulties persist, and that GenAI unfortunately can compound them and even introduce new metacognitive difficulties. Furthermore, struggling students often expressed cognitive dissonance about their problem solving ability, thought they performed better than they did, and finished with an illusion of competence. Based on our observations from both groups, we propose ways to scaffold the novice GenAI experience and make suggestions for future work. James Prather, Brent N. Reeves, Juho Leinonen 0001, Stephen MacNeil, Arisoa S. Randrianasolo, Brett A. Becker, Bailey Kimmel, Jared Wright, Ben Briggs |
ICER (1) | 2 |
| 2024 | Self-Regulation, Self-Efficacy, and Fear of Failure Interactions with How Novices Use LLMs to Solve Programming ProblemsabstractWe explored how undergraduate introductory programming students naturalistically used generative AI to solve programming problems. We focused on the relationship between their use of AI to their self-regulation strategies, self-efficacy, and fear of failure in programming. In this repeated-measures, mixed-methods research, we examined students' patterns of using generative AI with qualitative student reflections and their self-regulation, self-efficacy, and fear of failure with quantitative instruments at multiple times throughout the semester. We also explored the relationships among these variables to learner characteristics, perceived usefulness of AI, and performance. Overall, our results suggest that student factors affect their baseline use of AI. In particular, students with higher self-efficacy, lower fear of failure, or higher prior grades tended to use AI less or later in the problem-solving process and rated it as less useful than others. Interestingly, we found no relationship between students' self-regulation strategies and their use of AI. Students who used AI less or later in problem-solving also had higher grades in the course, but this is most likely due to prior characteristics as our data do not suggest that this is a causal relationship. Lauren E. Margulieux, James Prather, Brent N. Reeves, Brett A. Becker, Gozde Cetin Uzun, Dastyni Loksa, Juho Leinonen 0001, Paul Denny 0001 |
ITiCSE (1) | 3 |
| 2024 | How Instructors Incorporate Generative AI into Teaching ComputingabstractGenerative AI (GenAI) has seen great advancements in the past two years and the conversation around adoption is increasing. Widely available GenAI tools are disrupting classroom practices as they can write and explain code with minimal student prompting. While most acknowledge that there is no way to stop students from using such tools, a consensus has yet to form on how students should use them if they choose to do so. At the same time, researchers have begun to introduce new pedagogical tools that integrate GenAI into computing curricula. These new tools offer students personalized help or attempt to teach prompting skills without undercutting code comprehension. This working group aims to detail the current landscape of education-focused GenAI tools and teaching approaches, present gaps where new tools or approaches could appear, identify good practice-examples, and provide a guide for instructors to utilize GenAI as they continue to adapt to this new era. James Prather, Juho Leinonen 0001, Natalie Kiesler, Jamie Gorson Benario, Sam Lau, Stephen MacNeil, Narges Norouzi, Simone Opel, Virginia Pettit, Leo Porter 0001, Brent N. Reeves, Jaromír Savelka, David H. Smith IV, Sven Strickroth, Daniel Zingaro |
ITiCSE (2) | 11 |
| 2024 | Prompt Problems: A New Programming Exercise for the Generative AI EraabstractLarge language models (LLMs) are revolutionizing the field of computing education with their powerful code-generating capabilities. Traditional pedagogical practices have focused on code writing tasks, but there is now a shift in importance towards reading, comprehending and evaluating LLM-generated code. Alongside this shift, an important new skill is emerging -- the ability to solve programming tasks by constructing good prompts for code-generating models. In this work we introduce a new type of programming exercise to hone this nascent skill: 'Prompt Problems'. Prompt Problems are designed to help students learn how to write effective prompts for AI code generators. A student solves a Prompt Problem by crafting a natural language prompt which, when provided as input to an LLM, outputs code that successfully solves a specified programming task. We also present a new web-based tool called Promptly which hosts a repository of Prompt Problems and supports the automated evaluation of prompt-generated code. We deploy Promptly in one CS1 and one CS2 course and describe our experiences, which include student perceptions of this new type of activity and their interactions with the tool. We find that students are enthusiastic about Prompt Problems, and appreciate how the problems engage their computational thinking skills and expose them to new programming constructs. We discuss ideas for the future development of new variations of Prompt Problems, and the need to carefully study their integration into classroom practice. Paul Denny 0001, Juho Leinonen 0001, James Prather, Andrew Luxton-Reilly, Thezyrie Amarouche, Brett A. Becker, Brent N. Reeves |
SIGCSE (1) | 7 |
| 2024 | Solving Proof Block Problems Using Large Language ModelsabstractLarge language models (LLMs) have recently taken many fields, including computer science, by storm. Most recent work on LLMs in computing education has shown that they are capable of solving most introductory programming (CS1) exercises, exam questions, Parsons problems, and several other types of exercises and questions. Some work has investigated the ability of LLMs to solve CS2 problems as well. However, it remains unclear how well LLMs fare against more advanced upper-division coursework, such as proofs in algorithms courses. After all, while known to be proficient in many programming tasks, LLMs have been shown to have more difficulties in forming mathematical proofs. Seth Poulsen, Sami Sarsa, James Prather, Juho Leinonen 0001, Brett A. Becker, Arto Hellas, Paul Denny 0001, Brent N. Reeves |
SIGCSE (1) | 8 |
| 2024 | "It's Weird That it Knows What I Want": Usability and Interactions with Copilot for Novice ProgrammersabstractRecent developments in deep learning have resulted in code-generation models that produce source code from natural language and code-based prompts with high accuracy. This is likely to have profound effects in the classroom, where novices learning to code can now use free tools to automatically suggest solutions to programming exercises and assignments. However, little is currently known about how novices interact with these tools in practice. We present the first study that observes students at the introductory level using one such code auto-generating tool, Github Copilot, on a typical introductory programming (CS1) assignment. Through observations and interviews we explore student perceptions of the benefits and pitfalls of this technology for learning, present new observed interaction patterns, and discuss cognitive and metacognitive difficulties faced by students. We consider design implications of these findings, specifically in terms of how tools like Copilot can better support and scaffold the novice programming experience. James Prather, Brent N. Reeves, Paul Denny 0001, Brett A. Becker, Juho Leinonen 0001, Andrew Luxton-Reilly, Garrett B. Powell, James Finnie-Ansley, Eddie A. Santos |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2023 | Transformed by Transformers: Navigating the AI Coding Revolution for Computing Education: An ITiCSE Working Group Conducted by HumansabstractThe recent advent of highly accurate and scalable large language models (LLMs) has taken the world by storm. From art to essays to computer code, LLMs are producing novel content that until recently was thought only humans could produce. Recent work in computing education has sought to understand the capabilities of LLMs for solving tasks such as writing code, explaining code, creating novel coding assignments, interpreting programming error messages, and more. However, these technologies continue to evolve at an astonishing rate leaving educators little time to adapt. This working group seeks to document the state-of-the-art for code generation LLMs, detail current opportunities and challenges related to their use, and present actionable approaches to integrating them into computing curricula. James Prather, Paul Denny 0001, Juho Leinonen 0001, Brett A. Becker, Ibrahim Albluwi, Michael E. Caspersen, Michelle Craig, Hieke Keuning, Natalie Kiesler, Tobias Kohn, Andrew Luxton-Reilly, Stephen MacNeil, Andrew Petersen 0001, Raymond Pettit, Brent N. Reeves, Jaromír Savelka |
ITiCSE (2) | 15 |
| 2023 | Evaluating the Performance of Code Generation Models for Solving Parsons Problems With Small Prompt VariationsabstractThe recent emergence of code generation tools powered by large language models has attracted wide attention. Models such as OpenAI Codex can take natural language problem descriptions as input and generate highly accurate source code solutions, with potentially significant implications for computing education. Given the many complexities that students face when learning to write code, they may quickly become reliant on such tools without properly understanding the underlying concepts. One popular approach for scaffolding the code writing process is to use Parsons problems, which present solution lines of code in a scrambled order. These remove the complexities of low-level syntax, and allow students to focus on algorithmic and design-level problem solving. It is unclear how well code generation models can be applied to solve Parsons problems, given the mechanics of these models and prior evidence that they underperform when problems include specific restrictions. In this paper, we explore the performance of the Codex model for solving Parsons problems over various prompt variations. Using a corpus of Parsons problems we sourced from the computing education literature, we find that Codex successfully reorders the problem blocks about half of the time, a much lower rate of success when compared to prior work on more free-form programming tasks. Regarding prompts, we find that small variations in prompting have a noticeable effect on model performance, although the effect is not as pronounced as between different problems. Brent N. Reeves, Sami Sarsa, James Prather, Paul Denny 0001, Brett A. Becker, Arto Hellas, Bailey Kimmel, Garrett B. Powell, Juho Leinonen 0001 |
ITiCSE (1) | 1 |
| 2023 | Using Large Language Models to Enhance Programming Error MessagesabstractA key part of learning to program is learning to understand programming error messages. They can be hard to interpret and identifying the cause of errors can be time-consuming. One factor in this challenge is that the messages are typically intended for an audience that already knows how to program, or even for programming environments that then use the information to highlight areas in code. Researchers have been working on making these errors more novice friendly since the 1960s, however progress has been slow. The present work contributes to this stream of research by using large language models to enhance programming error messages with explanations of the errors and suggestions on how to fix them. Large language models can be used to create useful and novice-friendly enhancements to programming error messages that sometimes surpass the original programming error messages in interpretability and actionability. These results provide further evidence of the benefits of large language models for computing educators, highlighting their use in areas known to be challenging for students. We further discuss the benefits and downsides of large language models and highlight future streams of research for enhancing programming error messages. Juho Leinonen 0001, Arto Hellas, Sami Sarsa, Brent N. Reeves, Paul Denny 0001, James Prather, Brett A. Becker |
SIGCSE (1) | 4 |
| 2023 | First Steps Towards Predicting the Readability of Programming Error MessagesabstractReading a programming error message is the first step in understanding what it is trying to tell the programmer about how to fix an error in their code. However, these are often difficult to read, especially for novices which is not surprising given that error messages in many of the most popular languages in which novices learn to code were not written with readability in mind. As a result, novices frequently struggle to understand them. This is a long-standing problem, with researchers highlighting concerns about programming error message readability over the last six decades. Very recent work has put forward evidence of the need for measuring readability in error messages and a framework for doing so. This framework consists of four factors of readability for programming error messages: message length, vocabulary, jargon, and sentence construction. We use this framework to implement an approach to automatically assess the readability of programming error messages. Using established readability factors as predictors in a machine learning model, we train several models using a dataset of C and Java error messages. We examine the performance of these models, and apply the best performing model to a previously published set of messages evaluated for readability by experts, non-experts and students. Our results validate the previously proposed readability factors, and our model classifies messages similarly to human raters. Finally, we discuss future work needed to improve the accuracy of the model. James Prather, Paul Denny 0001, Brett A. Becker, Robert Nix, Brent N. Reeves, Arisoa S. Randrianasolo, Garrett B. Powell |
SIGCSE (1) | 5 |
| 2022 | Getting By With Help From My Friends: Group Study in Introductory Programming Understood as Socially Shared RegulationabstractBackground and Context. Metacognitive skills are important for all students learning to program and interest in applying pedagogical approaches in early programming courses that focus on metacognitive aspects is growing. However, most studies of such approaches are not rigorously based in theory, and when they are, almost always utilize foundational education and psychology theories from as far back as the 1970s. More recent theory is less tested, and not all relevant metacognitive theories have been explored in the computing education research literature. James Prather, Lauren E. Margulieux, Jacqueline L. Whalley, Paul Denny 0001, Brent N. Reeves, Brett A. Becker, Paramvir Singh, Garrett B. Powell, Nigel Bosch |
ICER (1) | 5 |
| 2022 | Novice Reflections During the Transition to a New Programming LanguageabstractAs computing students progress through their studies they become proficient with multiple programming languages. Prior work investigating language transitions for novices has tended to analyze program artifacts rather than explore the benefits and difficulties as perceived by students in their own words, and has often overlooked problems that may arise in switching paradigms or where familiar syntax has a different meaning in the new language. In this paper, we ask students to reflect on the transition from an interpreted language and environment (MATLAB) to a compiled language (C), prompting comments on the aspects of learning the new language that they found both easier and harder. Analysis of over 70,000 words written by 771 students revealed that the highest-performing students expressed more negative sentiments towards the language transition -- a surprising result that we hypothesize is explained by their generally stronger metacognitive skills. We also report the most common difficulties described by students, which include challenges with syntax, error messages, and the process of compilation, and suggest teaching practices that might help students as they transition to a new programming language. Paul Denny 0001, Brett A. Becker, Nigel Bosch, James Prather, Brent N. Reeves, Jacqueline L. Whalley |
SIGCSE (1) | 5 |
| 2007 | Zing'Em: a Web-based Likert-scale Student-Team Peer Evaluation ToolabstractWe describe the design and implementation of a peer evaluation system named Zing'Em, in use for 3 years and over 60 course sections. Zing'Em supports Likert-scale items, is easy for professors to set up, simple for students to use and provides analyses of teams for the professor to use in grading and possible interventions. This paper documents the design goals, implementation tradeoffs and experience of users. We critique the system and suggest improvements for future work. Brent N. Reeves |
ICALT | 1 |