VLDB 2026 Research / reviewers in the wild / expert
Sami Sarsa
dblp:278/2930
· DBLP profile ↗
18ranked-venue papers
3as first author
18since 2021 · last 2025
0000-0002-7277-9282ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 18 · 3 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating Language Models for Generating and Judging Programming FeedbackabstractThe emergence of large language models (LLMs) has transformed research and practice across a wide range of domains. Within the computing education research (CER) domain, LLMs have garnered significant attention, particularly in the context of learning programming. Much of the work on LLMs in CER, however, has focused on applying and evaluating proprietary models. In this article, we evaluate the efficiency of open-source LLMs in generating high-quality feedback for programming assignments and judging the quality of programming feedback, contrasting the results with proprietary models. Our evaluations on a dataset of students' submissions to introductory Python programming exercises suggest that state-of-the-art open-source LLMs are nearly on par with proprietary models in both generating and assessing programming feedback. Additionally, we demonstrate the efficiency of smaller LLMs in these tasks and highlight the wide range of LLMs accessible, even for free, to educators and practitioners. Charles Koutcheme, Nicola Dainese, Sami Sarsa, Arto Hellas, Juho Leinonen 0001, Syed Ashraf, Paul Denny 0001 |
SIGCSE (1) | 3 |
| 2024 | Evaluating Contextually Personalized Programming Exercises Created with Generative AIabstractProgramming skills are typically developed through completing various hands-on exercises. Such programming problems can be contextualized to students’ interests and cultural backgrounds. Prior research in educational psychology has demonstrated that context personalization of exercises stimulates learners’ situational interests and positively affects their engagement. However, creating a varied and comprehensive set of programming exercises for students to practice on is a time-consuming and laborious task for computer science educators. Previous studies have shown that large language models can generate conceptually and contextually relevant programming exercises. Thus, they offer a possibility to automatically produce personalized programming problems to fit students’ interests and needs. This article reports on a user study conducted in an elective introductory programming course that included contextually personalized programming exercises created with GPT-4. The quality of the exercises was evaluated by both the students and the authors. Additionally, this work investigated student attitudes towards the created exercises and their engagement with the system. The results demonstrate that the quality of exercises generated with GPT-4 was generally high. What is more, the course participants found them engaging and useful. This suggests that AI-generated programming problems can be a worthwhile addition to introductory programming courses, as they provide students with a practically unlimited pool of practice material tailored to their personal interests and educational needs. Evanfiya Logacheva, Arto Hellas, James Prather, Sami Sarsa, Juho Leinonen 0001 |
ICER (1) | 4 |
| 2024 | "Like a Nesting Doll": Analyzing Recursion Analogies Generated by CS Students Using Large Language ModelsabstractGrasping complex computing concepts often poses a challenge for students who struggle to anchor these new ideas to familiar experiences and understandings. To help with this, a good analogy can bridge the gap between unfamiliar concepts and familiar ones, providing an engaging way to aid understanding. However, creating effective educational analogies is difficult even for experienced instructors. We investigate to what extent large language models (LLMs), specifically ChatGPT, can provide access to personally relevant analogies on demand. Focusing on recursion, a challenging threshold concept, we conducted an investigation analyzing the analogies generated by more than 350 first-year computing students. They were provided with a code snippet and tasked to generate their own recursion-based analogies using ChatGPT, optionally including personally relevant topics in their prompts. We observed a great deal of diversity in the analogies produced with student-prescribed topics, in contrast to the otherwise generic analogies, highlighting the value of student creativity when working with LLMs. Not only did students enjoy the activity and report an improved understanding of recursion, but they described more easily remembering analogies that were personally and culturally relevant. Seth Bernstein, Paul Denny 0001, Juho Leinonen 0001, Lauren Kan, Arto Hellas, Matt Littlefield, Sami Sarsa, Stephen MacNeil |
ITiCSE (1) | 7 |
| 2024 | Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-JudgeabstractLarge language models (LLMs) have shown great potential for the automatic generation of feedback in a wide range of computing contexts. However, concerns have been voiced around the privacy and ethical implications of sending student work to proprietary models. This has sparked considerable interest in the use of open source LLMs in education, but the quality of the feedback that such open models can produce remains understudied. This is a concern as providing flawed or misleading generated feedback could be detrimental to student learning. Inspired by recent work that has utilised very powerful LLMs, such as GPT-4, to evaluate the outputs produced by less powerful models, we conduct an automated analysis of the quality of the feedback produced by several open source models using a dataset from an introductory programming course. First, we investigate the viability of employing GPT-4 as an automated evaluator by comparing its evaluations with those of a human expert. We observe that GPT-4 demonstrates a bias toward positively rating feedback while exhibiting moderate agreement with human raters, showcasing its potential as a feedback evaluator. Second, we explore the quality of feedback generated by several leading open-source LLMs by using GPT-4 to evaluate the feedback. We find that some models offer competitive performance with popular proprietary LLMs, such as ChatGPT, indicating opportunities for their responsible use in educational settings. Charles Koutcheme, Nicola Dainese, Sami Sarsa, Arto Hellas, Juho Leinonen 0001, Paul Denny 0001 |
ITiCSE (1) | 3 |
| 2024 | Solving Proof Block Problems Using Large Language ModelsabstractLarge language models (LLMs) have recently taken many fields, including computer science, by storm. Most recent work on LLMs in computing education has shown that they are capable of solving most introductory programming (CS1) exercises, exam questions, Parsons problems, and several other types of exercises and questions. Some work has investigated the ability of LLMs to solve CS2 problems as well. However, it remains unclear how well LLMs fare against more advanced upper-division coursework, such as proofs in algorithms courses. After all, while known to be proficient in many programming tasks, LLMs have been shown to have more difficulties in forming mathematical proofs. Seth Poulsen, Sami Sarsa, James Prather, Juho Leinonen 0001, Brett A. Becker, Arto Hellas, Paul Denny 0001, Brent N. Reeves |
SIGCSE (1) | 2 |
| 2023 | Automated Program Repair Using Generative Models for Code Infilling
Charles Koutcheme, Sami Sarsa, Juho Leinonen 0001, Arto Hellas, Paul Denny 0001 |
AIED | 2 |
| 2023 | Exploring the Responses of Large Language Models to Beginner Programmers' Help RequestsabstractBackground and Context: Over the past year, large language models (LLMs) have taken the world by storm. In computing education, like in other walks of life, many opportunities and threats have emerged as a consequence. Arto Hellas, Juho Leinonen 0001, Sami Sarsa, Charles Koutcheme, Lilja Koivuniemi, Juha Sorva |
ICER (1) | 3 |
| 2023 | Evaluating Distance Measures for Program RepairabstractBackground and Context: Struggling with programming assignments while learning to program is a common phenomenon in programming courses around the world. Supporting struggling students is a common theme in Computing Education Research (CER), where a wide variety of support methods have been created and evaluated. An important stream of research here focuses on program repair, where methods for automatically fixing erroneous code are used for supporting students as they debug their code. Work in this area has so far assessed the performance of the methods by evaluating the closeness of the proposed fixes to the original erroneous code. The evaluations have mainly relied on the use of edit distance measures such as the sequence edit distance and there is a lack of research on which distance measure is the most appropriate. Charles Koutcheme, Sami Sarsa, Juho Leinonen 0001, Lassi Haaranen, Arto Hellas |
ICER (1) | 2 |
| 2023 | Comparing Code Explanations Created by Students and Large Language ModelsabstractReasoning about code and explaining its purpose are fundamental skills for computer scientists. There has been extensive research in the field of computing education on the relationship between a student's ability to explain code and other skills such as writing and tracing code. In particular, the ability to describe at a high-level of abstraction how code will behave over all possible inputs correlates strongly with code writing skills. However, developing the expertise to comprehend and explain code accurately and succinctly is a challenge for many students. Existing pedagogical approaches that scaffold the ability to explain code, such as producing exemplar code explanations on demand, do not currently scale well to large classrooms. The recent emergence of powerful large language models (LLMs) may offer a solution. In this paper, we explore the potential of LLMs in generating explanations that can serve as examples to scaffold students' ability to understand and explain code. To evaluate LLM-created explanations, we compare them with explanations created by students in a large course (n ≈ 1000) with respect to accuracy, understandability and length. We find that LLM-created explanations, which can be produced automatically on demand, are rated as being significantly easier to understand and more accurate summaries of code than student-created explanations. We discuss the significance of this finding, and suggest how such models can be incorporated into introductory programming education. Juho Leinonen 0001, Paul Denny 0001, Stephen MacNeil, Sami Sarsa, Seth Bernstein, Joanne Kim, Andrew Tran, Arto Hellas |
ITiCSE (1) | 4 |
| 2023 | Evaluating the Performance of Code Generation Models for Solving Parsons Problems With Small Prompt VariationsabstractThe recent emergence of code generation tools powered by large language models has attracted wide attention. Models such as OpenAI Codex can take natural language problem descriptions as input and generate highly accurate source code solutions, with potentially significant implications for computing education. Given the many complexities that students face when learning to write code, they may quickly become reliant on such tools without properly understanding the underlying concepts. One popular approach for scaffolding the code writing process is to use Parsons problems, which present solution lines of code in a scrambled order. These remove the complexities of low-level syntax, and allow students to focus on algorithmic and design-level problem solving. It is unclear how well code generation models can be applied to solve Parsons problems, given the mechanics of these models and prior evidence that they underperform when problems include specific restrictions. In this paper, we explore the performance of the Codex model for solving Parsons problems over various prompt variations. Using a corpus of Parsons problems we sourced from the computing education literature, we find that Codex successfully reorders the problem blocks about half of the time, a much lower rate of success when compared to prior work on more free-form programming tasks. Regarding prompts, we find that small variations in prompting have a noticeable effect on model performance, although the effect is not as pronounced as between different problems. Brent N. Reeves, Sami Sarsa, James Prather, Paul Denny 0001, Brett A. Becker, Arto Hellas, Bailey Kimmel, Garrett B. Powell, Juho Leinonen 0001 |
ITiCSE (1) | 2 |
| 2023 | Using Large Language Models to Enhance Programming Error MessagesabstractA key part of learning to program is learning to understand programming error messages. They can be hard to interpret and identifying the cause of errors can be time-consuming. One factor in this challenge is that the messages are typically intended for an audience that already knows how to program, or even for programming environments that then use the information to highlight areas in code. Researchers have been working on making these errors more novice friendly since the 1960s, however progress has been slow. The present work contributes to this stream of research by using large language models to enhance programming error messages with explanations of the errors and suggestions on how to fix them. Large language models can be used to create useful and novice-friendly enhancements to programming error messages that sometimes surpass the original programming error messages in interpretability and actionability. These results provide further evidence of the benefits of large language models for computing educators, highlighting their use in areas known to be challenging for students. We further discuss the benefits and downsides of large language models and highlight future streams of research for enhancing programming error messages. Juho Leinonen 0001, Arto Hellas, Sami Sarsa, Brent N. Reeves, Paul Denny 0001, James Prather, Brett A. Becker |
SIGCSE (1) | 3 |
| 2023 | The Implications of Large Language Models for CS Teachers and StudentsabstractThe introduction of Large Language Models (LLMs) has generated a significant amount of excitement both in industry and among researchers. Recently, tools that leverage LLMs have made their way into the classroom where they help students generate code and help instructors generate learning materials. There are likely many more uses of these tools -- both beneficial to learning and possibly detrimental to learning. To help ensure that these tools are used to enhance learning, educators need to not only be familiar with these tools, but with their use and potential misuse. The goal of this BoF is to raise awareness about LLMs and to build a learning community around their use in computing education. Aligned with this goal of building an inclusive learning community, our BoF is led by globally distributed discussion leaders, including undergraduate researchers, to facilitate multiple coordinated discussions that can lead to a broader conversation about the role of LLMs in CS education. Stephen MacNeil, Joanne Kim, Juho Leinonen 0001, Paul Denny 0001, Seth Bernstein, Brett A. Becker, Michel Wermelinger, Arto Hellas, Andrew Tran, Sami Sarsa, James Prather, Viraj Kumar |
SIGCSE (2) | 10 |
| 2023 | Automatically Generating CS Learning Materials with Large Language ModelsabstractRecent breakthroughs in Large Language Models (LLMs), such as GPT-3 and Codex, now enable software developers to generate code based on a natural language prompt. Within computer science education, researchers are exploring the potential for LLMs to generate code explanations and programming assignments using carefully crafted prompts. These advances may enable students to interact with code in new ways while helping instructors scale their learning materials. However, LLMs also introduce new implications for academic integrity, curriculum design, and software engineering careers. This workshop will demonstrate the capabilities of LLMs to help attendees evaluate whether and how LLMs might be integrated into their pedagogy and research. We will also engage attendees in brainstorming to consider how LLMs will impact our field. Stephen MacNeil, Andrew Tran, Juho Leinonen 0001, Paul Denny 0001, Joanne Kim, Arto Hellas, Seth Bernstein, Sami Sarsa |
SIGCSE (2) | 8 |
| 2023 | Experiences from Using Code Explanations Generated by Large Language Models in a Web Software Development E-BookabstractAdvances in natural language processing have resulted in large language models (LLMs) that can generate code and code explanations. In this paper, we report on our experiences generating multiple code explanation types using LLMs and integrating them into an interactive e-book on web software development. Three different types of explanations -- a line-by-line explanation, a list of important concepts, and a high-level summary of the code -- were created. Students could view explanations by clicking a button next to code snippets, which showed the explanation and asked about its utility. Our results show that all explanation types were viewed by students and that the majority of students perceived the code explanations as helpful to them. However, student engagement varied by code snippet complexity, explanation type, and code snippet length. Drawing on our experiences, we discuss future directions for integrating explanations generated by LLMs into CS classrooms. Stephen MacNeil, Andrew Tran, Arto Hellas, Joanne Kim, Sami Sarsa, Paul Denny 0001, Seth Bernstein, Juho Leinonen 0001 |
SIGCSE (1) | 5 |
| 2022 | How to Help to Ask for Help? Help Request Prompt Structure Influence on Help Request Quantity and Course RetentionabstractFull research paper—Feedback and support are at the core of efficient learning. While the effect of feedback has been explored in a multitude of studies, the effect of asking for help and the effect of how that help is asked for is an under-explored area. In this work, we present the results of a randomized controlled trial organized in an introductory programming course for lifelong learners. In the study, we explored the usefulness of different types of help prompts used to guide learners asking for help. Gauging the effect of different combinations of three prompts used to scaffold writing out the help request, (1) "Describe the issue with your program", (2) "Explain how your program works", and (3) "What have you tried to do to resolve the issue", we study how prompts influence learners’ behavior. Using log data collected from the online platform where the prompts were explored, we study how the prompts affect whether learners end up sending a help request instead of simply considering to send one, how the scaffolding prompts affect whether learners send further help requests, and whether the questions have an effect on course retention. Our results show that the help request prompts have an impact on whether learners end up actually writing a help request, which in turn also influences whether the learner will ask for help on another occasion. Further, we observe that help request prompts may have an effect on whether learners figure out a solution to a problem on their own. Alarmingly, we also observe that the prompts can affect how far learners go in a course. Based on our results, we advise teachers and researchers to pay attention into how learners are guided into asking for help. Sami Sarsa, Jesper Pettersson, Arto Hellas |
FIE | 1 |
| 2022 | Automatic Generation of Programming Exercises and Code Explanations Using Large Language ModelsabstractThis article explores the natural language generation capabilities of large language models with application to the production of two types of learning resources common in programming courses. Using OpenAI Codex as the large language model, we create programming exercises (including sample solutions and test cases) and code explanations, assessing these qualitatively and quantitatively. Our results suggest that the majority of the automatically generated content is both novel and sensible, and in some cases ready to use as is. When creating exercises we find that it is remarkably easy to influence both the programming concepts and the contextual themes they contain, simply by supplying keywords as input to the model. Our analysis suggests that there is significant value in massive generative machine learning models as a tool for instructors, although there remains a need for some oversight to ensure the quality of the generated content before it is delivered to students. We further discuss the implications of OpenAI Codex and similar tools for introductory programming education and highlight future research streams that have the potential to improve the quality of the educational experience for both teachers and students alike. Sami Sarsa, Paul Denny 0001, Arto Hellas, Juho Leinonen 0001 |
ICER (1) | 1 |
| 2022 | Steps Learners Take when Solving Programming Tasks, and How Learning Environments (Should) Respond to ThemabstractEvery year, millions of students learn how to write programs. Learning activities for beginners almost always include programming tasks that require a student to write a program to solve a particular problem. When learning how to solve such a task, many students need feedback on their previous actions, and hints on how to proceed. In the case of programming, the feedback should take the steps a student has taken towards implementing a solution into account, and the hints should help a student to complete or improve a possibly partial solution. Only a limited number of learning environments for programming give feedback and hints on intermediate steps students take towards a solution, and little is known about the quality of the feedback provided. To determine the quality of feedback of such tools and to help further developing them, we create and curate data sets that show what kinds of steps students take when solving programming exercises for beginners, and what kind of feedback and hints should be provided. This working group aims to 1) select or create several data sets with steps students take to solve programming tasks, 2) introduce a method to annotate students' steps in these data sets, 3) attach feedback and hints to these steps, 4) set up a method to utilize these data sets in various learning environments for programming, and 5) analyse the quality of hints and feedback in these learning environments. Johan Jeuring, Hieke Keuning, Samiha Marwan, Dennis J. Bouvier, Cruz Izu, Natalie Kiesler, Teemu Lehtinen, Dominic Lohr, Andrew Petersen 0001, Sami Sarsa |
ITiCSE (2) | 10 |
| 2022 | Who Continues in a Series of Lifelong Learning Courses?abstractAlthough computing education research quite often targets within-university courses, an important role of universities is educating the public through open online lifelong learning offerings. Compared to within-university courses, in lifelong learning, the student population is often more diverse. For example, participants often have more varied motivations and aspirations as well as more varied educational backgrounds. In this work, we explore what kinds of learners attend open online lifelong learning programming courses and what characteristics of learners lead to completing courses and proceeding to subsequent courses. We examine student-related factors collected through surveys in our online course environment. These factors include motivation, previous experience, and demographics. Our results show that motivations, previous experience, and demographics by themselves only explain a small amount of the variance in completing courses or continuing to a subsequent course. At the same time, we identify individual factors that are more likely to lead to learners dropping out (or continuing) in the courses. Our study provides further evidence that lifelong learning benefits most the already educated part of the population with prior knowledge and high motivation. This calls for further studies that seek to identify means to engage and support participants less likely to continue in such courses. Sami Sarsa, Arto Hellas, Juho Leinonen 0001 |
ITiCSE (1) | 1 |