VLDB 2026 Research / reviewers in the wild / expert
Charles Koutcheme
dblp:313/4141
· DBLP profile ↗
10ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0002-2272-2763ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 9 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Aligning Small Language Models for Programming Feedback: Towards Scalable Coding Support in a Massive Global CourseabstractProviding timely and actionable feedback is essential for students learning to program. While large language models (LLMs) are increasingly used to automate this process, they remain costly to deploy and raise concerns around privacy and institutional control. Small language models (SLMs) offer a promising alternative: they can be run locally and integrated more flexibly into educational platforms. However, their out-of-the-box performance is often poor, requiring targeted training to be effective in classrooms. In this paper, we investigate whether a trained 3B-parameter SLM, guided by rubric-based prompting and a pipeline combining supervised and preference-based learning, can generate diagnostic feedback that approaches the quality of larger models. We deploy the model in a large-scale online programming course and compare its feedback to its base and fine-tuned variants, Llama-3.1-8B, and GPT-4.1, using human ratings from 53 teaching assistants and an automated LLM-as-a-judge analysis. Our results show that careful training narrows the feedback quality gap between an SLM and an LLM from over 80 to just 10 percentage points on key metrics. The trained SLM more rarely hallucinates errors, is often rated as helpful by educators, and only occasionally misses issues in student code. These findings suggest that small models can serve as practical and scalable targeted feedback solutions in large educational settings, while LLMs may remain necessary for more comprehensive diagnostic feedback. Charles Koutcheme, Juliette Woodrow, Chris Piech |
SIGCSE (1) | 1 |
| 2026 | Fine-Tuning Open-Source Models as a Viable Alternative to Proprietary LLMs for Explaining Compiler MessagesabstractCryptic compiler error messages continue to present a significant barrier for novice programmers, especially in foundational languages like C. Although large language models (LLMs) can generate accurate and comprehensible error explanations, their computational requirements, propensity for over-assistance, and privacy concerns constrain their suitability for widespread adoption in educational tools. This work investigates how Supervised Fine-Tuning (SFT) can enhance the performance of smaller, open-source models when explaining C compiler errors to students in introductory programming courses (CS1/2). We derive a training dataset of 40,000 input-output pairs from CS1/2 student C compiler errors to fine-tune three open-source models: Qwen3-4B, Llama-3.1-8B, and Qwen3-32B. Model performance was assessed through a dual evaluation framework involving expert human reviewers and a large-scale automated analysis of 8,000 responses using an ensemble of models as judges. Our results indicate that SFT significantly improves both expert and LLM-as-judge ratings in smaller open-source models, with reduced gains in the larger model. We analyse the trade-offs between model size and quality, and validate LLM-as-judge by demonstrating inter-rater agreement with experts. Our findings demonstrate that fine-tuning smaller models on high-quality data is a viable strategy for creating specialised pedagogical tools. We provide a replicable methodology for enabling broader access to advanced AI capabilities within educational contexts, especially with smaller, economical models. Lorenzo Lee Solano, Charles Koutcheme, Juho Leinonen 0001, Alexandra Vassar, Jake Renzella |
SIGCSE (1) | 2 |
| 2025 | Evaluating Language Models for Generating and Judging Programming FeedbackabstractThe emergence of large language models (LLMs) has transformed research and practice across a wide range of domains. Within the computing education research (CER) domain, LLMs have garnered significant attention, particularly in the context of learning programming. Much of the work on LLMs in CER, however, has focused on applying and evaluating proprietary models. In this article, we evaluate the efficiency of open-source LLMs in generating high-quality feedback for programming assignments and judging the quality of programming feedback, contrasting the results with proprietary models. Our evaluations on a dataset of students' submissions to introductory Python programming exercises suggest that state-of-the-art open-source LLMs are nearly on par with proprietary models in both generating and assessing programming feedback. Additionally, we demonstrate the efficiency of smaller LLMs in these tasks and highlight the wide range of LLMs accessible, even for free, to educators and practitioners. Charles Koutcheme, Nicola Dainese, Sami Sarsa, Arto Hellas, Juho Leinonen 0001, Syed Ashraf, Paul Denny 0001 |
SIGCSE (1) | 1 |
| 2024 | Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-JudgeabstractLarge language models (LLMs) have shown great potential for the automatic generation of feedback in a wide range of computing contexts. However, concerns have been voiced around the privacy and ethical implications of sending student work to proprietary models. This has sparked considerable interest in the use of open source LLMs in education, but the quality of the feedback that such open models can produce remains understudied. This is a concern as providing flawed or misleading generated feedback could be detrimental to student learning. Inspired by recent work that has utilised very powerful LLMs, such as GPT-4, to evaluate the outputs produced by less powerful models, we conduct an automated analysis of the quality of the feedback produced by several open source models using a dataset from an introductory programming course. First, we investigate the viability of employing GPT-4 as an automated evaluator by comparing its evaluations with those of a human expert. We observe that GPT-4 demonstrates a bias toward positively rating feedback while exhibiting moderate agreement with human raters, showcasing its potential as a feedback evaluator. Second, we explore the quality of feedback generated by several leading open-source LLMs by using GPT-4 to evaluate the feedback. We find that some models offer competitive performance with popular proprietary LLMs, such as ChatGPT, indicating opportunities for their responsible use in educational settings. Charles Koutcheme, Nicola Dainese, Sami Sarsa, Arto Hellas, Juho Leinonen 0001, Paul Denny 0001 |
ITiCSE (1) | 1 |
| 2024 | Propagating Large Language Models Programming FeedbackabstractLarge language models (LLMs) such as GPT-4 have emerged as promising tools for providing programming feedback. However, effective deployment of LLMs in massive classes and Massive Open Online Courses (MOOCs) raises financial concerns, calling for methods to minimize the number of calls to the APIs and systems serving such powerful models. In this article, we revisit the problem of 'propagating feedback' within the contemporary landscape of LLMs. Specifically, we explore feedback propagation as a way to reduce the cost of leveraging LLMs for providing programming feedback at scale. Our study investigates the effectiveness of this approach in the context of students requiring next-step hints for Python programming problems, presenting initial results that support the viability of the approach. We discuss our findings' implications and suggest directions for future research in optimizing feedback mechanisms for large-scale educational environments. Charles Koutcheme, Arto Hellas |
L@S | 1 |
| 2023 | Training Language Models for Programming Feedback Using Automated Repair Tools
Charles Koutcheme |
AIED | 1 |
| 2023 | Automated Program Repair Using Generative Models for Code Infilling
Charles Koutcheme, Sami Sarsa, Juho Leinonen 0001, Arto Hellas, Paul Denny 0001 |
AIED | 1 |
| 2023 | Exploring the Responses of Large Language Models to Beginner Programmers' Help RequestsabstractBackground and Context: Over the past year, large language models (LLMs) have taken the world by storm. In computing education, like in other walks of life, many opportunities and threats have emerged as a consequence. Arto Hellas, Juho Leinonen 0001, Sami Sarsa, Charles Koutcheme, Lilja Koivuniemi, Juha Sorva |
ICER (1) | 4 |
| 2023 | Evaluating Distance Measures for Program RepairabstractBackground and Context: Struggling with programming assignments while learning to program is a common phenomenon in programming courses around the world. Supporting struggling students is a common theme in Computing Education Research (CER), where a wide variety of support methods have been created and evaluated. An important stream of research here focuses on program repair, where methods for automatically fixing erroneous code are used for supporting students as they debug their code. Work in this area has so far assessed the performance of the methods by evaluating the closeness of the proposed fixes to the original erroneous code. The evaluations have mainly relied on the use of edit distance measures such as the sequence edit distance and there is a lack of research on which distance measure is the most appropriate. Charles Koutcheme, Sami Sarsa, Juho Leinonen 0001, Lassi Haaranen, Arto Hellas |
ICER (1) | 1 |
| 2022 | Exploring How Students Solve Open-ended Assignments: A Study of SQL Injection Attempts in a Cybersecurity CourseabstractResearch into computing and learning how to program has been ongoing for decades. Commonly, this research has been focused on novice learners and the difficulties they encounter, especially during CS1. Cybersecurity is a critical aspect in computing -- as a topic in university education as well as a core skill in the industry. In this study, we investigate how students solve open-ended assignments on a cybersecurity course offered to university students after two years of CS studies. Specifically, we looked at how students perform SQL injection attacks on an web application system, and study to what extent we can characterize the process in which they come up with successful injections. Our results show that there are distinguishable strategies used by individual students who seek to hack the system, where these approaches revolve around exploration and exploitation tactics. We also find evidence of learning due to a more pronounced use of exploitation in a subsequent similar assignment. Charles Koutcheme, Artturi Tilanterä, Aleksi Peltonen, Arto Hellas, Lassi Haaranen |
ITiCSE (1) | 1 |