EDBT 2026 Demo / reviewers in the wild / expert
Arto Hellas
dblp:184/5570
· DBLP profile ↗
80ranked-venue papers
4as first author
45since 2021 · last 2026
0000-0001-6502-209XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 65 · 4 first-author · 36 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 10 since 2021Software engineering, systems software and programming languages · 6 · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Systems, architecture and hardware · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
Rose Niousha, Samantha Boatright Smith, Bita Akram, Peter Brusilovsky, Arto Hellas, Juho Leinonen 0001, John DeNero, Narges Norouzi |
AIED | 5 |
| 2026 | Personalized Worked Example Generation from Student Code Submissions Using Pattern-based Knowledge ComponentsabstractAdaptive programming practice often relies on fixed libraries of worked examples and practice problems, which require substantial authoring effort and may not correspond well to the logical errors and partial solutions students produce while writing code. As a result, students may receive learning content that does not directly address the concepts they are working to understand, while instructors must either invest additional effort in expanding content libraries or accept a coarse level of personalization. We present an approach for knowledge-component (KC) guided educational content generation using pattern-based KCs extracted from student code. Given a problem statement and student submissions, our pipeline extracts recurring structural KC patterns from students' code through AST-based analysis and uses them to condition a generative model. In this study, we apply this approach to worked example generation, and compare baseline and KC-conditioned outputs through expert evaluation. Results suggest that KC-conditioned generation improves topical focus and relevance to students' underlying logical errors, providing evidence that KC-based steering of generative models can support personalized learning at scale. Griffin Pitts, Muntasir Hoq, Peter Brusilovsky, Narges Norouzi, Arto Hellas, Juho Leinonen 0001, Bita Akram |
L@S | 5 |
| 2026 | The Impostor Phenomenon and the Confidence GapabstractThe Impostor Phenomenon (IP) and the Confidence Gap describe gendered differences in self-assessment and perceived competence, with well-documented implications for learning, retention, and performance. While each has been studied independently, little research has investigated how these phenomena interact -- particularly within computing education contexts. This study addresses the gap by examining the relationship between IP and confidence among university students in web software development courses in Northern Europe. Drawing on 392 survey responses, we analyze how IP and self-reported confidence vary with gender, degree level, and prior programming experience, and expanding prior work, how these constructs relate to one another. Our findings confirm previously reported patterns -- women report higher IP and lower confidence than men -- and provide new insights: confidence partially mediates the relationship between gender and IP, indicating that lower confidence may partially help explain gender disparities in IP. Furthermore, while programming experience predicts confidence, its relationship to IP is more nuanced. Experience has a small direct positive effect on IP that is largely offset by an indirect confidence-driven reduction. These findings suggest that impostor feelings are shaped more by self-perception than skill, and they highlight the need for interventions that focus on both skill development and psychological support. Arto Hellas, Andrew Petersen 0001 |
SIGCSE (1) | 1 |
| 2026 | Knowledge Component-Driven Alignment of CS1 Textbooks and ExercisesabstractWe present a reproducible pipeline that aligns CS1 textbook sections with problems from a public dataset via a Knowledge Component (KC) -a single conceptual skill required for problem solving- ontology. It assigns KCs to sections and problems, respects the prerequisite order to avoid inserting problems too early, and generates tips for not-yet-taught concepts. We evaluate three KC assignment strategies: embedding-only, embedding with a Large Language Model (LLM) tie-breaker, and direct LLM assignment. We find direct assignment matches or exceeds human annotators. Our results show that constrained LLMs can enrich CS1 textbooks with curriculum-aware practice problems. Samantha Boatright Smith, Arun Balajiee Lekshmi Narayanan, Anurata Prabha Hridi, Rafaella Sampaio de Alencar, Bita Akram, Arto Hellas, Juho Leinonen 0001, Peter Brusilovsky, Narges Norouzi |
SIGCSE (2) | 6 |
| 2025 | Emergence of LLMs: (Not-so-)Significant Delving in Essay Answers in a MOOC on the Ethics of AI
Leo Leppänen, Lili Aunimo, Arto Hellas, Jukka K. Nurminen, Linda Mannila |
AIED (6) | 3 |
| 2025 | Evaluating Language Models for Generating and Judging Programming FeedbackabstractThe emergence of large language models (LLMs) has transformed research and practice across a wide range of domains. Within the computing education research (CER) domain, LLMs have garnered significant attention, particularly in the context of learning programming. Much of the work on LLMs in CER, however, has focused on applying and evaluating proprietary models. In this article, we evaluate the efficiency of open-source LLMs in generating high-quality feedback for programming assignments and judging the quality of programming feedback, contrasting the results with proprietary models. Our evaluations on a dataset of students' submissions to introductory Python programming exercises suggest that state-of-the-art open-source LLMs are nearly on par with proprietary models in both generating and assessing programming feedback. Additionally, we demonstrate the efficiency of smaller LLMs in these tasks and highlight the wide range of LLMs accessible, even for free, to educators and practitioners. Charles Koutcheme, Nicola Dainese, Sami Sarsa, Arto Hellas, Juho Leinonen 0001, Syed Ashraf, Paul Denny 0001 |
SIGCSE (1) | 4 |
| 2025 | The Potential of Serverless Edge-powered Islands for Web DevelopmentabstractWeb developers face two significant challenges when developing their applications and websites: latency and payload size. Given that web services rely on servers, the related communication incurs a cost in terms of latency. In contrast, the payload passed to the client incurs a communication cost, not to mention the computational cost to the client. The concept of serverless edge computing, built on top of content delivery networks (CDNs), is an approach that has begun to gain the attention of web developers for its promise of lower latencies due to its efficiencies in communication thanks to globally distributed networks and replication. Islands architecture is a technical approach that addresses payload size by giving developers easy ways to defer and potentially even avoid the cost of loading content. Combined, these two approaches form edge-powered islands and, in this article, we examine how the combination can help to address these two notable costs web developers have to consider in their daily work. Our findings indicate that edge-powered islands can provide a way to introduce interactivity to otherwise static websites while wrapping dynamic portions of a page within islands to gain the benefits of static approaches in more dynamic contexts, such as storefronts. In addition, islands can provide loading benefits even for more application-like websites, such as social networks, and give web developers an additional control layer in their development work. Juho Vepsäläinen, Petri Vuorimaa, Arto Hellas |
J. Web Eng. | 3 |
| 2024 | From Sparse to Smart: Leveraging AI for Effective Online Judge Problem Classification in Programming Education
Filipe D. Pereira, Maely Moraes, Marcelo Henrique Oliveira Henklain, Arto Hellas, Elaine Oliveira, Dragan Gasevic, Raimundo S. Barreto, Rafael Ferreira Leite de Mello |
EC-TEL (1) | 4 |
| 2024 | Leveraging Large Language Models for Next-Generation Educational Technologies
Neil T. Heffernan, Rose E. Wang, Christopher J. MacLellan, Arto Hellas, Chenglu Li, Candace A. Walkington, Joshua Littenberg-Tobias, David Joyner, Steven Moore, Adish Singla, Zachary A. Pardos, Maciej Pankiewicz, Juho Kim 0001, Shashank Sonkar, Clayton Cohn, Anthony Botelho, Andrew S. Lan, Mingyu Feng, Tanja Käser, Eamon Worden |
EDM | 4 |
| 2024 | Evaluating Contextually Personalized Programming Exercises Created with Generative AIabstractProgramming skills are typically developed through completing various hands-on exercises. Such programming problems can be contextualized to students’ interests and cultural backgrounds. Prior research in educational psychology has demonstrated that context personalization of exercises stimulates learners’ situational interests and positively affects their engagement. However, creating a varied and comprehensive set of programming exercises for students to practice on is a time-consuming and laborious task for computer science educators. Previous studies have shown that large language models can generate conceptually and contextually relevant programming exercises. Thus, they offer a possibility to automatically produce personalized programming problems to fit students’ interests and needs. This article reports on a user study conducted in an elective introductory programming course that included contextually personalized programming exercises created with GPT-4. The quality of the exercises was evaluated by both the students and the authors. Additionally, this work investigated student attitudes towards the created exercises and their engagement with the system. The results demonstrate that the quality of exercises generated with GPT-4 was generally high. What is more, the course participants found them engaging and useful. This suggests that AI-generated programming problems can be a worthwhile addition to introductory programming courses, as they provide students with a practically unlimited pool of practice material tailored to their personal interests and educational needs. Evanfiya Logacheva, Arto Hellas, James Prather, Sami Sarsa, Juho Leinonen 0001 |
ICER (1) | 2 |
| 2024 | "Like a Nesting Doll": Analyzing Recursion Analogies Generated by CS Students Using Large Language ModelsabstractGrasping complex computing concepts often poses a challenge for students who struggle to anchor these new ideas to familiar experiences and understandings. To help with this, a good analogy can bridge the gap between unfamiliar concepts and familiar ones, providing an engaging way to aid understanding. However, creating effective educational analogies is difficult even for experienced instructors. We investigate to what extent large language models (LLMs), specifically ChatGPT, can provide access to personally relevant analogies on demand. Focusing on recursion, a challenging threshold concept, we conducted an investigation analyzing the analogies generated by more than 350 first-year computing students. They were provided with a code snippet and tasked to generate their own recursion-based analogies using ChatGPT, optionally including personally relevant topics in their prompts. We observed a great deal of diversity in the analogies produced with student-prescribed topics, in contrast to the otherwise generic analogies, highlighting the value of student creativity when working with LLMs. Not only did students enjoy the activity and report an improved understanding of recursion, but they described more easily remembering analogies that were personally and culturally relevant. Seth Bernstein, Paul Denny 0001, Juho Leinonen 0001, Lauren Kan, Arto Hellas, Matt Littlefield, Sami Sarsa, Stephen MacNeil |
ITiCSE (1) | 5 |
| 2024 | Analyzing Students' Preferences for LLM-Generated AnalogiesabstractIntroducing students to new concepts in computer science can often be challenging, as these concepts may differ significantly from their existing knowledge and conceptual understanding. To address this, we employed analogies to help students connect new concepts to familiar ideas. Specifically, we generated analogies using large language models (LLMs), namely ChatGPT, and used them to help students make the necessary connections. In this poster, we present the results of our survey, in which students were provided with two analogies relating to different computing concepts, and were asked to describe the extent to which they were accurate, interesting, and useful. This data was used to determine how effective LLM-generated analogies can be for teaching computer science concepts, as well as how responsive students are to this approach. Seth Bernstein, Paul Denny 0001, Juho Leinonen 0001, Matt Littlefield, Arto Hellas, Stephen MacNeil |
ITiCSE (2) | 5 |
| 2024 | Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-JudgeabstractLarge language models (LLMs) have shown great potential for the automatic generation of feedback in a wide range of computing contexts. However, concerns have been voiced around the privacy and ethical implications of sending student work to proprietary models. This has sparked considerable interest in the use of open source LLMs in education, but the quality of the feedback that such open models can produce remains understudied. This is a concern as providing flawed or misleading generated feedback could be detrimental to student learning. Inspired by recent work that has utilised very powerful LLMs, such as GPT-4, to evaluate the outputs produced by less powerful models, we conduct an automated analysis of the quality of the feedback produced by several open source models using a dataset from an introductory programming course. First, we investigate the viability of employing GPT-4 as an automated evaluator by comparing its evaluations with those of a human expert. We observe that GPT-4 demonstrates a bias toward positively rating feedback while exhibiting moderate agreement with human raters, showcasing its potential as a feedback evaluator. Second, we explore the quality of feedback generated by several leading open-source LLMs by using GPT-4 to evaluate the feedback. We find that some models offer competitive performance with popular proprietary LLMs, such as ChatGPT, indicating opportunities for their responsible use in educational settings. Charles Koutcheme, Nicola Dainese, Sami Sarsa, Arto Hellas, Juho Leinonen 0001, Paul Denny 0001 |
ITiCSE (1) | 4 |
| 2024 | Propagating Large Language Models Programming FeedbackabstractLarge language models (LLMs) such as GPT-4 have emerged as promising tools for providing programming feedback. However, effective deployment of LLMs in massive classes and Massive Open Online Courses (MOOCs) raises financial concerns, calling for methods to minimize the number of calls to the APIs and systems serving such powerful models. In this article, we revisit the problem of 'propagating feedback' within the contemporary landscape of LLMs. Specifically, we explore feedback propagation as a way to reduce the cost of leveraging LLMs for providing programming feedback at scale. Our study investigates the effectiveness of this approach in the context of students requiring next-step hints for Python programming problems, presenting initial results that support the viability of the approach. We discuss our findings' implications and suggest directions for future research in optimizing feedback mechanisms for large-scale educational environments. Charles Koutcheme, Arto Hellas |
L@S | 2 |
| 2024 | Using Large Language Models for Teaching ComputingabstractIn the past year, large language models (LLMs) have taken the world by storm, demonstrating their potential as a transformative force in many domains including computing education. Computing education researchers have found that LLMs can solve most assessments in introductory programming courses, including both traditional code writing tasks and other popular tasks such as Parsons problems. As more and more students start to make use of LLMs, the question instructors might ask themselves is "what can I do?". We propose that one promising way forward is to integrate LLMs into teaching practice, providing all students with an equal opportunity to learn how to interact productively with LLMs as well as encounter and understand their limitations. In this workshop, we first present state-of-the-art research results on how to utilize LLMs in computing education practice, after which participants will take part in hands-on activities using LLMs. We end the workshop by brainstorming ideas with participants around adapting their classrooms to most effectively integrate LLMs while avoiding some common pitfalls. Juho Leinonen 0001, Stephen MacNeil, Paul Denny 0001, Arto Hellas |
SIGCSE (2) | 4 |
| 2024 | Discussing the Changing Landscape of Generative AI in Computing EducationabstractIn a previous Birds of a Feather discussion, we delved into the nascent applications of generative AI, contemplating its potential and speculating on future trajectories. Since then, the landscape has continued to evolve revealing the capabilities and limitations of these models. Despite this progress, the computing education research community still faces uncertainty around pivotal aspects such as (1) academic integrity and assessments, (2) curricular adaptations, (3) pedagogical strategies, and (4) the competencies students require to instill responsible use of these tools. The goal of this Birds of a Feather discussion is to unravel these pressing and persistent issues with computing educators and researchers, fostering a collaborative exploration of strategies to navigate the educational implications of advancing generative AI technologies. Aligned with this goal of building an inclusive learning community, our BoF is led by globally distributed leaders to facilitate multiple coordinated discussions that can lead to a broader conversation about the role of LLMs in CS education. Stephen MacNeil, Juho Leinonen 0001, Paul Denny 0001, Natalie Kiesler, Arto Hellas, James Prather, Brett A. Becker, Michel Wermelinger, Karen Reid |
SIGCSE (2) | 5 |
| 2024 | Solving Proof Block Problems Using Large Language ModelsabstractLarge language models (LLMs) have recently taken many fields, including computer science, by storm. Most recent work on LLMs in computing education has shown that they are capable of solving most introductory programming (CS1) exercises, exam questions, Parsons problems, and several other types of exercises and questions. Some work has investigated the ability of LLMs to solve CS2 problems as well. However, it remains unclear how well LLMs fare against more advanced upper-division coursework, such as proofs in algorithms courses. After all, while known to be proficient in many programming tasks, LLMs have been shown to have more difficulties in forming mathematical proofs. Seth Poulsen, Sami Sarsa, James Prather, Juho Leinonen 0001, Brett A. Becker, Arto Hellas, Paul Denny 0001, Brent N. Reeves |
SIGCSE (1) | 6 |
| 2024 | Instructor Perceptions of AI Code Generation Tools - A Multi-Institutional Interview StudyabstractMuch of the recent work investigating large language models and AI Code Generation tools in computing education has focused on assessing their capabilities for solving typical programming problems and for generating resources such as code explanations and exercises. If progress is to be made toward the inevitable lasting pedagogical change, there is a need for research that explores the instructor voice, seeking to understand how instructors with a range of experiences plan to adapt. In this paper, we report the results of an interview study involving 12 instructors from Australia, Finland and New Zealand, in which we investigate educators' current practices, concerns, and planned adaptations relating to these tools. Through this empirical study, our goal is to prompt dialogue between researchers and educators to inform new pedagogical strategies in response to the rapidly evolving landscape of AI code generation tools. Judithe Sheard, Paul Denny 0001, Arto Hellas, Juho Leinonen 0001, Lauri Malmi, Simon |
SIGCSE (1) | 3 |
| 2023 | Automated Program Repair Using Generative Models for Code Infilling
Charles Koutcheme, Sami Sarsa, Juho Leinonen 0001, Arto Hellas, Paul Denny 0001 |
AIED | 4 |
| 2023 | Exploring the Responses of Large Language Models to Beginner Programmers' Help RequestsabstractBackground and Context: Over the past year, large language models (LLMs) have taken the world by storm. In computing education, like in other walks of life, many opportunities and threats have emerged as a consequence. Arto Hellas, Juho Leinonen 0001, Sami Sarsa, Charles Koutcheme, Lilja Koivuniemi, Juha Sorva |
ICER (1) | 1 |
| 2023 | Evaluating Distance Measures for Program RepairabstractBackground and Context: Struggling with programming assignments while learning to program is a common phenomenon in programming courses around the world. Supporting struggling students is a common theme in Computing Education Research (CER), where a wide variety of support methods have been created and evaluated. An important stream of research here focuses on program repair, where methods for automatically fixing erroneous code are used for supporting students as they debug their code. Work in this area has so far assessed the performance of the methods by evaluating the closeness of the proposed fixes to the original erroneous code. The evaluations have mainly relied on the use of edit distance measures such as the sequence edit distance and there is a lack of research on which distance measure is the most appropriate. Charles Koutcheme, Sami Sarsa, Juho Leinonen 0001, Lassi Haaranen, Arto Hellas |
ICER (1) | 5 |
| 2023 | The Rise of Disappearing Frameworks in Web Development
Juho Vepsäläinen, Arto Hellas, Petri Vuorimaa |
ICWE | 2 |
| 2023 | Comparing Code Explanations Created by Students and Large Language ModelsabstractReasoning about code and explaining its purpose are fundamental skills for computer scientists. There has been extensive research in the field of computing education on the relationship between a student's ability to explain code and other skills such as writing and tracing code. In particular, the ability to describe at a high-level of abstraction how code will behave over all possible inputs correlates strongly with code writing skills. However, developing the expertise to comprehend and explain code accurately and succinctly is a challenge for many students. Existing pedagogical approaches that scaffold the ability to explain code, such as producing exemplar code explanations on demand, do not currently scale well to large classrooms. The recent emergence of powerful large language models (LLMs) may offer a solution. In this paper, we explore the potential of LLMs in generating explanations that can serve as examples to scaffold students' ability to understand and explain code. To evaluate LLM-created explanations, we compare them with explanations created by students in a large course (n ≈ 1000) with respect to accuracy, understandability and length. We find that LLM-created explanations, which can be produced automatically on demand, are rated as being significantly easier to understand and more accurate summaries of code than student-created explanations. We discuss the significance of this finding, and suggest how such models can be incorporated into introductory programming education. Juho Leinonen 0001, Paul Denny 0001, Stephen MacNeil, Sami Sarsa, Seth Bernstein, Joanne Kim, Andrew Tran, Arto Hellas |
ITiCSE (1) | 8 |
| 2023 | Seeing Program Output Improves Novice Learning GainsabstractIn this article, we report results from a randomized controlled trial where novice programmers completed code mimicking exercises -- writing and modifying code shown to them -- designed to help learn the basics of how variables work. Using a tailored code writing system with feedback on program correctness, we conducted a two-group design study where only one of the groups could see the program output and feedback on the correctness of the program they wrote, while the other group just saw feedback on correctness. Learning gain was measured using a code-reading multiple choice questionnaire as both a pretest and a posttest. Our data suggests that being able to see program output leads to higher learning gains for novices, when compared to just being able to see feedback on the correctness of the code. For more experienced students, we observed benefits from code mimicking in both groups, without a strong distinction between being able to see the output and not being able to see the output. Based on our experiment, we recommend that environments used by novices for learning programming should encourage -- or even require -- running the code before allowing submitting the program for assessment. Juho Leinonen 0001, Arto Hellas, John Edwards 0002 |
ITiCSE (1) | 2 |
| 2023 | Evaluating the Performance of Code Generation Models for Solving Parsons Problems With Small Prompt VariationsabstractThe recent emergence of code generation tools powered by large language models has attracted wide attention. Models such as OpenAI Codex can take natural language problem descriptions as input and generate highly accurate source code solutions, with potentially significant implications for computing education. Given the many complexities that students face when learning to write code, they may quickly become reliant on such tools without properly understanding the underlying concepts. One popular approach for scaffolding the code writing process is to use Parsons problems, which present solution lines of code in a scrambled order. These remove the complexities of low-level syntax, and allow students to focus on algorithmic and design-level problem solving. It is unclear how well code generation models can be applied to solve Parsons problems, given the mechanics of these models and prior evidence that they underperform when problems include specific restrictions. In this paper, we explore the performance of the Codex model for solving Parsons problems over various prompt variations. Using a corpus of Parsons problems we sourced from the computing education literature, we find that Codex successfully reorders the problem blocks about half of the time, a much lower rate of success when compared to prior work on more free-form programming tasks. Regarding prompts, we find that small variations in prompting have a noticeable effect on model performance, although the effect is not as pronounced as between different problems. Brent N. Reeves, Sami Sarsa, James Prather, Paul Denny 0001, Brett A. Becker, Arto Hellas, Bailey Kimmel, Garrett B. Powell, Juho Leinonen 0001 |
ITiCSE (1) | 6 |
| 2023 | Using Large Language Models to Enhance Programming Error MessagesabstractA key part of learning to program is learning to understand programming error messages. They can be hard to interpret and identifying the cause of errors can be time-consuming. One factor in this challenge is that the messages are typically intended for an audience that already knows how to program, or even for programming environments that then use the information to highlight areas in code. Researchers have been working on making these errors more novice friendly since the 1960s, however progress has been slow. The present work contributes to this stream of research by using large language models to enhance programming error messages with explanations of the errors and suggestions on how to fix them. Large language models can be used to create useful and novice-friendly enhancements to programming error messages that sometimes surpass the original programming error messages in interpretability and actionability. These results provide further evidence of the benefits of large language models for computing educators, highlighting their use in areas known to be challenging for students. We further discuss the benefits and downsides of large language models and highlight future streams of research for enhancing programming error messages. Juho Leinonen 0001, Arto Hellas, Sami Sarsa, Brent N. Reeves, Paul Denny 0001, James Prather, Brett A. Becker |
SIGCSE (1) | 2 |
| 2023 | Time-constrained Code Recall Tasks for Monitoring the Development of Programming PlansabstractProgrammers rely on the recognition and utilization of reoccurring code sequences to understand and create code. Knowledge of these sequences --programming plans -- has been shown to be a factor that differentiates novice programmers from experts. Although the information on the development of programming plans would be beneficial to both teachers and students, explicitly following their development over a longer time period is scarce. In this article, we describe an easy-to-apply methodology for monitoring the development of programming plans. The development of programming plans is evaluated with time-constrained code recall tasks, where students are shown snippets of code for a short period of time, after which they write the snippets they saw. To determine the existence of programming plans, the short duration is designed so that reading the shown code is not feasible in the given time period. We demonstrate the methodology through an experiment in which we studied the development of programming plans in students in a beginner web programming course. Ava Heinonen, Arto Hellas |
SIGCSE (1) | 2 |
| 2023 | The Implications of Large Language Models for CS Teachers and StudentsabstractThe introduction of Large Language Models (LLMs) has generated a significant amount of excitement both in industry and among researchers. Recently, tools that leverage LLMs have made their way into the classroom where they help students generate code and help instructors generate learning materials. There are likely many more uses of these tools -- both beneficial to learning and possibly detrimental to learning. To help ensure that these tools are used to enhance learning, educators need to not only be familiar with these tools, but with their use and potential misuse. The goal of this BoF is to raise awareness about LLMs and to build a learning community around their use in computing education. Aligned with this goal of building an inclusive learning community, our BoF is led by globally distributed discussion leaders, including undergraduate researchers, to facilitate multiple coordinated discussions that can lead to a broader conversation about the role of LLMs in CS education. Stephen MacNeil, Joanne Kim, Juho Leinonen 0001, Paul Denny 0001, Seth Bernstein, Brett A. Becker, Michel Wermelinger, Arto Hellas, Andrew Tran, Sami Sarsa, James Prather, Viraj Kumar |
SIGCSE (2) | 8 |
| 2023 | Automatically Generating CS Learning Materials with Large Language ModelsabstractRecent breakthroughs in Large Language Models (LLMs), such as GPT-3 and Codex, now enable software developers to generate code based on a natural language prompt. Within computer science education, researchers are exploring the potential for LLMs to generate code explanations and programming assignments using carefully crafted prompts. These advances may enable students to interact with code in new ways while helping instructors scale their learning materials. However, LLMs also introduce new implications for academic integrity, curriculum design, and software engineering careers. This workshop will demonstrate the capabilities of LLMs to help attendees evaluate whether and how LLMs might be integrated into their pedagogy and research. We will also engage attendees in brainstorming to consider how LLMs will impact our field. Stephen MacNeil, Andrew Tran, Juho Leinonen 0001, Paul Denny 0001, Joanne Kim, Arto Hellas, Seth Bernstein, Sami Sarsa |
SIGCSE (2) | 6 |
| 2023 | Experiences from Using Code Explanations Generated by Large Language Models in a Web Software Development E-BookabstractAdvances in natural language processing have resulted in large language models (LLMs) that can generate code and code explanations. In this paper, we report on our experiences generating multiple code explanation types using LLMs and integrating them into an interactive e-book on web software development. Three different types of explanations -- a line-by-line explanation, a list of important concepts, and a high-level summary of the code -- were created. Students could view explanations by clicking a button next to code snippets, which showed the explanation and asked about its utility. Our results show that all explanation types were viewed by students and that the majority of students perceived the code explanations as helpful to them. However, student engagement varied by code snippet complexity, explanation type, and code snippet length. Drawing on our experiences, we discuss future directions for integrating explanations generated by LLMs into CS classrooms. Stephen MacNeil, Andrew Tran, Arto Hellas, Joanne Kim, Sami Sarsa, Paul Denny 0001, Seth Bernstein, Juho Leinonen 0001 |
SIGCSE (1) | 3 |
| 2023 | Implications of Edge Computing for Static Site Generation
Juho Vepsäläinen, Arto Hellas, Petri Vuorimaa |
WEBIST | 2 |
| 2023 | The State of Disappearing Frameworks in 2023
Juho Vepsäläinen, Arto Hellas, Petri Vuorimaa |
WEBIST | 2 |
| 2023 | Synthesizing research on programmers' mental models of programs, tasks and concepts - A systematic literature reviewabstractProgrammers’ mental models represent their knowledge and understanding of programs, programming concepts, and programming in general. They guide programmers’ work and influence their task performance. Understanding mental models is important for designing work systems and practices that support programmers. Although the importance of programmers’ mental models is widely acknowledged, research on mental models has decreased over the years. The results are scattered and do not take into account recent developments in software engineering. In this article, we analyze the state of research on programmers’ mental models and provide an overview of existing research. We connect results on mental models from different strands of research to form a more unified knowledge base on the topic. We conducted a systematic literature review on programmers’ mental models. We analyzed literature addressing mental models in different contexts, including mental models of programs, programming tasks, and programming concepts. Using nine search engines, we found 3678 articles (excluding duplicates). Of these, 84 were selected for further analysis. Using the snowballing technique, starting from these 84, we obtained a final result set containing 187 articles. We show that the literature shares a kernel of shared understanding of mental models. By collating and connecting results on mental models from different fields of research, we provide a comprehensive synthesis of results related to programmers’ mental models. The research field on programmers’ mental models faces many challenges arising from a lack of a shared knowledge base and poorly defined constructs. By creating a unified knowledge base on the topic, this work provides a basis for future work on mental models. We also point to directions for future studies. In particular, we call for studies that examine programmers working with modern practices and tools. Ava Heinonen, Bettina Lehtelä, Arto Hellas, Fabian Fagerholm |
Inf. Softw. Technol. | 3 |
| 2022 | Evaluating CodeClusters for Effectively Providing Feedback on Code SubmissionsabstractFull research paper—Most introductory programming courses rely on the use of automated assessment for grading programming assignments. While such systems save teachers’ time by eliminating manual grading, the submissions may not be reviewed by the teachers at all, losing valuable insight into how students solve the assignments. In this paper, we introduce CodeClusters which provides teachers’ a quick overview of general patterns in code submissions. The main features of the system are full-text search and N-gram -based similarity detection model that can cluster and subset the code by various aspects such as AST similarity or software also has It has also an interface for streamlining the writing of feedback to multiple submissions at once. CodeClusters has been primarily designed to be used jointly with automated assessment systems where automated tests would assess the functionality of the code, and CodeClusters would be used for gaining a higher-level view of programming patterns and for writing feedback to students. CodeClusters was evaluated in a think-aloud study with university lecturers responsible for programming courses at two research-first universities. The lecturers were pleased by the possibility of providing better feedback to students quickly, saw it could improve the quality of their courses over solely automated assessment, and expressed interest in using CodeClusters in their own courses. Teemu Koivisto, Arto Hellas |
FIE | 2 |
| 2022 | Piloting Natural Language Generation for Personalized Progress FeedbackabstractFull research paper—We describe the results of a pilot study wherein we applied simple natural language generation methods to produce automated feedback for students of an online course based on student high-level progress data. Experimenting with both personalized and non-personalized feedback, we show that such feedback can be easily produced given access to even rudimentary data regarding student assignment submissions and their correctness. Our results suggest that students perceive automatically generated feedback generally positively and believe it to be useful. Our results also indicate that minor personalization and stylistic alterations in the feedback can have meaningful effects on how the feedback is interacted with and perceived. In particular, we observe that personalized feedback is perceived as being slightly easier to understand and as being better aligned with their progress. Students also felt better about the personalized feedback in comparison to non-personalized feedback. We conclude that the automated generation of personalized textual feedback shows promise as a low-threshold way of increasing student satisfaction. Further research is needed to assess the effect of different types of automated personalized feedback on student performance and behavior. Leo Leppänen, Arto Hellas, Juho Leinonen 0001 |
FIE | 2 |
| 2022 | On Things that Matter in Learning Programming: Towards a Scale for New Programming StudentsabstractFull research paper—In this paper, we report on the development of a succinct and easy-to-administer 11-item scale that quantifies students’ self-efficacy, social aspect, independence, and meaning of studies, with a focus on introductory programming studies. The scale has been constructed using exploratory factor analysis of survey response data collected from students attending introductory programming courses offered by two universities. We evaluate the scale by using it to examine differences between university contexts, and assess to what extent the scale relates to students’ perceived impact of the COVID-19 pandemic on studies, prior programming experience, self-assessed competence, and seeking help. Our evaluation of the scale suggests that social aspect was correlated with being more strongly influenced by the COVID-19 pandemic, while the perceived ability to work independently was correlated with reduced influence of the COVID-19 pandemic. Prior programming experience was positively correlated with self-perceived ability to work independently and with self-efficacy. Similarly, self-estimated competence was positively correlated with self-efficacy. Finally, social aspect and meaning of studies were positively correlated with help-seeking. Our evaluations show that the scale holds promise as a new tool for researchers and practitioners seeking to improve understanding of their study contexts. Hannu Pesonen, Arto Hellas |
FIE | 2 |
| 2022 | How to Help to Ask for Help? Help Request Prompt Structure Influence on Help Request Quantity and Course RetentionabstractFull research paper—Feedback and support are at the core of efficient learning. While the effect of feedback has been explored in a multitude of studies, the effect of asking for help and the effect of how that help is asked for is an under-explored area. In this work, we present the results of a randomized controlled trial organized in an introductory programming course for lifelong learners. In the study, we explored the usefulness of different types of help prompts used to guide learners asking for help. Gauging the effect of different combinations of three prompts used to scaffold writing out the help request, (1) "Describe the issue with your program", (2) "Explain how your program works", and (3) "What have you tried to do to resolve the issue", we study how prompts influence learners’ behavior. Using log data collected from the online platform where the prompts were explored, we study how the prompts affect whether learners end up sending a help request instead of simply considering to send one, how the scaffolding prompts affect whether learners send further help requests, and whether the questions have an effect on course retention. Our results show that the help request prompts have an impact on whether learners end up actually writing a help request, which in turn also influences whether the learner will ask for help on another occasion. Further, we observe that help request prompts may have an effect on whether learners figure out a solution to a problem on their own. Alarmingly, we also observe that the prompts can affect how far learners go in a course. Based on our results, we advise teachers and researchers to pay attention into how learners are guided into asking for help. Sami Sarsa, Jesper Pettersson, Arto Hellas |
FIE | 3 |
| 2022 | Automatic Generation of Programming Exercises and Code Explanations Using Large Language ModelsabstractThis article explores the natural language generation capabilities of large language models with application to the production of two types of learning resources common in programming courses. Using OpenAI Codex as the large language model, we create programming exercises (including sample solutions and test cases) and code explanations, assessing these qualitatively and quantitatively. Our results suggest that the majority of the automatically generated content is both novel and sensible, and in some cases ready to use as is. When creating exercises we find that it is remarkably easy to influence both the programming concepts and the contextual themes they contain, simply by supplying keywords as input to the model. Our analysis suggests that there is significant value in massive generative machine learning models as a tool for instructors, although there remains a need for some oversight to ensure the quality of the generated content before it is delivered to students. We further discuss the implications of OpenAI Codex and similar tools for introductory programming education and highlight future research streams that have the potential to improve the quality of the educational experience for both teachers and students alike. Sami Sarsa, Paul Denny 0001, Arto Hellas, Juho Leinonen 0001 |
ICER (1) | 3 |
| 2022 | Planning a Multi-institutional and Multi-national Study of the Effectiveness of Parsons ProblemsabstractProgramming is a complex task that requires the development of many skills including knowledge of syntax, problem decomposition, algorithm development, and debugging. Code-writing activities are commonly used to help students develop these skills, but the difficulty of writing code from a blank page can overwhelm many novices. Parsons problems offer a simpler alternative to writing code by providing scrambled code blocks that must be placed in the correct order to solve a problem. The extensive literature on Parsons problems documents numerous benefits to using them as both formative and summative assessments. These include more efficient learning, the possibility to dynamically adapt to learner needs, and more reliable grading. Despite these positive findings, further research is needed in order to draw broader inferences. Most work has been conducted at single institutions under unique conditions that are not easily replicated, and some prior studies have been inconclusive or had limitations that affected data validity. To address this, we propose a multi-institutional and multi-national study of the effectiveness of Parsons problems for novice programmers. We will focus on introductory programming courses (CS0/1/2) that use Java, Python, and C/C++ as these are the most common teaching languages. The working group will collaborate to refine the scope, methodology and research questions, and contribute to data collection and analysis. Barbara Ericson, Paul Denny 0001, James Prather, Rodrigo Duran 0001, Arto Hellas, Juho Leinonen 0001, Craig S. Miller, Briana B. Morrison, Janice L. Pearce, Susan H. Rodger |
ITiCSE (2) | 5 |
| 2022 | Exploring How Students Solve Open-ended Assignments: A Study of SQL Injection Attempts in a Cybersecurity CourseabstractResearch into computing and learning how to program has been ongoing for decades. Commonly, this research has been focused on novice learners and the difficulties they encounter, especially during CS1. Cybersecurity is a critical aspect in computing -- as a topic in university education as well as a core skill in the industry. In this study, we investigate how students solve open-ended assignments on a cybersecurity course offered to university students after two years of CS studies. Specifically, we looked at how students perform SQL injection attacks on an web application system, and study to what extent we can characterize the process in which they come up with successful injections. Our results show that there are distinguishable strategies used by individual students who seek to hack the system, where these approaches revolve around exploration and exploitation tactics. We also find evidence of learning due to a more pronounced use of exploitation in a subsequent similar assignment. Charles Koutcheme, Artturi Tilanterä, Aleksi Peltonen, Arto Hellas, Lassi Haaranen |
ITiCSE (1) | 4 |
| 2022 | Who Continues in a Series of Lifelong Learning Courses?abstractAlthough computing education research quite often targets within-university courses, an important role of universities is educating the public through open online lifelong learning offerings. Compared to within-university courses, in lifelong learning, the student population is often more diverse. For example, participants often have more varied motivations and aspirations as well as more varied educational backgrounds. In this work, we explore what kinds of learners attend open online lifelong learning programming courses and what characteristics of learners lead to completing courses and proceeding to subsequent courses. We examine student-related factors collected through surveys in our online course environment. These factors include motivation, previous experience, and demographics. Our results show that motivations, previous experience, and demographics by themselves only explain a small amount of the variance in completing courses or continuing to a subsequent course. At the same time, we identify individual factors that are more likely to lead to learners dropping out (or continuing) in the courses. Our study provides further evidence that lifelong learning benefits most the already educated part of the population with prior knowledge and high motivation. This calls for further studies that seek to identify means to engage and support participants less likely to continue in such courses. Sami Sarsa, Arto Hellas, Juho Leinonen 0001 |
ITiCSE (1) | 2 |
| 2022 | Time-on-Task Metrics for Predicting Performanceabstract\emphTime-on-task is one key contributor to learning. However, how time-on-task is measured often varies, and is limited by the available data. In this work, we study two different time-on-task metrics---derived from programming process data---for predicting performance in an introductory programming course. The first metric, coarse-grained time-on-task, is based on students' submissions to programming assignments; the second, fine-grained time-on-task, is based on the keystrokes that students take while constructing their programs. Both types of time-on-task metrics have been used in prior work, and are supposedly designed to measure the same underlying feature: time-on-task. However, previous work has found that the correlation between these two metrics is not as high as one might expect. We build on that work by analyzing how well the two metrics work for predicting students' performance in an introductory programming course. Our results suggest that the correlation between the fine-grained time-on-task metric and both weekly exercise points and exam points is higher than the correlation between the coarse-grained time-on-task metric and weekly exercise points and exam points. Furthermore, we show that the fine-grained time-on-task metric is a better predictor of students' future success in the course exam than the coarse-grained time-on-task metric. We thus propose that future work utilizing time-on-task as a predictor of performance should use as fine-grained data as possible to measure time-on-task if such data is available. Juho Leinonen 0001, Francisco Enrique Vicente Castro, Arto Hellas |
SIGCSE (1) | 3 |
| 2021 | Fine-Grained Versus Coarse-Grained Data for Estimating Time-on-Task in Learning Programming
Juho Leinonen 0001, Francisco Enrique Vicente Castro, Arto Hellas |
EDM | 3 |
| 2021 | Algorithm Visualization and the Elusive Modality EffectabstractThe modality effect in multimedia learning suggests that pictures are best accompanied by audio explanations rather than text, but this finding has not been replicated in computing education. We investigate which instructional modality works best as an accompaniment for algorithm visualizations. In a randomized controlled trial, learners were split into three conditions who viewed an instructional video on Dijkstra’s algorithm, with diagrams accompanied by audio, text, or both. We find neither a modality effect in favor of the audio condition nor a verbal redundancy effect in favor of using only a single modality rather than both. Taken together with earlier research, our findings suggest that the modality effect is difficult to apply reliably and computing educators should not rush to integrate audio into visualizations in expectation of the effect. We discuss theoretical viewpoints that future research should attend to; these include alternative part-explanations of the modality effect and attention-based models of working memory, among others. Albina Zavgorodniaia, Artturi Tilanterä, Ari Korhonen, Otto Seppälä, Arto Hellas, Juha Sorva |
ICER | 5 |
| 2021 | Does the Early Bird Catch the Worm? Earliness of Students' Work and its Relationship with Course OutcomesabstractIntuitively, it seems plausible that students who start their work earlier and work on more days than their peers should perform better in any course. But does the early bird really catch the worm? In this article, we examine introductory programming students' time management behavior as evidenced by data collected from a programming environment. We analyze: 1) the earliness of students' work, i.e. when they start working on their course assignments, 2) the number of days students work on course assignments, and 3) the relationship between earliness, the number of days worked, and course outcomes. Our results provide further support for the notion that, on average, students who start working on course assignments early perform slightly better in the course. At the same time, we found that starting early does not necessarily mean that students work on more days, and that starting early and working on many days does not necessarily mean that students get better grades. In addition, some students who start working early on the assignments in the first weeks of the course seem to start delaying when they begin working on assignments as the course progresses, while other students seem to be able to continue starting early throughout the course. Juho Leinonen 0001, Francisco Enrique Vicente Castro, Arto Hellas |
ITiCSE (1) | 3 |
| 2020 | Relation of Individual Time Management Practices and Time Management of TeamsabstractFull research paper-Team configuration, work practices, and communication have a considerable impact on the outcomes of student software projects. This study observes 150 college students who first individually solve exercises and then carry out a class project in teams of three. All projects had the same requirements. We analyzed how students' behavior on individual pre-project exercises predict team project outcomes, investigated how students' time management practices affected other team members, and analyzed how students divided their work among peers. Our results indicate that teams consisting of only low-performing students were the most dysfunctional in terms of workload balance, whereas teams with both low-and high-performing students performed almost as well as teams consisting of only high-performing students. This suggests that teams should combine students of varying skill levels rather than allowing teams with only low performers or letting students to form teams without constraints. We also observed that individual students' poor time management practices impair their teammates' time management. This underlines the importance of encouraging good time management practices. Most teams reported that they divided tasks in a way that is beneficial for the acquisition of technical skills rather than collaboration and communication skills. Only a few teams assigned tasks so that students would have worked only on tasks they already knew and thus felt most comfortable to work with. Tapio Auvinen, Nick Falkner, Arto Hellas, Petri Ihantola, Ville Karavirta, Otto Seppälä |
FIE | 3 |
| 2020 | On the Differences in Time That Students Take to Write Solutions to Programming ProblemsabstractFull research paper-In this work, we study productivity differences in an introductory programming course. Focusing on a set of students who completed all programming assignments in the course, we quantify differences in productivity, measured through the time spent on completing the assignments. We focus both on the overall time needed to complete all programming assignments in the course, as well as on time spent on individual programming assignments. In addition, the effect of previous programming experience and difficulty of the programming assignment is considered. Our results show significant productivity differences between students. In addition, while programming experience influences productivity, a proportion of students who have never programmed before are faster in completing the programming assignments than students with considerable amounts of previous programming experience. Our results suggest that the classic credit-based or lecture hour based workload estimates of a course fit poorly to the whole course population in programming, suggesting that programming courses and training should be adjusted based on the participant. Fabian Fagerholm, Arto Hellas |
FIE | 2 |
| 2020 | Deadlines and MOOCs: How Do Students Behave in MOOCs with and without DeadlinesabstractFull research paper—Online education can be delivered in many ways. For example, some MOOCs let students to proceed with their own pace, while others rely on strict schedules. Although the variety of how MOOCs can be organized is generally well understood, less is known about how the different ways of organizing MOOCs affect retention. In this work, we compare self-paced and fixed-schedule MOOCs in terms of retention and work-load. Using data from over 8.000 students participating in two versions of a massive open online course in programming, we observe that drop-out rates at the beginning of the courses are greater than towards the end of the courses, with self-paced MOOC being more extreme in this respect. Mostly because of different starts, the fixed-schedule course has a better overall retention rate (45%) than its self-paced counterpart (13%). We hypothesize that students initial investment of time and effort contributes to their persistence in their course, meaning that they do not want to let their initial investment go to waste. At the same time, in both self-paced and fixed-schedule MOOCs, there are students who receive almost full points from one week but fail to continue to the next week. This suggests that the issue of dropouts in MOOCs may also be related to participants struggling to take up new tasks or schedule their work over a longer time period. Our results support scheduling student activities in open online courses and opens up new research directions in engaging students in self-paced courses. Petri Ihantola, Ilenia Fronza, Tommi Mikkonen, Miska Noponen, Arto Hellas |
FIE | 5 |
| 2020 | Study Major, Gender, and Confidence Gap: Effects on Experience, Performance, and Self-Efficacy in Introductory ProgrammingabstractThe term Confidence Gap refers to the phenomenon of men being more confident in their ability to succeed in their studies and elsewhere. It is an acknowledged phenomenon both in Computer Science as well as STEM subjects at large, likely influencing students' career path choices and selection of study major. In this work, we analyze data from multiple introductory programming courses. We do this by looking at the interaction of (1) students' performance measured in terms of completed assignments, (2) self-reported confidence in the ability to succeed in the programming course, (3) major, and (4) gender. Aligned with prior research, we observe the existence of the Confidence Gap. At the same time, men and women who chose Computer Science as their major are more confident in their ability to succeed in their first programming course than their counterparts in other subjects. Nea Pirttinen, Arto Hellas, Lassi Haaranen, Rodrigo Duran 0001 |
FIE | 2 |
| 2020 | Programming Versus Natural Language: On the Effect of Context on Typing in CS1abstractAnalyzing keystroke data from students working on essay and programming tasks, we study to what extent the difference in task context influences performance in typing. Using data from two introductory programming courses offered at two separate institutions, we compare and contrast typing speed between programming and natural language tasks. We observe that students tend to be faster at typing (the same) character pairs when writing natural language text than when learning to write code. We show that students improve on typing character pairs that appear in frequently used words in programming languages, and that typing programming constructs also improves. We find that students are faster at detecting and erasing their mistakes when typing natural language text than when programming. Our results support theories regarding contextual memory, procedural memory, and practice, and have implications for course curriculum and pedagogy design. John Edwards 0002, Juho Leinonen 0001, Chetan Birthare, Albina Zavgorodniaia, Arto Hellas |
ICER | 5 |
| 2020 | Teaching Container-Based DevOps Practices
Jami Kousa, Petri Ihantola, Arto Hellas, Matti Luukkainen |
ICWE | 3 |
| 2020 | Capturing and Characterising Notional MachinesabstractA notional machine is a pedagogic device to assist the understanding of some aspect of programs or programming. It is typically used to support explaining a programming construct, or the user-understandable semantics of a program. For example, a variable is like a box with a label, and assignment copies or moves a value into that box. This working group will capture examples of notional machines from actual pedagogical practice, as expressed in textbooks (or other teaching materials) or used in the classroom. We will interview at least 30 teachers about their experience with, and perceptions of, the use of notional machines in teaching. Using the interviews, we will work on devising and refining a form to characterise essential features of notional machines. We will also attempt to relate them to each other to describe potential learning sequences or progressions. The working group report will contain descriptions of notional machines used at different levels in education, in different countries, by many teachers. Capturing and Characterising Notional Machines Sally Fincher, Johan Jeuring, Craig S Miller Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). ITiCSE 2020,,Trondheim, Norway © 2020 Copyright held by the owner/author(s). 978-1-4503-0000-0/18/06...$15.00 https://doi.org/10.1145/1234567890 The resulting catalogue of notional machines will allow a teacher to select a machine for a particular use, permit comparison between them, and provide a starting point for further categorization and analysis of notional machines. Additionally, we will make more theoretical explorations. We will explore a variety of presentational formats, examining what is necessary and what superfluous; we will look for dimensions of comparison and will examine how notional machines are instantiated across the discipline. We argue that the creation and use of notional machines is potentially a signature pedagogy for computing [1] and that creating and using notional machines represents a certain level of pedagogic sophistication that might be an indicator of pedagogic content knowledge (PCK). Sally Fincher, Johan Jeuring, Craig S. Miller, Peter Donaldson, Benedict du Boulay, Matthias Hauswirth, Arto Hellas, Felienne Hermans, Colleen M. Lewis, Andreas Mühling, Janice L. Pearce, Andrew Petersen 0001 |
ITiCSE | 7 |
| 2020 | Crowdsourcing Content Creation for SQL PracticeabstractCrowdsourcing refers to the act of using the crowd to create content or to collect feedback on some particular tasks or ideas. Within computer science education, crowdsourcing has been used -- for example -- to create rehearsal questions and programming assignments. As a part of their computer science education, students often learn relational databases as well as working with the databases using SQL statements. In this article, we describe a system for practicing SQL statements. The system uses teacher-provided topics and assignments, augmented with crowdsourced assignments and reviews. We study how students use the system, what sort of feedback students provide to the teacher-generated and crowdsourced assignments, and how practice affects the feedback. Our results suggest that students rate assignments highly, and there are only minor differences between assignments generated by students and assignments generated by the instructor. Juho Leinonen 0001, Nea Pirttinen, Arto Hellas |
ITiCSE | 3 |
| 2020 | Achievement Goal Orientation Profiles and Performance in a Programming MOOCabstractIt has been suggested that performance goals focused on appearing talented (appearance goals) and those focused on outperforming others (normative goals) have different consequences, for example, regarding performance. Accordingly, applying this distinction into appearance and normative goals alongside mastery goals, this study explores what kinds of achievement goal orientation profiles are identified among over 2000 students participating in an introductory programming MOOC. Using Two-Step cluster analysis, five distinct motivational profiles are identified. Course performance and demographics of students with different goal orientation profiles are mostly similar. Students with Combined Mastery and Performance Goals perform slightly better than students with Low Goals. The observations are largely in line with previous studies conducted in different contexts. The differentiation of appearance and normative performance goals seemed to yield meaningful motivational profiles, but further studies are needed to establish their relevance and investigate whether this information can be used to improve teaching. Kukka-Maaria Polso, Heta Tuominen, Arto Hellas, Petri Ihantola |
ITiCSE | 3 |
| 2020 | Gender Differences in Introductory Programming: Comparing MOOCs and Local CoursesabstractWe analyzed three introductory programming MOOCs and four introductory programming courses offered locally in a Finnish university. The course has been offered in all instances with roughly the same content, barring adjustments based on course feedback. We sought to understand how gender interacts with participating in the course in both instances. In particular, we looked at the differences in persistence, confidence, interest in CS, prior experience, and performance between men and women. Overall, we found that men have more prior experience in both instances and have a higher interest in a CS degree. Furthermore, men perform slightly better on the MOOC while there was no significant difference in performance when it came to gender in the local instance. Aligned with prior research, we found a considerable gap in confidence between male and female students in both instances. At the same time, while women are still underrepresented in CS, we observe a considerable increase in women attending the MOOC. Unfortunately, women are also more likely to drop out early on in the MOOC than men. Rodrigo Duran 0001, Lassi Haaranen, Arto Hellas |
SIGCSE | 3 |
| 2020 | A Study of Keystroke Data in Two Contexts: Written Language and Programming Language Influence Predictability of Learning OutcomesabstractWe study programming process data from two introductory programming courses. Between the course contexts, the programming languages differ, the teaching approaches differ, and the spoken languages differ. In both courses, students' keystroke data -- timestamps and the pressed keys -- are recorded as students work on programming assignments. We study how the keystroke data differs between the contexts, and whether research on predicting course outcomes using keystroke latencies generalizes to other contexts. Our results show that there are differences between the contexts in terms of frequently used keys, which can be partially explained by the differences between the spoken languages and the programming languages. Further, our results suggest that programming process data that can be collected non-intrusive in-situ can be used for predicting course outcomes in multiple contexts. The predictive power, however, varies between contexts possibly because the frequently used keys differ between programming languages and spoken languages. Thus, context-specific fine-tuning of predictive models may be needed. John Edwards 0002, Juho Leinonen 0001, Arto Hellas |
SIGCSE | 3 |
| 2019 | Exploring the Value of Student Self-Evaluation in Introductory ProgrammingabstractProgramming teachers have a strong need for easy-to-use instruments that provide reliable and pedagogically useful insights into student learning. Currently, no validated tools exist for rapidly assessing student understanding of basic programming knowledge. Concept inventories and the SCS1 questionnaire can offer great benefits; this article explores the additional value that may be gained from relatively simple self-evaluation metrics. We apply a lightweight self-evaluation instrument (SEI) in an introductory programming course and compare the results to existing performance measures, such as examination grades and the SCS1. We find that the SEI has a similar correlation with a program-writing examination as the SCS1 does, although both instruments correlate only moderately with the examination and each other. Furthermore, students are much more likely to voluntarily answer the lightweight SEI than SCS1. Overall, our results suggest that both the SEI and other instruments need to be greatly improved and outline future work towards that end. Rodrigo Duran 0001, Jan-Mikael Rybicki, Juha Sorva, Arto Hellas |
ICER | 4 |
| 2019 | Admitting Students through an Open Online Course in Programming: A Multi-year Analysis of Study SuccessabstractSince 2012, part of computer science student body at the University of Helsinki has been selected by using a massively open online version of the same introductory programming course that our freshmen take. In this multi-year study, we compare study success between students accepted through the online course (MOOC intake) and students accepted through the traditional entrance exam and high school matriculation exam based intake (normal intake). Our findings indicate that the MOOC intake perform better in computer science studies when looking at completed credits and grade point average, but there is no difference when considering other courses. Retention among the MOOC intake is better than among the normal intake. Additionally, students in the MOOC intake are more likely to complete their capstone project and Bachelor's thesis in the studied time-frame. However, the MOOC intake makes the already skewed gender balance more pronounced. Juho Leinonen 0001, Petri Ihantola, Antti Leinonen, Henrik Nygren, Jaakko Kurhila, Matti Luukkainen, Arto Hellas |
ICER | 7 |
| 2019 | Towards a Common Instrument for Measuring Prior Programming KnowledgeabstractComputing education researchers and educators use a wide range of approaches for measuring students' prior knowledge in programming. Such measurement can help adapt the learning goals and assessment tools for groups of learners at different skills levels and backgrounds. There seems to be no consensus on if and how prior programming knowledge should be measured. Traditional background surveys are often ad-hoc or non-standard, which do not allow comparison of results between different course contexts, levels, and learner groups. Moreover, surveys may yield inaccurate information and may not be useful due to lack of detail. In contrast, tests can provide much higher detail and accuracy than surveys about student knowledge or skills, but large-scale tests are typically very time-consuming or impractical to arrange. To bridge the gap between ad-hoc surveys and standardized tests, we propose and evaluate a novel self-evaluation instrument for measuring prior programming knowledge in introductory programming courses. This instrument investigates in higher detail typical course concepts in programming education considering the different levels of proficiency. Based on a sample of two thousand introductory programming course students, our analysis shows that the instrument is internally consistent, correlates with traditional background information metrics and identifies students of varying programming backgrounds. Rodrigo Duran 0001, Jan-Mikael Rybicki, Arto Hellas, Sanna Suoranta |
ITiCSE | 3 |
| 2019 | Non-restricted Access to Model Solutions: A Good Idea?abstractIn this article, we report an experiment where students in an introductory programming course were given the opportunity to view model solutions to programming assignments whenever they wished, without the need to complete the assignments beforehand or to wait for the deadline to pass. Our experiment was motivated by the observation that some students may spend hours stuck with an assignment, leading to non-productive study time. At the same time, we considered the possibility of students using the sample solutions as worked examples, which could help students to improve the design of their own programs. Our experiment suggests that many of the students use the model solutions sensibly, indicating that they can control their own work. At the same time, a minority of students used the model solutions as a way to proceed in the course, leading to poor exam performance. Henrik Nygren, Juho Leinonen 0001, Arto Hellas |
ITiCSE | 3 |
| 2019 | A Periodic Table of Computing Education Learning TheoriesabstractComputing education research is built on the use of suitable methods within appropriate theoretical frameworks to provide guidance and solutions for our discipline, in a way that is rigorous and repeatable. However, the scale of theory covered extends well beyond the CS discipline and includes educational theory, behavioural psychology, statistics, economics, and game theory, among others. A computing education researcher's journey towards appropriate and discipline relevant theory can be challenging and, when a researcher has learned one area of theory, it can be easy to return to familiar theory, as it may not be clear what the next step could be. The periodic table is a visual arrangement of the elements to group like with like, providing insight into how families of elements will react. Could we do the same with learning theories located in the domain of computer science education, and would it be useful? The working group will identify and survey existing literature on relationships between key areas of theory in computing education, identify ways of organising these research areas to show how knowledge of one could assist another, and produce initial graphical representations of theory and their relationship groupings to assist researchers in understanding how computing theory is currently used in the discipline and what theories might become of interest. Claudia Szabo, Nick Falkner, Andrew Petersen 0001, Heather Bort, Cornelia Connolly, Kathryn I. Cunningham, Peter Donaldson, Arto Hellas, Judithe Sheard |
ITiCSE | 8 |
| 2019 | Exploring the Applicability of Simple Syntax Writing Practice for Learning ProgrammingabstractWhen learning programming, students learn the syntax of a programming language, the semantics underlying the syntax, and practice applying the language in solving programming problems. Research has suggested that simply the syntax may be hard to learn. In this article, we study difficulty of learning the syntax of a programming language. We have constructed a tool that provides students code that they write character-by-character. When writing, the tool automatically highlights each character in code that is incorrectly typed, and through the highlight-based feedback directs students into writing correct syntax. We conducted a randomized controlled trial in an introductory programming course organized in Java. One half of the population had the tool in the course material immediately before programming exercises where the practiced syntax was used, while the other half of the course population did not have the tool, thus approaching the exercises in a traditional way. Our results imply that isolated syntax writing practice may not be a meaningful addition to the arsenal used for teaching programming, at least when the programming course utilizes a large set of small programming exercises. We encourage researchers to replicate our work in contexts where syntax seems to be an issue. Antti Leinonen, Henrik Nygren, Nea Pirttinen, Arto Hellas, Juho Leinonen 0001 |
SIGCSE | 4 |
| 2018 | Taxonomizing features and methods for identifying at-risk students in computing coursesabstractSince computing education began, we have sought to learn why students struggle in computer science and how to identify these at-risk students as early as possible. Due to the increasing availability of instrumented coding tools in introductory CS courses, the amount of direct observational data of student working patterns has increased significantly in the past decade, leading to a flurry of attempts to identify at-risk students using data mining techniques on code artifacts. The goal of this work is to produce a systematic literature review to describe the breadth of work being done on the identification of at-risk students in computing courses. In addition to the review itself, which will summarize key areas of work being completed in the field, we will present a taxonomy (based on data sources, methods, and contexts) to classify work in the area. Arto Hellas, Petri Ihantola, Andrew Petersen 0001, Vangel V. Ajanovski, Mirela Gutica, Timo Hynninen, Antti Knutas, Juho Leinonen 0001, Christopher H. Messom, Soohyun Nam Liao |
ITiCSE | 1 |
| 2018 | Crowdsourcing programming assignments with CrowdSorcererabstractSmall automatically assessed programming assignments are an often used resource for learning programming. Creating sufficiently large amounts of such assignments is, however, time consuming. As a consequence, offering large quantities of practice assignments to students is not always possible. CrowdSorcerer is an embeddable open-source system that students and teachers alike can use for creating and evaluating small automatically assessed programming assignments. While creating programming assignments, the students also write simple input-output -tests, and are gently introduced to the basics of testing. Students can also evaluate the assignments of others and provide feedback on them, which exposes them to code written by others early in their education. In this article we both describe the CrowdSorcerer system and our experiences in using the system in a large undergraduate programming course. Moreover, we discuss the motivation for crowdsourcing course assignments and present some usage statistics. Nea Pirttinen, Vilma Kangas, Irene Nikkarinen, Henrik Nygren, Juho Leinonen 0001, Arto Hellas |
ITiCSE | 6 |
| 2018 | A Study of Pair Programming Enjoyment and Attendance using Study Motivation and Strategy MetricsabstractWe explore educational pair programming in a university context with high student autonomy and individual responsibility. The data comes from two separate introductory programming courses with optional pair programming assignments. We analyze lab attendance and course outcomes to determine whether students' previous programming experience or gender influence attendance. We further compare these statistics to self-reported data on study motivation, study strategies, and student enjoyment of pair programming. The influence of grading systems on pair programming behavior and course outcomes is also examined. Our results suggest that gender and previous programming experience correlate with participation in pair programming labs. At the same time, there are no significant differences in self-reported enjoyment of pair programming between any of the groups, and the results from commonly used study motivation and strategy questionnaires provide little insight into students/ actual behavior. Onni Aarne, Petrus Peltola, Juho Leinonen 0001, Arto Hellas |
SIGCSE | 4 |
| 2018 | Supporting Self-Regulated Learning with Visualizations in Online Learning EnvironmentsabstractIn this article, we study how visualizations could be used to support students' self-regulation in online learning. We conducted a randomized controlled trial with three groups: one control group without visualization, one treatment group with textual visualization, and one treatment with graphical visualization with information on peers' average achievement. We studied how different visualizations affect students' academic performance and behavior. We focused on four factors; starting, scheduling, earliness and exercise points, where the first three are related to time management and self-regulation. The last factor measures course performance in terms of completed exercises. Our results suggest that the lowest performing students can benefit from a visualization, whereas the highest performing students are not affected by the presence or absence of a visualization. We also found that visualizations that do not provide the means to compare your own performance with others may even be harmful to performance oriented students. Kalle Ilves, Juho Leinonen 0001, Arto Hellas |
SIGCSE | 3 |
| 2018 | Subgoal Labeled Worked Examples in K-3 EducationabstractWorked examples are step-by-step instructions that are used to demonstrate and teach problem-solving processes. Subgoal labels are used to group the steps of worked examples into cohesive units that may help the learner to identify key information about the process. We conducted a study on the applicability of subgoal labeled worked examples with 9 and 10-year-old pupils (n=43) who were learning the principles of programming using LightBot. Using a between groups design, pupils in three classes were working with LightBot. One of the groups had no additional instructional materials for the LightBot environment, one of the groups had a set of worked examples without subgoal labels, and the last group had the same set of worked examples with subgoal labels. We measured pupils' success in terms of how many LightBot levels they completed during the class. In addition, pupils' beliefs and attitudes towards programming were assessed before and after the experiment. Our results indicate that in a programming environment such as LightBot, simple worked examples provide no significant benefit over no examples, but worked examples with subgoal labels can help pupils complete more levels. At the same time, the instructional materials in the study had no significant influence on the pupils' beliefs towards computer use or programming. Johanna Joentausta, Arto Hellas |
SIGCSE | 2 |
| 2018 | Social Help-seeking Strategies in a Programming MOOCabstractBeing able to seek help is a crucial part of any learning process. This includes both collaborative models such as asking for help from others as well as independent models such as using course materials and the vast resources provided by the Web. Currently, MOOC research has addressed social help-seeking within the MOOC course, either using MOOC platform tools (forum, chat) or arranging activities using external platforms (Google Hangout, Facebook groups). However, MOOC learning activities take place in a larger social ecology, including friends and teachers, general online communities and alumni communities. Using survey data from a programming MOOC, we show a typology of social learning strategies: non-use of social help-seeking, seeking help from friends and seeking help from alumni and teacher communities. We further show that students using social help-seeking strategies orient themselves more with a surface approach but are also less likely to drop the course. We conclude this work by addressing the various design possibilities identified by this work. Matti Nelimarkka, Arto Hellas |
SIGCSE | 2 |
| 2018 | Achievement Goals in CS1: Replication and ExtensionabstractReplication research is rare in CS education. For this reason, it is often unclear to what extent our findings generalize beyond the context of their generation. The present paper is a replication and extension of Achievement Goal Theory research on CS1 students. Achievement goals are cognitive representations of desired competence (e.g., topic mastery, outperforming peers) in achievement settings, and can predict outcomes such as grades and interest. We study achievement goals and their effects on CS1 students at six institutions in four countries. Broad patterns are maintained --- mastery goals are beneficial while appearance goals are not --- but our data additionally admits fine-grained analyses that nuance these findings. In particular, students' motivations for goal pursuit can clarify relationships between performance goals and outcomes. Daniel Zingaro, Michelle Craig, Leo Porter 0001, Brett A. Becker, Yingjun Cao, Phillip T. Conrad, Diana Cukierman, Arto Hellas, Dastyni Loksa, Neena Thota |
SIGCSE | 8 |
| 2018 | Transfer-Learning Methods in Programming Course Outcome PredictionabstractThe computing education research literature contains a wide variety of methods that can be used to identify students who are either at risk of failing their studies or who could benefit from additional challenges. Many of these are based on machine-learning models that learn to make predictions based on previously observed data. However, in educational contexts, differences between courses set huge challenges for the generalizability of these methods. For example, traditional machine-learning methods assume identical distribution in all data—in our terms, traditional machine-learning methods assume that all teaching contexts are alike. In practice, data collected from different courses can be very different as a variety of factors may change, including grading, materials, teaching approach, and the students. Transfer-learning methodologies have been created to address this challenge. They relax the strict assumption of identical distribution for training and test data. Some similarity between the contexts is still needed for efficient learning. In this work, we review the concept of transfer learning especially for the purpose of predicting the outcome of an introductory programming course and contrast the results with those from traditional machine-learning methods. The methods are evaluated using data collected in situ from two separate introductory programming courses. We empirically show that transfer-learning methods are able to improve the predictions, especially in cases with limited amount of training data, for example, when making early predictions for a new context. The difference in predictive power is, however, rather subtle, and traditional machine-learning models can be sufficiently accurate assuming the contexts are closely related and the features describing the student activity are carefully chosen to be insensitive to the fine differences. Jarkko Lagus, Krista Longi, Arto Klami, Arto Hellas |
ACM Trans. Comput. Educ. | 4 |
| 2018 | Designing and implementing an environment for software start-up education: Patterns and anti-patternsabstractToday’s students are prospective entrepreneurs, as well as potential employees in modern, start-up-like intrapreneurship environments within established companies. In these settings, software development projects face extreme requirements in terms of innovation and attractiveness of the end-product. They also suffer severe consequences of failure such as termination of the development effort and bankruptcy. As the abilities needed in start-ups are not among those traditionally taught in universities, new knowledge and skills are required to prepare students for the volatile environment that new market entrants face. This article reports experiences gained during seven years of teaching start-up knowledge and skills in a higher-education institution. Using a design-based research approach, we have developed the Software Factory, an educational environment for experiential, project-based learning. We offer a collection of patterns and anti-patterns that help educational institutions to design, implement and operate physical environments, curricula and teaching materials, and to plan interventions that may be required for project-based start-up education. Fabian Fagerholm, Arto Hellas, Matti Luukkainen, Kati Kyllonen, Sezin Gizem Yaman, Hanna Mäenpää |
J. Syst. Softw. | 2 |
| 2017 | Patterns for Designing and Implementing an Environment for Software Start-Up EducationabstractToday's students are prospective entrepreneurs and employees in modern, start-up like environments within established companies. In these settings, software development projects face extreme requirements in terms of innovation and attractiveness of the end-product. They also suffer severe consequences of failure such as termination of the development effort and bankruptcy. As the abilities needed in start-ups are not among those traditionally taught in universities, new knowledge and skills are required to prepare students for the volatile environment that new market entrants face. This paper reports experiences gained during seven years of teaching start-up knowledge and skills in a higher-education institution. We offer a collection of patterns that help educational institutions to design, implement and operate physical environments, curricula and teaching materials, and to plan interventions that may be required for project-based start-up education. Fabian Fagerholm, Arto Hellas, Matti Luukkainen, Kati Kyllonen, Sezin Gizem Yaman, Hanna Mäenpää |
SEAA | 2 |
| 2017 | Comparison of Time Metrics in ProgrammingabstractResearch on the indicators of student performance in introductory programming courses has traditionally focused on individual metrics and specific behaviors. These metrics include the amount of time and the quantity of steps such as code compilations, the number of completed assignments, and metrics that one cannot acquire from a programming environment. However, the differences in the predictive powers of different metrics and the cross-metric correlations are unclear, and thus there is no generally preferred metric of choice for examining time on task or effort in programming. In this work, we contribute to the stream of research on student time on task indicators through the analysis of a multi-source dataset that contains information about students' use of a programming environment, their use of the learning material as well as self-reported data on the amount of time that the students invested in the course and per-assignment perceptions on workload, educational value and difficulty. We compare and contrast metrics from the dataset with course performance. Our results indicate that traditionally used metrics from the same data source tend to form clusters that are highly correlated with each other, but correlate poorly with metrics from other data sources. Thus, researchers should utilize multiple data sources to gain a more accurate picture of students' learning. Juho Leinonen 0001, Leo Leppänen, Petri Ihantola, Arto Hellas |
ICER | 4 |
| 2017 | Searching for Early Developmental Activities Leading to Computational Thinking SkillsabstractDrawing on the long debate about whether computer science (CS) and computational thinking skills are innate or learnable, this working group is based on the following hypothesis: The apparent innate ability of some CS learners who succeed in CS courses despite no prior exposure to computing is a manifestation of early childhood experiences and learning outside formal education. Quintin I. Cutts, Peter Donaldson, Elizabeth Cole 0001, Bedour Alshaigy, Mirela Gutica, Arto Hellas, Edurne Larraza-Mendiluze, Robert McCartney, Elizabeth Ann Patitsas, Charles Riedesel |
ITiCSE | 6 |
| 2017 | Plagiarism in Take-home Exams: Help-seeking, Collaboration, and Systematic CheatingabstractDue to the increased enrollments in Computer Science education programs, institutions have sought ways to automate and streamline parts of course assessment in order to be able to invest more time in guiding students' work. Arto Hellas, Juho Leinonen 0001, Petri Ihantola |
ITiCSE | 1 |
| 2017 | Preventing Keystroke Based Identification in Open Data SetsabstractLarge-scale courses such as Massive Online Open Courses (MOOCs) can be a great data source for researchers. Ideally, the data gathered on such courses should be openly available to all researchers. Studies could be easily replicated and novel studies on existing data could be conducted. However, very fine-grained data such as source code snapshots can contain hidden identifiers. For example, distinct typing patterns that identify individuals can be extracted from such data. Hence, simply removing explicit identifiers such as names and student numbers is not sufficient to protect the privacy of the users who have supplied the data. At the same time, removing all keystroke information would decrease the value of the shared data significantly. Juho Leinonen 0001, Petri Ihantola, Arto Hellas |
L@S | 3 |
| 2017 | Progsnap: Sharing Programming Snapshots for Research (Abstract Only)abstractRecent years have seen increasing interest in using programming snapshot data for education research. One barrier to such research, especially for studies involving data from multiple institutions, is that the data is in a wide variety of native formats, and those formats may not be conducive to automated analysis. To overcome this barrier, we propose a structured data model and archival data format called Progsnap (https://cloudcoderdotorg.github.io/progsnap-spec/). Progsnap is designed to be a neutral export format, is currently supported by two open-source programming exercise systems, and we believe will be an easy target for data export from other systems. An open source Python library makes it easy to automate analysis of Progsnap datasets. David Hovemeyer, Arto Hellas, Andrew Petersen 0001, Jaime Spacco |
SIGCSE | 2 |
| 2017 | Stereotype Modeling for Problem-Solving Performance Predictions in MOOCs and Traditional CoursesabstractStereotypes are frequently used in real life to classify students according to their performance in class. In literature, we can find many references to weaker students, fast learners, struggling students, etc. Given the lack of detailed data about students, these or other kinds of stereotypes could be potentially used for user modeling and personalization in the educational context. Recent research in MOOC context demonstrated that data-driven learner stereotypes could work well for detecting and preventing student dropouts. In this paper, we are exploring the application of stereotype-based modeling to a more challenging task -- predicting student problem-solving and learning in two programming courses and two MOOCs. We explore traditional stereotypes based on readily available factors like gender or education level as well as some advanced data-driven approaches to group students based on their problem-solving behavior. Each of the approaches to form student stereotype cohorts is validated by comparing models of student learning: do students in different groups learn differently? In the search for the stereotypes that could be used for adaptation, the paper examines ten approaches. We compare the performance of these approaches and draw conclusions for future research. Roya Hosseini 0001, Peter Brusilovsky, Michael Yudelson, Arto Hellas |
UMAP | 4 |
| 2017 | A Contingency Table Derived Method for Analyzing Course DataabstractWe describe a method for analyzing student data from online programming exercises. Our approach uses contingency tables that combine whether or not a student answered an online exercise correctly with the number of attempts that the student made on that exercise. We use this method to explore the relationship between student performance on online exercises done during semester with subsequent performance on questions in a paper-based exam at the end of semester. We found that it is useful to include data about the number of attempts a student makes on an online exercise. Alireza Ahadi, Arto Hellas, Raymond Lister |
ACM Trans. Comput. Educ. | 2 |
| 2016 | Control-Flow-Only Abstract Syntax Trees for Analyzing Students' Programming ProgressabstractThe abstraction of student code for use in automated analysis is a key challenge. The code must be processed in a manner that reveals interesting properties while reducing the "noise" introduced by less important details. In this work, we investigate the importance of control flow as a property in the analysis of students' programming processes. David Hovemeyer, Arto Hellas, Andrew Petersen 0001, Jaime Spacco |
ICER | 2 |