Maciej Pankiewicz

dblp:278/7352 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
12since 2021 · last 2026
0000-0002-6945-0523ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 3 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Practice as the Key to Success: Understanding the Role of Prior Knowledge, Affective States and Learning Resources in Computer Science Education
abstract
Introductory programming courses (CS1) bring together students with diverse prior experiences, which shape their use of learning resources, emotional responses, and academic performance. This study employs structural equation modeling, self-reported affective data, and multimodal interaction logs to investigate how prior programming knowledge affects affect, resource utilization, and outcomes in a CS1 course that features an automated assessment tool (AAT), instructional videos, and worked examples. Students who persisted with practice and advanced beyond basic tasks achieved the strongest outcomes, though they followed different emotional pathways. By contrast, relying solely on videos or worked examples did not significantly lead to success, and disengagement was generally tied to weaker performance. Novices often reported confusion and frustration; while these emotions sometimes hindered learning, they also drove deeper engagement with the AAT, improving outcomes for those who persisted. Experienced students, however, more often reported boredom, which consistently reduced practice and led to poorer outcomes. These findings underscore the need for adaptive support that balances challenge with guidance to sustain engagement and promote success for diverse learners in CS1.
Andres Felipe Zambrano, Jiayi Zhang 0004, Maciej Pankiewicz, Ryan Baker 0001
LAK3
2025 Usage Patterns and Performance Gains in Gamified Online Judges: A Data-Driven Analysis Informed by Cognitive Psychology in CS1
Luiz A. L. Rodrigues, Andres Felipe Zambrano, Maciej Pankiewicz, Amanda Barany, Ryan Baker 0001
AIED (6)3
2025 srcML-DKT: Enhancing Deep Knowledge Tracing with Robust Code Representations from srcML
Maciej Pankiewicz, Yang Shi 0004, Ryan Baker 0001
EDM1
2024 Leveraging Large Language Models for Next-Generation Educational Technologies
Neil T. Heffernan, Rose E. Wang, Christopher J. MacLellan, Arto Hellas, Chenglu Li, Candace A. Walkington, Joshua Littenberg-Tobias, David Joyner, Steven Moore, Adish Singla, Zachary A. Pardos, Maciej Pankiewicz, Juho Kim 0001, Shashank Sonkar, Clayton Cohn, Anthony Botelho, Andrew S. Lan, Mingyu Feng, Tanja Käser, Eamon Worden
EDM12
2024 De-Identifying Student Personally Identifying Information with GPT-4
Shreya Singhal, Andres Felipe Zambrano, Maciej Pankiewicz, Xiner Liu, Chelsea Porter, Ryan Baker 0001
EDM3
2024 Navigating Compiler Errors with AI Assistance - A Study of GPT Hints in an Introductory Programming Course
abstract
We examined the efficacy of AI-assisted learning in an introductory programming course at the university level by using a GPT-4 model to generate personalized hints for compiler errors within a platform for automated assessment of programming assignments. The control group had no access to GPT hints. In the experimental condition GPT hints were provided when a compiler error was detected, for the first half of the problems in each module. For the latter half of the module, hints were disabled. Students highly rated the usefulness of GPT hints. In affect surveys, the experimental group reported significantly higher levels of focus and lower levels of confrustion (confusion and frustration) than the control group. For the six most commonly occurring error types we observed mixed results in terms of performance when access to GPT hints was enabled for the experimental group. However, in the absence of GPT hints, the experimental group's performance surpassed the control group for five out of the six error types.
Maciej Pankiewicz, Ryan Baker 0001
ITiCSE (1)1
2024 Comparison of Three Programming Error Measures for Explaining Variability in CS1 Grades
abstract
Programming courses can be challenging for first year university students, especially for those without prior coding experience.Students initially struggle with code syntax, but as more advanced topics are introduced across a semester, the difficulty in learning to program shifts to learning computational thinking (e.g., debugging strategies).This study examined the relationships between students' rate of programming errors and their grades on two exams.Using an online integrated development environment, data were collected from 280 students in a Java programming course.The course had two parts.The first focused on introductory procedural programming and culminated with exam 1, while the second part covered more complex topics and object-oriented programming and ended with exam 2. To measure students' programming abilities, 51095 code snapshots were collected from students while they completed assignments that were autograded based on unit tests.Compiler and runtime errors were extracted from the snapshots, and three measures -Error Count, Error Quotient and Repeated Error Density -were explored to identify the best measure explaining variability in exam grades.Models utilizing Error Quotient outperformed the models using the other two measures, in terms of the explained variability in grades and Bayesian Information Criterion.Compiler errors were significant predictors of exam 1 grades but not exam 2 grades; only runtime errors significantly predicted exam 2 grades.The findings indicate that leveraging Error Quotient with multiple error types (compiler and runtime) may be a better measure of students' introductory programming abilities, though still not explaining most of the observed variability.
Valdemar Svábenský, Maciej Pankiewicz, Jiayi Zhang 0004, Elizabeth B. Cloude, Ryan Baker 0001, Eric Fouh
ITiCSE (1)2
2024 Ordered Network Analysis in CS Education: Unveiling Patterns of Success and Struggle in Automated Programming Assessment
abstract
Computer science (CS) education at the university level is often challenging, particularly for students with no prior programming experience. To help scaffold students' CS learning, instructors often utilize systems for automated assessment of programming assignments, where students can individually learn online using automatically generated feedback. However, despite the growing usage of these systems, learning outcomes are often mixed and not all students benefit equally from using these applications. In this study, we utilize Ordered Network Analysis (ONA) to examine data from a system for automated assessment of programming assignments and compare platform activity between novice students (N=110) achieving high (N=43) and low (N=67) scores on the final test of an introductory CS course. We identify and visualize differences in the activity patterns between the groups. High performing novice students tend to request feedback more often, while low performing students more often leave the assignment unsolved after experiencing an unsuccessful attempt. These findings show that Ordered Network Analysis can serve as a useful tool for understanding student behaviors, facilitating the design of targeted interventions that might support learners at key moments in their programming engagement towards task success.
Andres Felipe Zambrano, Maciej Pankiewicz, Amanda Barany, Ryan Baker 0001
ITiCSE (1)2
2023 Measuring Self-regulated Learning Processes in Computer Science Education
abstract
Self-regulated learning (SRL) is important for computer science education. Yet, students often do not have SRL skills to benefit their learning. In this study, we examined 187 (n=187) students’ SRL behaviors while they built programs with an automated feedback tool. Anchored in Winne and Hadwin’s (1998) COPES model of SRL, our results showed that novices used more operators to debug compiler errors, while more experienced programmers used more operators to debug non-compiler errors. Finally, a random forest classifier showed that prior knowledge was the most important COPES feature predicting learning gain, followed closely by the student’s perceived programming ability, use of evaluations with the automated feedback tool, and operators used to debug non-compiler errors on failed programs.
Elizabeth B. Cloude, Ryan Baker 0001, Maciej Pankiewicz
ICCE3
2023 Large Language Models (GPT) for automating feedback on programming assignments
abstract
Addressing the challenge of generating personalized feedback for programming assignments is demanding due to several factors, like the complexity of code syntax or different ways to correctly solve a task. In this experimental study, we automated the process of feedback generation by employing OpenAI’s GPT-3.5 model to generate personalized hints for students solving programming assignments on an automated assessment platform. Students rated the usefulness of GPT-generated hints positively. The experimental group (with GPT hints enabled) relied less on the platform's regular feedback but performed better in terms of percentage of successful submissions across consecutive attempts for tasks, where GPT hints were enabled. For tasks where the GPT feedback was made unavailable, the experimental group needed significantly less time to solve assignments. Furthermore, when GPT hints were unavailable, students in the experimental condition were initially less likely to solve the assignment correctly. This suggests potential over-reliance on GPT- generated feedback. However, students in the experimental condition were able to correct reasonably rapidly, reaching the same percentage correct after seven submission attempts. The availability of GPT hints did not significantly impact students' affective state.
Maciej Pankiewicz, Ryan Baker 0001
ICCE1
2021 Assessing the Cold Start Problem in Adaptive Systems
abstract
This study presents a comparison of the following methods for estimating task difficulty in terms of the so-called "cold start" problem - during the initial phase of the introduction of the adaptive educational system to the public. The data originates from the item-based online programming course made available on the RunCode online learning platform. The dataset contains 50055 submissions on 76 tasks uploaded by 299 learners. The reference task difficulty values have been estimated with the Item Response Theory Graded Response Model on the full dataset. Results of this study show, that for the smallest size of the sample n = 5 learners, the Elo rating algorithm achieves a reasonably high correlation of 0.702. For small sizes of the sample n = 5, 10, and 20 learners, the Elo algorithm delivers slightly better estimates than the Glicko rating algorithm and outperforms the proportion correct (PC) method. For the size of the sample n = 50, the Glicko algorithm delivers better results than the Elo, and the difference between algorithms and the proportion correct method decreases.
Maciej Pankiewicz
ITiCSE (2)1
2021 On-the-fly Estimation of Task Difficulty for Item-based Adaptive Online Learning Environments
abstract
There are several methods aimed at assessing assignment difficulty and learner ability that originate to a great extent from the area of item response theory (IRT). Computational demands have been defined as the main hurdle in the operational usage of these models within adaptive online learning environments. Therefore, alternative methods of difficulty estimation have been analyzed. The aim of this study is to present and compare several methods of estimating the difficulty of assignments for adaptive item-based online learning environments. We compare the following methods of estimating task difficulty: learner feedback, proportion correct, and two rating algorithms: Elo and Glicko with reference values provided by the IRT graded response model. The data originates from an item-based online programming course consisting of 76 programming tasks of various difficulty, solved by 299 users. There have been 50055 attempts recorded in total, for which more than 300000 unit tests have been executed. Learners have provided their opinion on the difficulty of tasks and have generated 2558 ratings. The highest correlation has been observed for the Glicko rating algorithm (0.959), followed by proportion correct (0.943) and the Elo rating algorithm (0.927). The lowest value (0.827) has been observed for the learner feedback method. We conclude, that in the context of an adaptive online learning environment, where the trade-off between operability and accuracy needs to be accepted, the Glicko rating algorithm may be chosen as a fast method for task difficulty estimation.
Maciej Pankiewicz, Marcin Bator
ITiCSE (1)1
2020 Measuring task difficulty for online learning environments where multiple attempts are allowed - the Elo rating algorithm approach
Maciej Pankiewicz
EDM1
2020 A warm-up for adaptive online learning environments - the Elo rating approach for assessing the cold start problem
Maciej Pankiewicz
ICCE1