EDBT 2026 Demo / reviewers in the wild / expert
Filipe D. Pereira
dblp:241/8680 · also Filipe Dwan Pereira
· DBLP profile ↗
21ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0003-4914-3347ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 18 · 8 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 7 first-author · 11 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Translating XAI Into Actionable Feedback Using LLMs to Prevent Student Dropout
Filipe D. Pereira, George Zambonin, André C. A. Nascimento, Mario A. P. Santos, Mariana G. Mello, Tyagi M. Lima, Luiz A. L. Rodrigues, Cleon Xavier, Newarney Torrezão da Costa, Dragan Gasevic, Gabriel Alves 0001, Rafael Ferreira Leite de Mello |
AIED | 1 |
| 2025 | Automatic Short Answer Grading in the LLM Era: Does GPT-4 with Prompt Engineering beat Traditional Models?abstractAssessing short answers in educational settings is challenging due to the need for scalability and accuracy, which led to the field of Automatic Short Answer Grading (ASAG). Traditional machine learning models, such as ensemble and embeddings, have been widely researched in ASAG, but they often suffer from generalizability issues. Recently, Large Language Models (LLMs) emerged as an alternative to optimize ASAG systems. However, previous research has failed to present a comprehensive analysis of LLMs' performance powered by prompt engineering strategies and compare its capabilities to traditional models. This study presents a comparative analysis between traditional machine learning models and GPT-4 in the context of ASAG. We investigated the effectiveness of different models and text representation techniques and explored prompt engineering strategies for LLMs. The results indicate that traditional machine learning models outperform LLMs. However, GPT-4 showed promising capabilities, especially when configured with optimized prompt components, such as few-shot examples and clear instructions. This study contributes to the literature by providing a detailed evaluation of LLM performance compared to traditional machine learning models in a multilingual ASAG context, offering insights for developing more efficient automatic grading systems. Rafael Ferreira Leite de Mello, Cleon Pereira Junior, Luiz A. L. Rodrigues, Filipe D. Pereira, Luciano de Souza Cabral, Newarney Torrezão da Costa, Geber L. Ramalho, Dragan Gasevic |
LAK | 4 |
| 2024 | Can GPT4 Answer Educational Tests? Empirical Analysis of Answer Quality Based on Question Complexity and Difficulty
Luiz A. L. Rodrigues, Filipe D. Pereira, Luciano de Souza Cabral, Geber L. Ramalho, Dragan Gasevic, Rafael Ferreira Leite de Mello |
AIED (1) | 2 |
| 2024 | From Sparse to Smart: Leveraging AI for Effective Online Judge Problem Classification in Programming Education
Filipe D. Pereira, Maely Moraes, Marcelo Henrique Oliveira Henklain, Arto Hellas, Elaine Oliveira, Dragan Gasevic, Raimundo S. Barreto, Rafael Ferreira Leite de Mello |
EC-TEL (1) | 1 |
| 2024 | Solving the imbalanced data issue: automatic urgency detection for instructor assistance in MOOC discussion forumsabstractAbstract In MOOCs, identifying urgent comments on discussion forums is an ongoing challenge. Whilst urgent comments require immediate reactions from instructors, to improve interaction with their learners, and potentially reducing drop-out rates—the task is difficult, as truly urgent comments are rare. From a data analytics perspective, this represents a highly unbalanced (sparse) dataset . Here, we aim to automate the urgent comments identification process, based on fine-grained learner modelling —to be used for automatic recommendations to instructors. To showcase and compare these models, we apply them to the first gold standard dataset for U rgent i N structor I n TE rvention (UNITE) , which we created by labelling FutureLearn MOOC data. We implement both benchmark shallow classifiers and deep learning. Importantly, we not only compare, for the first time for the unbalanced problem, several data balancing techniques , comprising text augmentation, text augmentation with undersampling, and undersampling, but also propose several new pipelines for combining different augmenters for text augmentation . Results show that models with undersampling can predict most urgent cases; and 3X augmentation + undersampling usually attains the best performance. We additionally validate the best models via a generic benchmark dataset (Stanford). As a case study, we showcase how the naïve Bayes with count vector can adaptively support instructors in answering learner questions/comments, potentially saving time or increasing efficiency in supporting learners. Finally, we show that the errors from the classifier mirrors the disagreements between annotators. Thus, our proposed algorithms perform at least as well as a ‘super-diligent’ human instructor (with the time to consider all comments). Laila Alrajhi, Ahmed Alamri, Filipe D. Pereira, Alexandra I. Cristea, Elaine Harada T. de Oliveira |
User Model. User Adapt. Interact. | 3 |
| 2024 | The engage taxonomy: SDT-based measurable engagement indicators for MOOCs and their evaluationabstractAbstract Massive Online Open Course (MOOC) platforms are considered a distinctive way to deliver a modern educational experience, open to a worldwide public. However, student engagement in MOOCs is a less explored area, although it is known that MOOCs suffer from one of the highest dropout rates within learning environments in general, and in e-learning in particular. A special challenge in this area is finding early, measurable indicators of engagement. This paper tackles this issue with a unique blend of data analytics and NLP and machine learning techniques together with a solid foundation in psychological theories. Importantly, we show for the first time how Self-Determination Theory (SDT) can be mapped onto concrete features extracted from tracking student behaviour on MOOCs. We map the dimensions of Autonomy, Relatedness and Competence, leading to methods to characterise engaged and disengaged MOOC student behaviours, and exploring what triggers and promotes MOOC students’ interest and engagement. The paper further contributes by building the Engage Taxonomy, the first taxonomy of MOOC engagement tracking parameters, mapped over 4 engagement theories: SDT, Drive, ET, Process of Engagement. Moreover, we define and analyse students’ engagement tracking, with a larger than usual body of content (6 MOOC courses from two different universities with 26 runs spanning between 2013 and 2018) and students (initially around 218.235). Importantly, the paper also serves as the first large-scale evaluation of the SDT theory itself, providing a blueprint for large-scale theory evaluation. It also provides for the first-time metrics for measurable engagement in MOOCs, including specific measures for Autonomy, Relatedness and Competence; it evaluates these based on existing (and expanded) measures of success in MOOCs: Completion rate, Correct Answer ratio and Reply ratio. In addition, to further illustrate the use of the proposed SDT metrics, this study is the first to use SDT constructs extracted from the first week, to predict active and non-active students in the following week. Alexandra I. Cristea, Ahmed Alamri, Mohammad Alshehri, Filipe D. Pereira, Armando M. Toda, Elaine Harada T. de Oliveira, Craig D. Stewart |
User Model. User Adapt. Interact. | 4 |
| 2023 | Evaluation of a Hybrid AI-Human Recommender for CS1 Instructors in a Real Educational Scenario
Filipe D. Pereira, Elaine Harada T. de Oliveira, Luiz A. L. Rodrigues, Luciano de Souza Cabral, David B. F. Oliveira, Leandro S. G. Carvalho, Dragan Gasevic, Alexandra I. Cristea, Diego Dermeval, Rafael Ferreira Leite de Mello |
EC-TEL | 1 |
| 2023 | G is for Generalisation: Predicting Student Success from KeystrokesabstractStudent performance prediction aims to build models to help educators identify struggling students so they can be better supported. However, prior work in the space frequently evaluates features and models on data collected from a single semester, of a single course, taught at a single university. Without evaluating these methods in a broader context there is an open question of whether or not performance prediction methods are capable of generalising to new data. We test three methods for evaluating student performance models on data from introductory programming courses from two universities with a total of 3,323 students. Our results suggest that using cross-validation on one semester is insufficient for gauging model performance in the real world. Instead, we suggest that where possible future work in student performance prediction collects data from multiple semesters and uses one or more as a distinct hold-out set. Failing this, bootstrapped cross-validation should be used to improve confidence in models' performance. By recommending stronger methods for evaluating performance prediction models, we hope to bring them closer to practical use and assist teachers to understand struggling students in novice programming courses. Zac Pullar-Strecker, Filipe D. Pereira, Paul Denny 0001, Andrew Luxton-Reilly, Juho Leinonen 0001 |
SIGCSE (1) | 2 |
| 2022 | GARFIELD: A Recommender System to Personalize Gamified Learning
Luiz A. L. Rodrigues, Armando M. Toda, Filipe D. Pereira, Paula T. Palomino, Ana C. T. Klock, Marcela Pessoa, David B. F. Oliveira, Isabela Gasparini, Elaine Harada T. de Oliveira, Alexandra I. Cristea, Seiji Isotani |
AIED (1) | 3 |
| 2022 | Towards the understanding of cultural differences in between gamification preferences: A data-driven comparison between the US and Brazil
Armando M. Toda, Ana C. T. Klock, Filipe D. Pereira, Luiz A. L. Rodrigues, Paula T. Palomino, Vinícius Lopes, Craig D. Stewart, Elaine Harada T. de Oliveira, Isabela Gasparini, Seiji Isotani, Alexandra I. Cristea |
EDM | 3 |
| 2022 | Are They Learning or Playing? Moderator Conditions of Gamification's Success in Programming ClassroomsabstractStudents face several difficulties in introductory programming courses (CS1), often leading to high dropout rates, student demotivation, and lack of interest. The literature has indicated that the adequate use of gamification might improve learning in several domains, including CS1. However, the understanding of which (and how) factors influence gamification’s success, especially for CS1 education, is lacking. Thus, there is a clear need to shed light on pre-determinants of gamification’s impact. To tackle this gap, we investigate how user and contextual factors influence gamification’s effect on CS1 students through a quasi-experimental retrospective study ( \( N = 399 \) ), based on a between-subject design (conditions: gamified or non-gamified) in terms of final grade (academic achievement) and the number of programming assignments completed in an educational system (i.e., how much they practiced). Then, we evaluate whether and how user and contextual characteristics (e.g., age, gender, major, programming experience, working situation, internet access, and computer access/sharing) moderate that effect. Our findings indicate that gamification amplified to some extent the impact of practicing. Overall, students practicing in the gamified version presented higher academic achievement than those practicing the same amount in the non-gamified version. Intriguingly, those in the gamified version that practiced much more extensively than the average showed lower academic achievements than those who practiced comparable amounts in the non-gamified version. Furthermore, our results reveal gender as the only statistically significant moderator of gamification’s effect: in our data, it was positive for females but non-significant for males. These findings suggest which (and how) personal and contextual factors moderate gamification’s effects, indicate the need to further understand and examine context’s role, and show that gamification must be cautiously designed to prevent students from playing instead of learning. Luiz A. L. Rodrigues, Filipe D. Pereira, Armando M. Toda, Paula T. Palomino, Wilk Oliveira, Marcela Pessoa, Leandro S. G. Carvalho, David B. F. Oliveira, Elaine Harada T. de Oliveira, Alexandra I. Cristea, Seiji Isotani |
ACM Trans. Comput. Educ. | 2 |
| 2021 | MOOC Next Week Dropout Prediction: Weekly Assessing Time and Learning Patterns
Ahmed Alamri, Zhongtian Sun, Alexandra I. Cristea, Craig D. Stewart, Filipe D. Pereira |
ITS | 5 |
| 2021 | Urgency Analysis of Learners' Comments: An Automated Intervention Priority Model for MOOC
Laila Alrajhi, Ahmed Alamri, Filipe D. Pereira, Alexandra I. Cristea |
ITS | 3 |
| 2021 | A Recommender System Based on Effort: Towards Minimising Negative Affects and Maximising Achievement in CS1 Learning
Filipe D. Pereira, Hermino B. F. Junior, Luiz Rodriguez, Armando M. Toda, Elaine Harada T. de Oliveira, Alexandra I. Cristea, David B. F. Oliveira, Leandro S. G. Carvalho, Samuel C. Fonseca, Ahmed Alamri, Seiji Isotani |
ITS | 1 |
| 2021 | Towards a Human-AI Hybrid System for Categorising Programming ProblemsabstractAs programming skills are increasingly required world-wide and across disciplines, many students use online platforms that provide automatic feedback through a Programming Online Judge (POJ) mechanism. POJs are very popular e-learning tools, boasting large collections of programming problems. Despite their many benefits, students often struggle when solving problems not compatible with their prior knowledge. One important cause of this is that usually statements of problems are not classified according to programming topics (paradigms, data structures, etc.) and, hence, students waste time and effort in trying to solve exercises that are not tailored to their level and needs. Thus, to support students, we propose a new, "front-heavy" pipeline method to predict topics of POJ problems, using Bidirectional Encoder Representations from Transformers (BERT) for contextual text augmentation over the problem statements and further allowing for (lighter-weight) classical machine learning for classification. Our model outperformed all current state-of-the art, with an F1-score of 86% using stratified 10 fold cross-validation in a classically challenging multi-classification problem with seven categories. As a proof of concept, we conducted an experiment to show how our predictive model can be used as a human-AI hybrid complement for POJ, where learners would use AI-based recommendations to find the most appropriate problems. Filipe D. Pereira, Francisco Pires, Samuel C. Fonseca, Elaine Harada T. de Oliveira, Leandro S. G. Carvalho, David B. F. Oliveira, Alexandra I. Cristea |
SIGCSE | 1 |
| 2020 | Automatic Subject-based Contextualisation of Programming Assignment Lists
Samuel C. Fonseca, Filipe D. Pereira, Elaine Harada T. de Oliveira, David Fernandes, Leandro S. G. Carvalho, Alexandra I. Cristea |
EDM | 2 |
| 2020 | Prediction of Users' Professional Profile in MOOCs Only by Utilising Learners' Written Texts
Tahani Aljohani, Filipe D. Pereira, Alexandra I. Cristea, Elaine Harada T. de Oliveira |
ITS | 2 |
| 2020 | Can We Use Gamification to Predict Students' Performance? A Case Study Supported by an Online Judge
Filipe D. Pereira, Armando M. Toda, Elaine Harada T. de Oliveira, Alexandra I. Cristea, Seiji Isotani, Dion Laranjeira, Adriano Almeida, Jonas Mendonça |
ITS | 1 |
| 2019 | Early Dropout Prediction for Programming Courses Supported by Online Judges
Filipe D. Pereira, Elaine Harada T. de Oliveira, Alexandra I. Cristea, David Fernandes, Luciano Silva, Gene Aguiar, Ahmed Alamri, Mohammad Alshehri |
AIED (2) | 1 |
| 2019 | Early Performance Prediction for CS1 Course Students using a Combination of Machine Learning and an Evolutionary AlgorithmabstractMany researchers have started extracting student behaviour by cleaning data collected from web environments and using it as features in machine learning (ML) models. Using log data collected from an online judge, we have compiled a set of successful features correlated with the student grade and applying them on a database representing 486 CS1 students. We used this set of features in ML pipelines which were optimised, featuring a combination of an automated approach with an evolutionary algorithm and hyperparameter-tuning with random search. As a result, we achieved an accuracy of 75.55%, using data from only the first two weeks to predict the student final grades. We show how our pipeline outperforms state-of-the-art work on similar scenarios. Filipe D. Pereira, Elaine Harada T. de Oliveira, David Fernandes, Alexandra I. Cristea |
ICALT | 1 |
| 2019 | Predicting MOOCs Dropout Using Only Two Easily Obtainable Features from the First Week's Activities
Ahmed Alamri, Mohammad Alshehri, Alexandra I. Cristea, Filipe D. Pereira, Elaine Harada T. de Oliveira, Lei Shi 0003, Craig D. Stewart |
ITS | 4 |