VLDB 2026 Research / reviewers in the wild / expert
Rafael Ferreira Leite de Mello
dblp:205/0582 · also Rafael Fe Mello, Rafael Ferreira 0002, Rafael Ferreira Mello
· DBLP profile ↗
93ranked-venue papers
16as first author
50since 2021 · last 2026
0000-0003-3548-9670ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 58 · 7 first-author · 41 since 2021Human-computer interaction and ubiquitous computing · 56 · 7 first-author · 40 since 2021Artificial intelligence and machine learning · 18 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 17 · 4 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Understanding Teacher Revisions of Large Language Model-Generated Feedback
Conrad Borchers, Luiz A. L. Rodrigues, Newarney Torrezão da Costa, Cleon Xavier, Rafael Ferreira Leite de Mello |
AIED | 5 |
| 2026 | Translating XAI Into Actionable Feedback Using LLMs to Prevent Student Dropout
Filipe D. Pereira, George Zambonin, André C. A. Nascimento, Mario A. P. Santos, Mariana G. Mello, Tyagi M. Lima, Luiz A. L. Rodrigues, Cleon Xavier, Newarney Torrezão da Costa, Dragan Gasevic, Gabriel Alves 0001, Rafael Ferreira Leite de Mello |
AIED | 12 |
| 2026 | Automated Assessment of Handwritten Math Problems: A Comparison of Prompting Strategies for Open and Closed-source LLMsabstractAssessing handwritten mathematical solutions is essential for identifying students’ weaknesses and fostering personalized learning. However, scaling such assessment remains challenging for Learning Analytics, which has traditionally focused on digital or typed data. The current study investigated the potential of Large Language Models (LLMs) to automating the assessment of handwritten mathematical solutions and explore how they can be incorporated into large scale learning analytics pipelines. We curated 300 student solution images, annotated them using to a taxonomy of math error types, and compared open-source (Qwen2.5-7B and Gemma3 12B-IT) and closed-source (Gemini 2.0 Flash and GPT-4) LLMs. Two prompting strategies were tested: from adapted from related work and one tailed to the taxonomy of math error types using established prompt design principles. The results revealed that LLMs, particularly Gemini, achieved strong to moderate performance in diagnosing and classifying student errors, while exposing recurring model specific errors. These findings highlight both the promise and limitations of LLMs for integrating handwritten work in LA and recommend that learning analytics practitioners and researchers combine careful model selection, principled prompt design, and error-level analysis to develop AI-powered LA systems that are accurate, equitable, and pedagogically actionable. Daniel Carneiro Rosa, Andreza Falcão, Jamilla Lobo, Everton Souza, Moésio Wenceslau, Dragan Gasevic, Rafael Ferreira Leite de Mello, Luiz A. L. Rodrigues |
LAK | 7 |
| 2026 | From Solo Graders to Assisted Annotation: Integrating LLM Suggestions into the Educational Data Creation Pipeline
Cleon Xavier, Luiz A. L. Rodrigues, Ana Valdo, Ariadne Carvalho, Gabriela Matos, Lucas Kalinke, Ramon Vilela, Erika Resende, Thais Moraes, Nara Nobre-Silva, Newarney Torrezão da Costa, Fabíola Gonçalves C. Ribeiro, Anderson Pinheiro, Dragan Gasevic, Rafael Ferreira Leite de Mello |
LAK | 16 |
| 2025 | The Impact of Oversampling Techniques on the Detection of Cognitive Presence
Vitor Rolim, Cleon Xavier, Luiz A. L. Rodrigues, Newarney Torrezão da Costa, Rafael Dueire Lins, Dragan Gasevic, Rafael Ferreira Leite de Mello |
AIED (5) | 7 |
| 2025 | Designing Actionable and Interpretable Analytics Indicators for Improving Feedback in AI-Based SystemsabstractInternational audience Esther Félix, Elaine Harada T. de Oliveira, Ilmara M. M. Ramos, Mar Pérez-Sanagustín, Esteban Villalobos, Isabel Hilliger, Rafael Ferreira Leite de Mello, Julien Broisin |
CSEDU (1) | 7 |
| 2025 | Tutoria: Delivering Personalized Feedback at Scale with Artificial Intelligence
Newarney Torrezão da Costa, Cleon Xavier, Fabíola Gonçalves C. Ribeiro, Gabriel Alves 0001, Luiz A. L. Rodrigues, Taciana Pontual Falcão, Rafael Ferreira Leite de Mello |
EC-TEL (2) | 7 |
| 2025 | Scaffolding Learning Scenarios with a Socratic Chatbot: Insights from Educators
Isabel Hilliger, Mar Pérez-Sanagustín, Esteban Villalobos, Rafael Ferreira Leite de Mello |
EC-TEL (2) | 4 |
| 2025 | Escreva Mais: A Mobile Application to Enhancing Writing Skills in Resource-Constrained Classrooms
Rafael Ferreira Leite de Mello, Gabriel Barbosa, Silas Augusto, Lenon Anthony, Jamilla Lobo, Cleon Xavier, Newarney Torrezão da Costa, Luiz A. L. Rodrigues |
EC-TEL (2) | 1 |
| 2025 | Can GPT-4o Evaluate Usability Like Human Experts? A Comparative Study on Issue Identification in Heuristic Evaluation
Guilherme Corredato Guerino, Luiz A. L. Rodrigues, Bruna Santana Capeleti, Rafael Ferreira Leite de Mello, André Pimenta Freire, Luciana A. M. Zaina |
INTERACT (3) | 4 |
| 2025 | That's What RoBERTa Said: Explainable Classification of Peer FeedbackabstractContains fulltext : 317127.pdf (Publisher’s version ) (Open Access) Rafael Ferreira Leite de Mello, Cleon Pereira Junior, Luiz A. L. Rodrigues, Martine Baars, Olga Viberg |
LAK | 2 |
| 2025 | Automatic Short Answer Grading in the LLM Era: Does GPT-4 with Prompt Engineering beat Traditional Models?abstractAssessing short answers in educational settings is challenging due to the need for scalability and accuracy, which led to the field of Automatic Short Answer Grading (ASAG). Traditional machine learning models, such as ensemble and embeddings, have been widely researched in ASAG, but they often suffer from generalizability issues. Recently, Large Language Models (LLMs) emerged as an alternative to optimize ASAG systems. However, previous research has failed to present a comprehensive analysis of LLMs' performance powered by prompt engineering strategies and compare its capabilities to traditional models. This study presents a comparative analysis between traditional machine learning models and GPT-4 in the context of ASAG. We investigated the effectiveness of different models and text representation techniques and explored prompt engineering strategies for LLMs. The results indicate that traditional machine learning models outperform LLMs. However, GPT-4 showed promising capabilities, especially when configured with optimized prompt components, such as few-shot examples and clear instructions. This study contributes to the literature by providing a detailed evaluation of LLM performance compared to traditional machine learning models in a multilingual ASAG context, offering insights for developing more efficient automatic grading systems. Rafael Ferreira Leite de Mello, Cleon Pereira Junior, Luiz A. L. Rodrigues, Filipe D. Pereira, Luciano de Souza Cabral, Newarney Torrezão da Costa, Geber L. Ramalho, Dragan Gasevic |
LAK | 1 |
| 2025 | Exploring Human-AI Collaboration in Educational Contexts: Insights from Writing Analytics and Authorship Attribution
Hongchen Pan, Eduardo Oliveira 0001, Rafael Ferreira Leite de Mello |
LAK | 3 |
| 2025 | LLMs Performance in Answering Educational Questions in Brazilian Portuguese: A Preliminary Analysis on LLMs Potential to Support Diverse Educational NeedsabstractQuestion-answering systems facilitate adaptive learning and respond to student queries, making education more responsive. Despite that, challenges such as natural language understanding and context management complicate their widespread adoption, where Large Language Models (LLMs) offer a promising solution. However, existing research is predominantly focused on English, proprietary models, and often limited to a single question type, subject, or skill, leaving a gap in understanding LLMs' performance in languages like Brazilian Portuguese and across questions of various characteristics. This study investigates how LLMs could be integrated in an educational question-answering system efficiently to answer different question types (multiple-choice, cloze, open-ended), subjects (mathematics and Portuguese language), and skills (summation/subtraction, multiplication, interpretation, and grammar), evaluating answers by GPT-4 - the main LLM at the time of writing - and Sabiá - the open-source Brazilian Portuguese LLM - based on grades assigned by two experienced teachers. Overall, both LLMs demonstrated strong overall performance, with mean scores close to 9.8 out of 10. However, specific challenges emerged, with distinct strengths and weaknesses observed for each model, such as GPT-4's error in a multiple-choice subtraction question and Sabiá's misinterpretation of a cloze question. Luiz A. L. Rodrigues, Cleon Xavier, Newarney Torrezão da Costa, Hyan Batista, Luiz Felipe Bagnhuk Silva, Weslei Chaleghi de Melo, Dragan Gasevic, Rafael Ferreira Leite de Mello |
LAK | 8 |
| 2024 | Automatic Detection of Narrative Rhetorical Categories and Elements on Middle School Written Essays
Rafael Ferreira Leite de Mello, Luiz A. L. Rodrigues, Erverson B. G. de Sousa, Hyan Batista, Mateus Lins, André C. A. Nascimento, Dragan Gasevic |
AIED (1) | 1 |
| 2024 | Anticipating Student Abandonment and Failure: Predictive Models in High School Settings
Emanuel Marques Queiroga, Daniel Santana, Marcelo da Silva, Martim de Aguiar, Vinícius G. dos Santos, Rafael Ferreira Leite de Mello, Ig Ibert Bittencourt, Cristian Cechinel |
AIED (1) | 6 |
| 2024 | Can GPT4 Answer Educational Tests? Empirical Analysis of Answer Quality Based on Question Complexity and Difficulty
Luiz A. L. Rodrigues, Filipe D. Pereira, Luciano de Souza Cabral, Geber L. Ramalho, Dragan Gasevic, Rafael Ferreira Leite de Mello |
AIED (1) | 6 |
| 2024 | From Sparse to Smart: Leveraging AI for Effective Online Judge Problem Classification in Programming Education
Filipe D. Pereira, Maely Moraes, Marcelo Henrique Oliveira Henklain, Arto Hellas, Elaine Oliveira, Dragan Gasevic, Raimundo S. Barreto, Rafael Ferreira Leite de Mello |
EC-TEL (1) | 8 |
| 2024 | AI in Education Unplugged Support Equity Between Rural and Urban Areas in Brazil
Carlos S. Portela, Paula T. Palomino, Geiser Chalco Challco, Alvaro Sobrinho, Thiago D. Cordeiro, Rafael Ferreira Leite de Mello, Diego Dermeval, Ig Ibert Bittencourt, Seiji Isotani |
ICTD | 6 |
| 2024 | Towards Improving Rhetorical Categories Classification and Unveiling Sequential Patterns in Students' WritingabstractTo meet the growing demand for future professionals who can present information to an audience and create quality written products, educators are increasingly assigning writing assignments that require students to gather information from multiple sources, reorganise and reinterpret knowledge from source materials, and plan for rhetorical structure goals in order to meet the task requirements. When evaluating an essay coherence, scorers manually look for the presence of required rhetorical categories, which takes time. Supervised Machine Learning (ML) techniques have proven to be an effective tool for automatic detection of rhetorical categories that approximate students’ cognitive engagement with source information. Previous studies that addressed this problem used relatively small datasets and reported relatively low kappa scores for accuracy, limiting the use of such models in real-world scenarios. Moreover, to empower educators to effectively evaluate the overall quality of students’ writing, the associations between the sequential patterns of rhetorical categories in students’ writing and writing performance must be examined, which remains largely unexplored in educational domain. Therefore, to fill these gaps, our study aimed to i) investigate the impact of data augmentation approaches on the performance of deep learning algorithms in classifying rhetorical categories in student essays according to Bloom‘s taxonomy ii) and explore the sequential patterns of rhetorical categories in students’ writing that can influence writing performance. Our findings showed that deep learning-based model BERT on Easy Data Augmentation (EDA) based augmented data achieved 20% higher Cohen’s kappa than normal (non-augmented) data, and we discovered that students in different performance groups were statistically different in terms of rhetorical patterns. Our proposed study is valuable in terms of building a data analytic foundation that can be used to create formative feedback on students’ writings based on the patterns of rhetorical categories to improve essay quality. Sehrish Iqbal, Mladen Rakovic, Guanliang Chen, Tongguang Li, Jasmine Bajaj, Rafael Ferreira Leite de Mello, Yizhou Fan, Naif R. Aljohani, Dragan Gasevic |
LAK | 6 |
| 2024 | Visual Data Science with Blockly-DSabstractThe workshop will give educators an introduction to graphical data analysis techniques for exploring, summarizing, and effectively communicating data using Visual Blocks (Blockly) and Jupiter Notebooks. Participants will gain skills to create and interpret various graphs and plots, fostering insights into data patterns and relationships. The emphasis will be on adeptly selecting visualizations. Upon completion of this course, educators should be proficient in: Understanding the principles of graphical data analysis and their practical applications; and creating and interpreting diverse graph types and plots. The workshop will further delve into statistical methods essential for data analysis, guiding educators on employing descriptive statistics to explore data and derive meaningful business insights. The focus will be on fostering an understanding of statistical concepts and their practical application, as well as the effective interpretation and communication of findings. Additionally, the program will introduce educators to machine learning regressors, with a primary focus on linear regression. Participants will learn the skills to train, evaluate, and apply regressors to predict continuous target variables from input features. The emphasis will be on grasping and applying machine learning principles, as well as effectively interpreting and communicating model predictions. Furthermore, participants will engage in hands-on experience with machine learning classifiers, encompassing logistic regression, decision trees, naive Bayes, and neural networks. Educators will acquire the expertise to train, evaluate, and apply classifiers for predicting categorical target variables from input features. The emphasis here is on understanding and applying machine learning principles, as well as adeptly interpreting model predictions. Luiz Barboza, Rafael Ferreira Leite de Mello, Erico Souza Teixeira, Andrew Olney |
SIGCSE (2) | 2 |
| 2024 | Towards explainable automatic punctuation restoration for Portuguese using transformers
Tiago Barbosa de Lima, Vitor Rolim, André C. A. Nascimento, Péricles B. C. Miranda, Valmir Macario, Luiz A. L. Rodrigues, Elyda L. S. X. Freitas, Dragan Gasevic, Rafael Ferreira Leite de Mello |
Expert Syst. Appl. | 9 |
| 2024 | Assessing students' handwritten text productions: A two-decades literature review
Lenardo Chaves e Silva, Alvaro Sobrinho, Thiago D. Cordeiro, Alan Silva, Diego Dermeval, Ig Ibert Bittencourt, Jário José Santos, Rafael Ferreira Leite de Mello, Carlos S. Portela, Maurício R. de A. Souza, Rodrigo Lisbôa Pereira, Edson Koiti Kudo Yasojima, Seiji Isotani |
Expert Syst. Appl. | 9 |
| 2023 | Understanding Peer Feedback Contributions Using Natural Language ProcessingabstractAbstract Peer feedback has been widely used in computer-supported collaborative learning (CSCL) setting to improve students’ engagement with massive courses. Although the peer feedback process increases students’ self-regulatory practice, metacognition, and academic achievement, instructors need to go through large amounts of feedback text data which is much more time-consuming. To address this challenge, the present study proposes an automated content analysis approach to identify relevant categories in peer feedback based on traditional and sequence-based classifiers using TF-IDF and content-independent features. We use a data set from an extensive course (N = 231 students) in the setting of engineering higher education. In particular, a total of 2,444 peer feedback messages were analyzed. The CRF classification model based on the TF-IDF features achieved the best performance. The results illustrate that the ability to scale up the automatic analysis of peer feedback provides new opportunities for student-improved learning and improved teacher support in higher education at scale. Mayara Simões de Oliveira Castro, Rafael Ferreira Leite de Mello, Giuseppe Fiorentino, Olga Viberg, Daniel Spikol, Martine Baars, Dragan Gasevic |
EC-TEL | 2 |
| 2023 | Evaluation of a Hybrid AI-Human Recommender for CS1 Instructors in a Real Educational Scenario
Filipe D. Pereira, Elaine Harada T. de Oliveira, Luiz A. L. Rodrigues, Luciano de Souza Cabral, David B. F. Oliveira, Leandro S. G. Carvalho, Dragan Gasevic, Alexandra I. Cristea, Diego Dermeval, Rafael Ferreira Leite de Mello |
EC-TEL | 10 |
| 2023 | Automatic Simplification of Legal Texts in Portuguese Using Machine LearningabstractTexts produced by the Brazilian judiciary have a complex and technical vocabulary, with elaborate use of the Portuguese language and many legal terms difficult to be understood, generating a barrier in communication between the judiciary and the population. In this sense, the Automatic Text Simplification (ATS), activity of the Natural Language Processing (NLP) area, can be applied to improve the readability of these types of text using specialized algorithms, and promote scalability in simplifying them, in view of the great demand in the courts. In this context, this article presents an evaluation of four methods of state of the art in text simplification, evaluated according to readability metrics, to improve the quality of existing texts in the judicial summaries, dataset containing 100 summaries of the Federal Regional Court of the 5th Region (TRF5) and another 100 of the Federal Supreme Court (STF). The methods MUSS(EN), MUSS(PT), Transformers and NMT + Attention were tested, and the results of the simplifications exceeded the FRE readability index of the original texts, making them more readable. Alexandre Alves, Péricles B. C. Miranda, Rafael Ferreira Leite de Mello, André C. A. Nascimento |
JURIX | 3 |
| 2023 | Blockly-DS: Blocks Programming for Data Science with Visual, Statistical, Descriptive and Predictive AnalysisabstractInterest in data science has been growing across industries - both STEM and non-STEM. Non-STEM students often have difficulties with programming and data analysis tools. These entry barriers can be minimized, and these concepts can be easily absorbed when using visual tools. Thus, for this specific audience, the use of visual tools has been essential for teaching data science. Several of these tools are available, but they all have limitations. This work presents Blockly-DS: a new tool capable of assisting in teaching data science to a non-STEM audience. The Blockly-DS tool is being tested in two Brazilian higher education institutions, one, IBMEC, a business undergraduate university, and the other, FIAP, a STEM school that offers an MBA as well as corporate and undergraduate courses. The preliminary results presented in this article refers to a validation with two groups of training sessions for junior financial analysts of a major Brazilian bank in partnership with FIAP. Luiz Barboza, Rafael Ferreira Leite de Mello, Micah Gideon Modell, Erico Souza Teixeira |
LAK | 2 |
| 2023 | Towards Automated Analysis of Rhetorical Categories in Students Essay Writings using Bloom's TaxonomyabstractEssay writing has become one of the most common learning tasks assigned to students enrolled in various courses at different educational levels, owing to the growing demand for future professionals to effectively communicate information to an audience and develop a written product (i.e. essay). Evaluating a written product requires scorers who manually examine the existence of rhetorical categories, which is a time-consuming task. Machine Learning (ML) approaches have the potential to alleviate this challenge. As a result, several attempts have been made in the literature to automate the identification of rhetorical categories using Rhetorical Structure Theory (RST). However, RST do not provide information regarding students’ cognitive level, which motivates the use of Bloom’s Taxonomy. Therefore, in this research we propose to: i) investigate the extent to which classification of rhetorical categories can be automated based on Bloom’s taxonomy by comparing the traditional ML classifiers with the pre-trained language model BERT, ii) explore the associations between rhetorical categories and writing performance. Our results showed that BERT model outperformed the traditional ML-based classifiers with 18% better accuracy, indicating it can be used in future analytics tool. Moreover, we found a statistical difference between the associations of rhetorical categories in low-achiever, medium-achiever and high-achiever groups which implies that rhetorical categories can be predictive of writing performance. Sehrish Iqbal, Mladen Rakovic, Guanliang Chen, Tongguang Li, Rafael Ferreira Leite de Mello, Yizhou Fan, Giuseppe Fiorentino, Naif R. Aljohani, Dragan Gasevic |
LAK | 5 |
| 2023 | Learner-centred Analytics of Feedback Content in Higher EducationabstractFeedback is an effective way to assist students in achieving learning goals. The conceptualisation of feedback is gradually moving from feedback as information to feedback as a learner-centred process. To demonstrate feedback effectiveness, feedback as a learner-centred process should be designed to provide quality feedback content and promote student learning outcomes on the subsequent task. However, it remains unclear how instructors adopt the learner-centred feedback framework for feedback provision in the teaching practice. Thus, our study made use of a comprehensive learner-centred feedback framework to analyse feedback content and identify the characteristics of feedback content among student groups with different performance changes. Specifically, we collected the instructors’ feedback on two consecutive assignments offered by an introductory to data science course at the postgraduate level. On the basis of the first assignment, we used the status of student grade changes (i.e., students whose performance increased and those whose performance did not increase on the second assignment) as the proxy of the student learning outcomes. Then, we engineered and extracted features from the feedback content on the first assignment using a learner-centred feedback framework and further examined the differences of these features between different groups of student learning outcomes. Lastly, we used the features to predict student learning outcomes by using widely-used machine learning models and provided the interpretation of predicted results by using the SHapley Additive exPlanations (SHAP) framework. We found that 1) most features from the feedback content presented significant differences between the groups of student learning outcomes, 2) the gradient boost tree model could effectively predict student learning outcomes, and 3) SHAP could transparently interpret the feature importance on predictions. Jionghao Lin, Lisa-Angelique Lim, Yi-Shan Tsai, Rafael Ferreira Leite de Mello, Hassan Khosravi, Dragan Gasevic, Guanliang Chen |
LAK | 5 |
| 2023 | Towards explainable prediction of essay cohesion in Portuguese and EnglishabstractTextual cohesion is an essential aspect of a formally written text, related to linguistic mechanisms that connect elements such as words, sentences, and paragraphs. Several studies have proposed approaches to estimate textual cohesion in essays automatically. There is limited research that aims to study the extent to which the use of machine learning approaches can predict the textual cohesion of essays written in different languages (not just English). This paper reports on the findings of a study that aimed to propose and evaluate approaches that automatically estimate the cohesion of essays in Portuguese and English. The study proposed regression-based models grounded in conventional feature-based machine learning methods and deep learning-based pre-trained language models. The study also examined the explainability of automated approaches to scrutinize their predictions. We analyzed two datasets composed of 4,570 (Portuguese) and 7,101 (English) essays. The results demonstrate that a deep learning-based model achieved the best performance on both datasets with a moderate Pearson correlation with human-rated cohesion scores. However, the explainability of the automatic cohesion estimations based on conventional machine learning models offered a stronger potential than that of the deep learning model. Hilário Oliveira, Rafael Ferreira Leite de Mello, Bruno Alexandre Barreiros Rosa, Mladen Rakovic, Péricles B. C. Miranda, Thiago D. Cordeiro, Seiji Isotani, Ig Ibert Bittencourt, Dragan Gasevic |
LAK | 2 |
| 2023 | A novel multi-objective grammar-based framework for the generation of Convolutional Neural Networks
Cleber A. C. F. da Silva, Daniel Carneiro Rosa, Péricles B. C. Miranda, Filipe R. Cordeiro, Tapas Si, André C. A. Nascimento, Rafael Ferreira Leite de Mello, Paulo S. G. de Mattos Neto |
Expert Syst. Appl. | 7 |
| 2023 | Applications of convolutional neural networks in education: A systematic literature review
Lenardo Chaves e Silva, Alvaro Sobrinho, Thiago D. Cordeiro, Rafael Ferreira Leite de Mello, Ig Ibert Bittencourt, Diego Dermeval, Alan Silva, Seiji Isotani |
Expert Syst. Appl. | 4 |
| 2022 | Multi-Objective Optimization of Sampling Algorithms Pipeline for Unbalanced ProblemsabstractThe sequencing of sampling algorithms has shown to be a promising approach in generating balanced versions of unbalanced data. Sequencing allows different algorithms of under-sampling and/or over-sampling to be performed in sequence, producing a resulting balanced database. However, defining the most appropriate sequence of sampling algorithms is challenging. This article treats the sequencing problem as a combinatorial optimization task and proposes a multi-objective optimization method to seek promising solutions that maximize the performance of classifiers both in accuracy and in F1-score. The results showed that the proposed method was capable of finding optimized sequences that improved the performance of the classifiers, obtaining statistically better results, mainly in F1- score, when compared with competing methods, in most of the selected unbalanced problems. Péricles B. C. Miranda, Rafael Ferreira Leite de Mello, André C. A. Nascimento, Tapas Si |
CEC | 2 |
| 2022 | Enhancing Instructors' Capability to Assess Open-Response Using Natural Language Processing and Learning Analytics
Rafael Ferreira Leite de Mello, José Rodrigues Lima Neto, Giuseppe Fiorentino, Gabriel Alves 0001, Verenna Arêdes, João Victor Galdino Ferreira Silva, Taciana Pontual Falcão, Dragan Gasevic |
EC-TEL | 1 |
| 2022 | A Penny for your Thoughts: Students and Instructors' Expectations about Learning Analytics in BrazilabstractStakeholder engagement is a key aspect for the successful implementation of Learning Analytics (LA) in Higher Education Institutions (HEIs). Studies in Europe and Latin America (LATAM) indicate that, overall, instructors and students have positive views on LA adoption, but there are differences between their ideal expectations and what they consider realistic in the context of their institutions. So far, very little has been found about stakeholders’ views on LA in Brazilian higher education. By replicating the survey conducted in other countries, in seven Brazilian HEIs, we found convergences both with Europe and LATAM, reinforcing the need for local diagnosis and indicating the risk of assuming a ”LATAM identity”. Our findings contribute to building a corpus of knowledge on stakeholders expectations with a contextualised comprehension of the gaps between ideal and predicted scenarios, which can inform institutional policies for LA implementation in Brazil. Taciana Pontual Falcão, Rodrigo L. Rodrigues, Cristian Cechinel, Diego Dermeval, Elaine Harada T. de Oliveira, Isabela Gasparini, Rafael Dias Araújo, Tiago Thompsen Primo, Dragan Gasevic, Rafael Ferreira Leite de Mello |
LAK | 10 |
| 2022 | NASC: Network analytics to uncover socio-cognitive discourse of student rolesabstractRoles that learners assume during online discussions are an important aspect of educational experience. The roles can be assigned to learners and/or can spontaneously emerge through student-student interaction. While existing research proposed several approaches for analytics of emerging roles, there is limited research in analytic methods that can i) automatically detect emerging roles that can be interpreted in terms of higher-order constructs of collaboration; ii) analyse the extent to which students complied to scripted roles and how emerging roles compare to scripted ones; and iii) track progression of roles in social knowledge progression over time. To address these gaps in the literature, this paper propose a network-analytic approach that combines techniques of cluster analysis and epistemic network analysis. The method was validated in an empirical study discovered emerging roles that were found meaningful in terms of social and cognitive dimensions of the well-known model of communities of inquiry. The study also revealed similarities and differences between emerging and script roles played by learners and identified different progression trajectories in social knowledge construction between emerging and scripted roles. The proposed analytic approach and the study results have implications that can inform teaching practice and development techniques for collaboration analytics. Maverick Andre Dionisio Ferreira, Rafael Ferreira Leite de Mello, Vitomir Kovanovic, André C. A. Nascimento, Rafael Dueire Lins, Dragan Gasevic |
LAK | 2 |
| 2022 | Uncovering Associations Between Cognitive Presence and Speech Acts: A Network-Based ApproachabstractThis research aimed to explore the relationship between different indicators of the depth and quality of participation in computer-mediated learning environments. By using network analyses and statistical tests, we discovered significant associations between the cognitive presence phases of the Community of Inquiry framework and speech acts, and examined the impact of two different instructional interventions on these associations. We found that there are strong associations between some speech acts and cognitive presence phases. In addition, the study revealed that the association between speech acts and cognitive presence is moderated by external facilitation, but not affected by user role assignment. The results suggest that speech acts can plausibly be used to provide feedback in relation to cognitive presence and can potentially be used to increase the generalizability of cognitive presence classification. Sehrish Iqbal, Zach Swiecki, Srecko Joksimovic, Rafael Ferreira Leite de Mello, Naif R. Aljohani, Saeed-Ul Hassan, Dragan Gasevic |
LAK | 4 |
| 2022 | Towards automated content analysis of rhetorical structure of written essays using sequential content-independent features in PortugueseabstractBrazilian universities have included essay writing assignments in the entrance examination procedure to select prospective students. The essay scorers manually look for the presence of required Rhetorical Structure Theory (RST) categories and evaluate essay coherence. However, identifying RST categories is a time-consuming task. The literature reported several attempts to automate the identification of RST categories in essays with machine learning. Still, previous studies have focused on using machine learning algorithms trained on content-dependent features that can diminish classification performance, leading to over-fitting and hindering model generalisability. Therefore, this paper proposes: (i) the analysis of state-of-the-art classifiers and content-independent features to the task of RST rhetorical moves; (ii) a new approach that considers the sequence of the text to extract features – i.e. sequential content-independent features; (iii) an empirical study about the generalisability of the machine learning models and sequential content-independent features for this context; (iv) the identification of the most predictive features for automated identification of RST categories in essays written in Portuguese. The best performing classifier, XGBoost, based on sequential content-independent features, outperformed the classifiers used in the literature and are based on traditional content-dependent features. The XGBoost classifier based on sequential content-independent features also reached promising accuracy when tested for generalisability. Rafael Ferreira Leite de Mello, Giuseppe Fiorentino, Hilário Oliveira, Péricles B. C. Miranda, Mladen Rakovic, Dragan Gasevic |
LAK | 1 |
| 2021 | Analytics of Emerging and Scripted Roles in Online Discussions: An Epistemic Network Analysis Approach
Maverick Andre Dionisio Ferreira, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Dragan Gasevic |
AIED (2) | 2 |
| 2021 | Contrasting Automatic and Manual Group Formation: A Case Study in a Software Engineering Postgraduate Course
Giuseppe Fiorentino, Péricles B. C. Miranda, André C. A. Nascimento, Ana Paula C. Furtado, Henrik Bellhäuser, Dragan Gasevic, Rafael Ferreira Leite de Mello |
AIED (2) | 7 |
| 2021 | Aligning Expectations About the Adoption of Learning Analytics in a Brazilian Higher Education Institution
Samantha Garcia, Elaine Cristina Moreira Marques, Rafael Ferreira Leite de Mello, Dragan Gasevic, Taciana Pontual Falcão |
AIED (2) | 3 |
| 2021 | Towards Automatic Content Analysis of Rhetorical Structure in Brazilian College Entrance Essays
Rafael Ferreira Leite de Mello, Giuseppe Fiorentino, Péricles B. C. Miranda, Hilário Oliveira, Mladen Rakovic, Dragan Gasevic |
AIED (2) | 1 |
| 2021 | A Multi-Objective Grammatical Evolution Framework to Generate Convolutional Neural Network ArchitecturesabstractDeep Convolutional Neural Networks (CNNs) have reached the attention in the last decade due to their successful application to many computer vision domains. Several handcrafted architectures have been proposed in the literature, with increasing depth and millions of parameters. However, the optimal architecture size and parameters setup are dataset-dependent and challenging to find. For addressing this problem, this work proposes a Multi-Objective Grammatical Evolution framework to automatically generate suitable CNN architectures (layers and parameters) for a given classification problem. For this, a Context-free Grammar is developed, representing the search space of possible CNN architectures. The proposed method seeks to find suitable network architectures considering two objectives: accuracy and F1-score. We evaluated our method on CIFAR-10, and the results obtained show that our method generates simpler CNN architectures and overcomes the results achieved by larger (more complex) state-of-the-art CNN approaches and other grammars. Cleber A. C. F. da Silva, Daniel Carneiro Rosa, Péricles B. C. Miranda, Filipe R. Cordeiro, Tapas Si, André C. A. Nascimento, Rafael Ferreira Leite de Mello, Paulo S. G. de Mattos Neto |
CEC | 7 |
| 2021 | Towards automated content analysis of feedback: A multi-language study
Ikenna Osakwe, Alexander Whitelock-Wainwright, Guanliang Chen, Rafael Ferreira Leite de Mello, Anderson Pinheiro Cavalcanti, Dragan Gasevic |
EDM | 4 |
| 2021 | Reducing the size of training datasets in the classification of online discussionsabstractSupervised machine learning models have been widely used to address the classification of messages in online discussions. Supervised learning algorithms require a large set of annotated data to accurately create a predictive model. However, data annotation is a complex task due to three factors: (i) depends on specialists to accurately label data; (ii) it is often a time-consuming and labour-intensive work,and(iii) in educational settings, it is not always easy to collect a substantial volume of data required by the machine learning algorithms. This paper presents an active learning-based approach that can reduce the amount of annotated data required to build machine learning models for the classification of educational data. The results obtained show that with only 20% of the annotated data, the proposed approach achieved similar results to those presented in the previous works that used the complete databases to train the machine learning model. Vitor Rolim, Rafael Ferreira Leite de Mello, André C. A. Nascimento, Rafael Dueire Lins, Dragan Gasevic |
ICALT | 2 |
| 2021 | Fostering Autonomy through Augmentative and Alternative CommunicationabstractAccording to the World Health Organization, an estimated one billion people live with a disability. Millions of them are non-verbal and also experience motor-skill challenges. The restrictions on participation and communication caused by such disabilities often lead to discrimination and social exclusion, including the lack of access to formal education. Augmentative and Alternative Communication (AAC) is a method to afford communication for people with speech impairment. Software applications that implement AAC bring benefits of adaptability and personalization over traditional paper-based methods, but their usability needs improvement, particularly to increase user autonomy. This paper presents an interface redesign of the Livox AAC application, and a new user onboarding process based on user research to adjust the interface to user needs, contributing to user autonomy on AAC use. João P. C. Uchoa, Taciana Pontual Falcão, André C. A. Nascimento, Péricles B. C. Miranda, Rafael Ferreira Leite de Mello |
ICALT | 5 |
| 2021 | The impact of automatic text translation on classification of online discussions for social and cognitive presencesabstractThis paper reports the findings of a study that measured the effectiveness of employing automatic text translation methods in automated classification of online discussion messages according to the categories of social and cognitive presences. Specifically, we examined the classification of 1,500 Portuguese and 1,747 English discussion messages using classifiers trained on the datasets before and after the application of text translation. While the English model generated, with the original and translated texts, achieved results (accuracy and Cohen’s κ) similar to those of the previously reported studies, the translation to Portuguese led to a decrease in the performance. The indicates the general viability of the proposed approach when converting the text to English. Moreover, this study highlighted the importance of different features and resources, and the limitations of the resources for Portuguese as reasons of the results obtained. Arthur Barbosa, Maverick Andre Dionisio Ferreira, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Dragan Gasevic |
LAK | 3 |
| 2021 | A Teacher-facing Learning Analytics Dashboard for Process-oriented Feedback in Online LearningabstractIn online learning, teachers need constant feedback about their students’ progress and regulation needs. Learning Analytics Dashboards for process-oriented feedback can be a valuable tool for this purpose. However, few such dashboards have been proposed in literature, and most of them lack empirical validation or grounding in learning theories. We present a teacher-facing dashboard for process-oriented feedback in online learning, co-designed and evaluated through an iterative design process involving teachers and visualization experts. We also reflect on our design process by discussing the challenges, pitfalls, and successful strategies for building this type of dashboard. Raphael A. Dourado, Rodrigo L. Rodrigues, Nivan Ferreira, Rafael Ferreira Leite de Mello, Alex Sandro Gomes, Katrien Verbert |
LAK | 4 |
| 2021 | Student appreciation of data-driven feedback: A pilot study on OnTaskabstractFeedback plays a crucial role in student learning. Learning analytics (LA) has demonstrated potential in addressing prominent challenges with feedback practice, such as enabling timely feedback based on insights obtained from large data sets. However, there is insufficient research looking into relations between student expectations of feedback and their experience with LA-based feedback. This paper presents a pilot study that examined students’ experience of LA-based feedback, offered with the OnTask system, taking into consideration the factors of students’self-efficacy and self-regulation skills. Two surveys were carried out at a Brazilian university, and the results highlighted important implications for LA-based feedback practice, including leveraging the ‘partnership’ between the human teacher and the computer, and developing feedback literacy among learners. Yi-Shan Tsai, Rafael Ferreira Leite de Mello, Jelena Jovanovic 0001, Dragan Gasevic |
LAK | 2 |
| 2021 | ImageDataset2Vec: An image dataset embedding for algorithm selection
Lucas V. Dias, Péricles B. C. Miranda, André C. A. Nascimento, Filipe R. Cordeiro, Rafael Ferreira Leite de Mello, Ricardo B. C. Prudêncio |
Expert Syst. Appl. | 5 |
| 2020 | DocEng'2020 Competition on Extractive Text SummarizationabstractThe DocEng'2020 Competition on Extractive Text Summarization assessed the performance of six new methods and fourteen classical algorithms for extractive text sumarization. The systems were evaluated using the CNN-Corpus, the largest test set available today for single document extractive summarization using two different strategies and the ROUGE and the direct match measures. Rafael Dueire Lins, Rafael Ferreira Leite de Mello, Steven J. Simske |
DocEng | 2 |
| 2020 | An Assessment of Sentence Simplification Methods in Extractive Text SummarizationabstractThe unprecedented growth of textual content on the Web made essential the development of automatic or semi-automatic techniques to help people to find valuable information in such a huge heap of text data. Automatic text summarization is one of such techniques that is being pointed out as offering a viable solution in such a chaotic scenario. Extractive text summarization, in particular, selects a set of sentences from a text according to specific criteria. Strategies for extractive summarization can benefit from preprocessing techniques that emphasize the relevance or infor-mativeness of sentences with respect to the selection criteria. This paper tests such a hypothesis using sentence simplification methods. Four methods are used to simplify a corpus of news articles in English: a rule-based method, an optimization method, a supervised deep learning model and an unsupervised deep learning model. The simplified outputs are summarized using 14 sentence selection strategies. The combinations of simplification and summarization methods are compared with the baseline --- the summarized corpus without previous simplification --- with a quantitative analysis, which suggests sentence compression with restrictions and models learned from large parallel corpora tend to perform better and yield gains over summarization without prior simplification. Rafaella F. Vale, Rafael Dueire Lins, Rafael Ferreira Leite de Mello |
DocEng | 3 |
| 2020 | Towards a Maturity Model for Learning Analytics Adoption An Overview of its Levels and AreasabstractLearning Analytics is a new field in education whose adoption can bring benefits for teaching and learning processes. However, many higher education institutions may not be ready to start using learning analytics due to challenges such as organizational culture, infrastructure, and privacy. In this context, Maturity Models (MMs) can support institutions to systematize their processes, enabling them to progress successively in the learning analytics adoption. MMs are used in different fields to support the improvement of processes, describing them in terms of maturity levels, and identifying enhancements that could lead an organization to higher levels of such maturity. Thus, this paper presents an outline of a MM for Learning Analytics adoption in higher education institutions, describing its levels and areas, together with its development methodology. Elyda L. S. X. Freitas, Fernando da Fonseca de Souza, Vinicius Cardoso Garcia, Rafael Ferreira Leite de Mello, Dragan Gasevic |
ICALT | 4 |
| 2020 | Towards automatic cross-language classification of cognitive presence in online discussionsabstractThis paper presents a study that examined automated cross-language classification of online discussion messages for the levels of cognitive presence, a key construct from the widely used Community of Inquiry (CoI) model of online learning. Specifically, we examined the classification of 1,500 Portuguese language discussion messages using a classifier trained on a corpus of the 1,747 English language discussion messages. In the study, a random forest classifier was developed using a small set of 108 validated indicators of psychological processes, linguistic coherence, and online discussion structure. The classifier obtained 67% accuracy and Cohen's κ of 0.32, showing a moderate level of inter-rater agreement above chance and the general viability of the proposed approach. Most importantly, the findings suggest that certain aspects of cognitive presence construct are highly generalizable and transfer across different languages. Finally, the paper also presents a novel method for addressing class imbalance problem using a generic algorithm heuristic technique, which provided substantial improvements over the use of imbalanced dataset. Results and practical implications are further discussed. Gian Barbosa, Raissa Camelo, Anderson Pinheiro Cavalcanti, Péricles B. C. Miranda, Rafael Ferreira Leite de Mello, Vitomir Kovanovic, Dragan Gasevic |
LAK | 5 |
| 2020 | How good is my feedback?: a content analysis of written feedbackabstractFeedback is a crucial element in helping students identify gaps and assess their learning progress. In online courses, feedback becomes even more critical as it is one of the resources where the teacher interacts directly with the student. However, with the growing number of students enrolled in online learning, it becomes a challenge for instructors to provide good quality feedback that helps the student self-regulate. In this context, this paper proposed a content analysis of feedback text provided by instructors based on different indicators of good feedback. A random forest classifier was trained and evaluated at different feedback levels. The results achieved outcomes up to 87% and 0.39 of accuracy and Cohen's κ, respectively. The paper also provides insights into the most influential textual features of feedback that predict feedback quality. Anderson Pinheiro Cavalcanti, Arthur Diego, Rafael Ferreira Leite de Mello, Katerina Mangaroska, André C. A. Nascimento, Fred Freitas, Dragan Gasevic |
LAK | 3 |
| 2020 | Let's shine together!: a comparative study between learning analytics and educational data miningabstractLearning Analytics and Knowledge (LAK) and Educational Data Mining (EDM) are two of the most popular venues for researchers and practitioners to report and disseminate discoveries in data-intensive research on technology-enhanced education. After the development of about a decade, it is time to scrutinize and compare these two venues. By doing this, we expected to inform relevant stakeholders of a better understanding of the past development of LAK and EDM and provide suggestions for their future development. Specifically, we conducted an extensive comparison analysis between LAK and EDM from four perspectives, including (i) the topics investigated; (ii) community development; (iii) community diversity; and (iv) research impact. Furthermore, we applied one of the most widely-used language modeling techniques (Word2Vec) to capture words used frequently by researchers to describe future works that can be pursued by building upon suggestions made in the published papers to shed light on potential directions for future research. Guanliang Chen, Vitor Rolim, Rafael Ferreira Leite de Mello, Dragan Gasevic |
LAK | 3 |
| 2020 | Perceptions and expectations about learning analytics from a brazilian higher education institutionabstractSeveral tools to support learning processes based on educational data have emerged from research on Learning Analytics (LA) in the last few years. These tools aim to support students and instructors in daily activities, and academic managers in making institutional decisions. Although the adoption of LA tools is spreading, the field still needs to deepen the understanding of the contexts where learning takes place, and of the views of the stakeholders involved in implementing and using these tools. In this sense, the SHEILA framework proposes a set of instruments to perform a detailed analysis of the expectations and needs of different stakeholders in higher education institutions, regarding the adoption of LA. Moreover, there is a lacuna in research on stakeholders' expectations from LA outside the Global North. Therefore, this paper reports on the findings of the application of interviews and focus groups, based on the SHEILA framework, with students and teaching staff from a Brazilian public university, to investigate their perceptions of the potential benefits and risks of using LA in higher education in the country. Findings indicate that there is a high interest in using LA for improving the learning experience, in particular, being able to provide personalized feedback, to adapt teaching practices to students' needs, and to make evidence-based pedagogical decisions. From the analysis of these perspectives, we point to opportunities for using LA in Brazilian higher education. Taciana Pontual Falcão, Rafael Ferreira Leite de Mello, Rodrigo L. Rodrigues, Juliana R. Basto Diniz, Yi-Shan Tsai, Dragan Gasevic |
LAK | 2 |
| 2020 | Towards automatic content analysis of social presence in transcripts of online discussionsabstractThis paper presents an approach to automatic labeling of the content of messages in online discussion according to the categories of social presence. To achieve this goal, the proposed approach is based on a combination of traditional text mining features and word counts extracted with the use of established linguistic frameworks (i.e., LIWC and Coh-metrix). The best performing classifier obtained 0.95 and 0.88 for accuracy and Cohen's kappa, respectively. This paper also provides some theoretical insights into the nature of social presence by looking at the classification features that were most relevant for distinguishing between the different categories. Finally, this study adopted epistemic network analysis to investigate the structural construct validity of the automatic classification approach. Namely, the analysis showed that the epistemic networks produced based on messages manually and automatically coded produced nearly identical results. This finding thus produced evidence of the structural validity of the automatic approach. Maverick Andre Dionisio Ferreira, Vitor Rolim, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Guanliang Chen, Dragan Gasevic |
LAK | 3 |
| 2020 | A multi-objective optimization approach for the group formation problem
Péricles B. C. Miranda, Rafael Ferreira Leite de Mello, André C. A. Nascimento |
Expert Syst. Appl. | 2 |
| 2019 | I Wanna Talk Like You: Speaker Adaptation to Dialogue Style in L2 Practice Conversation
Arabella Sinclair, Rafael Ferreira Leite de Mello, Dragan Gasevic, Christopher G. Lucas, Adam Lopez |
AIED (2) | 2 |
| 2019 | DocEng'19 Competition on Extractive Text SummarizationabstractThe DocEng'19 Competition on Extractive Text Summarization assessed the performance of two new and fourteen previously published extractive text sumarization methods. The competitors were evaluated using the CNN-Corpus, the largest test set available today for single document extractive summarization. Rafael Dueire Lins, Rafael Ferreira Leite de Mello, Steven J. Simske |
DocEng | 2 |
| 2019 | The CNN-Corpus: A Large Textual Corpus for Single-Document Extractive SummarizationabstractThis paper details the features and the methodology adopted in the construction of the CNN-corpus, a test corpus for single document extractive text summarization of news articles. The current version of the CNN-corpus encompasses 3,000 texts in English, and each of them has an abstractive and an extractive summary. The corpus allows quantitative and qualitative assessments of extractive summarization strategies. Rafael Dueire Lins, Hilário Oliveira, Luciano de Souza Cabral, Jamilson Batista, Bruno Tenório Ávila, Rafael Ferreira Leite de Mello, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske |
DocEng | 6 |
| 2019 | The CNN-Corpus in Spanish: a Large Corpus for Extractive Text Summarization in the Spanish LanguageabstractThis paper details the development and features of the CNN-corpus in Spanish, possibly the largest test corpus for single document extractive text summarization in the Spanish language. Its current version encompasses 1,117 well-written texts in Spanish, each of them has an abstractive and an extractive summary. The development methodology adopted allows good-quality qualitative and quantitative assessments of summarization strategies for tools developed in the Spanish language. Rafael Dueire Lins, Hilário Oliveira, Luciano de Souza Cabral, Jamilson Batista, Bruno Tenório Ávila, Diego A. Salcedo, Rafael Ferreira Leite de Mello, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske |
DocEng | 7 |
| 2019 | Predictors of Student Satisfaction: A Large-scale Study of Human-Human Online Tutorial Dialogues
Guanliang Chen, David Lang, Rafael Ferreira Leite de Mello, Dragan Gasevic |
EDM | 3 |
| 2019 | An Analysis of the use of Good Feedback Practices in Online Learning CoursesabstractFeedback is an essential component of any learning experience. It allows students to identify gaps in their learning and improve their self-regulation. However, providing useful feedback is a challenging and time-consuming task. In digital learning environments, this challenge is even more significant due to a large number of students. Thus, this paper reports on the findings of an analysis of the quality of feedback provided by instructors in an online course. The paper also proposes a supervised machine learning algorithm that can identify the presence of good practices in feedback messages sent to students in a digital learning environment. The results reveal the most commonly used kinds of feedback and how to identify them automatically. The results of the study could potentially be used to improve the quality of the feedback provided by instructors in online education. Anderson Pinheiro Cavalcanti, Rafael Ferreira Leite de Mello, Vitor Rolim, Maverick Andre Dionisio Ferreira, Fred Freitas, Dragan Gasevic |
ICALT | 2 |
| 2019 | Students' Perceptions about Learning Analytics in a Brazilian Higher Education InstitutionabstractIn recent years, several tools to support learning processes have emerged from the research on Learning Analytics (LA). However, little attention has been given to the contexts where learning takes place, and to the stakeholders involved in implementing and using these tools. Following the SHEILA Framework, we present the results of interviews and focus groups with students from a Brazilian public university on their perceptions of the potential and risks of using LA in higher education. Findings indicate great difficulties to use the Learning Management System; lack of effective communication with teaching staff; lack of quality and timely feedback; high perceived value of personalization; rejection of competition and rankings; issues with self-confidence; and trustworthiness of the educational institution. From the analysis of students' perspectives, we point to opportunities to use LA in the context of a Brazilian public university. Taciana Pontual Falcão, Rafael Ferreira Leite de Mello, Rodrigo L. Rodrigues, Juliana R. Basto Diniz, Dragan Gasevic |
ICALT | 2 |
| 2019 | Using Social Network Analysis to Measure the Effect of Learning Analytics in Computing EducationabstractStudent retention and learning in STEM disciplines is a growing problem. The 2012 report by the US President's Council of Advisors on Science and Technology (PCAST) predicts a future deficit in science, engineering, and mathematics (STEM) in the following decade and emphasizes the importance of addressing this issue. With this as a motivating factor, the OSBLE+ Social Programming Environment (SPE) was used to leverage social and programming data for the basis of automatically generated prompts inserted into the SPE. These prompts were designed to stimulate help-seeking, help-giving, and social interaction in the learning environment. A social network analysis was performed in order to determine whether exposure to the automated interventions would positively affect the relationship among students over time. Results of this study suggest that students in the experimental treatment who were presented with automated prompts developed more connected and social networks than those in the control treatment. Daniel M. Olivares, Rafael Ferreira Leite de Mello, Olusola O. Adesope, Vitor Rolim, Dragan Gasevic, Christopher D. Hundhausen |
ICALT | 2 |
| 2019 | Identifying Students' Weaknesses and Strengths Based on Online Discussion using Topic ModelingabstractThis paper proposes a topic model-based approach to extract students' weaknesses and strength based on Latent Dirichlet Allocation (LDA). Our approach combines textual data extracted from online discussion forums written by students with external sources like Wikipedia. The results show the effectiveness of the proposed approach to create a user profile based on the topics covered by the students in discussion forums. Vitor Rolim, Rafael Ferreira Leite de Mello, Maverick Andre Dionisio Ferreira, Anderson Pinheiro Cavalcanti, Rinaldo Lima |
ICALT | 2 |
| 2019 | Analysing Social Presence in Online Discussions Through Network and Text AnalyticsabstractThis paper presents an approach to studying relationships between students' social presence and course topics from transcripts of asynchronous discussions in online learning environments. Specifically, the paper uses topic modeling and epistemic network analysis to investigate how students' social presence is expressed across different course topics. Finally, we show how this method can be adopted to examine how students' social presence changed due to an instructional intervention. The results of this study and its implications are further discussed. Vitor Rolim, Rafael Ferreira Leite de Mello, Vitomir Kovanovic, Dragan Gasevic |
ICALT | 2 |
| 2019 | Using Artificial Intelligence for Augmentative Alternative Communication for Children with Disabilities
Rodica Neamtu, André Câmara, Carlos Pereira, Rafael Ferreira Leite de Mello |
INTERACT (1) | 4 |
| 2018 | Towards Combined Network and Text Analytics of Student Discourse in Online Discussions
Rafael Ferreira Leite de Mello, Vitomir Kovanovic, Dragan Gasevic, Vitor Rolim |
AIED (1) | 1 |
| 2018 | Automated Analysis of Cognitive Presence in Online Discussions Written in Portuguese
Valter Neto, Vitor Rolim, Rafael Ferreira Leite de Mello, Vitomir Kovanovic, Dragan Gasevic, Rafael Dueire Lins, Rodrigo L. Rodrigues |
EC-TEL | 3 |
| 2018 | Combining sentence similarities measures to identify paraphrases
Rafael Ferreira Leite de Mello, George D. C. Cavalcanti, Fred Freitas, Rafael Dueire Lins, Steven J. Simske, Marcelo Riss |
Comput. Speech Lang. | 1 |
| 2016 | Mobile Summarizer and News Summary Navigator: Two Multilingual News Article Summarization Tools for Mobile DevicesabstractMobile devices such as smart phones and tablets are omnipresent in modern societies. Such devices allow browsing the Internet. This paper briefly describes two tools for news article summarization in mobile devices that attempts to automatically collect and sieve the most important information of news article in WebPages. Luciano de Souza Cabral, Manoel Neto, Artur Borges, Rafael Dueire Lins, Rinaldo Lima, Rafael Ferreira Leite de Mello, Marcelo Riss, Steven J. Simske |
DocEng | 6 |
| 2016 | Appling Link Target Identification and Content Extraction to improve Web News SummarizationabstractThe existing automatic text summarization systems whenever applied to web-pages of news articles show poor performance as the text is encapsulated within a HTML page. This paper takes advantage of the link identification and content extraction techniques. The results show the validity of such a strategy. Rodolfo Ferreira, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Hilário Oliveira, Marcelo Riss, Steven J. Simske |
DocEng | 2 |
| 2016 | A Hypermedia-based Adaptive Educational System for Assisting Students in Systems and Information Technology Domain for Accountability
Inés Maria González Vidal, Evandro de Barros Costa, Leandro Dias da Silva, Fabrísia Ferreira de Araújo, Rafael Ferreira Leite de Mello |
WorldCIST (2) | 5 |
| 2016 | Assessing sentence similarity through lexical, syntactic and semantic analysis
Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Steven J. Simske, Fred Freitas, Marcelo Riss |
Comput. Speech Lang. | 1 |
| 2016 | Assessing shallow sentence scoring techniques and combinations for single and multi-document summarization
Hilário Oliveira, Rafael Ferreira Leite de Mello, Rinaldo Lima, Rafael Dueire Lins, Fred Freitas, Marcelo Riss, Steven J. Simske |
Expert Syst. Appl. | 2 |
| 2015 | A Quantitative and Qualitative Assessment of Automatic Text Summarization SystemsabstractText summarization is the process of automatically creating a shorter version of one or more text documents. This paper presents a qualitative and quantitative assessment of the 22 state-of-the-art extractive summarization systems using the CNN corpus, a dataset of 3,000 news articles. Jamilson Batista, Rodolfo Ferreira, Hilário Tomaz, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Steven J. Simske, Gabriel Pereira e Silva, Marcelo Riss |
DocEng | 4 |
| 2015 | Automatic Document Classification using Summarization StrategiesabstractAn efficient way to automatically classify documents may be provided by automatic text summarization, the task of creating a shorter text from one or several documents. This paper presents an assessment of the 15 most widely used methods for automatic text summarization from the text classification perspective. A naive Bayes classifier was used showing that some of the methods tested are better suited for such a task. Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Luciano de Souza Cabral, Fred Freitas, Steven J. Simske, Marcelo Riss |
DocEng | 1 |
| 2015 | Automatic Text Document Summarization Based on Machine LearningabstractThe need for automatic generation of summaries gained importance with the unprecedented volume of information available in the Internet. Automatic systems based on extractive summarization techniques select the most significant sentences of one or more texts to generate a summary. This article makes use of Machine Learning techniques to assess the quality of the twenty most referenced strategies used in extractive summarization, integrating them in a tool. Quantitative and qualitative aspects were considered in such assessment demonstrating the validity of the proposed scheme. The experiments were performed on the CNN-corpus, possibly the largest and most suitable test corpus today for benchmarking extractive summarization strategies. Gabriel Pereira e Silva, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Luciano de Souza Cabral, Hilário Oliveira, Steven J. Simske, Marcelo Riss |
DocEng | 2 |
| 2014 | A Context Based Text Summarization SystemabstractText summarization is the process of creating a shorter version of one or more text documents. Automatic text summarization has become an important way of finding relevant information in large text libraries or in the Internet. Extractive text summarization techniques select entire sentences from documents according to some criteria to form a summary. Sentence scoring is the technique most used for extractive text summarization, today. Depending on the context, however, some techniques may yield better results than some others. This paper advocates the thesis that the quality of the summary obtained with combinations of sentence scoring methods depend on text subject. Such hypothesis is evaluated using three different contexts: news, blogs and articles. The results obtained show the validity of the hypothesis formulated and point at which techniques are more effective in each of those contexts studied. Rafael Ferreira Leite de Mello, Fred Freitas, Luciano de Souza Cabral, Rafael Dueire Lins, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske, Luciano Favaro |
Document Analysis Systems | 1 |
| 2014 | A platform for language independent summarizationabstractThe text data available on the Internet is not only huge in volume, but also in diversity of subject, quality and idiom. Such factors make it infeasible to efficiently scavenge useful information from it. Automatic text summarization is a possible solution for efficiently addressing such a problem, because it aims to sieve the relevant information in documents by creating shorter versions of the text. However, most of the techniques and tools available for automatic text summarization are designed only for the English language, which is a severe restriction. There are multilingual platforms that support, at most, 2 languages. This paper proposes a language independent summarization platform that provides corpus acquisition, language classification, translation and text summarization for 25 different languages. Luciano de Souza Cabral, Rafael Dueire Lins, Rafael Ferreira Leite de Mello, Fred Freitas, Bruno Tenório Ávila, Steven J. Simske, Marcelo Riss |
ACM Symposium on Document Engineering | 3 |
| 2014 | A new sentence similarity assessment measure based on a three-layer sentence representationabstractSentence similarity is used to measure the degree of likelihood between sentences. It is used in many natural language applications, such as text summarization, information retrieval, text categorization, and machine translation. The current methods for assessing sentence similarity represent sentences as vectors of bag of words or the syntactic information of the words in the sentence. The degree of likelihood between phrases is calculated by composing the similarity between the words in the sentences. Two important concerns in the area, the meaning problem and the word order, are not handled, however. This paper proposes a new sentence similarity assessment measure that largely improves and refines a recently published method that takes into account the lexical, syntactic and semantic components of sentences. The new method proposed here was benchmarked using a publically available standard dataset. The results obtained show that the new similarity assessment measure proposed outperforms the state of the art systems and achieve results comparable to the evaluation made by humans. Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Fred Freitas, Steven J. Simske, Marcelo Riss |
ACM Symposium on Document Engineering | 1 |
| 2014 | Transforming graph-based sentence representations to alleviate overfitting in relation extractionabstractRelation extraction (RE) aims at finding the way entities, such as person, location, organization, date, etc., depend upon each other in a text document. Ontology Population, Automatic Summarization, and Question Answering are fields in which relation extraction offers valuable solutions. A relation extraction method based on inductive logic programming that induces extraction rules suitable to identify semantic relations between entities was proposed by the authors in a previous work. This paper proposes a method to simplify graph-based representations of sentences that replaces dependency graphs of sentences by simpler ones, keeping the target entities in it. The goal is to speed up the learning phase in a RE framework, by applying several rules for graph simplification that constrain the hypothesis space for generating extraction rules. Moreover, the direct impact on the extraction performance results is also investigated. The proposed techniques outperformed some other state-of-the-art systems when assessed on two standard datasets for relation extraction in the biomedical domain. Rinaldo Lima, Jamilson Batista, Rafael Ferreira Leite de Mello, Fred Freitas, Rafael Dueire Lins, Steven J. Simske, Marcelo Riss |
ACM Symposium on Document Engineering | 3 |
| 2014 | A multi-document summarization system based on statistics and linguistic treatment
Rafael Ferreira Leite de Mello, Luciano de Souza Cabral, Fred Freitas, Rafael Dueire Lins, Gabriel Pereira e Silva, Steven J. Simske, Luciano Favaro |
Expert Syst. Appl. | 1 |
| 2013 | An Inductive Logic Programming-Based Approach for Ontology Population from the Web
Rinaldo Lima, Bernard Espinasse, Hilário Oliveira, Rafael Ferreira Leite de Mello, Luciano de Souza Cabral, Dimas Filho, Fred Freitas, Renê Gadelha |
DEXA (1) | 4 |
| 2013 | A Four Dimension Graph Model for Automatic Text SummarizationabstractText summarization is the process of automatically creating a shorter version of one or more text documents. In this context, word-based, sentence-based and graph-based methods approaches are largely used. Among these, graph based methods for automatic text summarization produce summaries based on the relationships between sentences. These relationships may also support the creation of several text processing applications such as extractive and abstractive summaries, question-answering and information retrieval systems, among others. A new graph model for text processing applications is proposed in this paper. It relies on four dimensions (similarity, semantic similarity, co reference, discourse information) to create the graph. The rationale behind the proposal presented here is resorting to more dimensions than previous works, and taking into account co reference resolution, taking into account to the role of pronouns in connecting the sentences. Co reference was not used in any previous graph based summarization technique. An experiment was performed using the Text Rank algorithm with the presented approach, on the CNN corpus. The results show that the model proposed here outperforms the current approaches both quantitatively and qualitatively. Rafael Ferreira Leite de Mello, Fred Freitas, Luciano de Souza Cabral, Rafael Dueire Lins, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske, Luciano Favaro |
Web Intelligence | 1 |
| 2013 | Assessing sentence scoring techniques for extractive text summarization
Rafael Ferreira Leite de Mello, Luciano de Souza Cabral, Rafael Dueire Lins, Gabriel Pereira e Silva, Fred Freitas, George D. C. Cavalcanti, Rinaldo Lima, Steven J. Simske, Luciano Favaro |
Expert Syst. Appl. | 1 |
| 2013 | RetriBlog: An architecture-centered framework for developing blog crawlers
Rafael Ferreira Leite de Mello, Fred Freitas, Patrick Henrique da S. Brito, Jean Melo, Rinaldo Lima, Evandro de Barros Costa |
Expert Syst. Appl. | 1 |
| 2012 | A Confidence-Weighted Metric for Unsupervised Ontology Population from Web Texts
Hilário Oliveira, Rinaldo Lima, Rafael Ferreira Leite de Mello, Fred Freitas, Evandro de Barros Costa |
DEXA (1) | 4 |
| 2012 | Towards an Ontology-Based System to Improve Usability in Collaborative Learning Environments
Endhe Elias, Dalgoberto Miquilino, Ig Ibert Bittencourt, Thyago Tenório, Rafael Ferreira Leite de Mello, Alan Silva, Seiji Isotani, Patrícia A. Jaques |
ITS | 5 |
| 2012 | A framework for building web mining applications in the world of blogs: A case study in product sentiment analysis
Evandro de Barros Costa, Rafael Ferreira Leite de Mello, Patrick Henrique da S. Brito, Ig Ibert Bittencourt, Olavo Holanda, Aydano Machado, Tarsis Marinho |
Expert Syst. Appl. | 2 |