Luciano de Souza Cabral

dblp:130/8106 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0002-4235-5753ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 9 · 2 first-authorArtificial intelligence and machine learning · 4Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2025 Automatic Short Answer Grading in the LLM Era: Does GPT-4 with Prompt Engineering beat Traditional Models?
abstract
Assessing short answers in educational settings is challenging due to the need for scalability and accuracy, which led to the field of Automatic Short Answer Grading (ASAG). Traditional machine learning models, such as ensemble and embeddings, have been widely researched in ASAG, but they often suffer from generalizability issues. Recently, Large Language Models (LLMs) emerged as an alternative to optimize ASAG systems. However, previous research has failed to present a comprehensive analysis of LLMs' performance powered by prompt engineering strategies and compare its capabilities to traditional models. This study presents a comparative analysis between traditional machine learning models and GPT-4 in the context of ASAG. We investigated the effectiveness of different models and text representation techniques and explored prompt engineering strategies for LLMs. The results indicate that traditional machine learning models outperform LLMs. However, GPT-4 showed promising capabilities, especially when configured with optimized prompt components, such as few-shot examples and clear instructions. This study contributes to the literature by providing a detailed evaluation of LLM performance compared to traditional machine learning models in a multilingual ASAG context, offering insights for developing more efficient automatic grading systems.
Rafael Ferreira Leite de Mello, Cleon Pereira Junior, Luiz A. L. Rodrigues, Filipe D. Pereira, Luciano de Souza Cabral, Newarney Torrezão da Costa, Geber L. Ramalho, Dragan Gasevic
LAK5
2024 Can GPT4 Answer Educational Tests? Empirical Analysis of Answer Quality Based on Question Complexity and Difficulty
Luiz A. L. Rodrigues, Filipe D. Pereira, Luciano de Souza Cabral, Geber L. Ramalho, Dragan Gasevic, Rafael Ferreira Leite de Mello
AIED (1)3
2023 Evaluation of a Hybrid AI-Human Recommender for CS1 Instructors in a Real Educational Scenario
Filipe D. Pereira, Elaine Harada T. de Oliveira, Luiz A. L. Rodrigues, Luciano de Souza Cabral, David B. F. Oliveira, Leandro S. G. Carvalho, Dragan Gasevic, Alexandra I. Cristea, Diego Dermeval, Rafael Ferreira Leite de Mello
EC-TEL4
2019 The CNN-Corpus: A Large Textual Corpus for Single-Document Extractive Summarization
abstract
This paper details the features and the methodology adopted in the construction of the CNN-corpus, a test corpus for single document extractive text summarization of news articles. The current version of the CNN-corpus encompasses 3,000 texts in English, and each of them has an abstractive and an extractive summary. The corpus allows quantitative and qualitative assessments of extractive summarization strategies.
Rafael Dueire Lins, Hilário Oliveira, Luciano de Souza Cabral, Jamilson Batista, Bruno Tenório Ávila, Rafael Ferreira Leite de Mello, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske
DocEng3
2019 The CNN-Corpus in Spanish: a Large Corpus for Extractive Text Summarization in the Spanish Language
abstract
This paper details the development and features of the CNN-corpus in Spanish, possibly the largest test corpus for single document extractive text summarization in the Spanish language. Its current version encompasses 1,117 well-written texts in Spanish, each of them has an abstractive and an extractive summary. The development methodology adopted allows good-quality qualitative and quantitative assessments of summarization strategies for tools developed in the Spanish language.
Rafael Dueire Lins, Hilário Oliveira, Luciano de Souza Cabral, Jamilson Batista, Bruno Tenório Ávila, Diego A. Salcedo, Rafael Ferreira Leite de Mello, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske
DocEng3
2016 Mobile Summarizer and News Summary Navigator: Two Multilingual News Article Summarization Tools for Mobile Devices
abstract
Mobile devices such as smart phones and tablets are omnipresent in modern societies. Such devices allow browsing the Internet. This paper briefly describes two tools for news article summarization in mobile devices that attempts to automatically collect and sieve the most important information of news article in WebPages.
Luciano de Souza Cabral, Manoel Neto, Artur Borges, Rafael Dueire Lins, Rinaldo Lima, Rafael Ferreira Leite de Mello, Marcelo Riss, Steven J. Simske
DocEng1
2015 Automatic Document Classification using Summarization Strategies
abstract
An efficient way to automatically classify documents may be provided by automatic text summarization, the task of creating a shorter text from one or several documents. This paper presents an assessment of the 15 most widely used methods for automatic text summarization from the text classification perspective. A naive Bayes classifier was used showing that some of the methods tested are better suited for such a task.
Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Luciano de Souza Cabral, Fred Freitas, Steven J. Simske, Marcelo Riss
DocEng3
2015 Automatic Text Document Summarization Based on Machine Learning
abstract
The need for automatic generation of summaries gained importance with the unprecedented volume of information available in the Internet. Automatic systems based on extractive summarization techniques select the most significant sentences of one or more texts to generate a summary. This article makes use of Machine Learning techniques to assess the quality of the twenty most referenced strategies used in extractive summarization, integrating them in a tool. Quantitative and qualitative aspects were considered in such assessment demonstrating the validity of the proposed scheme. The experiments were performed on the CNN-corpus, possibly the largest and most suitable test corpus today for benchmarking extractive summarization strategies.
Gabriel Pereira e Silva, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Luciano de Souza Cabral, Hilário Oliveira, Steven J. Simske, Marcelo Riss
DocEng4
2014 A Context Based Text Summarization System
abstract
Text summarization is the process of creating a shorter version of one or more text documents. Automatic text summarization has become an important way of finding relevant information in large text libraries or in the Internet. Extractive text summarization techniques select entire sentences from documents according to some criteria to form a summary. Sentence scoring is the technique most used for extractive text summarization, today. Depending on the context, however, some techniques may yield better results than some others. This paper advocates the thesis that the quality of the summary obtained with combinations of sentence scoring methods depend on text subject. Such hypothesis is evaluated using three different contexts: news, blogs and articles. The results obtained show the validity of the hypothesis formulated and point at which techniques are more effective in each of those contexts studied.
Rafael Ferreira Leite de Mello, Fred Freitas, Luciano de Souza Cabral, Rafael Dueire Lins, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske, Luciano Favaro
Document Analysis Systems3
2014 A platform for language independent summarization
abstract
The text data available on the Internet is not only huge in volume, but also in diversity of subject, quality and idiom. Such factors make it infeasible to efficiently scavenge useful information from it. Automatic text summarization is a possible solution for efficiently addressing such a problem, because it aims to sieve the relevant information in documents by creating shorter versions of the text. However, most of the techniques and tools available for automatic text summarization are designed only for the English language, which is a severe restriction. There are multilingual platforms that support, at most, 2 languages. This paper proposes a language independent summarization platform that provides corpus acquisition, language classification, translation and text summarization for 25 different languages.
Luciano de Souza Cabral, Rafael Dueire Lins, Rafael Ferreira Leite de Mello, Fred Freitas, Bruno Tenório Ávila, Steven J. Simske, Marcelo Riss
ACM Symposium on Document Engineering1
2014 A multi-document summarization system based on statistics and linguistic treatment
Rafael Ferreira Leite de Mello, Luciano de Souza Cabral, Fred Freitas, Rafael Dueire Lins, Gabriel Pereira e Silva, Steven J. Simske, Luciano Favaro
Expert Syst. Appl.2
2013 An Inductive Logic Programming-Based Approach for Ontology Population from the Web
Rinaldo Lima, Bernard Espinasse, Hilário Oliveira, Rafael Ferreira Leite de Mello, Luciano de Souza Cabral, Dimas Filho, Fred Freitas, Renê Gadelha
DEXA (1)5
2013 A Four Dimension Graph Model for Automatic Text Summarization
abstract
Text summarization is the process of automatically creating a shorter version of one or more text documents. In this context, word-based, sentence-based and graph-based methods approaches are largely used. Among these, graph based methods for automatic text summarization produce summaries based on the relationships between sentences. These relationships may also support the creation of several text processing applications such as extractive and abstractive summaries, question-answering and information retrieval systems, among others. A new graph model for text processing applications is proposed in this paper. It relies on four dimensions (similarity, semantic similarity, co reference, discourse information) to create the graph. The rationale behind the proposal presented here is resorting to more dimensions than previous works, and taking into account co reference resolution, taking into account to the role of pronouns in connecting the sentences. Co reference was not used in any previous graph based summarization technique. An experiment was performed using the Text Rank algorithm with the presented approach, on the CNN corpus. The results show that the model proposed here outperforms the current approaches both quantitatively and qualitatively.
Rafael Ferreira Leite de Mello, Fred Freitas, Luciano de Souza Cabral, Rafael Dueire Lins, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske, Luciano Favaro
Web Intelligence3
2013 Assessing sentence scoring techniques for extractive text summarization
Rafael Ferreira Leite de Mello, Luciano de Souza Cabral, Rafael Dueire Lins, Gabriel Pereira e Silva, Fred Freitas, George D. C. Cavalcanti, Rinaldo Lima, Steven J. Simske, Luciano Favaro
Expert Syst. Appl.2