Hilário Oliveira

dblp:117/7535 · also Hilário Tomaz Alves de Oliveira · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 10 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Large Language Models for chatbot applications handling sensitive information
Hilário Oliveira, Alvaro Sobrinho, Andrey dos Reis Cadima Dias, André Araújo 0003, Rafael Dias Araújo, Diego Dermeval, Sebastian Munoz-Najar Galvez
Expert Syst. Appl.1
2025 Exploring T5-Based Code QA Systems to Support Teaching Programming in Portuguese and English
abstract
Code question and answering (QA) systems have demonstrated potential as practical tools for supporting students in introductory programming courses by providing automated assistance with understanding and debugging source code snippets. Although several studies have examined such systems, they primarily focus on a single language, typically English, resulting in limited investigation of other languages, such as Portuguese, or multilingual settings. This work investigates the performance of three Transformer-based models, CodeT5, Flan-T5, and T5, on a QA task involving the Java programming language. Experiments were conducted using the CodeQA dataset to fine-tune and evaluate the models across three scenarios: (i) the original English data, (ii) a Brazilian Portuguese translation developed in this work, and (iii) a bilingual version combining English and Portuguese data. Model performance was assessed using the ROUGE-L and BERTScore evaluation metrics. The results show that the base version of the CodeT5 model consistently outperformed the other models across all evaluated scenarios. The main contributions of this work are the creation of a Portuguese version of the CodeQA dataset and a comparative analysis of the multilingual performance of T5-based models in code-related QA tasks.
Eduardo Dos S. Lopes, Hilário Oliveira, Kelly Assis de Souza Gazolli
CLEI2
2024 Assessing the Reliability and Validity of the Measures for Automatic Text Summarization
abstract
Automatic Text Summarization (ATS) is a research area that originated in the late 1950s and has gained increasing importance with the surging amount of text data available today. One of the key challenges in this area is how to quantitatively assess the quality of the summaries produced. The three most widely quantitative measures used for this task are: ROUGE, BLEU and BERTScore. This paper attempts to comparatively evaluate the validity and reliability of such measures. The concept of Shannon' entropy from information theory served as background for this work. Experiments were conducted using the CNN corpus, focusing on news articles written in English.
Rafael Dueire Lins, Hilário Oliveira, Steven J. Simske
DocEng2
2024 Assessing Abstractive and Extractive Methods for Automatic News Summarization
abstract
Automatic Text Summarization (ATS) is a research area that originated in the late 1950s and has gained increasing importance with the surge of text data available today. ATS approaches are generally classified into extractive and abstractive methods. Extractive summarization selects the most relevant sentences from a text and copies them verbatim into the summary. Abstractive text summarization aims to generate a more cohesive and concise version of the original text. Recently, pre-trained and large language models such as BERT, BART, Pegasus, Llama, Gemma, and GPT-3 have revolutionized numerous natural language tasks, including creating more humanlike summaries. Despite the progress in recent years, there is room for improvement in directly comparing different summarization tools using a corpus containing high-quality reference summaries. This paper evaluates the quality of summaries generated by such tools, comparing the results obtained from abstractive and extractive summarization methods. Experiments were conducted using the CNN-corpus, focusing on news articles written in English and employing four evaluation measures. The experimental results unveiled that the Pegasus model outperformed others in abstractive summarization. These findings highlight the potential for further advancements in the field, suggesting that using more specialized models explicitly tailored for the summarization task may yield superior results compared to large general-purpose language models.
Hilário Oliveira, Rafael Dueire Lins
DocEng1
2023 Towards explainable prediction of essay cohesion in Portuguese and English
abstract
Textual cohesion is an essential aspect of a formally written text, related to linguistic mechanisms that connect elements such as words, sentences, and paragraphs. Several studies have proposed approaches to estimate textual cohesion in essays automatically. There is limited research that aims to study the extent to which the use of machine learning approaches can predict the textual cohesion of essays written in different languages (not just English). This paper reports on the findings of a study that aimed to propose and evaluate approaches that automatically estimate the cohesion of essays in Portuguese and English. The study proposed regression-based models grounded in conventional feature-based machine learning methods and deep learning-based pre-trained language models. The study also examined the explainability of automated approaches to scrutinize their predictions. We analyzed two datasets composed of 4,570 (Portuguese) and 7,101 (English) essays. The results demonstrate that a deep learning-based model achieved the best performance on both datasets with a moderate Pearson correlation with human-rated cohesion scores. However, the explainability of the automatic cohesion estimations based on conventional machine learning models offered a stronger potential than that of the deep learning model.
Hilário Oliveira, Rafael Ferreira Leite de Mello, Bruno Alexandre Barreiros Rosa, Mladen Rakovic, Péricles B. C. Miranda, Thiago D. Cordeiro, Seiji Isotani, Ig Ibert Bittencourt, Dragan Gasevic
LAK1
2022 Towards automated content analysis of rhetorical structure of written essays using sequential content-independent features in Portuguese
abstract
Brazilian universities have included essay writing assignments in the entrance examination procedure to select prospective students. The essay scorers manually look for the presence of required Rhetorical Structure Theory (RST) categories and evaluate essay coherence. However, identifying RST categories is a time-consuming task. The literature reported several attempts to automate the identification of RST categories in essays with machine learning. Still, previous studies have focused on using machine learning algorithms trained on content-dependent features that can diminish classification performance, leading to over-fitting and hindering model generalisability. Therefore, this paper proposes: (i) the analysis of state-of-the-art classifiers and content-independent features to the task of RST rhetorical moves; (ii) a new approach that considers the sequence of the text to extract features – i.e. sequential content-independent features; (iii) an empirical study about the generalisability of the machine learning models and sequential content-independent features for this context; (iv) the identification of the most predictive features for automated identification of RST categories in essays written in Portuguese. The best performing classifier, XGBoost, based on sequential content-independent features, outperformed the classifiers used in the literature and are based on traditional content-dependent features. The XGBoost classifier based on sequential content-independent features also reached promising accuracy when tested for generalisability.
Rafael Ferreira Leite de Mello, Giuseppe Fiorentino, Hilário Oliveira, Péricles B. C. Miranda, Mladen Rakovic, Dragan Gasevic
LAK3
2021 Towards Automatic Content Analysis of Rhetorical Structure in Brazilian College Entrance Essays
Rafael Ferreira Leite de Mello, Giuseppe Fiorentino, Péricles B. C. Miranda, Hilário Oliveira, Mladen Rakovic, Dragan Gasevic
AIED (2)4
2019 The CNN-Corpus: A Large Textual Corpus for Single-Document Extractive Summarization
abstract
This paper details the features and the methodology adopted in the construction of the CNN-corpus, a test corpus for single document extractive text summarization of news articles. The current version of the CNN-corpus encompasses 3,000 texts in English, and each of them has an abstractive and an extractive summary. The corpus allows quantitative and qualitative assessments of extractive summarization strategies.
Rafael Dueire Lins, Hilário Oliveira, Luciano de Souza Cabral, Jamilson Batista, Bruno Tenório Ávila, Rafael Ferreira Leite de Mello, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske
DocEng2
2019 The CNN-Corpus in Spanish: a Large Corpus for Extractive Text Summarization in the Spanish Language
abstract
This paper details the development and features of the CNN-corpus in Spanish, possibly the largest test corpus for single document extractive text summarization in the Spanish language. Its current version encompasses 1,117 well-written texts in Spanish, each of them has an abstractive and an extractive summary. The development methodology adopted allows good-quality qualitative and quantitative assessments of summarization strategies for tools developed in the Spanish language.
Rafael Dueire Lins, Hilário Oliveira, Luciano de Souza Cabral, Jamilson Batista, Bruno Tenório Ávila, Diego A. Salcedo, Rafael Ferreira Leite de Mello, Rinaldo Lima, Gabriel Pereira e Silva, Steven J. Simske
DocEng2
2018 Automatic cohesive summarization with pronominal anaphora resolution
Jamilson Antunes, Rafael Dueire Lins, Rinaldo Lima, Hilário Oliveira, Marcelo Riss, Steven J. Simske
Comput. Speech Lang.4
2017 A Regression-Based Approach Using Integer Linear Programming for Single-Document Summarization
abstract
Most of the existing approaches for extractive singledocument summarization of news articles rely on a single method to summarize all input documents. Recent work demonstrated that this is a significant limitation, since no summarization technique can achieve high performance for all input articles. In this context, this paper proposes a new regression-based approach using Integer Linear Programming (ILP) for single-document summarization. The proposed solution relies on a concept-based ILP method to generate multiple candidate summaries for each input article exploring different concept weighting methods and representation forms. Afterward, a regression model enriched with several extracted features at summary, sentence and ngram level is applied to select among the candidates the most informative summary based on an estimation of the traditional ROUGE-1 score. The investigated features are derived from indicators of content importance such as frequency, position, and coverage. Experiments conducted on the DUC 2001-2002 and CNN corpora show that the proposed method statistically outperforms other state-of-the-art extractive summarization approaches in most scenarios regarding ROUGE-1 and ROUGE-2 recall measures.
Hilário Oliveira, Rafael Dueire Lins, Rinaldo Lima, Fred Freitas
ICTAI1
2016 Appling Link Target Identification and Content Extraction to improve Web News Summarization
abstract
The existing automatic text summarization systems whenever applied to web-pages of news articles show poor performance as the text is encapsulated within a HTML page. This paper takes advantage of the link identification and content extraction techniques. The results show the validity of such a strategy.
Rodolfo Ferreira, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Hilário Oliveira, Marcelo Riss, Steven J. Simske
DocEng4
2016 Assessing Concept Weighting in Integer Linear Programming based Single-document Summarization
abstract
Some of the recent state-of-the-art systems for Automatic Text Summarization rely on the concept-based approach using Integer Linear Programming (ILP), mainly for multi-document summarization. A study on the suitability of such an approach to single-document summarization is still missing, however. This work presents an assessment of several methods of concept weighing for a concept-based ILP approach on the single-document summarization scenario. The unigram and bigram representations for concepts are also investigated. The experimental results obtained on the DUC 2001-2002 and the CNN corpora show that bigrams are more suitable than unigrams for the representation of concepts. Among the concept scoring methods investigated, the sentence position method presented the best performance on all evaluation corpora.
Hilário Oliveira, Rinaldo Lima, Rafael Dueire Lins, Fred Freitas, Marcelo Riss, Steven J. Simske
DocEng1
2016 Assessing shallow sentence scoring techniques and combinations for single and multi-document summarization
Hilário Oliveira, Rafael Ferreira Leite de Mello, Rinaldo Lima, Rafael Dueire Lins, Fred Freitas, Marcelo Riss, Steven J. Simske
Expert Syst. Appl.1
2015 Automatic Text Document Summarization Based on Machine Learning
abstract
The need for automatic generation of summaries gained importance with the unprecedented volume of information available in the Internet. Automatic systems based on extractive summarization techniques select the most significant sentences of one or more texts to generate a summary. This article makes use of Machine Learning techniques to assess the quality of the twenty most referenced strategies used in extractive summarization, integrating them in a tool. Quantitative and qualitative aspects were considered in such assessment demonstrating the validity of the proposed scheme. The experiments were performed on the CNN-corpus, possibly the largest and most suitable test corpus today for benchmarking extractive summarization strategies.
Gabriel Pereira e Silva, Rafael Ferreira Leite de Mello, Rafael Dueire Lins, Luciano de Souza Cabral, Hilário Oliveira, Steven J. Simske, Marcelo Riss
DocEng5
2013 An Inductive Logic Programming-Based Approach for Ontology Population from the Web
Rinaldo Lima, Bernard Espinasse, Hilário Oliveira, Rafael Ferreira Leite de Mello, Luciano de Souza Cabral, Dimas Filho, Fred Freitas, Renê Gadelha
DEXA (1)3
2013 Information Extraction from the Web: An Ontology-Based Method Using Inductive Logic Programming
abstract
Relevant information extraction from text and web pages in particular is an intensive and time-consuming task that needs important semantic resources. Thus, to be efficient, automatic information extraction systems have to exploit semantic resources (or ontologies) and employ machine-learning techniques to make them more adaptive. This paper presents an Ontology-based Information Extraction method using Inductive Logic Programming that allows inducing symbolic predicates expressed in Horn clausal logic that subsume information extraction rules. Such rules allow the system to extract class and relation instances from English corpora for ontology population purposes. Several experiments were conducted and preliminary experimental results are promising, showing that the proposed approach improves previous work over extracting instances of classes and relations, either separately or altogether.
Rinaldo Lima, Bernard Espinasse, Hilário Oliveira, Laura Pentagrossa, Fred Freitas
ICTAI3
2013 Group Profiling for Understanding Educational Social Networking
Ricardo B. C. Prudêncio, Luciano Meira, Alexandre Azevedo Filho, André C. A. Nascimento, Hilário Oliveira
SEKE6
2012 A Confidence-Weighted Metric for Unsupervised Ontology Population from Web Texts
Hilário Oliveira, Rinaldo Lima, Rafael Ferreira Leite de Mello, Fred Freitas, Evandro de Barros Costa
DEXA (1)1