Hugo Gonçalo Oliveira

dblp:88/1408 · DBLP profile ↗
← Back
11ranked-venue papers in the field
7as first author
6since 2021 · last 2025
0000-0002-5779-8645ORCID · verified

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 5 (3 first)Other / Interdisciplinary · 4 (3 first)Information Retrieval & Web Search · 2 (1 first)
YearPublicationVenuePosition
2025 Exploring Medium-Sized LLMs for Knowledge Base Construction
abstract
Knowledge base construction (KBC) is one of the great challenges in Natural Language Processing (NLP) and of fundamental importance to the growth of the Semantic Web. Large Language Models (LLMs) may be useful for extracting structured knowledge, including subject-predicate-object triples. We tackle the LM-KBC 2023 Challenge by leveraging LLMs for KBC, utilizing its dataset and benchmarking our results against challenge participants. Prompt engineering and ensemble strategies are tested for object prediction with pretrained LLMs in the 0.5-2B parameter range, which is between the limits of tracks 1 and 2 of the challenge.Selected models are assessed in zero-shot and few-shot learning approaches when predicting the objects of 21 relations. Results demonstrate that instruction-tuned LLMs outperform generative baselines by up to four times, with relation-adapted prompts playing a crucial role in performance. The ensemble approach further enhances triple extraction, with a relation-based selection strategy achieving the highest F1 score. These findings highlight the potential of medium-sized LLMs and prompt engineering methods for efficient KBC.
Tomás Cerveira Da Cruz Pinto, Hugo Gonçalo Oliveira, Chris-Bennet Fleger
LDK2
2023 Adopting Linguistic Linked Data Principles: Insights on Users' Experience
Verginica Barbu Mititelu, Maria Pia di Buono, Hugo Gonçalo Oliveira, Blerina Spahiu, Giedre Valunaite Oleskeviciene
LDK3
2023 GPT3 as a Portuguese Lexical Knowledge Base?
Hugo Gonçalo Oliveira, Ricardo Rodrigues 0001
LDK1
2023 SmartEDU: Accelerating Slide Deck Production with Natural Language Processing
Maria João Costa, Hugo Amaro, Hugo Gonçalo Oliveira
NLDB3
2023 On the Acquisition of WordNet Relations in Portuguese from Pretrained Masked Language Models
abstract
This paper studies the application of pretrained BERT in the acquisition of synonyms, antonyms, hypernyms and hyponyms in Portuguese.Masked patterns indicating those relations were compiled with the help of a service for validating semantic relations, and then used for prompting three pretrained BERT models, one multilingual and two for Portuguese (base and large).Predictions for the masks were evaluated in two different test sets.Results achieved by the monolingual models are interesting enough for considering these models as a source for enriching wordnets, especially when predicting hypernyms of nouns.Previously reported performances on prediction were improved with new patterns and with the large model.When it comes to selecting the related word from a set of four options, performance is even better, but not enough for outperforming the selection of the most similar word, as computed with static word embeddings.
Hugo Gonçalo Oliveira
GWC1
2021 On the Utility of Word Embeddings for Enriching OpenWordNet-PT
abstract
The maintenance of wordnets and lexical knwoledge bases typically relies on time-consuming manual effort. In order to minimise this issue, we propose the exploitation of models of distributional semantics, namely word embeddings learned from corpora, in the automatic identification of relation instances missing in a wordnet. Analogy-solving methods are first used for learning a set of relations from analogy tests focused on each relation. Despite their low accuracy, we noted that a portion of the top-given answers are good suggestions of relation instances that could be included in the wordnet. This procedure is applied to the enrichment of OpenWordNet-PT, a public Portuguese wordnet. Relations are learned from data acquired from this resource, and illustrative examples are provided. Results are promising for accelerating the identification of missing relation instances, as we estimate that about 17% of the potential suggestions are good, a proportion that almost doubles if some are automatically invalidated.
Hugo Gonçalo Oliveira, Fredson Silva de Souza Aguiar, Alexandre Rademaker
LDK1
2019 Fast developing of a Natural Language Interface for a Portuguese WordNet: Leveraging on Sentence Embeddings
abstract
We describe how a natural language interface can be developed for a wordnet with a small set of handcrafted templates, leveraging on sentence embeddings.The proposed approach does not use rules for parsing natural language queries but experiments showed that the embeddings model is tolerant enough for correctly predicting relation types that do not match known patterns exactly.It was tested with OpenWordNet-PT, for which this method may provide an alternative interface, with benefits also on the curation process.
Hugo Gonçalo Oliveira, Alexandre Rademaker
GWC1
2016 An overview of Portuguese WordNets
abstract
Semantic relations between words are key to building systems that aim to understand and manipulate language.For English, the "de facto" standard for representing this kind of knowledge is Princeton's WordNet.Here, we describe the wordnet-like resources currently available for Portuguese: their origins, methods of creation, sizes, and usage restrictions.We start tackling the problem of comparing them, but only in quantitative terms.Finally, we sketch ideas for potential collaboration between some of the projects that produce Portuguese wordnets.
Valeria de Paiva, Livy Real, Hugo Gonçalo Oliveira, Alexandre Rademaker, Cláudia Freitas, Alberto Simões 0001
GWC3
2015 Automatic Generation of Poetry Inspired by Twitter Trends
Hugo Gonçalo Oliveira
IC3K1
2014 Onto.PT: recent developments of a large public domain Portuguese wordnet
abstract
This document describes the current state of Onto.PT, a new large wordnet for Portuguese, freely available, and created automatically after exploiting and integrating existing lexical resources in a wordnet structure.Besides an overview on Onto.PT, its creation and evaluation, we enumerate the developments of version 0.6.Moreover, we provide a quantitative view on this version, its comparison to other Portuguese wordnets, in terms of contents and size, as well as some details about its global coverage and availability.
Hugo Gonçalo Oliveira, Paulo Gomes
GWC1
2012 Integrating Lexical-Semantic Knowledge to Build a Public Lexical Ontology for Portuguese
Hugo Gonçalo Oliveira, Leticia Antón Pérez, Paulo Gomes
NLDB1