Hugo Gonçalo Oliveira

dblp:88/1408 · DBLP profile ↗
← Back
58ranked-venue papers
27as first author
27since 2021 · last 2026
0000-0002-5779-8645ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 41 · 20 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 8 first-author · 10 since 2021Databases, data management, data science and information retrieval · 11 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 CorEGe-PT: Compiling a Large Corpus of Academic Texts in Portuguese
abstract
This paper describes the creation of a large-scale corpus of academic texts in Portuguese, dubbed CorEGe-PT, extracted from the institutional repository of a Portuguese university. Its compilation methodology, which combined automatic and manual procedures, is detailed, together with challenges faced and proposed solutions. The process included a thorough analysis of the metadata, which will be publicly released together with the documents, extracted in a markdown format. CorEGe-PT covers five areas of knowledge and, with over 34,000 documents and 1B tokens, is the largest of corpus of its kind in Portuguese, which will enable in-depth linguistic studies while providing data for adapting Large Language Models to academic Portuguese and related tasks.
Tanara Zingano Kuhn, José Matos 0004, Bruno Neves, Daniela Pereira, Elisabete Cação, Ivo Simões, Jacinto Estima, Delfim Leão, Hugo Gonçalo Oliveira
LREC9
2026 RelEx-PT: A Portuguese Sentence-Level Relation Extraction Dataset
Tomás Pinto, Catarina Silva 0001, Hugo Gonçalo Oliveira
LREC3
2026 Reasoning or not? A comprehensive evaluation of reasoning LLMs for dialogue summarization
Keyan Jin, Yapeng Wang 0001, Leonel Santos, Xu Yang 0010, Sio Kei Im, Hugo Gonçalo Oliveira
Expert Syst. Appl.7
2026 Multi-task specialized expert model for hierarchical aspect-based sentiment analysis in consumer healthcare
Jiaxuan Li 0003, Jielong Guo, Patrick Pang 0001, Hugo Gonçalo Oliveira, Benjamin K. Ng, Tao Tan 0002
Expert Syst. Appl.4
2025 Refining Metrical Constraints in LLM-Generated Poetry with Feedback
Manex Agirrezabal, Hugo Gonçalo Oliveira
ICCC2
2025 A Full Pipeline for Context-Aware Pun Generation
Marcio Lima Inácio, Hugo Gonçalo Oliveira
ICCC2
2025 Cognitive Flow: An LLM-Automated Framework for Quantifying Reasoning Distillation
abstract
The ability of large language models (LLMs) to reason effectively is crucial for a wide range of applications, from complex decision-making to scientific research. However, it remains unclear how well reasoning capabilities are transferred or preserved when LLMs undergo Knowledge Distillation (KD), a process that typically reduces model size while attempting to retain performance. In this study, we explore the effects of model distillation on the reasoning abilities of various reasoning language models (RLMs). We introduce Cognitive Flow, a novel framework that systematically extracts meaning and map states in Chain-of-Thought (CoT) processes, offering new insights on model reasoning and enabling quantitative comparisons across RLMs. Using this framework, we investigate the impact of KD on CoTs produced by RLMs. We target DeepSeek-R1-671B and its distilled 70B, 32B and 14B versions, as well as QwenQwQ-32B from the Qwen series. We evaluate the models on three subsets of mathematical reasoning tasks with varying complexity from the MMLU benchmark. Our findings demonstrate that while distillation can effectively replicate a similar reasoning style under specific conditions, it struggles with simpler problems, revealing a significant divergence in the observable thought process and a potential limitation in the transfer of a robust and adaptable problem-solving capability.
José Matos 0004, Catarina Silva 0001, Hugo Gonçalo Oliveira
INLG3
2025 Exploring Medium-Sized LLMs for Knowledge Base Construction
abstract
Knowledge base construction (KBC) is one of the great challenges in Natural Language Processing (NLP) and of fundamental importance to the growth of the Semantic Web. Large Language Models (LLMs) may be useful for extracting structured knowledge, including subject-predicate-object triples. We tackle the LM-KBC 2023 Challenge by leveraging LLMs for KBC, utilizing its dataset and benchmarking our results against challenge participants. Prompt engineering and ensemble strategies are tested for object prediction with pretrained LLMs in the 0.5-2B parameter range, which is between the limits of tracks 1 and 2 of the challenge.Selected models are assessed in zero-shot and few-shot learning approaches when predicting the objects of 21 relations. Results demonstrate that instruction-tuned LLMs outperform generative baselines by up to four times, with relation-adapted prompts playing a crucial role in performance. The ensemble approach further enhances triple extraction, with a relation-based selection strategy achieving the highest F1 score. These findings highlight the potential of medium-sized LLMs and prompt engineering methods for efficient KBC.
Tomás Cerveira Da Cruz Pinto, Hugo Gonçalo Oliveira, Chris-Bennet Fleger
LDK2
2025 Twitter and Sentiment Analysis for Wildfire Heat Mapping
abstract
ABSTRACT Nowadays, automated intelligent systems play an increasingly vital role in aiding decision‐making processes across various fields. Firefighting represents a crucial area where accurate information gathering is paramount for efficient resource allocation. Social media platforms as Twitter (or X) have emerged as valuable sources of real‐time data, often referred to as ‘citizen science’, offering additional insights alongside traditional data sources. In this work, we introduce a novel pipeline that leverages Natural Language Processing (NLP) techniques and Twitter data, utilising transformer models to identify and monitor wildfire incidents. Expanding on this approach, we incorporate sentiment analysis to provide deeper insights into public perceptions and emotions related to fire events. Additionally, we present visual representations of geographic data through heat mapping, potentially aiding firefighters in making informed decisions. By integrating advanced NLP techniques with social media data, our approach presents a promising strategy for enhancing wildfire management efforts.
Catarina Silva 0001, Isabel Carvalho, João Cabral Pinto, Alberto Cardoso, Hugo Gonçalo Oliveira
Expert Syst. J. Knowl. Eng.6
2024 MultiLexBATS: Multilingual Dataset of Lexical Semantic Relations
abstract
Understanding the relation between the meanings of words is an important part of comprehending natural language. Prior work has either focused on analysing lexical semantic relations in word embeddings or probing pretrained language models (PLMs), with some exceptions. Given the rarity of highly multilingual benchmarks, it is unclear to what extent PLMs capture relational knowledge and are able to transfer it across languages. To start addressing this question, we propose MultiLexBATS, a multilingual parallel dataset of lexical semantic relations adapted from BATS in 15 languages including low-resource languages, such as Bambara, Lithuanian, and Albanian. As experiment on cross-lingual transfer of relational knowledge, we test the PLMs’ ability to (1) capture analogies across languages, and (2) predict translation targets. We find considerable differences across relation types and languages with a clear preference for hypernymy and antonymy as well as romance languages.
Dagmar Gromann, Hugo Gonçalo Oliveira, Lucia Pitarch, Elena Apostol, Jordi Bernad, Eliot Bytyci, Chiara Cantone, Sara Carvalho, Francesca Frontini, Radovan Garabík, Jorge Gracia, Letizia Granata, Anas Fahad Khan, Timotej Knez, Penny Labropoulou, Chaya Liebeskind, Maria Pia di Buono, Ana Ostroski Anic, Sigita Rackeviciene, Ricardo Rodrigues 0001, Gilles Sérasset, Linas Selmistraitis, Mahammadou Sidibé, Purificação Silvano, Blerina Spahiu, Enriketa Sogutlu, Ranka Stankovic, Ciprian-Octavian Truica, Giedre Valunaite Oleskeviciene, Slavko Zitnik, Katerina Zdravkova
LREC/COLING2
2024 Puntuguese: A Corpus of Puns in Portuguese with Micro-edits
abstract
Humor is an intricate part of verbal communication and dealing with this kind of phenomenon is essential to building systems that can process language at large with all of its complexities. In this paper, we introduce Puntuguese, a new corpus of punning humor in Portuguese, motivated by previous works showing that currently available corpora for this language are still unfit for Machine Learning due to data leakage. Puntuguese comprises 4,903 manually-gathered punning one-liners in Brazilian and European Portuguese. To create negative examples that differ exclusively in terms of funniness, we carried out a micro-editing process, in which all jokes were edited by fluent Portuguese speakers to make the texts unfunny. Finally, we did some experiments on Humor Recognition, showing that Puntuguese is considerably more difficult than the previous corpus, achieving an F1-Score of 68.9%. With this new dataset, we hope to enable research not only in NLP but also in other fields that are interested in studying humor; thus, the data is publicly available.
Marcio Lima Inácio, Gabriela Wick-Pedro, Renata Ramisch, Luís Espírito Santo, Xiomara S. Q. Chacón, Roney Lira de Sales Santos, Rogério F. de Sousa, Rafael T. Anchiêta, Hugo Gonçalo Oliveira
LREC/COLING9
2024 Zero-Shot Metrical Poetry Generation with Open Language Models: a Quantitative Analysis
Manex Agirrezabal, Hugo Gonçalo Oliveira
ICCC2
2024 Generation of Punning Riddles in Portuguese with Prompt Chaining
Marcio Lima Inácio, Hugo Gonçalo Oliveira
ICCC2
2024 Sentiment-Aware Dialogue Flow Discovery for Interpreting Communication Trends
abstract
Customer-support services increasingly rely on automation, whether full or with human intervention.Despite optimising resources, this may result in mechanical protocols and lack of human interaction, thus reducing customer loyalty.Our goal is to enhance interpretability and provide guidance in communication through novel tools for easier analysis of message trends and sentiment variations.Monitoring these contributes to more informed decision-making, enabling proactive mitigation of potential issues, such as protocol deviations or customer dissatisfaction.We propose a generic approach for dialogue flow discovery that leverages clustering techniques to identify dialogue states, represented by related utterances.State transitions are further analyzed to detect prevailing sentiments.Hence, we discover sentimentaware dialogue flows that offer an interpretability layer to artificial agents, even those based on black-boxes, ultimately increasing trustworthiness.Experimental results demonstrate the effectiveness of our approach across different dialogue datasets, covering both human-human and human-machine exchanges, applicable in task-oriented contexts but also to social media, highlighting its potential impact across various customer-support settings.
Patrícia Sofia Pereira Ferreira, Isabel Carvalho, Ana Alves 0001, Catarina Silva 0001, Hugo Gonçalo Oliveira
SIGDIAL5
2023 Leveraging Question Answering for Domain-Agnostic Information Extraction
Bruno Carlos Luís Ferreira, Hugo Gonçalo Oliveira, Catarina Silva 0001
CIARP2
2023 Evaluating the Extraction of Toxicological Properties with Extractive Question Answering
Bruno Carlos Luís Ferreira, Hugo Gonçalo Oliveira, Hugo Amaro, Ângela Laranjeiro, Catarina Silva 0001
EANN2
2023 Unsupervised Flow Discovery from Task-Oriented Dialogues
Patrícia Sofia Pereira Ferreira, Daniel Martins, Ana Alves 0001, Catarina Silva 0001, Hugo Gonçalo Oliveira
HIS (4)5
2023 Generating Wildfire Heat Maps with Twitter and BERT
João Cabral Pinto, Hugo Gonçalo Oliveira, Alberto Cardoso, Catarina Silva 0001
IDEAL2
2023 Adopting Linguistic Linked Data Principles: Insights on Users' Experience
Verginica Barbu Mititelu, Maria Pia di Buono, Hugo Gonçalo Oliveira, Blerina Spahiu, Giedre Valunaite Oleskeviciene
LDK3
2023 GPT3 as a Portuguese Lexical Knowledge Base?
Hugo Gonçalo Oliveira, Ricardo Rodrigues 0001
LDK1
2023 SmartEDU: Accelerating Slide Deck Production with Natural Language Processing
Maria João Costa, Hugo Amaro, Hugo Gonçalo Oliveira
NLDB3
2023 On the Acquisition of WordNet Relations in Portuguese from Pretrained Masked Language Models
abstract
This paper studies the application of pretrained BERT in the acquisition of synonyms, antonyms, hypernyms and hyponyms in Portuguese.Masked patterns indicating those relations were compiled with the help of a service for validating semantic relations, and then used for prompting three pretrained BERT models, one multilingual and two for Portuguese (base and large).Predictions for the masks were evaluated in two different test sets.Results achieved by the monolingual models are interesting enough for considering these models as a source for enriching wordnets, especially when predicting hypernyms of nouns.Previously reported performances on prediction were improved with new patterns and with the large model.When it comes to selecting the related word from a set of four options, performance is even better, but not enough for outperforming the selection of the most similar word, as computed with static word embeddings.
Hugo Gonçalo Oliveira
GWC1
2022 Exploring Transformers for Ranking Portuguese Semantic Relations
abstract
We explored transformer-based language models for ranking instances of Portuguese lexico-semantic relations. Weights were based on the likelihood of natural language sequences that transmitted the relation instances, and expectations were that they would be useful for filtering out noisier instances. However, after analysing the weights, no strong conclusions were taken. They are not correlated with redundancy, but are lower for instances with longer and more specific arguments, which may nevertheless be a consequence of their sensitivity to the frequency of such arguments. They did also not reveal to be useful when computing word similarity with network embeddings. Despite the negative results, we see the reported experiments and insights as another contribution for better understanding transformer language models like BERT and GPT, and we make the weighted instances publicly available for further research.
Hugo Gonçalo Oliveira
LREC1
2022 A Brief Survey of Textual Dialogue Corpora
abstract
Several dialogue corpora are currently available for research purposes, but they still fall short for the growing interest in the development of dialogue systems with their own specific requirements. In order to help those requiring such a corpus, this paper surveys a range of available options, in terms of aspects like speakers, size, languages, collection, annotations, and domains. Some trends are identified and possible approaches for the creation of new corpora are also discussed.
Hugo Gonçalo Oliveira, Patrícia Sofia Pereira Ferreira, Daniel Martins, Catarina Silva 0001, Ana Alves 0001
LREC1
2022 A survey of the extraction and applications of causal relations
abstract
Abstract Causationin written natural language can express a strong relationship between events and facts. Causation in the written form can be referred to as a causal relation where a cause event entails the occurrence of an effect event. A cause and effect relationship is stronger than a correlation between events, and therefore aggregated causal relations extracted from large corpora can be used in numerous applications such as question-answering and summarisation to produce superior results than traditional approaches. Techniques like logical consequence allow causal relations to be used in niche practical applications such as event prediction which is useful for diverse domains such as security and finance. Until recently, the use of causal relations was a relatively unpopular technique because the causal relation extraction techniques were problematic, and the relations returned were incomplete, error prone or simplistic. The recent adoption of language models and improved relation extractors for natural language such as Transformer-XL (Daiet al. (2019).Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860 ) has seen a surge of research interest in the possibilities of using causal relations in practical applications. Until now, there has not been an extensive survey of the practical applications of causal relations; therefore, this survey is intended precisely to demonstrate the potential of causal relations. It is a comprehensive survey of the work on the extraction of causal relations and their applications, while also discussing the nature of causation and its representation in text.
Brett Drury, Hugo Gonçalo Oliveira, Alneu de Andrade Lopes
Nat. Lang. Eng.2
2021 Exploring a Masked Language Model for Creative Text Transformation
Hugo Gonçalo Oliveira
ICCC1
2021 On the Utility of Word Embeddings for Enriching OpenWordNet-PT
abstract
The maintenance of wordnets and lexical knwoledge bases typically relies on time-consuming manual effort. In order to minimise this issue, we propose the exploitation of models of distributional semantics, namely word embeddings learned from corpora, in the automatic identification of relation instances missing in a wordnet. Analogy-solving methods are first used for learning a set of relations from analogy tests focused on each relation. Despite their low accuracy, we noted that a portion of the top-given answers are good suggestions of relation instances that could be included in the wordnet. This procedure is applied to the enrichment of OpenWordNet-PT, a public Portuguese wordnet. Relations are learned from data acquired from this resource, and illustrative examples are provided. Results are promising for accelerating the identification of missing relation instances, as we estimate that about 17% of the potential suggestions are good, a proportion that almost doubles if some are automatically invalidated.
Hugo Gonçalo Oliveira, Fredson Silva de Souza Aguiar, Alexandre Rademaker
LDK1
2020 Comparing Different Methods for Assigning Portuguese Proverbs to News Headlines
Rui Mendes 0004, Hugo Gonçalo Oliveira
ICCC2
2020 TECo: Exploring Word Embeddings for Text Adaptation to a given Context
Rui Mendes 0004, Hugo Gonçalo Oliveira
ICCC2
2020 WeirdAnalogyMatic: Experimenting with Analogy for Lyrics Transformation
Hugo Gonçalo Oliveira
ICCC1
2020 Amplifying the Range of News Stories with Creativity: Methods and their Evaluation, in Portuguese
abstract
Headlines are key for attracting people to a story, but writing appealing headlines requires time and talent.This work aims to automate the production of creative short texts (e.g., news headlines) for an input context (e.g., existing headlines), thus amplifying its range.Well-known expressions (e.g., proverbs, movie titles), which typically include word-play and resort to figurative language, are used as a starting point.Given an input text, they can be recommended by exploiting Semantic Textual Similarity (STS) techniques, or adapted towards higher relatedness.For the latter, three methods that exploit static word embeddings are proposed.Experimentation in Portuguese lead to some conclusions, based on human opinions: STS methods that look exclusively at the surface text, recommend more related expressions; resulting expressions are somewhat related to the input, but adaptation leads to higher relatedness and novelty; humour can be an indirect consequence, but most outputs are not funny.
Rui Mendes 0004, Hugo Gonçalo Oliveira
INLG2
2020 Corpora and Baselines for Humour Recognition in Portuguese
abstract
Having in mind the lack of work on the automatic recognition of verbal humour in Portuguese, a topic connected with fluency in a natural language, we describe the creation of three corpora, covering two styles of humour and four sources of non-humorous text, that may be used for related studies. We then report on some experiments where the created corpora were used for training and testing computational models that exploit content and linguistic features for humour recognition. The obtained results helped us taking some conclusions about this challenge and may be seen as baselines for those willing to tackle it in the future, using the same corpora.
Hugo Gonçalo Oliveira, André Clemêncio, Ana Alves 0001
LREC1
2020 AIA-BDE: A Corpus of FAQs in Portuguese and their Variations
abstract
We present AIA-BDE, a corpus of 380 domain-oriented FAQs in Portuguese and their variations, i.e., paraphrases or entailed questions, created manually, by humans, or automatically, with Google Translate. Its aims to be used as a benchmark for FAQ retrieval and automatic question-answering, but may be useful in other contexts, such as the development of task-oriented dialogue systems, or models for natural language inference in an interrogative context. We also report on two experiments. Matching variations with their original questions was not trivial with a set of unsupervised baselines, especially for manually created variations. Besides high performances obtained with ELMo and BERT embeddings, an Information Retrieval system was surprisingly competitive when considering only the first hit. In the second experiment, text classifiers were trained with the original questions, and tested when assigning each variation to one of three possible sources, or assigning them as out-of-domain. Here, the difference between manual and automatic variations was not so significant.
Hugo Gonçalo Oliveira, João Ferreira 0004, Pedro Fialho, Ricardo Rodrigues 0001, Luísa Coheur, Ana Alves 0001
LREC1
2019 Fast developing of a Natural Language Interface for a Portuguese WordNet: Leveraging on Sentence Embeddings
abstract
We describe how a natural language interface can be developed for a wordnet with a small set of handcrafted templates, leveraging on sentence embeddings.The proposed approach does not use rules for parsing natural language queries but experiments showed that the embeddings model is tolerant enough for correctly predicting relation types that do not match known patterns exactly.It was tested with OpenWordNet-PT, for which this method may provide an alternative interface, with benefits also on the curation process.
Hugo Gonçalo Oliveira, Alexandre Rademaker
GWC1
2018 Integrating a ETHNO-MUSIC and a Tra-la-Lyrics for Composing Popular Spanish Songs
María Navarro 0001, Hugo Gonçalo Oliveira
ICCC2
2018 A Set of Procedures Attempting at the Generation of Verbal Humor in Portuguese
Hugo Gonçalo Oliveira, Ricardo Rodrigues 0001
ICCC1
2017 A Survey on Intelligent Poetry Generation: Languages, Features, Techniques, Reutilisation and Evaluation
abstract
Poetry generation is becoming popular among researchers of Natural Language Generation, Computational Creativity and, broadly, Artificial Intelligence.To produce text that may be regarded as poetry, computational systems are typically knowledge-intensive and deal with several levels of language.Interest on the topic resulted in the development of several poetry generators described in the literature, with different features covered or handled differently, by a broad range of alternative approaches, as well as different perspectives on evaluation, another challenging aspect due the underlying subjectivity.This paper surveys intelligent poetry generators around a set of relevant axis -target language, form and content features, applied techniques, reutilisation of material, and evaluation -and aims to organise work developed on this topic so far.
Hugo Gonçalo Oliveira
INLG1
2017 Co-PoeTryMe: a Co-Creative Interface for the Composition of Poetry
abstract
Co-PoeTryMe is a web application for poetry composition, guided by the user, though with the help of automatic features, such as the generation of full (editable) drafts, as well as the acquisition of additional well-formed lines, or semantically-related words, possibly constrained by the number of syllables, rhyme, or polarity.Towards the final poem, the latter can replace lines or words in the draft.
Hugo Gonçalo Oliveira, Tiago Mendes, Ana Boavida
INLG1
2017 Multilingual extension and evaluation of a poetry generator
abstract
Abstract Poetry generation is a specific kind of natural language generation where several sources of knowledge are typically exploited to handle features on different levels, such as syntax, semantics, form or aesthetics. But although this task has been addressed by several researchers, and targeted different languages, all known systems have focused on a limited purpose and a single language. This article describes the effort of adapting the same architecture to generate poetry in three different languages – Portuguese, Spanish and English. An existing architecture is first described and complemented with the adaptations required for each language, including the linguistic resources used for handling morphology, syntax, semantics and metric scansion. An automatic evaluation was designed in such a way that it would be applicable to the target languages. It covered three relevant aspects of the generated poems, namely: the presence of poetic features, the variation of the linguistic structure and the semantic connection to a given topic. The automatic measures applied for the second and third aspect can be seen as novel in the evaluation of poetry. Overall, poems were successfully generated in the three languages addressed. Despite minor differences in different languages or seed words, poems revealed to have a regular metre, frequent rhymes, to exhibit an interesting degree of variation, and to be semantically-associated with the initially given seeds.
Hugo Gonçalo Oliveira, Raquel Hervás, Alberto Díaz 0001, Pablo Gervás
Nat. Lang. Eng.1
2016 Poetry from Concept Maps-Yet Another Adaptation of PoeTryMe's Flexible Architecture
Hugo Gonçalo Oliveira, Ana Alves 0001
ICCC1
2016 One does not simply produce funny memes! - Explorations on the Automatic Generation of Internet humor
Hugo Gonçalo Oliveira, Alexandre Miguel Pinto
ICCC1
2016 Computational Creativity Infrastructure for Online Software Composition: A Conceptual Blending Use Case
Martin Znidarsic, Amílcar Cardoso, Pablo Gervás, Pedro Martins 0003, Raquel Hervás, Ana Alves 0001, Hugo Gonçalo Oliveira, Ping Xiao, Simo Linkola, Hannu Toivonen, Janez Kranjc, Nada Lavrac
ICCC7
2016 Can Topic Modelling benefit from Word Sense Information?
Adriana Ferrugento, Hugo Gonçalo Oliveira, Ana Alves 0001, Filipe Rodrigues 0001
LREC2
2016 Discovering Fuzzy Synsets from the Redundancy in Different Lexical-Semantic Resources
Hugo Gonçalo Oliveira, Fábio Santos
LREC1
2016 TweetMT: A Parallel Microblog Corpus
Iñaki San Vicente, Iñaki Alegria, Cristina España-Bonet, Pablo Gamallo 0001, Hugo Gonçalo Oliveira, Eva Martínez Garcia, Antonio Toral, Arkaitz Zubiaga, Nora Aranberri
LREC5
2016 An overview of Portuguese WordNets
abstract
Semantic relations between words are key to building systems that aim to understand and manipulate language.For English, the "de facto" standard for representing this kind of knowledge is Princeton's WordNet.Here, we describe the wordnet-like resources currently available for Portuguese: their origins, methods of creation, sizes, and usage restrictions.We start tackling the problem of comparing them, but only in quantitative terms.Finally, we sketch ideas for potential collaboration between some of the projects that produce Portuguese wordnets.
Valeria de Paiva, Livy Real, Hugo Gonçalo Oliveira, Alexandre Rademaker, Cláudia Freitas, Alberto Simões 0001
GWC3
2016 Revisiting the ontologising of semantic relation arguments in wordnet synsets
abstract
Abstract Ontologising is the task of associating terms, in text, with an ontological representation of their meaning, in an ontology. In this article, we revisit algorithms that have previously been used to ontologise the arguments of semantic relations in a relationless thesaurus, resulting in a wordnet. For increased flexibility, the algorithms do not use the extraction context when selecting the most adequate synsets for each term argument. Instead, they exploit a term-based lexical network which can be established by knowledge extracted automatically, or obtained from the resource the relations are being ontologised to. On the latter idea, we made several experiments to conclude that the algorithms can be used both for wordnet creation and for their enrichment. Besides describing the algorithms with some detail, the aforementioned experiments, which target both English and Portuguese, and their results are reported and discussed.
Hugo Gonçalo Oliveira, Paulo Gomes
Nat. Lang. Eng.1
2015 Automatic Generation of Poetry Inspired by Twitter Trends
Hugo Gonçalo Oliveira
IC3K1
2015 In reality there are as many religions as there are papers - First Steps Towards the Generation of Internet Memes
Hugo Gonçalo Oliveira, Alexandre Miguel Pinto
ICCC2
2014 Adapting a Generic Platform for Poetry Generation to Produce Spanish Poems
Hugo Gonçalo Oliveira, Raquel Hervás, Alberto Díaz 0001, Pablo Gervás
ICCC1
2014 Exploiting Portuguese Lexical Knowledge Bases for Answering Open Domain Cloze Questions Automatically
Hugo Gonçalo Oliveira, Inês Coelho, Paulo Gomes
LREC1
2014 Onto.PT: recent developments of a large public domain Portuguese wordnet
abstract
This document describes the current state of Onto.PT, a new large wordnet for Portuguese, freely available, and created automatically after exploiting and integrating existing lexical resources in a wordnet structure.Besides an overview on Onto.PT, its creation and evaluation, we enumerate the developments of version 0.6.Moreover, we provide a quantitative view on this version, its comparison to other Portuguese wordnets, in terms of contents and size, as well as some details about its global coverage and availability.
Hugo Gonçalo Oliveira, Paulo Gomes
GWC1
2013 Towards the automatic enrichment of a thesaurus with information in dictionaries
abstract
Abstract Regarding that information in broad‐coverage knowledge bases, such as thesauri, is usually incomplete, merging information from different sources is an alternative to amplify coverage. We propose a method for the enrichment of a thesaurus with information acquired automatically from dictionaries. First, synonymy pairs are extracted. Then, these pairs are assigned to the most similar candidate synsets. Finally, the remaining pairs are the target of clustering to identify new synsets. After selecting the adequate experimentation settings, this method was applied to enrich a Portuguese thesaurus with synonyms extracted from three dictionaries, which resulted in TRIP, a larger and broader thesaurus with new words and concepts. The steps towards the creation of this new thesaurus and its evaluation are described here.
Hugo Gonçalo Oliveira, Paulo Gomes
Expert Syst. J. Knowl. Eng.1
2012 Folheador: browsing through Portuguese semantic relations
Hugo Gonçalo Oliveira, Hernani Pereira Costa, Diana Santos
EACL1
2012 Integrating Lexical-Semantic Knowledge to Build a Public Lexical Ontology for Portuguese
Hugo Gonçalo Oliveira, Leticia Antón Pérez, Paulo Gomes
NLDB1
2011 Automatic Discovery of Fuzzy Synsets from Dictionary Definitions
abstract
In order to deal with ambiguity in natural language, it is common to organise words, according to their senses, in synsets, which are groups of synonymous words that can be seen as concepts. The manual creation of a broad-coverage synset base is a time-consuming task, so we take advantage of dictionary definitions for extracting synonymy pairs and clustering for identifying synsets. Since word senses are not discrete, we create fuzzy synsets, where each word has a membership degree. We report on the results of the creation of a fuzzy synset base for Portuguese, from three electronic dictionaries. The resulting resource is larger than existing hancrafted Portuguese thesauri.
Hugo Gonçalo Oliveira, Paulo Gomes
IJCAI1
2010 Automatic Creation of a Conceptual Base for Portuguese using Clustering Techniques
abstract
When a semantic network is based on triples relating terms, ambiguity arises as a problem, so we have used these triples to identify clusters, which can be seen as synsets. We report the results of this approach on a synonymy network extracted from a dictionary and additional tests involving manually created thesaurus. Part of the resulting synsets were also evaluated by human subjects.
Hugo Gonçalo Oliveira, Paulo Gomes
ECAI1
2010 Second HAREM: Advancing the State of the Art of Named Entity Recognition in Portuguese
Cláudia Freitas, Cristina Mota, Diana Santos, Hugo Gonçalo Oliveira, Paula Carvalho 0001
LREC4