Esteban Rodríguez Betancourt

dblp:264/9486 · DBLP profile ↗
← Back
3ranked-venue papers in the field
3as first author
2since 2021 · last 2023
0009-0005-9002-6582ORCID · reported

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 3 (3 first)
YearPublicationVenuePosition
2023 Exploring the Limits of Large Language Models for Word Definition Generation: A Comparative Analysis
abstract
In this paper, we explore the ability of large language models (LLMs) to generate word definitions for newly invented words in Spanish through the task of Unknown Definition Modeling. The main goal of our study is to determine the extent to which LLMs can abstract meaning from context and compare the performance of different models for this task. To conduct our analysis, we created a dataset of 20 made-up words, usage examples, and their definitions in Spanish. We then evaluated several LLMs, including OpenAI GPT-3.5-turbo, OpenAI GPT-3, and Google Flan-T5, using automatic evaluation based on cosine similarity of sentence embeddings and qualitative human evaluation on a 4-point Likert scale. Our findings indicate that larger models tend to generate better definitions than smaller models, with the performance of the models generally aligning with their size. This study contributes to our understanding of LLMs' strengths and weaknesses in generating definitions for unknown words, and offers valuable insights for future research and applications in natural language processing.
Esteban Rodríguez Betancourt, Edgar Casasola Murillo
CLEI1
2022 Analysis of Semantic Shift Before and After COVID-19 in Spanish Diachronic Word Embeddings
abstract
Words can shift their meaning across time. This case study shows the results obtained by the exploratory analysis of the semantic shifting on Spanish vocabulary using Diachronic Words Embeddings. Diachronic data consists of a 2018 Spanish corpus, before the COVID-19 outbreak, and a second corpus with documents from 2021. We focused on the semantic shift of three of the topics: COVID-19, masks and vaccines. This paper addresses the construction of the diachronic Spanish word embeddings model, as well as the results obtained by the analysis using a non-supervised distance vector technique. The results allowed to identify shifts related to increase in COVID-19 content.
Esteban Rodríguez Betancourt, Edgar Casasola Murillo
CLEI1
2019 Deep Neural Network Comparison for Spanish Tweets Polarity Classification
abstract
Two deep neural network models were compared in the task of polarity classification in Spanish text retrieved from social networks. For each model accuracy, precision, recall and F1 was calculated over a particular corpus. Also, the effect of adding gaussian noise on the inputs over the classifier results was evaluated.
Esteban Rodríguez Betancourt, Pablo Sauma Chacón, Edgar Casasola Murillo
CLEI1