Felipe Bravo-Marquez

dblp:04/8612 · also Felipe José Bravo Márquez · DBLP profile ↗
← Back
30ranked-venue papers
14as first author
12since 2021 · last 2025
0000-0002-2153-4306ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 11 first-author · 8 since 2021Databases, data management, data science and information retrieval · 10 · 6 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Can Large Language Models Compete with Specialized Models in Lexical Semantic Change Detection?
abstract
In this paper, we present a comprehensive comparison between specialized Lexical Semantic Change Detection (LSCD) models and Large Language Models (LLMs) for the LSCD task. In addition to comparing models, we also investigate the role of automatic prompt selection for improving LLM performance. We evaluate three approaches: Average Pairwise Distance (APD), Word-in-Context (WiC), and Word Sense Induction (WSI). Using Spearman correlation as the evaluation metric, we assess the performance of Mixtral, Llama 3.1, Llama 3.3, and specialized LSCD models across English and Spanish datasets. Our results show that by using prompt optimization and LLMs, we achieve state-of-the-art performance for the English dataset and outperform specialized LSCD models at the annotation level in the same dataset. For Spanish, specialized models outperform LLMs across all three approaches—WiC, APD, and WSI—indicating that specialized LSCD models are still more effective for semantic change detection in Spanish.
Frank D. Zamora-Reina, Felipe Bravo-Marquez, Dominik Schlechtweg, Nikolay Arefyev
ECAI2
2025 Adapting Bias Evaluation to Domain Contexts using Generative Models
abstract
Numerous datasets have been proposed to evaluate social bias in Natural Language Processing (NLP) systems.However, assessing bias within specific application domains remains challenging, as existing approaches often face limitations in scalability and fidelity across domains.In this work, we introduce a domain-adaptive framework that utilizes prompting with Large Language Models (LLMs) to automatically transform template-based bias datasets into domain-specific variants.We apply our method to two widely used benchmarks-Equity Evaluation Corpus (EEC) and Identity Phrase Templates Test Set (IPTTS)-adapting them to the Twitter and Wikipedia Talk data.Our results show that the adapted datasets yield bias estimates more closely aligned with real-world data.These findings highlight the potential of LLM-based prompting to enhance the realism and contextual relevance of bias evaluation in NLP systems. 1
Tamara Quiroga, Felipe Bravo-Marquez, Valentin Barriere
EMNLP2
2025 WEFE: A Python Library for Measuring and Mitigating Bias in Word Embeddings
abstract
Word embeddings, which are a mapping of words into continuous vectors, are widely used in modern Natural Language Processing (NLP) systems. However, they are prone to inherit stereotypical social biases from the corpus on which they are built. The research community has focused on two main tasks to address this problem: 1) how to measure these biases, and 2) how to mitigate them. Word Embedding Fairness Evaluation (WEFE) is an open source library that implements many fairness metrics and mitigation methods in a unified framework. It also provides a standard interface for designing new ones. The software follows the object-oriented paradigm with a strong focus on extensibility. Each of its methods is appropriately documented, verified and tested. WEFE is not limited to just a library: it also contains several replications of previous studies as well as tutorials that serve as educational material for newcomers to the field. It is licensed under BSD-3 and can be easily installed through pip and conda package managers.
Pablo Badilla, Felipe Bravo-Marquez, María José Zambrano, Jorge Pérez 0001
J. Mach. Learn. Res.2
2025 Unsupervised Framing Analysis for Social Media Discourse in Polarizing Events
abstract
This study investigates the concept of frames in the realm of online polarization, with a focus on social media platforms. The research extends the understanding of how frames–emerging, complex, and often subtle concepts–become prominent in online conversations that are polarized. The study proposes a comprehensive methodology for identifying and characterizing these frames, integrating machine learning techniques, network analysis algorithms, and natural language processing tools. This method aims for generalizability across multiple platforms and types of user engagement. Two novel metrics, homogeneity and relevancy are introduced for the rigorous evaluation of identified frame candidates. Grounded in several foundational presumptions, including the role of topics and multi-word expressions in framing, the study sheds light on how frames emerge and gain significance within digital communities. The research questions explored include the methods for identifying frames, the variability and significance of these frames, and the effectiveness of different computational techniques in this context. To validate the approach, we present a case study of the 2021 Chilean presidential election, using data from both \(\mathbb {X}\) (formerly known as Twitter) and WhatsApp platforms. This real-world application allows for the examination of how frames fluctuate in response to events and the specific mechanisms of platforms. Overall, the study makes several key contributions to the field, offering new insights and methodologies for analyzing the complexities of online polarization. It serves as groundwork for future research on the dynamics of online communities, especially those associated with distinctly polarized events.
Hernan Sarmiento, Ricardo Córdova, Felipe Bravo-Marquez, Marcelo Luis Barbosa dos Santos, Sebastián Valenzuela
ACM Trans. Web4
2024 Unpacking Bias: An Empirical Study of Bias Measurement Metrics, Mitigation Algorithms, and Their Interactions
abstract
Word embeddings (WE) have been shown to capture biases from the text they are trained on, which has led to the development of several bias measurement metrics and bias mitigation algorithms (i.e., methods that transform the embedding space to reduce bias). This study identifies three confounding factors that hinder the comparison of bias mitigation algorithms with bias measurement metrics: (1) reliance on different word sets when applying bias mitigation algorithms, (2) leakage between training words employed by mitigation methods and evaluation words used by metrics, and (3) inconsistencies in normalization transformations between mitigation algorithms. We propose a very simple comparison methodology that carefully controls for word sets and vector normalization to address these factors. We conduct a component isolation experiment to assess how each component of our methodology impacts bias measurement. After comparing the bias mitigation algorithms using our comparison methodology, we observe increased consistency between different debiasing algorithms when evaluated using our approach.
Felipe Bravo-Marquez, María José Zambrano
LREC/COLING1
2023 RiverText: A Python Library for Training and Evaluating Incremental Word Embeddings from Text Data Streams
abstract
Word embeddings have become essential components in various information retrieval and natural language processing tasks, such as ranking, document classification, and question answering. However, despite their widespread use, traditional word embedding models present a limitation in their static nature, which hampers their ability to adapt to the constantly evolving language patterns that emerge in sources such as social media and the web (e.g., new hashtags or brand names). To overcome this problem, incremental word embedding algorithms are introduced, capable of dynamically updating word representations in response to new language patterns and processing continuous data streams.
Gabriel Iturra-Bocaz, Felipe Bravo-Marquez
SIGIR2
2022 Simple Yet Powerful: An Overlooked Architecture for Nested Named Entity Recognition
abstract
Named Entity Recognition (NER) is an important task in Natural Language Processing that aims to identify text spans belonging to predefined categories. Traditional NER systems ignore nested entities, which are entities contained in other entity mentions. Although several methods have been proposed to address this case, most of them rely on complex task-specific structures and ignore potentially useful baselines for the task. We argue that this creates an overly optimistic impression of their performance. This paper revisits the Multiple LSTM-CRF (MLC) model, a simple, overlooked, yet powerful approach based on training independent sequence labeling models for each entity type. Extensive experiments with three nested NER corpora show that, regardless of the simplicity of this model, its performance is better or at least as well as more sophisticated methods. Furthermore, we show that the MLC architecture achieves state-of-the-art results in the Chilean Waiting List corpus by including pre-trained language models. In addition, we implemented an open-source library that computes task-specific metrics for nested NER. The results suggest that metrics used in previous work do not measure well the ability of a model to detect nested entities, while our metrics provide new evidence on how existing approaches handle the task.
Matías Rojas, Felipe Bravo-Marquez, Jocelyn Dunstan
COLING2
2022 Identifying and Characterizing New Expressions of Community Framing during Polarization
Hernan Sarmiento, Felipe Bravo-Marquez, Eduardo Graells-Garrido, Barbara Poblete
ICWSM2
2022 Evaluation Benchmarks for Spanish Sentence Representations
abstract
Due to the success of pre-trained language models, versions of languages other than English have been released in recent years. This fact implies the need for resources to evaluate these models. In the case of Spanish, there are few ways to systematically assess the models’ quality. In this paper, we narrow the gap by building two evaluation benchmarks. Inspired by previous work (Conneau and Kiela, 2018; Chen et al., 2019), we introduce Spanish SentEval and Spanish DiscoEval, aiming to assess the capabilities of stand-alone and discourse-aware sentence representations, respectively. Our benchmarks include considerable pre-existing and newly constructed datasets that address different tasks from various domains. In addition, we evaluate and analyze the most recent pre-trained Spanish language models to exhibit their capabilities and limitations. As an example, we discover that for the case of discourse evaluation tasks, mBERT, a language model trained on multiple languages, usually provides a richer latent representation than models trained only with documents in Spanish. We hope our contribution will motivate a fairer, more comparable, and less cumbersome way to evaluate future Spanish language models.
Vladimir Araujo, Andres Carvallo, Souvik Kundu 0008, José Cañete, Marcelo Mendoza, Robert E. Mercer, Felipe Bravo-Marquez, Marie-Francine Moens, Alvaro Soto
LREC7
2022 ALBETO and DistilBETO: Lightweight Spanish Language Models
abstract
In recent years there have been considerable advances in pre-trained language models, where non-English language versions have also been made available. Due to their increasing use, many lightweight versions of these models (with reduced parameters) have also been released to speed up training and inference times. However, versions of these lighter models (e.g., ALBERT, DistilBERT) for languages other than English are still scarce. In this paper we present ALBETO and DistilBETO, which are versions of ALBERT and DistilBERT pre-trained exclusively on Spanish corpora. We train several versions of ALBETO ranging from 5M to 223M parameters and one of DistilBETO with 67M parameters. We evaluate our models in the GLUES benchmark that includes various natural language understanding tasks in Spanish. The results show that our lightweight models achieve competitive results to those of BETO (Spanish-BERT) despite having fewer parameters. More specifically, our larger ALBETO model outperforms all other models on the MLDoc, PAWS-X, XNLI, MLQA, SQAC and XQuAD datasets. However, BETO remains unbeaten for POS and NER. As a further contribution, all models are publicly available to the community for future research.
José Cañete, Sebastian Donoso, Felipe Bravo-Marquez, Andres Carvallo, Vladimir Araujo
LREC3
2022 Automatic Extraction of Nested Entities in Clinical Referrals in Spanish
abstract
Here we describe a new clinical corpus rich in nested entities and a series of neural models to identify them. The corpus comprises de-identified referrals from the waiting list in Chilean public hospitals. A subset of 5,000 referrals (58.6% medical and 41.4% dental) was manually annotated with 10 types of entities, six attributes, and pairs of relations with clinical relevance. In total, there are 110,771 annotated tokens. A trained medical doctor or dentist annotated these referrals, and then, together with three other researchers, consolidated each of the annotations. The annotated corpus has 48.17% of entities embedded in other entities or containing another one. We use this corpus to build models for Named Entity Recognition (NER). The best results were achieved using a Multiple Single-entity architecture with clinical word embeddings stacked with character and Flair contextual embeddings. The entity with the best performance is abbreviation , and the hardest to recognize is finding . NER models applied to this corpus can leverage statistics of diseases and pending procedures. This work constitutes the first annotated corpus using clinical narratives from Chile and one of the few in Spanish. The annotated corpus, clinical word embeddings, annotation guidelines, and neural models are freely released to the community.
Pablo Báez, Felipe Bravo-Marquez, Jocelyn Dunstan, Matías Rojas, Fabián Villena
ACM Trans. Comput. Heal.2
2021 PolyLM: Learning about Polysemy through Language Modeling
abstract
To avoid the "meaning conflation deficiency" of word embeddings, a number of models have aimed to embed individual word senses.These methods at one time performed well on tasks such as word sense induction (WSI), but they have since been overtaken by task-specific techniques which exploit contextualized embeddings.However, sense embeddings and contextualization need not be mutually exclusive.We introduce PolyLM, a method which formulates the task of learning sense embeddings as a language modeling problem, allowing contextualization techniques to be applied.PolyLM is based on two underlying assumptions about word senses: firstly, that the probability of a word occurring in a given context is equal to the sum of the probabilities of its individual senses occurring; and secondly, that for a given occurrence of a word, one of its senses tends to be much more plausible in the context than the others.We evaluate PolyLM on WSI, showing that it performs considerably better than previous sense embedding techniques, and matches the current stateof-the-art specialized WSI method despite having six times fewer parameters.Code and pre-trained models are available at https:// github.com/AlanAnsell/PolyLM.
Alan Ansell, Felipe Bravo-Marquez, Bernhard Pfahringer
EACL2
2020 WEFE: The Word Embeddings Fairness Evaluation Framework
abstract
Word embeddings are known to exhibit stereotypical biases towards gender, race, religion, among other criteria. Severa fairness metrics have been proposed in order to automatically quantify these biases. Although all metrics have a similar objective, the relationship between them is by no means clear. Two issues that prevent a clean comparison is that they operate with different inputs, and that their outputs are incompatible with each other. In this paper we propose WEFE, the word embeddings fairness evaluation framework, to encapsulate, evaluate and compare fairness metrics. Our framework needs a list of pre-trained embeddings and a set of fairness criteria, and it is based on checking correlations between fairness rankings induced by these criteria. We conduct a case study showing that rankings produced by existing fairness methods tend to correlate when measuring gender bias. This correlation is considerably less for other biases like race or religion. We also compare the fairness rankings with an embedding benchmark showing that there is no clear correlation between fairness and good performance in downstream tasks.
Pablo Badilla, Felipe Bravo-Marquez, Jorge Pérez 0001
IJCAI2
2020 An integrated model for textual social media data with spatio-temporal dimensions
Juglar Diaz, Barbara Poblete, Felipe Bravo-Marquez
Inf. Process. Manag.3
2019 AffectiveTweets: a Weka Package for Analyzing Affect in Tweets
abstract
AffectiveTweets is a set of programs for analyzing emotion and sentiment of social media messages such as tweets. It is implemented as a package for the Weka machine learning workbench and provides methods for calculating state-of-the-art affect analysis features from tweets that can be fed into machine learning algorithms implemented in Weka. It also implements methods for building affective lexicons and distant supervision methods for training affective models from unlabeled tweets. The package was used by several teams in the shared tasks: EmoInt 2017 and Affect in Tweets SemEval 2018 Task 1.
Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer, Saif M. Mohammad
J. Mach. Learn. Res.1
2019 WekaDeeplearning4j: A deep learning package for Weka based on Deeplearning4j
Steven Lang, Felipe Bravo-Marquez, Christopher Beckham, Mark A. Hall, Eibe Frank
Knowl. Based Syst.2
2018 Transferring sentiment knowledge between words and tweets
abstract
Message-level and word-level polarity classification are two popular tasks in Twitter sentiment analysis. They have been commonly addressed by training supervised models from labelled data. The main limitation of these models is the high cost of data annotation. Transferring existing labels from a related problem domain is one possible solution for this problem. In this paper, we study how to transfer sentiment labels from the word domain to the tweet domain and vice versa by making their corresponding instances compatible. We model instances of these two domains as the aggregation of instances from the other (i.e., tweets are treated as collections of the words they contain and words are treated as collections of the tweets in which they occur) and perform aggregation by averaging the corresponding constituents. We study two different setups for averaging tweet and word vectors: 1) representing tweets by standard NLP features such as unigrams and part-of-speech tags and words by averaging the vectors of the tweets in which they occur, and 2) representing words using skip-gram embeddings and tweets as the average embedding vector of their words. A consequence of our approach is that instances of both domains reside in the same feature space. Thus, a sentiment classifier trained on labelled data from one domain can be used to classify instances from the other one. We evaluate this approach in two transfer learning tasks: 1) sentiment classification of tweets by applying a word-level sentiment classifier, and 2) induction of a polarity lexicon by applying a tweet-level polarity classifier. Our results show that the proposed model can successfully classify words and tweets after transfer.
Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer
Web Intell.1
2016 Annotate-Sample-Average (ASA): A New Distant Supervision Approach for Twitter Sentiment Analysis
abstract
The classification of tweets into polarity classes is a popular task in sentiment analysis. State-of-the-art solutions to this problem are based on supervised machine learning models trained from manually annotated examples. A drawback of these approaches is the high cost involved in data annotation. Two freely available resources that can be exploited to solve the problem are: 1) large amounts of unlabelled tweets obtained from the Twitter API and 2) prior lexical knowledge in the form of opinion lexicons. In this paper, we propose Annotate-Sample-Average (ASA), a distant supervision method that uses these two resources to generate synthetic training data for Twitter polarity classification. Positive and negative training instances are generated by sampling and averaging unlabelled tweets containing words with the corresponding polarity. Polarity of words is determined from a given polarity lexicon. Our experimental results show that the training data generated by ASA (after tuning its parameters) produces a classifier that performs significantly better than a classifier trained from tweets annotated with emoticons and a classifier trained, without any sampling and averaging, from tweets annotated according to the polarity of their words.
Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer
ECAI1
2016 Determining Word-Emotion Associations from Tweets by Multi-label Classification
abstract
The automatic detection of emotions in Twitter posts is a challenging task due to the informal nature of the language used in this platform. In this paper, we propose a methodology for expanding the NRC word-emotion association lexicon for the language used in Twitter. We perform this expansion using multi-label classification of words and compare different word-level features extracted from unlabelled tweets such as unigrams, Brown clusters, POS tags, and word2vec embeddings. The results show that the expanded lexicon achieves major improvements over the original lexicon when classifying tweets into emotional categories. In contrast to previous work, our methodology does not depend on tweets annotated with emotional hashtags, thus enabling the identification of emotional words from any domain-specific collection using unlabelled tweets.
Felipe Bravo-Marquez, Eibe Frank, Saif M. Mohammad, Bernhard Pfahringer
WI1
2016 From Opinion Lexicons to Sentiment Classification of Tweets and Vice Versa: A Transfer Learning Approach
abstract
Message-level and word-level polarity classification are two popular tasks in Twitter sentiment analysis. They have been commonly addressed by training supervised models from labelled data. The main limitation of these models is the high cost of data annotation. Transferring existing labels from a related problem domain is one possible solution for this problem. In this paper, we propose a simple model for transferring sentiment labels from words to tweets and vice versa by representing both tweets and words using feature vectors residing in the same feature space. Tweets are represented by standard NLP features such as unigrams and part-of-speech tags. Words are represented by averaging the vectors of the tweets in which they occur. We evaluate our approach in two transfer learning problems: 1) training a tweet-level polarity classifier from a polarity lexicon, and 2) inducing a polarity lexicon from a collection of polarity-annotated tweets. Our results show that the proposed approach can successfully classify words and tweets after transfer.
Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer
WI1
2016 Building a Twitter opinion lexicon from automatically-annotated tweets
Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer
Knowl. Based Syst.1
2015 Positive, Negative, or Neutral: Learning an Expanded Opinion Lexicon from Emoticon-Annotated Tweets
Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer
IJCAI1
2015 From Unlabelled Tweets to Twitter-specific Opinion Words
abstract
In this article, we propose a word-level classification model for automatically generating a Twitter-specific opinion lexicon from a corpus of unlabelled tweets. The tweets from the corpus are represented by two vectors: a bag-of-words vector and a semantic vector based on word-clusters. We propose a distributional representation for words by treating them as the centroids of the tweet vectors in which they appear. The lexicon generation is conducted by training a word-level classifier using these centroids to form the instance space and a seed lexicon to label the training instances. Experimental results show that the two types of tweet vectors complement each other in a statistically significant manner and that our generated lexicon produces significant improvements for tweet-level polarity classification.
Felipe Bravo-Marquez, Eibe Frank, Bernhard Pfahringer
SIGIR1
2014 A novel deterministic approach for aspect-based opinion mining in tourism products reviews
Edison Marrese-Taylor, Juan D. Velásquez 0001, Felipe Bravo-Marquez
Expert Syst. Appl.3
2014 Meta-level sentiment models for big social data analysis
Felipe Bravo-Marquez, Marcelo Mendoza, Barbara Poblete
Knowl. Based Syst.1
2013 Identifying Customer Preferences about Tourism Products Using an Aspect-based Opinion Mining Approach
abstract
In this study we extend Bing Liu's aspect-based opinion mining technique to apply it to the tourism domain. Using this extension, we also offer an approach for considering a new alternative to discover consumer preferences about tourism products, particularly hotels and restaurants, using opinions available on the Web as reviews. An experiment is also conducted, using hotel and restaurant reviews obtained from TripAdvisor, to evaluate our proposals. Results showed that tourism product reviews available on web sites contain valuable information about customer preferences that can be extracted using an aspect-based opinion mining approach. The proposed approach proved to be very effective in determining the sentiment orientation of opinions, achieving a precision and recall of 90%. However, on average, the algorithms were only capable of extracting 35% of the explicit aspect expressions.
Edison Marrese-Taylor, Juan D. Velásquez 0001, Felipe Bravo-Marquez, Yutaka Matsuo
KES3
2012 A Zipf-Like Distant Supervision Approach for Multi-document Summarization Using Wikinews Articles
Felipe Bravo-Marquez, Manuel Manriquez
SPIRE1
2011 A Text Similarity Meta-Search Engine Based on Document Fingerprints and Search Results Records
abstract
The retrieval of similar documents from the Web using documents as input instead of key-term queries is not currently supported by traditional Web search engines. One approach for solving the problem consists of fingerprint the document's content into a set of queries that are submitted to a list of Web search engines. Afterward, results are merged, their URLs are fetched and their content is compared with the given document using text comparison algorithms. However, the action of requesting results to multiple web servers could take a significant amount of time and effort. In this work, a similarity function between the given document and retrieved results is estimated. The function uses as variables features that come from information provided by search engine results records, like rankings, titles and snippets. Avoiding therefore, the bottleneck of requesting external Web Servers. We created a collection of around 10,000 search engine results by generating queries from 2,000 crawled Web documents. Then we fitted the similarity function using the cosine similarity between the input and results content as the target variable. The execution time between the exact and approximated solution was compared. Results obtained for our approximated solution showed a reduction of computational time of 86% at an acceptable level of precision with respect to the exact solution of the web document retrieval problem.
Felipe Bravo-Marquez, Gaston L'Huillier, Sebastián A. Ríos, Juan D. Velásquez 0001
Web Intelligence1
2010 DOCODE-Lite: A Meta-Search Engine for Document Similarity Retrieval
Felipe Bravo-Marquez, Gaston L'Huillier, Sebastián A. Ríos, Juan D. Velásquez 0001, Luis A. Guerrero
KES (2)1
2010 Hypergeometric Language Model and Zipf-Like Scoring Function for Web Document Similarity Retrieval
Felipe Bravo-Marquez, Gaston L'Huillier, Sebastián A. Ríos, Juan D. Velásquez 0001
SPIRE1