Enrique Amigó

dblp:71/5174 · DBLP profile ↗
← Back
25ranked-venue papers in the field
17as first author
6since 2021 · last 2025
0000-0003-1482-824XORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 23 (16 first)Database Systems & Data Management · 1 (1 first)Data Mining & Knowledge Discovery · 1
YearPublicationVenuePosition
2025 EXIST 2025: Learning with Disagreement for Sexism Identification and Characterization in Tweets, Memes, and TikTok Videos
Laura Plaza, Jorge Carrillo de Albornoz, Iván Árcos, Paolo Rosso, Damiano Spina, Enrique Amigó, Julio Gonzalo 0001, Roser Morante
ECIR (5)6
2024 EXIST 2024: sEXism Identification in Social neTworks and Memes
Laura Plaza, Jorge Carrillo de Albornoz, Enrique Amigó, Julio Gonzalo 0001, Roser Morante, Paolo Rosso, Damiano Spina, Berta Chulvi, Alba Maeso, Víctor Ruiz
ECIR (5)3
2023 Overview of EXIST 2023: sEXism Identification in Social NeTworks
Laura Plaza, Jorge Carrillo de Albornoz, Roser Morante, Enrique Amigó, Julio Gonzalo 0001, Damiano Spina, Paolo Rosso
ECIR (3)4
2023 A unifying and general account of fairness measurement in recommender systems
abstract
Fairness is fundamental to all information access systems, including recommender systems. However, the landscape of fairness definition and measurement is quite scattered with many competing definitions that are partial and often incompatible. There is much work focusing on specific – and different – notions of fairness and there exist dozens of metrics of fairness in the literature, many of them redundant and most of them incompatible. In contrast, to our knowledge, there is no formal framework that covers all possible variants of fairness and allows developers to choose the most appropriate variant depending on the particular scenario. In this paper, we aim to define a general, flexible, and parameterizable framework that covers a whole range of fairness evaluation possibilities. Instead of modeling the metrics based on an abstract definition of fairness, the distinctive feature of this study compared to the current state of the art is that we start from the metrics applied in the literature to obtain a unified model by generalization. The framework is grounded on a general work hypothesis: interpreting the space of users and items as a probabilistic sample space, two fundamental measures in information theory (Kullback–Leibler Divergence and Mutual Information) can capture the majority of possible scenarios for measuring fairness on recommender system outputs. In addition, earlier research on fairness in recommender systems could be viewed as single-sided, trying to optimize some form of equity across either user groups or provider/procurer groups, without considering the user/item space in conjunction, thereby overlooking/disregarding the interplay between user and item groups. Instead, our framework includes the notion of statistical independence between user and item groups. We finally validate our approach experimentally on both synthetic and real data according to a wide range of state-of-the-art recommendation algorithms and real-world data sets, showing that with our framework we can measure fairness in a general, uniform, and meaningful way.
Enrique Amigó, Yashar Deldjoo, Stefano Mizzaro, Alejandro Bellogín
Inf. Process. Manag.1
2023 What is My Problem? Identifying Formal Tasks and Metrics in Data Mining on the Basis of Measurement Theory
abstract
The design and analysis of experimental research in Data Mining (DM) is anchored in a correct choice of the type of task addressed (clustering, classification, regression, etc.). However, although DM is a relatively mature discipline, there is no consensus yet about what is the taxonomy of DM tasks, which are their formal characteristics, and their corresponding metrics. In this paper, we formalize DM tasks in terms of Measurement Theory, which is a cornerstone of quantitative research in many disciplines, but has not yet been incorporated (in a consensual way) into some areas of Computer Science, including DM. The proposed formal framework provides a methodology to precisely define DM tasks for any given scenario and identify appropriate metrics. We validate this framework via (i) its coverage of existing DM tasks, (ii) its capability to group existing metrics into families, and (iii) its coverage of actual DM research problems, using about 250 papers from ACM KDD 2019 and IEEE ICDM 2019 conferences as reference sample.
Enrique Amigó, Julio Gonzalo 0001, Stefano Mizzaro
IEEE Trans. Knowl. Data Eng.1
2022 Ranking Interruptus: When Truncated Rankings Are Better and How to Measure That
abstract
Most of information retrieval effectiveness evaluation metrics assume that systems appending irrelevant documents at the bottom of the ranking are as effective as (or not worse than) systems that have a stopping criteria to 'truncate' the ranking at the right position to avoid retrieving those irrelevant documents at the end. It can be argued, however, that such truncated rankings are more useful to the end user. It is thus important to understand how to measure retrieval effectiveness in this scenario. In this paper we provide both theoretical and experimental contributions. We first define formal properties to analyze how effectiveness metrics behave when evaluating truncated rankings. Our theoretical analysis shows that de-facto standard metrics do not satisfy desirable properties to evaluate truncated rankings: only Observational Information Effectiveness (OIE) -- a metric based on Shannon's information theory -- satisfies them all. We then perform experiments to compare several metrics on nine TREC datasets. According to our experimental results, the most appropriate metrics for truncated rankings are OIE and a novel extension of Rank-Biased Precision that adds a user effort factor penalizing the retrieval of irrelevant documents.
Enrique Amigó, Stefano Mizzaro, Damiano Spina
SIGIR1
2020 Axiomatic thinking for information retrieval: introduction to special issue
Enrique Amigó, Hui Fang 0001, Stefano Mizzaro, ChengXiang Zhai
Inf. Retr. J.1
2020 On the foundations of similarity in information access
Enrique Amigó, Fernando Giner, Julio Gonzalo 0001, M. Felisa Verdejo
Inf. Retr. J.1
2020 On the nature of information access evaluation metrics: a unifying framework
Enrique Amigó, Stefano Mizzaro
Inf. Retr. J.1
2020 Integrating learned and explicit document features for reputation monitoring in social media
Fernando Giner, Enrique Amigó, M. Felisa Verdejo
Knowl. Inf. Syst.2
2019 A comparison of filtering evaluation metrics based on formal constraints
Enrique Amigó, Julio Gonzalo 0001, M. Felisa Verdejo, Damiano Spina
Inf. Retr. J.1
2018 Are we on the Right Track?: An Examination of Information Retrieval Methodologies
abstract
The unpredictability of user behavior and the need for effectiveness make it difficult to define a suitable research methodology for Information Retrieval (IR). In order to tackle this challenge, we categorize existing IR methodologies along two dimensions: (1) empirical vs. theoretical, and (2) top-down vs. bottom-up. The strengths and drawbacks of the resulting categories are characterized according to 6 desirable aspects. The analysis suggests that different methodologies are complementary and therefore, equally necessary. The categorization of the 167 full papers published in the last SIGIR (2016 and 2017) and ICTIR (2017) conferences suggest that most of existing work is empirical bottom-up, suggesting lack of some desirable aspects. With the hope of improving IR research practice, we propose a general methodology for IR that integrates the strengths of existing research methods.
Enrique Amigó, Hui Fang 0001, Stefano Mizzaro, ChengXiang Zhai
SIGIR1
2018 An Axiomatic Analysis of Diversity Evaluation Metrics: Introducing the Rank-Biased Utility Metric
abstract
Many evaluation metrics have been defined to evaluate the effectiveness ad-hoc retrieval and search result diversification systems. However, it is often unclear which evaluation metric should be used to analyze the performance of retrieval systems given a specific task. Axiomatic analysis is an informative mechanism to understand the fundamentals of metrics and their suitability for particular scenarios. In this paper, we define a constraint-based axiomatic framework to study the suitability of existing metrics in search result diversification scenarios. The analysis informed the definition of Rank-Biased Utility (RBU) -- an adaptation of the well-known Rank-Biased Precision metric -- that takes into account redundancy and the user effort associated to the inspection of documents in the ranking. Our experiments over standard diversity evaluation campaigns show that the proposed metric captures quality criteria reflected by different metrics, being suitable in the absence of knowledge about particular features of the scenario under study.
Enrique Amigó, Damiano Spina, Jorge Carrillo de Albornoz
SIGIR1
2017 A Formal and Empirical Study of Unsupervised Signal Combination for Textual Similarity Tasks
Enrique Amigó, Fernando Giner, Julio Gonzalo 0001, M. Felisa Verdejo
ECIR1
2017 Axiomatic Thinking for Information Retrieval: And Related Tasks
abstract
This is the first workshop on the emerging interdisciplinary research area of applying axiomatic thinking to information retrieval (IR) and related tasks. The workshop aims to help foster collaboration of researchers working on different perspectives of axiomatic thinking and encourage discussion and research on general methodological issues related to applying axiomatic thinking to IR and related tasks.
Enrique Amigó, Hui Fang 0001, Stefano Mizzaro, ChengXiang Zhai
SIGIR1
2017 EvALL: Open Access Evaluation for Information Access Systems
abstract
The EvALL online evaluation service aims to provide a unified evaluation framework for Information Access systems that makes results completely comparable and publicly available for the whole research community. For researchers working on a given test collection, the framework allows to: (i) evaluate results in a way compliant with measurement theory and with state-of-the-art evaluation practices in the field; (ii) quantitatively and qualitatively compare their results with the state of the art; (iii) provide their results as reusable data to the scientific community; (iv) automatically generate evaluation figures and (low-level) interpretation of the results, both as a pdf report and as a latex source. For researchers running a challenge (a comparative evaluation campaign on shared data), the framework helps them to manage, store and evaluate submissions, and to preserve ground truth and system output data for future use by the research community. EvALL can be tested at http://evall.uned.es.
Enrique Amigó, Jorge Carrillo de Albornoz, Mario Almagro-Cádiz, Julio Gonzalo 0001, Javier Rodríguez-Vidal, M. Felisa Verdejo
SIGIR1
2016 Tweet Stream Summarization for Online Reputation Management
Jorge Carrillo de Albornoz, Enrique Amigó, Laura Plaza, Julio Gonzalo 0001
ECIR2
2015 A Formal Approach to Effectiveness Metrics for Information Access: Retrieval, Filtering, and Clustering
Enrique Amigó, Julio Gonzalo 0001, Stefano Mizzaro
ECIR1
2014 ORMA: A Semi-automatic Tool for Online Reputation Monitoring in Twitter
Jorge Carrillo de Albornoz, Enrique Amigó, Damiano Spina, Julio Gonzalo 0001
ECIR2
2014 A general account of effectiveness metrics for information tasks: retrieval, filtering, and clustering
abstract
In this tutorial we will present, review, and compare the most popular evaluation metrics for some of the most salient information related tasks, covering: (i) Information Retrieval, (ii) Clustering, and (iii) Filtering. The tutorial will make a special emphasis on the specification of constraints for suitable metrics in each of the three tasks, and on the systematic comparison of metrics according to such constraints. The last part of the tutorial will investigate the challenge of combining and weighting metrics.
Enrique Amigó, Julio Gonzalo 0001, Stefano Mizzaro
SIGIR1
2014 Learning similarity functions for topic detection in online reputation monitoring
abstract
Reputation management experts have to monitor--among others--Twitter constantly and decide, at any given time, what is being said about the entity of interest (a company, organization, personality...). Solving this reputation monitoring problem automatically as a topic detection task is both essential--manual processing of data is either costly or prohibitive--and challenging--topics of interest for reputation monitoring are usually fine-grained and suffer from data sparsity. We focus on a solution for the problem that (i) learns a pairwise tweet similarity function from previously annotated data, using all kinds of content-based and Twitter-based features; (ii) applies a clustering algorithm on the previously learned similarity function. Our experiments indicate that (i) Twitter signals can be used to improve the topic detection process with respect to using content signals only; (ii) learning a similarity function is a flexible and efficient way of introducing supervision in the topic detection clustering process. The performance of our best system is substantially better than state-of-the-art approaches and gets close to the inter-annotator agreement rate. A detailed qualitative inspection of the data further reveals two types of topics detected by reputation experts: reputation alerts / issues (which usually spike in time) and organizational topics (which are usually stable across time).
Damiano Spina, Julio Gonzalo 0001, Enrique Amigó
SIGIR3
2013 An unsupervised transfer learning approach to discover topics for online reputation management
abstract
Microblogs play an important role for Online Reputation Management. Companies and organizations in general have an increasing interest in obtaining the last minute information about which are the emerging topics that concern their reputation. In this paper, we present a new technique to cluster a collection of tweets emitted within a short time span about a specific entity. Our approach relies on transfer learning by contextualizing a target collection of tweets with a large set of unlabeled "background" tweets that help improving the clustering of the target collection. We include background tweets together with target tweets in a TwitterLDA process, and we set the total number of clusters. In practice, this means that the system can adapt to find the right number of clusters for the target data, overcoming one of the limitations of using LDA-based approaches (the need of establishing a priori the number of clusters). Our experiments using RepLab 2012 data show that using the background collection gives a 20% improvement over a direct application of TwitterLDA using only the target collection. Our data also confirms that the approach can effectively predict the right number of target clusters in a way that is robust with respect to the total number of clusters established a priori.
Tamara Martín-Wanton, Julio Gonzalo 0001, Enrique Amigó
CIKM3
2013 A general evaluation measure for document organization tasks
abstract
A number of key Information Access tasks -- Document Retrieval, Clustering, Filtering, and their combinations -- can be seen as instances of a generic {\em document organization} problem that establishes priority and relatedness relationships between documents (in other words, a problem of forming and ranking clusters). As far as we know, no analysis has been made yet on the evaluation of these tasks from a global perspective. In this paper we propose two complementary evaluation measures -- Reliability and Sensitivity -- for the generic Document Organization task which are derived from a proposed set of formal constraints (properties that any suitable measure must satisfy).
Enrique Amigó, Julio Gonzalo 0001, M. Felisa Verdejo
SIGIR1
2009 A comparison of extrinsic clustering evaluation metrics based on formal constraints
Enrique Amigó, Julio Gonzalo 0001, Javier Artiles, M. Felisa Verdejo
Inf. Retr.1
2009 A comparison of extrinsic clustering evaluation metrics based on formal constraints
Enrique Amigó, Julio Gonzalo 0001, Javier Artiles, M. Felisa Verdejo
Inf. Retr.1