VLDB 2026 Research / reviewers in the wild / expert
Julio Gonzalo 0001
dblp:52/3559 · also Julio Antonio Gonzalo Arroyo
· DBLP profile ↗
31ranked-venue papers in the field
2as first author
6since 2021 · last 2025
0000-0002-5341-9337ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 30 (2 first)Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EXIST 2025: Learning with Disagreement for Sexism Identification and Characterization in Tweets, Memes, and TikTok Videos
Laura Plaza, Jorge Carrillo de Albornoz, Iván Árcos, Paolo Rosso, Damiano Spina, Enrique Amigó, Julio Gonzalo 0001, Roser Morante |
ECIR (5) | 7 |
| 2024 | The CLEF 2024 Monster Track: One Lab to Rule Them All
Nicola Ferro 0001, Julio Gonzalo 0001, Jussi Karlgren, Henning Müller |
ECIR (6) | 2 |
| 2024 | EXIST 2024: sEXism Identification in Social neTworks and Memes
Laura Plaza, Jorge Carrillo de Albornoz, Enrique Amigó, Julio Gonzalo 0001, Roser Morante, Paolo Rosso, Damiano Spina, Berta Chulvi, Alba Maeso, Víctor Ruiz |
ECIR (5) | 4 |
| 2023 | Overview of EXIST 2023: sEXism Identification in Social NeTworks
Laura Plaza, Jorge Carrillo de Albornoz, Roser Morante, Enrique Amigó, Julio Gonzalo 0001, Damiano Spina, Paolo Rosso |
ECIR (3) | 5 |
| 2023 | What is My Problem? Identifying Formal Tasks and Metrics in Data Mining on the Basis of Measurement TheoryabstractThe design and analysis of experimental research in Data Mining (DM) is anchored in a correct choice of the type of task addressed (clustering, classification, regression, etc.). However, although DM is a relatively mature discipline, there is no consensus yet about what is the taxonomy of DM tasks, which are their formal characteristics, and their corresponding metrics. In this paper, we formalize DM tasks in terms of Measurement Theory, which is a cornerstone of quantitative research in many disciplines, but has not yet been incorporated (in a consensual way) into some areas of Computer Science, including DM. The proposed formal framework provides a methodology to precisely define DM tasks for any given scenario and identify appropriate metrics. We validate this framework via (i) its coverage of existing DM tasks, (ii) its capability to group existing metrics into families, and (iii) its coverage of actual DM research problems, using about 250 papers from ACM KDD 2019 and IEEE ICDM 2019 conferences as reference sample. Enrique Amigó, Julio Gonzalo 0001, Stefano Mizzaro |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2021 | Authority and priority signals in automatic summary generation for online reputation managementabstractAbstract Online reputation management (ORM) comprises the collection of techniques that help monitoring and improving the public image of an entity (companies, products, institutions) on the Internet. The ORM experts try to minimize the negative impact of the information about an entity while maximizing the positive material for being more trustworthy to the customers. Due to the huge amount of information that is published on the Internet every day, there is a need to summarize the entire flow of information to obtain only those data that are relevant to the entities. Traditionally the automatic summarization task in the ORM scenario takes some in‐domain signals into account such as popularity, polarity for reputation and novelty but exists other feature to be considered, the authority of the people. This authority depends on the ability to convince others and therefore to influence opinions. In this work, we propose the use of authority signals that measures the influence of a user jointly with (a) priority signals related to the ORM domain and (b) information regarding the different topics that influential people is talking about. Our results indicate that the use of authority signals may significantly improve the quality of the summaries that are automatically generated. Javier Rodríguez-Vidal, Jorge Carrillo de Albornoz, Julio Gonzalo 0001, Laura Plaza |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2020 | On the foundations of similarity in information access
Enrique Amigó, Fernando Giner, Julio Gonzalo 0001, M. Felisa Verdejo |
Inf. Retr. J. | 3 |
| 2019 | Propagating sentiment signals for estimating reputation polarity
Anastasia Giahanou, Julio Gonzalo 0001, Fabio Crestani |
Inf. Process. Manag. | 2 |
| 2019 | A comparison of filtering evaluation metrics based on formal constraints
Enrique Amigó, Julio Gonzalo 0001, M. Felisa Verdejo, Damiano Spina |
Inf. Retr. J. | 2 |
| 2019 | Automatic detection of influencers in social networks: Authority versus domain signalsabstractGiven the task of finding influencers (opinion makers) for a given domain in a social network, we investigate (a) what is the relative importance of domain and authority signals, (b) what is the most effective way of combining signals (voting, classification, learning to rank, etc.) and how best to model the vocabulary signal, and (c) how large is the gap between supervised and unsupervised methods and what are the practical consequences. Our best results on the RepLab dataset (which improves the state of the art) uses language models to learn the domain‐specific vocabulary used by influencers and combines domain and authority models using a Learning to Rank algorithm. Our experiments show that (a) both authority and domain evidence can be trained from the vocabulary of influencers; (b) once the language of influencers is modeled as a likelihood signal, further supervised learning and additional network‐based signals only provide marginal improvements; and (c) the availability of training data sets is crucial to obtain competitive results in the task. Our most remarkable finding is that influencers do use a distinctive vocabulary, which is a more reliable signal than nontextual network indicators such as the number of followers, retweets, and so on. Javier Rodríguez-Vidal, Julio Gonzalo 0001, Laura Plaza, Henry Anaya-Sánchez |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2017 | A Formal and Empirical Study of Unsupervised Signal Combination for Textual Similarity Tasks
Enrique Amigó, Fernando Giner, Julio Gonzalo 0001, M. Felisa Verdejo |
ECIR | 3 |
| 2017 | Sentiment Propagation for Predicting Reputation Polarity
Anastasia Giahanou, Julio Gonzalo 0001, Ida Mele, Fabio Crestani |
ECIR | 2 |
| 2017 | EvALL: Open Access Evaluation for Information Access SystemsabstractThe EvALL online evaluation service aims to provide a unified evaluation framework for Information Access systems that makes results completely comparable and publicly available for the whole research community. For researchers working on a given test collection, the framework allows to: (i) evaluate results in a way compliant with measurement theory and with state-of-the-art evaluation practices in the field; (ii) quantitatively and qualitatively compare their results with the state of the art; (iii) provide their results as reusable data to the scientific community; (iv) automatically generate evaluation figures and (low-level) interpretation of the results, both as a pdf report and as a latex source. For researchers running a challenge (a comparative evaluation campaign on shared data), the framework helps them to manage, store and evaluate submissions, and to preserve ground truth and system output data for future use by the research community. EvALL can be tested at http://evall.uned.es. Enrique Amigó, Jorge Carrillo de Albornoz, Mario Almagro-Cádiz, Julio Gonzalo 0001, Javier Rodríguez-Vidal, M. Felisa Verdejo |
SIGIR | 4 |
| 2016 | Tweet Stream Summarization for Online Reputation Management
Jorge Carrillo de Albornoz, Enrique Amigó, Laura Plaza, Julio Gonzalo 0001 |
ECIR | 4 |
| 2015 | A Formal Approach to Effectiveness Metrics for Information Access: Retrieval, Filtering, and Clustering
Enrique Amigó, Julio Gonzalo 0001, Stefano Mizzaro |
ECIR | 2 |
| 2014 | ORMA: A Semi-automatic Tool for Online Reputation Monitoring in Twitter
Jorge Carrillo de Albornoz, Enrique Amigó, Damiano Spina, Julio Gonzalo 0001 |
ECIR | 4 |
| 2014 | A general account of effectiveness metrics for information tasks: retrieval, filtering, and clusteringabstractIn this tutorial we will present, review, and compare the most popular evaluation metrics for some of the most salient information related tasks, covering: (i) Information Retrieval, (ii) Clustering, and (iii) Filtering. The tutorial will make a special emphasis on the specification of constraints for suitable metrics in each of the three tasks, and on the systematic comparison of metrics according to such constraints. The last part of the tutorial will investigate the challenge of combining and weighting metrics. Enrique Amigó, Julio Gonzalo 0001, Stefano Mizzaro |
SIGIR | 2 |
| 2014 | SIGIR 2014 workshop on semantic matching in information retrievalabstractRecently, significant progress has been made in research on what we call semantic matching (SM), in web search, question answering, online advertisement, cross-language information retrieval, and other tasks. Advanced technologies based on machine learning have been developed. Let us take Web search as example of the problem that also pervades the other tasks. When comparing the textual content of query and documents, Web search still heavily relies on the term-based approach, where the relevance scores between queries and documents are calculated on the basis of the degree of matching between query terms and document terms. This simple approach works rather well in practice, partly because there are many other signals in web search (hypertext, user logs, etc.) that complement it. However, when considering the long tail of web searches, it can suffer from data sparseness, e.g., Trenton does not match New Jersey Capital. Query document mismatches occur when searcher and author use different terms (representations), and this phenomenon is prevalent due to the nature of human language. Julio Gonzalo 0001, Hang Li 0001, Alessandro Moschitti, Jun Xu 0001 |
SIGIR | 1 |
| 2014 | Learning similarity functions for topic detection in online reputation monitoringabstractReputation management experts have to monitor--among others--Twitter constantly and decide, at any given time, what is being said about the entity of interest (a company, organization, personality...). Solving this reputation monitoring problem automatically as a topic detection task is both essential--manual processing of data is either costly or prohibitive--and challenging--topics of interest for reputation monitoring are usually fine-grained and suffer from data sparsity. We focus on a solution for the problem that (i) learns a pairwise tweet similarity function from previously annotated data, using all kinds of content-based and Twitter-based features; (ii) applies a clustering algorithm on the previously learned similarity function. Our experiments indicate that (i) Twitter signals can be used to improve the topic detection process with respect to using content signals only; (ii) learning a similarity function is a flexible and efficient way of introducing supervision in the topic detection clustering process. The performance of our best system is substantially better than state-of-the-art approaches and gets close to the inter-annotator agreement rate. A detailed qualitative inspection of the data further reveals two types of topics detected by reputation experts: reputation alerts / issues (which usually spike in time) and organizational topics (which are usually stable across time). Damiano Spina, Julio Gonzalo 0001, Enrique Amigó |
SIGIR | 2 |
| 2013 | An unsupervised transfer learning approach to discover topics for online reputation managementabstractMicroblogs play an important role for Online Reputation Management. Companies and organizations in general have an increasing interest in obtaining the last minute information about which are the emerging topics that concern their reputation. In this paper, we present a new technique to cluster a collection of tweets emitted within a short time span about a specific entity. Our approach relies on transfer learning by contextualizing a target collection of tweets with a large set of unlabeled "background" tweets that help improving the clustering of the target collection. We include background tweets together with target tweets in a TwitterLDA process, and we set the total number of clusters. In practice, this means that the system can adapt to find the right number of clusters for the target data, overcoming one of the limitations of using LDA-based approaches (the need of establishing a priori the number of clusters). Our experiments using RepLab 2012 data show that using the background collection gives a 20% improvement over a direct application of TwitterLDA using only the target collection. Our data also confirms that the approach can effectively predict the right number of target clusters in a way that is robust with respect to the total number of clusters established a priori. Tamara Martín-Wanton, Julio Gonzalo 0001, Enrique Amigó |
CIKM | 2 |
| 2013 | A general evaluation measure for document organization tasksabstractA number of key Information Access tasks -- Document Retrieval, Clustering, Filtering, and their combinations -- can be seen as instances of a generic {\em document organization} problem that establishes priority and relatedness relationships between documents (in other words, a problem of forming and ranking clusters). As far as we know, no analysis has been made yet on the evaluation of these tasks from a global perspective. In this paper we propose two complementary evaluation measures -- Reliability and Sensitivity -- for the generic Document Organization task which are derived from a proposed set of formal constraints (properties that any suitable measure must satisfy). Enrique Amigó, Julio Gonzalo 0001, M. Felisa Verdejo |
SIGIR | 2 |
| 2009 | A comparison of extrinsic clustering evaluation metrics based on formal constraints
Enrique Amigó, Julio Gonzalo 0001, Javier Artiles, M. Felisa Verdejo |
Inf. Retr. | 2 |
| 2009 | A comparison of extrinsic clustering evaluation metrics based on formal constraints
Enrique Amigó, Julio Gonzalo 0001, Javier Artiles, M. Felisa Verdejo |
Inf. Retr. | 2 |
| 2008 | Workshop on Novel Methodologies for Evaluation in Information Retrieval
Mark Sanderson, Martin Braschler, Nicola Ferro 0001, Julio Gonzalo 0001 |
ECIR | 4 |
| 2008 | Web people search: results of the first evaluation and the plan for the secondabstractThis paper presents the motivation, resources and results for the first Web People Search task, which was organized as part of the SemEval-2007 evaluation exercise. Also, we will describe a survey and proposal for a new task, "attribute extraction", which is planned for inclusion in the second evaluation, planned for autumn, 2008. Javier Artiles, Satoshi Sekine, Julio Gonzalo 0001 |
WWW | 3 |
| 2008 | Interactive question answering: Is Cross-Language harder than monolingual searching?
Fernando López-Ostenero, Víctor Peinado, Julio Gonzalo 0001, M. Felisa Verdejo |
Inf. Process. Manag. | 3 |
| 2005 | A testbed for people searching strategies in the WWWabstractThis paper describes the creation of a testbed to evaluate people searching strategies on the World-Wide-Web. This task involves resolving person names' ambiguity and locating relevant information characterising every individual under the same name. Javier Artiles, Julio Gonzalo 0001, M. Felisa Verdejo |
SIGIR | 2 |
| 2005 | The impact of evaluation on multilingual text retrievalabstractWe summarize the impact of the first five years of activity of the Cross-Language Evaluation Forum (CLEF) on multilingual text retrieval system performance and show how the CLEF evaluation campaigns have contributed to advances in the state-of-the-art. Julio Gonzalo 0001, Carol Peters |
SIGIR | 1 |
| 2005 | Evaluating Hierarchical Clustering of Search Results
Juan M. Cigarrán, Anselmo Peñas, Julio Gonzalo 0001, M. Felisa Verdejo |
SPIRE | 3 |
| 2005 | Noun phrases as building blocks for cross-language Search Assistance
Fernando López-Ostenero, Julio Gonzalo 0001, M. Felisa Verdejo |
Inf. Process. Manag. | 2 |
| 2004 | Interactive Cross-Language Document Selection
Douglas W. Oard, Julio Gonzalo 0001, Mark Sanderson, Fernando López-Ostenero, Jianqiang Wang 0002 |
Inf. Retr. | 2 |