Gabriela González Sáez

dblp:274/1869 · also Gabriela Nicole González Sáez · DBLP profile ↗
← Back
7ranked-venue papers in the field
2as first author
7since 2021 · last 2025
0000-0003-0878-5263ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 7 (2 first)
YearPublicationVenuePosition
2025 LongEval at CLEF 2025: Longitudinal Evaluation of IR Model Performance
Matteo Cancellieri, Alaa El-Ebshihy, Tobias Fink, Petra Galuscáková, Gabriela González Sáez, Lorraine Goeuriot, David Iommi, Jüri Keller, Petr Knoth, Philippe Mulhem, Florina Piroi, David Pride, Philipp Schaer
ECIR (5)5
2024 LongEval: Longitudinal Evaluation of Model Performance at CLEF 2024
Rabab Alkhalifa, Hsuvas Borkakoty, Romain Deveaud, Alaa El-Ebshihy, Luis Espinosa Anke, Tobias Fink, Gabriela González Sáez, Petra Galuscáková, Lorraine Goeuriot, David Iommi, Maria Liakata, Harish Tayyar Madabushi, Pablo Medina-Alias, Philippe Mulhem, Florina Piroi, Martin Popel, Christophe Servan, Arkaitz Zubiaga
ECIR (6)7
2023 LongEval: Longitudinal Evaluation of Model Performance at CLEF 2023
Rabab Alkhalifa, Iman Munire Bilal, Hsuvas Borkakoty, José Camacho-Collados, Romain Deveaud, Alaa El-Ebshihy, Luis Espinosa Anke, Gabriela González Sáez, Petra Galuscáková, Lorraine Goeuriot, Elena Kochkina, Maria Liakata, Daniel Loureiro, Harish Tayyar Madabushi, Philippe Mulhem, Florina Piroi, Martin Popel, Christophe Servan, Arkaitz Zubiaga
ECIR (3)8
2023 LongEval-Retrieval: French-English Dynamic Test Collection for Continuous Web Search Evaluation
abstract
LongEval-Retrieval is a Web document retrieval benchmark that focuses on continuous retrieval evaluation. This test collection is intended to be used to study the temporal persistence of Information Retrieval systems and will be used as the test collection in the Longitudinal Evaluation of Model Performance Track (LongEval) at CLEF 2023. This benchmark simulates an evolving information system environment - such as the one a Web search engine operates in - where the document collection, the query distribution, and relevance all move continuously, while following the Cranfield paradigm for offline evaluation. To do that, we introduce the concept of a dynamic test collection that is composed of successive sub-collections each representing the state of an information system at a given time step. In LongEval-Retrieval, each sub-collection contains a set of queries, documents, and soft relevance assessments built from click models. The data comes from Qwant, a privacy-preserving Web search engine that primarily focuses on the French market. LongEval-Retrieval also provides a 'mirror' collection: it is initially constructed in the French language to benefit from the majority of Qwant's traffic, before being translated to English. This paper presents the creation process of LongEval-Retrieval and provides baseline runs and analysis.
Petra Galuscáková, Romain Deveaud, Gabriela González Sáez, Philippe Mulhem, Lorraine Goeuriot, Florina Piroi, Martin Popel
SIGIR3
2023 Exploratory Visualization Tool for the Continuous Evaluation of Information Retrieval Systems
abstract
This paper introduces a novel visualization tool that facilitates the exploratory analysis of continuous evaluation for information retrieval systems. We base our analysis on score standardization and meta-analysis techniques applied to Information Retrieval evaluation. We present three functionalities: evaluation overview, delta evaluation, and meta-analysis applied to three perspectives: evaluation rounds, queries, and systems. To illustrate the use of the tool, we provide an example using the TREC-COVID test collection.
Gabriela González Sáez, Petra Galuscáková, Romain Deveaud, Lorraine Goeuriot, Philippe Mulhem
SIGIR1
2022 Continuous Result Delta Evaluation of IR Systems
abstract
Classical evaluation of information retrieval systems evaluates a system in a static test collection. In the case of Web search, the evaluation environment (EE) is continuously changing and the hypothesis of using a static test collection is not representative of this changing reality. Moreover, the changes in the evaluation environment, as the document set, the topics set, the relevance judgments, and the chosen metrics, have an impact on the performance measurement [1, 4]. To the best of our knowledge, there is no way to evaluate two versions of a search engine with evolving EEs.
Gabriela González Sáez
SIGIR1
2021 CLEF eHealth Evaluation Lab 2021
Lorraine Goeuriot, Hanna Suominen, Liadh Kelly, Laura Alonso Alemany, Nicola Brew-Sam, Viviana Cotik, Darío Filippo, Gabriela González Sáez, Franco M. Luque, Philippe Mulhem, Gabriella Pasi, Roland Roller, Sandaru Seneviratne, Jorge Vivaldi, Marco Viviani 0001
ECIR (2)8