Jüri Keller

dblp:300/0770 · DBLP profile ↗
← Back
6ranked-venue papers in the field
3as first author
6since 2021 · last 2026
0000-0002-9392-8646ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 6 (3 first)
YearPublicationVenuePosition
2026 Evaluating Information Retrieval Models Along Time: The LongEval Lab at CLEF 2026
Timo Breuer 0002, Matteo Cancellieri, Alaa El-Ebshihy, Maik Fröbe, Petra Galuscáková, Lorraine Goeuriot, Gabriel Iturra-Bocaz, Jüri Keller, Petr Knoth, Andreas Konstantin Kruff, Philippe Mulhem, Florina Piroi, David Pride, Philipp Schaer, Didier Schwab
ECIR (4)8
2026 Formalized Information Needs Improve Large-Language-Model Relevance Judgments
abstract
Cranfield-style retrieval evaluations with too few or too many relevant documents or with low inter-assessor agreement on relevance can reduce the reliability of observations. In evaluations with human assessors, information needs are often formalized as retrieval topics to avoid an excessive number of relevant documents while maintaining good agreement. However, emerging evaluation setups that use Large Language Models (LLMs) as relevance assessors often use only queries, potentially decreasing the reliability. To study whether LLM relevance assessors benefit from formalized information needs, we synthetically formalize information needs with LLMs into topics that follow the established structure from previous human relevance assessments (i.e., descriptions and narratives). We compare assessors using synthetically formalized topics against the LLM-default query-only assessor on the~2019/2020~editions of TREC Deep Learning and Robust04. We find that assessors without formalization judge many more documents relevant and have a lower agreement, leading to reduced reliability in retrieval evaluations. Furthermore, we show that the formalized topics improve agreement between human and LLM relevance judgments, even when the topics are not highly similar to their human counterparts. Our findings indicate that LLM relevance assessors should use formalized information needs, as is standard for human assessment, and synthetically formalize topics when no human formalization exists to improve evaluation reliability.
Jüri Keller, Maik Fröbe, Björn Engelmann 0002, Fabian Haak 0001, Timo Breuer 0002, Birger Larsen, Philipp Schaer
SIGIR1
2025 LongEval at CLEF 2025: Longitudinal Evaluation of IR Model Performance
Matteo Cancellieri, Alaa El-Ebshihy, Tobias Fink, Petra Galuscáková, Gabriela González Sáez, Lorraine Goeuriot, David Iommi, Jüri Keller, Petr Knoth, Philippe Mulhem, Florina Piroi, David Pride, Philipp Schaer
ECIR (5)8
2025 Counterfactual Query Rewriting to Use Historical Relevance Feedback
Jüri Keller, Maik Fröbe, Gijs Hendriksen, Daria Alexander, Martin Potthast, Matthias Hagen, Philipp Schaer
ECIR (3)1
2025 Continuous Evaluation in Information Retrieval Across Methods and Time
abstract
Evaluating Information Retrieval (IR) systems is essential yet challenging. The dynamic nature of information and relevance further complicates IR. For example, Adar et al. observed that as early as 2009, websites frequently changed multiple times per hour. In response, recent IR systems have become more personalized, semantic, and context-aware. They evolved from lexical ranking functions to complex neural models embedded in feature-rich systems. While this often improves retrieval quality, it also challenges their reliability and robustness. The proposed research is motivated by the overarching goal of maintaining the trustworthiness of IR systems. To do so a rigorous evaluation is needed. Only by assessing the quality of a system can it be maintained and improved. Such an evaluation can not take place in isolation but must consider the dynamics of the search setting at all stages of a system-from development to maintenance and improvement. Based on the CRISP-DM methodology, these stages are sketched out as an ''IR Life Cycle'', which we will further define and oppose with a continuous evaluation framework in future work.
Jüri Keller
SIGIR1
2022 ir_metadata: An Extensible Metadata Schema for IR Experiments
abstract
The information retrieval (IR) community has a strong tradition of making the computational artifacts and resources available for future reuse, allowing the validation of experimental results. Besides the actual test collections, the underlying run files are often hosted in data archives as part of conferences like TREC, CLEF, or NTCIR. Unfortunately, the run data itself does not provide much information about the underlying experiment. For instance, the single run file is not of much use without the context of the shared task's website or the run data archive. In other domains, like the social sciences, it is good practice to annotate research data with metadata. In this work, we introduce \textttir\_metadata - an extensible metadata schema for TREC run files based on the PRIMAD model. We propose to align the metadata annotations to PRIMAD, which considers components of computational experiments that can affect reproducibility. Furthermore, we outline important components and information that should be reported in the metadata and give evidence from the literature. To demonstrate the usefulness of these metadata annotations, we implement new features in \textttrepro\_eval that support the outlined metadata schema for the use case of reproducibility studies. Additionally, we curate a dataset with run files derived from experiments with different instantiations of PRIMAD components and annotate these with the corresponding metadata. In the experiments, we cover reproducibility experiments that are identified by the metadata and classified by PRIMAD. With this work, we enable IR researchers to annotate TREC run files and improve the reuse value of experimental artifacts even further.
Timo Breuer 0002, Jüri Keller, Philipp Schaer
SIGIR2