EDBT 2026 Demo / reviewers in the wild / expert
Timo Breuer 0002
dblp:176/0934-2
· DBLP profile ↗
17ranked-venue papers in the field
7as first author
15since 2021 · last 2026
0000-0002-1765-2449ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 17 (7 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Information Retrieval Models Along Time: The LongEval Lab at CLEF 2026
Timo Breuer 0002, Matteo Cancellieri, Alaa El-Ebshihy, Maik Fröbe, Petra Galuscáková, Lorraine Goeuriot, Gabriel Iturra-Bocaz, Jüri Keller, Petr Knoth, Andreas Konstantin Kruff, Philippe Mulhem, Florina Piroi, David Pride, Philipp Schaer, Didier Schwab |
ECIR (4) | 1 |
| 2026 | Sim4IA-Bench: A User Simulation Benchmark Suite for Next Query and Utterance Prediction
Andreas Konstantin Kruff, Christin Kreutz, Timo Breuer 0002, Philipp Schaer, Krisztian Balog |
ECIR (4) | 3 |
| 2026 | Formalized Information Needs Improve Large-Language-Model Relevance JudgmentsabstractCranfield-style retrieval evaluations with too few or too many relevant documents or with low inter-assessor agreement on relevance can reduce the reliability of observations. In evaluations with human assessors, information needs are often formalized as retrieval topics to avoid an excessive number of relevant documents while maintaining good agreement. However, emerging evaluation setups that use Large Language Models (LLMs) as relevance assessors often use only queries, potentially decreasing the reliability. To study whether LLM relevance assessors benefit from formalized information needs, we synthetically formalize information needs with LLMs into topics that follow the established structure from previous human relevance assessments (i.e., descriptions and narratives). We compare assessors using synthetically formalized topics against the LLM-default query-only assessor on the~2019/2020~editions of TREC Deep Learning and Robust04. We find that assessors without formalization judge many more documents relevant and have a lower agreement, leading to reduced reliability in retrieval evaluations. Furthermore, we show that the formalized topics improve agreement between human and LLM relevance judgments, even when the topics are not highly similar to their human counterparts. Our findings indicate that LLM relevance assessors should use formalized information needs, as is standard for human assessment, and synthetically formalize topics when no human formalization exists to improve evaluation reliability. Jüri Keller, Maik Fröbe, Björn Engelmann 0002, Fabian Haak 0001, Timo Breuer 0002, Birger Larsen, Philipp Schaer |
SIGIR | 5 |
| 2026 | Now That Your System Has Been Reproduced, What Does This Mean for the Users?abstractReproducibility lies at the basis of the empirical method: a novel approach will be widely adopted if its experimental results can be validated and reproduced by the community. Previous work on reproducibility in Information Retrieval (IR) has mainly addressed the reproducibility and replicability of offline experiments, with a few exceptions that replicate user studies. To the best of our knowledge, no previous work has investigated how reproducibility affects real users. In this paper, we do that by evaluating and comparing the reproducibility of an IR system both offline and online. We consider a reference system and generate a constellation of reproduced systems with varying parameters. We select 6 systems with different degrees of offline reproducibility. We then run a between-subjects online experiment with 280 participants and collect clicks to evaluate online reproducibility. Results show that real users do not perceive moderate variations of the reproducibility degree of systems, while they become relevant when the difference with the original system increases. Furthermore, we trained a click model to evaluate online reproducibility with simulated clicks. Results are not consistent with those from the user study, suggesting that better click models are needed to evaluate online reproducibility. Our data and source code are publicly available: https://github.com/angelogeninatti/reproducibilityLogs . Angelo Geninatti Cossatin, Timo Breuer 0002, Noemi Mauro, Maria Maistro |
ACM Trans. Inf. Syst. | 2 |
| 2025 | "We Share Our Code Online": Why This Is Not Enough to Ensure Reproducibility and Progress in Recommender Systems ResearchabstractIssues with reproducibility have been identified as a major factor hampering progress in recommender systems research. In response, researchers increasingly share the code of their models. However, the provision of only the code of the proposed model is usually not sufficient to ensure reproducibility. In many works, the central claim is that a new model is advancing the state of the art. Thus, it is crucial that the entire experiment is reproducible, including the configuration and the results of the considered baselines. With this work, our goal is to gauge the level of reproducibility in algorithms research in recommender systems. We systematically analyzed the reproducibility level of 65 papers published at a top-ranked conference during the last three years. Our results are sobering. While the model code is shared in about two thirds of the papers, the code of the baselines is provided only in eight cases. The hyperparameters of the baselines are reported even less frequently, and how these were exactly determined is not explained in any paper. As a result, it is commonly not only impossible to reproduce the full result tables reported in the papers, it is also unclear if the claimed improvements over the state of the art were actually achieved. Overall, we conclude that the research community has not reached the required level of reproducibility yet. We therefore call for more rigorous reproducibility standards to ensure progress in this field. Faisal Shehzad, Timo Breuer 0002, Maria Maistro, Dietmar Jannach |
RecSys | 2 |
| 2025 | Evaluating Contrastive Feedback for Effective User SimulationsabstractThe use of Large Language Models (LLMs) for simulating user behavior in the domain of Interactive Information Retrieval has recently gained significant popularity. However, their application and capabilities remain highly debated and understudied. This study explores whether the underlying principles of contrastive training techniques, which have been effective for fine-tuning LLMs, can also be applied beneficially in the area of prompt engineering for user simulations. Andreas Konstantin Kruff, Timo Breuer 0002, Philipp Schaer |
SIGIR | 2 |
| 2025 | Second SIGIR Workshop on Simulations for Information Access (Sim4IA 2025)abstractSimulations in information access (IA) have recently gained interest, as shown by various tutorials and workshops around that topic. Simulations can be key contributors to central IA research and evaluation questions, especially around interactive settings when real users are unavailable, or their participation is impossible due to ethical reasons. In addition, simulations in IA can help contribute to a better understanding of users, reduce complexity of evaluation experiments, and improve reproducibility. Building on recent developments in methods and toolkits, the second iteration of our Sim4IA workshop aims to again bring together researchers and practitioners to form an interactive and engaging forum for discussions on the future perspectives of the field. An additional aim is to plan an upcoming TREC/CLEF campaign. Philipp Schaer, Christin Kreutz, Krisztian Balog, Timo Breuer 0002, Andreas Konstantin Kruff |
SIGIR | 4 |
| 2024 | Context-Driven Interactive Query Simulations Based on Generative Large Language Models
Björn Engelmann 0002, Timo Breuer 0002, Jana Isabelle Friese, Philipp Schaer, Norbert Fuhr |
ECIR (2) | 2 |
| 2024 | Browsing and Searching Metadata of TRECabstractInformation Retrieval (IR) research is deeply rooted in experimentation and evaluation, and the Text REtrieval Conference (TREC) has been playing a central role in making that possible since its inauguration in 1992. TREC's mission centers around providing the infrastructure and resources to make IR evaluations possible at scale. Over the years, a plethora of different retrieval problems were addressed, culminating in data artifacts that remained as valuable and useful tools for the IR community. Even though the data are largely available from TREC's website, there is currently no resource that facilitates a cohesive way to obtain metadata information about the run file - the IR community's de-facto standard data format for storing rankings of system-oriented IR experiments. Timo Breuer 0002, Ellen M. Voorhees, Ian Soboroff |
SIGIR | 1 |
| 2024 | SIGIR 2024 Workshop on Simulations for Information Access (Sim4IA 2024)abstractSimulations in various forms have been used to evaluate information access systems, like search engines, recommender systems, or conversational agents. In the form of the Cranfield paradigm, a simulation setup is well-known in the IR community, but user simulations have recently gained interest. While user simulations help to reduce the complexity of evaluation experiments and help with reproducibility, they can also contribute to a better understanding of users. Building on recent developments in methods and toolkits, the Sim4IA workshop aims to bring together researchers and practitioners to form an interactive and engaging forum for discussions on the future perspectives of the field. An additional aim is to plan an upcoming TREC/CLEF campaign. Philipp Schaer, Christin Kreutz, Krisztian Balog, Timo Breuer 0002, Norbert Fuhr |
SIGIR | 4 |
| 2023 | Simulating Users in Interactive Web Table RetrievalabstractConsidering the multimodal signals of search items is beneficial for retrieval effectiveness. Especially in web table retrieval (WTR) experiments, accounting for multimodal properties of tables boosts effectiveness. However, it still remains an open question how the single modalities affect user experience in particular. Previous work analyzed WTR performance in ad-hoc retrieval benchmarks, which neglects interactive search behavior and limits the conclusion about the implications for real-world user environments. Björn Engelmann 0002, Timo Breuer 0002, Philipp Schaer |
CIKM | 2 |
| 2023 | An in-depth investigation on the behavior of measures to quantify reproducibilityabstractScience is facing a so-called reproducibility crisis, where researchers struggle to repeat experiments and to get the same or comparable results. This represents a fundamental problem in any scientific discipline because reproducibility lies at the very basis of the scientific method. A central methodological question is how to measure reproducibility and interpret different measures. In Information Retrieval (IR), current practices to measure reproducibility rely mainly on comparing averaged scores. If the reproduced score is close enough to the original one, the reproducibility experiment is deemed successful, although the identical scores can still rely on entirely different result lists. Therefore, this paper focuses on measures to quantify reproducibility in IR and their behavior. We present a critical analysis of IR reproducibility measures by synthetically generating runs in a controlled experimental setting, which allows us to control the amount of reproducibility error. These synthetic runs are generated by a deterioration algorithm based on swaps and replacements of documents in ranked lists. We investigate the behavior of different reproducibility measures with these synthetic runs in three different scenarios. Moreover, we propose a normalized version of Root Mean Square Error (RMSE) to quantify reproducibility better. Experimental results show that a single score is not enough to decide whether an experiment is successfully reproduced because such a score depends on the type of effectiveness measure and the performance of the original run. This study highlights how challenging it can be to reproduce experimental results and quantify the amount of reproducibility. Maria Maistro, Timo Breuer 0002, Philipp Schaer, Nicola Ferro 0001 |
Inf. Process. Manag. | 2 |
| 2022 | Validating Simulations of User Query Variants
Timo Breuer 0002, Norbert Fuhr, Philipp Schaer |
ECIR (1) | 1 |
| 2022 | ir_metadata: An Extensible Metadata Schema for IR ExperimentsabstractThe information retrieval (IR) community has a strong tradition of making the computational artifacts and resources available for future reuse, allowing the validation of experimental results. Besides the actual test collections, the underlying run files are often hosted in data archives as part of conferences like TREC, CLEF, or NTCIR. Unfortunately, the run data itself does not provide much information about the underlying experiment. For instance, the single run file is not of much use without the context of the shared task's website or the run data archive. In other domains, like the social sciences, it is good practice to annotate research data with metadata. In this work, we introduce \textttir\_metadata - an extensible metadata schema for TREC run files based on the PRIMAD model. We propose to align the metadata annotations to PRIMAD, which considers components of computational experiments that can affect reproducibility. Furthermore, we outline important components and information that should be reported in the metadata and give evidence from the literature. To demonstrate the usefulness of these metadata annotations, we implement new features in \textttrepro\_eval that support the outlined metadata schema for the use case of reproducibility studies. Additionally, we curate a dataset with run files derived from experiments with different instantiations of PRIMAD components and annotate these with the corresponding metadata. In the experiments, we cover reproducibility experiments that are identified by the metadata and classified by PRIMAD. With this work, we enable IR researchers to annotate TREC run files and improve the reuse value of experimental artifacts even further. Timo Breuer 0002, Jüri Keller, Philipp Schaer |
SIGIR | 1 |
| 2021 | repro_eval: A Python Interface to Reproducibility Measures of System-Oriented IR Experiments
Timo Breuer 0002, Nicola Ferro 0001, Maria Maistro, Philipp Schaer |
ECIR (2) | 1 |
| 2020 | Reproducible Online Search Experiments
Timo Breuer 0002 |
ECIR (2) | 1 |
| 2020 | How to Measure the Reproducibility of System-oriented IR ExperimentsabstractReplicability and reproducibility of experimental results are primary concerns in all the areas of science and IR is not an exception. Besides the problem of moving the field towards more reproducible experimental practices and protocols, we also face a severe methodological issue: we do not have any means to assess when reproduced is reproduced. Moreover, we lack any reproducibility-oriented dataset, which would allow us to develop such methods. Timo Breuer 0002, Nicola Ferro 0001, Norbert Fuhr, Maria Maistro, Tetsuya Sakai, Philipp Schaer, Ian Soboroff |
SIGIR | 1 |