EDBT 2026 Demo / reviewers in the wild / expert
Philipp Schaer
dblp:67/8577
· DBLP profile ↗
27ranked-venue papers in the field
5as first author
18since 2021 · last 2026
0000-0002-8817-4632ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 27 (5 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Evaluating Information Retrieval Models Along Time: The LongEval Lab at CLEF 2026
Timo Breuer 0002, Matteo Cancellieri, Alaa El-Ebshihy, Maik Fröbe, Petra Galuscáková, Lorraine Goeuriot, Gabriel Iturra-Bocaz, Jüri Keller, Petr Knoth, Andreas Konstantin Kruff, Philippe Mulhem, Florina Piroi, David Pride, Philipp Schaer, Didier Schwab |
ECIR (4) | 14 |
| 2026 | LISP - A Rich Interaction Dataset and Loggable Interactive Search Platform
Jana Isabelle Friese, Andreas Konstantin Kruff, Philipp Schaer, Norbert Fuhr, Nicola Ferro 0001 |
ECIR (4) | 3 |
| 2026 | Validating Search Query Simulations: A Taxonomy of Measures
Andreas Konstantin Kruff, Nolwenn Bernard, Philipp Schaer |
ECIR (1) | 3 |
| 2026 | Sim4IA-Bench: A User Simulation Benchmark Suite for Next Query and Utterance Prediction
Andreas Konstantin Kruff, Christin Kreutz, Timo Breuer 0002, Philipp Schaer, Krisztian Balog |
ECIR (4) | 4 |
| 2026 | Formalized Information Needs Improve Large-Language-Model Relevance JudgmentsabstractCranfield-style retrieval evaluations with too few or too many relevant documents or with low inter-assessor agreement on relevance can reduce the reliability of observations. In evaluations with human assessors, information needs are often formalized as retrieval topics to avoid an excessive number of relevant documents while maintaining good agreement. However, emerging evaluation setups that use Large Language Models (LLMs) as relevance assessors often use only queries, potentially decreasing the reliability. To study whether LLM relevance assessors benefit from formalized information needs, we synthetically formalize information needs with LLMs into topics that follow the established structure from previous human relevance assessments (i.e., descriptions and narratives). We compare assessors using synthetically formalized topics against the LLM-default query-only assessor on the~2019/2020~editions of TREC Deep Learning and Robust04. We find that assessors without formalization judge many more documents relevant and have a lower agreement, leading to reduced reliability in retrieval evaluations. Furthermore, we show that the formalized topics improve agreement between human and LLM relevance judgments, even when the topics are not highly similar to their human counterparts. Our findings indicate that LLM relevance assessors should use formalized information needs, as is standard for human assessment, and synthetically formalize topics when no human formalization exists to improve evaluation reliability. Jüri Keller, Maik Fröbe, Björn Engelmann 0002, Fabian Haak 0001, Timo Breuer 0002, Birger Larsen, Philipp Schaer |
SIGIR | 7 |
| 2025 | LongEval at CLEF 2025: Longitudinal Evaluation of IR Model Performance
Matteo Cancellieri, Alaa El-Ebshihy, Tobias Fink, Petra Galuscáková, Gabriela González Sáez, Lorraine Goeuriot, David Iommi, Jüri Keller, Petr Knoth, Philippe Mulhem, Florina Piroi, David Pride, Philipp Schaer |
ECIR (5) | 13 |
| 2025 | Counterfactual Query Rewriting to Use Historical Relevance Feedback
Jüri Keller, Maik Fröbe, Gijs Hendriksen, Daria Alexander, Martin Potthast, Matthias Hagen, Philipp Schaer |
ECIR (3) | 7 |
| 2025 | REANIMATOR: Reanimate Retrieval Test Collections with Extracted and Synthetic ResourcesabstractRetrieval test collections are essential for evaluating information retrieval systems, yet they often lack generalizability across tasks. To overcome this limitation, we introduce REANIMATOR, a versatile framework designed to enable the repurposing of existing test collections by enriching them with extracted and synthetic resources. REANIMATOR enhances test collections from PDF files by parsing full texts and machine-readable tables, as well as related contextual information. It then employs state-of-the-art large language models to produce synthetic relevance labels. Including an optional human-in-the-loop step can help validate the resources that have been extracted and generated. We demonstrate its potential with a revitalized version of the TREC-COVID test collection, showcasing the development of a retrieval-augmented generation system and evaluating the impact of tables on retrieval-augmented generation. REANIMATOR enables the reuse of test collections for new applications, lowering costs and broadening the utility of legacy resources. Björn Engelmann 0002, Fabian Haak 0001, Philipp Schaer, Mani Erfanian Abdoust, Linus Netze, Meik Bittkowski |
SIGIR | 3 |
| 2025 | Evaluating Contrastive Feedback for Effective User SimulationsabstractThe use of Large Language Models (LLMs) for simulating user behavior in the domain of Interactive Information Retrieval has recently gained significant popularity. However, their application and capabilities remain highly debated and understudied. This study explores whether the underlying principles of contrastive training techniques, which have been effective for fine-tuning LLMs, can also be applied beneficially in the area of prompt engineering for user simulations. Andreas Konstantin Kruff, Timo Breuer 0002, Philipp Schaer |
SIGIR | 3 |
| 2025 | Second SIGIR Workshop on Simulations for Information Access (Sim4IA 2025)abstractSimulations in information access (IA) have recently gained interest, as shown by various tutorials and workshops around that topic. Simulations can be key contributors to central IA research and evaluation questions, especially around interactive settings when real users are unavailable, or their participation is impossible due to ethical reasons. In addition, simulations in IA can help contribute to a better understanding of users, reduce complexity of evaluation experiments, and improve reproducibility. Building on recent developments in methods and toolkits, the second iteration of our Sim4IA workshop aims to again bring together researchers and practitioners to form an interactive and engaging forum for discussions on the future perspectives of the field. An additional aim is to plan an upcoming TREC/CLEF campaign. Philipp Schaer, Christin Kreutz, Krisztian Balog, Timo Breuer 0002, Andreas Konstantin Kruff |
SIGIR | 1 |
| 2024 | Context-Driven Interactive Query Simulations Based on Generative Large Language Models
Björn Engelmann 0002, Timo Breuer 0002, Jana Isabelle Friese, Philipp Schaer, Norbert Fuhr |
ECIR (2) | 4 |
| 2024 | SIGIR 2024 Workshop on Simulations for Information Access (Sim4IA 2024)abstractSimulations in various forms have been used to evaluate information access systems, like search engines, recommender systems, or conversational agents. In the form of the Cranfield paradigm, a simulation setup is well-known in the IR community, but user simulations have recently gained interest. While user simulations help to reduce the complexity of evaluation experiments and help with reproducibility, they can also contribute to a better understanding of users. Building on recent developments in methods and toolkits, the Sim4IA workshop aims to bring together researchers and practitioners to form an interactive and engaging forum for discussions on the future perspectives of the field. An additional aim is to plan an upcoming TREC/CLEF campaign. Philipp Schaer, Christin Kreutz, Krisztian Balog, Timo Breuer 0002, Norbert Fuhr |
SIGIR | 1 |
| 2023 | Simulating Users in Interactive Web Table RetrievalabstractConsidering the multimodal signals of search items is beneficial for retrieval effectiveness. Especially in web table retrieval (WTR) experiments, accounting for multimodal properties of tables boosts effectiveness. However, it still remains an open question how the single modalities affect user experience in particular. Previous work analyzed WTR performance in ad-hoc retrieval benchmarks, which neglects interactive search behavior and limits the conclusion about the implications for real-world user environments. Björn Engelmann 0002, Timo Breuer 0002, Philipp Schaer |
CIKM | 3 |
| 2023 | An in-depth investigation on the behavior of measures to quantify reproducibilityabstractScience is facing a so-called reproducibility crisis, where researchers struggle to repeat experiments and to get the same or comparable results. This represents a fundamental problem in any scientific discipline because reproducibility lies at the very basis of the scientific method. A central methodological question is how to measure reproducibility and interpret different measures. In Information Retrieval (IR), current practices to measure reproducibility rely mainly on comparing averaged scores. If the reproduced score is close enough to the original one, the reproducibility experiment is deemed successful, although the identical scores can still rely on entirely different result lists. Therefore, this paper focuses on measures to quantify reproducibility in IR and their behavior. We present a critical analysis of IR reproducibility measures by synthetically generating runs in a controlled experimental setting, which allows us to control the amount of reproducibility error. These synthetic runs are generated by a deterioration algorithm based on swaps and replacements of documents in ranked lists. We investigate the behavior of different reproducibility measures with these synthetic runs in three different scenarios. Moreover, we propose a normalized version of Root Mean Square Error (RMSE) to quantify reproducibility better. Experimental results show that a single score is not enough to decide whether an experiment is successfully reproduced because such a score depends on the type of effectiveness measure and the performance of the original run. This study highlights how challenging it can be to reproduce experimental results and quantify the amount of reproducibility. Maria Maistro, Timo Breuer 0002, Philipp Schaer, Nicola Ferro 0001 |
Inf. Process. Manag. | 3 |
| 2022 | Validating Simulations of User Query Variants
Timo Breuer 0002, Norbert Fuhr, Philipp Schaer |
ECIR (1) | 3 |
| 2022 | ir_metadata: An Extensible Metadata Schema for IR ExperimentsabstractThe information retrieval (IR) community has a strong tradition of making the computational artifacts and resources available for future reuse, allowing the validation of experimental results. Besides the actual test collections, the underlying run files are often hosted in data archives as part of conferences like TREC, CLEF, or NTCIR. Unfortunately, the run data itself does not provide much information about the underlying experiment. For instance, the single run file is not of much use without the context of the shared task's website or the run data archive. In other domains, like the social sciences, it is good practice to annotate research data with metadata. In this work, we introduce \textttir\_metadata - an extensible metadata schema for TREC run files based on the PRIMAD model. We propose to align the metadata annotations to PRIMAD, which considers components of computational experiments that can affect reproducibility. Furthermore, we outline important components and information that should be reported in the metadata and give evidence from the literature. To demonstrate the usefulness of these metadata annotations, we implement new features in \textttrepro\_eval that support the outlined metadata schema for the use case of reproducibility studies. Additionally, we curate a dataset with run files derived from experiments with different instantiations of PRIMAD components and annotate these with the corresponding metadata. In the experiments, we cover reproducibility experiments that are identified by the metadata and classified by PRIMAD. With this work, we enable IR researchers to annotate TREC run files and improve the reuse value of experimental artifacts even further. Timo Breuer 0002, Jüri Keller, Philipp Schaer |
SIGIR | 3 |
| 2021 | repro_eval: A Python Interface to Reproducibility Measures of System-Oriented IR Experiments
Timo Breuer 0002, Nicola Ferro 0001, Maria Maistro, Philipp Schaer |
ECIR (2) | 4 |
| 2021 | Living Lab Evaluation for Life and Social Sciences Search Platforms - LiLAS at CLEF 2021
Philipp Schaer, Johann Schaible, Leyla Jael Castro |
ECIR (2) | 1 |
| 2020 | Living Labs for Academic Search at CLEF 2020
Philipp Schaer, Johann Schaible, Bernd Müller |
ECIR (2) | 1 |
| 2020 | How to Measure the Reproducibility of System-oriented IR ExperimentsabstractReplicability and reproducibility of experimental results are primary concerns in all the areas of science and IR is not an exception. Besides the problem of moving the field towards more reproducible experimental practices and protocols, we also face a severe methodological issue: we do not have any means to assess when reproduced is reproduced. Moreover, we lack any reproducibility-oriented dataset, which would allow us to develop such methods. Timo Breuer 0002, Nicola Ferro 0001, Norbert Fuhr, Maria Maistro, Tetsuya Sakai, Philipp Schaer, Ian Soboroff |
SIGIR | 6 |
| 2015 | Query Expansion for Survey Question Retrieval in the Social Sciences
Nadine Dulisch, Andreas Oskar Kempf, Philipp Schaer |
TPDL | 3 |
| 2014 | Bibliometric-Enhanced Information Retrieval
Philipp Mayr 0001, Andrea Scharnhorst, Birger Larsen, Philipp Schaer, Peter Mutschke |
ECIR | 4 |
| 2013 | A framework for specific term recommendation systemsabstractIn this paper we present the IRSA framework that enables the automatic creation of search term suggestion or recommendation systems (TS). Such TS are used to operationalize interactive query expansion and help users in refining their information need in the query formulation phase. Our recent research has shown TS to be more effective when specific to a certain domain. The presented technical framework allows owners of Digital Libraries to create their own specific TS constructed via OAI-harvested metadata with very little effort. Thomas Lüke, Philipp Schaer, Philipp Mayr 0001 |
SIGIR | 2 |
| 2012 | Integrating Interactive Visualizations in the Search Process of Digital Libraries and IR Systems
Daniel Hienert, Frank Sawitzki, Philipp Schaer, Philipp Mayr 0001 |
ECIR | 3 |
| 2012 | Improving Retrieval Results with Discipline-Specific Query Expansion
Thomas Lüke, Philipp Schaer, Philipp Mayr 0001 |
TPDL | 2 |
| 2012 | Extending Term Suggestion with Author Names
Philipp Schaer, Philipp Mayr 0001, Thomas Lüke |
TPDL | 1 |
| 2011 | A Novel Combined Term Suggestion Service for Domain-Specific Digital Libraries
Daniel Hienert, Philipp Schaer, Johann Schaible, Philipp Mayr 0001 |
TPDL | 2 |